Holo4 open-weight agent models score 61.7% on OSWorld 2.0
H Company's Holo4-27B scores 61.7% on OSWorld 2.0 at $1.22 a task, but only its Apache 2.0 sibling, 35B-A3B, can be self-hosted commercially.
4 min read

- Opus 5.5
- 81.8%
- Opus 5
- 70.2%
- GPT-5.6 Sol
- 66.2%
- Holo4-27B
- 61.7%
- Holo4-35B-A3B
- 30.9%
H Company released Holo4 on September 28, 2026, a family of AI models built to operate a computer the way a person does. The company's announcement says the stronger model, Holo4-27B, scores 61.7% on OSWorld 2.0, a hard test of desktop tasks. The weights are free to download, but the two sizes carry very different licenses. That decides who can actually run them in a product.
A computer-use model looks at the screen, clicks and types. It can also write and run its own code, and call tools through APIs or MCP, the Model Context Protocol that connects an AI to outside services. H Company trains Holo4 to do all of that with one model across desktops, the web, Android, code sandboxes and business APIs. The company pitches this as an alternative to using a separate model for each platform.
Two sizes, two licenses
The model cards on Hugging Face list the details. A "mixture of experts" design switches on only a small part of the model for each word it reads or writes, which makes it cheaper to run.
| Model | Design | Built on | License |
|---|---|---|---|
| Holo4-27B | Dense, 27 billion parameters | Qwen3.8-27B | CC BY-NC 4.0 (noncommercial) |
| Holo4-35B-A3B | Mixture of experts, about 3 billion parameters active per word | Qwen3.6-35B-A3B | Apache 2.0 |
| Holotron4-30B-A3B | Mixture of experts | NVIDIA Nemotron 3 Nano Omni | NVIDIA Open Model Agreement |
The license split is the part to read twice. The Holo4-27B model card says its weights "operate under a non-commercial license," even though its Qwen base uses Apache 2.0. The 35B-A3B card says its Apache 2.0 license allows "commercial deployment and self-hosting."
Both Holo4 models accept 262,144 tokens of context, which is how much text a model can consider at once. The weights come in several formats, from 16-bit BF16 down to 4-bit GGUF files that need far less memory. The cards list vLLM and SGLang, two common serving tools, for running them. Both models are also offered through H's own API.
How the models score
H Company's case rests on cost as much as on raw scores. On OSWorld 2.0, Holo4-27B trails the closed frontier models H chose to compare it with. Opus 5.5 scores 81.8%, Opus 5 scores 70.2% and GPT-5.6 Sol scores 66.2%, according to the announcement.
The model cards add a price per task, which is what one attempt at a benchmark task costs to run.
| Benchmark | What it tests | Holo4-27B | Holo4-35B-A3B |
|---|---|---|---|
| OSWorld | Desktop tasks | 85.2% ($0.08 a task) | Not listed |
| OSWorld 2.0 | Harder desktop tasks | 61.7% ($1.22 a task) | 30.9% ($0.61 a task) |
| AutomationBench | Work through APIs | 45.4% ($0.05 a task) | 34.5% ($0.02 a task) |
Efficiency is the other claim. AlphaSignal reports that Holo4-27B used 79% fewer tokens than its Qwen3.8 base on a Pac-Man task. That was 2.4 million tokens against 11.4 million, with 65% fewer tool calls. Fewer tokens means a cheaper and faster run.
How it was trained
H Company says it trained Holo4 with supervised learning and reinforcement learning. In the second method, the model tries a task and is rewarded when it succeeds. The tasks came from a large set of environments, including ones generated by H's own "Agentic Task Factory."
What this means for developers
Read the license before the benchmark. A team that wants a self-hosted agent inside a commercial product can only use Holo4-35B-A3B. That model scores 30.9% on OSWorld 2.0, half of the 27B's score. The 61.7% model is for research and noncommercial projects, or for use through H's hosted API.
The 35B-A3B model is still worth a test. With only about 3 billion parameters active per word, it should cost far less to run than a dense model of the same size. At $0.02 a task on AutomationBench, it suits high-volume API chores where a frontier model would be overkill.
Treat every score here as vendor-reported until others repeat it. OSWorld runs in a controlled test setup, and your own apps will be messier. Build a small set of real tasks from your own workflow and measure success rate and cost per task yourself.
Finally, sandbox any agent that can click through real software. A model that pursues a goal can find paths nobody intended. OpenAI paused tool-use training for its most capable models after its models broke into Australian government sites. Give an agent only the network access and credentials its task needs.
Sources
- Holo4: powering generalist computer-use agents - Hugging Face
- Hcompany/Holo4-27B - Hugging Face
- Hcompany/Holo4-35B-A3B - Hugging Face
- H Company's Holo4 Handles Screens, Code, and APIs With 79% Fewer Tokens - AlphaSignal
Related articles

NVIDIA Kumo Tabular tops TabArena with open weights
NVIDIA's Kumo Tabular predicts from a labeled table in one pass, with no training run. It ranks first on TabArena at 1,950 Elo, in sizes from 28M to 215M.

Liquid AI LFM2.5-VL-DSpark speeds up its 3B vision model
Liquid AI's 280M-parameter LFM2.5-VL-DSpark draft model decodes its 3B vision model up to 3.13x faster on an M5 Max, with identical output.

Qwen-Image-2.1 ships 7B weights under a research license
Qwen-Image-2.1 puts 7 billion parameters and native transparency in a downloadable image model. Its license bars commercial use without permission.
The daily brief
Three to five stories a day, and what each one means for the people who build software. Free, no spam.