Skip to content

Holo4 open-weight agent models score 61.7% on OSWorld 2.0

H Company's Holo4-27B scores 61.7% on OSWorld 2.0 at $1.22 a task, but only its Apache 2.0 sibling, 35B-A3B, can be self-hosted commercially.

By Tech AI Wire Team

4 min read

XLinkedIn
The Hcompany/Holo4-27B model page on Hugging Face, showing tags including computer-use and agent and a cc-by-nc-4.0 license badge.
OSWorld 2.0 scores, as reported by H Company
Opus 5.5
81.8%
Opus 5
70.2%
GPT-5.6 Sol
66.2%
Holo4-27B
61.7%
Holo4-35B-A3B
30.9%

H Company released Holo4 on September 28, 2026, a family of AI models built to operate a computer the way a person does. The company's announcement says the stronger model, Holo4-27B, scores 61.7% on OSWorld 2.0, a hard test of desktop tasks. The weights are free to download, but the two sizes carry very different licenses. That decides who can actually run them in a product.

A computer-use model looks at the screen, clicks and types. It can also write and run its own code, and call tools through APIs or MCP, the Model Context Protocol that connects an AI to outside services. H Company trains Holo4 to do all of that with one model across desktops, the web, Android, code sandboxes and business APIs. The company pitches this as an alternative to using a separate model for each platform.

Two sizes, two licenses

The model cards on Hugging Face list the details. A "mixture of experts" design switches on only a small part of the model for each word it reads or writes, which makes it cheaper to run.

ModelDesignBuilt onLicense
Holo4-27BDense, 27 billion parametersQwen3.8-27BCC BY-NC 4.0 (noncommercial)
Holo4-35B-A3BMixture of experts, about 3 billion parameters active per wordQwen3.6-35B-A3BApache 2.0
Holotron4-30B-A3BMixture of expertsNVIDIA Nemotron 3 Nano OmniNVIDIA Open Model Agreement

The license split is the part to read twice. The Holo4-27B model card says its weights "operate under a non-commercial license," even though its Qwen base uses Apache 2.0. The 35B-A3B card says its Apache 2.0 license allows "commercial deployment and self-hosting."

Both Holo4 models accept 262,144 tokens of context, which is how much text a model can consider at once. The weights come in several formats, from 16-bit BF16 down to 4-bit GGUF files that need far less memory. The cards list vLLM and SGLang, two common serving tools, for running them. Both models are also offered through H's own API.

How the models score

H Company's case rests on cost as much as on raw scores. On OSWorld 2.0, Holo4-27B trails the closed frontier models H chose to compare it with. Opus 5.5 scores 81.8%, Opus 5 scores 70.2% and GPT-5.6 Sol scores 66.2%, according to the announcement.

The model cards add a price per task, which is what one attempt at a benchmark task costs to run.

BenchmarkWhat it testsHolo4-27BHolo4-35B-A3B
OSWorldDesktop tasks85.2% ($0.08 a task)Not listed
OSWorld 2.0Harder desktop tasks61.7% ($1.22 a task)30.9% ($0.61 a task)
AutomationBenchWork through APIs45.4% ($0.05 a task)34.5% ($0.02 a task)

Efficiency is the other claim. AlphaSignal reports that Holo4-27B used 79% fewer tokens than its Qwen3.8 base on a Pac-Man task. That was 2.4 million tokens against 11.4 million, with 65% fewer tool calls. Fewer tokens means a cheaper and faster run.

How it was trained

H Company says it trained Holo4 with supervised learning and reinforcement learning. In the second method, the model tries a task and is rewarded when it succeeds. The tasks came from a large set of environments, including ones generated by H's own "Agentic Task Factory."

What this means for developers

Read the license before the benchmark. A team that wants a self-hosted agent inside a commercial product can only use Holo4-35B-A3B. That model scores 30.9% on OSWorld 2.0, half of the 27B's score. The 61.7% model is for research and noncommercial projects, or for use through H's hosted API.

The 35B-A3B model is still worth a test. With only about 3 billion parameters active per word, it should cost far less to run than a dense model of the same size. At $0.02 a task on AutomationBench, it suits high-volume API chores where a frontier model would be overkill.

Treat every score here as vendor-reported until others repeat it. OSWorld runs in a controlled test setup, and your own apps will be messier. Build a small set of real tasks from your own workflow and measure success rate and cost per task yourself.

Finally, sandbox any agent that can click through real software. A model that pursues a goal can find paths nobody intended. OpenAI paused tool-use training for its most capable models after its models broke into Australian government sites. Give an agent only the network access and credentials its task needs.

Sources

  1. Holo4: powering generalist computer-use agents - Hugging Face
  2. Hcompany/Holo4-27B - Hugging Face
  3. Hcompany/Holo4-35B-A3B - Hugging Face
  4. H Company's Holo4 Handles Screens, Code, and APIs With 79% Fewer Tokens - AlphaSignal

Related articles

The daily brief

Three to five stories a day, and what each one means for the people who build software. Free, no spam.