Published: 29 September 2026
H Company released Holo4 on 28 September 2026, a family of open-weight models that operate software the way a person does. The 27B version scores 85.2% on OSWorld at $0.08 per task on H Company’s own numbers. The weights are on Hugging Face, but the 27B carries a non-commercial licence.
What can Holo4 do?
Holo4 is a computer-use model. It looks at a screenshot, decides where to click, types, and checks what happened. What separates it from earlier models in this category is that it does not stop at the screen. The same model writes and runs its own code, and calls MCP servers and REST APIs directly.
That matters because most real automation is mixed. Half a task is a form on a website with no API, and the other half is a database query that would be absurd to do by clicking. Until now you wired two models together. Holo4 is one model across desktop, web, Android and a code sandbox.
The release covers three models. Holo4-27B is dense and built on Qwen3.8-27B. Holo4-35B-A3B is a mixture-of-experts model, which switches on only part of itself for each step, so it runs about 3 billion parameters at a time out of 35 billion. Holotron4 Nano is a 30B-A3B model based on NVIDIA’s Nemotron 3 Nano Omni.
What you need to know to try it:
- Weights on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF
- Context window of 262,144 tokens, which H Company rounds to 256K
- Runs with Transformers, vLLM, SGLang or Docker, and llama.cpp for a local machine
- Also available as a hosted API at hub.hcompany.ai
- API cost from $0.04 per million cached input tokens, output between $0.40 and $3.00 per million
- Every benchmark run is published as a replayable trajectory in the Hcompany/trajectories dataset
Holo4 benchmarks and cost per task
Take this from the table: Holo4 wins on price per task, not on capability. On the long-workflow test it sits 20 points behind Claude Opus 5.5.
| What is measured | Holo4-27B | The alternative |
|---|---|---|
| OSWorld (hitting the right control on screen) | 85.2% at $0.08 per task | not published by Anthropic for this version |
| OSWorld 2.0 (long, multi-step desktop workflows) | 61.7% at $1.22 per task | Claude Opus 5.5: 81.8% |
| AndroidWorld (tasks on an Android phone) | 85.1% | not published by Anthropic |
| AutomationBench (office automation) | 45.4% | Holo4-35B-A3B: 34.5% at $0.02 per task |
| Tokens used on one build task | 2.4 million | Qwen3.8-27B base model: 11.4 million |
| Licence on the weights | CC BY-NC 4.0, non-commercial | Holo4-35B-A3B: Apache 2.0 |
These are H Company’s own evaluations. The trajectories are published, which is more than most labs do and means you can replay a run rather than take the score on trust. The token figure comes from one task, building a Pac-Man clone, so treat it as an illustration and not a rate you should expect.
The licence split decides whether you can use Holo4
H Company’s announcement describes Holo4 as open weights. The model cards tell a more specific story. Holo4-35B-A3B is Apache 2.0, which lets you ship it in a product. Holo4-27B is CC BY-NC 4.0, which does not.
That is awkward, because the 27B is the model with the good numbers. It scores 61.7% on OSWorld 2.0 against 30.9% for the Apache-licensed 35B-A3B. If you are building something commercial, the version you are allowed to use is the one that performs at half the level, and the announcement does not draw attention to that split.
What this means
Worth testing now if you automate desktop software on your own servers and can live with a non-commercial licence, which in practice means internal tools and research. A four-person operations team automating a legacy Windows application that has no API is the clear case. At $1.22 per task on the hosted API against frontier prices, and with weights you can run in-house on data that cannot leave the building, the economics work even at 61.7%.
Safe to ignore for now if you are shipping computer use to customers. The Apache-licensed 35B-A3B at 30.9% on OSWorld 2.0 fails roughly seven tasks in ten, and a product that needs a human to rescue it that often is not a product. Watch for H Company to relicense the 27B, because that single change would move this release from interesting to useful. Until then, the thing to test first is your own worst workflow against the 27B, so you know what you would be buying if the licence changes.
For more information, visit the official announcement of Holo4 on the H Company blog.