Holo4 Open Computer-Use Agents Hit 61.7% on OSWorld 2.0 — Still 20 Points Behind Opus 5.5
H Company shipped Holo4 on September 28: open-weight 27B and 35B-A3B computer-use agents on the H Models API. Vendor OSWorld 2.0 scores, $0.40/$3 pricing, and the CC BY-NC vs Apache license split.
H Company released Holo4 on September 28, 2026 — a pair of open-weight computer-use agents that click GUIs, write code, call MCP tools, and hit business APIs from the same model. The flagship is a 27B dense model; the cheaper sibling is a 35B-A3B mixture of experts.
On H Company’s own OSWorld 2.0 run, Holo4 27B scores 61.7% at about $1.22 per task. That trails Claude Opus 5.5’s vendor-reported 81.8% — but sits well below the multi-dollar frontier cost-per-task numbers H Company charts next to it. Both sizes are live on the H Models API, with weights on Hugging Face.
Treat the scoreboard as vendor-reported until independent harnesses catch up. The more useful move from H Company is publishing 7,366 replayable trajectories behind those scores.
What Shipped Today
Per the Hugging Face announcement and the company newsroom post, Holo4 is two production models plus a smaller cousin:
- Holo4-27B — dense VLM on Qwen3.8-27B; API ID
holo4-27b - Holo4-35B-A3B — MoE on Qwen3.6-35B-A3B (3B active); API ID
holo4-35b-a3b - Holotron4 Nano — the same post-training stack applied to NVIDIA’s Nemotron 3 Nano Omni
H Company’s pitch is that most agent models specialize in one interface. Holo4 is trained to mix GUI actions, code execution, and MCP/API calls in one run — on desktop, web, Android, a code sandbox, or business APIs — without swapping model IDs. The Models API is OpenAI-compatible at https://api.hcompany.ai/v1, with 262,144-token context listed on the model cards.
Docs tell developers to start with holo4-35b-a3b for interactive loops and move to holo4-27b for long multi-step work. A free tier still points at the older holo3-1-35b-a3b.
The Rate Card From H Company
These are H Models API list prices from the launch page, in USD per 1 million tokens:
| Model | Input | Cache read | Output | Context |
|---|---|---|---|---|
Holo4 27B (holo4-27b) |
$0.40 | $0.04 | $3.00 | 256K |
Holo4 35B-A3B (holo4-35b-a3b) |
$0.30 | $0.03 | $2.00 | 256K |
For comparison, Claude Opus 5.5 lists at $4 / $20, and GPT-6 Sol at $2 / $10. Token price is not cost per finished desktop task — agent loops burn screenshots and tool traces — but the gap is large enough that H Company’s cost-per-task charts are the claim worth testing, not ignoring.
H Company also says the API defaults to zero data retention: prompts and responses are not stored; only request time, model, and token counts are logged.
Benchmarks: What H Company Claims
All numbers below are from H Company’s launch table unless noted. Holo4 scores are in H Company’s harness. Frontier scores are public figures across different harnesses and effort levels — H Company says so in the footnotes. Label them vendor-reported.
| Benchmark | Holo4 27B | Holo4 35B-A3B | Qwen3.8 27B (base) | Frontier (as cited) |
|---|---|---|---|---|
| OSWorld | 85.2% · $0.08 | 80.8% · $0.05 | 84.3% · $0.22 | Fable 5 86.0%; Qwen3.8 Max 86.1% |
| OSWorld 2.0 (avg partial / success) | 61.7% / 41.5% · $1.22 | 30.9% / 12.3% · $0.61 | 48.0% / 19.4% · $3.49 | Opus 5.5 81.8% / 48.7% · $8.48; GPT-6 Astra 73.5% · $9.07 |
| AutomationBench (public set) | 45.4% · $0.05 | 34.5% · $0.02 | 40.3% · $0.09* | Opus 5 50.3% · $3.05; GPT-5.6 Sol 45.8% · $0.67 |
| AndroidWorld | 85.1% · $0.08 | 77.6% · $0.07 | 81.9% · $0.13 | Fable 5 88.8% |
| ALE-CLI (105 Linux tasks) | 44.1% / 19.4% · $0.82 | 30.9% / 13.5% · $0.29 | 43.5% / 19.0% | Opus 5.5 63.7% / 34.3% · $8.22 |
*Qwen3.8 27B AutomationBench score measured in H Company’s harness where no public score existed. Costs under Holo4 rows use H Models API rates from H Company’s token counts.
The pattern is clear on H Company’s chart: Holo4 27B beats its Qwen base on long workflows (61.7% vs 48.0% on OSWorld 2.0) and undercuts the closed models on dollars per task. It does not close the absolute gap to Opus 5.5 or GPT-6 Astra on those long runs. AutomationBench still has an asterisk — 480 of the 600 public tasks overlap H Company’s training-data split; on the 120 held-out public tasks H Company reports 49.3% for 27B. The official private-set score is not out yet.
That is still a useful middle tier for teams that cannot pay frontier rates for overnight computer-use loops — closer to the open-weight lane we tracked when Qwen 3.8-27B landed than to Fable-class agents.
Open Weights — With a License Split You Should Not Miss
Both Holo4 sizes ship BF16, FP8, NVFP4, and 4-bit GGUF on Hugging Face. The licenses are not the same:
| Model | License (weights) | Base |
|---|---|---|
| Holo4-27B | CC BY-NC 4.0 (non-commercial) | Qwen3.8-27B (Apache 2.0) |
| Holo4-35B-A3B | Apache 2.0 | Qwen3.6-35B-A3B (Apache 2.0) |
If you are building a commercial product and want self-hosted weights, the 35B-A3B MoE is the one you can actually ship under Apache 2.0. The stronger 27B dense model is research/non-commercial unless you negotiate something else. H Company’s docs are explicit: check each model card. Trajectories and the evaluation dataset are Apache 2.0.
Holotron4 Nano is a separate story — H Company says applying the same recipe to Nemotron 3 Nano Omni lifts OSWorld CUA from 21.0% to 76.3% and AutomationBench from 19.4% to 35.6% (absolute points, vendor-reported). That is a post-training transfer claim, not a new frontier model.
How They Trained It (And Why the Harness Matters)
H Company describes an Agentic Task Factory that has produced about 10,000 verifiable tasks from docs and screenshots across web apps (~4,000), MCP servers (~3,000), and desktop/OS (~3,000). Training then goes:
- Supervised fine-tuning on 127B tokens — roughly three-quarters successful agent trajectories (desktop 45%, web 14%, MCP/API 12%, mobile 3%).
- Two RL LoRA experts — one for desktop/web, one for terminal/MCP/API.
- Equal-weight merge back into one model, no further training.
Alongside that, H Company rebuilt its harness after OSWorld 2.0 failure tags. The milestones it plots run from 6.1% on August 10 to 59.8% on September 8 before the 61.7% release — with the biggest fixes described as durable memory across hundreds of steps and a shell on the desktop machine itself. That is a reminder that computer-use leaderboards mix model quality and harness engineering. Comparing Holo4’s harness to Anthropic’s or OpenAI’s without reading the footnotes is how soft numbers become hard headlines.
What It Means for Indian Developers
For Bangalore, Hyderabad, and Pune teams automating browser QA, ERP clicks, or internal MCP tools, Holo4’s list prices matter more than the OSWorld brag. At $0.30–$0.40 input and $2–$3 output, you can run long agent loops that would be painful on Opus 5.5 — especially if most of the context is cacheable screenshots and tool traces at $0.03–$0.04 per million cached tokens.
Practical stack: use holo4-35b-a3b (Apache 2.0 weights if you self-host) for high-volume UI chores; reserve holo4-27b API calls for novel, long workflows; keep Opus 5.5 or GPT-6 Sol/Astra for failure-expensive coding agents. Replay trajectories at trajectories.hcompany.ai before you trust a benchmark line in a pitch deck.
Also budget sandboxing. Computer-use agents that drive real desktops inherit the same breakout class of risk we covered around NVIDIA’s OpenShell / Sentry stack and earlier agent sandbox incidents. Open weights do not remove that operational problem.
What’s Confirmed vs. Still Open
| Claim | Status |
|---|---|
| Holo4 27B and 35B-A3B live on H Models API; weights on Hugging Face | Confirmed by H Company / Hugging Face (Sep 28, 2026) |
| $0.40/$3.00 (27B) and $0.30/$2.00 (35B-A3B) per 1M tokens; cache $0.04 / $0.03 | Confirmed on H Company launch page |
| 27B weights CC BY-NC 4.0; 35B-A3B Apache 2.0 | Confirmed on Hugging Face model cards |
| OSWorld 2.0 61.7% at $1.22/task; AutomationBench 45.4% | Vendor-reported in H Company’s harness |
| Opus 5.5 / GPT-6 Astra / Fable comparator scores | Cited public / other-vendor figures — different harnesses |
| 7,366 open trajectories behind the scores | Confirmed dataset on Hugging Face + trajectory viewer |
| AutomationBench private-set score; DSpark drafter checkpoints | Promised / pending — not published yet |
Frequently Asked Questions
What is Holo4?
Holo4 is H Company’s September 28, 2026 release of generalist computer-use agent models: a 27B dense model and a 35B-A3B MoE, plus Holotron4 Nano built on Nemotron 3 Nano Omni. They act through GUIs, code, MCP, and APIs.
How much does the Holo4 API cost?
On the H Models API: Holo4 27B is $0.40 input / $3.00 output per million tokens ($0.04 cached). Holo4 35B-A3B is $0.30 / $2.00 ($0.03 cached). Both list a 256K context window on the launch page.
Are Holo4 weights open source?
Weights are downloadable. Holo4-35B-A3B is Apache 2.0. Holo4-27B is CC BY-NC 4.0 — non-commercial. Check the Hugging Face cards before you ship a product on local weights.
Does Holo4 beat Claude Opus 5.5 on computer use?
Not on H Company’s own OSWorld 2.0 comparison: 61.7% for Holo4 27B versus 81.8% for Opus 5.5 (Anthropic harness, max effort, as cited). H Company’s claim is cost-competitive middle-tier performance, not a frontier takeover.
Where can I try Holo4?
Models API quickstart at hub.hcompany.ai, managed Agents API, HoloDesktop CLI, and the free HoloTab Chrome extension. Evaluation trajectories are at trajectories.hcompany.ai and on the Hcompany/trajectories Hugging Face dataset.
1 comment