🔍 Read the full analysis: From AI To Action: Holo4 For Computer-Use Agents on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
H Company has released Holo4, a pair of open-weight models designed to handle software tasks through graphical interfaces, code and tools. The company reports a 61.7% OSWorld 2.0 score for the 27B model, but the results have not been independently verified, and the 35B-A3B model scored 30.9% on the same benchmark.
H Company has released Holo4, a series of open-weight models built to operate software through graphical interfaces, code and MCP or API tools, as detailed in the original analysis. The company reports that its 27B model scored 61.7% on OSWorld 2.0; the result could make the release relevant to developers exploring lower-cost software agents, though independent verification is still pending.
The release includes a 27B dense model and a 35B-A3B Mixture of Experts model. H Company says both are available through its H Models API and as downloads on Hugging Face in FP16, FP8 and GGUF formats. The models are intended to choose among screen interaction, code execution and tool calls within a task, rather than rely on just one interface.
On OSWorld 2.0, H Company reports scores of 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B. Its announcement compares the 27B result with 81.8% for Opus 5.5, which it identifies as the strongest closed model in the comparison. Those figures are company-reported; the announcement says benchmark releases, harnesses and task subsets differ across models.
H Company says it trained the models using supervised and reinforcement learning across tasks and environments, including tasks generated by its Agentic Task Factory. It has also published the trajectories behind its public benchmark scores at trajectories.hcompany.ai and on Hugging Face. These records allow outside reviewers to inspect the steps used to produce the reported results, though publication alone does not establish that the results will reproduce independently.
One Model Across Software Interfaces
Business software tasks often combine actions in a graphical application with code or API operations. A model that can move among those interfaces could reduce the need to connect separate agents for each stage. H Company presents Holo4 as a general-purpose computer-use agent that can click and type, write and run code, and call MCP or API tools.
If independent evaluations confirm the reported performance, open weights could give developers more flexibility to inspect, adapt or host an agent themselves. H Company also claims lower costs than larger closed models, but the comparison depends on its stated pricing assumptions and evaluation setup. The release is therefore a testable claim about capability and cost, not yet evidence of dependable performance across routine business workflows.
Top picks for "action holo4 computer"
As an affiliate, we earn on qualifying purchases.
From Holo1 to Holo4
Holo4 follows H Company’s earlier agentic model and is built on Qwen base models, according to the company’s benchmark notes: Qwen3.8 27B for the dense version and Qwen3.6 35B-A3B for the MoE version. The announcement also describes side-by-side examples in FreeCAD and Godot, using the same prompt and harness to show changes from the Qwen base model. Such examples illustrate selected tasks; they do not substitute for broad independent testing.
H Company’s cost comparisons use H Models API rates for a single Holo4 run and Alibaba Cloud list prices for Qwen, with cache hits priced at 20% of input for the MoE model. Comparisons involving GPT and Opus use effort sweeps from OpenAI launch data. The company notes that models may have been tested with different releases, harnesses and task subsets, making direct comparisons difficult.
“Real work is not siloed that way, and a single business task can require combining these different approaches.”
— H Company
Benchmark Results Await Replication
The published benchmark figures are self-reported by H Company, and independent reproductions are not yet described in the source material. On AutomationBench, the company says it used its internal harness, version 1.0.6, while some competing scores came from the public set and cost figures from a leaderboard running the private set. Holo4 has not yet been evaluated on that private set.
The reason for the large OSWorld 2.0 score gap between the two Holo4 models—61.7% for the 27B version and 30.9% for the 35B-A3B version—is not explained in the announcement. The material also does not establish how reliably Holo4 handles varied business workflows outside benchmark tasks and selected demonstrations. Published trajectories make review possible, but do not resolve those questions on their own.
Private Set Tests and Reproductions
H Company says it plans to report Holo4’s results on the AutomationBench private set once that evaluation is complete. The next evidence to watch for is independent benchmark submissions or reproductions using the released weights and task trajectories. Developers can access the models through the H Models API or download them from Hugging Face; broader deployment performance and costs remain to be established.
Key Questions
What is Holo4?
Holo4 is H Company’s series of open-weight agentic models designed to operate software using graphical interfaces, code, MCP and APIs.
Which Holo4 models are available?
The release includes a 27B dense model and a 35B-A3B Mixture of Experts model. H Company says they are available through its API and on Hugging Face in FP16, FP8 and GGUF formats.
How did Holo4 score on OSWorld 2.0?
H Company reports 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B. The scores are company-reported and have not been independently verified in the supplied material.
Can outside reviewers check the benchmark runs?
H Company says it has published the trajectories behind its public benchmark scores at trajectories.hcompany.ai and on Hugging Face. Reviewers can inspect those records, but independent reproduction remains a separate step.
What remains to be tested?
Independent benchmark results, Holo4’s performance on AutomationBench’s private set and its reliability on varied business workflows remain open questions. The company says it will report the private-set evaluation when complete.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
