Holo4 Tests a Generalist Route to Computer-Use Agents

Hcompany’s Holo4 release combines GUI control, code execution, MCP, and API calls in one agent model. The more important question is how teams should evaluate that flexibility across long, stateful workflows.

Source artwork for Holo4: powering generalist computer-use agents
Source artwork · Hugging Face / credited contributors ↗
THE SHORT VERSION

Holo4’s multi-interface design is promising, but production evaluation must treat the harness, permissions, and cross-interface state checks as part of the system.

Hcompany has introduced Holo4, a family of computer-use models designed to move between graphical interfaces, code, MCP tools, and APIs without requiring a separate specialist model for each interaction mode. The release includes a 27-billion-parameter dense model and a 35B-A3B mixture-of-experts model, plus an updated small model called Holotron4 Nano. Model weights and multiple quantized formats are available through Hugging Face, while hosted access is offered through Hcompany’s own API.

The broad-interface design matters because useful automation rarely stays inside one clean boundary. A workflow might begin in a browser, continue through a spreadsheet API, require a shell command, and finish by checking the visible application state. Holo4’s central proposition is that one policy can choose among those surfaces. That is more ambitious than optimizing only for clicking or only for structured tool calls, but it also expands the number of ways a run can fail.

Generality is a routing problem

Giving an agent several interfaces does not automatically make it effective. The agent must select the right interface at each step and preserve a coherent understanding of the task while switching between them. A GUI may be the only route through an undocumented dialog, whereas an API can be safer and more precise for a bulk update. Code may transform data efficiently, but the visual application may still be the authoritative place to confirm the result.

That turns interface selection into part of the reasoning task. Consider an expense-processing workflow: the agent could read invoices through a document interface, normalize line items with code, submit records through an API, then open the accounting application to verify status badges and exceptions. Evaluation should therefore measure more than whether the final screen looks plausible. It should record which surface was used, whether state remained consistent across surfaces, and whether the agent recovered when one route became unavailable.

Hcompany says the same Holo4 model can operate on desktops, the web, Android, coding sandboxes, and business APIs. This avoids an explicit model-selection layer, though a production system still needs orchestration around credentials, permissions, retries, and audit logs. A generalist model simplifies one architectural decision; it does not remove operational controls.

Read the benchmark claims with their harnesses

The announcement reports a 61.7% score for Holo4 27B on OSWorld 2.0 and 30.9% for Holo4 35B-A3B. It also compares performance and estimated cost across OSWorld 2.0 and AutomationBench. Those numbers are useful reference points, but the release itself notes important comparability limits: model results come from different task releases, harnesses, subsets, pricing assumptions, and cache policies. For AutomationBench, some Holo4 measurements use an internal harness and public-set scores, while other entries reference a private-set leaderboard.

This is especially important for computer-use systems because the harness is part of the effective product. Hcompany describes rebuilding its execution loop to provide persistent memory over hundreds of steps and a shell on the desktop machine. Changes like those can affect completion rates independently of model weights. Timeouts, screenshot resolution, action validation, context compression, and recovery rules can all change the outcome.

A fair internal evaluation should hold the environment and harness constant across candidate models. Teams should publish or retain at least four result layers: end-task success, failure category, total interaction cost, and number of risky or irreversible actions. Median performance is insufficient for unattended automation if a small fraction of runs can send the wrong message, overwrite a record, or become trapped in a costly loop.

Training breadth creates both value and uncertainty

Hcompany attributes the release to supervised and reinforcement learning across roughly 10,000 generated tasks spanning web applications, MCP servers, desktops, and hybrid environments. Its Agentic Task Factory reportedly constructs interactive environments and verifiable assignments from documentation and software artifacts. The company also exposes trajectories for public benchmark runs, which gives evaluators material for inspecting action sequences rather than relying only on aggregate scores.

Synthetic and generated tasks can widen coverage efficiently, particularly for rare UI states. The open question is how closely their distribution matches a specific organization’s work. Real desktops contain stale sessions, ambiguous labels, slow responses, permission prompts, and locally customized processes. A model trained broadly may still need careful qualification on those conditions. Public trajectories help diagnose behavior on the disclosed benchmarks, but they are not evidence that every business workflow will transfer unchanged.

A practical adoption checklist

Start with a bounded workflow whose completion can be verified from underlying state, not merely from a screenshot. Give the agent read-only access first, then introduce reversible writes behind explicit policy checks. Include tasks that force interface switching, because that is the core benefit being claimed. Also test degraded conditions: missing APIs, shifted controls, expired sessions, partial tool failures, and contradictory state between a GUI and its backend.

Cost should include the entire run. The examples in the announcement involve dozens of calls and, in some cases, more than a million tokens. Even when per-token pricing is attractive, long trajectories can accumulate latency and expense. Track retries, context growth, tool execution, and human review alongside model charges.

Finally, separate capability from authority. An agent that can operate many interfaces should receive narrower permissions, not broader default access. Use scoped credentials, sandboxed code execution, confirmation for irreversible steps, and an immutable event log. Holo4 makes a notable case for one model spanning the interfaces of modern work. Whether that becomes dependable automation will be decided by harness quality, verification, and permission design as much as by benchmark position.

Source: Holo4: powering generalist computer-use agents ↗. How we write

← Back to all articles