Overview
Holo4 is a family of general computer-use agent models released by H Company on September 28, 2026, including two sizes: 27B dense and 35B-A3B mixture-of-experts, both supporting a 262,144 token context. The family can complete tasks across desktop, web, mobile, and API, can look at the screen and then click and type, write and run its own code, call MCP or REST APIs, and choose the most suitable interface based on the task. The same model runs in the same way on desktop, web, Android, code sandboxes, and enterprise APIs, without needing to switch models for different platforms. The release of Holo4 was accompanied by the debut of Holotron4 Nano based on NVIDIA Nemotron 3 Nano Omni, which is 30B-A3B, follows the NVIDIA Open Model Agreement, and is only available in BF16 and FP8.
Key Features
- Dual sizes and long context: Holo4 offers two sizes: 27B dense (based on Qwen3.8-27B) and 35B-A3B mixture-of-experts (based on Qwen's MoE foundation), both supporting a 262,144 token context.
- Unified cross-platform operation: The same model runs in the same way on desktop, web, Android, code sandboxes, and enterprise APIs, without needing to switch models for different platforms.
- Autonomous selection of multiple interfaces: The model can look at the screen and then click and type, write and run its own code, call MCP or REST APIs, and choose the most suitable interface based on the task.
- Supervised fine-tuning plus online reinforcement learning: Training follows supervised fine-tuning plus online reinforcement learning: the SFT dataset is about 127 billion tokens, about three-quarters of which are successful agent trajectories; reinforcement learning trains two LoRA experts, one responsible for desktop and web, and one responsible for terminal, MCP, and API, then merges them with equal weight.
- Agentic Task Factory: The internal Agentic Task Factory constructs environments and verifiable tasks from documents, screenshots, and real software, and has already produced about ten thousand tasks.
- Replayable trajectories and multiple deployment formats: All benchmark runs are released as replayable trajectories in the trajectories dataset on Hugging Face; the weights are available on Hugging Face in BF16, FP8, NVFP4, and 4-bit GGUF, and can be deployed with Transformers, vLLM, SGLang, Docker, or llama.cpp, with an OpenAI-compatible hosted API also provided.
Use Cases
- Execute general computer operation tasks across desktop, web, Android, and code sandboxes, such as looking at the screen and then clicking and typing, and writing and running code.
- Call enterprise systems through MCP or REST APIs, and choose the most suitable interface based on the task to complete automated workflows.
- Handle multi-step, long-duration agent tasks in scenarios such as OSWorld 2.0 long workflows.
- Use verifiable tasks constructed by the Agentic Task Factory for agent capability evaluation and replay analysis.
- Integrate into enterprise applications through an OpenAI-compatible hosted API or local deployment.
Pros
- The same model runs across desktop, web, Android, code sandboxes, and enterprise APIs, with no need to switch models for different platforms.
- Both sizes support a 262,144 token context, suitable for long workflows and complex tasks.
- In official evaluations, 27B scores 85.2% on OSWorld at $0.08 per task, while 35B-A3B scores 80.8% at $0.05.
- 27B scores 85.1% on AndroidWorld and 61.7% on OSWorld 2.0 long workflows at $1.22 per task.
- 35B-A3B is Apache 2.0 and commercially usable, with weights available in BF16, FP8, NVFP4, and 4-bit GGUF formats.
- All benchmark runs are released as replayable trajectories in the trajectories dataset on Hugging Face, making reproduction and review easier.
Pricing
API pricing per million tokens: 27B input $0.40, output $3.00; 35B-A3B input $0.30, output $2.00; cached input as low as $0.04. Official evaluation cost per task: on OSWorld, 27B is $0.08 and 35B-A3B is $0.05; on OSWorld 2.0 long workflows, 27B is $1.22 and 35B-A3B is $0.61; on AutomationBench, 27B is $0.05 and 35B-A3B is $0.02. In terms of licensing, 35B-A3B is Apache 2.0 and commercially usable, while 27B is CC BY-NC 4.0 for non-commercial use; Holotron4 Nano is 30B-A3B, follows the NVIDIA Open Model Agreement, and is only available in BF16 and FP8. For pricing information not provided, refer to the official website.
Summary
Holo4 is a family of general computer-use agent models released by H Company on September 28, 2026, including two sizes: 27B dense and 35B-A3B mixture-of-experts, both supporting a 262,144 token context, and capable of completing tasks across desktop, web, mobile, and API. Official evaluations cover OSWorld, AndroidWorld, OSWorld 2.0, AutomationBench, and the internal Agentic Task Factory set, and also publish per-task costs and API pricing. 35B-A3B is Apache 2.0 and commercially usable, while 27B is CC BY-NC 4.0 for non-commercial use; the weights are available in multiple formats and support multiple deployment methods.
Version History
- H Company releases Holo4, a generalist computer-use agent model series (2026-09-28): It ships in 27B dense and 35B-A3B MoE sizes, both with a 262,144 token context, and the same model runs on desktop, web, Android, a code sandbox and enterprise APIs. Holo4 27B scores 85.2 percent on OSWorld at 0.08 dollars per task. Holotron4 Nano, built on Nemotron 3 Nano Omni, shipped alongside.