AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

H Company says its new Holo4 agent models can work through graphical interfaces, code, MCP tools and APIs, and are available in 27B and 35B-A3B versions. The company reports benchmark results and lower estimated task costs, but comparisons use different evaluation setups and do not establish performance across all real business workflows.

H Company has announced Holo4, a series of agentic models designed to carry out software tasks through graphical interfaces, code, MCP tools and APIs. The company says the models are available in 27B dense and 35B-A3B Mixture of Experts versions through its H Models API, a release aimed at business workflows that can cross between different kinds of software access.

Holo4 can click and type on a screen, write and run code, or call tools exposed through the Model Context Protocol (MCP) and APIs, according to H Company. The company says the same model can run on desktops, the web, Android, a code sandbox and business APIs, without requiring developers to select a separate model for each platform. It also released Holotron4 Nano, an updated version of Holotron 3.

The company reports that Holo4 27B scored 61.7% on OSWorld 2.0, a benchmark for computer use, while its 35B-A3B model scored 30.9%. H Company compares the 27B result with 81.8% for Opus 5.5. The supplied report does not give a corresponding Opus score for the 35B-A3B model. H Company also says Holo4 competes with frontier models on AutomationBench, which tests API use.

H Company presents these results alongside estimates that Holo4 costs less per task than larger competitors. Its cost figures are calculated from agent-run token use and model pricing, and the company notes that benchmark releases, harnesses and task subsets differ. For AutomationBench, Holo4 and two Qwen models were evaluated on version 1.0.6 in H Company’s internal harness, while other models’ cited scores and costs come from public-set results and a leaderboard using a private set.

At a glance
announcementWhen: Announced in the source report; no publ…
The developmentH Company announced Holo4, a series of agentic models designed to use multiple software interfaces, and released the models through its API.

One Agent Across Software Interfaces

Many work processes involve more than one way of interacting with software: a person or agent may need to navigate a screen, run a script and retrieve information from a service API. H Company’s central product claim is that one model can combine those methods, which could reduce the need to route tasks among separate agents built for individual interfaces.

The practical value depends on whether the system can complete varied workflows reliably, not only whether it can perform well on selected benchmarks. H Company says it trained Holo4 with supervised and reinforcement learning across environments and tasks, including tasks generated by its Agentic Task Factory. Its published examples include creating models in FreeCAD and building a game in Godot. These demonstrations illustrate the intended use, but the supplied material does not establish how often the models succeed across customer deployments or how much human review those tasks require.

From Benchmark Scores to Workflows

Holo4 builds on an earlier H Company model and is presented as a generalist agent for business software. The company contrasts its approach with models trained mainly for one interface type: a screen-oriented model may lack access when a task depends on an API, while an API-focused agent may not be able to operate software that offers no API. Holo4 is intended to switch among these channels as a task requires.

H Company says the models improve on their Qwen base models and has released benchmark trajectories for public inspection through a viewer and a downloadable dataset. Its report describes the OSWorld cost estimates as based on input and output tokens for each run, with Holo4 priced at H Models API rates. Comparisons for other models draw on different sources, including model cards, company-run harnesses and official leaderboard entries, so the figures are not all results from a single controlled evaluation.

““Real work is not siloed that way, and a single business task can require combining these different approaches.””

— H Company

Limits of the Published Evaluations

The supplied report does not identify its publication date, provide independent validation of the benchmark or cost claims, or show how results vary across repeated runs. It also does not quantify reliability in deployed business workflows, the need for human supervision, or the full operating costs beyond the token-based estimates described for some comparisons.

H Company says it will report Holo4 on AutomationBench’s private set once the model has been evaluated there. Until then, its AutomationBench comparison combines results drawn from different sets and evaluation sources. The company’s report also does not give enough detail to establish that its examples represent typical task success rates.

Private-Set Results and Deployment Evidence

The next stated benchmark milestone is H Company’s evaluation of Holo4 on AutomationBench’s private set, after which it says it will report the result. Developers can access the two Holo4 models through the H Models API, while the company lists model collections and benchmark trajectories for download or replay.

Further evidence about performance on varied business tasks, reliability over long workflows and costs under comparable evaluation conditions would help clarify how well the release’s claims translate beyond the published tests. The supplied material does not give a schedule for those results.

Key Questions

What did H Company announce?

H Company announced Holo4, a series of agentic models intended to work across graphical interfaces, code, MCP tools and APIs. It also announced Holotron4 Nano, an updated version of Holotron 3.

What Holo4 model sizes are available?

The announced series includes a 27B dense model and a 35B-A3B Mixture of Experts model. H Company says both are available through its H Models API.

How did Holo4 score on OSWorld 2.0?

H Company reports scores of 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B. It cites 81.8% for Opus 5.5 as a comparison for the 27B result.

Are the cost comparisons directly comparable?

Not fully, based on the supplied report. H Company says estimates draw on different model sources, pricing assumptions, harnesses and task subsets. It describes some figures as token-based estimates and says Holo4 has not yet been evaluated on AutomationBench’s private set.

Source: rss

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

IEEE Rolls Out Large Language Models Training Course

IEEE has announced a new training course focused on the development and deployment of large language models, aimed at professionals and researchers.

IdeaNavigator AI: One Evidence-Mined Idea a Day

IdeaNavigator AI now publicly releases one evidence-mined product idea daily, transforming idea validation by sourcing from real online complaints and feedback.

Cross-platform buyer history for multi-marketplace resellers

Resellers across eBay, Poshmark, and Mercari are testing a manual cross-platform buyer ledger to identify repeat customers and improve decision-making.

What AI Did To Stackoverflow In A Graph

A new graph illustrates how AI models have transformed Stack Overflow activity, highlighting shifts in question-answer dynamics and user engagement.