The Hermes Index Ranks Agent Models on Score and Cost. The Lab That Built It Just Raised $90M.
Nous Research's new leaderboard puts a dollar figure next to every agent score. The numbers are useful, but the harness owns them.

What the Hermes Index is
Nous Research, the open-source AI lab founded in 2023, launched the Hermes Index on October 6, 2026. It ranks 14 frontier models on how well they perform as agents inside the lab's own Hermes Agent framework, and reports a mean cost per task beside each score. The pairing is the whole point: agent leaderboards usually report capability, and capability without a price tag is how teams pick a model they cannot afford to run.
The methodology is straightforward. Four agent benchmark suites run through the same Hermes Agent runtime: the new Hermes Bench (150 tasks across 25 categories, built by Nous for this index), TerminalBench 4, TerminalBench Science, and SkillsBench. Each model gets one attempt per task (pass@1), and reasoning effort is set to high wherever the model supports it. Scores and per-task costs are averaged across the four suites. One day after launch, Nous announced a $90 million funding round, bringing its total raised to roughly $160 million at a $1.5 billion valuation, with Nvidia and Microsoft's venture arm M12 among the backers. The lab says the money goes toward enterprise deployments of its Hermes technology, while the core framework stays open source under the MIT license.
The leaderboard, with the column that matters
Here are the published figures, with a column I added myself: points per dollar, computed from Nous's reported score and mean cost per task.
| Rank | Model | Hermes Index | Mean cost / task | Points per dollar |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 63.31 | $4.99 | 12.7 |
| 2 | GPT-6 Astra | 56.25 | $11.61 | 4.8 |
| 3 | Claude Sonnet 5.5 | 53.14 | $2.82 | 18.8 |
| 4 | GPT-6 Sol | 44.10 | $2.23 | 19.8 |
| 5 | Grok 4.7 | 39.32 | $10.77 | 3.7 |
| 9 | DeepSeek V4.1 Flash | 36.91 | $0.259 | 142.5 |
| 14 | Ling 3.0 Flash | 21.56 | $0.054 | 399.3 |
Nous also names six models on the efficiency frontier, where no other entry is both cheaper and higher-scoring: Claude Opus 5.5, Claude Sonnet 5.5, GPT-6 Sol, DeepSeek V4.1 Flash, GPT-6 Luna, and Ling 3.0 Flash. If you are choosing a model this week, the frontier is the shortlist. Everything else is dominated on both axes.
The score-per-dollar math
Three comparisons carry the practical weight:
Sonnet 5.5 is the value pick of the top tier. It scores 53.14 against Opus 5.5's 63.31, which is 84% of the leader's score, and it costs $2.82 per task against $4.99, which is 57% of the leader's cost. If your tasks are the kind where a 16% score gap rarely changes the outcome, you keep most of the capability and cut the bill nearly in half.
GPT-6 Astra is the most expensive point on the board. At $11.61 per task for a score of 56.25, it costs more than twice per task what Opus 5.5 costs while scoring 7 points lower. Per index point, Astra runs about $0.206, against $0.079 for Opus 5.5 and $0.053 for Sonnet 5.5. That does not make Astra a bad model. It makes it a model whose agent runs are token-hungry under this harness, and you should know that before you route production traffic through it.
The open models redraw the curve. DeepSeek V4.1 Flash delivers 58% of the leader's score at 5% of the cost. Ling 3.0 Flash manages 21.56 points at $0.054 per task, which is roughly four hundred points per dollar. These are not the models you pick for the hardest agentic coding task in your company. They are the models you pick for the ten thousand routine tasks where a 36.91 is good enough and the invoice is what decides.
Why cost per task belongs next to every agent score
Teams that run agents in production already know the score is the cheap part of the evaluation. A model that scores two points higher and burns four times the tokens is not better for your workload. It is better for the leaderboard and worse for your budget. The Hermes Index is the first major agent leaderboard I have seen that treats the invoice as a first-class metric, and it should not be the last.
This is also the correct unit of comparison. Price per million tokens tells you the rate; cost per task tells you the bill. A verbose model at a low rate can cost more per completed task than a terse model at a high one. Yesterday's piece on this site about benchmark validity argued that benchmarks smuggle in their designers' assumptions; the Hermes Index at least makes one of its assumptions, that cost matters, explicit and measurable.
The number is a property of the harness, not the model
Here is the part that keeps me from treating the index as a verdict. Per-task cost is not a property of the model alone. It is a joint product of the model, the harness, the reasoning effort setting, and the task mix. Artificial Analysis measured Claude Opus 5.5 at $5.98 per Intelligence Index task at max reasoning with fallback, and GPT-6 Astra at $3.26 on the same harness. Under Hermes Agent with reasoning effort set to high and one attempt per task, Astra costs $11.61 per task and Opus 5.5 costs $4.99. Same models, different rigs, opposite cost order.
That is not a contradiction. It is the whole lesson. A per-task cost figure tells you what running that model costs through that harness on those tasks. Change the harness, change the effort setting, change the task distribution, and the number moves. Anyone quoting $11.61 as "the cost of GPT-6 Astra" is misreading the index. It is the cost of GPT-6 Astra as an agent inside Hermes Agent, with the dials set where Nous set them.
In favour
- Cost per task is the right column. Agent economics are the part of model selection most teams get wrong, and putting the invoice next to the score is a genuine improvement over score-only leaderboards.
- The methodology is legible. Four named suites, one harness, pass@1, high reasoning effort. You can disagree with the choices because you can see them.
- The Pareto framing is honest. Publishing the efficiency frontier instead of just a ranked list tells teams that "best" depends on the budget constraint, which is how procurement actually works.
- Open source stays open. Keeping the Hermes framework under the MIT license while raising enterprise money is the arrangement that lets outsiders audit the rig.
Against
- One harness, one verdict. Every number is conditioned on Hermes Agent. Models may order differently in other agent frameworks or in your production stack, and the index cannot tell you which.
- Fourteen models, vendor-selected. The field is whoever Nous chose to run. Missing models and missing suites are invisible in a leaderboard.
- No variance reported. A single mean cost per task with no spread hides how much task difficulty varies. Averages over 150 tasks can conceal a distribution where a model is cheap on the easy ones and ruinous on the hard ones.
- The timing invites skepticism. A leaderboard launch followed within a day by a $90M raise at a $1.5B valuation is, at minimum, a marketing event. Treat the marketing and the measurement separately, and verify the measurement.
What would make me wrong
- Independent replication outside the Hermes Agent stack reproduces the same ranking and similar per-task costs, especially the Astra $11.61 figure.
- Teams running agents at scale report that cost-per-task rankings predicted their actual invoices better than price-per-token comparisons did.
- Nous publishes per-task cost distributions, not just means, and the ordering survives.
- The opposite: a credible independent run shows the cost ordering flips across harnesses, which would confirm that per-task cost is too harness-dependent to standardize, and the index becomes a curiosity rather than a tool.
Sources
- Nous Research launches Hermes Index to benchmark agentic AI models — launch details, leaderboard figures, $90M round, MIT license
- Nous Research's Hermes Index Ranks AI Agents by Score and Real Cost — four-suite methodology (Hermes Bench, TerminalBench 4, TerminalBench Science, SkillsBench), pass@1, high reasoning effort, Pareto frontier
- Claude Opus 5.5 Leaps Forward — Artificial Analysis per-task costs ($5.98 Opus 5.5, $3.26 GPT-6 Astra) on the Intelligence Index, showing cost is harness-dependent
- Gemini 4 Argon vs Claude Opus 5.5 — price per token vs price per completed task distinction
Related on Everyday Data Science:
- Yesterday: 56 Benchmarks, 53 Models, and an Uncomfortable Result — Stanford's validity audit is the lens to read the Hermes Index through
- MCP, Explained for Data People
- Your Agent's Memory Is Too Long. Cutting It Improves Both Accuracy and Cost.
- The 100-Line Eval Harness That Catches Agent Regressions
The leaderboard says Opus 5.5 wins. The invoice says Sonnet 5.5, or DeepSeek, or Ling, depending on your budget. Which number would you optimize for, and what would it take for you to trust it?
About the writer
Data Scientist & AI Researcher
Data scientist and AI researcher at Pace University. I coined Artificial Frictional Unemployment, and built the first machine learning model for crop yield prediction in Sierra Leone. Author of Understanding Agentic AI. I write about agentic systems and applied ML, with a bias toward what actually works, and who gets left out when it doesn't.
Found this useful? Passing it on to someone who builds is the best way to help the publication grow.
Built something worth sharing? Write it up for us →