Nous Research Has Launched Hermes Benchmarking Index

The company also secured $90 million in new funding, bringing its total valuation to $1.5 billion.

Updated on Oct. 7, 2026 in Artificial Intelligence

Isometric editorial illustration of a steel scale weighing silicon wafers, symbolizing AI performance and cost efficiency benchmarks.
Nous Research launched its Hermes Index benchmarking tool to evaluate AI model performance, coinciding with a $90 million funding round led by Nvidia and M12. AI Illustration. Upload story photo >

Live Poll

Do you prioritize benchmark scores when choosing AI tools for your work or business?

Nous Research has officially launched the Hermes Index, a new benchmarking tool designed to evaluate the performance and cost-efficiency of 14 frontier agentic AI models. The launch arrives alongside the announcement of $90 million in new funding led by Nvidia and M12.

Why it matters

This new benchmarking tool provides a standardized framework for businesses to compare AI capabilities against operational costs, while the fresh capital will be used to scale enterprise deployments of Hermes technology.

The Hermes Index ranks 14 models using Hermes Bench, Terminal-Bench 4.0, and SkillsBench, with Claude Opus 5.5 scoring 63.31 at $4.99 per task compared to GPT-6 Astra, which scored 56.25 at an average cost of $11.61 per task.

The players

Nous Research

An artificial intelligence laboratory focused on developing agentic AI models and standardized benchmarking tools.

Nvidia

A global technology company that designs graphics processing units and provides the hardware infrastructure powering modern AI development.

M12

The venture capital arm of Microsoft that invests in early-stage technology companies with a focus on enterprise solutions.

The details

The evaluation harness, which began development in early 2026, measures models across four distinct test suites to determine their effectiveness in agentic workflows. By providing concrete performance metrics alongside task costs, Nous Research aims to clarify the value proposition of competing AI systems for professional users.

Timeline

  1. Nous Research was founded in 2023.

  2. The Hermes Agent evaluation harness began development on February 25, 2026.

  3. The Hermes Index was launched on October 6, 2026.

  4. Nous Research secured $90 million in new funding on October 7, 2026.

The Tech Race

This development shifts the competitive landscape by prioritizing standardized performance-to-cost metrics over the raw model capability race typical of the last two years. It establishes a new paradigm where developers must balance agentic autonomy with predictable per-task operating expenses.

Enterprises and software developers can now use these standardized benchmarks to select AI models that better align with their specific budgetary requirements and performance needs. This transition reduces the guesswork previously involved in choosing between premium models like Claude Opus and cost-effective alternatives like DeepSeek.

The takeaway

The move toward transparent benchmarking highlights a growing industry focus on the economic viability of AI agents in real-world workflows. Professionals should prioritize evaluating models based on cost-per-task metrics to ensure long-term operational sustainability.

Further reading

For more information on the current state of AI tools, see our Artificial Intelligence section.

Source note: This article includes information reported by Crypto Briefing.

Live Poll

Do you prioritize benchmark scores when choosing AI tools for your work or business?