Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
Podcast39 min 45 sec
Listen to Episode
Note: AI-generated summary based on third-party content. Not financial advice. Read more.
Quick Insights

Investors should target the emerging AI Evaluation & Benchmarking sector, focusing on picks-and-shovels testing platforms like VALS that provide independent verification for corporate AI adoption.

Capitalize on enterprise cost-management infrastructure by watching Stripe, which is strategically positioned to capture multi-model billing and traffic layers through its acquisition of OpenRouter.

Prioritize investments in token-optimization tools and fixed-subscription developer agents like Cognition’s Devin that help enterprises curb soaring computing costs.

Exercise near-term caution with Meta Platforms, Inc. (META), as discrepancies between Llama 4's public benchmark scores and its actual real-world performance challenge its open-source competitive moat.

Weigh the margin risks of private frontier labs like Anthropic, where massive compute overhead and heavy token consumption across Claude Sonnet and Claude Opus continue to compress provider profitability.

Detailed Analysis

Meta Platforms, Inc. (META)

  • Meta released Llama 4, which showed strong performance on major open-source public benchmarks where questions and rubrics are public.
    • However, independent testing on private, held-out benchmarks indicated that the model underperformed expectations.
    • This highlighted the growing disconnect between self-reported AI performance using public benchmarks and actual real-world capability.

Takeaways

  • Investors should be cautious when evaluating foundation model progress based solely on marketing or public leaderboards, as models can inadvertently overfit or hack public benchmarks.
  • Meta’s competitive moat in open-source AI depends heavily on real-world utility rather than synthetic benchmark scores.

Anthropic (Private)

  • Anthropic’s models, including Claude, Opus, and Sonnet, are seeing rapid adoption inside Fortune 10 enterprises for software engineering workflows (e.g., Claude Code).
  • Anthropic operates on narrow gross margins to serve these frontier models due to massive compute and serving costs.
    • In specific enterprise coding tests, Claude Sonnet proved more expensive in practice than Claude Opus in certain contexts because it was significantly more token-hungry.

Takeaways

  • While Anthropic continues to capture enterprise mindshare, high inference costs and margin compression remain key financial risks for leading frontier model providers.
  • Model pricing dynamics are non-intuitive; higher-tier intelligence models can sometimes be more cost-effective than mid-tier models if they complete tasks with fewer total tokens.

Stripe (Private)

  • Stripe acquired OpenRouter, a prominent platform that functions as a model gateway and aggregator for routing API requests across various AI models.
  • Model routing is becoming essential as enterprises seek to optimize costs by sending simple requests to cheaper models and complex tasks to frontier models.

Takeaways

  • Stripe's acquisition positions it well to capture value from the infrastructure and billing layers of multi-model enterprise architectures.

AI Evaluation & Benchmarking (Investment Theme)

  • As AI becomes a multi-trillion-dollar industry, an independent ecosystem of rating, testing, and auditing firms (similar to credit rating agencies or accounting audit firms) is emerging.
  • Startups like VALS are building specialized infrastructure to evaluate complex, multi-day agentic workflows across domains such as finance, software engineering (VibeCode, ValSmith), and recursive self-improvement (RSI).
    • Business model integrity requires a strict separation between evaluating models and selling training data to avoid conflicts of interest reminiscent of historical accounting scandals.

Takeaways

  • Independent AI verification and safety auditing represent a critical, high-growth picks-and-shovels infrastructure category in the AI value chain.
  • Look for investment opportunities in testing, verification, and governance platforms that serve both frontier AI labs and enterprise buyers seeking objective performance validation.

Enterprise AI Spend & Model Economics (Investment Theme)

  • Enterprise AI budgets are shifting rapidly, with instances where developer token spend has approached or even exceeded base salary costs (e.g., teams consuming $100 to $300+ per engineer daily, or spending over $1.5 million in tokens during heavy internal testing).
  • Enterprises are struggling to quantify the exact return on investment (ROI) of their AI spending and require internal benchmarking tools to determine the optimal model for their specific proprietary data and workflows.
    • Specialized, token-efficient agents (such as Cognition’s Devin) and fixed-subscription pricing models often provide better unit economics than unmanaged, raw token-based API usage.

Takeaways

  • As enterprise AI spending balloons, tools that provide cost visibility, token optimization, and performance benchmarking will see strong enterprise demand.
  • Companies that fail to monitor and optimize their model routing risk massive cost overruns without corresponding productivity gains.
Ask about this postAnswers are grounded in this post's content.
Episode Description
a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult problems: how do you actually measure whether a model is getting better? As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. They unpack why self-reported model scores can be misleading, how VALS evaluates models in the hours before a release, and why measuring increasingly agentic systems means testing work that can unfold over hours, days, or even weeks. They also explore why evals are becoming critical for enterprises trying to understand the ROI of AI, what happens if token spend begins to rival employee salaries, and how evaluations could eventually provide a shared language for everything from model routing and recursive self-improvement to AI policy and international coordination. Resources: Follow Rayan Krishnan on X: https://x.com/RayanKrishnan Follow Ben Horowitz on X: https://x.com/bhorowitz Follow Jennifer Li on X: https://x.com/JenniferHli   Stay Updated: Find a16z on YouTube: YouTube Find a16z on X Find a16z on LinkedIn Listen to the a16z Show on Spotify Listen to the a16z Show on Apple Podcasts Follow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
About The a16z Show
The a16z Show

The a16z Show

By Andreessen Horowitz

The a16z Podcast discusses tech and culture trends, news, and the future – especially as ‘software eats the world’. It features industry experts, business leaders, and other interesting thinkers and voices from around the world. This podcast is produced by Andreessen Horowitz (aka “a16z”), a Silicon Valley-based venture capital firm. Multiple episodes are released every week; visit a16z.com for more details and to sign up for our newsletters and other content as well!