AI Efficiency Is Repricing The Compute Market | Steve Hou
AI Efficiency Is Repricing The Compute Market | Steve Hou
Podcast43 min 50 sec
Listen to Episode
Note: AI-generated summary based on third-party content. Not financial advice. Read more.
Quick Insights

NVIDIA (NVDA) is a strong buy as long-term GPU rental rates for its H100 and A100 chips are rising, signaling robust AI computing demand that undermines peak-cycle fears.
Favor the compute layer over model builders—avoid pure bets on frontier labs like OpenAI and Anthropic, which face margin erosion from cheaper open-weight models.
Look to invest in the emerging AI orchestration layer, where platforms that smartly route tasks between models are set to capture significant value.
The shift toward cost-efficient AI inference is broadening demand, making NVDA a core holding for AI exposure.

Detailed Analysis

AI Compute Market (General Theme)

• The AI compute market is currently dominated by a few major players, with OpenAI and Anthropic accounting for roughly half of the demand, while hyperscalers provide most of the supply. • A major shift is expected toward an "orchestration layer" where companies use smart routing to substitute between expensive frontier models for high-value tasks and cheaper, open-weight models for simpler tasks. • This transition is expected to fragment the market, moving away from the current concentrated paradigm and creating a need for financial instruments like futures contracts to hedge compute costs. • The market is in a transitional phase that could be "rocky," with a risk of an "air pocket" where the ROI from the current CapEx paradigm doesn't materialize quickly enough before the new, broader-based adoption phase picks up.

Takeaways

• The AI build-out is not a monolithic trend; it's evolving from a concentrated, CapEx-heavy phase to a more fragmented, efficiency-driven one. • This transition creates both risks and opportunities. The risk is a short-term gap in demand, while the opportunity lies in the long-term explosion of inference demand from millions of enterprises. • The companies that enable this orchestration and smart routing between different models are positioned to accrue significant value.


AI Model Token Expenditure Index

• This index, created by Silicon Data, is an expenditure-weighted price index for AI model tokens, similar to the PCE (Personal Consumption Expenditures) index in economics. • It tracks the blended price (input and output) of various large language models, weighted by their usage volume on public routing platforms. • The index is not a measure of total token demand, but rather a gauge of consumer behavior, specifically the substitution between expensive, high-quality models and cheaper, less capable ones. • The index plateaued and began to mean-revert in June, which the creator interprets as a sign of "token efficiency" rather than a collapse in demand. Users are rationally substituting away from frontier models for many tasks. • This shift coincided with a sell-off in AI-related stocks, partly due to fears that it challenges the current CapEx funding model and the profit margins of frontier AI labs.

Takeaways

• The move toward "token efficiency" is a healthy, rational development for the long-term adoption of AI, as it makes it more affordable for enterprises to experiment and integrate AI into their workflows. • The index's recent trend is a leading indicator that the value may be shifting away from the model layer (frontier labs) and toward the compute layer or the end-user. • This does not signal a secularly bearish trend for AI, but rather a necessary and expected evolution of the market.


NVIDIA (NVDA) and GPU Compute Market

• The demand for GPUs is analyzed through Silicon Data's rental indices for different chips, including the H100, A100, and B200. • While the on-demand price for H100 chips has fluctuated, the rental rate for the older A100 chip has remained strong and is even going up, indicating robust and growing inference demand. • The forward curve for GPU rentals is a critical indicator. As of late July, the one-year contract price for H100s has been "monotonically going up," with multiple cloud providers raising prices. • This suggests a firming of demand and a supply shortage for long-term compute contracts, contrary to some market fears of excess capacity. • The strong demand for older A100 chips suggests that as newer chips like the H100 become "workhorses" for inference, their demand will also remain robust.

Takeaways

• Despite short-term fluctuations in on-demand pricing, the fundamental demand for GPU compute, especially for long-term contracts, remains very strong. • The strength in older chip rentals (A100) is a powerful signal that inference demand is broad-based and growing, not just limited to cutting-edge training. • The forward curve data directly contradicts the narrative that the AI build-out is peaking, showing instead that cloud providers have pricing power and are raising long-term rates.


Memory and Storage Market

• Demand for memory (DRAM) and data storage is surging due to AI models' hunger for longer contexts and the generation of massive amounts of data. • High prices and shortages are driving algorithmic innovation, such as the memory efficiency improvements seen in models like China's Kimi, which reduces the need for memory to grow linearly with context length. • This is compared to the "DeepSeek moment," where an efficiency breakthrough caused a temporary scare but did not structurally reduce demand. • The long-term demand trajectory for memory is clear, driven by new modalities like advanced voice interaction and, eventually, video, which will generate exponentially more data.

Takeaways

• While efficiency gains may pressure the price (P) of memory in the short term, the quantity (Q) of memory demanded is expected to grow so much that the overall market will continue to expand significantly. • The "Jevons paradox" applies: making AI more memory-efficient will broaden its adoption and use cases, ultimately driving more total demand for memory and storage. • Short-term stock volatility in memory makers is likely due to high valuations and positioning, but the long-term demand trend remains powerfully upward.


Frontier AI Labs (e.g., OpenAI, Anthropic)

• The profit margins of frontier AI labs are being challenged by the rise of powerful, much cheaper open-weight models, particularly from China. • The market is moving away from a "duopoly" of frontier labs toward a model where the orchestration layer allows for easy substitution. • This doesn't mean frontier labs will be unprofitable, but it suggests that the bulk of the value may accrue more to the compute layer or the end-user than previously thought. • The analogy is made to the pharmaceutical industry, where generic drugs create competition but the frontier innovators can still be highly profitable.

Takeaways

• The era of "token maxing" exclusively on expensive frontier models is giving way to a more cost-conscious approach, which could pressure the near-term revenue growth and valuations of these labs. • The long-term winners at the model layer will be those that can maintain a performance edge for the most complex, high-value tasks while the market for simpler tasks becomes commoditized. • The perceived value is shifting on the margin away from the model layer, making pure-play investments in this area potentially riskier than investments in the compute or orchestration layers.


Enterprise AI Adoption & Orchestration Layer

• Genuine enterprise adoption is the next major phase to watch for, and it will be enabled by the very trend of cheaper, more efficient models. • Cheaper models allow for true "token maxing" by enterprises, letting them experiment freely and integrate AI into their core workflows without breaking their budgets. • This adoption is expected to follow a J-curve, being slower than many hope but with increasingly evident signs of ROI. • The value is expected to accrue to the orchestration layer—the platforms that allow companies to smartly route tasks to different models, retain data sovereignty, and build AI into their products.

Takeaways

• The biggest investment opportunity may lie in companies building the orchestration and routing infrastructure that enables enterprises to adopt AI in a cost-effective, sovereign way. • Look for companies that help enterprises manage, route, and integrate multiple AI models into their specific workflows, rather than betting on a single model provider to win. • The productivity gains from AI are currently concentrated in small, agile companies. The next wave will come when larger enterprises successfully navigate this integration, driving a massive new source of demand for compute.

Ask about this postAnswers are grounded in this post's content.
Episode Description
AI’s next phase hinges on a paradox: falling costs could threaten today’s winners while unlocking far greater demand. Steve Hou, head of research at Silicon Data and former Bloomberg strategist, joins us to examine the changing economics of AI compute. We discuss token efficiency, model routing, GPU pricing, memory bottlenecks, and when enterprise adoption may finally deliver measurable returns. Enjoy! TIMESTAMPS: 00:00 Intro 01:01 Why AI Compute Needs Hedging 06:55 What The Token Index Really Shows 14:04 Token Maxing Meets Efficiency 18:35 Who Captures AI’s Value? 22:12 Old GPUs Reveal Surging Demand 27:10 GPU Markets Keep Tightening 32:03 The Memory Bottleneck 37:07 AI’s Next Phase FOLLOW STEVE › X/Twitter – https://x.com/stevehou › Silicon Data – https://www.silicondata.com/ FOLLOW THE SHOW › Forward Guidance – https://x.com/ForwardGuidance › Felix – https://x.com/fejau_inc › Telegram – https://t.me/+CAoZQpC-i6BjYTEx › Blockworks – https://x.com/Blockworks EVENTS › Join us at Digital Asset Summit 2026 Asia October 7th & Digital Asset 2026 London November 10-11th https://blockworks.com/events DISCLAIMER Nothing said on Forward Guidance is a recommendation to buy or sell securities or tokens. This podcast is for informational purposes only. Any views expressed are opinions, not financial advice. Hosts and guests may hold positions in the companies, funds, or projects discussed.
About Forward Guidance
Forward Guidance

Forward Guidance

By Blockworks

The laws of macro investing are being re-written, and investors who fail to adapt to the rapidly changing monetary environment will struggle to keep pace. Felix Jauvin interviews the brightest minds in finance about which asset classes they think will thrive in the financial future that they envision. Follow Felix: https://twitter.com/fejau_inc Follow Forward Guidance: https://twitter.com/ForwardGuidance  Subscribe on YouTube: https://www.youtube.com/@ForwardGuidanceBW Follow Blockworks: https://twitter.com/Blockworks_ Forward Guidance Newsletter: https://blockworks.co/newsletter/forwardguidance Forward Guidance Telegram: https://t.me/+nSVVTQITWSdiYTIx