
Investors should prioritize NVIDIA (NVDA) as the foundational "supply chain" play, while remaining cautious of hardware depreciation risks as chip release cycles accelerate to three times per year. Look for opportunities in the Open Source AI ecosystem (such as Llama or Mistral integrations), which currently offers 15x better cost-effectiveness for 90% of enterprise workflows compared to closed models. Focus on private or public companies that "own their intelligence" through proprietary tuned models rather than those simply renting general APIs, as token costs are projected to drop 10x while usage explodes 100x. The most immediate growth is shifting from coding tools to "co-work" AI agents in high-value sectors like Legal, Finance, and Sales. Finally, monitor the Energy and Infrastructure sectors, as power availability and specialized data center designs are now the primary bottlenecks for AI scaling.
• Fireworks is a specialized intelligence platform focusing on the inference market, sitting between chip providers (NVIDIA) and model providers. • The company has scaled to $800 million in ARR (Annual Recurring Revenue) in just a few years, with projections to double by the end of the year. • They process over 40 trillion tokens per day, primarily from customized models rather than off-the-shelf general models. • The platform emphasizes "quality first," achieving zero KLD (bit-wise equivalence between training and inference) to ensure no accuracy is lost when deploying models.
• Specialization over AGI: The core investment thesis is that the future belongs to millions of specialized models rather than one dominant AGI (Artificial General Intelligence) model. • Enterprise Data Moat: Most valuable data is private enterprise data. Fireworks enables companies to turn this proprietary data into "private intelligence" that general models like GPT-4 cannot access. • High Growth vs. Margins: The company currently operates with lower margins (30-40%) to prioritize hyper-growth and innovation, with plans to optimize for gross margins once the systems stabilize.
• Mentioned as a critical partner and a "supply chain" solver for the AI ecosystem. • NVIDIA is training its own models (e.g., NemoTron) to ensure there is a U.S.-native open-source supply, preventing a "blockage" in the AI stack. • CEO Jensen Huang’s "five-layered AI cake" (Application, Model, Infrastructure, Chips, Energy) is the framework for the industry's growth.
• Supply Chain Bottlenecks: The industry is currently bottlenecked by the lower layers: Energy and Chips. • Hardware Depreciation Risks: The speed of model development is outstripping hardware cycles. A chip SKU might be released three times a year, potentially making older hardware less valuable for the newest, most efficient models.
• Open-source models have crossed a "quality threshold" where they can solve 90% of enterprise workflows. • They offer 15x more cost-effectiveness than frontier closed models in many scenarios. • Open-source provides "control" and "ownership," allowing companies to tune models to their specific "taste" or "judgment."
• Bullish on Open Ecosystems: The transcript suggests that the "power lines" (frontier models like OpenAI/Anthropic) will not replace everything. Specialized appliances (open-source tuned models) will dominate the application layer. • National Security: There is a growing debate regarding the quality of Chinese open-source models (which currently lead some leaderboards) and the need for a strong U.S. open-source ecosystem.
• An AI-powered coding platform cited as a pioneer in model tuning. • They have achieved "escape velocity," scaling from single-digit millions to massive revenue by focusing on product innovation while outsourcing platform infrastructure to Fireworks.
• The "Year of Co-work": While last year was the "year of coding" (led by companies like Cursor), this year is the "year of co-work," where AI agents are being applied to legal, finance, recruiting, and sales.
• Price Prediction: Token costs are expected to drop 10x in the next three years. • Usage Explosion: This 10x cost reduction is predicted to drive a 100x increase in usage. • Efficiency: Future ROI will come from "token maxing" (using fewer, more precise tokens to solve a task) rather than just buying more compute.
• The discussion suggests that OpenAI and Anthropic may be overvalued if the market shifts toward specialized, smaller models. • Actionable Insight: Investors should look for companies that "own their intelligence" (proprietary tuned models) rather than those just "renting" general APIs, as the latter lacks a durable moat and cost control.
• Heterogeneous Design: Future data centers will likely move away from "one size fits all" to specialized setups (e.g., combining NVIDIA GPUs for processing with ASIC accelerators like Grok for generation). • Sovereign AI: Nations and large corporations are increasingly seeking "sovereign models" to avoid the risk of being "cut off" by a single provider or administration.
• Scaling to Bankruptcy: Startups with product-market fit may fail if they cannot control inference costs as they scale. • Supply Chain: Global shortages in everything from high-end chips to basic electricians and power infrastructure remain the primary drag on the AI revolution.

By Harry Stebbings
The Twenty Minute VC (20VC) interviews the world's greatest venture capitalists with prior guests including Sequoia's Doug Leone and Benchmark's Bill Gurley. Once per week, 20VC Host, Harry Stebbings is also joined by one of the great founders of our time with prior founder episodes from Spotify's Daniel Ek, Linkedin's Reid Hoffman, and Snowflake's Frank Slootman. If you would like to see more of The Twenty Minute VC (20VC), head to www.20vc.com for more information on the podcast, show notes, resources and more.