Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper
Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper
YouTube1 hr 23 min
Watch on YouTube
Note: AI-generated summary based on third-party content. Not financial advice. Read more.
Quick Insights

Maintain NVIDIA (NVDA) as a foundational holding for near-term AI hardware leadership, while adding Advanced Micro Devices (AMD) to capture upside as hyperscalers diversify toward cost-effective inference chips.

Invest in Micron Technology (MU) to capitalize on strong pricing power driven by severe, ongoing bottlenecks in High-Bandwidth Memory (HBM) and DRAM.

Anchor manufacturing exposure in Taiwan Semiconductor Manufacturing Company (TSM) for critical advanced chip fabrication, while utilizing Intel Corporation (INTC) as a strategic hedge against Asian geopolitical supply chain risks.

Target renewable energy operators (solar and wind) and distributed power providers, which are poised to monetize surplus power by servicing regional, low-overhead AI data centers.

Shift software allocations toward enterprise automation and background Agentic AI developers, which are positioned to expand profit margins as AI inference computing costs fall by multiple orders of magnitude.

Detailed Analysis

NVIDIA (NVDA)

  • NVIDIA maintains an industry-leading position in AI hardware and interconnect technology like NVLink, making their GPUs the gold standard for low-latency and interactive AI applications.
    • The company is shipping entire rack-scale systems such as the NVL72 (Grace Blackwell), shifting programming from individual chips to entire data center racks.
    • While demand for chips like Blackwell remains exceptionally high, generational improvements in performance-per-watt (from Hopper to Blackwell to Rubin) show smaller incremental gains on standard precision math than marketing suggests.
    • NVIDIA strategically avoids competing directly with its cloud and inference customers, instead cultivating an ecosystem of "neocloud" partners like CoreWeave to drive sustained hardware demand.

Takeaways

  • NVDA remains the safest core holding for near-term AI infrastructure exposure due to its software ecosystem and advanced networking moat (NVLink).
  • Investors should monitor potential margin pressure over the long term as customers explore cheaper compute alternatives for high-throughput background inference workloads.

Advanced Micro Devices (AMD)

  • AMD chips offer strong raw computational power (FLOPS per dollar) that can rival NVIDIA when paired with optimized software kernels and alternative parallelism strategies.
    • Market perception often underestimates AMD's software viability, creating an arbitrage opportunity for efficient inference providers.
    • Major tech companies, including Meta and OpenAI, are actively securing large allocations of AMD compute to diversify away from exclusive reliance on NVIDIA.

Takeaways

  • AMD is well-positioned as a primary secondary-source beneficiary as hyperscalers and inference providers look for cost-effective hardware to satisfy soaring inference demand.

Micron Technology (MU) & High-Bandwidth Memory Producers

  • High-Bandwidth Memory (HBM) and dynamic RAM (DRAM) manufactured by Micron, SK Hynix, and Samsung are critical bottlenecks for AI accelerators.
    • Unlike logic processors, memory cannot easily be integrated on-die at massive scale, making off-chip memory capacity vital for storing the expanding KV cache (the working memory of long AI conversations).
    • Memory suppliers remain disciplined with capital expenditures due to historical boom-and-bust cycles, prolonging the memory supply crunch and keeping memory component costs elevated.

Takeaways

  • Memory producers like MU possess strong structural pricing power in the AI supply chain, as expanding AI context windows and agent memory directly drive dense DRAM and HBM demand.

Taiwan Semiconductor Manufacturing Company (TSM)

  • TSMC remains the primary manufacturing choke point for advanced AI silicon, with packaging and wafer allocation constraining the entire hardware startup ecosystem.
    • Foundries maintain tight process controls to limit variation between chips, though specialized inference buyers may increasingly accept higher variance or lower-tier binning to lower silicon procurement costs.
    • While geopolitical risks around Taiwan are frequently cited, Western fab capabilities are closer in performance-per-watt (within roughly 2x) than popular market narratives suggest.

Takeaways

  • TSM holds an indispensable role as the foundry of choice for leading AI chip designers, though investors should watch packaging capacity as a key limiter of near-term revenue realization.

Intel Corporation (INTC)

  • Western semiconductor fabrication capabilities, spearheaded by Intel, face a narrower performance-per-watt gap against leading Asian foundries than broad market sentiment implies.
    • In a potential geopolitical supply disruption, domestic fabs provide a viable floor for Western compute capacity without catastrophic performance degradation.

Takeaways

  • INTC represents a strategic hedge on geopolitical concentration in the semiconductor supply chain, with domestic manufacturing assets providing downside protection against Asian supply shocks.

AI Energy & Distributed Infrastructure (Sector Theme)

  • The AI infrastructure model is beginning to bifurcate between centralized mega-clusters (for model training) and distributed small-footprint sites (for model inference).
    • Training requires massive concentrated power (100MW to 1GW) with high networking redundancy, while inference can be served across distributed, modular 1-megawatt sites.
    • Asynchronous background AI agents do not require 99.99% uptime, enabling operators to use cheap, intermittent renewable power (solar and wind) and lower-cost data centers with ~95% uptime without backup diesel generators.

Takeaways

  • Power constraints will accelerate demand for distributed, small-scale power solutions and low-overhead regional data centers rather than relying exclusively on monolithic multi-gigawatt facilities.
  • Renewable energy operators with intermittent surplus generation could see new revenue streams by supplying non-latency-critical AI inference workloads.

Agentic AI & Long-Horizon Inference (Sector Theme)

  • The dominant computing workload is shifting rapidly from model training to inference, and from real-time interactive chatbots to background, long-horizon autonomous agents.
    • Background workloads (deep research, automated cybersecurity pen-testing, coding) are projected to grow from roughly 50% of inference workloads today to 90% over time.
    • Background agents operate asynchronously, prioritizing low cost per token and high throughput over instantaneous latency, drastically lowering the cost threshold required to deploy intelligence.
    • Cost per trillion tokens is expected to drop from roughly $5 million toward thousands or tens of thousands of dollars, opening vast new product categories.

Takeaways

  • Software companies positioned around autonomous background workflows and enterprise automation stand to capture significant value as inference costs decline by multiple orders of magnitude.
  • Pure low-latency inference providers face commoditization pressure from high-throughput, low-cost token factories tailored for asynchronous AI agents.
Ask about this postAnswers are grounded in this post's content.
Video Description
Neil Movva, co-founder of Sail Research, joins Patrick to explain why the next era of AI may be defined not by faster chatbots, but by background agents that work autonomously for hours, days, or even weeks—and what it will take to make that intelligence cheap enough to become abundant. Neil lays out Sail’s “token factory” strategy across software, chips, data centers, and power; the lessons he learned chasing “speed of light” performance at Nvidia; why he’ll buy almost any chip at the right price; and why unreliable, distributed data centers powered by sources like solar and wind could actually be an advantage. They also discuss transformers and memory, the future of AI training data, the economics of the chip shortage, open versus closed models, and Neil’s belief that there will always be demand for more intelligence. TIMESTAMPS 0:00 Intro 0:38 Building a “Token Factory” 4:21 The Future of Background Agents 13:09 Nvidia and the GPU Stack 23:27 Chips, Memory, and Transformers 36:14 The Future of AI Training Data 44:32 Chip Scarcity and Compute Arbitrage 52:44 Reinventing the AI Data Center 59:01 Power and the “Scavenger Strategy” 1:10:10 Open vs. Closed AI #ArtificialIntelligence #AIAgents #NVIDIA #AIInfrastructure #InvestLikeTheBest Presented by Ramp: https://ramp.com/invest Sponsored by Vanta, WorkOS, Rogo, and Ridgeline: https://www.vanta.com/invest https://workos.com/ https://rogo.ai/invest https://www.ridgelineapps.com/ ****** Patrick O'Shaughnessy is the CEO of Positive Sum. All opinions expressed by Patrick and podcast guests are solely their own and do not reflect the opinion of Positive Sum. This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions. Clients of Positive Sum may maintain positions in the securities discussed in this podcast. To learn more, visit psum.vc
About Invest Like The Best
Invest Like The Best

Invest Like The Best

By @iltb_podcast

Conversations with the best investors and business leaders in the world. We explore their ideas, methods, and stories to help you ...