148. 对游凯超3小时访谈:开源Infra、模型Co-design 、“如果vLLM失败,我们会后悔一辈子”
148. 对游凯超3小时访谈:开源Infra、模型Co-design 、“如果vLLM失败,我们会后悔一辈子”
Podcast2 hr 6 min
Listen to Episode
Note: AI-generated summary based on third-party content. Not financial advice. Read more.
Quick Insights

Investors should monitor foundational open-source AI infrastructure projects like vLLM, which act as critical ecosystem gatekeepers across various hardware architectures like Nvidia and AMD.

The commercialization of vLLM through its newly formed entity, Infraact, mirrors the enterprise success of companies like Red Hat and Databricks, making it a key trend to watch for institutional AI backing.

Investors focusing on the AI deployment market should prioritize firms that successfully co-design their models with hardware constraints, as this dramatically reduces cost-per-token metrics.

Look for opportunities tied to high-performance open-source model ecosystems, particularly those inspired by the explosive global adoption of DeepSeek-R1.

Detailed Analysis

vLLM (Open Source AI Inference Ecosystem)

  • vLLM is an open-source large language model (LLM) inference engine and ecosystem, originally initiated as a research project at UC Berkeley (SkyLab) that grew out of the PagedAttention algorithm.
  • It serves as a foundational open-source software layer that bridges the gap between hardware resources and large language models, efficiently managing the intermediate states (KV Cache) during text generation.
  • The project has achieved massive global adoption, with prominent AI players and model providers—such as DeepSeek, Kimi, Meta, and others—utilizing or drawing inspiration from vLLM for their inference pipelines.
  • vLLM was officially donated to the PyTorch Foundation to ensure long-term neutral governance, intellectual property protection (trademark security), and community continuity.
  • A commercial entity, Infraact, was founded by core maintainers (including You Kai Chao and Woosuk Kwon) to provide institutional backing, handle enterprise agreements, secure dedicated compute clusters, and drive large-scale deployments.

Takeaways

  • Open Source Infrastructure Moat: Open-source AI infrastructure projects that solve fundamental bottlenecks (like serving speed and memory efficiency) capture massive ecosystem value, as every new model release depends heavily on robust inference engines.
  • Enterprise Support & Commercialization: Projects transitioning from university labs to foundational industry standards often require dedicated corporate entities (analogous to Red Hat for Linux or Databricks for Spark) to handle high-level enterprise requirements, legal compliance (NDAs), and sustained cluster-level testing.
  • Actionable Outlook: Investors looking at AI infrastructure should monitor ecosystems anchored by foundational tools like vLLM, as they hold strategic gatekeeper status across diverse hardware architectures (Nvidia, AMD) and open-source model releases.

DeepSeek (DEEPSEEK)

  • DeepSeek is recognized as having a top-tier infrastructure team globally, with their engineering advancements tightly co-designed alongside their model architectures (such as DeepSeek-MoE, DeepSeek-V3, and DeepSeek-R1).
  • DeepSeek's internal engineering teams actively push the boundaries of inference optimization, developing advanced variations of speculative decoding and mixture-of-experts (MoE) routing that interact closely with engines like vLLM.
  • The rapid release cycle and explosive popularity of models like DeepSeek-R1 catalyzed a massive wave of open-source model deployment globally, stressing and expanding the capabilities of underlying inference infrastructure.

Takeaways

  • Model-Infra Co-Design: The performance ceiling of modern AI models relies heavily on tight cooperation between algorithm researchers and infrastructure engineers. Models designed with hardware and inference constraints in mind (such as optimized attention mechanisms or hardware-friendly routing) achieve vastly superior real-world throughput.
  • Actionable Outlook: Companies demonstrating excellence in both algorithmic efficiency and low-level infrastructure optimization (reducing the cost-per-token while maintaining output quality) possess a sustainable competitive advantage in the AI deployment market.
Ask about this postAnswers are grounded in this post's content.
Episode Description
今天我们的嘉宾是游凯超,他是创业公司Inferact联合创始人兼首席科学家——这个公司很特殊,是从伯克利一个校园开源项目vLLM演化而来。这个社区的维护者用了将近3年时间,把它从一篇算法论文变成了一个开源社区,又变成了一个公司的正常运营。 这些维护者们放弃了大量的个人诱惑,坚持着他们所相信的开源精神。 今年初,Inferact获得了1.5亿美元种子轮融资。我们讨论了当一个寄托着社区情怀的开源项目,变成一个商业化的组织运营之后,它所面临的抉择、转变与思考。 另外一方面,我们也非常希望在AI Infra技术前沿上,给大家带来一些新的思路。所以我也和游凯超聊了聊,在AI Infra、模型结构,甚至Harness Engineering都要联合设计的时代背景下,多方的联合设计应该怎么做? 接下来,就是我对游凯超的访谈。 OUTLINE: 00:02:18 从算法到机器学习系统 00:37:56 开源项目vLLM的诞生 01:07:25 “如果vLLM失败了,我们会后悔一辈子” 01:20:11 “仁慈的独裁者” 01:37:16 从社区到创业 01:52:56 模型与Infra的Co-design 02:15:10 Token VS 电力 02:35:41 技术预测 LINKS: 我们的播客在小宇宙、Apple Podcast、Spotify等全音频平台播出; 我们的视频播客在Bilibili、小红书、视频号、抖音等全视频平台播出; 如果你想服用文字版,请搜索我们工作室的公众号:语言即世界language is world。 DISCLAIMER: 本内容不作为投资建议。 CONTACT: xiaojunzhang@lisw.ai Jump into the new world-and explore with us!😉
About 张小珺Jùn|商业访谈录
张小珺Jùn|商业访谈录

张小珺Jùn|商业访谈录

By 张小珺

努力做中国最优质的科技、商业访谈。 张小珺:财经作者,写作中国商业深度报道,范围包括AI、科技巨头、风险投资和知名人物,也是播客《张小珺Jùn | 商业访谈录》制作人。 如果我的访谈能陪你走一段孤独的未知的路,也许有一天可以离目的地更近一点,我就很温暖:)