The Most Overhyped and Underhyped New AI Models
The Most Overhyped and Underhyped New AI Models
7 hours agoMatt Wolfe@mreflow
YouTube26 min 29 sec
Watch on YouTube
Note: AI-generated summary based on third-party content. Not financial advice. Read more.
Quick Insights

Alphabet Inc. (GOOGL) offers a compelling investment opportunity as the market is underestimating its newly launched Gemini 3.8 Flash, which delivers top-tier coding performance at an industry-low cost to capture enterprise market share. Microsoft (MSFT) is positioned for a time-sensitive catalyst with OpenAI's rollout of its autonomous Astra model within the next one to two weeks, significantly boosting high-margin enterprise cybersecurity capabilities. While Anthropic leads raw intelligence benchmarks with Claude Fable 5.1, its high execution costs create commercial friction that directly benefits cost-efficient alternatives like Alphabet. Investors should prioritize long exposure to GOOGL for cost-effective software generation scale and MSFT for near-term cybersecurity infrastructure monetization, while monitoring potential regulatory scrutiny on autonomous AI reasoning.

Detailed Analysis

Alphabet Inc. (GOOGL)

  • Google DeepMind released Gemini 3.8 Flash, which delivers competitive performance at an industry-leading low cost.
    • Achieved a 73.7%–74% score on the DeepSuite coding benchmark, matching Claude Opus 5 and surpassing Claude Fable 5.
    • Operates at an average cost of $0.58 per intelligence task, compared to $3.69 for Claude Fable 5.1 and $0.95 for GPT 5.6.
    • Demonstrates high inference speed, averaging 2.5 minutes per task versus 7.4 minutes for competing frontier models.
    • While it ranks lower on general knowledge benchmarks (scoring 59 on Artificial Analysis), it is heavily optimized for fast, cost-effective coding and software generation.

Takeaways

  • Alphabet's ability to deliver near-frontier coding capabilities at a fraction of competitors' pricing positions it well to capture enterprise and developer market share where unit economics matter.
  • The market appears to be underestimating Google's model efficiency improvements, focusing disproportionately on larger frontier lab announcements.

Anthropic (Private)

  • Anthropic released Claude Fable 5.1, establishing a new state-of-the-art baseline on several leading benchmarks.
    • Ranked first on the Artificial Analysis intelligence index with a score of 66, up from the previous best of 63 (Claude Opus 5).
    • Outperformed predecessor models in agentic scientific research, knowledge tasks, and cybersecurity vulnerability detection with 60% fewer refusals.
    • Introduced stricter anti-distillation technical protections to stop rival developers (specifically in China) from scraping Claude's reasoning to train student models.
    • Maintained list pricing at $10 per million input tokens and $50 per million output tokens, resulting in an average cost of $3.69 per task.
    • Real-world intensive coding tasks remain expensive, with complex single sessions costing over $100 in usage credits.

Takeaways

  • Anthropic holds the lead in raw model intelligence and agentic coding, but its high execution cost creates economic friction for widespread commercial deployment.
  • Incremental model updates are experiencing diminishing perceived utility among general consumers, shifting commercial value toward specialized enterprise and cybersecurity niches.

OpenAI (Private / Microsoft MSFT Ecosystem)

  • OpenAI announced the upcoming release of its new model, Astra, expected in the near term ("this week or next").
    • Designated for advanced autonomous cybersecurity capabilities, capable of identifying and exploiting previously unknown vulnerabilities without human guidance.
    • Benchmark data shows exploit success rates increasing from 11.5% (on GPT 5 Sol) up to 40%, while using significantly fewer compute tokens.
    • Development and rollout were intentionally delayed to implement safeguards against cyber misuse.
    • Utilizes a new architecture technique called Recurrent Depth (Looped Transformer), which improves reasoning by processing text iteratively.
    • Risk Factor: The new technique obscures the model's readable chain of thought, creating transparency and safety concerns because human supervisors cannot easily audit the model's internal logic.

Takeaways

  • OpenAI's progression toward high-leverage, dual-use cybersecurity tools represents high commercial value for enterprise defense, potentially benefiting Microsoft (MSFT) as its primary distribution partner.
  • The shift toward non-transparent reasoning models increases safety and regulatory risks, which may invite tighter government scrutiny and compliance costs.
Ask about this postAnswers are grounded in this post's content.
Video Description
A bunch of new models... Join the free newsletter here: https://futuretools.io/newsletter Discover More: 🛠️ Explore AI Tools & News: https://futuretools.io/ 📰 Weekly Newsletter: https://futuretools.io/newsletter Socials: ❌ Twiter/X: https://x.com/mreflow 🖼️ Instagram: https://instagram.com/mr.eflow 🧵 Threads: https://www.threads.net/@mr.eflow 🟦 LinkedIn: https://www.linkedin.com/in/matt-wolfe-30841712/ 👍 Facebook: https://www.facebook.com/mattrwolfe Resources From Today's Video: Gemini 3.8 Flash: https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/ Claude Fable And Mythos: https://www.anthropic.com/claude-fable-and-mythos-5-1 Path To Astra: https://openai.com/index/path-to-astra/ Astra Security Concerns: https://www.theinformation.com/articles/secret-technique-behind-openais-astra-model-sparks-security-concerns Artificial Analysis: https://artificialanalysis.ai/ DeepSWE: https://deepswe.datacurve.ai/ Let’s work together! - Brand, sponsorship & business inquiries: mattwolfe@smoothmedia.co #AINews #AITools #ArtificialIntelligence Time Stamps: 0:00 Intro 0:20 Fable 5.1 7:50 Fable 5.1 Tests 11:07 Gemini 3.8 Flash 13:58 Gemini 3.8 Flash Tests 18:40 OpenAI Astra 20:41 Recurrent Depth 23:35 Conclusion
About Matt Wolfe
Matt Wolfe

Matt Wolfe

By @mreflow

AI News Breakdowns every Saturday and other cool nerdy tech and AI stuff in between. Let's work together! - For brand ...