Key Takeaways
- Claude Fable 5 demonstrates unprecedented capabilities: migrating 50 million lines of code in a day, generating complex 3D renderings, and performing 10x optimizations over GPT 5.5 at a fraction of the cost.
- NVIDIA's Nemotron 3 Ultra achieves legal domain frontier performance with 1/8 to 1/50 the cost of leading models, while DiffusionGemma delivers 4x speedup over Gemma 4 with up to 1000+ TPS.
- Anthropic releases a policy paper on AI exponential growth and launches a 1000-person fellowship, while ethical debates around AI safety and open-source intensify with LeCun and Mollick weighing in.
- xAI's Grok Voice reaches Pareto frontier on EVA-Bench with accuracy and user experience leadership at lower cost, and Grok Build plugin marketplace enters beta.
- Runway deepens partnership with Lionsgate and announces sold-out AI Film Festival, signaling growing mainstream adoption of AI-generated content.
1. Model and Infrastructure Breakthroughs
- Claude Fable 5 is now available as an orchestration model in Computer for Pro and Max users. Deedy reports it can migrate 50 million lines of code from Stripe in a day (human equivalent ~2 months), create complex 3D scenes (detailed aircraft, space sims with 5000+ objects), and deliver 10x optimization over GPT 5.5 for interactive network evaluators, all at a price similar to GPT 5.5 and 6x cheaper than GPT 5.5 Pro. It also produces pixel-perfect documents, slides, and websites, marking the biggest quality leap since o3. — via 1 2 3
- NVIDIA and Google DeepMind release DiffusionGemma, a text diffusion model that generates 256 tokens per step, achieving 150+ TPS on DGX Spark and 1000+ TPS on a single H100. Demis Hassabis notes it is 4x faster than Gemma 4. NVIDIA provides BF16/NVFP4 checkpoints and vLLM support from day one. — via 1 2
- NVIDIA Nemotron 3 Ultra, post-trained with Trajectory Labs, achieves frontier performance in legal domain: Legal Agent Benchmark pass rate from 0% to 5.8% (between Sonnet 4.6's 4.2% and Opus 4.6's 6.6%). Retained task pass rate jumps from ~70% to ~95%, and operating cost is 1/8 to 1/50 of Sonnet/Opus 4.6. — via 1
- xAI Grok Voice reaches Pareto frontier on EVA-Bench with superior accuracy and user experience, priced far below competitors. — via 1
2. Industry, Policy and Safety Debates
- Anthropic CEO Dario Amodei publishes "Policy on the AI Exponential," arguing AI development far outpaces policymaking and announces three new initiatives to bridge the gap. Additionally, Anthropic launches the Claude Corps national fellowship, training 1,000 people at nonprofits to use Claude with pay, targeting early-career professionals. — via 1 2
- Yann LeCun criticizes the "safety" narrative as a guise for censorship and control, urging open-source to capture both model and interface layers to avoid a surveillance economy. He also calls out Anthropic for quietly degrading Claude's performance, labeling it opaque and harmful to independent researchers. Ethan Mollick defends Anthropic's rollout of Fable guardrails but acknowledges they failed to explain them effectively. — via 1 2 3 4
- xAI model is deployed in eToro's Agent Tori, using real-time data to analyze market sentiment for consumers. — via 1
- Runway deepens collaboration with Lionsgate, co-developing original IP. Its 2026 AI Film Festival NYC premiere sold out, highlighting growing momentum in AI-generated entertainment. — via 1 2
3. Tools, Platforms and Practical Insights
- xAI Grok Build plugin marketplace enters beta, enabling terminal-based development with MongoDB, Vercel, Sentry, and more. — via 1
- Factory Desktop launches Missions feature (Ben Tossell). — via 1
- Ethan Mollick advises against simply switching to cheaper models; instead, use a tiered architecture where smart models orchestrate and audit cheap ones. He also questions whether open-weight models can remain free and profitable as costs rise, and whether they are safe enough to avoid government intervention. — via 1 2
- swyx highlights Satya Nadella's redefinition of Microsoft AI strategy—moving from a single model to an ecosystem, emphasizing evaluations as IP and agent traces as balance-sheet assets. He also promotes the "model labs vs. agent labs" framework as the clearest lens for where AI value lies. — via 1 2
- swyx introduces PoeticHQ, an AI system for complex multi-step tasks with 99%+ accuracy, 10x fewer tokens than agents, and $50M in funding (used by AIG, SoFi, etc.). — via 1
