Key Takeaways
- In a benchmark of 7 frontier models for automated research tasks, Fable-5 leads overall, but open-source Kimi K2.7-Code surpasses frontier models in ML engineering tasks.
- AI agents are still nascent; best practices for rebuilding companies around them remain unclear, with only months of practical experience.
- U.S. AI policy risks banning open-source models, potentially making Chinese AI the default for 6 billion people globally, warns Yann LeCun.
- Anthropic's Fable model faces controversy over a government block, with researchers petitioning for its restoration.
- xAI's Grok Build now supports LaTeX rendering and an Agent Dashboard for managing multiple agents.
- Model self-training (training on own outputs) inherits quirks and is hard to filter, explaining similarities within model families.
1. Model Capabilities and Benchmarking
- swyx tested 7 frontier models on automated research tasks: Fable-5 (Anthropic) won overall, but open-source Kimi K2.7-Code (Moonshot) outperformed all frontier models on ML engineering tasks. — via 1
- Ethan Mollick challenges a headline claiming AI “fell short” on math: solving 7/10 novel problems is a significant leap forward. — via 1
- NVIDIA's Nemotron 3 Ultra AI shows fast performance on agent, math, and science tasks, but struggles with hard-coded tasks. — via 1
- A model trained on outputs from a previous model inherits unusual habits that are hard to filter, potentially explaining similarities within model families (Google DeepMind research). — via 1
2. AI Agents and Workflows
- Ethan Mollick notes that rebuilding companies around AI agents is still uncertain; agents have only been in practice for a few months, requiring experimentation and failure. — via 1
- Aravind Srinivas discusses how AI agents will disrupt advertising models, and highlights common token budget management misconceptions. — via 1
- swyx argues that traditional code review is dead; dynamic workflows apply beyond coding to knowledge work requiring judgment. — via 1 2
- Ethan Mollick points out that models are relatively weak on vision, making visual steps the largest error-accumulation points in workflows. — via 1
3. Policy, Open Source, and Controversies
- Yann LeCun warns that U.S. AI policy could ban open-source models, making Chinese AI the default for 6 billion people, with the U.S. isolated—time is short. — via 1
- Anthropic's Fable model was blocked by the Trump administration amid communication issues and security disagreements; researchers are petitioning for its restoration. — via 1
- Ethan Mollick reflects on AI regulation: clear boundaries are complex because tools amplify model capabilities, and risk varies between weak open systems and strong closed ones. — via 1
4. Infrastructure, Tools, and Breakthroughs
- xAI's Grok Build now renders LaTeX in terminal and features an Agent Dashboard to manage multiple agents simultaneously. — via 1 2
- Radical Numerics emerged from stealth with a $50M seed round and released Omnii, a genomic language model. — via 1
- NVIDIA highlights World-Action Models (WAMs) as the second dominant approach for robot foundation models alongside classical VLA. — via 1
