Key Takeaways
- OpenAI and Sam Altman paused some frontier RL training to improve safety and monitoring, with the largest runs still halted; Ethan Mollick calls the step a sign that alignment problems are severe. — via 1 2 3
- DeepSeek V4 Pro is now available on US-hosted Perplexity Computer, posting a 0.359 WANDR score at $0.75 per task and claiming to be 62% cheaper than the next cost-performance frontier model. — via 1
- Miles v0.1, an open-source RL framework for LLMs and multimodal models, has 72 contributors, 1,326 commits, and adoption by frontier model developers. — via 1
- Anthropic says Claude designed novel protein binders for 14 of 15 targets; Yann LeCun argues open-source protein-design tools deserve most of the credit and have shared blind spots. — via 1 2
- NVIDIA open-sourced SkillEvaluator after measuring that adding skills improves agent correctness by 41 points, effectiveness by 39 points, and efficiency by 35 points. — via 1
- Ethan Mollick says OpenAI's willingness to spend 20% of research reasoning compute on chain-of-thought monitoring shows alignment issues are serious and labs need common standards. — via 1
1. AI Safety and Frontier Training Pause
- OpenAI and Sam Altman announced a temporary pause on parts of frontier RL training, including a two-week halt intended to harden research environments and expand monitoring before deploying the latest model. Altman said model progress is outpacing safety alignment, and he will act unilaterally if needed; OpenAI also said the largest frontier RL runs remain paused until small-scale training and evaluations validate safety measures. — via 1 2
- Ethan Mollick interprets OpenAI committing 20% of its research reasoning compute to chain-of-thought monitoring as evidence that alignment is already a serious problem. He argues labs need common policies and standards, because monitoring reasoning at that scale is not business-as-usual. — via 1
- OpenAI also highlighted Replit Free Mode, now powered by GPT-5.6 Luna, framing it as making intelligent coding tools broadly accessible. — via 1
2. Models, Agents, and Open-Source Tools
- DeepSeek V4 Pro is now hosted in the US on Perplexity Computer. According to Aravind Srinivas, it scores 0.359 on WANDR, costs $0.75 per task, and is 62% cheaper than the next model on the cost-performance frontier, making it a notable price-performance signal. — via 1
- Miles v0.1 is an open-source RL framework aimed at LLM and multimodal model training, supporting debugging, hardware efficiency, and scaling. It has attracted 72 contributors, logged 1,326 commits, and is already used by multiple frontier model developers, according to Aravind Srinivas. — via 1
- NVIDIA released SkillEvaluator after benchmarking 300+ verified skills: adding a skill lifts agent correctness by 41 points, effectiveness by 39 points, and efficiency by 35 points on the same model and task. The tool is open-sourced so users can measure their own agents. — via 1
- Sentence Transformers v6 introduces support for multi-vector late-interaction models such as ColBERT, using MaxSim for fine-grained semantic matching. Elastic, Qdrant, and other vector databases already support the approach, which could push the industry beyond bi-encoder retrieval. — via 1
- Runway made MiniMax H3 unlimited for Max plan users for a limited time, allowing limitless video generation. — via 1
- NVIDIA RTX Spark brings creative workflows, local AI tools, and RTX gaming together in one PC, aimed at users who want both AI productivity and gaming in a single machine. — via 1
3. AI in Science: Protein Design and Its Limits
- Anthropic says Claude autonomously designed novel protein binders from scratch, hitting 14 of 15 targets, with independent construction and testing by Adaptyv Bio and Twist Bioscience. The demo highlights AI's potential in early-stage drug discovery. — via 1
- Yann LeCun pushes back on the framing: Claude's orchestration is impressive and requires biological judgment, but most of the credit belongs to open-source protein design tools from Baker Lab, Columbia, MIT, ByteDance Seed, and others. He also warns that tools sharing the same training distribution can hit collective blind spots on out-of-distribution targets, so the claim should be limited to domains where the stack is already validated. — via 1 2
4. Independent Signals and Observations
- Ethan Mollick questions “sovereign AI”: if sovereign models lag far behind the frontier, using them in the highest-stakes frontier applications, such as cybersecurity, may carry large costs. — via 1
- Mollick also warns that Qwen 27B, despite being a good local model, clearly underperforms on agentic tasks measured by GDPval-AA, and recommends that teams run their own benchmarks before relying on local models. — via 1
- On generative media, swyx focuses on AI Engineer World's Fair 2026's Generative Media Track and argues that the hard parts of generative media happen before and after model inference, not inside the model. — via 1
