Key Takeaways
- OpenAI disclosed performance results for its custom "Jalapeño" inference chip, showing efficiency and latency gains — via 1 2 3.
- ChatGPT Work can now log into sites and apps without seeing credentials, and OpenAI introduced $100-per-seat Business Premium Seats — via 1 2 3 4 5.
- Perplexity's local-first Portable Computer research scores 82.6% with a 27B on-device model (85.4% post-trained) and now runs on NVIDIA DGX Spark — via 1 2 3 4.
- NVIDIA's Vera Rubin NVL72 production racks have arrived, built with automated one-tray-per-minute manufacturing — via 1.
- Andrew Ng released a new open-source OpenWorker version with built-in cybersecurity agents and local model support — via 1.
- Developers should avoid Codex's "locked use" feature because it can lock macOS keychains due to a known Apple bug — via 1.
1. OpenAI silicon and agentic workspace
- OpenAI has built its own inference chip, "Jalapeño", with test results showing better per-watt compute, faster responses, higher throughput, and lower latency. Sam Altman says the chip "is fast," while Greg Brockman argues AI's role in chip design is underrated and highlights OpenAI's lean chip team and AI-driven improvement loops. — via 1 2 3.
- OpenAI introduced ChatGPT Business Premium Seats at $100 per seat for small businesses and startups, promising more powerful tools and more efficient workflows. Greg Brockman adds that the seats remove the 5-hour limit and provide more usage. — via 1 2.
- ChatGPT Work can now log into websites and mobile apps on the user's behalf without ChatGPT ever seeing usernames or passwords, enabling tasks like booking appointments, reimbursements, billing, and opening accounts. Sam Altman and Greg Brockman both emphasize the password-isolation design. — via 1 2 3.
2. Perplexity and NVIDIA double down on local and production AI
- Perplexity released its Portable Computer research: a local-first, private agent whose on-device 27B model scores 82.6% on real knowledge-work tasks, beating open harnesses Pi and Hermes; post-trained PPLX 27B reaches 85.4%. The research is supported by NVIDIA and can run locally on DGX Spark, with an open ecosystem around low-cost open-weight models, inference frameworks, and unified-memory hardware. For Max users, Perplexity Computer adds a persistent Dream agent and Brain memory system, improving correctness by 9.3 points, timeliness by 8.0, recall by 8.9, and reducing token use by 15%. — via 1 2 3 4.
- NVIDIA's Vera Rubin NVL72 production racks have arrived, with compute trays made by fully automated lines at one per minute; Microsoft is the first to deploy, and Foxconn's production line is now live. At Hot Chips 2026, NVIDIA also framed Vera CPU, Vera Rubin, Groq 3 LPX, Spectrum-X Multiplane, and BlueField-4 Scale-In as a full-stack agentic AI platform. — via 1 2.
- NVIDIA Dynamo adds Shadow engine recovery in preview, keeping a standby engine warm so an LLM crash avoids cold-restart capacity loss. In GLM-5.2 tests, recovery took 7.3 seconds—nearly 39x faster than a cold restart. — via 1.
- NVIDIA introduced the Jetson Orin Nano 2 for entry-level edge AI, claiming 2x the inference performance of Orin Nano Super and 40% power reduction at the same performance in 15W mode, while keeping the same compact size. — via 1.
3. Open-source agents, models, and developer tooling
- Andrew Ng released a new version of OpenWorker, an open-source AI agent that can complete real tasks on a user's computer. New features include secure workflows and built-in cybersecurity agents that scan code for vulnerabilities, dependency supply-chain injection risks, and cloud security attack surfaces; the framework's code is fully auditable, and users can run local open-weight models to protect sensitive code. — via 1.
- Hugging Face highlighted new open models and tools, including GLM-5.3-Flash: a native multimodal, 1M-context model with 320B-A18B parameters, released under MIT and running entirely on Chinese AI chips. It also showcased Apodex 1.1 for async agent-team collaboration and promoted the open-weight Breeze TTS 2 as the top-ranked TTS model on the Hub. — via 1 2 3.
- OpenAI's new WebMCP and a companion notebook format were open-sourced, targeting scenarios where agents and humans need to collaborate through a UI (e.g., co-editing notebook cells). Unlike MCP/API, WebMCP lives directly in the browser; the related notebook uses markdown files and supports bringing your own coding agent. Hamel Husain used the stack to run models on his own infrastructure and shared the resulting runbook. — via 1.
- Developers should avoid Codex's "locked use" feature for now, warns swyx, because it depends on a fragile macOS feature and caused two macOS keychain lockouts this week. Apple developer forums reportedly confirm the bug; cloud-based alternatives are not mature yet. — via 1.
