Key Takeaways
- OpenAI's first custom reasoning chip Jalapeño outperformed commercial systems on InferenceX and is planned for deployment by end of year. — via 1
- Perplexity launched Portable Computer, a fully local agent runtime on NVIDIA DGX Spark, extending its push into on-device AI. — via 1 2 3 4
- NVIDIA is moving AI into orbit with SpaceX's Vera Rubin NVL72 and into edge robotics with Jetson Orin Nano 2. — via 1 2
- MiniMax H3 on Sol-Engine delivered 22x–28x faster video inference, while NVIDIA's Nemotron 3.5 Lightning reached the top four open-weight models. — via 1 2
- Mistral AI partnered with HUMAIN to build localized frontier models for Saudi Arabia, and Runway shipped WAN 3.0. — via 1 2
1. AI Hardware and Infrastructure
- OpenAI's first custom inference chip, Jalapeño, delivered stronger results than commercial comparison systems on the InferenceX benchmark, with higher per-kilowatt peak throughput and better token latency. The team used AI-assisted design to accelerate development and plans to deploy the chip by end of year while continuing work on future generations. — via 1
- NVIDIA and SpaceX designed a space-optimized Vera Rubin NVL72 system, with launch into orbit scheduled for Q4 next year and scaled deployment expected by 2028. This marks a major step toward running AI infrastructure in space. — via 1
- NVIDIA introduced Jetson Orin Nano 2, a compact robot computer for entry-level edge AI that doubles the inference performance of Orin Nano Super and reduces power consumption by 40% at equivalent performance in 15W mode. It gives edge robotics a meaningful performance-per-watt upgrade. — via 1
- NVIDIA's Nemotron 3.5 Lightning is now among the top four open-weight models, averaging 86.4% success on the standard OpenClaw agent benchmark, while Nemotron 3 Ultra remains #1. The result shows NVIDIA pushing further into open-weight model competitiveness, not just hardware. — via 1
2. Local AI and Agentic Computing
- Perplexity launched Portable Computer on NVIDIA DGX Spark, a local-first agent runtime where the orchestrator LLM, sub-agent models, and agent harness all run on-device with no cloud dependency. It includes one-click local inference setup; Aravind Srinivas also said he received a DGX Station from Jensen after a demo, with local hardware now capable of running frontier models like GLM 5.3. — via 1 2 3 4 5
- NVIDIA highlighted MiniMax H3 running on Sol-Engine with 5-second video inference sped up 22.2x (from 152.3s to 6.85s) and 10-second video inference 27.7x (from 414.1s to 14.93s) versus the SGLang baseline. NVIDIA says the speedup changes the unit economics of video generation: at reference API prices, a fully loaded GB200 could produce roughly $210 of output value per hour, with a GPU-only gross margin above 97%. — via 1
- Perplexity is also building infrastructure for large-scale RL systems, focusing on multi-agent collaboration, continuous learning, and reinforcement learning, with plans to open-source much of the work. The company also cut GPT-5.6 Sol credits on Computer by 20% for all platform users until November 21, 2026. — via 1 2
- A direct coding-agent comparison shared by Ben Tossell shows Factory Droid completing a task in 8 minutes and 20 tool calls at $1.60, versus Claude Code's 19 minutes, 40 calls, and $1.87. The data point adds practical cost and latency evidence for teams choosing agent builders. — via 1
3. Models, Partnerships, and Ecosystem
- Mistral AI struck a strategic partnership with HUMAIN covering AI infrastructure, advanced model development, and AI solution deployment in Saudi Arabia and the region. The two will co-develop localized frontier models, initially focusing on cybersecurity, voice, and Arabic capabilities. — via 1
- Runway brought WAN 3.0 to its platform, enabling video and audio generation from multiple image, video, and audio reference inputs. The release expands Runway's multimodal creation workflow. — via 1
- Hugging Face surfaced several open-ecosystem developments: DeepSeek-V4-Flash-0731 GGUF quantized models are nearing 20k downloads, Anthropic released a generative protein binder dataset, and Semantic Scholar added a CLI and Skill that lets agents find papers citing a given paper. These signal continued momentum across open-weight models, scientific AI, and agent-reader tooling. — via 1 2 3
- Sundar Pichai predicted that within three years content will convert freely between any modalities, making today's formats look as primitive as flip phones. Rowan Cheung interpreted this as ending the 'more content wins' strategy, since quality-focused creators will benefit once quantity is no longer scarce. — via 1
