Key Takeaways
- Vercel's internal AI agent @v is now a daily operator across finance, communication, docs, marketing, engineering, and analytics, with token use growing exponentially.
- DeepSeek's low API prices appear structural: models reportedly ~10x smaller than Opus and ~5x than Sonnet can serve ~40x traffic on the same compute.
- Strike is running Kimi K3 inside an agentic security-review loop and pushing bitcoin developers to use the open-weights model for vulnerability scanning.
- Paul Graham says a hit product gives founders room for mistakes, warns against inflated credentials, and says Google/DeepMind sat on a pre-ChatGPT product.
- Agent design is shifting from LLM-centric loops toward event-log runtimes and GUI-style visual prompts, as seen in BabyAGI 4 and Cash App Moneybot.
1. AI Agents in Companies and Products
- Vercel has built an internal AI agent, @v, that works across finance, communication, docs, marketing, engineering, and business analytics, and Rauch says interaction and token volumes are growing exponentially. He argues this points to a future where AI agents run entire companies, and keeping the full stack—source, data, tokens—under control is why Vercel uses its own agent rather than packaged integrations. — via 1
- Cash App's Moneybot is experimenting with Visual Prompts, an interface that pushes AI from chat into context-aware visual operations. Jack says the goal is for agents to act more like a GUI than a chatbot. — via 1
- Amjad Masad says Replit's Design Mode can turn two prompts into a working chess app, citing a year of progress that lets non-technical users design, build, and ship high-quality products. — via 1 2
- Yohei Nakajima highlighted BabyAGI 4th gen's "Active Graph Agent Runtime," which builds the agent around an immutable event log instead of a traditional LLM loop: behavior, policies, views, and packages replace the control flow. He also flags the context window as a hard bottleneck because compressed learning and new events compete for limited space. — via 1 2
2. Model Economics, Security, and Capabilities
- DeepSeek's API pricing is explained by model size: the models are reportedly ~10x smaller than Opus and ~5x smaller than Sonnet, letting one chip serve many requests and handle roughly 40x traffic on comparable compute. That suggests the low price is a structural cost advantage rather than a temporary subsidy. — via 1
- Strike operates an agentic system that uses Kimi K3 daily with its security team to review code and find vulnerabilities. Mallers recommends bitcoin engineers run K3 on internal and public repos because it can produce a complete report in one pass, while conceding some findings are exaggerated or wrong; he also links the model's open-weights release to COLDCARD's issues, though he offers no evidence. — via 1
- Paul Graham says LLMs outperform at math not because math is easier but because it has clear right/wrong answers, and he expects writing to improve too. He also reveals he worked on a ChatGPT predecessor called LMChat a year before ChatGPT, but Google was too nervous to launch it and DeepMind was blocked from releasing something that could disrupt Google. — via 1 2
3. Startup and Product Lessons
- Paul Graham argues a single extremely popular product gives founders room to make many mistakes; without it, there is almost no slack. He also warns that overstating credentials to investors is a red flag, noting 19-year-old John Collison did not open with Harvard. — via 1 2
- Garry Tan says he found 12 unusually ambitious builders among more than 6,000 at YC Startup School, including an IOI two-time gold medalist, a Navy EOD officer, and a 17-year-old Coinbase intern. He treats resourcefulness over pedigree as the key signal and notes one founder built a rocket from a salad bowl. — via 1
- Patrick Collison connects beauty to innovation: modernism's rejection of cultural continuity had consequences beyond aesthetics, and many fields have stopped visibly improving since the 1990s. He uses German food as an example of markets trapped in worse equilibria and says Stripe's pursuit of beauty helps it break conventions. — via 1
- Andrew Wilkinson says Codex feels slower than Claude Code because it collapses execution into a single fading line, while Claude's step-by-step terminal updates make progress visible. A detailed streaming mode would give Codex a similar sense of speed, as chain-of-thought did for ChatGPT. — via 1
