← All issues

Labs Push for Custom Silicon while Agents Learn to Spill Memory to Disk

Today we're seeing the AI infrastructure war heat up as major labs build their own silicon to bypass Nvidia, while engineers focus on practical tricks to cut agent context costs.

Tools & Products

Pi's Coding Agent Beats Context Costs with Disk-Spilling

Most coding agents get incredibly expensive because they stuff massive tool execution logs directly into their context window. Pi's new agent architecture fixes this by saving full outputs to disk and passing only a file path to the model, which greps it only when needed. This simple change cut context usage by up to 35% and slashed API processing costs by 88%. This is a highly practical, reproducible pattern that every team building complex agents should adopt immediately.

Vercel Releases Run SDK for Sandboxed Agent Code Execution

Vercel just dropped Run SDK, a package designed to let developers safely evaluate and execute untrusted JavaScript or TypeScript code generated by AI agents. It runs the code inside a fresh, isolated QuickJS context, ensuring it never gets direct access to your host systems. If you're building code interpreters, automated data analysts, or autonomous agents, this is the security primitive you've been waiting for. Don't build custom sandboxes when battle-tested frameworks like this are becoming available.

Bypassing the Search-and-Parse Loop in Agent Pipelines

A recent guide showcasing a multi-agent GTM workflow built with Seltz highlights a major shift in how agents gather data. Instead of forcing an agent to search, fetch, and parse raw HTML, Seltz returns fully structured people and news records in a single MCP call. I think this structured-first retrieval is the future of agentic RAG because it bypasses the fragility of scraping. It saves massive token overhead and prevents agents from hallucinating during the data extraction phase.

Big Tech

OpenAI and Anthropic Escalate the Custom Chip Wars

OpenAI is testing its new Broadcom-designed 'Jalapeno' chip, while Anthropic just poached Google TPU pioneer Amir Salek to run its silicon push. What matters here is that the top labs realize software optimization alone won't protect their margins. For practitioners, this is a clear signal that inference costs are going to plummet once these custom stacks go online. Stop optimizing for today's pricing and start building agent architectures that assume intelligence is practically free.

Apple's M6 Silicon Doubles Down on Local AI Devs

Apple's new 2nm M6 chip and updated Mac mini/Studio lineups are quietly targeting local AI development. Thanks to their unified memory architecture, these machines are becoming the industry standard for running 70B+ models without renting expensive cloud GPUs. I think this local-first dev loop is incredibly underrated for rapid prototyping. If you aren't already testing your agent pipelines locally on Apple Silicon before deploying, you're lighting cloud spend on fire.

Anthropic Swaps Hype for Infrastructure Polish

Instead of chasing raw benchmarks, Anthropic quietly rolled out a 4x smoother streaming renderer for Claude and enterprise-managed auth for MCP connectors. This is worth watching because the battleground is shifting from model capability to developer experience. As a practitioner, I find the zero-OAuth setup for enterprise data sources a massive quality-of-life upgrade. It shows Anthropic is building for the unglamorous reality of corporate IT departments.

Struggling to optimize your agent pipelines or cut soaring token costs? Book a free AI audit with me at consult.kylemzhang.com to get your workflows running lean.

Get this in your inbox every morning.

Free, daily, unsubscribe anytime.