← All issues

Giant Open Models, Agent Governance, and the Cost of Rushing AI

Today we are seeing the realities of scaling AI in production, from Moonshot's massive 2.8T parameter model architecture to DoorDash's practical blueprint for governing autonomous agents. Meanwhile, Google's latest product rollback reminds us of the dangers of deploying generative tools without rigorous safety boundaries.

Tools & Products

DoorDash Builds Centralized Gateway for AI Agent Tool Access

DoorDash is showing everyone the blueprint for enterprise agent infrastructure by putting a governance layer over Model Context Protocol tools. Everyone talks about building agents, but no one wants to talk about how you manage API credentials, rate limits, and security when an LLMs is running wild in your systems. By centralizing this into an Agent Gateway, they've turned fragile prompt-based integrations into managed, observable infrastructure. If you're building agentic workflows for your team, this is the exact architecture you should copy.

Big Tech

Google Earth Instantly Yanks AI Image Generation Feature

Google Earth launched an AI feature allowing users to generate hyperrealistic edits on satellite imagery, only to pull it within 24 hours due to immediate deepfake exploits. It turns out that letting anyone easily generate realistic bomb craters or military buildups on actual map data is a major security risk. This is a classic case of product teams rushing cool AI features to market without running basic adversarial testing first. If you are building user-facing generative tools, let this be a warning to implement strict input and output filtering before your users turn your product into a disinformation machine.

Research

Moonshot AI Releases 2.8T Parameter Kimi K3 Model

Shipping a 2.8-trillion parameter open model is a massive engineering flex, but actually running a 5-terabyte file is a nightmare for most dev teams. What I find fascinating here is the architectural pragmatism, specifically their use of LatentMoE and FP4 Quantization-Aware Training to keep memory usage semi-reasonable. This is a clear signal that the future of frontier AI isn't just about scaling compute, but how aggressively we can compress these models for production. If you are building enterprise apps, keep an eye on how these quantization techniques trickle down to smaller, highly-efficient local models.

OpenAI Teases Astra Model with Core Math Breakthroughs

OpenAI quietly previewed its new Astra model, which reportedly solved ten long-standing mathematical and theoretical computer science problems. What's crucial here isn't just the raw reasoning, but the fact that the solutions were automatically verified using the Lean theorem-proving language. By combining LLM generation with deterministic formal verification, OpenAI is bypassing the typical hallucination bottleneck of LLM math. I think this hybrid approach is the real path forward for high-reliability software engineering and complex reasoning tasks.

Industry

Chime Lays Off 10% of Workforce in Push for AI Efficiency

Chime laying off 150 workers to restructure around AI is the quiet part being said out loud across the fintech industry. We are officially past the era of using AI just to write marketing copy; it's now about restructuring headcount to optimize operational margins before an IPO. I expect to see a lot more legacy and mid-stage tech companies use the AI efficiency playbook to justify leaner engineering and operations teams. For practitioners, this means the highest leverage roles right now are in workflow orchestration and system automation rather than maintaining status-quo headcount.

Want to design secure, highly-efficient AI workflows for your engineering or ops teams? Book a free AI audit with me at consult.kylemzhang.com to get started.

Get this in your inbox every morning.

Free, daily, unsubscribe anytime.