← All issues

GPT-6 'Astra' Drops, But the Real Story is How We Control and Route These Beasts

Today's drop of OpenAI's Astra—the first of the GPT-6 line—has everyone talking, but building real workflows right now requires looking past the benchmark hype. I'm focusing on the infrastructure: specifically how we route, secure, and customize these massive models without burning through our entire runway.

Tools & Products

OpenAI Drops GPT-6 'Astra' to Mixed Reviews

OpenAI just dropped Astra, pitching it as the first in their GPT-6 family, and the initial reaction is incredibly polarized. It absolutely crushes the ARC-AGI-3 benchmark and shows insane speed in computer-use tasks, but early builders are finding it to be a massive token hog. In my experience, these initial frontier releases are always 'show horses' rather than stable production workhorses; they require heavy optimization and burn cash fast. If you're building agentic workflows today, don't rush to migrate your entire stack to Astra just yet. Treat it as a preview of where reasoning is going, but stick to more predictable models for your core production pipelines until we see the inevitable, more optimized iterations.

Anthropic Invites Customization with Claude Code Plugins

Anthropic is quietly testing a plugin system for Claude Code, which is a massive win for engineering teams wanting tighter control over agentic behavior. By letting developers write plugins to log actions, restrict agent permissions, or tweak the CLI interface, they are addressing the biggest barrier to enterprise adoption: trust. In my workflow consulting, the number one fear teams have is an agent running wild on their codebase or violating compliance. Providing a structured way to enforce local guardrails directly within the developer's terminal is exactly how you move AI coding tools from 'novelty' to 'standard issue.' This is the kind of boring, practical infrastructure improvement that actually moves the needle for engineering velocity.

DigitalOcean Solves the AI Agent Token Drain with Session Pinning

We've all heard that LLM routing is the holy grail for saving money, but in agentic loops, naive routing actually ends up costing more. When you switch models mid-conversation, you completely wipe out your prompt prefix cache, forcing the new model to re-process the entire chat history at full price. DigitalOcean’s new Inference Router tackles this head-on by introducing 'session pinning,' which locks an agent to a specific model after the initial routing decision is made. It uses a highly optimized 1.5B classifier model that runs at the proxy level, resolving intent without adding a second expensive API call. For anyone building multi-turn agents, this is a masterclass in why infrastructure-level routing is vastly superior to DIY application-layer logic.

DeepTeam Demystifies Red Teaming with Open-Source Automation

Evaluating your LLM for correctness is useless if a user can easily prompt-inject their way into leaking your database. Red teaming has historically been a luxury reserved for big labs with massive budgets, but the open-source DeepTeam framework is changing that by simulating adversarial attacks in production. It dynamically tests your application against over 40 vulnerabilities—like PII leakage, bias, and jailbreaks—without requiring you to hand-craft complex test datasets. For engineering teams, running this as a standard CI/CD step is the easiest way to catch security gaps before they hit the real world. If you are deploying user-facing LLMs or agents, tools like this are no longer optional—they are a baseline requirement.

Ready to stop burning tokens and start building production-ready AI workflows? Book a free AI audit at consult.kylemzhang.com to get your stack optimized.

Get this in your inbox every morning.

Free, daily, unsubscribe anytime.