← All issues

Runbooks Over Raw Memory: Why Your AI Agents are Failing in Production

Today's data highlights a massive shift from pure agent hype to cold, hard engineering reality. We are finally getting the empirical data needed to build cost-effective, reliable AI workflows rather than just impressive demos.

Tools & Products

Anthropic Updates the Model Context Protocol (MCP) Roadmap

This is a massive signal for anyone building agentic workflows. Anthropic's MCP is quickly becoming the default open standard for how LLMs safely connect to data sources and tools. The new focus areas—like agentic messaging primitives and HTTP-native transport—mean we are moving away from ad-hoc API wrappers and toward a unified transport layer. If you are building custom tool integrations today, align them with MCP now or prepare to rewrite them in six months.

TrueFoundry Open-Sources TrueForge for Enterprise AI Agents

I've been saying for months that context bloat is the silent killer of startup margins. TrueForge tackles this head-on with built-in context compaction and delayed tool-schema loading, which keeps your system fast and your API bills low. It's a highly pragmatic, self-hostable alternative to heavier framework abstractions. If you're building governed, enterprise-ready agents that need to run in your own VPC, this belongs on your evaluation list.

DuckDB v2.0 Swaps SQL Parser to Enable Runtime Grammar Extensions

This might look like a low-level database update, but it's a huge deal for local-first AI architectures. By replacing its PostgreSQL-derived parser with a PEG-based one, DuckDB makes it incredibly easy to extend SQL syntax at runtime. For developers building AI agents that generate and execute SQL on the fly, this eliminates a ton of parsing errors and enables much tighter, custom dialect integration. Local analytics engines are fast becoming the default runtime for LLM data tools.

Research

Empirical Study of 8,100 Trials Reveals How AI Agent Skills Actually Work

This is the most important research paper I've read this week. The data proves that "skills" don't actually teach agents new facts—instead, they function as procedural anchors that keep the agent from getting derailed. Standardizing your agent's experience into a step-by-step checklist (a "runbook") improves task success by over 6% compared to just dumping raw execution logs into the context. Stop trying to make your LLMs smarter through raw memory; focus on building rigid operational runbooks.

Scaffolding Choices Drive Agent Costs More Than Interface Protocols

We are officially in the optimization era of AI engineering, where architectural decisions have immediate financial consequences. This study revealed that bare-bones CLI agent scaffolds were between 5x and 28x cheaper to run than complex protocol-heavy setups, regardless of whether they used MCP. The lesson here is clear: don't over-engineer your transport layer. Keep your agent scaffolding as lightweight as possible to avoid burning your runway on unnecessary token overhead.

Pragmatic Design: Why Not Every Workflow Needs an AI Agent

The hype cycle is cooling, and practical engineering is taking over. Trying to solve every software workflow with a fully autonomous agent is a recipe for high latency, unpredictable behavior, and massive API bills. The winning architectural pattern right now is hybrid: use LLMs strictly for cognitive reasoning and translation, but keep your routing, business logic, and guardrails strictly deterministic. If you can write a standard if/else block to handle a step, do not let an LLM decide it.

Ready to optimize your AI infrastructure and cut down on waste? Book a free AI audit with me at consult.kylemzhang.com and let's streamline your workflows.

Get this in your inbox every morning.

Free, daily, unsubscribe anytime.