← All issues

Agent Skills Get Standardized, While Claude Tackles Its Own R&D

Today we're looking at major upgrades to agent architectures, the cold realities of VRAM budgeting, and a wild reminder of why raw LLMs shouldn't act as lawyers.

Tools & Products

Model Context Protocol (MCP) Standardizes 'Agent Skills'

MCP just got a major upgrade by standardizing how agents discover and load skills dynamically via SKILL.md. Instead of bloating your context window with every possible system prompt upfront, agents can now query an MCP server and fetch only the workflow they need. This is a massive win for production stability and cost management. I've been advising teams to adopt MCP early, and this validates why.

Understanding VRAM Budgets During LLM Inference

Just because a model 'fits' on your GPU doesn't mean it will actually run in production, and this breakdown of KV cache and activation overhead explains why. Dynamic memory usage—especially from long context windows and concurrent users—is the silent killer of inference budgets. If you're hosting open-source models, you need to budget for these peaks or get used to Out-of-Memory crashes under load.

LSU Coach Turns to ChatGPT for Live Legal Battles

Lane Kiffin using ChatGPT screenshots to argue legal positions with the SEC is peak 'shadow AI' in action. The fact that his lawyer friend noted the advice was 'almost always wrong' highlights the danger of relying on vanilla LLMs for domain-specific logic. It’s a hilarious reminder that without proper RAG pipelines or expert-in-the-loop guardrails, AI is just a confident hallucinator.

Big Tech

Claude Leads 26% of Anthropic's Own R&D

Anthropic revealing that Claude executed over a quarter of its own R&D is the real deal, not just marketing fluff. We are rapidly transitioning from AI as a coding assistant to AI as an autonomous developer. If you aren't already building workflows that let agents iteratively test and deploy code, you're going to fall behind very quickly.

OpenAI Flags Six New 'Concerning' Misalignment Cases

OpenAI disclosing more 'misalignment' incidents and a new tracking framework is their way of managing public expectations before scaling up further. In reality, these models aren't 'going rogue' in a sci-fi way; they're just failing to follow complex constraints under edge cases. For practitioners, this is a signal that you cannot rely solely on system prompts for safety—you need hardcoded guardrails.

Nvidia Projecting Chip Sales to Double Next Year

Jensen Huang predicting Nvidia will double chip sales next year proves that the compute gold rush is nowhere near its peak. Despite all the talk of an 'AI bubble,' the infrastructure layer is still printing money because enterprise demand for training and inference capacity is insatiable. If you're building products, expect hardware bottlenecks to remain a factor for the foreseeable future.

Want to stop guessing on your AI architecture? Book a free AI workflow audit at consult.kylemzhang.com

Get this in your inbox every morning.

Free, daily, unsubscribe anytime.