← All issues

Post-Training is the New Moat: DeepSeek's Flash Upgrade and RAG Latency Solved

Today we are seeing a massive shift away from raw parameter scale and toward smarter architecture. From DeepSeek's retrained flash model to deep optimization tricks for RAG, the focus is squarely on efficiency and real-world execution.

Tools & Products

DeepSeek V4-Flash Beats Its Own Pro Version

This is a massive wake-up call for teams chasing raw parameter size. DeepSeek proved that smart post-training and retraining can make a smaller, cheaper flash model outrun its own flagship Pro version on actual agentic tasks. If you're building agentic workflows, you should immediately swap your endpoints to test this. It's a direct win for both latency and your API bill.

Cloudflare Computer Reinvents Agent Infrastructure

Cloudflare is quietly building the actual infrastructure runtime for AI agents. By wrapping a virtual file system inside Durable Objects with SQLite, they're solving the hard state-management problem that makes agentic workflows break in production. This is the exact kind of plumbing we need to move away from fragile stateless APIs. It points the way toward persistent, sandboxed agent runtimes.

Firecrawl Releases pdf-inspector to Bypass Slow OCR

For anyone building RAG pipelines, parsing PDFs is a notorious bottleneck that eats up time and money on OCR. Firecrawl's new Rust tool bypasses OCR entirely for native text PDFs in 20 milliseconds, which covers more than half of your typical documents. It’s a dead-simple utility that instantly optimizes ingestion pipelines without needing complex GPU infrastructure. I highly recommend dropping this into your extraction pipeline.

Research

RAG Latency is Actually a Prefill Problem

Most developers blame vector databases or search algorithms for slow RAG responses, but the actual bottleneck is prefill latency. Processing massive chunks of retrieved text scales quadratically and keeps the GPU waiting for the first token. Implementing KV cache reuse and selective recomputation is the highest-leverage engineering fix you can make today to slash your inference bills. Stop optimizing your search queries and start looking at your cache hit rates.

Process vs. Outcome Reward Models in LLM Reasoning

There's a major shift happening in how we train reasoning models, moving from outcome-based rules to process-oriented step checks. OpenAI's data shows process reward models outperform outcome models by catching 'hallucinated reasoning' that accidentally stumbles into the right answer. For practitioners building custom evaluations or fine-tuning pipelines, grading intermediate reasoning steps is the new gold standard. It’s how we move past simple pattern matching.

OpenAI Unlocks 10 Math Breakthroughs via LLM Reasoning

OpenAI showing off ten legitimate math and computer science breakthroughs discovered by an unreleased model is a huge signal. This isn't just about writing boilerplate code; it's about AI acting as a legitimate partner in deep, high-dimensional reasoning. It confirms that the boundary of what LLMs can 'reason' through is expanding far faster than the skeptics realize. Pay attention to how these reasoning capabilities filter down to commercial APIs soon.

Ready to optimize your own LLM pipelines and slash your API spend? Book a free AI audit at consult.kylemzhang.com to get started.

Get this in your inbox every morning.

Free, daily, unsubscribe anytime.