← All issues

Agent Engines, Speed Tiers, and the Real Cost of Inference

Today we're looking at a fundamental shift in AI architecture. As models optimize for raw speed and scale, the real battle is moving from weights to the surrounding infrastructure.

Tools & Products

xAI Launches Grok Bot for Autonomous Browser Operations

xAI's Grok Bot bypasses APIs entirely by logging directly into your browser tools, representing a massive shift toward GUI-based agents. At $120 a seat, it's a steep entry price, but the real value is automating legacy software that doesn't have an API. For workflow builders, this means we can finally stop writing fragile integration connectors and start treating the browser as the universal interface. It's early, but this is the practical direction agentic workflows are heading.

DeepSeek Open-Sources MIT-Licensed Agent Framework

DeepSeek's release of Harness v0.1 is a direct shot at closed agent platforms, offering a completely modular, plugin-first architecture. Making everything from memory to the file system a swappable plugin is exactly how developers actually want to build. This MIT-licensed approach gives teams full ownership over their agent stacks without vendor lock-in. It's rough around the edges, but it's a great playground for anyone testing multi-step agent behaviors today.

OpenAI Ships GPT-5.6 Sol Ultrafast Powered by Cerebras

Generating 750 tokens per second on Cerebras hardware completely changes what's possible for real-time agent loops. Usually, speed means sacrificing model intelligence, but keeping the weights on-chip bypasses the memory transfer bottleneck entirely. If you're building voice agents or multi-step reasoning chains, this reduces latency from minutes to seconds. This is the hardware optimization we've been waiting for to make real-time agentic UX viable.

InsForge Challenges Supabase with Agent-First Backend Design

Standard backends are designed for human eyes, which is why agents waste millions of tokens querying tables and trying to parse generic errors. InsForge is an open-source alternative designed specifically for agents, returning a full topology in just 500 tokens. In benchmarks, it cut token consumption by nearly two-thirds compared to Supabase for the same RAG app. If you're building agentic apps, optimizing your backend structure is the lowest-hanging fruit to slash your API bills.

Big Tech

Google Releases Gemini 3.7 Flash with Temporary Pricing Cuts

Google is shipping models at a frantic pace, dropping Gemini 3.7 Flash just three weeks after 3.6. The 50% price cut through the end of the year makes it an incredibly cheap option for agentic workloads. They're clearly targeting developers who are sensitive to token costs for high-volume tasks. It’s a smart land-grab, but teams need to watch if the rapid-fire release schedule introduces breaking behavior in their prompts.

Google Rolls Out Sheets Canvas for Gemini Mini-Apps

Sheets Canvas brings interactive Gemini-powered UI layers directly on top of spreadsheets, essentially letting non-technical users build mini-apps on the fly. This is a massive play for enterprise productivity because it keeps the data where business teams already work. Instead of export-import cycles, you get a dynamic, natural language interface on top of your existing sheets. It proves that the future of enterprise AI isn't separate portals, but deeply integrated canvas spaces.

Chief Revenue Officer and COO Part Ways with OpenAI

With Chief Revenue Officer Denise Dresser leaving just days after COO Brad Lightcap's exit, OpenAI is seeing a major reshuffling of its business leadership right before its highly anticipated IPO. While the technical research side remains dominant, managing enterprise sales and commercial growth is getting highly competitive. For enterprise customers, this leadership churn might introduce some friction in enterprise agreement negotiations. Dali Rajic from Wiz has his work cut out for him.

Research

Agent Scaffolds Outperform Raw Frontier Model Upgrades

Recent benchmarks show that custom agent scaffolds produce up to 22x larger performance swings than swapping frontier models. This is a crucial lesson for builders: throwing a more expensive model at a problem is rarely the solution. The secret sauce is the 'agent harness'—how you manage context, verify tool execution, and handle errors. Focus your engineering effort on building a robust orchestration loop rather than waiting for the next flagship model to save you.

How Production LLMs Scale Reasoning at Inference Time

We are moving away from purely training-time compute toward test-time (inference) compute to drive reasoning. Techniques like chain-of-thought, majority voting, and tree search essentially trade tokens and latency for accuracy. But as DeepSeek's R1 work showed, simple rule-based rewards often beat complex Monte Carlo Tree Search systems at scale. For practitioners, the key is knowing when to pay the inference premium for hard math vs. keeping the loop lean for simple text tasks.

Small Cheap Runs Can Predict Large-Scale Model Scaling

A new study proves that tuning hyperparameters correctly allows researchers to predict scaling laws from tiny, inexpensive training runs. This is massive for smaller labs and startups trying to train custom models on a budget. It democratizes the pre-training process, taking the guesswork out of multi-million dollar compute runs. While big tech has the capital to burn, this research levels the playing field for efficient, targeted model training.

Startups & Funding

Anthropic Targets $2 Trillion Valuation with Surging Revenue

With projections of $100-120 billion in annualized revenue by the end of 2026, investors are whispering about a potential $2 trillion IPO for Anthropic. What's interesting is how they financed their massive compute buildout before this revenue spike, proving that institutional capital is highly comfortable backing AI infrastructure. This massive valuation puts immense pressure on Anthropic to maintain its lead in the enterprise agent space. The commercial war between them and OpenAI is only going to escalate.

Selena Gomez's Mental Health Startup Wondermind Hit with Fraud Lawsuit

Investors are suing Selena Gomez and her co-founders over failed mental health startup Wondermind, alleging they hid functional issues and failed to deliver a promised app. This is a classic cautionary tale of celebrity-backed AI and tech plays where distribution is prioritized over actual product development. For the startup ecosystem, it highlights why due diligence on execution capability is more important than follower counts. Founders need to build real utility, not just count on social media hype.

Trump-Backed Crypto Startup World Liberty Financial Approved for Bank

World Liberty Financial, the crypto venture backed by the Trump family, has received preliminary approval to launch a bank. This represents a major convergence of decentralized finance, regulatory policy, and political influence. While the technical details are still thin, it's a clear signal of the shifting regulatory environment toward crypto and digital assets in the US. Keep an eye on how this affects liquidity and transaction rails for decentralized web3 agents.

Tired of burning budget on fragile AI integrations? Let's design a high-leverage agent architecture for your business. Book a free AI workflow audit at consult.kylemzhang.com

Get this in your inbox every morning.

Free, daily, unsubscribe anytime.