Google's AI guard changes as on-device agents and MoE optimization go open-source
Today is a massive transition day. Google's foundational AI brains are spinning out to build the next wave of self-improving models, while the open-source community is dropping insane optimization keys for running multi-model agents on a budget.
Tools & Products
Cursor Open-Sources "Mixture-of-Kittens" (MoK) MoE Kernel
MoE architectures are great for inference, but the training communication bottleneck on NVL72s has been a massive headache. Cursor open-sourcing a fused kernel that hits a 2.37x speedup by overlapping math and networking is a huge win for anyone building custom MoEs. It shows that the best developer tooling companies are winning not just on UX, but by solving deep CUDA-level bottlenecks. This is a massive unlock for teams looking to train custom expert mixtures without enterprise-scale budgets.
Liquid AI Ships On-Device 2.6B Agentic Model
Running a capable agent fully on-device used to be a pipe dream, but LFM2.5-2.6B is changing the equation. It fits in under 2.5 GB of RAM while beating Qwen-9B on tool-use benchmarks, making it viable for phones and edge devices. For enterprise teams worried about data privacy and token-based billing, this is a massive step toward zero-marginal-cost local pipelines. This model proves that small, optimized weights are rapidly closing the gap with massive cloud endpoints.
Superlinked Inference Engine (SIE) Tackles Multi-Model GPU Sharing
Most real-world AI pipelines use a daisy-chain of small models—like an OCR parser, a reranker, and then an LLM—but keeping dedicated GPUs active for all of them kills your budget. SIE solves this by orchestrating multiple models on a single shared GPU pool, dynamically loading and evicting weights based on demand. If you're struggling with low GPU utilization in production, this is the architecture you should be copying. It bridges the gap between running specialized local models and keeping infrastructure costs sustainable.
Cloudflare OS for AI Productivity
Cloudflare open-sourcing their internal 'Cloudflare OS' shows where the enterprise workspace is heading. It treats files as custom applications that can be dynamically generated or manipulated by sandboxed AI agents on the fly. By open-sourcing it, they're giving teams a solid blueprint for building highly customized, agent-driven intranets without starting from scratch. It is a highly practical starting point for anyone trying to implement secure, agent-first developer workflows.
Cloudflare Programmable Wallets for AI Agents
The biggest bottleneck for autonomous agents right now isn't intelligence—it's that they don't have credit cards. Cloudflare's new Programmable Wallets solve this by giving agents stable cryptographic identities and strictly capped spending accounts. Once agents can safely pay for their own API calls and premium content, we'll see a massive explosion in truly autonomous B2B workflows. This is a crucial infrastructure piece that makes agentic commerce safe and auditable.
Big Tech
The Google AI Leadership Shakeup
Jeff Dean leaving Google after 27 years alongside Demis Hassabis stepping into a Chairman role is the end of an era. This isn't a talent crisis for Google, but it is a clear shift from academic research dominance to raw commercial execution. The fact that Google is immediately investing in Dean's new venture tells you everything about how they plan to keep a foot in the door of whatever next-gen architecture he builds. For practitioners, this highlights that the frontier of AI is rapidly shifting from corporate labs to highly agile, specialized startups.
Google Cloud API Gateway Adds Model Routing
Google putting an OpenAI-compatible routing layer directly in their Cloud API Gateway is a slick move to capture enterprise ingress. Teams are tired of writing custom wrapper services just to failover from Gemini to Claude or GPT-4. By baking this into the serverless gateway level, Google makes it dead simple to run cost-optimized routing without adding latency-inducing middleware. If you are building multi-model architectures, this significantly lowers your infrastructure overhead.
Meta Doubles Ads Foundation Model Training Efficiency
Meta's GEM training framework proves that the biggest gains in AI aren't coming from bigger datasets, but from raw engineering optimization. By utilizing jagged flash attention to kill padding waste and adopting MXFP8 precision, they doubled their efficiency for their recommendation models. If you are training at scale, ignoring these low-level compute-saving techniques is essentially throwing millions of dollars of GPU time directly into the furnace. It's a clear signal that software efficiency, not just hardware scale, is the real competitive moat.
Startups & Funding
Google Veterans Launch Discovery Loop
When the minds behind modern deep learning—including Jeff Dean and Sanjay Ghemawat—leave Google to launch an AI startup, you pay attention. Operating as a public benefit corporation, Discovery Loop is targeting the holy grail of self-improving AI models. Google backing them with compute guarantees they're keeping a tight relationship with the team that essentially built their entire tech stack. For the industry, this is the ultimate validation that self-improving architectures are the next major frontier.
Anthropic Signs $10B Cloud Capacity Deal with Volta
Anthropic securing $10 billion in compute capacity from Volta—specifically tied to a new Norway data center—proves that the infrastructure land grab is far from over. Everyone is trying to de-risk their hardware supply chain before next-generation chips drop. For startups, this is a stark reminder that the baseline cost to compete at the frontier remains astronomical. If you aren't backed by multi-billion-dollar compute alliances, your focus has to be on workflow and application layer value.
Former OpenAI Researcher Launches Conduit for Telepathy AI
Leaving OpenAI to build non-invasive 'thought-to-text' models is a wildly ambitious play. Conduit is betting that the ultimate interface bottleneck for AI isn't typing or speaking, but human cognitive throughput. It's a high-risk, high-reward hardware-software play that feels reminiscent of early Neuralink, but focused purely on building the ultimate natural interface for AI. While it is early days, the transition of talent from top labs into sci-fi interfaces is worth watching.
Ready to stop wasting money on idle GPUs and start building optimized AI workflows? Book a free AI audit at consult.kylemzhang.com and let's get to work.