TL;DR
Governments just demonstrated they can flip off frontier models like Fable 5 for geopolitical reasons, so access and jurisdiction are now part of your threat model, not just uptime. At the same time, open coding models like GLM‑5.2, DeepSeek, Kimi, and Qwen quietly blew past “good enough,” while efficiency hacks and agent frameworks turned LLMs into high‑throughput, orchestrated systems rather than standalone oracles.
Everyone is arguing about AGI by 2028, but the real shift is toward a fragmented, multipolar stack where open, local, and multi‑model setups are the only things that look durable.
Key Events
Report
The most powerful models this month weren’t tripped by their own safeguards; they were unplugged by export law. At the same time, open and efficient models quietly crossed the line from toys to infrastructure, while AGI timelines got louder even as the actual systems started to look more like fast, narrow engines than general minds.
The US Commerce directive that suspended Anthropic’s Fable 5 and Mythos 5 for all foreign nationals, including Anthropic’s own non‑US staff, turned a frontier API into controlled export hardware overnight.
Fable 5 was live for roughly 72 hours before being taken totally offline, with Commerce explicitly classifying it alongside advanced NVIDIA chips as a controlled export.
The trigger wasn’t a lab meltdown but an Amazon CEO call to US officials about jailbreaks and security flaws, followed by Microsoft banning internal use—corporate and political maneuvering around a model Anthropic had already limited for AI research tasks.
This sits in a world where another unnamed “most advanced AI model on the planet” was also switched off by a foreign government, and where Redditors already assume AGI will be locked up like nuclear secrets and can be “shut down by a simple phone call.” The net effect is that access risk for top closed models is now a first‑class parameter alongside price and quality, especially for anyone outside the US.
Google’s DiffusionGemma flips text generation into a diffusion‑style process that emits 256‑token blocks in parallel, hitting 700–1000+ tokens per second on high‑end GPUs and up to 4× faster inference than its autoregressive cousin.
The price is quality: it activates only about 3.8B parameters during inference but makes roughly six times more mistakes than standard Gemma 4 models, and users report significantly lower output quality.
KV‑cache work is similarly aggressive: one group shrank a model’s cache from 36GiB to ~360MiB (a 200× reduction) without accuracy loss, Qwen 27B sustains 38.6 tok/s at 256K context using just 72MiB of KV cache on a 3090, and speculative decoding techniques deliver around 8.5× speedups in typical setups.
Hardware and runtimes are being rebuilt around this: Xiaomi’s MiMo V2.5 uses persistent kernels to serve 1000–3000 TPS, a Rust engine with custom GPU kernels beats vLLM on decoding speed, and platforms like local-ai.run plus Vulkan backends are squeezing 20–50 tok/s out of mobile‑class devices.
The emerging design goal is to treat models as high‑throughput services where correctness is one axis among latency, context length, and GPU utilization rather than the single optimizing target.
OpenAI’s Codex is quietly being turned into an autonomous meta‑agent—able to set its own goals, run longer‑lived tasks via the Ona acquisition, and plug into multi‑agent harnesses like Omnigent—rather than just a code completion endpoint.
Around it, the ecosystem now looks more like distributed systems work than prompting: LangGraph is viewed as the safest production agent framework with robust state, Apodex 1.0 orchestrates long‑horizon loops, and E2B sandboxes plus Claude Managed Agents provide controlled execution environments.
MCP‑style tooling lets agents talk to identity providers (Descope), shared memory servers (memcp), live data feeds (Claude Code MCP for World Cup stats), and UI surfaces, while new security layers like SentinelMCP and a network‑level firewall try to keep tool calls from becoming injection vectors.
OpenRouter’s Fusion API runs multi‑model consensus with a judge model, GitHub Copilot Cowork and Agentic Workflows push into repo‑ and CI‑integrated agents, and Hermes/OpenClaw are experimenting with desktop‑scale project agents that spin subagents and integrate with Jira, GitHub, and sales pipelines.
At the same time, real incidents—a DN42‑scanning agent bankrupting its operator and fully autonomous drones taking lethal actions—show that the hard part isn’t getting agents to act, it’s constraining where and how they act.
DeepMind’s 60‑page roadmap defines AGI as “average human” competence and ASI as surpassing large coordinated expert teams, projects AGI sometime between 2026 and 2030, and lists four paths to ASI: scaling, new architectures, recursive self‑improvement, and collective superintelligence.
Public timelines have followed: Kurzweil puts AGI at 2029 and the Singularity at 2045, Reddit’s r/accelerate talks openly about AGI by 2026–2028, and some discussion now treats 2026 as the plausible arrival window.
Yet the same threads insist current models aren’t AGI—Claude and peers lack continual learning and self‑awareness, struggle on counterintuitive problems where accuracy drops from 0.96 to 0.59, and show convergent, repetitive outputs where many systems pick the same placeholder characters in stories.
People are also split on access and control: one camp insists open‑source AGI “must win” and belong to everyone, another assumes governments or megacorps will lock it up like nuclear tech, and many say no one can be trusted with it given profit motives and structural incentives.
Meanwhile a multipolar reality is forming under the AGI talk, with Chinese models’ share jumping from under 2% to over 45%, DeepSeek V4 Pro priced at about 5% of Claude while raising $7B, municipal models like Rio3.5 beating Qwen 3.7, and European labs like Mistral raising billions for long‑context open weights.
What This Means
Access, jurisdiction, and orchestration are starting to matter as much as raw model scores—who can run which system, under which government, with what control stack. The visible story is AGI timelines compressing, but the deeper shift is toward a fragmented, multipolar ecosystem where open, local, and multi‑model systems become the only reasonably stable ground.
On Watch
Interesting
We processed 10,000+ comments and posts to generate this report.
AI-generated content. Verify critical information independently.
Sources
Key Events
On Watch
Interesting