TL;DR
Frontier models like GPT‑5.6 are now being rationed by governments, while export controls on Anthropic's models are loosening and Google is getting blocked in return, so access to top AI is turning into a geopolitical patchwork.
Under that ceiling, Chinese and open-weight models plus speculative decoding and quantization are making high-end capabilities cheap enough to run locally, which is colliding hard with a post-tokenmaxxing era where FinOps teams finally care what all those tokens are buying.
Key Events
Report
The center of gravity in AI just slipped away from who has the biggest model to who controls access, infra, and harnesses. Frontier models are being gated like weapons at the same time that Chinese and open-weight stacks and ultra-efficient decoding quietly eat the rest of the market.
GPT‑5.6 Sol ships as a state-of-the-art model, topping Terminal-Bench 2.1 and beating GPT‑5.5 on Sol Ultra scores, but its rollout is restricted to a small set of government-approved partners at explicit U.S. request.
The White House is effectively deciding who gets to use GPT‑5.6, approving access customer by customer and treating the model as sensitive infrastructure.
In parallel, the Department of Commerce just lifted export controls on Claude Fable 5 and Mythos 5, restoring access to over 100 organizations and enabling broader deployment after years of cybersecurity-driven restrictions.
Regulators also blocked Gemini 3.5 Pro from the U.S. market, underscoring how access to top-tier models now depends as much on jurisdiction as on capability.
GLM‑5.2 is widely described as the first Chinese model to match or beat American public models, surpassing Claude on benchmarks, scoring 1524 Elo on coding and agent tests, and becoming the most-liked model on Hugging Face.
Enterprises are already securing compute to post-train GLM‑5.2 in-house, while 60% of companies tracking AI budgets report shifting toward cheaper models and open-source Chinese options.
Qwen 3.6 27B is called the local sweet spot for coding, pushing around 100 tokens per second on a single RTX 3090 and handling summarization and dev workflows for many users.
LongCat‑2.0 extends this with a 1.6‑trillion-parameter MoE language model, whose Owl Alpha variant is now the most-used model on Hermes Agent at about 10.1 trillion monthly tokens and ranks among the top three globally by volume on OpenRouter.
Ornith‑1.0 then shows an open MIT-licensed stack matching Claude Opus 4.7 on agentic coding tasks, with even its 9B variant beating Qwen 3.6 35B on some benchmarks.
Community experiments now show that harness configuration alone can swing coding-task accuracy by over 11 percentage points, and many practitioners report that the harness often matters more than the base model for code work.
ZCode acts as a dedicated harness for GLM‑5.2, while Harness APIs and similar stacks expose multi-agent presets and reproducible A/B tests that effectively turn orchestration into a configurable layer above models.
Hermes Agent's Mixture-of-Agents presets behave like virtual models that beat Claude Opus 4.8 by 8% and GPT‑5.5 by 11% on internal benchmarks, with the latest release closing 692 high-priority issues.
In sharp contrast, LangChain- and LangGraph-based agents have shown 30% silent failure rates in production, test suites with 87% pass rates that still miss regressions, and support bots that can issue incorrect refunds despite hitting the right APIs, while developers complain about CI pipelines slowed by 18 minutes of agent evaluation.
Real-world users of Lovable-style AI app builders describe impressive three-week MVPs that later buckle under deployment and maintenance, echoing the pattern of strong demos but fragile systems.
DeepSeek's DSpark speculative decoding reports throughput gains between 51% and 400% over traditional MTP while cutting single-token inference FLOPs to roughly 27% for a 1M context variant.
Qwen3.5‑MoE with speculative decoding reached about 1.22× performance with 91% draft acceptance, and NVIDIA's inference software stack is credited with up to 5× speed-ups that directly lower token costs.
On the hardware side, NVFP4 quantization lets Qwen3.6‑27B hit around 130 tokens per second on an RTX 6000 Blackwell, while an official GLM‑5.2 NVFP4 build delivers mid-teens tokens per second at 128K context on four DGX Sparks.
ComfyUI's new convrot INT8 models run more than twice as fast as FP16 and GGUF on many NVIDIA GPUs, and community benchmarks of Krea2 INT8 ConvRot confirm large speed gains in image generation.
Against that backdrop, tokenmaxxing is now widely described as a dead idea. Meta logged about 73.7 trillion internal tokens in a month at an estimated cost of $221 million.
One company accrued a £300,000 bill for AI usage and then halted most tools. A four-agent loop ran for 11 days and spent roughly $47,000 before anyone noticed.
Around 60% of enterprises now report hard guardrails on AI spending, and 98% of FinOps teams say AI bills fall explicitly under their remit.
What This Means
Policy is turning frontier models into gated utilities just as Chinese and open-weight stacks, aggressive quantization, and agent harnesses make everything below that layer cheaper, faster, and more chaotic. The power center in AI is drifting from individual models toward whoever owns the routing, eval, and spend controls across that stack.
On Watch
Interesting
We processed 10,000+ comments and posts to generate this report.
AI-generated content. Verify critical information independently.
Sources
Key Events
On Watch
Interesting