The Big Ways AI Just Changed

July 3, 2026 · Episode Links & Takeaways

MAIN STORY

The Big Ways AI Just Changed: Why June Was the Most Significant Month in Years

Just in time for July 4th weekend, we're closing the books on what I'd argue is one of the most significant months in the post-ChatGPT history of AI. If the beginning of 2026 was all about the explosion of real agentic use cases, the middle of the year has been about recognizing the consequences and challenges of that increasing capacity — and the rest of 2026 is going to be all about figuring it out from here.

The Setup: From Subsidy Era to Token Scarcity
May and June are a matched pair telling one story.
The story of May was the shift from the AI subsidy era to the token scarcity era, as providers moved from seat-based subscriptions to usage-based models — the inevitable consequence of agentic workloads consuming massively more intelligence than the queries of '24. By early June it got real: Walmart moved from unlimited internal tool usage to token budgets, and Uber, after burning through its AI budget in the first four months of the year, set a $1,500 per month cap on AI spend. Token efficiency and token discipline became the idée du jour, at least for the vanguard.

The Infrastructure Adapts
Quiet moves that signaled the efficiency era arriving.
Artificial Analysis shifted metrics in its core intelligence index to better reflect agentic usage. And in a story I still think is wildly under-discussed, Microsoft pushed not only a new set of proprietary models trained from the ground up, but a product to post-train models to the specific requirements of a particular enterprise customer — a story that got lost in a million other Microsoft announcements, arriving just before token efficiency became the conversation.

Artificial Analysis Intelligence index update
Mustafa Suleyman (X) Introducing Microsoft Frontier Tuning

Fable 5 Arrives
"Finding your most complex problems and letting Fable 5 rip."
On June 10th, Anthropic released Fable 5 — and unlike past disappointing jumps to new numerical categories, it was immediately and clearly much more powerful, particularly for technical and coding use cases (though I've found the improvement extraordinarily clear in every area). The best way I can describe the difference: previous coding models lowered the activation energy to start big projects, but Fable 5 was the first to obviate the completion energy to actually finish them. The new AI Daily Brief website that chunks every episode into shareable components exists because of exactly that. Those first 48 hours brought infinite examples of more complex, more complete work: Riley Brown one-shotting a Replit-style app-building app, creators testing 3D worlds, and one story of Fable 5 building a requested product feature while the customer call was still happening.

The First Cracks
Enterprises said "absolutely not" to 30-day data retention.
Not everything was hunky-dory. Beyond questions about overaggressive guardrails on topics like biology, Anthropic instituted a 30-day retention policy for Mythos-class models — prompts and outputs retained for trust and safety review — which immediately made many enterprises balk. It was a preview of a broader power issue: companies realizing how much their access to one of the most important assets in business is mediated by a single or small handful of companies.

The Government Steps In
A precedent for direct government intervention in frontier AI.
By Friday of that first week, the story transformed: the US government used an export control directive to demand Anthropic suspend Fable 5 and Mythos 5 access for foreign nationals, and Anthropic said the only way to comply was to shut down access for everyone. We'd later learn a narrow jailbreak report from Amazon triggered the flurry of activity — though it served more as a catalyst for parts of the government waking up to how much more powerful this class of models is. The ban extended as GPT-5.6 got delayed too, with OpenAI announcing it would be a set of three models and the government approving access in waves. To many, it felt like the beginning of a messy, ad hoc AI licensing regime — not based in legal precedent, just shooting from the hip.

The Alternatives Boom
Cost and sovereignty both now argue for diversification.
With Fable offline, companies had a second major reason beyond cost to look at alternatives to frontier closed-source models. The month brought a ton of experimentation with routing companies building architectures that send tasks to the right level of model — and enormous interest in new models, none more than Z.ai's GLM 5.2. Since the January 2025 DeepSeek moment, someone has proclaimed a new "DeepSeek moment" every few months, but GLM 5.2 is the first that legitimately earns the label: not as good as Fable 5, but exceeding the Opus 4.6/GPT-5.2 level that initiated the agentic era. For many, it was the first open weight model that made the fallback strategy feel less like compromise and more like genuine frontier competition.

Custom Models and Integrated Architectures
Open weights as raw material, not just fallback.
It wasn't just raw open models getting attention, but custom post-trained models built on top of them — like Cursor's Composer 2.5, built off Kimi — and integrated multi-model systems. Harvey and Fireworks paired an open weight GLM worker with an Opus advisor for legal tasks, beating Opus alone at a fraction of the cost, while OpenRouter's Fusion used a panel of models, a judge, and a synthesizer for hard tasks. It would be wildly overstating it to say everyone switched en masse, but for the first time since I've been doing this show, local AI became a serious boardroom question worldwide.

The Harness Era Accelerates
No new model to play with meant ecosystems got the spotlight.
In the Fable pause, emphasis shifted to the harnesses and ecosystems around models. Both Anthropic and OpenAI pushed dedicated HTML/website artifact builders, getting knowledge workers to rethink the spreadsheets and slide decks they used to use. And in the immediate aftermath of Fable 5 going offline, Microsoft CEO Satya Nadella wrote a long post on X arguing every company needs a learning loop around its AI usage — firms don't just need the right model, they need to own the compounding context, decisions, evaluations, and institutional memory surrounding it.

Claude Tag
65% of Anthropic's product team code now starts in Slack.
One feature announcement worth specific note: Claude Tag, which isn't just another way to talk to Claude in Slack — it lets anyone anywhere in Slack call on the power of Claude Code, democratizing advanced technical capability, giving Claude Code more persistent context, and shifting AI from an individual to a group experience. People took notice partly because of the near-reverence with which Anthropic's own team talked about it, headlined by the claim that 65% of its product team code is now produced by initiating Claude Code from Slack rather than the app or terminal.

Compute Becomes a Market
Memory shortages, neo-clouds, and data center politics brewing.
The token scarcity shift is also about physical limits. One of the big market themes this month was the outperformance of memory companies as the memory shortage came into focus. Compute itself is becoming a market of its own — led by SpaceX, which expanded its Anthropic deal and struck similar deals with Google and Reflection AI, and now reporting suggests Meta and Zuckerberg are following Elon into that accidental neo-cloud space. Meanwhile, data centers become more of a political hot button every month: June was a fairly low ebb, but when you've got Erin Brockovich on one hand and former Tea Party conservatives on the other mobilizing against the same thing, it's going to be part of the discourse.

Meanwhile, in the Average Enterprise
"Bot sitting": 6.4 hours a week making AI usable.
While a tiny sliver of early adopters deals with token efficiency, companies in the average band of adoption are uncovering new challenges around agentic work — like the phenomenon of "bot sitting" identified in a Glean report, which found workers spending an average of 6.4 hours per week feeding agents context, checking outputs, and rerunning underwhelming results. June reinforced that the capability overhang won't be solved by new models — in fact, new models make it worse — only by real change management. On that front, the latest KPMG quarterly pulse survey showed growth in CEOs actively owning AI as a strategic priority, and found organizations where CEOs are accountable for AI are more than twice as likely to report meaningful business value.

The Big Questions for July
Fable is back, but nothing is actually resolved.
Despite Fable's return, we don't know how the government's agreement with Anthropic affects the release of GPT-5.6, or how this ad hoc licensing regime handles the even more advanced models reportedly waiting in the wings — so expect more and more questions on the policy side. For companies, I think the lasting legacy is a significant Overton window shift away from being locked into whatever the state-of-the-art closed frontier model is. These aren't short-term changes; they're setting a map for the rest of the year and beyond. In the immediate term, though: big chunks of the corporate world turn off for this part of the summer. If you don't — if you use July and August to really see what this new class of models can do — you have a chance to significantly increase your value to whoever you need to be valuable for.