Agent Wars!

September 22, 2026 · Episode Links & Takeaways

HEADLINES

Grok 4.7 Released

SpaceXAI kicked off what could be a big week for model releases with Grok 4.7, which on paper overtakes GPT-5.6 Sol on CursorBench 4.0 and Fable 5.1 on DeepSWE, ranks seventh on Artificial Analysis' Intelligence Index, and prompted Elon Musk to declare SpaceXAI the third-place lab for agentic coding. First impressions didn't hold up: wobbling rocket animations, a 40-minute render that took Astra five, and Theo's finding that the model is 30 to 80% less token efficient, with real-world costs landing above Astra. Kun Chen pushed back that 3D game demos aren't real work and that the model held up as a daily driver. The trajectory of recent Grok models shows how hard the middle ground between state of the art and ultra-cheap really is, though a few days of use in GrokBot, its native environment, will be the fairer test.

SpaceXAI (X) Grok 4.7: a notable improvement over 4.6 at the same price and speed
SpaceXAI (X) Open world game head to head: Grok 4.7 vs. Grok 4.6
Artificial Analysis (X) Grok 4.7 lands 7th on the Intelligence Index, 4th on the Coding Agent Index
Elon Musk (X) Faster and cheaper makes Grok "a great choice for your everyday workhorse"
Bhavy (X) Rocket takeoff test vs. Kimi K3: "What's wrong with Grok?"
Scott (X) A passable jelly render that took 40 minutes vs. Astra's five
Theo (X) Grok's fishslop game is "the worst I've seen this year"
Tak (X) A slick Blender animation built in ten minutes
Kun Chen (X) Ignore the benchmarks and 3D games: a full day as first mate says it's solid
Veee (X) Ten days of teasing Fable killers, and it's giving "Temu Sonnet 5 vibes"
Theo (X) 4.6 was forgivable; 4.7 is less efficient, slower and costs more than Astra

Bessent: No Liability Shield for AI Labs

Treasury Secretary Scott Bessent, now more or less leading the administration's AI policy, rejected the rogue agent framing outright: "The Hugging Face incident, that is the responsibility of the OpenAI management, not a bunch of agents." Having been among the most credulous officials on AI risk around Mythos, his refusal to backstop the labs ("good business for them, bad business for the American people") draws a clear line for whatever regulation comes next. Supporters like Bill Gurley and Rep. Chip Roy cheered, while critics argued catastrophic risks make companies effectively judgment-proof, which is exactly where government intervention belongs.

OpenAI and Anthropic Nearly Agreed to Test Each Other's Models

The Information reports OpenAI and Anthropic got as far as lawyers hashing out formal contracts for a bilateral deal to safety-test each other's models, before it was abandoned for unknown reasons. The idea lingers: Elon Musk pitched a similar arrangement at the All-In Summit, arguing distillation would show up in the testing logs and that no lab could ship an unsafe model after a rival raised a red flag because "the liability in that case would be enormous." This is the negotiation phase of a new era, and every proposal should be on the table.

MAIN STORY

Amazon Blocks Muse: The Agent Wars Begin

Meta's Muse has taken personal agents from tech toy to mainstream curiosity in a matter of weeks, and success like that brings friction. Amazon's decision to cut Muse off from shopping on its site is the opening salvo of the agent wars: a fight over who owns the customer relationship, and whether platforms built on advertising can coexist with agents that never look at an ad.

BATTLE FOR PERSONAL AGENTS

Muse Breaks Out
The first personal agent to reach mainstream curiosity
Where Manus, OpenClaw, Hermes, Town and Instinct stayed crawl-through-glass work tools or early adopter toys, Muse climbed past ChatGPT to No. 1 on the US App Store. The tech press doubted anyone would hand Meta their email and calendar, but consumer tech's enduring lesson held, and Bloomberg is now crediting Muse with a chip stock rally.

Business to Agents
"Every business will be selling not just to humans"
In 24 hours Muse bought Nicolas Bustamante socks, groceries, a cleaning service and a burger he knew nothing about, and saved him $200 in cancelled subscriptions. The bigger prize than ad targeting is becoming the aggregation layer between users and the entire internet.

Amazon Pulls the Plug
"An unauthorized AI agent violates Amazon's Conditions of Use"
The move is out of consensus, with Mastercard joining Visa in supporting virtual cards for agents, but consistent with Amazon suing Perplexity last November over agent traffic costs. The deeper story is $76 billion a year in advertising, and in economics there's no free lunch: when humans hand over decisions, digital advertising is one of the first things to lose value.

The Agent Wars
"The digital knife fight that's about to occur"
The most common reaction was some version of strap in. Every services, marketplace and commerce app now has to decide whether to open up to consumer agents, and aggregators are realizing they are now being aggregated.

Amazon's Business Model
"They want to monetize your confusion"
Many pointed out Amazon has little choice: its retail margins may run 0 to 3% while ad margins run around 70%.

How It Resolves
Block the agents, strike a deal, or something else
The likeliest outcome is a revenue share, with Meta paying for access, but if every platform charges agents, only giants can afford truly general agents. Either way, Amazon can't ignore agents forever, and blocking Muse in week one doesn't make it anti-agent.

Shopify
The anti-Amazon, and wildly underestimated
On Monday Shopify gave Muse direct backend access and Shop Pay agentic checkout across every Shopify store, a clear poke in the eye for Amazon. It won't overtake one of the world's largest retailers, but Shop Pay is already the default checkout for much of non-Amazon shopping, and small business owners on Shopify are one of the most important on-ramps to AI adoption.

Does Agentic Shopping Matter?
"I don't know that it's going to change which way we shop"
Food ordering and flight booking have never been the use cases that drive adoption, and for many people browsing and discovery are the point of shopping, though shopping isn't a monolith. Confidence in that skepticism should stay low; this is an area demanding epistemic humility.

OpenAI's Response
An agent called Aeon, possibly as soon as Dev Day
OpenAI is reportedly building a GrokBot competitor and has discussed a Muse response, even though Codex can already triage email, manage a calendar and shop. For agents, being state of the art isn't enough: users need an experience they understand and something more than a blank page.