How the 4 New AI Models Change How You Work

Jul 9, 2026 · Episode Links & Takeaways

MAIN STORY

How the 4 New AI Models Change How You Work

The government-imposed pause on model releases is over, and the backlog is arriving all at once. This week alone brought GPT Live, Grok 4.5, Cognition's SWE 1.7, and a fresh round of extended GPT-5.6 Sol vs. Fable 5 comparisons — on top of Fable's return earlier this month. The throughline isn't just raw capability: it's a shift in how these models are architected to work together, from voice models that hand off reasoning to background LLMs, to coding models built explicitly to serve as implementation agents under a smarter orchestrator.

GPT Live
"It feels magical and real," says Sam Altman
OpenAI's new voice model drops the old turn-based system for full-duplex processing — it can listen and generate speech at the same time, deciding many times a second whether to talk, pause, interrupt, or quietly hand a task off to GPT-5.5 (soon 5.6) running in the background. It comes in two flavors: GPT Live 1 for paid subscribers, GPT Live Mini for free users, both with that background reasoning hookup enabled. The consumer-facing "grannies" demo became one of the week's most talked-about pieces of AI marketing, showing the model handling live translation, multi-part fact-checking, and the exact frustration-proofing that's dogged Siri and Alexa for years. Early testers ran the gamut: Derya Unutmaz and Riley Brown described genuine skeptic-to-daily-user conversions, Simon Smith and Ethan Mollick framed it as part of a broader shift from AI-as-tool to AI-as-colleague, and one X poster called it the moment AI "stops feeling like software and starts feeling like a live cognitive presence." But Gail Weiner pushed back hard that the voice is gorgeous while the reasoning underneath stays shallow for real work, and a viral clip from Husk — where the model confidently miscounted the E's in "seventeen" — was a reminder that the voice layer's polish doesn't guarantee frontier intelligence.

OpenAI Introducing GPT-Live
OpenAI (X) Intro thread with the "grandmas" demo video
Sam Altman (X) GPT Live feels "magical and real"
Riley Brown (X) GPT Live's similarity to Thinking Machines' interaction model
Simon Smith (X) Awesome for learning, especially language
Derya Unutmaz (X) Early access impressions of GPT Live
Riley Brown (X) This one was unexpected… I was note excited for voice models
Simon Smith (X) GPT Live feels like "Her” or “Jarvis"
Ethan Mollick (X) GPT Live and Claude Tag as new modes of working with AI
Max Weinbach (X) Using GPT Live inside Codex with thread management
Prinz (X) The AI "genie" analogy
Gail Weiner (X) Just a pretty voice, intelligence is lacking
Husk (X) Not so smart, how many E’s in seventeen?
Sightbringer (X) "Interface collapse" and voice moving AI from tool-use to relationship-use

Grok 4.5
Near-Opus performance, priced like Haiku
Grok 4.5 is the first release from the new SpaceXAI-Cursor partnership, framed explicitly as "the first model trained specifically for coding and agents" and built for real-world engineering across large, multi-repo codebases. It roughly matches Opus 4.8 and GPT-5.5 on benchmarks like Terminal-Bench 2.1, SWE-Bench Pro, and DeepSWE 1.0, and topped Artificial Analysis's SaaS-workflow agent benchmark outright. The bigger story is cost: Artificial Analysis found it ran about a third the price of GPT-5.5, a fifth of Opus 4.8, and a ninth of Fable 5 per task, using less than half the tokens of Opus's runs. Reviewers like Theo were won over by how pleasant it was to work with on heavy multi-target tasks, even while conceding it doesn't match Fable or GPT-5.6 on thoroughness, and Elon Musk himself admitted Fable is still ahead on raw capability — "but most tasks don't require Fable-level capability." The emerging use case isn't Grok replacing the frontier orchestrator model, but sitting underneath Fable or GPT-5.6 as the cheap, fast implementation agent.

SWE 1.7
Cognition's near-frontier coder, built for speed over polish
Cognition's SWE 1.7, built on a Kimi K2.7 base, lands just behind frontier coding models on standard benchmarks but at roughly a third to half the cost — and it's fast, serving at 1,000 tokens per second in its Lightning mode. Cognition framed the release as evidence that RL post-training hasn't hit a ceiling even on an already heavily fine-tuned base model, calling out breakthroughs in multi-region training and self-compaction for long-horizon tasks. Early users focused less on the benchmarks than on what the speed itself changes: Nader Dabit noted tasks that used to justify walking away now finish before you've mentally moved on, and JP Zenoware simply called it too fast to watch. That speed plays into a gap swyx (Sean Wang) has written about previously — the uncomfortable middle ground between tasks fast enough to sit with and complex enough to fully delegate — which SWE 1.7 seems built to close.

GPT-5.6 Sol vs. Fable 5
"A wise owl" versus "a rottweiler," per Peter Gostev
The extended reviews of GPT-5.6 Sol kept rolling in, and a consensus is forming: it isn't smarter than Fable 5, but it's a far more relentless executor. Peter Gostev's now widely-shared framing — Fable as a thoughtful "wise owl," Sol as a "rottweiler" that won't let go of a task until it's done — captured that split, and he ultimately concluded the two are different enough in feel that it's worth using both and learning where each one earns its keep. Dan Shipper's team at Every, who'd had internal access for over a month, landed on their own analogy: "GPT-5.6 is like a Porsche, Fable is like a warp drive" — reaching for Fable on the loosest, longest assignments and Sol for the work that fills the rest of the day. Teammate Austin Tedesco said going back to 5.5 after 5.6 felt like "trying to shoot a basketball that's twice as heavy as the one I'm used to using." Lawyer Prinz reported that 5.6 can now effectively replace an associate for legal research where the relevant authorities are all publicly available online, crediting its needle-in-a-haystack search over Fable's.

Claire Vo (X) Things GPT-5.6 Sol is significantly better at
Dan Shipper (X) GPT-5.6 is a much better writer than Fable
Prinz (X) GPT-5.6 Sol Pro saturates Prinzbench on legal tasks
Prinz (X) GPT 5.6 replaces associate-level legal research
Dan Shipper (X) "GPT-5.6 is like a Porsche, Fable is like a warp drive"
Katie Parrott (X) Losing access to GPT-5.6 during the government crackdown felt like shooting hoops with a heavier basketball
Every (X) Full GPT-5.6 Vibe Check
Peter Gostev (X) Fable vs. GPT-5.6-Sol, the "wise owl vs. rottweiler" thread
Peter Gostev (X) a note of caution on token spend with 5.6-Sol at high settings
Dean Ball (X) Two very different frontier models that are completely distinct
Ethan Mollick (X) Both Sol and Fable are a big jump from prior models, these are the only choices for work where intelligence matters

GPT-6 and Fable 5.1 Rumors
Both labs' next flagships may be closer than expected
Rumor source Leo (@synthwavedd) reports that GPT-6 could arrive within about a month — possibly before the end of July — built on a significantly larger pretrain than the roughly 4-trillion-parameter "Spud" base underlying 5.5 and 5.6, explicitly timed to compete with Fable 5.1, which he says is in late-stage development at Anthropic for a release "in the coming weeks." Andrew Curran corroborated much of this independently, adding that even a ready GPT-6 will likely be held back initially by government review, and that both labs are confident in what they have internally, seeing "nothing above us but air. No ceiling." Elon's rumored 10-trillion-parameter Grok in training suggests the entire field is chasing bigger pretrains at once.