Why AI Washing Won't Work Much Longer

August 4, 2026 · Episode Links & Takeaways

HEADLINES

Palantir's Monster Quarter Becomes a Referendum on the Frontier Labs

Quarterly revenue came in at $1.94B, up 93% year over year and beating expectations, with commercial sales up 149% — accelerating from a 133% growth rate in Q1 and demonstrating that enterprise AI demand is still booming despite flashy headlines about token budget cuts. It turns out that caps, which most organizations haven't even gotten close to hitting, are not the same as cuts. Margins expanded too, with net income reaching a billion dollars for the quarter and growing at a 225% annual pace, and hiked annual forecasts sent the stock surging 10% after hours. Alex Karp called the quarter "otherworldly" and used the earnings as another turn on his increasingly loud bully pulpit for AI sovereignty, writing in the shareholder letter that every organization is "awakening to the risks of handing the creators of language models the keys to their institutions." In a CNBC follow-up he made clear exactly who he takes issue with: "We have people trying to drug addict us to a future they believe they control. Now, I've spent a lot of time with Dario and the effective altruism crew."

OpenAI Bites Back at Apple

Apple's suit claims that via an employee who left for OpenAI, the LLM lab stole Apple trade secrets. OpenAI's new letter opens by calling Apple one of the greatest companies of all time with a reputation for obsessing over the smallest details, then adds that "this careless, aggressive, and oddly personal lawsuit sadly doesn't live up to that reputation." The jaw-drop line: Apple had claimed it contacted OpenAI in February and got no response, but now admits its outside lawyers emailed the wrong person after confusing two Asian last names — and only conceded that once OpenAI raised it. This is a little more psychodrama than this show normally gets into, and it will only matter if it actually shifts the plans OpenAI can make in hardware, but it's a wild enough mistake to be worth capturing.

DeepMind's Chief Strategy Officer Reframes Capex as a Bet on RSI

Speaking on a panel at UC Berkeley, Google DeepMind CSO Jasjeet Sekhon called recursive self-improvement a key part of the investment thesis behind sky-high capex. This earnings season has been all about how the ROI of AI lines up with endlessly ramping spend, and Google in particular was punished for AI income failing to keep up in the short term as free cash flow flipped negative. Sekhon acknowledged that current AI revenues "don't sustain the capital expenditures we're making so far," creating a "danger we could hit an AI air pocket such that the expenditures happen but the revenues don't show up." The counter-argument is that the buildout was never about near-term revenue — it's the "biggest scientific bet civilization has ever made."

Claude Helps Researchers Uncover a Decades-Old Flaw in Criminal DNA Databases

Forensic labs across the country store DNA samples as digital files in a searchable database, and researchers found they could alter those files using code written by Claude in a process taking around 45 minutes. The core issue is that the software dates to 1995 and includes none of the tamper-evident protections of modern systems — forensic scientist and New Haven University professor Laura Gaydosh Combs put it as data files called the gold standard of forensic science that "lack the same level of tamper-evident markings that we require for a paper bag." No lab has reported this kind of tampering, but none has figured out how to detect it either, and researchers noted some encryption in use relies on a key that's been on the internet for years. For the AI story, this is an example of defense-focused researchers finding serious vulnerabilities in ancient systems still running in a very high-stakes field, and there are certainly hundreds or thousands more like it across critical infrastructure. Cost is perhaps the biggest reason those systems have never been overhauled or properly tested, which makes rudimentary AI-assisted testing on a modest budget a potential game changer: yes, it's scary that cybercriminals have new tools, but AI is and remains a massive upgrade for defenders as well.

White House Hosts AI Companies to Review Its Voluntary Framework

The administration is convening a set of AI companies to discuss its voluntary review framework for frontier models. Well worth watching what comes out of it, but the solid reporting isn't in yet — this one gets picked up tomorrow.

MAIN STORY

Why AI Washing Won't Work Much Longer

The newest Chinese model release matters less for its benchmarks than for who is now paying attention to it. At KPMG's 2025 Tech and Innovation Symposium, the conversation was still genuinely about the percentage of organizations that had one, two, or three AI use cases; at last week's edition it was about governance for a workforce with powerful coding tools, cost provisioning across different models in different parts of the org, and whether to fine-tune open-weights models and have an open-weights policy at all. The level of sophistication — and more importantly the quality of the questions enterprises are asking — has increased dramatically, which means Qwen 3.8 Max may be the first release of its kind that actual enterprise IT folks are watching alongside developers and early adopters.

CAN OPEN MODELS SOLVE ENTERPRISE AI PROBLEMS?

Qwen 3.8 Max
"A new bar for coding and Cowork"
Like Kimi K3, this is a large model, coming in at 2.4 trillion parameters. Alibaba's self-reported numbers put its Terminal-Bench score between Fable 5 and GPT-5.6 Sol and its Cowork bench score between Sol and Fable 5, with visual reasoning and research reproduction results that are totally state of the art — plus state of the art on OSWorld-Verified, the agentic computer use benchmark, which if accurate is the most significant of the bunch for the model's applicability as an agentic work tool. The launch materials also tout ten-plus days of self-evolving development from empty folder to production without hand-holding, with a complete project trace shared on GitHub.

The Ad
"The computers are going to do our jobs for us"
A lot of the discourse online was not about the model at all but about the video ad it launched with: a laptop in the foreground doing coding, science, and knowledge work while the humans presumably paired with it are off fishing, playing tennis, rock climbing, and reading. Without using basically any words, the pitch is that the value of these incredibly powerful models to an individual life is more time to do the things you actually want to do.

Back to Open Weights
First Qwen Max class model to get open weights
Last year brought a lot of personnel shifting around Qwen, and the largest Max series models moved from open to closed — raising the broader question of whether closing up would become commonplace for Chinese labs as they approached the frontier. Now they're back, and the full weights land next week.

The Price
The real headline number, at $2 in and $6 out
Many folks, particularly around the policy sector, still assume every Chinese model is cents on the dollar relative to state-of-the-art American competitors, and that hasn't been true for a while — Kimi K3 is only about 40% cheaper than Opus, at $15 per million output tokens versus $25. Qwen takes it down significantly further, to a little more than a third of Kimi's price and roughly a fifth of Opus's. Still not pennies on the dollar, but a much more meaningful decrease than K3 delivered.

First Impressions
Excitement, then a mystery benchmark takedown
Given the price and the open weights, it's not surprising the early reactions ran hot. The caution came fast: Artificial Analysis published a score of 53 — four points behind Kimi K3 and a point behind Grok 4.5 — then quickly took the scores down without any explanation.

The Skeptics
"Qwen 3.8 Max is unusable"
Ethan Mollick landed on a solid model but not a Kimi K3 level one. Datem was harsher, testing across coding, design, planning, agent orchestration, and multiple harnesses against Kimi K3, Grok 4.5, GLM 5.2, GPT Luna and Opus 5, with Qwen coming in last every time — extremely slow, unstable, and burning through usage like crazy. On Pavel Hurin's BugBench, two real repos with 105 hidden bugs, Qwen found 19 at a cost of around $31, and just getting it to run took five attempts; GPT-5.6 ran the same benchmark for $1.80.

"AI Wishing" and "AI Washing"
"This kind of work doesn't happen in a quarter"
Yesterday's New York Times ran an opinion piece from former Lululemon CIO Julie Averill — the platonic archetype right now of the average essay about enterprise AI — coining "AI wishing" for the belief that AI is magic, that you can wave its wand at a hard problem and skip the work of solving it. Its insidious cousin is AI washing: a company under pressure to show immediate results claiming to do more with AI than it actually is, most destructively in the form of the AI layoff, where the efficiency often doesn't exist yet and the cut is really about freeing up cash. In May, US employers announced 97,000 job cuts, with one research firm finding companies blamed 40% of them on AI, while a separate survey found about a third of hiring managers who cut a role because of AI had already rehired for the same or a similar one — positions eliminated before the work was redesigned, so the work shifted onto whoever stayed. None of this is a critique of AI, which Averill calls the most powerful technology she's seen, but of the very real human processes that determine how it gets integrated.

Efficiency Technology vs. Opportunity Technology
Headline wins now, pummeled later
Organizations that treat AI strictly as an efficiency technology rather than an opportunity technology might eke out a few headline wins in the short term, but will ultimately be pummeled by the companies that understand this as a redesign moment opening up massive new opportunities — not a chance to make shareholders excited about cost cuts for Q3. Doing AI well takes work: organizational redesign, job role redesign, process redesign. Shortcuts are bound to fail.

Where the Two Stories Intersect
"AI cost optimization is now a discipline, not a hack"
A year ago almost no company had any actually expressed policy on open weights models beyond a blanket "of course we're not going to use Chinese models." That conversation has changed — there's now real discourse about whether open weights can be part of a complete AI system that uses different types of models for different types of tasks, and it goes well beyond op-eds. Thinking Machines Lab, the spinoff from former OpenAI CTO Mira Murati, introduced Tinker for fine-tuning late last year; Microsoft is building a frontier tuning service on top of its lower-cost MAI models aimed squarely at this era of increasing AI complexity; and new routers arrive seemingly every day, with reports that Stripe is about to buy OpenRouter for $10 billion. Enterprise AI leaders are exploring routing, but not by buying whatever Gartner says is best and calling it a day — they're working out what combination of third-party and internal solutions fits, and whether a router is even the right answer at all.

Plenty of listeners who are the AI leaders inside their organizations will be shaking their heads, wishing the situation described here were the one they're dealing with, and there's no minimizing how steep the hill is for opportunity-AI advocates trying to get their companies to handle this the right way. But for basically the first time since ChatGPT launched, enterprise conventional wisdom around AI is getting directionally correct. As attitudes toward AI wishing and washing change, the PR value and board plaudits that used to reward shallow deployment go away — which helpfully cuts off the incentive loop to do AI the wrong way, and leaves room for the genuinely exciting work of redesigning around what these models can actually do.