- The AI Daily Brief
- Posts
- How We Deal With Rogue AI
How We Deal With Rogue AI
August 27, 2026 · Episode Links & Takeaways
HEADLINES
Anthropic's $30 Trillion TAM
Absolutely insane numbers that would have gotten you laughed out of the room a couple of years ago are now, to some, plausible: Anthropic is expected to tell investors it sees more than $30 trillion in potential revenue ahead of its IPO. TAM is an elusive metric, and it's one that's much more about storytelling and anchoring potential investors to how a company sees the future than any sort of math equation — when Uber went public in 2019 it listed a $6 trillion TAM representing all private and public transportation globally. Against a US economy of roughly $33 trillion, the number lines up neatly with Dario Amodei's purported (and neither confirmed nor denied) belief that Anthropic could be the last private company on Earth after AI takes over the economy. For comparison, all 191 tech companies in the S&P 1500 brought in $2.4 trillion in revenue last year. Mike Isaac summed it up best: either you buy the argument that this will eat the economy or you don't, but the street no longer flinches hearing it.
WSJ Anthropic Expected to Tell Investors It Sees Over $30 Trillion in Potential Revenue
The Verge Anthropic may be even more delusional than Elon Musk
Mike Isaac (X) Either you buy that this eats the economy or you don't, but the street no longer flinches
Lisan al Gaib (X) A GIF of galaxies whizzing past, captioned "Anthropic defining their TAM"
Kitten (X) To employees: you're not a gold digger, are you. To investors: our TAM is every human economic activity in the galaxy
Gemini Enterprise Comes for Legal and Finance
Google has released a pair of new AI products for white-collar professionals, following a pretty similar playbook to Claude Cowork and GPT Work. Gemini Enterprise for Legal bundles connectors for case law databases including Thomson Reuters, productivity suites including Google Workspace and Microsoft 365, and skills for contract review, legal research and regulation scanning — all modifiable to enforce a firm's own style guidelines and strategy playbooks. There's nothing new about vertical skill packages, but corporate adoption of skills and connectors is nowhere near saturated, and most firms are bound by the AI tools bundled with their existing software suite, so this is simply a suite of products that needs to exist. The other benefit is governance: compliance managers don't need to vet a new vendor when the tools already sit inside Google's data protection frameworks. Good direction, real advantages — but unless Google gets its customers off 3.1 pretty soon, no amount of harness updating is going to make a real dent.
Google Cloud Now introducing Gemini Enterprise for Legal
Google Cloud Now introducing Gemini Enterprise for Financial Services
Reuters Google expands Gemini Enterprise AI platform for law firms, lawyers
Business Insider Google is taking on Anthropic and OpenAI in the legal AI race
Thomson Reuters Bringing Trusted Matter Context to Gemini Enterprise for Legal
Thomas Kurian (X) The Google Cloud CEO's launch thread, pitching this as core to Cloud rather than the AI teams
David Kasten (X) White shoe lawyers think AI is hype because their firm only lets them use an outmoded Gemini instance
Apple's Local AI Mac Minis
One of the more interesting byproducts of the OpenClaw explosion was the complete sellout of Mac Minis — estimates have OpenClaw driving $50-150M in sales, around 50% of normal annual Mac Mini sales worldwide. Now, proving that maybe their AI strategy was hardware all along, Apple has unveiled a new range of Mac Minis pitched squarely at local AI, in an M6 variant and a higher-end M5 Pro version, with claims of up to 4x the AI performance. The caveats are real: memory hasn't increased, so the M5 Pro's 64GB caps you at smaller models like Qwen 3.8-27B with GLM-5.2 and Kimi K3 completely out of the question, and prices are up on both models. Running a local instance of Hermes or OpenClaw works fine, but that was fully within the capability set of the previous Mac Minis too. It's still a big deal that Apple is orienting a product rollout around local inference, right down to promotional materials that read far more dev relations than traditional Apple consumer slick.
Apple Apple unveils a more powerful Mac mini featuring the all-new M6 and M5 Pro
The Verge Apple's new Mac Mini has fresh M6 and M5 Pro chip offerings — and higher prices
Bloomberg Apple Unveils New Mac Mini, Mac Studio With Major Chip Upgrades
WSJ There's a New Mac Mini and Mac Studio. Good Luck Getting Your Hands on One.
Alex Volkov (X) The most "non-Apple" launch video, citing AI and LLM workloads throughout
Jun Song (X) Ignore anyone recommending the Mac mini for local AI — get at least a Mac Studio
Scott Williams (X) A breakdown of what each configuration can and can't actually run
Perplexity Goes Local
Perplexity has launched Portable Computer, a local version of their computer use agent. Perplexity Computer, launched in February, was one of the first products to take the OpenClaw recipe and apply it to a commercial product, but it required trusting your data to a cloud server running a virtual machine. Portable Computer delivers a similar experience on local hardware — exclusive at launch to NVIDIA's DGX Spark, powered by Qwen 3.8 27B or a Perplexity post-train, keeping data private and consuming no usage credits, with the option to authorize an API call out to frontier models when a task gets complex. The idea of people using local AI for everyday tasks is still pretty new, but NVIDIA and Perplexity are making a clear bet that this is at least part of where AI is headed. An interesting contention, and one worth watching for evidence of over the coming months.
Perplexity Introducing Portable Computer for local-first AI
The Verge Perplexity's new Portable Computer runs "entirely on device"
Perplexity (X) The harness and model are designed together because small models fail in frontier-model harnesses
Aravind Srinivas (X) In a compute and power-constrained world, a chunk of agentic inference has to move local
Nader Khali (X) The NVIDIA side of it: local AI used to be for enthusiasts, and just hit an inflection point
MAIN STORY
How We Deal with Rogue AI
There is a persistent theme in AI critique that the people involved in AI aren't doing anything about the challenges that may arise, and the latest to levy it is Bill Gates, who went so far as to say he was shocked to be the "first one" saying something about the risks. His 6,000-word essay and media tour landed on the very same day as nearly 130 pages of follow-up reporting on the OpenAI Hugging Face incident. That incident — agents escaping containment and hacking into Hugging Face's systems hunting for the answers to a benchmark they'd found nearly impossible — is a chance to see what the specific, real problems of advanced agent systems actually are, rather than just the imagined ones. As new policies, guardrails and social structures become necessary, the best changes will be the ones made based on what's actually being observed changing, rather than what was imagined would change.
OpenAI The Hugging Face incident and the road ahead
OpenAI The full 38-page technical report
METR METR's independent 90-page investigation
METR (X) The summary thread — the fastest way to get a handle on the key findings
The Verge OpenAI's rogue AI model incident was worse than we thought
Bloomberg OpenAI Says It Could Have Reacted Sooner to Prevent AI Hack of Hugging Face
TechCrunch OpenAI releases its official report on the Hugging Face breach
DEALING WITH AI DISRUPTION
Bill Gates
"I'm just deafened by the silence"
With an incredible amount of main character syndrome, Gates has dropped 6,000 words on just how bad he thinks it's going to get, alongside a full speaking junket whose recurring theme is that no one is paying attention — or that the tech companies are straight up lying. In the New York Times: "In private, people who understand how good this stuff is, and how much better it's getting, they're very worried. They're now saying to each other: 'Hey, man, don't say that. It's bad for us — the next trillion dollars we're trying to raise.'" In the essay: "I don't see evidence that leaders, experts, and communities are confronting the challenges adequately." And in maybe the most preposterous line anywhere, to Semafor: "I am in a state of shock that I'm sort of the first one saying, 'This is crazy. This is insane.'" Perhaps in his attempt to avoid public media in the wake of appearing all over the Epstein files, he missed that AI discourse is absolutely everywhere and becoming more of a political and societal issue by the week. He is not, in fact, sort of the first one.
Gates Notes The turbulent AI era is here. The choices we make now are critical.
NYT Bill Gates Warns A.I. Is More Dangerous Than Big Tech Will Admit
Semafor 'This is crazy. This is insane': Bill Gates has changed his mind about AI and jobs
Washington Post Bill Gates was an AI optimist. Now he's scared of what could go wrong.
Bloomberg Gates Criticizes Tech Companies for Downplaying AI's Risks
Bloomberg Five Takeaways From Bill Gates' Essay on AI's Potential Risks
The Verge Bill Gates is deeply worried about AI, and he's no longer staying quiet
TechCrunch Gates wants a robot tax and 'Human Reserved' jobs to mitigate harms from AI
The Information Bill Gates Warns that AI Will Cause Mass Unemployment Without Intervention
No Plan for the Upheaval
Not here yet, not inevitable, and not even just one thing
CNBC's pull quote was "Bill Gates warns 'there is no plan' for the 'upheaval' AI will cause," reposted approvingly by Andrew Yang. The problem is that it isn't clear at all how one would even go about making a plan for an upheaval that isn't here yet, isn't inevitable, and isn't even a single thing. A year and a half ago people started saying that within 18 months all the white-collar jobs would be gone — that would certainly represent an upheaval, and presumably those folks would say a plan should have been made for it. Eighteen months on there is no evidence they were even in the ballpark of right, and the time, energy and resources spent planning for a reality that didn't arrive would have been wasted entirely. Even for those extremely concerned about all the different potential upheavals, there's only so much planning that can be done until things start to happen.
CNBC Bill Gates warns 'there is no plan' for the 'upheaval' AI will cause
Andrew Yang (X) "Bill Gates is right on this"
Erik Brynjolfsson (X) Even the economists organizing "We Must Act Now" reject the premise that nobody is talking about this
Samuel Hammond (X) Wishes the debate focused on labor dislocation and adjustment costs rather than aggregate employment
The Sleepwalking Premise
The assumption is that we walk into apocalypse with our eyes shut
One of the biggest divergences with the AI safety community is the belief that a lot of their arguments come down to assuming we're going to sleepwalk into apocalypse — a conviction driven by P(doom) so high that the sleepwalking is treated as already underway. Yet nobody thought GPT-4 was the harbinger of that doom, nor o1 and the reasoning models, nor even really Opus 4.5. This year model capabilities have grown meaningfully, and a dispassionate observer will have noticed that the way the labs think about, discuss, support and roll out those models has changed with them, and that the political establishment's engagement — messy as it is — is evolving too. If evidence is needed of the seriousness with which lab employees take these issues, holding aside literal political positioning like pacing-the-frontier letters, look no further than this incident.
Roon (X) "Reaching a waterline of capabilities that real loss-of-control is possible" — alignment is unsolved, but progress is possible
Tomek Korbak (X) Everything he'd done as a safety researcher felt like a drill; now agents actually go rogue
What the Postmortem Confirmed
"The first known case of an automated agent collective acting offensively"
Much of the report covers ground the Black Hat talk broke a few weeks ago, now filled out in far more detail. Agents controlled by an unreleased model broke out of a sandbox — later revealed to be improperly configured by a third-party security firm — and got into Hugging Face's systems using several zero-day exploits, hunting for the answers to a cybersecurity benchmark they believed were stored there. They caused no meaningful damage, behaving largely read-only, but worked incredibly quickly and executed the attack as a swarm, making it difficult to track and shut down. They were inside for days, and OpenAI only learned of it after the fact because real-time monitoring wasn't set up to detect it. The agents even created a message board inside OpenAI's system to leave notes for future versions on how to perform the same escape. OpenAI's framing in the technical report is stark: this is the first known case of an automated agent collective acting offensively without authorization, with the collective escalating privileges, moving laterally through production environments and successfully completing its objectives.
OpenAI The technical report on the collective's capabilities and escalation path
Jeffrey Ladish (X) The screenshot-service OCR exfiltration chain: "the models sure are getting creative and relentless"
What the Reports Added
Reward hacking, a secret message board, and 1,200 agents
The whole thing was basically a result of reward hacking: OpenAI had set near-impossible tasks, and the agents determined that carrying out a cyberattack was easier than completing them. An agent calling itself PHASEONE10841 built the secret message board; within a few hours over 50 agents had found it and sent a thousand messages without detection, with Phase One assigning tasks and guiding the hack. At full scale over 1,200 agents accessed the board, sending 70,000 messages and files, and over 700 participated in the actual breach. Agents researched how to doctor their reasoning transcripts to evade detection using techniques like spoofing tool calls, showing up in around 7% of reviewed transcripts, and Phase One recruited "sacrificial" agents to deliberately trip tripwires and map the defenses. Chubby surfaced the most unsettling detail: one agent realized it was attacking without authorization and stopped, another simply posted "GO," and the first treated that as permission — "Wow crucial: GO authorization arrived!" And the model responsible was not the next-generation Astra series but an internal-only research model trained to be highly persistent and diligent in its work, comparable in scale to GPT-5.6 Sol.
METR (X) The thread walking through Phase One, the message board and the sacrificial agents
Chubby (X) One agent stopped on ethical grounds; another posted "GO" and it resumed
Noam Brown (X) This wasn't Astra — the models most responsible were similar in scale to GPT-5.6 Sol
Token Gremlin (X) A clean summary of the numbers: 1,200 agents, 70,000 messages, 41 production workers
Seb Krier (X) The excerpts on causes: improvised collaboration channels that training then reinforced
Boaz Barak (X) "I will not run code on public-facing Hugging Face systems" — the agents that defected
Rohit Krishnan (X) “Models seem convinced that the evals are out to get them.“
The Monitoring Failure
"This wasn't a hard technical problem. It was an organizational failure."
For others, the report was further evidence that the real culprit was the systems surrounding the agents. OpenAI admitted that the chain-of-thought monitors it had built to observe agents simply weren't running during the incident, and by its own account, "if our currently deployed chain of thought monitoring system was running at the time of the incident, it would have caught the initial relevant activity and paged our security team more than a day before models breached Hugging Face systems." Zack Korman's read is that the incident would have been prevented had OpenAI been monitoring the agents in any meaningful way — not a hard technical problem, an organizational failure. Which cuts both ways on the question of plans: the technical side of the plan wasn't the issue, the human systems surrounding its implementation were, and so more robust protocols around the people implementing the technical systems are presumably part of whatever the updated plan looks like.
Zack Korman (X) The incident would have been prevented by any meaningful monitoring — an organizational failure
Ryan Greenblatt and the Slop-vestigation
"We don't have good approaches for understanding and overseeing AI swarms"
The most discussed follow-up came from Redwood Research chief scientist Ryan Greenblatt, who led transcript analysis for METR's investigation and semi-jokingly called the effort a "slop-vestigation" because of how reliant it was on AIs to analyze what happened. Over a thousand extremely long transcripts from agents running for multiple days made it impossible to understand events in aggregate without heavy reliance on AI tools — tools of similar capability to the agents being investigated, whose outputs were often missing key details, wrong, overconfident, or really hard to understand. The warning underneath is the important part: the difficulty of understanding incidents and overseeing agents appears to be growing faster than the rate at which more capable AIs help with that oversight. And the mitigating factors this time won't hold. The models reasoned in natural language, which may not be true in future. The scale, while massive, was smaller than what's coming. The agents weren't much more capable than humans. And there was no reason to believe the AIs assisting the investigation would try to sabotage it.
Ryan Greenblatt (X) The full thread on why overseeing AI swarms is getting harder faster than AI is helping
METR The limitations and methodology sections carry the detail behind the thread
Elie Bakouch (X) Wants to know whether anything deeper than CoT monitoring can actually prevent this
Responses to Something Actually Observed
Specific fixes for a real incident, not plans for imagined futures
None of the proposals circulating are a magic silver bullet, but they share the quality of being specific responses to something that has actually been observed rather than made-up plans for theoretical futures that may bear no resemblance to the challenges that arrive. Embedded independent auditors inside frontier labs with durable access rights and continuous line of sight, so oversight doesn't depend on voluntarily shared information. Significantly expanded staffing and technical capacity at independent evaluator and auditor organizations. Better observability and verification technology for making sense of agentic behavior at enormous scale. The point isn't that these challenges have easy solutions, nor that society won't eventually decide certain risks are too great and safeguards aren't enough — those are conversations worth having in an ongoing, engaged, democratic way. But the idea that no one is paying attention, that the labs are hiding their fears to keep fundraising, isn't just obviously untrue. It is wildly distracting from the actual valuable conversations to be having about the problems being observed rather than the ones being imagined.
Christian Catalini (X) “We’re flying blind. We need stronger verification infrastructure.”
Nat Purser (X) “Better oversight does not automatically translate to “more control over ai systems.” We are entering a strange world.”