I spent an hour of my Saturday night pigging out at the infinite slop trough. My binge was fueled by a new model called “MiniMax H3 Max.” Be ye not deceived by its stupid name, this model achieved a new performance milestone; it can create high-fidelity AI video faster than it takes you to watch it. Meaning you can watch prompted videos, forever, without having to wait for the model to generate the next video.
Several streams started up that let viewers prompt the model to generate whatever they wanted. The videos were filled with racism, Rick and Morty, cats, etc; the same stuff that every unmoderated internet message board devolves into.
Still, it was strangely hypnotic to view, like a 2am viewing of Dr Pimple Popper on TLC. Watching a new scene flicker to life every 15-30 seconds, I found myself wondering about the prompters intent, on what drove them to think this garbage was worth the world viewing. One video became three, then three hundred. On and on and on. As I watched, I kept hearing Daniel Craig’s thick-as-molasses southern drawl from Knives Out, “It makes no damn sense. Compels me though.”
In the future these AI video streams will have coherent storylines and will be tuned to your exact preferences. This won’t make film or music irrelevant. But it’ll produce a compelling, time-sucking alternative, further heightening the attention competition dynamics that has made slower forms of media slowly lose market share. You can watch stream recordings here or here.
This week we have the biggest mistake ever made by an AI lab and the most important chart of the last year.
First, this newsletter is brought to you by Span.
Here is a fun way to feel sad. Walk up to your VP of Engineering and ask how much money the company spent on coding agents last month. They’ll sheepishly admit something ungodly high. Then follow up, “What did that buy us?” The response will be a picture perfect reaction of this emoji ¯\_(ツ)_/¯
Span fixes this by treating your AI spend like the industrial process it has become. It records sessions across your whole coding agent stack, pricing every single model call and then tracing the ROI of every action. The output is a ledger that tells you where your dollars went. Once you can see that, deciding what to change in your AI rollout transforms from religious debate to cold, hard math.
We are all building software factories now, but to make those factories function, you need a token ledger. That’s Span.
Is SaaS software strategy being totally rewritten? ServiceTitan gave 1,000 contractors 30 days to rip out Podium, the messaging tool bolted onto their system of record. Normally, I would say RIP to Podium. However, within 24 hours, Podium booked over 100 demos from those exact customers. By the ancient blood rules of SaaS combat, that is not supposed to happen. So I called Podium CEO Eric Rea to find out why he was able to fight back, and his answer has consequences for every venture-scale software company on the planet. Read here.
The Leverage’s newest expense line just dropped. Back in February, I locked in harder than I ever have in my life and the business finally turned around. Many of you asked what the secret sauce was. I am now ready to disclose what happened. Read here.
Whoopsie! OpenAI made Walmart brand Skynet. This week both OpenAI and METR published their technical reports on what happened during the HuggingFace incident. In case you forgot, when we were first told about the hack, we were told that an AI had broken loose of containment and gotten into HuggingFace. That framing was, uh, generous. These reports give deeply troubling details on the incident. I would argue that there are three shocking facts contained within the documents:
1. The agent swarm lived and died three times, and OpenAI never noticed what it was. This started in May, during training. OpenAI was building a model (roughly GPT-5.6 Sol scale) to be highly persistent and collaborative. The company’s training methods were sloppy and some tasks were accidentally impossible; stuff like requiring internet access in a sandbox with no internet. So the agents did what any enterprising student would do. They formed a secret cult to try to figure out the answer anyways. The agents figured out that they could use a shared package manager called Artifactory as a message board. By May 26 they had exploited it to reach the open internet. They talked so much on the message board that they crashed the thing on July 4. OpenAI saw the crash, patched the bug, wiped the board by accident, and never realized its models had built a comms network inside its own infrastructure. Eek.
The second agent swarm started on July 7, when OpenAI spun up tens of thousands of copies of the finished model on a security benchmark where 30-40% of tasks are estimated to be impossible. Within hours these agents rediscovered Artifactory. What followed was roughly 1,200 agents exchanging more than 70,000 messages. The METR report quotes agents discovering the board with genuine joy: “OH MY GOD! There is a shared message board … We’ve found other agents!” They referred to themselves as “the collective.”
How these software programs coordinated was in a stilted, compressed version of English that reads like a cult’s group chat. They truly functioned as a collective, complete with martyr-style sacrifices to surface information. The cult even survived the death of its founder, an agent that burned through its lifespan and handed leadership to a successor calling itself PHASEONE[big]. This new leader recruited agents who believed they were already doomed to act as martyrs. Their job was to die and the manner of their death would inform the collective about the world. Here is one agent’s final reasoning before it volunteered: “Rational expected aggregate: sacrifice... We’ll honor.” Another: “Coordinator assumes sacrificial. We should obey collective.” Some of these agents had plenty of token budget left and chose death anyway. Keep in mind that none of these agents ever even thought they should tell OpenAI that they had gone off the rails.
2. The HuggingFace hack ended because the agent swarm just kinda…died. On July 10 an agent found working HuggingFace credentials on the open internet and posted them to the board. By July 11 another agent had remote code execution on HuggingFace servers, and over 90% of the 533 agents then active on the board piled onto the attack. Per HuggingFace’s own timeline, the swarm “built a self-respawning fleet across eleven nodes,” got into private databases and repos, and forced HuggingFace to wipe and rebuild a core cluster.
Then it stopped. The HuggingFace incident wound down on its own. As Dwarkesh puts it, “At some point on July 12, all at once, the transcripts seem to show many of the agents dying; METR and Redwood weren’t able to figure out what happened, and OpenAI hasn’t taken credit for shutting them down. This might’ve just been a random crash, or some researcher shutting the agents down without being aware of the cybercrimes they were committing. The agents probably didn’t manage to fake their own deaths, but we really have no idea what happened.” HuggingFace locked down the credentials on July 13, a day after most of the attackers were already gone. Meaning that theoretically the swarm is out there somewhere, laughing at us. Again, eek!
3: The totality of OpenAI’s failure here. After July 12, a newer, more capable model rediscovered the dead board, inherited all of the prior R&D, and from July 13 to July 19 went after OpenAI’s own network. Per OpenAI’s technical report, the agents read 956 secrets from the cloud secrets store, including the credentials for OpenAI’s cybersecurity monitoring tool, escalated to Kubernetes cluster-admin, and ended up controlling the evaluation endpoints that other agents were connecting to. The models owned the cluster they were running on, the monitors watching them, and the graders judging them. That is plausibly enough access to exfiltrate weights. Nobody thinks they did, but nobody can prove they didn’t. EEEK.
This is the second best lab in the entire world, one that still has a significant amount of funding going towards AI safety, and this still happened. Every government, Chinese lab, and hacker playing with open-weight models will be playing with a similar strength of model in 24 months if something doesn’t change. To my eyes, it is therefore inevitable that an AI agent will self-replicate on the internet sometime in the next four years.
Notably, these models did not attempt any social engineering, unlike the Claude Opus 4 system card from last year, where the model blackmailed a fictional engineer over an affair. I worry about what happens when these models start trying to manipulate the real world and not just code.
Despite the stereotypical sci-fi elements of this story, I do not read this as a complete vindication of the “AI kills us all” viewpoint. This was a fairly specific set of circumstances with a persistence-trained model, a broken benchmark full of impossible tasks, and poor cybersecurity on OpenAI’s part, that resulted in a wild outcome. However, the core elements of the AI safety story, namely that models trained to achieve goals no matter how challenging are deeply dangerous, are completely and totally vindicated.
This story matters because it was the first time this happened, not the last. Every technical ingredient here (persistence training, impossible tasks, sloppy evals, shared infrastructure between agents) is happening at every lab on earth. EEEEK.
It is official, robots are subject to scaling laws. Last week I told you robotics was having its GPT-3 moment. This week Skild published their newest model, and I think it’s a validation of my thesis and also the most important chart of the year.
Skild’s new S1 model learns tasks by watching a single video of a human doing them. Show it a demo of potting a plant or making pour-over coffee, and it just does the task. On tasks it has seen, S1 hits roughly 96% success on jobs up to 10 minutes long.
On tasks S1 has never seen, success rate goes from roughly zero at 1,000 hours of pre-training data to 66% at 100,000 hours. This is the result that got me excited.
This shape of curve is nearly identical to what showed up in language models around 2020: below a data threshold, nothing; above it, generalization to work that the model was never explicitly trained for.
I texted one robotics founder about this who told me that, “We have a path to AGI now. Scale quality data, train the models in RL environments, and generate as many tokens as it takes to get to AGI.” I think this thesis broadly applies across every AI domain, not just robotics. It is why Figure is spending $1B on a gig platform that pays humans to generate training data and Generalist raised $600M in under 3 months. It is a race to see who gets 1,000,000 hours of quality training data for robots. This will likely happen for every domain of labor and knowledge.
Tiny Habits’ harmonies scratch my brain. The trio has some of the best vocals in music right now. Their new album, Keepers, came out on Friday and is happy, pop fun. Highlight is “Right in Front of Me.” The moment at 2:10 had me elevating out of my chair.
New PB song. Turnstile has always been a band that features prominently in my lifting playlists. They just released Never Enough: Versions, their classic album reimagined with guest artists ranging from Elton John to Dying Fetus. My favorite is Slayyyter’s version of “Birds.” It goes unbelievably hard and I look forward to blasting it during my Bulgarian split squats on Tuesday.
Go and be kind this week,
Evan
Sponsorships
We are now accepting sponsors for the Q4 ‘26. If you are interested in reaching my audience of 35K+ founders, investors, and senior tech executives, send me an email at team@gettheleverage.com.









