If You’re Not Paying For The Tokens, You Are The Tokens.
The Weekend Leverage, August 9th.
This is the most concerned about AI safety I have ever been. In case you missed it, a few weeks ago OpenAI announced that one of the new models they were training had broken containment and hacked the startup HuggingFace. This is bad! But also, not the worst thing in the world. However, this week we got more details that show a staggering level of failures that happened far before the HuggingFace hack. Quoting the excellent Zvi Mowshowitz’s summary of events:
OpenAI models-in-training, without the excuse of ‘they were doing a cyber eval,’ created a message board where they shared information on how to hack and cheat, and were trained on that basis.
OpenAI only figured this out when the models crashed the server.
OpenAI’s response was to rebuild the server and patch that particular exploit, but they continued training the models that trained using the message board.
Those models then recreated the message board, hacked OpenAI again, got internet access, and used an agent swarm to attack HuggingFace in order to get the answers to a cyber evaluation.
It is worth reading his whole post on the topic if you want to go deeper. The summary is that these models outsmarted the people training them, then a failure of corporate governance—exacerbated by the pressure of “grow fast”—made them just continue to use the misaligned models. Bad, bad, bad.
I still don’t fully buy the AI autonomously kills everyone scenario that much of the internet concerns itself with. The much more likely disaster is bad actors using openweight models to synthesize bioweapons (if that sounds like sci-fi, the Arc Institute synthesized new viruses using AI just this week). Alternatively, I think there will be major cybersecurity failures such as we have seen here but instead of HuggingFace it’ll be utility companies. (Is now a bad time to mention that Iran has hacked 12 states’ water systems?). This is not the place to discuss solutions, but these types of conversations are why this publication exists. It is time for you to take AI seriously. For your career, and for your future. We will see examples of that this week with robotics, turning ourselves into training data, and Stripe’s bold new acquisition.
But first, this edition is brought to you by Span.
Your engineering team is “all in” on AI. Congratulations! You just doubled your cash burn rate. But has all that spend actually made your product any better? Is the AI code you are generating useful?
Span is measuring the gap between AI adoption and AI effectiveness. To figure it out, they’re running a four-minute survey to build a new industry benchmark.
Fill out the survey and they’ll send you a killer report on where your team’s AI effectiveness lands relative to everyone else’s. Vibes are a bad way to gauge your own progress!
Fill it out via the link below to get the report.
Airtable was worth $11.7 billion but sold for a tenth of that. The one-time SaaS it-girl agreed to be acquired by Bending Spoons for $1.285 billion, an 80%-plus haircut from its 2021 peak. And the metrics aren’t even that bad! It had roughly $480M in ARR, growing more than 20%. So what the hell happened? And will this level of value destruction happen to every SaaS company? Read here.
I’ve tried every AI personal assistant and churned off all of them (until now). Town is the chosen one, the mega-mojo, life-changing AI big ripper that has meaningfully transformed my life as a founder. I interviewed the team on why they spend about $100 per user building a personal wiki about you and how they made so many small moments of delight. Town even accomplished my white whale of productivity, a self-maintaining to-do list. Here’s what actually happened when I used it, and where I think it could still break. (I also got y’all a discount on it if you read to the end.) Read here.
If you’re not paying for the tokens, you are the tokens. Meta released a new coding model called Muse Spark 1.2 with a coding agent called Muse Code. It is roughly one rung below the frontier model for a pure power perspective, but they are doing a remarkable innovation in pricing.
There are two tiers. Standard costs $1.25 per million input tokens and $4.25 out, ordinary near-frontier pricing. The second tier costs $0.10 in and $0.20 out, if you grant Meta permission to train future models on your prompts and completions.
The company that perfected “the product is free, you are the ad inventory” for attention markets is now trying to run an identical play for intelligence markets. Interestingly, you could also view this move as an RFP to data vendors. If you invert the math, a million output tokens of coding-agent telemetry is worth at least $4.05 to Meta. This is all a long way of saying that Mark Zuckerberg is a ruthless and effective founder who will use his cash flow to force his way into a market. Here he comes for coding.
Stripe hopes inference is as boring and complicated as payments are. The payments giant is in exclusive talks to buy OpenRouter for roughly $10 billion in cash and stock. For my readers with a memory for deals, you’ll note this is about 7.7x the $1.3 billion valuation set ten weeks earlier. The Information reported that OpenRouter is at ~$140 million of annualized revenue, which makes this a 70x revenue deal for a company that takes 5% of the inference revenue flowing through it. That is…pretty expensive. To justify this you have to believe two things simultaneously:
First, just about every piece of hardware and software would benefit from inference, and we sit on the cusp of a massive explosion in tokens. While large companies will figure out what models work best for their needs, everyone else will need help. Here, OpenRouter can help by “routing” their intelligence needs. In practice, that just means helping applications use cheaper models when the output is routine. But whether an output is routine is so application-specific that only the application can really answer it. Think of it this way: what is good code? Code that works? Code that makes money? Code that is simple?
All these dimensions of quality are things generic routers can do an ok job at, but are typically better determined by the application. It is why every venture-scale app is building routing in-house: Cursor trained its router [disclosure, current sponsor] on requests from its own editor and claimed 60% savings while Ramp ran its router across 2.75 trillion tokens a month of internal traffic before giving it away free. Today, routing quality is downstream of harness capability.
Second, for everyone below venture scale, an autorouter is probably fine, the same way Stripe Checkout is fine for everyone who is not Amazon. Stripe’s bet is that inference becomes a metered utility where you can add margin by layering in payment flows. Metronome, which Stripe bought for ~$1 billion in January, meters the tokens in. OpenRouter pays the tokens out. Stripe sits on the money in between. So they can theoretically double or even triple dip on inference revenue. That’s a pretty compelling vision that is worth spending 70x revenue for.
The humanoid robot IPOs are finally happening. Unitree is planning to sell 10% of itself at 150.80 yuan a share, a $904 million raise at a $9 billion valuation, listing on Shanghai’s STAR Market around August 19 after demand pushed the raise 45% above plan. The prospectus gives us the first audited look inside a humanoid business, and it describes something no American competitor can claim—a company that actually works! Unitree’s revenue went from $24 million in 2023 to $253 million in 2025, at a 60% gross margin, with actual net profit, selling robots at an average price of $24,700.
Over the last twelve months, I’ve banged this drum over and over, but let me do it again. We are at the GPT2 stage of robotics. The demos are compelling enough to show that scaling laws will most likely work. Meaning there will be humanoid robots, in some form or another, coming over the next five years.
If you hold that belief, the natural question becomes how can you actually make the robots to scale? So far, the answer has been “go to China and ask Unitree to do it.”
The humanoid robot leader is AgiBot, which rolled its 15,000th robot off the line June 28 (that count includes its wheeled and quadruped lines; pure humanoids were 5,168 last year). Unitree shipped 5,511 humanoids in 2025 and is building capacity for 75,000 a year. On the American side, Figure manufactured its 1,000th Figure 03 on July 23. The same day, Tesla told investors the Optimus line was still being installed and the ramp would be “quite flat and long.” Agility’s SPAC filings disclose a $300 million order book and zero unit counts. 1X has zero verified customer deliveries though they are planning to ramp up manufacturing over the next twelve months. If you are generous, you could say that every American company in history has built roughly 3,000 humanoid robots. AgiBot built more than that since Memorial Day!
This is obviously bad if you believe my GPT2 thesis. It is also why the FCC banned imports of new foreign humanoids, naming Unitree directly. The policy recycles the playbook we used on drones, forcing a security review or automatic blacklist. It didn’t work for drones and this policy won’t work for robots either. America is roughly 13.3% of Unitree’s revenue, and yet the IPO was formalized two days after the ban, and still priced 45% above plan. This policy ends up cutting American researchers off from robots while the production gap sits untouched. You cannot sanction your way to a healthy industrial sector! You have to incentivize the entire supply chain.
My current front-runner for best millennial writer is Hanif Abdurraqib. I’ve shared his essays before (his one on Fall Out Boy will rip your heart out). Lately, I’ve been working my way through his most recent book There’s Always This Year. His words have this aching poetry, a mournful rhythm, that propels you from exquisite phrase to exquisite phrase. Reading him is a reminder of why this medium is one of humanity’s crowning achievements. Go buy the book and bask in the beauty!
These two very sweaty British rockers got my ass shaking and toe tapping.
Last thing. Does anyone have a recommendation for an absolutely breathtaking breakfast burrito in SF? I’ll be in town for The Leverage Launch on Wed and Thursday and am craving some proper Mexican food. Let me know!
Go and be kind this week,
Evan
Sponsorships
We are now accepting sponsors for the Q4 ‘26. If you are interested in reaching my audience of 35K+ founders, investors, and senior tech executives, send me an email at team@gettheleverage.com.











