Author’s note: This is the latest in my “What’s the Bet” series. AI has the potential to change everything, and founders have to make company-defining choices in response. That is not easy. In this series, companies pay me to analyze their bet and explain it to you. I retain complete editorial independence. Today’s bet is brought to you by Span, which tracks engineering effort and AI spending and connects them to software outcomes. If you have any feedback, just reply to this email.
Span’s bet is that companies will pay an independent evaluator to tell them how to deploy their coding agents—and their engineers—more effectively.
In ye olden times, before datacenters carpeted the US interior like McDonald’s, and before any of us were asked “what’s your p(doom)” at a kid’s birthday party, an engineer wrote code by hand. They would bundle that code into a pull request, a proposed change that other engineers would supposedly check. Eventually, the feature would go live. To track it all, managers used tickets, delivery metrics, planning meetings, and, in my favorite example at a startup I worked at, a whiteboard labeled “ship this or the company dies.”
Contrast that to today. As I type this column, I have 4 different AI agents building products. I have no clue what the code looks like. I have no idea how long it will be until they finish or how much it’ll cost me. This is (somehow) progress. My software ambitions have expanded considerably faster than my ability to manage them. The operating philosophy is YEET.
The issue is that a pull request alone doesn’t explain what coding agents did to produce it. For example, I’m hacking on a new consumer app right now. My coding agent is reading my PRD, running tests on the code it writes, trying things that fail, trying again, and then, finally, producing a change. There is still a pull request at the end, just like the before times. However, that final merge doesn’t tell me about all the dead ends and tokens consumed along the way.
In this new world of coding, CFOs, founders, and engineering leaders have to ask themselves a crucial question: “Is the money I’m spending on Claude Code, on top of what I already pay the engineers, buying me anything?”
Span wants to answer that question across a product organization’s entire suite of AI tools. They have a two part bet here. First, understanding agent sessions will lead to better engineering decisions. And second, a third party can provide a more useful analysis than the companies actually selling the coding tools.
The bill is growing. What did it buy?
Anthropic reports average enterprise consumption of $150 to $250 per developer per month. At that usage level, 1,000 developers would consume $1.8M to $3M annually. My experience with the very best AI engineers is that the number is close to $10K a month per developer, so I think at scale, a fully pilled AI product organization is likely increasing their operating costs by up to 80%. That amount of cash being burned generates heat sufficient to singe your board’s eyebrows off.
At that level of spending, engineering leaders need to know where the next dollar should go. A better model? Another engineer? Fixing the environment their agents keep getting stuck in? There is always more AI to buy but figuring out which of those investments would help is the management problem.
More code can mean more work
The first problem with measuring AI productivity is that your productivity metric can go up while your engineering organization gets worse.
Span’s analysis of 248,099 pull requests showed that AI written code was longer, submitted bigger PRs, and broke far more frequently.
The implication is that faster code generation moves work into review and repair. Counting shipped code misses both the effort required to produce it and the work it creates for everyone else. So, a “is AI making our code more expensive and also worse” tool has to go beyond mere code examination. It needs to be something larger.
What the session tells you
One of my biggest frustrations with using coding agents is when something breaks. It is unclear if the model simply can’t do what I want, or if I’m just shit at prompting it. Maybe there is something legacy in the code base? Most of the time I am stuck asking “Why broke? Plz fix.” Which, like sometimes works, but it all adds up quickly. At the speed and sprawl of a growing startup, founders needs to know where to invest; on better models or better instructions or something else entirely. Essentially, better information means better product investments.
Span built its Context Layer by combining code, tickets, and other work signals to show where engineering spending goes. The company has an “AI Effectiveness suite” that extends that context layer into the agent session itself. By combining AI spending with estimated human effort, they can get the actual cost. You have probably experienced this. Ever use a worse model trying to save money but then end up needing to regenerate the result with the frontier because the results suck? The same thing is happening company wide.
Span offered up a study of 103 engineering teams as proof that their product is working to me. They have found three things to be associated with quality:
Prompt clarity: Each one-point increase was associated with 27.2% lower token cost per merged AI line.
Environment readiness: Across scores of 2.5 to 4.5, merged AI lines per human turn rose from roughly 5 to 17.
Quality stewardship: Each one-point increase was associated with 39% fewer review cycles per 1,000 merged AI lines.
Now, cheap code can still be useless. But the need for these tools, and the insights they have already generated, suggest that some expensive AI problems are fixable management problems.
In fact, when they first pitched me on sponsoring The Leverage I was a little hesitant. This problem seemed so obvious that I thought for sure this would be a feature, not a startup. Maybe you think so too. Here’s why that instinct could be wrong.
Why wouldn’t someone else own this?
There are 2 groups who could. The first is the dashboards engineering leaders already pay for: stuff like DX, Jellyfish, Swarmia, Faros, Weave, and LinearB. These all plug into GitHub and Jira, count what shipped, etc etc. Plus, nearly all of them have bolted on an “AI impact” module in the last 18 months that pulls usage numbers from stuff like Claude Code.
And to be fair, until this April, Span was at least somewhat similar. The AI-detection model it launched last September still read finished code, the same as everyone else. The difference is where Span chose to do their measurement. The incumbents measure AI at the point everything is “done.” Span records the session itself, through a program on every engineer’s computer. To my eyes, measuring everything at the point of pull request is an idea made for a pre-AI world. You have to be able to measure AI agents as effectively as employees, because increasingly, they’ll be more productive then them anyways.
The second group is the coding tools themselves. After all, they already have ways to measure coding progress! This is like saying “my credit card lists all of my expenses, why do I need a finance department?” The context layer that Span has built allows it to elevate the level of analysis and complexity beyond one tool or one agent paradigm. Right now, coding agents are akin to religion, every practitioner has their preference, and to an outsider they all kinda sound the same.
Meaning that the finance department doesn’t care which church you attend. It needs one central ledger by which to measure AI ROI. Existing measurement tools measure too late in the process while coding agents are too siloed. Worse, none of the competition knows what the engineers cost, which is the bigger of the 2 bills anywho. (Well, it is bigger for now at least.) This is why this company has earned the right to exist and win the market. To account for the costs, you actually have to account for all of them together.
In my conversations with engineering managers, they tell me that “what did this AI session cost and what did it ship” is becoming as normal a question to ask as “how many hires do you need.” Something will have to answer that question. Span’s bet is that the answer comes whatever keeps the record across every tool a company uses.





