Over the two months, it feels like every company within a 20 mile radius of Stanford has launched some type of “personal agent” product. Grok Bots, Muse, OpenAI’s Dots. Reflexively, whenever I see that much consensus in Silicon Valley, I want someone to argue the opposite.
That person is my friend Nathan Baschez. Nathan is currently a product designer at Notion, and in a previous life was a founder at Lex and Every. Before that he spent some time at Gimlet Media and Substack. Fun career fact about him is that he designed and built the original version of Product Hunt. He is a great writer and was one of the earliest believers in my potential as an blogger. In this post, he argues that personal agents will matter, but the future of work is more multiplayer, more automated, and more complex than any individual agent can handle.
First though, this newsletter is brought to you by Lovable.
Last week I built an app for my daughter’s care team without opening Lovable’s editor, and I want to explain how, because it’s one of the most magical experiences I’ve had with AI tools this year.
Lovable, a software creation platform, has an MCP server, which means an AI agent can operate Lovable the way a person would. So I talked to Claude about my problems. Claude created the project, sent my spec in plan mode, showed me the plan, and once I said “approve,” had Lovable build it. In between, Claude read the database access rules, ran a query to check that my sign-in had actually made me a parent, and, when the first test email said my daughter had 0 words, found the 1-second bug in the date logic and had Lovable fix it. I opened Lovable twice: to sign in and to buy a domain. By the end I had a functioning app, perfect for my life circumstances.
Welcome to the wild future. If your AI assistant has the context on your problems, it can use Lovable to solve them. Subscribers to The Leverage get $5 off their first month of Lovable Pro.
And with that, here’s Nathan.
Technological revolutions don’t arrive fully formed. They evolve over years, decades even. Early on, crude architectures generate excitement and reveal limitations. Then they eventually make way for newer and better layers of the stack. Take the internet, for example. Decades passed between the first packets sent over TCP and the birth of the web and HTTP; another decade separated the web from the dominance of interactive web applications.
We should expect the LLM revolution to look no different. It would be really weird if, just four years after the ChatGPT moment, we had already figured everything out. We should be open to new paradigms and architectures. In fact, we should be obsessively searching for them, because every paradigm shift creates winners and losers — even the mini-shifts within a mega-shift like AI.
If I had to put my money on the most underrated nascent paradigm shift, it would be factories. Factories are a logical next step after skills and personal agents.
This essay is meant to help you win the next paradigm. In it, I’ll define “factory” in a slightly more broad way than most people use it, make the economic argument for factories’ inevitable dominance, and show you how to build one. If you want to get ahead of the next paradigm shift within the AI revolution, or just use AI more effectively within your business, this is for you.
What is a factory?
Factory (noun) — a shared AI system designed to handle specific types of tasks within an organization.
Usage: “I am rolling out an update to the slide deck factory so it adheres better to our new positioning, along with a few design fixes for issues we observed last week. Let me know what y’all think!”
Imagine two consulting firms that directly compete with each other. Let’s call them BigBrain, Inc. and DumboCo.
BigBrain and DumboCo each have about 100 employees. They do similar revenue and are equally excited about using AI to empower their teams. But they have two very different philosophies for how to use AI: BigBrain is factory-pilled, whereas at DumboCo everyone uses their own personal agent to do everything.
At DumboCo, most people use Codex or Claude Cowork. They have a Slack channel where people share skills and “AI wins,” but everyone is free to pick and choose what they want to adopt. Some people invest far more effort than others in their setups and see corresponding gains in their agents’ performance and productivity. Sometimes they will help other members of the team or do a “show and tell” session, but it’s still hard to get everyone to their level of productivity.
From my conversations with founders and folks who work in tech, I would estimate that 99% of companies today are operating this way, with varying degrees of intensity. But there is a better way.
BigBrain, on the other hand, is a bit more centralized in its approach to AI. Everyone still uses their own personal agent more or less constantly, but there’s also a layer of shared factories for anything complex and frequent. For example: slide deck creation, employee and customer onboarding, data questions, software bug fixes, app designs, and prototypes, etc. They even have a “factory factory” that makes it easy to spin up new factories as needed.
The biggest AI enthusiasts on each team tend to be the ones who create and manage the factories that serve that team. For example, the CFO happens to be a big AI nerd and created a “monthly close” factory for her whole team to use. Meanwhile, people who just need to use the factories don’t have to worry about the details of how they are created and maintained in order to benefit from them.
There are three key attributes that define a factory and differentiate it from a personal agent.
Factories are:
Shared — multiple people on the team can use them.
Specific — they are optimized to handle certain types of recurring tasks, unlike a general-purpose agent.
Scrutinized — some people on the team can see the work flowing through the factory to observe and fix failures.
Most of BigBrain’s factories are fairly simple: each is a Notion database of tasks, with agents hooked up to act on those tasks as they are created, moving them through a sequence of stages. It can be as simple as todo → doing → done.
Some factories are purely autonomous; others involve deep collaboration with humans at strategic points. For example, the customer support factory handles emails as they come in, fully autonomously, with human intervention in rare cases. The slide deck factory, by contrast, involves collaboration at many steps (project brief, research, writing, and design). If you ask most people what an “AI factory” is on X, they’d probably emphasize the lack of humans in the loop, but for me this is the least interesting part about the systems commonly called “software factories.” The more important part is that it is a purpose-built, shared system that can be optimized by some people for the benefit of everyone else.
You can buy factories from vendors, or you can build them. The most mature vendor market is for customer support agents, as sold by companies like Sierra, Fin, and Decagon. These all meet my definition of a factory even if they don’t say they are selling factories. Cognition, Cursor, and others are explicitly selling software factories. Amplitude is pivoting from product analytics to selling a “product factory” — a system to autonomously identify, build, and ship potential product improvements by looking at metrics. XBOW, Novee, Snyk, and others are building cybersecurity factories.
“What about skills?” you might ask. Skills can be sort of factory-like, in that they help agents perform specific tasks, and the prompts are shared. But the actual runtime environment is not shared, so people running a skill with different agents on different machines may run into configuration issues or other hiccups. And, crucially, skill runs are not easily scrutinized by skill owners. So the feedback loop is broken.
The economic logic that makes factories inevitable
It is as simple as Adam Smith: factories are key to specialization and gains from trade — the root of all productivity and prosperity.
DumboCo in the example above is sort of like a hunter-gatherer tribe. Everyone is responsible for reinventing their own wheel (and hand-axe, shelter, clothing, first-aid—you get the idea) . Over the course of a month, let’s say each of the 100 employees spends anywhere from 5 to 50 hours tweaking their agent’s behavior across 5–10 different types of tasks. That is an average of about 3.7 hours per task that month spent moving up the learning curve. (And, let’s be honest, not all hours are created equal. Some people are much better than others at optimizing agents for different kinds of tasks.)
BigBrain, on the other hand, is like a modern society with different specialized organizations that build on and depend on each other. One person or team can set up a factory that everyone else can benefit from, without having to learn anything about how the sausage is made. With a good “factory router,” employees’ personal agents can even use the factories on their behalf, getting the job done without the employees having to know the factories exist! For example, the HR team could have a factory to answer questions about benefits, and to the user it just feels like they asked their AI a question and quickly got a good answer.
BigBrain has the same budget of 5 to 50 hours of “AI tinkering time” per person but concentrates it on just 1–2 tasks per person. This yields 18.3 hours of optimization time per task per month, which means going 5x deeper into the learning curve on each task. This is the magic of specialization and gains from trade! And this 5x improvement is before taking into account the difference in quality of hours between a specialist and everyone else.
But the “hours into the learning curve” idea is pretty abstract. What benefits do teams actually see when they spend more time optimizing each task?
Higher quality
More reliable
Cheaper
Faster
More secure
In general, these things make the difference between a half-baked skill someone threw together once (unused, out of date) and an active workflow that a business can rely on.
How to build a factory: a pattern language
There are many ways to build a factory, and, as noted above, you can also buy one. But here is the simplest starting point for most people building their own. Let’s use the example of a slide deck factory.
Set up a database inside Notion called “Slide Decks.” Add a status column. To start, you can keep it as simple as “todo,” “doing,” and “done.”
Hook up an agent with access to the database via the Notion MCP, CLI, etc. Or you could use a Notion Custom Agent.
Set up the agent to be triggered when a new card is created or when the status changes. You can use any agent or hosting platform you want for this, as long as it’s not one that goes to sleep when your laptop is shut.
Write some instructions for the agent to follow. When the agent is triggered to handle a new task, have it set the status to “doing” and, when it finishes, to “done.” Give the agent a template it can start from and instructions for how you like your decks created.
You may want to host the instructions in Notion for ease of reading and collaboration, plus versioning and comments. Or you could store the prompt in GitHub or anywhere else.
Observe runs. Fix anything that fails. Repeat! (Agents can help you do this!)
Share your factory with your team. If it’s a Notion database, it’s as simple as giving them access to that page. They (or their agents) can put tasks in and watch them get completed.
For discoverability, you can put a “factory list” in your team’s system prompt. You could also make a Notion page with a list of factories or a skill for each factory; there are many ways to do it.
This is the minimum viable factory setup. In most cases, you should start here and add complexity only as you actually need it, not ahead of time. But there are a few things most serious factory attempts will need sooner rather than later. Here are the most essential patterns I have observed.
First, at some point you will want to run your factory on an experimental new version of the prompt without messing up the existing flow of traffic. For this, you will need a basic prompt versioning system. You could solve it the software engineering way, with separate development and production environments and prompts deployed as code through a version control system like Git, but that adds overhead.
The easier way is to use something like a Notion database. You can have a “prompts” database for the factory, with a “status” column that can be set to testing, production, or retired. And if you have different types of prompts for different stages in the factory process (e.g., research, writing, and design), you can make another column for that. Then the agents can read their instructions from a central source of truth that is easy for everyone on your team to inspect and understand.
Second, once you have different versions of a prompt, at some point you will want to know whether an experimental new version is actually an improvement. You can test this yourself using your own creativity and judgment, but that takes a lot of time and work, and quality will vary depending on who is doing the testing and what they remember to test for. What you need is basically a checklist of the different tasks you want the factory to be able to handle, plus a rubric for judging whether each task was completed successfully. These are commonly known in the industry as “evals,” and they are much easier to create and use than you may think.
Bottom line
Factories — or whatever they end up being called — are AI systems that are shared, specific, and scrutinized. They are inevitable because they are economically rational. They enable specialization and gains from trade.
Do not be surprised if you start to hear much more about them over the coming months.





