In a previous article, I argued that AI-native companies get their advantage by redesigning work around agents rather than adding AI to existing processes. This article is about the question that follows: what limits how fast such a company can grow.
Compute, model quality and budget are all real constraints, and for the frontier labs they are the dominant ones. But in the agentic work I run every day, they are never the constraint I hit first. The constraint I hit first is supplying enough human judgment to keep every loop aligned with its outcome.
What an outcome owner actually does
The useful unit of work in the operating model I am describing is not the task or even the agent. It is the outcome engine: a bounded, governed agentic loop aimed at a verifiable operational outcome. Agents perform the work inside defined policies, evidence requirements and control points, and a named person remains accountable for the outcome. That person is the outcome owner.
This is a broader job than human-in-the-loop. Human-in-the-loop describes a control point: a person reviews, approves or redirects a specific decision, whether for regulatory reasons, quality, safety or an irreversible action. Those control points matter and belong in the design. Outcome ownership sits above them. The owner stays accountable for whether the outcome is still defined correctly, whether the policies and tolerances still fit, whether the evidence shows the loop is actually working, how novel exceptions get resolved, and when the engine needs to be changed, paused or retired. Approval is a checkpoint. Ownership is continuous.
Continuous accountability does not mean continuous watching. In a well-built engine, the state lives in telemetry, decision records and evaluations, and the owner’s attention is pulled in by exceptions. If an owner has to mentally reconstruct an engine every morning, that is a design failure in the engine, and it was one of mine for a while. But even with good instrumentation, someone still has to be the person the exceptions escalate to, and someone still has to notice when the world changed in a way no policy anticipated.
The span-of-control question
Greg Meyer recently asked the right question about this: how many agents can one human manage before the human becomes the bottleneck? He borrows the factory concept of span of control, and his conclusion is that the binding constraint was never agent intelligence. It is “how much human judgment you can supply per unit of time.” He is also clear that the span is variable. With deterministic work and agents that escalate well, one person can hold many more loops than with ambiguous work and agents that interrupt badly.
My own experience is consistent with both halves. For the engines I run today, which are complex and still maturing, two or three is the realistic number if I want each outcome genuinely owned rather than glanced at. Four or five is my ceiling, and at five I am dropping things. I offer that as an operating observation from one practitioner, on one class of engine, at one point in maturity. A stable, well-instrumented, low-stakes engine costs far less judgment than a volatile, ambiguous, newly deployed one. One person might own a single engine of the second kind or ten of the first.
Which points to the real metric. The constraint is not a fixed number of engines per person. It is judgment load: the amount of human attention an engine consumes through decisions that cannot safely be settled by policy, evidence or automated evaluation. It depends on how often those decisions arise, how ambiguous and consequential they are, and how much context they require. Engines differ enormously on that measure, and a well-designed engine’s judgment load should fall as it matures. Maturity alone does not produce that. Better policies, better evaluation and better instrumentation do.
Others are seeing the same bottleneck
Ethan Mollick at Wharton has been approaching this from the organizational side. He notes that our organizations were built around human intelligence because that was the only kind available. He has also argued that excessive AI-generated output can overwhelm approvals, staffing models and coordination systems. And he observes that agentic systems increasingly resemble mini-organizations and that directing them looks like management. Management of agentic work is becoming the scarce capability.
MindStudio illustrates one version of the failure mode with what it calls the piling problem: a team reviews twenty or thirty outputs during a pilot, approves the agent, then discovers the production system can generate hundreds a day. The agent works, but the queue grows faster than people can absorb it.
The numbers being reported point the same way, though they measure a different unit. OpenAI CFO Sarah Friar has relayed an account from a leader at a large consulting firm describing part of its back-end operation at roughly one human to five agents. That is an anecdote about agents, and an engine may contain several agents plus policies, evaluators and control points, so it does not confirm any engine ratio. What it shows is that at least some organizations are beginning to count humans per unit of supervised AI work. Deloitte’s 2026 State of AI in the Enterprise survey of 3,235 business and IT leaders fits the same picture: 74% expect their organizations to be using AI agents by 2027, while only 21% report a mature governance model for them. Deployment ambition is outrunning governance readiness, and the human capacity to supervise, redirect and remain accou
ntable for these systems is part of that gap.
These sources do not establish a universal human-to-engine ratio. They do show the same underlying risk: AI output can scale faster than the management, governance and review capacity around it. My conclusion from running these systems is that accountable human judgment becomes a governing constraint on scale.
What this does to the growth model
Sam Altman and Dario Amodei have both floated the one-person billion-dollar company. The judgment-load view does not disprove it. It states the condition for it: a single founder would need to design the company so that agents and deterministic controls absorb almost all execution and coordination, leaving a residual judgment load that fits inside one person’s capacity. That is plausible for some businesses and structurally implausible for others, and knowing which kind you are building is now a strategy question.
For everyone else, workforce planning changes shape. For work that agents can absorb, task volume becomes a much weaker basis for sizing teams. The planning question shifts to the outcomes being op
erated and the judgment load they create. Do not divide engines by a fixed number. Weight each engine by its ambiguity, consequence, volatility, exception rate and maturity, and size outcome ownership to the total. Expect two opposing forces over time: judgment load per engine should fall as engines are engineered to maturity, while better agents make more engines economically viable. If you add engines faster than you reduce the judgment load of each one, your demand for capable owners rises even as your agents improve. That has been exactly my experience.
Compute and model access can usually be expanded faster than outcome-level judgment can be developed. But the combination of domain context, agent fluency and authority required to own an outcome cannot be developed or transferred at short notice. It is one of the slowest capabilities an AI-native company can grow, and it sets the pace for everything else. Start now
.



