Output to Outcome: Why the Hours Metric Finally Broke
Hours logged was never a good proxy for value created. Knowledge work does not scale linearly with time in a seat, and it never did. What kept the proxy alive for as long as it survived was a quieter fact underneath it: producing a unit of output used to take roughly the time it took. A feature took a week because it took a week. That relationship held loosely enough, for long enough, that a lot of organizations never had to confront how weak the underlying assumption actually was. That relationship is now broken, and the break is not subtle.
Mik Kersten’s new book, “Output to Outcome”, is the clearest account I have read of why the break happened and what it costs an organization that doesn’t respond to it. Kersten, the researcher behind the Flow Framework and the author of Project to Product, argues that AI has shifted the constraint in knowledge work from producing output to delivering outcomes, and that most companies are still managing, measuring, and rewarding the layer that no longer matters. I wrote about the visibility problem with time tracking back in April, when I argued that hours are activity data dressed up as impact data. Kersten’s book gives that argument a sharper edge and a name.
The Cost of Output Just Collapsed
Kersten names the mechanism plainly. AI has driven the cost of producing output toward zero for a wide range of knowledge work. An engineer who used to need a week to draft an integration, write the tests, and stand up a first version can now produce something comparable in an afternoon, sometimes an hour, with an AI system doing the first-draft labor. That is not a marginal productivity gain. It is a change in the unit economics of output itself.
Once that is true, an hours-based system stops being merely imperfect and becomes something closer to actively hostile to the behavior you want. A utilization dashboard built on logged hours will read a compressed delivery time as underperformance. The engineer who used AI well and finished in two hours what used to take two days now shows up as underutilized, low-effort, or worse, idle. The system is punishing the exact efficiency gain it exists to capture. That is not a tuning problem you fix with a better dashboard. It is a sign the metric was measuring the wrong layer of the organization all along, and AI simply removed the cover that let everyone avoid noticing.
Hours Were Always Measuring the Wrong Layer
Hours measure activity inside a team. They do not measure value moving through the system. Those are different layers, and leaders who conflate them end up managing the wrong one, confidently.
What actually determines whether an organization converts effort into results is flow, and flow breaks down into two numbers that matter far more than anything a timesheet captures. The first is queue time: how long does a piece of work sit waiting before anyone actually touches it. The second is time to deliver: once work starts, how long does it take to become a landed outcome the business can use. Together those two numbers tell you whether your organization is a system that moves work through to completion or a system that accumulates work in the middle of the pipeline while everyone stays busy at the edges.
Hours logged tells you none of this. It never did. Nobody noticed because output used to take roughly as long as it took, so hours and delivery time moved together closely enough to fake a correlation. AI severed that correlation. Now you can have an organization logging fewer hours and delivering dramatically more, or an organization logging the same hours it always has while its actual queue times balloon because nobody restructured how work moves once the production step got fast. The timesheet cannot distinguish between those two organizations. Flow metrics can, and that difference is the whole argument.
Why Most Organizations Aren’t Seeing the Gains
Kersten’s data points to something leaders should find uncomfortable: most companies deploying AI are not seeing the productivity gains they expected, and the reason is not that the AI underperforms. The reason is that the organization is still measuring, managing, and rewarding the old layer. Output, hours, utilization. Those metrics were built for a world where the constraint sat inside individual task execution. AI moved the constraint. It now sits in how fast work moves between people, teams, and decision points, and most organizations have not rebuilt their measurement systems to see that constraint, let alone manage it.
This is the same pattern I have written about with process debt and organizational decoupling. You cannot layer a faster engine onto a structure built around the old bottleneck and expect the structure to get out of the way on its own. If work still has to sit in a queue waiting for a weekly steering meeting, or wait on a manual handoff between teams that were never asked to coordinate more tightly, the AI-generated draft sits in that queue exactly as long as the human-generated draft used to. The production step got faster. The system around it did not. Kersten’s finding, that most companies aren’t capturing the gains, is what you would expect when the bottleneck moves and the org chart doesn’t.
Queue Time Is a Structural Problem, Not a Willpower Problem
It is worth being precise about where queue time actually comes from, because the instinct in most organizations is to treat it as a people problem. Someone is too slow to pick up the ticket. Someone is sitting on the approval. Someone didn’t prioritize it. That instinct is usually wrong, and it leads leaders to apply pressure to individuals for a delay that the structure created.
Queue time accumulates at handoffs. Work waits when it moves from one team to another and has to clear a different set of priorities, a different manager’s attention, or a different review cycle before it can proceed. It waits when a decision has to travel up a chain of approval that was designed for a slower era of production, back when the thing being approved took long enough to build that a week of review didn’t change the overall timeline much. It waits when a single centralized function, a security review board, a data governance council, an architecture committee, becomes the mandatory gate for every team’s output regardless of how fast that output can now be produced. None of those delays show up as low individual effort. They show up as structure, and structure is exactly what most utilization dashboards are blind to.
This is why AI adoption without organizational redesign tends to produce a strange result: individual throughput goes up and overall delivery time barely moves. The engineer finishes faster. The draft then sits exactly where it always sat, waiting for the same review board that used to receive work once a week and now receives it five times as often without any corresponding increase in its own capacity to process it. You have made one stage of a five-stage pipeline dramatically faster and left the other four stages untouched. The system’s overall speed is still set by its slowest stage, and AI did nothing to that stage except increase the pressure on it.
Measuring Outcomes Without Losing the Thread
None of this means abandoning measurement in favor of vibes. It means measuring further downstream than most organizations currently do. Queue time and delivery time are a starting point, not the whole system, and they need to be paired with a measure of whether the delivered work actually held up: how often it required rework, how often it failed once it reached customers or internal users, how often the outcome had to be revisited because the first version missed the mark. Speed without a quality check is just a faster way to produce the wrong thing, and any executive who has lived through a rushed release knows how expensive that can be.
The instrumentation for this already exists in most engineering organizations, even if nobody is looking at it the right way. Ticketing systems capture timestamps for when work enters a queue and when it moves between states. Deployment systems capture how often something ships and how often it gets rolled back or patched. The data to build a genuine flow picture is usually sitting in tools the organization already pays for. What’s missing is the decision to report on it instead of reporting on hours, and that decision has to come from leadership, because nobody below the executive level has the authority to retire a utilization dashboard that a board has come to expect.
What to Actually Change in the Organization
Kersten’s book is not just a diagnosis. It lays out an operating model for making this shift real, and the pieces are worth translating into what a leadership team should do this quarter rather than treat as a future roadmap item. The core move is what he calls the Outcome Loop: a single value stream’s feedback path from strategy and objectives, through roadmap and delivery, to a measurable result that gets checked against the original intent. Every product line, every major internal system, should have one of these loops defined explicitly, with a named owner accountable for the outcome, not just the output.
The second piece is the Outcome Tree, which is how that single loop scales across a whole organization instead of staying a good idea in one team. Rather than a functional hierarchy where work climbs through layers of management review, the Outcome Tree structures accountability around outcomes at every node, with humans still making the calls but AI amplifying what each node can produce. This is the direct answer to the review-board bottleneck I described above. Instead of routing every team’s output through one centralized gate, the organization is restructured into modular value streams that each own their outcomes end to end, with governance built into the loop rather than bolted on as a separate queue.
The shift underneath both of these is what Kersten calls Functions to Flow: moving people out of siloed functional groups, where a request has to travel between departments to get anything done, and into cross-functional teams organized around a value stream that can carry work from intake to delivered outcome without leaving the team. This is not a reorg for its own sake. It is the structural change that lets faster production actually turn into faster delivery, instead of piling up in front of the next functional gate.
None of this requires waiting for a perfect plan. Pick one value stream, likely the one under the most competitive pressure, and rebuild its Outcome Loop first: name the owner, define the outcome metric, and remove the handoffs that don’t need to exist. Prove the model there before asking the rest of the organization to restructure around it.
What This Means for the Board Conversation
For a CEO, COO, or CIO, the practical shift is a change in the question you ask. Stop asking how much time a piece of work took. Start asking two different questions: how long did this sit in queue before anyone started it, and how long did it take from the moment someone started to the moment it landed as a usable outcome. Those two questions replace an entire category of management theater, the kind built on utilization reports and time allocation reviews, with something that actually predicts whether the business is converting its AI investment into results.
This also reframes what an AI productivity report should look like when it reaches the board. A report that shows individual engineers completing tasks faster is not evidence of organizational gain. It is evidence of task-level gain, which is real but incomplete, and it can coexist with a system that is no faster overall because the compressed task time simply produces a bigger pile of finished work sitting in the same queue it always sat in. The gain you actually care about, the one that shows up in revenue, in time to market, in cost, only appears when queue time and delivery time both compress. If your AI reporting doesn’t track those two numbers, you don’t know whether you’re getting outcomes or just faster output that’s piling up somewhere you’re not looking.
A weak proxy that quietly misled leaders in a world of linear output becomes an actively dangerous one once output can be manufactured at near zero cost and the real constraint has moved somewhere the proxy was never built to see. The organizations that recognize this first will stop rewarding activity and start managing flow. The ones that don’t will keep asking how many hours were logged this quarter, and they will keep getting an answer to a question that stopped mattering the moment AI changed what output actually costs to produce.