Theory of Constraints in IT: why shipping faster still made the system slower
Marcus Dalton
September 18, 2026
We sped up the busiest team and the company got slower. I have done this more than once. We hired another frontend. We cut the PR queue. We bought a CI that was actually fast. Throughput at that station went up. Lead time for a customer-visible change did not. Goldratt would have been bored. We were sprinting at a non-constraint and piling inventory — tickets, half-merged branches, features waiting on legal, features waiting on a migration only two people understood — in front of the real gate.
Theory of Constraints is a factory idea: the system’s throughput is the throughput of the bottleneck. IT is a factory whether we like the metaphor or not. Work-in-progress is real. I still use a cheap version of ToC when a team is “busy” and nothing is arriving. Here is why local speed can be a lie, and how I find the gate I am actually paying for.
The last time faster meant slower
A product squad got good at shipping UI. Velocity — I do not love the word — went up. Design was in the room. Backend was “keeping up” by opening PRs that sat on a review from the one person who knew billing. Those PRs aged. Frontend started a second feature that also needed billing. We now had two features in flight, a reviewer who was also on-call, and a Slack channel of “just a quick look.” The constraint was not React. The constraint was a single head that could say yes to money code. Speeding up everyone else increased the queue in front of that head. The system slowed. Incidents went up because the same head was rushing reviews at 6 p.m.
The fix was not another reviewer trained on nothing. The fix was: one billing change at a time, a written path so a second person could review, and a stop on starting UI that had no billing design. Throughput of finished billing features went up when we stopped starting. That is ToC in a hoodie. It is also rude, because the busy people feel punished for being fast.

How I find the constraint without a consulting workshop
I walk a finished change backward from the user. Where did it wait longest? Not “where were people busy.” Wait. A week in code review. A week in QA because the environment is a rumor. A week in app-store review. A week in legal. A week because staging data is a lie. The longest wait that is structural — it happens to most changes of that type — is a candidate constraint.
I also look at inventory. If we have forty tickets in “in progress” and two in “released this month,” we are optimizing stations and ignoring the gate. If we have a huge backlog and a fast team, the constraint might be intake: we cannot say no, so the system is flooded. Flooding a factory does not make it produce more. It makes it thrash.
Common IT constraints I have actually named:
- The person who can approve a production data change.
- A shared staging that is always dirty.
- A flaky suite that makes CI a coin flip, so people batch work.
- A vendor (Apple, a bank, a payments partner) with a calendar we do not own.
- A monolith deploy window that is once a week because we are scared.
- A product decision that never lands, so engineering starts three options.
I do not start with “the developers are slow.” That is the cheapest story and the least often true once the team is competent. I start with queues.
Why “ship faster” at a non-constraint makes the system worse
WIP has a carrying cost. A half-built feature is a merge conflict, a stale assumption, and a context switch. If you accelerate the station before the constraint, you grow WIP. WIP hides in Jira as “almost done.” It also hides as mental load. People look busy. The user sees nothing. Then someone starts a fifth almost-done because the first four are blocked. Local utilization looks great. Flow looks dead.
This is why I am suspicious of “we need more engineers on X” when X is not waiting on engineers. More engineers on a non-constraint is a bigger queue at the constraint. I have hired into that lie. I have also been the hire.
It is also why sprinting every team to 100% utilization is a ToC own-goal. A constraint needs slack in front of it — a small buffer — and slack after it so it never starves. If every team is at 100%, the constraint is starved whenever anyone sneezes. I will take an “idle” platform engineer over a fully booked one if the platform is the gate. Idle is a dirty word in IT. It is sometimes the only way the gate stays fed.
What I do once I think I know the gate
Exploit it: make sure the constraint only does constraint work. The billing reviewer does not sit in three status meetings. The person who can run the migration does not also own the newsletter. I have failed this by “just this once” adding them to a hiring loop.
Subordinate everything else: other teams start less. Frontend does not start a screen that will rot. We sequence. This is the part people hate. It feels like we are slowing them. We are protecting the gate.
Elevate: now, and only now, add capacity at the constraint. Train a second reviewer. Buy a second staging. Automate the part of legal that is a checklist. Split the weekly deploy if fear was the only reason. Do not elevate a non-constraint because it is easier to buy a tool than to tell a fast team to wait.
Then look again. The constraint moves. After we doubled billing reviewers, the gate became app-store review, then it became “we cannot decide the copy.” I do not install a permanent ToC office. I install a habit of asking where the last five releases waited.

Kanban, Scrum, and the factory words I still use
WIP limits are ToC for people who do not want the book. A column that cannot grow is a way to stop flooding the gate. I would rather have a small limit and an argument than a sprint that “committed” to twelve items and finished the three that did not need the constraint.
Story points at a non-constraint team are a way to celebrate local speed. I have used them. I now prefer lead time of a finished slice and a count of items waiting on the named gate. If the gate is legal, I want a metric legal will look at. If they will not look, the constraint is not a process problem. It is a power problem. ToC will not repeal power. It will name it. Naming it is still useful in a 1:1 with a founder.
DevOps as “everyone deploys whenever” is great when deploy is not the constraint. If deploy is the constraint because the suite is a coin flip, “deploy more” is flooding. Fix the suite. That is elevate. I have watched companies buy a deploy dashboard and keep a coin-flip suite. The dashboard is a non-constraint tool.
A personal rule I use in planning
Before I add people or a process to a busy team, I ask: what would happen if this team shipped twice as many PRs next week? If the honest answer is “they would wait twice as long on X,” I work on X. If the honest answer is “users would see twice as much,” then maybe they are the constraint and I should help them. Busy is not the tell. The tell is whether their extra output would arrive.
The same pattern shows up in incidents. We page the busiest on-call because they are “the person who knows.” We make them faster with better dashboards. We do not add a second brain. The constraint is knowledge, not CPU. Elevating means a runbook and a shadow rotation, not a prettier graph. I have bought the graph. I have also written the runbook. Only one of those moved the wait.
I still like Goldratt’s novels more than the software-consulting versions. The software versions often become a new ceremony. I do not need a new ceremony. I need a named gate, a smaller queue in front of it, and the social courage to tell a fast team to start less. Shipping faster at the wrong station still makes the system slower. I have the Jira history to prove it. I also have the scar of being proud of that station while the user waited on a review I was too busy to give. Next time I feel proud of a local speed-up, I will ask where the extra work will sit. If I cannot name the shelf, I am not done thinking.