Feature Flags vs Shipping the Agent’s Whole Change at Once
Marcus Dalton
September 30, 2026
The new search shipped on a Tuesday. It was the biggest single change my team had merged that quarter: about 2,400 lines, most of them written by a coding agent over two days, reviewed by two engineers, and covered by a healthy set of new tests. It replaced the search in our B2B reporting product with a faster, typo-tolerant version that customers had been asking for.
For most customers it was faster. For our largest customer, whose account holds about forty times more records than the median, search went from under a second to nine seconds at the 95th percentile. Their operations team noticed within an hour, and their account manager noticed ten minutes after that.
We rolled back. Because the search change had gone out in the same deploy as six smaller merges from other engineers, the rollback took those six with it, including a fix for an export bug that another customer had been waiting on. Reapplying everything cleanly took most of Wednesday.
Two weeks later we shipped the same search again, behind a feature flag, to five percent of accounts. It went well. But the flag version brought its own problems, and working through both taught me more about when an agent’s change should be flagged than any amount of reading about flags had.
Why agent changes arrive as one big piece
Before agents, a change like the new search would have arrived over several weeks as a series of pull requests: the new index, then the query layer, then the interface, each merged separately, each small enough to reason about. That was not a deliberate release strategy. It was a side effect of how long it takes a person to write 2,400 lines.
An agent writes the whole thing in one session. It is often better code for being written as a whole, because it is internally consistent. But it means the natural unit of delivery has changed from “a small step” to “the entire feature”, and our release process was built for small steps. We had never needed to decide how to release a big change, because big changes never arrived at once.
So the question is not really “feature flags or not”. It is: now that a complete feature can land in one merge, how do you separate putting the code into production from exposing the behaviour to customers?

The first flag attempt was a mess
For the second attempt, an engineer asked the agent to “put the new search behind a per-account feature flag”. The agent did exactly that, in the most literal way. The resulting diff had twenty-three separate checks of the flag spread across fourteen files: in the API handler, the query builder, the result serialiser, three React components, the analytics events, and the tests.
Every one of those checks was a fork in the code. Reviewing it meant reasoning about two versions of fourteen files at once. Worse, the new tests all ran with the flag on, because the agent had written them for the new search. The old search path, which almost every customer would still be using on day one, was now covered only by the old tests, several of which the agent had modified to account for shared code it had changed.
We threw that version away. The second flagged version took a different shape, and it is the one I now describe to every team that asks.
What a good flag around an agent’s change looks like
One seam, not twenty-three. The flag should be checked in exactly one place: the point where the application chooses which implementation to use. For search, that meant a single function that picks the old or new search service based on the account’s flag, with both services exposing the same interface. Everything downstream of that choice is either entirely old or entirely new. We asked the agent to restructure its change that way, and it did so well once asked. The flag diff went from twenty-three checks to one.
Both paths tested in CI. The test suite runs the search tests twice, once with the flag off and once with it on. It adds a couple of minutes. It is the only way to know that the old path, the one most customers are on, still works after the new code lands next to it.
Default off, everywhere, including in tests you forget about. The agent’s first version defaulted the flag to on in the development environment “for convenience”. That is harmless until someone copies the development config to create a new environment. Default off in every environment, and turned on deliberately.
Targeting you can explain to support. We roll out by account, starting with a small percentage of small accounts, then larger ones, and only then the largest. The rollout plan is written in the pull request description. Support can see which accounts have the new search, so when a customer calls, the first question has an answer.
A removal plan from day one. Every flag has an owner and a removal date in the flag’s description. The removal pull request, deleting the old implementation and the flag check, is prepared by the agent at the same time as the flagged change and left open as a draft. When the rollout reaches one hundred percent and has been stable for two weeks, the owner rebases and merges it. That draft pull request is the most useful thing we added, and it costs almost nothing, because the agent writes it while the context is fresh.

What the flagged rollout caught
At five percent of accounts, all small, the new search was better on every measure. At twenty-five percent, including a few medium accounts, still better. When we enabled it for the first account with more than a million records, the 95th percentile latency for that account jumped to six seconds.
This time it affected one account, for eight minutes, and turning it off was a toggle in the flag dashboard rather than a rollback that took other people’s work with it. The cause was an index the new search relied on that had not been built for accounts above a certain size, because the background job that built it had a batch limit the agent had copied from an older job. The fix took a day. The rollout resumed the following week and reached every account without incident.
The same bug was in the Tuesday release. The difference was that the flag let us find it with one account instead of all of them, and undo it without undoing anything else.
When not to flag
After the search rollout, a couple of engineers started asking the agent to flag everything. That is its own mistake, and we measured it.
Three months later, I counted the flags in our codebase. There were thirty-one. Nineteen had been at one hundred percent for more than a month and were still in the code. Seven had no clear owner. Two were fully off and guarding code nobody remembered writing. Every one of those is a branch in the code that someone has to understand, and a combination of states that the tests may or may not cover.
So we wrote down when a flag is not worth it:
- Small, self-contained changes that a revert fully undoes. If the change touches one area, has no data migration, and reverting the merge commit restores the previous behaviour exactly, a flag adds complexity without adding safety. Ship it whole and revert if needed.
- Pure refactors with no intended behaviour change. A flag around a refactor means maintaining two implementations of the same behaviour. The safety comes from tests and a careful review, not from a toggle.
- Database schema changes. A flag cannot undo a migration. Schema changes need to be backward compatible on their own, by adding the new structure, moving data, and removing the old structure in separate releases. The flag can control which code uses the new structure, but it is not a substitute for making the migration safe.
- Internal tools used by a handful of colleagues. Tell them, ship it, fix what they report.
The rule we use now
Our rule is two questions long. Does the change alter behaviour that customers will see? And would reverting it be expensive, or would seeing it on a few accounts first tell us something useful? If both answers are yes, the change goes behind a flag with one seam, both paths tested, and a removal pull request drafted at the same time. If either answer is no, it ships whole.
For agent-written changes, the first question is yes more often than you might expect, because agents tend to deliver whole features rather than steps. The second question is where most of the judgement goes. A small interface change that is trivially reverted does not need a flag, even if customers see it. A new search engine, a new pricing calculation, or a change to how permissions are evaluated does, even when it looks finished.
We also added one line to our agent instructions: when asked to add a feature flag, check it in exactly one place, keep both implementations behind a common interface, and write tests for both states. Since then, the agent’s flagged changes have arrived looking like our second search attempt rather than our first.
What changed in how we think about releases
The Tuesday rollback was expensive mainly because it was all or nothing, for the search change and for everyone else’s changes that happened to ship in the same deploy. The flag did not make the new search better. It made the release smaller than the change, which is what our old way of working had done by accident when changes arrived in pieces.
Agents write complete features. The release does not have to be complete at once. Deciding when to separate the two, and keeping the separation temporary, is now part of reviewing an agent’s work, alongside reading the code.