Rolling Back an Agent Deploy vs Debugging the Live Site

Andre Okonkwo

Andre Okonkwo

September 30, 2026

Rolling Back an Agent Deploy vs Debugging the Live Site

The first alert came at 6:52 a.m. from a customer, not from my monitoring. “Our sync has been failing since about 5. Getting 422s on every page after the first. Anything change?”

Something had changed. I run a small B2B API that lets accounting tools pull invoice data from a handful of e-commerce platforms. The night before, I had left a coding agent working through a list of improvements, and one of them was “make pagination more consistent across endpoints.” It had done that. It had also, as part of being consistent, changed the pagination cursor from an opaque string to a base64-encoded JSON object, and changed the validator to reject the old format. Every client that stored a cursor from before the deploy and sent it back was now getting a validation error.

What I did next is the part I want to talk about, because I did it wrong, and the way I did it wrong is very natural when an agent is sitting right there offering to help.

Forty minutes of fixing forward

I opened the agent, pasted the customer’s message and the error from the logs, and asked it to fix the problem. It understood immediately. It proposed accepting both cursor formats, wrote the change, and ran the tests. I merged and deployed. Seven minutes.

The 422s stopped. Then a different customer reported duplicate invoices. The agent’s compatibility shim decoded old cursors correctly but interpreted their offset one page off, so clients resuming from an old cursor re-fetched a page they already had. Their accounting tool imported the same forty invoices twice.

I pasted that error back. The agent found the off-by-one and fixed it. Deploy. Nine more minutes.

That fix corrected new requests but did nothing for clients that had already advanced to a new-format cursor based on the wrong offset. Their next page was now skipping forty records. I asked the agent to handle that case. It proposed a migration to normalise stored cursors, except we do not store cursors; clients do. It then proposed a heuristic. I stopped.

By 7:35 I had deployed three times, each fix creating a narrower but stranger problem, and I had customers with duplicate data, customers with missing data, and customers who were fine. If I had rolled back at 6:55, every client would have been sending old-format cursors to code that understood them, and the only damage would have been two hours of failed syncs that their tools would retry automatically.

Founder at a kitchen counter early in the morning with a phone, coffee, and an open laptop

Why the agent makes fixing forward feel right

Before I had an agent, fixing forward during an incident was hard. I had to understand the bug, write the fix, test it, and deploy it, all while stressed. Rolling back was obviously faster, so I usually rolled back.

With an agent, fixing forward feels nearly as fast as rolling back. You paste an error, get a plausible fix in a minute, and deploy. Each individual step looks cheap. What that hides is that every fix is a new deploy of code nobody has had time to think about, pushed into a system that is already in a strange state. The agent is good at fixing the error in front of it. It is not good at noticing that each fix changes what the next error will be, because it only sees the part of the incident you paste in.

The agent also has no stake in the incident ending. It will cheerfully propose a fifth fix. It will not say, “We have been doing this for forty minutes; maybe put things back the way they were.” That judgement has to come from you, and at 7 a.m. with customers emailing, your judgement is not at its best.

The rule I use now

I wrote this on an index card and stuck it to my monitor.

If the last deploy is the likely cause and rolling back does not lose data, roll back first. Debug afterwards, somewhere that is not production.

That rule has three conditions, and each one matters.

“The last deploy is the likely cause.” If errors started within minutes of a deploy, assume the deploy did it until proven otherwise. Do not start by reading logs to understand the bug. Start by checking the timeline. My deploy was at 4:58. Errors started at 5:02. That should have been the end of the diagnosis.

“Rolling back does not lose data.” This is the condition that makes rollback unsafe, and the only one that should slow you down. If the bad deploy ran a migration that dropped or transformed data, or if new code has been writing data in a format the old code cannot read, rolling back the code may make things worse. In my case, the new code had not written anything persistent. The cursors lived in client storage and were still in the old format for most clients. Rollback was safe.

“Debug afterwards, somewhere that is not production.” Once the site is stable again, the bug is still there, on a branch, waiting. Now you can reproduce it on staging with real request samples, let the agent investigate without customers waiting, and think about the fix instead of reacting to it.

What makes rollback actually available

The rule only works if rollback is fast, boring, and safe. After the cursor incident, I spent a weekend making sure it was.

One command, or one button. My host keeps previous builds, so rollback is instant: point traffic at the previous build. If your host rebuilds from Git on every deploy, rollback means reverting the commit and waiting for a new build, which is slower and can fail for unrelated reasons. For anything customers depend on, I want the previous artifact kept and promotable.

Database changes that do not trap you. The biggest reason rollback becomes unsafe is a schema change that the old code cannot live with. I now require that every migration is backward compatible with the previous release: add columns before code uses them, stop using columns before dropping them, never rename in one step. That discipline is sometimes called expand and contract. It means the code can always roll back one release without touching the database.

Knowing which deploy is live. Every response from my API now includes a header with the commit hash. When a customer reports a problem, their request logs tell me exactly which version they hit.

Monitoring that notices before customers do. I had error tracking, but it grouped the 422s as “validation errors,” which I had configured to ignore because clients send bad input all the time. I added an alert on error rate change after a deploy: if any status code jumps sharply within fifteen minutes of a release, I get paged.

Mechanic's workbench with a disassembled part laid out on a cloth in labelled trays

When fixing forward is the right call

I do not want to pretend rollback is always correct. There are real cases where you should fix forward.

  • The bad deploy already changed data in a way the old code cannot handle, and there is no clean way back.
  • The problem is not in the deploy. A third-party API changed, a certificate expired, a disk filled. Rolling back your code does nothing.
  • The fix is tiny and certain. A typo in a config value, a missing environment variable, a wrong URL. One line, obviously right, and faster than a rollback.
  • Rollback would reintroduce something worse, such as a security fix that went out in the same deploy.

Even then, I now fix forward differently. I stop asking the agent to “fix the error” in a loop. I describe the whole incident: what changed, what the symptoms are, which customers are affected, and what state their data is in. Then I ask for a plan before any code. And I cap it: if the first fix does not resolve it cleanly, I stop and reconsider rollback, even if it is harder.

How the cursor incident actually ended

At 7:40 I rolled back to the build from the previous afternoon. Old-format cursors worked again immediately. That fixed the 422s for everyone who had not yet received a new-format cursor.

For the dozen or so clients who had received new-format cursors during the two and a half hours the bad code was live, I wrote a small temporary endpoint that translated them back, and I emailed each affected customer with the exact time window and a suggestion to re-run their sync from a known date. Two of them had duplicate invoices to clean up. One had missing invoices, which a re-sync fixed. It took most of the day.

Then, calmly, on a branch, with the agent and a set of recorded requests from the incident, I redid the pagination change properly. The new cursor format shipped a week later, accepting both formats for ninety days, with a deprecation header on old-format responses and an email to every customer explaining the change.

The agent wrote most of that too. The difference was that nobody was waiting on it, and I had time to ask the question I should have asked at 6:55: what would happen if I just put it back?

More articles for you