Pairing With an Agent vs Handing It the Ticket and Walking Away

Marcus Dalton

Marcus Dalton

September 30, 2026

Pairing With an Agent vs Handing It the Ticket and Walking Away

In the same week, I handed two tickets to a coding agent and walked away from both. One came back finished and better than I would have done it. The other came back finished, tested, well documented, and wrong in a way that took longer to unpick than the task would have taken to do properly.

The first was “add CSV export to the monthly reports page.” The second was “invited users should skip the onboarding survey.” On paper, the second was smaller: a condition in one flow, a couple of tests. In practice it was a minefield of product rules that lived in three people’s heads and nowhere in the repository. The agent made reasonable guesses at every one of them, and three of its guesses were wrong.

I redid the survey ticket the next morning, this time sitting with the agent and steering it step by step. It took forty minutes. That contrast, the same tool, the same week, two very different outcomes, is what finally made me think clearly about when to pair with an agent and when to hand it the ticket.

What went right with the CSV export

The CSV ticket had three properties that made delegation work, and I only understood them in hindsight.

The acceptance criteria were checkable. Click export, get a CSV with the same columns and rows as the table on screen, respecting the current filters. Either the file matches the table or it does not. The agent could write a test for that, and I could verify it in a minute.

The repository already knew how to do it. Two other pages had CSV export. The agent found them, copied the pattern, and adapted it. It even used our streaming response helper for large reports, which I had forgotten existed.

A wrong turn would be obvious and cheap. If the export had missed a column or ignored a filter, the test or a quick click would show it. Nothing about the task could quietly go wrong in production without someone noticing on the first use.

I wrote a five-line ticket, started the agent, went to a planning meeting, and came back to a clean PR. Review took ten minutes. Good delegation.

A hand pinning a new work order card to a cork board next to other cards

What went wrong with the survey

The survey ticket failed all three tests, though I did not see it at the time.

“Invited users should skip the onboarding survey” sounds precise. It is not. Who counts as invited: someone who joined through an invite link, or anyone added to an existing organisation, including by an admin through the API? What if an invited user later creates their own organisation? Should they see the survey then? What about the analytics event the survey completion fires, which three downstream dashboards depend on? Does skipping the survey mean firing the event with an empty payload, firing a different event, or not firing it at all?

Every one of those questions had an answer. The answers lived with our product manager, our data analyst, and one engineer who built the invite flow. None of them were in the code in a form the agent could discover. The agent picked sensible defaults: invited means joined via invite link; skip the survey everywhere for those users; do not fire the completion event. The first and third choices were wrong. Invited includes admin-added users, and the dashboards needed the event fired with a flag. Nobody would have noticed the missing events for weeks.

The agent did nothing a new engineer would not have done. The problem was mine: I handed a ticket full of hidden decisions to a worker who could not ask the people who knew the answers, and then I left the room.

What pairing looked like

The next morning I opened a session and worked through the ticket with the agent, one step at a time, with Slack open in another window.

I asked it to find every place invited users are created. It found the invite-link path and missed the admin API path; I pointed it at the second. I asked what the survey completion triggers. It listed the analytics event and a welcome email. I checked with our analyst, who said fire the event with skipped: true. I told the agent. It asked, sensibly, whether invited users who later create an organisation should see the survey; I did not know, so I asked the product manager, who said yes. Each of these took a minute or two. The agent wrote the code between my answers.

Pairing, in this sense, is not me typing and the agent watching, or the reverse. It is short turns: the agent does a small piece of work, I look at it, I add a constraint or a correction or a fact from outside the repository, and it does the next piece. I am the one who can walk over to the product manager’s desk. The agent is the one who can find all eleven call sites in ten seconds.

Forty minutes, and the change was right. The previous day’s walk-away attempt had cost about the same in agent time, plus an hour of my review to find the problems, plus a revert.

A rally car co-driver reading pace notes from a notebook while the driver steers on a gravel road

How I decide now

Over the following months my team and I settled on a few questions that predict which mode a ticket needs. None of them is about how big the task is. Size turned out to be a poor predictor.

How many decisions does the ticket hide? Read the ticket and try to list the choices someone will have to make that the ticket does not settle. For the CSV export: basically none, because existing exports answer them. For the survey: at least four. If the list has more than one or two items and you cannot answer them in the ticket, pair, or answer them first.

Where do the answers live? If the answers are in the code, tests or docs, an agent can find them on its own. If they live in people’s heads, a Slack thread, or a product spec in another tool, the agent cannot get them. Either put them in the ticket, or be there to supply them.

Can a wrong result be seen? If a mistake would show up immediately in a test, a type error, or the first click, walking away is safe; you will catch it on review. If a mistake would be silent, like missing analytics events, subtly wrong permissions, or data written with the wrong flag, you need to be present for the decisions, because review after the fact is exactly where silent mistakes get approved.

Is there a pattern to follow? Tasks that say “do it like we did it over there” delegate well. Tasks that need a new pattern, a new abstraction, or a first-of-its-kind integration pair better, because the design choices are the work.

Writing a ticket you can walk away from

Sometimes a ticket with hidden decisions still needs to go to an agent unattended, because I have a day of meetings or because it is one of eight similar tickets. In those cases the work shifts to the ticket itself. A walk-away ticket on my team now has:

  • The decisions, already made. “Invited includes users added via the admin API. Fire onboarding_completed with skipped: true. Users who later create their own organisation see the survey.”
  • Where to start. The two or three files or modules most relevant, so the agent does not have to guess which of several similar flows is the right one.
  • What done means, in checkable terms. Specific tests that must exist and pass, and ideally one behaviour I can verify by hand in a minute.
  • What not to touch. Especially shared modules or anything with downstream consumers.
  • What to do when stuck. “If you find a case this ticket does not cover, stop and write the question in the PR description rather than choosing.”

That last line is the one that converts silent wrong guesses into visible questions. It does not always work, but when it does, I come back to a PR that says “unclear whether admin-added users count as invited; I have assumed not,” which is a thirty-second fix instead of a week of missing data.

Writing a ticket like this takes me ten to fifteen minutes. That is often most of the value of pairing, done up front. If I find I cannot write the decisions down because I do not know them, that tells me the ticket is not ready for anyone to walk away from, human or agent.

The cost people forget when they walk away

Walking away feels free. The agent works while you do something else. But the cost reappears at review, and it is larger than it looks.

When I pair, I review continuously in small pieces, each one fresh in my mind. When I walk away, I come back to a finished diff and have to rebuild the context from scratch: what the ticket was, what the code around it does, why the agent made each choice. For a well-specified ticket like the CSV export, that is ten minutes. For a ticket full of decisions, it is a slow reconstruction of choices I was not present for, and it is where I am most likely to approve something I did not really understand.

So the real comparison is not “forty minutes of my attention versus zero.” It is forty minutes of attention during the work versus some amount of attention afterwards, plus the risk that review misses what pairing would have caught. On the easy tickets, walking away wins clearly. On the hard ones, pairing is usually cheaper, even though it does not feel that way at the time.

What the survey taught the team

The invited-user change shipped correctly after the pairing session. The dashboards kept their events. And we added a column to our backlog template: “decisions not yet made.” If that column is empty, the ticket can go to an agent unattended. If it is not, someone either fills in the answers or pairs on it.

The agent did not get better at guessing product rules. We got better at noticing which tickets require guessing, and at making sure that when a guess is needed, a person who knows the answer is in the room.

More articles for you