Logs for an Agent-Built App: What You Need When You Did Not Write the Handler

Elena Vasquez

Elena Vasquez

September 30, 2026

Logs for an Agent-Built App: What You Need When You Did Not Write the Handler

The message came on a Sunday evening. A student at my friend’s ceramics studio had paid for a six-week wheel-throwing course, received the Stripe receipt, and never got a booking confirmation. Her name was not on the class list. My friend forwarded the receipt with one line: “Did the website eat her?”

The booking app was the one a coding agent wrote for the studio earlier this year: Next.js, Prisma, Stripe Checkout, a nightly reminder email, now running in Docker on a small VPS. I had read most of it by then. I had not read the Stripe webhook handler closely, because it worked, and because it was the agent’s code and it had tests.

So I opened the logs. For that afternoon, the handler had produced exactly one line:

Webhook received

No event ID. No event type. No booking ID. No error. The line was the same for every successful payment that week. It told me the endpoint had been called. It did not tell me whether anything happened next.

I found the bug in about forty minutes, mostly by reading code I had not written and guessing. The rest of this is what I changed so that the next time, the logs do the reading for me.

What agent-written logging tends to look like

After that Sunday I searched the whole repository for console. calls. The pattern was consistent, and I have since seen the same thing in two other agent-built projects.

Development breadcrumbs left in. Lines like “Webhook received”, “Creating booking”, “Done”. They were useful while the agent was debugging its own first draft and useless afterwards, because they carry no identifiers. When fifty requests produce “Creating booking”, you cannot tell which is which.

Errors logged without context. console.error("Error:", err.message). The message was often something like “Unique constraint failed on the fields: (slotId,studentId)”, with no indication of which request, which student or which slot.

Catch, log, and return success. This was the bug. The webhook handler wrapped everything in a try/catch that logged the error and returned a 200 response. The agent had clearly learned that Stripe retries webhooks that fail, and it wanted to avoid a retry storm. So when the booking insert failed, the handler reported success to Stripe, Stripe never retried, and the error message went to a log line nobody would read.

Whole objects dumped. In a couple of places the agent logged the entire request body or the full Stripe session object “for debugging”. That meant customer names and email addresses in plain text in a log file on a server, retained for as long as Docker felt like keeping it.

None of this is unusual for code written quickly by a person either. The difference is that when I write a handler, I remember roughly what it does when it breaks. When an agent writes it, the logs are the only memory there is.

Hand holding a flashlight whose beam cuts through a dark stone corridor

The actual bug

For completeness: the student had tried to book, abandoned the checkout, and come back an hour later to try again. The first attempt had created a pending booking row before sending her to Stripe. The second attempt created another pending row for the same slot. When the payment succeeded, the handler marked every pending booking for that student and slot as confirmed in one transaction. The unique constraint on confirmed bookings rejected the second update, and the whole transaction rolled back. The catch block logged the error message, returned 200, and her booking stayed pending forever. The nightly cleanup deleted stale pending rows the next morning.

It is a reasonable bug. Anyone could write it. What made it expensive was that the evidence was spread across three places and none of them named the student, the slot, or the Stripe event.

What I need from logs when I did not write the code

I did not install a logging platform. The app has perhaps two hundred requests on a busy day. What I needed was a short list of rules that turn a log file into something I can search when a specific person says something went wrong.

One ID that follows each request. Every incoming request gets a request ID, generated at the edge of the app or taken from the proxy header, and every log line during that request includes it. Now when I find one interesting line, I can pull every other line from the same request with a single search. This alone would have halved my Sunday.

Business identifiers on every line that matters. The booking ID, the slot ID, the Stripe event ID, the Stripe checkout session ID. Not names or emails; the IDs. My friend can give me a receipt, the receipt has a payment reference, the payment reference leads to the checkout session, and with the session ID in the logs I can go from “a student emailed” to “here is exactly what happened” in one search.

Log at the boundaries. I do not need a line for every function call. I need a line when something enters the app, when it calls something outside the app, and when it leaves:

  • Incoming webhook: event ID, event type, whether the signature checked out.
  • Outgoing calls to Stripe or the email provider: what was called, for which booking, and whether it succeeded, with the duration.
  • End of each request: route, status code, duration, and the IDs involved.

Log the decisions that drop work. This is the one I would not have thought of before this bug. Any branch where the code decides not to do something, like skipping a duplicate, treating a booking as already confirmed, or ignoring an event type, gets a line saying so and why. Silent skips are where agent-written code hides most of its assumptions, because the agent made a judgement call there and never mentioned it.

Errors with the stack and the context, not the payload. The full error, the stack trace, the request ID, and the business IDs. Never the whole object that was being processed.

Structured, one JSON object per line. I switched the app to a small structured logger that writes one JSON object per line with a level, a timestamp, a message and fields. It sounds like overkill for a studio booking app. It is the difference between searching for a booking ID with a quick filter command and scrolling.

The catch block, fixed

The agent’s instinct to avoid retry storms was not wrong; it was just applied to the wrong failures. The fix separates two kinds:

  • If the event is malformed or of a type we do not handle, log it at info level with the event ID and type and return 200. Retrying will never help.
  • If processing fails for a reason that might be temporary, or for a reason a person needs to look at, log it at error level with every ID we have and return a 500. Stripe retries over the following hours, and the error shows up in Stripe’s dashboard as well as ours.

The handler also became idempotent on the Stripe event ID, so a retry of an event that already succeeded does nothing and logs “event already processed”. That line has appeared a handful of times since. Each one is Stripe doing its job and the app correctly ignoring it, which is reassuring to see in writing.

Rows of glazed ceramic bowls drying on wooden boards by a window

Where the logs live on a small VPS

Two practical points I had also skipped.

Rotation. By default, Docker’s json-file log driver keeps container logs without a size limit. On a small VPS that is a slow-motion disk-full incident. I set a maximum size and a small number of rotated files per container in the Compose file. For this app that keeps roughly three weeks, which is longer than anyone takes to notice a missing booking.

Reading them. For an app this size, reading them is docker compose logs with a time window, piped through a JSON filter for the ID I am looking for. That is enough. I thought about shipping logs somewhere with a nicer interface, and decided that the time to do that is when I find myself searching more than once a month.

Asking the agent to fix its own logging

I did not rewrite the logging by hand. I gave the agent the rules above as a short section in the repository’s instructions file and asked it to apply them to every route and job. Then I reviewed the result as carefully as any other change, because logging is code and it can be wrong in the same ways.

Two things came up in review. The agent added the customer’s email to the “booking confirmed” line, which I asked it to replace with the booking ID. And in one route it logged the error and rethrew it, and the framework logged it again, so every failure appeared twice. Both were quick to fix and both would have been annoying later.

Then I tested the failure paths, not just the happy one. The Stripe CLI can send test webhook events to a local server, so I replayed a checkout completion for a booking that already existed, sent a malformed event, and triggered the duplicate pending-booking case that caused the original bug. For each one I asked a single question: from the logs alone, without opening the code, can I tell what happened and to whom?

The first time through, the answer for the duplicate case was still “almost”. The skip was logged but did not say which of the two pending rows was kept. One more field fixed that.

The test I use now

For any app an agent wrote that I now have to keep running, I pick the last real support question, like “I paid but did not get a confirmation”, and try to answer it from the logs alone. If I have to open the code to understand what happened, the logs are missing something, and I add it before the next Sunday message arrives.

The student got her place and a free glaze session for the trouble. The app now writes about four times as many log lines as before. Every one of them has an ID in it.

More articles for you