Serverless Functions vs a Long-Running Process When the Agent Wrote Both
Daniel Park
September 30, 2026
The repository looked complete. There was an app/ folder full of API routes, a form for tenants to report a broken boiler or a leaking tap, an upload handler for photos, and an admin page. And there was a second folder called worker/ with a single file, index.ts, that started with while (true). It picked jobs off a table, sent SMS messages to contractors, resized photos, and at 6 a.m. built a digest email for the property manager.
A coding agent had written all of it in an afternoon for a small property-management firm, and the owner’s nephew had deployed it to a serverless platform over a weekend. The web part worked beautifully. Tenants filed requests. Photos uploaded. The admin page showed everything.
Nine days later the firm called me because no contractor had received a single text. There were 212 maintenance requests sitting in the jobs table with a status of pending. The worker had never run. Not once. It was not on the platform, because serverless platforms do not run while (true). They run functions when a request arrives, and nothing ever sent a request to worker/index.ts.
Nobody had done anything stupid. The agent had written two kinds of program, which is often the correct architecture, and the deploy had only picked up one of them.
Agents write both kinds of code, and rarely say which is which
I have seen this shape in maybe half the agent-generated backends I have been asked to look at. The agent correctly figures out that some work should not happen inside a web request: sending messages to third parties, processing images, anything slow or retry-prone. So it writes a queue and a worker. That is good engineering.
What it often does not do is spell out, in the deploy instructions, that there are now two things to run with two different lifecycles. The README says “deploy to Vercel” or “deploy to Netlify” because that is the most common instruction for the web framework it chose. The worker sits in the repo like a second engine nobody bolted into the car.
So the first real question with any agent-built backend is not “serverless or server?” It is: how many processes did the agent actually write, and what does each one expect about time?

How I sort the pieces
I read through the code and put every entry point into one of two buckets, using a few blunt questions.
Does it start because something asked? A form submission, an API call, a webhook from Stripe or Twilio, a page view. If yes, it is a candidate for a function.
Does it finish in a few seconds, reliably, even on a bad day? Not “usually.” Reliably. If a third-party API is slow, does this still finish in time? If it involves resizing a twelve-megapixel phone photo, is that still fast when the function is cold?
Can it be run twice without harm? Serverless platforms and webhook senders both retry. A handler that sends an SMS every time it runs will send two if it is retried after a timeout.
Does it need to remember anything between runs? An in-memory cache, a counter, a connection it wants to reuse, a lock.
Anything that starts on request, finishes quickly, is safe to repeat, and remembers nothing belongs in a function. Anything that loops, polls, batches, runs on a clock, holds a connection, or takes long enough that you would be nervous about a timeout belongs in a long-running process.
For the maintenance app, the sort looked like this:
- Functions: the tenant form, the photo upload (just storing the original), the admin API, the inbound SMS webhook for contractor replies
- Long-running process: sending outbound SMS with retries and backoff, photo resizing, the 6 a.m. digest, and a sweeper that re-queues jobs stuck in
processingfor more than ten minutes
That is the architecture the agent had written. It was right. It just needed both halves to exist in production.
Three ways to run both halves
Once you know you have two kinds of process, there are three honest options.
Put the worker on a small server, keep the web on the platform. This is what I did for the property firm. The web app stayed where it was. The worker went onto a small VPS as a systemd service with Restart=always, reading from the same Postgres database. Total cost: a few euros a month and about two hours, most of which was making sure the worker’s logs went somewhere I could read them and that it restarted after a reboot.
Rewrite the worker into platform-native pieces. Most serverless platforms now offer scheduled functions and some form of queue with a consumer function. The digest becomes a cron-triggered function. Each outbound SMS becomes a queue message consumed by a function. This works well, but it is a real rewrite. The agent’s while (true) loop assumed it could hold state across iterations, back off by sleeping, and sweep stuck jobs on its own schedule. All of that has to be re-expressed in the platform’s terms, and the retry semantics change.
Put everything on one server. Run the web app and the worker as two services on the same VPS. This is the simplest to reason about, and for a firm with a few hundred tenants it is plenty. The cost is that the web app now inherits the server’s availability. If the box reboots for a kernel update, tenants see an error page for a minute.
I briefly tried the third option as an experiment, asking the agent to fold the whole thing into one deployable. It did, and the result was tidy, but it also moved the inbound SMS webhook onto the box. The first time the VPS restarted, Twilio’s retries hit a closed port and contractor replies were delayed. On the serverless platform, the webhook had simply always been there. That convinced me that the split was not just an accident of the agent’s defaults. The request-driven parts genuinely benefit from a platform that is always listening, and the clock-driven parts genuinely need something that stays awake.

The subtle problems when both halves share a database
Running functions and a worker against the same database introduces a few problems that neither half has on its own. These are the ones I hit.
Connection counts. Each serverless function instance can open its own database connection. Under a burst, say sixty tenants filing requests after a building-wide power cut, you can end up with dozens of short-lived connections on top of the worker’s steady one. A small managed Postgres plan can run out. Whether you need a connection pooler in front of Postgres comes down to exactly this mix of many brief connections and a few long ones, and in this case the answer was yes. I put the functions behind the provider’s pooled endpoint and left the worker on a direct connection.
Job claiming. The agent’s worker used SELECT ... FOR UPDATE SKIP LOCKED to claim jobs, which is correct and safe if you run more than one worker. But it marked jobs as done only after the SMS was sent. If the worker crashed between sending and marking, the job would be re-claimed and the contractor would get a duplicate text. I added an idempotency key stored with the provider’s message ID so a retry could check whether the message had already gone out.
Who writes the status. The function wrote pending. The worker wrote processing and sent. The inbound SMS webhook wrote acknowledged. The admin page let the manager write closed. Four writers, one column, no rules about transitions. A contractor replying “on my way” to a request the manager had already closed flipped it back to acknowledged. I asked the agent to add a small state machine with allowed transitions and to reject anything else. It did it in ten minutes. Nobody had asked before.
Time zones. The digest was scheduled for “6 a.m.” in the worker using the server’s clock, which was UTC. The firm is in Lisbon, so this was fine for half the year and an hour off for the other half. Serverless cron schedules usually run in UTC too. Whichever half owns the clock, write the time zone down explicitly.
What I tell people now
When someone hands me an agent-built backend and asks where to host it, I do three things before answering.
- List every entry point. Routes, webhooks, scripts, anything with
while (true),setInterval,cron, or a queue consumer. The agent will have written some of each. Count them. - Sort them by what starts them. Requests go to functions. Clocks, queues, and loops go to a process. Anything that must be always reachable from the outside, like webhooks, should not depend on a single box that reboots.
- Make the deploy match the sort. If there are two kinds of process, there are two deploy targets, two sets of logs, and two things to monitor. Write that into the README yourself, because the agent’s README almost certainly only describes one.
It also helps to tell the agent up front. “The web app will run as serverless functions. Background work must run in a separate long-lived worker on a small VM. Document how to deploy both.” That one sentence produces a noticeably different repository: explicit deploy steps for each half, a health endpoint on the worker, and usually a note about time zones and retries.
The 212 requests went out over about forty minutes once the worker started, throttled so contractors did not get a wall of texts at once. The property manager spent the next day apologising to tenants about boilers that had been broken for a week. None of that was a serverless problem or a server problem. It was a two-process app deployed as a one-process app, and a README that did not know the difference.