Preview Deployments vs Deploying to Production From the Agent
Marcus Webb
September 30, 2026
I was proud of the setup. Every branch the agent pushed got its own preview URL. The agent would finish a task, push, wait for the build, open the preview in a headless browser, click through the page it had changed, and post a screenshot in the pull request. I would look at the screenshot over coffee, click the link if I cared, and merge. Production only ever received code that had already been looked at on a real URL. It felt like the grown-up version of letting an agent near a live site.
Then a customer emailed the shop I build for, a small online store that sells refurbished film cameras, to ask why there were three listings for a “Demo Product — Leica M6 (Sample)” at €1.
The agent had been working on the product page layout. To test it, it ran the project’s seed script against the preview, which inserted some sample products so the page had something to render. The seed script used DATABASE_URL. And DATABASE_URL, in the hosting dashboard, was set once for “all environments,” because when I first set up the project there was only one database and I clicked the checkbox that made the warning go away.
The preview was a separate deployment. It was not a separate world.
What a preview actually isolates
This is the thing I had not thought through, and I suspect a lot of people using agents with preview deployments have not either. A preview deployment isolates code. It builds your branch, runs it on its own URL, and keeps it away from your production domain. That is genuinely useful.
What it does not automatically isolate is everything the code talks to:
- The database, unless you gave previews a different connection string
- Object storage: uploads from a preview land in the same bucket as production if the keys are shared
- Payments: live keys in a preview mean real charges from a test click
- Email: a preview with production mail credentials will happily send “your order has shipped” to real customers from seeded or copied data
- Webhooks and third-party callbacks, which usually point at production and nothing else
- Analytics: preview traffic polluting the real numbers
When a human developer uses a preview, most of these risks stay theoretical because a human clicks around carefully. An agent does not click carefully. It clicks thoroughly. It fills in every form, submits it, runs the seed script, tries the checkout, and triggers the password-reset flow to see what the email looks like. That is exactly what you want from a tester and exactly what you do not want pointed at live data.

So why not skip previews and let the agent deploy to production?
After the demo-camera incident, a friend suggested the opposite approach: since the preview was effectively production anyway, drop the pretence. Let the agent deploy straight to production behind a feature flag or at low traffic, watch the error rate, and roll back if anything goes red.
For some teams that is a legitimate strategy, and I do not want to dismiss it. If you have good observability, real canary infrastructure, and changes that are fully reversible, deploying small changes to production quickly is often safer than long-lived staging environments that drift.
But I am one frontend developer maintaining a shop for two people who repair cameras. I do not have canary infrastructure. I have an error tracker and a phone. And the specific thing I want from the agent’s deploy step is visual verification: does the product page look right on a phone, does the cart drawer open, did the new filter break the grid on narrow screens. Production cannot give me that without showing it to customers first.
So the answer for me was not “previews or production.” It was “previews that are actually separate, and production that the agent never touches directly.”
What I changed
Environment variables got scoped, one by one. I went through every variable in the dashboard and asked whether a preview should ever see the production value. The answer was no for the database, storage keys, payment keys, mail credentials, and the analytics ID. Each of those now has a separate preview value. The only variables shared across environments are things like the public site name and feature toggles that do not reach outside the app.
Previews got their own database. I use a managed Postgres that supports branching, so each preview can get a copy-on-write branch of a small, sanitised snapshot. If you do not have branching, a single shared “preview” database with seed data works fine. The point is that nothing a preview does can reach a real order.
Payments use test mode, always, in previews. Stripe’s test keys exist for exactly this. The agent can run the whole checkout on a preview and I get a test receipt, not a refund request.
Mail goes to a catch-all inbox. Preview mail credentials point at a sandbox service that captures messages instead of delivering them. The agent can trigger every email flow it wants, and I can read what would have gone out.
Previews are not public. I turned on the host’s deployment protection so preview URLs require a login or a bypass token. The agent’s headless browser gets the token as a secret. Search engines and curious customers get a login wall. Before this, one preview URL had been indexed from a pull-request link that leaked into a public issue.
The agent has no production deploy permission. Its token can create previews. It cannot promote a deployment, change production environment variables, or push to the production branch. Production moves when I merge.

The things previews still cannot tell you
Once previews were properly separate, I ran into their honest limits, which are worth knowing before you lean on them too hard.
OAuth callbacks break. “Sign in with Google” needs registered redirect URLs, and every preview has a new random subdomain. You can register a wildcard on some providers, set up a stable preview alias, or skip social login on previews and use a test account. I chose the test account. The agent logs in with email and password on previews. It means the social login button is the one piece of the UI that only gets tested in production, which I now check by hand after any auth change.
Webhooks do not arrive. Stripe, the shipping provider, and the inventory sync all send webhooks to a production URL. Previews never see them. If the agent’s change touches webhook handling, the preview can only prove the handler compiles. I use the provider’s CLI to forward test events to a preview when it matters, but I do it myself, not through the agent.
Data shape differs. A sanitised snapshot or seed data is cleaner than real data. Real product descriptions have odd characters, missing images, and prices entered in the wrong currency by someone in a hurry. The preview will not show you what the new grid does with a product title that is 140 characters long unless your seed data includes one. I added a few deliberately ugly products to the seed file after the new layout clipped a real listing in production.
Caching and performance lie. Previews run on the same platform but with cold caches and no real traffic. They are useful for “does it render” and useless for “is it fast.”
What the agent’s loop looks like now
The agent finishes a task and pushes a branch. The host builds a preview with preview-scoped variables. The agent opens the preview with its bypass token, logs in with the test account, clicks through the affected pages at two viewport sizes, runs a test-mode checkout if the change is anywhere near the cart, and posts screenshots plus a short list of what it checked and what it could not check. That last part is my favourite addition. A typical note says something like: “Could not verify Google sign-in or the Stripe webhook on preview; both unchanged in this diff.”
I look at the screenshots. If the change is purely visual and the diff is small, I merge and production deploys. If the change touches something the preview cannot prove, such as auth, webhooks, or real-data edge cases, I check that part in production myself right after the deploy, while I still remember what changed.
The actual comparison
If I boil it down, deploying the agent’s work to a preview first buys you visual verification, a link you can look at before customers do, and a place for the agent to be thorough without consequences. It costs you the effort of making the preview genuinely separate, and it leaves blind spots around auth, webhooks, and messy data.
Letting the agent deploy straight to production buys you speed and real-world data, and it costs you your customers as the test audience. That can be a good trade if you have the tooling to catch problems in seconds and undo them cleanly. For a small shop with one developer, it mostly means finding out from an email.
The failure I actually had was neither of those. It was a preview that looked separate and was not. If you take one thing from this, go into your hosting dashboard today and look at which environment variables are checked for “Preview.” Anything that can write to a database, charge a card, send an email, or upload a file should have its own value there, or it should not be there at all.
The three demo Leicas were deleted within the hour. Nobody bought one at €1, which is the only reason this is a funny story and not a refund story.