A Staging URL vs Testing on Localhost When the Agent Cannot See Production

Casey Holt

Casey Holt

September 30, 2026

A Staging URL vs Testing on Localhost When the Agent Cannot See Production

The pull request description was confident in a way I have learned to read carefully. “Implemented session-based login with secure cookies. Tested locally: sign-in, sign-out, and protected routes all work as expected.” There was even a short screen recording of the agent logging in on http://localhost:3000, landing on the dashboard, and logging out again.

I merged it. It deployed. And in production, every login bounced the user straight back to the login page. No error. No message. Enter your password, press the button, see the login form again. A perfect loop.

The agent had not lied. Login did work on localhost. The problem was that localhost is not a smaller version of production. It is a different place with different rules, and the agent had no way to see the place where the code actually had to run.

What “tested locally” actually proves

The login loop took me forty minutes to diagnose, and the cause was two lines apart. The session cookie was set with secure: true, which is correct: browsers only send secure cookies over HTTPS. In production, the app sat behind a reverse proxy that terminated TLS and forwarded plain HTTP to the Node process. The framework, not told to trust the proxy, saw an HTTP request and refused to set a secure cookie on it. On localhost there was no proxy, and the agent’s dev configuration had quietly turned secure off when NODE_ENV was not production. So the one code path that mattered was never exercised anywhere the agent could reach.

Once I started looking, I found that almost every production-only bug I had seen from agent-written code fell into a small set of gaps between a laptop and a server:

  • HTTPS and proxies. Secure cookies, X-Forwarded-Proto, redirect loops between HTTP and HTTPS, HSTS, and mixed content warnings. Localhost is plain HTTP with no proxy in front.
  • Real domains. CORS rules, cookie domains, SameSite behaviour across subdomains, OAuth redirect URLs. localhost gets special treatment in browsers that a real hostname does not.
  • Production builds. npm run dev is not npm run build && npm start. Dev servers are forgiving about things production builds are strict about: missing environment variables at build time, server-only code leaking into client bundles, dynamic imports, and static generation that fetches data during the build.
  • The operating system. My laptop and the agent’s sandbox both used case-insensitive filesystems. The server is Linux. import Header from './header' works when the file is Header.tsx on one and fails on the other.
  • Resource limits. A laptop with 32 GB of RAM does not show you what happens when the container has 512 MB.
  • Time zones and locales. The laptop is in my time zone. The server is in UTC.
  • Real data volume. Queries that are instant on 40 seeded rows and slow on 400,000 real ones.

None of these are exotic. All of them are invisible from localhost. And an agent that cannot see production will report success on every one of them in good faith.

Glass revolving door spinning at a building entrance with nobody passing through

Why not just give the agent production access?

The obvious response is to let the agent see production. Give it read access to logs, let it hit the live site, maybe let it SSH in. Then it can verify its own work where it actually matters.

I considered it and decided against it for this project, and I think the reasons generalise. Production is where real users are, with real data and real side effects. An agent verifying a login flow on production creates real sessions. An agent verifying a checkout creates real orders. An agent reading logs sees real email addresses and, if the logging is sloppy, worse. And an agent with shell access to a production box is one confident command away from an outage.

Read-only log access is the most defensible of these, and I do allow it on some projects, through a filtered log stream that strips personal data. But logs only tell you what happened after users hit the problem. I wanted something that would catch the login loop before a user saw it.

The staging URL I actually built

What I set up was modest: a second copy of the app on the same VPS, at staging. on the real domain, behind the same reverse proxy with a real TLS certificate, built with the same production build command, running under the same service manager with the same memory limit. It has its own database, seeded with a sanitised copy of production from the previous week. It has its own environment variables, pointing at sandbox versions of the payment and email providers. It is protected by HTTP basic auth, and the agent gets the credentials as a secret.

That last point is the important one. The agent can reach staging. It cannot reach production. So its verification step changed from “run it locally and click around” to “deploy the branch to staging and click around there.”

When I replayed the login pull request against staging, the loop showed up immediately. The agent saw the redirect back to the login page, noticed that no Set-Cookie header survived, and proposed the fix itself: tell the framework to trust the proxy, and remove the development override that disabled secure cookies. It was the same fix I had found in forty minutes of production debugging. On staging it took the agent about three.

Empty theatre stage under work lights with a single chair and backstage ropes

What makes staging worth having rather than just another server

A second environment only helps if it matches production in the ways that matter, and keeping it matched is ongoing work. Plenty of small sites are better off without one. For an agent that cannot see production, the calculation shifts, because staging is not mainly for me or the client. It is the agent’s eyes.

For that job, staging has to share the things localhost cannot fake:

  • Same proxy configuration. I generate both proxy blocks from one template, so a header rule or an upload size limit cannot drift between them.
  • Same build and start commands. No dev server, ever, on staging.
  • Same OS and runtime versions. Same box, same Node, same system libraries.
  • Same resource limits. The staging service has the same memory cap as production.
  • A real hostname with HTTPS. A subdomain of the production domain, so cookie and CORS behaviour is as close as possible.
  • Realistic, sanitised data. Refreshed weekly from production with emails, names, and addresses scrambled.

And it must differ from production in exactly the ways that keep it safe: separate database, sandbox payment and email credentials, no public access, and no ability to send anything to a real customer.

Localhost still has a job

None of this means the agent stopped running things locally. Localhost is still where the fast loop lives: unit tests, type checks, layout tweaks, and the fifty small iterations it takes to get a component right. Deploying to staging takes a couple of minutes, and nobody wants to wait that long to see whether a margin changed.

What changed is the rule about when localhost is enough. The agent now checks its own diff against the list of gaps above. If the change touches cookies, sessions, authentication, redirects, CORS, the proxy, build configuration, environment variables, file paths or imports, or anything time-related, a local run is only the first step and staging is required before it calls the work done. If the change is a copy edit or a CSS adjustment, localhost is fine and the pull request says so. That split keeps staging for the changes that need it, rather than turning it into a ritual that slows everything down.

What staging still does not catch

I do not want to oversell this. Staging on the same box catches environment bugs well. It is weaker on:

Traffic. Staging gets one agent clicking around. Production gets hundreds of people at once. Race conditions and connection pool exhaustion rarely show up on staging.

Third-party webhooks. Payment and shipping providers send webhooks to production URLs. Sandbox webhooks can be pointed at staging, but they behave slightly differently from live ones.

Data that is weird in ways sanitisation removes. Scrambling names also removes the customer whose surname has an apostrophe that breaks a query. I keep a handful of deliberately awkward records in the seed.

Drift over time. The template approach helps, but anything edited by hand on one environment and not the other will eventually bite. Once a month I diff the two service configurations and proxy blocks.

How the agent’s reports read now

The biggest change is not technical. It is the wording of the pull request descriptions. I asked the agent to separate what it verified from where it verified it, and to name what it could not verify. A typical note now reads something like:

“Verified on staging (HTTPS, production build): sign-in, sign-out, session persistence across reloads, protected route redirect. Verified locally only: unit tests for password hashing. Not verified: behaviour under concurrent logins, Google sign-in (no staging OAuth app configured).”

That is a much more honest document than “tested locally, works as expected,” and it tells me exactly where to look after the production deploy. It also means that when I read “tested locally” now, I know it is a statement about a laptop, not a promise about the product.

The login loop cost one afternoon and about a dozen confused users. A staging URL the agent can reach costs a few extra megabytes of RAM on the same server and an hour of setup. For an agent that is otherwise working blind, it is the cheapest pair of glasses I have ever bought.

More articles for you