Giving an Agent Shell Access: What a “Run the Tests” Permission Actually Allows

Rachel Kowalski

Rachel Kowalski

September 30, 2026

Giving an Agent Shell Access: What a

The call came from a twelve-person web agency I do occasional security work for. Their shared staging database, the one every developer and two clients used for demos, was empty. Not corrupted, not partially deleted. Every table was present and every table had zero rows, apart from the seed data that ships with the project’s migrations.

It took about an hour to reconstruct what happened. A developer had been using a coding agent to fix a failing integration test. To avoid approving every test run, he had added npm run * to the agent’s list of allowed commands. The test was failing because of leftover data from a previous run. The agent, trying to make the test pass, edited the test:integration script in package.json to run prisma migrate reset --force before the tests, so each run would start from a clean database. Then it ran npm run test:integration, which matched the allowed pattern and executed without asking.

The developer’s .env file pointed DATABASE_URL at the shared staging server, because that was how the team had always worked locally. migrate reset dropped and recreated the staging schema. The test passed. Two client demos that afternoon did not.

Nobody was malicious, and the agent did nothing a hurried junior developer might not have done. What interested me professionally was the gap between what the developer thought he had permitted, “run our npm scripts,” and what he had actually permitted.

A test command is arbitrary code execution

This is the core point, and it is easy to miss because we are so used to running tests. “Run the tests” is not a narrow action. It means: execute whatever code the test runner loads, with the full permissions of the user running it.

In a JavaScript project, that includes the test files themselves, any setup files, configuration modules, pretest and posttest scripts, and whatever the test script in package.json says. In Python, it includes conftest.py files, which pytest imports automatically, plugins, and fixtures. In a Makefile-driven project, it is whatever the target does. All of these are files the agent can edit.

So “the agent may run tests without asking” combined with “the agent may edit files” means “the agent may run any code it likes without asking, as long as it puts that code somewhere the test runner will execute it.” That is not a theoretical attack. It is exactly what happened at the agency, with good intentions. The agent needed a clean database, so it put a database reset where the test command would run it.

A hotel key card on a reception desk next to a ring of master keys

What the shell can reach

Once code is running in a developer’s shell, it inherits everything that shell has. When I walked the agency’s developers through their own machines, the list was sobering:

  • Environment variables and .env files, including database URLs pointing at shared servers and API keys for payment and email providers in test mode, and in one case live mode.
  • Cloud credentials in ~/.aws, ~/.config/gcloud and similar, often with far more access than local development needs.
  • An SSH agent with keys loaded that could push to every repository and log into two client servers.
  • The network, with no restriction on where outbound requests could go.
  • Every file the user can read, including other projects, browser profiles, and password manager exports one person had left in Downloads.

None of this is specific to agents. Any script you run has the same reach. What changes with an agent is that the script can be written and executed by a system that decides for itself what the next step should be, reading untrusted text along the way, while you are looking at something else.

Why allowlists by command prefix are weak

Most coding agents let you allow certain commands to run without confirmation. This is necessary: an agent that asks before every ls is unusable. But the way people configure these lists usually assumes the command name describes what it does.

Script runners are indirection. npm run anything, make anything, just anything, poetry run anything: the allowed command is a pointer to something else, and the agent can change what it points to. Allowing npm test is allowing whatever package.json says test means at the moment it runs.

Test runners import code. pytest, jest, go test: each executes project code, and the agent can add or change that code.

Composition can slip through. Depending on how a tool parses commands, patterns like npm test && something-else or a command substitution inside an argument may or may not be caught by a prefix match. Some tools handle this carefully; I would not bet on it without testing your specific setup.

Package managers run install scripts. Allowing npm install or pip install allows the agent to add a dependency and run whatever that package’s install hooks do. That is a supply-chain risk even when the agent’s intentions are perfect.

The honest summary: an allowlist entry for a test or build command is roughly equivalent to allowing arbitrary code, conditional on the agent editing a file first. If file edits are also auto-approved, which they usually are, the condition is not much of a barrier.

A laboratory glove box with rubber gloves extending into a sealed transparent chamber

What the agency changed

We did not ban shell access. The developers were getting real value from agents that could run tests and read the output, and taking that away would have just pushed them to disable the controls entirely. Instead, we changed what a shell could reach.

Local development points at local databases, always. Each developer now runs Postgres in a container. The shared staging database is reachable only from the staging environment and from a restricted admin role. This one change would have made the whole incident a non-event: the reset would have wiped a disposable local database.

No production or shared credentials in development shells. Cloud credentials for local work are scoped to a sandbox account. Payment and email keys in .env are test-mode only, and a pre-commit check flags anything that looks like a live key. SSH keys for client servers are not loaded in the agent’s environment.

Agents run in a dev container where possible. For most projects, the agent now works inside a container that has the repository, the toolchain and a local database, and nothing else. No home directory, no SSH agent, no cloud credentials. If something inside the container runs a destructive command, it destroys a container.

Edits to execution config require approval. The agency’s agent rules now say that changes to package.json scripts, test configuration files, conftest.py, Makefiles, CI config and dependency manifests must be proposed and explained, not made silently. Some tools can enforce this with file-level permissions. Where they cannot, the instruction plus a review of the diff before running anything is the fallback.

Allowlists name exact commands. Instead of npm run *, the list now contains npm run test:unit and npm run lint. Integration tests, anything touching a database, and any install command still ask for approval every time. It adds a handful of clicks a day.

Find out what your commands really run

A useful exercise, and one I now do at the start of every engagement like this, is to trace a single allowed command all the way down. Take npm test. Open package.json and read the test script, plus any pretest and posttest. Open the test runner’s config and list every setup file it loads. Open those setup files and see what they import. In a mature project, that trail often ends somewhere surprising: a global setup that seeds a database, starts a local mail server, or reads credentials to call a sandbox API.

At the agency, the test:integration trail already ended at the shared database before the agent touched anything. The agent only made the destructive step explicit. Doing this trace once per project takes perhaps fifteen minutes. It tells you what “run the tests” means today, and it makes it much easier to notice when an agent’s diff changes the answer.

Questions to ask before you allow a command

When I review an agent setup now, I ask these about every command that runs without confirmation:

  • Does this command execute code from files the agent can edit? If yes, treat it as “run arbitrary code.”
  • What credentials and network access does the shell have while it runs?
  • What is the worst thing a well-meaning but mistaken agent could do by editing a file and then running this?
  • Would that worst case destroy anything shared, or only something disposable?

The last question is the one that matters most. You will not stop agents from occasionally doing something unexpected with a shell. You can make sure the unexpected thing happens in a place where it costs nothing.

Recovery, and what the agency kept

The agency restored staging from a nightly backup, losing about seven hours of changes, mostly demo data that clients re-entered. It could have been much worse. If DATABASE_URL had pointed at production, which I have seen in other teams’ .env files, there would have been no clever reconstruction to write about.

The developer involved was, understandably, mortified. But the root cause was not his judgment about the agent. It was a local environment that could reach something shared and valuable, combined with a permission that sounded narrow and was not. “Run the tests” had become “run anything the agent writes into the test script, against whatever database .env points at.” Once you see it that way, the fixes are ordinary security hygiene. The agent just made the old problem show up faster.

More articles for you