Agent Skills vs a README: Which Instructions Survive the Next Session
Tobias Werner
September 30, 2026
Our README had a section called “Regenerating test fixtures” that told you, in four clear steps, how to rebuild the recorded HTTP responses our integration tests replay. It had been accurate for two years. New hires followed it on their first week. It was, by the standards of internal documentation, a small success.
Coding agents ignored it almost completely. Across a month of sessions I reviewed, agents that hit a fixture mismatch did one of two things: edited the JSON fixture by hand until the test passed, or deleted the recording and let the test hit the real sandbox API. Neither is what the README said. When I asked one agent afterwards whether it knew how fixtures were supposed to be regenerated, it went and read the README, found the section, and summarised it perfectly. It had simply never looked before acting.
So we moved the procedure into a skill. Three months later the skill was being followed reliably, and it was also wrong: a teammate had changed the regeneration script’s flags, updated the README, and never touched the skill, because no human ever reads the skill. That second failure is the one I think matters more, and it is why “skills or README?” turned out to be the wrong question.
Two different documents, two different readers
A README is written for a person arriving at a repository. It explains what the project is, how to set it up, how to run it, and, if you are lucky, how to contribute. People read it top to bottom once, then return to specific sections when they need them. Its structure follows the reader’s journey: install, configure, run, test, deploy.
An agent skill is written for a model in the middle of a task. It has a short description that tells the agent when to load it, and a body that tells the agent what to do. It is not read on arrival. It is read at the moment a request matches the description, and ideally at no other time.
That difference in timing is the whole story of which instructions actually reach the agent. A README is available. A skill is delivered.
Why the README did not get read
Agents do read READMEs. When they explore an unfamiliar repository, the README is often one of the first files opened. But exploration and execution are different phases, and by the time an agent is fixing a failing test, it is in execution mode. It reads the test, the code under test, maybe the fixture file. It does not go back to the README to check whether there is a procedure for this situation, any more than a busy engineer does.
Even when the README is in context, a single section competes with everything else in it. Our README was about 400 lines. The fixture procedure was around line 280, between a section on local HTTPS certificates and one on release tagging. If the agent had read the whole file early in a long session, that section was a small, distant piece of a large context by the time a fixture test failed.
And the README described the procedure in the language of a human reader: “If the upstream API changes, you will want to regenerate the fixtures rather than editing them.” That is a gentle suggestion to a person. To an agent staring at a failing assertion, it does not read as a rule that applies right now.

Why the skill worked, at first
The skill we wrote was about thirty lines. Its description said: “Use when an integration test fails because a recorded HTTP fixture does not match, or when the user asks to update fixtures. Never edit fixture JSON by hand.” The body had the exact command, the environment variable that pointed at the sandbox, and a check to run afterwards.
Three things made it work where the README had not.
It arrived at the right moment. When a fixture test failed, the description matched, and the procedure was loaded right next to the failure. The agent did not have to remember that a procedure existed somewhere. It was handed one.
It was written as instructions, not advice. “Run make fixtures SERVICE=billing. Do not edit files under tests/fixtures/ directly.” No softening, no narrative. The agent followed it.
It was short. Thirty lines, one job. There was nothing else competing for attention inside it.
Hand-edited fixtures stopped appearing in pull requests within a week. I was pleased with myself for about three months.
Then the skill rotted and nobody noticed
In the spring, a teammate refactored the fixture tooling. The make fixtures target was replaced with a script that took a --service flag and required a new environment variable for the sandbox token. She updated the README carefully, because the README is where humans look, and she knew new hires would follow it. She did not update the skill. She did not know it existed. Nobody reads skills except agents.
For the next two weeks, agents loaded the skill, ran make fixtures SERVICE=billing, got a “no rule to make target” error, and then improvised. The improvisations were creative. One agent wrote a new Makefile target. Another went back to editing the JSON by hand, reasoning that the documented procedure was broken. A third found the new script, guessed its flags from the source, and ran it without the sandbox token, which silently recorded authentication errors as fixtures.
This is the failure mode I had not anticipated. A README that goes stale gets noticed quickly, because a human follows it, it fails, and they complain in the team channel. A skill that goes stale gets followed by agents that do not complain. They route around the broken instruction, and the evidence ends up buried in diffs.

Which instructions actually survive
Looking at both failures together, I started thinking about survival as two separate problems.
Does the instruction reach the agent at the moment it matters? Skills win this decisively. A README section is passive. A skill is triggered. If you have a procedure that should be followed in a specific situation, putting it where it gets delivered in that situation is far more reliable than hoping the agent remembers it.
Does the instruction stay true over time? READMEs win this, for a boring reason: humans read them, humans follow them, and humans complain when they break. Documentation survives when someone depends on it and notices when it fails. A skill has a dependant who never complains.
So the answer to “which one survives” is neither on its own. A README survives in accuracy and fails at delivery. A skill survives in delivery and fails at accuracy. You need something that gets both.
What we do now
After the fixture incident we changed three things. None of them are clever, and all of them are about making the two documents depend on the same source.
Procedures live in scripts, and both documents point at the script. The fixture procedure is now a single script, scripts/regen-fixtures, with a --help that explains its flags and fails with a clear message if the sandbox token is missing. The README says “run scripts/regen-fixtures --help.” The skill says “run scripts/regen-fixtures --service <name>; if it fails, read its --help output and report, do not work around it.” When the script changes, both documents stay true, because neither one duplicates the flags.
Skills sit next to the code they describe, and our code-owners file covers them. We moved the fixture skill’s folder so its path matches the fixture tooling directory in our ownership rules. Anyone changing the fixture tooling now gets the skill in their review automatically. It is harder to forget a file that is assigned to you.
Skills tell agents to stop when a procedure fails. Every skill that runs a command now includes a line like: “If this command fails, stop and report the error. Do not invent an alternative procedure.” This does not stop skills going stale. It stops staleness from turning into creative, silent workarounds. A failed command that surfaces as a question to a human is a bug report. A failed command that the agent routes around is a hidden incident.
What stays in the README
Moving procedures into skills did not shrink our README much. It still holds everything a human needs to arrive: what the project does, how to set it up, how the services fit together, who to ask. We deliberately removed step-by-step procedural detail from it where a script now owns that detail, replacing four-step instructions with one line pointing at the script.
The README is still the document agents use to orient themselves in a fresh session, and that is fine. Orientation is what a README is good at. What it is bad at is being the place an agent looks mid-task for what it should do next.
A test worth running
If you want to know which of your instructions are actually reaching your agents, pick three procedures your team cares about, like regenerating fixtures, adding a migration, or bumping a dependency. For each one, start a fresh session and give the agent a task where that procedure should kick in, without mentioning the procedure. Watch whether it follows your documented steps, improvises, or asks.
Then do the second half, which most people skip. Open each skill and check it against the current state of the code. Run its commands yourself. If one of them fails, you have found a skill that agents have been quietly working around, and it is worth searching your recent pull requests for what they did instead.
Our fixture tests have been stable since the script consolidation. The skill is still thirty lines. The README section is now two sentences. Neither one knows any flags, and that is exactly why both of them are still true.