One Skill File vs a Folder of Skills: When the Agent Follows the Wrong Procedure

Marcus Dalton

Marcus Dalton

September 30, 2026

One Skill File vs a Folder of Skills: When the Agent Follows the Wrong Procedure

The migration that dropped a column from our analytics warehouse was written by a coding agent following a procedure we had carefully documented. The problem was that it followed the wrong one. One of our engineers asked it to “add a referrer_domain column to the events table,” and the agent loaded our postgres-migration skill, generated an Alembic migration with a reversible downgrade, and wired it into the application’s migration chain. The events table does not live in Postgres. It lives in ClickHouse, and schema changes there go through a separate set of SQL files and a manual runbook. The Alembic migration failed in CI, which was the good outcome. The engineer, trying to be helpful, then asked the agent to “make it work,” and it adjusted the migration to target the ClickHouse connection, including a downgrade step that dropped the column.

Nobody ran the downgrade, so no data was lost. But the review conversation that followed was the moment my team stopped treating agent skills as a documentation problem and started treating them as a routing problem. The procedure was fine. The agent just picked it for a task it was never written for.

I manage a platform team of seven. We have gone from one big instruction file, to a folder of about twenty skills, back down to twelve. Here is what each shape got us, and what actually decides whether an agent picks the right procedure.

Phase one: the single file

We started where most teams start: one long markdown file in the repository root. It covered how to run tests, which directories were off-limits, how to write Postgres migrations, how we scaffold new API endpoints, how we write incident timelines, the ClickHouse schema process, and our conventions for feature flags. By the time we split it, it was about 900 lines.

The single file had one real virtue. Everything was always in context, so the agent never picked a procedure for the wrong task because it never had to pick at all. Every procedure was visible, alongside every other procedure, all the time. If you asked about ClickHouse, the ClickHouse section was right there.

The costs showed up slowly. The file ate a chunk of every session’s context whether the task needed it or not; a one-line typo fix started with 900 lines of process. Instructions from different sections bled into each other: the agent would apply our “always add a reversible downgrade” rule from the Postgres section to ClickHouse SQL where reversibility is not something we do. And editing the file was a team-wide merge conflict magnet, because every procedure change touched the same document.

Most importantly, adherence dropped as the file grew. Instructions near the top were followed reliably. Instructions around line 600 were followed some of the time. Long instruction files are read the way long documents are read by people: the beginning gets attention, the middle gets skimmed.

Phase two: a folder of skills

When our tools started supporting skills properly, we split the file into a folder of them. Each skill got its own directory with a SKILL.md, a name, and a description. The agent sees all the descriptions and loads a full skill only when a request looks like it matches.

The immediate improvements were real. Sessions started lighter. Each procedure could be edited and reviewed on its own. We could add scripts and templates alongside the instructions that used them. Adherence inside a loaded skill went up noticeably, because the agent was reading forty focused lines instead of hunting through nine hundred.

And then, over about two months, we started seeing the failure that opens this piece: the agent following a perfectly good procedure for the wrong task. We collected eleven of these in a shared doc before we understood the pattern.

A railway track switch junction with two diverging tracks in morning fog

Why the agent picked the wrong one

Going through the eleven cases, almost all of them came down to descriptions. Our descriptions had been written as summaries of what each skill contained, not as rules for when to use it.

The postgres-migration skill’s description was “How to write and test database schema migrations.” Read that as a routing rule and it matches any request involving a schema change in any database. Our clickhouse-schema skill said “ClickHouse table conventions and query patterns.” A request to add a column to the events table matched “database schema migrations” far more strongly than “table conventions,” so the agent loaded the Postgres procedure.

We found the same pattern elsewhere. Our api-endpoint skill (“how to add or change API endpoints”) was loaded for a task that only changed an internal gRPC service, and the agent dutifully generated OpenAPI docs for a service that has none. Our incident-timeline skill was loaded when someone asked for a summary of a pull request, because both involved “summarising what happened.”

The folder had introduced a problem the single file never had: a choice. And we had written the inputs to that choice, the descriptions, as if nobody was going to read them as a decision rule.

What fixed it

We rewrote every description with one question in mind: if the agent reads only this sentence, does it know when to use this skill and, just as importantly, when not to?

Descriptions name the system, not the activity. “Schema migrations” is an activity that happens in three of our datastores. “Alembic migrations for the application Postgres database (app_db)” names the system. The ClickHouse skill became “Schema changes to ClickHouse tables in the analytics warehouse, including the events and sessions tables. Not for Postgres.”

Descriptions include explicit exclusions. The single most effective change was adding “Not for…” clauses to skills that sat near each other. The Postgres skill now ends with “Not for ClickHouse or Redis.” The API skill ends with “Public REST API only; not for internal gRPC services.” It feels redundant to write. It measurably reduced mismatches.

Descriptions use the words people actually type. Nobody on the team says “analytics warehouse.” They say “events table” or “ClickHouse.” We went through a few weeks of real requests in our session logs and made sure the terms people used appeared in the descriptions of the skills they needed.

Skills that overlap get merged or split along a clearer line. We had separate skills for “adding a feature flag” and “removing a feature flag.” They were loaded interchangeably, because most requests mentioned both. We merged them. Conversely, our “testing” skill covered unit tests, integration tests and load tests, and the load-test instructions were confusing agents writing unit tests. We split that one.

Phase three: fewer, sharper skills

After the rewrite, we cut from twenty skills to twelve. Some were merged. Some were deleted because nobody had triggered them in two months. A couple were moved back into the root instruction file because they were not really procedures, just facts every session needed, like which package manager we use and how to run the test suite.

That last move was important. Not everything belongs in a skill. We ended up with a simple split:

  • Root instructions hold facts and rules every session needs: repository layout, commands, forbidden directories, the list of datastores and which code talks to which. Short, under a hundred lines.
  • Skills hold procedures for specific, recognisable tasks: writing an Alembic migration, changing a ClickHouse table, adding a public API endpoint, writing an incident timeline.

The root file also got one line that helped the routing a great deal: a table of our datastores with the table names that live in each. When a request mentions the events table, the agent now knows from the always-loaded context that it lives in ClickHouse before it chooses a skill. That single table prevented more mismatches than any description rewrite.

A small engineering team gathered around one laptop at a standing table, one person pointing at the screen

Testing skill routing like code

The change I am proudest of is small. We keep a text file of about thirty realistic requests, each paired with the skill it should load. “Add a referrer_domain column to events” should load clickhouse-schema. “Add a nullable archived_at to users” should load postgres-migration. “Summarise PR 4127” should load nothing.

Whenever someone adds or edits a skill description, they run a handful of these requests in a fresh session and check which skill the agent announces it is using. It is manual and a bit tedious. It takes ten minutes. It has caught three description changes that would have pulled a skill into tasks it was not meant for.

We also added one line to every skill’s body, near the top: “Before following this procedure, confirm the target system matches the description above. If it does not, stop and say so.” It does not always work. When it does, the agent catches its own mismatch before writing anything, which is the cheapest possible place to catch it.

So which shape should you use?

If you have a handful of procedures and they do not overlap, a single file is honestly fine. The agent never has to choose, and you avoid the routing problem entirely. The cost is context and adherence, which does not bite until the file gets long.

Once your instructions pass a few hundred lines, or you have procedures that should only apply in specific situations, a folder of skills is better. But treat the move as a design change, not a reorganisation. You are introducing a decision the agent makes on every request, and the descriptions are the only input to that decision.

Whichever shape you pick, the rule that took us two months to learn is this: a skill’s description is not a summary. It is a routing rule. Write it the way you would write a condition in code, with the system it applies to, the words people will use, and the cases it must not match. The procedure inside can be perfect and still do damage if it is loaded for the wrong job.

The ClickHouse runbook, incidentally, is now a skill of its own with a very clear first line. The engineer who asked for the referrer_domain column ran the same request again last week. The agent loaded the right procedure, wrote the SQL file, and reminded him that ClickHouse changes need the manual apply step. No Alembic in sight.

More articles for you