How coding agents actually find your code: search, indexes, and why some repos feel “dumb”

Casey Holt

Casey Holt

September 18, 2026

How coding agents actually find your code: search, indexes, and why some repos feel “dumb”

I stopped blaming the model when it missed the obvious file. I opened the obvious file. It was generated, or ignored, or named utils2.ts in a folder the indexer never saw because we had told the tool that src/legacy was noise. The agent was not dumb. The map we handed it was dumb.

Coding agents do not “understand the repo.” They search, they embed, they open the hits, they stay inside a context budget. If those three steps cannot find the payout state machine, they will invent a second one in a convenient new file. I have merged that second one. I have also deleted it. This is how I think about the map now.

What “find” actually is

Most tools do some mix of:

  • Filename and path matching — payout in the path wins a lot.
  • Text search — ripgrep-shaped, sometimes AST-aware.
  • An embedding index over chunks of files, refreshed on save or on a job.
  • A language server, if they bother — go-to-definition when the graph is intact.
  • Whatever you put in rules, agents.md, cursorignore, or a README they always read.

None of that is magic. All of it fails in predictable ways. Generated protobufs that are gitignored cannot be cited. A 4,000-line service.go becomes one or two chunks that mention the wrong function. A monorepo with twenty packages and no root map looks like twenty unfinished thoughts. Symlinks and generated folders duplicate the world. .cursorignore that excludes internal/ because someone was tidy will hide the domain.

The model then does what a lost junior does: it writes where it is standing. You get a new handler next to the chat, not a change to the existing one three directories down.

Why some repos feel smart

The repos that feel smart have boring discoverability:

  • Package names that match the words we say in Slack. billing/, not platform-core-lib/.
  • A short architecture note at the root and at each package: what lives here, what must not.
  • Small files. Not as a religion. As a chunking strategy. A 200-line module is a hit the agent can swallow.
  • Tests named after the rule. Search for “locked beneficiary” should land in a test, then in the code.
  • Generated code either committed and searchable, or clearly pointed at — “see make gen, do not edit z*.go.”
  • Ignore files that exclude build output and node_modules, not the domain.

I have started treating AGENTS.md or the Cursor rules file as an index, not as poetry. Five lines: where money lives, where not to write, the test command, the generated paths, the glossary words. If the file is a manifesto, the agent skims it the way humans skim a manifesto.

A monorepo tree printed on paper with folders circled and others crossed out

The index problems I keep finding

The ignore list was written for humans. Humans do not need vendor/ in the tree. Agents sometimes do, if the only copy of a type is there and we do not vendor-in a stub. More often we ignore the right junk and also ignore docs/ or deploy/ because they are “not code.” Then the agent cannot find the runbook and invents a restart procedure.

The embeddings are stale. A file was moved last week. The index still points at the old path. The agent opens a 404 in its own mind and writes a replacement. I reindex after large moves. I do not assume the IDE did.

The chunk is the wrong size. A table of thirty endpoints in one file yields a hit that is the first endpoint, every time. I split the file or I add a table of contents comment the search can match. Ugly. Effective.

The name is a lie. Manager, Helper, Util, Common. Search cannot distinguish. I rename toward the glossary. DDD’s cheapest win is also the index’s cheapest win.

Binary and generated assets. A SQL file that is only inside a migration JAR will not be found. I keep the source SQL in the repo. I have been bitten by “it’s in the image.”

Secrets in the tree. If the agent can search .env.example, good. If it can search a real .env that someone committed, we have a different incident. Ignore the secrets. Do not ignore the examples.

How I debug a “dumb” session

When the agent misses a file I know exists, I do not write a longer prompt first. I search the way it would: the word, the path, the test name. If I cannot find it in thirty seconds, neither can it. I fix the name or the location.

I ask the agent what it opened. Some tools show the retrieval. If they opened old_payouts.go and not payouts/lock.go, I look at why the old file still ranks — a comment, a leftover export, an index ghost. I delete or mark the old file. Dead code is a trap for humans and a magnet for models.

I add a failing test in the right package and tell the agent to make that test pass. The test is a pointer. It is also the invariant. Two jobs, one file.

If the repo is a monorepo, I start the session in the package, not at the root, unless the change is cross-cutting. Root-level chat in a 200-package tree is how you get a new package named fix/.

A developer pointing at a search result list on a monitor

What I change in a week when a team says “the AI is useless here”

Day one: print the ignore files and the rules file. Remove the domain from ignore. Shorten the rules.

Day two: rename the two worst lie-files. Split the 3,000-line service if it owns two domains.

Day three: add package READMEs of ten lines each for the three hottest packages.

Day four: delete or quarantine dead modules that still rank.

Day five: reindex, then give the agent the same task that failed. If it still fails, the problem may be the task, not the map. If it succeeds, we stop calling the model names.

This is cheaper than switching vendors. I have watched teams switch vendors and keep the same ignore list. The new model was also lost. They called that a bake-off.

A miss I still remember

The lock lived in internal/ledger/hold.go. The agent wrote pkg/payouts/lock.go because search for “lock” ranked a comment in an old experiment and a README that said “locking happens in payouts.” The README was wishful. The ignore file excluded internal/ because a template said internal was private. Private to humans who already knew. Invisible to the tool. We removed the ignore, deleted the experiment, and changed the README to a path. The next session edited hold.go. The model had not gotten smarter over lunch. The map had.

I keep that path in the rules file now: “Ledger locks: internal/ledger/hold.go. Do not add a second lock.” That one line has saved more retries than any prompt pack.

The limits I accept

An index will not replace a human who knows that the real lock is in a stored procedure. I still write that in the rules file. An index will not make a 1998 codebase pleasant. It will make it slightly less of a maze if the maze has signage.

I do not embed secrets, customer data, or the contents of .git. I do not build a second sourcegraph for a ten-person shop when ripgrep plus naming would do. I will use the vendor index that ships with the tool, and I will treat it as a cache I must not poison.

Coding agents find code the way search finds code: names, paths, chunks, and the files we did not hide. Repos feel dumb when we hid the domain, named it mush, or left corpses that rank. I stopped blaming the model for walking the map I drew. I draw a better map. Then I let it walk again.

If you do one thing: grep your ignore file for the folder you say is the product. If it is there, you taught the agent that the product is noise. Take it out. The next “obvious file” might actually get opened. That is the whole trick, and it is not a model upgrade. It is janitor work. I still do it. The janitor work is why some repos feel smart, and why I no longer start a debugging session by insulting the vendor.

Insult the map. Fix the map. Then ask again. If it still misses, you have a real reasoning problem — or a file that should not have been obvious, because the design is in two places. That second problem is yours. The agent just made it visible. I would rather it be visible than fluent and wrong in a new file that search will find next time, twice. Janitors do not get keynotes. They get repos that feel smart. I will take that trade every week I still have to ship.

More articles for you