A checked-in prompt file vs chat history: what a coding agent forgets on the next session

Casey Holt

Casey Holt

September 23, 2026

A checked-in prompt file vs chat history: what a coding agent forgets on the next session

Chat history feels like memory. You and the coding agent argued about the auth middleware yesterday, found the off-by-one, and left a trail of “yes, use the existing helper.” Today you open a fresh session and watch it invent a second helper with a confident smile. The conversation did not travel. The repository did—unless you never wrote anything down except in the chat pane.

A checked-in prompt file—AGENTS.md, .cursorrules, a CONTRIBUTING chunk for bots, a repo-local system prompt—looks unglamorous next to a long thread. It is also the difference between an agent that inherits your constraints and one that rediscovers your codebase like a tourist every morning. This is not a purity contest between files and chats. It is an inventory of what survives the session boundary.

What chat history actually stores

Chat history stores turns: your wording, the model’s wording, tool traces, maybe file snippets that were inlined for a while. Inside one session, that context is gold. It carries the bug narrative, the dead ends you already rejected, and the emotional tone of “we are not rewriting the billing module today.”

Across sessions, history is optional product behavior. Some tools resume threads. Some start clean. Some summarize aggressively and drop the sharp edges. Even when history resumes, it is tied to a product account and a UI affordance—not to git clone, not to a teammate’s laptop, not to CI. The next person (or you, on a second machine) does not inherit the thread by fetching main.

Chat also stores noise. Jokes, false starts, Secrets you should not have pasted, and temporary hypotheses that later proved wrong. Reloading all of that can help or can bias the model toward a discarded approach. Memory is not the same as a specification.

Version control diff view concept on a monitor in a calm office

What a checked-in prompt file actually stores

A prompt file stores durable intent: stack conventions, forbidden refactors, test commands, architectural maps, “ask before changing X,” link stubs to deeper docs. It is small on purpose. It is reviewed in PRs. It diffs. It can be blamed.

It does not store yesterday’s stack trace. It will not remember that the flaky test is named test_invoice_rounding unless you put that in a tracked doc. It is a constitution, not a diary. People who expect AGENTS.md to replace conversation are disappointed. People who expect chat to replace a constitution are repeatedly surprised.

The file’s power is bootstrapping. Every new session can be instructed to read it first. Every clone of the repo carries the same starting constraints. That is how teams get agents that stop suggesting a different package manager every Tuesday. Think of it as onboarding docs for a teammate who rereads only the first page unless you force more—except this teammate is tireless and literal.

Concrete contents that earn their keep

High-yield lines tend to be operational, not poetic: how to run unit tests; which package manager is canonical; where generated code is allowed to live; which directories are frozen; how to name PRs; whether the agent may touch migrations; what “done” means (tests green, types clean, no drive-by refactors). Low-yield lines are vibe essays—“be elegant,” “think carefully”—that every model already heard in its base training.

Point to deeper docs instead of pasting them. “Read /docs/auth.md before changing session code” beats inlining three pages of auth lore into the prompt file. The agent can fetch; your job is to make the fetch obvious. Likewise, a failing test is a stronger teacher than a paragraph forbidding a bug class. Prefer executable truth when you can.

What the agent forgets on the next session (almost always)

  • The informal decisions you made mid-thread (“leave the legacy adapter; we delete it next quarter”).
  • The exact files you already ruled out.
  • The reproduction steps you typed once and never added to a test.
  • Your preferred patch style for this repo if it only lived in chat examples.
  • Secrets and environment specifics—which is good—unless you also forgot to document the non-secret setup path.

What it can still “know” without history: whatever is in the tree. Code, tests, types, comments, and the prompt file. That is the durable brain. If a constraint matters twice, it belongs in the tree.

Open notebook beside a laptop with sticky notes on a programming desk

Failure modes of prompt files alone

Prompt files rot. They describe a monorepo layout from six months ago. They ban a library you already adopted. They grow into novels nobody reads, including the agent, because context windows and attention both have limits. A 4,000-line rules file is not seriousness; it is a junk drawer with markdown fences.

They also cannot capture episodic debugging. “The staging webhook fails only after 14:00 UTC because of a cron on the other service” is operational folklore. Put folklore in a runbook doc if it recurs; do not pretend a one-line agents rule replaces an on-call note.

Over-constraining is real. If every session begins with forty contradictory mandates, the model spends tokens on compliance theater. Prefer short, ranked rules: hard constraints first, preferences second, essays linked elsewhere.

Failure modes of chat history alone

History creates false continuity. You assume the agent remembers the auth decision; it assumes you want a greenfield. History also personalizes knowledge that should be institutional. When only your thread knows how releases work, the agent becomes a single-player game.

Privacy and leakage belong here too. Chat logs may leave the building even when the repo cannot. Treating chat as the system of record for architecture is how proprietary reasoning ends up in a vendor UI with a weaker retention story than git.

And history does not PR-review well. You cannot leave a comment on line 12 of last Thursday’s vibe. You can on a prompt file change.

A working division of labor

Checked-in prompt file: evergreen constraints, commands to run, maps of where things live, “never do X,” links to ADRs.

Checked-in docs/tests: anything that must remain true for humans and bots—API contracts, reproduction fixtures, architecture decisions.

Chat history: the active incident, the exploratory spike, the rhetorical back-and-forth that would clutter git.

Session starter habit: paste or auto-attach the prompt file; point at the issue; paste only the minimal logs. Do not paste the entire yesterday thread “just in case” unless you are resuming the same incident.

When a chat decision will matter next week, promote it. One paragraph in the prompt file or an ADR beats hoping the product’s memory feature feels loyal.

Team norms that make this boring (in a good way)

Require prompt-file changes in the same PR as convention changes. If you migrate to pnpm, update the agent rules in that PR—not three fires later. Keep a “hot paths” section listing the five directories that absorb most agent edits. Delete rules that no longer match CI.

For multi-agent or multi-tool shops, prefer one repo-owned instruction surface over six product-specific rule formats when you can. Translation layers drift. The source of truth should be fetchable with the code.

Measure with annoyance, not dogma. If every Monday starts with the same correction (“stop rewriting our error types”), that correction is incomplete until it is in the file. If the file is unreadably long, split it and link. If two rules fight, the agent will pick one at random and you will blame “hallucination” for a documentation bug.

Solo developers benefit too. Future-you is a teammate with amnesia. The prompt file is how Monday respects Friday’s decisions without requiring Friday’s chat product to still exist, still sync, or still price that memory feature the same way.

What to do tomorrow morning

Open your last surprising agent failure. Ask whether the missing knowledge was episodic or constitutional. Episodic stays in chat or a ticket. Constitutional becomes a checked-in rule or a test. Then start the next session with a clean thread on purpose, armed with the file, and see what still breaks. Whatever still breaks is the next line to write down.

Coding agents do not forget because they are rude. They forget because session memory is not repository memory. Chat history is a scratch buffer. A checked-in prompt file is a passport. Travel with the passport. Keep the scratch buffer for the trip you are on—not as the only map of the country. When in doubt, commit the sentence you are tired of repeating. That is the whole practice.

A compact decision rule

If losing the thread would cost you an hour of re-explaining policy, put the policy in the repo. If losing the thread would cost you five minutes of re-pasting a stack trace, keep it in chat. If you are unsure, write the one-sentence rule into the prompt file anyway; sentences are cheaper than rediscovery. And if the agent keeps violating a rule that is already written, your problem is not memory—it is priority, context bloat, or a rule that contradicts the code. Fix the conflict in git, where both humans and agents can see it.

More articles for you