A system for working with AI in 2026: the developer workflow I’d actually follow

Casey Holt

Casey Holt

September 18, 2026

A system for working with AI in 2026: the developer workflow I’d actually follow

I stopped opening Cursor and hoping. That was 2024 energy: a prompt, a shrug, a pull request that looked locally fine and was globally slightly wrong. The model did not run the project because it was smart. It ran the project because I had not given the work a loop with a human checkpoint that could say no.

The workflow I actually follow in 2026 is boring on purpose. It looks like a small operating system: a written job, a bounded context, a generated draft, a test that can fail, a review that hunts invariants, and a place the model is not allowed to touch without me. Tools change — Claude, Copilot, Cursor, Codex, whatever ships next quarter. The loop does not.

The job is a file, not a chat

I start in the repo, not in the sidebar. A short markdown file or a ticket that names the outcome and the thing that must remain true. “Add CSV export” is a prompt. “Export rows match the ledger for that period, column order is stable, and a mismatch pages us” is a job. I paste the job into the agent. I do not ask the agent to invent the job. Inventing the job is how you get a confident spec that is also a hallucination.

If I cannot write the invariant in one or two sentences, I am not ready for a model. I am ready for a walk or a conversation with the person who will be on-call. I have burned evenings letting Claude interview me into a design I did not believe. The transcript was long. The production change was still mine to own.

I keep a docs/ai/ note in services I touch often: the three invariants a model may not change, the packages it should not rewrite, the test command that is the source of truth. That note is not a manifesto. It is a preflight. Agents forget. Files do not, if I put them in the context.

Context is a budget

I do not dump the monorepo into the window. I name the files that matter: the module, the test, the migration, the runbook. If the agent needs to discover the tree, I let it search — and I watch what it opens. A repo that feels “dumb” is usually an index problem, not a model problem. I fix .gitignore, I add a small architecture note at the package root, I stop generating 400-line files with no heading.

I treat the context window like RAM on a small box. The hottest bytes are the invariant, the failing test, and the current diff. Chat history is a liability after about twenty turns. I start a new thread with the job file and the latest failure. Continuity is a feeling. Fresh context is a tool.

Secrets stay out. I have seen people paste .env because “it will help with the local setup.” It helps the model, and anyone who later exports the chat. I use a redacted compose override and a sentence that says which variable names exist, not the values.

A developer desk with a job note beside an editor and a terminal

The draft is guilty until a test says otherwise

I ask for the test first when the work is a rule. When the work is a dirty integration, I ask for a spike behind a flag, then I pin what I learned. I run the test myself. I do not trust the “all tests passed” line in the chat. I have watched that line lie when the agent ran a subset, or ran against a mock, or ran in a directory that was not the service.

My default commands live in the README: task test, make test-int, the Testcontainers suite. The agent is told to use those, not to invent npm test -- --passWithNoTests. Generated tests that only hit Mockito get deleted or promoted. I want the dangerous path on real Postgres or a real contract file.

I read the diff like a junior who is fluent and slightly reckless. I hunt:

  • Migrations that drop or widen columns “to match the type.”
  • Idempotency that became a comment.
  • Retries on non-idempotent POSTs.
  • Authz checks that moved into the client.
  • Log lines that include tokens or PII.

If the diff is large, I ask the agent to split it before I review. A 40-file “cleanup” is how invariants die in a rename. I have started refusing those PRs even when I wrote the first prompt.

Where the model is not allowed

Money, deletions, and permission changes need a human-written test before I accept a generated implementation. I put that in the job file. I will let the model draft the mapper and the boring DTO. I will not let it be the first author of a payout state transition.

I also keep production credentials, deploy keys, and “just run this against staging” out of the loop unless I am watching. An agent with shell access is an intern with root. Useful. Not unsupervised. I use the sandbox. I read the command. I have cancelled git reset --hard more than once.

On-call is human. If a generated change pages, I treat it as my change. The postmortem does not mention the vendor. It mentions the missing test.

The daily loop

  1. Write or update the job file. One outcome, one invariant, out of scope in three bullets.
  2. Gather the smallest file set. Open a fresh thread.
  3. Ask for a failing test or a spike, not a rewrite.
  4. Run the real test command. Paste the failure, not a paraphrase.
  5. Review the diff for the hunt list. Request a split if it is fat.
  6. Ship behind a flag when the blast radius is unclear.
  7. Update the job file with what we learned so the next thread does not start from folklore.

I do this even for “small” tickets. Small tickets are how a model quietly changes a default. The loop is faster than cleaning a clever Tuesday.

Pairing with a human still beats pairing only with a model when the problem is a product question. I use the agent to draft. I use the person to decide. Mixing those roles is how meetings become prompt reviews.

What I do not do anymore

I do not keep an all-day chat that becomes the design. I do not accept “I updated the tests” without seeing the command. I do not let the model write the commit message from a glance — I write the sentence that I want in git blame. I do not measure the day in accepted hunks. I measure it in invariants that stayed true and tickets that closed without a revert.

I do not use the same workflow for a spike and a money path. A spike can be sloppy if it dies in a branch. A money path that is sloppy is a Sev2 with a fluent explanation.

I do not collect prompts as if they were a skill tree. The skill is the checkpoint. Prompts rot. Checkpoints compound.

How I handle the team version

On a team I do not make this a branding program. I ask that every generated PR link the job file or the ticket invariant. I ask that CI run the real suite, not the agent’s memory of the suite. I do not ban assistants. I ban unsupervised merges to the money path. People who want a house style can copy the hunt list. People who want a prompt library can keep it in their notes. The shared artifact is the invariant and the test command.

If two engineers and a model disagree, the test wins. If there is no test, we write one before we argue taste. That rule has ended more Slack threads than any style guide I have published.

A week this saved

We needed a new export for finance. The first agent draft added a column and reordered two others because “the schema was cleaner.” The job file said the column order was the product. The test I wrote first — a golden CSV — failed. The second draft kept the order and added the column at the end. Finance never knew. The first draft would have broken their VLOOKUP on a Monday. That is the whole system: the job named the thing that could hurt, the test could fail, the model did not get to redefine the product in the name of cleanliness.

I would follow this loop with any assistant that can see the repo. I would not follow a vibe. Vibe coding a demo is a weekend. Vibe coding a ledger is how you get a demo that invoices. The workflow is how I keep those from becoming the same branch.

Write the job. Bound the context. Generate the draft. Fail a real test. Hunt the invariant. Leave the model off the pager. That is the system. It is not exciting. It is the only one I have trusted after a year of watching exciting chats ship slightly wrong systems that looked like they were written by someone competent — because they were, locally, and competence is not ownership.

If you take one piece: stop prompting ad hoc. Put the invariant in a file the agent must read. The rest of the tools will change. That file is the workflow.

More articles for you