A system for working with AI in 2026: the developer workflow I’d actually follow
Casey Holt
September 18, 2026
I stopped opening Cursor and hoping. That was 2024 energy: a prompt, a shrug, a pull request that looked locally fine and was globally slightly wrong. The model did not run the project because it was smart. It ran the project because I had not given the work a loop with a human checkpoint that could say no.
The workflow I actually follow in 2026 is boring on purpose. It looks like a small operating system: a written job, a bounded context, a generated draft, a test that can fail, a review that hunts invariants, and a place the model is not allowed to touch without me. Tools change — Claude, Copilot, Cursor, Codex, whatever ships next quarter. The loop does not.
The job is a file, not a chat
I start in the repo, not in the sidebar. A short markdown file or a ticket that names the outcome and the thing that must remain true. “Add CSV export” is a prompt. “Export rows match the ledger for that period, column order is stable, and a mismatch pages us” is a job. I paste the job into the agent. I do not ask the agent to invent the job. Inventing the job is how you get a confident spec that is also a hallucination.
If I cannot write the invariant in one or two sentences, I am not ready for a model. I am ready for a walk or a conversation with the person who will be on-call. I have burned evenings letting Claude interview me into a design I did not believe. The transcript was long. The production change was still mine to own.
I keep a docs/ai/ note in services I touch often: the three invariants a model may not change, the packages it should not rewrite, the test command that is the source of truth. That note is not a manifesto. It is a preflight. Agents forget. Files do not, if I put them in the context.
Context is a budget
I do not dump the monorepo into the window. I name the files that matter: the module, the test, the migration, the runbook. If the agent needs to discover the tree, I let it search — and I watch what it opens. A repo that feels “dumb” is usually an index problem, not a model problem. I fix .gitignore, I add a small architecture note at the package root, I stop generating 400-line files with no heading.
I treat the context window like RAM on a small box. The hottest bytes are the invariant, the failing test, and the current diff. Chat history is a liability after about twenty turns. I start a new thread with the job file and the latest failure. Continuity is a feeling. Fresh context is a tool.
Secrets stay out. I have seen people paste .env because “it will help with the local setup.” It helps the model, and anyone who later exports the chat. I use a redacted compose override and a sentence that says which variable names exist, not the values.

The draft is guilty until a test says otherwise
I ask for the test first when the work is a rule. When the work is a dirty integration, I ask for a spike behind a flag, then I pin what I learned. I run the test myself. I do not trust the “all tests passed” line in the chat. I have watched that line lie when the agent ran a subset, or ran against a mock, or ran in a directory that was not the service.
My default commands live in the README: task test, make test-int, the Testcontainers suite. The agent is told to use those, not to invent npm test -- --passWithNoTests. Generated tests that only hit Mockito get deleted or promoted. I want the dangerous path on real Postgres or a real contract file.
I read the diff like a junior who is fluent and slightly reckless. I hunt:
- Migrations that drop or widen columns “to match the type.”
- Idempotency that became a comment.
- Retries on non-idempotent POSTs.
- Authz checks that moved into the client.
- Log lines that include tokens or PII.
If the diff is large, I ask the agent to split it before I review. A 40-file “cleanup” is how invariants die in a rename. I have started refusing those PRs even when I wrote the first prompt.
Where the model is not allowed
Money, deletions, and permission changes need a human-written test before I accept a generated implementation. I put that in the job file. I will let the model draft the mapper and the boring DTO. I will not let it be the first author of a payout state transition.
I also keep production credentials, deploy keys, and “just run this against staging” out of the loop unless I am watching. An agent with shell access is an intern with root. Useful. Not unsupervised. I use the sandbox. I read the command. I have cancelled git reset --hard more than once.
On-call is human. If a generated change pages, I treat it as my change. The postmortem does not mention the vendor. It mentions the missing test.
The daily loop
- Write or update the job file. One outcome, one invariant, out of scope in three bullets.
- Gather the smallest file set. Open a fresh thread.
- Ask for a failing test or a spike, not a rewrite.
- Run the real test command. Paste the failure, not a paraphrase.
- Review the diff for the hunt list. Request a split if it is fat.
- Ship behind a flag when the blast radius is unclear.
- Update the job file with what we learned so the next thread does not start from folklore.
I do this even for “small” tickets. Small tickets are how a model quietly changes a default. The loop is faster than cleaning a clever Tuesday.
Pairing with a human still beats pairing only with a model when the problem is a product question. I use the agent to draft. I use the person to decide. Mixing those roles is how meetings become prompt reviews.
What I do not do anymore
I do not keep an all-day chat that becomes the design. I do not accept “I updated the tests” without seeing the command. I do not let the model write the commit message from a glance — I write the sentence that I want in git blame. I do not measure the day in accepted hunks. I measure it in invariants that stayed true and tickets that closed without a revert.
I do not use the same workflow for a spike and a money path. A spike can be sloppy if it dies in a branch. A money path that is sloppy is a Sev2 with a fluent explanation.
I do not collect prompts as if they were a skill tree. The skill is the checkpoint. Prompts rot. Checkpoints compound.
How I handle the team version
On a team I do not make this a branding program. I ask that every generated PR link the job file or the ticket invariant. I ask that CI run the real suite, not the agent’s memory of the suite. I do not ban assistants. I ban unsupervised merges to the money path. People who want a house style can copy the hunt list. People who want a prompt library can keep it in their notes. The shared artifact is the invariant and the test command.
If two engineers and a model disagree, the test wins. If there is no test, we write one before we argue taste. That rule has ended more Slack threads than any style guide I have published.
A week this saved
We needed a new export for finance. The first agent draft added a column and reordered two others because “the schema was cleaner.” The job file said the column order was the product. The test I wrote first — a golden CSV — failed. The second draft kept the order and added the column at the end. Finance never knew. The first draft would have broken their VLOOKUP on a Monday. That is the whole system: the job named the thing that could hurt, the test could fail, the model did not get to redefine the product in the name of cleanliness.
I would follow this loop with any assistant that can see the repo. I would not follow a vibe. Vibe coding a demo is a weekend. Vibe coding a ledger is how you get a demo that invoices. The workflow is how I keep those from becoming the same branch.
Write the job. Bound the context. Generate the draft. Fail a real test. Hunt the invariant. Leave the model off the pager. That is the system. It is not exciting. It is the only one I have trusted after a year of watching exciting chats ship slightly wrong systems that looked like they were written by someone competent — because they were, locally, and competence is not ownership.
If you take one piece: stop prompting ad hoc. Put the invariant in a file the agent must read. The rest of the tools will change. That file is the workflow.