From IDEs to AI agents: the history of dev tooling — and what I’d bet on next
Casey Holt
September 18, 2026
I learned to program in a world where the tool was a file and a compiler. Then the tool was an IDE that knew my symbols. Then the tool was a language server that any editor could rent. Now people talk about the tool as if it were a coworker who types. I have used all of those layers in production. I do not think we replaced the previous one. We stacked them, and we forgot which layer still holds the blast radius.
This is the history I actually lived through — not a museum of logos — and the bet I would make for the next few years. I am writing as someone who still opens Neovim for a quick patch, still pays for JetBrains when a Java service needs a debugger that does not lie, and still spends most product days in Cursor with Claude or Codex on a leash. The leash is the part of the story vendors skip.
The editor was a keyboard. Then it became a database of your program
The first tools that mattered to me were not “IDEs.” They were make, gdb, and an editor that did not crash when a file hit a few thousand lines. Vim and Emacs were religions because the machine was slow and the network was optional. You learned the buffer. You learned to grep. You were the indexer.
Visual Studio, IntelliJ, Eclipse — the fat IDEs — won when the program stopped fitting in a head. Refactor rename across a Java package is not a luxury. It is how you stay honest in a codebase that predates you. I still open IntelliJ IDEA for a Spring Boot service when I need to step through a transaction and see what Hibernate actually issued. VS Code will get me close. JetBrains still wins when the language is a platform, not a script.
The important shift was not the GUI. It was the project model: a tool that parsed the whole program, not just the file. That is the ancestor of everything that came next. Copilot does not understand your repo because it is magic. It understands a slice of your repo because something already built an index, a language server, a graph of imports. Agents that “know the codebase” are standing on IntelliJ’s shoulders whether they say so or not.
Language servers quietly ate the IDE monopoly
When Microsoft shipped the Language Server Protocol, they did something more lasting than shipping VS Code. They turned “the IDE knows your types” into a socket. Rust-analyzer, gopls, pylsp, typescript-language-server — I could stay in Neovim or move to VS Code or Zed and keep the same diagnostics. The editor became a client. The smarts became a service.
That split is why I do not panic when a new editor ships. Zed is fast. VS Code is the default because of extensions and because GitHub lives there. Neovim is still the place I can think in a terminal on a jump box. The editor brand is taste. The language server is infrastructure. If I am betting on a layer that survives the next five years, I bet on LSP, debug adapters, and a real test runner — not on which window chrome I like in 2026.
I also bet on formatters and compilers as the last judges. rustfmt and cargo clippy, gofmt, prettier or biome, tsc --noEmit. An agent that skips the compiler is a content generator. An agent that cannot get to green tests is a junior who will not stay late.

Autocomplete became a model. Then the model wanted a ticket
GitHub Copilot was the first time a lot of us felt the editor guess a whole function and be right often enough to change muscle memory. Tab-complete for boilerplate is a real productivity gain. I will not pretend I do not use it. I will also not pretend it is architecture.
The jump from Copilot-in-the-editor to Cursor, Windsurf, Claude Code, Codex, Aider, Continue — the agent that can open files, run tests, and propose a diff — is a different product. It is closer to a junior pair who has read too much Stack Overflow and will happily delete your migration if the prompt was vague. I have watched an agent “fix” a failing test by weakening the assertion. I have watched one rewrite a working sqlc query into an ORM call because the surrounding files used GORM. The tool is not evil. The tool is literal and eager.
What I want from an agent is not a personality. I want a tight loop: read the failing test, change the smallest surface, run the same test, show me the diff. Cursor’s agent mode and Claude Code in a repo with a Makefile get close when I write the ticket like a ticket. “Make it faster” is how you get a rewrite. “The N+1 is in InvoiceController#index; add a preload; do not change the JSON shape” is how you get a patch I will merge.
I still keep a human editor in the loop. I read the diff. I run the tests I did not ask the agent to run. I do not let an agent push to main. Those are not Luddite rules. They are the same rules I had for a contractor with a laptop I did not image.
The toolchain I would still pay for in 2026
If I were stocking a new machine for product work this year, I would pay for or install this, in this order:
- A real editor with LSP. VS Code or Cursor if I want the agent in the same window. JetBrains if the language is Java, Kotlin, or a heavy TypeScript monorepo that already lives there. Neovim as the always-available layer.
- A compiler and a test command I trust. One script:
just testormake test. If the agent cannot run one command and see red/green, I am the test runner, and I will get tired. - A typechecker or linter in CI that the agent cannot bypass. GitHub Actions that run
pytest,cargo test, orviteston the PR. The agent can lie to me. It has a harder time lying to CI if I do not give it the token that skips checks. - A debugger. I still need to stop on a line. Agents explain. Debuggers show. When a race only happens with two Temporal workers, I want a breakpoint, not a paragraph.
- Git, used like a seatbelt. Small commits. The agent’s work in a branch.
git reflogwhen it gets clever. I have started more recoveries with reflog than I have with an agent “undo.”
I would not pay for a second AI subscription until I am actually hitting the rate limit on the first. I would not install five agent CLIs. One editor agent and one terminal agent is plenty. The context window is not the scarce resource. My attention while reviewing the diff is.

What I think is already a dead end
Chat-only coding — the tab where you paste a file and hope — is a dead end for anything bigger than a snippet. The model cannot see the rest of the repo unless you built that loop. If your “AI workflow” is ChatGPT in a browser and a lot of copying, you are doing 2023. Wire the tool to the repo or stop calling it a toolchain.
Fully autonomous overnight agents that “just ship” are a dead end for any system with customers. I will let an agent draft. I will not let it own production credentials, cloud consoles, or the database. The demo videos skip the part where the agent runs a migration it invented. I have enough scars from humans doing that.
IDE lock-in as a personality is a dead end. If your whole skill is “I am a 10x Cursor user” and you cannot read a stack trace without the sidebar, you have rented your competence. The inverse is also true: refusing agents to prove seriousness is a dead end if the rest of the team ships with them. I use the agent for the boring slice. I keep the judgment. That is the job now, the same way using IntelliJ did not mean you stopped understanding Java.
Multi-agent swarms that debate in Slack are, so far, a dead end I have only seen in talks. One agent with a good test command beats three agents summarizing each other. If that changes, it will be because the tests got better, not because the personas got cuter.
A short history with names, because the names mattered
cTags and cscope were the first “go to definition” I trusted in C. Visual Studio’s IntelliSense made C# feel like a place you could refactor. IntelliJ made Java navigable for people who did not want to become Emacs lawyers. VS Code plus LSP made that navigability cheap. Copilot made generation cheap. Cursor and Claude Code made generation-plus-tools cheap.
Each step reduced the cost of a certain move: jump to a symbol, rename a method, write a boilerplate test, apply a mechanical refactor across twelve files. None of them reduced the cost of deciding whether the refactor should happen. That cost moved onto the human who still signs the commit. I think that is the stable pattern. Tools get better at motion. Someone still owns direction.
Sourcegraph and similar code-intel products were the enterprise version of “search that understands languages.” Internal RAG over the monorepo is the 2026 version. I have used both. They help when the question is “where do we already do this?” They hallucinate when the question is “what should we do?” I treat repo search as a better grep. I do not treat it as an architect.
What I would bet on next
I would bet on agents that are boring: they run your tests, they stay in a sandbox, they produce small diffs, they cite the files they read. I would bet on evals for those agents the way we eventually bet on CI — a suite of tasks from your own repo, not a vendor leaderboard. I would bet on language servers and build graphs remaining the ground truth, with models as a layer on top, not a replacement for go build.
I would bet that the winning UX is still an editor. The terminal agent is great for a well-specified chore. The editor is where I still see the program. If the next wave is “you never open a file,” I will be a late adopter. I have seen too many wrong files get confidently edited.
I would not bet on a single model vendor. I already switch: Claude for long reasoning and refactors, a faster model for rename-and-fill, local models only for snippets when the repo cannot leave the machine. The toolchain that survives is the one that can change the model without changing how I run tests.
And I would bet that junior hiring gets weirder, not easier. The agent writes the first draft that used to prove a junior could finish a ticket. What I now need to see is whether they can tell a bad draft from a good one — whether they run the tests the agent skipped, whether they notice the weakened assertion. That skill looks a lot like the skill we always needed. The IDE did not kill it. The agent will not either. It just makes the people who never had it faster at shipping a mess.
I started this career treating the editor as the tool. I now treat the editor plus an agent plus CI as the tool. The through-line is not intelligence. It is feedback. Grep was feedback. A red squiggle was feedback. A failing test is feedback. An agent that cannot be failed is not a step forward in that history. It is a chatbot with a write permission. I will keep the write permission on a branch, and I will keep the part of the toolchain that can still say no.