Hiring in the agent era: I wouldn’t staff like it’s 2021. Here’s the smaller team I’d build
Marcus Dalton
September 18, 2026
In 2021 I hired like the constraint was headcount approval. Open the req, fill the squad, add a contractor for the mobile app, plan a team offsite. The work expanded to the size of the Zoom gallery. Some of that was real growth. Some of it was people who existed so other people would not have to read a log.
I would not staff that way now. Not because “AI replaced engineers.” Because the mix of work changed, the cost of a wrong hire went up, and a lot of the tickets we used to give juniors are now a bad afternoon with Cursor plus a worse afternoon if nobody reviews the result.
Here is the smaller team I would build for a product that already has customers — say a B2B workflow app, a Spring and TypeScript stack, a real on-call. Not a fantasy five-person unicorn. A headcount I would defend to a CFO who has been reading vendor blogs.
What I no longer hire for first
I no longer open with three mid-level “full stack” seats and a promise that they will grow into ownership. In 2021 that was how you built a bench. In 2026 that is how you build a review queue. Agents are loud on the first draft and quiet on the invariant. If the only people who can see the invariant are already drowning in PRs, the bench does not grow. It generates.
I also no longer hire a dedicated “frontend village” and a dedicated “backend village” for a product that is one domain. Two timezones of handoff for a form that writes to one Postgres table was always a luxury. It is an unaffordable luxury when a single engineer can scaffold both sides and the risk has moved to the parts the scaffold gets wrong: authz, money, migrations, the job that retries twice and charges twice.
QA-as-a-phase hiring is the other 2021 habit I would drop. I want people who test. I do not want a separate team whose first look is a build that already escaped. A single QA engineer who owns exploratory and a Playwright suite can still be a force multiplier. A QA department that exists because engineers do not write regression tests is a process smell I used to staff around. I will not staff around it again.
The seats I would actually open
For a product team that needs to ship and stay up, I want this shape. Numbers assume one product, not a platform org.
- Two senior ICs who can own a slice end to end. Not “full stack” as a euphemism for junior. People who have been paged for their own decisions. One can lean backend, one can lean product UI. Both can read the other side’s PR without a translator.
- One tech lead who still writes the thin, sharp code — the first Temporal workflow, the authorization model, the migration that cannot be undone. This is not a meeting chair.
- One engineering manager if we are more than five, including the lead. If we are four, I am the manager and I am already oversubscribed. I would rather have a fractional EM from a sister team than a lead who pretends 1:1s are optional.
- One product manager who can write a spec the model cannot fake. Acceptance that names the failure. “User can export CSV” is a prompt. “Export is the same rows finance used last month, and a mismatch pages us” is a job.
- One designer who owns the system, not a stream of one-off Figma frames. Agents will generate UI. Someone has to refuse the third date picker.
- A shared SRE or platform person — maybe 50% — for the path from commit to production and the pages at 3am. I will not make every IC invent a Helm chart.
That is six and a half people, not twelve. I would add a mid-level only when one of the seniors has time to apprentice them on work the agent cannot be trusted with: reading a query plan, walking an incident, saying no to a product shortcut. A mid-level without that apprenticeship is how you get faster at generating debt.

What I need each person to do with agents
I do not hire for “prompt skill” as a primary signal. I hire for taste about when to throw the output away. In interviews I now include a short pairing session where we let Copilot or Claude draft a change, then I watch what the candidate deletes. The people I want delete the plausible tests and keep the ugly ones. The people I do not want compliment the model and add a comment.
The seniors I want can:
- Use an agent to draft the boring mapper and still write the transaction boundary themselves.
- Refuse a generated migration that drops a column “because the type changed.”
- Write a eval or a golden file for the one place we do ship model output to a user — support summaries, say — instead of vibe-checking production.
The PM I want can write a spec that is a contract, not a novel. Spec-driven development is fashionable. Most of what I have seen is a Google Doc the model summarizes into tickets the model then implements, and nobody notices the loop is closed around a hallucination. A good PM in this era is a person who can break that loop with a test only a human could have demanded.
The designer I want treats generated screens as a mood board. If they cannot hold a system in their head — type, spacing, empty states, the error that happens when Stripe returns card_declined — the agent will give us forty screens and no product.
What I would stop measuring
Lines of code and PR count were always vanity. In 2026 they are worse: they reward the person who accepts the most generated text. I have watched a dashboard light up green the week we turned on agents and watched escaped defects follow two sprints later. I will not take “velocity” as a compliment until I see change-fail rate and time-to-restore in the same slide.
I still want DORA-ish numbers. I want them on a team this small, not as a corporate OKR wallpaper. If we ship twice a week and page twice a week, we are not a high-performing small team. We are a generator with a pager.
I would also stop using “AI fluency” as a filter that means “uses ChatGPT in the take-home.” I would rather see a candidate explain a production bug they caused. Agents do not make that story less relevant. They make it more relevant, because the next bug will look like it was written by someone competent.

Contractors, agencies, and the seats I will rent
I will rent a specialist for a bounded problem: a nasty iOS store rejection, a SOC2 evidence week, a data backfill that needs a person who has done it. I will not rent a “pod” to own a core domain. Agents make pods look cheaper because the pod can produce more screenshots. The integration tax did not go down. It went up, because the pod’s code looks locally fine and is globally slightly wrong.
I will not hire an “AI engineer” as a third tribe unless we actually ship a model. A retrieval feature in support search is a product engineer plus a measured eval. A title that exists to reassure the board is 2021 energy with a new noun.
Juniors, still — just not as cheap throughput
I am not giving up on juniors. I am giving up on juniors as a way to buy more tickets per sprint. If I hire one, the deal is explicit: they sit with a senior on the paths that teach judgment, they use agents in the open, and we review the review. Their job is to become someone I can leave alone with the ledger. That takes calendar time from a senior. If I cannot afford that time, I cannot afford the junior. Hiring them anyway is how we created the 2026 market’s surplus of people with two years of generated React and no incidents they understand.
Internally I would rather promote a curious mid who has been on-call than hire a “staff-shaped” stranger who has only led in a company with a platform org of 200. On a six-person team, staff means you pick up the ugly operational work without a ticket. The title from a big company does not always transfer. I have learned that the expensive way.
The first ninety days of the smaller team
Week one: freeze vanity metrics. Put error budget and restore time on the same wall as ship dates.
Week two: write down the three invariants an agent is not allowed to touch without a human test — money, authz, deletions. Put them in the repo, not in a wiki.
Week three: kill one meeting that existed to coordinate the twelve-person version of this team. If the meeting still has a purpose, it will come back. Most of them do not.
Week four: pair every engineer with the PM on one spec that includes a failure. Ship that spec. If we cannot, the team is not small. It is underpowered, and I hire the missing senior before I hire a tool.
I would staff like it is 2026: fewer people, sharper seats, agents as interns who do not get merge rights. I would not staff like it is 2021 with a chatbot in the all-hands. The chatbot does not come to the incident. The smaller team does. That is the whole hiring plan.