The short version
- Let them use it, and keep handing the choosing down. Both halves or neither. This is the right call if you've got a review gate to make them answer at, you're forming your own seniors rather than hiring them, and you can't take a senior off delivery for a quarter.
- AI didn't take the debugging that used to form juniors. Claude's draft still breaks and they still have to fix it. What it took was the choosing - and it took it at the senior's keyboard, the moment a senior self-served a task they'd once have handed down.
- Banning the tool restores the typing, which never formed anyone. You've always been able to type code you didn't understand. Juniors always did.
- All of it rests on one contestable claim, and if you don't buy that claim, the ban is right. I'd rather say so here than have you find out in three years.
- If you can take a senior off delivery for a quarter, pair them instead. That beats what I'm recommending, it does both halves at once, and it isn't close.
AI didn't take the debugging. It took the choosing
AI didn't take the thing that formed your juniors. It took a different thing, and the argument you're having has named the wrong one.
Break the old loop into its parts and ask which one actually went:
| The part | What it formed | What AI did to it |
|---|---|---|
| Typing | Nothing | Automated. No loss |
| Debugging | The model of how things break | Survives. Claude's draft still breaks, and the loop is faster, so there's arguably more of this |
| Choosing the approach | Judgment. This is the one that mattered | Automated. Claude supplies the approach before your junior has one |
| Defending the choice | Judgment, under a consequence | Untouched. It was always a person asking why |
So I'll concede the whole case against me before I make mine. Yes, judgment is built by choosing badly and paying for it. That was true before any of this and it's true now, and the people telling you to keep juniors off these tools are right about how people learn. I'm not going to pretend otherwise to win an argument.
What changed is that the choosing became skippable. And the compiler never checked it anyway. A compiler tells you your syntax is wrong. It has never once told anyone their approach was wrong. Only a person does that.
Which means I don't need the ban's advocates to be wrong about the mechanism. I need them to notice their lever isn't attached to it. Stop a junior using Claude Code and you've restored the typing.
There's a randomised trial that speaks to this, and the interesting part isn't the headline. Researchers at Wharton and Penn ran nearly a thousand Turkish high-school students through three arms: a plain GPT-4 interface, a version built to give hints instead of answers, and no AI at all. Access improved practice performance a lot - 48% for the plain interface, 127% for the tutor version. Then they took it away for the exam, and the plain-interface students scored 17% worse than students who'd never had it. The authors' word for what happened is crutch.
Now look at which arm got hurt. The plain one: practice problems, no consequence, answer on request. The arm with a forcing function built into it didn't take the damage - the authors report the harm was largely mitigated there. That's the shape I'm arguing for. It's teenagers doing maths, and I'm not claiming it tells you what your engineers will do. It's an inference.
The strongest version of the case against me
Here's the objection at the strength its advocates would put it, because if my version of their argument doesn't match what you know it to be, you shouldn't trust anything else on this page:
Debugging survives, sure. But it's the wrong debugging. You learn to choose by debugging your own choice and finding out it was wrong. Debugging Claude's choice teaches you that Claude's choice was wrong - that's a fact about Claude, not a model of your own judgment. The loop needs the choice and the consequence to both be yours.
That's a real argument and it lands exactly where I live.
My answer is that the gate converts an acceptance into a choice. Asked "what approach didn't you take, and why," a junior who accepted has to either own the acceptance or admit there wasn't a choice - and being sent back to consider the alternatives is choosing, late.
Two things I won't hide. The junior defends an approach that's already sitting in their head, and whether that's choosing or rationalising is genuinely open. And no research settles it. The closest evidence is a meta-analysis of self-explanation - 69 effect sizes across 64 reports, g \= 0.55 - whose coded task types include studying worked problems, which is learners explaining a solution they didn't produce. That's structurally your junior's exact situation. It's also maths and text rather than code, and those worked examples were correct and built to teach, which a model's draft is neither. Structure fits. Domain doesn't.
So this whole post rests on one contestable claim, and I'd rather you knew that here than worked it out in three years. What rides on it is worse than "my advice gets weaker," and I'll show you the arithmetic further down rather than asking you to take it on trust.
Here's how you check it yourself, and it costs a minute. Take a junior's AI-assisted pull request and ask them for the approach they didn't take, and why. If they can't name one, they didn't choose. They accepted.
The senior's keyboard is where the choosing actually went
The choosing didn't leave at your junior's keyboard. It left at your seniors'.
Every time a senior self-serves a task they'd once have handed down - because it's twenty minutes with Claude and forty to explain it to someone - a junior doesn't get to choose. That's the rung. Nobody removed it on purpose. It just stopped being worth the forty minutes, one task at a time.
I didn't go to university for this. I came up through an apprenticeship, which is a system whose entire content is that somebody hands you work they could have finished faster themselves.
The economists have a version of this and I want to be careful how far I take it. Stanford's Digital Economy Lab tracked ADP payroll records across millions of workers and found early-career workers aged 22 to 25 in the most AI-exposed occupations saw a 16% relative decline in employment, while employment for experienced workers in the same occupations held up. For software developers aged 22 to 25, employment was down nearly 20% from its late-2022 peak by September 2025. Read the authors' own caution before you do anything with that: "we caution that the facts we document may in part be influenced by factors other than generative AI."
The hypothesis is the part that matters, and it's theirs, not mine. They propose that "AI may be automating the codifiable, checkable tasks that historically justified entry-level headcount, while complementing the judgment-, client-, and process-intensive tasks performed by experienced workers." Their phrase for the shape of it is task substitution at the apprentice margin.
One level down, on your team: the tasks that used to be the ladder were the codifiable, checkable ones. Those are the ones your seniors now finish themselves in twenty minutes. That's my inference from their hypothesis, which they offer as one explanation among several, and you should hold it that loosely.
This is why the recommendation has two halves. Run only the first - the junior uses the tool and defends it at the gate - while your seniors keep self-serving everything with a decision in it, and your junior defends lockfile bumps. The gate grips on nothing, because nothing with a choice in it ever reached them.
I'm not making a prediction about the industry, and I'm not interested in the version of this where AI comes for everyone's job. This is narrower and more boring. Your list of people who can review the changes that can take down production is short, and it's short because seniors take years to make. That's a fact about your staffing in 2029, and it gets decided by what your seniors do with their next twenty-minute task.
Five ways teams handle this, put the way their advocates put them
You've already considered all five. I've written each one the way its own advocate would.
| Option | The case for it, at full strength | My honest counter | Verdict |
|---|---|---|---|
| Let them, no extra rules | Your juniors will work with these tools for their whole careers. Every hour spent not using them trains them for a job that doesn't exist. And a ceremony you apply to their pull requests and nobody else's tells them exactly that you don't trust them, which is how you lose the ones worth keeping. Treat them like engineers and they'll become engineers | Nothing comes back. They accept an approach, ship it, and never find out. Neither do you, until it matters | Right if you rent your seniors rather than form them |
| Ban it | You can't review code you couldn't have written. Judgment is built by writing it badly and paying for it, and there's no other input - no craft in history has produced an expert who never did the work. Hand a junior a tool that supplies the answer and you get someone whose PRs look fine for three years and who then can't hold a Tier 1 review, and you won't find out until an incident | I agree with your mechanism. Your lever isn't attached to it, and you're training them for an exam they'll never sit | Right when a junior has no repertoire yet. Conceded below, without hedging |
| No AI for the first six months | Clean. Bounded. Natural end date. Maps onto a probation period you already run. And unlike what I'm recommending, it costs zero senior attention - the one thing this whole series says you've run out of | Nothing that isn't already an argument against the ban, plus one: you're using a clock to approximate something you already measure directly | Nowhere, and I have to earn that |
| Let them, and make them answer for it | What I'm recommending | Two costs, priced below. Both real | Right for a team with a review gate, that forms its own seniors, and that can't take a senior off delivery for a quarter |
| Pair them with a senior | A senior beside a junior on real work is the highest-bandwidth way anyone has moved judgment from one head to another. It's how apprenticeship has worked for centuries. It catches the wrong approach at the moment it's chosen, and it does the thing I'm recommending - a person asking why - about twenty times as often. It also does both my halves at once, and the expensive one for free, because a senior pairing on real work is handing the choosing down | Price. That's the entire counter | Beats what I'm recommending, outright, under the fourth condition below |
Why "no AI for six months" wins nowhere, and why that isn't a technicality
I could beat this one cheaply by saying you can't enforce it. That would be unfair and its advocates would be right to reject it, because some of you genuinely can. If you're air-gapped, or you're in a locked-down regulated environment with no egress, you can enforce a tool restriction perfectly well.
The real answer is that a time gate is the ban with a clock, and the ban always beats it. In every version of this argument, including the ones where I'm wrong.
The clock is the wrong variable. Six months of tests-and-docs work forms nothing. Six weeks of real work with a reviewer who actually asks why forms a lot. Time is a proxy for exposure and it's a bad one, and you already have a better one, because you already route work by what it can break. Anyone who wants what the time gate wants should run the ban and use a readiness test instead of a date.
And its cost point is real and I don't have an answer to it. A time gate costs zero senior attention. What I'm recommending costs two kinds. All I've got against that is a claim that it buys nothing, which is the crux again.
Why making them answer for it wins, and the one thing it loses to
Making them answer for it wins because a person asking why is the only forcing function that puts the choosing back - and it loses to a person doing that twenty times more often. That's the whole ruling, and it holds for a team with a review gate, forming its own seniors, that can't take a senior off delivery for a quarter.
Two arguments do the work, and a third does less than it looks.
The choosing. Against just letting them get on with it, the case is the crux above: nothing comes back, so the choosing never happens and nobody finds out, including them.
The exam isn't unaided. This is the argument against the ban and the time gate, and it's the only one on this page with a citation rather than an inference underneath it. Look again at that Wharton experiment's design: "the third part is an unassisted evaluation, where students take a closed-book, closed-laptop exam." They took the tool away for the test. That's why the result reads the way it does. Your engineers don't sit a closed-laptop exam. They keep the tool, on the job, for their entire careers. The Stanford authors put the same point from the other end, in a footnote: "one of the practical skills more likely to be learned on the job than in university computer science classes may be how to use AI for software development."
The best counter to that is a good one. The exam is unaided the moment the model is confidently wrong, and that's the moment I'm forming them for. True. And spotting a confidently wrong model is a verification act, not a production act, and the gate is where verification gets trained. Which routes straight back through the crux, as everything here does.
You can't enforce a ban anyway, as a standing policy. Scope that carefully, because I concede the opposite below. Across a team's normal work, a tool restriction keys on "did AI write this," and nobody can verify that answer - not you, and not the junior, who genuinely can't draw the line between a completion they accepted, a block they edited, and a function they wrote after reading a suggestion. In a bounded, observed setting, twenty minutes in a room with the laptop closed, it's perfectly enforceable. Which is exactly why the second condition below prescribes an exercise and not a policy.
Notice what that argument does and doesn't do. It says a ban won't work. It doesn't say a ban is wrong. If you believe judgment requires production experience, an enforceability objection just tells you to try harder at enforcing.
What actually discriminates, and against what
| Let them, no rules | Ban it | Six-month gate | Pair them | |
|---|---|---|---|---|
| A person asking why puts the choosing back | This wins | Tie. They want the choosing back too, they've just picked a lever that doesn't reach it | Tie, same reason | Pairing wins outright |
| The exam isn't unaided | Tie. They keep the tool too | This wins | This wins | Tie. Pairing keeps the tool too |
| Enforceable as standing policy | Tie | Constrains it. Doesn't beat it | Constrains it. Doesn't beat it | Tie |
Read that honestly and four things fall out.
Against unrestricted use, only the first row does any work, and that row is the crux. So my case against just letting them get on with it is exactly as strong as a claim I've already told you is unsettled. If the crux is false, that row goes, the ban supplies precisely what the crux would then say judgment requires, and the second row dies too, because its defence routes through the crux. If the crux is false, the ban is right. Not "let them get on with it." The option I've spent this page arguing against.
Against the ban and the time gate, only the second row wins. Notice its shape: it says they've aimed at the wrong exam, not that they've misunderstood learning. On learning we agree.
Against pairing, I lose the first row and tie the rest. I beat it on cost. That's all.
And the rows aren't independent supports. Three arguments, one claim underneath. If it goes, they all go together.
What pairing costs, and why I'm not pretending it's a different thing
A senior beside a junior is more of exactly what I'm selling you. Same forcing function, higher bandwidth: a person asking why at the moment of the choice, twenty times a week, instead of once, a day later, at the gate.
I'm not going to tell you it's a different product. It's more of the same one. Call it a senior for every two or three juniors, against roughly ten minutes of senior attention per pull request. A factor of about twenty, and that factor is my entire remaining argument.
There's a reading where each of my events bites harder than each of theirs, because being wrong and finding out teaches more than being redirected before you were wrong. It cuts my way and I'm not going to lean on it. It's a speculation, there's no evidence for it on your population, and twenty times the events is twenty times the events.
One thing this isn't: the sensible middle. "Let them use it but make them explain themselves" has the shape of a compromise, and it isn't one, because it restricts nothing. It takes nothing from the ban. It adds a requirement from a different direction entirely.
Four times I'd tell you to do something else
Four situations where I'd tell you not to run this, and three of them end in you buying nothing at all. Each has a test you can run this week.
You don't have a review gate
The test: name the last pull request a reviewer returned unread. Not "could have." Returned. If you can't name one, you don't have a gate. You have a queue with a stamp on it.
The first half of this is a gate mechanism. Without a gate it collapses into unrestricted use with a paperwork step nobody enforces, which is the plain-interface arm of that experiment, and you already know how that went.
Build the gate. It's free, it's yours, and it has nothing to do with juniors. And do the second half anyway while you're at it: stop letting your seniors self-serve the ladder. That half needs a decision, not infrastructure, and you can make it today.
The junior has no repertoire yet
This is where the people telling you to keep juniors off these tools are right, and I'd rather concede it plainly than route around it.
The test: hand them a hundred-line diff in your codebase, in a language they know, with a bug in it. Sit with them, laptop closed, twenty minutes. Can they find it?
If not, the gate has nothing to grip on. Ask a junior in week three what approach they didn't take and there's no answer, because there were never any candidates in their head. Ask why this approach and you'll get a plausible reply generated by the same model that wrote the code, and neither of you can tell. The gate can't separate understanding from fluency when there's no repertoire underneath it. A ban at least produces a signal you can't fake: the code works or it doesn't, and they fixed it or they didn't.
The route isn't a ban, because a ban still doesn't survive contact with a team's normal week. It's exercises where the signal can't be faked, in a room, tool closed, until they can pass that test. Costs you twenty minutes.
You rent your seniors, you don't form them
The test: when did you last promote someone onto your Tier 1 reviewer list from inside, rather than hire one?
If the answer is that you hire them, do the first option on that table. Let them use it, skip the ceremony, don't spend the senior minutes. The exam argument still holds for you and the enforceability point still holds. What disappears is the payoff - you're not collecting the formation, so there's nothing to buy with the ten minutes. And the second half goes with it, because handing the choosing down is a slower way to ship a task and you're not collecting on that either.
Then be honest about what renting is. It works. At the level of your team, this quarter, it genuinely works. It's a collective action problem and the Stanford authors name it in those terms, citing Becker: "inefficient incentives to train entry-level workers who may move firms." The obvious counter is that the pool you're renting from is thinning. That's an argument about 2030 and it isn't an argument about your Q3, and I won't dress it up as a reason to spend senior hours now.
Buy nothing. Not from me, not from anyone.
You can take a senior off delivery for a quarter
The test: could you take one senior off delivery for a quarter and have them do nothing but sit with two juniors on real work? Not name the senior - naming is free, and anyone with a headcount spreadsheet does it in four seconds. Take them off.
If yes, pair them, and don't do what I'm recommending. Pairing beats it on the only row that discriminates in my favour, ties the rest, and does both halves at once, the expensive one for free, because a senior pairing on real work is handing the choosing down by construction. I need two clauses and a cost table to get where pairing arrives by existing.
This is the one that routes toward what I sell, so give it the suspicion it deserves. When you can genuinely fund it, all I've got left is that I'm cheaper, and you've just told me cost isn't your binding constraint.
And don't split the difference with "the gate thing now, pairing later." If you can fund pairing today, doing the cheap version first is a quarter of senior attention spent on a smaller dose of what you're about to buy properly.
What this costs you, and the way it fools you
Claude can write the annotation.
Paste the diff in, ask for the inline comments, and you get a pull request that looks exactly like the thing I'm recommending is working. If the annotation were the mechanism, this would be theatre with extra steps, and I'd be selling you a process that makes your dashboard prettier and your engineers no better.
So the annotation isn't the load-bearing part. The reviewer's live question is. A model will write a plausible comment on a hunk. It cannot sit in your review and answer the follow-up.
Which means this only works if a human actually asks. Adopt the annotation requirement without the question and you've bought the paperwork and none of the formation.
Here's the count, and it's a baseline rather than a verdict. Look at your last ten junior pull requests and count the ones where a reviewer asked a question the author visibly had to think about. Most teams are at zero or one, and that's the point - the number is supposed to move. It measures where you are, not whether you can. At ten minutes a pull request, zero becomes ten inside a month. That's different from the fourth condition above, which is a capacity test: you either have the slack or you don't, and no amount of intent changes it.
The two costs, and the second is the one that will get argued about:
| The half | What it costs | Who pays |
|---|---|---|
| Make them answer for it | Roughly ten minutes of senior attention per junior pull request | The reviewer |
| Keep handing the choosing down | A senior who hands a task down ships it slower than a senior who does it themselves. This quarter. Measurably | The senior with the least time. The same person this whole series says is your constraint |
I'm not going to soften the second one. Handing down a task you could finish yourself in twenty minutes is a decision to be slower now so that in two years your Tier 1 reviewer list has three names on it instead of one. The payback is measured in years and you will not feel it this quarter. That's the trade, and it isn't a comfortable one.
And the uncomfortable line, which costs me nothing and you something: junior formation was never free. The reason your team stopped doing it isn't that AI replaced it. It's that you stopped paying for it, and AI turned up with an excuse.
What to do on Monday
Run the four tests in order. Stop at the first yes.
- Name the last pull request a reviewer returned unread. Can't? Build the gate first, and stop your seniors self-serving the ladder while you do, because that half is free.
- Hand your greenest junior a hundred-line diff with a bug in it, laptop closed, twenty minutes. Can't find it? Exercises where the signal can't be faked, until they can.
- When did you last promote onto your Tier 1 list from inside? Never? Let them use it, skip the ceremony, buy nothing.
- Could you take a senior off delivery for a quarter? Yes? Pair them, and don't do the rest of this.
Four noes and you're who this is for. Let them use Claude Code. Route their work by what it can break, like everyone else's. Make them annotate what isn't obvious, then have somebody ask them what they didn't do and why. And when a senior is about to self-serve a twenty-minute task with a real decision in it, make them hand it down instead.
All of that is free and none of it needs me.
The one place I'd spend money is the fourth condition: you've got the slack to pair someone and nobody internal to spare. An outside senior is capacity you don't have, and 1-on-1 Claude Code mentoring is one way to buy it. Read the honest counter first. An outside mentor isn't in your codebase and isn't in your review queue. They do the asking-why half of pairing and none of the knows-your-monorepo half, and that's a real gap, not a quibble.
And if you're in the first condition, with no gate, don't buy a mentor. You'd be paying to form judgment in a system with nowhere to apply it.
Questions engineering leaders ask
Should we let junior engineers use Claude Code?
Yes, if you've got a review gate to make them answer at, you're forming your own seniors rather than hiring them, and you can't take a senior off delivery for a quarter to pair with them instead. Then it's the same terms as everyone else - routed by what the change can break - plus one requirement: they annotate what isn't obvious, and a reviewer asks what approach they didn't take and why. Restricting the tool restores the typing, and typing never formed anyone.
Our juniors say AI is teaching them. Our seniors say they don't understand what they submit. Who's right?
Nobody's measuring, which is the actual finding. BairesDev's Q2 2026 survey of 1,569 developers across 77 countries reports that 85% of juniors say AI tools improved their understanding, while 16% of seniors say juniors fully understand the code they submit. Three things before you quote that. BairesDev sells staff augmentation, so its commercial interest points at "juniors can't be trusted," which is the same direction as mine. Its methodology is available on request rather than published. And the 16% is what seniors think about juniors, not a measurement of what juniors know. The gap is what it's good for: both sides are confident and neither is checking.
What does team-wide AI use do to our senior engineers' time?
It quietly cancels the delegation. A task worth handing to a junior is often twenty minutes with Claude and forty to explain, so it stops getting handed down, one task at a time, with nobody deciding. That's the mechanism worth watching, and reversing it costs you nothing except shipping speed this quarter. The review-load question is a different one with a different answer, and it's about capacity rather than formation.
Should we just hire seniors instead of forming them?
If that's already your model, yes, and this post isn't for you. The argument that you'll regret it is an argument about the market in 2030, not your next two quarters, and I'm not going to sell you senior hours on the strength of a forecast. The catch is a real collective action problem - the Stanford paper names "inefficient incentives to train entry-level workers who may move firms." Everyone renting and nobody forming is how the pool empties. True, and still not a reason for your team to go first.
What if we want something more structured than any of this?
MentorCruise runs a 90-day sprint for teams at mentorcruise.com/teams/ai. I'm not going to describe what's in it here, because this post is an argument you're meant to be able to check and that would be me selling in the middle of it. Go and read the page and make your own call. For the individual engineers on your team who want the personal discipline underneath all this, the validation playbook is written for them rather than for you.