You can run all of them this afternoon, on data you already have and probably haven't looked at. Not one of them needs a vendor, a consultant, or me.
I started out writing code and now I run MentorCruise with a team of five, so here's the part I'd want said to me first, before the meeting. A tool got bought. A system didn't get built around it. These five checks tell you which piece of the system is missing, and three of the five end with nothing for you to buy.
The short version
- Start with the document that authorised the spend. If there's no number in it that was measured before the tools arrived, you don't have a failed rollout. You have an unmeasurable one, and it may have worked.
- Check seat utilisation before you conclude anything about people. If almost nobody's using it, that's a fit problem or a permissions problem, and both are cheaper than anything you could buy.
- If usage is high and merged pull requests per engineer are flat, generation was never your constraint. One company tracked over two years went to 2.09 times its baseline only after it mandated a doubling. Adoption alone didn't do it, and no training will.
- If throughput went up and the wait for review went up with it, the bottleneck relocated onto the people who were already your constraint. Route by blast radius. That's free, and it's most of the fix.
- If throughput went up and the wait for review went down while incidents climbed, your reviewers are approving without reading. Automation bias "cannot be prevented by training or instructions" [Parasuraman and Manzey, Human Factors, 2010]. Change the load, not the people.
Why the tools and the engineers are the wrong place to look first
The tools and the engineers are the wrong place to look first because both halves are measurably working. That's not a compliment to anyone. It's the reason the real answer is somewhere less obvious, and it's why this post is a tree rather than an opinion.
Take the tools. DORA's 2025 report found 90% of respondents using AI at work and more than 80% of them believing it had increased their productivity. Jellyfish, working with OpenAI across millions of pull requests at more than 500 companies, found that going from zero to full AI adoption correlated with a 113% increase in pull request throughput. The generation half of the machine works. Whatever went wrong, "it can't write code" isn't it.
Take your engineers. Same 90%. And in DX's June 2026 data from more than 400 companies, developers now self-report that 51.9% of their code is AI-authored. DX are careful about that number and so am I: it's self-reported, and they flag "potential for bias in both directions." Read it as roughly how much of the work people are handing to AI, not as a literal line count. Either way, it isn't resistance.
Now the honest part, and I'd rather you get it here than catch me at it later. Two of the five checks below can rule for the tool explanation. Check 2 can tell you the tool genuinely doesn't fit your work. Check 3 can tell you the tool was fine and you aimed it at a constraint you didn't have. So I'm not claiming the tools are never the answer. I'm claiming they're the wrong place to look first, because looking there first is how you end up buying a second tool to fix the first one.
What's actually going on is duller and more fixable. Somebody bought a tool. Nobody built the system around it: the measurement, the approval route, the question of what the constraint even was, or the gate. Which piece is missing is what the tree is for. I'm not going to guess it for you from here.
The five checks, in order
Run them in order and stop at the first one that matches. They're ordered cheapest-first, and each answer forecloses the ones below it.
| What you observed | What it means | What to do | Do you buy anything |
|---|---|---|---|
| No number anywhere that predates the tools | You can't tell yet. It may have worked | Write down four numbers. Re-ask in 90 days | No |
| Few seats active, and it's uniform across the team | The tool doesn't fit the work, or nobody was told they could use it | Cancel the seats, or fix the approval route | No |
| A few heavy users, everyone else at zero | A capability question, and a real one | That's a buying decision, and it's a different post | Not yet, and maybe never |
| Usage high, merged PRs per engineer flat | Generation was never your constraint | Find the real one. It isn't in this post | No, and not from anyone selling AI enablement |
| Throughput up, wait for review up, 3+ people approve your riskiest changes | The bottleneck relocated onto your seniors | Route by blast radius | No. It's free |
| Throughput up, wait for review up, 1 or 2 people approve your riskiest changes | A staffing problem wearing a process problem's clothes | Hire, or grow a second senior | Not training |
| Throughput up, wait for review down, incidents up | The gate went nominal | Change the load, not the people | Not training |
Check 1 - is there a number, and did it exist before the tools did?
Open the document that authorised the spend and look for a number with a date on it, measured before the tools arrived. Not the number you were promised. The number that was true. If it isn't there, you don't have a failed rollout. You have an unmeasurable one, and it may well have worked.
This is the check people skip because the answer feels insulting. Run it anyway, because the alternative is arguing about a number nobody wrote down.
Here's why your instinct isn't admissible either. METR ran a randomised trial in 2025 with 16 experienced open-source developers across 246 real issues on repositories they already knew. The developers expected AI to speed them up by 24%. Measured against the clock, they took 19% longer. Afterwards, having just been slowed down, they still believed AI had sped them up by 20%.
Don't carry that 19% around as a fact about your team, and don't let anyone else do it either. It's a sample of 16, METR explicitly disclaim it as evidence that AI fails to speed up most developers, and their own follow-up puts the number at -18% for returning developers and -4% for newly recruited ones, with confidence intervals that cross zero in both cases. They also found 30% to 50% of developers were declining to submit tasks because they didn't want to do them without AI, which means they're "systematically missing tasks which have high expected uplift." METR themselves think developers are likely more sped up now, in early 2026, than their early-2025 estimates suggest. The effect size is unstable and I'm not going to pretend otherwise in a post about measurement.
The durable finding is the other one. In that trial, people's sense of their own speed pointed the opposite way to the stopwatch. Self-report was the thing that broke. So "the team says it didn't help" and "the exec says it didn't land" are the same class of evidence, and it's the class that failed.
So write down four numbers this week. Seat utilisation. Merged pull requests per engineer per month. Median time from pull request opened to first review comment. Reverts or incidents per month. They're free, they're already sitting in systems you own, and they're the four this tree runs on. Re-ask in 90 days.
Don't buy anything. Don't cancel anything either. You don't know enough yet to do either one, and that's the finding.
(How to build a real measurement practice around those numbers - which metrics survive AI, how to baseline properly, what the research supports - is a bigger question than this post should answer, and there's a companion post that does.)
Check 2 - did the seats get used?
Open your AI vendor's admin console and look at the last 30 days. What share of the seats you're paying for had any activity in the last week? You're already paying for this number and it's on the billing page. Then look at the shape of it, not the size. The shape is the diagnosis.
If it's uniformly low across the team, ask one person who isn't using it a single question: what did you try it on, and what happened?
If the answer is about the tool being confidently wrong about your codebase, your stack, or your conventions, then the tool doesn't fit the work. Cancel the seats. That's the whole recommendation, and it's the cheapest thing you'll do this quarter.
If the answer is "I never got a licence", or "it's not on our SSO", or "I wasn't sure I was allowed to point it at our repo", you don't have a tool problem or a people problem. You have an approval route that's slower than your engineers' deadlines, and that's free to fix. There's a companion post on the policy side of this - the default-permit list, the response SLA, the amnesty - and it's the one to read next.
If it's bimodal - a few heavy users, and a long tail sitting at zero - that's a capability question, and it's a fair one. It's also a buying decision rather than a diagnosis, and it has its own post, which arbitrates five options honestly and tells a good share of readers to buy nothing.
One thing I want to be precise about, because it's an inference and not a finding. DORA's 90% is a measurement. "Therefore your low number is about fit or route rather than about your people" is me reasoning from it, and DORA doesn't say that. It's what I'd bet on, and you can check it in ten minutes with the question above, which is more than I can offer for most inferences.
Check 3 - did output actually go up?
Pull merged pull requests per engineer per month for the six months before the tools arrived, and for last month. GitHub Insights gives you this for nothing. If seat usage is high and that number is flat, generation was never your constraint, and this is the most expensive misdiagnosis in the post to get wrong in either direction.
You bought throughput for a team that wasn't blocked on throughput. More AI capability does nothing here. Training on AI does nothing here.
Two pieces of evidence say this branch is real rather than theoretical, and I'm laying them out because a diagnostic that only ever routes toward the author's own product is worth exactly nothing.
The first followed the same developers for over two years. Researchers tracked 802 developers across 196,212 pull requests between January 2024 and April 2026 at a company that "committed to doubling merged pull requests per engineer since mid-2025." It worked: "per-capita throughput eventually doubled, reaching 2.09x the pre-mandate baseline in April 2026." Read that again with the dates in it. A company with full AI adoption needed an explicit executive mandate to move throughput. Adoption alone didn't do it.
The second is DORA, describing your situation before you got to it: "AI doesn't fix a team; it amplifies what's already there. Strong teams use AI to become even better and more efficient. Struggling teams will find that AI only highlights and intensifies their existing problems." A team getting no throughput out of full adoption is a case DORA already predicted.
Go and find the actual constraint, and know in advance that it's probably requirements churn - the work changes shape before it ships - or friction in your environments and deploys. I'm not going to diagnose either one, because neither is what this post is about and neither has anything to do with AI. Here's the question that tells you, and you can answer it this afternoon: pull the last ten changes you shipped and, for each one, name the longest single wait between someone deciding to do it and it being live in production. If that wait is rarely "waiting for review", your constraint isn't in this post, and no AI purchase from anyone touches it.
Don't buy anything. Not from an AI enablement vendor, not from a training provider, and not from me.
Check 4 - did the wait for review get longer?
Take median time from pull request opened to first review comment, before the tools and now. Then count the distinct people who approved a change touching auth, payments, data deletion, or permissions in the last quarter. If throughput grew and that wait grew with it, the bottleneck didn't disappear when you bought the tools. It relocated onto the people who were already your constraint.
This is the one people mean when they say the rollout didn't deliver, and it has a name now. The 802-developer study is titled "AI Writes Faster Than Humans Can Review", and its finding is one sentence: "per-reviewer load roughly doubled and automated review overtook human review, while merge and revert rates held steady." The team doubled its output. The reading did not double. It got reassigned to machines.
Two caveats, because I'm about to ask you to change how your team works. It's one company, adoption wasn't randomly assigned, and the authors read their own evidence carefully rather than as clean causation. And "merge and revert rates held steady" is genuinely good news that cuts against the alarmist version of this. Nothing fell over there.
What didn't move is your reviewers' ceiling. SmartBear's guidance, drawn from their Cisco study, is that "developers should review no more than 200 to 400 lines of code (LOC) at a time" and "do not review for more than 60 minutes at a time." That describes how a person reads, not who typed it, and no model release has moved it.
So a bounded review against a doubled queue can't be fixed inside the review. That argument belongs to the companion standard and it makes it properly; what matters for your diagnosis is narrower. The lever is which pull requests get the senior hour, and it moves before anyone opens a diff.
Now the count you did. It matters more than the wait.
Three or more people approve your riskiest changes: you have a routing problem, and it's free to fix. Your seniors are spending attention uniformly across a queue that isn't uniform. Sort by what breaks if the change is wrong rather than by who or what wrote it, and put the required gate on the top tier only. There's a companion post that ships the standard as a copy-pasteable artifact, and it costs nothing.
One or two people approve your riskiest changes: stop. That's a staffing problem wearing a process problem's clothes, and no document fixes it, no training fixes it, and I don't have anything to sell you that fixes it either. Your options are to hire, or to grow a second person to that bar, which takes about as long as it takes and can't be compressed by buying something. Route the queue anyway - it'll help - but don't let anyone tell you the routing is the answer.
Check 5 - did the wait for review get shorter?
If the wait for review got shorter while throughput climbed, and your incidents went up alongside it, your reviewers are approving without reading. Take the same two numbers as check 4 - throughput, and median time to first review - and put reverts, hotfixes, and incidents next to them. This is the check whose dashboard looks like success, which is why it's last and why it's the one that gets missed.
The gate is nominal. It exists in your process diagram and it doesn't exist in anyone's attention.
Here's the part that costs me money to write. Your instinct now is to train the reviewers, and the research says don't bother. Parasuraman and Manzey reviewed the human-factors literature on automation in Human Factors in 2010, and their findings are unusually blunt for a review paper. Automation complacency "occurs under conditions of multiple-task load, when manual tasks compete with the automated task for the operator's attention." It is "found in both naive and expert participants and cannot be overcome with simple practice." Automation bias "occurs in both naive and expert participants, cannot be prevented by training or instructions, and can affect decision making in individuals as well as in teams."
Read that last clause again, slowly, if you were about to book something. Cannot be prevented by training or instructions. Not "is difficult to prevent." Your seniors aren't being lazy and they aren't undertrained. They're under load, and this is what people do under load. It's the most heavily replicated finding in this entire post and it says the obvious purchase doesn't work. The policy companion post builds its whole accountability clause on that mechanism, and it's the better place to read it out in full.
What makes it expensive is what AI output looks like. In the 2025 Stack Overflow Developer Survey, 66% of developers said their biggest frustration was "AI solutions that are almost right, but not quite". Almost-right is exactly what a fast approval misses. It passes every check a machine can run and it reads fine at speed.
So change the load, not the people. Cap the size. Cap the session. Make automated review comment rather than approve on anything with real blast radius. Make the author annotate before a human reads it. Same companion post as check 4, same artifact, still free.
What this diagnostic can't tell you
This diagnostic can't tell you whether the spend was worth it, whether a rollout you can't measure actually worked, or which of the five checks you'll land on. Four limits, and the third is the one I'd want on the record.
It can't tell you whether the spend was worth it. That's a different question with a different method and this post doesn't touch it.
If you landed on check 1, it can't tell you whether the rollout worked. Only that you can't know yet. Those aren't the same thing, and treating "unmeasured" as "failed" is how a working rollout gets cancelled in a budget meeting.
It can't tell you which check you'll land on, and I'm not going to guess. Every number in this post is about the teams in the studies, not about yours. I don't know how common each of these is, nobody has published it, and if I told you "most teams land on check 4" I'd be making it up in a direction that happens to suit me. The tree routes. That's its entire job, and a diagnostic that told you the answer before you ran it would be an opinion piece with a table in it.
And all five checks assume the tools actually reached a team. If the licences got bought and never rolled out, that isn't a failed rollout. It's a purchase, and nothing here diagnoses it.
Where this stops being a diagnosis
If you landed on check 1, 2, or 3, we're done, and there's nothing here for you to buy. That's most of what this post does. Write down the four numbers, cancel the seats, fix the approval route, or go find the constraint that was there before AI arrived and will be there after. None of that needs me.
If you landed on check 4 with one or two people approving your riskiest changes, we're also done. Hire, or grow someone. I'd be lying if I said a purchase compresses that.
If you landed on check 4 or 5 and you haven't routed by blast radius yet, do that first. It's free, it's most of the fix, and buying anything before you've tried it means you won't know what you bought.
That leaves one situation, and it's specific enough that you can tell whether you're in it. You ran check 4 or 5. You took the routing. And you're still stuck in one place: what "correct" actually looks like on your auth path, or your payment flow, or the subsystem that pages someone at 2am, lives in two people's heads and nowhere else. There's nothing to route by, because the standard was never written down.
Then here's the test. Could you take those two people off the review queue for a fortnight and have them write it down? Not schedule it. Take them off. If you can, do that - it's free, they know your codebase better than any outsider will, and it's better than anything I could sell you.
If you genuinely can't, that's a real question about what to do next, and it's a buying decision rather than a diagnosis. There's a post that arbitrates it across five options - workshop, internal training, online course, consultant, hire - and it recommends one and then names the situations where it loses, including the ones where the answer is to spend nothing. I'd rather you read that than hear it from me at the end of a post you came to for a diagnosis.
One thing I won't do, even here. None of this fixes automation bias. Nothing does - the research above is about as settled as this field gets. The routing is what changes the load, and the routing is free. What a purchase can do is get a standard out of two people's heads, and whether that's worth buying is exactly the question the other post exists to answer honestly.
Questions you'll get from whoever signed the invoice
How do I tell my exec the rollout didn't fail, it just isn't measurable?
Show them the document that authorised the spend and point at the absence. If there's no pre-tool number in it with a date, then "it didn't deliver" and "it delivered and we can't see it" are both consistent with everything you know, and picking the pessimistic one isn't rigour. Then hand over the four numbers you're now recording and a date 90 days out. Executives take "I'll have an answer on the 14th of October" better than most engineering leaders expect. What they don't take well is a second opinion that's also a guess.
Our AI tools are being used and nothing got faster. Did we waste the money?
Possibly not, and the distinction matters for what you do next. You bought a capability aimed at a constraint you didn't have, which is a targeting error, not a purchasing error, and the tools will still be useful once the actual bottleneck moves. What would waste the money is buying a second thing in the same direction. Run check 3's question first: for your last ten shipped changes, where did the longest wait actually sit?
Should we cancel the licences?
Yes, if your seat utilisation is uniformly low and the people not using it tell you the tool is wrong about your codebase. That's a fit answer and paying for it monthly won't change it. No, if the non-users tell you they never got access, weren't sure they were allowed, or couldn't get it approved for the repo that matters. Those cost you nothing to fix and cancelling would be solving the wrong problem loudly.
Our dashboards look fine but incidents are up. Is that the AI?
It's more likely your review gate went nominal, and AI made that visible without causing it. Check the direction of your approval times: if they got faster while pull request volume climbed, your reviewers are keeping up by reading less, and merge rates and revert rates were never going to show you that. It's the same mechanism DORA describes when they say AI "amplifies what's already there" - the gate was probably always thin, and nothing was pressing on it before.
Do we need to retrain the team?
Almost certainly not, and I say that as someone who sells training. If you're on check 5, the finding is that automation bias "cannot be prevented by training or instructions, and can affect decision making in individuals as well as in teams" [Parasuraman and Manzey, Human Factors, 2010]. Your engineers know how to review code. They're under a load that doubled without anyone deciding it should. If you're on check 2 with a bimodal usage curve, there's a genuine capability gap and it's worth a real answer - but that's a buying decision with five options and it deserves more than a line in an FAQ.