Over the last couple of years in my role as a software development lead, I have been working on pivoting from our traditional SDLC process to one that effectively leverages AI tools such as Claude Code/Cursor for increased productivity. We wanted better quality and shorter delivery times. We got some of both, and in a few places we got slower.
In this blog post, I am going to break down what processes we tried, what failed and what did we learn. If your team is being asked to "adopt AI" and nobody has said what that means in practice, I hope this saves you a few months.
Some context
I lead three teams, ten engineers in total, building a product for enterprise security teams. Following two key facts were central to our AI adoption.
First, our customers are enterprises, and mistakes in our product are expensive for them and for us. That makes us conservative by default. Second, many of our engineers are senior and have been with the company for years.
Process Before AI SDLC
Our now legacy software development process consisted of the following phases:
- Ideation and inception: Product owners and UX designers worked with our enterprise clients to work out what a feature needed to do.
- System Design: Once we understood what to build, we assigned a full-stack team to produce a high-level design, usually led by a software architect working with the developers.
- Building: Once the design was approved and accepted, the engineering team would then break down the work and start building the software.
- Continuous deployment: Finished work shipped in a two-week window, behind a feature flag, until it was ready for general availability.
What we changed
We decided to keep the four phases and add AI to each one. We didn't want to redesign the process and introduce new tools at the same time. If we changed both and delivery got worse, we wouldn't know which change caused it
In each phase, AI now produces the first draft and people review it. The details:
- AI-driven ideation: Product owners and UX designers use Claude's research and design abilities to draft customer journeys and mockups. We show these to current and prospective customers, and their feedback becomes the requirements.
- AI-driven system design: Our senior architects built a Claude skills harness for system design. A skill is a packaged set of instructions, conventions, and examples the model loads for a specific task. Ours encodes how we design systems: our architecture patterns, the questions we expect a design to answer, and the format we want.
Getting it right took several rounds of feedback, and it took longer than we expected. Now Claude Code produces the first high-level design from a set of requirements, and the architects and the team review it. - AI-driven building: We built skills for common development tasks, such as adding a frontend component or making a backend change. Developers run the harness to generate the change, then review it for correctness, quality, and fit with the codebase.
- Continuous Deployment: Mostly unchanged so far. We're exploring how AI could help with observability and monitoring, but we haven't changed anything in production yet..
What we learned
1. Reviewing everything the AI produces is exhausting
Because we're conservative, we required the people running the AI tools (we call them AI operators) to review every output in full. In practice this caused a lot of cognitive fatigue.
The reason is that reviewing and writing are different kinds of work. When you write a design or a piece of code yourself, you build an understanding of it as you go. When you review AI output, you have to rebuild that understanding from a finished artifact that someone else wrote.
AI outputs made this harder. They were dense and long, and they weren't organised for a person to read. A developer could spend more energy understanding a generated design doc than they'd have spent writing a shorter one.
2. AI can be overkill for simple tasks
Our senior engineers can ship a small, well-defined change quickly. When we sent those changes through the AI process, delivery slowed down.
The AI process has a fixed cost on every task: setting up the context, running the harness, and reviewing the output. For a large or unfamiliar task, that cost is small next to the time saved. For a change an experienced engineer could make in fifteen minutes, the fixed cost is most of the work.
We had treated the AI process as the default path for everything. It's more useful as a tool that engineers choose when the task is big enough, or unfamiliar enough, to pay back the overhead.
3. Knowing what to review is a skill
AI produced very different outputs across the phases: UX mockups, design docs, and code. Each needs a different kind of review.
Experienced engineers know where problems tend to hide. In a design, that might be data flows, trust boundaries, and failure modes. In code, it might be whether the tests actually check the behaviour or just exercise it. People who were newer to their role, or to the team, found it hard to know what to look at. They either reviewed everything with the same attention, which made the fatigue worse, or they missed the parts that mattered.
We found we had to write down, at a high level, the key things to check at each review step. That knowledge used to live in senior engineers' heads and got passed on slowly through code review comments. With AI producing more output than ever, it needed to be written down.
What we're changing next
This was a v1 attempt to try and implement an AI focused SDLC. In our next iteration, we are planning to improve the AI harness to provide easily human consumable artifacts and to start documenting quality attributes that a human in the loop needs to evaluate the AI generated artifacts.
This was version one. In the next version we're focusing on two things.
First, we're changing the harness so it produces artifacts written for people to read. That means a short summary first, a list of the decisions made and why, what changed, and the risks the reviewer should check. The goal is to lower the cost of review without lowering how carefully we review.
Second, we're documenting the quality attributes a reviewer should check for each type of artifact. This gives newer engineers a starting point and gives everyone a shared standard to review against.
If you're trying this with your team
Here's what I'd tell a team starting out, based on the above:
- Keep your existing process and add AI to it one phase at a time. You'll be able to see what each change actually did.
- Budget for building the harness. Skills that encode your team's conventions take several rounds to get right, and that time is the real cost of adoption.
- Plan for review from day one. If people review everything and the output isn't built for reading, you'll trade writing time for fatigue.
- Let people skip AI for small changes. Use it where the task is big enough to pay back the overhead.
- Write down what reviewers should check. It helps newer engineers most, and they're the ones most likely to trust AI output too easily.
We're still learning, and I'll write about the next version when we have results. If your team is going through the same shift and you want to talk it through, you can find me on MentorCruise.