Assessment

The AI SDLC Maturity Model

This is not a scale of how much AI your teams use. It is a scale of how safely and effectively that use is integrated into the way you build software. Score yourself against five levels below, then take the twelve-question self-assessment to find where you actually stand today.

Why Redesign Is the Real Gap

McKinsey's State of AI series puts the shape of the problem in numbers. As of its November 2025 edition, 88% of organizations use AI in at least one business function, yet only about 6% qualify as high performers, meaning 5% or more of enterprise EBIT is attributable to AI. High performers are nearly three times more likely to have fundamentally redesigned their workflows around AI rather than bolting a tool onto the old ones.

McKinsey's August 2026 edition holds the same pattern: 80% of respondents report individual productivity gains, but only 37% report any enterprise-level EBIT impact, and the share of true high performers is still about 6%. Roughly 75% of high performers have redesigned their workflows, versus about 25% of everyone else.

Usage is universal. Redesign is rare. Redesign is where the return is. The five levels below describe what redesign looks like as a maturing operating model, not as a count of tools installed.

The Five Levels

Level 0. No AI

What it looks like

No AI in the dev workflow. Everything is manual.

Risks

Falling behind the adoption curve, a growing productivity disadvantage, and trouble attracting talent.

Next step

Assess readiness, find low-risk entry points, and capture the baseline now while it is still clean.

Level 1. Ad hoc

What it looks like

Individuals use tools on their own initiative, often on personal accounts. No shared standards, no visibility, wide variance in results.

Risks

Shadow AI, IP and data exposure, quality that tracks individual prompting skill, and no measurement at all.

Next step

An approved tool list, data classification rules, minimum review standards for AI-generated code, and company-managed accounts with training opt-out.

Level 2. Guardrails introduced

What it looks like

Approved tools and usage guidelines exist. Basic data classification, review standards on paper, some usage logging, and repo instruction files in the main repos.

Risks

Guardrails exist but adoption is uneven, teams interpret the same rules differently, and measurement stays thin.

Next step

Introduce real measurement - adoption patterns, review quality, bug escape rates - against the pre-AI baseline.

Level 3. Measured and standardized

What it looks like

Consistent cross-team standards, tracked metrics for adoption, quality, and risk, regular review cycles, and AI in onboarding. Risk-tiered review is in place; agents run under their own identity with budgets.

Risks

Metrics that capture activity rather than quality, standards rigid enough to slow real work, and measurement overhead of its own.

Next step

Formalize the policy, automate compliance checks, and build the learn loop where every escaped defect becomes a new rule.

Level 4. Governed and optimized

What it looks like

Governance embedded in the engineering workflow. Automated compliance and quality checks on AI output, improvement loops driven by measured outcomes, and clear ownership. Evidence is generated at release; agents are delegated bounded work classes unattended.

Risks

Governance overhead that eats the speed benefit, complacency, and new capability outrunning the framework built for the last one.

Next step

Reassess continuously, adapt to new capability as it appears, and share what you learn.

Score Your Own Maturity

Answer yes or no to each question below by checking the box. The count updates live, along with the level it maps to and that level's next step.

0 of 12

Level 1 - Ad hoc

Next step: an approved tool list, data classification rules, minimum review standards for AI-generated code, and company-managed accounts with training opt-out.

Yes answers Level
0 to 3 Level 1 - Ad hoc
4 to 6 Level 2 - Guardrails introduced
7 to 9 Level 3 - Measured and standardized
10 to 12 Level 4 - Governed and optimized

Where It Fits in AIDLC

Maturity is not owned by one phase - it is a lens on all five. Teams typically start experimenting with AI during Analyze and Ideate, and their level climbs as Develop and Launch add guardrails, budgets, and measurement. Curate is where the learn loop that separates Level 3 from Level 4 actually lives - it is the phase that turns an escaped defect into a permanent rule.

Frequently Asked Questions

Level 0 is no AI in the workflow at all. Level 1 is ad hoc, individual use with no shared standards. Level 2 introduces approved tools, basic data rules, and review standards on paper. Level 3 adds consistent cross-team metrics, risk-tiered review, and scoped agent identities. Level 4 embeds governance directly in the engineering workflow with automated checks and continuous improvement loops.

Adoption counts how many people use AI tools. Maturity measures how safely and effectively that use is integrated: whether guardrails, measurement, and governance exist and actually work. McKinsey's own numbers show the gap - most organizations report AI usage, but only a small share qualify as high performers with redesigned workflows and measured enterprise impact.

The model treats reassessment as ongoing work, not an annual audit. Level 3 already folds measurement and escaped-defect review back into the standard on a schedule, and Level 4 makes continuous reassessment - adapting to new capability and sharing what was learned - the operating norm itself, not a periodic check.

Know Your Level. Now Move Up It.

A score is a starting point. The rollout playbook and the governance pillars turn it into a plan.