Insights
/
Systems & Execution
/
17 AI Maturity Frameworks Compared: What Each One Measures
Systems & Execution

17 AI Maturity Frameworks Compared: What Each One Measures

Alina Vasile
|
Updated
Aug 2026
|
16
 min read
Share
CONTENTS

Key takeaways

  • There is no single "best" AI maturity model. Enterprise frameworks like McKinsey's and PwC's assume a budget and a bench of specialists that most companies don't have, while SME-built tools like Singapore's AIRI or the EU's DMA tool assume a size and geography that most enterprises have outgrown.
  • The spread between digital and AI leaders and laggards has widened 60% over three years, and leaders now outperform laggards by two to six times on total shareholder returns, according to McKinsey's research.
  • Every framework in this list, regardless of who built it, converges on the same warning: a strong average score can hide a fatal weak spot in data quality or governance, and scaling on top of that gap is where most of the damage happens.
  • Most of these models were built for enterprises with dedicated transformation teams. A handful, mainly the SME-specific ones, are usable by a business without one.

Seventeen organizations have each decided the world needs their own way of scoring how "AI-mature" a business is. Gartner has one. So does McKinsey, BCG, PwC, IBM, AWS, MIT, the OECD, and the government of Singapore. Some of these are rigorous, evidence-based diagnostics built on survey data. Others are thinly documented marketing artifacts attached to a consulting pitch or a cloud contract.

This is a walk through all seventeen: what each one measures, what it's good for, and where it falls apart. Then a comparison across all of them, including the dimension most of them get wrong, and a look at where a much simpler self-assessment fits among a set of tools mostly built for organizations ten times the size of the ones reading this.

Two numbers are worth holding onto going in. AI adoption could add roughly $13 trillion to global economic output by 2030, about 1.2% of additional annual productivity growth, more than double what the steam engine or industrial robots added in their respective eras. And most estimates put the share of AI pilots that never make it past a proof of concept somewhere between 70% and 95%. Maturity frameworks exist to close that gap between the upside and the failure rate. Whether any given one does that for your business depends entirely on which one you pick.

Enterprise-class frameworks

These ten are built for organizations with multiple business units, dedicated data and platform teams, and enough budget to run a formal audit. They tend to be the most rigorously researched of the seventeen, and also the least usable without outside help.

Gartner AI Maturity Model

Gartner's model evaluates a business across seven capability pillars, strategy, governance, data, product, engineering, operating models, and culture, through a guided questionnaire that outputs a visual heat map of where current capability sits against a target. It maps organizations to one of five stages, from Awareness to Transformational, and, notably, pushes users to map out process friction and manual workarounds before recommending any automation, on the theory that automating a broken process just makes the mess move faster.

Strengths Limitations
A genuinely strategic roadmap that ties technical infrastructure to cultural and organizational change, backed by a large, continuously updated research library. Heavily qualitative. There’s no open scoring algorithm or public benchmarking database, so two assessors can reasonably reach different scores.

McKinsey "Rewired"

McKinsey's assessment, built from its "Rewired" research and QuantumBlack's client work, evaluates six capabilities: a transformation roadmap tied to real value, a bench of specialist talent, a fast-moving operating model, a flexible distributed technology environment, data embedded across the business, and a genuine adoption and scaling motion. Getting there involves auditing budgets, governance policies, and org charts, followed by structured interviews with six to ten senior leaders. McKinsey's own research is blunt about the state of play: almost every company is investing in AI, but only about 1% consider themselves fully mature.

Strengths Limitations
Grounded in real client evidence rather than self-reported scores, and it treats workflow redesign, not tool purchasing, as the actual driver of financial impact. Resource-intensive. It assumes weeks of senior leadership time and a team capable of running structured interviews and document audits.

BCG "10-20-70" and Deploy-Reshape-Invent

BCG runs two related frameworks. The first is the 10-20-70 rule: in BCG's research, only 10% of AI value creation comes from the algorithm itself, 20% from technical infrastructure and data, and 70% from people, process redesign, and change management. It sorts companies into four brackets, from Stagnating to Future-Built, based on capability across 53 digital and AI dimensions.

The second is Deploy-Reshape-Invent, three value plays that sequence how a business should expect returns:

  • Deploy targets quick productivity gains (roughly 10-15%) through automating existing tasks,
  • Reshape targets deeper efficiency gains (30-50%) by re-engineering whole workflows,
  • Invent is about building genuinely new, AI-native products and business models.

BCG's own research suggests that for most companies, over 90% of eventual AI value comes from Reshape and Invent, not Deploy.

Strengths Limitations
Directs executive attention and budget toward the change management work that actually produces value, and gives a clear sequencing logic for where returns show up over time. Lighter on deep technical infrastructure audits than some peers, which can mean under-diagnosing a genuinely broken data pipeline.

PwC AI Readiness Assessment

PwC's model produces a precise 0 to 100% score across twelve operational domains, mapped to five bands: AI-Vulnerable (0-29%), AI-Aware (30-48%), AI-Developing (49-66%), AI-Advanced (67-83%), and AI-Driven Leader (84-100%). The score comes from senior leader interviews benchmarked against a database of more than 500 organizations, and it's specifically designed to surface the gap between what a CEO and a CTO each believe about the company's own readiness, which PwC reports is often a double-digit-point spread.

Strengths Limitations
An unusually detailed, auditable scorecard that ties capability directly to peer benchmarks and P&L impact, not just technical readiness. Full execution depends on a PwC advisory engagement, which makes it hard to run as an independent, repeatable self-check.

MITRE AI Maturity Model

MITRE's model spans six pillars (ethical and responsible use, strategy and resources, organization, technology enablers, data, and performance) broken into twenty specific metrics, scored through a free multiple-choice questionnaire called the Organizational Assessment Tool. It maps to five levels, Initial through Optimized, and deliberately doesn't require every organization to hit Level 5 on every dimension. A defense contractor and a mid-market retailer can set very different target profiles and both be "done."

Strengths Limitations
Entirely open-source, vendor-neutral, and free under a royalty-free license. It also weighs ethics and fairness as seriously as engineering capability. Twenty metrics across six pillars is a lot of ground to cover without a dedicated process-improvement person driving the exercise.

CMMI AI Maturity (AIM) Model

CMMI Institute, the standards body behind decades of software process maturity work, launched AIM in mid-2026, extending its existing 31 practice areas with AI-specific guidance across eight domains: data, development, people, safety, security, services, suppliers, and virtual. Getting scored requires a formal appraisal from a certified Lead Appraiser, and the model includes crosswalks that map directly onto ISO/IEC 42001, ISO 23053, ISO 23894, and ISO 31000, letting one appraisal double as evidence for multiple compliance regimes at once.

Strengths Limitations
Audit-grade rigor with genuine international standards recognition, and the ISO crosswalks cut down on duplicated compliance work. Expensive by design: practitioner training, licensing, and a formal Lead Appraiser engagement are all required, not optional.

IBM GenAI Adoption Model

IBM's original AI Ladder (Collect, Organize, Analyze, Infuse) has evolved into a generative-AI-specific model that scores seven dimensions on a 0-3 scale, rolling up into Silver, Gold, or Platinum. For generative AI specifically, IBM lays out five phases, from simply consuming an off-the-shelf model through to building and running custom models securely across environments, with explicit guidance on things like Model Context Protocol integration and real-time bias detection along the way.

Strengths Limitations
Genuinely deep technical guidance, down to specific MLOps, RAG pipeline, and agent orchestration patterns, and it ties capability directly to compute cost management. Built around IBM’s own watsonx and hybrid cloud stack, so its usefulness as neutral advice thins out fast if you’re not on that platform.

AWS Generative AI Maturity Model

AWS extends its existing Cloud Adoption Framework into four levels: Envision, Experiment, Launch, and Scale, evaluated across five pillars (business, people, governance, platform, operations). It's explicitly built to move a team from mapping theoretical use cases to running production workloads with SLAs, and eventually to a self-service marketplace of reusable internal components.

Strengths Limitations
Highly actionable, concrete engineering guidance on integration patterns, automated pipelines, and cost controls. Every recommendation points toward AWS-native services like Bedrock, SageMaker, and Q Business, which limits its value if you run multi-cloud or on-premise.

Accenture AI Maturity Framework

Accenture's research sorts companies into four groups rather than a strict ladder: AI Experimenters, Builders, Innovators, and Achievers, based on foundational versus differentiating capability. In its widely cited study, 63% of companies land in the Experimenter group, with a maturity score around 29 out of 100, while a smaller group of Achievers scores roughly double that and correlates with meaningfully higher revenue growth.

Strengths Limitations
Backed by large-scale benchmark research, and the four-group framing is simple enough to put in front of a board without translation. Mostly a lead-in to a consulting engagement rather than an open, self-serve diagnostic with a public scoring rubric.

MIT Sloan CISR Enterprise AI Maturity Model

Built by MIT's Center for Information Systems Research from a survey of 721 companies plus follow-up executive interviews, this model maps four stages: Experiment and Prepare, Build Pilots and Capabilities, Industrialize, and Become AI Future Ready. The distribution is telling: 28% of companies sit at Stage 1, 34% at Stage 2, and only around 7% ever reach the top stage. Financial performance tracks the stages closely, with Stage 1 and 2 companies underperforming their industry averages and Stage 3 and 4 companies well ahead of them.

Strengths Limitations
Grounded in genuine academic survey data with a direct, measurable link to financial performance, not just a capability narrative. Reads more like an academic diagnostic than a step-by-step roadmap you can execute against directly.

SME and public-sector frameworks

Smaller businesses and public bodies run into a different problem than enterprises: not too little rigor, but frameworks that assume resources they don't have. These four were built with that constraint in mind.

EU EDIH Digital Maturity Assessment (DMA) Tool

Built under the EU's Digital Europe Programme and delivered free through the European Digital Innovation Hubs network, the DMA evaluates six dimensions, including digital strategy, human-centric digitalization, data management, and AI and automation. Rather than a self-scored questionnaire, an EDIH representative runs a structured interview and maps the business's actual practices to predefined behavioral indicators, producing a 0-100 score benchmarked against similar-sized regional and EU peers. Completing it can also unlock direct digitization grants for micro and small enterprises.

Strengths Limitations
The expert-interview format removes the self-scoring bias that undermines most questionnaires, and it’s directly tied to real funding. It’s a broad digital-maturity tool first. The AI-specific questions are a small slice of a much wider assessment.

Singapore AI Readiness Index (AIRI)

Built by AI Singapore, AIRI scores organizations across five pillars and twelve dimensions on a 1-4 scale, landing them in one of four bands: AI Unaware, AI Aware, AI Ready, and AI Competent. It's designed to be run without outside consultants, and results map directly to Singapore's state-funded training and engineering support programs, closing what the index's own research calls the SME "execution gap."

Strengths Limitations
Genuinely practical for a small team to self-administer, with a clear path into public funding once you’ve got a score. Calibrated specifically to Singapore and ASEAN market conditions and regulation, so it travels poorly outside that region.

Dynamic Adaptive Maturity Assessment Model (DAMA-AHP)

An academic model built for industrial SMEs, DAMA-AHP scores 66 distinct elements across six dimensions, from people and expertise to production processes. What sets it apart is the Analytic Hierarchy Process it borrows from decision science, which lets a company weight each dimension and element according to its own strategic priorities rather than accepting a fixed, generic weighting.

Strengths Limitations
Highly customizable, and it avoids the trap of a static, one-size-fits-all scoring model. The weighting math is genuinely non-trivial. Setting it up well without academic or consulting support is a real barrier.

OECD Framework for the Classification of AI Systems

Less a maturity ladder than a classification lens, the OECD's framework maps an AI system's characteristics against five values-based principles: inclusive growth, human-centered fairness, transparency, robustness and safety, and accountability. It's built for policymakers and risk assessors as much as business leaders, tracing data handling, traceability, and human oversight across an AI system's full lifecycle.

Strengths Limitations
Highly adaptable across sectors and pairs well with sector-specific rules, with strong grounding in ethics, transparency, and public trust. No financial ROI scorecard, no engineering roadmap. It answers “is this system classified correctly,” not “how mature are we.”

Niche and compliance-focused frameworks

Forrester AI Readiness Model

Forrester's model checks whether the baseline structures, initial setup, technical tooling, organizational preparation, are in place before a company launches a large-scale AI initiative. It functions less as an ongoing maturity ladder and more as a one-time, pre-project readiness check, which is exactly what it's good for: catching a gap before it becomes an expensive mistake. Much of Forrester's own methodology sits behind paid research, so public documentation on it is thinner than most of the frameworks above.

Strengths Limitations
Effective as a pre-launch gap check, surfacing capability holes before they turn into an expensive mid-project surprise. A checklist for a single point in time, not a framework for tracking maturity as it develops.

Salesforce AI Ethics Maturity Model

Salesforce's model tracks the maturity of an organization's responsible-AI practices specifically, across four stages: Ad hoc, Organized and repeatable, Managed and sustainable, and Optimized and innovative. It's narrower by design than most of the frameworks here, focused entirely on ethical design, bias review, and workforce alignment around responsible AI use rather than technical or financial capability.

Strengths Limitations
A focused, well-structured tool for organizations where brand trust and ethical design are the priority, not a side concern. Says nothing about infrastructure, MLOps, or financial performance. It’s one dimension, not a full picture.

Element AI Maturity Model

Element AI's model, from a whitepaper published before the company was acquired by ServiceNow in 2020, evaluates five dimensions, strategy, data, technology, people, and governance, across five levels from Exploring to Transforming. It's worth knowing this one is a legacy artifact: with Element AI no longer operating independently, nobody is actively updating it for multi-agent systems the way an active vendor would.

Strengths Limitations
An evenly balanced, multi-dimensional view that covers every standard readiness area without over-weighting any one of them. No modern platform integration guidance or automated monitoring content, unsurprising given how long it’s gone unmaintained.

How the seventeen compare

Line them up and three clusters emerge, and which cluster a framework sits in tells you more about whether it'll work for your business than any individual feature does.

Cluster Frameworks What it actually requires of you
Enterprise / consulting-grade Gartner, McKinsey, BCG, PwC, CMMI AIM, Accenture Weeks of senior leadership time, formal interviews or audits, often an outside advisory engagement to complete.
Hyperscaler / engineering-grade AWS, IBM A technical team already building on that vendor’s cloud, willing to accept platform-specific recommendations.
SME / public-sector EU DMA, Singapore AIRI, DAMA-AHP A single interview or self-assessment session, often tied to a specific region’s funding or training programs.
Open / academic MITRE, MIT Sloan CISR Genuine engagement with a lengthy questionnaire, but no cost and no vendor lock-in.
Niche / single-issue Forrester, Salesforce, Element AI, OECD Useful for one specific question (pre-launch readiness, ethics, classification) rather than a full maturity picture.

What almost none of them share is a scoring philosophy that guards against the most common way these assessments go wrong: a simple average across dimensions. If a business scores well on strategy and technology but poorly on governance and data quality, a plain average can still produce a healthy-looking overall number, one that clears the business to scale. Several of the more rigorous frameworks here, McKinsey's among them, instead anchor the overall score to the weakest dimensions rather than the mean, on the logic that a business with brilliant algorithms and no data governance isn't moderately mature, it's one incident away from a real problem. It's a principle worth borrowing even if you're not running any of these formally: don't let a strong score in one column paper over a failing one in another.

What's different about AI maturity versus ordinary IT maturity

A fair question, especially for a team that's already run a CMMI or TOGAF assessment before: why does AI need its own maturity model at all? The honest answer is that AI systems behave in ways traditional IT governance was never built to catch.

Dimension Traditional IT (CMMI / TOGAF) AI-specific (CMMI AIM / MITRE)
Data governance Schema validation, relational integrity, batch processing. Data lineage, vector embeddings, feature stores, real-time low-latency streaming.
Lifecycle management Software versioning, standard CI/CD pipelines. Model drift detection, continuous retraining, dataset lineage, model registries.
System behavior Deterministic, rule-based logic with predictable outputs. Non-deterministic outputs; measures bias, hallucination rate, toxic language.
Security Identity access management, firewalls, endpoint protection. Prompt injection defenses, model poisoning, adversarial input detection.
Accountability System owners, uptime SLAs. Explainability, interpretability, defined human-in-the-loop review points.

That last row is the one worth sitting with longest. A traditional IT system either works or it doesn't, and when it fails, you can usually trace exactly why. An AI system can produce a confident, plausible, wrong answer with no error message at all, which is why every serious framework in this list treats a defined human-in-the-loop checkpoint as a maturity requirement, not an optional extra.

The specific way this bites smaller businesses

Enterprise maturity models spend most of their attention on scaling risk. For a smaller business, the more immediate risk is continuity. It's a familiar pattern: a business "adopts AI," but what that means is one or two employees using consumer AI tools in browser tabs, with no shared documentation of how those tools are configured or what they're being trusted to do. If that person leaves, the knowledge usually leaves with them, and if there's no defined process for catching a wrong or questionable AI output before it reaches a client, nobody's checking until something has already gone out the door. This both a capability gap as well as a business continuity risk wearing an AI-shaped costume.

This is exactly the gap the SME-specific frameworks above are built to close, and it's a large part of why we built the process redesign work we do at Orbflo around documentation and handoff points first, rather than treating a slick AI tool as evidence that the underlying process is sound.

Designing for what comes after: agentic readiness

Every framework above was built to evaluate businesses using AI as a tool a person prompts. The frontier that's already arriving is businesses running networks of autonomous agents that coordinate entire processes with far less human prompting in the loop, and most of these seventeen models haven't caught up to it yet. A handful, mainly AWS's and Microsoft's newer Agentic Adoption Model, have started building in three specific checks worth watching for regardless of which framework you use: whether your systems support modern integration standards like the Model Context Protocol so agents can connect to your tools and databases directly, whether your data can be delivered as a clean, low-latency stream rather than an overnight batch job, and whether governance checks (bias detection, drift monitoring, prompt-injection defenses) are built directly into your deployment pipeline instead of happening as a manual review after the fact. A framework that doesn't ask any of these three questions yet is evaluating yesterday's version of the problem.

Picking one instead of drowning in seventeen

The honest starting point is scale, not preference. A large, regulated enterprise chasing audit-grade compliance has real reasons to reach for CMMI AIM or PwC's assessment. A company mainly trying to fix its change-management and workforce problem is better served by BCG's 10-20-70 rule or McKinsey's approach. A cloud-native engineering team is going to get more out of AWS's or IBM's model than out of anything built for a boardroom. And a small or mid-sized business without a transformation budget is generally better off with something built for that constraint, like Singapore's AIRI or the EU's DMA tool, assuming it operates in one of those regions.

Outside those two regions, there isn't an obvious equivalent, which is part of why we built the AI Operating System Scorecard the way we did: a short, self-administered diagnostic across nine dimensions rather than a twelve-domain enterprise audit. It doesn't compete with McKinsey's or PwC's depth, and it isn't trying to. It's closer in spirit to AIRI or the DMA tool: something a business without a dedicated transformation team can run itself, in an afternoon, to get a rough read on where the real gaps are before deciding whether a heavier framework is worth the investment.

Whichever one you use, two habits carry across all seventeen. Score conservatively rather than averaging, so a strong result in one column can't quietly cover for a weak one in another.

If you want a starting number

The AI Operating System Scorecard takes about the same amount of time as reading one section of this article, and gives you a baseline read on decision authority, data readiness, and process clarity, three of the dimensions every framework above treats as foundational, before you decide whether a heavier assessment is worth running.

Further reading & sources

  1. McKinsey, "Rewired and running ahead: Digital and AI leaders are leaving the rest behind," mckinsey.com
  2. McKinsey Global Institute, "Notes from the AI frontier: Modeling the impact of AI on the world economy," mckinsey.com
  3. Gartner, "Gartner AI Maturity Model and AI Roadmap Toolkit," gartner.com
  4. BCG, "AI @ Scale," bcg.com; BCG, "AI Transformation Is a Workforce Transformation," bcg.com
  5. PwC, "AI Readiness Assessment for business leaders," pwc.com
  6. MITRE, "MITRE AI Maturity Model," aimaturitymodel.mitre.org; MITRE, "The MITRE AI Maturity Model and Organizational Assessment Tool Guide," mitre.org
  7. CMMI Institute, "CMMI Institute Launches New AI Maturity (AIM) Model," cmmiinstitute.com; CMMI Institute, "CMMI AIM: Artificial Intelligence Maturity," cmmiinstitute.com
  8. IBM, "The IBM Maturity Model for GenAI Adoption: A 5-Phase Framework," ibm.com
  9. AWS Prescriptive Guidance, "Overview of the generative AI maturity model," docs.aws.amazon.com
  10. Accenture, "More Than 60% of Companies Are Only Experimenting with AI," newsroom.accenture.com
  11. MIT Sloan, "What's your company's AI maturity level?," mitsloan.mit.edu; MIT CISR, "New MIT CISR research finds companies with advanced enterprise AI outpace industry peers," mitsloan.mit.edu
  12. European Digital Innovation Hubs Network, "DMA Tool," european-digital-innovation-hubs.ec.europa.eu; European Commission, "Commission unveils new tool to help SMEs self-assess their digital maturity," digital-strategy.ec.europa.eu
  13. AI Singapore, "AI Readiness Index (AIRI)," aisingapore.org
  14. MDPI, "A Dynamic Assessment of Digital Maturity in Industrial SMEs (DAMA-AHP)," mdpi.com
  15. OECD, "AI principles," oecd.org; OneTrust, "Approaching the OECD Framework for the Classification of AI Systems," onetrust.com
  16. Salesforce, "AI Ethics Maturity Model," salesforceairesearch.com
  17. Element AI, referenced via Towards Data Science, "Inside AI Maturity Model," towardsdatascience.com
  18. Microsoft Learn, "Introduction to the Agentic AI adoption maturity model," learn.microsoft.com
  19. Orbflo, "AI Adoption Failure: Why 70-95% of Pilots Never Scale and What Can You Do Differently," orbflo.com/insights
  20. Orbflo, "How to Redesign Business Processes for AI: A Practical Blueprint for Teams," orbflo.com/insights

Frquently Asked Questions

Is there one AI maturity framework that works for every business?

No, and treating any of these seventeen as universal is the fastest way to waste the exercise. Enterprise frameworks assume a transformation budget and specialist staff most companies don't have. Hyperscaler models assume you're building on that specific cloud. SME frameworks assume a size and geography that a mid-sized enterprise has usually outgrown. Match the framework to your actual constraints before you match it to its reputation.

Can a small business realistically use an enterprise-grade model like McKinsey's or PwC's?

Technically, yes, but it's rarely a good use of the effort. These models are built around weeks of leadership interviews, formal document audits, and in PwC's case, an actual advisory engagement to complete properly. A smaller business is almost always better served starting with a lighter, self-administered diagnostic and reserving a heavier framework for after it has a specific, well-scoped problem worth that level of rigor.

What's the biggest mistake companies make when using any of these frameworks?

Averaging dimension scores instead of anchoring to the weakest one. A business can look mature on paper, strong strategy, strong technology, while its data governance or human-review process is genuinely broken, and an averaged score will hide that. Most of the frameworks that have caused real damage when followed too literally are the ones used this way. The fix costs nothing: score conservatively, and fix the weakest dimension before scaling anything built on top of it.

AUTHOR
Alina Vasile

Founder of Orbflo.

Exploring how AI-native companies can become faster, leaner, and more effective than ever before.

START WITH A DIAGNOSIS

Find out exactly where your business is losing speed and leverage

Decision Authority Icon
Decision Authority
AI Adoption Icon
AI Adoption
Process clarity icon
Process Clarity
Strategic Direction icon
Strategic Direction
Team Capability icon
Team Capability
AI Integration icon
Output
AI Integration icon
AI Integration
Coordination icon
Coordination
background gradientbackground gradient
Data readiness icon
Data Readiness

The AI Operating System Scorecard is a diagnostic tool that measures whether your business is structurally built to make AI compound, across nine dimensions including how decisions get made, how clearly your processes are defined and how your team is using and integrating AI.

The output is a clear view of where your biggest leverage gaps are and where to focus first.

Get your free diagnosis
background gradient grid floor

Get the weekly
AI Operating System Brief

One practical AI operating-system insight bi-weekly.

No fluff, no spam.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
background gradient