VIII.5The Security Governancegovern
OWASP AIMA - running an AI maturity assessment
A maturity score is only worth having if two assessors would give the same organization the same number, and if the result tells someone what to fix first. This chapter is how to get both out of OWASP AIMA, the open model for measuring how well an organization builds, secures and runs AI.
VIII.4 · ISO/IEC 42001, verification & maturity places AIMA in the OWASP standards chain and offers a quick four-rung ladder for a first conversation. This chapter is the working method behind it: what the model contains, how to score it without fooling yourself, and what the output looks like in a client engagement and in an internal program.
What AIMA is, and what it is not
The OWASP AI Maturity Assessment (AIMA) is an open, community-built maturity model for AI. It adapts OWASP SAMM, OWASP’s long-established software assurance maturity model, to the AI lifecycle, and extends it to data provenance, model robustness, privacy, fairness and transparency. Version 1.0 was released on 11 August 2025 under project co-leads Matteo Meucci and Philippe Schrettenbrunner. It ships as a PDF and an Excel toolkit (v1.0.1), is licensed CC BY-SA 4.0, and is an OWASP Incubator project. Everything is in the project repository.
The project describes V1.0 as coming with “detailed criteria and a worksheet for internal use or third-party evaluations”. Those are the two uses this chapter covers. Two boundaries keep expectations honest:
- It measures the program, not a system. AIMA asks whether the organization reliably does the things that keep AI safe: strategy, policy, data handling, threat modeling, testing, monitoring, incident response. Whether one model or agent is actually secure is a question for testing (VI.4 · AI red-team playbook) and for verification against requirements such as AISVS.
- It is not a certification. Nobody certifies an organization “AIMA Level 2”. The certifiable standard is ISO/IEC 42001 (VIII.4 · ISO/IEC 42001, verification & maturity). AIMA tells you how far you are from being ready for that audit.
The model on one page
AIMA groups its practices into eight domains, three practices each. The right-hand column is the plain-language question an assessor is really asking.
| Domain | Its three practices | What you are really asking |
|---|---|---|
| Responsible AI Principles | Ethical & Societal Impact · Transparency & Explainability · Fairness & Bias | Do you assess who the AI affects, can you explain what it does, and do you find and fix bias? |
| Governance | Strategy & Metrics · Policy & Compliance · Education & Awareness | Is there a strategy with owners and measures, a policy people follow, and training that reaches the people building AI? |
| Data Management | Data Quality & Integrity · Data Governance & Accountability · Data Training | Is data quality controlled, is someone accountable for the data, and is training data collected, labeled and licensed properly? |
| Privacy | Data Minimization & Purpose Limitation · Privacy by Design & Default · User Control & Transparency | Do you collect only what the use case needs, build privacy in by default, and give people control over their data? |
| Design | Threat Assessment · Security Architecture · Security Requirements | Are AI systems threat-modeled and built to written, verified security requirements? |
| Implementation | Secure Build · Secure Deployment · Defect Management | Are build and release controlled for AI systems, and are flaws tracked to closure? |
| Verification | Security Testing · Requirement-based Testing · Architecture Assessment | Do you test AI systems, against their requirements and their architecture, on a schedule? |
| Operations | Incident Management · Event Management · Operational Management | Would you detect an AI incident, respond to it with a plan, and run the system under control? |
Each practice has three maturity levels and two streams. Stream A, which AIMA calls Create & Promote, is about doing the activity. Stream B, Measure & Improve, is about measuring it and acting on what you measure. In plain terms, Level 1 means the activity happens informally, Level 2 means it is defined, documented and repeatable, and Level 3 means it is embedded across the lifecycle and continuously improved.
One question sits at each level of each stream. This is the full grid for Threat Assessment, taken from the V1.0 toolkit:
| Level | Stream A | Stream B |
|---|---|---|
| 1 | Is there basic awareness or informal identification of threats specific to AI systems? | Are informal threat mitigation strategies occasionally discussed or implemented? |
| 2 | Are threats systematically identified and documented for AI systems? | Are documented mitigation strategies developed and periodically reviewed? |
| 3 | Is comprehensive threat assessment consistently performed and integrated across AI lifecycle? | Are proactive and comprehensive mitigation strategies continuously implemented and refined? |
Six questions per practice, 24 practices, 144 questions. A lightweight pass is interviews and document review, not an audit.
Two scoring methods, and why you must pick one
The document and the toolkit score differently, and the difference is large enough to change a board slide.
The document method (gated, SAMM-style). Answer each question yes or no. A “yes” to every Level 1 question earns Level 1; a “yes” to every Level 1 and Level 2 question earns Level 2, and so on. Partial progress into the next level adds a ”+”, so a practice with every Level 1 criterion met and one of the two Level 2 criteria met scores 1+. Scores run 0, 1, 2, 3, with ”+” variants; 0 means no appreciable activity.
The toolkit method (averaged). The Excel toolkit asks for a number per question: 0 (“no maturity”), 1 (“initial maturity”), 2 (“not full maturity”) or 3 (“maturity”). The practice score is the sum of the six answers divided by six, so it comes out as a decimal. A Results sheet lists all 24 practice scores, and a separate sheet charts them. Each question also has an interview-note column for recording what you were told and what you saw.
Here is one practice, Strategy & Metrics, scored both ways from the same interview. The evidence: the strategy exists as an unapproved slide deck, the chatbot’s containment rate and escalation volume are reviewed monthly, nothing is measured on risk, and nothing links the AI strategy to the enterprise strategy.
| Question | Document (yes/no) | Toolkit (0-3) |
|---|---|---|
| Level 1, Stream A | yes | 3 |
| Level 1, Stream B | yes | 2 |
| Level 2, Stream A | no | 1 |
| Level 2, Stream B | yes | 2 |
| Level 3, Stream A | no | 0 |
| Level 3, Stream B | no | 0 |
| Practice score | 1+ | 8 / 6 = 1.33 |
The two agree roughly here. They diverge when a team does advanced work without the basics: an organization with an automated Level 3 practice but no documented Level 2 process stays capped by the gated method, while the average pulls it up.
- Use the gated method when the score will be compared year on year, reported to a board or regulator, or produced by a third party. One advanced team cannot inflate it.
- Use the average for a fast internal baseline and for tracking a trend between formal assessments.
- Either way, state the method in the report, and never compare a gated score with an averaged one.
Running an assessment, step by step
- Set the scope. The whole AI program, one business unit, or one AI product. When an activity is handled outside your scope, for example by a group-level privacy office, record that rather than scoring it “no”. AIMA’s own guidance warns against marking such items not applicable too early.
- Choose the depth. A lightweight assessment uses the worksheets, interviews and document review to give a provisional score. A detailed assessment adds verification: you sample the evidence to confirm the activity really happens, “not just paper compliance”. A practical mix is lightweight across all eight domains and detailed for the few practices leadership cares most about.
- Choose the scoring method and write it down, with how you read Level 1 and how you split overlapping evidence.
- Map questions to people. Each domain has natural owners, listed in the table below. Send the questions ahead so people bring artifacts, not opinions.
- Interview and collect evidence. Record what you were told and what you saw in the interview-note column. A policy that exists but nobody has read is a Level 1 answer, not Level 2.
- Score conservatively, then calibrate. No evidence means the lower level. Where two assessors are involved, run a calibration session on a few practices before finalizing, so “systematic” means the same thing across domains.
- Set targets from risk, not from the scores. Not every practice needs to reach 3. A company whose only AI is an internal writing assistant can accept Level 1 on Fairness & Bias; a lender using AI in credit decisions cannot.
- Turn the gaps into a roadmap. The gap between current score and target, per practice, is the backlog. Sequence it, give each item an owner and a quarter, and date the reassessment.
| Domain | Who usually answers |
|---|---|
| Responsible AI Principles | Product owner of each AI system, legal or ethics lead, data science lead |
| Governance | CISO or AI governance lead, risk and compliance, HR or learning for training |
| Data Management | Data owner or data governance lead, ML engineering lead |
| Privacy | Data protection officer, product owner |
| Design | Security architect, ML or platform engineering lead |
| Implementation | Platform or DevSecOps lead, ML engineering |
| Verification | Application security or red team lead, QA |
| Operations | SOC and incident response lead, SRE or operations |
Worked example 1: an assessment for a client
This scenario is illustrative.
A mid-size insurer runs two AI systems: a customer-support chatbot built on a hosted LLM with retrieval over policy documents, and an internal model that triages incoming claims. The CISO asks two questions: where do we stand, and what do we fix first?
Scope: the insurer's AI program - the support chatbot and the claims-triage model.Outside scope (recorded, not scored "no"): the group privacy office; Privacy evidence comes from its reports.Model: OWASP AIMA V1.0, all eight domains.Depth: lightweight everywhere; detailed, with evidence sampling, for Governance, Verification and Operations.Scoring: document method, gated yes/no with "+", because next year's result will be compared with this one.Level 1 reading: at least an informal version of the activity exists.Overlap: deployment evidence counts under Secure Deployment only.Targets: set in a workshop after calibration, from the client's risk appetite.Six of the 24 results, with the evidence behind each:
| Practice | What the evidence showed | Score | Target | First move |
|---|---|---|---|---|
| Strategy & Metrics | Unapproved strategy deck; chatbot metrics reviewed monthly, nothing on risk | 1+ | 2 | Approve the AI strategy with three risk indicators, reviewed quarterly by the risk committee |
| Policy & Compliance | AI acceptable-use policy approved and communicated; AI-related legal requirements documented and reviewed yearly | 2 | 2 | None this cycle; hold the level |
| Threat Assessment | Chatbot threat-modeled at launch, claims model never; mitigations never reviewed | 1 | 2 | Threat-model every AI system at design and on material change (VI.3) |
| Security Testing | Yearly web pentest touched the chatbot’s interface and its findings were fixed; no prompt-injection or jailbreak testing | 1 | 2 | A documented AI test plan run every release (VI.4) |
| Incident Management | IR plan covers IT outages only; one prompt-injection report handled in an email thread | 1 | 2 | An AI incident playbook plus one tabletop exercise (VII.3) |
| Fairness & Bias | Bias discussed informally; claims handlers spot-check a sample of triage decisions | 1 | 2 | Define fairness metrics for the claims model and test them before each retrain |
What the client receives:
- The scoping note above, so next year’s assessment can use the same rules.
- The full profile, 24 current scores against 24 targets, one page.
- The top three moves with owners and quarters. Here: the strategy with risk indicators, the AI test plan and the incident playbook, because each one closes a gap on both systems at once.
- A reassessment date six to twelve months out, when the first moves should have landed.
Resist the temptation to recommend “Level 3 everywhere”. A roadmap that closes three gaps this year is worth more to the client than a slide showing 24 red cells.
Worked example 2: adopting AIMA inside your own organization
Also illustrative. A 300-person software company ships LLM features in its product and runs an internal coding assistant. The head of security wants one measure of the AI program that engineering will actually use.
- Name an owner per domain, using the table above, and give each owner their six questions per practice.
- Baseline with the toolkit average. It is fast, and the decimals make progress visible quarter to quarter. Keep the evidence for each answer in the interview-note column, so the next quarter’s answers can be checked against it.
- Tie targets to work the teams already plan. Security Testing rises when AI tests run in CI (VII.1 · The secure AI SDLC); Event Management rises when model and tool calls reach the SIEM (VII.3 · Detection, IR & forensics for AI).
- Reassess every quarter, and calibrate once a year. A yearly external or peer assessment, using the gated method, keeps the internal numbers honest.
- Re-baseline when scope changes. Launching an agent product adds new surface; last quarter’s score for Security Testing may no longer describe the system you run.
What the trend looks like after three quarters (illustrative numbers, toolkit method):
| Practice | Q1 | Q2 | Q3 | What changed |
|---|---|---|---|---|
| Security Testing | 0.83 | 1.50 | 2.00 | Injection and jailbreak probes added to CI, results tracked per release |
| Event Management | 0.50 | 0.83 | 1.67 | Model and tool-call logs shipped to the SIEM with two detections |
| Incident Management | 1.00 | 1.00 | 1.83 | AI playbook written and exercised once |
| Data Training | 1.17 | 1.33 | 1.33 | Dataset licensing review started, not yet systematic |
Report the profile, not an overall average. A program average of 1.8 can hide a 0.5 in incident management, and that is the number that matters on the day something goes wrong.
Evidence that supports each level
AIMA defines the criteria, not the artifacts. This is typical evidence for four practices, as a starting point for your evidence request.
| Practice | Level 1 | Level 2 | Level 3 |
|---|---|---|---|
| Strategy & Metrics | A strategy note or deck, even unapproved; some AI metrics tracked | An approved strategy, communicated, with a defined set of risk indicators reviewed on a schedule | Risk indicators in board reporting; the strategy revised from what the metrics show |
| Threat Assessment | A threat list for at least one AI system | A documented threat model for every AI system, with mitigations reviewed periodically | Threat modeling is a lifecycle gate, updated on every material change |
| Security Testing | An occasional assessment that touched an AI feature | A documented AI test plan, covering injection, jailbreaks and data leakage, run on a schedule | AI tests run in CI; results tracked as metrics and audited |
| Incident Management | Ad hoc handling of an AI-related report | An AI incident procedure used consistently, with incidents documented and reviewed | Procedures exercised through tabletops; lessons fed back into design and testing |
Where AIMA sits next to the other frameworks
AIMA V1.0 does not ship a formal mapping to other frameworks. The table below is this book’s reading of how they fit together.
| If you also use | What it answers | How AIMA fits |
|---|---|---|
| NIST AI RMF (VIII.3) | The risk-management process: govern, map, measure, manage | AIMA measures how consistently you perform the activities your RMF profile commits to |
| ISO/IEC 42001 (VIII.4) | A certifiable AI management system | A readiness check before the audit, and the internal measure between audits; not a substitute for certification |
| EU AI Act (VIII.6) | Legal obligations by risk tier | Program maturity, not legal compliance; map specific obligations separately |
| AISVS and AIVSS (VIII.4) | Testable requirements, and the severity of a finding | Test results become evidence for the Design, Implementation and Verification scores |
| The four-rung ladder (VIII.4) | A quick headline for a first conversation | Move to AIMA when the result must be repeatable and auditable |