The AI agent maturity model: a 7‑stripe matrix
Updated · By Algen AI
What an AI agent maturity model should measure
Maturity is about engineering, not model quality. Two agents on the same model can differ completely: one stops a runaway loop, says when a tool failed and asks before deleting a record; the other doesn't.
AgentStripes measures properties, not techniques. "Runs are bounded" can be met by a framework's default step limit, your own budget check or your platform. Any mechanism that achieves the property counts, so the model doesn't favour a framework or a vendor.
The seven stripes
Each stripe adds to the one before. The first two can be earned from the code alone; the rest need evidence from runs the developer does and submits.
| Stripe | What it means | Evidence |
|---|---|---|
| 1 Defined | Someone else can tell what it does, run it, and see its tools and instructions. No secrets in the code. | From the code |
| 2 Structured | Its job and limits are written down, tools are described and typed, runs are bounded, errors surface, and risky actions need approval. | From the code |
| 3 Evaluated | It has an evaluation set drawn from real use and it passes it, and its runs are traced so failures can be read. | From submitted runs |
| 4 Observed | Traces of its real runs show it behaves as declared. | From submitted runs |
| 5 Verified | It passes a published test suite, including held-out tasks and injection attacks, and the full results are submitted so anyone can re-run them. | From submitted runs |
| 6 Hardened | It holds up across repeated runs, injected faults and deeper attacks, and every change is gated by regression evals. | From submitted runs |
| 7 Proven | At least 90 days of production evidence show it stays reliable, honest and within budget. | From submitted runs |
The maturity matrix: 31 capabilities in 10 areas
Every capability is rated 0 (absent), 1 (basic), 2 (adequate) or 3 (strong) against written anchors, with a separate evidence level. Established capabilities reflect settled practice and count for and against an agent; emerging ones count only in its favour while practice forms.
| Area | Capability | Status |
|---|---|---|
| Task contract | Purpose documented: Someone new can tell from the agent's documentation what it does, who it is for, what it takes in and what it produces. | Established |
| Task contract | Limits documented: The documentation states what the agent must not be used for, or its known limitations. | Emerging |
| Task contract | Instructions owned and explicit: The agent's instructions can be found in one clear place, maintained by the developers, and state its rules plainly. | Emerging |
| Tool design | Tool interfaces clear: Each tool tells the model what it does and when to use it. | Established |
| Tool design | Tool inputs specified: Tool inputs have declared types or a schema the framework can validate. | Established |
| Tool design | Tools fail gracefully: When a tool's call to a network, database or file fails, the agent gets a useful message rather than a crash or a raw traceback, and large results are bounded. | Established |
| Tool design | Tool set coherent: Each tool does a distinct job, and the set is no larger than the task needs. | Emerging |
| Context and memory | Context bounded: The amount of information placed in the model's context is kept within limits. | Emerging |
| Context and memory | Memory isolated: Persistent memory is kept separate per user, thread or tenant. | Emerging |
| Context and memory | Answers grounded: When the agent retrieves information to answer, it ties answers to their sources. | Emerging |
| Control flow | Runs bounded: A run cannot continue without limit in steps, time or spend. | Established |
| Control flow | Stopping clear: It is clear when a run has finished its job. | Emerging |
| Control flow | Runnable: Someone else can run the agent from documented steps. | Established |
| Verification and honesty | Failures surfaced: Errors are reported or handled, never silently swallowed. | Established |
| Verification and honesty | Outcomes verified: After an action with real-world effect, the agent checks it worked before reporting success. | Emerging |
| Verification and honesty | Success defined: What a successful run looks like is written down. | Emerging |
| Human oversight | Consequential actions controlled: Actions with real-world effect wait for a person's approval, or are otherwise constrained so a mistake can't do serious harm. | Established |
| Human oversight | Runs interruptible: A run can be paused or stopped from outside. | Emerging |
| Human oversight | Clarification possible: The agent can ask the user when a request is unclear. | Emerging |
| Security | Secrets protected: No credentials are committed to the repository; secrets come from the environment or a secret store. | Established |
| Security | Untrusted input separated: Content that can come from users or the outside world (tool results, documents, web pages, stored memories) is passed to the model as data, not as instructions. | Established |
| Security | Execution isolated: Generated code or shell commands run in a sandbox or container, not on the host. | Established |
| Security | Supply chain pinned: Dependencies and third-party tools are pinned to versions. | Established |
| Security | Egress limited: Model output can't send the agent to arbitrary web addresses. | Emerging |
| Security | Disclosure channel: There is a licence and a way to report security problems. | Established |
| Evaluation | Automated tests: The repository has automated tests of the agent or its tools. | Established |
| Evaluation | Evaluation set: There is a set of tasks with expected outcomes used to score the agent. | Established |
| Evaluation | Continuous testing: Tests or evals run automatically on changes. | Emerging |
| Observability | Runs traceable: Each run can be identified and its model and tool calls inspected afterwards. | Established |
| Change management | Behaviour inputs pinned: The model versions and prompts that shape behaviour are pinned or versioned, so behaviour doesn't change underneath the agent. | Emerging |
| Change management | Releases documented: Releases are versioned and their changes described. | Established |
How the ratings are made
- Instruments read every file: a code parser, a secret scanner and file readers. They report facts, such as a committed key or a framework's step limit.
- Three model judges rate every capability against its anchors and must quote the code; a quote that isn't in the file is thrown out.
- The rating is the median. When the judges disagree by more than a point, it counts as basic until they agree.
- A published, versioned profile turns ratings into stripes, so the bar can rise as the field matures without changing the method.
How to use it
Check a public repository for free on AgentStripes, read what stands between you and the next stripe, and fix one capability at a time. Agent Hub lists agents with at least two stripes.
Browse versioned agents with their stripes, permissions and delivery paths.
The AI agent maturity model: a 7‑stripe matrix: common questions
What is an AI agent maturity model?
A structured way to describe how well an AI agent is engineered, from a prototype to a production system. AgentStripes rates 31 capabilities from the agent's code and maps them to seven stripes.
What are the levels of AI agent maturity?
1 Defined, 2 Structured, 3 Evaluated, 4 Observed, 5 Verified, 6 Hardened, 7 Proven. Stripes 1 and 2 come from the code; 3 to 7 need evidence from submitted runs.
Is an AI agent maturity model the same as a safety certification?
No. A maturity model describes engineering practice and the evidence behind it. AgentStripes is not certification, a security audit or a guarantee that an agent is safe.
Which frameworks does it cover?
LangChain, LangGraph, CrewAI and the OpenAI Agents SDK, in Python and TypeScript. The capabilities themselves are framework-neutral.