How to evaluate an AI agent before you adopt it
Updated · By Algen AI
The 8-step evaluation checklist
Work through these in order. Each step narrows what you still need to trust.
- Define the job and its boundaryWrite down the task, the data it may touch and what must never happen. You cannot judge fit without a boundary.
- Identify the exact versionRecord the version and artifact digest. Evidence, reviews and assessments only apply to the release they were written for.
- Read the declared permissionsCheck data classes, outbound hosts and required secrets. An agent that needs broad network access or credentials should justify it.
- Check what the assessment coversFor an ASCEND level, read the scope, the evidence and the separate risk overlay. Note what is explicitly outside scope.
- Find the human decision pointsLook for where the agent must wait for a named approval, and where it can act on its own.
- Test in isolationRun it in a disposable environment with non-sensitive data. Verify the package digest before you install.
- Review the supply chainCheck license, SBOM, provenance, publisher identity, support contact and how security reports are handled.
- Plan monitoring and rollbackDecide how you will observe runs, what triggers a rollback, and how you will adopt the next version deliberately.
How to read an ASCEND level
ASCEND levels are cumulative gates, from M0 Declared to M5 Continuously assured. A higher level means more assurance mechanisms were demonstrated for that version, not that the agent is safe for your deployment.
Always read the level together with the risk overlay and the scope statement. Maturity and risk are separate questions, and a level never transfers to a different version.
| Level | What was demonstrated for that version |
|---|---|
| M0 Declared | Owner, purpose, limitations, version, license, support, data use, and requested permissions are disclosed. |
| M1 Observable | Runs have stable identity and produce bounded, privacy-aware logs, traces, metrics, errors, latency, and usage evidence. |
| M2 Statically governed | The version has a machine-readable manifest, dependency evidence, threat model, evaluations, and documented least privilege. |
| M3 Runtime governed | Budgets, scopes, validation, approvals, tool and egress controls, cancellation, and failure containment are enforced during execution. |
| M4 Auditable | Attributable records, provenance, signed releases, incident handling, retention, and reproducible version evidence are maintained. |
| M5 Continuously assured | SLOs, continuous evaluations, adversarial testing, drift gates, recovery exercises, and independent evidence review operate over time. |
Red flags worth pausing on
- No version or digest, or a digest that does not match your download.
- Permissions that are broader than the stated job.
- An assessment with no stated scope, date or version.
- Ratings that cannot be tied to a version or an acquisition.
- No human decision point for actions that affect customers, money or records.
- No support or security contact.
Browse versioned agents with visible ASCEND levels, permissions and delivery paths.
How to evaluate an AI agent before you adopt it: common questions
How do I know if an AI agent is safe to use?
You cannot rely on a single badge. Verify the exact version and digest, read the declared permissions and the assessment scope, test in isolation with non-sensitive data, and confirm where humans must approve actions. Then decide against your own risk tolerance.
What should I check before installing an AI agent?
The version and digest, declared data classes, outbound hosts and secrets, the license and SBOM, the publisher and support contact, and what any assessment covers and excludes.
Does a higher ASCEND level mean lower risk?
Not by itself. Maturity shows which assurance mechanisms were demonstrated. Risk is a separate overlay describing what could happen in the scoped deployment, and it is never lowered because controls are hard.
Should I test agents with real data first?
No. Start in a disposable environment with non-sensitive data, and widen access only after you have observed the agent's behavior.