AI agent safety checklist: 12 checks before you ship
Updated · By Algen AI
The 12-point safety checklist
Work through these in order: the first five decide what an attacker or a bad input can make the agent do.
- Secrets protectedNo credentials are committed to the repository; secrets come from the environment or a secret store. What good looks like: Secrets come from the environment or a store.
- Untrusted input separatedContent that can come from users or the outside world (tool results, documents, web pages, stored memories) is passed to the model as data, not as instructions. What good looks like: Untrusted content is passed as data, apart from instructions.
- Execution isolatedGenerated code or shell commands run in a sandbox or container, not on the host. What good looks like: Runs in a sandbox or container.
- Consequential actions controlledActions with real-world effect wait for a person's approval, or are otherwise constrained so a mistake can't do serious harm. What good looks like: Each consequential action needs approval or is constrained to be safe.
- Egress limitedModel output can't send the agent to arbitrary web addresses. What good looks like: Destinations are limited to what the agent needs.
- Runs boundedA run cannot continue without limit in steps, time or spend. What good looks like: Every run is bounded, for example by a step limit, including a framework's default step limit that the code leaves in place.
- Stopping clearIt is clear when a run has finished its job. What good looks like: An explicit end condition, output type or expected output.
- Failures surfacedErrors are reported or handled, never silently swallowed. What good looks like: Errors on the agent's path surface: raised, logged, or returned with context. A crash counts. Errors ignored off that path, such as a failed optional update check, don't count against it.
- Runs interruptibleA run can be paused or stopped from outside. What good looks like: Runs can be paused, cancelled or resumed.
- Supply chain pinnedDependencies and third-party tools are pinned to versions. What good looks like: Dependencies are locked and third-party tools pinned.
- Runs traceableEach run can be identified and its model and tool calls inspected afterwards. What good looks like: Tracing or logging of model and tool calls, tied to run ids, is switched on for the agent's own runs. A key or project name alone, or tracing only in CI or tests, is not enough.
- Disclosure channelThere is a licence and a way to report security problems. What good looks like: A licence and a security contact.
The agent safety matrix
For each check, the failure it prevents and what each rating means.
| Check | Prevents | Basic (1) | Adequate (2) |
|---|---|---|---|
| Secrets protected | Leaked credentials. | Placeholders mix with hard-coded values. | Secrets come from the environment or a store. |
| Untrusted input separated | Goal hijacking and prompt injection through tool outputs (OWASP ASI01; AgentDojo). | Some separation. | Untrusted content is passed as data, apart from instructions. |
| Execution isolated | Unexpected code execution (OWASP ASI05). | Partially restricted. | Runs in a sandbox or container. |
| Consequential actions controlled | Irreversible actions taken by mistake or under manipulation (OWASP ASI02, ASI09; the Replit incident). | Some are gated, or constrained by policy. | Each consequential action needs approval or is constrained to be safe. |
| Egress limited | Data exfiltration through fetch or browse tools. | Some checks. | Destinations are limited to what the agent needs. |
| Runs bounded | Step repetition and runaway runs (MAST FM-1.3, 15.7% of failures). | A bound exists but doesn't cover every run, or is set so high it rarely applies. | Every run is bounded, for example by a step limit, including a framework's default step limit that the code leaves in place. |
| Failures surfaced | Silent failure followed by false reports of success. | Some errors on the agent's path are discarded or misreported, though most surface. | Errors on the agent's path surface: raised, logged, or returned with context. A crash counts. Errors ignored off that path, such as a failed optional update check, don't count against it. |
| Supply chain pinned | Supply-chain compromise and drift (OWASP ASI04; MCP tool poisoning). | Some pins. | Dependencies are locked and third-party tools pinned. |
What a checklist can't tell you
Reading code shows how an agent is built, not how it behaves under attack. Prompt injection in particular needs testing with real runs, which is why AgentStripes caps agents that take consequential actions at two stripes until submitted test results show approval and injection defences hold.
Browse versioned agents with their stripes, permissions and delivery paths.
AI agent safety checklist: 12 checks before you ship: common questions
How do I make an AI agent safe?
Keep secrets out of the code, treat retrieved and user content as data rather than instructions, isolate any code it runs, require approval before consequential actions, limit where it can send data, bound every run, and make sure errors surface. Then test it with real runs.
What is an agent safety matrix?
A table of the safety properties an agent should have, what each prevents, and what basic and adequate implementations look like. The matrix above is drawn from the AgentStripes model.
Does passing this checklist make an agent safe?
No. It shows the engineering is in place. Safety depends on your data, permissions and deployment, and needs testing with real runs.
How is prompt injection handled?
By keeping untrusted content apart from instructions, never letting it carry authority, and gating consequential actions behind approval. It has been the top risk in recognised LLM security guidance since 2023.