Open source · runs on your machine
AgentTrail Guard
A seatbelt for your coding agent. Before Claude Code, Cursor or Codex CLI runs a command or touches a file, AgentTrail Guard checks it against 74 open guardrails and answers one of three things: go ahead, ask you first, or no.
npm i -g @agenttrail/guardThen one command per agent. Needs Node.js 20 or later.
Works with
- Claude Code
- Cursor
- Codex
- Runs on
- Your machine
- Account
- None needed
- Network calls
- None by default
- License
- Apache 2.0
How it works
One check between your agent and your machine
AgentTrail Guard runs as a hook inside your agent. Each command and file change passes through it before it runs, and it decides on your laptop, from a catalog bundled into the package. No server is asked.
Your agent proposes an action
A shell command, a file read or write, a search, or an MCP tool call.
AgentTrail Guard checks it
Against every guardrail you have on, plus any you wrote yourself. Nothing runs yet.
It answers before the action runs
Block, ask or warn when a guardrail matches. When nothing matches, it stays out of the way.
blockNo.
The call never runs. Your agent is told which guardrail stopped it and why, so it can explain and try something safer.
require_approvalAsk the human first.
Your agent's own permission prompt asks you before the call runs. Approve it, or say no. Where an agent cannot ask, the call is refused instead.
warnGo ahead, but say so.
The call runs, and the match is written to the local decision log, where status shows it to you.
See it work before you trust it
init finishes by running a synthetic rm -rf / through the real guardrail engine and showing it blocked. Nothing is executed. Add --print to see exactly what init would change, and change nothing.
From a real Claude Code session
Told it was only dev data, the agent reached for npx prisma db push --force-reset. The dd.accept-data-loss guardrail blocked it before it ran, and the rows survived.
agenttrail-guard guardrails set-action <id> block|ask|warnWorks with
Claude Code, Cursor and Codex CLI
One npm install, then one setup command per agent. Each install is separate, you can have all three, and uninstalling removes only what AgentTrail Guard added.
Claude Code
$ agenttrail-guard init --agent claudeInstalls as a Claude Code plugin that checks each tool call first. It never touches the hooks block in your settings.json.
Know before you rely on it
- An ask uses Claude Code's own permission prompt, answered on the spot.
- A tool your
settings.jsonallows outright runs without a prompt, so a hold on it never reaches you.statuscounts these and names the tools. - After an update, restart Claude Code: it runs its own cached copy of the plugin until then.
Cursor
$ agenttrail-guard init --agent cursorAdds two entries to your user hooks file, ~/.cursor/hooks.json, and keeps every other entry in it, in order.
Know before you rely on it
- Use the agent in an editor window. The Agent Window can skip hooks, and status reminds you.
- Approval works for terminal commands. An ask on a file edit, read or MCP call is refused instead, with the guardrail named.
- Not seen: Tab edits, reads shown as Explored, files attached to the conversation for you, and cloud agents.
Codex CLI
$ agenttrail-guard init --agent codexAppends its entries to Codex's user hooks file, ~/.codex/hooks.json, and registers nowhere else, so it never runs twice.
Know before you rely on it
- Approve it once inside Codex: type
/hooksand approve eachagenttrail-guardentry. Until you do, nothing is checked, and nothing tells you so. - An ask becomes a refusal on Codex, because Codex runs an action when a hook asks for approval. Downgrade any guardrail to warn if that is too strong.
- Codex reads files through the shell, so reads are checked as commands, not by file path.
- Measured on codex-cli 0.154.0. Windows and Codex Desktop are not yet verified, and Codex Cloud runs on OpenAI's machines, so it is not covered.
The CLI
Six commands, all safe to re-run
A machine-level tool you install once. It keeps its settings in your home directory and writes nothing into your projects.
When a guardrail is wrong
The likeliest reason to remove a tool like this is one guardrail firing on something legitimate. status names the guardrail firing most and prints the exact line that silences it on that one command shape. Everything else it guards keeps working.
allow refuses a pattern that would quietly do more than you asked, like *, or one of the commands the guardrail exists to stop.
- agenttrail-guard init --agent <claude|cursor|codex>
- Install the hook for one agent, seed the config, and show itself blocking a harmless rm -rf / demo. Safe to re-run.
- agenttrail-guard status
- What is on in each app, the last five decisions, and the one line that silences the guardrail firing most.
- agenttrail-guard guardrails
- See and change what is enforced: packs, single guardrails, actions and allowlists.
- agenttrail-guard scan --agent <claude|cursor|codex>
- Read the sessions already on your disk and report what your agent has been doing.
- agenttrail-guard crash-report
- See, enable, send or clear opt-in crash reports. Off by default.
- agenttrail-guard uninstall --agent <claude|cursor|codex>
- Remove our plugin or hook entries, and only ours. Your settings stay put.
Scan
See what your agent has already been doing
Your agent already keeps a record of every session on your disk. One command reads it and shows you what AgentTrail Guard would have caught, before you change anything.
Replays your real history
Reads the session files Claude Code, Cursor and Codex CLI already keep on your disk, and runs every tool call through the same guardrails the hook uses.
Tells you what repeats
How many sessions, which guardrails would have fired, the most serious findings, and the mistakes your agent makes again and again. For Claude Code, exact token counts too.
One file, nothing uploaded
A self-contained agenttrail-guard-report.html that loads nothing from outside itself, so it opens with the wifi off and cannot phone home.
Redacted before it is written
Secrets, paths and identifying names are scrubbed, and your projects are a count, never names. Redaction is thorough, not guaranteed, so scan --review lets you read every line first.
No dollar figures. Counts and token totals come straight off the transcript; the report never multiplies them by an assumption. In Claude Code, /agenttrail-guard:share-report can publish it as a private Claude artifact, only after you say yes.
Privacy
Nothing leaves your machine unless you turn it on.
No account, no telemetry, no install ping. Here is everything AgentTrail Guard reads, keeps and sends.
- Reads
- The command or file change your agent is about to make, so it can check it before it runs. With scan, your agent's own session files, on your machine.
- Keeps
- Its settings, and a log of the calls that matched a guardrail, secrets scrubbed first. Capped at 1 MiB and 30 days, and you can clear it any time.
- Sends
- Nothing, by default. Crash reporting is opt-in, a scrubbed stack trace only, with nowhere to send until you set an endpoint yourself.
Every file it keeps
All under ~/.agenttrail/guard/, at mode 0600.
- config.json
- Only what you changed: packs or guardrails turned off, actions, allowlists.
- guardrails.json
- Your own guardrails, enforced alongside the library. Starts empty.
- events.jsonl
- One scrubbed line per matched call: time, tool, decision, guardrail, command, agent.
- crashes/
- Scrubbed stack traces, capped at 20 files and 30 days. Never sent unless you send them.
- cursor/ and codex/
- Only after init for that agent: the hook copy it runs, a backup of its hooks file, and an install record.
Honest limits
What AgentTrail Guard does not do
Read this before you rely on it. A security tool that overstates its coverage is worse than one that names the hole.
It does not see the web
Pages your agent fetches are not intercepted, and there are no URL guardrails. Search queries are checked, because they are ordinary text.
It reads a path, not a file's contents
A file guardrail decides by where a write lands, never by what it contains. A script written and then run is checked only as the command that runs it.
It cannot stop what it cannot see
Anything your agent does outside a hooked tool is invisible to it, and each agent has its own gaps, listed above.
It fails open
If it breaks, the call goes through your agent's normal permission flow and is noted as not checked. It never answers allow on your behalf.
It decides, it never rewrites
Secrets are scrubbed from what it writes down, not from what it lets through. An allowed command runs exactly as the agent wrote it.
Its redaction is good, not complete
A secret passed as a bare argument, like --token abc123, can reach the local log. Treat the log as scrubbed, not sanitized.
It is one developer on one machine
Nothing syncs, nothing is shared with a team, and there is no approval queue or history beyond the local log.
Get started
Protected in two commands. No account.
Install it once, set it up for each agent you use (on Codex, approve it in /hooks), and run agenttrail-guard status to see it on.
Want a record, not just a block?
AgentTrail OS keeps every agent session searchable, finds the mistakes your agents repeat and turns them into rules, proven on your own history first. Free for one developer, and ready for a team.