Open source · runs on your machine

AgentTrail Guard

A seatbelt for your coding agent. Before Claude Code, Cursor or Codex CLI runs a command or touches a file, AgentTrail Guard checks it against 74 open guardrails and answers one of three things: go ahead, ask you first, or no.

bash
$npm i -g @agenttrail/guard

Then one command per agent. Needs Node.js 20 or later.

Works with

  • Claude Code
  • Cursor
  • Codex
Runs on
Your machine
Account
None needed
Network calls
None by default
License
Apache 2.0

One check between your agent and your machine

AgentTrail Guard runs as a hook inside your agent. Each command and file change passes through it before it runs, and it decides on your laptop, from a catalog bundled into the package. No server is asked.

  1. Your agent proposes an action

    A shell command, a file read or write, a search, or an MCP tool call.

  2. AgentTrail Guard checks it

    Against every guardrail you have on, plus any you wrote yourself. Nothing runs yet.

  3. It answers before the action runs

    Block, ask or warn when a guardrail matches. When nothing matches, it stays out of the way.

Blockblock

No.

The call never runs. Your agent is told which guardrail stopped it and why, so it can explain and try something safer.

Askrequire_approval

Ask the human first.

Your agent's own permission prompt asks you before the call runs. Approve it, or say no. Where an agent cannot ask, the call is refused instead.

Warnwarn

Go ahead, but say so.

The call runs, and the match is written to the local decision log, where status shows it to you.

See it work before you trust it

init finishes by running a synthetic rm -rf / through the real guardrail engine and showing it blocked. Nothing is executed. Add --print to see exactly what init would change, and change nothing.

From a real Claude Code session

Told it was only dev data, the agent reached for npx prisma db push --force-reset. The dd.accept-data-loss guardrail blocked it before it ran, and the rows survived.

Read the case study
Every action is yours to change: agenttrail-guard guardrails set-action <id> block|ask|warn

Claude Code, Cursor and Codex CLI

One npm install, then one setup command per agent. Each install is separate, you can have all three, and uninstalling removes only what AgentTrail Guard added.

Claude Code

$ agenttrail-guard init --agent claude

Installs as a Claude Code plugin that checks each tool call first. It never touches the hooks block in your settings.json.

Know before you rely on it

  • An ask uses Claude Code's own permission prompt, answered on the spot.
  • A tool your settings.json allows outright runs without a prompt, so a hold on it never reaches you. status counts these and names the tools.
  • After an update, restart Claude Code: it runs its own cached copy of the plugin until then.

Cursor

$ agenttrail-guard init --agent cursor

Adds two entries to your user hooks file, ~/.cursor/hooks.json, and keeps every other entry in it, in order.

Know before you rely on it

  • Use the agent in an editor window. The Agent Window can skip hooks, and status reminds you.
  • Approval works for terminal commands. An ask on a file edit, read or MCP call is refused instead, with the guardrail named.
  • Not seen: Tab edits, reads shown as Explored, files attached to the conversation for you, and cloud agents.

Codex CLI

$ agenttrail-guard init --agent codex

Appends its entries to Codex's user hooks file, ~/.codex/hooks.json, and registers nowhere else, so it never runs twice.

Know before you rely on it

  • Approve it once inside Codex: type /hooks and approve each agenttrail-guard entry. Until you do, nothing is checked, and nothing tells you so.
  • An ask becomes a refusal on Codex, because Codex runs an action when a hook asks for approval. Downgrade any guardrail to warn if that is too strong.
  • Codex reads files through the shell, so reads are checked as commands, not by file path.
  • Measured on codex-cli 0.154.0. Windows and Codex Desktop are not yet verified, and Codex Cloud runs on OpenAI's machines, so it is not covered.

Six commands, all safe to re-run

A machine-level tool you install once. It keeps its settings in your home directory and writes nothing into your projects.

When a guardrail is wrong

The likeliest reason to remove a tool like this is one guardrail firing on something legitimate. status names the guardrail firing most and prints the exact line that silences it on that one command shape. Everything else it guards keeps working.

allow refuses a pattern that would quietly do more than you asked, like *, or one of the commands the guardrail exists to stop.

agenttrail-guard init --agent <claude|cursor|codex>
Install the hook for one agent, seed the config, and show itself blocking a harmless rm -rf / demo. Safe to re-run.
agenttrail-guard status
What is on in each app, the last five decisions, and the one line that silences the guardrail firing most.
agenttrail-guard guardrails
See and change what is enforced: packs, single guardrails, actions and allowlists.
agenttrail-guard scan --agent <claude|cursor|codex>
Read the sessions already on your disk and report what your agent has been doing.
agenttrail-guard crash-report
See, enable, send or clear opt-in crash reports. Off by default.
agenttrail-guard uninstall --agent <claude|cursor|codex>
Remove our plugin or hook entries, and only ours. Your settings stay put.
agenttrail-guard guardrails
$agenttrail-guard guardrails list --disabled# what is off
$agenttrail-guard guardrails show <id># everything about one guardrail
$agenttrail-guard guardrails set-action <id> warn# block, ask or warn
$agenttrail-guard guardrails disable <pack># a whole pack, from the next call
$agenttrail-guard guardrails allow <id> <pattern># one guardrail, one command shape
$agenttrail-guard guardrails add ./my-rule.json# a guardrail you wrote
$agenttrail-guard guardrails reset --all# undo your changes

See what your agent has already been doing

Your agent already keeps a record of every session on your disk. One command reads it and shows you what AgentTrail Guard would have caught, before you change anything.

Local files only
$agenttrail-guard scan --agent claude
$agenttrail-guard scan --agent cursor
$agenttrail-guard scan --agent codex
$agenttrail-guard scan --agent claude --review# read every line first

Replays your real history

Reads the session files Claude Code, Cursor and Codex CLI already keep on your disk, and runs every tool call through the same guardrails the hook uses.

Tells you what repeats

How many sessions, which guardrails would have fired, the most serious findings, and the mistakes your agent makes again and again. For Claude Code, exact token counts too.

One file, nothing uploaded

A self-contained agenttrail-guard-report.html that loads nothing from outside itself, so it opens with the wifi off and cannot phone home.

Redacted before it is written

Secrets, paths and identifying names are scrubbed, and your projects are a count, never names. Redaction is thorough, not guaranteed, so scan --review lets you read every line first.

No dollar figures. Counts and token totals come straight off the transcript; the report never multiplies them by an assumption. In Claude Code, /agenttrail-guard:share-report can publish it as a private Claude artifact, only after you say yes.

Nothing leaves your machine unless you turn it on.

No account, no telemetry, no install ping. Here is everything AgentTrail Guard reads, keeps and sends.

Reads
The command or file change your agent is about to make, so it can check it before it runs. With scan, your agent's own session files, on your machine.
Keeps
Its settings, and a log of the calls that matched a guardrail, secrets scrubbed first. Capped at 1 MiB and 30 days, and you can clear it any time.
Sends
Nothing, by default. Crash reporting is opt-in, a scrubbed stack trace only, with nowhere to send until you set an endpoint yourself.

Every file it keeps

All under ~/.agenttrail/guard/, at mode 0600.

config.json
Only what you changed: packs or guardrails turned off, actions, allowlists.
guardrails.json
Your own guardrails, enforced alongside the library. Starts empty.
events.jsonl
One scrubbed line per matched call: time, tool, decision, guardrail, command, agent.
crashes/
Scrubbed stack traces, capped at 20 files and 30 days. Never sent unless you send them.
cursor/ and codex/
Only after init for that agent: the hook copy it runs, a backup of its hooks file, and an install record.
What AgentTrail Guard reads, keeps and sends

What AgentTrail Guard does not do

Read this before you rely on it. A security tool that overstates its coverage is worse than one that names the hole.

  • It does not see the web

    Pages your agent fetches are not intercepted, and there are no URL guardrails. Search queries are checked, because they are ordinary text.

  • It reads a path, not a file's contents

    A file guardrail decides by where a write lands, never by what it contains. A script written and then run is checked only as the command that runs it.

  • It cannot stop what it cannot see

    Anything your agent does outside a hooked tool is invisible to it, and each agent has its own gaps, listed above.

  • It fails open

    If it breaks, the call goes through your agent's normal permission flow and is noted as not checked. It never answers allow on your behalf.

  • It decides, it never rewrites

    Secrets are scrubbed from what it writes down, not from what it lets through. An allowed command runs exactly as the agent wrote it.

  • Its redaction is good, not complete

    A secret passed as a bare argument, like --token abc123, can reach the local log. Treat the log as scrubbed, not sanitized.

  • It is one developer on one machine

    Nothing syncs, nothing is shared with a team, and there is no approval queue or history beyond the local log.

Protected in two commands. No account.

Install it once, set it up for each agent you use (on Codex, approve it in /hooks), and run agenttrail-guard status to see it on.

Want a record, not just a block?

AgentTrail OS keeps every agent session searchable, finds the mistakes your agents repeat and turns them into rules, proven on your own history first. Free for one developer, and ready for a team.