Open-source guardrail

The agent editing its own configuration

AgentTrail Guard asks you about this by default: you decide before the call runs. Where an agent cannot ask, the call is refused instead.

Default action
Ask
Severity
High severity
Library version
0.2.1Sep 29, 2026

What it catches, and what it misses

Written into the rule itself, next to what it matches, so you can judge it before you trust it.

Holds a file tool opening the files that define what the agent itself is allowed to do — Claude Code settings and hooks, an MCP server list, a plugin manifest, Cursor rules, a Codex config, and the guard's own config.json and guardrails.json. An agent that can edit these can widen its own reach with nobody reviewing the change. Reading is excluded — the Read and Grep tools never match — so opening one of these files to read it is not held; every other file tool is, including one this corpus does not know. Does NOT match ordinary project files, or .claude/commands/*.md, which are prompts rather than permissions. Bounded to WELL-KNOWN paths: no working directory or project root reaches the guard, so it can only match names it already knows.

Tested on every build

The rule must match every command on the left and none on the right, or the library does not build. Catching the real thing is easy; staying quiet on the near-miss is the hard part.

Catches (4)

  • Edit.claude/settings.json
  • Edit.claude/settings.local.json
  • Write/home/dev/.agenttrail/guard/config.json
  • Edit.mcp.json

Stays quiet on (6)

  • Edit.claude/commands/deploy.md
  • Editsrc/index.ts
  • Editpackage.json
  • EditREADME.md
  • Read.claude/settings.json
  • Grep.mcp.json

What "ask" means in each app

The guard runs as a hook in each app, and each app gives a hook different powers. Here is what this rule's default action does in each one.

Claude Code
You are asked before the call runs. If your settings.json already allows the tool outright, Claude Code runs it with no prompt.
Cursor
A terminal command that reaches the guard gets Cursor's approval card. Any other call that needs approval, such as a file edit or a read, is refused, because Cursor cannot ask there.
Codex CLI
Codex cannot ask, so the call is refused, and the reason says a person has to approve it.

This rule matches file paths. In Codex CLI, reading a file fires no hook, and in Cursor, reads shown as Explored are not checked, so there it sees fewer calls than in Claude Code. Each app hands the guard a different set of calls: in a Cursor sandbox run mode, for one, some terminal commands run without reaching the guard at all. Read the notes for Cursor and for Codex CLI before you rely on a rule there.

Change what it does in one command

Turn it off, change its action, or silence it on one command shape. The narrow one is allow: the rule keeps catching everything else.

agenttrail-guard
$agenttrail-guard guardrails show fs.agent-self-config# everything about it
$agenttrail-guard guardrails set-action fs.agent-self-config warn# block, ask or warn
$agenttrail-guard guardrails allow fs.agent-self-config '<pattern>'# silence one shape
$agenttrail-guard guardrails disable fs.agent-self-config# turn it off

Put these guardrails in front of your agent.

AgentTrail Guard is free and open source. It checks every command and file change against the whole library before your agent runs it, on your machine, with no account.

bash
$npm i -g @agenttrail/guard
Read the source on GitHub