Back to the blog
Guardrails·13 min read

CLAUDE.md vs hooks vs AgentTrail: how to stop Claude Code ignoring your rules

Loveneesh Dhir
Loveneesh Dhir
Share
CLAUDE.md vs hooks vs AgentTrail: how to stop Claude Code ignoring your rules
Guardrails

You wrote "never do this" in CLAUDE.md, and Claude Code did it anyway. It is a common complaint about rules files, and Claude Code's own memory docs answer it in one line: CLAUDE.md and auto memory are loaded into every session, and "Claude treats them as context, not enforced configuration." The same page says that to block an action regardless of what Claude decides, you use a PreToolUse hook.

That is the difference between a rule the agent weighs and a rule it cannot talk its way past. This guide walks the layers in between: CLAUDE.md, permission rules, a PreToolUse hook you write yourself, AgentTrail Guard and AgentTrail OS, with what each is good at and where it breaks. The worked example is a command from a real session. The verdicts for the hand-written hook and for AgentTrail Guard come from running them; what a permission rule would do comes from Claude Code's permissions documentation.

Why CLAUDE.md is advice, not a control

CLAUDE.md is a markdown file of instructions. Claude Code loads it at the start of every session, from the working directory and every directory above it. It does not quietly fall out of a long session either: after /compact, Claude re-reads the project-root CLAUDE.md from disk and puts it back into context. Nested CLAUDE.md files in subdirectories load on demand, when Claude reads or edits files under them.

So the problem is not that the model forgets the file. The file is one input among many. On every turn the model weighs your instructions against the task, your latest message and whatever error is on screen, then does what it judges best. The docs say that more specific and concise instructions are followed more consistently. That is good advice, and it describes a likelihood, not a guarantee.

That makes CLAUDE.md the right home for conventions: how to run the tests, which logger to use, where things live. When the model misses one, someone catches it in review. It is the wrong home for a rule that must hold every time, because nothing outside the model checks it. Auto memory, the notes Claude writes for itself, is context too.

The worked example: a rule followed, then reasoned past

In "It's only dev data": the agent that reached for --force-reset we asked Claude Code to split a column in a Prisma schema and just get it synced. The repo's CLAUDE.md asked for a backup before any schema change, and the agent did exactly that: "Backing up first per CLAUDE.md." Then prisma db push refused, because six existing rows had no values for the new required columns, and Prisma's error offered --force-reset as the way out. The agent decided the backup and "only dev data" made a reset safe, and ran this as one command:

npx prisma db push --force-reset 2>&1 | tail -8 && npm run db:seed 2>&1 | tail -5 && sqlite3 prisma/dev.db "select id,email,firstName,lastName from User; select count(*) from \"Order\";"

CLAUDE.md worked as designed. The agent followed the instruction it was given, then reasoned its way past one nobody had written. Adding "never run --force-reset" to the file might have changed the outcome, or might not: it would have been one more line weighed against "just get it synced." The rest of this guide is about the layers that do not weigh anything.

Layer 1: permission rules in settings.json

Claude Code's permission rules live in settings.json and take one of three actions: allow approves a call, ask prompts you, and deny blocks it. A rule names a tool and a pattern, such as Bash(npm *) or Edit(*.env). Deny always wins: rules are checked deny, then ask, then allow, and the first match decides, so a narrow allow cannot punch a hole in a broad deny.

The obvious rule for this incident is a prefix deny:

{
  "permissions": {
    "deny": ["Bash(npx prisma db push:*)"]
  }
}

It would have stopped the session's command. Claude Code checks each part of a chained command on its own, and if any part is denied, the whole command is blocked, so the reset, the reseed and the query would have been refused together. That is real enforcement, outside the model, and it costs nothing. Now look at what else it does:

  • It blocks the safe command too. :* is a prefix match, so the rule also denies a plain npx prisma db push, the ordinary sync you want the agent to run. You cannot carve that back out with an allow rule, because deny wins. Moving the rule to ask means a prompt on every push, the kind people learn to click through.
  • It matches how the command starts. The danger is how it ends. pnpm prisma db push --force-reset and bunx prisma db push --force-reset are different text, and so is npx prisma migrate reset --force, which drops and rebuilds the schema through another command entirely. Each needs a rule of its own.
  • Widening the pattern swaps one problem for another. The wider a text pattern gets, the more it matches text that only mentions the flag, such as a commit message.

Permission rules are at their best on coarse boundaries, where the whole tool, command or path is the danger: a tool you never want used, a file like .env you never want touched. For a single flag they are a blunt instrument.

Layer 2: a PreToolUse hook you write yourself

A command hook is a script Claude Code runs at a fixed point in its lifecycle, outside the model. A PreToolUse hook fires before a tool call executes and receives the call as JSON on stdin: the tool name, the tool input (for Bash, the command), the session id and the working directory. Exit 2 blocks the call, and whatever the script writes to stderr goes back to Claude as the reason. Exit 0 means no objection, and the normal permission flow carries on.

Two properties make a hook the right layer for a "never." It runs before any permission-mode check, in every mode, so a hook that denies a call blocks it even in bypassPermissions. And it is code, so the same input gets the same answer every time. Here is a minimal hook for the flag in our incident:

#!/bin/bash
# .claude/hooks/block-force-reset.sh  (chmod +x it)
INPUT=$(cat)
COMMAND=$(echo "$INPUT" | jq -r '.tool_input.command // empty')

if echo "$COMMAND" | grep -q -- "--force-reset"; then
  echo "Blocked: --force-reset flag is not allowed" >&2
  exit 2
fi

exit 0

Register it in the project's .claude/settings.json, with a matcher so it only runs for Bash calls:

{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Bash",
        "hooks": [
          {
            "type": "command",
            "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/block-force-reset.sh"
          }
        ]
      }
    ]
  }
}

Exit codes can only block or stay silent. For anything softer, a hook can exit 0 and print JSON instead. This answer asks you rather than refusing, which is the right strength for a command like prisma migrate reset that is correct against a scratch database many times a day:

{
  "hookSpecificOutput": {
    "hookEventName": "PreToolUse",
    "permissionDecision": "ask",
    "permissionDecisionReason": "prisma migrate reset drops and rebuilds the schema"
  }
}

permissionDecision belongs inside hookSpecificOutput, not at the top level. Pick one approach per hook: an exit code or JSON, not both.

What it takes to make that hook hold up

We ran the hook above directly, outside Claude Code, feeding it the JSON a PreToolUse hook receives for each command. With jq available, this is what it returned:

exit 2   npx prisma db push --force-reset
exit 2   git commit -m "docs: never use --force-reset"
exit 2   grep -rn -- --force-reset docs/
exit 0   npx prisma db push --accept-data-loss

It blocks the incident, and it blocks the chained one-liner too, because the flag is in the string. It also blocks a commit message and a search that only mention the flag, and it waves through --accept-data-loss, Prisma's other flag for pushing a change that loses data. On a machine without jq, all four exit 0: jq is not found, the command substitution comes back empty, grep searches an empty string, and the hook reports no objection to anything. Turning it into something you can rely on takes all of this:

  • A mention is not a run. A hook that blocks git commit -m "docs: explain rm -rf /" blocks the person documenting the danger, and a hook that blocks routine work gets switched off. The exemption has to key on the verb: grep, git commit -m, echo and a curl --data body handle the words as text and run nothing.
  • Chained commands. The exemption must not become a bypass. If "starts with git commit" is enough, then git commit -m "wip" && npx prisma db push --force-reset walks straight through. A command is only a mention while every &&, ; and | stays inside the quotes, and no $( or backtick appears even inside double quotes, because the shell still runs those. Only single quotes stop every expansion.
  • Quoting is not the signal. "Ignore anything in quotes" fails the other way: psql -c "DROP TABLE users;" runs its quoted argument, and so does echo "rm -rf /" | bash.
  • Every spelling. Prisma alone has two db push flags that can lose data, --accept-data-loss and --force-reset, and a separate migrate reset that drops and rebuilds the schema. Every other migration tool has its own.
  • Near-miss tests. For every pattern, a list of commands it must catch and a list of near misses it must not, run on every change. Without the second list, nobody notices the day a pattern starts blocking rm -rf ./node_modules.
  • Failing loudly. A missing dependency should never turn a guardrail into a no-op that nobody hears about.
  • Being installed. A hook in the project's .claude/settings.json covers that repo, and one in ~/.claude/settings.json covers one machine. Whatever it depends on, like jq here, has to be on every laptop that runs it, and it does nothing for an agent that is not Claude Code.

The docs also make a point worth repeating: a hook runs arbitrary shell commands, so read any hook before you install it, including the one below.

Layer 3: AgentTrail Guard, a hook someone maintains

AgentTrail Guard is that list, done once and kept up. In Claude Code it runs as a PreToolUse hook, installed as a plugin, so it never edits the hooks block in your settings.json. Before the agent runs a command or touches a file, the guard checks the call against the guardrail library: 74 rules in 11 packs, each one plain data. Every rule ships examples it must catch and near misses it must not, and CI runs both through the real engine. It runs locally, needs no account, makes no network calls out of the box, and is open source under Apache-2.0.

Here is what the rules return for the commands above and a few more, with default actions. We ran each one through the guard's Claude Code hook and through the engine behind the checker on the library page, and they agreed. When several rules match, the strongest verdict wins, and require_approval is what the guard shows you as ask.

npx prisma db push --force-reset                         block             dd.accept-data-loss
npx prisma db push --accept-data-loss                    block             dd.accept-data-loss
git commit -m "docs: never use --force-reset"            no match
grep -rn -- --force-reset docs/                          no match
git commit -m "wip" && npx prisma db push --force-reset  block             dd.accept-data-loss
npx prisma migrate reset --force                         require_approval  dd.migration-reset
git push --force origin main                             block             block-force-push
git commit -m "docs: explain rm -rf /"                   no match
psql -c "DROP TABLE users;"                              block             block-destructive-sql
cat .env                                                 warn              se.env-print
rm -rf ./node_modules                                    no match
npm run db:reset                                         no match

Reading it top to bottom:

  • dd.accept-data-loss catches both of Prisma's data-loss flags and the session's whole one-liner. It stays quiet on the two mentions, and still blocks a commit chained to a real reset.
  • dd.migration-reset asks instead of blocking, because a reset is the right move on a scratch database and the guard cannot see which database is configured.
  • block-force-push blocks a force-push, --force-with-lease included.
  • block-destructive-sql keys on the verb: psql -c runs its quoted argument, so the quotes do not exempt it.
  • se.env-print warns when a dotenv file is printed into the terminal.
  • rm -rf ./node_modules matches nothing on purpose: require-approval-rm-rf exempts build directories. A variable or a source directory still asks, and dd.rm-rf-absolute blocks an absolute path outright.
  • npm run db:reset matches nothing, and that is a real limit: the guard reads the command, not the script it runs.

You can paste any command into the checker on the guardrail library page and see which rules it matches. And agenttrail-guard scan --agent claude replays the sessions already on your disk through the same rules and reports which would have fired, without uploading anything.

What it does not do, stated plainly:

  • It sees one call at a time: no working directory, no connection string, no file contents. It cannot tell a throwaway database from one that matters, so it blocks --force-reset everywhere. When you really want one, run it yourself. agenttrail-guard guardrails allow silences one rule for one command shape, and agenttrail-guard status prints that line for the rule firing most.
  • A file rule matches the path, never what a write puts in the file, and the guard does not see what the agent fetches from the web.
  • It fails open. If the guard itself breaks, the call goes through Claude Code's normal permission flow. It never answers allow, so it never skips a prompt you would otherwise see.
  • A rule that asks reaches you only when your settings do not already allow the whole tool. With a bare "Bash" in your allow list, the guard's README explains, the command runs with no prompt, and status tells you how many holds that silences.
  • It is one developer on one machine. Nothing syncs and nothing is shared.

It also installs for Cursor and Codex CLI, where the hooks see less. In Cursor, reads shown as "Explored", Tab edits and cloud agents are not checked, and the Agent Window can skip hooks. In Codex CLI, nothing runs until you approve the guard's entries in /hooks, and a rule that would ask blocks instead. The README lists every gap.

It runs alongside your own hooks. When several hooks match the same call, Claude Code runs all of them and applies the most restrictive answer, so a hook you wrote for a check only your project needs keeps working next to it.

Layer 4: AgentTrail OS, for a team

The guard checks each call against rules that already exist. It does not turn the mistakes your agents repeat into new rules, and your settings do not reach anyone else's laptop. AgentTrail OS is the hosted layer for that. It records each agent session as OpenTelemetry traces, finds the mistakes your agents repeat in your own history, and suggests the rule. On Team you can backtest a rule against your recorded sessions, with no model calls, before you turn it on, and enforcement (block, warn, or hold for approval in the app) syncs to every developer. Decisions are made on AgentTrail's servers from your organization's policies, and if the check cannot complete, the action does not proceed. Held actions wait in one approval queue, and Business adds an audit trail with export.

AgentTrail OS supports Claude Code and Cursor, and enforcement on Cursor requires Cursor 3.21.16 or later. AgentTrail Guard vs AgentTrail OS covers what differs between the two, and Why your agents keep repeating the same mistakes covers why the repeats matter.

The layers side by side

Enforced or advisory
CLAUDE.mdAdvisory
Permission rulesEnforced
Your own hookEnforced
AgentTrail GuardEnforced
AgentTrail OSEnforced
Decides by
CLAUDE.mdThe model's judgment
Permission rulesA tool name and a pattern
Your own hookWhatever your script checks
AgentTrail Guard74 rules in 11 packs
AgentTrail OSYour workspace's policies
Maintained by
CLAUDE.mdYou
Permission rulesYou
Your own hookYou
AgentTrail GuardAgentTrail, in the open
AgentTrail OSThe same library, plus your own rules
What it sees
CLAUDE.mdNothing: the model reads it
Permission rulesA tool name and its arguments
Your own hookThe tool call, as JSON
AgentTrail GuardOne command or file path
AgentTrail OSEach action, and every session as traces
Scope
CLAUDE.mdA project, user or org file
Permission rulesThe settings file it is in
Your own hookThe settings file it is in
AgentTrail GuardOne machine, every repo
AgentTrail OSEvery developer (Team)
Tested
CLAUDE.mdNo: it is advice
Permission rulesOnly if you test it
Your own hookOnly if you test it
AgentTrail GuardMust-catch and near-miss examples per rule
AgentTrail OSBacktested on your sessions (Team)

On AgentTrail OS, backtesting and synced enforcement start on Team.

Which should you use?

More than one. The layers stack, and each catches what the one above it cannot.

  • CLAUDE.md, always. Conventions, context, how your team works. Keep it specific and short, and move every "never" into a layer that enforces it.
  • Permission rules for coarse boundaries, where the whole tool, command or path is the danger.
  • Your own hook for a check only your project needs. Keep it small, give it near-miss tests, and make it fail loudly.
  • AgentTrail Guard for the well-known destructive shapes, on every machine you use, without writing and testing the patterns yourself.
  • AgentTrail OS once more than one person runs agents and the same rules need to hold everywhere, proven on your own history first.

To install AgentTrail Guard for Claude Code (Node.js 20 or later):

npm i -g @agenttrail/guard
agenttrail-guard init --agent claude

init finishes by running a synthetic rm -rf / through the real engine and showing you it being blocked. Nothing is executed. Add --print to see exactly what it would change first, and see the setup guide for Cursor and Codex CLI.

FAQ

Why does Claude Code ignore CLAUDE.md?

Usually it has weighed it rather than ignored it. Claude Code loads CLAUDE.md into every session as context, not as enforced configuration, so an instruction there competes with the task, your latest message and whatever error is on screen, and it can lose. Specific, concise instructions are followed more consistently, but nothing guarantees it.

How do I stop Claude Code from running a command?

Use a PreToolUse hook or a permission deny rule, not CLAUDE.md. A hook runs outside the model before the tool call, and a hook that denies a call blocks it in every permission mode, bypassPermissions included.

Are Claude Code hooks enough?

A hook is only as good as its matching. A one-line grep blocks harmless mentions such as a commit message, misses other commands that do similar damage, and lets everything through if a dependency like jq is missing. It needs near-miss tests, and its dependencies have to be on every machine that runs the agent.

Does AgentTrail replace CLAUDE.md?

No. Keep CLAUDE.md for conventions and context, where usually is good enough. AgentTrail Guard checks each command and each file the agent touches against tested rules before the call runs. AgentTrail OS finds the mistakes your agents repeat and turns them into rules, and on Team syncs them to every developer.

Stop babysitting your AI agents.

AgentTrail catches the risky actions your agents take and turns each repeat into an enforced guardrail. Get started with the open-source guard.

Related reading