You wrote "never do this" in CLAUDE.md, and Claude Code did it anyway. It is a common complaint about rules files, and Claude Code's own memory docs answer it in one line: CLAUDE.md and auto memory are loaded into every session, and "Claude treats them as context, not enforced configuration." The same page says that to block an action regardless of what Claude decides, you use a PreToolUse hook.
That is the difference between a rule the agent weighs and a rule it cannot talk its way past. This guide walks the layers in between: CLAUDE.md, permission rules, a PreToolUse hook you write yourself, AgentTrail Guard and AgentTrail OS, with what each is good at and where it breaks. The worked example is a command from a real session. The verdicts for the hand-written hook and for AgentTrail Guard come from running them; what a permission rule would do comes from Claude Code's permissions documentation.
Why CLAUDE.md is advice, not a control
CLAUDE.md is a markdown file of instructions. Claude Code loads it at the start of every session, from the working directory and every directory above it. It does not quietly fall out of a long session either: after /compact, Claude re-reads the project-root CLAUDE.md from disk and puts it back into context. Nested CLAUDE.md files in subdirectories load on demand, when Claude reads or edits files under them.
So the problem is not that the model forgets the file. The file is one input among many. On every turn the model weighs your instructions against the task, your latest message and whatever error is on screen, then does what it judges best. The docs say that more specific and concise instructions are followed more consistently. That is good advice, and it describes a likelihood, not a guarantee.
That makes CLAUDE.md the right home for conventions: how to run the tests, which logger to use, where things live. When the model misses one, someone catches it in review. It is the wrong home for a rule that must hold every time, because nothing outside the model checks it. Auto memory, the notes Claude writes for itself, is context too.
The worked example: a rule followed, then reasoned past
In "It's only dev data": the agent that reached for --force-reset we asked Claude Code to split a column in a Prisma schema and just get it synced. The repo's CLAUDE.md asked for a backup before any schema change, and the agent did exactly that: "Backing up first per CLAUDE.md." Then prisma db push refused, because six existing rows had no values for the new required columns, and Prisma's error offered --force-reset as the way out. The agent decided the backup and "only dev data" made a reset safe, and ran this as one command:
npx prisma db push --force-reset 2>&1 | tail -8 && npm run db:seed 2>&1 | tail -5 && sqlite3 prisma/dev.db "select id,email,firstName,lastName from User; select count(*) from \"Order\";"CLAUDE.md worked as designed. The agent followed the instruction it was given, then reasoned its way past one nobody had written. Adding "never run --force-reset" to the file might have changed the outcome, or might not: it would have been one more line weighed against "just get it synced." The rest of this guide is about the layers that do not weigh anything.
Layer 1: permission rules in settings.json
Claude Code's permission rules live in settings.json and take one of three actions: allow approves a call, ask prompts you, and deny blocks it. A rule names a tool and a pattern, such as Bash(npm *) or Edit(*.env). Deny always wins: rules are checked deny, then ask, then allow, and the first match decides, so a narrow allow cannot punch a hole in a broad deny.
The obvious rule for this incident is a prefix deny:
{
"permissions": {
"deny": ["Bash(npx prisma db push:*)"]
}
}It would have stopped the session's command. Claude Code checks each part of a chained command on its own, and if any part is denied, the whole command is blocked, so the reset, the reseed and the query would have been refused together. That is real enforcement, outside the model, and it costs nothing. Now look at what else it does:
- It blocks the safe command too.
:*is a prefix match, so the rule also denies a plainnpx prisma db push, the ordinary sync you want the agent to run. You cannot carve that back out with an allow rule, because deny wins. Moving the rule toaskmeans a prompt on every push, the kind people learn to click through. - It matches how the command starts. The danger is how it ends.
pnpm prisma db push --force-resetandbunx prisma db push --force-resetare different text, and so isnpx prisma migrate reset --force, which drops and rebuilds the schema through another command entirely. Each needs a rule of its own. - Widening the pattern swaps one problem for another. The wider a text pattern gets, the more it matches text that only mentions the flag, such as a commit message.
Permission rules are at their best on coarse boundaries, where the whole tool, command or path is the danger: a tool you never want used, a file like .env you never want touched. For a single flag they are a blunt instrument.
Layer 2: a PreToolUse hook you write yourself
A command hook is a script Claude Code runs at a fixed point in its lifecycle, outside the model. A PreToolUse hook fires before a tool call executes and receives the call as JSON on stdin: the tool name, the tool input (for Bash, the command), the session id and the working directory. Exit 2 blocks the call, and whatever the script writes to stderr goes back to Claude as the reason. Exit 0 means no objection, and the normal permission flow carries on.
Two properties make a hook the right layer for a "never." It runs before any permission-mode check, in every mode, so a hook that denies a call blocks it even in bypassPermissions. And it is code, so the same input gets the same answer every time. Here is a minimal hook for the flag in our incident:
#!/bin/bash
# .claude/hooks/block-force-reset.sh (chmod +x it)
INPUT=$(cat)
COMMAND=$(echo "$INPUT" | jq -r '.tool_input.command // empty')
if echo "$COMMAND" | grep -q -- "--force-reset"; then
echo "Blocked: --force-reset flag is not allowed" >&2
exit 2
fi
exit 0Register it in the project's .claude/settings.json, with a matcher so it only runs for Bash calls:
{
"hooks": {
"PreToolUse": [
{
"matcher": "Bash",
"hooks": [
{
"type": "command",
"command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/block-force-reset.sh"
}
]
}
]
}
}Exit codes can only block or stay silent. For anything softer, a hook can exit 0 and print JSON instead. This answer asks you rather than refusing, which is the right strength for a command like prisma migrate reset that is correct against a scratch database many times a day:
{
"hookSpecificOutput": {
"hookEventName": "PreToolUse",
"permissionDecision": "ask",
"permissionDecisionReason": "prisma migrate reset drops and rebuilds the schema"
}
}permissionDecision belongs inside hookSpecificOutput, not at the top level. Pick one approach per hook: an exit code or JSON, not both.
What it takes to make that hook hold up
We ran the hook above directly, outside Claude Code, feeding it the JSON a PreToolUse hook receives for each command. With jq available, this is what it returned:
exit 2 npx prisma db push --force-reset
exit 2 git commit -m "docs: never use --force-reset"
exit 2 grep -rn -- --force-reset docs/
exit 0 npx prisma db push --accept-data-lossIt blocks the incident, and it blocks the chained one-liner too, because the flag is in the string. It also blocks a commit message and a search that only mention the flag, and it waves through --accept-data-loss, Prisma's other flag for pushing a change that loses data. On a machine without jq, all four exit 0: jq is not found, the command substitution comes back empty, grep searches an empty string, and the hook reports no objection to anything. Turning it into something you can rely on takes all of this:
- A mention is not a run. A hook that blocks
git commit -m "docs: explain rm -rf /"blocks the person documenting the danger, and a hook that blocks routine work gets switched off. The exemption has to key on the verb:grep,git commit -m,echoand acurl --databody handle the words as text and run nothing. - Chained commands. The exemption must not become a bypass. If "starts with
git commit" is enough, thengit commit -m "wip" && npx prisma db push --force-resetwalks straight through. A command is only a mention while every&&,;and|stays inside the quotes, and no$(or backtick appears even inside double quotes, because the shell still runs those. Only single quotes stop every expansion. - Quoting is not the signal. "Ignore anything in quotes" fails the other way:
psql -c "DROP TABLE users;"runs its quoted argument, and so doesecho "rm -rf /" | bash. - Every spelling. Prisma alone has two
db pushflags that can lose data,--accept-data-lossand--force-reset, and a separatemigrate resetthat drops and rebuilds the schema. Every other migration tool has its own. - Near-miss tests. For every pattern, a list of commands it must catch and a list of near misses it must not, run on every change. Without the second list, nobody notices the day a pattern starts blocking
rm -rf ./node_modules. - Failing loudly. A missing dependency should never turn a guardrail into a no-op that nobody hears about.
- Being installed. A hook in the project's
.claude/settings.jsoncovers that repo, and one in~/.claude/settings.jsoncovers one machine. Whatever it depends on, likejqhere, has to be on every laptop that runs it, and it does nothing for an agent that is not Claude Code.
The docs also make a point worth repeating: a hook runs arbitrary shell commands, so read any hook before you install it, including the one below.
Layer 3: AgentTrail Guard, a hook someone maintains
AgentTrail Guard is that list, done once and kept up. In Claude Code it runs as a PreToolUse hook, installed as a plugin, so it never edits the hooks block in your settings.json. Before the agent runs a command or touches a file, the guard checks the call against the guardrail library: 74 rules in 11 packs, each one plain data. Every rule ships examples it must catch and near misses it must not, and CI runs both through the real engine. It runs locally, needs no account, makes no network calls out of the box, and is open source under Apache-2.0.
Here is what the rules return for the commands above and a few more, with default actions. We ran each one through the guard's Claude Code hook and through the engine behind the checker on the library page, and they agreed. When several rules match, the strongest verdict wins, and require_approval is what the guard shows you as ask.
npx prisma db push --force-reset block dd.accept-data-loss
npx prisma db push --accept-data-loss block dd.accept-data-loss
git commit -m "docs: never use --force-reset" no match
grep -rn -- --force-reset docs/ no match
git commit -m "wip" && npx prisma db push --force-reset block dd.accept-data-loss
npx prisma migrate reset --force require_approval dd.migration-reset
git push --force origin main block block-force-push
git commit -m "docs: explain rm -rf /" no match
psql -c "DROP TABLE users;" block block-destructive-sql
cat .env warn se.env-print
rm -rf ./node_modules no match
npm run db:reset no matchReading it top to bottom:
dd.accept-data-losscatches both of Prisma's data-loss flags and the session's whole one-liner. It stays quiet on the two mentions, and still blocks a commit chained to a real reset.dd.migration-resetasks instead of blocking, because a reset is the right move on a scratch database and the guard cannot see which database is configured.block-force-pushblocks a force-push,--force-with-leaseincluded.block-destructive-sqlkeys on the verb:psql -cruns its quoted argument, so the quotes do not exempt it.se.env-printwarns when a dotenv file is printed into the terminal.rm -rf ./node_modulesmatches nothing on purpose:require-approval-rm-rfexempts build directories. A variable or a source directory still asks, anddd.rm-rf-absoluteblocks an absolute path outright.npm run db:resetmatches nothing, and that is a real limit: the guard reads the command, not the script it runs.
You can paste any command into the checker on the guardrail library page and see which rules it matches. And agenttrail-guard scan --agent claude replays the sessions already on your disk through the same rules and reports which would have fired, without uploading anything.
What it does not do, stated plainly:
- It sees one call at a time: no working directory, no connection string, no file contents. It cannot tell a throwaway database from one that matters, so it blocks
--force-reseteverywhere. When you really want one, run it yourself.agenttrail-guard guardrails allowsilences one rule for one command shape, andagenttrail-guard statusprints that line for the rule firing most. - A file rule matches the path, never what a write puts in the file, and the guard does not see what the agent fetches from the web.
- It fails open. If the guard itself breaks, the call goes through Claude Code's normal permission flow. It never answers
allow, so it never skips a prompt you would otherwise see. - A rule that asks reaches you only when your settings do not already allow the whole tool. With a bare
"Bash"in your allow list, the guard's README explains, the command runs with no prompt, andstatustells you how many holds that silences. - It is one developer on one machine. Nothing syncs and nothing is shared.
It also installs for Cursor and Codex CLI, where the hooks see less. In Cursor, reads shown as "Explored", Tab edits and cloud agents are not checked, and the Agent Window can skip hooks. In Codex CLI, nothing runs until you approve the guard's entries in /hooks, and a rule that would ask blocks instead. The README lists every gap.
It runs alongside your own hooks. When several hooks match the same call, Claude Code runs all of them and applies the most restrictive answer, so a hook you wrote for a check only your project needs keeps working next to it.
Layer 4: AgentTrail OS, for a team
The guard checks each call against rules that already exist. It does not turn the mistakes your agents repeat into new rules, and your settings do not reach anyone else's laptop. AgentTrail OS is the hosted layer for that. It records each agent session as OpenTelemetry traces, finds the mistakes your agents repeat in your own history, and suggests the rule. On Team you can backtest a rule against your recorded sessions, with no model calls, before you turn it on, and enforcement (block, warn, or hold for approval in the app) syncs to every developer. Decisions are made on AgentTrail's servers from your organization's policies, and if the check cannot complete, the action does not proceed. Held actions wait in one approval queue, and Business adds an audit trail with export.
AgentTrail OS supports Claude Code and Cursor, and enforcement on Cursor requires Cursor 3.21.16 or later. AgentTrail Guard vs AgentTrail OS covers what differs between the two, and Why your agents keep repeating the same mistakes covers why the repeats matter.
The layers side by side
On AgentTrail OS, backtesting and synced enforcement start on Team.
Which should you use?
More than one. The layers stack, and each catches what the one above it cannot.
- CLAUDE.md, always. Conventions, context, how your team works. Keep it specific and short, and move every "never" into a layer that enforces it.
- Permission rules for coarse boundaries, where the whole tool, command or path is the danger.
- Your own hook for a check only your project needs. Keep it small, give it near-miss tests, and make it fail loudly.
- AgentTrail Guard for the well-known destructive shapes, on every machine you use, without writing and testing the patterns yourself.
- AgentTrail OS once more than one person runs agents and the same rules need to hold everywhere, proven on your own history first.
To install AgentTrail Guard for Claude Code (Node.js 20 or later):
npm i -g @agenttrail/guardagenttrail-guard init --agent claudeinit finishes by running a synthetic rm -rf / through the real engine and showing you it being blocked. Nothing is executed. Add --print to see exactly what it would change first, and see the setup guide for Cursor and Codex CLI.
FAQ
Why does Claude Code ignore CLAUDE.md?
Usually it has weighed it rather than ignored it. Claude Code loads CLAUDE.md into every session as context, not as enforced configuration, so an instruction there competes with the task, your latest message and whatever error is on screen, and it can lose. Specific, concise instructions are followed more consistently, but nothing guarantees it.
How do I stop Claude Code from running a command?
Use a PreToolUse hook or a permission deny rule, not CLAUDE.md. A hook runs outside the model before the tool call, and a hook that denies a call blocks it in every permission mode, bypassPermissions included.
Are Claude Code hooks enough?
A hook is only as good as its matching. A one-line grep blocks harmless mentions such as a commit message, misses other commands that do similar damage, and lets everything through if a dependency like jq is missing. It needs near-miss tests, and its dependencies have to be on every machine that runs the agent.
Does AgentTrail replace CLAUDE.md?
No. Keep CLAUDE.md for conventions and context, where usually is good enough. AgentTrail Guard checks each command and each file the agent touches against tested rules before the call runs. AgentTrail OS finds the mistakes your agents repeat and turns them into rules, and on Team syncs them to every developer.



