Exact outbound authorization for AI coding agents, starting with
git push. Your agent writes code and runs tests freely; agent-guard makes the outbound intent visible and gives the host a decision before code leaves the machine.
Your agent runs git push origin main. The hook stops it and names the command
that will do it properly. You run that, and see the effect rather than the
command line:
remote: origin (git@github.com:you/project.git)
branch: main
update: fast-forward
remote is at 5863769554d56461ba6f3ee3c1c1ee54985d9a6c
would move it to e52e547c023fa4bd1a7d08577dc4efdca3581278
adds 2 commit(s):
e52e547c023fa4bd1a7d08577dc4efdca3581278
c12504e4bacb12aad388149c4712d0af1df547eb
Push this? [y/N] y
Pushed e52e547c023fa4bd1a7d08577dc4efdca3581278 to origin/main.
git push origin main cannot tell you any of that. Which URL origin is, what
the remote holds right now, which commits it would gain, whether anything would
be discarded — those come from asking the repository and the remote, which is
what agent-guard push does before asking you.
It then performs the push itself, re-resolving the transaction first: if the repository moved while you were reading — your agent committed again, or someone else advanced the remote — the authorization no longer matches and the push is refused rather than performed against a state nobody approved.
Here is that path actually running — the hook refusing the push and naming the command, the effect resolved and approved, the receipt, and a force push refused rather than rerouted:
Nothing in that recording is staged: it is
demos/push-broker/demo.sh driving the real
binaries against a real repository. Run it yourself — it uses a throwaway
repository and a local remote, so it needs no credentials and touches nothing
of yours:
./demos/push-broker/demo.shagent-guard is for developers running AI coding agents — Claude Code, Cursor,
Codex CLI, Aider — who want a narrow, inspectable control at the point where
local code becomes a remote Git change.
Two layers of outbound control, one decision surface:
- Action layer (today): parse recognized Git push forms, conservatively flag embedded Git argv candidates, and evaluate shell, file-write, and HTTP tool calls against policy
- Content layer (experimental, opt-in): detect credentials and PII in tool inputs and outputs before they reach the LLM provider or external API
- Evidence layer: JSONL records for decisions; optional Ed25519-signed receipts only when the Guard owns execution and a signing key is configured
Best fit: solo and small-team devs running coding agents in real workflows. Local-first by default — network export occurs only when the host explicitly configures an outbound action or SIEM destination.
Why now: EU AI Act enforcement begins 2026-08-02. Claude Code's PreToolUse hook has known gaps with MCP tools (#33106). DNS-tunnel credential exfiltration exploits (CVE-2025-55284) are already in the wild. The cost of "the agent did something irreversible" is no longer hypothetical.
The Claude Code plugin, which installs the hook and a starter policy:
npx agent-guard-plugin initThe command-line tools:
cargo install guard-hook --locked
cargo install agent-guard-cli --locked
cargo install guard-verify --lockedagent-guard-cli is what performs the push shown above. From a repository:
mkdir -p ~/.agent-guard
(umask 077; set -C; : > ~/.agent-guard/broker.gitconfig)
agent-guard push --remote origin --branch mainConfig creation refuses to overwrite an existing file; if already configured, keep the existing trusted config and run only the push command.
That is the command the hook names when it stops a push, and it runs as
printed after the host-owned broker config has been created. The config path
defaults to $AGENT_GUARD_BROKER_GIT_CONFIG, then
~/.agent-guard/broker.gitconfig; it is deliberately separate from the
agent-writable repository and ordinary global Git config. The policy defaults
to $AGENT_GUARD_POLICY, then to the one
npx agent-guard-plugin init installs — the same policy that refused the
push. Name another with --policy. Whichever is used is printed in the
preview, so the rules a push was judged by are never left implied.
It prints the preview and asks before doing anything.
As a Rust library:
[dependencies]
agent-guard-sdk = "0.2"From Python — the distribution is agent-guard-python, the import name is
agent_guard:
pip install agent-guard-pythonThe Node binding (@agent-guard/node) is not published; build it from a
checkout with npm ci --prefix crates/agent-guard-node && npm run build --prefix crates/agent-guard-node.
- Source version:
v0.2.8 - Latest published release:
v0.2.8— crates.io, PyPI, npm - Announcement: GitHub Discussions #1
If you are touching the repository itself, use the shared verification entrypoint:
./scripts/verify.sh fullUseful narrower paths:
./scripts/verify.sh rust./scripts/verify.sh lint./scripts/verify.sh python./scripts/verify.sh node
The verification script uses temporary directories for Python build/test work so routine verification does not leave venv_* style residue in the repository root.
The fastest adoption path is the zero-config outbound preset. It covers all five action-layer categories (code egress, package release, artifact egress, remote mutation, destructive shell) with sensible defaults, so you do not have to write your first rule.
cargo install --path crates/guard-hook --locked
guard-hook check \
--policy presets/coding-agent-outbound.yaml \
--agent-id smoke-test < event.jsonA real git push from the agent then surfaces as an ask decision; a git push --force is denied outright; a cargo build passes through with no friction. See presets/README.md for adoption with the Rust SDK, Node binding, or Claude Code PreToolUse hook, and for the contributing guide on new presets.
For a runnable decision preview of the preset — an agent finishing a feature, then the gate firing on git push — use the bundled demo:
npm ci --prefix crates/agent-guard-node
npm run build:debug --prefix crates/agent-guard-node
npm run demo:outbound --prefix crates/agent-guard-nodeIf you prefer a runnable end-to-end demo of the multi-side-effect runtime, the Node side-effect wedge is also wired up:
npm ci --prefix crates/agent-guard-node
npm run build:debug --prefix crates/agent-guard-node
npm run demo:wedge --prefix crates/agent-guard-nodeWhat you should see:
=== agent-guard side-effect wedge ===
[1] shell decision: execute
[2] file decision: execute
[3] http decision: execute
[4] remote publish decision: ask_for_approval
That path is documented in Side-Effect Wedge Demo. For the fastest shell-only proof, use Three-Minute Proof.
The core runtime decision now looks like this:
agent action (outbound moment)
-> agent-guard
-> execute | deny | ask_for_approval | handoff
-> optional guard-owned execution
-> optional Ed25519-signed execution receipt
This is the difference between:
- hoping the model behaves
- and putting an explicit gate in front of every outbound action
Today, the runtime can already own execution for:
- shell / terminal
- file write
- outbound mutation HTTP
Together those three surfaces cover the action-layer categories the preset bundles (code egress, package release, artifact egress, remote mutation, destructive shell).
agent-guard push is the one place where the Guard performs the outbound
action rather than advising on it:
- Policy, on the equivalent command. A push your policy denies never reaches you — being asked to approve what policy already refused teaches people to click through refusals.
- Preview, resolved from the repository and the remote: the URL, both object ids, the update kind, and the commits the remote would gain.
- Your decision, on that.
- Execution, which re-resolves and spends a one-use authorization against what it just resolved. The push pins the approved object id rather than the branch name, and leases the approved remote object id, so neither end can move between your decision and the push.
- A receipt once the broker execution stage is entered, including Git or
authorization refusals.
--receipt <path>persists it. Policy denials, preview failures and a human declining before execution are not execution attempts and do not produce a receipt.
Ordinary non-force pushes of one branch are what it performs today. Force, mirror, remote branch removal, tags and multiple refspecs fail closed, and the hook says so rather than pointing you at a command that would refuse.
The broker treats the checkout as hostile input: it resolves one push URL, copies regular refs and objects into a temporary bare repository, and executes there without repository hooks or config. Keeping the host-owned broker config and credentials away from the agent remains a deployment decision. The Claude Code hook is still fail-open advisory; an agent with its own credential can push without consulting the broker.
Credential isolation lists the code and deployment requirements together and gives you a check that tells you whether the agent can authenticate independently.
The new fixed Linux Docker reference implements host-controlled setup and approval without a new daemon or RPC. Its configuration tests are not isolation proof: native authenticated container acceptance is a separate required gate. The accepted plan keeps Shell in bounded maintenance and credential isolation as the next product milestone. The maintainer will pilot it first; the separate 0.2.8 delivery checkpoint records the cancelled old publication and verified successor delivery, not completion of real pilot feedback.
- One narrow outbound decision: recognized direct
git pushspellings, plumbing-levelgit send-packcalls, and explicitly modeled wrappers are normalized before policy matching, including repository selectors, destructive flags, and force/delete refspec shorthand. Unknown outer commands containing adjacent standalone Git argv tokens are governed by a conservative, explicitly unverified check so they cannot weaken that decision. - Zero-config preset: a copy-able policy that covers the five action-layer categories on day one — no rule-writing required.
- Small integration surface: wrap existing LangChain-style tools or OpenAI-style handlers, or hook into Claude Code's PreToolUse via
guard-hook. No runtime rewrite. - Truthful evidence: decisions are recorded as JSONL; executions can carry a signed receipt when the Guard owns the action and has an explicit key.
- solo and small-team devs running Claude Code / Cursor / Codex CLI / Aider against real codebases
- shell-enabled coding agents that publish, push, deploy, or otherwise produce outbound effects
- teams that want a local forensic decision trail and optional signed execution receipts
- chat-only assistants with no tool execution
- teams looking for a full orchestration framework
- teams expecting a finished enterprise control plane on day one
agent-guard controls the outbound side effect on each tool call. It deliberately does not govern the surrounding autonomous loop — budget caps, verifier gates, retry admission, and JSONL run records are a different failure mode (a 47-retry overnight bill vs. a single rogue git push).
For that layer, see MartinLoop: it wraps autonomous coding agents with budgets, verifier gates, and run records. The two layers compose — MartinLoop decides whether the next attempt is admitted; agent-guard decides whether the side effects inside that attempt are allowed to leave.
What is strong today (action layer):
- the push broker resolves one exact push URL, snapshots branch refs and primary objects into an isolated bare repository, revalidates both object ids at execution, and never loads repository hooks or execution config
- recognized direct Git push entry points and explicitly modeled wrappers are normalized into one policy decision; force, mirror, delete, and destructive refspec forms cannot fall back to a weaker raw string match
- adjacent standalone Git argv candidates under an unknown outer command are conservatively governed at the same decision strength and labeled as unverified; argv inspection cannot establish whether an arbitrary program will execute those arguments
- the broader zero-config preset covers five outbound action categories as an advisory policy, not as a credential-isolated containment boundary
- shell / terminal, file write, and outbound mutation HTTP are the underlying runtime proof surfaces
- HTTP policy rules are method-aware: a rule can carry a
method:constraint (e.g. denyPOST/DELETEto a host) instead of matching the URL alone - normalized runtime decisions, a local single-user approval workflow, JSONL decision records, and optional Ed25519-signed execution receipts are available now
- the SDK already includes policy signing, execution receipts, metrics, anomaly detection, and SIEM export beyond the narrow wedge
What is experimental and opt-in (content layer):
- credential / PII detection on outbound content —
write_filecontent andhttp_requestbody — behind the off-by-defaultcontentfeature, with three enforcement modes (block/mask/warn). See Content layer below. - the same detection on input text (prompts) before it reaches the LLM provider, via the top-level
input_content:policy block andGuard::check_content
What is roadmap (primary product boundary):
- productized host separation for the existing broker path, so its dedicated Git configuration, credentials, SSH setup and signing key are unreachable from the agent rather than merely documented as deployment prerequisites
What to understand before integrating:
- raw runtime APIs expose
execute | deny | ask_for_approval | handoff - adapter
enforceis still strongest on shell-like execution paths today - Bash has the deepest validator path;
read_file/write_filenormalize paths and fail closed on symlink escapes; HTTP policy rules can match on URL and method - Python and Node bindings default to the SDK's platform sandbox selection; both also accept an explicit
backendargument onexecute/run, resolved truthfully (a backend that is not compiled in or not functional yields thenonebackend, never a false isolation claim) - default builds carry no OS sandbox feature and resolve to
none; enabling a compiled backend does not by itself isolate the complete agent's credentials - broader capability coverage is intentionally narrow, not generic
- broader policy workflow and control-plane ideas are future expansion paths, not the phase-one hook
- the advisory shell layer cannot prove arbitrary launcher semantics or stop a process that bypasses the hook; that class-level guarantee requires the broker to run in a credential-isolated deployment
Where agent-guard sits on the OWASP Top 10 for Agentic Applications (ASI01–ASI10). It is an execution-control layer, so it is a primary control for the side-effect risks and a containment backstop for the autonomy ones — not a full-stack agentic-security platform.
- ✅ Primary control: ASI02 Tool Misuse, ASI05 Unexpected Code Execution
- 🟡 Containment / accountability: ASI01 Goal Hijack, ASI03 Privilege Abuse, ASI08 Cascading Failures, ASI09 Human-Agent Trust, ASI10 Rogue Agents
- ⬜ Out of scope (by design): ASI04 supply-chain / MCP scanning, ASI06 memory poisoning, ASI07 inter-agent comms
For Guard-owned executions configured with a signing key, Ed25519 receipts add cryptographic provenance. Decision-only hooks do not create those receipts and remain dependent on the host honoring the decision. Full mapping: Framework Support Matrix §10.
The action layer decides whether a call may leave. The content layer inspects
what leaves with it. It is off by default — opt in with the content
feature flag — and scans three surfaces: write_file content, http_request
body, and host-supplied input text (prompts) via Guard::check_content.
Add a content block to any tool rule:
tools:
http_request:
mode: full_access
content:
mode: block # block | mask | warn
detect: [secrets, pii] # optional; defaults to bothThe three modes:
| Mode | Effect |
|---|---|
block |
Deny the call when sensitive content is detected (SENSITIVE_CONTENT_BLOCKED). |
mask |
Execute a redacted copy — each finding becomes [REDACTED:<label>] — and emit a ContentFinding audit record. |
warn |
Execute unchanged, but emit a ContentFinding audit record. |
For input text the Guard never performs the downstream call, so the host
consumes the outcome directly — configure a top-level input_content: block
and call check_content on the text before forwarding it:
input_content:
mode: mask # block | mask | warn
detect: [secrets, pii] # optional; defaults to bothuse agent_guard_sdk::{Context, Guard};
let guard = Guard::from_yaml_file("policy.yaml")?;
let outcome = guard.check_content(prompt, &Context::default());
if outcome.blocked { /* refuse to forward the prompt */ }
let safe_prompt = outcome.masked_text.as_deref().unwrap_or(prompt);Findings only ever expose the kind of data (e.g. AWS Access Key, Email),
never the raw matched substring — audit records carry labels and counts, not secrets.
Run the example:
cargo run -p agent-guard-sdk --example content_policy --features contentThis is a spike-grade detector set (named patterns + entropy fallback for secrets, regex + Luhn for PII), not a compliance-grade DLP engine. Treat it as a safety net, not the primary control.
- Outbound preset: the zero-config policy for coding-agent users — start here
- Claude Code plugin: one-command install —
/plugin marketplace add XuebinMa/agent-guard, then/plugin install agent-guard@agent-guard - Claude Code PreToolUse hook: wire
guard-hookinto your live Claude Code session manually - Node Quickstart: shortest programmatic path for a new developer
- Side-Effect Wedge Demo: runnable proof of the multi-side-effect runtime
- Secure Shell Tools: first integration when shell is the dominant risk
- Check vs Enforce: when to keep your handler vs when to move execution into
agent-guard - Framework Support Matrix: current Node / Python / Rust adoption surfaces
- User Manual: install, policy basics, and SDK integration
Additional references:
- Latest published release
- Join the Discussion
- Deployment Guide
- Roadmap: what's shipped, partial, and planned
- Documentation Archive
- Documentation Hub
- Claude Code: the
guard-hookPreToolUse adapter is the lowest-friction entry — point one--policyflag at the outbound preset - Node: strongest programmatic surface, with wrappers for LangChain-style tools and OpenAI-style handlers
- Python:
wrap_langchain_tool/wrap_openai_toolare available and the real-package CI matrix runs; the adapter surface remains beta - Rust SDK: most direct integration path for hosts that want explicit control over side-effect decisioning and execution
We welcome security research and contributions. Please see CONTRIBUTING.md for details.
Copyright © 2026 agent-guard team. Distributed under the MIT License.
