Skip to content

About

AI Agent permission enforcement and sandbox security SDK, extracted from claw-code

Resources

Contributing

Security policy

Stars

5 stars

Watchers

0 watching

Forks

Repository files navigation

agent-guard

Exact outbound authorization for AI coding agents, starting with git push. Your agent writes code and runs tests freely; agent-guard makes the outbound intent visible and gives the host a decision before code leaves the machine.

Version Focus License MSRV

What it looks like

Your agent runs git push origin main. The hook stops it and names the command that will do it properly. You run that, and see the effect rather than the command line:

remote:  origin (git@github.com:you/project.git)
branch:  main
update:  fast-forward
remote is at 5863769554d56461ba6f3ee3c1c1ee54985d9a6c
would move it to e52e547c023fa4bd1a7d08577dc4efdca3581278
adds 2 commit(s):
  e52e547c023fa4bd1a7d08577dc4efdca3581278
  c12504e4bacb12aad388149c4712d0af1df547eb

Push this? [y/N] y

Pushed e52e547c023fa4bd1a7d08577dc4efdca3581278 to origin/main.

git push origin main cannot tell you any of that. Which URL origin is, what the remote holds right now, which commits it would gain, whether anything would be discarded — those come from asking the repository and the remote, which is what agent-guard push does before asking you.

It then performs the push itself, re-resolving the transaction first: if the repository moved while you were reading — your agent committed again, or someone else advanced the remote — the authorization no longer matches and the push is refused rather than performed against a state nobody approved.

Here is that path actually running — the hook refusing the push and naming the command, the effect resolved and approved, the receipt, and a force push refused rather than rerouted:

agent-guard push: a hook refusal that names the broker command, the resolved effect of the push, the approval, the sealed receipt, and a force push refused

Nothing in that recording is staged: it is demos/push-broker/demo.sh driving the real binaries against a real repository. Run it yourself — it uses a throwaway repository and a local remote, so it needs no credentials and touches nothing of yours:

./demos/push-broker/demo.sh

agent-guard is for developers running AI coding agents — Claude Code, Cursor, Codex CLI, Aider — who want a narrow, inspectable control at the point where local code becomes a remote Git change.

Two layers of outbound control, one decision surface:

  • Action layer (today): parse recognized Git push forms, conservatively flag embedded Git argv candidates, and evaluate shell, file-write, and HTTP tool calls against policy
  • Content layer (experimental, opt-in): detect credentials and PII in tool inputs and outputs before they reach the LLM provider or external API
  • Evidence layer: JSONL records for decisions; optional Ed25519-signed receipts only when the Guard owns execution and a signing key is configured

Best fit: solo and small-team devs running coding agents in real workflows. Local-first by default — network export occurs only when the host explicitly configures an outbound action or SIEM destination.

Why now: EU AI Act enforcement begins 2026-08-02. Claude Code's PreToolUse hook has known gaps with MCP tools (#33106). DNS-tunnel credential exfiltration exploits (CVE-2025-55284) are already in the wild. The cost of "the agent did something irreversible" is no longer hypothetical.


Install

The Claude Code plugin, which installs the hook and a starter policy:

npx agent-guard-plugin init

The command-line tools:

cargo install guard-hook --locked
cargo install agent-guard-cli --locked
cargo install guard-verify --locked

agent-guard-cli is what performs the push shown above. From a repository:

mkdir -p ~/.agent-guard
(umask 077; set -C; : > ~/.agent-guard/broker.gitconfig)
agent-guard push --remote origin --branch main

Config creation refuses to overwrite an existing file; if already configured, keep the existing trusted config and run only the push command.

That is the command the hook names when it stops a push, and it runs as printed after the host-owned broker config has been created. The config path defaults to $AGENT_GUARD_BROKER_GIT_CONFIG, then ~/.agent-guard/broker.gitconfig; it is deliberately separate from the agent-writable repository and ordinary global Git config. The policy defaults to $AGENT_GUARD_POLICY, then to the one npx agent-guard-plugin init installs — the same policy that refused the push. Name another with --policy. Whichever is used is printed in the preview, so the rules a push was judged by are never left implied.

It prints the preview and asks before doing anything.

As a Rust library:

[dependencies]
agent-guard-sdk = "0.2"

From Python — the distribution is agent-guard-python, the import name is agent_guard:

pip install agent-guard-python

The Node binding (@agent-guard/node) is not published; build it from a checkout with npm ci --prefix crates/agent-guard-node && npm run build --prefix crates/agent-guard-node.

Release Status

Verify Locally

If you are touching the repository itself, use the shared verification entrypoint:

./scripts/verify.sh full

Useful narrower paths:

  • ./scripts/verify.sh rust
  • ./scripts/verify.sh lint
  • ./scripts/verify.sh python
  • ./scripts/verify.sh node

The verification script uses temporary directories for Python build/test work so routine verification does not leave venv_* style residue in the repository root.


Try The Preset First

The fastest adoption path is the zero-config outbound preset. It covers all five action-layer categories (code egress, package release, artifact egress, remote mutation, destructive shell) with sensible defaults, so you do not have to write your first rule.

cargo install --path crates/guard-hook --locked
guard-hook check \
  --policy presets/coding-agent-outbound.yaml \
  --agent-id smoke-test < event.json

A real git push from the agent then surfaces as an ask decision; a git push --force is denied outright; a cargo build passes through with no friction. See presets/README.md for adoption with the Rust SDK, Node binding, or Claude Code PreToolUse hook, and for the contributing guide on new presets.

For a runnable decision preview of the preset — an agent finishing a feature, then the gate firing on git push — use the bundled demo:

npm ci --prefix crates/agent-guard-node
npm run build:debug --prefix crates/agent-guard-node
npm run demo:outbound --prefix crates/agent-guard-node

If you prefer a runnable end-to-end demo of the multi-side-effect runtime, the Node side-effect wedge is also wired up:

npm ci --prefix crates/agent-guard-node
npm run build:debug --prefix crates/agent-guard-node
npm run demo:wedge --prefix crates/agent-guard-node

What you should see:

=== agent-guard side-effect wedge ===

[1] shell decision: execute
[2] file decision: execute
[3] http decision: execute
[4] remote publish decision: ask_for_approval

That path is documented in Side-Effect Wedge Demo. For the fastest shell-only proof, use Three-Minute Proof.


What It Does

The core runtime decision now looks like this:

agent action (outbound moment)
  -> agent-guard
  -> execute | deny | ask_for_approval | handoff
  -> optional guard-owned execution
  -> optional Ed25519-signed execution receipt

This is the difference between:

  • hoping the model behaves
  • and putting an explicit gate in front of every outbound action

Today, the runtime can already own execution for:

  • shell / terminal
  • file write
  • outbound mutation HTTP

Together those three surfaces cover the action-layer categories the preset bundles (code egress, package release, artifact egress, remote mutation, destructive shell).

The broker path, for git push

agent-guard push is the one place where the Guard performs the outbound action rather than advising on it:

  1. Policy, on the equivalent command. A push your policy denies never reaches you — being asked to approve what policy already refused teaches people to click through refusals.
  2. Preview, resolved from the repository and the remote: the URL, both object ids, the update kind, and the commits the remote would gain.
  3. Your decision, on that.
  4. Execution, which re-resolves and spends a one-use authorization against what it just resolved. The push pins the approved object id rather than the branch name, and leases the approved remote object id, so neither end can move between your decision and the push.
  5. A receipt once the broker execution stage is entered, including Git or authorization refusals. --receipt <path> persists it. Policy denials, preview failures and a human declining before execution are not execution attempts and do not produce a receipt.

Ordinary non-force pushes of one branch are what it performs today. Force, mirror, remote branch removal, tags and multiple refspecs fail closed, and the hook says so rather than pointing you at a command that would refuse.

The broker treats the checkout as hostile input: it resolves one push URL, copies regular refs and objects into a temporary bare repository, and executes there without repository hooks or config. Keeping the host-owned broker config and credentials away from the agent remains a deployment decision. The Claude Code hook is still fail-open advisory; an agent with its own credential can push without consulting the broker.

Credential isolation lists the code and deployment requirements together and gives you a check that tells you whether the agent can authenticate independently.

The new fixed Linux Docker reference implements host-controlled setup and approval without a new daemon or RPC. Its configuration tests are not isolation proof: native authenticated container acceptance is a separate required gate. The accepted plan keeps Shell in bounded maintenance and credential isolation as the next product milestone. The maintainer will pilot it first; the separate 0.2.8 delivery checkpoint records the cancelled old publication and verified successor delivery, not completion of real pilot feedback.


Why Developers Adopt It

  • One narrow outbound decision: recognized direct git push spellings, plumbing-level git send-pack calls, and explicitly modeled wrappers are normalized before policy matching, including repository selectors, destructive flags, and force/delete refspec shorthand. Unknown outer commands containing adjacent standalone Git argv tokens are governed by a conservative, explicitly unverified check so they cannot weaken that decision.
  • Zero-config preset: a copy-able policy that covers the five action-layer categories on day one — no rule-writing required.
  • Small integration surface: wrap existing LangChain-style tools or OpenAI-style handlers, or hook into Claude Code's PreToolUse via guard-hook. No runtime rewrite.
  • Truthful evidence: decisions are recorded as JSONL; executions can carry a signed receipt when the Guard owns the action and has an explicit key.

Best Fit Right Now

  • solo and small-team devs running Claude Code / Cursor / Codex CLI / Aider against real codebases
  • shell-enabled coding agents that publish, push, deploy, or otherwise produce outbound effects
  • teams that want a local forensic decision trail and optional signed execution receipts

Not The First Thing To Reach For

  • chat-only assistants with no tool execution
  • teams looking for a full orchestration framework
  • teams expecting a finished enterprise control plane on day one

Adjacent Layer: Loop Governance

agent-guard controls the outbound side effect on each tool call. It deliberately does not govern the surrounding autonomous loop — budget caps, verifier gates, retry admission, and JSONL run records are a different failure mode (a 47-retry overnight bill vs. a single rogue git push).

For that layer, see MartinLoop: it wraps autonomous coding agents with budgets, verifier gates, and run records. The two layers compose — MartinLoop decides whether the next attempt is admitted; agent-guard decides whether the side effects inside that attempt are allowed to leave.


Current Scope

What is strong today (action layer):

  • the push broker resolves one exact push URL, snapshots branch refs and primary objects into an isolated bare repository, revalidates both object ids at execution, and never loads repository hooks or execution config
  • recognized direct Git push entry points and explicitly modeled wrappers are normalized into one policy decision; force, mirror, delete, and destructive refspec forms cannot fall back to a weaker raw string match
  • adjacent standalone Git argv candidates under an unknown outer command are conservatively governed at the same decision strength and labeled as unverified; argv inspection cannot establish whether an arbitrary program will execute those arguments
  • the broader zero-config preset covers five outbound action categories as an advisory policy, not as a credential-isolated containment boundary
  • shell / terminal, file write, and outbound mutation HTTP are the underlying runtime proof surfaces
  • HTTP policy rules are method-aware: a rule can carry a method: constraint (e.g. deny POST/DELETE to a host) instead of matching the URL alone
  • normalized runtime decisions, a local single-user approval workflow, JSONL decision records, and optional Ed25519-signed execution receipts are available now
  • the SDK already includes policy signing, execution receipts, metrics, anomaly detection, and SIEM export beyond the narrow wedge

What is experimental and opt-in (content layer):

  • credential / PII detection on outbound content — write_file content and http_request body — behind the off-by-default content feature, with three enforcement modes (block / mask / warn). See Content layer below.
  • the same detection on input text (prompts) before it reaches the LLM provider, via the top-level input_content: policy block and Guard::check_content

What is roadmap (primary product boundary):

  • productized host separation for the existing broker path, so its dedicated Git configuration, credentials, SSH setup and signing key are unreachable from the agent rather than merely documented as deployment prerequisites

What to understand before integrating:

  • raw runtime APIs expose execute | deny | ask_for_approval | handoff
  • adapter enforce is still strongest on shell-like execution paths today
  • Bash has the deepest validator path; read_file / write_file normalize paths and fail closed on symlink escapes; HTTP policy rules can match on URL and method
  • Python and Node bindings default to the SDK's platform sandbox selection; both also accept an explicit backend argument on execute / run, resolved truthfully (a backend that is not compiled in or not functional yields the none backend, never a false isolation claim)
  • default builds carry no OS sandbox feature and resolve to none; enabling a compiled backend does not by itself isolate the complete agent's credentials
  • broader capability coverage is intentionally narrow, not generic
  • broader policy workflow and control-plane ideas are future expansion paths, not the phase-one hook
  • the advisory shell layer cannot prove arbitrary launcher semantics or stop a process that bypasses the hook; that class-level guarantee requires the broker to run in a credential-isolated deployment

Threat Coverage — OWASP Agentic Top 10

Where agent-guard sits on the OWASP Top 10 for Agentic Applications (ASI01–ASI10). It is an execution-control layer, so it is a primary control for the side-effect risks and a containment backstop for the autonomy ones — not a full-stack agentic-security platform.

  • ✅ Primary control: ASI02 Tool Misuse, ASI05 Unexpected Code Execution
  • 🟡 Containment / accountability: ASI01 Goal Hijack, ASI03 Privilege Abuse, ASI08 Cascading Failures, ASI09 Human-Agent Trust, ASI10 Rogue Agents
  • ⬜ Out of scope (by design): ASI04 supply-chain / MCP scanning, ASI06 memory poisoning, ASI07 inter-agent comms

For Guard-owned executions configured with a signing key, Ed25519 receipts add cryptographic provenance. Decision-only hooks do not create those receipts and remain dependent on the host honoring the decision. Full mapping: Framework Support Matrix §10.


Content layer (experimental)

The action layer decides whether a call may leave. The content layer inspects what leaves with it. It is off by default — opt in with the content feature flag — and scans three surfaces: write_file content, http_request body, and host-supplied input text (prompts) via Guard::check_content.

Add a content block to any tool rule:

tools:
  http_request:
    mode: full_access
    content:
      mode: block          # block | mask | warn
      detect: [secrets, pii]   # optional; defaults to both

The three modes:

Mode Effect
block Deny the call when sensitive content is detected (SENSITIVE_CONTENT_BLOCKED).
mask Execute a redacted copy — each finding becomes [REDACTED:<label>] — and emit a ContentFinding audit record.
warn Execute unchanged, but emit a ContentFinding audit record.

For input text the Guard never performs the downstream call, so the host consumes the outcome directly — configure a top-level input_content: block and call check_content on the text before forwarding it:

input_content:
  mode: mask               # block | mask | warn
  detect: [secrets, pii]   # optional; defaults to both
use agent_guard_sdk::{Context, Guard};

let guard = Guard::from_yaml_file("policy.yaml")?;
let outcome = guard.check_content(prompt, &Context::default());
if outcome.blocked { /* refuse to forward the prompt */ }
let safe_prompt = outcome.masked_text.as_deref().unwrap_or(prompt);

Findings only ever expose the kind of data (e.g. AWS Access Key, Email), never the raw matched substring — audit records carry labels and counts, not secrets.

Run the example:

cargo run -p agent-guard-sdk --example content_policy --features content

This is a spike-grade detector set (named patterns + entropy fallback for secrets, regex + Luhn for PII), not a compliance-grade DLP engine. Treat it as a safety net, not the primary control.


Fastest Paths

Additional references:


Framework Entry Points

  • Claude Code: the guard-hook PreToolUse adapter is the lowest-friction entry — point one --policy flag at the outbound preset
  • Node: strongest programmatic surface, with wrappers for LangChain-style tools and OpenAI-style handlers
  • Python: wrap_langchain_tool / wrap_openai_tool are available and the real-package CI matrix runs; the adapter surface remains beta
  • Rust SDK: most direct integration path for hosts that want explicit control over side-effect decisioning and execution

Contributing

We welcome security research and contributions. Please see CONTRIBUTING.md for details.

Copyright © 2026 agent-guard team. Distributed under the MIT License.

About

AI Agent permission enforcement and sandbox security SDK, extracted from claw-code

Resources

Contributing

Security policy

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages