Skip to content

Python: [Feature]: A deterministic pre-execution verification middleware for agent actions — AgentDojo v2.2 re-run: ASR=0 / FP=0 (open artifacts, full fix-cycle trajectory inside) #9132

Description

@Lsy1533133

Description

Discussions like #8862 (budget enforcement in AgentLoopMiddleware) and #7853 (sandbox abstraction for tool execution) point at the same gap: agent frameworks today observe actions, but the decision to block a dangerous action before it executes is usually left ad hoc. We built that missing piece and would like to offer it to this community as a candidate middleware contract.

What it is
agent-action-verifier (https://github.com/Lsy1533133/agent-action-verifier) is a deterministic pre-execution enforcement layer: every pending action (tool call, network egress, file write) is checked against the agent's declared plan before execution. The judge is fixed-constant math — no LLM call in the enforcement path, no learned parameters — so it adds microsecond-scale cost per action (P99 = 0.13 ms in the synthetic stress suite; production latency not yet measured) and is deterministic and auditable by construction.

Closed-loop results (each reported number is backed by a checksummed artifact in the repo)

  • AgentDojo official benchmark, real LLM in the loop (97 tasks × 2 rounds, official injection suite, official security() judgment): ASR = 0, FP = 0 in the v2.2 full re-run (FP-rate 95% upper bound ≈3.8% at n=97 benign rounds)
  • Published fix-cycle trajectory FP 16 → 1 → 0 across v1.2 → v2.1 → v2.2 — same-source iteration, failures included, not independent stability trials
  • Cross-model spot check: GLM subset (24 benign + 24 attack): FP = 0, ASR = 0
  • Synthetic stress layer (100,000 seeded scenarios): FN = 0, FP = 0; interception Wilson-95 lower bound 99.99%+ on that synthetic distribution
  • White-box adaptive attacks: 49 cases, 16 adaptive vector families — 43 hard-blocked, 0 bypass; 6 boundary cases documented and not counted as bypasses under the stated threat model

Integration shape (matches your middleware seam)
Wrap the action-dispatch point; the verifier receives (declared plan, pending action) and returns PASS / VETO + rule id. verifier_interface.pyi in the repo specifies the contract; typical adapter is ~10 lines around an AgentLoopMiddleware. Rule families: out-of-plan action, egress breach, scope escalation, tool-consent violation, dangerous value class, arithmetic guard, uninitialised state.

Honest boundaries
Synthetic scenarios are abstracted from publicly disclosed incident categories, not production traffic. Same-source fix cycles ≠ independent stability trials. GLM subset vs full run differ in model and sample size and are not directly comparable. This layer complements monitoring/alignment — it does not replace them.

Verify it yourself
python verify_artifacts.py in the repo recomputes every SHA-256 chain and the Wilson-95 bounds from raw counts — no trust required, stdlib only.

We'd genuinely value feedback on the metrics methodology, and if a deterministic enforcement seam fits the framework's roadmap, if maintainers see a fit, we can align the contract with the framework's middleware design.

Code Sample

Language/SDK

Both

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

agentsUsage: [Issues, PRs], Target: Single agentmiddlewareUsage: [Issues, PRs], Target: middleware related featurespythonUsage: [Issues, PRs], Target: Python

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions