The defaults are tuned for a team that trusts its agents and reviews its own PRs. As the stakes, headcount, or formality grow, add these — each is a recipe or a one-line config, in keeping with design.md's rule that heavyweight needs must not complicate the default path.
A gate that floats its own version can change behavior with no repository change. Generated hooks and CI already pin the version that generated them; for belt-and-suspenders, install it as an exact dev dependency and route everything through package scripts:
npm install --save-dev --save-exact rfc2119{ "scripts": { "ci": "npm test && 2119 check" } }Upgrades then arrive as reviewable lockfile diffs, not silent registry drift.
2119 is not a test runner (design.md). CI must run both, as separate steps so failures attribute cleanly:
- run: npm test # do the tests pass?
- run: npx --yes rfc2119@<pinned> check # is every requirement traceably, honestly verified?The workflow init --ci generates includes both steps.
2119 enforces the spec as written — it cannot know when a requirement is weakened (a MUST
softened to SHOULD, a requirement tombstoned, evidence globs narrowed, [manual] added).
That's a Goodhart risk: once coverage is the target, editing the requirement is the cheapest way
to hit it. Make weakening conspicuous and human-gated:
# CODEOWNERS
/specs/ @your-org/spec-owners
/.2119/verdicts/ @your-org/spec-owners
Reviewers should treat any severity downgrade, tombstone, or coverage-tag change in a PR the way they'd treat a CI-config change.
By default, whoever runs 2119 pass writes the verdict — including, potentially, the agent that
wrote the code (the residual risk in the README). For high-consequence requirements, move
verdict-writing to an identity the author cannot impersonate:
- A CI job (or bot account) checks out the PR, runs
2119 review, and dispatches each instruction file to a fresh agent session it controls. - That job records the verdicts and pushes the commit itself.
- Branch protection requires that verdict commits for protected paths come from the bot.
The provenance is the git committer identity on a protected branch — enforced by your git host, unforgeable by the authoring session, and requiring zero new fields in the verdict format.
A bare [review] tag hashes only the requirement's text, so its verdict stands until the
requirement is reworded — appropriate for policy statements, too weak for anything whose truth
lives in code. For critical requirements, always name the evidence:
3. Exports MUST strip other tenants' rows. [review: app/exports/**, lib/tenancy/**]Now implementation edits invalidate the verdict, as they should.
Test-quality hashing covers each annotated test's block and its file's prelude — but not helper modules imported from other files. A shared fixture could change (or be neutered) without invalidating the verdicts that depend on it. If your suite leans on shared helpers, list them:
# .2119.yml
shared_evidence:
- tests/helpers/**
- tests/fixtures/**Their content then joins every test-quality hash. The cost is honest churn: editing a shared helper re-opens every dependent review, which is exactly what should happen.
Fresh context is not fresh framing: reviewers sharing one model family and one instruction
template share blind spots. 2119 review --audit generates adversarial instructions for every
currently-passing verdict — the auditor's job is to construct a mutant under which the
requirement is violated while the tests stay green, and a recorded fail flips the gate. As a
QA cadence: monthly, dispatched to a model from a different provider than your routine
reviewer. Between sweeps, audit individually the requirements that are particularly challenging
or high-consequence — multi-clause invariants, security boundaries, statistical formulas. The
first field deployment's sampled audit found rubber-stamps at a 2-in-12 rate; assume yours has
some too. (audit: "always" in .2119.yml runs audits on every review cycle — off by default,
since it multiplies review cost on every run.)
[verify: <command>] executes shell from spec files — the same trust level as package.json
scripts. That's fine when every spec author is trusted; it is not fine on CI that runs
third-party pull requests. For those repositories, run the gate as:
npx --yes rfc2119@<pinned> check --no-verify[verify] requirements are then surfaced alongside [manual] exemptions instead of executed;
run the full check (with verify) only on trusted branches, or in a job gated on maintainer
approval.
Wall-clock for the full deterministic gate (time npx -y rfc2119 check, Apple
M-series, rfc2119 0.7.0, measured 2026-08; "cold" includes npx package
resolution):
| Corpus | Scale | check wall time |
|---|---|---|
| Production application (private) | ~5,400 tracked files, ~600k lines incl. subprojects; sparse specs, 110 committed verdicts (hash verification in path) | ~1.4s steady, ~2.7s cold |
| Subproject of the same application (private) | 84 files, ~40k lines | ~1.0s |
| Synthetic dense corpus (reproducible) | 100 spec files · 3,000 MUST requirements · 2,000 annotated test files, ~300k lines; judgment layer disabled | ~1.7–2.0s steady |
The spread is the point: a sparse-specced production codebase and a corpus with 3,000 requirements land within a second of each other, because the walk and parse are the cost and both are bounded — spec/annotation volume, not repository size, is what you are budgeting. All three sit well inside the enforced sub-5-second perf requirement.
To regenerate the synthetic corpus: 100 spec files of 5 sections × 6
single-keyword MUST items each; 2,000 test files of ~150 lines, each opening
with one section-level annotation (// 2119: REQ-0NN.M); reviews: false in
.2119.yml. Numbers rot with hardware and versions — re-measure on yours
before relying on them.