Owner: DreamLab AI Status: Active Date: 2026-09-14 Version: 1.0 Governed by: ADR-2010, ADR-2011 Bounded context: DDD: Augmentation Conditions Source analysis: Judgment Broker Augmentation Audit against arXiv 2609.12482 (CIVIC-AI 2026, When Does AI Augment Work?) Method: build-with-quality (EDD → TDD, executed evidence, cross-family audit)
The Judgment Broker loop is cryptographically complete and semantically hollow. Every human decision is a signed, supersession-chained Nostr event, but the decision card never shows the proposed change, the recorded rationale is fabricated by the UI, the boundary between agent and human is set by the requesting agent's own risk declaration, the human never learns whether an approval took effect, and nothing in the estate measures whether the humans doing the reviewing stay capable of doing it.
This PRD adopts the six augmentation conditions from arXiv 2609.12482 as the canon's audit lens for human judgement, and lands seven changes across four repositories so that the estate can be graded against them from code rather than declared compliant.
| Goal | Outcome |
|---|---|
| G1 Audit lens | The six conditions are rows in the compatibility matrix and the acceptance frame for closeout packages CP-05 and CP-07 |
| G2 Non-vacuous verification | A 31403 records a judgement a human formed with the proposal in view, in their own words |
| G3 Task-set boundary | Escalation posture derives from operator-declared task properties, never solely from the requesting agent's self-tier |
| G4 Closed recovery loop | A human sees whether their decision applied; a stalled case ages visibly; a restarted actor recovers its pending cases |
| G5 Intent kept | Declared intent is persisted and compared against the act, and that comparison feeds HITL Precision |
| G6 Humans instrumented | Reviewer time, volume and override rate are measured; juniors get substantive review work; reviewers keep exposure to agent failure |
| G7 Continue during outage | An operator can execute an approved action manually with a signed receipt when the mesh is down |
| # | Condition | Layer | Grading question |
|---|---|---|---|
| C1 | Durable net value | snapshot | Is verification, exception and repair effort counted against the productivity claim? |
| C2 | Meaningful human control | snapshot | Does the reviewer have the competence, time, information and authority to detect and override, and can they continue during unavailability? |
| C3 | Accountability and recovery | snapshot | Are authority, provenance, escalation and fallback explicitly assigned and observable? |
| C4 | Deepening learning | longitudinal | Is human learning measured, not only agent learning? |
| C5 | Career pathways | longitudinal | Do junior reviewers get substantive review work? |
| C6 | Job purpose | longitudinal | Is the human's judgement captured, not only their click? |
Boundary properties: verifiability, reversibility, stakes. The paper's design rule: assign agents work where gains scale, actions reverse and errors stay inspectable; retain human authority where decisions set goals, validity, interpretation or consequential use.
Each FR names its owning repository per ADR-2006 (canon owns the cross-repo view, substrates own implementation) and the expectation ids that prove it.
docs/architecture/compatibility-matrix.mdgains a Augmentation conditions table with one row per condition per substrate, each cell a status (absent | partial | measured) with afile:linecitation.docs/estate-review/closeout/README.mdCP-05 and CP-07 acceptance columns reference C1–C6 by id.docs/terminology.mddefines augmentation condition, task-property triple, vacuous verification, calibration sample.- The judgment-broker PRD and DDD carry an amendment pointing here.
- The forum
ActionRowrendersActionRequest.fields(the proposed change) andcontext_urlin full, with the agent'sreasoning, below the reviewer's controls. The agent'srisk_tierandconfidencerender only after the reviewer has opened the rationale input, or below the controls, never above Approve. - The forum client collects a typed human rationale. For effective tier
highorcriticalthe rationale is mandatory (minimum 20 characters); the publish button is disabled until present. The published 31403reasoningis the human's text verbatim. The string"Human {action} via governance UI"is removed from the codebase. - VisionClaw
AcspCaseQueuerenders the full proposal payload, proposal URN, reasoning hash, and generation provenance where present (model,source_excerpt,confidenceonly when not the hardcoded default), collects a typed rationale under the same rule, and never fabricates a rationale. - VisionClaw replaces the hardcoded
confidence: 0.5withNoneso that absence renders as absence (mirrors the D7 honesty rule).
nostr-bbs-coregainsTaskProperties { verifiability: Inspectable|Partial|Opaque, reversibility: Reversible|Compensable|Irreversible, stakes: Bounded|Significant|Critical }serialised as tags onPanelDefinition(panel default, operator-set) and optionally onActionRequest(may only tighten, never loosen, the panel default).effective_tier(panel, request)is a pure function:IrreversibleorCriticalstakes ⇒ at leasthigh;Opaqueverifiability ⇒ at leastmediumand never member-suppressed; otherwise the agent's declared tier bounded below by the panel default. The relay stores the effective tier onbroker_casesand the client reads only the effective tier.- The relay enforces its advertised
ESCALATION_DEFAULT_TIER/ESCALATION_DEFAULT_POSTURE: an unlabelled request receives the default tier, and a request whose effective tier ishighorcriticalcannot be resolved by anything other than a human 31403. - agentbox
governance_request_actionacceptstask_propertiesand derives a default from the manifestauthority_class(zero-tolerance⇒Irreversible,recoverable⇒Compensable), and the authority gate stamps the triple on its own 31402.
- The forum auth worker exposes
POST /api/governance/receipts/{response_event_id}/application(NIP-98, registered agent or admin) accepting{stage: consumer-received|applied|not-applied|applied-manually, acknowledgement}; the relay-sidegovernance_receiptsrow advances to that stage atomically and the decision chain in the UI shows it. - agentbox
broker-bridgeposts the receipt afterApplicationReceiptStore.beginandfinish; failure to post is journalled, never silent. - The relay cron computes pending-case age; a case older than the panel's
max_pending_hours(default 72) receives anescalated-on-agereceipt and the UI badges it. Age is visible on every pending card (created_atdifferenced client-side). - agentbox authority-gate denials are appended to the execution journal as
authority.denyrecords with stage and reason, readable at/v1/agent-events. - VisionClaw
ElevationActorgainsOPEN_CASE_TTL(14 days, same asdecision_elevation_actor.rs), boot reconciliation ofpendingrows, and anexpiredreceipt.
kpi_agent_eventsandNewAgentTrajectorygainintent TEXT NULL; the hub tap persists the envelope'sintentverbatim and never synthesises one.GET /api/tracereturnsintentand a per-rowintent_match: true|false|nullcomputed by comparing the declared operation and target with the recordedaction_type_nameandtarget_urn.kpi_compute.rsHITL Precision is computed from decided cases: an escalation is warranted when the human outcome differs from the agent's requested action (reject, amend, delegate, or approve with a rationale that names a change) or whenintent_match == false. Theawaiting_data_sourcestub is removed; the KPI reports its denominator.- Whelk gate rejections carry
decided_by: system:whelk-gateand are excluded from the Trust Variance human-outcome series.
- Forum auth worker
GET /api/governance/reviewers(admin NIP-98) returns per-reviewer: decisions, median and p90 time-to-decision, override rate (decisions ≠ agent's request), superseded count, calibration-sample catch rate. - Reviewer tier via Delegate. A governance role
reviewerexists. An admin publishesDelegate { to: <did> }on a case; the relay then admits a 31403 from that delegatee for that case only, with the delegatee's outcome recorded and the admin's delegation preserved in the chain. Reviewers otherwise remain read-only. - Calibration sampling.
PanelDefinition.calibration_sample_rate(default 0.1). A keyed selection — HMAC-SHA256 of the request id under a relay-held secret, compared against the rate — marks aLow/Reversiblerequestcalibration_sample: true; the key is never published, so the requesting agent (which chooses the request id) cannot search for an id the sampler never selects; such requests are shown to reviewers rather than suppressed. The reviewer telemetry reports the share of sampled requests that received a decision. - Seeded probes. An
ActionRequestmay carry tagprobe=<sha256>; the client hides the tag until decided; telemetry counts probes rejected as the catch rate. Probes are published only by a registered probe agent named in panel metadata. Known limit (2026-09-15): the tag is on the signed event, so the measured agent can see it; the tagless design in the DDD §9.5 supersedes this once implemented. - VisionClaw
CaseViewgainscreatedAt; the case queue shows age and sorts oldest first. - The dream-cycle ledger schema (VisionFlow
docs/dream-cycle/LEDGER.md, agentboxdream-ledger.js) gainsReviewerandReview-minutescolumns, populated from the PR merge event.
- agentbox MCP tool
governance_manual_continue { case_id, executed_by, evidence }records an operator-executed action: it writes anapplied-manuallyapplication receipt bound to the approved operation digest, mints a PROV-O activity withexecuted_bya human DID, and posts the receipt to the forum. - The relay accepts
applied-manuallyonly from an admin pubkey and only for a case alreadyDecided: Approve. - The authority gate's
no-decision-surfacedeny path returns a structured hint naming the manual-continuation tool so an operator learns the option at the moment of denial.
- No fabricated human text. No surface may write a rationale, intent or confidence the human or agent did not author. Absence renders as absence.
- Backward compatibility. New tags and columns are additive; legacy events without task properties parse to the panel default; legacy receipts without application stages remain valid.
- Evidence. Every FR ships with an
EXP-AC-NNNexpectation, executed evidence with command, raw output, timestamp and git SHA, an auditor on a different model family, and a stabilising test named instabilized_by. - Security gate.
deepsec-gate.sh --diffruns per repository on the change set; exit 78 is recorded as SKIPPED, never passed.
| Milestone | Delivers | Exit |
|---|---|---|
| M1 Canon | FR1, this PRD, ADR-2010, ADR-2011, DDD, EXP-AC-001..007 | Matrix table present; ledger index regenerated |
| M2 Substrates | FR2–FR7 on branches feat/augmentation-conditions in each repo |
Tests green; evidence audited; deepsec receipts |
| M3 Publish | nostr-bbs-core and nostr-bbs-mesh republished; docs updated in all four repos |
crates.io versions resolvable; diagram docs re-cited |
| M4 Live | Probe suite run against the live edge forum (relay, auth API) | Receipt of each probe; deploy status recorded honestly |
- Multi-relay federation of receipts.
- Payment enforcement.
- A precedent system (deleted per agentbox; when it returns it must satisfy C4 by design).
- Replacing
RiskTier— it remains as the agent's declaration; only its authority changes.
- Kinds 31400–31405 unchanged; new data rides tags and
outcome_detail. - Receipt stages:
signed → relay-accepted → projection-committed → consumer-received → applied | not-applied | applied-manually, plusescalated-on-ageandexpiredas side receipts. - Identity:
did:nostr:<hex>for humans and agents;system:whelk-gatefor reasoner outcomes.