Full suites should almost never run. A general request to test, release, or merge does not approve one. Failures, shared-file changes, absent baselines, and uncertain coverage require diagnosis or selection review, never automatic full expansion.
-
Identify affected operator journeys and integration boundaries. Select individual Python/frontend files, exact Rust test names, and catalogued Playwright entries. Keep required browser/device coverage for these journeys, not unrelated features.
-
Collect/list those tests and tell the user the expected count, runtime boundary, rationale, and exclusions before execution. List-only operations are not full runs.
-
Stage new source/test files, then obtain a diff digest:
python scripts/test_selection.py --baseline origin/main --candidate WORKTREE --digest
The digest binds that diff alone. It pins git's object-name abbreviation length, which git otherwise scales to how many objects the local clone holds, so the same baseline and candidate hash identically in a working clone and in CI's fresh one.
-
Write
.github/test-selection.jsonwithchange_digest,reviewed_by,reason,exclusions,expected_tests, and the listspython,frontend,native,playwright. Empty lists explicitly mean no relevant tests at that layer; explain them inexclusions. Frontend paths start withui/src/; Rust names must be exactmodule::testnames. Playwright selectors come from:python scripts/playwright_impact.py catalog --manifest ui/playwright-impact.json
expected_testsrecords counts per layer. Python defaults to 3.12; opt into other supported versions withpython_versionsonly when needed. Setpython_browser: trueonly when the selected Python tests require Chromium.python_node: truerequests Node for cross-language guard/tool tests. -
Validate the plan, commit it with the change, and review CI's selection receipts. Changing the diff or base invalidates the digest. Update and review the selection again; do not merely refresh its digest without considering coverage.
CI validates the plan before test jobs start. It does not run tests for omitted layers, and never defaults empty target arrays to whole suites. Failed selection validation blocks the PR rather than pretending that zero tests proves coverage. The plan is human/agent judgment, not cryptographic proof of adequate coverage. Normal code review must assess its rationale; an agent must not invent approval. Existing branch-protection Python check names remain present: unselected versions only report their omission without installing or executing tests. Invalid selection makes these checks fail, so skipped test layers cannot hide a missing review.
Focused local examples:
poetry run pytest -q tests/test_playwright_impact.py tests/test_test_selection.py
npm --prefix ui test -- src/api/runtimeDefaults.test.ts
npm --prefix ui run test:e2e -- tests/interface.spec.ts --project=mobile-webkit --grep 'configured harness default'Pytest and the npm entry points reject unscoped runs; Playwright's configuration also guards
direct npx playwright test. Real-Core no longer implicitly runs other projects.
These guards prevent accidents, not a determined actor editing/bypassing the tools.
Release tags do not authorize full coverage. Supply reviewed selection and
review_reason through manual dispatch if automatic impact needs review. none
is allowed only with an explicit explanation. Missing baseline/shared paths block
for review; downstream build and sandbox preparation wait for selected coverage.
The daily stable driver automatically supplies conservative catalog coverage for
shared/unmapped changes: desktop/compact interface, small/wide mobile Chromium and
WebKit, real-Core contracts, plus all matched feature areas. It validates that
selection before tagging and retains the existing downstream test gates. Missing
or invalid baselines still stop the daily release. This exception applies only to
daily coverage selection, not routine PR/local runs or exceptional full suites.
Release packaging runs its focused package/updater contract tests and compile
checks, not another complete backend, frontend, or Rust test suite. Package
installation/smoke checks remain required; they are not full product-suite runs.
Only after a fresh explicit user approval, manually dispatch Playwright with
scope=full, full_approval=RUN_FULL_SUITE, and a nonempty review_reason recording
who approved it and why focused coverage is insufficient. PRs and tag pushes cannot
authorize this mode. Full selection via a union of project selectors is rejected.
No scheduled, push-triggered, or retry-triggered full-suite execution is permitted.
For an explicitly approved local full UI run, both NEBULA_FULL_SUITE_APPROVAL
(RUN_FULL_SUITE) and NEBULA_FULL_SUITE_REASON must be set for that one command.
Do not persist these variables in agent profiles or CI defaults. Backend full runs
likewise need explicit user approval; ordinary CI only accepts individual files.