medarc_verifiers/: Core Python package (CLI entrypoints, parsers, rewards, orchestration utilities).environments/<env>/: Individual Verifiers environments (each is a small Python package with<env>.pyand its ownpyproject.toml).configs/: TOML configs formedarc-eval bench, endpoint registries, and environment/judge configs.docs/: Usage docs formedarc-evaland related workflows.tests/:pytestsuite.
- IMPORTANT: Read
docs/medarc-verifiers-architecture.mdbefore writing or modifying any code. - Quick workflow: eval → process → winrate
- eval outputs:
runs/evals/<model>/<env>/<variant>/... - processed parquet:
runs/processed/<model>/<env>.parquet+runs/processed/env_index.json - winrate outputs:
runs/processed/winrate/latest.jsonandruns/processed/winrate/latest.csv
- eval outputs:
medarc-evalCLI entrypoint/router: (medarc_verifiers/cli/main.py; docs:docs/medarc-eval.md)medarc-orchestrateCLI entrypoint: (medarc_verifiers/orchestrate/cli.py; docs:docs/medarc-orchestrate.md)- Old YAML-runner
runs/rawartifacts must be converted withscripts/convert_legacy_raw_runs.pybefore processing. - Environment
load_environment()params become CLI flags (seemedarc-eval <env> --help). - Environment authoring utilities (used by
environments/*):- parsing/prompts:
medarc_verifiers/parsers/,medarc_verifiers/prompts.py(XML preferred; BOXED supported) - MCQ grading:
medarc_verifiers/rewards/multiple_choice_accuracy.py - deterministic shuffling:
medarc_verifiers/utils/randomize_multiple_choice.py(useshuffle_seed) - judging:
medarc_verifiers/judging/,medarc_verifiers/utils/judge_helpers.py - multi-judge path:
medarc_verifiers/judging/multi_judge.py(judge cache is namespaced bybase_url::modelto avoid collisions)
- parsing/prompts:
uv venv --python 3.12 && source .venv/bin/activate: Create/activate a local venv.uv pip install -e .: Installmedarc-verifiersin editable mode.vf-install <env>: Install an environment fromenvironments/<env>/in editable mode.uv run medarc-eval <ENV> -m <MODEL> -n 5: Run a small evaluation.uv run medarc-eval bench --config configs/medmarks-smoke.toml: Run a batch evaluation from a TOML config.uv run pytest tests/: Run the full test suite.uv run ruff check medarc_verifiers/ && uv run ruff format medarc_verifiers/: Lint/format.
- Create a new environment package:
prime env init my-new-env(createsenvironments/my_new_env/). - Ensure the environment
pyproject.tomlhas[tool.prime.environment]with a correctloader(e.g.,my_new_env:load_environment). - Smoke test:
vf-install my-new-env && medarc-eval my-new-env -m gpt-4.1-mini -n 5.
- Python
>=3.11; prefer type hints for new code. - Formatting/linting: Ruff (configured in
pyproject.toml, line length120). - Naming:
snake_casefor functions/vars,PascalCasefor classes,test_*.pyfor tests.
- Frameworks:
pytest+pytest-asynciofor async code. - Example:
uv run pytest tests/test_cli/test_process_winrate.py -k winrate.
- Commit messages: Use a single short sentence (imperative mood, no period). Example: "Remove duplicate cli_env_args module"
- PRs: Short, scoped subjects with clear descriptions.
- Never commit secrets; use env vars like
OPENAI_API_KEY,PRIME_API_KEY, andPRIME_TEAM_ID. - Prime Inference usage reporting:
MEDARC_INCLUDE_USAGE=true/falseormedarc-eval ... --include-usage/--no-include-usage. - Keep local evaluation artifacts out of git (e.g.,
outputs/).