Micro-benchmark suites for Haxe. Write measures once, run them across targets with travix, compare commits, and publish charts.
haxelib install why-benchkit
lix install haxelib:why-benchkitRequires Haxe 4.3+.
- Depend on
why-benchkitand install your project dependencies. - Add a
bench.hxmlwith-lib why-benchkitand a-mainthat callsRunner.run:
# bench.hxml
-cp tests
-lib why-benchkit
-main BenchSuiteimport why.benchkit.Runner;
class BenchSuite {
static function main() {
Runner.run([
new MySuite(),
]);
}
}
@:name("my_lib") // optional: defaults to class name
class MySuite {
public function new() {}
// Omit @:warmup / @:iterations for adaptive measure (see below).
@:warmup(50) // optional: fixed warmup count
@:iterations(100000) // optional: fixed timed iterations
@:name("do-work") // optional: defaults to method name
public function doWork() {
return work(); // return a value so DCE cannot erase the work
}
}Public instance methods on each suite are discovered as measures.
- Run across targets:
haxelib run why-benchkit run --targets interp,node
# or
lix run why-benchkit run --targets interp,node --json-dir out/When @:warmup / @:iterations (or the same fields on Measure.run opts) are omitted, measurement adapts: warmup stabilizes, then a time budget (default ~150 ms via targetMs) chooses how many timed iterations to run. Set either or both for fixed counts; explicit values always override.
After warmup and iteration resolution, each timed measure runs N independent loops (sampleCount, default 5). Measure.run returns a RawMeasureResult — { metric, headline, ?evidence, ?timing } — and the Runner stamps name to form MeasureResult (raw & { name }). The default metric is Latency (ns/op mean); evidence is Samples (per-loop ns/op) or Summary; timing records iterations / warmup / loop wall samples. Host --samples applies only when measure opts omit sampleCount.
Metrics at a glance. Throughput · Latency (ns/op) · Duration · Bytes (Float byte counts) · Count · Custom(value, unit, polarity). Only Custom carries an explicit polarity (HigherIsBetter | LowerIsBetter); other variants have fixed polarity. headline tags what the scalar is (Mean, Median, …) without wrapping it.
@:custom. Mark a public method @:custom and return a RawMeasureResult yourself (e.g. Bytes or Custom). The Runner stamps name; do not combine with @:warmup / @:iterations. Host --samples does not apply.
JSON output is tink_json of BenchmarkResult (no separate hand schema).
TODO: Noise model / statistical comparison — v1 compare and reporters use the projected headline value ± a relative threshold. Sample variance / median / hypothesis tests are not used for compare yet; evidence may store samples for that future work.
haxelib run why-benchkit run --targets interp,neko,python,node,js,lua,cpp,jvm [--json-dir out/] [--samples 5]
# or
lix run why-benchkit run --targets interp,node --json-dir out/ --samples 5--targets is required. Known targets: interp,neko,python,node,js,lua,cpp,jvm.
| Flag | Default | Notes |
|---|---|---|
--targets |
— | Required. Comma-separated list |
--json-dir |
off | Write nested JSON under <dir>/<sha>/ or <dir>/_dirty/ |
--samples |
5 |
Independent timed loops per measure after warmup (>= 1) |
Install project deps before invoking the host. Console reporting always runs; with --json-dir, results land under a clean-commit SHA folder (or _dirty/ when the tree is dirty / git is unavailable), plus folder and root manifests.
Run the same suite at two git SHAs, diff polarity-aware relative value deltas (not ops/sec-only), and print a verdict table. Same measure key with different metric tags is Incompatible (no delta; not treated as a degradation).
haxelib run why-benchkit compare --base <sha-or-ref> --head <sha-or-ref> --targets node
# or
lix run why-benchkit compare --base origin/main --head HEAD --targets interp,node --samples 5 --threshold 0.10| Flag | Required | Default | Notes |
|---|---|---|---|
--base |
yes | — | Baseline commit (full or unambiguous short SHA / ref) |
--head |
yes | — | Candidate commit |
--targets |
yes | — | Same as run |
--samples |
no | 5 |
Passed through to both SHA runs |
--threshold |
no | 0.10 |
Relative value delta for “major” (default = 10%) |
--fail-on-missing |
no | off | Non-zero exit if any measure exists on only one side |
--post-pr-comment |
no | off | Post/update a markdown summary on the current GitHub PR |
Exit 0 when there is at least one paired measure and no major degradations (improvements / unchanged / incompatible are fine). Exit 1 on major degradation, orchestration failure, or zero paired measures. Missing-side counts are always printed; they only fail the process with --fail-on-missing.
TODO: Allow customizing the per-worktree install command (today hard-coded lix download).
TODO: Noise model / statistical comparison (v1 is polarity-aware value ± threshold only — see Measuring).
Posts (or updates) a single markdown summary on the current GitHub PR after the table. Prefer the gh CLI when available; otherwise uses Actions env (GITHUB_* + token). Comment failure is warn-only and does not override the compare exit code.
permissions:
contents: read
pull-requests: write
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- run: lix run why-benchkit compare --base origin/main --head HEAD --targets node --post-pr-commentGenerate a Chart.js page that fetches manifests and result JSON at runtime:
haxelib run why-benchkit html --out=bin/out.html --json-dir=out/
# or
lix run why-benchkit html --out=bin/out.html --json-dir=out/Writes sibling .css / .js next to the HTML. Default fetch base is the relative path from the HTML file to --json-dir; override with --json-base when the deploy layout differs. Preview over HTTP (file:// fetch will fail).
Rebuild the root catalog from commit folders under --json-dir (skips _dirty):
haxelib run why-benchkit manifest --json-dir out/
# or
lix run why-benchkit manifest --json-dir out/Additively copy a local JSON tree onto another git branch (skips _dirty; commits by default; --push optional):
haxelib run why-benchkit sync --source-dir=out/ --dest-branch=gh-pages --dest-dir=bench-data/
lix run why-benchkit sync --source-dir=out/ --dest-branch=gh-pages --dest-dir=bench-data/ --pushPublish the viewer and JSON under the same origin so relative fetch works. Exclude _dirty/ from anything you publish.
- Accumulate history with a clean tree:
why-benchkit run --targets … --json-dir docs/bench-data/ - Generate the viewer:
why-benchkit html --out=docs/index.html --json-dir=docs/bench-data/ - Push
docs/(or your Pages folder) and open over HTTPS.
docs/
index.html
index.css
index.js
bench-data/
manifest.json
<full-sha>/…
When JSON should live on another branch (e.g. gh-pages):
why-benchkit sync --source-dir=out/ --dest-branch=gh-pages --dest-dir=bench-data/
# Add --push to also push origin/gh-pages| Requirement | Why |
|---|---|
actions/checkout with a token that can write |
Push needs write access |
permissions: contents: write |
Workflow token scope |
Fetch dest branch history (fetch-depth: 0 or git fetch origin <dest-branch>) |
Worktree / orphan detection needs the remote ref |
Configure user.name / user.email before sync |
Commit fails otherwise |
permissions:
contents: write
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- run: |
git config user.name "github-actions[bot]"
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
# … run benches into e.g. out/ …
- run: lix run why-benchkit sync --source-dir=out/ --dest-branch=gh-pages --dest-dir=bench-data/ --pushPure / logic suite (tests.hxml → RunTests / TestBatch) via travix:
lix run travix interp
lix run travix nodeGitHub Actions (.github/workflows/tests.yml) runs the same two travix targets on every push and pull request. Integration mains are not part of CI.
Integration mains (tests/integration/*Main.hx) are not in TestBatch. Run them manually with their hxmls, e.g. haxe hostrun.hxml, haxe hostcompare.hxml, haxe sync.hxml, haxe manifest.hxml.
JSON layout. With --json-dir, results land as <json-dir>/<sha|/_dirty>/<haxeVersion>/<target>.json plus folder and root manifest.json; _dirty/ is local-only and never listed in the root catalog.
Browser JS. For target js, the host injects reporter config via packaged travix hooks; you do not need a local .travix. See .travix/README.md.
Standalone reporter config. Set WHY_BENCHKIT_CONFIG (native/node) or window.why.benchkit (browser) to a JSON object with reporters (default console; optional json + outputDir). Root target is required when using the json reporter.