Problem
When fail_on_error=False, samples that error during execution are completely excluded from metric computation — both numerator and denominator. This inflates accuracy scores.
Example: 100 samples, 10 error, 80 of remaining 90 are correct → reported accuracy is 88.9% (80/90) rather than 80% (80/100).
Current behavior
task_run_sample returns None for errored samples
completed_scores filters to only isinstance(score_dict, dict) results
scorer_for_metrics further filters out NaN values
accuracy() computes sum(values) / len(scores) — only over successfully scored samples
- A warning is printed to the display (
"WARNING: N of M executed samples had errors and were not scored") but this is display-only and not reflected in the metric value
Suggested improvements
- Add a task-level option that controls how errored samples are treated in metrics at eval time.
- Document the current behavior explicitly — the Errors and Limits page documents
fail_on_error but does not mention that errored samples are excluded from accuracy computation.
Environment
- inspect_ai version: 0.3.74
- Python: 3.12
Problem
When
fail_on_error=False, samples that error during execution are completely excluded from metric computation — both numerator and denominator. This inflates accuracy scores.Example: 100 samples, 10 error, 80 of remaining 90 are correct → reported accuracy is 88.9% (80/90) rather than 80% (80/100).
Current behavior
task_run_samplereturnsNonefor errored samplescompleted_scoresfilters to onlyisinstance(score_dict, dict)resultsscorer_for_metricsfurther filters out NaN valuesaccuracy()computessum(values) / len(scores)— only over successfully scored samples"WARNING: N of M executed samples had errors and were not scored") but this is display-only and not reflected in the metric valueSuggested improvements
fail_on_errorbut does not mention that errored samples are excluded from accuracy computation.Environment