Skip to content

fail_on_error=False inflates accuracy by excluding errored samples from denominator #3707

Description

@xiangxl-a

Problem

When fail_on_error=False, samples that error during execution are completely excluded from metric computation — both numerator and denominator. This inflates accuracy scores.

Example: 100 samples, 10 error, 80 of remaining 90 are correct → reported accuracy is 88.9% (80/90) rather than 80% (80/100).

Current behavior

  1. task_run_sample returns None for errored samples
  2. completed_scores filters to only isinstance(score_dict, dict) results
  3. scorer_for_metrics further filters out NaN values
  4. accuracy() computes sum(values) / len(scores) — only over successfully scored samples
  5. A warning is printed to the display ("WARNING: N of M executed samples had errors and were not scored") but this is display-only and not reflected in the metric value

Suggested improvements

  1. Add a task-level option that controls how errored samples are treated in metrics at eval time.
  2. Document the current behavior explicitly — the Errors and Limits page documents fail_on_error but does not mention that errored samples are excluded from accuracy computation.

Environment

  • inspect_ai version: 0.3.74
  • Python: 3.12

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions