Goal
Kill the curriculum-learning hypothesis fast, or earn the right to run the full experiment matrix.
Scope
- Task: GSM8K (kill task), XSum (anchor)
- Model: Llama-3.2 1B
- Budget: ~40–50 A100-hours across 5 configs × 3 seeds
Deliverables
- Dataset prep, exposure-bias diagnostic Δ(k), task-aware pass@1 scoring, kill-phase experiment configs (K-P1, K-1..K-5), Databricks notebooks for prep / training / diagnostic
- MLflow run for K-P1 pilot + per-run Δ(k) JSON
Gates
- A (pilot): SFT loss decreases, pass@1 > 10%, extract rate > 80%, Δ(k) > 0 and roughly monotonic in k
- B (headline): K-2 beats K-1 on GSM8K pass@1 by ≥ 2 points across 3 seeds
- C (anchor): K-4 does not regress vs K-3 on XSum ROUGE-L
Out of scope
Alternate schedules, replacement-rate ablations.
Goal
Kill the curriculum-learning hypothesis fast, or earn the right to run the full experiment matrix.
Scope
Deliverables
Gates
Out of scope
Alternate schedules, replacement-rate ablations.