Distill teacher chains-of-thought into a LoRA adapter via a strict boxed-answer format contract + two-phase Train→Nudge (silver-medal NVIDIA Nemotron reasoning recipe, as a tested library).
kaggle nvidia lora mamba reasoning fine-tuning peft sft mixture-of-experts trl llm chain-of-thought nemotron trace-distillation
-
Updated
Jul 27, 2026 - Python