Skip to content

Latest commit

 

History

History
313 lines (236 loc) · 8.87 KB

File metadata and controls

313 lines (236 loc) · 8.87 KB

Nemotron Training Recipes

Open and efficient models for agentic AI. Reproducible training pipelines with transparent data, techniques, and weights.

<iframe width="560" height="315" src="https://www.youtube.com/embed/_y9SEtn1lU8" title="Nemotron Overview" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen></iframe>

Quick Start

// Install the Nemotron training recipes
$ git clone https://github.com/NVIDIA-NeMo/Nemotron
$ cd Nemotron && uv sync

// Run a tiny SFT job on your cluster
$ uv run nemotron steps run sft/automodel -c tiny --run YOUR-CLUSTER

// Run the Nano3 pipeline stage by stage
$ uv run nemotron nano3 data prep pretrain --run YOUR-CLUSTER
$ uv run nemotron nano3 pretrain --run YOUR-CLUSTER
$ uv run nemotron nano3 data prep sft --run YOUR-CLUSTER
$ uv run nemotron nano3 sft --run YOUR-CLUSTER
$ uv run nemotron nano3 data prep rl --run YOUR-CLUSTER
$ uv run nemotron nano3 rl --run YOUR-CLUSTER

Note: The --run YOUR-CLUSTER flag submits jobs to your configured Slurm cluster via NeMo-Run. See Execution through NeMo-Run for setup instructions.

Sample Deployments and Applications

::::{grid} 1 2 2 2 :gutter: 3

:::{grid-item-card} Deployment Guides :link: deployment-guides :link-type: doc

Deployment guides for Nemotron models: TensorRT-LLM, vLLM, SGLang, NIM, and Hugging Face. :::

:::{grid-item-card} Sample Applications :link: application-examples :link-type: doc

End-to-end applications: RAG agents, ML agents, and multi-agent systems. :::

::::

Customization Workflows with Nemotron Steps

::::{grid} 1 2 2 2 :gutter: 3

:::{grid-item-card} Translation :link: translation/index :link-type: doc

Translate JSONL or Parquet corpora with translate/nemo_curator, NeMo Curator backends, and optional FAITH quality scoring. :::

:::{grid-item-card} Build MCQ Benchmarks :link: build-benchmarks/index :link-type: doc

Generate and translate custom multiple-choice benchmarks with byob/mcq. :::

:::{grid-item-card} Data Curation :link: curate/index :link-type: doc

Filter JSONL text with curate/nemo_curator before translation or training data preparation. :::

:::{grid-item-card} Synthetic Data Generation :link: sdg/index :link-type: doc

Use sdg/data_designer to produce SFT, tool-use, and preference datasets. :::

:::{grid-item-card} Model Evaluation :link: model-eval/index :link-type: doc

Evaluate hosted endpoints or checkpoints with eval/model_eval. :::

::::

Training Recipes

::::{grid} 1 2 2 2 :gutter: 3

:::{grid-item-card} Nemotron 3 Nano :link: nemotron/nano3/README :link-type: doc

31.6B total / 3.6B active parameters, 25T tokens, up to 1M context. Hybrid Mamba-Transformer with sparse MoE.

Stages: Pretraining → SFT → RL :::

:::{grid-item-card} Nemotron 3 Omni :link: nemotron/omni3/README :link-type: doc

GA-checkpoint multimodal post-training recipe with stage-local container builds and a three-step RL stack.

Stages: SFT → RL MPO → RL text → RL vision → Eval :::

:::{grid-item-card} Embedding Fine-Tuning :link: nemotron/embed/README :link-type: doc

Fine-tune Llama-Nemotron-Embed-1B-v2 on domain-specific data with synthetic data generation, evaluation, and NIM deployment.

Stages: SDG → Data Prep → Finetune → Eval → Export → Deploy :::

::::

Recipe Layout

Nemotron keeps data-producing recipes separate from model-family training recipes:

Path Purpose Example
src/nemotron/recipes/data/curation/ Filter, dedup, and curate existing corpora Nemotron-CC
src/nemotron/recipes/data/sdg/ Generate synthetic datasets that can feed multiple families Long-document SDG feeding Omni3 SFT
src/nemotron/recipes/<family>/ Family-specific training, RL, evaluation, and model lifecycle commands Nano3, Omni3

Training Pipeline

Each recipe family has its own stage layout, and all of them can be tracked through artifact lineage:

Family Stage layout
Nano3 Pretraining → SFT → RL
Omni3 SFT → RL MPO → RL text → RL vision → Eval
Super3 Pretraining → SFT → RL → Quantization → Eval
Embed SDG → Data Prep → Finetune → Eval → Export → Deploy

Why Nemotron?

Open Models Transparent training data, techniques, and weights for community innovation
Compute Efficiency Model pruning enabling higher throughput via TensorRT-LLM
High Accuracy Built on frontier open models with human-aligned reasoning
Flexible Deployment Deploy anywhere: edge, single GPU, or data center with NIM

Features

  • End-to-end pipelines from raw data to deployment-ready models
  • Artifact lineage via W&B from data to model
  • Built on NVIDIA's NeMo stack (Megatron-Bridge, NeMo-RL)
  • Reproducible with versioned configs, data blends, and checkpoints

Resources

:caption: Nemotron
:hidden:

Home <self>
application-examples.md
deployment-guides.md
:caption: Nemotron Step Basics
:hidden:

About <steps/index.md>
Basics <steps/basics.md>
Getting Started <steps/getting-started.md>
Airgap Environment <steps/airgap.md>
:caption: Data Curation
:hidden:

About <curate/index.md>
Getting Started <curate/getting-started.md>
Tasks <curate/how-to/index.md>
Reference <curate/reference/index.md>
:caption: Synthetic Data Generation
:hidden:

About <sdg/index>
Getting Started <sdg/getting-started>
Tips for Using Agents <sdg/using-skills>
Planning <sdg/planning>
Tasks <sdg/how-to/index>
Reference <sdg/reference/index>
:caption: Translation
:hidden:

About <translation/index.md>
Getting Started <translation/getting-started.md>
Tips for Using Agents <translation/using-skills.md>
translation/explanation/index.md
Tasks <translation/how-to/index.md>
Reference <translation/reference/index.md>
:caption: Build MCQ Benchmarks
:hidden:

About <build-benchmarks/index.md>
Getting Started <build-benchmarks/getting-started.md>
Concepts <build-benchmarks/explanation/index.md>
Tasks <build-benchmarks/how-to/index.md>
Reference <build-benchmarks/reference/index.md>
:caption: Model Training
:hidden:

About <train-models/index.md>
Getting Started <train-models/getting-started.md>
Tips for Using Agents <train-models/using-skill.md>
Concepts <train-models/explanation/index.md>
Tasks <train-models/how-to/index.md>
Reference <train-models/reference/index.md>
:caption: Model Evaluation
:hidden:

About <model-eval/index.md>
Getting Started <model-eval/getting-started.md>
Tips for Using Agents <model-eval/using-skills.md>
Concepts <model-eval/explanation/index.md>
Tasks <model-eval/how-to/index.md>
Reference <model-eval/reference/index.md>
:caption: Training Recipes
:hidden:

Nemotron 3 Nano <nemotron/nano3/README.md>
Nemotron 3 Omni <nemotron/omni3/README.md>
Nemotron 3 Super <nemotron/super3/README.md>
Llama Nemotron Embed <nemotron/embed/README.md>
nemotron/artifacts.md
:caption: Nemotron Kit
:hidden:

nemotron/kit.md
nemotron/nvidia-stack.md
nemo_runspec/package-readme.md
nemo_runspec/nemo-run.md
nemo_runspec/omegaconf.md
nemo_runspec/artifacts.md
nemotron/wandb.md
nemotron/cli.md
nemotron/data-prep.md
nemotron/xenna-observability.md
:caption: Data Recipes
:hidden:

nemotron/data/curation/nemotron-cc.md
nemotron/data/sdg/long-document.md
:caption: Architecture
:hidden:

architecture/README.md
architecture/design-philosophy.md
architecture/cli-architecture.md
runspec/v1/spec.md