Nemotron-3-Nano-4B

NVIDIA

Parameters

4.0B

Architecture

Mamba2-Transformer Hybrid (Nemotron-Hybrid)

Released

16.03.2026

License

NVIDIA Nemotron Open Model License

Open Weights Commercial Use Multimodal BF16 Nemotron en

Input Modalities

text

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

262,144 tokens

Openness Index Score 70.0/100

About

NVIDIA-Nemotron-3-Nano-4B is a small language model (SLM) trained from scratch by NVIDIA and designed as a unified model for both reasoning and non-reasoning tasks: it answers queries by first generating a reasoning trace, then concluding with a final response, with the reasoning behavior controllable via a system prompt. It was compressed from NVIDIA-Nemotron-Nano-9B-v2 using the Nemotron Elastic framework (arXiv 2511.16664).

Its Mamba2-Transformer hybrid (Nemotron-Hybrid) architecture consists primarily of Mamba2 SSM and MLP layers combined with just four Attention layers, giving 3.97B parameters with a 262K-token context window and high inference efficiency on NVIDIA hardware (A10G, A100, H100, GeForce RTX; NeMo 25.07 runtime). Input and output are text, English and coding languages.

Pre-training used more than 10 trillion tokens (data cutoff September 2024). The post-training corpus spans English and multilingual text across code, legal, math, science and finance domains, and includes synthetic reasoning traces distilled from DeepSeek R1/R1-0528, Qwen3-235B-A22B, Nemotron 4 340B and Qwen2.5 models - the model notes it was "Improved using Qwen".

It targets edge-ready Agentic AI on Jetson Thor, GeForce RTX and DGX Spark: AI gaming NPCs (teammates/companions), local voice assistants, and IoT automation. Model dates Dec 2025 - Jan 2026; released on Hugging Face March 16, 2026 under the NVIDIA Nemotron Open Model License; ready for commercial use.

Training Data 10+ trillion tokens pre-training, compressed from NVIDIA-Nemotron-Nano-9B-v2

Benchmark Scores

Benchmark Score Date
Multi-IF
instruction_following
68.27%
—
BFCL-V4
general_agent
32.14%
—
MMLU-Pro
knowledge
18.39%
—
SWE-bench Verified
coding_agent
2.98%
—
Humanity's Last Exam
stem_reasoning
5.30%
—
SWE-bench Pro
coding_agent
0.12%
—
GPQA Diamond
stem_reasoning
18.70%
—
Terminal Bench 2.1
coding_agent
3.77%
—
SuperGPQA
knowledge
26.13%
—
BrowseComp-zh
general_agent
3.30
—
AA-LCR
long_context
21.62%
—
BrowseComp
general_agent
1.82%
—
NoLiMa
long_context
0.89%
—
Gaia2
general_agent
26.50
—
LongBenchPro
long_context
39.24%
—
GDPVal-AA v2
general_agent
0.00
—
Claw-Eval Avg
coding_agent
38.90%
—
LongBench v2
long_context
16.06%
—
WildClawBench
coding_agent
6.47%
—
TAU3-Bench
general_agent
1.20
—
QwenClawBench
coding_agent
25.15%
—
TAU2-Bench
general_agent
14.81%
—
LiveCodeBench v6
stem_reasoning
41.61%
—
LCB-Pro 25Q2 (Easy)
stem_reasoning
71.58%
—
LCB-Pro 25Q2 (Medium)
stem_reasoning
30.29%
—
OJBench
stem_reasoning
25.55%
—
SciCode
30.77%
—
AIME 2025
stem_reasoning
42.45%
—
AIME 26
stem_reasoning
52.68%
—
HMMT Feb 26
stem_reasoning
39.90%
—
MATH-500
math
45.59%
—
IFBench
instruction_following
59.23%
—
MMLU-Redux
knowledge
33.61%
—
IFEval
instruction_following
88.37%
—

Model Tree, Spaces and Papers

Model tree for nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16

Base model

nvidia/NVIDIA-Nemotron-Nano-12B-v2-Base

Finetuned

nvidia/NVIDIA-Nemotron-Nano-12B-v2

Finetuned

nvidia/NVIDIA-Nemotron-Nano-9B-v2

Finetuned

(22)

this model

Adapters

13 models

Finetunes

19 models

Quantizations

48 models

Ethical Considerations (NVIDIA Trustworthy AI)

Ethical Considerations

NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our Trustworthy AI terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.

We advise against circumvention of any provided safety guardrails contained in the Model without a substantially similar guardrail appropriate for your use case.For more details: Safety and Explainability Subcards.

For more detailed information on ethical considerations for this model, please see the Model Card++ Bias, and Privacy Subcards.

Please report security vulnerabilities or NVIDIA AI Concerns here.

Safetensors

Model size

4B params

Tensor type

BF16

·

Inference Engines and Test Hardware

Inference

  • Engines: HF, vLLM, llama-cpp, TRT-LLM, SGLang
  • Test Hardware: NVIDIA GeForce RTX, H100 80GB, DGX Spark, Jetson Thor/Orin Nano

Evaluation Datasets

Dataset Collection Period
Problems in Elementary Mathematics for Home Study 4/23/2025
GSM8K 4/23/2025

Evaluation Dataset:

  • Data Collection Method by dataset: Hybrid: Human, Synthetic
  • Labeling Method by dataset: Hybrid: Automated, Human, Synthetic

NVIDIA-Sourced Synthetic Datasets

NVIDIA-Sourced Synthetic Datasets

Dataset Modality Dataset Size (Tokens) Seed Dataset Model(s) used for generation
Synthetic Art of Problem Solving from DeepSeek-R1 Text 25.5B Art of Problem Solving; American Mathematics Competitions 8; American Mathematics Competitions 10; DeepSeek-R1
Synthetic Moral Stories and Social Chemistry from Mixtral-8x22B-v0.1 Text 327M social-chemestry-101; Moral Stories Mixtral-8x22B-v0.1
Synthetic Social Sciences seeded with OpenStax from DeepSeek-V3, Mixtral-8x22B-v0.1, and Qwen2.5-72B Text 83.6M OpenStax - CC BY-SA subset DeepSeek-V3; Mixtral-8x22B-v0.1; Qwen2.5-72B
Synthetic Health Sciences seeded with OpenStax from DeepSeek-V3, Mixtral-8x22B-v0.1, and Qwen2.5-72B Text 9.7M OpenStax - CC BY-SA subset DeepSeek-V3; Mixtral-8x22B-v0.1; Qwen2.5-72B
Synthetic STEM seeded with OpenStax, Open Textbook Library, and GSM8K from DeepSeek-R1, DeepSeek-V3, DeepSeek-V3-0324, and Qwen2.5-72B Text 175M OpenStax - CC BY-SA subset; GSM8K; Open Textbook Library - CC BY-SA & GNU subset DeepSeek-R1, DeepSeek-V3; DeepSeek-V3-0324; Qwen2.5-72B
Nemotron-PrismMath Text 4.6B Big-Math-RL-Verified; OpenR1-Math-220k Qwen2.5-0.5B-instruct, Qwen2.5-72B-Instruct; DeepSeek-R1-Distill-Qwen-32B
Synthetic Question Answering Data from Papers and Permissible Books from Qwen2.5-72B-Instruct Text 350M arXiv; National Institutes of Health ExPorter; BioRxiv; PMC Article; USPTO Backgrounds; peS2o; Global Regulation; CORE; PG-19; DOAB CC BY & CC BY-SA subset; NDLTD Qwen2.5-72B-Instruct
Synthetic FineMath-4+ Reprocessed from DeepSeek-V3 Text 9.2B Common Crawl DeepSeek-V3
Synthetic FineMath-3+ Reprocessed from phi-4 Text 27.6B Common Crawl phi-4
Synthetic Union-3+ Reprocessed from phi-4 Text 93.1B Common Crawl phi-4
Refreshed Nemotron-MIND from phi-4 Text 73B Common Crawl phi-4
Synthetic Union-4+ Reprocessed from phi-4 Text 14.12B Common Crawl phi-4
Synthetic Union-3+ minus 4+ Reprocessed from phi-4 Text 78.95B Common Crawl phi-4
Synthetic Union-3 Refreshed from phi-4 Text 80.94B Common Crawl phi-4
Synthetic Union-4+ Refreshed from phi-4 Text 52.32B Common Crawl phi-4
Synthetic AGIEval seeded with AQUA-RAT, LogiQA, and AR-LSAT from DeepSeek-V3 and DeepSeek-V3-0324 Text 4.0B AQUA-RAT; LogiQA; AR-LSAT DeepSeek-V3; DeepSeek-V3-0324
Synthetic AGIEval seeded with AQUA-RAT, LogiQA, and AR-LSAT from Qwen3-30B-A3B Text 4.2B AQUA-RAT; LogiQA; AR-LSAT Qwen3-30B-A3B
Synthetic Art of Problem Solving from Qwen2.5-32B-Instruct, Qwen2.5-Math-72B, Qwen2.5-Math-7B, and Qwen2.5-72B-Instruct Text 83.1B [Art of Pr

Private and Online Dataset Sources

Private Non-publicly Accessible Datasets of Third Parties

Dataset
Global Regulation
Workbench

Online Dataset Sources

The English Common Crawl data was downloaded from the Common Crawl Foundation (see their FAQ for details on their crawling) and includes the snapshots CC-MAIN-2013-20 through CC-MAIN-2025-13. The data was subsequently deduplicated and filtered in various ways described in the Nemotron-CC paper.

Additionally, we extracted data for fifteen languages from the following three Common Crawl snapshots: CC-MAIN-2024-51, CC-MAIN-2025-08, CC-MAIN-2025-18. The fifteen languages included were Arabic, Chinese, Danish, Dutch, French, German, Italian, Japanese, Korean, Polish, Portuguese, Russian, Spanish, Swedish, and Thai. As we did not have reliable multilingual model-based quality classifiers available, we applied just heuristic filtering instead—similar to what we did for lower quality English data in the Nemotron-CC pipeline, but selectively removing some filters for some languages that did not work well. Deduplication was done in the same way as for Nemotron-CC.

The GitHub Crawl was collected using the GitHub REST API and the Amazon S3 API. Each crawl was operated in accordance with the rate limits set by its respective source, either GitHub or S3. We collect raw source code and subsequently remove any having a license which does not exist in our permissive-license set (for additional details, refer to the technical report).

Dataset Modality Dataset Size (Tokens) Collection Period
English Common Crawl Text 3.360T 4/8/2025
Multilingual Common Crawl Text 812.7B 5/1/2025
GitHub Crawl Text 747.4B 4/29/2025
English Common Crawl 1.1 Text Not disclosed 10/2/2025

Public Datasets (with collection periods)

Public Datasets

Dataset Collection Period
Problems in Elementary Mathematics for Home Study 4/23/2025
GSM8K 4/23/2025
PRM800K 4/23/2025
CC-NEWS 4/23/2025
Common Crawl 4/23/2025
Wikimedia 4/23/2025
Bespoke-Stratos-17k 4/23/2025
tigerbot-kaggle-leetcodesolutions-en-2k 4/23/2025
glaive-function-calling-v2 4/23/2025
APIGen Function-Calling 4/23/2025
LMSYS-Chat-1M 4/23/2025
Open Textbook Library - CC BY-SA & GNU subset and OpenStax - CC BY-SA subset 4/23/2025
Advanced Reasoning Benchmark, tigerbot-kaggle-leetcodesolutions-en-2k, PRM800K, and SciBench 4/23/2025
FineWeb-2 4/23/2025
Court Listener Legacy Download
peS2o Legacy Download
OpenWebMath Legacy Download
BioRxiv Legacy Download
PMC Open Access Subset Legacy Download
OpenWebText2 Legacy Download
Stack Exchange Data Dump Legacy Download
PubMed Abstracts Legacy Download
NIH ExPorter Legacy Download
arXiv Legacy Download
BigScience Workshop Datasets Legacy Download
Reddit Dataset Legacy Download
SEC's Electronic Data Gathering, Analysis, and Retrieval (EDGAR) Legacy Download
Public Software Heritage S3 Legacy Download
The Stack Legacy Download
mC4 Legacy Download
Advanced Mathematical Problem Solving Legacy Download
MathPile Legacy Download
NuminaMath CoT Legacy Download
PMC Article Legacy Download
FLAN Legacy Download
Advanced Reasoning Benchmark Legacy Download
SciBench Legacy Download
WikiTableQuestions Legacy Download
FinQA Legacy Download
Riddles Legacy Download
Problems in Elementary Mathematics for Home Study Legacy Download
MedMCQA Legacy Download
Cosmos QA Legacy Download
MCTest Legacy Download
AI2's Reasoning Challenge Legacy Download
OpenBookQA Legacy Download
MMLU Auxiliary Train Legacy Download
social-chemestry-101 Legacy Download
Moral Stories Legacy Download
The Common Pile v0.1 Legacy Download
FineMath Legacy Download
MegaMath Legacy Download
FastChat 6/30/2025
MultiverseMathHard 10/2/2025
SWE-Gym 10/2/2025
WorkBench 10/2/2025
WildChat-1M 10/2/2025
OpenCodeReasoning-2 10/2/2025
HelpSteer3 10/2/2025
opc-sft-stage2 10/2/2025
Big-Math-RL-Verified 10/2/2025
NuminaMath CoT 10/2/2025
MetaMathQA 10/2/2025
simple-arithmetic-problems 10/2/2025
arithmetic 10/2/2025
Skywork-OR1-RL-Data 10/2/2025
News Commentary 10/2/2025
FastChat 10/2/2025
Essential-Web 10/2/2025
finepdfs 10/2/2025
HotpotQA 10/2/2025
SQuAD2.0 10/2/2025
NLTK Words Lists 10/2/2025

Training Datasets

Training, Testing, and Evaluation Datasets

Training datasets

  • Data Modality: Text
  • Text Training Data Size: More than 10 Trillion Tokens
  • Train/Test/Valid Split: We used 100% of the corpus for pre-training and relied on external benchmarks for testing.
  • Data Collection Method by dataset: Hybrid: Automated, Human, Synthetic
  • Labeling Method by dataset: Hybrid: Automated, Human, Synthetic

Properties: The post-training corpus for NVIDIA-Nemotron-3-Nano-4B consists of English and multilingual text (German, Spanish, French, Italian, Korean, Portuguese, Russian, Japanese, Chinese and English). Our sources cover a variety of document types such as: webpages, dialogue, articles, and other written materials. The corpus spans domains including code, legal, math, science, finance, and more. We also include a small portion of question-answering, and alignment style data to improve model accuracies. For several of the domains listed above we used synthetic data, specifically reasoning traces, from DeepSeek R1/R1-0528, Qwen3-235B-A22B, Nemotron 4 340B, Qwen2.5-32B-Instruct-AWQ, Qwen2.5-14B-Instruct, Qwen 2.5 72B.

More details on the datasets and synthetic data generation methods can be found in the technical report NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model .

Software Integration (NeMo 25.07, supported hardware)

Software Integration

  • Runtime Engine(s): NeMo 25.07
  • Supported Hardware Microarchitecture Compatibility: NVIDIA A10G, NVIDIA H100-80GB, NVIDIA A100, GeForce RTX
  • Operating System(s): Linux

The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.

Model Architecture, Input and Output

Model Architecture

  • Architecture Type: Mamba2-Transformer Hybrid
  • Network Architecture: Nemotron-Hybrid

Input

  • Input Type(s): Text
  • Input Format(s): String
  • Input Parameters: One-Dimensional (1D): Sequences
  • Other Properties Related to Input: Context length up to 262K. Supported languages include English.

Output

  • Output Type(s): Text
  • Output Format: String
  • Output Parameters: One-Dimensional (1D): Sequences
  • Other properties Related to Output: Sequences up to 262K

Our models are designed and optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA’s hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.

Use Case: Edge-Ready Agentic AI

Use Case

NVIDIA-Nemotron-3-Nano-4B is an edge-ready small language model intended for Agentic AI in edge platforms (Jetson Thor, GeForce RTX, DGX Spark). It targets key-uses including AI gaming NPCs (teammates / companions), local voice assistants (for devices, apps, and games), and IoT automation. It is to be used in English and coding languages.

Release Date: 3/16/2026

Huggingface 3/16/2026 via https://huggingface.co/

Evaluation Results (Reasoning-off and Reasoning-on)

Evaluation Results:

We evaluated our model in **Reasoning-off** mode across these benchmarks

Benchmark NVIDIA-Nemotron-3-Nano-4B-BF16
BFCL v3 61.1
IFBench-Prompt 43.2
IFBench-Instruction 44.2
Orak 22.9
IFEval-Prompt 82.8
IFEval-Instruction 88
HaluEval 62.2
RULER (128k) 91.1
Tau2-Airline 28.0
Tau2-Retail 34.8
Tau2-Telecom 24.9
EQ-Bench3 63.2

We also evaluated our model in **Reasoning-On** mode across these benchmarks.

Benchmark NVIDIA-Nemotron-3-Nano-4B-BF16
AIME25 78.5
MATH500 95.4
GPQA 53.2
LCB 51.8
BFCL v3 61.1
IFEVAL-Prompt 87.9
IFEVAL-Instruction 92
Tau2-Airline 33.3
Tau2-Retail 39.8
Tau2-Telecom 33

All evaluations were done using NeMo-Skills & Orak. For Orak we evaluated on three games (Super Mario, Darkest Dungeon & StarDew Valley)

Deployment Geography: Global

License / Terms of Use (NVIDIA Nemotron Open Model License)

License/Terms of Use

Governing Terms: Use of this model is governed by the NVIDIA Nemotron Open Model License.

Model Overview: Unified Reasoning/Non-Reasoning SLM

Model Overview

NVIDIA-Nemotron-3-Nano-4B-BF16 is a small language model (SLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be controlled via a system prompt. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so, albeit with a slight decrease in accuracy for harder prompts that require reasoning. Conversely, allowing the model to generate reasoning traces first generally results in higher-quality final solutions to queries and tasks.

The model has been compressed from NVIDIA-Nemotron-Nano-9B-v2 using the Nemotron Elastic framework. The details of the parent model NVIDIA-Nemotron-Nano-9B-v2 can be found in (Nemotron-H tech report). The model uses a hybrid architecture consisting primarily of Mamba-2 and MLP layers combined with just four Attention layers.

The supported languages include: English. Improved using Qwen.

This model is ready for commercial use.

Architecture

Decoder Block input Embedding vocab 131K · d 3136 Linear / Recurrent Hybrid 40:8 · dₕ 128 ×42 Dense FFN SwiGLU · d 13K Final RMSNorm LM Head vocab 131K output
Attention
Hybrid Attention (40:8)
Layers
42
Hidden size
3136
Context
262K tokens
Parameters
3970M

Source: Hugging Face config.json · NemotronHForCausalLM · model repo

Type: Mamba2-Transformer hybrid (Nemotron-Hybrid): mostly Mamba2 SSM + MLP layers + 4 attention layers
Attention: Only 4 full-attention layers; remaining positions handled by Mamba2 selective state-space blocks
Decoder: Hybrid SSM-Transformer decoder
Total parameters 3970M
Context length 262K
Extended context 262K
Attention Layers 4
Arch Family

Nemotron-Hybrid

Compression

Nemotron Elastic framework (arXiv 2511.16664)

Hardware

A10G, A100, H100-80GB, GeForce RTX, Jetson Thor, DGX Spark

Input

text

Output

text

Parent Model

NVIDIA-Nemotron-Nano-9B-v2

Runtime

NeMo 25.07

Training Pipeline

  1. 1
    pretraining

    Pre-training (>10T tokens)

    Trained from scratch on more than 10 trillion tokens (data cutoff September 2024).

  2. 2
    other

    Elastic compression from Nemotron-Nano-9B-v2

    Compressed from the 9B parent via the Nemotron Elastic framework (arXiv 2511.16664).

  3. 3
    sft

    Post-training SFT with synthetic reasoning traces

    English + multilingual corpus across code, legal, math, science, finance; synthetic reasoning traces distilled from DeepSeek R1/R1-0528, Qwen3-235B-A22B, Nemotron 4 340B, Qwen2.5-32B/14B/72B; improved using Qwen.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
AI2's Reasoning Challenge finetune — —
APIGen Function-Calling finetune — —
Bespoke-Stratos-17k finetune — —
Big-Math-RL-Verified finetune — —
Cosmos QA finetune — —
Essential-Web pretraining — —
FineMath pretraining — —
FineWeb-2 pretraining — —
HelpSteer3 finetune — —
LMSYS-Chat-1M finetune — —
MCTest finetune — —
MMLU Auxiliary Train finetune — —
MedMCQA finetune — —
MegaMath pretraining — —
MetaMathQA finetune — —
Moral Stories finetune — —
MultiverseMathHard finetune — —
NuminaMath CoT finetune — —
OpenCodeReasoning-2 finetune — —
OpenWebMath pretraining — —
SWE-Gym finetune — —
Skywork-OR1-RL-Data finetune — —
The Stack pretraining — —
WikiTableQuestions finetune — —
WildChat-1M finetune — —
arithmetic finetune — —
finepdfs pretraining — —
glaive-function-calling-v2 finetune — —
mC4 pretraining — —
opc-sft-stage2 finetune — —
peS2o pretraining — —
simple-arithmetic-problems finetune — —
social-chemestry-101 finetune — —
tigerbot-kaggle-leetcodesolutions-en-2k finetune — —
Private third-party datasets (Global Regulation, Workbench) training — —
Online dataset sources (Common Crawl Foundation) pretraining — —
NVIDIA-sourced synthetic datasets training — —

Related Models