Back to Developer Zone

NVIDIA Models

Explore all 4 models from NVIDIA with detailed pricing, pros & cons, and developer recommendations.

4
Models
$0.050
Lowest Input
1M
Max Context
3
Quality Tiers

Quick Recommendations

Best Value: Nemotron-3-4B ($0.050/1M)
Best Quality: Nemotron-3-Ultra

Nemotron-3-Ultra

Flagship

Agentic reasoning, orchestration, complex planning

Official Pricing

When to use: Frontier agentic workloads: multi-step planning, tool orchestration, complex reasoning chains, and production agent deployment.

Upgrade Highlights

  • 550B MoE with 55B active — open frontier reasoning model
  • Hybrid Mamba-Transformer: 300+ tok/s, 1M context
  • Purpose-built for agentic orchestration and tool calling
  • Open-weight: deploy with NeMo, NIM, or self-hosted
  • $0.50/1M input — frontier capability at open-model pricing
Input Price
$0.500
per 1M tokens
Output Price
$2.00
per 1M tokens
Cached Input
$0.130
per 1M tokens
Batch Input
per 1M tokens
Context Window: 1M
Max Output: 32,768 tokens
Knowledge Cutoff: 2026-04
VisionFunction CallingFine-tuningJSON ModeFree Tier

Pros

  • 550B total / 55B active — open frontier reasoning
  • Hybrid Mamba-Transformer-MoE architecture
  • 1M context for long-horizon agent tasks
  • 300+ tokens/sec throughput
  • Open-weight with NeMo customization

Cons

  • No vision support
  • 55B active params still requires significant compute for self-hosting
  • Smaller ecosystem vs OpenAI/Anthropic

Performance

Output Speed~90 tok/s
Rate Limit5,000 RPM

Multimodal

Image InputImage OutputAudio InputAudio Output

Benchmarks

MMLU
88.5%
MT-Bench
9.4
BFCL
78.2%
AgentBench
72.6%

Nemotron-3-Nano-Omni

Lite

Multimodal agents, video/audio/image understanding, edge deployment

Official Pricing

When to use: Multimodal AI agents needing real-time video, audio, and image understanding at the edge. Document intelligence, customer support with screen sharing, and audio-visual content analysis.

Upgrade Highlights

  • First NVIDIA omni-model: vision + audio + language unified
  • 30B-A3B MoE — deploys on Jetson and DGX Spark
  • 9x throughput: same performance, far fewer resources
  • Native 1080p screen reasoning for GUI agents
  • Open-weight: customize with NeMo for domain tasks
Input Price
$0.100
per 1M tokens
Output Price
$0.400
per 1M tokens
Cached Input
per 1M tokens
Batch Input
per 1M tokens
Context Window: 262K
Max Output: 8,192 tokens
Knowledge Cutoff: 2026-04
VisionFunction CallingFine-tuningJSON ModeFree Tier

Pros

  • Omni-modal: video + audio + image + text in one model
  • 30B total / 3B active — runs on 25GB RAM
  • 9x throughput vs comparable open omni-models
  • Open-weight with commercial use rights
  • Native 1920x1080 visual reasoning

Cons

  • 3B active params limits complex reasoning
  • 8K max output
  • Smaller context window (256K)

Performance

Output Speed~200 tok/s
Rate Limit15,000 RPM

Multimodal

Image InputImage OutputAudio InputAudio Output

Benchmarks

MMLU
71.2%
MMMU
65.8%
OSWorld
42.3%
DocVQA
92.1%

Nemotron-3-8B

Mid-tier

Synthetic data generation, reward modeling

Official Pricing

When to use: Generating synthetic training data, reward modeling for RLHF, and NVIDIA ecosystem workflows.

Upgrade Highlights

  • Purpose-built for synthetic data generation — train other models
  • Strong reward model: ranks responses better than GPT-4 judge
  • NVIDIA NIM deployment — optimized inference on GPUs
  • Fine-tuning with NeMo framework
  • $0.10/1M input — cost-effective for data generation at scale
Input Price
$0.100
per 1M tokens
Output Price
$0.400
per 1M tokens
Cached Input
per 1M tokens
Batch Input
per 1M tokens
Context Window: 128K
Max Output: 4,096 tokens
Knowledge Cutoff: 2025-03
VisionFunction CallingFine-tuningJSON ModeFree Tier

Pros

  • Optimized for synthetic data generation
  • Strong reward model capabilities
  • NVIDIA ecosystem integration
  • Fine-tuning support

Cons

  • 4K max output
  • No vision
  • Smaller model size limits complex tasks

Performance

Output Speed~100 tok/s
Rate Limit10,000 RPM

Multimodal

Image InputImage OutputAudio InputAudio Output

Benchmarks

MMLU
73.8%
MT-Bench
8.2

Nemotron-3-4B

Lite

Lightweight synthetic data, edge deployment

Official Pricing

When to use: Edge deployment, lightweight synthetic data generation, and cost-sensitive batch processing.

Upgrade Highlights

  • 4B params — deploys on Jetson and edge GPUs
  • Synthetic data generation at $0.05/1M input
  • NVIDIA NIM: optimized containerized deployment
  • Open-source: full weights for customization
Input Price
$0.050
per 1M tokens
Output Price
$0.200
per 1M tokens
Cached Input
per 1M tokens
Batch Input
per 1M tokens
Context Window: 128K
Max Output: 4,096 tokens
Knowledge Cutoff: 2025-03
VisionFunction CallingFine-tuningJSON ModeFree Tier

Pros

  • Compact 4B — runs on edge GPUs
  • Synthetic data generation
  • NVIDIA NIM optimized
  • Open-source weights

Cons

  • Limited reasoning capability
  • 4K max output
  • No vision

Performance

Output Speed~180 tok/s
Rate Limit20,000 RPM

Multimodal

Image InputImage OutputAudio InputAudio Output

Benchmarks

MMLU
66.5%
MT-Bench
7.1

Side-by-Side Comparison

ModelTierInputOutputContext
Nemotron-3-UltraFlagship$0.500$2.001M
Nemotron-3-Nano-OmniLite$0.100$0.400262K
Nemotron-3-8BMid-tier$0.100$0.400128K
Nemotron-3-4BLite$0.050$0.200128K