Back to AI models

Llama-3.1-Nemotron-Ultra-253B-v1

An NVIDIA-tuned Llama model for advanced reasoning and enterprise agent workflows.

Nemotron Reasoning LLM Open weights Open weights: Yes API: NVIDIA NIM and self-hosted

Plain-English Summary

An NVIDIA-tuned Llama model for advanced reasoning and enterprise agent workflows.

Technical Notes

253B-parameter reasoning model derived from Llama 3.1, optimized for instruction following and agentic tasks.

Best For

Reasoning, enterprise agents, synthetic data, self-hosted research

Watch Out For

Large hardware requirement and older than Nemotron 3 Ultra.

Model Specs

Maker
NVIDIA
Country
USA
Family
Nemotron
Type
Reasoning LLM
Release date
2025-04-08
Date confidence
Exact
Status
Open weights
Context window
131,072 tokens
Max output
Not listed
Parameters
253B
Active parameters
253B
Architecture
Dense Llama-derived transformer
Reasoning
Yes
Tool calling
Yes / host dependent
Structured output
Host dependent
API available
NVIDIA NIM and self-hosted
Open weights
Yes
Self-hostable
Yes
License
NVIDIA Open Model License / Llama terms
Inputs
Text
Outputs
Text
Approx. price
See provider pricing
How to access
Provider website
Model / API ID
Llama-3.1-Nemotron-Ultra-253B-v1
Knowledge cutoff
Not publicly disclosed
Fine-tuning
Fine-tuning, NIM, self-hosting
Completeness
95 %
Verified
Jul 18, 2026
Snapshot date
Jul 18, 2026
Source
Official source