Instructions to use nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="nvidia/Llama-3_1-Nemotron-Ultra-253B-v1", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("nvidia/Llama-3_1-Nemotron-Ultra-253B-v1", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nvidia/Llama-3_1-Nemotron-Ultra-253B-v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/Llama-3_1-Nemotron-Ultra-253B-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/nvidia/Llama-3_1-Nemotron-Ultra-253B-v1
- SGLang
How to use nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "nvidia/Llama-3_1-Nemotron-Ultra-253B-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/Llama-3_1-Nemotron-Ultra-253B-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "nvidia/Llama-3_1-Nemotron-Ultra-253B-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/Llama-3_1-Nemotron-Ultra-253B-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 with Docker Model Runner:
docker model run hf.co/nvidia/Llama-3_1-Nemotron-Ultra-253B-v1
AWQ OR GPTQ Quant
huggingface/modules/transformers_modules/Llama3.1-253B/configuration_decilm.py", line 22, in
from .block_config import BlockConfig
ModuleNotFoundError: No module named 'transformers_modules.Llama3'
I tried transformers version mentioned in the readme, along with the latest. I can not get either to proceed with quantization. Is there something I am missing? Would be awesome if you could point me in the correct direction or someone would upload a quantized version.
from transformers import AutoTokenizer, AutoModelForCausalLM
MODEL_ID = "models/Llama3.1-253B"
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID, device_map="auto", torch_dtype="auto", trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
from datasets import load_dataset
NUM_CALIBRATION_SAMPLES = 512
MAX_SEQUENCE_LENGTH = 2048
Load dataset.
ds = load_dataset("HuggingFaceFW/fineweb", split="train_sft")
ds = ds.shuffle(seed=42).select(range(NUM_CALIBRATION_SAMPLES))
Preprocess the data into the format the model is trained with.
def preprocess(example):
return {
"text": tokenizer.apply_chat_template(
example["messages"],
tokenize=False,
)
}
ds = ds.map(preprocess)
Tokenize the data (be careful with bos tokens - we need add_special_tokens=False since the chat_template already added it).
def tokenize(sample):
return tokenizer(
sample["text"],
padding=False,
max_length=MAX_SEQUENCE_LENGTH,
truncation=True,
add_special_tokens=False,
)
ds = ds.map(tokenize, remove_columns=ds.column_names)
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import GPTQModifier
Configure the quantization algorithm to run.
recipe = GPTQModifier(targets="Linear", scheme="W4A16", ignore=["lm_head"])
Apply quantization.
oneshot(
model=model,
dataset=ds,
recipe=recipe,
max_seq_length=MAX_SEQUENCE_LENGTH,
num_calibration_samples=NUM_CALIBRATION_SAMPLES,
)
Save to disk compressed.
SAVE_DIR = MODEL_ID.split("/")[1] + "-W4A16-G128"
model.save_pretrained(SAVE_DIR, save_compressed=True)
tokenizer.save_pretrained(SAVE_DIR)
Little chefs (miniature people), full of energy and enthusiasm, prepare traditional Syrian cheese dessert in a bright, imaginative kitchen. The scene is full of movement, vitality, and a joyful atmosphere—the little chefs pull and stretch cheese dough with team spirit and joy.
They stand on spoons, climb on mixing bowls, and roll the dough on cutting boards that appear enormous compared to their size. The kitchen atmosphere is playful and dreamy, with glowing pastel colors: soft white, warm beige, light turquoise, and pistachio green.
The ingredients—cylindrical halawa jibneh, cheese, sweet syrup, and ground pistachios—glow under the dim lighting, creating a magical, storybook atmosphere. The whole scene feels like a fun and lively culinary adventure in a miniature world!