Instructions to use kaitchup/Llama-3.3-70B-Instruct-AutoRound-GPTQ-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use kaitchup/Llama-3.3-70B-Instruct-AutoRound-GPTQ-4bit with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="kaitchup/Llama-3.3-70B-Instruct-AutoRound-GPTQ-4bit") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("kaitchup/Llama-3.3-70B-Instruct-AutoRound-GPTQ-4bit") model = AutoModelForCausalLM.from_pretrained("kaitchup/Llama-3.3-70B-Instruct-AutoRound-GPTQ-4bit", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use kaitchup/Llama-3.3-70B-Instruct-AutoRound-GPTQ-4bit with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "kaitchup/Llama-3.3-70B-Instruct-AutoRound-GPTQ-4bit" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kaitchup/Llama-3.3-70B-Instruct-AutoRound-GPTQ-4bit", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/kaitchup/Llama-3.3-70B-Instruct-AutoRound-GPTQ-4bit
- SGLang
How to use kaitchup/Llama-3.3-70B-Instruct-AutoRound-GPTQ-4bit with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "kaitchup/Llama-3.3-70B-Instruct-AutoRound-GPTQ-4bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kaitchup/Llama-3.3-70B-Instruct-AutoRound-GPTQ-4bit", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "kaitchup/Llama-3.3-70B-Instruct-AutoRound-GPTQ-4bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kaitchup/Llama-3.3-70B-Instruct-AutoRound-GPTQ-4bit", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use kaitchup/Llama-3.3-70B-Instruct-AutoRound-GPTQ-4bit with Docker Model Runner:
docker model run hf.co/kaitchup/Llama-3.3-70B-Instruct-AutoRound-GPTQ-4bit
Your quants are not listed in the base model
Hey Kaitchup,
First I want to thank you for the work you do.
I have 2xAMD MI100 and your quants are the fastest (which work) on that device. Especially in prompt processing.
For Example for your quantization of the Llama-3.3-70B-Instruct in vllm I get 391 t/s of prompt eval and 19 t/s generation.
The other closest quant gives me only 320 t/s of the prompt eval.
Big + is the the fact that you do not use bf16 which is not supported on AMD MI60 and MI100
Back to the topic.
I'm looking for all quants going through the base model card in Hugginface and then clicking on "Quaternizations"
And your quants are not in the list.
There is another account which uses exactly the same names for quants. But he is using BF16 and they always fail on the MI100.
Your quants were recommended on Reddit without providing full link and i was puzzled when originally they didn't work.
I could find your version only using Google.
Hi,
I don't know how this is listed by Hugging Face. I thought this was done automatically but if my model is not there, something must misconfigured.
Thank you for letting me know.
