Unsloth Docs
Train your own model with Unsloth, an open-source framework for LLM fine-tuning and reinforcement learning.
At Unsloth, our mission is to make AI as accurate and accessible as possible. Train, run, evaluate and save gpt-oss, Llama, DeepSeek, TTS, Qwen, Mistral, Gemma LLMs 2x faster with 70% less VRAM.
Our docs will guide you through running & training your own model locally.
Get started Our GitHub
{% columns %} {% column %} {% content-ref url="fine-tuning-llms-guide" %} fine-tuning-llms-guide {% endcontent-ref %}
{% content-ref url="unsloth-notebooks" %} unsloth-notebooks {% endcontent-ref %}
{% endcolumn %}
{% column %} {% content-ref url="all-our-models" %} all-our-models {% endcontent-ref %}
{% content-ref url="../models/tutorials-how-to-fine-tune-and-run-llms" %} tutorials-how-to-fine-tune-and-run-llms {% endcontent-ref %} {% endcolumn %} {% endcolumns %}
🦥 Why Unsloth?
- Unsloth streamlines model training locally and on Colab/Kaggle, covering loading, quantization, training, evaluation, saving, exporting, and integration with inference engines like Ollama, llama.cpp, and vLLM.
- We directly collaborate with teams behind gpt-oss, Qwen3, Llama 4, Mistral, Google (Gemma 1–3) and Phi-4, where we’ve fixed critical bugs in models that greatly improved model accuracy.
- Unsloth is the only training framework to support all model types: vision, text-to-speech (TTS), BERT, reinforcement learning (RL) while remaining highly customizable with flexible chat templates, dataset formatting and ready-to-use notebooks.
⭐ Key Features
- Supports full-finetuning, pretraining, 4-bit, 16-bit and 8-bit training.
- The most efficient RL library, using 80% less VRAM. Supports GRPO, GSPO etc.
- Supports all models: TTS, multimodal, BERT and more. Any model that works in transformers works in Unsloth.
- 0% loss in accuracy - no approximation methods - all exact.
- MultiGPU works already but a much better version is coming!
- Unsloth supports Linux, Windows, Colab, Kaggle, NVIDIA and AMD & Intel. See:
{% content-ref url="beginner-start-here/unsloth-requirements" %} unsloth-requirements {% endcontent-ref %}
Quickstart
Install locally with pip (recommended) for Linux or WSL devices:
pip install unsloth
Use our official Docker image: unsloth/unsloth. Read our Docker guide.
For Windows install instructions, see here.
{% content-ref url="install-and-update" %} install-and-update {% endcontent-ref %}
What is Fine-tuning and RL? Why?
Fine-tuning an LLM customizes its behavior, enhances domain knowledge, and optimizes performance for specific tasks. By fine-tuning a pre-trained model (e.g. Llama-3.1-8B) on a dataset, you can:
- Update Knowledge: Introduce new domain-specific information.
- Customize Behavior: Adjust the model’s tone, personality, or response style.
- Optimize for Tasks: Improve accuracy and relevance for specific use cases.
Reinforcement Learning (RL) is where an "agent" learns to make decisions by interacting with an environment and receiving feedback in the form of rewards or penalties.
- Action: What the model generates (e.g. a sentence).
- Reward: A signal indicating how good or bad the model's action was (e.g. did the response follow instructions? was it helpful?).
- Environment: The scenario or task the model is working on (e.g. answering a user’s question).
Example use-cases of fine-tuning or RL:
- Train LLM to predict if a headline impacts a company positively or negatively.
- Use historical customer interactions for more accurate and custom responses.
- Train LLM on legal texts for contract analysis, case law research, and compliance.
You can think of a fine-tuned model as a specialized agent designed to do specific tasks more effectively and efficiently. Fine-tuning can replicate all of RAG's capabilities, but not vice versa.
{% content-ref url="beginner-start-here/faq-+-is-fine-tuning-right-for-me" %} faq-+-is-fine-tuning-right-for-me {% endcontent-ref %}
{% content-ref url="reinforcement-learning-rl-guide" %} reinforcement-learning-rl-guide {% endcontent-ref %}
Beginner? Start here!
If you're a beginner, here might be the first questions you'll ask before your first fine-tune. You can also always ask our community by joining our Reddit page.
Unsloth Requirements
Here are Unsloth's requirements including system and GPU VRAM requirements.
System Requirements
- Operating System: Works on Linux and Windows.
- Supports NVIDIA GPUs since 2018+ including Blackwell RTX 50 and DGX Spark.
Minimum CUDA Capability 7.0 (V100, T4, Titan V, RTX 20 & 50, A100, H100, L40 etc) Check your GPU! GTX 1070, 1080 works, but is slow. - The official Unsloth Docker image
unsloth/unslothis available on Docker Hub. - Unsloth works on AMD and Intel GPUs! Apple/Silicon/MLX is in the works.
- If you have different versions of torch, transformers etc.,
pip install unslothwill automatically install all the latest versions of those libraries so you don't need to worry about version compatibility. - Your device should have
xformers,torch,BitsandBytesandtritonsupport.
{% hint style="info" %} Python 3.13 is now supported! {% endhint %}
Fine-tuning VRAM requirements:
How much GPU memory do I need for LLM fine-tuning using Unsloth?
{% hint style="info" %} A common issue when you OOM or run out of memory is because you set your batch size too high. Set it to 1, 2, or 3 to use less VRAM.
For context length benchmarks, see here. {% endhint %}
Check this table for VRAM requirements sorted by model parameters and fine-tuning method. QLoRA uses 4-bit, LoRA uses 16-bit. Keep in mind that sometimes more VRAM is required depending on the model so these numbers are the absolute minimum:
| Model parameters | QLoRA (4-bit) VRAM | LoRA (16-bit) VRAM |
|---|---|---|
| 3B | 3.5 GB | 8 GB |
| 7B | 5 GB | 19 GB |
| 8B | 6 GB | 22 GB |
| 9B | 6.5 GB | 24 GB |
| 11B | 7.5 GB | 29 GB |
| 14B | 8.5 GB | 33 GB |
| 27B | 22GB | 64GB |
| 32B | 26 GB | 76 GB |
| 40B | 30GB | 96GB |
| 70B | 41 GB | 164 GB |
| 81B | 48GB | 192GB |
| 90B | 53GB | 212GB |
| 405B | 237 GB | 950 GB |
FAQ + Is Fine-tuning Right For Me?
If you're stuck on if fine-tuning is right for you, see here! Learn about fine-tuning misconceptions, how it compared to RAG and more:
Understanding Fine-Tuning
Fine-tuning an LLM customizes its behavior, deepens its domain expertise, and optimizes its performance for specific tasks. By refining a pre-trained model (e.g. Llama-3.1-8B) with specialized data, you can:
- Update Knowledge – Introduce new, domain-specific information that the base model didn’t originally include.
- Customize Behavior – Adjust the model’s tone, personality, or response style to fit specific needs or a brand voice.
- Optimize for Tasks – Improve accuracy and relevance on particular tasks or queries your use-case requires.
Think of fine-tuning as creating a specialized expert out of a generalist model. Some debate whether to use Retrieval-Augmented Generation (RAG) instead of fine-tuning, but fine-tuning can incorporate knowledge and behaviors directly into the model in ways RAG cannot. In practice, combining both approaches yields the best results - leading to greater accuracy, better usability, and fewer hallucinations.
Real-World Applications of Fine-Tuning
Fine-tuning can be applied across various domains and needs. Here are a few practical examples of how it makes a difference:
- Sentiment Analysis for Finance – Train an LLM to determine if a news headline impacts a company positively or negatively, tailoring its understanding to financial context.
- Customer Support Chatbots – Fine-tune on past customer interactions to provide more accurate and personalized responses in a company’s style and terminology.
- Legal Document Assistance – Fine-tune on legal texts (contracts, case law, regulations) for tasks like contract analysis, case law research, or compliance support, ensuring the model uses precise legal language.
The Benefits of Fine-Tuning
Fine-tuning offers several notable benefits beyond what a base model or a purely retrieval-based system can provide:
Fine-Tuning vs. RAG: What’s the Difference?
Fine-tuning can do mostly everything RAG can - but not the other way around. During training, fine-tuning embeds external knowledge directly into the model. This allows the model to handle niche queries, summarize documents, and maintain context without relying on an outside retrieval system. That’s not to say RAG lacks advantages as it is excels at accessing up-to-date information from external databases. It is in fact possible to retrieve fresh data with fine-tuning as well, however it is better to combine RAG with fine-tuning for efficiency.
Task-Specific Mastery
Fine-tuning deeply integrates domain knowledge into the model. This makes it highly effective at handling structured, repetitive, or nuanced queries, scenarios where RAG-alone systems often struggle. In other words, a fine-tuned model becomes a specialist in the tasks or content it was trained on.
Independence from Retrieval
A fine-tuned model has no dependency on external data sources at inference time. It remains reliable even if a connected retrieval system fails or is incomplete, because all needed information is already within the model’s own parameters. This self-sufficiency means fewer points of failure in production.
Faster Responses
Fine-tuned models don’t need to call out to an external knowledge base during generation. Skipping the retrieval step means they can produce answers much more quickly. This speed makes fine-tuned models ideal for time-sensitive applications where every second counts.
Custom Behavior and Tone
Fine-tuning allows precise control over how the model communicates. This ensures the model’s responses stay consistent with a brand’s voice, adhere to regulatory requirements, or match specific tone preferences. You get a model that not only knows what to say, but how to say it in the desired style.
Reliable Performance
Even in a hybrid setup that uses both fine-tuning and RAG, the fine-tuned model provides a reliable fallback. If the retrieval component fails to find the right information or returns incorrect data, the model’s built-in knowledge can still generate a useful answer. This guarantees more consistent and robust performance for your system.
Common Misconceptions
Despite fine-tuning’s advantages, a few myths persist. Let’s address two of the most common misconceptions about fine-tuning:
Does Fine-Tuning Add New Knowledge to a Model?
Yes - it absolutely can. A common myth suggests that fine-tuning doesn’t introduce new knowledge, but in reality it does. If your fine-tuning dataset contains new domain-specific information, the model will learn that content during training and incorporate it into its responses. In effect, fine-tuning can and does teach the model new facts and patterns from scratch.
Is RAG Always Better Than Fine-Tuning?
Not necessarily. Many assume RAG will consistently outperform a fine-tuned model, but that’s not the case when fine-tuning is done properly. In fact, a well-tuned model often matches or even surpasses RAG-based systems on specialized tasks. Claims that “RAG is always better” usually stem from fine-tuning attempts that weren’t optimally configured - for example, using incorrect LoRA parameters or insufficient training.
Unsloth takes care of these complexities by automatically selecting the best parameter configurations for you. All you need is a good-quality dataset, and you'll get a fine-tuned model that performs to its fullest potential.
Is Fine-Tuning Expensive?
Not at all! While full fine-tuning or pretraining can be costly, these are not necessary (pretraining is especially not necessary). In most cases, LoRA or QLoRA fine-tuning can be done for minimal cost. In fact, with Unsloth’s free notebooks for Colab or Kaggle, you can fine-tune models without spending a dime. Better yet, you can even fine-tune locally on your own device.
FAQ:
Why You Should Combine RAG & Fine-Tuning
Instead of choosing between RAG and fine-tuning, consider using both together for the best results. Combining a retrieval system with a fine-tuned model brings out the strengths of each approach. Here’s why:
- Task-Specific Expertise – Fine-tuning excels at specialized tasks or formats (making the model an expert in a specific area), while RAG keeps the model up-to-date with the latest external knowledge.
- Better Adaptability – A fine-tuned model can still give useful answers even if the retrieval component fails or returns incomplete information. Meanwhile, RAG ensures the system stays current without requiring you to retrain the model for every new piece of data.
- Efficiency – Fine-tuning provides a strong foundational knowledge base within the model, and RAG handles dynamic or quickly-changing details without the need for exhaustive re-training from scratch. This balance yields an efficient workflow and reduces overall compute costs.
LoRA vs. QLoRA: Which One to Use?
When it comes to implementing fine-tuning, two popular techniques can dramatically cut down the compute and memory requirements: LoRA and QLoRA. Here’s a quick comparison of each:
- LoRA (Low-Rank Adaptation) – Fine-tunes only a small set of additional “adapter” weight matrices (in 16-bit precision), while leaving most of the original model unchanged. This significantly reduces the number of parameters that need updating during training.
- QLoRA (Quantized LoRA) – Combines LoRA with 4-bit quantization of the model weights, enabling efficient fine-tuning of very large models on minimal hardware. By using 4-bit precision where possible, it dramatically lowers memory usage and compute overhead.
We recommend starting with QLoRA, as it’s one of the most efficient and accessible methods available. Thanks to Unsloth’s dynamic 4-bit quants, the accuracy loss compared to standard 16-bit LoRA fine-tuning is now negligible.
Experimentation is Key
There’s no single “best” approach to fine-tuning - only best practices for different scenarios. It’s important to experiment with different methods and configurations to find what works best for your dataset and use case. A great starting point is QLoRA (4-bit), which offers a very cost-effective, resource-friendly way to fine-tune models without heavy computational requirements.
{% content-ref url="../fine-tuning-llms-guide/lora-hyperparameters-guide" %} lora-hyperparameters-guide {% endcontent-ref %}
Unsloth Notebooks
Explore our catalog of Unsloth notebooks:
Also see our GitHub repo for our notebooks: github.com/unslothai/notebooks
GRPO (RL)Text-to-speechVisionUse-caseKaggle
Colab notebooks
Standard notebooks:
- gpt-oss (20b) • Inference • Fine-tuning
- DeepSeek-OCR - new
- Qwen3 (14B) • Qwen3-VL (8B) - new
- Qwen3-2507-4B • Thinking • Instruct
- Gemma 3n (E4B) • Text • Vision • Audio
- IBM Granite-4.0-H - new
- Gemma 3 (4B) • Text • Vision • 270M - new
- Phi-4 (14B)
- Llama 3.1 (8B) • Llama 3.2 (1B + 3B)
GRPO (Reasoning RL) notebooks:
- gpt-oss-20b (automatic kernels creation) - new
- gpt-oss-20b (auto win 2048 game) - new
- Qwen3-VL (8B) - Vision GSPO - new
- Qwen3 (4B) - Advanced GRPO LoRA
- Gemma 3 (4B) - Vision GSPO - new
- DeepSeek-R1-0528-Qwen3 (8B) (for multilingual usecase)
- Gemma 3 (1B)
- Llama 3.2 (3B) - Advanced GRPO LoRA
- Llama 3.1 (8B)
- Phi-4 (14B)
- Mistral v0.3 (7B)
Text-to-Speech (TTS) notebooks:
- Sesame-CSM (1B) - new
- Orpheus-TTS (3B)
- Whisper Large V3 - Speech-to-Text (STT)
- Llasa-TTS (1B)
- Spark-TTS (0.5B)
- Oute-TTS (1B)
Speech-to-Text (SST) notebooks:
- Whisper-Large-V3
- Gemma 3n (E4B) - Audio
Vision (Multimodal) notebooks:
- Qwen3-VL (8B) - new
- DeepSeek-OCR - new
- Gemma 3 (4B) - vision
- Gemma 3n (E4B) - vision
- Llama 3.2 Vision (11B)
- Qwen2.5-VL (7B)
- Pixtral (12B) 2409
- Qwen3-VL - Vision GSPO - new
- Qwen2.5-VL - Vision GSPO
- Gemma 3 (4B) - Vision GSPO - new
Large LLM notebooks:
Notebooks for large models: These exceed Colab’s free 15 GB VRAM tier. With Colab’s new 80 GB GPUs, you can fine-tune 120B parameter models.
{% hint style="info" %} Colab subscription or credits are required. We don't earn anything from these notebooks. {% endhint %}
- gpt-oss-120b - new
- Qwen3 (32B) - new
- Llama 3.3 (70B) - new
- Gemma 3 (27B) - new
Other important notebooks:
- Customer support agent - new
- Automatic Kernel Creation with RL - new
- ModernBERT-large - new as of Aug 19
- Synthetic Data Generation Llama 3.2 (3B) - new
- Tool Calling - new
- Customer support agent - new
- Mistral v0.3 Instruct (7B)
- Ollama
- ORPO
- Continued Pretraining
- DPO Zephyr
- Inference only
- Llama 3 (8B)
Specific use-case notebooks:
- Customer support agent - new
- Automatic Kernel Creation with RL - new
- DPO Zephyr
- BERT - Text Classification - new as of Aug 19
- Ollama
- Tool Calling - new
- Continued Pretraining (CPT)
- Multiple Datasets by Flail
- KTO by Jeffrey
- Inference chat UI
- Conversational
- ChatML
- Text Completion
Rest of notebooks:
- Qwen2.5 (3B)
- Gemma 2 (9B)
- Mistral NeMo (12B)
- Phi-3.5 (mini)
- Phi-3 (medium)
- Gemma 2 (2B)
- Qwen 2.5 Coder (14B)
- Mistral Small (22B)
- TinyLlama
- CodeGemma (7B)
- Mistral v0.3 (7B)
- Qwen2 (7B)
Kaggle notebooks
Standard notebooks:
- gpt-oss (20B) - new
- Gemma 3n (E4B)
- Qwen3 (14B)
- Magistral-2509 (24B) - new
- Gemma 3 (4B)
- Phi-4 (14B)
- Llama 3.1 (8B)
- Llama 3.2 (1B + 3B)
- Qwen 2.5 (7B)
GRPO (Reasoning) notebooks:
- Qwen2.5-VL - Vision GRPO - new
- Qwen3 (4B)
- Gemma 3 (1B)
- Llama 3.1 (8B)
- [Phi-4 (14B)](https://www.kaggle.com/notebooks/welcome?src=https%3A%2F%2Fgithub.com%2Funslothai/notebooks/blob/main/nb/
…(truncated)