Unsloth is an optimised fine-tuning library that makes LoRA training 2-5× faster with 70% less VRAM. This guide walks through fine-tuning a Llama or Mistral model on your own data using Unsloth on a free Google Colab GPU.
Fine-tuning is the right choice when: you have 500+ high-quality examples of the task; the task requires a consistent style or format that prompting cannot reliably produce; you need the lowest possible inference cost for a specialised task; or you want to inject domain knowledge not present in the base model.
Fine-tuning is overkill when: a well-crafted system prompt and few-shot examples achieve 90%+ of the quality you need; you have fewer than 200 training examples; or you are still exploring the problem space. Always try prompt engineering first.
Low-Rank Adaptation (LoRA) fine-tunes a model by adding small trainable adapter matrices to the attention layers, rather than updating all weights. The base model is frozen; only the adapters (0.1-2% of total parameters) are trained. This drastically reduces VRAM requirements and training time.
QLoRA combines LoRA with 4-bit quantisation of the base model, enabling fine-tuning of a 7B model on a single 16 GB GPU, or a 13B model on a 24 GB GPU. Unsloth implements an optimised version of QLoRA that is significantly faster than the reference implementation.
Open a new Colab notebook with a T4 (free) or A100 (Pro) GPU. Install Unsloth: `!pip install 'unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git'` followed by `!pip install --no-deps trl peft accelerate bitsandbytes`.
Load a model: `from unsloth import FastLanguageModel; model, tokenizer = FastLanguageModel.from_pretrained(model_name='unsloth/Phi-4-mini-instruct', max_seq_length=2048, dtype=None, load_in_4bit=True)`. Unsloth supports Llama 3, Mistral, Phi, Gemma, and Qwen families.
Format your data as a list of instruction-response pairs. Each example should be a dict with an `instruction` (the input) and `output` (the expected response). For chat models, use the model's chat template to format examples correctly.
Use `datasets` to load and process your data: `from datasets import Dataset; data = Dataset.from_list(your_list)`. Add a formatting function that applies the chat template and stores the result in a `text` field. Aim for at least 500 examples; 2000+ is better for reliable results.
Add LoRA adapters: `model = FastLanguageModel.get_peft_model(model, r=16, target_modules=['q_proj', 'v_proj'], lora_alpha=16, lora_dropout=0)`. `r=16` is a good starting value; higher r means more parameters but better fit.
Set up the Trainer with SFTTrainer from the `trl` library: `trainer = SFTTrainer(model=model, train_dataset=dataset, dataset_text_field='text', max_seq_length=2048, args=TrainingArguments(per_device_train_batch_size=2, num_train_epochs=3, ...))`. Start training with `trainer.train()`. On a T4 GPU, 1000 examples with a 7B model takes about 20-30 minutes.
Save as GGUF for local deployment: `model.save_pretrained_gguf('my-model', tokenizer, quantization_method='q4_k_m')`. This exports a GGUF file you can run directly with Ollama or LM Studio.
Push to Hugging Face Hub: `model.push_to_hub_gguf('your-username/my-model', tokenizer, quantization_method='q4_k_m', token='hf_...')`. Once uploaded, you can pull it with Ollama from any machine.
The quality of your training data is the biggest factor in fine-tuning success. This guide covers data collection strategies, formatting standards, quality filtering, and the minimum viable dataset size for different tasks.
Read guideKnowing when your fine-tuned model is actually better than the base model requires systematic evaluation. This guide covers benchmark datasets, LLM-as-judge evaluation, and metrics for task-specific assessment.
Read guideDPO trains models to prefer good responses over bad ones using human preference data, without the complexity of reinforcement learning. This guide covers collecting preference data, training with TRL's DPO trainer, and evaluating results.
Read guide