Fine-Tuning Qwen3-4B with Unsloth Kaggle Notebook
Status: 馃殌 Completed
Tech Stack: Python, PyTorch, Unsloth, Hugging Face (Transformers, TRL, Datasets), QLoRA, Groq API, Weights & Biases, ipywidgets.
ProTutor is a custom 4-billion parameter AI tutor fine-tuned to teach software engineering concepts. This project demonstrates an end-to-end Large Language Model (LLM) distillation pipeline, leveraging a 120B parameter teacher model to generate synthetic training data, which is then used to fine-tune a smaller student model (Qwen3-4B).
Built entirely on a free Kaggle T4 GPU, this pipeline utilizes QLoRA (Quantized Low-Rank Adaptation) and the Unsloth framework to optimize memory usage and accelerate training speed, completing the entire fine-tuning process in under 30 minutes.
This project is designed to run natively on a Kaggle Notebook with a T4 GPU (脳2) or any local Jupyter environment with at least 16GB VRAM.
- Environment Setup (Kaggle): Ensure Internet and GPU T4 are enabled in notebook settings.
- API Keys: Add your free Groq API key to Kaggle Secrets as
GROQ_API_KEY. - Install Dependencies:
Run the following command to install the required stack (Note: Unsloth requires specific PyTorch/CUDA wheels):
pip install -r requirements.txt
- Synthetic Data Generation (Distillation): - Utilized the
openai/gpt-oss-120bmodel via the Groq API to asynchronously generate ~300 high-quality question-answer pairs.- Enforced a strict zero-shot persona (standard English, universal analogies, specific summary formatting) without relying on system prompts at inference.
- Formatting & Tokenization: - Serialized raw JSON data into the exact
ChatMLtemplate (<|im_start|>,<|im_end|>) expected by the Qwen instruct model. - Model Initialization & QLoRA: - Loaded the
Qwen3-4B-Instructbase model in 4-bit precision using Unsloth.- Attached trainable LoRA matrices (Rank = 32) to the attention and MLP projections, targeting roughly 1.6% (66M) of the total model parameters.
- Training Loop: - Executed supervised fine-tuning using TRL's
SFTTrainer.- Applied
train_on_responses_onlyto mask the loss on user prompts, forcing the model to strictly optimize for the assistant's persona. - Tracked training loss and metrics via Weights & Biases.
- Applied
- Interactive Evaluation & Export: - Built a custom, interactive chat UI within the notebook using
ipywidgetsto test the held-out validation set.- Exported the final trained weights to GGUF format for local/edge deployment.