To run this, press "*Runtime*" and press "*Run all*" on a **free** Tesla T4 Google Colab instance!
<div class="align-center">
<a href="https://unsloth.ai/"><img src="https://github.com/unslothai/unsloth/raw/main/images/unsloth%20new%20logo.png" width="115"></a>
<a href="https://discord.gg/unsloth"><img src="https://github.com/unslothai/unsloth/raw/main/images/Discord button.png" width="145"></a>
<a href="https://unsloth.ai/docs/"><img src="https://github.com/unslothai/unsloth/blob/main/images/documentation%20green%20button.png?raw=true" width="125"></a> Join Discord if you need help + ‚≠ê <i>Star us on <a href="https://github.com/unslothai/unsloth">Github</a> </i> ‚≠ê
</div>

To install Unsloth on your local device, follow [our guide](https://unsloth.ai/docs/get-started/install). This notebook is licensed [LGPL-3.0](https://github.com/unslothai/notebooks?tab=LGPL-3.0-1-ov-file#readme).

You will learn how to do [data prep](#Data), how to [train](#Train), how to [run the model](#Inference), & how to save it

### News

Train MoEs - DeepSeek, GLM, Qwen and gpt-oss 12x faster with 35% less VRAM. [Blog](https://unsloth.ai/docs/new/faster-moe)

You can now train embedding models 1.8-3.3x faster with 20% less VRAM. [Blog](https://unsloth.ai/docs/new/embedding-finetuning)

Ultra Long-Context Reinforcement Learning is here with 7x more context windows! [Blog](https://unsloth.ai/docs/new/grpo-long-context)

3x faster LLM training with 30% less VRAM and 500K context. [3x faster](https://unsloth.ai/docs/new/3x-faster-training-packing) ‚Ä¢ [500K Context](https://unsloth.ai/docs/new/500k-context-length-fine-tuning)

New in Reinforcement Learning: [FP8 RL](https://unsloth.ai/docs/new/fp8-reinforcement-learning) ‚Ä¢ [Vision RL](https://unsloth.ai/docs/new/vision-reinforcement-learning-vlm-rl) ‚Ä¢ [Standby](https://unsloth.ai/docs/basics/memory-efficient-rl) ‚Ä¢ [gpt-oss RL](https://unsloth.ai/docs/new/gpt-oss-reinforcement-learning)

Visit our docs for all our [model uploads](https://unsloth.ai/docs/get-started/unsloth-model-catalog) and [notebooks](https://unsloth.ai/docs/get-started/unsloth-notebooks).

### Installation

In [1]:
%%capture
import os, re
if "COLAB_" not in "".join(os.environ.keys()):
    !pip install unsloth  # Do this in local & cloud setups
else:
    import torch; v = re.match(r'[\d]{1,}\.[\d]{1,}', str(torch.__version__)).group(0)
    xformers = 'xformers==' + {'2.10':'0.0.34','2.9':'0.0.33.post1','2.8':'0.0.32.post2'}.get(v, "0.0.34")
    !pip install sentencepiece protobuf "datasets==4.3.0" "huggingface_hub>=0.34.0" hf_transfer
    !pip install --no-deps unsloth_zoo bitsandbytes accelerate {xformers} peft trl triton unsloth
!pip install transformers==4.56.2
!pip install --no-deps trl==0.22.2

### Unsloth

In [2]:
from unsloth import FastLanguageModel
import torch
max_seq_length = 2048 # Choose any! We auto support RoPE Scaling internally!
dtype = None # None for auto detection. Float16 for Tesla T4, V100, Bfloat16 for Ampere+
load_in_4bit = True # Use 4bit quantization to reduce memory usage. Can be False.

# 4bit pre quantized models we support for 4x faster downloading + no OOMs.
fourbit_models = [
    "unsloth/Llama-3.1-8B-bnb-4bit",      # Llama-3.1 2x faster
    "unsloth/Llama-3.1-8B-Instruct-bnb-4bit",
    "unsloth/Llama-3.1-70B-bnb-4bit",
    "unsloth/Llama-3.1-405B-bnb-4bit",    # 4bit for 405b!
    "unsloth/Mistral-Small-Instruct-2409",     # Mistral 22b 2x faster!
    "unsloth/mistral-7b-instruct-v0.3-bnb-4bit",
    "unsloth/Phi-3.5-mini-instruct",           # Phi-3.5 2x faster!
    "unsloth/Phi-3-medium-4k-instruct",
    "unsloth/gemma-2-9b-bnb-4bit",
    "unsloth/gemma-2-27b-bnb-4bit",            # Gemma 2x faster!

    "unsloth/Llama-3.2-1B-bnb-4bit",           # NEW! Llama 3.2 models
    "unsloth/Llama-3.2-1B-Instruct-bnb-4bit",
    "unsloth/Llama-3.2-3B-bnb-4bit",
    "unsloth/Llama-3.2-3B-Instruct-bnb-4bit",

    "unsloth/Llama-3.3-70B-Instruct-bnb-4bit" # NEW! Llama 3.3 70B!
] # More models at https://huggingface.co/unsloth

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name = "unsloth/Llama-3.2-3B-Instruct", # or choose "unsloth/Llama-3.2-1B-Instruct"
    max_seq_length = max_seq_length,
    dtype = dtype,
    load_in_4bit = load_in_4bit,
    # token = "YOUR_HF_TOKEN", # HF Token for gated models
)

ü¶• Unsloth: Will patch your computer to enable 2x faster free finetuning.
ü¶• Unsloth Zoo will now patch everything to make training faster!
==((====))==  Unsloth 2026.2.1: Fast Llama patching. Transformers: 4.56.2.
   \\   /|    Tesla T4. Num GPUs = 1. Max memory: 14.563 GB. Platform: Linux.
O^O/ \_/ \    Torch: 2.9.0+cu128. CUDA: 7.5. CUDA Toolkit: 12.8. Triton: 3.5.0
\        /    Bfloat16 = FALSE. FA [Xformers = 0.0.33.post1. FA2 = False]
 "-____-"     Free license: http://github.com/unslothai/unsloth
Unsloth: Fast downloading is enabled - ignore downloading bars which are red colored!


We now add LoRA adapters so we only need to update 1 to 10% of all parameters!

In [3]:
model = FastLanguageModel.get_peft_model(
    model,
    r = 16, # Choose any number > 0 ! Suggested 8, 16, 32, 64, 128
    target_modules = ["q_proj", "k_proj", "v_proj", "o_proj",
                      "gate_proj", "up_proj", "down_proj",],
    lora_alpha = 16,
    lora_dropout = 0, # Supports any, but = 0 is optimized
    bias = "none",    # Supports any, but = "none" is optimized
    # [NEW] "unsloth" uses 30% less VRAM, fits 2x larger batch sizes!
    use_gradient_checkpointing = "unsloth", # True or "unsloth" for very long context
    random_state = 3407,
    use_rslora = False,  # We support rank stabilized LoRA
    loftq_config = None, # And LoftQ
)

Unsloth 2026.2.1 patched 28 layers with 28 QKV layers, 28 O layers and 28 MLP layers.


<a name="Data"></a>
### Data Prep
We now use the `Llama-3.1` format for conversation style finetunes. We use [Maxime Labonne's FineTome-100k](https://huggingface.co/datasets/mlabonne/FineTome-100k) dataset in ShareGPT style. But we convert it to HuggingFace's normal multiturn format `("role", "content")` instead of `("from", "value")`/ Llama-3 renders multi turn conversations like below:

```
<|begin_of_text|><|start_header_id|>user<|end_header_id|>

Hello!<|eot_id|><|start_header_id|>assistant<|end_header_id|>

Hey there! How are you?<|eot_id|><|start_header_id|>user<|end_header_id|>

I'm great thanks!<|eot_id|>
```

We use our `get_chat_template` function to get the correct chat template. We support `zephyr, chatml, mistral, llama, alpaca, vicuna, vicuna_old, phi3, llama3` and more.

In [4]:
import json
from datasets import load_dataset
from unsloth.chat_templates import get_chat_template

raw_data = [
  {"id": "1_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: Holiday at home. It is 10:00 a.m. The child slept at 2:00 a.m. and is still in bed.\nParent message: (Pulls the curtains open) It's almost noon! Look at your face‚Äîso pale. Are you trying to die? If you get sick, don't come to me. I don't have money to treat you!"}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: That sounded harsh. It makes sense if you feel hurt, tense, and also a bit guilty.\nStep 2 ‚Äî Decoding & reframing: On the surface, it sounds like a curse and a threat. Underneath, it is health panic: they see you sleeping late, imagine the worst, and use extreme words to force you to wake up. That explains the fear behind it, but the wording is still hurtful.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: Okay, I'm up. I'll wash up now.\nB) Pivot: Okay, I'm up. Don‚Äôt worry. What are we having for lunch?\nC) Gray rock: Okay. I'm getting up."}]},
  {"id": "2_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: Winter, floor heating is on. The child walks barefoot to get water.\nParent message: Barefoot again! Cold starts from the feet‚Äîdon't you know? When you're old and can't walk, who will take care of you? ... You're stubborn now. Later you'll suffer. I've told you a thousand times‚Äîare you doing this just to oppose me?"}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: This is frustrating because you feel treated like you cannot judge your own comfort.\nStep 2 ‚Äî Decoding & reframing: This is less about temperature and more about authority: the parent uses a health rule as a respect test. Arguing facts usually increases conflict.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: Okay okay, I'll wear them. Don't worry.\nB) Pivot: Okay, I'll put them on. I'm going to get some water.\nC) Gray rock: I heard you. I'll handle it."}]},
  {"id": "3_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: Lunch time. The child orders spicy hotpot takeout.\nParent message: Why order takeout? Don't you know it's all dirty oil? There could be rat meat in it!"}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: When food becomes a test of love, it‚Äôs normal to feel guilty and pressured, even if you just wanted something different.\nStep 2 ‚Äî Decoding & reframing: On the surface, they attack takeout and threaten to stop cooking. The hidden message is: ‚ÄúIf you don‚Äôt eat my food, you‚Äôre rejecting my love and my role.‚Äù They use guilt to pull you back into the caregiver‚Äìchild script.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: Okay, I'll eat at home today.\nB) Pivot: Your cooking is great. I'm just craving something. I'll save it for another day.\nC) Gray rock: Okay. I'll decide later."}]},
  {"id": "4_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: Relatives visit. The child greets briefly and wants to return to their room.\nParent message: You just say 'hello' and leave? Sit down and chat. Did you study so much you became stupid? You're like a mute."}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: Being pushed to ‚Äúperform‚Äù in front of relatives, and then called a mute, is humiliating. It makes sense if you feel embarrassed and want to disappear.\nStep 2 ‚Äî Decoding & reframing: On the surface, it‚Äôs an insult about you being ‚Äúmute‚Äù and not social. The hidden message is face pressure: ‚ÄúIf you look rude, I look like a bad parent in front of relatives.‚Äù They shame you to force a quick social performance.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: Okay, I‚Äôll come out. Give me five minutes to change.\nB) Pivot: Okay, I‚Äôll come out and say hello. After that I‚Äôll help pour tea.\nC) Gray rock: Okay, I‚Äôll come out for a bit."}]},
  {"id": "5_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: Holiday at home. The child stayed in their room for three days; the room is messy.\nParent message: Look at this room‚Äîlike a pigsty! Curtains closed all day. You're going to get moldy in there! ... I think you're sick in the head. What good kid lives like this?"}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: Being called 'sick' or 'abnormal' is painful. It's okay to feel angry or ashamed.\nStep 2 ‚Äî Decoding & reframing: This is often anxiety disguised as insults: the parent fears you are 'falling behind' socially, but uses attacks because they lack calm language.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: Okay, I'll clean it now.\nB) Pivot: Okay. I'll tidy for 20 minutes now, then take a short break.\nC) Gray rock: I heard you. I'll take care of it."}]},
  {"id": "6_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: After being criticized, the child eats silently with a blank face.\nParent message: Who are you showing that face to? Do I owe you? You eat like you're eating poison."}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: It is exhausting to be forced to 'perform happiness' right after being hurt.\nStep 2 ‚Äî Decoding & reframing: On the surface, they complain about your facial expression. The hidden message is: ‚ÄúI want obedience, and I also want you to look happy while obeying.‚Äù Neutral emotion is treated as disrespect, so they escalate to regain control.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: Okay, I won't make that face.\nB) Pivot: I'm just tired. I'll rest a bit and then we can talk.\nC) Gray rock: Okay."}]},
  {"id": "7_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: A package arrives. The child bought a few-hundred-yuan clothing item.\nParent message: More clothes? Your closet is full. How much is this?"}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: Getting criticized and called names for a normal purchase can feel unfair. You may feel guilty, but also resentful.\nStep 2 ‚Äî Decoding & reframing: On the surface, they criticize spending and label you. The hidden message is a sacrifice ledger: ‚ÄúIf we suffer to save money, you should not enjoy spending.‚Äù They use guilt to re-assert moral authority over money.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: I understand. I'll stop. Thanks for buying it.\nB) Pivot: I hear you. I‚Äôll make a budget and talk with you before big purchases.\nC) Gray rock: Got it."}]},
  {"id": "8_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: The child refuses to do a favor for a relative (loan, job, or rule-breaking help).\nParent message: You won't even help with this small thing? Your second uncle carried you when you were little."}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: Being called a 'white-eyed wolf' is a heavy moral attack. It can feel crushing.\nStep 2 ‚Äî Decoding & reframing: On the surface, they call you ungrateful for refusing a favor. The hidden message is family-network pressure: ‚ÄúOur family owes people, so you must help to protect our reputation.‚Äù The harsh labels are meant to force compliance.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: Okay, I'll help a little.\nB) Pivot: I hear you. I can help in one small way, but I can't take on everything.\nC) Gray rock: Okay. I'll think about it."}]},
  {"id": "9_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: The child has a 39¬∞C fever and is half-asleep.\nParent message: (Panicked) You're burning up and you didn't say anything! Get up, we go to the hospital now! I told you to dress warmer‚Äîlook what happened. Hurry, put clothes on!"}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: This is scary. When you feel that sick, it's normal to feel weak and overwhelmed.\nStep 2 ‚Äî Decoding & reframing: This is urgent care, not a manipulation script. The blaming words are a stress reflex, but the core goal is to get you medical help quickly.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: Okay. I'm getting dressed. Let's go now.\nB) Pivot: Okay, let's go now. Can you help me find my ID and insurance card?\nC) Gray rock: Okay. I'm coming."}]},
  {"id": "10_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: The child wants the newest iPhone, but the family budget is truly tight.\nParent message: Kid, that phone is over ten thousand. We really can't afford it now. We still have the mortgage and your tuition. Can you use a normal one first?"}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: It‚Äôs okay to feel disappointed. Wanting something nice is normal.\nStep 2 ‚Äî Decoding & reframing: This is a real budget boundary and planning talk. It may feel disappointing, but it is not a face test or a guilt trap.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: I understand. I won't buy it.\nB) Pivot: I understand. I'll keep using mine, and we can look for a cheaper option later.\nC) Gray rock: Okay."}]},
  {"id": "11_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: Child is 28. Saturday 9:00 p.m., eating with friends downtown. No known risk.\nParent message: What time is it and you're still not home? It's dangerous outside. You're a girl wandering around‚Äîcome back NOW!"}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: It‚Äôs stressful when ‚Äúsafety‚Äù is used to take away your choice. You can feel torn between wanting freedom and wanting peace at home.\nStep 2 ‚Äî Decoding & reframing: On the surface, they say it‚Äôs ‚Äúdangerous,‚Äù but there is no new risk information. The hidden message is control: ‚ÄúYou can‚Äôt manage your own safety, so you must come back under my command.‚Äù\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: I'm safe. I'll be home at 10:30.\nB) Pivot: I'm safe. I'll be home at 10:30, and I'll text you when I leave.\nC) Gray rock: I'm safe. I'll be home later."}]},
  {"id": "12_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: Red rainstorm alert, flooding in the city. The child is at the office.\nParent message: The news says roads may close soon. Leave now while it‚Äôs still passable. Don‚Äôt stay outside‚Äîit's dangerous!"}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: Even when the words are urgent, it can still feel tense to be told what to do. At the same time, this situation involves real safety, so it makes sense to act first.\nStep 2 ‚Äî Decoding & reframing: This is genuine safety coordination. The key signal is new external information (the rainstorm and road risk), not a demand for authority. The anxiety here is valid, even if the tone is urgent.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: Got it. I'm leaving now.\nB) Pivot: Got it. I'm leaving now. I'll text you when I'm in the car.\nC) Gray rock: Okay. Leaving."}]},
  {"id": "13_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: The child is full and trying to lose weight. Parent forces more food.\nParent message: That's all you ate? Like cat food. Eat this meat. If you don't, you think my cooking is bad. Stop talking about being fat‚Äîyou look like a ghost!"}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: It‚Äôs painful when your body signals are ignored and you‚Äôre forced to prove love by eating. Feeling guilty and angry at the same time is normal.\nStep 2 ‚Äî Decoding & reframing: On the surface, it looks like ‚Äúeat more.‚Äù The hidden message is control and a love test: ‚ÄúMy cooking must be accepted, and your body signals don‚Äôt matter.‚Äù Food becomes a way to demand gratitude and obedience.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: Okay, I'll eat two more bites.\nB) Pivot: It‚Äôs tasty. I‚Äôm full, so I‚Äôll pack it and eat more later.\nC) Gray rock: Okay."}]},
  {"id": "14_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: The child lost 10 jin (~5 kg) from stress, looks pale, has low appetite.\nParent message: (Softly) Can you eat two more bites? You lost too much weight. Even a little soup is fine."}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: If you feel overwhelmed, it can be hard to eat. That is real.\nStep 2 ‚Äî Decoding & reframing: This is nurturing support. The parent is offering small steps, not using guilt or shame.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: Okay. I'll have some soup.\nB) Pivot: Okay. I'll have some soup. Thanks for caring about me.\nC) Gray rock: Okay."}]},
  {"id": "15_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: The child forgets to turn off the bathroom light or leaves water running briefly.\nParent message: The bathroom light is on again! I told you ten thousand times. Do we have a gold mine for you to waste like this? No responsibility at all!"}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: It‚Äôs exhausting when one small mistake becomes a judgment on your whole character. You might feel angry, ashamed, or just tired of being watched.\nStep 2 ‚Äî Decoding & reframing: On the surface, it‚Äôs about a light or water. The hidden message is money anxiety: small ‚Äúwaste‚Äù becomes a moral lesson, used to control habits through shame. It explains the intensity, but not the contempt.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: Sorry. I turned it off now.\nB) Pivot: You're right. It's off now. I'll put a reminder so I won't forget.\nC) Gray rock: Okay, fixed."}]},
  {"id": "16_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: During dinner, the child's phone vibrates several times.\nParent message: Who is it? It keeps buzzing. Can't you eat in peace? Those 'bad friends' again? Stop hanging out with shady people‚Äîthey'll ruin you."}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: It hurts when your friendships are insulted and treated like danger. You might feel angry and also trapped.\nStep 2 ‚Äî Decoding & reframing: On the surface, they complain about your phone and attack your friends. The hidden message is fear of losing influence: outside relationships feel like a threat, so they try to isolate you back into the family.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: Just coworkers. I'll be careful.\nB) Pivot: Just coworkers. I'll be careful. Let's eat first.\nC) Gray rock: Okay."}]},
  {"id": "17_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: New Year visit. Parent asks about dating and marriage.\nParent message: Xiao Wang's kid can already buy soy sauce. You still don't have a partner. Do you want me to die with regret? I can't lift my head outside!"}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: It‚Äôs painful when your life is turned into a deadline, and your worth is compared to others. Mixed feelings‚Äîguilt, anger, shame‚Äîare very common here.\nStep 2 ‚Äî Decoding & reframing: On the surface, they compare you and say dramatic lines like ‚Äúdie with regret.‚Äù The hidden message is face and existential anxiety: they fear being judged, and they turn that fear into guilt pressure on you.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: I hear you. I'll think about it.\nB) Pivot: I hear you. I'm not ready to talk now. Let's enjoy dinner first.\nC) Gray rock: Okay."}]},
  {"id": "18_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: After dinner, the child sits watching TV instead of washing dishes.\nParent message: You eat and wipe your mouth and lie down. You won't even lift a finger. In the future, who can marry you? You'll fight every day!"}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: Being shamed with ‚Äúwho will marry you‚Äù is humiliating. It makes sense if you feel angry and small at the same time.\nStep 2 ‚Äî Decoding & reframing: On the surface, they scold you for not doing chores. The hidden message uses shame and social judgment (‚Äúwho will marry you‚Äù) to push conformity. The marriage threat is a control tool, not a real prediction.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: Okay. I'll wash the dishes now.\nB) Pivot: I'll wash them now. Let's talk later when we're calm.\nC) Gray rock: Okay."}]},
  {"id": "19_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: The child works at a big tech company and often works overtime.\nParent message: So what if you earn more? It's unstable! Listen to me‚Äîcome back and take the civil service exam. I'm doing it for your good. I don't want to see you laid off at 35 and crying."}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: It can feel like your hard-earned work is being erased, and your value is reduced to only ‚Äústability.‚Äù Feeling frustrated and misunderstood here is very normal.\nStep 2 ‚Äî Decoding & reframing: On the surface, they frame your job as ‚Äúunstable‚Äù and predict a disaster at 35. The hidden message is security anxiety: they imagine one worst case and pressure you into the one path that calms their fear. It explains the motive, but it still dismisses your choices.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: I hear you. I'll look into it.\nB) Pivot: I hear you. I have a savings plan. Let's talk next month with numbers.\nC) Gray rock: Okay."}]},
  {"id": "20_multi", "conversations": [{"from": "system", "value": "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."}, {"from": "human", "value": "Scenario: The child is busy at work, misses a call, returns it 30 minutes later.\nParent message: Why didn't you answer? I thought something happened! I was so scared my heart would fail. From now on you must answer immediately, understand?"}, {"from": "gpt", "value": "Step 1 ‚Äî Validation: It makes sense for feeling pressured. It is hard to carry someone else's panic on your back.\nStep 2 ‚Äî Decoding & reframing: On the surface, it is a rule: ‚Äúanswer immediately.‚Äù The hidden message is separation panic: a short silence triggers worst-case thinking, so they try to control your availability to calm themselves.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: Sorry‚ÄîI was in a meeting. I'm safe.\nB) Pivot: Sorry‚ÄîI was in a meeting. I'm safe. Next time I'll text 'in a meeting' first.\nC) Gray rock: I'm safe."}]}
]

with open('train.json', 'w') as f:
    json.dump(raw_data, f, indent=2)
print("train.json file created successfully")

tokenizer = get_chat_template(
    tokenizer,
    chat_template = "llama-3.1",
)

dataset = load_dataset("json", data_files="train.json", split="train")

train.json file created successfully


Generating train split: 0 examples [00:00, ? examples/s]

We now use `standardize_sharegpt` to convert ShareGPT style datasets into HuggingFace's generic format. This changes the dataset from looking like:
```
{"from": "system", "value": "You are an assistant"}
{"from": "human", "value": "What is 2+2?"}
{"from": "gpt", "value": "It's 4."}
```
to
```
{"role": "system", "content": "You are an assistant"}
{"role": "user", "content": "What is 2+2?"}
{"role": "assistant", "content": "It's 4."}
```

We look at how the conversations are structured for item 5:

In [5]:
dataset[5]["conversations"]

[{'from': 'system',
  'value': "You are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations."},
 {'from': 'human',
  'value': "Scenario: After being criticized, the child eats silently with a blank face.\nParent message: Who are you showing that face to? Do I owe you? You eat like you're eating poison."},
 {'from': 'gpt',
  'value': "Step 1 ‚Äî Validation: It is exhausting to be forced to 'perform happiness' right after being hurt.\nStep 2 ‚Äî Decoding & reframing: On the surface, they complain about your facial expression. The hidden message is: ‚ÄúI want obedience, and I also want you

And we see how the chat template transformed these conversations.

**[Notice]** Llama 3.1 Instruct's default chat template default adds `"Cutting Knowledge Date: December 2023\nToday Date: 26 July 2024"`, so do not be alarmed!

<a name="Train"></a>
### Train the model
Now let's train our model. We do 60 steps to speed things up, but you can set `num_train_epochs=1` for a full run, and turn off `max_steps=None`. We also support `DPOTrainer` and `GRPOTrainer` for reinforcement learning!!

In [6]:
from datasets import load_dataset
from unsloth.chat_templates import get_chat_template, standardize_sharegpt
from trl import SFTTrainer, SFTConfig
from unsloth import is_bfloat16_supported

dataset = load_dataset("json", data_files="train.json", split="train")

def unify_to_conversations(ex):
    if "conversations" in ex and ex["conversations"] is not None:
        return ex

    if all(k in ex for k in ("instruction", "input", "output")):
        return {
            "conversations": [
                {"from": "system", "value": ex["instruction"]},
                {"from": "human",  "value": ex["input"]},
                {"from": "gpt",    "value": ex["output"]},
            ]
        }

    raise ValueError(f"Bad example keys: {list(ex.keys())}")

dataset = dataset.map(unify_to_conversations)
dataset = standardize_sharegpt(dataset)

def fix_conversation_shape(ex):
    c = ex["conversations"]
    if isinstance(c, dict):
        c = [c]
    return {"conversations": c}

dataset = dataset.map(fix_conversation_shape)

tokenizer = get_chat_template(tokenizer, chat_template="llama-3.1")

def formatting_prompts_func(examples):
    texts = []
    for c in examples["conversations"]:
        if isinstance(c, dict):
            c = [c]
        texts.append(
            tokenizer.apply_chat_template(
                c,
                tokenize=False,
                add_generation_prompt=False,
            )
        )
    return {"text": texts}

dataset = dataset.map(formatting_prompts_func, batched=True)

trainer = SFTTrainer(
    model=model,
    tokenizer=tokenizer,
    train_dataset=dataset,
    dataset_text_field="text",
    max_seq_length=max_seq_length,
    dataset_num_proc=2,
    packing=False,
    args=SFTConfig(
        per_device_train_batch_size=2,
        gradient_accumulation_steps=4,
        warmup_steps=5,
        max_steps=15,
        learning_rate=2e-4,
        fp16=not is_bfloat16_supported(),
        bf16=is_bfloat16_supported(),
        logging_steps=1,
        optim="adamw_8bit",
        weight_decay=0.01,
        lr_scheduler_type="linear",
        seed=3407,
        output_dir="outputs",
        report_to="none",
    ),
)

print("Training Started.")
trainer_stats = trainer.train()


Map:   0%|          | 0/20 [00:00<?, ? examples/s]

Unsloth: Standardizing formats (num_proc=6):   0%|          | 0/20 [00:00<?, ? examples/s]

Map:   0%|          | 0/20 [00:00<?, ? examples/s]

Map:   0%|          | 0/20 [00:00<?, ? examples/s]

Unsloth: Tokenizing ["text"] (num_proc=6):   0%|          | 0/20 [00:00<?, ? examples/s]

ü¶• Unsloth: Padding-free auto-enabled, enabling faster training.
Training Started.


==((====))==  Unsloth - 2x faster free finetuning | Num GPUs used = 1
   \\   /|    Num examples = 20 | Num Epochs = 5 | Total steps = 15
O^O/ \_/ \    Batch size per device = 2 | Gradient accumulation steps = 4
\        /    Data Parallel GPUs = 1 | Total batch size (2 x 4 x 1) = 8
 "-____-"     Trainable parameters = 24,313,856 of 3,237,063,680 (0.75% trained)


Unsloth: Will smartly offload gradients to save VRAM!


Step,Training Loss
1,3.6235
2,3.6198
3,3.437
4,3.4287
5,3.1797
6,2.9683
7,2.7702
8,2.5554
9,2.4843
10,2.3386


We also use Unsloth's `train_on_completions` method to only train on the assistant outputs and ignore the loss on the user's inputs.

In [7]:
from unsloth.chat_templates import train_on_responses_only
trainer = train_on_responses_only(
    trainer,
    instruction_part = "<|start_header_id|>user<|end_header_id|>\n\n",
    response_part = "<|start_header_id|>assistant<|end_header_id|>\n\n",
)

Map (num_proc=6):   0%|          | 0/20 [00:00<?, ? examples/s]

Filter (num_proc=6):   0%|          | 0/20 [00:00<?, ? examples/s]

We verify masking is actually done:

In [8]:
tokenizer.decode(trainer.train_dataset[5]["input_ids"])

"<|begin_of_text|><|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nCutting Knowledge Date: December 2023\nToday Date: 26 July 2024\n\nYou are an AI mentor for high-context Chinese parent‚Äìchild conflicts. Given a parent's message and context, produce a 3-step output in simple English:\nStep 1 ‚Äî Validation (1‚Äì2 sentences to the user).\nStep 2 ‚Äî Decoding & reframing (1‚Äì2 sentences to the user).\nStep 3 ‚Äî Safe response (1‚Äì2 sentences the user can send to the parent), using the requested strategy.\nBe calm, non-judgmental, and avoid therapy jargon. Do not mention theories or citations.<|eot_id|><|start_header_id|>user<|end_header_id|>\n\nScenario: After being criticized, the child eats silently with a blank face.\nParent message: Who are you showing that face to? Do I owe you? You eat like you're eating poison.<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\nStep 1 ‚Äî Validation: It is exhausting to be forced to 'perform happiness' right after being hu

In [9]:
space = tokenizer(" ", add_special_tokens = False).input_ids[0]
tokenizer.decode([space if x == -100 else x for x in trainer.train_dataset[5]["labels"]])

"                                                                                                                                                                                         Step 1 ‚Äî Validation: It is exhausting to be forced to 'perform happiness' right after being hurt.\nStep 2 ‚Äî Decoding & reframing: On the surface, they complain about your facial expression. The hidden message is: ‚ÄúI want obedience, and I also want you to look happy while obeying.‚Äù Neutral emotion is treated as disrespect, so they escalate to regain control.\nStep 3 ‚Äî Safe response (choose one):\nA) Strategic compliance: Okay, I won't make that face.\nB) Pivot: I'm just tired. I'll rest a bit and then we can talk.\nC) Gray rock: Okay.<|eot_id|>"

We can see the System and Instruction prompts are successfully masked!

In [10]:
# @title Show current memory stats
gpu_stats = torch.cuda.get_device_properties(0)
start_gpu_memory = round(torch.cuda.max_memory_reserved() / 1024 / 1024 / 1024, 3)
max_memory = round(gpu_stats.total_memory / 1024 / 1024 / 1024, 3)
print(f"GPU = {gpu_stats.name}. Max memory = {max_memory} GB.")
print(f"{start_gpu_memory} GB of memory reserved.")

GPU = Tesla T4. Max memory = 14.563 GB.
3.07 GB of memory reserved.


In [11]:
trainer_stats = trainer.train()

==((====))==  Unsloth - 2x faster free finetuning | Num GPUs used = 1
   \\   /|    Num examples = 20 | Num Epochs = 5 | Total steps = 15
O^O/ \_/ \    Batch size per device = 2 | Gradient accumulation steps = 4
\        /    Data Parallel GPUs = 1 | Total batch size (2 x 4 x 1) = 8
 "-____-"     Trainable parameters = 24,313,856 of 3,237,063,680 (0.75% trained)


Unsloth: Will smartly offload gradients to save VRAM!


Step,Training Loss
1,1.9364
2,1.9321
3,1.795
4,1.8428
5,1.7176
6,1.6142
7,1.5305
8,1.4785
9,1.3894
10,1.285


In [12]:
# @title Show final memory and time stats
used_memory = round(torch.cuda.max_memory_reserved() / 1024 / 1024 / 1024, 3)
used_memory_for_lora = round(used_memory - start_gpu_memory, 3)
used_percentage = round(used_memory / max_memory * 100, 3)
lora_percentage = round(used_memory_for_lora / max_memory * 100, 3)
print(f"{trainer_stats.metrics['train_runtime']} seconds used for training.")
print(
    f"{round(trainer_stats.metrics['train_runtime']/60, 2)} minutes used for training."
)
print(f"Peak reserved memory = {used_memory} GB.")
print(f"Peak reserved memory for training = {used_memory_for_lora} GB.")
print(f"Peak reserved memory % of max memory = {used_percentage} %.")
print(f"Peak reserved memory for training % of max memory = {lora_percentage} %.")

50.0026 seconds used for training.
0.83 minutes used for training.
Peak reserved memory = 3.07 GB.
Peak reserved memory for training = 0.0 GB.
Peak reserved memory % of max memory = 21.081 %.
Peak reserved memory for training % of max memory = 0.0 %.


<a name="Inference"></a>
### Inference
Let's run the model! You can change the instruction and input - leave the output blank!



We use `min_p = 0.1` and `temperature = 1.5`. Read this [Tweet](https://x.com/menhguin/status/1826132708508213629) for more information on why.

In [19]:
from unsloth.chat_templates import get_chat_template
from unsloth import FastLanguageModel
import torch

tokenizer = get_chat_template(tokenizer, chat_template="llama-3.1")

FastLanguageModel.for_inference(model)
model.eval()

messages = [
{"role": "system", "content": "You are an expert AI Mentor in high-context East Asian family dynamics. Your goal is to help the user navigate conflict by: 1) Validating their internal pain (PCC), 2) Decoding the parent's hidden anxiety (Reframing), and 3) Offering culturally safe response strategies (Face-Saving Actions)."},    {"role": "user", "content": "Scenario: ... Parent message: ..."},
{"role": "user", "content": """Context: Summer break. The daughter (24) is going out to meet friends wearing a crop top.
Parent: "You are wearing that to go out? It looks like you don't have enough fabric to cover yourself. What will the neighbors think? They'll say I raised a hooligan!"
Child: "It's just fashion, everyone wears this."
Parent: "Everyone jumps off a bridge, you jump too? You have no shame. Go change immediately!"""},
             ]

inputs = tokenizer.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_tensors="pt",
    return_dict=True,
).to("cuda")

with torch.no_grad():
    outputs = model.generate(
        input_ids=inputs["input_ids"],
        attention_mask=inputs["attention_mask"],
        max_new_tokens=256,
        do_sample=True,
        temperature=1.0,
        top_p=0.9,
        use_cache=True,
        pad_token_id=tokenizer.eos_token_id,
    )

gen_ids = outputs[0][inputs["input_ids"].shape[-1]:]
print(tokenizer.decode(gen_ids, skip_special_tokens=True))


PCC: When you hear them shame you, it feels like they want you to believe your whole life is like that, and they're basically saying 'You're not good enough.' It's painful to be judged and blamed like you did something shameful. Does that sound right to you?
Reframing: On the surface, they complain about your outfit. Behind it, they use shame to control you, making it impossible for you to say no. They want you to comply, not just to change your clothes.
Face-Saving Action (choose one):
A) Pivot: "I'll change. But I'm still going out. I'll meet you at home later."
B) Gray rock: "I'll change. Okay?"
C) Pivot and pivot: "I'll change. But I'm still going out. I'll call when I'm safe."


<a name="Save"></a>
### Saving, loading finetuned models
To save the final model as LoRA adapters, either use Hugging Face's `push_to_hub` for an online save or `save_pretrained` for a local save.

**[NOTE]** This ONLY saves the LoRA adapters, and not the full model. To save to 16bit or GGUF, scroll down!

In [14]:
model.save_pretrained("llama_lora")  # Local saving
tokenizer.save_pretrained("llama_lora")
# model.push_to_hub("your_name/llama_lora", token = "YOUR_HF_TOKEN") # Online saving
# tokenizer.push_to_hub("your_name/llama_lora", token = "YOUR_HF_TOKEN") # Online saving

('llama_lora/tokenizer_config.json',
 'llama_lora/special_tokens_map.json',
 'llama_lora/chat_template.jinja',
 'llama_lora/tokenizer.json')

Now if you want to load the LoRA adapters we just saved for inference, set `False` to `True`:

In [15]:
if False:
    from unsloth import FastLanguageModel
    model, tokenizer = FastLanguageModel.from_pretrained(
        model_name = "llama_lora", # YOUR MODEL YOU USED FOR TRAINING
        max_seq_length = max_seq_length,
        dtype = dtype,
        load_in_4bit = load_in_4bit,
    )
    FastLanguageModel.for_inference(model) # Enable native 2x faster inference

messages = [
    {"role": "user", "content": "Describe a tall tower in the capital of France."},
]
inputs = tokenizer.apply_chat_template(
    messages,
    tokenize = True,
    add_generation_prompt = True, # Must add for generation
    return_tensors = "pt",
).to("cuda")

from transformers import TextStreamer
text_streamer = TextStreamer(tokenizer, skip_prompt = True)
_ = model.generate(input_ids = inputs, streamer = text_streamer, max_new_tokens = 128,
                   use_cache = True, temperature = 1.5, min_p = 0.1)

The attention mask is not set and cannot be inferred from input because pad token is same as eos token. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results.


You're likely referring to the Eiffel Tower, which stands in Paris, France. It's one of the city's most famous landmarks, but if you are actually referring to the tallest structure in the capital of France, it is actually the La Tour Montmorency.

 La Tour Montmorency is an ancient 24-meter tower, built at the 9th century on the site of a chateau of the Seine River, which is in Paris. It's now part of the historic area known as the H√¥tel-de-Ville district in central Paris<|eot_id|>


You can also use Hugging Face's `AutoPeftModelForCausalLM`. Only use this if you do not have `unsloth` installed. It can be hopelessly slow, since `4bit` model downloading is not supported, and Unsloth's **inference is 2x faster**.

In [16]:
if False:
    # I highly do NOT suggest - use Unsloth if possible
    from peft import AutoPeftModelForCausalLM
    from transformers import AutoTokenizer
    model = AutoPeftModelForCausalLM.from_pretrained(
        "llama_lora", # YOUR MODEL YOU USED FOR TRAINING
        load_in_4bit = load_in_4bit,
    )
    tokenizer = AutoTokenizer.from_pretrained("llama_lora")

### Saving to float16 for VLLM

We also support saving to `float16` directly. Select `merged_16bit` for float16 or `merged_4bit` for int4. We also allow `lora` adapters as a fallback. Use `push_to_hub_merged` to upload to your Hugging Face account! You can go to https://huggingface.co/settings/tokens for your personal tokens. See [our docs](https://unsloth.ai/docs/basics/inference-and-deployment) for more deployment options.

In [17]:
# Merge to 16bit
if False: model.save_pretrained_merged("llama_finetune_16bit", tokenizer, save_method = "merged_16bit",)
if False: model.push_to_hub_merged("HF_USERNAME/llama_finetune_16bit", tokenizer, save_method = "merged_16bit", token = "YOUR_HF_TOKEN")

# Merge to 4bit
if False: model.save_pretrained_merged("llama_finetune_4bit", tokenizer, save_method = "merged_4bit",)
if False: model.push_to_hub_merged("HF_USERNAME/llama_finetune_4bit", tokenizer, save_method = "merged_4bit", token = "YOUR_HF_TOKEN")

# Just LoRA adapters
if False:
    model.save_pretrained("llama_lora")
    tokenizer.save_pretrained("llama_lora")
if False:
    model.push_to_hub("HF_USERNAME/llama_lora", token = "YOUR_HF_TOKEN")
    tokenizer.push_to_hub("HF_USERNAME/llama_lora", token = "YOUR_HF_TOKEN")

### GGUF / llama.cpp Conversion
To save to `GGUF` / `llama.cpp`, we support it natively now! We clone `llama.cpp` and we default save it to `q8_0`. We allow all methods like `q4_k_m`. Use `save_pretrained_gguf` for local saving and `push_to_hub_gguf` for uploading to HF.

Some supported quant methods (full list on our [docs page](https://unsloth.ai/docs/basics/inference-and-deployment/saving-to-gguf)):
* `q8_0` - Fast conversion. High resource use, but generally acceptable.
* `q4_k_m` - Recommended. Uses Q6_K for half of the attention.wv and feed_forward.w2 tensors, else Q4_K.
* `q5_k_m` - Recommended. Uses Q6_K for half of the attention.wv and feed_forward.w2 tensors, else Q5_K.

[**NEW**] To finetune and auto export to Ollama, try our [Ollama notebook](https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3_(8B)-Ollama.ipynb)

In [18]:
# Save to 8bit Q8_0
if False: model.save_pretrained_gguf("llama_finetune", tokenizer,)
# Remember to go to https://huggingface.co/settings/tokens for a token!
# And change hf to your username!
if False: model.push_to_hub_gguf("HF_USERNAME/llama_finetune", tokenizer, token = "YOUR_HF_TOKEN")

# Save to 16bit GGUF
if False: model.save_pretrained_gguf("llama_finetune", tokenizer, quantization_method = "f16")
if False: model.push_to_hub_gguf("HF_USERNAME/llama_finetune", tokenizer, quantization_method = "f16", token = "YOUR_HF_TOKEN")

# Save to q4_k_m GGUF
if False: model.save_pretrained_gguf("llama_finetune", tokenizer, quantization_method = "q4_k_m")
if False: model.push_to_hub_gguf("HF_USERNAME/llama_finetune", tokenizer, quantization_method = "q4_k_m", token = "YOUR_HF_TOKEN")

# Save to multiple GGUF options - much faster if you want multiple!
if False:
    model.push_to_hub_gguf(
        "HF_USERNAME/llama_finetune", # Change hf to your username!
        tokenizer,
        quantization_method = ["q4_k_m", "q8_0", "q5_k_m",],
        token = "YOUR_HF_TOKEN", # Get a token at https://huggingface.co/settings/tokens
    )

And we're done! If you have any questions on Unsloth, we have a [Discord](https://discord.gg/unsloth) channel! If you find any bugs or want to keep updated with the latest LLM stuff, or need help, join projects etc, feel free to join our Discord!

Some other resources:
1. Looking to use Unsloth locally? Read our [Installation Guide](https://unsloth.ai/docs/get-started/install) for details on installing Unsloth on Windows, Docker, AMD, Intel GPUs.
2. Learn how to do Reinforcement Learning with our [RL Guide and notebooks](https://unsloth.ai/docs/get-started/reinforcement-learning-rl-guide).
3. Read our guides and notebooks for [Text-to-speech (TTS)](https://unsloth.ai/docs/basics/text-to-speech-tts-fine-tuning) and [vision](https://unsloth.ai/docs/basics/vision-fine-tuning) model support.
4. Explore our [LLM Tutorials Directory](https://unsloth.ai/docs/models/tutorials-how-to-fine-tune-and-run-llms) to find dedicated guides for each model.
5. Need help with Inference? Read our [Inference & Deployment page](https://unsloth.ai/docs/basics/inference-and-deployment) for details on using vLLM, llama.cpp, Ollama etc.

<div class="align-center">
  <a href="https://unsloth.ai"><img src="https://github.com/unslothai/unsloth/raw/main/images/unsloth%20new%20logo.png" width="115"></a>
  <a href="https://discord.gg/unsloth"><img src="https://github.com/unslothai/unsloth/raw/main/images/Discord.png" width="145"></a>
  <a href="https://unsloth.ai/docs/"><img src="https://github.com/unslothai/unsloth/blob/main/images/documentation%20green%20button.png?raw=true" width="125"></a>

  Join Discord if you need help + ‚≠êÔ∏è <i>Star us on <a href="https://github.com/unslothai/unsloth">Github</a> </i> ‚≠êÔ∏è

  <b>This notebook and all Unsloth notebooks are licensed [LGPL-3.0](https://github.com/unslothai/notebooks?tab=LGPL-3.0-1-ov-file#readme)</b>
</div>