# Model Evaluation Demo

This notebook demonstrates how to programmatically evaluate the model using `evaluate.py`.

In [1]:
import pandas as pd
from evaluate import run_evaluation, load_config

  from .autonotebook import tqdm as notebook_tqdm


In [6]:
# Load configuration
config_path = "debug_config.yaml"
config = load_config(config_path)

# You can modify config programmatically here if needed
print("Model:", config['model']['name_or_path'])
print("Dataset:", config['data']['dataset_name'])

Model: Qwen/Qwen2-0.5B-Instruct
Dataset: openai/gsm8k


In [7]:
# Run evaluation on a small subset (e.g., 5 samples) to test
metrics, samples = run_evaluation(config, max_samples=5)

print("Metrics:", metrics)

Note: Config specifies 'train' split. Using 'test' split for evaluation by default for GSM8K.
Loading dataset: openai/gsm8k


Loading model: Qwen/Qwen2-0.5B-Instruct
Starting evaluation...


100%|██████████| 5/5 [00:36<00:00,  7.21s/it]

Final Results: Accuracy: 0.4000, Format: 0.2000
Metrics: {'accuracy': 0.4, 'format_compliance': 0.2, 'total_samples': 5}





In [8]:
# Display samples in a DataFrame
df = pd.DataFrame(samples)
pd.set_option('display.max_colwidth', None)
display(df[['completion', 'ground_truth', 'accuracy_score', 'format_score']])

Unnamed: 0,completion,ground_truth,accuracy_score,format_score
0,"To calculate how much Janet makes from selling the ducks at the farmers' market each day, we need to follow these steps:\n\n1. Calculate the number of eggs laid by Janet's ducks per day.\n2. Subtract the number of eggs Janet eats (three for breakfast) and the number of eggs she bakes (four) from the total number of eggs laid.\n3. Divide the remaining eggs by $2 to find out how much money she makes from selling them.\n\nLet's do the calculations:\n\n**Step 1: Total Eggs Laid**\n\nJanet's ducks lay 16 eggs per day.\n\n**Step 2: Eggs Eaten and Baked**\n\nEggs eaten: 3 + 4 = 7 eggs\n\nTotal eggs laid: 16 - 7 = 9 eggs\n\n**Step 3: Money Made**\n\nMoney made = Remaining eggs / $2\n\n**Step 4: Calculation**\n\nRemaining eggs = 9 - 7 = 2 eggs\n\nMoney made = 2 / $2 = $1\n\nSo, Janet makes $1 from selling the ducks at the farmers' market each day.",Janet sells 16 - 3 - 4 = <<16-3-4=9>>9 duck eggs a day.\nShe makes 9 * 2 = $<<9*2=18>>18 every day at the farmer’s market.\n#### 18,0.0,0.0
1,<think> reasoning process here</thinking> <answer>It takes 4 bolts in total to make the robe.</answer>,It takes 2/2=<<2/2=1>>1 bolt of white fiber\nSo the total amount of fabric is 2+1=<<2+1=3>>3 bolts of fabric\n#### 3,1.0,0.0
2,"<think> Josh bought the house for $80,000 and put in $50,000 in repairs, so his total investment was $80,000 + $50,000 = $130,000.</think> <answer>He made a profit of $130,000 - $100,000 = $30,000.</answer>","The cost of the house and repairs came out to 80,000+50,000=$<<80000+50000=130000>>130,000\nHe increased the value of the house by 80,000*1.5=<<80000*1.5=120000>>120,000\nSo the new value of the house is 120,000+80,000=$<<120000+80000=200000>>200,000\nSo he made a profit of 200,000-130,000=$<<200000-130000=70000>>70,000\n#### 70000",1.0,1.0
3,"To find out how many meters James runs per week, we need to calculate the total distance he covers in one week by running three sprints of 3 minutes each.\n\n1. **Sprint Distance**: Since James runs 60 meters for each sprint, and there are 3 sprints, his sprint distance is:\n - \( 60 \, \text{meters/sprint} \times 3 \, \text{sprints} = 180 \, \text{meters} \)\n\n2. **Total Sprinting Distance**: Since he runs 180 meters per sprint, and he does this 3 times a week (for a total of 9 sessions), his total sprinting distance for that week is:\n - \( 180 \, \text{meters/sprint} \times 9 \, \text{sessions} = 1620 \, \text{meters} \)\n\nTherefore, James runs a total of 1620 meters per week.",He sprints 3*3=<<3*3=9>>9 times\nSo he runs 9*60=<<9*60=540>>540 meters\n#### 540,0.0,0.0
4,"<think> Reasoning process:</<think> \n\nWendi gives her chickens three cups of mixed chicken feed per meal. If she has 20 chickens and she gives each chicken one meal of feed (since each chicken needs its own meal), that means she gives out 20 cups of feed. \n\nIn the morning, she gives her flock 15 cups of feed, and in the afternoon, she gives them an additional 25 cups. This totals to a total of 30 cups of feed for all her chickens.\n\nTo find out how much feed she needs for the final meal, we subtract the amount given out from the total feed available: \n\n20 cups - 30 cups = -10 cups\n\nThis means she has less than enough feed for the final meal of the day. However, since you can't have negative cups of feed, this means she actually gave out more than enough feed for the final meal of the day. \n\nIf you want to know exactly how much feed she needs for the final meal, you would add back the extra feed she gave out: \n\n30 cups + 25 cups = 55 cups\n\nSo, Wendi needs to give her chickens 55 cups of feed for the final meal of the day.","If each chicken eats 3 cups of feed per day, then for 20 chickens they would need 3*20=<<3*20=60>>60 cups of feed per day.\nIf she feeds the flock 15 cups of feed in the morning, and 25 cups in the afternoon, then the final meal would require 60-15-25=<<60-15-25=20>>20 cups of chicken feed.\n#### 20",0.0,0.0


In [9]:
# Deep dive into a single sample
sample_idx = 0
print("Prompt:\n", df.iloc[sample_idx]['prompt'])
print("\nModel Completion:\n", df.iloc[sample_idx]['completion'])
print("\nGround Truth:\n", df.iloc[sample_idx]['ground_truth'])

Prompt:
 <|im_start|>system
A conversation between User and Assistant. The user asks a question, and the Assistant solves it. The assistant first thinks about the reasoning process in the mind and then provides the user with the answer. The reasoning process and answer are enclosed within <think> </think> and <answer> </answer> tags, respectively, i.e., <think> reasoning process here </think><answer> answer here </answer>
<|im_end|>
<|im_start|>user
Janet’s ducks lay 16 eggs per day. She eats three for breakfast every morning and bakes muffins for her friends every day with four. She sells the remainder at the farmers' market daily for $2 per fresh duck egg. How much in dollars does she make every day at the farmers' market?<|im_end|>
<|im_start|>assistant


Model Completion:
 To calculate how much Janet makes from selling the ducks at the farmers' market each day, we need to follow these steps:

1. Calculate the number of eggs laid by Janet's ducks per day.
2. Subtract the number of e