Distillation Guides
- Distilling the Knowledge in a Neural Network — Hinton, Vinyals, Dean, the OG paper (2015)
- Knowledge Distillation: A Survey — Gou et al., the canonical knowledge distillation survey (2020)
- A Comprehensive Survey on Knowledge Distillation — diffusion, transformer, LLM distillation (2025)
- OpenAI Distillation Guide — built-in distillation pipeline
- OpenAI Cookbook Walkthrough — step-by-step distillation to fine-tune
- Snorkel AI Complete Guide — LLM distillation demystified
- HuggingFace Knowledge Distillation — everything you need to know about distillation
- DataCamp Tutorial — practical distillation guide with examples
- NVIDIA NeMo Pruning + Distillation — prune Llama 8B → 4B then distill
Open Source Toolkits
- DataClaw — share Claude Code and Codex conversations as HuggingFace datasets
- DistillKit — production-ready, online/offline distillation
- EasyDistill — black-box and white-box methods, data synthesis
- HuggingFace TRL GKD Trainer — generalized knowledge distillation
Fine-tuning Claude
- AWS Bedrock Fine-tuning Claude 3 Haiku — fine-tune Claude on Bedrock
- Survey on Knowledge Distillation of LLMs — comprehensive academic survey
AI Infrastructure
- Clanker Cloud — AI-powered DevOps agent for production ops
- Prime Intellect — train, evaluate, and deploy your own agentic models

