Use these after activating the decoder-only-runner environment. For Decoder-Chinese-SLM
checkpoints, keep MODEL_KIND=hf; loading is local/offline by default.
Dual SCENIC further training only:
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
NPROC_PER_NODE=8 \
MODEL_KIND=hf \
bash run_scenic_further_training_from_base.sh /PATH/TO/MY/BASE_DECODER_SLMFull SCENIC training plus 20-row pruning matrix:
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
NPROC_PER_NODE=8 \
SPARSITY_GPU_IDS=0,1,2,3,4,5,6,7 \
MODEL_KIND=hf \
bash run_linear_sparsity_revision_from_base.sh /PATH/TO/MY/DECODER_SLM_CHECKPOINTIf a run fails, first check the preflight and stage logs:
decoder-diagnose /PATH/TO/MY/DECODER_SLM_CHECKPOINT --model-kind hf
tail -n 120 outputs/scenic_further_training/logs/checkpoint_preflight.log
tail -n 120 outputs/scenic_further_training/logs/contrastive_sft.log
tail -n 120 outputs/scenic_further_training/logs/regular_sft.log
tail -n 120 outputs/decoder_pruning_full_matrix/logs/dense_regular_sft.logThat checkpoint path must be the final/full model folder, not only the tokenizer folder. For
Decoder-Chinese-SLM, it should contain config.json, tokenizer.json,
tokenizer_config.json, special_tokens_map.json, and model.safetensors or
pytorch_model.bin in the same directory. If CPU preflight loading is too slow, set
CHECKPOINT_LOAD_MODEL=0 after the folder layout has been confirmed.
Small public-safe repository for continuing training and sampling from a trained decoder-only language model.
It supports two common checkpoint layouts:
- Hugging Face causal language model directories with
config.json, tokenizer files, and model weights. - Simple custom PyTorch decoder-only checkpoints with a compatible
config.jsonandmodel.pt/checkpoint.pt.
Large checkpoint, dataset, and output files are ignored by git so this repo can stay public without accidentally publishing your trained SLM or private data.
For local CPU/MPS testing:
conda env create -f environment.yml
conda activate decoder-only-runner
pip install -e .If you prefer plain pip:
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
pip install -e .For a Linux CUDA server such as an 8x NVIDIA H20 box, install CUDA PyTorch instead:
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements-cu124.txt
pip install -e .Place the trained model under checkpoints/, for example:
checkpoints/my-model/
config.json
tokenizer.json
model.safetensors
For a custom PyTorch checkpoint, use:
checkpoints/my-model/
config.json
model.pt
vocab.json
The custom config.json should include at least:
{
"vocab_size": 256,
"block_size": 256,
"n_layer": 6,
"n_head": 6,
"n_embd": 384
}If vocab_size is 256 and no tokenizer is provided, the runner uses a UTF-8 byte tokenizer.
The repo includes the final SCENIC artifacts needed for the decoder SLM runs:
data/scenic/SCENIC_full_training_dataset.json
data/scenic/SCENIC_full_anchor_positive_negative.json
data/benchmarks/iot_instruction_benchmark_200.json
Supported additional formats:
.txt,.md,.text: entire file is used as training text..jsonl: each line can containtext, orpromptandcompletion.
Example:
{"text": "A complete training example goes here."}
{"prompt": "Question: ...\nAnswer:", "completion": " ..."}Run both 5-epoch SCENIC training jobs from a base decoder checkpoint:
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
NPROC_PER_NODE=8 \
MODEL_KIND=hf \
bash run_scenic_further_training_from_base.sh /PATH/TO/MY/BASE_DECODER_SLMThis writes:
outputs/scenic_further_training/contrastive_sft_5epoch/checkpoint-final/
outputs/scenic_further_training/regular_sft_5epoch/checkpoint-final/
Defaults:
- contrastive triplet SFT: 5 epochs on
data/scenic/SCENIC_full_anchor_positive_negative.json - regular SFT: 5 epochs on
data/scenic/SCENIC_full_training_dataset.json - regular SFT starts from the base model. To chain regular SFT after contrastive SFT, set
REGULAR_START=contrastive. MODEL_KIND=hfis the right setting forDecoder-Chinese-SLMcheckpoints. They are local Hugging Face-style Llama causal LM checkpoints, even after repairing the tokenizer withPreTrainedTokenizerFast.LOCAL_ONLY=1is the default, so model/tokenizer loading uses local files only.MODEL_KIND=customis only for the oldermodel.ptcustomDecoderOnlyTransformerformat.
Planner check without training:
DRY_RUN=1 bash run_scenic_further_training_from_base.sh /PATH/TO/MY/BASE_DECODER_SLMSingle-process test run:
decoder-train \
--model-path checkpoints/my-model \
--train-data data/train.jsonl \
--output-dir outputs/my-model-further-trained \
--block-size 1024 \
--batch-size 1 \
--gradient-accumulation-steps 8 \
--learning-rate 3e-5 \
--mixed-precision no8x H20 GPU run:
accelerate launch --config_file configs/accelerate_8xh20.yaml \
-m decoder_only.train \
--model-path checkpoints/my-decoder-slm \
--train-data data/train.jsonl \
--validation-data data/valid.jsonl \
--output-dir outputs/my-decoder-slm-further-trained \
--block-size 2048 \
--batch-size 1 \
--gradient-accumulation-steps 16 \
--learning-rate 2e-5 \
--mixed-precision bf16 \
--gradient-checkpointing \
--save-every 500The final checkpoint is written to:
outputs/my-decoder-slm-further-trained/checkpoint-final/
To scan a local decoder SLM checkpoint and create a JSON file shaped like
all_sparsity_results-2.json:
./generate_sparsity_results.sh checkpoints/my-decoder-slm all_sparsity_results.jsonOptional labels can be supplied as environment variables:
MODEL_TRAINING=regular_sft RUN_LABEL=magnitude_0p5 TARGET_SPARSITY=0.5 \
./generate_sparsity_results.sh checkpoints/my-decoder-slm all_sparsity_results.jsonThe script measures checkpoint sparsity only. Training and benchmark EM fields are emitted as
null unless they are produced by a separate evaluation workflow. The sparsity calculation follows
the encoder-only implementation: it measures prunable nn.Linear weights, excludes output heads by
default, and records encoder-style pruning_config fields.
To start from one decoder checkpoint, create dense SFT baselines, run the one-shot pruning matrix, run progressive magnitude pruning with recovery, and write the final JSON:
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
NPROC_PER_NODE=8 \
SPARSITY_GPU_IDS=0,1,2,3,4,5,6,7 \
MODEL_KIND=hf \
bash run_linear_sparsity_revision_from_base.sh /PATH/TO/MY/DECODER_SLM_CHECKPOINTThe script uses the bundled SCENIC regular, SCENIC contrastive, and IoT benchmark JSON files by default.
Expected final JSON rows:
- dense baselines: 2
- regular original one-shot: 7
- contrastive original one-shot: 7
- regular progressive magnitude: 2
- contrastive progressive magnitude: 2
- total JSON rows: 20
For a planner/schema check without training or pruning:
DRY_RUN=1 bash run_linear_sparsity_revision_from_base.sh /PATH/TO/MY/DECODER_SLM_CHECKPOINTpython -m decoder_only.generate \
--model-path outputs/my-decoder-slm-further-trained/checkpoint-final \
--prompt "Once upon a time" \
--max-new-tokens 80Useful options:
python -m decoder_only.generate --helpcheckpoints/is intentionally ignored except for.gitkeep.data/andoutputs/are intentionally ignored except for lightweight placeholders.- Use GitHub Releases, Hugging Face Hub, or another artifact store for large public weights.
- Keep tokenizer files with the checkpoint whenever the model was trained with a tokenizer other than UTF-8 bytes.