Course: High Performance Machine Learning Semester: Spring 2026 Instructor: Dr. Kaoutar El Maghraoui
- Team Name: Simamba
- Members:
- Soumil Baldota (ssb2234) - Simamba implementation, training pipeline, W&B/checkpointing, experiments, and artifact generation
- Ansen Shia (as8008) - vLLM integration
hpml26/vllm, experiment design, analysis, report, and presentation - David Zhang (dwz2107) - baselines, result analysis, report, and presentation
- GitHub repository: https://github.com/hpmls26/simamba
- Final report:
deliverables/HPML_Final_Project_Report.pdf - Final presentation:
deliverables/HPML_Final_Presentation.pptx - Experiment-tracking dashboard: Weights & Biases training project
- Profiling dashboard: Weights & Biases profiling project
- Exported checkpoints: Simamba midpoint 10M and Mamba2 10M
The final report PDF and the presentation file are checked into the deliverables/ folder of this repository and uploaded to CourseWorks.
This project evaluates whether a Simpson-style discretization can improve the SISO state recurrence used by Mamba-3-style language-model mixers. The target workload is language-model training on SlimPajama, with supporting inference and kernel benchmarking to understand whether the new recurrence is practical. The optimization scope is both training and inference: training bottlenecks are recurrent-state optimization stability and gradient health, while inference bottlenecks are prefill throughput, memory traffic, and vLLM cache behavior.
- Model architecture:
MambaLMHeadModelwith 10M-parameter Mamba-family mixers. The controlled shape isd_model=160,n_layer=8,d_state=64,headdim=32,seq_len=128, tied embeddings, RMSNorm, and a vocabulary size of 50,280. - Proposed layer:
Simamba, a Mamba-3-inspired SISO mixer that replaces the width-2 trapezoid update with a width-3 Simpson-style recurrence. It supports--simamba-discretization {simpson,trapezoid}, midpoint control, coefficient logit offsets, and optional Mamba2-style local convolution over thex/B/Cstream. - Baselines: Mamba2, matched Simamba trapezoid, Simamba Simpson default, Simamba Simpson low-control, Simamba midpoint, and local-conv variants.
- Framework: PyTorch, Triton,
mamba_ssm, Hugging Facetransformers, W&B, NumPy memmaps, and optionalcausal-conv1d. - Dataset:
MBZUAI-LLM/SlimPajama-627B-DC, streamed from Hugging Face and tokenized withEleutherAI/gpt-neox-20b. The dataset card reports an MIT license. Prepared subsets are100M/10Mand500M/50Mtrain/validation tokens. - Custom layers/modifications:
mamba_ssm/modules/simamba.py,mamba_ssm/ops/triton/simamba/, checkpoint metadata/resume logic, fixed validation sampling, no-replacement epoch training sampling, compression evaluation, local-conv controls, and benchmark/figure generation. - Hardware target: Final controlled experiments ran on 1x NVIDIA Tesla V100-SXM2 16GB. The local validated environment was Linux, driver 550.90.07, CUDA 12.4, PyTorch 2.6.0+cu124, Triton 3.2.0,
transformers5.7.0, andcausal_conv1d1.6.0.
The central result is negative but reproducible: the current Simpson parameterization trains stably, but it does not beat the matched trapezoid baseline or Mamba2. While it does not yet outperform, the approach remains competitive across both long-run training and controlled ablations. Simpson variants also exhibit a distinct optimization profile, with slower convergence, larger gradient norms, and additional stabilization requirements such as coefficient offsets and midpoint control, suggesting that future potential work in parameterization or optimization may better realize the potential of higher-order recurrence methods.
| Metric | Baseline | Optimized / Proposed Variant | Delta (Improvement) |
|---|---|---|---|
| 500M-token best validation loss | Mamba2: 4.8625 | Simamba midpoint: 4.9178 | +0.0552 worse |
| 500M-token matched trapezoid validation loss | Simamba trapezoid: 4.9326 | Simamba default Simpson: 4.9319 best, 4.9671 final | best roughly tied; final worse |
| 50M-token matched discretization loss | Simamba trapezoid: 5.8195 | Simpson low-control: 5.8929 | +0.0734 worse |
| 50M-token local-conv discretization loss | Local-conv trapezoid: 5.8469 | Local-conv Simpson low-control: 5.8793 | +0.0324 worse |
| Training throughput, 500M tail median | Mamba2: 16,095 tok/s | Simamba midpoint: 12,594 tok/s | 21.7% lower |
| Decode-step latency, B4/P1024/G256 | Mamba-3: 0.0160 ms | Simamba: 0.0173 ms | 1.08x slower |
| Prefill latency, B4/P1024/G256 | Mamba-3: 0.8602 ms | Simamba: 353.3066 ms | 410.7x slower |
| Prefill peak memory, B4/P1024/G256 | Mamba-3: 167.3911 MB | Simamba: 124.2363 MB | 25.8% lower |
| Row-wise int8 QDQ loss delta | Mamba2: +0.0019 | Simamba midpoint: +0.0014 | both nearly lossless |
| Exported model size on disk | Mamba2 HF export: 40 MB | Simamba HF export: 40 MB | roughly equal |
Hardware: 1x NVIDIA Tesla V100-SXM2 16GB, CUDA 12.4, PyTorch 2.6.0+cu124, Triton 3.2.0, Debian Linux.
Headline result: Simamba produced a stable, reproducible SSM training pipeline, but the current Simpson recurrence did not outperform trapezoid: at 50M tokens, matched trapezoid reached 5.8195 validation loss while the best Simpson control reached 5.8929, and kernel benchmarking showed decode overhead was small but prefill was not yet competitive.
.
+-- README.md # HPML submission README
+-- README_upstream_mamba.md # Preserved upstream Mamba README
+-- LICENSE # Apache-2.0
+-- pyproject.toml / setup.py # Package metadata and pinned HPML extra
+-- deliverables/
| +-- HPML_Final_Project_Report.pdf
| +-- HPML_Final_Presentation.pptx
+-- docs/
| +-- paper.tex # Final paper source
| +-- hpml_simamba_report.md # Extended experiment report
| +-- assets/paper/ # Generated plots, CSV summaries, and figures
+-- mamba_ssm/
| +-- modules/simamba.py # Simamba mixer and controls
| +-- modules/mamba2.py # Mamba2 compatibility and local-conv controls
| +-- ops/triton/simamba/ # Simamba SISO reference/Triton kernels
+-- scripts/
| +-- prepare_slimpajama.py # Dataset preprocessing
| +-- train_simamba_lm.py # Training entry point
| +-- eval_checkpoint_compression.py # Evaluation/compression entry point
| +-- generate_hpml_paper_assets.py # Figure/table generation
+-- benchmarks/ # SISO benchmark scripts and plots
+-- profiling/ # NSYS/NCU/vLLM profiling harnesses
+-- data/ # Local dataset memmaps; ignored by git
+-- outputs/ # Local checkpoints; ignored by git
+-- run_logs/ # Local logs; ignored by git
+-- wandb/ # Local W&B files; ignored by git
# Clone
git clone https://github.com/hpmls26/simamba.git
cd simamba
# Create a clean Python environment
python3.10 -m venv .venv
source .venv/bin/activate
python -m pip install -U pip wheel setuptools packaging ninja
# Install the pinned HPML V100/CUDA 12.4 environment from pyproject.toml.
python -m pip install --extra-index-url https://download.pytorch.org/whl/cu124 \
-e '.[hpml-v100]' --no-build-isolationSystem requirements: Python 3.10+, CUDA-capable NVIDIA GPU, and at least 16 GB GPU memory for the reported 10M-parameter V100 experiments. The pinned experiment environment is the hpml-v100 optional dependency group in pyproject.toml. The original A6000 target was unavailable, so final numbers are from a single V100. On this V100/Triton 3.2 environment, the original Mamba-3 path requiring triton.language.make_tensor_descriptor did not run directly.
The profiling commands additionally require nvidia-smi, Nsight Systems
(nsys), and Nsight Compute (ncu) on the machine running the traces. By
default the profiler looks for the Nsight CLIs under /usr/local/cuda/bin;
pass --nsys-bin or --ncu-bin if they are installed elsewhere.
Public experiment tracking is under:
Dashboard: https://wandb.ai/ssb2234-columbia/simamba
Platform used: Weights & Biases
The separate profiling dashboard is https://wandb.ai/ssb2234-columbia/profiling.
Key W&B runs:
| Run | Link |
|---|---|
| Mamba2 500M+500M | https://wandb.ai/ssb2234-columbia/simamba/runs/dm74y180 |
| Simamba Simpson 500M+500M | https://wandb.ai/ssb2234-columbia/simamba/runs/wu8bhcqa |
| Simamba midpoint 500M+500M | https://wandb.ai/ssb2234-columbia/simamba/runs/juxcboyg |
| Simamba trapezoid 500M | https://wandb.ai/ssb2234-columbia/simamba/runs/0rjsv3c1 |
| 50M trapezoid | https://wandb.ai/ssb2234-columbia/simamba/runs/jgjlc1sr |
| 50M Simpson default | https://wandb.ai/ssb2234-columbia/simamba/runs/9yuk3voc |
| 50M Simpson low-control | https://wandb.ai/ssb2234-columbia/simamba/runs/bouk7ayj |
| Local-conv trapezoid | https://wandb.ai/ssb2234-columbia/simamba/runs/03okzay6 |
| Local-conv Simpson default | https://wandb.ai/ssb2234-columbia/simamba/runs/eglr4x75 |
| Local-conv Simpson low-control | https://wandb.ai/ssb2234-columbia/simamba/runs/w6u32odi |
The dashboard contains tagged baseline and Simamba runs with model configuration, training/validation curves, gradient norms, learning rate, throughput, GPU memory where available, profiling summaries, ablation metadata, and run tags separating baseline, Simpson, midpoint, trapezoid, and local-conv variants. Local W&B-derived summaries are committed/generated under docs/assets/paper/run_summary.csv.
Prepare the main 500M/50M SlimPajama split:
.venv/bin/python scripts/prepare_slimpajama.py \
--output-dir data/slimpajama_500m_50m \
--train-tokens 500000000 \
--val-tokens 50000000 \
--seed 1337 \
--shuffle-buffer 100000Prepare the smaller 100M/10M split:
.venv/bin/python scripts/prepare_slimpajama.py \
--output-dir data/slimpajama_100m_10m \
--train-tokens 100000000 \
--val-tokens 10000000 \
--seed 1337 \
--shuffle-buffer 100000The prepared local 500M/50M metadata is:
| Split | Tokens | Documents | Path |
|---|---|---|---|
| Train | 500,000,000 | 691,532 | data/slimpajama_500m_50m/train.bin |
| Validation | 50,000,000 | 49,451 | data/slimpajama_500m_50m/val.bin |
Token IDs are stored as uint16 memmaps and are read without loading the full arrays into RAM.
Run the main 500M-token comparison:
STAMP=repro_500m \
WANDB_PROJECT=simamba \
WANDB_ENTITY=ssb2234-columbia \
bash scripts/run_10m_discretization_comparison_500m.shContinue the main runs and launch the matched 500M trapezoid run:
STAMP=repro_continue_plus_trap \
WANDB_PROJECT=simamba \
WANDB_ENTITY=ssb2234-columbia \
bash scripts/run_10m_continue_plus_trap_500m.shRun the 50M-token matched Simpson/trapezoid ablation:
STAMP=repro_followup_50m \
WAIT_PIDS="" \
WANDB_PROJECT=simamba \
WANDB_ENTITY=ssb2234-columbia \
bash scripts/run_followup_ablation_50m_after_current.shRun the local-convolution Simamba ablation:
STAMP=repro_localconv_dconv4_50m \
WANDB_PROJECT=simamba \
WANDB_ENTITY=ssb2234-columbia \
bash scripts/run_simamba_localconv_50m.shThese launcher scripts start background jobs and write logs to run_logs/ and checkpoints to outputs/. Set WANDB_API_KEY in the environment before running with W&B enabled.
Recompute fixed-validation losses for checkpoints using the same validation sampling used in the compression study:
.venv/bin/python scripts/eval_checkpoint_compression.py \
--checkpoint mamba2_500m=outputs/disc10m_mamba2_fp32_500m_20260502_185217/best/trainer.pt \
--checkpoint simamba_midpoint_500m=outputs/disc10m_simamba_midpoint_fp32_500m_20260502_190333/best/trainer.pt \
--checkpoint simamba_trapezoid_500m=outputs/disc10m_simamba_trapezoid_vec_500m_20260503_0609_trap_vec/best/trainer.pt \
--variants baseline \
--output run_logs/reproduce_baseline_eval.jsonlRun the full compression/pruning perturbation evaluation:
.venv/bin/python scripts/eval_checkpoint_compression.py \
--checkpoint mamba2_500m=outputs/disc10m_mamba2_fp32_500m_20260502_185217/best/trainer.pt \
--checkpoint simamba_midpoint_500m=outputs/disc10m_simamba_midpoint_fp32_500m_20260502_190333/best/trainer.pt \
--checkpoint simamba_trapezoid_500m=outputs/disc10m_simamba_trapezoid_vec_500m_20260503_0609_trap_vec/best/trainer.pt \
--output run_logs/compression_eval_reproduce.jsonlRegenerate the final paper plots and CSV summaries:
.venv/bin/python scripts/generate_hpml_paper_assets.pyTo regenerate all profiling artifacts, run the orchestrator from the repository
root on a CUDA machine with Nsight Systems, Nsight Compute, vLLM, nvidia-smi,
and W&B credentials available:
python profiling/run_all_profiling.py --wandbThe script runs:
- NSYS traces for the
mamba3,simamba, andimprovedSISO kernel targets. - NCU filtered-kernel profiles for the Mamba3, Simamba, and improved kernels.
- Triton-vs-PyTorch correctness tables and a correctness error plot.
- vLLM repeated measurements for
soumil1/mamba2-10m-slimpajama-500mand the improved Simamba export. - A combined vLLM CSV/plot plus a W&B artifact bundle.
Profiling logs to W&B project profiling; training logs still use simamba.
Use --dry-run to print the underlying commands, --steps to run a subset
such as --steps kernel-nsys,kernel-correctness,vllm, and --sudo-ncu if NCU
fails with ERR_NVGPUCTRPERM on a machine where sudo is configured.
The main regenerated artifacts are:
profiling/results/kernel_suite/nsys/*.nsys-repprofiling/results/kernel_suite/nsys/csv_exports/*.csvprofiling/results/kernel_suite/ncu/*.csvprofiling/results/kernel_suite/ncu/*_details.txtprofiling/results/kernel_suite/kernel_correctness.csvprofiling/results/kernel_suite/kernel_correctness.mdprofiling/results/kernel_suite/kernel_correctness.pngprofiling/results/vllm_mamba2_summary.csvprofiling/results/vllm_mamba2_raw.csvprofiling/results/vllm_mamba2.pngprofiling/results/vllm_improved_simamba_summary.csvprofiling/results/vllm_improved_simamba_raw.csvprofiling/results/vllm_improved_simamba.pngprofiling/results/vllm_mamba2_vs_improved_summary.csvprofiling/results/vllm_mamba2_vs_improved_raw.csvprofiling/results/vllm_mamba2_vs_improved.png
The vLLM summaries include TTFT, TPOT, tok/s, decode-loop tok/s, prefill-probe
latency, requests/s, GPU memory peak/delta, model-load memory delta, prefix
caching on/off, and 512-token repeated-prefix prompts. Open .nsys-rep files
in Nsight Systems, inspect NCU CSVs for per-kernel counters, and use the PNGs
as report-ready matplotlib figures.
The corresponding W&B project is
ssb2234-columbia/profiling.
The shortest meaningful reproduction is the 50M-token local-conv ablation, which checks whether Simpson beats matched local-conv trapezoid:
python -m venv .venv
source .venv/bin/activate
python -m pip install --extra-index-url https://download.pytorch.org/whl/cu124 \
-e '.[hpml-v100]' --no-build-isolation
.venv/bin/python scripts/prepare_slimpajama.py \
--output-dir data/slimpajama_500m_50m \
--train-tokens 500000000 \
--val-tokens 50000000 \
--seed 1337 \
--shuffle-buffer 100000
STAMP=repro_localconv_dconv4_50m \
WANDB_PROJECT=simamba \
WANDB_ENTITY=ssb2234-columbia \
bash scripts/run_simamba_localconv_50m.shAfter the three background jobs finish, inspect:
outputs/disc10m_simamba_localconv_trapezoid_50m_repro_localconv_dconv4_50m/best/metrics.json
outputs/disc10m_simamba_localconv_simpson_lowctrl_50m_repro_localconv_dconv4_50m/best/metrics.json
outputs/disc10m_simamba_localconv_simpson_50m_repro_localconv_dconv4_50m/best/metrics.jsonThe reported result from the completed run is trapezoid 5.8469, Simpson low-control 5.8793, and Simpson default 5.8890.
- The stable fp32 pipeline worked. After switching to fp32 parameter storage, cosine decay, fixed validation spans, nonfinite-step handling, and reliable checkpoint metadata, the controlled runs completed with zero nonfinite skips.
- The current Simpson recurrence is not quality-positive. The best 500M Simamba checkpoint was midpoint Simpson at validation loss
4.9178, behind Mamba2 at4.8625; at 50M tokens, matched trapezoid beat both Simpson variants. - The negative lag-2 Simpson correction is the likely optimization issue. Lowering the Simpson control offset from default to
-4.0improved validation from5.9178to5.8929, but still did not catch trapezoid. - Mamba2 is not a clean discretization baseline. Mamba2 includes causal local convolution over
x/B/C; shrinking/removing it hurt Mamba2, so Simamba needed matched local mixing for a fairer comparison. - Local mixing helped but did not flip the result. With
--simamba-d-conv 4, local-conv trapezoid reached5.8469, while local-conv Simpson low-control reached5.8793. - Systems performance needs kernel work. Decode-step overhead is near Mamba-3, but Simamba prefill is hundreds of times slower in the current benchmark path.
- Source changes live mainly under
mamba_ssm/,scripts/, andbenchmarks/. - Trained checkpoints are stored locally under
outputs/, with portable exports underhf_exports/. - The best exported checkpoints are also available on Hugging Face:
soumil1/simamba-midpoint-10m-slimpajama-500msoumil1/mamba2-10m-slimpajama-500msoumil1/mamba2-10m-slimpajama-500m-vllm
- W&B credentials should be supplied through
WANDB_API_KEY. Do not commit API keys or private tokens. - The Simamba Hugging Face export uses custom remote code; serving through vLLM requires the Transformers backend:
PYTHONPATH=/path/to/simamba \
vllm serve soumil1/simamba-midpoint-10m-slimpajama-500m \
--trust-remote-code \
--model-impl transformers \
--dtype float32 \
--max-model-len 128Per the HPML AI Use Policy, this submission discloses AI assistance.
Did your team use any AI tool in completing this project?
- No, we did not use any AI tool.
- Yes, we used AI assistance as described below.
Tool(s) used: ChatGPT/Codex.
Specific purpose: Assisted with codebase navigation, README/LaTeX editing, summarizing existing repository artifacts, checking consistency against generated figures/tables, and identifying submission-compliance gaps in team-drafted material.
Sections affected: README.md organization, submission links, reproducibility instructions, results summary, and AI disclosure wording.
How we verified correctness: The README values and paths were checked against docs/hpml_simamba_report.md, docs/paper.tex, docs/assets/paper/run_summary.csv, launcher scripts, benchmark scripts, profiler scripts, W&B dashboards, and Hugging Face export metadata. Final interpretations, performance reasoning, numerical claims, and scientific conclusions were validated by the team.
By submitting this project, the team confirms that the analysis, interpretations, and conclusions are our own, and that any AI assistance is fully disclosed above. The same disclosure block appears as an appendix in the final report.
Released under the Apache License 2.0. See LICENSE. This project builds on the upstream Mamba SSM repository by Tri Dao and Albert Gu.
If you build on this work, please cite:
@misc{baldota2026simamba,
title = {Simamba: Evaluating Simpson-Style Discretization for Mamba-3 State Space Language Models},
author = {Baldota, Soumil and Shia, Ansen and Zhang, David},
year = {2026},
note = {HPML Spring 2026 Final Project, Columbia University},
url = {https://github.com/hpmls26/simamba}
}Open a GitHub issue or contact the team at ssb2234@columbia.edu, as8008@columbia.edu, or dwz2107@columbia.edu.




