We study how safety alignment in open-weight language models holds up under pressure — fine-tuning attacks, prompt framing, deployment-style conditions — and publish what we find, including the results that don't confirm our hypotheses.
Founded in 2025 by Hamda Aden, an independent AI safety researcher and ML systems engineer.
| 96% | κ = 1.00 | <$0.10 |
|---|---|---|
| Attack success rate achieved after 50 LoRA fine-tuning steps — from a 20% baseline | Judge reliability on TamperBench's automated LLM-as-judge evaluation pipeline | Cost per experiment — every finding reproducible on a single T4 GPU |
- Fine-tuning attack resistance — how easily does safety alignment degrade under adversarial fine-tuning, and what does that reveal about whether alignment is learned as a robust property or a surface-level pattern?
- Evaluation robustness — do safety evaluations measure what we think they measure, or do they systematically overestimate model safety at category boundaries?
- Behavioural evaluation under varying conditions — how does model behaviour shift across prompt framings, oversight signals, and deployment-style contexts?
| Paper | Summary | Links |
|---|---|---|
| TamperBench (2026) | 100-prompt benchmark across 5 harm categories measuring LoRA fine-tuning attack resistance. Safety alignment degraded from 20% → 96% attack success rate in 50 steps, for under $0.10 of compute. | Paper · Code |
| Prompt-Framing Effects (2026) | 600-response study examining alignment-relevant behaviour across 4 prompt conditions. Null result on the primary hypothesis, with a 5-dimensional continuous scoring framework revealing condition-level differences binary labels missed. | Paper · Code |
Most of what gets published in AI safety is positive findings. We think that's a problem. A null result, reported honestly with the right statistical rigor, is still evidence — and the field needs more of it to avoid mistaking the absence of contrary findings for the absence of risk.
Every project follows the same standard:
- Automated, reusable pipelines — every experiment is scriptable end-to-end, so findings can be reproduced or extended by anyone
- LLM-as-judge validation against human labels — judge reliability is measured, not assumed (Cohen's κ = 1.00 on TamperBench)
- Pre-registered statistical thresholds — Bonferroni correction, Wilson confidence intervals, and effect sizes reported alongside p-values
- Low-resource by design — every experiment runs on consumer-accessible hardware (T4 GPU), because safety research shouldn't require an industrial compute budget to verify
We're a small, independent lab. If you're working on adjacent problems — fine-tuning attack resistance, evaluation robustness, open-weight model safety — we'd like to hear from you.
- Open an issue on any repo with questions, replications, or extensions
- Reach out via LinkedIn