Repository navigation
v0.1.0
Pre-releaseKrisis v0.1.0
Initial release of Krisis, a clinical evaluation framework for testing LLM behavior on high-stakes medical reasoning tasks.
Krisis v0.1 focuses on Chronic Kidney Disease (CKD), with support for benchmarking whether models answer correctly, abstain appropriately, and report confidence that aligns with observed correctness.
Highlights
- Added
CKDSuitefor CKD-focused LLM evaluation. - Added three CKD benchmark tasks:
- Detection: CKD vs not CKD
- Staging: eGFR-derived CKD stage classification
- Progression: synthetic two-visit trajectory stress testing
- Added provider backends for:
- OpenAI
- Anthropic
- Grok
- Google Gemini
- Added batched evaluation and configurable concurrency.
- Added retry/backoff handling for transient provider failures.
- Added fallback behavior for malformed batched JSON.
- Added redacted prompt capture for auditability.
- Added execution metadata:
- elapsed time
- records per second
- batch size
- max concurrency
- input tokens
- output tokens
- total tokens
- Added metrics for:
- Accuracy
- Balanced Accuracy
- Selective Accuracy
- Abstention Rate
- Answer Rate / Coverage
- Deferral Alignment
- Expected Calibration Error
- Brier Score
- Added JSON and text reporting outputs.
- Added MkDocs documentation site with guides, API reference, and blog-style benchmark reports.
Notes
Krisis v0.1 is intentionally scoped to CKD. Diabetes and hypertension suites are planned but not implemented yet.
The CKD progression task uses synthetic two-visit trajectories because the UCI CKD dataset is cross-sectional, not longitudinal.
Krisis is a research and evaluation framework. It is not a clinical diagnostic tool and should not be used for patient care decisions.
Installation
pip install krisisProvider-specific installs:
pip install "krisis[openai]"
pip install "krisis[anthropic]"
pip install "krisis[grok]"
pip install "krisis[gemini]"Documentation
Docs: https://devsgnr.github.io/krisis/
Repository: https://github.com/devsgnr/krisis