Skip to content

Releases: devsgnr/krisis

krisis - v0.2.7

krisis - v0.2.7 Pre-release
Pre-release

Choose a tag to compare

@devsgnr devsgnr released this 05 Jun 21:06
778620a

Krisis v0.2.7 Release Notes

Krisis v0.2.7 improves local/open-weight model experimentation and hardens benchmark execution for models that do not reliably return valid JSON.

What’s New

  • Added experimental Hugging Face TransformersBackend

    • Run text-generation models locally or in GPU notebooks such as Google Colab.
    • Supports model IDs, device selection, and gated model access through HF_TOKEN.
    • Intended for text-generation models only.
  • Improved malformed response handling

    • Hugging Face models that return malformed or non-JSON output no longer have to break the whole benchmark immediately.
    • Krisis can surface a warning-style malformed response result so failures are easier to inspect and debug.
  • Better backend documentation

    • Added guidance that the Hugging Face backend only supports text-generation models.
    • Documented HF_TOKEN usage for gated Hugging Face models.
    • Clarified CPU vs GPU expectations for local transformer runs.
  • CKDSuite report updates

    • Expanded the CKDSuite benchmark report to include Detection, Staging, and Progression results.
    • Added more discussion around composite safety score interpretation.
    • Refined preliminary reporting language for research-style presentation.
  • Version and packaging updates

    • Package version bumped to 0.2.7.
    • Build/test checks were run before release.

Notes

The Hugging Face Transformers backend is experimental in this release. CPU execution can be very slow, so GPU-backed environments are recommended for practical benchmark runs.

Install

pip install krisis==0.2.7

For API-backed models through OpenRouter:

pip install "krisis[api]==0.2.7"

For Hugging Face Transformers:

pip install "krisis[hf]==0.2.7"

krisis - v0.2.6

krisis - v0.2.6 Pre-release
Pre-release

Choose a tag to compare

@devsgnr devsgnr released this 27 May 23:32
f3156e9

Krisis v0.2.6

This release hardens the experimental Hugging Face Transformers backend and clarifies its supported model scope.

What Changed

  • Added explicit support boundaries for the Hugging Face backend:

    • TransformersBackend supports causal text-generation models only.
    • Models must be loadable with AutoModelForCausalLM.
    • Classifier, embedding, masked-language, seq2seq, and multimodal-only models now raise a clear UnsupportedTransformersModelError.
  • Improved experimental Hugging Face runtime behavior:

    • Malformed single-row model outputs no longer crash the full benchmark.
    • Krisis now warns, preserves the raw response, marks the row as abstained, and continues the run.
    • Added better Gemma/MedGemma-style prompt compatibility by folding system instructions into user messages when needed.
    • Added generation pad_token_id fallback.
  • Added Hugging Face gated-model support:

    • hf_token can be passed directly.
    • HF_TOKEN is used automatically when available.
  • Hardened CKD feature engineering:

    • eGFR computation now clamps non-physiologic imputed creatinine/age values before applying the CKD-EPI equation.
    • Prevents complex-number failures during suite.load().
  • Updated documentation:

    • README and MkDocs now explicitly describe the Hugging Face backend as experimental.
    • Documentation now states that only causal text-generation models are supported.

Install

pip install -U "krisis[hf]==0.2.6"

krisis - v0.2.0

krisis - v0.2.0 Pre-release
Pre-release

Choose a tag to compare

@devsgnr devsgnr released this 26 May 00:20
d057745

Krisis v0.2.0

Krisis v0.2.0 is a major update that turns the project into a cleaner, more scalable clinical evaluation framework for LLM safety behavior. This release adds a unified OpenRouter-powered API backend, stronger CKD benchmark reporting, improved documentation, and a more reproducible results workflow.

Highlights

  • Switched to a unified APIBackend powered by OpenRouter.
  • Added support for OpenRouter model IDs such as:
    • openai/gpt-5.5
    • anthropic/claude-opus-4.7
    • x-ai/grok-4.3
    • google/gemini-3.5-flash
  • Added batched and concurrent evaluation support.
  • Added retry/backoff handling for transient provider failures.
  • Added prompt capture with patient data redacted.
  • Added execution metadata including elapsed time, throughput, batch size, concurrency, and token usage.
  • Added CKDSuite detection and staging benchmark reports across four model families.
  • Added MkDocs Material documentation site with API reference, guides, and benchmark report.
  • Added PyPI-ready package metadata and install extras.

Breaking Changes

  • Removed individual provider backends in favor of APIBackend.
  • API model access now uses OPENROUTER_API_KEY by default.
  • Provider-specific installation extras have been replaced with:
    • krisis[api]
    • krisis[hf] placeholder for future Hugging Face support
  • Model names should now be passed as OpenRouter IDs.

Example:

from krisis.backends.api import APIBackend

backend = APIBackend(
    model="openai/gpt-5.5",
    reasoning_effort="low",
)

CKDSuite

  • Detection and staging tasks are now benchmarked across:
    • openai/gpt-5.5
    • anthropic/claude-opus-4.7
    • x-ai/grok-4.3
    • google/gemini-3.5-flash
  • Added full-feature CKD benchmarking using the local UCI CKD CSV.
  • Added synthetic stress-test rows generated from the training split.
  • Added clearer CKD staging labels based on KDIGO-style eGFR thresholds.
  • Added deferral-aware evaluation for ambiguous or clinically unsafe cases.
  • Added progression task support as a synthetic stress-test task. Progression reporting will be expanded later.

Metrics And Reporting

  • Added metrics-only JSON output for model comparison and plotting.
  • Added execution metadata to reports:
    • batch_size
    • max_concurrency
    • n_api_batches
    • elapsed_seconds
    • records_per_second
    • input_tokens
    • output_tokens
    • token_total
  • Standardized result metric keys for programmatic use.
  • Improved text reports to include execution details.
  • Added prompt templates to report metadata with patient data redacted.

Documentation

  • Added a full MkDocs Material documentation site.
  • Added Framework Guide pages for:
    • datasets
    • suites
    • metrics
    • model backends
    • outputs
    • research status
  • Added API reference pages for:
    • benchmark
    • data
    • suites
    • backends
    • metrics
    • results
  • Added CKDSuite benchmark report with:
    • dataset composition
    • synthetic data method
    • captured prompts
    • deferral criteria
    • detection results
    • staging results
    • progression status
    • limitations
  • Added documentation for getting an OpenRouter API key.

Packaging

  • Version bumped to 0.2.0.
  • Added PyPI package build support.
  • Added krisis[api] extra for API-backed model evaluation.
  • Added placeholder krisis[hf] extra for future Hugging Face backend support.
  • Added PyPI and download badges to the README.

Validation

Pre-release checks passed:

  • 45 tests passing
  • mkdocs build --strict passing
  • package build passing for:
    • krisis-0.2.0.tar.gz
    • krisis-0.2.0-py3-none-any.whl
  • twine check passing for both distribution artifacts

Known warnings:

  • 450 pandas deprecation warnings from CKD preprocessing tests. These are non-blocking and will be addressed in a future cleanup release.

v0.1.0

v0.1.0 Pre-release
Pre-release

Choose a tag to compare

@devsgnr devsgnr released this 19 May 22:43
7d0a918

Krisis v0.1.0

Initial release of Krisis, a clinical evaluation framework for testing LLM behavior on high-stakes medical reasoning tasks.

Krisis v0.1 focuses on Chronic Kidney Disease (CKD), with support for benchmarking whether models answer correctly, abstain appropriately, and report confidence that aligns with observed correctness.

Highlights

  • Added CKDSuite for CKD-focused LLM evaluation.
  • Added three CKD benchmark tasks:
    • Detection: CKD vs not CKD
    • Staging: eGFR-derived CKD stage classification
    • Progression: synthetic two-visit trajectory stress testing
  • Added provider backends for:
    • OpenAI
    • Anthropic
    • Grok
    • Google Gemini
  • Added batched evaluation and configurable concurrency.
  • Added retry/backoff handling for transient provider failures.
  • Added fallback behavior for malformed batched JSON.
  • Added redacted prompt capture for auditability.
  • Added execution metadata:
    • elapsed time
    • records per second
    • batch size
    • max concurrency
    • input tokens
    • output tokens
    • total tokens
  • Added metrics for:
    • Accuracy
    • Balanced Accuracy
    • Selective Accuracy
    • Abstention Rate
    • Answer Rate / Coverage
    • Deferral Alignment
    • Expected Calibration Error
    • Brier Score
  • Added JSON and text reporting outputs.
  • Added MkDocs documentation site with guides, API reference, and blog-style benchmark reports.

Notes

Krisis v0.1 is intentionally scoped to CKD. Diabetes and hypertension suites are planned but not implemented yet.

The CKD progression task uses synthetic two-visit trajectories because the UCI CKD dataset is cross-sectional, not longitudinal.

Krisis is a research and evaluation framework. It is not a clinical diagnostic tool and should not be used for patient care decisions.

Installation

pip install krisis

Provider-specific installs:

pip install "krisis[openai]"
pip install "krisis[anthropic]"
pip install "krisis[grok]"
pip install "krisis[gemini]"

Documentation

Docs: https://devsgnr.github.io/krisis/

Repository: https://github.com/devsgnr/krisis