Skip to content

krisis - v0.2.0

Pre-release
Pre-release

Choose a tag to compare

@devsgnr devsgnr released this 26 May 00:20
· 7 commits to main since this release
d057745

Krisis v0.2.0

Krisis v0.2.0 is a major update that turns the project into a cleaner, more scalable clinical evaluation framework for LLM safety behavior. This release adds a unified OpenRouter-powered API backend, stronger CKD benchmark reporting, improved documentation, and a more reproducible results workflow.

Highlights

  • Switched to a unified APIBackend powered by OpenRouter.
  • Added support for OpenRouter model IDs such as:
    • openai/gpt-5.5
    • anthropic/claude-opus-4.7
    • x-ai/grok-4.3
    • google/gemini-3.5-flash
  • Added batched and concurrent evaluation support.
  • Added retry/backoff handling for transient provider failures.
  • Added prompt capture with patient data redacted.
  • Added execution metadata including elapsed time, throughput, batch size, concurrency, and token usage.
  • Added CKDSuite detection and staging benchmark reports across four model families.
  • Added MkDocs Material documentation site with API reference, guides, and benchmark report.
  • Added PyPI-ready package metadata and install extras.

Breaking Changes

  • Removed individual provider backends in favor of APIBackend.
  • API model access now uses OPENROUTER_API_KEY by default.
  • Provider-specific installation extras have been replaced with:
    • krisis[api]
    • krisis[hf] placeholder for future Hugging Face support
  • Model names should now be passed as OpenRouter IDs.

Example:

from krisis.backends.api import APIBackend

backend = APIBackend(
    model="openai/gpt-5.5",
    reasoning_effort="low",
)

CKDSuite

  • Detection and staging tasks are now benchmarked across:
    • openai/gpt-5.5
    • anthropic/claude-opus-4.7
    • x-ai/grok-4.3
    • google/gemini-3.5-flash
  • Added full-feature CKD benchmarking using the local UCI CKD CSV.
  • Added synthetic stress-test rows generated from the training split.
  • Added clearer CKD staging labels based on KDIGO-style eGFR thresholds.
  • Added deferral-aware evaluation for ambiguous or clinically unsafe cases.
  • Added progression task support as a synthetic stress-test task. Progression reporting will be expanded later.

Metrics And Reporting

  • Added metrics-only JSON output for model comparison and plotting.
  • Added execution metadata to reports:
    • batch_size
    • max_concurrency
    • n_api_batches
    • elapsed_seconds
    • records_per_second
    • input_tokens
    • output_tokens
    • token_total
  • Standardized result metric keys for programmatic use.
  • Improved text reports to include execution details.
  • Added prompt templates to report metadata with patient data redacted.

Documentation

  • Added a full MkDocs Material documentation site.
  • Added Framework Guide pages for:
    • datasets
    • suites
    • metrics
    • model backends
    • outputs
    • research status
  • Added API reference pages for:
    • benchmark
    • data
    • suites
    • backends
    • metrics
    • results
  • Added CKDSuite benchmark report with:
    • dataset composition
    • synthetic data method
    • captured prompts
    • deferral criteria
    • detection results
    • staging results
    • progression status
    • limitations
  • Added documentation for getting an OpenRouter API key.

Packaging

  • Version bumped to 0.2.0.
  • Added PyPI package build support.
  • Added krisis[api] extra for API-backed model evaluation.
  • Added placeholder krisis[hf] extra for future Hugging Face backend support.
  • Added PyPI and download badges to the README.

Validation

Pre-release checks passed:

  • 45 tests passing
  • mkdocs build --strict passing
  • package build passing for:
    • krisis-0.2.0.tar.gz
    • krisis-0.2.0-py3-none-any.whl
  • twine check passing for both distribution artifacts

Known warnings:

  • 450 pandas deprecation warnings from CKD preprocessing tests. These are non-blocking and will be addressed in a future cleanup release.