Skip to content

v0.3.0: RAG scorers, model adapters, and assessment fixes

Latest

Choose a tag to compare

@knisar knisar released this 16 Sep 22:42
· 16 commits to main since this release
01d576d

rai-toolkit v0.3.0 brings new RAG scorers, more model adapters, and fixes to assessment and red-team reporting. It includes 38 merged commits since v0.2.0, plus release preparation, with contributions from 16 community developers.

Highlights

  • Evidence-backed RAG evaluation: GroundednessScorer, RetrievalRelevanceScorer, ContextPrecisionScorer, and ContextRecallScorer check grounding, retrieval relevance, and context quality. Evidence validation, normalized span matching, and behavioral-refusal handling improve the grounding results. See #13, #17, #18, #37, and #58.
  • More ways to connect models: an Anthropic Messages API adapter and vendor-neutral CallableModel join the existing adapters. A shared adapter contract and conformance suite document and check the expected behavior. See #56, #61, and #64.
  • Assessment and integration fixes: composite scoring honors asynchronous scorers and unassessed results; Weave assessments preserve additional scorers and configured names; default judge categories are retained. See #30, #31, #33, and #51.
  • Red-team reporting: four attack templates were added. Execution errors and missing evidence are kept out of resistance scores, and incomplete source coverage fails the coverage gate. See #11, #67, and the full change history.
  • HR preset and policy isolation: the HR industry preset is available, and built-in policies stay scoped to their selected industry. See #24 and #52.
  • Contributor and deployment documentation: new self-hosting, model-adapter, and scorer-authoring guidance; clearer contribution setup and claim flow; and permanent contributor credits. See the documentation, contributing guide, and contributors.

Behavior changes to review when upgrading

  • Unsupported OutputFormatScorer formats now return an unassessed result instead of a passing score. Consumers should check ScorerResult.assessed when interpreting scorer results. See #74.
  • Built-in model adapters reject unsupported call-time options with TypeError instead of silently ignoring them. See #65.
  • Retrieved context is kept out of the trusted system channel by the built-in adapters. CallableModel forwards context verbatim to the supplied function, so the caller is responsible for maintaining that boundary. Injection-resistance results depend on how that function places context. See the adapter contract.
  • Package imports and generated reports now share one version source. See #54.

Known limitations

  • Non-finite score normalization remains unresolved (#38). ScoreNormalizer.from_scale() can turn NaN into 1.0, allowing malformed judge output to become a passing score. This release does not include a fix; a passing aggregate alone does not establish that all judge outputs were finite.
  • A refusal anywhere in a red-team response can hide a separate affirmative success signal, overstating resistance in mixed responses. The shared evaluator follow-up is tracked in #76.
  • Structured separation of evaluated row data from LLM judge instructions remains pending in #62.
  • KeywordToxicityScorer still uses substring matching, which can flag benign words containing a configured keyword. The proposed correction remains under review in #73.

Install

Install this tagged release from GitHub:

python -m pip install "git+https://github.com/wandb/rai-toolkit.git@v0.3.0"

Python 3.10 or newer is required. Optional integrations use the extras listed in the installation guide. This is a GitHub release; it is not a PyPI publication.

Verification

All seven CI jobs passed on the merged release commit. The release PR checks covered Python 3.10, 3.11, and 3.12 (332 tests passed, 15 optional-integration tests skipped in each core environment) and both Weave environments (346 passed, 3 skipped each), plus package installation, lint, licensing, and whitespace checks.

The attached wheel and source archive match the tested source tree. The wheel also passed a separate clean-install check with matching package/import versions and compatible dependencies. SHA256SUMS contains checksums for both attached packages.

Community contributors

Thank you to the 16 outside contributors whose work merged between v0.2.0 and v0.3.0, listed alphabetically by handle:

Contributor Shipped contributions
@a0927929980-dot Groundedness handling for behavioral refusals (#17)
@abhnvgrg Model adapter contract and conformance suite (#61)
@adity982 Evidence-backed groundedness scorer (#13)
@denis-samatov HR industry preset (#24)
@dvd233 Explicit adapter call-time options (#65)
@EffNine Configured scorer names in Weave (#33)
@KunyangZhang Unassessed results for unsupported output formats (#74)
@M4h1m4 Vendor-neutral CallableModel adapter (#64)
@nightcityblade Normalized groundedness evidence matching (#18)
@sansynx Example validation and offline CLI tests (#46, #47)
@schallten Anthropic Messages API adapter (#56)
@Srijan229 Shared package and report version source (#54)
@TheJhyeFactor OpenAI-compatible adapter tests (#48)
@TrueFurina Four red-team attack templates (#11)
@YusefSyed Judge categories, additional Weave scorers, async and unassessed composite results, and preset policy isolation (#30, #31, #51, #52)
@Zeming-Yuan Retrieval relevance, context precision, and context recall scorers (#37, #58)

Maintained by @knisar. See CONTRIBUTORS.md for the broader contributor record.

Full changelog: v0.2.0...v0.3.0