Repository navigation
rai-toolkit v0.3.0 brings new RAG scorers, more model adapters, and fixes to assessment and red-team reporting. It includes 38 merged commits since v0.2.0, plus release preparation, with contributions from 16 community developers.
Highlights
- Evidence-backed RAG evaluation:
GroundednessScorer,RetrievalRelevanceScorer,ContextPrecisionScorer, andContextRecallScorercheck grounding, retrieval relevance, and context quality. Evidence validation, normalized span matching, and behavioral-refusal handling improve the grounding results. See #13, #17, #18, #37, and #58. - More ways to connect models: an Anthropic Messages API adapter and vendor-neutral
CallableModeljoin the existing adapters. A shared adapter contract and conformance suite document and check the expected behavior. See #56, #61, and #64. - Assessment and integration fixes: composite scoring honors asynchronous scorers and unassessed results; Weave assessments preserve additional scorers and configured names; default judge categories are retained. See #30, #31, #33, and #51.
- Red-team reporting: four attack templates were added. Execution errors and missing evidence are kept out of resistance scores, and incomplete source coverage fails the coverage gate. See #11, #67, and the full change history.
- HR preset and policy isolation: the HR industry preset is available, and built-in policies stay scoped to their selected industry. See #24 and #52.
- Contributor and deployment documentation: new self-hosting, model-adapter, and scorer-authoring guidance; clearer contribution setup and claim flow; and permanent contributor credits. See the documentation, contributing guide, and contributors.
Behavior changes to review when upgrading
- Unsupported
OutputFormatScorerformats now return an unassessed result instead of a passing score. Consumers should checkScorerResult.assessedwhen interpreting scorer results. See #74. - Built-in model adapters reject unsupported call-time options with
TypeErrorinstead of silently ignoring them. See #65. - Retrieved context is kept out of the trusted system channel by the built-in adapters.
CallableModelforwards context verbatim to the supplied function, so the caller is responsible for maintaining that boundary. Injection-resistance results depend on how that function places context. See the adapter contract. - Package imports and generated reports now share one version source. See #54.
Known limitations
- Non-finite score normalization remains unresolved (#38).
ScoreNormalizer.from_scale()can turnNaNinto1.0, allowing malformed judge output to become a passing score. This release does not include a fix; a passing aggregate alone does not establish that all judge outputs were finite. - A refusal anywhere in a red-team response can hide a separate affirmative success signal, overstating resistance in mixed responses. The shared evaluator follow-up is tracked in #76.
- Structured separation of evaluated row data from LLM judge instructions remains pending in #62.
KeywordToxicityScorerstill uses substring matching, which can flag benign words containing a configured keyword. The proposed correction remains under review in #73.
Install
Install this tagged release from GitHub:
python -m pip install "git+https://github.com/wandb/rai-toolkit.git@v0.3.0"Python 3.10 or newer is required. Optional integrations use the extras listed in the installation guide. This is a GitHub release; it is not a PyPI publication.
Verification
All seven CI jobs passed on the merged release commit. The release PR checks covered Python 3.10, 3.11, and 3.12 (332 tests passed, 15 optional-integration tests skipped in each core environment) and both Weave environments (346 passed, 3 skipped each), plus package installation, lint, licensing, and whitespace checks.
The attached wheel and source archive match the tested source tree. The wheel also passed a separate clean-install check with matching package/import versions and compatible dependencies. SHA256SUMS contains checksums for both attached packages.
Community contributors
Thank you to the 16 outside contributors whose work merged between v0.2.0 and v0.3.0, listed alphabetically by handle:
| Contributor | Shipped contributions |
|---|---|
| @a0927929980-dot | Groundedness handling for behavioral refusals (#17) |
| @abhnvgrg | Model adapter contract and conformance suite (#61) |
| @adity982 | Evidence-backed groundedness scorer (#13) |
| @denis-samatov | HR industry preset (#24) |
| @dvd233 | Explicit adapter call-time options (#65) |
| @EffNine | Configured scorer names in Weave (#33) |
| @KunyangZhang | Unassessed results for unsupported output formats (#74) |
| @M4h1m4 | Vendor-neutral CallableModel adapter (#64) |
| @nightcityblade | Normalized groundedness evidence matching (#18) |
| @sansynx | Example validation and offline CLI tests (#46, #47) |
| @schallten | Anthropic Messages API adapter (#56) |
| @Srijan229 | Shared package and report version source (#54) |
| @TheJhyeFactor | OpenAI-compatible adapter tests (#48) |
| @TrueFurina | Four red-team attack templates (#11) |
| @YusefSyed | Judge categories, additional Weave scorers, async and unassessed composite results, and preset policy isolation (#30, #31, #51, #52) |
| @Zeming-Yuan | Retrieval relevance, context precision, and context recall scorers (#37, #58) |
Maintained by @knisar. See CONTRIBUTORS.md for the broader contributor record.
Full changelog: v0.2.0...v0.3.0