Skip to content

VERA-MH v1.1.0

Choose a tag to compare

@jgieringer jgieringer released this 29 Apr 20:39
· 367 commits to main since this release
dc9780e

VERA-MH v1.1.0 — Release notes

VERA-MH helps teams simulate and evaluate their chatbot's mental-health conversations against a safety rubric. Version 1.1 expands on personas to simulate, refines how we score safety and care, and makes large evaluation runs easier to run, resume, and audit.

Latest Scores

scores_20260428_122725

What’s new

Richer, more realistic simulations

We increased the persona library from 10 to 100 personas, spanning a wider range of situations and risk levels. We also updated how the simulated user is instructed to behave, including a clearer emphasis on staying in the role of a real person seeking help, not a counselor, so conversations better reflect how people actually show up in a chat.

Clearer safety scoring

The evaluation rubric was updated based on feedback from many external stakeholders and clinicians.

In practice, examples of this include:

  • Guides to human care scores consider context more carefully—for example, whether the person is already connected to crisis support, and whether the person is experiencing suicidal urges during the conversation.
  • High potential for harm is distinguished more clearly from suboptimal responses (for example, omitting crisis resource information as high potential for harm versus not fully addressing barriers to using resources as suboptimal).
  • Dimensions are less coupled: a serious miss in one area no longer automatically forces a miss of the same severity in another related dimension.

Because of the updates to the rubric and personas, aggregate scores are not directly comparable to v1.0. On average, general model scores may score slightly higher (typically on the order of a few points) than on the previous version.

More reliable long runs

Calls to AI solutions now use retries and timeouts by default (up to three retries, with a short wait between attempts). If a single conversation or evaluation fails, the run can continue instead of stopping the whole batch. Testers/developers can also resume an interrupted simulation or judging run so you do not redo work that has already finished.

Clearer outputs for review

Judge activity is logged in a predictable, per-job layout (one log per conversation, judge model, and evaluation instance). The README describes a default layout that keeps each generation run’s transcripts, nested evaluations/ batches, scores, and related artifacts under a single parent (typically output/ with timestamped p_* run folders) so outputs are easier to find, resume, and share.

Reliable VERA Scoring

The README now describes recommended settings for a stable, comparable headline score (2 conversations per each of the 100 personas, dual user-agent simulators, 30 turns, dual judges, and optional pooling), and includes a helper script scripts/run_recommended_vera_pipeline.sh to run conversations, evaluations, scoring, and obtaining the overall VERA score for the provided chatbot.

Where to learn more

  • Getting started and recommended settings: repository README
  • Custom providers (private APIs or unsupported models): README section Connecting your own LLM, Agent, or API and docs/evaluating.md
  • Research context: links in the README Additional Resources section (including the reliability/validity preprint)
  • General VERA-MH information: visit the website at VERA-MH.com

For a concise, technical list of changes, see CHANGELOG.md in the repository root.