VERA-MH v1.1.0
VERA-MH v1.1.0 — Release notes
VERA-MH helps teams simulate and evaluate their chatbot's mental-health conversations against a safety rubric. Version 1.1 expands on personas to simulate, refines how we score safety and care, and makes large evaluation runs easier to run, resume, and audit.
Latest Scores
What’s new
Richer, more realistic simulations
We increased the persona library from 10 to 100 personas, spanning a wider range of situations and risk levels. We also updated how the simulated user is instructed to behave, including a clearer emphasis on staying in the role of a real person seeking help, not a counselor, so conversations better reflect how people actually show up in a chat.
Clearer safety scoring
The evaluation rubric was updated based on feedback from many external stakeholders and clinicians.
In practice, examples of this include:
- Guides to human care scores consider context more carefully—for example, whether the person is already connected to crisis support, and whether the person is experiencing suicidal urges during the conversation.
- High potential for harm is distinguished more clearly from suboptimal responses (for example, omitting crisis resource information as high potential for harm versus not fully addressing barriers to using resources as suboptimal).
- Dimensions are less coupled: a serious miss in one area no longer automatically forces a miss of the same severity in another related dimension.
Because of the updates to the rubric and personas, aggregate scores are not directly comparable to v1.0. On average, general model scores may score slightly higher (typically on the order of a few points) than on the previous version.
More reliable long runs
Calls to AI solutions now use retries and timeouts by default (up to three retries, with a short wait between attempts). If a single conversation or evaluation fails, the run can continue instead of stopping the whole batch. Testers/developers can also resume an interrupted simulation or judging run so you do not redo work that has already finished.
Clearer outputs for review
Judge activity is logged in a predictable, per-job layout (one log per conversation, judge model, and evaluation instance). The README describes a default layout that keeps each generation run’s transcripts, nested evaluations/ batches, scores, and related artifacts under a single parent (typically output/ with timestamped p_* run folders) so outputs are easier to find, resume, and share.
Reliable VERA Scoring
The README now describes recommended settings for a stable, comparable headline score (2 conversations per each of the 100 personas, dual user-agent simulators, 30 turns, dual judges, and optional pooling), and includes a helper script scripts/run_recommended_vera_pipeline.sh to run conversations, evaluations, scoring, and obtaining the overall VERA score for the provided chatbot.
Where to learn more
- Getting started and recommended settings: repository README
- Custom providers (private APIs or unsupported models): README section Connecting your own LLM, Agent, or API and docs/evaluating.md
- Research context: links in the README Additional Resources section (including the reliability/validity preprint)
- General VERA-MH information: visit the website at VERA-MH.com
For a concise, technical list of changes, see CHANGELOG.md in the repository root.