ModelGauge

Goal: Make it easy to automatically and uniformly measure the behavior of many AI Systems.

Warning

This repo is still in beta with a planned full release in Fall 2024. Until then we reserve the right to make backward incompatible changes as needed.

ModelGauge is an evolution of crfm-helm, intended to meet their existing use cases as well as those needed by the MLCommons AI Safety project.

Summary

ModelGauge is a library that provides a set of interfaces for Tests and Systems Under Test (SUTs) such that:

Each Test can be applied to all SUTs with the required underlying capabilities (e.g. does it take text input?)
Adding new Tests or SUTs can be done without modifications to the core libraries or support from ModelGauge authors.

Currently ModelGauge is targeted at LLMs and single turn prompt response Tests, with Tests scored by automated Annotators (e.g. LlamaGuard). However, we expect to extend the library to cover more Test, SUT, and Annotation types as we move toward full release.

Docs

Developer Quick Start
Tutorial for how to create a Test
Tutorial for how to create a System Under Test (SUT)
How we use plugins to connect it all together.

Name		Name	Last commit message	Last commit date
Latest commit History 321 Commits
.github		.github
demo_plugin		demo_plugin
docs		docs
modelgauge		modelgauge
plugins		plugins
tests		tests
.gitignore		.gitignore
.readthedocs.yaml		.readthedocs.yaml
CONTRIBUTING.md		CONTRIBUTING.md
LICENSE.md		LICENSE.md
README.md		README.md
conftest.py		conftest.py
mkdocs.yml		mkdocs.yml
poetry.lock		poetry.lock
publish_all.py		publish_all.py
pyproject.toml		pyproject.toml
tox.ini		tox.ini

License

mlcommons/modelgauge

Folders and files

Latest commit

History

Repository files navigation

ModelGauge

Summary

Docs

About

Resources

License

Stars

Watchers

Forks

Languages