Augmentation of ChatGPT with clinician-informed tools improves performance on medical calculation tasks

Abstract: Prior work has shown that large language models (LLMs) have the ability to answer expert-level multiple choice questions in medicine, but are limited by both their tendency to hallucinate knowledge and their inherent inadequacy in performing basic mathematical operations. Unsurprisingly, early evidence suggests that LLMs perform poorly when asked to execute common clinical calculations. Recently, it has been demonstrated that LLMs have the capability of interacting with external programs and tools, presenting a possible remedy for this limitation. In this study, we explore the ability of ChatGPT (GPT-4, November 2023) to perform medical calculations, evaluating its performance across 48 diverse clinical calculation tasks. Our findings indicate that ChatGPT is an unreliable clinical calculator, delivering inaccurate responses in one-third of trials (n=212). To address this, we developed an open-source clinical calculation API (openmedcalc.org), which we then integrated with ChatGPT. We subsequently evaluated the performance of this augmented model by comparing it against standard ChatGPT using 75 clinical vignettes in three common clinical calculation tasks: Caprini VTE Risk, Wells DVT Criteria, and MELD-Na. The augmented model demonstrated a marked improvement in accuracy over unimproved ChatGPT. Our findings suggest that integration of machine-usable, clinician-informed tools can help alleviate the reliability limitations observed in medical LLMs.

Find our preprint on medrXiv.

Repo structure

├── LICENSE
├── README.md
└── results-data                       
    ├── calculators-list.csv            - list of calculators by calculator type
    ├── exploratory-analysis.csv        - results from exploratory analysis
    └── focused-analysis-results.csv    - results from focused analysis (MELD, 
                                          Caprini, Wells DVT)

Name		Name	Last commit message	Last commit date
Latest commit History 6 Commits
results-data		results-data
.gitignore		.gitignore
LICENSE		LICENSE
README.md		README.md

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

results-data

results-data

.gitignore

.gitignore

LICENSE

LICENSE

README.md

README.md

Repository files navigation

Augmentation of ChatGPT with clinician-informed tools improves performance on medical calculation tasks

Repo structure

About

Releases 2

Packages

License

stanfordaimlab/llm-as-clinical-calculator

Folders and files

Latest commit

History

Repository files navigation

Augmentation of ChatGPT with clinician-informed tools improves performance on medical calculation tasks

Repo structure

About

Resources

License

Stars

Watchers

Forks