Elliot "Li" Bearden (he/him), an AI safety research engineer. I'm an MSCS candidate at Grand Canyon University, where my capstone is on run-level reproducibility of LLM sycophancy benchmarks. Before this I spent five years as an Applied AI Engineer at Deepgram, building evaluation infrastructure, data pipelines, and 150+ enterprise ML deployments.
A replication audit of published LLM sycophancy benchmarks — specifically, whether single-run confidence intervals understate true measurement uncertainty. The PARROT phase is complete and published: K=5 across 5 models and 1,302 items, run through a deterministic offline replay harness. The SycEval phase is in progress. Capstone defense is July 2026; arXiv preprint targeting August 2026.
Python, PyTorch, FastAPI, vLLM, Docker, Kubernetes, W&B. LLM evaluation, sycophancy benchmarking, statistical reproducibility, ML infrastructure.
Site: libearden.dev



