Skip to content
View LiBearden's full-sized avatar
🔧
Exploring new tools and libraries!
🔧
Exploring new tools and libraries!

Highlights

  • Pro

Block or report LiBearden

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
libearden/README.md

Who I am

Elliot "Li" Bearden (he/him), an AI safety research engineer. I'm an MSCS candidate at Grand Canyon University, where my capstone is on run-level reproducibility of LLM sycophancy benchmarks. Before this I spent five years as an Applied AI Engineer at Deepgram, building evaluation infrastructure, data pipelines, and 150+ enterprise ML deployments.

What I'm working on

A replication audit of published LLM sycophancy benchmarks — specifically, whether single-run confidence intervals understate true measurement uncertainty. The PARROT phase is complete and published: K=5 across 5 models and 1,302 items, run through a deterministic offline replay harness. The SycEval phase is in progress. Capstone defense is July 2026; arXiv preprint targeting August 2026.

Stack

Python, PyTorch, FastAPI, vLLM, Docker, Kubernetes, W&B. LLM evaluation, sycophancy benchmarking, statistical reproducibility, ML infrastructure.

Reach me

Site: libearden.dev

Pinned Loading

  1. sycophancy-benchmark-reliability sycophancy-benchmark-reliability Public

    Run-level reproducibility audit of the PARROT sycophancy benchmark

    Python