Skip to content

Repository files navigation

ElectroBench

Evaluating frontier AI models on code-graded electrical engineering simulation tasks.

Portfolio hub: https://pvpprojects.netlify.app/electrobench

Blind evaluation (frozen)

Indicator Result
Frozen blind tasks 12
Attempts per condition 5
Blind baseline 18.3%
With official documentation 50.0%
Documentation + repair 60.0%

Technical Case Study

This repository hosts the academic-style Technical Case Study and reproducibility evidence for ElectroBench.

Classification: portfolio technical case study (academic format). Not a peer-reviewed publication.

Domains and tools

  • Power systems — pandapower
  • Control systems — python-control
  • Circuit simulation — SciPy-based grading; ngspice CLI integration smoke-tested in Docker

Precise ngspice wording

Circuit tasks are SciPy-graded. Ngspice is validated by a command-line integration smoke test; it is not claimed as the candidate grading backend.

Author

Pedro Vitor Brunello Pagliarin — Electrical Engineer | Python | AI Model Evaluation

About

ElectroBench — reproducible code-graded benchmark for LLM-generated electrical engineering simulations

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages