Evaluating frontier AI models on code-graded electrical engineering simulation tasks.
Portfolio hub: https://pvpprojects.netlify.app/electrobench
| Indicator | Result |
|---|---|
| Frozen blind tasks | 12 |
| Attempts per condition | 5 |
| Blind baseline | 18.3% |
| With official documentation | 50.0% |
| Documentation + repair | 60.0% |
This repository hosts the academic-style Technical Case Study and reproducibility evidence for ElectroBench.
Classification: portfolio technical case study (academic format). Not a peer-reviewed publication.
- PDF: Pedro_Vitor_ElectroBench_Technical_Case_Study.pdf
- Markdown companion: CASE_STUDY.md
- Release: ElectroBench v1.0 — Frozen Blind Evaluation
- Power systems —
pandapower - Control systems —
python-control - Circuit simulation — SciPy-based grading; ngspice CLI integration smoke-tested in Docker
Circuit tasks are SciPy-graded. Ngspice is validated by a command-line integration smoke test; it is not claimed as the candidate grading backend.
Pedro Vitor Brunello Pagliarin — Electrical Engineer | Python | AI Model Evaluation