Software Developer · AI Evaluation & Data Quality — based in Maharashtra, India.
Full-stack developer with an MCA and 3+ years of dedicated Python development, now focused on large-language-model evaluation — designing tasks and tests that reveal where AI coding agents fail. I build in Python, .NET, and SQL, and I write about it.
🔗 Portfolio · ✍️ Technical Writing · 💼 LinkedIn · ✉️ zahid23saim@gmail.com
Plus: LLM evaluation & benchmark authoring, factuality & STEM auditing, test writing, database optimization.
| Project | What it is | Built with |
|---|---|---|
| llm-eval-harness | A dependency-free harness for scoring LLM answers against a gold set — per-question match rules, validation, CI-friendly. Live demo → | Python · pytest |
| aspnet-minimal-api | A tested ASP.NET Core (.NET 8) task API — full CRUD, validation, in-memory integration tests (7/7 passing). | C# · ASP.NET Core · xUnit |
| plsql-bulk-examples | Turning a slow row-by-row Oracle load into fast batched BULK COLLECT + FORALL, with SAVE EXCEPTIONS. |
Oracle PL/SQL |
- Automating LLM Answer Evaluation with a Small Python Scoring Script
- Speeding Up a Slow PL/SQL Routine with BULK COLLECT and FORALL
- Building a Tested ASP.NET Core Minimal API in One File
- What I Learned Evaluating LLMs Across Four Languages
🌐 English · Hindi · Tamil · Marathi | Open to remote, project-based software & AI-evaluation work.