Product Manager & Builder. I build products for people, and people are not simple.
Evidence before roadmap. Evals before demos. Monitoring after launch.
- orderbuddy-eval — an end-to-end LLM evaluation pipeline: 10k synthetic support tickets, intent discovery with MiniLM + HDBSCAN, validation against Banking77 and CLINC150, and a three-tier Claude bake-off with binary safety and groundedness gates plus cost per 1,000 classifications. The headline finding was that the model tier barely mattered and one ambiguous line in my prompt mattered a lot: budget-tier safety failures dropped from 48 to 5 after a one-sentence fix.
- Bharat 2047 — a prototype of four civic systems: blockchain voting, an AI grievance desk, auditable banking, and verifiable credentials.
- Steerly — AI career discovery for Indian students.
- The BackBench — an independent publication I research, write and build solo.
Founded and ran SkillAid (2020–2022) end to end — raised $110K, reached $150K+ revenue, then led a responsible wind-down. Amazon (2025). Growth at NNIIT (2023–2024). I mentor 70+ MBA aspirants and 20+ early-stage founders.
I care about the gap between a demo and a system somebody can rely on. That usually means golden sets, written rubrics, gates that fail loudly, and cost math that survives contact with a finance team. When an eval disagrees with what I hoped, the eval wins and I publish it.
- Website — pawanchander.com
- Writing — The eval practice