Eval protocol for a 0.5B coder LoRA (no LLM-as-judge in serve) #13
YauhenBichel
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Constraint
PythonVibeGuardis deterministic. We will not add a second model inscripts/serve.pyto "score" drafts. Eval can live indocs/or a script that never ships in the sidecar.What we do not have yet
all_pairs())pass, lengthqwen2.5-coder:0.5bwith the same system prompt, no LoRAPropose a protocol
Keep it cheap: 10–20 prompts, Mac or Linux, Ollama optional. If you need a GPU cluster, say so, but the default should run on a laptop.
Related: #8 (guard evasion), #9 (data vs style prior).
All reactions