lcode 0.3.1
Which local model codes best on your machine? lcode bench measures it.
Added
-
lcode benchscores models on this machine with eight small coding tasks:- fixing a bug
- finding code
- writing a script from a spec
- renaming across files
- a one-line edit in a long file
- recovering from failing commands
- fixing a function to match its spec
- adding a feature across files
Each task runs in a temporary folder and has an automatic check, some with hidden tests. It reports tasks passed, time, generation and prompt speed, tool-call errors and memory use.
lcode bench model-a model-bcompares models,--jsonsaves the results (a documented, versioned format) and--markdownprints a table for a model test report.
lcode bench qwen3.6-35b qwen3.5-9b --markdownResults on an RTX 4080 Laptop GPU (12 GB) with 31 GB of RAM, 32K context:
qwen3.6-35b |
qwen3.5-9b |
qwen3.5-4b |
|
|---|---|---|---|
| Passed | 8/8 | 6/8 | 6/8 |
| Time | 3m 57s | 11m 55s | 5m 20s |
| Generation | 71.4 tok/s | 60.1 tok/s | 91.3 tok/s |
| Prompt reading | 855 tok/s | 3,507 tok/s | 5,186 tok/s |
| Tool-call errors | 0 | 9 | 1 |
Please share yours, especially from Apple Silicon Macs and other GPUs: paste the --markdown table into a model test report. Details: Benchmarking models.
Upgrade
uv tool upgrade lcode-cli # or: pipx upgrade lcode-cliFull changelog: v0.3.0...v0.3.1