Skip to content

lcode 0.3.1

Choose a tag to compare

@nasser1941 nasser1941 released this 30 Sep 17:36
968ba99

Which local model codes best on your machine? lcode bench measures it.

Added

  • lcode bench scores models on this machine with eight small coding tasks:

    • fixing a bug
    • finding code
    • writing a script from a spec
    • renaming across files
    • a one-line edit in a long file
    • recovering from failing commands
    • fixing a function to match its spec
    • adding a feature across files

    Each task runs in a temporary folder and has an automatic check, some with hidden tests. It reports tasks passed, time, generation and prompt speed, tool-call errors and memory use. lcode bench model-a model-b compares models, --json saves the results (a documented, versioned format) and --markdown prints a table for a model test report.

lcode bench qwen3.6-35b qwen3.5-9b --markdown

Results on an RTX 4080 Laptop GPU (12 GB) with 31 GB of RAM, 32K context:

qwen3.6-35b qwen3.5-9b qwen3.5-4b
Passed 8/8 6/8 6/8
Time 3m 57s 11m 55s 5m 20s
Generation 71.4 tok/s 60.1 tok/s 91.3 tok/s
Prompt reading 855 tok/s 3,507 tok/s 5,186 tok/s
Tool-call errors 0 9 1

Please share yours, especially from Apple Silicon Macs and other GPUs: paste the --markdown table into a model test report. Details: Benchmarking models.

Upgrade

uv tool upgrade lcode-cli        # or: pipx upgrade lcode-cli

Full changelog: v0.3.0...v0.3.1