Nanotune v1.7.0: benchmarks are reproducible, chat streams, training is flag-driven #106
will-lamerton
announced in
Annoucements
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Built by the Nano Collective, a community collective building AI tooling not for profit, but for the community.
v1.7.0 is the largest Nanotune release since 1.0. It lands in three areas that came up again and again in issues: the trustworthiness of benchmark numbers, the responsiveness of
nanotune chat, and the friction of editingconfig.jsonby hand to tune a run. Every change below ships in this version, onnpm, today.Benchmarks are reproducible by default
Up to 1.6.x,
nanotune benchmarkdefaulted totemperature: 0.8with no seed and a single sample per test. Running the same suite twice against the same model produced different pass rates. For a tool whose output is a score, every number was ambiguous: a 3-point move could be a real regression, or just sampling noise.Benchmarks now default to
temperature: 0with a fixed seed of42. Two consecutive runs with default flags produce identical results. If you want the old sampling behaviour, it is one flag:Existing numbers will move. Scores recorded before this release were sampled at temperature 0.8; re-running the same suite will produce a different, and from now on stable, figure. Re-baseline any pass rates you are tracking.
--samples <n>A new flag runs each test
ntimes and records the per-test pass rate and variance in both reports, instead of reporting a single coin flip:A test passes when the majority of its samples pass. Each sample uses
seed + sample_index, so runs stay reproducible while the samples differ from one another.--temperature,--seed, and--samplesnow reject a value they cannot parse instead of falling back to the default.--samples 5xused to run a single sample; it now fails with a message naming the flag, so a score is never reported under settings you did not ask for.nanotune benchmark compareA score on its own does not tell you whether fine-tuning helped.
comparediffs two saved runs and reports what moved, per test and per category:nanotune benchmark compare # the two most recent runs nanotune benchmark compare before.json after.jsonWith one argument, it compares that run against the latest. With none, it picks the two most recent.
nanotune benchmark --baseBenchmarks the base model, before any fine-tuning, as a control. This is the number your fine-tuned score is only meaningful against:
The base model has to be downloaded, converted, and quantized before it can be benchmarked, which is slow. The result is cached under
~/.nanotune/models/base-cache/, keyed by model id and quantization, so every project fine-tuning from the same base model pays that cost once. The cache is written to a temp file and renamed into place, so a run killed mid-quantize cannot leave a corrupt GGUF that later runs treat as valid.Saved reports record whether the run was a base-model control, along with the temperature, seed, and sample count used.
Chat streams as it generates
nanotune chatwaited for the entire reply before printing anything, so a long answer looked like a hung terminal. Tokens now stream in over SSE as the model produces them.Press
Escto cancel a generation in flight. Whatever had been generated is kept and stays in the conversation history, so the transcript on screen and the history the model sees never disagree. Cancelling a turn that produced nothing rolls it back entirely./saveand/keepThe best training examples tend to surface during a chat, and there was no way to capture them without retyping into
data add./savewrites to.nanotune/chats/by default and refuses to overwrite an existing file unless you pass--force, since a transcript cannot be recovered once it is gone.Training hyperparameters are settable from the CLI
Tuning a run meant hand-editing
config.json. Every training hyperparameter now has a flag:New flags:
--batch-size,--num-layers,--steps-per-eval,--save-every,--fine-tune-type(lora,dora, orfull),--lora-rank,--lora-alpha,--lora-dropout,--max-seq-length,--grad-checkpoint/--no-grad-checkpoint,--val-batches, and--train-seed.Flags override
config.json; anything you do not pass keeps its configured value. Values are validated against the config schema before the run starts, so a bad number fails immediately with a message naming the flag you typed rather than dying insidemlx_lmminutes later.--fine-tune-type fullskips the LoRA parameter block entirely.Ctrl+C stops training gracefully
Ctrl+C used to tear the app down while MLX was still writing, and the "checkpoint saved" hint it printed was never verified. Training now stops on a signal, MLX flushes its checkpoint, and the summary reports the actual last-saved iteration and how to resume:
A second Ctrl+C gives up on the checkpoint rather than trapping you behind a trainer that will not exit. Interrupted runs exit
130.The train/validation split is visible
ensureValidationSet()silently moved ~10% oftrain.jsonlintovalid.jsonlon the first run, so the training-example count dropped with no explanation. The split is now reported when it happens, and both counts are shown on every run.--seed <n>makes the split reproducible. Because the split only happens when no validation set exists yet, passing a seed to an already-split project now says so instead of letting you believe the split was re-rolled.Working with the validation set
Every
datasubcommand takes-e, --evalto operate onvalid.jsonlinstead oftrain.jsonl:nanotune data add --eval nanotune data list --eval nanotune data import extra.jsonl --eval nanotune data validate --eval nanotune data export valid-backup.jsonl --evaldata validateno longer warns that a validation set has fewer than 50 examples. A validation set is a slice of the training data, so that floor never applied to it and the warning was noise on a correct split.nanotune data exportExports training data to JSONL, CSV, or JSON, chosen by the output file's extension:
nanotune data export backup.jsonlJSONL and JSON round-trip exactly: re-importing the output reproduces the original data, including multi-turn conversations and per-example context messages. CSV cannot represent a multi-turn example, so those are skipped rather than silently truncated, and reported in the summary. Existing files prompt before being overwritten;
--yesskips the prompt for scripts and CI.Editing and repairing training data
data listgainseto edit the selected example in place. Only the first user/assistant turn is editable; every other message is preserved untouched, and the count of preserved turns is shown so a multi-turn example cannot be quietly flattened.data validategains two repair flags:--fixonly removes examples that are identical across every message. This is deliberately stricter than the duplicate warning, which compares first user messages only: two examples sharing an input but differing in output are real data.--rewrite-contextfixes examples whose context message has drifted from the current config, and never inserts one where an example did not have it. When both are passed, the rewrite runs first so examples that only become identical after normalization are still caught.Config typos are reported
ConfigSchemasilently dropped unknown keys, soloraLayersinstead ofnumLayerswas discarded with no warning and the default used instead. Unknown keys are now reported with a suggestion:Nested objects are checked too. An invalid
config.jsonnow fails with the offending path and what was wrong with it, instead of the raw Zod issue array serialized as a wall of JSON.Fixes
data list --evalcorrupted training data. Editing a validation example wrote totrain.jsonlat the same index, overwriting a real training example and then displaying training data in the validation view. Silent and unrecoverable."explain the input parameter","..."was swallowed as a header. Detection now requires the row to be exactly the two column names.normalizeTextcanonicalized",', and`to a single character, so a test expecting one quote style passed on any other. Quote characters are now kept distinct.llama-serverdying could kill the CLI. On Node 22 an unhandled rejection from a killed child process is fatal, so a server that died mid-run took the whole CLI down with a raw stack dump. The handler is now attached at spawn.nanotune exportprogress jumped. Each sub-step's progress is now mapped onto its slice of the overall export.data liststranded you on an empty page. Pagination now steps back.judge.jsoncould briefly exist with loose permissions. It is now written to a fresh0600temp file and renamed into place, which is atomic and carries the mode with it.train --save-everywas ignored in the stop summary. The last-checkpoint iteration was computed fromconfig.jsonrather than the value the run actually used; it now reports what was actually saved.Object.prototypekeys were not reported as unknown config keys. A key likeconstructorinconfig.jsonpassed the known-key check via the prototype chain.--baseruns could delete each other's work. The stale-artifact sweep removed every.tmp-*.ggufin the cache directory, including one belonging to a running export. It now only sweeps files whose owning process is gone.filesallowlist was trimmed, and spec files are excluded fromdist.Documentation
benchmark compare,benchmark --base,data export, chat streaming,/saveand/keep, the training hyperparameter flags, and the--evalflag across the data commands.docs/testing-guide.mdcovering the Ink component test harness.judge.jsonneeds both a file mode and an explicitchmod.Internals
src/commands/previously had no test harness, which is why several of the bugs above shipped. Commands are now rendered throughink-testing-libraryand driven with real keypresses.knipnow enforcesexports: "error". Dead exports (runGGUFInference,parseLlamaCppStderr,updateExample) were removed rather than kept on the assumption a caller existed..tsxand into testable helpers:benchmark-compare.ts,benchmark-utils.ts,model-cache.ts,chat-helpers.ts, plusbuildTrainingArgs,mergeEditedTurn,clampPagination, andscaleProgress.Toolchain
ai7.0.77,@ai-sdk/anthropic4.0.36,commander15.0.0,execa10.0.1,typescript7.0.2,c812.0.0,@biomejs/biome2.5.7,@types/node26.2.0,reactand@types/react.$schemato the installed Biome version.Install
Full source, issues, and the changelog are on the project repo at github.com/Nano-Collective/nanotune. Every one of this release's contributors is a first-time Nanotune contributor; thank you to @addyCooks, @rohanshrma222, @akramcodez, @yashksaini-coder, @Aryagarg23, and @floze-the-genius.
All reactions