Skip to content

0.7.0 — decisions on a recurrent hybrid reuse the state, CoreAI.systemOne and 255 options

Choose a tag to compare

@github-actions github-actions released this 24 Sep 02:08
· 61 commits to main since this release
1592a1e

Decisions on a recurrent hybrid reuse the state, the hosted System One call is one op (CoreAI.systemOne), and the runtime pin moves to coreai-models 0.2.7-zoo. exact: "0.6.0" resolvers move to 0.7.0 (a minor: the catalog id apus-openjev-v1-4b is renamed, and a request that repeats a question id is refused).

Start here: installation and smallest working example · System One, on device · Examples/Decide

brew install john-rocky/tap/systemone && systemone serve     # http://127.0.0.1:8090/v1/systemone
brew upgrade systemone                                        # from 0.6.0
.package(url: "https://github.com/john-rocky/coreai-kit", exact: "0.7.0")

Added

  • CoreAI.systemOne — a /v1/systemone request in, its response out, as one op. The response holds the typed answers in request order, usage, stateTokens, milliseconds, and the wire object byte for byte as systemone serve returns it. TypedDecisions.systemOne(_:) is the model-level call. On the zoo's 13-request fixture, all 36 decider-0.8b answers name the author's option, every number within 0.02 of the author's fp32 assembly.
  • A choice lists up to 255 options, the hosted API's width, where the model can read that many: minicpm5-2b reads numbers past 26 options, decider-0.8b its author's labels A–Z, AA, AB, …, and system-one-scorer-4b scores 255 rows. TypedDecisions.maxOptions is each model's own count.
  • Three catalog decision models. laya-multilingual is an encoder read at its option markers (Decision.Format.encoder): 201/201 fixture rows on the iPhone 17 Pro, 47 ms per decision. qwen3.5-2b-decision is argmax-identical to the fp32 reference on 58/58 rows. system-one-scorer-4b has a scalar scoring head (Decision.Format.scalar): 48/48, CC BY-NC 4.0, Mac only. CatalogEntry.license names a restrictive weights license, and Decision.Format.letterList reads a lettered option list (up to 52).
  • Catalog-fitted temperatures: CatalogEntry.calibration, DecisionCalibration, decide-cli calibrate. minicpm5-2b on SemIf's authored144: ECE 0.167 → 0.071, accuracy 0.701 either way.
  • Token-level scoring on TypedDecisions — logits(for:), prefill(tokens:), tokenizer — for a caller with its own readout.
  • Examples/Decide/conformance/check.py <base_url>: 22 requests in the hosted forms and the shape each answer must come back in, for any server that speaks the route.

Changed

  • Decisions on a recurrent hybrid reuse the state. TypedDecisions checkpoints the state after its prefix with coreai-models 0.2.7-zoo's InferenceEngine.checkpoint(); each later question prefills only its own tokens. Eight questions on one state, M4 Max, per decision, checkpointed vs from scratch: 186 vs 825 ms (decider-0.8b), 120 vs 696 ms (openthai-systemone), 220 vs 1,297 ms (qwen3.5-2b-decision), 1,211 vs 3,081 ms (apus-decision-v1-4b), 719 vs 5,851 ms (system-one-scorer-4b). The answers are bit-identical to the from-scratch path on the checked rows.
  • coreai-models 0.2.4-zoo → 0.2.7-zoo. The move also brings 0.2.5-zoo: the engine stops at a stop sequence instead of decoding to maxTokens.
  • apus-openjev-v1-4b is now apus-decision-v1-4b: catalog id, display name and Hub repo (the old repo URL redirects).
  • A /v1/systemone request that keys two questions by one id gets a 422; before, the second answer overwrote the first. SystemOne.maxOptions is 255 (was 16).
  • minicpm5-2b, minicpm5-1b and qwen3-0.6b read decisions at their catalog temperature: new probabilities, the same answers.

Removed

  • openjev-27b, withdrawn (its Hub repo is now private) pending the source model's training-data provenance. It entered the catalog after 0.6.0, so no tag carries it.

Fixed

  • Concurrent decide or logits(for:) calls on one TypedDecisions could drive its engine at once and crash in CoreAISequentialEngine or come back without logits. Engine calls now wait their turn in a FIFO async lock; a task cancelled while it waits throws CancellationError (#43, @mattt).
  • decider-0.8b read a described option twice over the wire (repair: repair: …), and P(repair) on the fixture's one described choice came out 0.941 against the author's 0.971. The chosen option did not change.

Verified environment

  • Mac: M4 Max, macOS 27.0 26A428, Xcode 27 27A266a. CI: XCTest 86 (17 skipped, 0 failures), Swift Testing 211 in 43 suites. The ChatDemo entry-check on this release's tree passes chat (from an empty cache), fm, japanese, japanese-fm and hybrid with the 0.4.1 validation record's outputs; the README's chat-cli command answers Tokyo.
  • systemone-0.7.0-macos-arm64.zip: Developer ID signed and notarized; the notary ticket names this binary's cdhash. From the tap, brew upgrade systemone moves 0.6.0 to 0.7.0 in 3.4 s and brew test passes. systemone serve answers /health 200 in 3.4 s (weights in the file cache), and Examples/Decide/clients/systemone.sh gets the answer printed in Examples/Decide/README.md (billing 0.5078) in 0.27 s.
  • Not on the phone: the checkpoint is not measured on the iPhone yet, and the phone rows in the docs predate it.