Releases: JS-PACKAGE/Neko.js
Release list
Neko.js v1.6.1
- Performance: on Node, each pinned ONNX payload file is hashed once per installed runtime instead of four times per load (prefetch plus three loader requests); unchanged files are recognised by device, inode, size, mtime and ctime, and new processes still verify every byte. Measured warm model load fell from about 2.3 s to about 1.3 s on macOS arm64. The cache no longer calls
chmodon files that already have mode0600, which would have advanced ctime. - Performance: Node ONNX sessions now use
os.availableParallelism()intra-op threads. On a 4 performance + 6 efficiency core Apple machine this raised prefill about 25% (about 213-238 to about 290 tokens/s) and lowered decode about 12% (about 28 to about 25 tokens/s); the best value depends on hardware. - Performance: structured calls reuse an engine-private decoder state for their fixed system instruction (text-only, at least 32 tokens, no explicit
reuse). Time to first token for a 115-token structured prompt fell from about 460-540 ms to about 130-160 ms on a hit, with identical output in the exercised cases.reuseCacheInfo()gainsprefixEntries,prefixBytes,prefixHitsandprefixMisses. - Performance: the constrained-decoding token trie uses typed arrays; building it for the 248,056-piece vocabulary took about 270 ms (was about 337 ms) and about 54 MB of external/array-buffer memory instead of about 230 MB of JavaScript heap.
Neko.js v1.6
- Fix (affects 1.5.0): a first-use model download from Hugging Face could not complete. Observed against the live Hub on 2026-10-03: non-LFS files redirect same-origin to
/api/resolve-cache/models/<id>/<revision>/<file>, which the default network policy denied (POLICY_DENIED), and interrupted CDN downloads could never resume because the signed redirect URL changes per request and was compared as part of the object location. The exact pinned model/revision/fileresolve-cachepath is now trusted (size/SHA-256 verification unchanged) and resume matches by origin and path, including partial staging written by 1.5.0. Verified by a complete live prefetch of the pinned default profile with every file SHA-256 checked. - Fix (CI, Windows): the 1.5.0 remote CI run failed 3 of 22 jobs, all
windows-2025, innpm run test:quality. The CLI test built its script path fromURL.pathname(D:\D:\a\...), and Git's CRLF conversion changed the byte-hashed quality fixtures (LOCAL_FIXTURE_INTEGRITY). The path now usesfileURLToPath, and.gitattributesmarksscripts/quality/fixtures/**as-text. The full remote CI matrix (19 of 22 jobs; the other 3 are opt-in real-model lanes that are skipped) passes on commitef501ed, including all threewindows-2025jobs. - Fix (Windows): Node image inputs such as
D:\photo.pngwere mistaken for URLs with a one-letter scheme and rejected as non-HTTP; drive-letter paths are now read as local paths (the same canonical file authorization applies). The cache permission test no longer asserts POSIX mode bits on Windows. - Add opt-in hybrid document retrieval: caller-supplied
DocumentEmbedder,DocumentIndex.searchHybridandaskDocuments({ embedder, embedding })combine BM25 with cosine ranking by reciprocal-rank fusion. Neko.js bundles no embedding model; embedder output is validated as untrusted, vectors are cached per index by content version under a fixed 128 MiB bound, and retrieval remains non-exhaustive. - Add
runToolLoop: a bounded select, approve/execute, feed-back and re-infer loop overinferToolsandexecuteToolCalls. Approval stays mandatory and sequential; calls requested by the terminal selection are returned, never executed. - Add opt-in
onEventobservability: content-free request lifecycle and engine-load events (ids, timings, token counts, error codes, model identity). Observer failures never change request results; worker forwarding is best-effort. - Document every 1.5.0 API and the items above in the English and Traditional Chinese usage guides.
- Packaging checks now fail on sync-conflict duplicate files (the working tree is iCloud-synced and produced
* 2/* 3copies) and bundleneko.js,neko.js/documents,neko.js/webandneko.js/reportwith esbuild under browser conditions. - Verified on macOS arm64, Node.js 26.7.0 CPU, default profile: 217 Node tests, 27 quality-tool tests, 21 browser contracts (Chromium, Firefox, WebKit), site smoke, packed-artifact check, the state-reuse smoke against the real model (cached continuation equal to baseline text and usage, branch isolation, abort leaves the cache unchanged, vision feature hits), one synthetic image OCR, one synthetic scanned-PDF OCR and native PDF extraction without inference. Remote CI passes its contract matrix on Linux, macOS and Windows for Node 22, 24 and 26 (the real-model lanes are opt-in and were not run). Observed, not asserted as quality: document QA on the PDF fixture returned
insufficient-evidence, and a 128-token budget truncated structured JSON (STRUCTURED_OUTPUT). Not verified: browser WebGPU, browser PDF extraction, the new hybrid retrieval with a real embedding model, and real-model inference on other platforms.
Neko.js v1.5
1.5.0 — 2026-10-03
- Add concurrent, resumable pinned-model downloads with validator-bound staging, cross-process/browser installation locks and explicit cache progress; preserve size/SHA-256 verification before promotion.
- Add engine-local, entry/byte-bounded decoder-state and vision-feature reuse, opaque retained-state handles and explicit release/clear controls. Add opt-in generation/report diagnostics with bounded output capture; reuse does not imply backend parity or output-quality guarantees.
- Add typed tool definitions, structured tool selection and separately invoked, approval-gated tool execution; model-selected calls never authorize handlers.
- Add inert native PDF extraction, bounded PDF rendering and model-backed OCR composition, provenance-aware document indexes, local retrieval and multi-document questions with exact validated citations. Require Node.js 22.13 or newer and package browser PDF assets.
- Add bounded structured/report/web-question/document-question streams and transactional session streaming; commit conversation history only after successful stream consumption.
- Add bounded worker pools, cancellation-aware scheduling and batch results with independent worker ownership; this is not tensor batching.
- Add a bilingual static project website with bounded preview routes. Expand CI contract coverage across OS/Node/browser matrices, packed-artifact checks and explicitly opt-in real CPU/WebGPU/quality lanes.
- Verify typecheck, lint, build, 191 Node tests, 27 quality-tool tests, the static-site HTTP smoke, packed-consumer artifact checks and built public native-PDF extraction on macOS ARM64 with Node.js 26.7.0. These checks do not exercise real model inference, model-backed OCR, browser PDF extraction or the new remote CI matrix; no fresh model-quality or cross-platform parity pass is claimed.
Installation / 安裝
Requires Node.js 22.13 or newer. This project is not published to npm. Download Neko.js-1.5.0.tgz and SHA256 into the same directory, verify the archive, then install it locally:
shasum -a 256 -c SHA256
npm install /absolute/path/to/Neko.js-1.5.0.tgz需要 Node.js 22.13 以上;本專案不發佈 npm。下載套件與 SHA256、驗證後從本機安裝。模型權重不包含於套件;首次設定須下載並驗證固定模型,亦可使用既有的已驗證快取。
Browser CPU/WASM is unsupported for the pinned model and never falls back. Browser WebGPU support and model-backed OCR quality require separate real inference verification. See the bilingual README for native dependency install-script approvals, model licensing and runtime setup.
Neko.js v1.4
1.4.0 — 2026-10-03
- Add tokenizer-aware JSON constrained decoding for an explicit supported schema subset, inferred
SchemaValueresults and explicit validation-only fallback. Reject unsupported constrained schemas before model acquisition; handle integral exponent constants, finite unique-enum domains and UTF-8 token boundaries without weakening runtime validation. - Add bounded AsyncIterable inference streams, transactional conversation sessions, context eviction, branching/reset and model-bound snapshots. Cache only exact rendered-prompt tokenization; do not claim growing-prefix or model KV-cache reuse.
- Make text-only inference planning tokenizer/configuration-only, without ONNX sessions; add staged report planning with explicitly unknown reduction, duration and total-cost bounds.
- Breaking persisted contracts: migrate reports/checkpoints to version 3 and
evidence-first-v3. Add conservative claim–evidence audits, zero-inference extractive reports, typed partial reports, cumulative failed-attempt accounting, increased-budget/retry authorization and bounded per-stage retries. Budget-induced incomplete JSON isBUDGET_EXCEEDED, not an output retry. - Add inert main-content extraction, structured tables/paragraph relations and document questions with SDK-constructed exact citations; reject unsupported claims without pretending citations certify truth or relevance.
- Add EXIF-oriented image regions, bounded tiling, preprocessing provenance and byte/count-bounded owned pixel caches; reauthorize/read/validate before reuse. Support browser raw images with all four channel layouts.
- Add integrity-checked streaming offline bundles with staged imports/cancellation, installation/quota diagnostics, privacy-safe health/diagnostics, explicit worker restart and realm-wide hard deadlines without replay. Bound transferred chunk backing buffers rather than cloning whole upstream allocations; tighten POSIX cache-ancestor trust checks.
- Register immutable Qwen3.5-2B revision
2ea7886f48b926aca97de8b0e041ffca7e3ebaa9alongside the default 0.8B model, with pinned default/all-q4 assets and provenance; exercise both 2B profiles with actual Node text/image/report inference. - Migrate callers, packed-consumer declarations and quality hierarchy contracts; advance the finite evaluator to
quality-claims-v5without lowering acceptance thresholds. Consolidate English/Traditional Chinese API and security documentation after implementation. - Verify 130 Node/27 quality-tool/21 browser contracts, packed-consumer inference, full Node streaming bundle transfer and Chromium 153 WebGPU cold Blob import followed by actual local-only inference, budget-only resume and deadline/restart. Runtime assets remain separately deployed/cached. The full four-fixture quality gate still fails; retain its diagnostics without lowering policy or claiming model-quality acceptance.
Neko.js v1.3
1.3.0 — 2026-10-03
- Validate cheap inference/report options before acquiring a model in inline and worker execution; expose
planInference()with exact chat/schema/image-expanded input tokens, output capacity and context-fit metadata. Valid cold planning calls still load the model/processor. - Stop structured generation at a deterministic single-JSON-value boundary and retain fail-closed Draft-07 runtime validation. Expose
json-boundary-runtime-validationevidence; no schema grammar constraints, repair, retry or factual guarantees. - Preserve all selected paragraph text in an exact source-quote ledger independent of generated summaries; bound section planning to multiple evidence-linked claims and expose structural coverage, conclusion basis and explicitly unmeasured semantic retention.
- Breaking persisted contracts: reports use
schemaVersion: 2; checkpoints useversion: 2andevidence-first-v2. Add validated report/checkpoint serialize/parse helpers and report integrity checksums. Reject older, unversioned and unknown versions without automatic migration; checksums are not authentication. - Record trusted worker execution identity before report checksum generation, so packed Node and browser-worker reports pass validated persistence without post-generation metadata mutation.
- Add quality-regression CI gates separate from informational benchmark output; do not equate valid schemas/provenance with model quality.
- Complete bilingual first-use Node cache/offline and browser mirror/runtime/CORS instructions; derive local tarball filenames from
npm pack. Separate platform API contracts, historical real inference and fresh verification instead of claiming Node/browser/OS parity. - Exercise current changes on macOS/arm64 Node 22.23.3 and 26.7.0 CPU, the packed Node consumer, and Chromium 153 WebGPU inline/worker inference. Document that the strict four-fixture quality gate still fails for missing format prefixes and malformed hierarchical model JSON; no quality or cross-platform parity pass is claimed.
Upgrade and known limitations / 升級與已知限制
- Package version:
1.3.0; Git tag:v1.3. This is a GitHub/local-tarball release, not an npm registry publication. - Persisted report/checkpoint v1 data is rejected. Reports require
schemaVersion: 2; checkpoints requireversion: 2andevidence-first-v2. No automatic migration. - The real-model strict quality gate remains failing: the text/image/boundary cases omit the required
FACTprefix, and the 300-paragraph hierarchy run fails on malformed model JSON atsection:67. The SDK rejects it and preserves all 300 source quotes and 67 completed stages in the checkpoint. This is not a quality pass or a completed 300-paragraph report. - Node CPU and Chromium WebGPU evidence is limited to the tested environments. Browser CPU/WASM is unsupported, and there is no automatic fallback or guarantee that every operator ran on GPU.
- 持久化契約升至 v2,舊版資料不會自動遷移;真實模型品質 gate 仍未通過,不宣稱模型品質或跨平台一致性達標。
Install / 安裝
Download Neko.js-1.3.0.tgz and SHA256 from this release, verify the SHA-256 checksum, then install the local archive:
shasum -a 256 -c SHA256
npm install ./Neko.js-1.3.0.tgzNode.js 22+ is required. Model assets are not bundled; follow the English or Traditional Chinese usage guide for verified model-cache prefetch and offline inference.
Neko.js v1.2
1.2.0 — 2026-10-03
- Upgrade the pinned Node image-processing dependency to
sharp@0.35.4to include upstream libheif and libvips security fixes; retain an exact-version install-script approval. - Add source-aware Page selection, full content snapshots, versioned citations and claim-evidence auditing; add aggregate report token/time budgets, resumable checkpoints, events, and deadline-bound async selector/resource/event/checkpoint hooks.
- Add bounded FIFO request admission, real Node/browser workers, cancellation and streaming, queue status, and actual load/warmup/runtime readiness APIs.
- Apply instance-scoped fail-closed policy to network and local inputs; preserve bound class policy hooks without freezing caller objects, capture model-source getters once before worker serialization, strip sensitive headers on cross-origin hops without restoring them, and align trusted worker URLs across source and bundled runtime layouts. Fail closed on opaque browser redirects; support pinned model mirrors and a verified loopback cache-mirror helper.
- Add typed conversations, joint multi-image inference, bounded generation/sampling/stop controls, and Draft-07 structured-output runtime validation; this does not claim constrained decoding.
- Register the pinned all-q4 profile with native/browser execution evidence without implying output-quality guarantees.
- Seed offline browser quality runs from the verified Node cache into the SDK CacheStorage contract, then enforce
localFilesOnlyand block external network requests; report seed and zero-HF-request evidence. - Seed headed browser UI smokes from the verified Node-cache mirror into browser Cache Storage, then require local-only inference and block model-host requests.
- Correct quality claim scoring so supported cross-ID context does not invalidate a matching fact, and accept both equivalent circle-left-square / square-right-circle relations while rejecting opposite or negated positions.
Neko.js v1.1
Neko.js 1.1.0
This release replaces the prototype with the createNeko() SDK for text/image inference and structured reports from URLs or HTML. It adds conditional Node/browser bundles, browser WASM assets in locally packed archives, model/backend inspection, explicit cache controls, cancellation, and streaming.
Report contracts validate structure, paragraph/image provenance, and generated-field language. English and Traditional Chinese script mismatches return LANGUAGE_MISMATCH; incomplete generation returns INCOMPLETE_GENERATION without automatic retry. imageFailurePolicy: 'omit' applies only to image inference failures; caller callback exceptions and aborts propagate.
Verification documented for this release includes packed-consumer checks, Node native CPU inference, headed browser inference, and offline reload from the same browser profile. The inference suite and browser contract checks were verified separately.
Known limitations
- Model output can be inaccurate: observed image descriptions confused saturated red/blue regions with pink/pinkish-red; a hierarchical report omitted facts and contradicted its source in its conclusion. Schema and provenance validity do not guarantee factual fidelity.
- Browser CPU/WASM is unsupported for the pinned model (
GatherBlockQuantized(1)); no automatic fallback is provided. - URL fetching can create SSRF risk; callers must apply destination validation and outbound network policy.
- Verified Node inference used the native CPU backend; backend availability remains runtime-dependent. Browser execution requires WebGPU. No Ollama parity or universal GPU execution is claimed.
- This project is not published to npm; this release is distributed via GitHub only. TTS is excluded.
Neko.js v1.0
繁體中文
Neko.js v1.0 是以固定 Qwen 模型進行本機圖片+提示推理的原型。Node.js 使用原生 ONNX Runtime 1.30.0;瀏覽器使用 WebGPU 與 ONNX Runtime Web 1.26.0-dev.20260416-b7804b056c。另使用 Transformers.js 4.2.0、Qwen3.5-0.8B Q4 embedding/decoder 與 FP16 vision encoder。已驗證 13 個模型檔的完整性及快取離線重新載入;另提供 inert 網頁擷取、來源資訊與 Markdown 報告 helper。
這不是完整的「輸入網址並產生結構化報告」產品。瀏覽器 CPU/WASM 對此模型不支援(GatherBlockQuantized(1));Node CPU 尚未驗證。不保證自動回退,也不宣稱所有運算子皆在 GPU 執行。尚未提供相同原始圖片與提示的 Ollama 執行結果作為參考,故無法宣稱 parity。請參閱中英文使用指南了解已驗證行為及限制。
English
Neko.js v1.0 is a local image-plus-prompt inference prototype using the pinned Qwen model. Node.js uses native ONNX Runtime 1.30.0; the browser uses WebGPU with ONNX Runtime Web 1.26.0-dev.20260416-b7804b056c. It also uses Transformers.js 4.2.0, Qwen3.5-0.8B Q4 embedding/decoder, and an FP16 vision encoder. Integrity of 13 model files and cached offline reload were verified. It also includes inert web extraction, provenance, and Markdown report helpers.
This is not a complete URL-to-structured-report product. Browser CPU/WASM is unsupported for this model (GatherBlockQuantized(1)); Node CPU is unverified. Automatic fallback is not claimed, nor is execution of every operator on GPU. No Ollama run using the same original image and prompt was provided as a reference, so parity is not claimed. See the bilingual usage guides for verified behavior and limitations.