Releases: srnoob2570/opencode-ollama-cloud
Releases · srnoob2570/opencode-ollama-cloud
Release list
v0.1.12
- Support for opencode V2 (2.x). The entry module now exports the dual shape
{id, setup, server}: opencode 1.x reads.server(the classic factory, unchanged) and opencode 2.x validatesid + setupand runssetup(ctx)against the v2 plugin context. Under V2 the plugin injects the catalog with official pricing throughctx.provider.transform→draft.models.set(verified end-to-end against opencode 2.0.15: the 20 catalog models appear and generation works). The transform is only registered after the catalog loads — on failure, models.dev'sollama-cloudfallback stays alive. Full findings indocs/research/soporte-v1-v2.md.
v0.1.11
The TUI plugin no longer registers its own /model command. It shadowed opencode's native /model picker, which now opens as expected. The model card goes with it: family, quantization, capabilities, limits, release date, official rates and token sizes. The stats line, /stats and the pricing the plugin reports are unchanged.
v0.1.10
Changed
- The catalog is no longer built in this repo. The plugin consumes a single
catalog.jsonfrom srnoob2570/ollama-cloud-catalog, where GitHub Actions build it behind a hash gate: a change in Ollama's/v1/modelslist triggers a full re-extraction, specs come from/api/showand a model parses the rate card. A failed scrape aborts upstream instead of publishing a degraded catalog. The localscripts/, the schemas and the pricing workflows are gone. - Pricing ships embedded in each model's
costblock instead of a separatecatalog/pricing.json. The plugin keeps showing the standard off-peak rate; peak windows live in the artifact'sx_ollamaextension. The cost counter behaves the same as before. - The model card always labels quantization as "(declared)". The value comes from Ollama's
/api/showthrough the upstream pipeline, so the old "(implicit)" source distinction no longer exists. catalogUrlno longer means "catalog without rates". A custom URL is tried first; when it fails or serves an invalid document, the official mirrors stay as fallback.- Documentation: the READMEs claimed the stats line renders "next to the token counter". It does not, and never did: the line renders on the right side of the prompt row (the row with the model name), one row above opencode's own context/cost counter, which is a separate sidebar line the plugin does not touch. The example line now shows only what the plugin renders (
38.2 tok/s · TTFT 380 ms · Session average).
Full changelog: v0.1.9...v0.1.10
v0.1.9
Added
- The plugin now updates itself, the way
@tarquinen/opencode-dcpdoes. On every boot the server entry checks npm once (10 s timeout, a failed check is ignored). When a newer release exists and the plugin came from npm with an unpinned spec, opencode reinstalls the latest on its next start, a toast says "Updated … Restart opencode to finish.", and the stats line shows an↑ <version>badge until then. Repo checkouts and pinned specs are never touched. The check runs even withstats: "off"because it is plugin infrastructure, not stats. tui: "ensure"plugin option (off by default). The TUI host of opencode 1.18+ reads its plugin list only fromtui.json, so an npm install left the stats line dead until a manual edit. With the option on, the server entry registers the TUI entry itself: it patches the tui.json opencode actually reads ($OPENCODE_TUI_CONFIGwhen set, else the global one), adds the spec only when it is missing, keeps comments and existing entries (like["…", { "stats": "off" }]) intact, backs up the file and writes it atomically. Dev installs write nothing. The change lands on the next TUI launch.jsonc-parseris a new dependency.
Fixed
- The stats line no longer stays on the dashes placeholder when the TUI starts with no open session (a fresh start, not a restored one). The slot renders once at mount, so the session id used to stay empty for the whole run; the TUI plugin now falls back to
api.route.currentand picks up the session after it opens.
Changed
- Documentation:
opencode plugin @srnoob2570/opencode-ollama-cloudis the primary install route (it registers the provider and TUI entries in one command). Manualtui.jsonedits and the newtui: "ensure"option are the alternatives.
Full changelog: v0.1.8...v0.1.9
v0.1.8
Fixed
- Cancelling a request no longer prints
no usage chunk seen; steps will be dropped (include_usage missing?). Aborted and cancelled streams never receive the final usage chunk by design, so they are dropped silently. The one-time hint now fires only when a stream completes naturally without a usage chunk, the diagnostic case it exists for.
Changed
- Documentation: the TUI plugin entry must be registered in
tui.json. Since opencode 1.18, the TUI host only loads its plugins from~/.config/opencode/tui.json(or the project's), never from thepluginarray inopencode.json. The README examples now show the verified setup (provider entry inopencode.json, plain file path fortui.tsxintui.json), and the tested version is updated to opencode 1.18.27.
Full changelog: v0.1.7...v0.1.8
v0.1.7
Added
update-pricingGitHub Actions workflow. Refreshing the rate card is now a manual run from the Actions tab (orgh workflow run update-pricing): the workflow fetches Ollama's live pricing page, prints every rate that changed, commits the refreshedcatalog/pricing.jsonand purges the jsDelivr cache. It never runs on a schedule, and the localbun run update-pricingstill works the same way.statsDebugplugin option. When set, the server appends one line per claim attempt (pendings seen, result, per-step source and timestamps) to a boundedstats-debug.login the plugin cache dir, default off — the tool for diagnosing a missing step.
Fixed
- The plugin could stop loading entirely. opencode's legacy loader calls every exported function of a plugin entry module as a plugin factory, so the new
createStatsDebugSinkexport (sorting beforedefault) received the loader's input object as its directory argument and crashed the whole load — no hooks, no measurements, dashes forever. The helpers now live inplugin/models.tsandplugin/debug-sink.ts; the entry module exports only the plugin factory. (The oldtoModelV2export had the same latent bug, harmless only becausedefaultran first.) - Stats timing. A response that took longer than 30 s used to vanish from the session average because the pending window was anchored at request start; it now anchors at stream end. Single-chunk responses (
wire-nostream) no longer fold the whole wait into the session TPS: decode time is measured from first chunk to stream end, as the metric is defined, and those rows carry a(direct)tag in/stats. Zero-token steps (completion_tokens: 0) are rejected consistently instead of entering the average. - Cross-session leaks. The handoff is one file per session (
stats-<sessionID>.json), so a concurrent session can no longer clobber another session's stats, the TUI shows dashes instead of another session's numbers at startup, and stale files older than 24 h plus the legacy single-slotstats.jsonare cleaned up. - Claim correlation. Pending wire measurements now correlate to their assistant message by time (largest
tsat or before the message'stime.created, 2 s tolerance) instead of blind newest-wins, so an early update for an aborted attempt can't count a stale measurement, and an overlapping compaction pending is consumed and dropped. /statsno longer attributes a mixed-model session average to one model — the header readsSession · last model <id>— and the dialog body refreshes every second while it stays open.- The TUI unsubscribes its event handlers and stops its poll timer on dispose (reload-safe), and refresh errors are reported through opencode's log too, not only the private debug file.
Changed
- Stats measure ollama-cloud only. The event-based estimation for other providers is retired, together with its
(event)tag and the untyped runtime fields it relied on; every measured step now comes off the wire. Ratified in the spec (decisions D1–D3, 2026-09-03). - Token sizes in the
/modelcard use the same decimal base (1 000) as the stats instead of binary (1 024). - Server-side bookkeeping is bounded: at most 500 session collectors and 500 steps per session, with exact running totals so the session summary is unchanged. When the opencode seam stops delivering the session header or the usage chunk, the server logs a one-time warning instead of silently freezing the live line on dashes.
v0.1.6
Added
bun run update-pricing. You refresh the pricing table by hand, never on a schedule. The command fetches Ollama's live rate card, prints every rate that changed and rewritescatalog/pricing.json. When the page and the catalog disagree, say a model Ollama just added or retired, it stops with a report and writes nothing.- Pre-commit hooks. Contributors get prettier formatting and a typecheck gate;
pre-commit installsets them up.
Changed
- Pricing now shows the official Ollama Cloud rate by default, the number your credits actually pay per million tokens. The rates come from Ollama's public rate card and ship in
catalog/pricing.json. No configuration needed.pricing: "off"turns it off; the oldpricing: "reference"still works and means on. - The
/modelcard lists the three rates: input, cached input and output. Cache reads are priced at the cached-input rate, so sessions with cache hits cost what Ollama actually charges.
Removed
- The upstream "reference price" and
catalog/pricing-overrides.json. That estimate existed because Ollama published no rates; with an official rate card there is nothing left to correct. Pricing no longer ships insidecatalog/catalog.jsoneither. The table is the only home for rates, and the automated catalog update cannot touch it.
Full changelog: v0.1.5...v0.1.6
v0.1.5
Fixed
- The
/statsand/modeldialogs render in English now. They shipped with Spanish labels (Sesión, Cuantización, "hace 1m") against an otherwise English TUI; every presented string matches the core UI language. The live line was already English and is unchanged. - The catalog's provenance strings are English too, and two rows keep their real source on every refresh: glm-5.2 and nemotron-3-ultra have always had an empty registry blob, so their researched implicit provenance (the checkpoint, the library README) no longer degrades to "previous run" with each catalog update.
v0.1.4
Added
- Streaming stats, opt-out. The plugin measures TTFT and tokens per second for every LLM step, on your machine. For ollama-cloud it wraps the provider fetch and reads the final usage chunk, so the timing is wire-accurate; for every other provider it estimates from opencode's own events. The line next to the token counter shows the session average, like
38.2 tok/s · TTFT 380 ms · Session average. Three things to know. It counts only your main conversation, so subagents, title generation and compaction stay out. The average belongs to the session rather than the model. Nothing is stored and nothing leaves your machine. Idea by @adilfaisal01. - Setup. The stats UI ships as a second plugin entry,
"opencode-ollama-cloud/tui". The API it uses exists in opencode 1.18.25 but isn't documented, so treat the UI as best effort; if a future opencode moves that API, the stats simply disappear and the rest of the plugin keeps working.stats: "off"on both entries leaves the plugin exactly as before. - Two dialogs.
/statsopens the session summary plus the last responses, and you can tell wire measurements apart from the event estimates./modelopens the active model's card: quantization, family, capabilities, limits, release date, and the reference price whenpricing: "reference"is set. - Quantization. The updater now extracts the value Ollama declares for each model, read from the registry config blob and cross-checked against
/api/show. Fifteen of the nineteen models come straight from the registry.glm-5.2andnemotron-3-ultracarry values researched from public sources, with the source noted.minimax-m3andminimax-m2.7stayunknownbecause the public signals contradict each other or are simply absent. Treat it as declared, not guaranteed. Ollama doesn't document the precision the remote inference actually runs at. The field is optional and has no closed enum, so a new Ollama format can't break the updater.
Removed
- Type-level break. The unused
pricingfield is gone from the exportedPluginOptstype. Thepricingoption in your opencode config is unchanged and works the same as before. Only code that setspricingon the internalPluginOptstype will fail typecheck.
Changed
- Internal cleanup. Test fixtures are shared,
familyis now the single vocabulary for model bases, the models.dev seed URL lives in one place, and the updater CLI rejects unknown arguments instead of running a full update on typos.
CI
catalog.jsonnow validates against the published JSON schema, alongside the hand-mirrored checks, so the two contracts can't drift apart.
Full changelog: v0.1.3...v0.1.4
v0.1.3
Added
- Reference prices (opt-in): the catalog ships an upstream-API reference price per model. Pricing resolves automatically from models.dev first-party entries, with manual corrections in
catalog/pricing-overrides.json(overrides win; override-vs-seed disagreements are recorded in the catalog'sconflicts). pricing: "reference"plugin option: opencode's session cost counter shows the reference rate. Defaultoff— and models without pricing data stay at $0 (no partial estimates).
Fixed
- Updater hardening: transient models.dev outages abort loudly instead of publishing a regressed catalog; catalog writes are atomic and self-heal from torn files; malformed overrides are ignored with a warning; marketplace price ties now resolve deterministically; override
asOfkeeps the date the value was taken.
Docs
- Reference pricing and the
pricingoption documented in both READMEs;update --forcefor regenerating after enrichment changes.
Full changelog: v0.1.2...v0.1.3