Repository navigation
v0.8.49
Mostly a reliability cycle driven by three field reports from a pip-installed
deployment: plugin services that crashed or stalled at boot, a reranker whose
1.2GB of weights could not be downloaded or found, model-download status files
that collided across processes, and a built-in ASR card that was never shipped
in the wheel. Two long-standing invisible-config gaps went with them — a
downloaded model and an edited voice setting both now reach the running
service/agent without a restart.
Added
- Test builds (alpha / beta / rc) are now installable the way a real user
installs OpenSquad. Both release workflows used to exclude those tags, so
taggingv0.8.49-beta.1only built desktop installers — nothing reached PyPI
or npm, and the Release that was created was not flagged pre-release (i.e.
the update check would have offered it to stable users). Everyv*tag now
runs the full pipeline: the PyPI upload carries the PEP 440 spelling
(0.8.49b1, still invisible topip install opensquad), npm publishes under
thenextdist-tag (npm install -g opensquad-ai@next), Docker is skipped so
:latestcannot move, and the GitHub Release is flaggedprerelease. Testers
install withpip install --pre opensquad==0.8.49b1; stable users keep
getting the last stable release from every channel. The tag↔version pair is
validated up front byscripts/sync_version.py --check-tag. - The model-provider list can be refreshed, and a configured provider's API key
can be replaced. The "connect provider" dialog offered about a dozen vendors
while the catalog holds 225 providers / 8292 models: a v1localStoragecopy
won over the backend, the in-memory copy had no TTL, and the refresh button had
been deleted (its i18n keys were still in the bundle, unused) — so a stale
short list could never heal. There is a refresh action again, and the dialog
now states the provider count, where the list came from and whether it is
stale. Each configured vendor gets a "change API key" action that rewrites
api_keyon the vendor's existing model cards in place (merge semantics — the
parameters a user edited by hand survive) instead of asking them to edit JSON.
On the backend, a failed upstream refresh no longer quietly keeps an old disk
cache: after 7 days without a successful fetch the bundled catalog is preferred
and the response says so (source,cache_age_seconds,cache_stalein
meta).
Changed
- The
whisperplugin is no longer shipped. Its declared dependency
openai-whisperwas never what provided thewhisperimport: that name came
fromwhisper1.1.10, an unrelated round-robin-database package, whose import
raisesTypeError: argument of type 'NoneType' is not iterable. Installing
openai-whisperon top would not have fixed it either — both distributions
claim the same top-level module. So the launcher's dependency self-check
reported the dependency missing on every start and refused to start the
service: a plugin nobody could use, surfacing as a service that would not run.
The plugin, its admin panel (plugin-views/whisper/) and its tests are gone.
The ASR surface is unchanged — sensevoice is still the built-in ASR card,
services.whisper_urland thebuiltin_service: whisperhandling are
untouched, and a Whisper-compatible service you run yourself on that port
still works.
Fixed
- Plugin services crash-looped on a pip install with
ModuleNotFoundError: No module named 'pydantic_core._pydantic_core'.external_api,feishuand
telegram(adapter + config), the feishu/telegramsend_tools,
websearch/websearch.pyandplugins/plugin_manager.pyput a computed project
root at the front ofsys.path. In a pip layout that root is a
site-packages (…/site-packages/plugins/external_api/adapter.py→ three
levels up), and plugin services are executed by the bundled Agent Python (3.11)
even when the tree was installed by 3.12 — soinsert(0)shadowed the runtime's
own compiled packages with cp312 binaries, andpydantic_core/__init__.pythen
could not find its own_pydantic_coreextension. They all append now, which
keepsplugins.*/opensquadimportable without giving them priority.
Reproduced against the real bundled 3.11 runtime with a 3.12-tagged
pydantic_coretree: exactly the reported error withinsert(0), clean import
withappend. - Every plugin service waited ~66s at boot because of a dependency no running
service used. The startup batch answered "is this declared pip dependency
installed?" by importing it in the plugin interpreter. A cold
import lark_oapimeasures 10-13s and blew the 60s probe budget under startup
load, and the batch is a global gate (_plugin_deps_ready), so one disabled
feishu's dependency kept websearch, sensevoice and the rest in "Starting…" while
the agent's calls to 127.0.0.1:9001 were refused. The batch now answers from
importlib.metadata— one cached subprocess for the whole run, an installed
distribution is "present" with no import — and no longer pre-installs
dependencies for services that are disabled or not set to auto-start; those are
installed on demand when the user starts the service. Measured here: 29 light
dependencies went from 26.3s to 3.6s with identical verdicts. The import probe
stays as the fallback for a dependency that is not installed (it fails fast),
and the per-service check keeps using it. - The reranker weights could not finish downloading through the mirror, and a
stale "Download failed" outlived the download. The legacyhuggingface_hub
fallback setHF_ENDPOINTandHF_HUB_DISABLE_XETinside the download thread
— i.e. afterimport huggingface_hub, which freezes both into
huggingface_hub.constantsat its own import time (verified on
huggingface_hub 1.26.0: they keep their old values whatever the environment
says later). So the download went to huggingface.co rather than hf-mirror.com,
and with Xet enabled the 1.19GBmodel.safetensorsdied at ~79% on a 401 from
cas-server.xethub.hf.co, which the mirror does not proxy. Both variables are
now set before the import. Separately, once the weights had been completed by
another path nothing ever cleared the persistedstate: error, so the card kept
showing "Download failed: HTTP 502" for a model that was already loaded —
reranker_model_store.get_status()now reconciles a complete model toready
instead of reporting a failure that no longer applies. - Plugin services were told their dependencies were not installed when they
were, so websearch / sensevoice / feishu / telegram refused to start. The
launcher decides by importing each declared pip dependency in the plugin
interpreter. Two things made that decision wrong. (1) The probe guessed the
import name from the distribution name, sopython-telegram-botwas probed as
python_telegram_bot— a module that has never existed under any version. The
package was reinstalled on every startup and reported missing every time, the
circuit breaker opened, and the service never ran; the live launcher log shows
six such failures. The import name now comes from the interpreter's own
metadata (importlib.metadata.packages_distributions(), re-read after every
install), with the static map as an explicit override and the dash-to-underscore
guess only as a last resort. (2) The probe budget was 15 s and a timeout was
read as "missing". Cold imports are not fast — measured on the bundled Agent
Python,lark_oapitakes 10-13 s andtorch~4 s — and the box is busiest
exactly when the launcher starts a dozen services, so a healthy dependency
timed out and blocked its service. Probes now have 60 s, verified modules are
cached (a positive cannot become a negative without an uninstall), and a probe
that still runs out of time is reported as inconclusive: it is installed to
be safe and the service is started, with a warning, instead of being refused.
A realModuleNotFoundErrorstill blocks the start, as before. Separately,
plugin children were handed the launcher's ownPYTHONPATH(its 3.12 package
tree) although they run on a different interpreter;_build_child_process_env
now takes the child's interpreter and clearsPYTHONPATHwhenever it differs
from the launcher's, mirroring what the frozen branch already did. Agent
children are unaffected — they resolve to the launcher's interpreter. - An agent boot could freeze the whole process stack, with nothing in any log
saying why.opensquad startpipes each child's stderr and only read it
after the child exited. A pipe holds ~4 KiB on Windows; once it filled, the
child's next write blocked forever — and here that write was a log call made
while holdinglogging's handler lock. Every thread that logs then queued
behind the lock, so the launcher never reached the line that starts the
log-forwarding thread for the agent it had just spawned, so the agent's own
stdout pipe was never drained, so the agent's event loop froze inside its own
log write: agents never registered and the UI showed "reconnecting" until the
tree was killed. Three independent guards now break the chain, any one of
which is sufficient: (1)opensquad startdrains every child's stderr from
the moment it is spawned, viaStreamTail— a daemon reader with a bounded
ring buffer, whose tail is what gets printed if the child dies; (2) the
launcher starts a child's log-forwarding thread before logging anything
about that child (agent and plugin-service paths both); (3) the console copy
of every logger is a bounded queue drained by a daemon thread
(log_setup.nonblocking_console_handler) whose console writes go through a
private duplicate of the file descriptor, so a stalled console drops records
instead of blocking the thread that logs — or hanging the process on its way
out — while the rotating file handler, now added first, still receives every
line. Reproduced with a child that logs 1.4 MB into a pipe nobody reads: the
old console handler wedges it, the new one lets it exit 0.--verbose
(inherited stdio) still avoids the pipe, but is no longer needed. - No agent could start from a pip or npm install: the artifacts shipped zero
prompt templates.0.8.47and0.8.48contained no file under
prompts/(the publishedopensquad-0.8.48wheel has 578 entries and not one
of them is a prompt), soagents_boot.build_system_promptraised
FileNotFoundError: Base prompt not found: …\site-packages\src\opensquad\prompts\thought_fc.mdon the first boot — the
coder crash-looped and the pm agent never bound its web port while the UI
showed "crashed" / "reconnecting".prompts/is a builtin-resource sibling of
theopensquadpackage, likeskills/oragents/, and had been left out of
packages.find.include,package-dataandMANIFEST.in. All four templates
and the 49parts/fragments they include by name now ship, and
scripts/verify_release_artifacts.pyrequiresprompts/base_fc.md,
prompts/thought_fc.mdand ≥50 files underprompts/before a release is
allowed to upload. npm install -g opensquad-aithenopensquad startno longer does nothing
at all. The wrapper forwarded by re-resolvingopensquadonPATH, which
cannot work on Windows: the shim it is running from is itselfopensquad.cmd
(Node refuses to spawn a.cmdwithout a shell) andpip install --user
puts the real script in%APPDATA%\Python\Python3xx\Scripts, which is
normally not onPATH. The failed spawn leftstatusnull, and the wrapper
droppedres.errorand exited 1 printing nothing — no banner, no error,
indistinguishable from a hung command. It now runs the CLI through the very
interpreter that owns the package (python -m opensquad …, newly enabled by
opensquad/__main__.py) and reports the spawn failure instead of swallowing
it.opensquad startno longer dies on a machine whose SQLAlchemy predates
2.0.38. The gateway passedpool_size/max_overflow/pool_timeoutto
create_async_engineand relied on the dialect's default pool class: up to
2.0.37 a file-backedsqlite+aiosqliteengine getsNullPool, which rejects
those three arguments withTypeErrorwhile the module is imported — the
gateway exited 1 five times in a row and port 9555 never bound, so a
pip install opensquadon such a machine could not start at all. The pool
class is now pinned toAsyncAdaptedQueuePool(the 2.0.38+ default), which
keeps the intended pool on every 2.0.x.sqlalchemy>=2.0.0had been satisfied
by the older release already present, so pip never upgraded it.pip install opensquadshipped no built-in ASR card, so no fresh install
could transcribe voice.ensure_builtin_model_cards()copies
builtin-sensevoice-asr.jsonout of the installedmodel_cards/, and
workspace_utils.BUILTIN_MODEL_CARD_FILESnames that card — but the card was
in no artifact (ignored by.gitignore, absent fromMANIFEST.in, absent
from[tool.setuptools.package-data] model_cards), so the copy loop matched
nothing and returned[]without a word. The agent-voice ASR picker had no
"系统内置 SenseVoice" entry, and 1:1 voice plus group voice failed with
Agent has no ASR configured/ "内置语音转文本不可用". The card is now
tracked and packaged in all four places,ensure_builtin_model_cards()logs a
warning when the install carries no such card (a packaging bug used to be
indistinguishable from "this deployment just has no ASR"), and a test fails if
any name inBUILTIN_MODEL_CARD_FILESis untracked — that guard was the one
missing.builtin-whisper-asr.jsonis deliberately not shipped: it points
at the removed whisper plugin's port 5001, so offering it would advertise a
service that no longer exists.- A model download could abort at 0% with
[WinError 5] 拒绝访问: '…\\download_status.json.tmp' -> '…\\download_status.json'. Status writes
usedtmp + os.replace, andos.replaceneeds DELETE access to the target —
which Windows denies while another process holds the file open for reading,
and this file is read by the launcher, the gateway and the Electron UI on
every poll. The write also ran once per 256 KB chunk inside the download
loop, so the first collision killed the download. The SenseVoice store and the
shared_model_downloaderstore (reranker) now persist status through one
helper that serialises writers per process, retries the replace (10 attempts,
50 ms → 500 ms backoff) and then rewrites the target in place — a plain
open/write needs no DELETE access, so a reader cannot block it — and never
raises. The model-file rename uses the same retry, the websearch setup-status
write goes through the same helper, andtests/test_plugin_model_status.py
reproduces the original failure against the old form. - A running download could be marked "Download interrupted" by a reader in
another process.read_status()andModelStore.get_status()treated
"state == downloadingwith no thread in this process" as an interruption
and persisted it, so the launcher reading the file while the plugin service
downloaded flipped the UI to a failure (and its retry button) mid-download.
Readers are now side-effect free, and only a record that stopped advancing for
150 s counts as dead — a download stalled inside one long read is still
"downloading".get_status()also reconciles the other direction: once the
weights are on disk, a staleerror/idlebecomesready. - A reranker model that had been downloaded could never be loaded, and an
upgrade could cost a fresh 1.2GB download. Two halves of the same path
disagreement. The store downloads the Qwen3-Reranker weights into the
workspace ({workspace}/data/plugins/websearch/reranker) — user data that
survives an upgrade — while the sidecar resolved its model directory as
plugins/websearch/service/reranker/models/…, inside the installed tree.
(1) A model the store fetched was therefore never found: the spawn printed
"model missing at …; auto-downloading", returned without starting anything,
and search silently kept Bing order while the UI reported the model as ready.
The sidecar now falls back to the store's active snapshot. (2) A deployment
that followed the old manual deploy (weights in the plugin tree — the only
path the sidecar looked at) had to re-fetch 1.2GB after every reinstall, which
replaces that tree. The store now carries an install-dir copy into the
workspace once per process: a rename when both are on one volume, a background
copy otherwise (the install-dir copy is never deleted on the copy path), and a
no-op whenever the workspace already has the weights. - A loaded websearch model store could make the telegram plugin unimportable
in the same process.plugins/websearch/reranker_model_store.pyput the
plugin tree at the front ofsys.pathso it could import
_model_downloader. That tree contains atelegram/package, so any later
import telegramin the process resolved to the plugin directory instead of
the installedpython-telegram-botand raisedImportError— and the launcher
imports plugin modules in-process for its status routes, so the telegram
plugin could be reported unavailable depending on which module was imported
first. It appends now, like the ten modules fixed in this class last cycle,
and it is covered by the_PATH_FIXEDguard in
tests/test_plugin_runtime_paths.py. - A model downloaded after its service had started stayed invisible to that
service. Both plugin services resolve their model at boot: websearch decides
whether to spawn the reranker sidecar, and sensevoice opens its ONNX session —
so weights fetched later from the admin UI were never picked up (search kept
Bing order, transcription kept failing) until the user restarted the service by
hand. A plugin's download action now reports the status file it writes
(download_status_path) and the launcher watches it: when the weights become
ready it restarts the owning service (the plugin itself), and it leaves a
service the user stopped alone. A failed or cancelled download ends the watch
without a restart. This complements the sidecar's own fallback to the store's
active snapshot, which covers the auto-download-at-boot path. - Voice settings changed in the UI did not reach a running agent.
PUT /api/agents/{name}/configwrites the newvoice.*cards to config.json,
but the agent's config hot-reload only carriedtools/tool_levels/
model, so a running agent kept the boot-time voice cards and the ASR/TTS
tools kept calling the previous endpoint until a restart. The mtime poll now
also applies avoicechange to the agent runtime context and the injected
ASR/TTS tool config — the same two calls the WebSocketset_voice_config
path already made, so an edit from either surface takes effect the same way.
An unchanged voice is not re-applied, and a failure in one half does not stop
the other.
Full Changelog: v0.8.48...v0.8.49