Token Terminator v0.11.0 — Context, learning & observability
Keep the evidence. Terminate the redundant tokens.
This release brings full-history context ownership, a learned outer policy around
JEV, and a per-profile view of the results. It includes the unshipped 0.10.0
milestone; the previous published version is 0.9.0.
Three additions, one exact-evidence foundation
Select Token Terminator as your Hermes ContextEngine
Use the supported Hermes plugin selector to make TT the context owner instead of
LCM or the built-in compressor. TT starts from full available history, stores exact
pinned evidence, batches JEV relevance/guard/salience and can rediscover omitted
sources. The existing compiler, SkillGate, recovery and Context IR remain in place.
There is no generative history summarizer and no Hermes core patch. Middleware-only
mode remains available for other engines and runtimes.
Learn outside JEV
The opt-in policy-train workflow turns independently labelled omission experiments
into CatBoost omission-harm and recovery-token predictors. A System-2 model revises
semantic feature questions using development errors; JEV supplies batched Noul/Score
probabilities. Grouped validation keeps related tasks/sources together. Questions,
models and thresholds freeze before the separate final holdout. JEV weights do not
change. Runtime prediction needs no CatBoost or pickle: it loads bounded numeric
JSON approved by SHA-256. The policy may veto an omission, not relax safety rules.
off is the default, shadow audits and active requires explicit approval.
See savings per bot
Run token-terminator dashboard at localhost:7474 for profile totals, separate
input/output usage, exact input savings, a fourteen-day trend and local model-rate
valuation. The supported Hermes Desktop extension adds TT ↓ … to the bottom
status bar; clicking opens a summary above it. The Desktop popover works without
the standalone server. Both views read accounting, never prompts or vault contents,
and make no model calls. Profile/connection identities remain separate.
Safety and correctness
- Purpose before reduction: embeddings, reranking, classifiers, helpers,
tokenizer work and internal JEV calls bypass conversational optimization.
Internal callbacks cannot recursively capture or rewrite the main request. - Actual target measurement: ContextEngine and IR accept only exact-token
and character decreases across complete canonical request JSON, including
tool schemas, receipts and real vault references. This is not billed-token framing. - Conservative JEV-backed wrapper policy: retain supplied history unless
measured, reversible compression or known budget pressure justifies a change.
Native JEV remains a typed-decision service, not a supplied chat model wrapper. - Exact evidence first: protected or uncertain material is not force-cut.
Missing/corrupt evidence, invalid learned features and unsupported scopes fail open. - Profile isolation: Hermes task-local profile home takes precedence over the
process launch home; ambiguous historical accounting is not reassigned silently. - Cross-platform artifact integrity: policy hashes cover exact UTF-8 bytes on
Windows and Unix. Existing user configuration is not rewritten by the installer.
Install and enable deliberately
Use the Python interpreter from the environment that runs Hermes:
python -m pip install --upgrade \
'git+https://github.com/AronAxe/Token-Terminator.git@v0.11.0'
python -m rtk_hermes_plus.cli install-context-engineRemove the older rtk-hermes-plus distribution first if present. Enable the general
token-terminator plugin, then select Provider Plugins → Context Engine →
Token Terminator. Preserve other enabled entries in your configuration:
plugins:
enabled:
- token-terminator
context:
engine: token-terminatorRestricted toolsets must allow context_engine. Existing JEV provider/key settings
remain valid. Enable JEV and Context IR explicitly as needed, then restart Hermes.
For the Desktop counter, enable its Desktop half under Capabilities → Plugins.
For dashboard-only installation without choosing an engine:
token-terminator install-dashboard
token-terminator dashboardThe Rust package is updated in step with this release:
cargo add token-terminator@0.11.0It remains the artifact-identity / verification / reduction-invariant companion,
not a Rust port of the Python ContextEngine, dashboard or training system.
Validation and limits
The implementation baseline passed 611 tests, with two optional external
integrations skipped, plus pinned-Hermes contracts, existing Context IR/engine
regressions, real CatBoost fitting and localhost Chromium desktop/mobile checks.
Final workflow status is recorded on the release commit. No paid inference tests
are required for publication; no live pricing or calibration claim is made.
The synthetic learned-policy demonstration retained 8/8 required evidence spans
where deliberately weak fixed-gate scores retained 0/8, while still using about
59% fewer tokens than full input. Its prompt is intentionally larger than the
incorrectly pruned fixed result. Fixture service responses and narrow synthetic
labels do not establish real-model quality, production error rates or net savings.
No production policy ships or activates automatically.
Output savings are not measured. Generated output and input savings are separate;
API-equivalent valuation is not a subscription discount or JEV/recovery-adjusted
net saving. Missing prices and ambiguous usage remain unknown. The localhost server
is loopback-only but is not authenticated against other local users/processes.
Protected/unscored history may still exceed a model limit; unresolved overflow
requires host/provider enforcement. Archive-only rediscovery is bounded lexical
search, not an LCM archive importer. Plan capacity for persistent evidence pins.
Complete Electron click-through and gateway-authentication end-to-end coverage are
not claimed. See the guides for the verified integration boundaries.
Read more
ContextEngine ·
Learned policy ·
Dashboard ·
Migration ·
Security ·
Wiki