Releases: tanishqbaweja/trebellcode
Release list
Trebell Code 1.3.6
Trebell Code 1.3.6
Version 1.3.6 is a focused Native reliability and efficiency follow-up to v1.3.5. It keeps the v1.3.5 architecture intact while removing several avoidable recovery turns caused by conventional model path/tool-call mistakes.
Native recovery without guessing
- Relative
workspace/...paths now receive the same bounded workspace-root fallback as/workspace/...only when the literal path does not exist. A realworkspacedirectory still wins, so Trebell does not silently reinterpret valid user paths. - When
trebell_repo/read_sourcehits a known "file is not indexed" boundary andtrebell_workspace/read_fileis already exposed, Native can perform that bounded read internally rather than spending another model turn asking for the same file through a second tool. The model still sees the original repository tool lifecycle plus explicit fallback provenance. - Provider-flattened Trebell tool aliases are repaired only when they identify exactly one exposed tool. Ambiguous names are left untouched instead of being fuzzy-guessed.
- If a later same-turn
trebell_terminal/runcall loses only its executable, Native may restore the command only when exactly one earlier successful terminal call in that same turn has identical args and cwd. Failed executions, timed-out commands, signal-killed runs, empty-args calls, and ambiguous command matches are never used as repair templates. - Recipe tool allowlists now also filter the schemas sent to the model. Disallowed tools remain blocked by the execution gateway even if a provider hallucinates a hidden call, so this reduces prompt cost without weakening the policy boundary.
- Large virtualized tool output keeps its persistent retrieval handle and signal-aware failure lines while using a smaller hot preview, reducing repeated noisy context without discarding the full stored output.
Measured behavior
- The earlier
/workspace/...normalization had already reduced one failure-repair regression from 31,920 input tokens / 10 model turns / 17 tool calls to 8,826 input / 4 turns / 5 calls while independently passing verification. - On the current v1.3.6 candidate, the same failure-repair benchmark again completed successfully in 4 model turns / 6 tool calls / 9,191 provider input tokens, with 0 failed tool calls and independent verification passing. Model/provider behavior remains stochastic; this is a measured run, not a fixed guarantee.
- A repository-tool-description compression experiment was rejected: although it reduced the provider-visible schema from 5,117 to 4,793 JSON characters, a reproduced three-scenario run still used 44,968 input tokens / 17 model turns / 21 tool calls, worse than the cleaner 35,397 input / 14 turns / 17 calls baseline. The fuller descriptions remain in place.
- The live Native benchmark now exposes per-turn input/output/cache usage and request-cost evidence so future optimizations can be judged from the expensive turn instead of only aggregate totals.
- On the same 92 KB noisy-output repair task, the earlier full-schema / 6,000-character-preview run used 18,745 provider input tokens / 6 model turns / 4 tool calls. The current combined allowlist-filtering + bounded-preview build used 15,345 input / 6 turns / 4 calls with independent verification passing, about 18% less provider input in that measured pair. On the first recipe-restricted request specifically, the visible tool surface fell from 11 functions / 5,117 schema characters / ~1,280 estimated schema tokens to 1 function / 772 characters / ~193 estimated schema tokens. Both runs reported zero cached input tokens; provider/model behavior is stochastic, so these are measurements rather than a fixed guarantee.
Validation status before release
- Focused Native loop/session/path coverage passed, including negative cases that prove recovery does not cross permission/tool exposure boundaries or reuse ambiguous/failed terminal evidence.
- Latest full deterministic suite on the candidate: 838 passed, 0 failed.
- v1.3.5 remains untouched; v1.3.6 is a new release and will receive its own installer, updater metadata, blockmap, release metadata, and Git tag.
The full deterministic suite was completed before the final release-hardening commits; the corrected packaging run then reran package-specific Codex, desktop/browser, real-model Native, NSIS, and updater gates on the published commit.
Final release hardening
- Explicit required-tool recovery now narrows the provider-visible schema to the one required tool for the bounded retry, then restores the normal tool surface afterward.
- Packaged Windows PowerShell failures now decode CLIXML envelopes into readable errors, and clean release bootstrap installs dependencies before checking for Electron.
- The final corrected packaging run used -SkipTests because the deterministic suite had already completed successfully. Focused Native loop/session coverage passed 42/42 and release-critical PowerShell/updater coverage passed 9/9; packaged Codex, desktop/browser/SnapShot, real-model deepseek-v4.1, NSIS, and updater metadata/blockmap gates were run again on the exact published commit.
Trebell Code 1.3.5
Trebell Code 1.3.5
Version 1.3.5 separates Trebell Native cleanly from external coding harnesses, expands direct provider support, and adds measured context/tool-loop optimizations without trading away coding quality or verification.
Native vs external harness architecture
- Trebell Native now owns its own agent loop, prompt, context, tools, persistence and direct provider transport instead of treating Codex as an inference bridge.
- Codex uses the real Codex app-server, native account/config/model catalog and native continuation identity. Selecting a Trebell Native provider no longer changes Codex inference.
- Claude Code, OpenCode and ACP-backed runtimes keep their native prompt/session semantics and receive only a small Trebell application-context addition.
- Cross-runtime thread ownership stays truthful: saved external-harness threads resume through the harness that owns them rather than being silently converted to another runtime.
- AgentRouter still receives its required Codex-compatible client fingerprint at the provider-transport layer; this is compatibility metadata, not Codex harness substitution.
Direct provider support
- Trebell Native has first-class direct adapters for official OpenAI, Anthropic and Gemini APIs in addition to AgentRouter, JustWorker, HCNSec and VyceAI.
- Official OpenAI uses the Responses API and preserves Trebell namespaced function-tool identity.
- Official OpenAI Responses requests now carry a deterministic
prompt_cache_keyderived from the stable model/instructions/tool/first-task prefix, so later turns in the same task keep the same cache-routing/accounting key while materially different task prefixes separate cleanly. OpenAI still performs prompt caching automatically where supported; Trebell does not claim a cache hit unless the provider reports one. - Anthropic uses the Messages API with Anthropic-native authentication semantics.
- Gemini currently uses Google's OpenAI-compatible endpoint; native cached-content behavior remains intentionally unclaimed until measured.
- Provider telemetry records request/response bytes, latency, normalized usage, cached-input fields when exposed, endpoint and wire API without persisting secrets.
Measured Native efficiency improvements
- ACP runtime context is injected once per external ACP session instead of being repeated on every prompt.
- The obsolete provider compatibility bridge and Responses-to-Chat bridge path were removed from normal architecture.
- Conventional
/workspace/...model paths are normalized to the active Trebell workspace, eliminating repeated failed tool calls on Windows. A failure-repair benchmark dropped from 31,920 input tokens / 10 model turns / 17 tool calls during the regression to 8,826 input / 4 turns / 5 calls after the path fix while still independently passing verification. - Repeated virtualized tool output is cooled across later user turns to a compact receipt plus searchable handle and high-signal failure lines. On the 92 KB noisy-output benchmark this measured 23,063 input tokens / 7 turns / 6 calls, versus 29,696 input immediately before the change and 37,526 input in the preserved v1.3.3 baseline.
- Native now performs one bounded recovery when the user explicitly names an exact exposed tool and the model tries to answer without calling it. Negated, vague and stale-context tool mentions do not trigger this behavior.
- A tool-order experiment showed the requested tool was called with both normal and reordered schemas, so Trebell keeps schema ordering stable rather than shuffling tools and breaking stable schema hashes.
- A proposed reduction of the baseline repository manifest to only
search_codewas rejected after it caused three focused regressions.search_symbols,search_filesandread_sourceremain directly available until a lower-cost migration is proven without behavior loss.
Current prompt/context telemetry
One real Vyce Native coding smoke on the release candidate measured:
- 4 model turns and 8,082 provider input tokens total;
- 582 estimated system-prompt tokens;
- 11 baseline tool functions and 1,280 estimated tool-schema tokens;
- one stable prefix hash and one stable tool-schema hash across all four requests;
- 0 provider-reported cached input tokens;
- real workspace edit plus terminal verification;
- independent verification passed;
- no Codex app-server in the Native request path.
These are measurements, not fixed guarantees; model behavior and provider routing remain stochastic.
Live/runtime validation
- Real Codex live smoke passed with the Codex-owned model catalog and
gpt-6-luna, while Trebell Native remained configured for a different provider. - OpenCode 1.18.32 passed a real tool/coding smoke with
opencode/muse-spark-1.3-contributor-free. - Antigravity ACP 1.2.1 passed a real tool/coding smoke with
gemini-3.8-flash-high. - Grok reached its real runtime but was blocked by an upstream
Rate limitedresponse during the latest rerun. - Claude Code remained installed but unauthenticated; Cursor's launcher remained unavailable on this machine, so neither is claimed as a live pass.
- Real Trebell Native provider smoke passed for JustWorker, HCNSec and VyceAI. AgentRouter reached the real route but remained blocked by upstream HTTP 402 budget-pool exhaustion.
- Official OpenAI, Anthropic and Gemini adapters are contract-tested; no live success is claimed because those API keys were not configured in the test environment.
Packaged Windows validation
- A real unpacked Electron build passed the installed Codex relay and command-execution smoke.
- Packaged Codex reported its bundled runtime as authenticated through ChatGPT and kept a Codex-owned model catalog even while the Trebell Native provider was switched to VyceAI.
- The packaged desktop/browser smoke passed browser navigation, typing/clicking, cookies, screenshots, video recording, desktop capture, SnapShot persistence, zoom controls, notifications and background/startup toggling.
- A real packaged Trebell Native
deepseek-v4.1run called packaged computer screenshot plus browser open/snapshot/type/click capabilities and created/verified a filesystem proof file. - Windows packaging now builds unpacked and NSIS artifacts in unique isolated temp directories, uses the already-installed Electron distribution to avoid archive-extraction races, retries once in a fresh isolated directory for observed Windows materialization races, and only copies completed installer/update artifacts into
desktop-dist. - The second packaged smoke now waits for the first Electron process tree to fully terminate and uses fresh GUI/app-server/CDP/fixture ports before relaunching, avoiding a Windows single-instance/socket teardown race between the desktop smoke and the real-model Native smoke.
Automated regression baseline
npm test: 829 passed, 0 failed on the consolidated pre-release tree.- Latest full Playwright UI suite in headless Google Chrome: 193 passed, 0 failed.
- Production UI build succeeds.
The release pipeline still performs its own deterministic tests and packaged smoke before publishing.
Trebell Code 1.3.4
Trebell Code 1.3.4
Version 1.3.4 is a reliability and runtime-integration release built from the post-1.3.3 verification work.
Runtime and harness reliability
- OpenCode, Grok Build, and Antigravity runtime detection/authentication now report their real state instead of borrowing readiness from another harness.
- Managed Antigravity ACP installation is pinned to the curated supported runtime and launches from the correct process directory.
- OpenCode live validation prefers the free
opencode/muse-spark-1.3-contributor-freemodel when available, while still preserving connected-provider discovery. - External ACP/OpenCode permission behavior, runtime environment filtering, and process teardown were hardened and covered by integration tests.
UI and thread reliability
- Settings and secondary pages scale more consistently against the sidebar at narrower effective viewport sizes.
- Runtime capability surfaces now stay truthful about unsupported controls, including background-process support.
- Signed-in-but-blocked Freebuff sessions are shown as unavailable instead of falsely ready.
- Stale Codex catalog rows whose underlying rollout no longer exists are hidden from normal history navigation.
- Long-history pagination and virtualized conversation anchoring were tightened and verified in headless Google Chrome.
- Diagnostics refresh failures preserve the last valid log while surfacing the new error.
Persistence and security hardening
- Thread catalog/state handling was tightened around SQLite-backed metadata and runtime identity.
- Secret redaction and runtime child-process environments were hardened so unrelated host credentials are not inherited or persisted.
- Terminal command construction and runtime shutdown behavior received additional Windows coverage.
Verification performed for this release
npm test: 824 passed, 0 failed.- Full Playwright UI suite in headless Google Chrome: 193 passed, 0 failed.
- Live coding-tool smoke passed through:
- Grok Build 1.0.41 with
grok-4.7 - Antigravity ACP 1.2.1 with
gemini-3.8-flash-high - OpenCode 1.18.32 with
opencode/muse-spark-1.3-contributor-free
- Grok Build 1.0.41 with
- Live provider transport smoke passed for JustWorker, HCNSec, and VyceAI.
- AgentRouter transport remained blocked by its upstream HTTP 402 budget-pool quota state.
Release-pipeline hardening
- Packaged Windows validation now runs from an isolated deterministic Trebell home instead of inheriting whichever harness/provider the developer last selected.
- The native Codex queue integration test now validates stable real-app-server behavior without racing Codex's own idle queue auto-dispatch.
Trebell Code 1.3.3
Trebell Code 1.3.3
Version 1.3.3 hardens Trebell Native as a real first-party coding harness and adds live-provider proof that Native is independent from the Codex app-server path.
Trebell Native real-model validation
- Added
npm run test:vyce:native, a real-provider end-to-end Native harness gate. - The live test wires
attachAgentRelaydirectly toProviderManager.turn; it never starts the Codex app-server. - The real model must receive the Trebell Native system prompt and Native tool schemas, inspect source, edit a file, run terminal verification, produce a final answer, and leave output that an independent Node process verifies afterward.
- The v1.3.3 validation passed with VyceAi +
deepseek-v4.1. - Live instrumentation records model turns, provider request size, schema size, tool calls, arguments, and token usage without exposing credentials.
Native coding-agent prompt
- Added a first-party Trebell Native system prompt instead of relying on tool schemas and optional thread developer instructions alone.
- The prompt teaches evidence-driven coding, permission boundaries, verification discipline, untrusted tool-data handling, minimal coherent edits, and honest failure reporting.
- It tells the model to use exact paths/commands supplied by the user or tools instead of rediscovering them.
- It encourages batching independent read-only tool calls to reduce unnecessary model round trips.
- The current release candidate measures the system prompt at roughly 532 estimated tokens; prompt size is not the dominant Native cost.
Progressive repository tools
- Reduced the baseline Native repository tool surface to high-value search/read primitives plus stable
trebell_repo/discover+trebell_repo/invoke. - Advanced capabilities such as semantic rename/code actions, call hierarchy, Git history/blame, deterministic verification control, and durable repository knowledge are exposed only when specifically requested.
- Generic discovery queries such as “repository structure and project layout” no longer inflate the advanced schema set.
- Discovery returns capability metadata as data instead of mutating the provider-visible schema manifest.
- Native MCP follows the same stable-manifest idea with generic discovery/call/resource tools while policy still resolves the real target tool.
- On the measured coding path, baseline functions dropped from 31 to 11 and first-request tool-schema JSON dropped from about 14.3 KB to about 4.9 KB.
Native context efficiency
- Native UI turns now keep scoped repository instructions but replace large preloaded source excerpts with a compact untrusted repository seed map.
- On the Trebell repository audit task, the repository evidence preamble dropped from about 2,778 estimated tokens to about 340 tokens while preserving selected paths and orientation.
- Exact source is fetched on demand through Native repository/workspace tools instead of being replayed in every model/tool round trip.
- Byte-identical repeated source reads are replaced with a compact unchanged-observation marker.
- Large redacted tool output is stored outside hot model context behind bounded
trebell_output/searchandtrebell_output/readhandles, with a signal-aware preview that retains likely failure/assertion lines.
Tool-loop reliability
- Added one bounded recovery attempt for empty terminal model responses; a second empty completion now fails visibly.
- Normalized common safe argv mistakes such as
args: "verify.mjs"andcommand: "node verify.mjs"without implicitly interpreting shell syntax. - Normalized model-style workspace-root paths such as
/src/app.jsbefore both policy evaluation and execution while preserving traversal/absolute-path boundaries. - Added schema validation for Trebell-owned tool arguments before policy evaluation/execution.
- Invalid or unexpected tool arguments fail closed instead of reaching the executor.
- Long-running server/watcher control lives behind lazy
trebell_processtools instead of bloating the always-present one-shot terminal schema. - Project-list background refreshes no longer erase unrelated user-action errors, so failed cross-environment activation remains visibly reported after Trebell safely rolls the environment back.
- Agent Browser recording now prefers WebM codecs, flushes a final MediaRecorder chunk before stop, and fails visibly on empty output. The packaged hidden-window smoke temporarily exposes the capture source to the compositor and now verifies non-empty encoded video.
Native inference telemetry and provider capabilities
- Every Native model step now records a stable inference id, prompt-section estimates, tool-schema/stable-prefix hashes, normalized provider usage, cache-read/write tokens, request/response bytes, latency, provider endpoint/wire API, and context-window utilization where known.
- Trebell preserves internal provenance so actual user text can be measured separately from Trebell application/untrusted working context without changing provider-visible serialization.
- Provider capability metadata distinguishes protocol compatibility from verified prompt caching, explicit cache control, stateful continuation, persistent connections, native compaction, and cache telemetry.
- A controlled Vyce repeated-prefix experiment observed 0 cached input tokens, so Trebell does not pretend the current Vyce Chat Completions route has a cache/stateful continuation feature it has not demonstrated.
Live efficiency result
The first measured real Native coding run consumed roughly 40.4k input tokens for the small validation fixture. The latest equivalent run completed successfully in 4 model turns / 11,997 input tokens / 373 output tokens, with no failed tool calls, one stable prefix, one stable tool-schema hash, and independent post-turn verification.
A new bench:vyce:native live benchmark also covers multi-file refactoring, failing-test repair, and large noisy command output in disposable repositories with independent verification and per-scenario token/latency/tool metrics. The final v1.3.3 benchmark passed all three scenarios:
- multi-file refactor: 5 turns / 16,349 input tokens;
- failing-test repair: 4 turns / 13,148 input tokens;
- 92 KB noisy-output repair: 7 turns / 37,526 input tokens with virtualized command output.
The three tasks totaled 16 model turns / 67,023 input tokens. Large-output work remains an optimization area: virtualization materially bounds retained context and preserves useful failure signals, but model/tool-loop decisions can still dominate total cost.
Release validation
- Explicit
TREBELL_HOMEvalues now isolate ElectronuserDatatoo, preventing single-instance/prefs collisions in isolated release and desktop-browser validation. - Unknown
/api/*routes now return an explicit 404 instead of falling through to frontend assets; the retired mobile-device API is verified absent regardless of whetherui/distalready exists. - 808 / 808 deterministic tests passed.
- 138 / 138 full visual/screenshot tests passed in offline provider-safe mode.
- 2 / 2 targeted Native UI E2E tests passed (compact context seed + thread-owned background process/runtime UI).
- The real Vyce Native coding gate passed with
deepseek-v4.1, including independent post-turn verification.
Trebell Code 1.3.2
Trebell Code 1.3.2
Version 1.3.2 is a repository and release-hardening update built on Trebell's current multi-harness desktop architecture.
Repository cleanup
- Promoted the actual Trebell Code application to the GitHub repository root.
- Removed the obsolete upstream Codex source tree from the Trebell repository root.
- Added goal.md to the repository as the product/architecture contract.
- Added an ignored release-artifacts folder for locally retained installers and updater metadata.
- Clarified that .env is developer/live-test scaffolding only; normal users configure API keys in Trebell Settings.
Architecture and reliability audit
- Retired obsolete mobile-device control and mobile-companion surfaces.
- Moved runtime-specific behavior behind capability contracts where appropriate.
- Kept managed-inference behavior provider-aware without coupling conversation identity to provider.
- Made SQLite authoritative over the JSONL event mirror and fixed byte-bounded fallback compaction.
- Added explicit verified/unverified trust labels to repository knowledge.
- Added MCP result redaction regression coverage.
- Fixed truthful runtime-readiness reporting in Settings.
- Fixed exhausted thread-history pagination.
- Hardened UI tests against asynchronous navigation/readiness races.
- Made Playwright build the current UI automatically before local runs.
Documentation and manual validation
- Rebuilt README.md around the current multi-harness product instead of the older Codex-only architecture.
- Added MANUAL_TESTING.md with real-world validation steps for providers, runtimes, WSL/SSH, forges, MCP, packaging, updater, persistence, secrets and long-running workflows.
- Tightened third-party attribution and provenance documentation.
Automated validation checkpoint
At the release-housekeeping audit checkpoint:
- 782 / 782 deterministic tests passed.
- 138 / 138 visual/screenshot tests passed.
- Windows packaging/update tests passed.
- Installed-app release smoke tests remain part of the Windows release script.
Trebell Code 1.3.1
Trebell Code 1.3.1
This patch release fixes the desktop theme/runtime issues discovered during final Windows validation.
Fixes
- Trebell now defaults to Dark instead of inheriting Windows Light through the old implicit System setting.
- Existing version-1 profiles using the old implicit
Systemappearance migrate once to Dark; users can still explicitly choose System or Light afterward. - The modern Trebell workspace now keeps sidebar, canvas, header, provider footer, and composer in one coherent appearance instead of mixing dark chrome with a white workspace.
- Release validation now keeps Trebell Code and the Agent Browser hidden while native desktop/browser capabilities are tested, avoiding repeated windows appearing during builds.
- Windows release validation relaunches a fresh packaged Trebell instance for model-driven agent checks and cleans stale packaged process trees safely.
Validation
- Headless Playwright verifies Dark is the initial appearance and checks the exact sidebar/workspace/header/composer colors.
- State migration tests verify legacy System defaults migrate only once.
- Native packaged validation remains enabled for Codex, Vyce
deepseek-v4.1, computer use, Agent Browser actions, snapshots, recording, and updater metadata.
Trebell Code 1.3.0
Trebell Code Desktop v1.3.0
Trebell Code 1.3 is the large post-1.2 parity and reliability release. It expands Trebell from a Codex-centered desktop shell into a multi-harness coding workspace while tightening the workflows around background agents, worktrees, browsers, terminals, source control, updates, and provider compatibility.
Multi-harness runtime support
- Runtime profiles for Codex, Claude Code, Cursor, Grok Build, OpenCode, and Antigravity.
- Multiple configured instances per harness, with runtime-specific probing, authentication/status reporting, model discovery, and isolated homes where supported.
- Local, WSL, and SSH environment support for agent execution and workspace access.
- Provider-agnostic usage tracking and restart recovery across supported runtimes.
Background and multi-model work
Ctrl+Enter/Cmd+Enterstarts a new task in the background so the composer is immediately free for another prompt.- Shift-click models in a new-thread model picker to send the same task to multiple models.
- Each selected model receives its own isolated Git worktree and thread.
- Fan-out validates that the project is a Git checkout with a real base branch and rejects detached-HEAD launches.
- Partial failures preserve the failed prompt in Trebell Stash. Ambiguous RPC failures attempt to recover the possibly-started thread by its unique worktree path and warn before retrying, reducing duplicate background work.
Project and worktree parity
- Per-project identity with automatic monograms, emoji, custom monograms, colors, or imported images.
- Project model, permissions, workspace, worktree-submodule, cleanup, and project-action overrides.
- T3-compatible worktree submodule modes:
recursive,top-level, andnone, includingt3.jsoninheritance. - Safe managed-worktree cleanup with inactivity, merged-branch, last-thread-deletion, and unchanged-from-base rules.
- Dirty worktrees, active agent paths, running terminal paths, and non-Trebell worktrees are never automatically removed.
- Cleaned managed worktrees retain their branch metadata and are recreated automatically when reopened.
- Project actions can be saved, imported from
t3.json/package.json, run during worktree setup, wait for setup completion, and open preview URLs.
Codex capability surfaces
- MCP server status, skills, plugins, plugin marketplaces, plugin sharing, apps/connectors, hooks, experimental features, configuration layers, account/usage information, and provider capabilities are exposed through Trebell's Harness Tools UI.
- Connector detail inspection includes tool summaries instead of presenting connectors as dead list items.
- Plugin sharing supports save, checkout, delete, discoverability, and target updates using the bundled Codex app-server's real RPC contracts.
Browser, desktop, and device tools
- Isolated Agent Browser with navigation, DOM snapshots, click/type, screenshots, responsive viewport controls, cookies, recording permission, annotations, and localhost preview discovery.
- One-time browser-profile import for Firefox and Helium on Windows. Other Chromium browsers remain intentionally excluded because their cookies use app-bound encryption that Trebell does not bypass.
- Desktop screenshot context plus explicit full-access mouse, keyboard, scroll, and typing controls for agent computer use.
- Android emulator and iOS simulator discovery/control surfaces, with separate opt-in agent access.
- Global SnapShot shortcut with optional accessibility-text capture and local pending-capture recovery.
Terminal, source control, and review workflow
- Persistent terminal scrollback survives app restarts as stopped history rather than pretending stale PTYs are still running.
- Git branches, worktrees, status, commit/push/pull/fetch, PR creation/review/comments/merge/update/checkout, multi-forge detection, and diagnostics remain integrated in the workspace.
- Review comments, checkpoints, rewind/revert support where the active harness exposes it, linked pull requests, and diff/file review state are retained across the streamlined workspace UI.
Reliability and desktop experience
- Restart continuation can recover supported active threads after Trebell restarts; it remains opt-in to avoid silently resuming work.
- Provider switching no longer leaves a stale model catalog visible while the harness restarts.
- Non-Freebuff compatibility bridges now use per-instance dynamic loopback ports, preventing multiple Trebell instances/tests from fighting over one hard-coded socket.
- Firefox/Helium import, attachment ceilings, remote attachment streaming, conditional keybindings, open-in-editor integration, license browsing, custom theme import/export, appearance modes, usage/cost views, and background tray mode are covered by the current test suite.
- One-click packaged desktop updates now check GitHub releases, download through
electron-updater, show progress, and install/restart only after explicit user confirmation. Normal app quit never silently installs a downloaded update.
Provider validation
- Freebuff, AgentRouter, Vyce AI, JustWorker, and HCNSec integrations remain available to the Codex harness where configured.
- AgentRouter keeps the Codex-compatible request fingerprint required by its current gateway and uses the Responses API.
- Vyce live model discovery is tested against the configured account.
- A live two-model fan-out smoke test ran
deepseek-v4.1andagnes-3.0-flashthrough Trebell's real bundled Codex app-server in separate Git worktrees. Both threads completed and returned the expected proof marker.
Run the reusable live fan-out check from trebell-code:
npm run test:vyce:fanoutPre-release validation
- Node unit/integration suite: 103/103 passing.
- Playwright desktop-style harness flow: passing.
- Real bundled Codex app-server browser relay integration: passing.
- Real Windows DPAPI round-trip and Firefox/Helium cookie-import tests: passing.
- Real temporary Git repository tests for worktree creation, cleanup protection, merge/inactivity rules, and restoration: passing.
- Live Vyce multi-model fan-out through the Trebell + Codex harness: passing.
The Windows release pipeline additionally builds and launches the unpacked packaged app, validates the bundled Codex relay and desktop surfaces, runs a packaged model-driven harness check when Vyce credentials are available, produces NSIS update metadata (latest.yml + blockmap), and then publishes the installer and updater assets to GitHub.
Trebell Code Desktop v1.2.0
Trebell Code Desktop v1.2.0
This development release rolls the large post-1.1 feature batch into the next minor version and validates it directly on Windows.
Highlights
- Project actions with saved commands,
t3.jsonimport, package-script discovery, preview URLs, and worktree setup actions. - Live worktree setup flow with completion-aware PTYs, blocking setup support, visible progress, success/failure states, and terminal access.
- Delegated-agent fleet backed by real Codex thread state plus live per-thread activity, waiting flags, token/context usage, last activity, and errors.
- Rich workspace previews for images, video, audio, PDF, Markdown, CSV, and TSV.
- Open-in-editor integration for common desktop editors and file-manager fallback.
- Automatic localhost preview discovery and one-click project previews.
- Multi-forge source control, remote WSL/SSH environments, conditional keybindings, browser automation, computer use, review workflows, and Codex capability surfaces retained from the earlier parity work.
- Windows development now launches the bundled native
codex.exedirectly, avoiding.cmdpath failures in folders containing spaces. - Windows packaging uses
node-pty's shipped Windows prebuilds instead of forcing a local C++ rebuild, so creating an installer does not require Visual Studio Spectre libraries. - Playwright validation is isolated from the user's real Trebell profile and can use an installed browser channel such as Chrome.
- Vyce supports both
VYCEAI_API_KEYandVYCE_API_KEY, and local validation loads the repository.envautomatically.
Validation
- Node unit/integration suite: 59/59 passing.
- Playwright desktop-style harness flow: passing in installed Chrome.
- Live Vyce
deepseek-v4.1model discovery and direct inference: passing. - Live end-to-end Vyce + Codex + Trebell browser tool loop: passing (
open -> snapshot -> click -> snapshot), including writing and reading the observed proof token from the workspace. - The Windows release build also launches the packaged EXE and requires a model-driven packaged-app check:
deepseek-v4.1must invoke Trebell's desktop screenshot and isolated-browser tools, then use Codex shell/filesystem tools to write and verify a proof file.
Scope note
Trebell Code 1.2 is a Codex-centered harness with multiple model API providers. It should not be described as feature-for-feature identical to current T3 Code: T3 also ships independent Claude Code, Cursor, Grok Build, OpenCode, and Antigravity harnesses, mobile clients, multi-environment pairing/load balancing, multi-account provider instances, and cross-provider usage/cost aggregation.
Run the live checks from trebell-code:
npm run test:vyce
$env:TREBELL_E2E_BROWSER_CHANNEL="chrome"
npm run test:vyce:agentThe live scripts automatically load ../.env when it exists.
Trebell Code Desktop v1.0.1
Trebell Code Desktop v1.0.1
Hotfix release for AgentRouter authentication and desktop scaling.
AgentRouter
- AgentRouter no longer uses a stale hard-coded model catalog
- models are discovered live from the authenticated AgentRouter /v1/models endpoint for the saved key
- saving an AgentRouter key now validates it immediately and rolls back an invalid replacement instead of marking the harness ready
- accidentally pasted wrapping quotes/backticks around provider keys are normalized
- AgentRouter chat requests keep bearer authentication but no longer override the upstream User-Agent, avoiding AgentRouter client-fingerprint rejection
- the model picker now reflects the models actually enabled for the specific AgentRouter account/key
Desktop zoom
- Ctrl + mouse-wheel up zooms the entire Trebell desktop UI in
- Ctrl + mouse-wheel down zooms out
- Ctrl + + / - also changes desktop zoom
- Ctrl + 0 resets to 100%
- desktop zoom is persisted across app restarts
- zoom is clamped to a safe 70%–250% range
- installed Windows validation exercises the real Ctrl + mouse-wheel gesture
Validation
- full Node harness and provider-manager tests
- Responses compatibility bridge tests
- Chromium UI E2E
- Windows NSIS installer build, silent install and launch
- bundled Codex app-server validation
- agent browser and desktop snapshot checks
- SHA-256 installer checksum
This build is unsigned, so Windows SmartScreen may display an Unknown publisher warning.
Trebell Code Desktop v1.0.0
Trebell Code Desktop v1.0.0
Trebell Code 1.0 turns the Codex-backed desktop runtime into a provider-agnostic agentic coding workbench with a calmer T3-inspired control surface.
Agent harness and browser use
- Codex remains the local harness for threads, turns, shell, files, diffs, approvals, checkpoints, worktrees, MCP and orchestration
- the isolated Trebell agent browser is model-callable through dynamic tools for open, snapshot, click, type and screenshot
- browser element annotations, screenshots and cookie import become normal agent context
- new desktop Snap Shot capture attaches the primary display as explicit image context without granting silent global mouse/keyboard control
- real PTY terminals, workspace editing, Git/PR workflows, delegated-agent activity and remote control remain integrated
Selectable inference providers
- Freebuff, AgentRouter, JustWorker.icu, HCNSec.cn and VyceAi
- AgentRouter exposes its supported catalog before authentication
- JustWorker uses its documented Anthropic-compatible messages transport
- HCNSec and other chat-only providers route through Trebell's Responses compatibility bridge
- VyceAi uses authenticated model discovery and OpenAI-compatible chat transport
- Trebell Remote now follows the selected provider instead of silently forcing Freebuff
- runtime User-Agent identity is derived from the package version
Live external-model validation
- a temporary isolated Railway validation used VyceAi deepseek-v4.1
- the model successfully executed Trebell-compatible browser tool calls in sequence: open -> snapshot -> click -> snapshot
- the revealed browser value was returned only after the tool-result round trip
- public CI still keeps provider credentials out of the repository and release artifacts
v1 interface
- simplified thread-first sidebar with less permanent chrome
- aligned workspace/right-panel hierarchy
- denser, calmer conversation and composer layout
- provider configuration is consolidated instead of permanently occupying navigation space
- Browser, Files, Diff, Git, Agents and Runtime remain contextual surfaces rather than dashboard clutter
Release validation
- full Node harness tests and Responses bridge tests
- Chromium UI E2E
- exact Windows installer build, silent install and launch
- bundled Codex app-server startup
- real PTY validation
- isolated agent-browser navigation/snapshot/type/click/screenshot and cookie import
- desktop Snap Shot PNG capture through the packaged Electron preload/IPC path
- provider-switch compatibility against the bundled Codex binary
- SHA-256 installer checksum
This build is unsigned, so Windows SmartScreen may display an Unknown publisher warning.