Releases: KumarNavish/navish-commander
Release list
Personal connector 0.4.5 — the agent survives its local server
Personal ChatGPT connector 0.4.5 on unchanged runtime 1.6.0-rc.5.
The Mac agent opened its local Commander server once at module scope, with no retry and the MCP SDK's default 60-second request timeout. A slow or failed start rejected outside every handler and ended the process — three such exits are recorded in the agent log, each dropping the channel connection and whatever request was in flight. The quieter failure was worse: a stdio child that died after startup left the agent running against a server it could no longer reach, so every later job failed and there was no crash for launchd to restart.
The agent now retries the local connection with capped backoff and an explicit connect timeout, reconnects on demand when the child is gone, shares a single connection attempt across concurrent jobs so they never race up competing servers against the same state directory, and installs unhandledRejection and uncaughtException handlers ahead of all startup work so no exit is silent. The disk journal remains the execution owner; recovery semantics, the catalog, transport code and the runtime are unchanged.
test/connector-agent-resilience.test.mjs drives the real agent binary against a local server that fails its first two starts: the patched agent retries, backs off and proceeds, while the previous build hangs until the test times out. The suite passes 150 tests.
Verified live after deployment: killing the running Commander stdio child left the agent up with launchd's run count unchanged, and the next job over the deployed ChatGPT OAuth/MCP path started a fresh server and completed.
No new hosted Remote Desktop Commander throughput comparison was run. The recorded ratio remains the 0.4.1 result of 2.079x median with a 95% interval of 1.988-2.165x, which does not pass the strict lower-bound gate. OpenAI's automatic safety review remains outside Commander's control.
Evidence: evidence/connector-agent-resilience-20260922.json.
Navish Commander 1.6.0-rc.5 — Claude Desktop host parity
Runtime 1.6.0-rc.5 candidate: Claude Desktop host parity
Real use from a Claude conversation on 21 September 2026 exposed that the Claude Desktop extension, which runs as an Electron utility process on Claude's built-in Node, could not start pipe workers, batches or searches: the runtime spawned its detached supervisors with process.execPath, which is the Claude application binary in that host, so every launch ended PROCESS_LAUNCH_OUTCOME_UNCERTAIN within 70 ms while PTY workers kept working. The runtime now resolves a plain Node.js 22.16+ executable before any launch claim (NAVISH_NODE, then a plain node execPath, then PATH and the standard install locations), reports it in commander_devices and commander_get_config, and treats a missing runtime or a spawn error as a definite failure that releases the session identity instead of leaving an uncertain claim.
Further parity changes against hosted Remote Desktop Commander: commander_list_sessions reads only the newest requested page (previously every session record: 7–19 s with 1,683 records in that host); document engines are warmed after connect so the first DOCX/XLSX/PDF call of a chat does not pay module loading; commander_process_output accepts wait_ms; commander_start_process waits up to 10 s inline; commander_read_multiple_files returns up to 64 KiB per file; blockedCommands (RDC's default list) and defaultShell are enforced before any claim and set through commander_set_config_value; a failed commander_edit_block reports the closest text with a character diff; commander_get_file_info line counts follow wc -l; commander_ping and commander_reconcile (owner-verified release of a quarantined resource without re-execution) are new; commander_force_terminate reports operationState: terminated instead of an error; browser steps after an open no longer need expectedUrl. The local catalog has 35 tools (36 through the personal connector).
Six new regressions drive the real MCP server through a Node binary named like an application, an unusable NAVISH_NODE, and the new policies; the suite passes 149 tests. The connector catalog is regenerated; connector transport code is unchanged. Publication and a fresh hosted RDC comparison of this candidate are pending.
Live Claude Desktop probe record against hosted RDC: evidence/claude-host-parity-probe-20260921.json. Merged in #12 with all 22 checks passing; the extracted archive passes 56 MCP tests.
Personal connector 0.4.4 — runtime 1.6.0-rc.5 deployed and verified
Personal ChatGPT connector 0.4.4 on runtime 1.6.0-rc.5.
The connector now serves the rc.5 runtime, whose Claude Desktop host fixes also apply to the connector's Mac agent: pipe workers, batches and durable search start correctly, document engines are warmed, and the catalog is 35 tools (36 over the connector, including its receipt tool).
Verified on the deployed endpoint with a complete owner OAuth/PKCE flow: 36 tools listed, commander_ping returned 1.6.0-rc.5, an exact write was independently read back, and a pipe worker completed through the connector — the operation class that failed entirely on rc.4.
This deployment also repaired a live outage: the channel worker's AGENT_TOKEN no longer matched the Mac agent configuration rotated overnight, so the agent had been failing CHANNEL_CONNECT_FAILED since about 10:25 on 21 September. The pairing was restored and both the channel worker and the Netlify front door redeployed; connector transport code is unchanged.
No new hosted Remote Desktop Commander throughput comparison was run for this candidate. The recorded ratio remains the 0.4.1 result of 2.079x median with a 95% interval of 1.988-2.165x, which does not pass the strict lower-bound gate. OpenAI's automatic safety review remains outside Commander's control.
Evidence: evidence/rc5-live-verification-20260921.json, evidence/claude-host-parity-probe-20260921.json.
Navish Commander 1.6.0-rc.4 — direct chat document workflows
Direct execution from ChatGPT or Claude: native Word, Excel, PDF and image workflows, file and search tools, durable processes, browser workflows and parallel command batches. Chat supplies reasoning; Commander does not invoke Codex, Claude Code, another model runner or a paid model API.
Install
- Claude Desktop: download the
.mcpband install it in Extensions. The current runtime was installed and enabled on macOS, with all 33 tool permissions read back after reopening settings. - Claude Code or standard local MCP: extract the runtime ZIP; use its root with
claude --plugin-dir /absolute/path/to/extraction, or runnode src/server.mjsfrom an MCP client. - Personal ChatGPT connector: use the connector source ZIP and
docs/PERSONAL_CONNECTOR.md. It contains source, lockfiles and deployment instructions, not owner credentials. Existing free Cloudflare and Netlify hosting is supported; account quotas still apply. - Verify downloads against
SHA256SUMS.txt. Node >=22.16 is required; PTY workers additionally require Python >=3.9. Document rendering uses installed Chrome; Linux CI explicitly setsNAVISH_CHROMIUM=/usr/bin/google-chrome.
Verified
All 11 final CI jobs passed, including Node 22.16/24 on Linux and macOS. 142 local tests and 49 tests against a fresh runtime ZIP extraction passed. All 6,757 executable and dependency files in the final archive match that tested extraction. Strict Claude plugin validation and the real Claude Code MCP handshake passed without a model call. Actual Latest / Extra High ChatGPT conversations completed native document work and, in a separate fresh chat, invoked Commander by name without a plugin chip to recover exact metrics and hashes. Synthetic artifacts and sanitized tool-call exports are in evidence/.
The predeclared 50-pair hosted RDC comparison verified all 400 worker outcomes: 2.0793x median throughput (808.325 ms vs 1,680.730 ms), paired 95% interval 1.9883–2.1652x. The strict lower-bound-at-least-2 gate fails. Controlled delivery recorded 0/30 Commander failures vs 10/30 RDC duplicate effects; normal delivery and reconnect passed for both. These fixed workloads are not general model-productivity or population-reliability claims.
Candidate limits
This is an evaluation release, not full unattended or provider certification. During the document conversation, two Word writes were blocked by OpenAI's safety checks, and the client subsequently retried denied intents despite stop instructions. Artifact completion does not excuse that failure. The fresh read-only recovery conversation had no denial or mutation. No Claude model conversation or public directory approval is claimed.
See docs/ACCEPTANCE.md, docs/CHAT_DOCUMENTS.md, docs/DOCUMENTS.md and docs/RDC_CAPABILITIES.md for the exact evidence, negative results and format limits. Earlier release assets and failed measurements remain unchanged.
Navish Commander 1.6.0-rc.3 — direct chat core tools
Commander 1.6.0-rc.3 expands the direct chat interface from 13 to 32 tools. ChatGPT supplies reasoning and code; Commander executes files, processes, browser workflows and command batches without launching Codex or another model runner.
Added directory browsing, multi-file reading, directory creation and moves, file metadata and SHA-256, precise text edits with stale-source protection, durable progressive search, interactive process input, process/session recovery and control, configuration and audit inspection.
Validation: 137 runtime/MCP tests; 44 tests against the actual extracted archive; all 20 GitHub checks passed. Archive and source integrity verified. The personal ChatGPT connector is deployed at connector 0.3.0 and its refreshed catalog contains the 32 runtime tools plus its receipt tool.
This is an evaluation candidate. It does not claim full RDC capability parity, unattended certification or a new 2× performance result. Rich document/PDF operations and RDC-specific account/device integration remain gaps. The earlier Codex-backed job design was removed; historical records remain readable.
Use the ZIP for Claude Code or a generic local MCP client; use the MCPB for Claude Desktop. Node.js 22.16+ is required. Python 3.9+ is needed for PTY workers only. Both archives contain the same bytes; SHA-256 checksums are attached.
See capability coverage and implementation PR.
Live ChatGPT Latest / Extra High check: with Navish Commander explicitly selected, a chat repaired a telemetry summarizer, left its tests unchanged, ran 4/4 tests successfully and wrote a verified report. Independent host checks passed 100 additional cases. A follow-up completed interactive input/acknowledgement/owned termination and four parallel arithmetic workers, all verified. Ten durable mutation receipts and eight executed commands were inspected; none launched a model runner.
The earlier natural-language-only attempt offered a Work handoff, failed its Stay in Chat transition and lost a continuation to a connection interruption before Commander received any operation. That failure remains in the evidence. These bounded successes do not establish friction-free routing, universal reliability or full unattended certification. The ChatGPT tool-call download timed out; the evidence distinguishes rendered chat observations from independent files and durable receipts.
Download direct-chat-core-evidence.json and verify-direct-chat-evidence.mjs. After inspecting the embedded fixture code, run node verify-direct-chat-evidence.mjs direct-chat-core-evidence.json to check record consistency and rerun the four fixture tests in an isolated temporary directory. It does not replay hosted actions or remeasure performance.
Withdrawn architecture — Navish Commander 1.6.0-rc.1
Architecture withdrawn: This candidate introduced a Codex-backed executor and consumed Codex allowance. That contradicts the intended ChatGPT-to-tools workflow. Do not use its repository-job launcher for Codex-free chat work. Runtime 1.6.0-rc.2 / connector 0.2.1 removes that executor. The records below are historical and are not acceptance of the user's intended architecture.
Commander can now run a repository engineering goal as a durable job. Execution continues after the submitting chat ends; a later chat can find the job, inspect independent check results and recover the patch without reconstructing the conversation.
This candidate includes runtime 1.6.0-rc.1, personal connector 0.2.0, a general local MCP server, Claude Code plugin and Claude Desktop .mcpb package.
Delivered
commander_start_job,commander_job_status,commander_jobs, andcommander_cancel_job.- Isolated Git checkouts, immutable job IDs, captured patches and independent declared checks.
- Local recovery through
navish jobsandnavish job JOB_ID --patch. - Existing ChatGPT-authenticated Codex execution with managed workspace restrictions. No new paid API or hosting service.
- The personal ChatGPT deployment matches the released runtime and connector digests.
Actual chat evidence
A ChatGPT Latest + Extra High conversation submitted a README improvement that completed in 188.481 seconds with a passing independent check. A fresh conversation recovered its job ID, changed file, check result and exact patch hash without being given the ID. The unaltered 13,619-byte patch is public and its check was reproduced against the original base.
The initial recovery incorrectly inferred that the patch had not been applied or published. That adverse observation remains preserved. The deployed fix reports downstream integration as unobserved, and the same chat correctly reported that distinction on a fresh status call.
The earlier real engineering job produced the terminal recovery commands but remained needs_attention because its full validation was incomplete. Its patch was separately reviewed and checked before integration; its original outcome was not relabelled.
Job record and patch · Repository job guide
Validation and remaining gates
- 138/138 runtime/MCP tests across Linux/macOS and Node 22.16/24; 17/17 Chromium fixture checks.
- 45/45 tests against the extracted release archive; Claude plugin validation and authored-file/checksum verification passed.
- Current-source comparison: all 160 workers verified, 1.9296× RDC throughput, paired 95% interval 1.8083–1.9865×. The strict 2× throughput gate fails.
- Controlled delivery: 0/30 Commander failures versus 10/30 RDC duplicate effects. Normal and reconnect scenarios passed for both. This fixed fault mixture is not a production failure rate.
General unattended ChatGPT certification, public ChatGPT directory approval and a current Claude Desktop conversation remain unestablished. This is an evaluation candidate, not a stable or independently certified release. Jobs retain failed/uncertain results and do not automatically replay, merge or publish work.
Full acceptance ledger · CI · Merged PR
Downloads
Use the .zip as a Claude Code/local MCP plugin or the .mcpb with Claude Desktop. Both include dependencies and the public evidence. Verify either against SHA256SUMS.txt; the ZIP and MCPB contain identical bytes. Source installers and the optional personal ChatGPT connector are in the tagged repository.
Navish Commander rc.3 — live Extra High evaluation and reader repair
A real ChatGPT Latest + Extra High evaluation exposed a reader defect that hid later lines beyond the response byte cap. Runtime 1.5.0-rc.3 fixes line-offset pagination, Unicode/CRLF boundaries and continuation metadata; personal connector 0.1.2 carries the repair.
The evaluation includes 12 actual conversations and 116 exported Commander calls, preserving ten baseline trials and two separate successful read-only repair confirmations. Both confirmation chats verified actual host contents without writes or process launches. Six planned mutation tasks were not submitted. Three reported denials are uncorroborated; UI-export omissions and misattribution are documented. This is an evaluation candidate, not unattended certification.
Read the complete evaluation and sanitized evidence.
Validation: 117 runtime/MCP tests locally, 24 extracted-bundle tests, strict Claude manifest validation, and an actual Claude Code MCP connection. All 52 baseline fixture inputs regenerate identically. The evidence verifier rejects altered answers, dropped cases, stale source, hidden export failures and altered host observations. Historical >=2x deterministic RDC results remain tied to rc.2/0.1.1 and are not reassigned to this changed candidate.
Downloads:
navish-commander-1.5.0-rc.3.zip: Claude Code/local MCP package with production dependencies.navish-commander-1.5.0-rc.3.mcpb: validated Claude Desktop bundle; GUI installation and a Claude model conversation remain unobserved.navish-commander-source-1.5.0-rc.3.zip: complete source, connector 0.1.2, deployment configuration, fixtures, verifier, tests and evidence; install dependencies and supply your own private configuration.SHA256SUMS.txt: SHA-256 digests for all three archives.
From the matching source checkout, verify the local package with python3 scripts/verify_release.py --integrity-only PACKAGE.zip SHA256SUMS.txt and the recorded conversations with node scripts/verify-chat-efficacy.mjs. Successful verification means package/record consistency, not product certification. No credentials, installed Mac state, or private control repository are included.
Navish Commander 1.5.0-rc.2 — verified local plugin and RDC benchmarks
Navish Commander 1.5.0-rc.2 provides a downloadable local Claude Code plugin and MCP bundle with durable pipe workers, stable operation identities, restart recovery, and artifact verification.
Measured against the actual hosted Remote Desktop Commander service on the same Mac:
- 2.80× staged throughput, 95% paired interval 2.63–2.90×.
- 3.11× inline throughput, interval 2.95–3.39×.
- 0/30 versus 10/30 failed workflows in the fixed delivery-fault suite. Both products passed normal delivery and reconnect; RDC's ten duplicate-delivery cases produced duplicate effects.
Each throughput mode used 20 paired rounds with four workers. Every outcome verified. Source revision, runtime digest, harnesses, all samples, negative development results, and limits are public in evidence/ and the benchmark protocol.
The real Claude Code host loaded the extracted plugin and connected. An existing ChatGPT-authenticated Codex client completed an eight-call conversation plus four calls after a client restart; independent checks verified single effects, all worker hashes, unchanged identities, and correct exit-7 failure reporting. Inspect the sanitized tool trace.
Validation includes 92 runtime/MCP tests, 17 Chromium fixture tests, four integration tests against the extracted ZIP, strict plugin/MCPB validation, zero reported dependency vulnerabilities, and six passing final CI jobs. The archive verifier checks every authored packaged file, the runtime digest, and the recorded gates; it also rejected an intentionally corrupted disposable archive.
Download and extract the ZIP, then run:
claude --plugin-dir /absolute/path/to/extracted/navish-commanderUse the archive root containing .claude-plugin/plugin.json. JavaScript dependencies are bundled. Node.js 22.16+ is required; optional interactive PTY work also needs Python 3.9+. Select transport: "pipe" for noninteractive workers. The .mcpb is also supplied, but Claude Desktop GUI installation and a Claude model conversation were not observed; its quota prompted the user-authorized Codex substitution.
To verify a download, check out this release tag and run:
python3 scripts/verify_release.py /path/to/navish-commander-1.5.0-rc.2.zip /path/to/SHA256SUMS.txtThis is an evaluation candidate, with defined local-plugin acceptance passed. Results compare local stdio with RDC's hosted relay and deterministic workers; network costs contribute. They do not establish 2× general language-model productivity, a production chat failure rate, independent certification, a public ChatGPT connector, or store approval. Earlier negative results and adverse CI runs remain recorded. See the acceptance ledger.
Connector 0.4.3 — recover oversized results without replay
Fixes a reproduced result-delivery deadlock on the personal ChatGPT connector. Runtime 1.6.0-rc.4 and its installed Claude Desktop/MCP packages are unchanged.
A valid interactive command emitted 700,000 bytes, creating a 1.4 MB MCP response. The old channel could repeatedly reconnect without delivering that result. Connector 0.4.3 now sends a bounded notice containing execution state, stable identity, response size and hash. The full original response remains in the local journal. The client can recover process or batch output through bounded reads without repeating the command.
Validation: 143 local tests passed. The actual local SQLite channel test covers dropped upload, reconnect and duplicate suppression. The deployed personal endpoint recovered all 700,000 bytes in 11 reads, verified exactly one side effect, preserved the journal, stopped its test worker and confirmed zero pending requests. Cloud and installed Mac source hashes match 53696b0d899dafa7a92ad863ea92ffabbe561c833324fba41d9e83965f8563eb.
The source ZIP contains code, lockfiles, regression tests, deployment instructions and sanitized evidence. Install dependencies and use your own private configuration as documented in docs/PERSONAL_CONNECTOR.md. Check the ZIP against connector-0.4.3-SHA256SUMS.txt. No credential, model runner, paid API or new subscription is included.
This specific recovery fix is not a new hosted RDC throughput measurement, a new ChatGPT conversation, or full unattended certification. Prior source-bound measurements and platform-denial failures remain in the acceptance ledger. Existing free hosting and chat account quotas apply.
Personal connector 0.1.3 — response recovery
Connector 0.1.3 adds durable response identity and receipt-recovery guidance. It is deployed on the personal HTTPS frontend, channel and pinned Mac agent. Local runtime 1.5.0-rc.3 is unchanged; its existing Claude Code ZIP and Desktop MCPB assets remain immutable.
Eight additional Latest + Extra High conversations produced 69 exported calls. Four of six previously unsubmitted workflows completed, including tested code repair and both four-worker batches. One of two recovery confirmations passed with both receipt lookups independently verified; the other correctly reported partial file effects but retained an unverified lookup claim.
Validation: 120 runtime/MCP tests, 27 extracted-bundle tests, source-bound record verification, and ten rejected corruption controls. CI includes macOS/Linux with Node 22.16/24, a Chromium fixture, package checks and immutable historical evidence checks.
Fresh hosted RDC comparison on the same Mac: 160/160 workers verified; 2.0119× measured throughput, with a 95% interval of 1.7668–2.1605×. The strict 2× throughput gate did not pass. Controlled delivery passed for Commander: 0/30 failures versus RDC's 10/30 duplicate effects. These are deterministic execution workloads, not model productivity or universal unattended certification.
Detailed workflow and recovery evidence · Source and installation
This source archive includes the connector, local MCP/plugin source, lockfiles, tests, deployment definitions and sanitized evidence. It contains no owner credentials, private execution state or installed dependencies. Follow the documented installation instructions. Verify the archive with the accompanying SHA-256 checksum.
Evaluation candidate. Full unattended ChatGPT certification and public directory approval remain unestablished. Existing hosting and account allowances were used; no new paid service or model API was purchased.