Repository navigation
v1.1.0 - 2026-07-21
Release Notes
This release turns hedos into a working seat for coding agents and rounds out the command surface. The gateway speaks a third wire format and enforces tool calling per model, hedos launch runs a coding harness against a local model with nothing to configure, and four new capabilities land on the command line. The default posture is unchanged: the gateway still binds loopback and trusts every local caller.
- Coding harnesses in one command.
hedos launchruns Claude Code, OpenCode, Aider, Goose, or Crush against a model on your shelf. The gateway starts inside the same process on a free port, the harness is wired to it, and both stop together; your own harness config is never touched, so running the tool directly afterwards behaves exactly as before. Name the harness and model, or pick each interactively. - The Anthropic messages dialect. Alongside OpenAI (
/v1) and Ollama (/api), the gateway now serves Anthropic's/v1/messages, so Claude Code and anything built on the Anthropic SDK can point at hedos and reach every model you own. - Tool calling, end to end. hedos reads each model's chat template to decide whether it can call tools, shows that as a capability on the shelf, and the gateway advertises and enforces it: a request that asks a model to use tools is served when the model can and refused with a clear reason when it cannot, instead of failing deep in a runtime.
hedos transcribeturns an audio file into text through a local whisper model — the inverse ofspeak, and the last capability that lacked a first-class command. Point it at a WAV and it streams the transcript back;--languageand--translatemap to what the runtime honors.hedos run --imagefeeds a local image to a vision (see) model, so you can ask a model about a picture from the shell. The picker offers only vision-capable models, and naming a model that cannot see says so up front instead of silently answering blind.- A fit verdict on the shelf.
hedos lsnow has a FIT column that says whether each model will actually run in this machine's memory — fits, tight, too big, or unknown — using the same assessment the install catalog uses, and--jsoncarries it too. hedos statsreads the gateway's audit log back and reports per-model request counts, the rejection rate, and p50/p90/p99 latency, as a table or under--json.- The Ollama daemon starts itself. A cold connection to Ollama now starts
ollama serveand waits for it, so a served request or a pull no longer fails just because the daemon was not already running. - Gated Hugging Face repositories authenticate. hedos reads
HF_TOKEN,HF_TOKEN_PATH, or the token thathuggingface-cli loginwrites when it plans and downloads a gated model — the install path used to ignore it and refuse every gated repo. When access is still denied, the error points at the model's terms page instead of only telling you to add a token. - The interactive
hedos pullpicker is navigable. A blank search shows the models that fit this machine's memory, any query searches Hugging Face, and a "search again" row moves between the two, so switching from recommendations to a search, or trying another query, no longer means re-running the command. No matches or an all-installed recommendation list returns to the prompt instead of ending the command. - Hardening. Approved runtimes receive an allowlisted environment rather than the full ambient one; a Hugging Face revision that is not a safe path component is rejected before it reaches the filesystem; registry and settings writes serialize through an advisory file lock so two hedos processes cannot clobber each other; the audit log and generated launch configs are written owner-only; and an install request's headers are redacted when it is formatted, so an access token cannot leak into a log or an error string. An off-by-default token-authentication prototype is present but inert — multi-client access remains a later, opt-in decision.
- Faster and lighter. The shelf is served from a shared cached snapshot instead of being cloned per request, one identification cache is reused across resolution passes, provisioned Python bundles are stamped so they are not rechecked every command, and the bundles themselves ship as a single compressed archive that unpacks on first use, shrinking the binary.
Install
curl -fsSL https://hedos.ai/install | bashCargo
cargo install hedosHomebrew
brew install theiskaa/tap/hedos| File | Platform | Checksum |
|---|---|---|
| hedos-aarch64-apple-darwin.tar.xz | Apple Silicon macOS | checksum |
| hedos-x86_64-apple-darwin.tar.xz | Intel macOS | checksum |
| hedos-aarch64-unknown-linux-gnu.tar.xz | ARM64 Linux | checksum |
| hedos-x86_64-unknown-linux-gnu.tar.xz | x64 Linux | checksum |