Skip to content

v1.1.0 - 2026-07-21

Choose a tag to compare

@github-actions github-actions released this 21 Jul 20:43

Release Notes

This release turns hedos into a working seat for coding agents and rounds out the command surface. The gateway speaks a third wire format and enforces tool calling per model, hedos launch runs a coding harness against a local model with nothing to configure, and four new capabilities land on the command line. The default posture is unchanged: the gateway still binds loopback and trusts every local caller.

  • Coding harnesses in one command. hedos launch runs Claude Code, OpenCode, Aider, Goose, or Crush against a model on your shelf. The gateway starts inside the same process on a free port, the harness is wired to it, and both stop together; your own harness config is never touched, so running the tool directly afterwards behaves exactly as before. Name the harness and model, or pick each interactively.
  • The Anthropic messages dialect. Alongside OpenAI (/v1) and Ollama (/api), the gateway now serves Anthropic's /v1/messages, so Claude Code and anything built on the Anthropic SDK can point at hedos and reach every model you own.
  • Tool calling, end to end. hedos reads each model's chat template to decide whether it can call tools, shows that as a capability on the shelf, and the gateway advertises and enforces it: a request that asks a model to use tools is served when the model can and refused with a clear reason when it cannot, instead of failing deep in a runtime.
  • hedos transcribe turns an audio file into text through a local whisper model — the inverse of speak, and the last capability that lacked a first-class command. Point it at a WAV and it streams the transcript back; --language and --translate map to what the runtime honors.
  • hedos run --image feeds a local image to a vision (see) model, so you can ask a model about a picture from the shell. The picker offers only vision-capable models, and naming a model that cannot see says so up front instead of silently answering blind.
  • A fit verdict on the shelf. hedos ls now has a FIT column that says whether each model will actually run in this machine's memory — fits, tight, too big, or unknown — using the same assessment the install catalog uses, and --json carries it too.
  • hedos stats reads the gateway's audit log back and reports per-model request counts, the rejection rate, and p50/p90/p99 latency, as a table or under --json.
  • The Ollama daemon starts itself. A cold connection to Ollama now starts ollama serve and waits for it, so a served request or a pull no longer fails just because the daemon was not already running.
  • Gated Hugging Face repositories authenticate. hedos reads HF_TOKEN, HF_TOKEN_PATH, or the token that huggingface-cli login writes when it plans and downloads a gated model — the install path used to ignore it and refuse every gated repo. When access is still denied, the error points at the model's terms page instead of only telling you to add a token.
  • The interactive hedos pull picker is navigable. A blank search shows the models that fit this machine's memory, any query searches Hugging Face, and a "search again" row moves between the two, so switching from recommendations to a search, or trying another query, no longer means re-running the command. No matches or an all-installed recommendation list returns to the prompt instead of ending the command.
  • Hardening. Approved runtimes receive an allowlisted environment rather than the full ambient one; a Hugging Face revision that is not a safe path component is rejected before it reaches the filesystem; registry and settings writes serialize through an advisory file lock so two hedos processes cannot clobber each other; the audit log and generated launch configs are written owner-only; and an install request's headers are redacted when it is formatted, so an access token cannot leak into a log or an error string. An off-by-default token-authentication prototype is present but inert — multi-client access remains a later, opt-in decision.
  • Faster and lighter. The shelf is served from a shared cached snapshot instead of being cloned per request, one identification cache is reused across resolution passes, provisioned Python bundles are stamped so they are not rechecked every command, and the bundles themselves ship as a single compressed archive that unpacks on first use, shrinking the binary.

Install

curl -fsSL https://hedos.ai/install | bash

Cargo

cargo install hedos

Homebrew

brew install theiskaa/tap/hedos
File Platform Checksum
hedos-aarch64-apple-darwin.tar.xz Apple Silicon macOS checksum
hedos-x86_64-apple-darwin.tar.xz Intel macOS checksum
hedos-aarch64-unknown-linux-gnu.tar.xz ARM64 Linux checksum
hedos-x86_64-unknown-linux-gnu.tar.xz x64 Linux checksum