Repository navigation
Releases: at0m-b0mb/Hermes-Agent-Console
Release list
Hermes 1.7.0 — five bugs found auditing the codebase
A full pass over the codebase looking for bugs and vulnerabilities. Five real ones, each reproduced against a running instance before being fixed, and each now held by a test.
Stored XSS through a filename — the serious one
The Files view rendered a filename into a JavaScript string inside an onclick attribute. esc() is the right escaping for an attribute value, but a value spliced into JS inside an on* attribute is entity-decoded by the HTML parser before the script is compiled — so ' becomes a quote again and the filename closes the string and runs as code.
Agents choose filenames, and an agent reading a hostile page or inbox is the exact case the threat model is built around. This turned "an agent read something nasty" — contained by design — into arbitrary script in the operator's console, with localStorage['hermes.token'] one read away. Demonstrated end to end: the agent wrote the file, the API served the name, the browser executed it.
Paths now travel in data- attributes and are read back with dataset, so the handler is a constant and the value is never compiled.
An agent could fill the disk with its own work
plan caps each call at twelve steps, but nothing capped depth — each of the twelve could plan twelve more: 1728 queued tasks in three generations, measured. delegate blocked only self-delegation, leaving A→B→A. The spend ceiling is not the backstop it looks like: a local model costs 0.00, so the daily cap never trips. Now depth-limited to 3 with a 60-task ceiling, and the model is told which limit it hit.
An interrupted run wedged its agent permanently
A run row is only finalised at the end of run_task, so Ctrl-C or a reboot mid-run left it saying running forever — and the dispatcher reads exactly that to decide who is busy. That agent never got work again, across restarts. Runs now record their owning process and reconcile at boot; liveness is checked per row so a second terminal cannot close out the live server's runs.
The key vault failed silently
decrypt() returned "" for nothing-stored, corrupt data, and "stored but this machine cannot open it" alike — so a restored backup made every key read as absent and re-entering one silently overwrote a key that was never lost. The three states are distinguishable now. The cipher also gained domain separation, while still reading blobs written by the old format.
The events table grew forever
Written on every emit, read by nothing, pruned by nothing: ~15 KB per run, ~127 MB a year from one hourly duty. Now bounded by a retention window and a row ceiling. The tamper-evident audit chain is deliberately untouched.
107 tests, up from 96. Full detail in CHANGELOG.md.
Hermes 1.6.0 — outbound connections actually work
Nothing that needed TLS worked, and it looked like four separate bugs
An agent fetching a web page, the email connector, and every cloud AI backend all failed with the same line:
[SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: unable to get local issuer certificate
One cause. On a macOS python.org build, OpenSSL is pointed at a cert.pem inside the framework that only exists once somebody runs the Install Certificates.command shipped beside the interpreter — and nobody runs it. The default SSL context loads zero roots, so every TLS connection fails at once. Ollama kept working because it is plain http on localhost, which is exactly what made this look like several unrelated problems.
The machine was not short of trust material — it had the system keychain and two OpenSSL bundles. The interpreter simply was not pointed at any of them.
hermes/tlstrust.pyfinds roots that actually load: the interpreter's own store,SSL_CERT_FILE, the usual OpenSSL bundle locations, the bundle your Python shipped with but never wired up, and finally the macOS system keychain — cached after the first look.http_fetch, IMAP, SMTP andprovidersall share it.- Verification is never turned off. The tempting one-line "fix" turns a loud failure into a silent one, in the component that reads pages strangers wrote. Tests assert
CERT_REQUIRED,check_hostname, and that no module builds its own default context. - The error explains itself — it says it is a trust problem on this machine and gives the exact command for your platform.
hermes doctorhas a Certificates line — how many roots, from where, and the full remedy when there are none.
Verified against Gmail, Outlook and iCloud IMAP; Gmail SMTP_SSL and Outlook STARTTLS; a cloud AI endpoint; and ordinary web fetches.
Project site
at0m-b0mb.github.io/Hermes-Agent-Console — deployed automatically whenever the site or its screenshots change. Same design tokens as the console; no fonts, no scripts, no dependencies.
83 tests. Full detail in CHANGELOG.md.
Hermes 1.5.0 — hardened for a server
Hardening for anything reachable from a network, plus the UI fixes that came out of finally looking at the console at full size.
Denial of service — found by attacking a running instance
- A client that never authenticated could park every worker thread. Two hundred half-finished requests took the thread count from 3 to 203, and nothing timed them out. There is now a 20-second socket timeout, a 64-connection ceiling, and — the part that matters — a per-client share of 8, so one client saturating the port cannot take the console away from everyone else.
- The brute-force lockout could be turned against the operator. It keyed on client IP, so somebody spraying wrong tokens locked out the real user — and behind a reverse proxy, where every client shares one address, it was worse than it looks. The token is now checked before the lockout is consulted.
- A body that was announced and never sent held a thread open. Bodies are capped at 2 MB and read with a timeout.
- A malformed
Content-Lengthkilled the handler thread with an unhandledValueErrorand a traceback full of absolute paths.
What the server tells the network
- Security headers on every response:
nosniff,X-Frame-Options: DENY,Referrer-Policy: no-referrer, aPermissions-Policydisabling camera/microphone/geolocation/payment/USB, and a CSP pinning every fetch, script, style and frame ancestor to this origin. - The Python version is no longer advertised in
Server:. - The session token is a header everywhere except the live event stream, which cannot set one — it used to be accepted in the query string on every endpoint.
- A 500 returns a reference, not an exception string.
X-Forwarded-Foris read only from an address you name inserver.trusted_proxy.
hermes doctor reports the security posture
Vault and home permissions, the Host allowlist, whether a proxy is trusted, which agents hold shell access or a wide filesystem scope, and whether the audit chain still verifies.
The console
- Task cards were unreadable on the board — the card was a two-column flex, so the buttons claimed their width first and left the title four characters. Actions now wrap underneath when there is no room.
- Keyboard focus is visible —
:focus-visiblerings, a skip link past the sidebar, and labels on icon-only buttons. prefers-reduced-motionandprefers-contrast: moreare respected, and layouts hold at narrower widths.
Documentation
Six real screenshots of the running console, and a security section describing the ceilings and headers as they actually are.
75 tests, up from 60. Full detail in CHANGELOG.md.
Hermes 1.4.0 — know which model can do the job
hermes bench — which of your models can actually be an agent
Being good at conversation and being able to drive a tool loop are different skills, and the gap is enormous at small sizes. So ask directly:
hermes benchTwo fixed scenarios per model, graded on what ended up on disk — not on what the model said it did:
| Scenario | What it really tests |
|---|---|
| Basics | Can it call a tool at all, and put the result where it was told? |
| Assistant | Read several files, apply a rule, use the real date instead of inventing one, produce an artefact in an exact format without disturbing the sources. |
The first cut of this graded every model "excellent", which meant it was measuring nothing. The assistant scenario exists because a benchmark that cannot separate your models is not worth running. On three local models it now spreads them 17/17, 16/17 and 15/17, with the slowest 2.3× the fastest.
It cleans up completely — no throwaway agents, tasks or runs left in your history.
Tools daily assistant work actually needs
append_file— add to a file instead of replacing it. With onlywrite_file, an agent building something up across steps either loses the earlier work or resends the whole file each time.move_file— move and rename, scope-checked at both ends, refusing to clobber unless told to.now— the real date, time and weekday. A model's idea of today comes from its training data and is confidently wrong.calc— exact arithmetic, which matters the moment an agent is adding up invoices. Parsed to a literal expression tree and refused unless it is pure arithmetic, so nothing here can reach a name or call a function.
Agents made before a tool existed can now use it
A grant dict is a closed list, so an agent created before a tool shipped could never call it and nothing hinted why. Missing entries are filled with that tool's default on startup — ask for anything that writes, deny for outbound mail — widening what an agent asks about, never what it may do unsupervised.
A Files view
See what your agents actually produced, with breadcrumbs, a preview drawer and a download. It enforces the same floor as the agents: protected paths are marked in the listing and refuse to open, verified against a .env sitting in the workspace.
Fixed
The benchmark could hang fifteen minutes per model. Its throwaway agent was supervised, so the moment a model reached for a tool left on "ask", the run parked on an approval nobody was there to give. It is now allowed exactly the tools its scenario needs and denied everything else.
60 tests, up from 46. Full detail in CHANGELOG.md.
Hermes 1.3.0 — see what a run is doing
Aimed at the two things that were actually painful in use: not knowing what a run was doing, and small models fighting the tools.
You can see what a run is doing
- A live elapsed clock on every running task and run row, ticking each second. A run that has been going twenty minutes and one that started ten seconds ago looked identical — and that difference is the whole question you are asking when you look at the board.
- A "waiting on you" badge on any run blocked on your approval. The reason a run looks frozen is usually you, and nothing said so.
hermes tasksshows the same board from a terminal, with elapsed time and what needs a decision.
Less retyping, less scrolling
- Re-run any finished task with ↻ — it opens the composer prefilled with the original brief, agent and priority, because the reason you re-run something is usually that the brief needed a word changing.
- Filter boxes on the Work board and Runs list, matching title, brief, agent, model and status. The board filter survives the auto-refresh instead of eating what you were typing.
Tools small models can actually use
read_filetakesfrom_lineandmax_lines. A big file used to be a flat refusal that left the model with nowhere to go — it would usually just call the same thing again. The error now tells it how to ask for a slice.list_dirtakesdepth. Walking a tree one call per directory burns a step limit fast. Confirmed in a live run:depth 3saw the whole workspace in a single call.- Numeric arguments accept strings, because models send
"3"about half the time.
The console cannot ship broken
CI now runs node --check on the console. It is a single plain script tag with no build step, so one stray bracket takes the whole UI down while every Python check still passes. That happened while building this release — the UI came up blank with App undefined — and this check would have caught it in seconds.
46 tests, up from 42. Every view re-verified rendering clean with no console errors, and the new tools exercised by a real agent against a local model. Full detail in CHANGELOG.md.
Hermes 1.2.0 — verified end to end
Found by running the published build rather than the source tree: a fresh curl | bash install from this repo, real multi-step agent tasks against a local model, the dispatcher working an unattended queue, and the security floor probed from inside a live run.
What held up: the agent loop, the quality gate (it rejected an agent's work twice and made it redo it), the approval path, the filesystem sandbox, and the tamper-evident audit chain. A queued task was picked up with nobody pressing Run, and a protected .env stayed unreadable while an agent worked in the same folder.
These are the gaps that showed up around them.
The record was missing the decisions that matter most
- A human approving or denying a tool call was never audited — the one action the whole security model rests on was absent from the chain that claimed to cover it.
- Storing or clearing an API key was not audited either. Now recorded — the fact it moved, never the value.
- Pausing the workforce left no trace in the live feed, while starting it did.
An API call could fail silently
POST /api/agents/<id>/<anything> answered 200 OK and did nothing — an unrecognised subresource fell through to a plain GET, so a script setting an agent's autonomy got a success and an unchanged agent. Now 404 with the right route named, and 405 for an unsupported method.
search_files could not find files
A query like *.md returned "no matches" because the tool only grepped contents, never filenames. In a real run that made an agent conclude a folder full of markdown was empty and write "No Markdown Files Found" over it. Globs now search names, plain queries search contents and names. The same task afterwards: 4 files found, correct answer, 3 steps instead of 25.
hermes run hid its best feature
The quality gate ran, rejected work and sent agents back to redo it — and the CLI printed none of it. It now shows the gate checking, passing, and exactly what it sent back. hermes doctor also stopped reporting "0 configured" agents on a fresh install.
Automatic scoring admits that it is off
Scoring needs a judge model and is off by default (it costs a second model call per run). Nothing said so, so Performance showed blanks forever. It now explains the blank and links to the setting, and the README stops implying grading is automatic — the always-on quality gate is described separately from opt-in scoring.
42 tests, up from 38. Full detail in CHANGELOG.md.
Hermes 1.1.0 — a keyboard for the console
The console gained a keyboard
- Command palette on ⌘K / Ctrl K / / — every view, every agent, and the actions worth reaching in one keystroke. Results are ranked, so a loose match on another row's description never outranks the row whose name you actually typed.
- Single-key navigation — g then a letter to jump between views, n to assign work, t for the theme, esc to close, ? for a sheet listing all of it.
Light theme
A full light palette beside the dark one, toggled with t and remembered per browser. With no stored choice it follows the operating system, and it is applied before the first paint so there is no flash of the wrong ground.
Elsewhere in the console
- Desktop notifications, opt-in, for the two things worth interrupting you: an agent waiting on your decision, and a run finishing.
- Export a run as Markdown — metadata, full transcript, final result.
- Agent icons render. They are emoji, and the font stack had no emoji family in it.
hermes run finishes what it starts
- Approvals are answered in the terminal. A supervised agent asks before every write, so
hermes runused to sit in silence for the full fifteen-minute approval timeout with no way to say yes. It prompts where you are now. --yesapproves tool calls as they come up — and deliberately stops short of the actions that leave the machine.email_sendstill asks, every time.- With no terminal attached, requests are denied immediately with the remedy printed rather than hanging until timeout.
Documentation
The README is rebuilt around "How do I…?" — tables mapping a thing you want to do to the way to do it and the file that implements it, a code map for where to change each behaviour, and a diagram of how a tool call travels from model output to a guarded, audited, redacted result.
Full detail in CHANGELOG.md.
Hermes 1.0.0
Hire AI agents. Give them jobs. Watch them work.
A local-first console for running a team of AI agents that behave like employees — they pick up their own queue, do the work with real tools, and a quality gate checks what they actually did before any of them is allowed to call a task finished.
curl -fsSL https://raw.githubusercontent.com/at0m-b0mb/Hermes-Agent-Console/main/install.sh | bash
hermesWhat you get
- Agents that own a queue. A dispatcher hands work to whoever is free and on shift. Nobody presses Run.
- A quality gate. A second model checks whether the work was performed or merely described, and names the specific gap when it was not.
- Two front ends, one engine — a graphical console at
localhost:4317, andhermes shellfor SSH sessions with no browser. - Six backends — Ollama, Groq, Gemini, Anthropic, OpenAI, or any OpenAI-compatible endpoint, chosen per agent.
- Runs fully offline. With Ollama and no API key, every feature works.
- Zero third-party dependencies. Python standard library only — asserted by a test that walks every import in the package.
Security
The floor is enforced in code, not in a prompt: 23 protected path patterns, 18 blocked command patterns, and Hermes' own vault, token and database off-limits to agents. email_send requires a human decision on that specific message at every autonomy level and every capability setting, so a prompt injection can hijack an agent completely and still not get one byte out.
Token auth with rate-limited lockout, loopback-only binding by default, an encrypted key vault, secret redaction, spend ceilings, and a hash-linked tamper-evident audit chain.
Verification
python3 tests/test_hermes.py37 checks covering protected paths, sandbox escapes, blocked commands, the outbound-mail gate, the redirect allowlist, tool arguments, injection scanning, redaction, the key vault, audit tamper-detection and autonomy. CI runs them on Linux and macOS across Python 3.9, 3.11 and 3.13, plus shellcheck, a cold install, and a job that pipes the installer into bash.
Full detail in CHANGELOG.md, including the six bugs found and fixed while auditing the first cut of the code.