Repository navigation
Releases: timmo001/triage
Release list
0.7.1
The issue list's State filter now works like the others: tick any mix of states. It starts with regressed, new and ongoing ticked, so quiet, resolved and muted issues stay out of the way, and clearing it goes back to that. States are listed in the same order everywhere, most pressing first.
The API's state can now be given more than once, like host, kind and label. A single state still works.
0.7.0
Issues that have stopped happening now say so.
Issues
- Open issues with no events for a week are quiet instead of ongoing. If one happens again it goes back to ongoing. Nothing is resolved for you.
- The web UI has a Quiet option and count in Filters, and a Quiet group when grouping by state, so you can go through them and resolve or mute the ones you're done with.
Upgrading
statecan now bequiet, andGET /api/issues/countscounts it. Older CLIs and workers don't know it and can fail to read issues from a0.7.0server, so update them along with the server.
Libraries
Issue.Statehasquiet, andIssueCounts.statesincludes it.
0.6.0
A better issue list in the web UI: the server pages, filters, sorts and groups it, and the page loads more as you scroll.
Web UI
- The issue list loads as you scroll, a page at a time, with skeleton rows while it loads. If a page fails, the issues already loaded stay, with a Retry button at the end of the table.
- Sort by last seen, first seen, worth, events or title, and group by state, kind or label. Groups collapse, and the page remembers which.
- Filters open in a panel, on the left on wider screens and from the bottom on phones. It shows how many issues are in each state, and lists of hosts, kinds and labels. Everything starts ticked; untick what you don't want to see.
- Search issue titles.
- Select several issues and resolve, mute or label them together.
- Requests retry for a few seconds when the server can't be reached, so a restart doesn't break the page.
- Narrow screens get a one-column list and a compact toolbar.
Server
GET /api/issuestakesoffset,state,sort,order,groupandsearch.host,kindandlabelcan be given more than once, matching any of them, andlabel=nonematches issues without a label. A singlehostworks as before.GET /api/issues/countscounts the issues matching the same filters, in all and in each state.- Issue lists include each issue's label.
- The default daily decision limit is now 100.
Home Assistant
- The app has an icon and a logo.
Libraries
- The API has
IssueFilters,IssueCounts,IssueSort,SortOrder,IssueGroupingandLabelFilter, and acountsendpoint in theissuesgroup.IssueSummaryhas an optionallabel. - The
listendpoint'shostis now an array.
0.5.0
Know where every issue comes from: which hosts it happened on, what each event was read from, which package and system it came from, and which worker made each decision.
Hosts
- Log errors and out-of-memory kills now get an issue per host, since the same message often has a cause particular to that machine. Crashes and failed units still group across hosts.
- Existing log error and out-of-memory issues are split by host when the server updates. Each keeps the original's state, decisions, labels and suggestions. Their issue IDs change, so old links to them won't work.
- The issue list has a Hosts column and a host filter, and each issue's page shows how many events came from each host, with when it was first and last seen there.
GET /api/issues?host=filters by host, andGET /api/hostslists every host with its event and issue counts.
Events
- Hosts now record the package that owns the program, unit file or kernel behind an event, with its version, plus the OS and kernel. Only events from the current boot get these, since older ones could have run different versions. Packages come from pacman. Hosts without it record the OS and kernel only.
- Events carry a
source,journalfor now. Older events read asjournal. - Each event card shows everything stored about it: host, source, program, unit, executable, signal, result, killed process, package, OS, kernel and boot.
Decisions
- Decisions and suggestions record who asked for them: the worker's name, or
serverorcli. The issue page shows it. Older ones show as unknown.
Libraries
Eventhas asourcefield (decoded asjournalwhen missing) and optionalpackageandsystemfields. The API hasHostCount,HostSummaryand ahostsgroup, andIssueDecisionandIssueSuggestionhaveby.
0.4.1
A fix-up for 0.4.0, whose npm and JSR packages didn't publish. Everything in 0.4.0 still applies, and this is the version to update the Home Assistant app to.
Fixes
- The libraries publish again.
lit-analyzeruses TypeScript without listing it as a dependency, and once TypeScript 7 (which has no JavaScript API) was installed it could pick that up and crash. TypeScript stays on 6 for it now, and installs always give it the same copy. mise run checknow checks the web UI's Lit templates too, so CI catches this kind of failure before a release.
0.4.0
The server now has a web page for your issues, which opens from Home Assistant's sidebar when it runs as the app.
Web UI
- The server serves a web page at its own URL, such as
http://localhost:7171/. It lists the server's issues, filtered by state, and you can resolve, mute, reopen or unmute each one. - The list has a Worth column with the latest decision's verdict, and can be sorted by it.
- Each issue's page shows its hosts and latest events, what each decision model made of it (worth fixing, severity, likely cause), and any suggested fixes. Suggestions are rendered from Markdown and sanitised, with images removed so model output can't make your browser fetch anything.
- Worth fixing and Noise buttons label an issue by hand, the same labels
triage agreementuses to compare decision models. - Pages have real paths, such as
/issues/<id>, so links and the back button work. - Sign in with an admin token. It's kept in that browser until you sign out.
Home Assistant
- Triage appears in Home Assistant's sidebar for admins, through ingress, with no token needed. It's served on its own port that only answers Home Assistant's Supervisor, so port 7171 still needs a token.
- New settings for running it yourself:
serve --ingress-port(TRIAGE_INGRESS_PORT) and--ingress-from(TRIAGE_INGRESS_FROM). - The app moved to
home-assistant/appin the repository. Nothing changes for an installed app.
Smaller changes
hosts list,admins listandworkers listsay when there are none, instead of printing nothing, and point at--serverwhen you're looking at this machine's own database.- The docs moved to https://triage.timmo.dev, with pages on planning your hardware, privacy and the hosting and model choices.
Packages
- Server image:
ghcr.io/timmo001/triage:0.4.0 - Home Assistant app image:
ghcr.io/timmo001/triage-app:0.4.0 @timmo001/effect-triageand@timmo001/effect-triage-client0.4.0on npm and JSR.issues.listreturnsIssueSummary(anIssuewith an optionalworth),issues.getreturnsIssueReview(theIssueDetailfields plusdecisions,suggestionsand an optionallabel), and there's a newPUT /api/issues/:id/labeltakingworthornoise.
Full Changelog: 0.3.0...0.4.0
0.3.0
Issues now have states, so you can resolve or mute them, and see when a fixed one comes back.
Issue states
- Each issue is new, ongoing, regressed, resolved or muted. Issues are new for a week after they're first seen, then ongoing.
triage resolve <issue>...resolves issues once they're fixed. If one happens again after that, it opens as regressed for a week, and decision models look at it again.triage mute <issue>...hides an issue from decision and language models however often it happens.triage reopen <issue>...undoes either.- Add
--server <url>to change a remote server's issues as the admin inTRIAGE_ADMIN_TOKEN, such as the Home Assistant app. triage issuesand the API show each issue's state, and the API hasPUT /api/issues/:id/status.- The server's database is migrated on start. Existing issues start open.
Fixes
collect --follow, which the agent runs, only read the current boot, becausejournalctl --followdoes. A new agent got the last few events instead of its history, and one whose machine was off overnight skipped the end of the previous boot. It now catches up with a plain read before following.
Packages
- Server image:
ghcr.io/timmo001/triage:0.3.0 - Home Assistant app image:
ghcr.io/timmo001/triage-app:0.3.0 @timmo001/effect-triageand@timmo001/effect-triage-client0.3.0on npm and JSR.Issuehas astate, which decodes asongoingfrom older servers, and the API has the new status endpoint.
Full Changelog: 0.2.1...0.3.0
0.2.1
The server image now listens on IPv6 as well as IPv4.
Fixes
- The server image listened on
0.0.0.0, IPv4 only. Home Assistant forwards app ports over IPv6 too, so a client that tried IPv6 first, such as one usinghomeassistant.local, had its connection reset. The image now listens on::, which takes both. Outside the image,TRIAGE_HOSTNAMEis unchanged.
Packages
- Server image:
ghcr.io/timmo001/triage:0.2.1 - Home Assistant app image:
ghcr.io/timmo001/triage-app:0.2.1 @timmo001/effect-triageand@timmo001/effect-triage-client0.2.1, with no changes, on npm and JSR
Full Changelog: 0.2.0...0.2.1
0.2.0
Every part of triage can now run on whichever machine suits it. A small always-on box such as a Home Assistant Green can run the server, while the machine with the models decides on issues and suggests fixes for it. This release also brings the Home Assistant app, Arch packages and a Linux binary.
Workers
triage work --decideand--suggestfetch issues from a server, ask your local models about them, and send the answers back, so the models only need to be reachable from the worker. It uses the same settings asserve --decideandserve --suggest, so a machine's settings mean the same whichever role it has.- Add a worker with
triage workers add <name>. Worker tokens can only fetch work and send back answers. - The server's daily limits (
TRIAGE_DECIDE_DAILYandTRIAGE_SUGGEST_DAILY) now cover the server and all its workers together. Nothing queues up while a worker is off.
Managing a server remotely
hosts,adminsandworkerstake--server <url>to manage a server over its API as an admin, with the token inTRIAGE_ADMIN_TOKEN, so you don't need a shell on the server.TRIAGE_SERVER_ADMIN_TOKENgives a server an admin token from its settings, at least 32 characters, to get started with. It isn't stored or listed.- Any setting can also come from a JSON file at
TRIAGE_OPTIONS, with lowercase keys and noTRIAGE_prefix. Environment variables win over the file.
Home Assistant
The Triage app runs the server on Home Assistant OS, for amd64 and aarch64. Add this repository to the app store, install Triage, set its admin token and start it. Its documentation tab explains adding hosts and workers.
Arch Linux
triage-bin(each release) andtriage-git(main) are in the timmo pacman repository.- They install three user services, none of them enabled:
triage-agent.servicecollects and uploads,triage-server.serviceruns the server, andtriage-worker.servicerunstriage work. See the README for setting each one up. - Each release also has the Linux binary attached.
Capture
- Long crash messages are kept whole.
- Blank error lines are skipped.
- Transient scopes and generated units, which get a new name each time, are grouped as one issue.
Packages
- Server image:
ghcr.io/timmo001/triage:0.2.0, now for amd64 and arm64 - Home Assistant app image:
ghcr.io/timmo001/triage-app:0.2.0, which runs as root, as the Supervisor needs @timmo001/effect-triage: shared schemas and the API, now with the work and token endpoints, on npm and JSR@timmo001/effect-triage-client: the Effect client, on npm and JSR
Full Changelog: 0.1.0...0.2.0
0.1.0
The first release of triage. It captures crashes and errors from your Linux machines, groups them into issues, and can ask a decision model which are worth fixing and a language model how to fix them.
Capture
triage collectreads the systemd journal for crashes, unit failures and errors, and groups them into issues by fingerprint.--follow --uploadkeeps watching and sends events to a server.- Everything is redacted when it's captured, before it's stored or sent anywhere: credentials, emails, home paths, user and host names, MAC and IP addresses, machine IDs, serials and Wi-Fi names. Add your own names to hide with
TRIAGE_REDACT_NAMES. - Failures and crashes keep the last 10 lines the unit logged, redacted the same way, and whether it was a system or user unit.
Server
triage servecollects events from your hosts over HTTP. Add hosts withtriage hosts add, which prints a token that can only send events, and admins withtriage admins addto read issues from the API.- Run it with Docker Compose, with optional Caddy or Cloudflare Tunnel files for HTTPS. See the README.
Decisions and suggestions
AI only runs when you ask for it. Automatic runs are off by default and stop at a daily limit when turned on.
- Decision models rate how likely an issue is worth fixing, through any TypeSafe System One API (such as Ollaya or Ollama) or Clef on Workers AI.
triage labelandtriage agreementcompare them against your own labels. triage suggestasks a language model how to fix an issue, through any OpenAI- or Anthropic-compatible API or Workers AI. Each suggestion records the events it was based on.serve --decideandserve --suggestdo this on a schedule, capped at 20 decisions and 5 suggestions a day by default.
Packages
- Server image:
ghcr.io/timmo001/triage:0.1.0(amd64 for now) @timmo001/effect-triage: shared schemas and the API, on npm and JSR@timmo001/effect-triage-client: the Effect client, on npm and JSR
Full Changelog: https://github.com/timmo001/triage/commits/0.1.0