Releases: gaborini/monitrs
Release list
v1.0.1
[1.0.1] - 2026-08-02
The soak 1.0.0 deferred, run and passed — after the first attempt found that the fault
was in the test harness rather than the program. Plus two small fixes that only surfaced
by using the released binary and the released documentation for real.
Fixed
-
The twelve-hour soak measured its own bookkeeping and called it monitrs.
Injector
incrates/monitrs/tests/soak.rskept every acknowledged keypress — a triple of
Durations, 48 bytes — so it could compute latency percentiles at the end. Over the
twelve-hour run on 2026-08-01 that was 856 644 keypresses and 40 155 KiB, against
§16.1's allowance of 16 384 KiB. The gate failed on the harness alone, and could
never have passed at any realistic input rate. The sample is now bounded at 32 768
triples, decimated so it stays spread evenly across the run, with the keypress count
and the worst case kept exactly — the latter because it is the stall detector, and a
stall the decimation skipped would be a stall the harness failed to see.Subtracting the harness analytically from the failed run predicted a residual of 21 KiB
on x86_64 and 785 KiB on aarch64. The re-run measured 593 KiB and 115 KiB — the right
magnitude, the wrong split, because at that scale a measurement minus a model is
noise-dominated. Good enough to conclude there was no leak; not good enough to publish
a figure, which is why the run was repeated rather than the arithmetic reported. The
measured result is under Known limitations below. -
monitrs snapshot --format json | headno longer reports an error. Closing the
pipe is the reader's decision, and monitrs treated it as a failure: it printed
monitrs: Broken pipe (os error 32)and exited non-zero, so under the
set -o pipefailthat careful scripts use, an ordinary pipeline failed. It now exits
quietly and successfully, and records the reason in--debug-logat debug level.
Onlysnapshotcould reach this in practice — its export is about 800 KB, larger than
any pipe buffer, so it is still writing after the reader has gone.completions
(17 KB),manpage(4 KB) andconfig(under 100 bytes) fit in the buffer and finish
writing first, which is why nobody had seen it there. Their handling is fixed too:
configprinted withprintln!andcompletionsused a helper that writes straight
to stdout, and both of those panic rather than return an error when the write fails,
so had the pipe ever broken the result would have been a panic report. Both now route
through the same path assnapshot.The recognition is deliberately narrow — the failure must itself be the broken pipe.
A broken pipe reachable only as some other error's cause is still reported and still
exits non-zero, because that one did not come from writing monitrs' own output, and
silently exiting 0 on a run that really failed would be worse than the message this
replaces.
Documentation
docs/soak-testing.md's independent memory check was measuring the wrong process.
The page asks for a resident-size reading taken from outside the soak, so the figure
does not rest solely on monitrs' own measurement, and then gavepgrep -f soak-—
which also matches theteethe same page tells you to pipe through. It is not a coin
flip:pgrepprints ascending PIDs,head -1takes the first, theteeexists from
the moment the pipeline starts, and cargo spawns the test binary minutes later once
the build finishes. So the corroborating reading was of the one process guaranteed not
to grow. The pattern is now anchored to^[^ ]*/deps/soak-, verified on Amazon Linux
2023 and macOS, and the page says how to confirm it matched exactly one process.
Known limitations
1.0.0 listed four. Two are closed by measurement, one stands unchanged, and one is
new — and the new one is the reason the first cannot simply be called fixed.
-
Closed: the budget has now been read on §16.1's own reference workload.
1.0.0said
every idle figure came from a host carrying 981–1007 processes, five times the 200 the
budget names, and that the run was owed. It has been taken — on 8-vCPU EC2 instances
padded to 199–200 processes, both Tier 1 Linux architectures, three runs per
configuration, twice over on two independent applies. Twelve runs agree to the digit.
The result is worse than the macOS figures, not better: Overview median 2.66% against
a 1% budget and p95 3.99% against 2%, so on Linux the median misses too, where on
macOS it passed at 0.60–0.85%. The Battery screen reads median 1.33%, p95 2.66–3.99%.
The padding was 6–16 synthetic processes against a baseline of 184–194, so this is a
97%-real process table, andmeasure-overhead.py's own gate reportedworkload matches §16.1's referenceon every run. -
New: the budget asks for more precision than the instrument has. Every reading above
is a multiple of 1.33% — the observed values are 0, 1.33, 2.66, 3.99, 5.32 and
nothing between, because one 10 ms scheduler tick over the script's 0.75 s poll is
1.33%. A median "below 1%" is therefore not a value this measurement can report: the
only reading under 1.33% is 0.00%. The p95 budget of 2% can be met only at 0.00% or
1.33%. Whether the answer is a coarser budget, a longer sample or a finer clock is
undecided; what is settled is that the budget as written cannot be evaluated at the
resolution it implies. -
Unchanged: the cause
1.0.0was designed around is still the wrong one, and the
medium tier's two filesystem-capacity reads are still the measured cost at 13.2–35.0 ms
of CPU per tick. Nothing in this release addresses it, and the Linux figures above do
not separate the two reads either. -
Closed: §16.1's twelve-hour soak is met. Run 2026-08-02 against the fixed harness,
on both Tier 1 Linux architectures, and it passes with room to spare: resident memory
moves 14 490 → 15 083 KiB on x86_64 and 16 097 → 16 212 KiB on aarch64, first
quartile to last, against an allowance of 16 384 KiB. Growth of 593 KiB and 115 KiB
over twelve hours. The history ring held one size across all 720 measurements and the
descriptor count held at 4, and an independent sampler reading/proc/<pid>/status
from outside the process agrees. Peak resident size fell from 54 724 KiB to 15 688 KiB
for the same work — 857 476 keypresses and 2 646 520 snapshots — which is the forty
megabytes the harness had been accumulating and calling monitrs'.The first attempt, on 2026-08-01 against
v1.0.0, failed. It failed in the harness; see
Fixed above.soak-10kpassed in both attempts.
Verifying a download
sha256sum --check --ignore-missing SHA256SUMS
gh attestation verify monitrs-*.tar.gz --repo gaborini/monitrsv1.0.0
[1.0.0] - 2026-08-01
The stability promise. Four surfaces are frozen and each has a machine guard behind it.
On the way there the sensors came off the one-size-fits-all schedule, and a reading
carried over from an earlier tick now says how old it is instead of being republished as
though it had just been measured.
Added
Sensors read on their own cadence, and a carried-over reading is dated
- Temperatures and the battery are read every 30 seconds instead of every 5 — and
every 5 seconds while the Battery screen is visible, plus immediately when you open it.
They used to ride the 5-second medium tier whatever was on screen. On macOS one such
read costs about 85 ms of wall clock, which is a long time to spend on a number nobody
is looking at. It did not fix the budget it was aimed at, and Known limitations
below says so rather than leaving you to find out. - A reading carried over from an earlier tick is published as stale, carrying its age.
Until now a value read four seconds ago was republished as freshly measured, and nothing
downstream could tell the difference. At a 30-second cadence that would have become a
half-minute-old temperature presented as current, which is the one thing this project
will not do. The rule reaches the interface, the JSON export, and a library consumer
readingSensorSnapshotalike. - The header dates its hottest reading:
temp 62.5C ~00:28. The alternative was a
field that vanished for most of every 30 seconds, which is what filtering to freshly
measured readings would have done — a dated number beats a blank. - The Battery screen states the age once per panel, in the panel's own trailing label
—82% discharging ~00:28,2 sensors ~00:28. A dozen fields unwrapped from one
retained snapshot are all exactly the same age, and~00:28printed beside each of them
would be one fact rendered twelve times; the fields themselves are drawn in the stale
style rather than as measurements. The charge meter is the exception, because every
meter annotates its own value where it draws it.
The Time Lens caret says what changed, not only when
- A caret parked on a historical sample now carries the sample's wall clock and how its
CPU compares with two baselines, then how far back it is:22:14:07Z, then
cpu prev +41 points, then30s -6 points, then-00:37 selected. §2.5's comparisons
have been in the core since0.1.0with nothing calling them; this is where they reach
a screen. A baseline the history cannot reach reads30s no baselinerather than+0,
because a zero delta would say the metric did not change when the truth is that there
was nothing to compare it against. - The segments are ordered by what nothing else carries, and the note drops from the
end when the caret's side of the row runs out of room. The wall clock leads: between 80
and 99 columns this note is the only place the selected sample's clock appears at all,
because the compact header's one-line strip reads the live snapshot. The relative offset
goes last, since the[<HISTORY -MM:SS]badge repeats it at every width.
What 1.0.0 froze, and the four guards behind it
- Four surfaces are frozen: the public API of
monitrs-core,monitrs-collectorsand
monitrs-tui; the JSON export; the configuration keys; and the default keymap.
CONTRIBUTING.md's What 1.0.0 froze is the reference rather than
this list — it states the terms of each, what is deliberately not frozen (layout,
wording, colour, glyph choice and panel arrangement are presentation, not API), and why
no one of the four guards can see another's blind spot. - Each surface has a mechanism, not a paragraph. The API is checked by
cargo-semver-checksin CI; the export bydocs/schema/v2.json, an inventory of all 292
field paths version 2 produces, andcrates/monitrs/tests/schema_contract.rs; the
configuration keys bydocs/schema/config-v1.json, all 32 of them, and
crates/monitrs/tests/config_contract.rs; the keymap by exact key→action assertions in
crates/monitrs-tui/src/keymap.rs. Three of the four fail a build today. The semver job
reports without blocking until the commit that tagsv1.0.0, which is where it becomes
the gate.
Changed
Breaking for library users, one at a time:
- Six public enums are now
#[non_exhaustive]:Effect,Action,ViewIdand
SortField(monitrs-tui), the paletteCommand(monitrs-tui), andHistoryMetric
(monitrs-core). Matching one of them from another crate now needs a wildcard arm,
which is itself the break — and this is the last moment it can be taken, because the
attribute is what buys the whole of1.xa new screen, a new effect, a new sortable
column, a new palette command or a new retained metric without a major bump.
MetricStateis deliberately excluded: there a consumer's exhaustive match is the
protection rather than an inconvenience, and a new availability state should cost a
major bump, because every one of them has to be handled deliberately. (FilterPattern
inmonitrs-corecarried the attribute already, so seven public enums carry it in
total;CONTRIBUTING.mdlists all seven with their file paths.) SparklineCaret::with_noteis replaced bywith_note_segments, which takes the note
as priority-ordered segments instead of one finished string. Only the widget knows which
side of the caret the note lands on and how much room that side has, so only the widget
can decide what fits; a caller sizing one string against the row's full width produced
notes that vanished outright for most scrub positions on a 160-column Overview.Effect::SetSensorInterest(bool)is a new variant — how the visible screen tells the
sampler to read the sensors every 5 seconds instead of every 30.SensorSnapshot::hottestis deprecated rather than removed. Itsfresh()filter
returnsNonefor exactly the retained readings that became normal in this release, so
it is now a trap rather than a convenience; the header's own use of it was the first
casualty. It still compiles and still means what it always meant, and its deprecation
note names the replacement, which yields the value together with its age. Removal
waits for a later major version, per the deprecation policy inCONTRIBUTING.md.
And two changes that break nothing and still need saying:
- The JSON export's usual shape changed, and
schema_versionis still2. No field
was removed and none was renamed, so the version does not move — but the variant
frequency inverted. Sensors are stale most of the time now, so a consumer of the command
palette'sexport snapshot <path>that readssensors.temperatures.availableused to
find it on every tick and will now usually findsensors.temperatures.staleinstead,
with the reading under.valueand its age beside it. The same goes for
sensors.battery. Both shapes are recorded indocs/schema/v2.jsonand both are
guarded. Themonitrs snapshot --format jsonsubcommand is unaffected: it takes a single
sample with every tier due, so its sensors are always freshly measured. sampling.slow_intervalandsampling.medium_intervalgovern different things now,
with no key added, removed or renamed.slow_interval(30 s) sets the sensor read while
nobody is looking at a sensor panel;medium_interval(5 s) sets it only while the
Battery screen is open, where it used to set it always. A configuration file written
before1.0.0still loads and still means what it says about the tiers.
docs/configuration.mdis the record, because a frozen key
whose meaning moves is precisely the thing none of the four guards can see.
Fixed
- Below 92 columns the tab strip lost every screen name, including the active one's —
a regression0.2.0listed in its own Known limitations. The active screen keeps its
title, bracketed, and only the other six condense to bare digits, so the one piece of
chrome that says where you are survives the narrow bands. The threshold is computed from
the room the footer's key hints leave rather than compared against a constant, so it
moves with them: 92 columns with the two hints a live view shows, 100 while the timeline
is frozen and a third appears.
Known limitations
Compiled by re-reading 0.2.0's list and checking each item against the code and
docs/benchmarks.md, rather than by copying it forward — plus what
this release itself leaves open. Every figure below comes from that file.
- The idle self-CPU p95 still misses §16.1's budget, and closing it is what this release
was built for. Median 0.60–0.85% against 1%, which passes; p95 4.30–9.50%
against 2%, which does not. The previous figures were 0.5–1.1% and 6–11%, so the p95
improved and the two median ranges overlap — "barely moved" is the honest reading of the
median. With the Battery screen visible, where the sensors return to five seconds, the
median is clearly worse (1.20–1.70%) while the p95 (6.00–8.30%) does not separate
from the Overview row's at all. §16.1's gate is about idle, and idle is the Overview row. - The cause this release was designed around turned out to be the wrong one, which is
the most useful thing this section can tell you. The plan assumed the ~85 mssysinfo
temperature read on the 5-second tier was what the budget was paying for. A thread- and
process-CPU instrument built afterwards bounded that read at about 4 ms of CPU: the
85 ms is a blocking wait on in-kernel IOKit calls, and §16.1 budgets CPU. What is
measured to cost CPU is the medium tier's remaining work — its two filesystem-capacity
reads — at 13.2–35.0 ms beyond a fast-only tick, positive in 15 of 15 runs, against
a whole-tick ...
v0.2.0
[0.2.0] - 2026-07-30
Two more screens, the container facts the collector was already reading, and a way to
watch one process together with everything it spawned.
Added
Two new screens
- A CPU screen at
3, which renumbers Storage to4, Network to5and Inspect to
6. It groups the per-core meters by core class —PERFORMANCEandEFFICIENCYon
Apple Silicon, and whatever else a platform names — because four efficiency cores at 90%
with four performance cores idle is a machine doing very little, and the opposite is a
machine working hard, from the same eight numbers. The load average is shown per logical
CPU beside it, since11.4means different things on 4 cores and on 128. A machine that
reports one kind of core gets one panel, which is a fact rather than a missing feature. - A Battery screen at
7. Charge, state, time remaining only where the platform
reports it — never derived from a rate — cycle count, current capacity against design
capacity with the resulting wear, pack temperature, and draw in watts. Plus a thermal
sensor panel: those readings previously reached the screen as a single figure in the
Overview header. A machine with no battery says so once, in the panel label, rather than
filling a screen with placeholders. On macOS, cycles, capacity, wear, pack temperature
and watts all readn/a: they live in the undocumentedAppleSmartBatteryregistry
properties that §9.3 forbids this build from reading. The charge, the state and the time
remaining are real; the rest is honest absence, and the Linux collector reports all of it. - The Linux battery collector, from
/sys/class/power_supply/*on the medium tier.
docs/platform-support.mdhad been promising it. The system battery is picked by
type == Batterywithscopeabsent orSystem, which is what excludes the charger and
a wireless mouse's cell; amp-hour drivers are converted throughvoltage_min_design, and
without that the capacity stays unavailable rather than being guessed.
Follow one process with its children
Fscopes the process table to the selected process and everything beneath it, and
Fagain lifts the scope. The palette hasfollow [pid]andunfollow. The panel title
saysfollowing 410while it is on, because four rows out of a thousand with no reason on
screen is indistinguishable from a monitor that has lost the other processes.- The trailing label carries what the whole family costs —
4 of 10 total, cpu >=107%, rss 479M. That figure is the point: a build's compilers come
and go every second, so no individual row ever answers "what is this build using", and a
text filter onccfinds every compiler on the machine including someone else's. >=marks a sum as a lower bound. A member whose CPU the OS refused is still counted
as a member and still shown withdeniedin its cell, but it cannot contribute to the
total. A sum with no contributors reports the members' own state, so an all-refused
family readsdeniedrather than0%.- Membership is downwards only — the process and its descendants, never its parent — and
the root is a(pid, start key)pair, so a recycled PID cannot quietly become the thing
being followed. When the root exits, monitrs stops following and says so rather than
silently following the orphans that init inherited. Losing only children changes nothing.
Container awareness on Linux
- The cgroup CPU ceiling is shown beside the host's CPU count, never instead of it:
8 logical, 8 physical, cgroup 1.5 CPUson Inspect,cgroup 1.5 cpuin the header, and
the same figure in the CPU screen's title. A group limited to 1.5 CPUs on a 64-CPU machine
is not "2% of the machine"; it is a wall its processes are throttled against. An unlimited
group is unsupported rather than a very large number of CPUs, and acpu.maxthat cannot
be parsed is unavailable rather than "no limit", because those mean opposite things to
someone trying to explain a stall. - The cgroup's own memory charge, from
memory.current— the counter the kernel compares
againstmemory.maxwhen it decides to OOM-kill — beside the limit it is enforced against:
cgroup limit 2.0G, 512M used (25%)./proc/meminfois not namespaced, so thememory
row is the host's: a process in a 2 GiB group on a 64 GiB host reads the host's 40 GiB
and concludes it is nearly out of memory when it has used 300 MiB of its allowance. Both
halves of that ratio now come from the group, or neither does. - The container is named where the cgroup path identifies one —
container docker 3f4a1b2c9d8e, orkubernetes/containerd …— on the Inspect screen's environment row,
ahead of the evidence, so a narrow panel truncates "how we guessed" rather than "what this
is". A/.dockerenvfile is evidence of a container that names no container, and that case
stays unnamed rather than being filled in with a placeholder. - The load average is deliberately not divided by the cgroup quota.
/proc/loadavgis
not namespaced either: inside a container it counts every runnable process on the machine,
including other tenants'. The divisor stays the host's CPU count and the label becomes
over 8 host cores, so the figure is visibly about the machine. Inventing a container's
saturation by division would be worse than not having it.
The Storage screen
It had two panels and thirty blank rows. It now has four.
- A
TOP DISK I/Opanel ranking processes by read and write rate, with cumulative
totals. Two ordering rules do the work: a process whose counters were refused sorts below
a measured idle one, because a refusal is not zero; and where rates tie — which on a real
machine is nearly every process — the tie-break is the cumulative totals rather than the
PID. Without the second rule the panel is twenty-eight rows of0B/sin launch order. - Inode usage per filesystem, from
getfsstaton macOS andstatfs(2)on Linux. A
filesystem with no inode table reports unsupported, never0 of 0, and a refused read
readsdeniedrather than collapsing into then/athat means "no inode table". - Mounts that share a device are marked, with the panel stating that
SIZE is not additive. An APFS Mac lists the same disk under several mount points, and adding those
figures up is a mistake the screen used to invite. - A throughput history panel from the retained ring, labelled as the machine aggregate,
because the ring keeps no per-device series and fabricating one is not allowed.
Open files and sockets
- The process detail overlay names the descriptors, not just counts them: file, socket,
pipe, event queue, shared memory, semaphore, with the path where there is one. Capped at
256 per process, with the number not listed stated rather than the list silently ending. - Socket counts on macOS flipped from unsupported to available. The old comment claimed
a syscall per descriptor;PROC_PIDLISTFDSalready carries the type of every descriptor,
so the count costs 3 µs where reading the paths costs 216 µs — measured, on a
442-descriptor process.
Pressure escalations are announced
- A Pressure Radar signal changing state now produces a notice quoting the diagnostics
engine's own rule text, not a paraphrase. Recoveries are announced too, atinfo
severity, because a signal that goes quiet without a word looks the same as one that
stopped being collected. An unavailable signal is
never a de-escalation: it clears the remembered state, so the samples either side of a
permission error cannot be stitched into a change that never happened. diagnostics.bell_on_critical(defaultfalse) rings the terminal bell once per
sample on escalation into critical. Configuring it whilediagnostics.enabled = falseis
reported as the contradiction it is.
Changed
-
TemperatureReading::high_celsiusis nowpeak_celsius— a breaking change for
library users. It was documented as the sensor's declared high threshold, but the
underlying interface reports a high-water mark on macOS, and the two are
indistinguishable from inside the collector. A bar scaled against a high-water mark sits
at 100% forever. Onlycritical_celsiusis a declared ceiling, and it is now the only
thing a thermal scale may use. -
BatterySnapshot::healthis a method rather than a field, derived from the new
capacity, so a wear percentage can no longer disagree with the two capacities printed
beside it. Also breaking for library users. -
LinuxEnrichment::cgroup_cpu_limitreturns aMetricState<CpuQuota>instead of an
Option<CpuMax>, andMeminfoSnapshot::to_snapshottakes the group's charge alongside
its limit. The fields those changes added are in the consolidated list below. -
MemorySnapshot::effective_limit_bytesnow accepts a stale cgroup reading, which is
the one place this codebase deliberately breaks its own "fresh values for calculations"
rule. A limit is configuration, not a measurement: if the last good read said 2 GiB and
this tick's failed, falling back to the host's 64 GiB advertises 62 GiB of headroom that
does not exist.CpuSnapshot::effective_coresis its new counterpart and follows the same
rule; it is new in this release rather than changed. -
The JSON export schema is version
2.monitrs snapshot --format jsonreports
"schema_version": 2, and a consumer that checks it — which is what the field is for —
will now refuse rather than misread. Two fields were removed:
sensors.battery.healthandsensors.temperatures[].high_celsius. The second's
replacement,peak_celsius, deliberately means something else, so a script reading the
old name as a declared limit would have silently reinterpreted a high-water mark as one.
The oldhealth...
v0.1.0
0.1.0 - 2026-07-30
First release. monitrs shows what a machine is doing now and what it was doing a
few minutes ago, and it is explicit about the difference between a metric it
measured, one it is still warming up, one the OS refused, and one this platform
cannot produce.
Added
The interface
- Five screens — Overview, Processes, Storage, Network, Inspect — and six
overlays: help, command palette, filter edit, sort selector, process detail, and
the process-action confirmation that covers both signals and renice. Notices are a
panel rather than an overlay, and the spike attribution is the body of the Time Lens
screen rather than something you open. Keyboard-first throughout; the mouse is
optional and off by default. - The Time Lens.
Spacepauses the visible timeline without stopping
collection,[and]scrub through the retained history, andLreturns to
live. Selecting a historical sample shows which processes accounted for the
change, with an explicit statement of how much of the change the named processes
account for — correlation, described as correlation. - The Pressure Radar. Each signal carries the rule text that produced it, the
raw metric it was derived from, and how long the state has held. Hysteresis stops
it flapping. No rule claims a diagnosis the data cannot support: no OOM, no memory
leak, no disk failure, no malware, no thermal throttling. - A stable process table. Sorting is total, with a
(pid, start key)
tie-breaker, so a hundred rows all reporting0%keep their order between
refreshes. The cursor follows the process you chose rather than the row it was
on; until you choose one it tracks the top of the table, so the busiest process is
what a fresh session shows. Tree mode re-attaches the children of a filtered-out
parent instead of scattering them. - Four layout bands from 160×48 down to 80×24, and below that a minimal process
list rather than a broken frame. Every panel's width is reserved from the layout,
so no value can push a column out of alignment. - Strict ASCII mode (
--ascii) whose output is seven-bit by assertion, a
Unicode mode, three themes, and--color off/NO_COLOR. Colour is never the only
carrier of meaning: every state also has a distinct character, and every
unavailable value says which kind of unavailable it is —warming up,
permission denied,n/a, or a specific reason — abbreviated but never merged
as the column narrows.
What it measures
- Per-metric availability.
MetricState<T>makes "no value" a state rather
than a zero, and a retained value cannot be displayed without its age. - Native enrichment by default. On Linux,
/procand/sys: PSI, cgroup v2
limits reported separately from host totals, device busy time, per-process I/O,
and a start key with clock-tick resolution. On macOS, documented APIs only —
sysctl,host_statistics64,host_processor_info,proc_pidinfo,
getifaddrs, and IOKit's public power interfaces — with wired and compressed
pages reported separately. No external commands and no private APIs on either. - Rates are computed from measured intervals, never from an assumed one, and a
counter that went backwards reports a reset rather than a spike. - Process identity is a PID plus a start key, so a reused PID inherits nothing:
not the selection, not a pin, and not a pending action.
Doing something about it
- Signals —
xopens the dialog,Tproposes SIGTERM,Kproposes SIGKILL —
behind a confirmation that names the process. A signal never follows from one
keypress, and the forceful ones demand a distinct key rather thanEnter, so
leaning on the confirm key cannot escalate. The identity is rechecked immediately
before delivery, so a PID reused between the dialog and the write is refused
instead of signalled. - Renice (
R) on both platforms, with the same revalidation. A dry run says in
advance whether the value will be permitted; lowering a nice value needs
privileges monitrs never acquires, and it says so rather than failing silently. - Process actions are refused while the timeline is away from live, because the
process on a frozen screen is not necessarily the process that PID names now. - No privilege escalation, ever. monitrs never invokes
sudo, never asks for a
password, and reports what it could not read.
Around the edges
monitrs snapshot --format jsonwith command arguments redacted by default,
monitrs config path/init/check, shell completions for five shells, and a
manpage — the last two generated from the same definition the program parses.- Versioned TOML configuration that rejects an unknown key rather than ignoring
it, points at the exact key for an out-of-range value, suggests the near miss for
a typo, and detects key conflicts. Reload is atomic and applies to the running
interface and the sampler; the settings that need a restart are named. --debug-logon every subcommand, carrying collector durations and the
dropped/coalesced counts, and never writing to a terminal that is showing the
interface.- The terminal is restored on quit, on
Ctrl-C, onSIGTERM, onSIGHUP, on
error, and on panic — before the panic report is printed, and before slow workers
are joined. A signal reaches the ordinary shutdown path rather than killing the
process where it stands, sokillgives the terminal back instead of leaving it
needingreset. Verified on a real pty by
scripts/verify-terminal-restoration.py, which checks the escape sequences and
the pty'stermiosstate for each case.
Measured
§16.1's budgets, measured rather than asserted. Frame render, input latency and
collection come from crates/monitrs/tests/capture.rs; self CPU, resident memory
and descriptors from scripts/measure-overhead.py, which observes the running
binary from outside. Full numbers and the per-read breakdown are in
docs/benchmarks.md.
| Budget | Measured on a 12-core Mac, ~1000 processes |
|---|---|
| frame render below 16 ms at 160×48 | median 200 µs, p95 353 µs |
| input-to-visible-response below 50 ms | median 417 µs, p95 486 µs |
| sample collection below 200 ms p95 | ordinary tick p95 15–21 ms; every fifth tick 121–161 ms |
| resident memory below 50 MiB | median 29 MiB, peak 31 MiB |
| no unbounded growth | 30-minute soak: resident size fell, descriptors flat, nothing dropped |
2224 tests. Among them: 86 recorded frames of the rendered interface across three
snapshot suites, covering ASCII, Unicode, no-colour, and the empty,
permission-denied, stale and warming-up states; 120 Linux /proc and /sys fixtures
that run on every platform because each parser takes bytes rather than a path;
property tests over the formatters and the sort; 13 integration tests over the
assembled application; and a soak harness that drives the real worker threads.
Known limitations
Named because §16.1's own last line asks for measurement rather than claims:
- The idle self-CPU budget is not met. Median 1.3–2.7% against a 1% target, and
p95 11–15% against 2%, on a host with about a thousand processes — five times
§16.1's reference workload. The cost is OS reads, not monitrs' own computation,
which is three orders of magnitude smaller;docs/benchmarks.mdlocates it read by
read and says what would close it. - No twelve-hour soak is on record. A 30-minute run with the shipped collector
is, and shows no growth. The twelve-hour run is the actual gate, and nothing has
been soaked on Linux. - A slow terminal can block a frame for an unbounded time, and no instrument here can
see it: the frame-time measurement renders through ratatui'sTestBackend, which
stops short of the write to the terminal, and the soak harness has no renderer. - Per-process socket counts and device busy time are unsupported on macOS: the first
costs one syscall per descriptor, the second has no documented API. Both say
n/arather than guessing. - PSI is Linux-only, and says
n/aon macOS rather than promising a value that will
never arrive. - Timestamps are UTC and labelled
Z: no time-zone database is bundled. - Only the
aarch64-apple-darwinrelease archive has been assembled and run by
hand. The other five targets are built by CI and have not been run on their
hardware.
Changed
- Relicensed from GPL-3.0 to dual MIT OR Apache-2.0 before any release, matching
the Rust ecosystem norm and this project's dependency licence policy. Anyone who
cloned the repository at its first commit saw the earlier licence.
Verifying a download
sha256sum --check --ignore-missing SHA256SUMS
gh attestation verify monitrs-*.tar.gz --repo gaborini/monitrs