Skip to content

Releases: xinbetween/flowlight-linux

Flowlight 0.5.13

Choose a tag to compare

@blessdyb blessdyb released this 01 Oct 08:37
780a620

0.5.12 did not fix the bug it claimed to. This does, and this time the check can fail.

What rows in the window do now

Every row carries something that came off the network — a host, a path, a process name. libadwaita reads a row's title as Pango markup unless told otherwise, so until now a request with two query parameters rendered as an empty row:

GET merino.services.mozilla.com/api/v1/suggest?q=&providers=…     ← showed nothing at all

and a path containing <span> would have been rendered rather than shown.

Why 0.5.12 missed

It set use_markup(false) on the builder. A builder applies a row's title as the object is constructed and the property afterwards, whatever order the chain lists them in — so the title was still parsed as markup, and the property still read false at the end. The check shipped with it counted those declarations: thirty-two rows, thirty-two declarations, none of them working.

Measured instead of reasoned about:

built with the builder:   uses markup: false   — and GTK logged "Failed to set text … from markup"
set after construction:   uses markup: false   — and said nothing at all

What changed

  • One helper constructs the row, sets the property, and only then sets the text. All thirty-two call sites go through it.
  • The helper moved into a small library beside the binary, so the test exercises what the window uses rather than a copy that can drift.
  • A structural check refuses any ActionRow::builder() in the window, because a builder cannot set the property first.
  • A behavioural check runs under xvfb, catches what GTK logs through g_log_structured — a plain log handler sees nothing, which is its own small lesson — and fails if a row's text could not be set. It was run against the old construction and failed there before being trusted here.

Nothing else changed. If you are on 0.5.12, this is worth taking: it is the difference between seeing your agents' requests and seeing blank rows where the interesting ones were.

v0.5.12 — What the window shows is data, not markup

Choose a tag to compare

@blessdyb blessdyb released this 01 Oct 07:58
4515b21

Found by opening the window on a real desktop, which this project had never done.

The bug, seen in situ

200 · 511 B              wget · pid 24142 · 1m ago
GET www.kernel.org/      wget · pid 24142 · 1m ago
GET api.github.com/zen   python3.12 · pid 24141 · 1m ago
                         firefox · pid 9886 · 3m ago      ← no request at all

That last row is Firefox asking merino.services.mozilla.com/api/v1/suggest?q=&providers=accuweather&…. libadwaita parses a row's title and subtitle as Pango markup unless it is told otherwise, and every row here carries text that came off the network. A query string with two parameters was enough to make the row empty.

The window had been saying so in its own log the whole time, once anybody looked:

Gtk-WARNING: Failed to set text '…' from markup due to error parsing markup: Entity did not end with
a semicolon; most likely you used an ampersand character without intending to start an entity

The half that is worse than a blank row

<span foreground="white"> in a path would have been rendered rather than shown. The window's own text was something a visited URL could write — in a program whose entire purpose is to say truthfully what a machine did.

The fix

Every row says its text is not markup, and so does the banner that carries error strings. Both setters are available at this project's libadwaita floor: v1_2 for rows, v1_3 for the banner, against the v1_5 already required for AlertDialog.

CI counts them — thirty-two rows, thirty-two declarations — and fails if those numbers ever differ. A grep is a blunt instrument, and the invariant is exactly "every row that carries data says so", which is what it checks.

One thing that is not this program's fault, written down anyway

On the machine this was found on, the window first drew its header bar and nothing else, which looks exactly like a program with nothing to say. It is GTK 4 rendering with the GPU on a machine whose GL is broken — MESA: error: ZINK: failed to choose pdev. GSK_RENDERER=cairo flowlight renders in software and shows everything, and the README now says so.

What else that session confirmed

v0.5.11's --socket-owner works where it was meant to: under the systemd unit the socket becomes the person's (srw------- blessdyb blessdyb) and the window connects, where before it was root's and the kernel refused. The window then shows real traffic, attributed, with byte counts and ages — which is the first time anybody has seen this program draw.

v0.5.11 — A window that can reach the daemon it was built for

Choose a tag to compare

@blessdyb blessdyb released this 01 Oct 07:13
8b2d0ab

Found by being asked how about the gui, and realising the window had only ever been asked its version.

The window could not work the way the README says to run it

under the service:   srw------- root root         → PermissionError [Errno 13]
started with sudo:   srw------- blessdyb blessdyb → {"ok":{"version":"0.5.9","enforcing":true,…}}

The socket is mode 0600 and belongs to SUDO_UID when there is one. Under a systemd unit there is no SUDO_UID, so it belongs to root — and the window, which runs as a person, is refused by the kernel. systemctl enable --now flowlightd is what the README tells people to do, and it was the one configuration in which the window could not work at all.

--socket-owner

# in the unit, which now carries this commented with the reason beside it:
ExecStart=/usr/sbin/flowlightd --socket-owner your-name
  • Resolved by reading /etc/passwd, not through getpwnam, for the reason the address lookup speaks DNS itself: this binary is statically linked against musl, which reads that file and does not consult NSS. Parsing it here means the answer does not depend on what it was linked against — and it can be tested against a file written out rather than against whoever happens to exist on the machine running the tests.
  • A name nobody has is refused. A socket belonging to a uid nobody has is a socket nothing can open.
  • A bare number is taken at its word, because a container with no password file is a normal place for this to run.
  • Fatal rather than warned about, because it is an explicit instruction, and the alternative is a machine that watches everything and offers no way to look at it.

Both guards now run before the kernel is touched

The database lock from v0.5.9 sat beside Store::open, which is late. The evidence was in its own output:

not intercepting: listening on 127.0.0.1:7891 …: Address in use (os error 98). Watching continues.
Error: another flowlightd (pid 19788) is already watching with …

A daemon about to refuse to run had already loaded its programs, attached its probes and tried to bind the proxy port. The lock and the socket owner are resolved immediately after the arguments now, before tracefs is read and before anything is loaded.

Verified

Six unit tests over the password-file parse and owner_from — by name, by number, comments and malformed lines stepped over, a missing file, an explicit name winning, and a name nobody has refused. The smoke test asserts that a socket asked to belong to root belongs to root, and that a name nobody has is refused.

Two of my own mistakes on the way, both caught rather than shipped: the flag's declaration never landed because the script that wrote it exited early, and the first version of the smoke check started a second daemon against the database the one under test was already watching — where v0.5.9's lock refused it, which was the feature working and the test being wrong.

v0.5.10 — A fresh index before installing what was just built

Choose a tag to compare

@blessdyb blessdyb released this 01 Oct 06:48
cfc92e2

A regression introduced in v0.5.8, found by v0.5.9's release.

E: Failed to fetch …/libgstreamer-plugins-base1.0-0_1.24.2-1ubuntu0.4_amd64.deb  404  Not Found

Installing the window's .deb pulls GTK from the archive, and the index a runner image ships with is as old as the image. Against a mirror mid-publish, the index names a pool file that is no longer there.

Why it reached us, which is the part that is mine

apt-get update was the first line of "Install the interface's libraries" — a step that went to the binaries job when v0.5.8 split the compiling out of the packaging. The new packages job was left installing against whatever the image had.

Splitting a job takes the things it quietly depended on with it. This was one of them.

CI was unaffected: its check job still refreshes the index early, and the rpm, Arch and tarball checks each run in containers that refresh their own.

What v0.5.8 did while this happened

This was the first time the resilient attach was asked for anything, and it did its job:

  • the arm64 half was installable throughout — .deb, .rpm, tarball, checksum and both window rpms were uploaded;
  • the release page carried its own notice naming the six missing files;
  • the check went red.

Re-running the one failed job completed the release and took the notice off the page again, which is the half of that mechanism nobody had seen work yet.

Before v0.5.8 the same failure threw away a complete arm64 build and published a page with nothing on it. That is what happened to v0.5.7, and is why the fix exists.

v0.5.9 — One watcher per database

Choose a tag to compare

@blessdyb blessdyb released this 01 Oct 06:24
b5552a1

Found on a real machine, while verifying something else.

The bug

Two daemons against one database is not a crash and not an error. Both attach uprobes to the same libraries, both read every call, and both write it down — so every number this reports is doubled, and nothing says so. In a tool whose whole claim is that its numbers are what happened, that is the worst shape a bug can take.

It is easy to arrive at: start the service, then run sudo flowlightd to look at something.

Which is how it surfaced. A stray daemon left over from one test doubled the output of the next on an Ubuntu 24.04 arm64 machine — three requests, twelve records — and it was nearly reported as a capture bug. The arithmetic was right both times: one daemon, one request, two records; two daemons, four.

The fix

An advisory flock on <database>.watching, taken before anything is loaded and held for as long as the process lives.

another flowlightd (pid 16538) is already watching with /var/lib/flowlight/flowlight.db. Two of them
would each read every request and write it down, which doubles every number this reports — so this one
is stopping instead. Stop that one, or name another database with --database.
  • flock, not a pid file alone. A pid file left behind by a daemon that was killed is a pid file that locks somebody out of their own machine. The kernel releases this one when the process goes, however it goes — which a test asserts.
  • The pid is written through the handle that already holds the lock. Opening the path again would be a second file description, which is a second lock, which is the thing being prevented.
  • LOCK_NB, because the point is to be told rather than to wait: a daemon that blocked there would look like one that had started.
  • Beside the database rather than in /run. The thing being protected is the database, so --database is a real way out and two daemons with two databases are not in conflict.
  • --no-store takes no lock, because it is not writing anything down.

Verified

Four unit tests in flowlight-platform: the second watcher is refused and told the pid, two databases do not conflict, and the lock is released when its holder goes.

The smoke test starts a second daemon against the running one's database and asserts both halves of the refusal — that it happens, and that it names the process, because "somebody" is not something anybody can act on.

Worth saying plainly

This was in every release before this one. It took installing on a machine and being careless enough to leave a daemon running to find it; no amount of CI across four distributions would have.

v0.5.8 — A release that attaches what it has

Choose a tag to compare

@blessdyb blessdyb released this 01 Oct 06:00
d6fd4b7

The two things v0.5.7 taught us when an arm64 runner was lost mid-build — no failed step, no log, GitHub holds no log blob for it at all.

A release now attaches what it has

attach needed every job above it, so it was skipped, and the result was worse than a red check:

  • the amd64 packages had been built, and were thrown away;
  • the release page for v0.5.7 had nothing on it;
  • the apt workflow never ran, because it waits for a Release that succeeded.

A release that exists and offers nothing is a release somebody arrives at and leaves.

It now uploads whatever was built, lists by name what was not, writes that list into the release's own notes under a marker — so a later run takes the sentence away again — and then fails. The check stays red, which is correct, and whoever comes to the page gets the half that works plus a sentence about the half that does not.

Both paths were run before being committed, against a stand-in for gh: half a release uploaded five files, named seven as missing, and exited 1; a whole one uploaded twelve, wrote no notice, and exited 0.

And the compiling is its own job

It shared a job with the packaging, so losing one runner took the packaging and every container check down with it — none of which had anything to do with the loss, and all of which had already passed on the other architecture.

job what it does timeout
binaries the musl daemon and the GTK window, per architecture 75 min
packages .deb, .rpm, tarball — and installs each where it is for. No compiler. 30 min
window the window compiled inside Fedora and openSUSE Tumbleweed 60 min
attach uploads, and says what is missing 15 min

Losing a runner now costs the expensive half alone, and gh run rerun --failed repeats only that.

Two details worth the words: an artifact does not carry the executable bit, so the packaging job puts it back — a package built from a file that is not executable installs a file that is not executable. And every job has a timeout now, because the lost one ran for forty-seven minutes before anybody could tell it was not going to finish.

Nothing in the program changed

This release is the machinery that publishes it. v0.5.7 is what runs.

v0.5.7 — What a real machine said

Choose a tag to compare

@blessdyb blessdyb released this 01 Oct 04:46
38a740d

Installed on an Ubuntu 24.04.5 arm64 machine, kernel 7.0.0, and started as a service — which is how almost everybody will run this, and the one path nothing here had ever taken. Two bugs, both in released code, and a hint the validator had been printing all along.

Interception could never start from the systemd unit

not intercepting: preparing the certificate authority: … creating /usr/local/share/flowlight:
Read-only file system (os error 30). Watching continues.

The unit sets ProtectSystem=full, so /usr is read-only to the service, and the daemon published its certificate authority to /usr/local/share/flowlight by default. The feature was unavailable the normal way round, and said so in a line of the journal nobody reads.

The unit already granted write access to exactly one directory — ConfigurationDirectory=flowlight — and its own comment said the certificate belonged there. The binary disagreed with it. The default is now /etc/flowlight.

Why nothing caught it, which is the part that needed fixing more than the path did:

  • the smoke test always passes --certificates explicitly, so the default was the one path nothing exercised;
  • CI installed the package and asserted the unit had not started, which is not the same as finding out whether it could.

CI now starts it, insists it is watching, and fails on any read-only filesystem error in its journal.

flowlightd check told a person their machine could not be watched

On Ubuntu /sys/kernel/tracing is drwx------. That refuses a person entry, so looking inside it to find out whether it exists answers "it does not" — and the report read that as "no tracefs" rather than "you are not root", then concluded the machine was unwatchable. It watches perfectly well.

The v0.5.4 fix for exactly this was verified on a runner whose directory is searchable, so CI agreed with the mistake. It is now decided by what the machine actually said: an io::ErrorKind::PermissionDenied anywhere in the error chain, and not root, is unknown.

flowlight-platform has a test that creates a directory nobody may enter and asserts a refusal and an absence are told apart — and asserts it is not running as root rather than passing quietly if it is. tracefs::mounted reads /proc/mounts as well, because the mount table is readable by everybody and is the thing that actually knows.

A hint that had been printed and ignored

Categories named three main categories, and desktop-file-validate said so: "application might appear more than once in the application menu". It exits zero for a hint, so CI was reading the status and ignoring the sentence. It now fails on any output — and earned its keep immediately by catching a second hint inside the fix for the first: Security asks for Settings or System beside it, and adding either brings back the original complaint. So Security goes. Network;Monitor is what a network monitor is.

What the machine confirmed, for the record

A static binary with no declared dependencies; installed disabled; check as root saying everything is here; probes loaded into kernel 7.0.0 and six requests read in the clear through OpenSSL (python3) and GnuTLS (wget), each attributed to the right process and pid; Firefox's NSS found inside its snap; Coverage reporting nothing unread and nothing dropped; the window resolving all 104 of its shared libraries against libadwaita 1.5.0 and GTK 4.14.5.

Four distributions' worth of CI missed both bugs, because both lived in the gap between "the package installs" and "the service runs". One machine, four minutes.

v0.5.6 — A window a launcher can find, and Arch's built from source

Choose a tag to compare

@blessdyb blessdyb released this 01 Oct 03:42
783ff0f

Roadmap item 0.5.6, and the last of the 0.5 milestone. Two gaps, both of which were written down rather than quietly left.

the desktop entry is valid
OK: on Arch the window builds from source, declares what it links against, installs with an entry a
    launcher can find, and comes off again

The window was invisible to the desktop it was written for

Installed, runnable from a command line, and in no launcher — because no package shipped a .desktop file or an icon. Every package that carries the window now carries both.

  • The entry is validated by desktop-file-validate, before it is installed and again afterwards. A .desktop file with a mistake in it is one every desktop environment ignores in silence, which is the worst way for this to be wrong.
  • The .rpm rebuilds the desktop and icon caches in %post; the .deb deliberately does not — Debian has dpkg triggers that desktop-file-utils and hicolor-icon-theme already own, so a package calling the tools itself would be doing the work twice. RPM has no equivalent.
  • The icon is a plain mark: three flows crossing a lens, one of them stopped, drawn with four shapes so that it still reads at sixteen pixels. Nothing in the packaging depends on what is in the file, only on its name — it is there to be replaced by something better rather than defended.

Arch's window builds from source

Which is what the AUR is for, and the opposite of flowlight-bin beside it. On a rolling distribution it is the only shape that works: something linking against the system's GTK, libadwaita and glibc that was built elsewhere is something that may not start, and on Arch elsewhere means "last week" as much as "on Fedora".

CI builds it with makepkg from a git archive of the working tree, so what is built is the code in the change. It installs the daemon's package first — because the window declares that it needs it, and that declaration is one of the things being checked — then asserts that the declarations name gtk4, libadwaita and flowlight, that ldd finds every library, that it answers --version, that the entry and the icon are where a launcher looks, and that removing the package takes them away again.

The release renders PKGBUILD-gui against GitHub's own tag archive, so the hash in it is the hash of the file somebody will actually fetch rather than one computed from whatever tree happened to be checked out.

One trap found by running the renderer by hand

A tarball whose version did not match the one being rendered produced arch=(flowlight-0.5.3-x86_64) — a PKGBUILD wrong in a way makepkg accepts and no architecture ever matches. It is refused now, by name. And uname -m says arm64 where Arch says aarch64, which only the local modes could ever have hit.

Where Linux support stands

daemon window
Ubuntu, Debian .deb .deb
Fedora .rpm .rpm, built on Fedora
openSUSE Tumbleweed .rpm .rpm, built on Tumbleweed
openSUSE Leap 15.6 .rpm — its libadwaita is older than 1.5
RHEL, Alma, Rocky .rpm from source
Arch PKGBUILD PKGBUILD-gui, from source
Alpine, Void, anything tarball and install.sh from source

Every row of that table is checked by CI on every push, inside a container of the distribution it names.

v0.5.5 — The window, built where it will run

Choose a tag to compare

@blessdyb blessdyb released this 01 Oct 03:14
3a45dd9

Roadmap item 0.5.5. The one package that cannot be built once and carried.

sudo rpm -i flowlight-gui-0.5.5-1.fedora41.x86_64.rpm
OK: on Fedora Linux 41 the window compiles, names that distribution's own libraries, installs, and runs
OK: on openSUSE Tumbleweed the window compiles, names that distribution's own libraries, installs, and runs

Why this one is different

The daemon is statically linked: the machine that compiles it has nothing to do with the machine that runs it. The window is the opposite — GTK, libadwaita and glibc, all the system's — so one built on Ubuntu 24.04 does not start where glibc is older. v0.5.2 said so and shipped no window outside Debian and Ubuntu, rather than shipping one that installs and does not start.

So it is compiled inside a container of each distribution it is packaged for, and RPM's own dependency scan records what it found there:

Requires: libadwaita-1.so.0()(64bit) libgtk-4.so.1()(64bit) libgio-2.0.so.0()(64bit)
          libc.so.6(GLIBC_2.34)(64bit) … flowlight = 0.5.5

Not one of those lines is written down anywhere in this repository. AutoReqProv is left alone in this spec for exactly the reason it is turned off in the daemon's: there, there is nothing to find; here, the scan is the entire point.

The package names the distribution that built it

Two files called flowlight-gui-0.5.5-1.x86_64.rpm that link against different libraries would be two files nobody can tell apart, so the release tag is filled in from /etc/os-release inside the container that did the building. The daemon's package deliberately has no such tag — the same file installs everywhere, and a .fc41 in its name would claim otherwise.

What is checked, in the order that matters

  1. It compiles there at all.
  2. libadwaita is at least 1.5 — asked before the build, so a distribution whose libadwaita is too old is named as exactly that in one sentence rather than arriving as a wall of compiler output.
  3. The scan named GTK, and named the daemon the window talks to.
  4. Installed, ldd finds every library it asks for. not found is the failure this release exists to prevent.
  5. It runs as far as --version, which is before GTK wants a display. A loader error happens before a display is ever wanted, so a container with no screen catches it.

The window learnt --version and --help on the way, answered before GTK is touched: a window that can only be asked its version by opening it is a window nobody can check from a script.

The compiler comes from rustup; the libraries come from the distribution

Fedora 41's packaged rustc is 1.91 and this workspace asks for 1.92, so the first attempt stopped before it reached a library at all — and the failure message blamed the distribution's libadwaita, which was a guess and was wrong.

The split that matters is narrower than "build it on Fedora": the point is that the window finds Fedora's GTK, libadwaita and glibc, not that Fedora's rustc compiled it.

Fedora and openSUSE Tumbleweed, and why not Leap

Leap 15.6's libadwaita is older than this window needs. That is a finding, so it is written down rather than left as a silence — and the daemon's rpm still installs on Leap, which CI has proven since v0.5.2, because it is static.

Two smaller things

  • The tarball checksums are attached now. The tarball build writes a .sha256 beside each archive and the release was not uploading it, which left the only published hash inside the PKGBUILD — fine for Arch, no use to anybody checking a download by hand.
  • Arch's window and a desktop entry are the next item. Arch's is a source PKGBUILD, which is what the AUR is for; and no package yet ships a .desktop file, which means the window does not appear in a launcher anywhere. Both recorded rather than half-done.

v0.5.4 — What this machine can and cannot do

Choose a tag to compare

@blessdyb blessdyb released this 01 Oct 01:24
bea0536

Roadmap item 0.5.4.

$ flowlightd check
ok  kernel                 6.8.0-45-generic — new enough for the probes
?   permission to load     not running as root. The daemon needs root or CAP_BPF to load anything; whether
                           this machine would allow it cannot be answered from here.
?   tracefs                /sys/kernel/tracing is mounted, and the file describing the tracepoint is
                           readable only by root — so whether its layout can be read cannot be answered
                           from here.
ok  refusing connections   cgroup v2 at /sys/fs/cgroup
ok  TLS libraries          /usr/lib/x86_64-linux-gnu/libssl.so.3 (openssl), …
ok  SELinux                SELinux is not enforcing anything here
ok  BPF in this kernel     present

Every other way of finding this out is a failure: a probe that will not attach, a rule that never bites, an empty screen. This asks the same questions in advance and says what each answer costs — the difference between "it does not work here" and "it does not work here because".

  • Read-only. It loads nothing, attaches nothing, writes nothing, and does not even open the database: a report about a machine must not leave a file on it.
  • Runnable by anybody, because the person reading it when something is wrong is not always the person with root.
  • Essential and inessential are separated. No cgroup v2 hierarchy means watching works and refusing does not, which is a line in the report rather than a failure. A kernel older than 4.18, no readable tracepoint layout, or a kernel built without CONFIG_BPF_SYSCALL are failures, and the command exits non-zero so a script can ask.

The fix that writing it demanded

CGROUP_ROOT was hard-coded to /sys/fs/cgroup. That is right on everything current and wrong on a hybrid hierarchy — RHEL 8's default, and Ubuntu's before 21.10 — where the cgroup v2 tree is one directory down at /sys/fs/cgroup/unified. Blocking there failed with a path in the message and no explanation of it.

The tree is now found rather than assumed: unified, hybrid, or neither. The third says plainly that connections cannot be refused, that watching is unaffected, and what to boot with to change it.

Being root is not the same as something being absent

The first version of this told a person running it that their machine could not be watched — because the file describing the tracepoint is readable only by root, and the report read "permission denied" as "no tracefs". Those two are a sudo and a mount command.

Asking whether the directory exists is a question anybody may ask, so that is what is asked; the layout file stays root's alone. A person gets unknown, with the reason, and no verdict about the machine on the strength of who is asking.

A crate whose functions take the root they read

flowlight-platform. hierarchy("/sys/fs/cgroup") on a machine, hierarchy(a_temporary_directory) in a test — because the only way to have a test for the hybrid hierarchy on a machine that does not have one is to write the directory out.

Seven tests do that: four release strings as four distributions actually write them (6.8.0-45-generic, 5.14.0-427.el9.x86_64, 6.6.9-arch1-1, 4.18.0-513.el8.x86_64), the 4.18 boundary from both sides, and all three cgroup layouts. It builds and its tests run anywhere, including the machines the rest of this cannot be compiled on.

Verified

Seven unit tests in the new crate. The smoke test asserts that as root every essential finding is present and positive, that no finding says nothing — a report of bare yeses is one nobody can act on — that refusing connections is reported as not essential, and that as a person the two things that cannot be answered come back unknown while the verdict stays ready.