Skip to content

v0.5.9 — One watcher per database

Choose a tag to compare

@blessdyb blessdyb released this 01 Oct 06:24
· 21 commits to main since this release
b5552a1

Found on a real machine, while verifying something else.

The bug

Two daemons against one database is not a crash and not an error. Both attach uprobes to the same libraries, both read every call, and both write it down — so every number this reports is doubled, and nothing says so. In a tool whose whole claim is that its numbers are what happened, that is the worst shape a bug can take.

It is easy to arrive at: start the service, then run sudo flowlightd to look at something.

Which is how it surfaced. A stray daemon left over from one test doubled the output of the next on an Ubuntu 24.04 arm64 machine — three requests, twelve records — and it was nearly reported as a capture bug. The arithmetic was right both times: one daemon, one request, two records; two daemons, four.

The fix

An advisory flock on <database>.watching, taken before anything is loaded and held for as long as the process lives.

another flowlightd (pid 16538) is already watching with /var/lib/flowlight/flowlight.db. Two of them
would each read every request and write it down, which doubles every number this reports — so this one
is stopping instead. Stop that one, or name another database with --database.
  • flock, not a pid file alone. A pid file left behind by a daemon that was killed is a pid file that locks somebody out of their own machine. The kernel releases this one when the process goes, however it goes — which a test asserts.
  • The pid is written through the handle that already holds the lock. Opening the path again would be a second file description, which is a second lock, which is the thing being prevented.
  • LOCK_NB, because the point is to be told rather than to wait: a daemon that blocked there would look like one that had started.
  • Beside the database rather than in /run. The thing being protected is the database, so --database is a real way out and two daemons with two databases are not in conflict.
  • --no-store takes no lock, because it is not writing anything down.

Verified

Four unit tests in flowlight-platform: the second watcher is refused and told the pid, two databases do not conflict, and the lock is released when its holder goes.

The smoke test starts a second daemon against the running one's database and asserts both halves of the refusal — that it happens, and that it names the process, because "somebody" is not something anybody can act on.

Worth saying plainly

This was in every release before this one. It took installing on a machine and being careless enough to leave a daemon running to find it; no amount of CI across four distributions would have.