Skip to content

retention

sloth wiki-sync edited this page Oct 1, 2026 · 2 revisions

Retention — what is actually deleted

Summary: --db ages rows out on three tiered windows and prunes observation rows when the database file exceeds --db-max-mb. That is an investigative tradeoff, not a deletion guarantee: it is not a "30-day deletion" promise, it is not a hard disk cap, and row deletion is not secure erasure. Separately, the one artifact class that is offline-crackable — EAPOL / PMKID exports — is opt-in behind --collect-handshakes and swept on a 7-day default window (§2c). JSONL, pcap and report artifacts still have no retention at all.

Sources: src/db.c (db_maintain, prune_tier, prune_oldest_observations, db_size_bytes), src/db.h (DB_DEFAULT_RETAIN_DAYS, DB_DEFAULT_MAX_MB), src/db_schema.c, tests/test_db.c, issue #96; src/secure_file.c for §4.1 (#87); src/eapol_log.c (eapol_sweep, eapol_maintain), include/eapol_log.h (EAPOL_DEFAULT_RETENTION_DAYS), tests/test_eapol_log.c for §2c (#87).

Last updated: 2026-10-01 (sloth 1.8.2).


1. The short version

If you need one paragraph for a risk register:

Sloth's database sink deletes rows whose last_seen timestamp falls outside a per-tier window, and prunes the oldest telemetry rows when the database file grows past a configured size. Both run only while sloth is running, at most once an hour. Neither is a guaranteed deletion deadline, a guaranteed size ceiling, or a secure wipe. Handshake exports — the only offline-crackable artifact class — are not written at all unless --collect-handshakes is given, and are deleted on a 7-day default window swept at startup and daily. Every other artifact sloth can write — JSONL, pcap, reports — has no retention mechanism whatsoever and grows without bound until the operator removes it.

2. What --db retention does

Retention exists only when --db FILE is given. Without it sloth holds state in bounded in-memory ring buffers and writes nothing durable, so there is nothing to retain.

--db-retain-days N (default 30, DB_DEFAULT_RETAIN_DAYS in src/db.h) sets the base window. Three tiers are derived from it:

Tier Window Tables What it holds
observation 1× (30 d) 14 bgp_sessions, ssh_flows, rdp_flows, snmp_flows, mqtt_flows, ldap_events, kerb_events, smb_sessions, deauth_events, seqnum_correlations, twin_episodes, eapol_events, scan_entries, scan_entry_ports
entity 3× (90 d) 21 devices, pnl_clients, pnl_ssids, probe_clients, beacon_aps, beacon_ap_ssids, ssid_akm_history, wifi_aps, wifi_stas, assocs, assoc_reqs, wifi_merged, arp, dhcp_leases, top_hosts, mdns_services, nbns_names, ssdp_devices, ndp_ras, ndp_ra_prefixes, sensors
finding 12× (360 d) 5 alerts, cleartext_creds, karma_candidates, rogue_radius, btm_requests

The ordering is the point: what fired outlives who was here, which outlives the individual observations that established it.

--db-retain-days 0 does not disable retention — a value of zero or less is replaced by the 30-day default (db_set_retain_days). There is no "keep everything" setting for the age-out pass.

2.1 Age is measured from last_seen, not from first record

The pass is literally DELETE FROM <table> WHERE last_seen < cutoff. So a device, AP or credential exposure that keeps being observed is never aged out, however long ago it first appeared. "30-day retention" describes the window after a thing stops being seen, not a cap on how long a record can exist. A permanently-installed AP on a sensor that runs for a year will still be in beacon_aps after that year.

2.2 When it runs

db_maintain() is called from db_tick() only, after a successful write commit, and at most once per DB_MAINT_INTERVAL_S = 3600 s. g_last_maint is seeded on the first tick, so the first maintenance pass happens no sooner than one hour after the first successful write.

Consequences, stated plainly:

  • Retention does not run while sloth is stopped. A database left on disk for six months with sloth off is exactly as it was left.
  • A run shorter than an hour never prunes anything.
  • If a database write fails, the sink is disabled for the rest of the process (db_fail() — one line on stderr, capture continues). Once that has happened, retention has also stopped, and there is no further warning.

2.3 Two tables are never pruned

The schema has 42 tables; the three tiers cover 40. sessions (one row per run, with --site-label) and meta (schema version) are in no tier and are never deleted by retention or by the size guard. sessions therefore grows by one row per sloth run, forever.

2c. Handshake exports: the opt-in and its sweep

Owner decision, 2026-09-30 (#87). Handshake material is the one artifact class where the file itself is the attack: a PMKID or a paired M1+M2 supports offline password guessing by anyone who gets a copy. It is therefore the one class with both an opt-in and a default retention window, and both are independent of --db.

Setting Default Effect
--collect-handshakes off Required before anything crackable is written. Without it --eapol-dir exits 2 and no export directory is created.
--handshake-retention DAYS 7 Delete exported artifacts last written before the window. 0 = keep forever. Rejected outside 0..36500 rather than coerced.

What is swept. Only the names sloth itself writes, inside the directory --eapol-dir pinned at startup: eapol.22000, the per-handshake <bssid>_<sta>.pcap files, and .<name>.tmp partials left by a crash mid-write. A file you put in that directory yourself is not touched — the sweep deletes sloth's artifacts, not the directory's contents.

When it runs. Once at startup, then once per EAPOL_SWEEP_INTERVAL_S (24 h) from the poll loop. Startup as well as daily because a sensor restarted more often than once a day would otherwise never age anything out.

Granularity is the whole artifact, by mtime. That is exact for the per-handshake pcaps — one file per handshake. It is not exact for eapol.22000: that is a single run-spanning file hashcat reads whole, its mtime is its last append, and the 22000 format carries no per-line timestamp, so there is nothing to expire a line against. Consequence, stated plainly: a collection that is still appending keeps lines older than the window, and the file goes only once nothing has been added for the entire window. If per-line expiry matters for your policy, roll the file outside sloth (a dated directory per --eapol-dir, rotated by a systemd timer).

Nothing is followed, and nothing is silent. Entries are fstatat(..., AT_SYMLINK_NOFOLLOW)-ed and only regular files are unlinked, so a symlink planted at an artifact name cannot redirect a deletion out of the export directory. Such an entry is left exactly as it is; it, and any unlinkat that fails, is counted and surfaced through the same path as an export failure — one stderr line, the running count and latest reason in the [e] view header, and storage_eapol_failures in the sensor_health record. A sweep that cannot honour the window says so.

What the gate does not do. It does not blind the detector. With --collect-handshakes absent, the EAPOL-Key parser, the M1..M4 state machine, the replay-counter pairing verdict, the PTK-generation counter the FragAttacks rule reads, association evidence from M3, the [e] view and the JSONL / --db eapol_events records all behave exactly as with the gate open. Only the .22000 lines and the per-handshake pcaps are withheld — the material that is dangerous because it is on disk.

Known limit: retention needs a live collection. The sweep runs against the directory --eapol-dir validated at startup. A run without --collect-handshakes has no such directory, so material left by an earlier opted-in run is not swept — sloth no longer knows where it is. Turning collection off stops new material; it does not clean up old material. Delete it yourself, or restart with the flags and let the startup sweep do it.

3. What the size guard actually does

--db-max-mb N (default 512, DB_DEFAULT_MAX_MB; 0 = unlimited) is best read as a pruning trigger, not a cap. What it does:

  1. Measures PRAGMA page_count × PRAGMA page_size — the main database file only. The -wal and -shm sidecars are not counted, so sloth's actual on-disk footprint can exceed --db-max-mb while the guard considers the file to be under it.

  2. If over, deletes the 512 oldest rows (by last_seen) from each of the 14 observation tables, then runs PRAGMA incremental_vacuum.

  3. Repeats, at most 64 rounds per maintenance pass — an upper bound of 64 × 512 × 14 = 458 752 rows. If the file is still over target after 64 rounds, the pass simply ends, with no message, and the next hourly pass continues where it left off.

  4. If a round deletes nothing — no prunable telemetry left — it logs once and stops:

    sloth: db over --db-max-mb (N MiB > M MiB) with no prunable telemetry left; findings are never dropped
    

    The file is then allowed to stay over target indefinitely.

Entity, finding and detector-evidence rows are never dropped by this guard. That is deliberate: a sensor that fills its disk should lose telemetry, not the findings the disk was being kept for. The direct consequence is that --db-max-mb cannot be relied on as a disk-capacity control — a database that is mostly alerts and credential exposures will sail past it and say so once.

PRAGMA incremental_vacuum only returns pages to the filesystem on a database created with auto_vacuum=INCREMENTAL. Sloth requests that pragma at open, but SQLite honours it only for a new file; on a database created by an older build the pages are reused rather than returned, so the file stops growing but never shrinks. One offline VACUUM; fixes that.

3b. In-memory retention: device correlation (#94)

One in-memory table has retention of its own, and it is here because it is the one whose contents can be personal data rather than telemetry.

src/seqnum_track.c holds a per-MAC sequence-counter trail and links addresses across MAC rotations. Two horizons apply, both while sloth is running and independent of --db:

Knob Default What it bounds
--correlate-retain SECS 300 Evidence window for a correlation. A pair is reported only while both addresses have been heard inside it, counted from now.
SEQNUM_CLIENT_RETAIN_S 3600 Per-MAC trails, dropped on snapshot once this stale. Effective value is max(3600, --correlate-retain).

--no-correlate switches the linkage off entirely: trails still render, no pair is linked, exported, or written to seqnum_correlations.

Before #94 neither horizon existed — the only way out of the table was eviction once it hit 256 entries, so a quiet sensor retained every address it had ever heard for the life of the process, and stale trails kept producing current correlations. Two things follow that are worth stating for a risk register:

  • Stated purpose. Correlation exists so device counts, alert dedup and transit passes are not inflated by MAC rotation. The five-minute evidence default is set for that purpose, not for building a movement history.
  • Not a deletion guarantee, same caveat as §1. Expiry runs on snapshot while sloth is running. Anything already exported to JSONL, the data socket, or --db seqnum_correlations is governed by those sinks' retention (the observation tier, 1×, for the DB) and not by these horizons.

4. What retention does not cover

Apart from the handshake exports in §2c, nothing outside the SQLite file is managed. These grow until the operator deletes them:

Artifact Flag Retention
JSONL forensic log -o FILE none — append-only, no rotation, no size cap. A -o run writes on the order of tens of GB/day
Per-alert pcaps --pcap-dir DIR none — one file per alert flow, kept forever
EAPOL / PMKID export --eapol-dir DIR + --collect-handshakes 7 days by default — see §2c. The one managed artifact outside the database, because it is the one that is offline-crackable
Per-handshake pcaps --eapol-dir DIR + --collect-handshakes 7 days by default (§2c)
Posture reports --report, --report-json none — overwritten per run at the path you name, never aged
Packets-view manual export w key none
Wi-Fi AP snapshot --snapshot-out FILE none — overwritten per run at the path you name
SQLite WAL / SHM --db not measured by the size guard; checkpointed by SQLite, not by sloth
Filesystem copies, backups, snapshots — outside sloth entirely

If your data-handling policy needs those bounded, bound them outside sloth — logrotate, a tmpfiles.d rule, a systemd timer. Apart from the crackable material in §2c, sloth deliberately ships no deletion logic for artifacts the operator asked for by name.

4.1 File permissions on those artifacts (current behaviour)

Retention does not bound these files; their permissions at creation are enforced. This subsection describes what the code does as of 1.8.2 (src/secure_file.c, #87). It records behaviour, not a policy.

  • Created private, independent of the umask. Files are opened with openat(…, O_CREAT | O_NOFOLLOW | O_CLOEXEC, 0600). A directory sloth creates (--eapol-dir, --pcap-dir) is made with mkdir(…, 0700) and then opened O_DIRECTORY | O_NOFOLLOW. A umask can only clear bits from those modes.
  • What is checked, and how. For files and directories sloth opens itself, every check below is an fstat of the opened descriptor, not a second lookup of the path. The SQLite side files (-wal, -shm, -journal) are the exception: sloth does not open them, so an existing one is checked by lstat of the path before SQLite opens it by path. That check is not pinned to a descriptor, and a swap between the check and SQLite's open is not detected.
  • An existing path is refused, not repaired — for append and truncate targets and for directories. That covers --eapol-dir and --pcap-dir, DIR/eapol.22000, -o FILE, --db FILE and its side files, --report and --report-json. sloth refuses such a path, with a reason on stderr, when it is the wrong type, is owned by a uid other than sloth's effective uid, has any group or other permission bit (mode & 077), is a file with more than one hard link, or is a symlink as the final component. It never chmods or chowns the path. The owning gid is not checked. Parent directories of an operator-named file are neither created nor checked.
  • Exclusive-create artifacts never open an existing name. Per-alert pcaps and the packets-view w export are created O_EXCL; an existing file at the name is left untouched and sloth moves to the next suffix (_2 … _99).
  • The per-handshake pcap is replaced, not inspected. It is written to an exclusive temp file DIR/.<name>.tmp and renamed over DIR/<bssid>_<sta>.pcap. Whatever sits at that name — any owner, mode or type, a symlink included — is replaced without being checked or written through. A stale temp file of that name is unlinked first, on the basis that only sloth writes the private directory.
  • No group mode exists for the artifacts in the table below. No flag or setting makes sloth create one of them group-readable or accept a group-accessible existing one. Whether a group mode should exist is an open question on #87; nothing here answers it. The table also lists one export that does not go through these checks at all (--snapshot-out).
Artifact Created as Write If refused
--eapol-dir DIR, --pcap-dir DIR directory 0700 — startup stops. --eapol-dir additionally needs --collect-handshakes (§2c), checked before the directory is created
DIR/eapol.22000 0600 append; a failed write is truncated back to the previous length counted, shown in the EAPOL view header; later exports still attempted
DIR/<bssid>_<sta>.pcap 0600 exclusive temp file in DIR, renamed over the old name a temp-file create, write or rename failure is counted and shown in the EAPOL view header; the previous capture stays
Per-alert pcaps in --pcap-dir 0600 exclusive create, _2 … _99 on a same-second name clash counted, retried next tick
Packets-view w export 0600 exclusive create in the working directory, same suffixing export fails
-o FILE JSONL 0600 append startup stops
--db FILE 0600, created before SQLite opens it (SQLITE_OPEN_NOFOLLOW where available) SQLite startup stops
FILE-wal, FILE-shm, FILE-journal by SQLite, mode copied from the main file SQLite an existing one failing the checks stops startup
--report, --report-json 0600 validated, then truncated and rewritten report skipped; the old file is left as it was
--snapshot-out FILE not covered — plain fopen(path, "w"): mode is 0666 & ~umask (0644 under the usual 022, so group- and world-readable), follows a symlink, existing file not validated truncated and rewritten write fails with one stderr line

Detail for the EAPOL export, including the refusal message, is in docs/views/eapol.md.

5. Deletion is logical, not secure erasure

DELETE FROM … marks pages free inside the database file. It does not overwrite the bytes. Sloth never sets PRAGMA secure_delete, so a stock SQLite build leaves the old content in place until those pages are reused. Specifically, after retention has "deleted" a row:

  • the bytes may remain in free pages of the main file;
  • the pre-delete page images may remain in the -wal file;
  • any backup, filesystem snapshot or copied file taken earlier is untouched;
  • incremental_vacuum returns free pages to the filesystem but does not zero them, and the filesystem does not zero them either.

Treat an aged-out database as not queryable through sloth's schema, not as erased. If an artifact must be unrecoverable, destroy the media or the file with a tool built for that; sloth does not claim to.

6. Known gaps

Recorded here rather than implied away:

  • No filesystem capacity control. Sloth does not check free space before writing, and there is no "cannot persist" error surfaced to the operator when the volume fills — the write simply fails and the sink disables itself with one stderr line. Tracked as remaining work on issue #96; not implemented as of 1.8.2.
  • The size guard measures the wrong number for a disk budget — the main database file only, excluding -wal/-shm — and gives up for the hour after 64 pruning rounds (§3).
  • Retention is process-local. It runs only while sloth runs. For the handshake sweep that also means it needs a live opted-in collection: material from an earlier run is not swept by a run that omits --collect-handshakes (§2c).
  • eapol.22000 expires whole-file, not per line. The format carries no per-line timestamp, so a file still being appended to keeps lines older than the window (§2c).
  • --snapshot-out bypasses the #87 file checks. It is written with plain fopen, so its mode follows the umask and a symlink at the path is followed (§4.1). Not fixed as of 1.8.2; a #87 follow-up.

Related pages

  • sqlite-schema — the tables themselves, and the MISSION §2 guardrails the schema enforces.
  • pcap-export — per-alert pcap files and their permissions.
  • log — the JSONL stream.
  • posture-report — --report output.

Clone this wiki locally