New Layer 0 — byte→line UTF-8 decoding, owned by the library.
mod_logfile's 2 KiB buffer can truncate mid-record and chop a multi-byte
codepoint, leaving an incomplete UTF-8 sequence. Until now decoding happened
entirely in the caller, so this benign truncation was indistinguishable from
genuine binary corruption, and the documented stdin recipe (BufRead::lines)
panicked on it.
Added (public API):
- classify_utf8(&[u8]) -> Utf8Decode — pure classifier, no I/O.
- Utf8Decode { Clean, TruncatedCodepoint { at }, InvalidBytes { at } }
(#[non_exhaustive]) — truncated codepoint (recoverable) typed apart from a
byte that cannot be part of any UTF-8 sequence (corruption).
- DecodedLine { text, decode } — text is lossy-recovered (U+FFFD) so the record
survives into LogStream, where collision detection can split it back out.
- read_log_lines<R: BufRead> — owns read_until(b'\n'); real I/O errors stay
io::Error, UTF-8 invalidity is reported per line, never an io::Error.
Changed:
- Library example/doctest now uses read_log_lines instead of the strict
BufRead::lines() recipe that panicked on truncated codepoints.
- fslog warns on genuine invalid UTF-8 (recovered with U+FFFD) and stays silent
for the benign truncation; its lossy reader now shares read_log_lines.
No breaking changes — LogStream input contract is unchanged; additions only.