v2.0.6 — True zero-allocation, 2-3× faster, format unified
Major correctness + performance pass. Every logging path is now genuinely zero-alloc, parallel scales near-linearly, and the binary log format is unified across logger types.
Highlights
| Bench (Apple M5) | v2.0.5 | v2.0.6 |
|---|---|---|
UltimateLogger |
45 ns / 128B / 1 alloc | 16 ns / 0 / 0 |
UltimateLogger parallel |
56 ns / 128B / 1 alloc | 5 ns / 0 / 0 |
StructuredLogger |
100 ns / 512B / 1 alloc | 36 ns / 0 / 0 |
StructuredLogger parallel |
54 ns / 512B / 1 alloc | 10 ns / 0 / 0 |
Structured + 5 fields |
113 ns / 512B / 1 alloc | 45 ns / 0 / 0 |
Structured + 10 fields |
145 ns / 512B / 1 alloc | 77 ns / 0 / 0 |
Structured → TerminalWriter (real-world) |
~150 ns | 27 ns / 0 / 0 |
Parallel is now faster than serial on several benches because the cross-core contention point was removed.
Bug fixes
-
Zero-alloc claim was not actually true on any path. Two unrelated bugs combined:
- The "small message stack buffer" optimization in
Logger.log,StructuredLogger.logFields, andUltimateLogger.logdeclaredvar stackBuf [N]byteand passed it towriter.Write. BecauseWriteis an interface method, escape analysis forced the array to the heap on every call (confirmed via-gcflags=-m=2). The "fast path" was strictly slower than the pool path it tried to skip. - The custom
leadingZeros64inbuffer_pool_generic.gowas broken — it summed eight 6-bit lookup values where seven of eight always returned 64, so it returned 512 for any input. The pool'sPutrejected every buffer because of the bogus index calculation, so everyGetallocated fresh.
Both fixed; all logger paths now hit the pool's reuse path.
- The "small message stack buffer" optimization in
-
Wall-clock timestamps were wrong.
runtime.nanotime()(monotonic since process start) was being treated as Unix nanoseconds in the binary header, so rendered timestamps came out as e.g.1970-01-01T21:35:59. Switched to a once-captured wall+mono offset (~5 ns/log, portable across darwin/linux/windows). -
AsyncWriterV2was racy under concurrent writers.RingBuffer.Putwas single-producer-only, butAsyncWriterV2.Writeis the package'sio.Writerand gets called from many goroutines. Two producers could overwrite the same slot. Fixed with CAS on head. -
MMapWriterhad a wrap-around race.offset.Addfollowed byoffset.Storecould interleave across writers and corrupt entries near the end of the file. Fixed with a CAS-loop on offset. -
Goroutine-per-flush in mmap writers removed.
MS_ASYNC(Linux/macOS) andFlushViewOfFile(Windows) are non-blocking; the previousgo w.syncRange(...)was both racy and a goroutine fan-out hazard. -
LogfmtWritercorrupted UTF-8.appendQuotedranged over runes then wrotebyte(c), truncating any multi-byte sequence to its low byte. Fixed; full UTF-8 round-trips. -
Float formatting was lossy. Hand-rolled formatter only emitted three decimals, mishandled NaN/±Inf, overflowed at 2⁶⁴. Replaced with
strconv.AppendFloat.
Performance
- Atomic sequence counter dropped. It was never read by any writer in the package, but every
Add(1)was the dominant cross-core cache-line contention point — the entire reason parallel was 2.5× slower than serial. Removed. - Direct-text fast path in
StructuredLogger. When the writer is*TerminalWriter, format text directly into the pooled buffer and skip the binary encode → re-decode round trip. Type assertion is ~1 ns; saves ~30-40 ns. The most common case (humans reading colored logs) is now also the fastest. - Native byte order in field encoding. Each int/float was eight individual byte stores. Now a single
*(*uint64)(unsafe.Pointer(&buf[pos])) = f.numper numeric field. The binary form is internal — only this package's writers consume it — so big-endian bought nothing. TerminalWritertimestamp cache.time.AppendFormatruns once per second, not once per log. Manual digit formatting via shift+mask.- Pre-built level prefix LUTs.
[6][]byte{"\x1b[32mINFO \x1b[0m", ...}so the line prefix is one append, not three. - 256-byte classifier table for escape detection. Branchless OR-loop; Go compiles byte-LUT loads to NEON
tbl/ SSSE3pshufbon supported architectures. defer mu.UnlockinTerminalWriter.Writereplaced with explicit unlock (~6 ns/log).
Breaking changes
The on-disk / on-the-wire binary log format changed:
old structured: magic(4) ver(1) lvl(1) seq(8) ts(8) msgLen(1) msg ...
new (unified): magic(4) ver(1) lvl(1) ts(8) msgLen(2) msg [fieldCount(1) fields...]
- 16-byte header instead of 22 (no sequence counter).
msgLenwidened fromuint8touint16.- Field values are stored in native byte order. Anything reading binary logs across machines with different endianness needs to handle this; in practice only
TerminalWriterandLogfmtWriterdecode the binary form, both in-package.
If you've persisted binary logs from a previous version and want to render them, do it before upgrading.
Other changes
- README rewritten — fresh benchmark numbers, accurate description of how zero-alloc actually works (the previous "stack-allocated buffers" claim was the bug we just fixed).
runtime.exitreplaced byos.Exitfor cross-compilation support (carried over from earlier release).
Compatibility
- Go 1.23+ (unchanged from v2.0.5).
- Public API surface is unchanged —
Logger,StructuredLogger,UltimateLogger, allFieldconstructors, all writers behave identically. Only the on-the-wire binary format changed.