Skip to content

v2.0.6 — True zero-allocation, 2-3× faster, format unified

Choose a tag to compare

@semihalev semihalev released this 25 Apr 13:26
· 3 commits to main since this release

Major correctness + performance pass. Every logging path is now genuinely zero-alloc, parallel scales near-linearly, and the binary log format is unified across logger types.

Highlights

Bench (Apple M5) v2.0.5 v2.0.6
UltimateLogger 45 ns / 128B / 1 alloc 16 ns / 0 / 0
UltimateLogger parallel 56 ns / 128B / 1 alloc 5 ns / 0 / 0
StructuredLogger 100 ns / 512B / 1 alloc 36 ns / 0 / 0
StructuredLogger parallel 54 ns / 512B / 1 alloc 10 ns / 0 / 0
Structured + 5 fields 113 ns / 512B / 1 alloc 45 ns / 0 / 0
Structured + 10 fields 145 ns / 512B / 1 alloc 77 ns / 0 / 0
Structured → TerminalWriter (real-world) ~150 ns 27 ns / 0 / 0

Parallel is now faster than serial on several benches because the cross-core contention point was removed.

Bug fixes

  • Zero-alloc claim was not actually true on any path. Two unrelated bugs combined:

    1. The "small message stack buffer" optimization in Logger.log, StructuredLogger.logFields, and UltimateLogger.log declared var stackBuf [N]byte and passed it to writer.Write. Because Write is an interface method, escape analysis forced the array to the heap on every call (confirmed via -gcflags=-m=2). The "fast path" was strictly slower than the pool path it tried to skip.
    2. The custom leadingZeros64 in buffer_pool_generic.go was broken — it summed eight 6-bit lookup values where seven of eight always returned 64, so it returned 512 for any input. The pool's Put rejected every buffer because of the bogus index calculation, so every Get allocated fresh.

    Both fixed; all logger paths now hit the pool's reuse path.

  • Wall-clock timestamps were wrong. runtime.nanotime() (monotonic since process start) was being treated as Unix nanoseconds in the binary header, so rendered timestamps came out as e.g. 1970-01-01T21:35:59. Switched to a once-captured wall+mono offset (~5 ns/log, portable across darwin/linux/windows).

  • AsyncWriterV2 was racy under concurrent writers. RingBuffer.Put was single-producer-only, but AsyncWriterV2.Write is the package's io.Writer and gets called from many goroutines. Two producers could overwrite the same slot. Fixed with CAS on head.

  • MMapWriter had a wrap-around race. offset.Add followed by offset.Store could interleave across writers and corrupt entries near the end of the file. Fixed with a CAS-loop on offset.

  • Goroutine-per-flush in mmap writers removed. MS_ASYNC (Linux/macOS) and FlushViewOfFile (Windows) are non-blocking; the previous go w.syncRange(...) was both racy and a goroutine fan-out hazard.

  • LogfmtWriter corrupted UTF-8. appendQuoted ranged over runes then wrote byte(c), truncating any multi-byte sequence to its low byte. Fixed; full UTF-8 round-trips.

  • Float formatting was lossy. Hand-rolled formatter only emitted three decimals, mishandled NaN/±Inf, overflowed at 2⁶⁴. Replaced with strconv.AppendFloat.

Performance

  • Atomic sequence counter dropped. It was never read by any writer in the package, but every Add(1) was the dominant cross-core cache-line contention point — the entire reason parallel was 2.5× slower than serial. Removed.
  • Direct-text fast path in StructuredLogger. When the writer is *TerminalWriter, format text directly into the pooled buffer and skip the binary encode → re-decode round trip. Type assertion is ~1 ns; saves ~30-40 ns. The most common case (humans reading colored logs) is now also the fastest.
  • Native byte order in field encoding. Each int/float was eight individual byte stores. Now a single *(*uint64)(unsafe.Pointer(&buf[pos])) = f.num per numeric field. The binary form is internal — only this package's writers consume it — so big-endian bought nothing.
  • TerminalWriter timestamp cache. time.AppendFormat runs once per second, not once per log. Manual digit formatting via shift+mask.
  • Pre-built level prefix LUTs. [6][]byte{"\x1b[32mINFO \x1b[0m", ...} so the line prefix is one append, not three.
  • 256-byte classifier table for escape detection. Branchless OR-loop; Go compiles byte-LUT loads to NEON tbl / SSSE3 pshufb on supported architectures.
  • defer mu.Unlock in TerminalWriter.Write replaced with explicit unlock (~6 ns/log).

Breaking changes

The on-disk / on-the-wire binary log format changed:

old structured: magic(4) ver(1) lvl(1) seq(8) ts(8) msgLen(1) msg ...
new (unified):  magic(4) ver(1) lvl(1) ts(8) msgLen(2) msg [fieldCount(1) fields...]
  • 16-byte header instead of 22 (no sequence counter).
  • msgLen widened from uint8 to uint16.
  • Field values are stored in native byte order. Anything reading binary logs across machines with different endianness needs to handle this; in practice only TerminalWriter and LogfmtWriter decode the binary form, both in-package.

If you've persisted binary logs from a previous version and want to render them, do it before upgrading.

Other changes

  • README rewritten — fresh benchmark numbers, accurate description of how zero-alloc actually works (the previous "stack-allocated buffers" claim was the bug we just fixed).
  • runtime.exit replaced by os.Exit for cross-compilation support (carried over from earlier release).

Compatibility

  • Go 1.23+ (unchanged from v2.0.5).
  • Public API surface is unchanged — Logger, StructuredLogger, UltimateLogger, all Field constructors, all writers behave identically. Only the on-the-wire binary format changed.