Repository navigation
Zero-copy decoding #66
Replies: 5 comments
|
documentation would be very important to this. to let folks know a zero copy version exists, but you have to use our special buffer class rather than Reader and that could impact interop with other classes. |
|
i don't see the need to a special zero copy streaming one. that really is an implementation detail. we can just make the streaming one always use the zero copy buffering. |
|
Re: Standalone reader section, which notes:
Array is the right choice here. The O(1) shift advantage of List doesn't matter in practice because the chunk count is small — network I/O delivers data in reasonably sized chunks, and they get consumed roughly as fast as they arrive. Looking at the full access pattern from the sketch:
The dominant operations (front read, front update, sequential scan) all favor Array or are neutral. Shift is the only one that favors List, and with small n it's not meaningful. |
|
Re: Chunk consolidation option section. Going with multi-chunk with fallback. No copy on input, zero-copy when the read falls within a single chunk (common case with typical network I/O), copy fallback when spanning chunks — which is no worse than what Consolidation is a potential future optimization. Opened #74 to track it. |
Uh oh!
There was an error while loading. Please reload this page.
The problem
Every decode of a variable-length value (str, bin, ext) copies the data out of the reader's buffer. The stdlib
Reader.block(len)allocates a newArray[U8]of sizelenand copies bytes into it from internal chunks — even when the requested range falls entirely within a single chunk. For the msgpack decoder, this means every decoded string, binary blob, and extension value triggers an allocation and memcpy proportional to its size.For small values (field names, short strings), the cost is negligible. For workloads that decode large payloads — multi-KB binary data, long strings, bulk ext values — the copies add up. Zero-copy decoding eliminates this overhead by returning views into the reader's existing buffer instead of copies.
How other libraries handle this
MessagePack-CSharp
TryReadStringSpanreturns aReadOnlySpan<byte>that references the input buffer directly. Only works when the data falls within a contiguous segment of the input; returnsfalseotherwise, and the caller falls back to a copying read.msgpack-ruby
Uses a size threshold. Below
read_reference_thresholdbytes, the decoder copies (because small copies are cheaper than reference management overhead). Above the threshold, it returns a reference into the input buffer. The threshold is configurable.rmp-serde (Rust)
Zero-copy via the
ReadSlicetrait. The deserializer can borrow&[u8]from the input slice. The deserialized values have lifetime bounds tied to the input — they can't outlive the buffer they reference. Requires the input to be a contiguous slice.Common pattern
All implementations share the same constraint: zero-copy only works when the requested data is contiguous in the input buffer. When data spans buffer boundaries, a copy is unavoidable.
Why this is natural in Pony
Pony's stdlib provides two building blocks that make zero-copy straightforward:
Array[U8] val.trim(from, to)returns a shared portion of the array. Both the original and the new array are immutable (val), as they share memory. The operation does not allocate a new array pointer nor copy elements. From the stdlib documentation:String.from_array(data: Array[U8] val)creates aString valthat reuses the underlying data pointer of the input array. No copy.Chained together, these give a complete zero-copy path from input buffer to decoded value:
No data copies at any step. Both
trimandfrom_arrayallocate small header objects (array/string metadata), but the element data is shared — no memcpy of the payload bytes.Why the stdlib Reader can't do this
The stdlib
Reader.block(len)returnsArray[U8] iso^. Theisocapability means isolated — the caller has exclusive access, and no other reference to that memory can exist. This is useful when the caller needs to mutate the result, but it meansblock()must copy data into a freshly allocated array rather than sharing the reader's internal buffer.The reader internally stores chunks as
Array[U8] val(immutable, shareable). The data is already in the right capability for sharing —block()copies it out intoisounnecessarily for callers who only needval.A zero-copy reader that returns
Array[U8] valinstead ofArray[U8] iso^can usetrimdirectly on its internal chunks, avoiding the copy entirely.Design
Standalone reader
A new
ZeroCopyReaderclass in the msgpack package. Not a replacement forbuffered.Reader— a separate type with its own API. The existingMessagePackDecoderandMessagePackStreamingDecodercontinue using the stdlibReaderunchanged.The reader uses chunk-based storage similar to the stdlib
Readerbut returnsvalinstead ofiso^fromblock(). The stdlib uses a doubly-linkedListfor O(1) head removal; the sketch above usesArrayfor simplicity but an implementation should considerListfor the same reason. It also needs the same supporting methods the msgpack decoder uses:u8(),u16_be(),u32_be(),u64_be(),i8(),i16_be(),i32_be(),i64_be(),f32_be(),f64_be(),peek_u8(offset),peek_u16_be(offset),peek_u32_be(offset),size(), andskip(n).For the fixed-size numeric reads (u8, u16_be, etc.), zero-copy doesn't matter — they read 1-8 bytes. These can copy.
Decoder integration
The existing
MessagePackDecodertakesbuffered.Reader(a concrete class). Without interface abstraction, a new decoder primitive is needed for the zero-copy reader:The decode logic is identical to
MessagePackDecoder— only the reader type and return capability differ. The key change isString.from_array(b.block(len)?)instead ofString.from_iso_array(b.block(len)?).This duplication is the cost of not abstracting over reader types. The alternative — a common interface — is deferred per the substitutability decision.
Which formats benefit
Only variable-length formats benefit from zero-copy:
String iso^String valArray[U8] iso^Array[U8] val(U8, Array[U8] val)Fixed-size formats (nil, bool, integers, floats) read at most 9 bytes. Zero-copy overhead (trim bookkeeping) would exceed the copy cost at these sizes. These methods can copy unconditionally.
Single-chunk vs. multi-chunk
The zero-copy fast path only works when the requested range falls within a single internal chunk. When it spans chunks, the reader must copy. How often the fast path hits depends on how data arrives:
Real-world network I/O typically delivers data in chunks of hundreds to thousands of bytes, so the fast path should hit for the majority of values in typical use.
Chunk consolidation option
An alternative to accepting multi-chunk fallbacks: consolidate on append. When the caller appends data, copy it into a single growing buffer. This is one copy on input, but then all reads are guaranteed zero-copy (single chunk).
The tradeoff: one copy on input (into the consolidated buffer) vs. potentially zero copies on read. Whether this wins depends on how many values are decoded from each appended chunk. If many values are decoded per chunk, consolidation amortizes the input copy across many zero-copy reads. If few values are decoded per chunk, it may be worse than the current approach.
There's a subtlety:
_buffermust bevalfortrimto work. But we're appending to it, which requires mutability. The pattern would be: accumulate into areforisoarray, then freeze it tovalbefore serving reads. This means consolidation happens in batches — append data, freeze, serve reads, append more, freeze again. Each freeze produces a new val array that covers the latest data.Streaming decoder variant
A zero-copy streaming decoder would pair the
ZeroCopyReaderwith the streaming decode logic. The existingMessagePackStreamingDecoderusesbuffered.Readerinternally and returnsMessagePackValue(which includesString valandArray[U8] val). A zero-copy variant would useZeroCopyReaderinstead and return the same types — but theString valandArray[U8] valvalues would reference the reader's buffer instead of being independent copies.The caller-facing API would be identical. The only observable difference is performance and the lifetime relationship between decoded values and the reader's buffer.
Lifetime considerations
Zero-copy means decoded values hold references to the reader's internal chunks. As long as any decoded value is live, the chunk it references stays in memory (Pony's GC handles this —
valreferences are traced normally).In the normal case, this is fine — decoded values are processed and discarded, and the chunks become eligible for GC. But if the caller holds decoded values for a long time (e.g., caching them), the entire original chunk stays alive even if only a few bytes from it are referenced. This is the classic "small slice holds large buffer alive" problem.
For most msgpack use cases (decode, process, discard), this isn't a concern. For callers who cache decoded values, the tradeoff is explicit: use zero-copy for throughput, accept that cached values pin their source chunks.
Scope
This proposal covers:
ZeroCopyReaderclass in the msgpack packageMessagePackZeroCopyDecoderprimitive (or equivalent) for decoding from itArray.trimandString.from_arrayThis proposal does not cover:
buffered.Reader(deferred)MessagePackDecoderorMessagePackStreamingDecoderAll reactions