Repository navigation
UTF-8 validation for str format values #69
Replies: 2 comments
Design decisions resolvedThe three open design decisions have been resolved and the document updated to reflect them:
Additionally, a follow-up section was added: after this work merges, an issue will be opened to track designing an RFC for adding UTF-8 validation to the Pony stdlib's |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
The problem
The MessagePack spec defines the str format family as storing
"a byte array as a UTF-8 string." The library currently treats
str values as opaque byte sequences -- the encoder accepts any
ByteSeqin its str methods, the decoders return aStringconstructed directly from the raw bytes via
String.from_iso_arrayorString.from_array, and neitherside checks whether the bytes are valid UTF-8.
This is a spec compliance gap. A caller reading the spec would
expect str values to contain valid UTF-8, and the library does
nothing to enforce or verify that. The spec does acknowledge
the issue:
So the spec explicitly leaves validation policy to the
implementation, but recommends that deserializers provide access
to the raw bytes. This means the library needs a deliberate
policy rather than the current accidental permissiveness.
The gap affects both sides:
MessagePackDecoderoperates onbuffered.Readerand returnsString iso^viaString.from_iso_array.MessagePackZeroCopyDecoderoperates onZeroCopyReaderandreturns
String valviaString.from_array(zero-copy path).Neither performs UTF-8 validation. The streaming decoder
(
MessagePackStreamingDecoder) delegates toMessagePackZeroCopyDecoderinternally.MessagePackEncoder.str()and all format-specificstr methods accept any
ByteSeq. A caller can encode arbitrarybinary data using the str format, producing messages that
violate the spec's intent. The spec provides the bin format
family specifically for raw binary data.
Pony's
Stringtype is a raw byte array with no UTF-8guarantee.
String.from_iso_arrayandString.from_arrayreusethe underlying data pointer without any validation. The
utf32method can decode individual codepoints and returns
0xFFFD(theUnicode replacement character) with a length of 1 for invalid
byte sequences, but there is no dedicated
is_valid_utf8methodin the stdlib. Validation would require walking the string byte
by byte using
utf32and checking for replacement characters, orimplementing a standalone UTF-8 validator.
How other libraries handle this
Python msgpack
The
rawparameter controls whether str format values aredecoded as Python
bytes(raw bytes) or Pythonstr(Unicode).The default is
raw=False, which decodes str values as Pythonstrby applying UTF-8 decoding. If the bytes are not validUTF-8, a
UnicodeDecodeErroris raised unless the caller setsunicode_errorsto a non-strict handler (e.g.,'replace'or'ignore').The
strict_map_keyparameter (defaultTrue) restricts mapkeys to specific types. For encoding,
use_bin_type=True(default since 1.0) encodes Python
stras msgpack str andPython
bytesas msgpack bin.The result: Python msgpack validates UTF-8 by default on decode,
with an escape hatch for invalid data. On encode, the type
system naturally separates str from bytes.
JavaScript @msgpack/msgpack
By default, str format values are decoded as JavaScript strings
using
TextDecoder, which performs UTF-8 decoding. TherawStrings: truedecode option skips UTF-8 decoding and returnsUint8Arrayinstead ofstringfor str values.JavaScript's
TextDecoderreplaces invalid UTF-8 sequences withthe replacement character by default, or can be configured to
throw on invalid input via the
fataloption.Rust rmp / rmp-serde
The
rmpcrate provides low-levelread_strandread_str_reffunctions that decode str format values as Rust
&str, whichrequires valid UTF-8. Attempting to decode invalid UTF-8 returns
a
DecodeStringError::InvalidUtf8error.For cases where the data might contain invalid UTF-8, the crate
provides
RawandRawRefwrapper types that can hold either avalid
Stringor raw bytes with an associated UTF-8 error.Callers choose which type to deserialize into based on whether
they expect valid UTF-8.
The result: Rust enforces UTF-8 validity by default because the
language's
strtype requires it. TheRaw/RawReftypes arethe escape hatch for working with potentially invalid data.
Go vmihailenco/msgpack
Go's
stringtype is a byte sequence with no UTF-8 guarantee(similar to Pony). The library decodes str format values directly
into Go strings without UTF-8 validation. Go provides
utf8.Valid()in the stdlib for callers who want to validateafter decoding, but the msgpack library does not call it.
Summary
Libraries in languages with distinct string/bytes types (Python,
Rust) enforce UTF-8 by default. Libraries in languages where
strings are byte arrays (Go) do not. Pony falls into the latter
category.
Design
Decision: opt-in validation
UTF-8 validation should be opt-in, not the default. Reasons:
Backward compatibility. The library has been returning
unvalidated strings since its creation. Changing the default
to validate would break existing callers that rely on the
current behavior, including callers handling pre-2013 msgpack
data where the same wire formats (now called str) carried
raw bytes before the str/bin split was introduced.
Pony's type system does not distinguish validated strings.
Unlike Rust, where
&strand&[u8]are different types,Pony's
Stringis always a byte array. Validation would bea runtime check with no compile-time representation of the
result. The caller receives a
Stringeither way.Performance. UTF-8 validation requires examining every
byte of every decoded string. For workloads that decode many
strings but do not require UTF-8 correctness (e.g., binary
protocol keys that happen to be ASCII), mandatory validation
is pure overhead.
Precedent. Go's msgpack library, the closest analog to
Pony's situation, also does not validate by default.
Where validation lives
Validation is available at all decoder layers and the encoder:
MessagePackDecoder: New methods alongside the existingstr methods.
MessagePackZeroCopyDecoder: Same new methods, mirroringthe existing parity between the two decoder primitives.
MessagePackStreamingDecoder: A configuration option thatenables validation for all str values decoded via
next().validation in the "decode then validate" pattern.
Decoder API: validating str methods
Add a parallel set of str decoding methods that validate UTF-8
after reading the bytes. The naming convention appends
_utf8todistinguish them from the existing non-validating methods.
For
MessagePackDecoder(operates onbuffered.Reader, returnsString iso^):For
MessagePackZeroCopyDecoder(operates onZeroCopyReader,returns
String val):Each
_utf8variant calls the corresponding existing method todecode the string, then validates the result. If validation
fails, the method errors. This means the reader has already
consumed the bytes -- the caller cannot recover the raw bytes
after a validation failure through these methods.
The alternative -- returning a union like
(String iso^ | Array[U8] iso^)-- would change the return typeand complicate the common case where the caller just wants a
valid string or an error.
UTF-8 validation implementation
The validator uses
String.utf32to walk the string byte bybyte and detect invalid sequences.
utf32returns(0xFFFD, 1)for invalid sequences -- the replacement characterwith a consumed length of 1. A legitimately encoded U+FFFD
character (bytes
0xEF 0xBF 0xBD) returns(0xFFFD, 3). Thelength field disambiguates.
The validator is a public primitive so callers can use the
"decode then validate" pattern (see "Accessing raw bytes on
decode failure" below):
This implementation depends on
String.utf32's error-reportingconvention. This is how Pony represents invalid UTF-8 throughout
the stdlib. The implementation should document this dependency
at the point of use.
Note: Pony's stdlib does not currently provide a built-in UTF-8
validation method on
String. After this work is merged, afollow-up issue will be opened to design an RFC for adding
UTF-8 validation to the stdlib's
Stringtype. If the stdlibgains native validation in the future, this library should
migrate to use it and potentially deprecate
MessagePackValidateUTF8.Streaming decoder configuration
The streaming decoder supports a configuration option for UTF-8
validation of str values. The constructor currently accepts a
limitsparameter;validate_utf8is added alongside it:When
_validate_utf8istrue, the_decode_fixstrand_decode_strmethods validate the decoded string beforereturning it. These methods delegate to
MessagePackZeroCopyDecoderinternally; validation occurs afterthe zero-copy decoder returns the
String val. On validationfailure, they return
InvalidUtf8.This is a constructor-time decision rather than per-call because
the streaming decoder already has a simple
next(): DecodeResultinterface with no per-call options. Adding a parameter to
nextwould change the API for all callers. A constructor flag keeps
the interface clean.
Streaming decoder error type for UTF-8 failures
A new
InvalidUtf8primitive is added to theDecodeResultunion, distinct from
InvalidData.InvalidDatameans "the stream contains an invalid MessagePackformat byte -- the stream is corrupt and decoding should stop."
UTF-8 validation failure is categorically different: the
MessagePack framing is perfectly valid, only the string content
violates the spec's UTF-8 expectation. The stream is not corrupt
and decoding can safely continue.
Conflating these under
InvalidDatawould force callers to useout-of-band knowledge to distinguish two failure modes with
different recovery semantics. Distinct semantics deserve distinct
representations.
The updated
DecodeResultunion:Encoder validation
The encoder adds
_utf8variants of the str encoding methodsthat validate before writing, leaving the existing methods
unchanged:
The encoder methods accept
ByteSeq, which is(String val | Array[U8] val). The validation implementationhandles both cases: for
String val, the validator usesString.utf32directly viaMessagePackValidateUTF8; forArray[U8] val, the bytes are wrapped in aStringviaString.from_arraybefore validation.Accessing raw bytes on decode failure
The spec recommends that deserializers provide access to the
original byte array. The current non-validating methods already
satisfy this -- they return the raw bytes as a
String. Forcallers using the validating methods who also want the raw bytes
on failure, the pattern would be:
_utf8variant.bytes (but this requires re-reading from the buffer, which
is not possible after the bytes have been consumed).
This is a limitation. Once a validating method errors, the bytes
are consumed and gone. The recommended mitigation is the "decode
then validate" pattern -- callers who need the raw bytes on
failure should use the non-validating method and then validate
the result themselves using
MessagePackValidateUTF8:This is the more Pony-idiomatic approach -- the caller controls
the flow and decides what to do with invalid data.
Follow-up
After this work is merged, open an issue on this repository to
track designing an RFC for adding UTF-8 validation to the Pony
stdlib's
Stringtype. UTF-8 validation is a general-purposeneed that does not belong permanently in a msgpack library. If
the stdlib gains
String.is_valid_utf8(or equivalent), thislibrary should migrate to use it.
Scope
This proposal covers:
_utf8variants of str decoding methods onMessagePackDecoder_utf8variants of str decoding methods onMessagePackZeroCopyDecodervalidate_utf8constructor option onMessagePackStreamingDecoderInvalidUtf8primitive in theDecodeResultunion_utf8variants of str encoding methods onMessagePackEncoderMessagePackValidateUTF8primitive for caller-sidevalidation
This proposal does not cover:
non-validating behavior is preserved)
this without runtime wrappers)
bin format methods for raw byte data)
with U+FFFD rather than erroring) -- this could be a separate
future enhancement
All reactions