JSOM 3.0.1 — formatter fixes: wrapping no longer drops commas, escape_unicode escapes codepoints
Fixes
Three bugs in the formatting engine, all found by writing the first real test suite for it
(the engine was 18.9% line-covered; the tests exercised JsonDocument::to_json() and barely
touched JsonFormatter at all).
1. Intelligent wrapping emitted invalid JSON — a corrupted document.
When a line break was inserted, the comma was left out:
[
"alpha", "bravo", "charlie", ..., "juliet"
"kilo", "lima"
]
FormatPresets::Pretty enables intelligent_wrapping with a small max_inline_array_size,
so any array of simple values long enough to wrap hit this. Config, Api and Debug
either keep traditional wrapping or inline, so they were not affected. Fixed: the comma is
written before the line break.
2. escape_unicode emitted invalid JSON for any non-ASCII text.
The escape wrote \u followed by zero padding and a raw byte — "h\u000<byte>h" — so
the output could not be parsed at all. FormatPresets::Debug enables escape_unicode, so
any document with non-ASCII text was affected. The bug's second half was worse than the
first: escaping the bytes of UTF-8 gives \u00c3\u00a4 for ä, which is valid JSON
meaning two completely different characters. Fixed: the UTF-8 sequence is decoded and the
codepoint is escaped — ä → \u00e4, 😀 → \ud83d\ude00 (surrogate pair), which
reads back as the same text.
3. Control characters were written raw.
A document holding a control character (U+0000–U+001F) produced output no JSON reader
accepts. Reachable for documents built in memory — the parser rejects raw control characters
in input. Now always escaped (\n, \t, \r, \b, \f, \u00XX for the rest),
independently of escape_unicode, exactly as the serializer does.
If you used JsonFormatOptions and parsed the result back, check your data — that is
what the first two bugs corrupt, silently.
// Affected: long arrays of simple values with intelligent_wrapping on (Pretty default)
auto bad = doc.to_json(FormatPresets::Pretty);
// Affected: any non-ASCII text with escape_unicode on (Debug default)
auto bad2 = doc.to_json(FormatPresets::Debug);new
include/jsom/utf8.hpp — a validating UTF-8 decoder (utf8::decode, utf8::encode) with
the accept/reject boundary pinned by tests: stray and truncated continuation bytes, overlong
encodings, UTF-8-encoded surrogate halves and values above U+10FFFF are rejected, and nothing
is thrown. It is what makes fix 2 correct.
Testing
The formatting engine now has real tests instead of almost none:
- invariants — formatting never changes what a document means (re-parsing the output
gives an equal document) across 5 presets × 34 documents; formatting is idempotent; every
preset emits valid JSON; Compact emits no newline; indenting presets really indent;
doc.to_json(options)andJsonFormatter{options}agree. - options — one test per documented switch: indent size, colon spacing (0/1/2), bracket
spacing including empty containers, array/object inline limits, "a container holding a
container goes multiline" (decided per container, not inherited),max_line_widthforcing
a wrap and0meaning no limit, intelligent wrapping packing several elements per line,
align_valuescolumn alignment, sorted keys,quote_keys = false,trailing_comma,
number fidelity (1.500stays1.500), preserved\uXXXX, control characters,
escape_unicodecodepoints and surrogate pairs, invalid UTF-8,max_depththrowing, and
100-level nesting. - decoder — the UTF-8 boundary case by case, plus encode/decode inverse.
| before | after | |
|---|---|---|
| test suite | 188 | 227 |
| library line coverage | 66.4% | 83.7% |
json_formatter.hpp |
18.9% | 93.5% |
All gates green: 227/227 tests, ASan+UBSan, ThreadSanitizer, C++17/C++20/C++23, 22 CLI
smoke checks, RFC 8259 conformance (y_ 95/95, n_ 188/188 with --validation=numbers),
clang-tidy 0 findings, pristine build.
Docs
README gains Writing Escapes (Formatter): what escape_unicode does, that it escapes
codepoints rather than bytes, that control characters are always escaped, and that reading
the text back exactly requires a decoding reader (ParsePresets::Unicode) because JSOM's
default parse mode deliberately keeps \uXXXX as literal text.