Skip to content

JSOM 3.0.1 — formatter fixes: wrapping no longer drops commas, escape_unicode escapes codepoints

Choose a tag to compare

@HarryPehkonen HarryPehkonen released this 19 Sep 17:31
· 24 commits to main since this release

Fixes

Three bugs in the formatting engine, all found by writing the first real test suite for it
(the engine was 18.9% line-covered; the tests exercised JsonDocument::to_json() and barely
touched JsonFormatter at all).

1. Intelligent wrapping emitted invalid JSON — a corrupted document.
When a line break was inserted, the comma was left out:

[
  "alpha", "bravo", "charlie", ..., "juliet"
  "kilo", "lima"
]

FormatPresets::Pretty enables intelligent_wrapping with a small max_inline_array_size,
so any array of simple values long enough to wrap hit this. Config, Api and Debug
either keep traditional wrapping or inline, so they were not affected. Fixed: the comma is
written before the line break.

2. escape_unicode emitted invalid JSON for any non-ASCII text.
The escape wrote \u followed by zero padding and a raw byte — "h\u000<byte>h" — so
the output could not be parsed at all. FormatPresets::Debug enables escape_unicode, so
any document with non-ASCII text was affected. The bug's second half was worse than the
first: escaping the bytes of UTF-8 gives \u00c3\u00a4 for ä, which is valid JSON
meaning two completely different characters. Fixed: the UTF-8 sequence is decoded and the
codepoint is escaped — ä → \u00e4, 😀 → \ud83d\ude00 (surrogate pair), which
reads back as the same text.

3. Control characters were written raw.
A document holding a control character (U+0000–U+001F) produced output no JSON reader
accepts. Reachable for documents built in memory — the parser rejects raw control characters
in input. Now always escaped (\n, \t, \r, \b, \f, \u00XX for the rest),
independently of escape_unicode, exactly as the serializer does.

If you used JsonFormatOptions and parsed the result back, check your data — that is
what the first two bugs corrupt, silently.

// Affected: long arrays of simple values with intelligent_wrapping on (Pretty default)
auto bad  = doc.to_json(FormatPresets::Pretty);
// Affected: any non-ASCII text with escape_unicode on (Debug default)
auto bad2 = doc.to_json(FormatPresets::Debug);

new

include/jsom/utf8.hpp — a validating UTF-8 decoder (utf8::decode, utf8::encode) with
the accept/reject boundary pinned by tests: stray and truncated continuation bytes, overlong
encodings, UTF-8-encoded surrogate halves and values above U+10FFFF are rejected, and nothing
is thrown. It is what makes fix 2 correct.

Testing

The formatting engine now has real tests instead of almost none:

  • invariants — formatting never changes what a document means (re-parsing the output
    gives an equal document) across 5 presets × 34 documents; formatting is idempotent; every
    preset emits valid JSON; Compact emits no newline; indenting presets really indent;
    doc.to_json(options) and JsonFormatter{options} agree.
  • options — one test per documented switch: indent size, colon spacing (0/1/2), bracket
    spacing including empty containers, array/object inline limits, "a container holding a
    container goes multiline" (decided per container, not inherited), max_line_width forcing
    a wrap and 0 meaning no limit, intelligent wrapping packing several elements per line,
    align_values column alignment, sorted keys, quote_keys = false, trailing_comma,
    number fidelity (1.500 stays 1.500), preserved \uXXXX, control characters,
    escape_unicode codepoints and surrogate pairs, invalid UTF-8, max_depth throwing, and
    100-level nesting.
  • decoder — the UTF-8 boundary case by case, plus encode/decode inverse.
before after
test suite 188 227
library line coverage 66.4% 83.7%
json_formatter.hpp 18.9% 93.5%

All gates green: 227/227 tests, ASan+UBSan, ThreadSanitizer, C++17/C++20/C++23, 22 CLI
smoke checks, RFC 8259 conformance (y_ 95/95, n_ 188/188 with --validation=numbers),
clang-tidy 0 findings, pristine build.

Docs

README gains Writing Escapes (Formatter): what escape_unicode does, that it escapes
codepoints rather than bytes, that control characters are always escaped, and that reading
the text back exactly requires a decoding reader (ParsePresets::Unicode) because JSOM's
default parse mode deliberately keeps \uXXXX as literal text.