Releases: HarryPehkonen/JSOM
Release list
JSOM 3.1.2 - crash fix: deep indentation could kill the process
Crash fix: a deeply indented document could kill the process
to_json() / JsonFormatter::format() threw std::length_error when a document was nested
deeper than the indentation could fit on one line, and a process that let that exception
escape died. The cause is unsigned arithmetic, not formatting:
available_width = options_.max_line_width - line_prefix.length(); // wrapped to ~2^64The indent prefix grows with depth, so once it exceeds max_line_width the subtraction
wrapped and the line buffer asked for a reserve of about 2^64 bytes.
Reachability. Pretty (indent 2, width 100) breaks at 51 levels of nesting; Debug
(indent 4, width 80) at 21. The parser accepts up to 256 levels, so a 320-byte document
is enough — that is the input the fuzzer produced. The line is present in every release
back to 2.0.0, so if you format untrusted or generated JSON with the pretty presets, this
is worth taking.
Fixed by clamping: with no room left on the line, every element goes on its own line,
which is what "nothing else fits" means. The two branches that computed the width computed
the same value, so there is one now, and an audit of the formatter found no other unsigned
width - length subtraction that is not bounded by its own maximum.
Pinned by tests/test_formatter_options.cpp (60 levels, 12-element array at the bottom,
across Pretty/Config/Api/Debug: no throw, and the output still parses back to the same
document) and archived as fuzz/regressions/deep-nesting-exceeds-line-width.json, which the
fuzz stage now replays on every run.
How it was found
The nightly fuzz campaign crashed 48 seconds into its JSOM window (2026-09-20 21:00) and
wrote an artifact; the fuzz-report watchdog reported it the next morning. Both halves
worked. The campaign's corpus keeps the input, so the next run replays it against the fixed
build.
Also in this release
The fuzz stage now feeds fuzz/regressions/ to the fuzzer (as well as fuzz/seeds/), so
every archived finding is replayed by every gate instead of only being kept on file. 15 CI
stages green.
JSOM 3.1.1 — documentation fixes (no behaviour change)
Documentation fixes
3.1.0 shipped with documentation that had drifted from the code — a reader following the
README would have been misled in three ways:
- the README advertised a path cache that was deleted in 3.0.0, and framed numbers as
"lazy evaluation", implying they were never inspected (the §6 grammar has been enforced
since 3.1.0); - FORMATTING.md documented a
prettyoption that does not exist inJsonFormatOptions,
gavemax_depthas 100 (it is 256), listed per-preset inline sizes that disagreed with the
constants, and all nine of its example outputs were stale; - the CI stage list quoted in the README and CLAUDE.md had drifted twice.
Nothing about the code changed: this release is the documentation being corrected.
The drift now fails a gate
A new docs CI stage runs tools/check_docs.py, which fails when:
- a retired identifier (the opt-in number switch,
PathCache,StreamingParser,
NavigationResult, …) appears in documentation or headers — unless the line is explicitly
about the removal; - a generated block disagrees with the code: the
JsonFormatOptionsdefaults table, the
per-preset settings table, and nine example outputs produced by running the builtjsom
overdocs/formatting-sample.json(a real file, so the examples derive from it instead of
being typed and re-typed); - a stage list quoted in the docs differs from
CI_DEFAULT_STAGES.
tools/check_docs.py --write regenerates those parts. The examples check fails rather than
passing silently when no built binary is available — a skip that passes is a fake gate.
The pre-push hook no longer hardcodes a stage list either: the default set is defined once,
in tools/ci.sh.
Gates
15 CI stages green (233 tests, ASan/UBSan, TSan, fuzz, C++17/20/23, CLI smoke, conformance
n_ 188/188 by default, tidy, docs, pristine).
JSOM 3.1.0 — number grammar enforced by default (contains breaking changes)
This release contains breaking changes — read this table first
The version moves by a minor number, but two of these break compilation and one changes
which inputs parse. If you depend on JSOM, treat 3.1.0 like a major:
| what | 3.0.1 | 3.1.0 |
|---|---|---|
| number grammar | opt-in | enforced by default — 01, 1., -.5, 0e+, 1+2 now fail |
| option name | options.validate_numbers = true |
options.allow_loose_numbers = true for the old leniency |
| preset | ParsePresets::Validate |
ParsePresets::Loose |
| CLI | --validation=lazy / --validation=numbers |
--validation=loose / --validation=numbers (now the default; lazy is gone) |
If you only ever parsed well-formed JSON, nothing changes for you except the version number.
If you were relying on the lenient default, set allow_loose_numbers.
Why: strictness turned out to be free — and faster
-01, 1., 2.e+3, 0e+ and the rest of the malformed numbers in the nst/JSONTestSuite
corpus used to be accepted as extensions, and rejecting them cost a second pass over the
collected number text (~1 ns/digit) — which is why the switch existed at all.
The RFC 8259 §6 grammar is now checked inside the scan the parser already performs,
with tight per-state digit loops. Interleaved A/B, same probe source, median of 5, Release,
-O3 -march=native:
| input | before | after | |
|---|---|---|---|
2,000 short numbers (123.45e3) |
0.2005 ms | 0.1963 ms | 0.98x |
| 1,000 × 17-digit numbers with exponents | 0.1100 ms | 0.1015 ms | 0.92x |
| 1,000 realistic mixed records | 1.0772 ms | 1.0451 ms | 0.97x |
Two effects pulling the same way: the second pass is gone, and a digit run costs one
comparison per character instead of a 6-way test per character. So the default got
stricter and slightly faster.
// Before
JsonParseOptions options;
options.validate_numbers = true; // was required to be strict; now pointless
auto doc = parse_document(json, options);
// After: strict by default, nothing to set. For the old tolerance:
JsonParseOptions loose_options;
loose_options.allow_loose_numbers = true; // accepts 01, 1., 1eE2, 1+2
auto loose_doc = parse_document(json, loose_options);Rejections name the whole token, so the message says what the reader got wrong rather than
where the grammar stopped:
Invalid number: 01
Invalid number: 1+2
Conformance
y_ 95/95 and n_ 188/188 by default — it previously took a flag to reach 188/188
(162/188 by default). --validation=loose reports the old 162/188 for comparison.
Loose mode is narrow on purpose
It relaxes the number grammar and nothing else. Escapes, control characters, whitespace and
structure are rejected in both modes, and so are non-number tokens (+1, .5, 0x1F,
1_000).
One consequence worth knowing: to_json() still preserves number text byte-for-byte, so a
document accepted in loose mode can serialize to text a strict reader rejects (1.0. in,
1.0. out). That is inherent to accepting non-JSON forms, and the round-trip guarantee is
per configuration — under the default it now holds in its strongest form (RFC 8259 in,
RFC 8259 out, byte identical).
Tests and gates
233 tests (was 227), 23 CLI smoke checks, 14 CI stages green. Includes the 3.0.1 formatter
fixes.
JSOM 3.0.1 — formatter fixes: wrapping no longer drops commas, escape_unicode escapes codepoints
Fixes
Three bugs in the formatting engine, all found by writing the first real test suite for it
(the engine was 18.9% line-covered; the tests exercised JsonDocument::to_json() and barely
touched JsonFormatter at all).
1. Intelligent wrapping emitted invalid JSON — a corrupted document.
When a line break was inserted, the comma was left out:
[
"alpha", "bravo", "charlie", ..., "juliet"
"kilo", "lima"
]
FormatPresets::Pretty enables intelligent_wrapping with a small max_inline_array_size,
so any array of simple values long enough to wrap hit this. Config, Api and Debug
either keep traditional wrapping or inline, so they were not affected. Fixed: the comma is
written before the line break.
2. escape_unicode emitted invalid JSON for any non-ASCII text.
The escape wrote \u followed by zero padding and a raw byte — "h\u000<byte>h" — so
the output could not be parsed at all. FormatPresets::Debug enables escape_unicode, so
any document with non-ASCII text was affected. The bug's second half was worse than the
first: escaping the bytes of UTF-8 gives \u00c3\u00a4 for ä, which is valid JSON
meaning two completely different characters. Fixed: the UTF-8 sequence is decoded and the
codepoint is escaped — ä → \u00e4, 😀 → \ud83d\ude00 (surrogate pair), which
reads back as the same text.
3. Control characters were written raw.
A document holding a control character (U+0000–U+001F) produced output no JSON reader
accepts. Reachable for documents built in memory — the parser rejects raw control characters
in input. Now always escaped (\n, \t, \r, \b, \f, \u00XX for the rest),
independently of escape_unicode, exactly as the serializer does.
If you used JsonFormatOptions and parsed the result back, check your data — that is
what the first two bugs corrupt, silently.
// Affected: long arrays of simple values with intelligent_wrapping on (Pretty default)
auto bad = doc.to_json(FormatPresets::Pretty);
// Affected: any non-ASCII text with escape_unicode on (Debug default)
auto bad2 = doc.to_json(FormatPresets::Debug);new
include/jsom/utf8.hpp — a validating UTF-8 decoder (utf8::decode, utf8::encode) with
the accept/reject boundary pinned by tests: stray and truncated continuation bytes, overlong
encodings, UTF-8-encoded surrogate halves and values above U+10FFFF are rejected, and nothing
is thrown. It is what makes fix 2 correct.
Testing
The formatting engine now has real tests instead of almost none:
- invariants — formatting never changes what a document means (re-parsing the output
gives an equal document) across 5 presets × 34 documents; formatting is idempotent; every
preset emits valid JSON; Compact emits no newline; indenting presets really indent;
doc.to_json(options)andJsonFormatter{options}agree. - options — one test per documented switch: indent size, colon spacing (0/1/2), bracket
spacing including empty containers, array/object inline limits, "a container holding a
container goes multiline" (decided per container, not inherited),max_line_widthforcing
a wrap and0meaning no limit, intelligent wrapping packing several elements per line,
align_valuescolumn alignment, sorted keys,quote_keys = false,trailing_comma,
number fidelity (1.500stays1.500), preserved\uXXXX, control characters,
escape_unicodecodepoints and surrogate pairs, invalid UTF-8,max_depththrowing, and
100-level nesting. - decoder — the UTF-8 boundary case by case, plus encode/decode inverse.
| before | after | |
|---|---|---|
| test suite | 188 | 227 |
| library line coverage | 66.4% | 83.7% |
json_formatter.hpp |
18.9% | 93.5% |
All gates green: 227/227 tests, ASan+UBSan, ThreadSanitizer, C++17/C++20/C++23, 22 CLI
smoke checks, RFC 8259 conformance (y_ 95/95, n_ 188/188 with --validation=numbers),
clang-tidy 0 findings, pristine build.
Docs
README gains Writing Escapes (Formatter): what escape_unicode does, that it escapes
codepoints rather than bytes, that control characters are always escaped, and that reading
the text back exactly requires a decoding reader (ParsePresets::Unicode) because JSOM's
default parse mode deliberately keeps \uXXXX as literal text.
JSOM 3.0.0 — path cache removed (measured), const reads race-free, C++20/23 and the CLI gated
Breaking changes
The path cache is gone, and with it its public API. Every JsonDocument used to own a
three-level path cache (exact paths, prefixes, recent prefixes with 10-minute aging) that
at(), find(), exists() and at_multiple() fed on every lookup. It was measured for
the first time and it lost, so it was deleted rather than repaired.
Removed: PathCache, JsonDocument::precompute_paths(), warm_path_cache(),
clear_path_cache(), get_path_cache_stats(), NavigationEngine::navigate_with_cache(),
navigate_simple(), NavigationResult, and the --cache-warm / --cache-stats flags.
| was | now |
|---|---|
NavigationEngine::navigate_simple(root, path) |
NavigationEngine::find(root, path) |
NavigationEngine::navigate_with_cache(root, path, cache) |
NavigationEngine::find(root, path) |
doc.precompute_paths(depth) / doc.warm_path_cache(paths) |
delete the call — navigation reads the document directly |
doc.get_path_cache_stats() |
delete the call |
--cache-warm, --cache-stats |
removed; unknown --options on pointer now exit 1 instead of being ignored |
at(), find(), exists(), at_multiple(), set_at(), remove_at() and list_paths()
are unchanged in signature and semantics.
CLI: jsom pointer … now rejects unknown --options with a message and exit 1. They
used to be collected into a vector nothing read, so a typo (or a deleted flag) exited 0.
Security
Reading one document from several threads is race-free now. The cache was reached from
const methods through const_cast<JsonDocument*>(this) and lived in mutable members, so
doc.at(path) on a const document wrote to it. Four threads reading one shared document
produced 71 ThreadSanitizer data races and then a SEGV inside memmove. The test was
written first and watched fail; it is clean now, and a tsan CI stage (~6 s) keeps it that
way. The contract is documented: any number of concurrent readers, one writer — a
returned reference dies at the next mutation.
2.0.0 is superseded. That race is present in the 2.0.0 release (it predates this work),
so anyone sharing a document across threads should move to 3.0.0 rather than stay on 2.0.0.
Also gone with the cache: a raw owning pointer (new PathCache() / delete), a
process-global mutation counter bumped by every mutation, cached raw pointers into document
storage that could dangle when a child vector grew, and wall-clock eviction inside a core
data structure.
The nesting-depth guard is unchanged and verified with the input that used to kill the
process: a 60 KB document with 30,000 nested arrays now answers
Maximum nesting depth exceeded (limit 256) and exits 1.
Performance
The cache made the common access pattern slower, not faster. -O3 -march=native, 100 KB
document of 1000 records, median of 5 runs:
| access pattern | with cache | without | |
|---|---|---|---|
| every path once (7003 paths) | 33.810 ms | 1.881 ms | 17.97× faster without |
| repeat one shallow path ×10000 | 2.640 ms | 2.649 ms | unchanged |
| shared-prefix sweep (1000 leaves) | 0.303 ms | 0.282 ms | 1.07× faster without |
| write + read ×2000 | 9.547 ms | 8.334 ms | 1.15× faster without |
| repeat one 200-deep path ×10000 | 42.909 ms | 69.835 ms | 1.6× faster with |
The single case the cache helped — repeating the same deep path — is served by holding the
pointer from the first find(), and the README shows how. It also retained ~72 KB per
1000-record document and had no hit/miss counters, so its benefit had never been
measurable at all.
JSOM remains faster than nlohmann/json in the 17 paired benchmarks the repo ships
(1.26×–1.98×), measured with tools/perf_probe.cpp; object parse is the slowest shape
(~19 MB/s) and object storage is the next known target.
Testing and tooling
New gates: tsan (framework-free probe, ~6 s), std (compiles and runs
tools/std_probe.cpp as C++17, C++20 and C++23 — that claim was previously untested),
cli (22 smoke checks over the jsom binary: exit codes, validation modes, pointer
operations, rejected flags), coverage (opt-in, gcov; reports 66.4% library line
coverage, weakest being the formatting engine at 18.9%). Local CI is git-hooks-only, 12
stages, ~7 minutes.
Fixed: the conformance-corpus test resolved its data path relative to the working
directory, so ctest — which runs from the build directory — failed while the same binary
passed from the project root. The path is baked in at configure time now. Removed 18
constants that nothing referenced.
Docs
Documentation describes only what JSOM offers now. CODING_STANDARDS.md gains rule 11
(const means "changes nothing", including hidden state) and the thread/standard gates in
its checklist. The version lives in exactly one place (project(JSOM VERSION …) in
CMakeLists.txt).
JSOM 2.0.0 — RFC 8259 conformance, bounded recursion, one parser
JSOM 2.0.0 — every input is safe, the spec is enforced, one parser.
Breaking changes
- Recursion is bounded. Documents nested deeper than 256 levels are rejected with
Maximum nesting depth exceeded (limit 256)instead of exhausting the C++ stack. The
limit isJsonParseOptions::max_depth(defaultlimits::MAX_NESTING_DEPTH= 256) and
every traversal honours it: parsing, serialization (compact and pretty), comparison,
path listing and formatting. - The RFC 8259 lexical rules are enforced on every parse, whatever the options:
a raw control character (U+0000–U+001F) inside a string,\uwithout exactly four hex
digits, any escape outside" \ / b f n r t u, and formfeed/vertical tab used as
whitespace are all syntax errors. Input that previously parsed leniently may now be
rejected — that is the point. - The event-based streaming API is gone:
StreamingParser,ParseEvents,PathNode,
DocumentBuilderandparse_document_streaming().parse_document()is the single
entry point (include/jsom/parse_document.hpp).
Security
- Unbounded recursion (CWE-674) fixed. ~60 KB of nested
[was enough to kill a
process with SIGSEGV — not an exception, so nothing could catch it. That is a remote
denial of service for anything parsing JSON from off-machine, and on a 1 MB
worker-thread stack about 3 KB sufficed. Now rejected, with the limit configurable so
callers on small stacks can lower it.
Conformance
- The nst/JSONTestSuite corpus is vendored and runnable in-tree
(cmake --build build --target run_conformance), parsing each file in a forked child
so that a crash is a result rather than a lost run. y_95/95 must-accept ·n_188/188 must-reject · 0 crashes with
--validation=numbers. By defaultn_is 162/188, because the number grammar stays
lenient unless asked.- Reference point: nlohmann/json 3.11.3 scores
y_95/95,n_187/188 on the same
corpus.
Added
JsonParseOptions::max_depthfor per-parse resource limits.- Opt-in number-grammar validation:
JsonParseOptions::validate_numbers,
ParsePresets::Validate, and--validation=lazy|numbersonjsom validateand
jsom format. tools/nesting_probe.cpp,tools/perf_probe.cpp— re-runnable measurement probes, so
the numbers inOPTIMIZATIONS.mdcan be reproduced.- Single-source versioning:
project(JSOM VERSION ...)generates<jsom/version.hpp>
(JSOM_VERSION,JSOM_VERSION_MAJOR/MINOR/PATCH); nothing else hard-codes a version.
Performance
- Faster than nlohmann/json 3.11.3 in all 17 paired benchmarks (1.02×–1.98×).
- Cost of the always-on lexical rules: +1.5% on string-heavy parsing, unmeasurable
elsewhere. Number validation: +6% (short numbers), +12% (17-digit), +0.5% (mixed) —
which is why it is opt-in.
Testing
- 185 tests (unit + regression), plus the same suite under ASan+UBSan, plus a fuzz target
that drives the parser in both configurations on every input and asserts
parse(to_json(doc)) == doc. jsom versionreports the version;CODING_STANDARDS.mddefines the gates
(zero warnings, tests, sanitizers, fuzzing, measured performance).