JSOM 3.1.0 — number grammar enforced by default (contains breaking changes)
This release contains breaking changes — read this table first
The version moves by a minor number, but two of these break compilation and one changes
which inputs parse. If you depend on JSOM, treat 3.1.0 like a major:
| what | 3.0.1 | 3.1.0 |
|---|---|---|
| number grammar | opt-in | enforced by default — 01, 1., -.5, 0e+, 1+2 now fail |
| option name | options.validate_numbers = true |
options.allow_loose_numbers = true for the old leniency |
| preset | ParsePresets::Validate |
ParsePresets::Loose |
| CLI | --validation=lazy / --validation=numbers |
--validation=loose / --validation=numbers (now the default; lazy is gone) |
If you only ever parsed well-formed JSON, nothing changes for you except the version number.
If you were relying on the lenient default, set allow_loose_numbers.
Why: strictness turned out to be free — and faster
-01, 1., 2.e+3, 0e+ and the rest of the malformed numbers in the nst/JSONTestSuite
corpus used to be accepted as extensions, and rejecting them cost a second pass over the
collected number text (~1 ns/digit) — which is why the switch existed at all.
The RFC 8259 §6 grammar is now checked inside the scan the parser already performs,
with tight per-state digit loops. Interleaved A/B, same probe source, median of 5, Release,
-O3 -march=native:
| input | before | after | |
|---|---|---|---|
2,000 short numbers (123.45e3) |
0.2005 ms | 0.1963 ms | 0.98x |
| 1,000 × 17-digit numbers with exponents | 0.1100 ms | 0.1015 ms | 0.92x |
| 1,000 realistic mixed records | 1.0772 ms | 1.0451 ms | 0.97x |
Two effects pulling the same way: the second pass is gone, and a digit run costs one
comparison per character instead of a 6-way test per character. So the default got
stricter and slightly faster.
// Before
JsonParseOptions options;
options.validate_numbers = true; // was required to be strict; now pointless
auto doc = parse_document(json, options);
// After: strict by default, nothing to set. For the old tolerance:
JsonParseOptions loose_options;
loose_options.allow_loose_numbers = true; // accepts 01, 1., 1eE2, 1+2
auto loose_doc = parse_document(json, loose_options);Rejections name the whole token, so the message says what the reader got wrong rather than
where the grammar stopped:
Invalid number: 01
Invalid number: 1+2
Conformance
y_ 95/95 and n_ 188/188 by default — it previously took a flag to reach 188/188
(162/188 by default). --validation=loose reports the old 162/188 for comparison.
Loose mode is narrow on purpose
It relaxes the number grammar and nothing else. Escapes, control characters, whitespace and
structure are rejected in both modes, and so are non-number tokens (+1, .5, 0x1F,
1_000).
One consequence worth knowing: to_json() still preserves number text byte-for-byte, so a
document accepted in loose mode can serialize to text a strict reader rejects (1.0. in,
1.0. out). That is inherent to accepting non-JSON forms, and the round-trip guarantee is
per configuration — under the default it now holds in its strongest form (RFC 8259 in,
RFC 8259 out, byte identical).
Tests and gates
233 tests (was 227), 23 CLI smoke checks, 14 CI stages green. Includes the 3.0.1 formatter
fixes.