Releases: kestrelcommerce/php-jsonc-parser
Release list
0.3.0
Fixed
Scanning is linear in the length of the document instead of quadratic.
The scanner addresses text by character offset, and resolving a character offset against a UTF-8 string means walking it from the start — so calling mb_substr() once per character made the cost grow with the square of the document length. Both the per-character reads and the per-token reads paid it. The document is now split into characters once in the scanner's constructor and indexed from there.
A parse/modify/serialize round trip, before and after:
| document | before | after |
|---|---|---|
| 8 KB | 55 ms | 9.3 ms |
| 33 KB | 720 ms | 22.4 ms |
| 90 KB | 5,307 ms | 61.4 ms |
| 183 KB | 21,384 ms | 128 ms |
Scaling is linear after the change: twice the input costs 2.1x the time. Memory cost is the character array, roughly 46x the source size, held per scanner and released with it.
Added
- Scanner complexity tests, asserting a ratio between two input sizes rather than a wall-clock budget so they carry the same meaning on any machine.
- Scanner multibyte tests covering offsets, token values, lengths, line and column numbers and the token stream across 2-, 3- and 4-byte characters, Greek, Thai, Cyrillic, CJK and emoji. The scanner previously had none, despite character-vs-byte offsets being what 0.2.0 had to fix.
Compatibility
No public API change, and no output change. All 31 storefront locale files of a production Shopify theme — including Thai, Greek, Cyrillic, Japanese, Korean and Chinese — round-trip byte-identically against 0.2.0 and this release.
Full changelog: 0.2.0...0.3.0
v0.2.0
Full Changelog: 0.1.0...0.2.0
Initial release
Full Changelog: https://github.com/kestrelwp/php-jsonc-parser/commits/0.1.0