Skip to content

v1.2.0

Latest

Choose a tag to compare

@iw4p iw4p released this 11 Sep 22:47
· 2 commits to main since this release
7e5dcac

No API changes. Every input that 1.1.0 parsed successfully parses to the same
value in 1.2.0, verified by tests/test_compat_1_1_0.py, which runs the frozen
1.1.0 parser next to the current one over every prefix of a corpus of documents.

Fixed

  • Numbers with an exponent (1e5, 2.5E-3) inside an incomplete array or
    object raised JSONDecodeError. They now parse; an exponent that has not
    received its digits yet (1e, 1e-) is dropped until it is complete.
  • In strict mode an unterminated string whose tail was an incomplete escape
    ("foo\, "foo\u00) returned "", discarding text that had already
    streamed. It now returns "foo"; only the unfinished escape is held back
    (issue #8).
  • In strict mode a string cut between the two halves of a surrogate pair
    ("\ud83d, half of an emoji) returned a lone surrogate, which raises
    UnicodeEncodeError as soon as it is encoded. The high half is now held
    back until its partner arrives.
  • The JSON5 parser raised on partial literals ({"a": tr, [fals, [Inf)
    and on exponent numbers, and treated a comment that had only streamed its
    first / as an unknown token. It now behaves like the JSON parser.
  • JSON5 string decoding no longer depends on whether the optional json5
    package is installed; the same escapes (\x41, \', line continuations,
    surrogate pairs) decode the same way either way.
  • bytes and bytearray input, which json.loads accepts, no longer crash
    the fallback parser with AttributeError. A chunk that ends in the middle
    of a multi-byte UTF-8 character drops the incomplete bytes.
  • A leading UTF-8 byte-order mark no longer causes a JSONDecodeError.
  • JSONParser.strict, .on_extra_token and .last_parse_reminding are
    readable and assignable again (assigning strict on a 1.x parser was
    silently ignored), and the 0.x method names parse_string, parse_number,
    parse_array, parse_object, parse_true, parse_false, parse_null
    and parse_space are callable again.

Changed

  • The scanner works on string indexes instead of re-slicing the input at
    every token, so a parse is linear in the input size. A 900 KB partial
    document went from about 1 s to about 60 ms per parse() call.
  • A literal that is not a prefix of true/false/null (for example
    [trap], which 1.1.0 returned as [True]) now raises, matching
    json.loads. Prefixes such as [t, [tru still parse.
  • _JSON5Parser is now a subclass of the JSON parser instead of a copy of it.
  • Packaging moved to pyproject.toml with requires-python >= 3.8,
    classifiers and a py.typed marker; the package is fully type-annotated.
  • CI runs on Python 3.8 through 3.14, with and without the optional json5
    dependency.

Full diff: #10