Skip to content

openbim-step 0.7.0

Choose a tag to compare

@GeneralPawz GeneralPawz released this 24 Sep 22:02
· 12 commits to main since this release

Added

  • parse_parallel_with(input, options, threads): parses the data section
    on several threads and returns exactly what parse_with returns -- the
    same exchange, the same diagnostics in the same order, the same error.
    Slices start at guessed record boundaries; each slice's parser must land
    exactly on its end offset, otherwise (a guess inside a string, comment or
    damaged record, or any error) the file is parsed sequentially, so errors
    and edge cases always come from the sequential parser. On seven IFC
    files (18-109 MB) 8 threads parse 3-5x faster than one; 16 threads with
    mimalloc reach 600-1000 MB/s. The worst case is one extra sequential
    parse. Parse small files with parse_with.
  • parse_events_borrowed: the same event stream as parse_events_with
    (events, order, diagnostics, errors), with text as Cow<'a, str> borrowed
    from the input wherever it needs no rewriting -- names, numbers,
    enumerations, binaries, and strings without escapes or quotes. Names keep
    their source case (the owned API upper-cases them). A consumer that
    converts every value into its own model allocates once per value instead
    of twice.

Changed

  • InstanceId stores ids of up to 22 digits inline (every u64 fits), so
    a parse no longer allocates once per id and reference. Equality, hashing,
    ordering, Debug and Display are unchanged; longer ids still work.
  • String decoding copies an escape-free body once instead of scanning and
    re-appending it, and number and name values are built without a second
    UTF-8 validation pass.
  • Faster tokenizing, output unchanged: string bodies jump to the next \ or
    ' instead of testing every byte for a print directive; whitespace is only
    checked for a directive or comment when it is followed by \ or /;
    numbers, ids and names without ignored controls are borrowed from the input
    without a second scan. On six real IFC exports (18-109 MB) the tokenizer
    runs at 407-541 MB/s, up from 183-317, and a full parse is 1.1-1.4x
    faster. Every token, record, diagnostic and error span is identical to
    0.6.2 on 800 real and 3,000 generated files.
  • New dependency: memchr (vectorized byte search).