You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
scan(input): a lazy record index. It parses the header exactly as parse does, then yields each data record's id, name and byte span
without tokenizing its parameters; decode_record / decode_record_borrowed (or Scan::decode) parse one record on demand.
Everything between records -- trivia, ENDSEC;, the end marker -- is read
with the lexer, and a decode fails unless the record ends exactly at its
span, so if the scan and every decode succeed, parse succeeds with the
same records: a framing slip is an error, never a different record. Junk
between records, a missing ENDSEC and a second file appended after the
end marker are errors, not skipped. Measured against ifc-lite 8.1.1 on
nine STEP and IFC files (0.1-415 MB, user cycles, min of 3): the scan
costs 0.90-1.06x ifc-lite's non-validating entity scanner (1.26x on the
escape-heavy IFC4_ADD2 sample), about 1.4-1.7 GB/s on large IFC files;
scanning and decoding every record with borrowed text costs 0.69-0.97x
ifc-lite decoding every entity (1.08x on IFC4_ADD2). Peak memory for a
fully decoded file is still 1.2-1.6x ifc-lite's.
Changed
Faster tokenizer, same output. Per-byte loops use a byte-class table
instead of range compares, and the per-token lexers are forced inline into Lexer::next_spanned, which removes a copy of every token result through
the stack. Borrowed events validate names and numbers with str::from_utf8
before falling back to the lossy conversion, and instance ids are built
from the lexer's bytes without a separate UTF-8 pass. Measured against
0.7.0 on four real IFC files (user-space instructions; the build VM was
under load, so cycles are indicative): tokenizing -18..-35% instructions
and -41..-58% cycles; parse_events_borrowed -8..-11% instructions and
-6..-16% cycles. The full IFC model read is -6..-8% instructions: the
tokenizer is now about a quarter of it.
Parameter lists are allocated once, at their exact length: the parser
collects each list on a reusable scratch stack and moves it into a Vec
of the final size, instead of growing a Vec by doubling and keeping the
slack. A typed value with one parameter is boxed straight off the stack. parse_with gives the record array's growth slack back once at the end.
Output unchanged (identical to 0.6.2 on 800 real and 3,000 generated
files). On a 109 MB Revit IFC: peak memory of parse 756 -> 614 MB
(-19%), parse_parallel_with 784 -> 642 MB, scan + decode of every
record 636 -> 496 MB (-22%), at 1-2% fewer cycles.
String literals are skipped with two fast paths: a closing quote whose
next byte can neither double it nor be skipped, and a backslash whose next
byte cannot start \\, \S\ or a directive, bypass the
control-and-directive-aware matchers. Output identical to 0.6.2 on 800
real and 3,000 generated files (tokens, records, partitions, errors).