Skip to content

Releases: openbimrs/step

openbim-step 0.9.0

Choose a tag to compare

@GeneralPawz GeneralPawz released this 26 Sep 15:55

Changed (breaking)

The owned model is reshaped to allocate less; parse output is unchanged in
content (canonical dumps identical to 0.8.0 on 4,316 files, strict and
recovering, writer output included). On three files, owned parse against
0.8.0 (user cycles, drop included): peak memory -28% (109 MB Revit IFC,
600 -> 431 MB), -34% (IFC4X3 with large coordinate lists), -32% (97 MB AP214);
cycles -23%, -18%, -30%. Heap allocations for the Revit model 7.5 M -> 2.6 M.

  • Str is the owned model's string and the new default for every S
    (Exchange, DataRecord, Record, HeaderRecord, Parameter,
    Event, EventSink), replacing String. Values of up to 22 bytes --
    nearly every number, GUID and entity name -- are stored inline without an
    allocation; longer ones are one shared Arc<str>, and parse allocates
    each distinct long name (record, typed-value and enumeration names) once
    per parse and shares it. Str derefs to str, compares equal to &str,
    and converts from &str, String and Arc<str>. The borrowed APIs keep
    Cow<'a, str> and do no interning.
  • DataRecord { id, records: Vec<Record> } is now
    DataRecord { id, instance: Instance } with
    Instance::Simple(Record) / Instance::Complex(Box<[Record]>): a simple
    instance holds its record inline instead of in a one-element Vec.
    DataRecord::records() returns the records as a slice for either shape;
    as_simple() now returns Some only for the simple form. New
    constructor DataRecord::complex.
  • A complex instance stays complex: #1=(A(1)); now writes back as
    written, where the writer used to turn a one-record complex instance into
    #1=A(1);. An empty Instance::Complex is refused by the writer.
  • Parameter lists are Box<[Parameter]> instead of Vec<Parameter>
    (Record::parameters, HeaderRecord::parameters, Parameter::List): a
    list never grows after parsing, so it carries no capacity. Constructors
    (Record::new, DataRecord::simple) accept anything Into<Box<[_]>>,
    including a Vec or an array.

Migrating: read record.records() instead of record.records; build text
with Str::from(..) or .into(); consume a boxed list by value with
.into_vec() (edition 2021's .into_iter() on a Box<[T]> iterates by
reference); compare names with name == "IFCWALL" or &*name.

openbim-step 0.8.0

Choose a tag to compare

@GeneralPawz GeneralPawz released this 26 Sep 14:21

Added

  • scan(input): a lazy record index. It parses the header exactly as
    parse does, then yields each data record's id, name and byte span
    without tokenizing its parameters; decode_record /
    decode_record_borrowed (or Scan::decode) parse one record on demand.
    Everything between records -- trivia, ENDSEC;, the end marker -- is read
    with the lexer, and a decode fails unless the record ends exactly at its
    span, so if the scan and every decode succeed, parse succeeds with the
    same records: a framing slip is an error, never a different record. Junk
    between records, a missing ENDSEC and a second file appended after the
    end marker are errors, not skipped. Measured against ifc-lite 8.1.1 on
    nine STEP and IFC files (0.1-415 MB, user cycles, min of 3): the scan
    costs 0.90-1.06x ifc-lite's non-validating entity scanner (1.26x on the
    escape-heavy IFC4_ADD2 sample), about 1.4-1.7 GB/s on large IFC files;
    scanning and decoding every record with borrowed text costs 0.69-0.97x
    ifc-lite decoding every entity (1.08x on IFC4_ADD2). Peak memory for a
    fully decoded file is still 1.2-1.6x ifc-lite's.

Changed

  • Faster tokenizer, same output. Per-byte loops use a byte-class table
    instead of range compares, and the per-token lexers are forced inline into
    Lexer::next_spanned, which removes a copy of every token result through
    the stack. Borrowed events validate names and numbers with str::from_utf8
    before falling back to the lossy conversion, and instance ids are built
    from the lexer's bytes without a separate UTF-8 pass. Measured against
    0.7.0 on four real IFC files (user-space instructions; the build VM was
    under load, so cycles are indicative): tokenizing -18..-35% instructions
    and -41..-58% cycles; parse_events_borrowed -8..-11% instructions and
    -6..-16% cycles. The full IFC model read is -6..-8% instructions: the
    tokenizer is now about a quarter of it.
  • Parameter lists are allocated once, at their exact length: the parser
    collects each list on a reusable scratch stack and moves it into a Vec
    of the final size, instead of growing a Vec by doubling and keeping the
    slack. A typed value with one parameter is boxed straight off the stack.
    parse_with gives the record array's growth slack back once at the end.
    Output unchanged (identical to 0.6.2 on 800 real and 3,000 generated
    files). On a 109 MB Revit IFC: peak memory of parse 756 -> 614 MB
    (-19%), parse_parallel_with 784 -> 642 MB, scan + decode of every
    record 636 -> 496 MB (-22%), at 1-2% fewer cycles.
  • String literals are skipped with two fast paths: a closing quote whose
    next byte can neither double it nor be skipped, and a backslash whose next
    byte cannot start \\, \S\ or a directive, bypass the
    control-and-directive-aware matchers. Output identical to 0.6.2 on 800
    real and 3,000 generated files (tokens, records, partitions, errors).

openbim-step 0.7.0

Choose a tag to compare

@GeneralPawz GeneralPawz released this 24 Sep 22:02

Added

  • parse_parallel_with(input, options, threads): parses the data section
    on several threads and returns exactly what parse_with returns -- the
    same exchange, the same diagnostics in the same order, the same error.
    Slices start at guessed record boundaries; each slice's parser must land
    exactly on its end offset, otherwise (a guess inside a string, comment or
    damaged record, or any error) the file is parsed sequentially, so errors
    and edge cases always come from the sequential parser. On seven IFC
    files (18-109 MB) 8 threads parse 3-5x faster than one; 16 threads with
    mimalloc reach 600-1000 MB/s. The worst case is one extra sequential
    parse. Parse small files with parse_with.
  • parse_events_borrowed: the same event stream as parse_events_with
    (events, order, diagnostics, errors), with text as Cow<'a, str> borrowed
    from the input wherever it needs no rewriting -- names, numbers,
    enumerations, binaries, and strings without escapes or quotes. Names keep
    their source case (the owned API upper-cases them). A consumer that
    converts every value into its own model allocates once per value instead
    of twice.

Changed

  • InstanceId stores ids of up to 22 digits inline (every u64 fits), so
    a parse no longer allocates once per id and reference. Equality, hashing,
    ordering, Debug and Display are unchanged; longer ids still work.
  • String decoding copies an escape-free body once instead of scanning and
    re-appending it, and number and name values are built without a second
    UTF-8 validation pass.
  • Faster tokenizing, output unchanged: string bodies jump to the next \ or
    ' instead of testing every byte for a print directive; whitespace is only
    checked for a directive or comment when it is followed by \ or /;
    numbers, ids and names without ignored controls are borrowed from the input
    without a second scan. On six real IFC exports (18-109 MB) the tokenizer
    runs at 407-541 MB/s, up from 183-317, and a full parse is 1.1-1.4x
    faster. Every token, record, diagnostic and error span is identical to
    0.6.2 on 800 real and 3,000 generated files.
  • New dependency: memchr (vectorized byte search).

openbim-step 0.6.2

Choose a tag to compare

@GeneralPawz GeneralPawz released this 24 Sep 19:40

Added

  • SchemaGraph::resolve_complex (#5): resolves a complex entity instance
    (#1=(A(..)B(..)C(..));, the Part 21 external mapping, ISO 10303-21:2016
    §12.2.5.3) against the schema. Returns a ComplexLayout with one
    ComplexPart per partial record, each holding only the explicit
    attributes its own entity declares (not the inherited ones, unlike
    attributes), plus the instance's full supertype closure in types.
    ComplexSlot::is_derived marks slots another type in the instance
    redeclares as derived, which a conforming file writes as *.
  • ComplexIssue: unknown partial types, partial records out of ascending
    name order, repeated partial records, and supertypes missing their own
    partial record are reported, never silently accepted. Parts that resolve
    are still returned alongside the issues.

Checked against the real data: all 314 complex instances in OCCT's AP214
test files linkrods.step (255) and screw.step (59) resolve against
AP214e3 with no issues and slot counts equal to parameter counts, and every
* falls in a slot marked derived.

Full changelog: v0.6.1...v0.6.2

cargo add openbim-step@0.6.2

openbim-step 0.6.1

Choose a tag to compare

@GeneralPawz GeneralPawz released this 24 Sep 17:32

Added

  • Opt-in reference-integrity diagnostics (#4):
    ParseOptions::check_references(true) reports every data record whose
    instance id was already defined (DiagnosticKind::DuplicateId, on the
    later record) and each distinct id a record references that no record in
    DATA defines (DiagnosticKind::DanglingReference, on the referencing
    record). Forward references are legal (ISO 10303-21:2016 §11.2) and are
    resolved at ENDSEC. Ids compare numerically, so #07 is #7, including
    ids beyond 64 bits. Works for parse_with and the streaming
    parse_events_with alike. Nothing is dropped, so these diagnostics keep
    ParseOutcome::is_lossless true.
  • Diagnostic::kind and Diagnostic::instance, and the DiagnosticKind
    enum (SkippedRecord, DuplicateId, DanglingReference).

Changed

  • ParseOutcome::is_lossless is now true unless a record was skipped;
    reference diagnostics do not count as loss. Only possible to observe with
    the new option enabled.

Measured on 100 MB files (1.0–1.5 M records, median of 5): enabling the
check adds 26–33% to parse_with. Off by default, so existing callers pay
nothing. Across the 757-file on-disk corpus it flags exactly one file:
buildingSMART's IFC4 Add2 annex example wall-elemented-case.ifc, whose
#154 references #161, which the file never defines.

openbim-step 0.6.0

Choose a tag to compare

@GeneralPawz GeneralPawz released this 24 Sep 14:32

Changed

  • Breaking: EntityDef::supertype: Option<String> is replaced by
    EntityDef::supertypes: Vec<String>, every direct supertype in
    SUBTYPE OF order. EntityDef::supertype() returns the first, for
    single-inheritance callers; with_supertype now appends. Migration:
    def.supertype.clone() becomes def.supertype().map(str::to_owned).
  • Breaking: EntityDef gained public supertypes and redeclared
    fields, so struct-literal construction must add them.

cargo semver-checks against the published 0.5.1 reports exactly these two
breaks (struct_pub_field_missing, constructible_struct_adds_field) and
182 of 184 checks passing.

Added

  • EntityDef::redeclared, EntityDef::is_redeclared, and Redeclaration:
    explicit SELF\X.a : T; redeclarations, kept apart from attributes.
  • SchemaGraph::direct_supertypes.

Fixed

  • Multiple inheritance (#2). Every supertype after the first was dropped, so
    SchemaGraph::attributes omitted inherited slots and is_a missed
    ancestors. Layouts now follow ISO 10303-21:2016 §12.2.5.2: supertypes in
    SUBTYPE OF order, higher supertypes first, and a supertype reached twice
    through a diamond counted once. supertypes, subtypes, and is_a walk
    every parent. AP242 has 248 multi-parent entities; IFC has none.
  • Explicit redeclarations no longer add a phantom positional slot (#3).
    SELF\styled_item.item : plane_or_planar_box; was parsed as a new
    attribute named SELF\styled_item.item; ISO 10303-21:2016 §12.2.8 says it
    has no effect on the encoding. This affected 481 AP242 entities, including
    advanced_face.
  • A \S\ page escape followed by an apostrophe no longer ends the string
    literal. \S\ takes exactly one following LATIN_CODEPOINT, which includes
    the apostrophe (ISO 10303-21:2016 §5.2, §6.4.3.1), so 'Stra\S\'e' is one
    string decoding to Stra§e. Previously it failed to lex. The recovery
    resynchronizer applies the same rule, and \\ is consumed as one escaped
    backslash in both so its second byte cannot open a page escape (#1).

Checked against OCCT's per-entity parameter counts (CheckNbParams in its
generated RWStep* readers): AP242 mismatches fell from 107 to 2, AP203e2
from 24 to 1. The residuals are OCCT departing from Part 21 (common_datum
diamond read twice; characterized_representation dropping derived *
slots) and are pinned in tests/schema_graph.rs.

Full changelog: v0.5.1...v0.6.0

crates.io: https://crates.io/crates/openbim-step/0.6.0

openbim-step 0.5.1

Choose a tag to compare

@GeneralPawz GeneralPawz released this 23 Sep 03:57

Added

  • SchemaGraph::direct_subtypes and SchemaGraph::subtypes: the downward
    counterpart of supertypes. subtypes(x) is every entity y != x for
    which is_a(y, x) holds, at any depth, in a deterministic sorted
    pre-order. A test checks that equivalence for every ordered entity pair.
    Answering "every IfcElement" needs this; EntityDef records only the
    upward edge, so the child index is built once at construction.