Releases: openbimrs/step
Release list
openbim-step 0.9.0
Changed (breaking)
The owned model is reshaped to allocate less; parse output is unchanged in
content (canonical dumps identical to 0.8.0 on 4,316 files, strict and
recovering, writer output included). On three files, owned parse against
0.8.0 (user cycles, drop included): peak memory -28% (109 MB Revit IFC,
600 -> 431 MB), -34% (IFC4X3 with large coordinate lists), -32% (97 MB AP214);
cycles -23%, -18%, -30%. Heap allocations for the Revit model 7.5 M -> 2.6 M.
Stris the owned model's string and the new default for everyS
(Exchange,DataRecord,Record,HeaderRecord,Parameter,
Event,EventSink), replacingString. Values of up to 22 bytes --
nearly every number, GUID and entity name -- are stored inline without an
allocation; longer ones are one sharedArc<str>, andparseallocates
each distinct long name (record, typed-value and enumeration names) once
per parse and shares it.Strderefs tostr, compares equal to&str,
and converts from&str,StringandArc<str>. The borrowed APIs keep
Cow<'a, str>and do no interning.DataRecord { id, records: Vec<Record> }is now
DataRecord { id, instance: Instance }with
Instance::Simple(Record)/Instance::Complex(Box<[Record]>): a simple
instance holds its record inline instead of in a one-elementVec.
DataRecord::records()returns the records as a slice for either shape;
as_simple()now returnsSomeonly for the simple form. New
constructorDataRecord::complex.- A complex instance stays complex:
#1=(A(1));now writes back as
written, where the writer used to turn a one-record complex instance into
#1=A(1);. An emptyInstance::Complexis refused by the writer. - Parameter lists are
Box<[Parameter]>instead ofVec<Parameter>
(Record::parameters,HeaderRecord::parameters,Parameter::List): a
list never grows after parsing, so it carries no capacity. Constructors
(Record::new,DataRecord::simple) accept anythingInto<Box<[_]>>,
including aVecor an array.
Migrating: read record.records() instead of record.records; build text
with Str::from(..) or .into(); consume a boxed list by value with
.into_vec() (edition 2021's .into_iter() on a Box<[T]> iterates by
reference); compare names with name == "IFCWALL" or &*name.
openbim-step 0.8.0
Added
scan(input): a lazy record index. It parses the header exactly as
parsedoes, then yields each data record's id, name and byte span
without tokenizing its parameters;decode_record/
decode_record_borrowed(orScan::decode) parse one record on demand.
Everything between records -- trivia,ENDSEC;, the end marker -- is read
with the lexer, and a decode fails unless the record ends exactly at its
span, so if the scan and every decode succeed,parsesucceeds with the
same records: a framing slip is an error, never a different record. Junk
between records, a missingENDSECand a second file appended after the
end marker are errors, not skipped. Measured against ifc-lite 8.1.1 on
nine STEP and IFC files (0.1-415 MB, user cycles, min of 3): the scan
costs 0.90-1.06x ifc-lite's non-validating entity scanner (1.26x on the
escape-heavy IFC4_ADD2 sample), about 1.4-1.7 GB/s on large IFC files;
scanning and decoding every record with borrowed text costs 0.69-0.97x
ifc-lite decoding every entity (1.08x on IFC4_ADD2). Peak memory for a
fully decoded file is still 1.2-1.6x ifc-lite's.
Changed
- Faster tokenizer, same output. Per-byte loops use a byte-class table
instead of range compares, and the per-token lexers are forced inline into
Lexer::next_spanned, which removes a copy of every token result through
the stack. Borrowed events validate names and numbers withstr::from_utf8
before falling back to the lossy conversion, and instance ids are built
from the lexer's bytes without a separate UTF-8 pass. Measured against
0.7.0 on four real IFC files (user-space instructions; the build VM was
under load, so cycles are indicative): tokenizing -18..-35% instructions
and -41..-58% cycles;parse_events_borrowed-8..-11% instructions and
-6..-16% cycles. The full IFC model read is -6..-8% instructions: the
tokenizer is now about a quarter of it. - Parameter lists are allocated once, at their exact length: the parser
collects each list on a reusable scratch stack and moves it into aVec
of the final size, instead of growing aVecby doubling and keeping the
slack. A typed value with one parameter is boxed straight off the stack.
parse_withgives the record array's growth slack back once at the end.
Output unchanged (identical to 0.6.2 on 800 real and 3,000 generated
files). On a 109 MB Revit IFC: peak memory ofparse756 -> 614 MB
(-19%),parse_parallel_with784 -> 642 MB, scan + decode of every
record 636 -> 496 MB (-22%), at 1-2% fewer cycles. - String literals are skipped with two fast paths: a closing quote whose
next byte can neither double it nor be skipped, and a backslash whose next
byte cannot start\\,\S\or a directive, bypass the
control-and-directive-aware matchers. Output identical to 0.6.2 on 800
real and 3,000 generated files (tokens, records, partitions, errors).
openbim-step 0.7.0
Added
parse_parallel_with(input, options, threads): parses the data section
on several threads and returns exactly whatparse_withreturns -- the
same exchange, the same diagnostics in the same order, the same error.
Slices start at guessed record boundaries; each slice's parser must land
exactly on its end offset, otherwise (a guess inside a string, comment or
damaged record, or any error) the file is parsed sequentially, so errors
and edge cases always come from the sequential parser. On seven IFC
files (18-109 MB) 8 threads parse 3-5x faster than one; 16 threads with
mimalloc reach 600-1000 MB/s. The worst case is one extra sequential
parse. Parse small files withparse_with.parse_events_borrowed: the same event stream asparse_events_with
(events, order, diagnostics, errors), with text asCow<'a, str>borrowed
from the input wherever it needs no rewriting -- names, numbers,
enumerations, binaries, and strings without escapes or quotes. Names keep
their source case (the owned API upper-cases them). A consumer that
converts every value into its own model allocates once per value instead
of twice.
Changed
InstanceIdstores ids of up to 22 digits inline (everyu64fits), so
a parse no longer allocates once per id and reference. Equality, hashing,
ordering,DebugandDisplayare unchanged; longer ids still work.- String decoding copies an escape-free body once instead of scanning and
re-appending it, and number and name values are built without a second
UTF-8 validation pass. - Faster tokenizing, output unchanged: string bodies jump to the next
\or
'instead of testing every byte for a print directive; whitespace is only
checked for a directive or comment when it is followed by\or/;
numbers, ids and names without ignored controls are borrowed from the input
without a second scan. On six real IFC exports (18-109 MB) the tokenizer
runs at 407-541 MB/s, up from 183-317, and a fullparseis 1.1-1.4x
faster. Every token, record, diagnostic and error span is identical to
0.6.2 on 800 real and 3,000 generated files. - New dependency:
memchr(vectorized byte search).
openbim-step 0.6.2
Added
SchemaGraph::resolve_complex(#5): resolves a complex entity instance
(#1=(A(..)B(..)C(..));, the Part 21 external mapping, ISO 10303-21:2016
§12.2.5.3) against the schema. Returns aComplexLayoutwith one
ComplexPartper partial record, each holding only the explicit
attributes its own entity declares (not the inherited ones, unlike
attributes), plus the instance's full supertype closure intypes.
ComplexSlot::is_derivedmarks slots another type in the instance
redeclares as derived, which a conforming file writes as*.ComplexIssue: unknown partial types, partial records out of ascending
name order, repeated partial records, and supertypes missing their own
partial record are reported, never silently accepted. Parts that resolve
are still returned alongside the issues.
Checked against the real data: all 314 complex instances in OCCT's AP214
test files linkrods.step (255) and screw.step (59) resolve against
AP214e3 with no issues and slot counts equal to parameter counts, and every
* falls in a slot marked derived.
Full changelog: v0.6.1...v0.6.2
cargo add openbim-step@0.6.2
openbim-step 0.6.1
Added
- Opt-in reference-integrity diagnostics (#4):
ParseOptions::check_references(true)reports every data record whose
instance id was already defined (DiagnosticKind::DuplicateId, on the
later record) and each distinct id a record references that no record in
DATAdefines (DiagnosticKind::DanglingReference, on the referencing
record). Forward references are legal (ISO 10303-21:2016 §11.2) and are
resolved atENDSEC. Ids compare numerically, so#07is#7, including
ids beyond 64 bits. Works forparse_withand the streaming
parse_events_withalike. Nothing is dropped, so these diagnostics keep
ParseOutcome::is_losslesstrue. Diagnostic::kindandDiagnostic::instance, and theDiagnosticKind
enum (SkippedRecord,DuplicateId,DanglingReference).
Changed
ParseOutcome::is_losslessis now true unless a record was skipped;
reference diagnostics do not count as loss. Only possible to observe with
the new option enabled.
Measured on 100 MB files (1.0–1.5 M records, median of 5): enabling the
check adds 26–33% to parse_with. Off by default, so existing callers pay
nothing. Across the 757-file on-disk corpus it flags exactly one file:
buildingSMART's IFC4 Add2 annex example wall-elemented-case.ifc, whose
#154 references #161, which the file never defines.
openbim-step 0.6.0
Changed
- Breaking:
EntityDef::supertype: Option<String>is replaced by
EntityDef::supertypes: Vec<String>, every direct supertype in
SUBTYPE OForder.EntityDef::supertype()returns the first, for
single-inheritance callers;with_supertypenow appends. Migration:
def.supertype.clone()becomesdef.supertype().map(str::to_owned). - Breaking:
EntityDefgained publicsupertypesandredeclared
fields, so struct-literal construction must add them.
cargo semver-checks against the published 0.5.1 reports exactly these two
breaks (struct_pub_field_missing, constructible_struct_adds_field) and
182 of 184 checks passing.
Added
EntityDef::redeclared,EntityDef::is_redeclared, andRedeclaration:
explicitSELF\X.a : T;redeclarations, kept apart fromattributes.SchemaGraph::direct_supertypes.
Fixed
- Multiple inheritance (#2). Every supertype after the first was dropped, so
SchemaGraph::attributesomitted inherited slots andis_amissed
ancestors. Layouts now follow ISO 10303-21:2016 §12.2.5.2: supertypes in
SUBTYPE OForder, higher supertypes first, and a supertype reached twice
through a diamond counted once.supertypes,subtypes, andis_awalk
every parent. AP242 has 248 multi-parent entities; IFC has none. - Explicit redeclarations no longer add a phantom positional slot (#3).
SELF\styled_item.item : plane_or_planar_box;was parsed as a new
attribute namedSELF\styled_item.item; ISO 10303-21:2016 §12.2.8 says it
has no effect on the encoding. This affected 481 AP242 entities, including
advanced_face. - A
\S\page escape followed by an apostrophe no longer ends the string
literal.\S\takes exactly one followingLATIN_CODEPOINT, which includes
the apostrophe (ISO 10303-21:2016 §5.2, §6.4.3.1), so'Stra\S\'e'is one
string decoding toStra§e. Previously it failed to lex. The recovery
resynchronizer applies the same rule, and\\is consumed as one escaped
backslash in both so its second byte cannot open a page escape (#1).
Checked against OCCT's per-entity parameter counts (CheckNbParams in its
generated RWStep* readers): AP242 mismatches fell from 107 to 2, AP203e2
from 24 to 1. The residuals are OCCT departing from Part 21 (common_datum
diamond read twice; characterized_representation dropping derived *
slots) and are pinned in tests/schema_graph.rs.
Full changelog: v0.5.1...v0.6.0
crates.io: https://crates.io/crates/openbim-step/0.6.0
openbim-step 0.5.1
Added
SchemaGraph::direct_subtypesandSchemaGraph::subtypes: the downward
counterpart ofsupertypes.subtypes(x)is every entityy != xfor
whichis_a(y, x)holds, at any depth, in a deterministic sorted
pre-order. A test checks that equivalence for every ordered entity pair.
Answering "everyIfcElement" needs this;EntityDefrecords only the
upward edge, so the child index is built once at construction.