A read-only BIM query server for agents. Point it at an IFC or gbXML model and it answers structured questions — fire-rated doors on level 3, elements with no material assigned, spaces below the minimum daylight area — bounded by a policy file, with deterministic results and a citation back to the GlobalId and source line behind every row.
Zero dependencies. Python 3.11+. MCP server over stdio, plus a CLI that answers the same questions so you can check a policy before you trust an agent to it.
bimq query elements_by_property model.ifc \
type=IfcDoor storey="Level 3" property=FireRating op=existsid type name tag storey_name source match
---------------------- ------- -------- ---- ----------- -------------- -------------------------------
0XBbD$nZDLuRru91_CQ_xe IfcDoor Door-302 D302 Level 3 office.ifc:314 Pset_DoorCommon.FireRating=EI60
31kamnSrrNbf3eF0_vhXrJ IfcDoor Door-301 D301 Level 3 office.ifc:304 Pset_DoorCommon.FireRating=EI60
2 row(s)
digest: sha256:06321f331735417dd149d649b8e26de71a63ff2bb26a31c38cbf66a4f0314b77
Then check it, because a citation you cannot follow is just a confident-looking string:
bimq cite model.ifc 31kamnSrrNbf3eF0_vhXrJ31kamnSrrNbf3eF0_vhXrJ (IfcGloballyUniqueId)
office.ifc:304 #297
#297= IFCDOOR('31kamnSrrNbf3eF0_vhXrJ',#5,'Door-301',$,$,$,$,'D301',2100.0,900.0,.DOOR.,.SINGLE_SWING_LEFT.,$);
The current instinct is to dump IFC text into a context window. That fails immediately at real model sizes, and it fails quietly: a 300 MB model is roughly 95% geometry, so what fits in the window is a truncated arbitrary slice, and the model answers from it anyway. The failure looks like a fluent paragraph about a door that does not exist.
bimq inverts it. The model stays on disk. Queries are structured, the answers are small, and every row carries the id and line it came from — so a claim can be checked against the file instead of trusted.
Three properties hold for every answer:
Bounded. A TOML policy file says what is readable — which files, which queries, which types, which storeys, which properties. The engine reduces the model to the visible set before the query runs, so a query primitive cannot reach what the policy hides even by accident.
Deterministic. Same model, same query, same bytes. Every answer carries a
digest you can pin in a test. Element order, group order and float rounding are
all fixed; the read block size and the file's name do not change a finding.
Cited. Every row carries {id, id_kind, source, ref, line}. bimq cite
resolves it back to the original statement. The test-suite re-reads the recorded
line for every element of every fixture and fails if the id is not there.
pip install bimqOr run it from a clone with no install at all — there is nothing to build:
python -m bimq describe tests/fixtures/office.ifc{
"mcpServers": {
"bimq": {
"command": "bimq",
"args": ["serve", "/srv/bim/tower.ifc", "-p", "/srv/bim/policy.toml"]
}
}
}Every query primitive becomes a bim_* tool, all annotated readOnlyHint, plus
bim_cite. Omit the model path to let each call name its own file — then
allow_sources is what stands between a path argument and your filesystem.
The server tells the agent how to behave on initialize: call bim_model_summary
first, quote a GlobalId for anything you assert, treat truncated as "there are
more", and read notes — because no results and no data recorded are
different findings and only the notes distinguish them.
A policy refusal comes back as a successful tool result carrying
policy_denied and the rule that fired, not as a protocol error. An agent that
receives a protocol error retries; an agent told "this policy does not expose
costs" reports the limit and moves on.
| Primitive | Answers |
|---|---|
model_summary |
schema, units, storeys, entity types, property-set names |
spatial_tree |
project → site → building → storey → space |
elements_by_type |
elements of a type, subtypes included |
elements_by_property |
property comparison; the fire-door workhorse |
elements_missing_property |
data completeness: who has no value for this field |
elements_missing_material |
no material through any of IFC's five ways of saying so |
spaces_by_area |
rooms inside an area range, always in m² |
property_values |
distinct values with counts — run this before guessing names |
quantity_rollup |
totals grouped by type, storey or PredefinedType |
element_detail |
expand specific GlobalIds to every pset and quantity |
bimq queries prints their parameters. List queries return compact rows on
purpose; element_detail is the drill-down, and keeping those separate is what
stops a query from becoming the context dump it replaced.
Aggregates report their own coverage. A roll-up over 200 walls where 160 carry no
quantity says so in summary and notes, because a total over 40 of 200 is not
a total.
name = "consultant-readonly"
[allow_sources]
roots = ["/srv/bim"]
max_bytes = 536870912
[allow_queries]
queries = ["model_summary", "spatial_tree", "elements_by_type", "elements_by_property"]
[scope_storeys]
names = ["Level 2", "Level 3"]
include_unplaced = false
[allow_types]
types = ["IfcBuiltElement", "IfcSpace", "IfcBuildingStorey"]
[deny_properties]
properties = ["*Cost*", "Pset_Tender.*"]
[redact_properties]
properties = ["*.Owner*", "*SerialNumber*"]
placeholder = "[redacted]"
[max_results]
limit = 200bimq policy check policy.toml # validate before shipping
bimq rules # every rule, with an exampleNotes on the design:
- An unknown table is a hard error, not a warning. A file whose job is to withhold data must not fail open because of a typo.
denyandredactare different tools. A denied property is gone; a redacted one is present with a placeholder. The distinction matters to an agent: redaction says this exists and you are not being shown it, so the agent reports a gap instead of concluding nobody entered the data.- Denial covers the query side too. You cannot filter on a denied property,
because
op=gt value=1000repeated a few times reconstructs it. - Withholding is reported, never silent. Answers carry
policy.elements_withheldand a note. Truncation setstruncated: true. - Every answer is capped even with no policy at all. "Unlimited" is not a sane default for something feeding a context window.
| Format | Notes |
|---|---|
IFC-SPF (.ifc, .ifczip) |
IFC2X3 / IFC4 / IFC4X3, streaming reader, no dependencies |
gbXML (.gbxml) |
energy models; ids are stamped gbXMLId, never confused with GlobalIds |
Wanted, one per PR: Revit export (pyRevit/Dynamo JSON), Speckle stream, IFC-JSON, COBie. See CONTRIBUTING.md.
bimq/sources/spf.py is a complete ISO 10303-21 reader in under 400 lines. The
parts that matter:
- The file is scanned in 4 MB blocks, so a 300 MB model is never one string. A block boundary can land inside a string literal, so the scanner explicitly matches unterminated literals and carries them forward. Tested at block sizes down to one byte, where the result must still be byte-identical.
- A
;inside'a;b'does not end a statement,''is an escaped quote, and\X2\...\X0\decodes to UTF-16 — soPhòng họpsurvives the round trip. - Comments appear between statements, inside parameter lists, and around section markers. All three are handled; the reported line still points at the entity.
- Geometry is never loaded. An instance is kept only if its first attribute
is a syntactically valid GlobalId — making it an
IfcRootsubtype — or if it is one of ~30 unrooted carriers of property, quantity, material or unit data. The test is applied to the raw text before tokenising, which is where the parse time on a real file actually goes.
Check the throughput claim yourself without needing a model of your own — this writes a file shaped like a real export (a modest element count buried in geometry), parses it, and reports:
$ bimq bench --synthetic 20000
synthetic model: 20000 elements among 820006 instances
tmp6l05ix9x.ifc: 32.2 MiB, 20001 elements
parse: 1.71 s · 18.8 MiB/s · 11,680 elements/s
peak rss: 107 MiB (3.3x file size)820,006 instances go in; 20,001 elements stay resident. That ratio is the whole argument — resident size tracks how many things the building has, not how many points were needed to draw them.
Model files are treated as untrusted input. The gbXML reader refuses entity
declarations outright, so a file cannot carry a billion-laughs expansion or an
external entity pointing at /etc/passwd.
Every number bimq returns is SI: metres, m², m³. A model authored in millimetres
with areas in square metres (what Revit exports) and one authored in feet with
areas in square feet both answer spaces_by_area max_m2=8 correctly.
IfcConversionBasedUnit chains are resolved, not guessed.
The fixtures are synthetic — stated plainly, because a fixture pretending to be a
real project is one nobody can check. What makes them useful is that the defects
are deliberate and enumerated: a wall with no material, a fire door with no
rating, a room below 8 m², a room with no area quantity at all, a door whose
rating is inherited from its type, and a Vietnamese room name written with \X2\
escapes.
python scripts/make_fixture.py tests/fixtures
bimq describe tests/fixtures/office.ifc
bimq query elements_missing_material tests/fixtures/office.ifc type=IfcWall
bimq query spaces_by_area tests/fixtures/office.ifc max_m2=8
bimq query property_values tests/fixtures/office.ifc type=IfcDoor property=FireRating
bimq query spaces_by_area tests/fixtures/legacy-imperial.ifc max_m2=8 # authored in feet
bimq query spaces_by_area tests/fixtures/clinic.gbxml max_m2=8 # gbXML, same primitivemake test # the suite
make fixtures # regenerate fixtures (byte-identical; CI checks this)
make bench # parse throughput on a fixtureMIT