Releases: graphlagoon/gsql2rsql
Release list
v0.12.2
v0.12.2 (2026-08-11)
Documentation
- docs: multi-source procedural BFS — why it fails at runtime and the options
A property start filter transpiles in procedural mode (cardinality is
data-dependent) and hits the embedded single-seed RAISE_ERROR when it
matches several nodes. Document the architectural reason (no seed column
in frontier/visited/result), the two workarounds (CTE mode; per-seed
loop), and — in the developer notes — the two designs to lift it, with
the real cost being re-A/B of every memory-optimization flag rather than
the rewrite itself.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> (6ed6450)
Fix
- fix: silent wrong answers in joins, aggregates, naming + loud error contracts
Fixes the ten defects catalogued in docs_help_dev/testing/test-failures-report.md
(F1-F10), each backed by new tests that failed before and pass after.
Silent wrong answers:
- Repeated pattern variable ((a)-[:T]->(a), closed cycles) never emitted the
closing SINK join pair, returning every path instead of every cycle. The
shared-variable branch in match_tree now binds the node to the preceding
relationship's far endpoint before skipping the duplicate data source. - percentileCont/percentileDisc vanished: the visitor dropped arguments past
the first (new extra_parameters on the aggregation AST node) and the
renderer had no pattern (PERCENTILE_CONT({1}) WITHIN GROUP (ORDER BY {0})).
Pattern-less aggregates now raise instead of emitting the bare operand. - LIMIT $n / SKIP $s were silently discarded; _extract_int_value now raises
TranspilerNotSupportedException for non-literal expressions. - RETURN a.name, b.name emitted two columns both named "name"; colliding
auto-derived aliases are qualified as a_name/b_name (explicit AS untouched). - RETURN * / WITH * leaked internal gsql2rsql* columns and anonymous edge
plumbing; the star now expands to the named variables in scope.
Invalid SQL made valid or loud:
- size(<string>) renders LENGTH() via type dispatch from the operator's input
schema (AST data_type is often unset at render time). - list + list renders CONCAT(), with list literals as ARRAY(...) operands.
- WHERE a:Person label predicates raise with a move-into-pattern hint instead
of rendering a non-boolean WHERE (both grammar paths). - Unknown functions raise TranspilerNotSupportedException naming what the
user wrote (raw_name) instead of leaking NotImplementedError(INVALID). - add_edge derives missing source/sink id properties from the descriptor's
node_id_columns and rejects only when underivable, so hand-built schemas
can no longer produce dangling join keys.
Documented (not fixed): relationship uniqueness (edge isomorphism) is not
enforced in fixed-length patterns - guarded by a strict xfail; unbounded VLP
depth cap of 10 is now documented and locked by tests.
Test infrastructure: get_base_spark() probes SparkContext liveness so modules
using the shared session survive legacy modules that call spark.stop().
Suites: non-PySpark 1801 passed / 7 skipped / 1 xfailed; PySpark 261 passed /
6 xfailed; pyright and mypy at 0 errors.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> (41c10d3)
- fix: actionable errors for unknown edge types + user-facing limitations docs
Planner (data_source.py, recursive_traversal.py):
- Unknown relationship types (single-hop AND VLP) used to either fail with
a generic "Failed to bind relationship" (single-hop) or a confusing
"Internal error: No enriched data for RecursiveTraversalOperator" three
phases later in Enrichment (VLP) -- the latter reading as a transpiler
bug when it's a schema/config mismatch. A typo inside an OR list
([:OWNS|KNOWS]) silently transpiled, dropping the bad type instead of
raising. - Added validate_edge_verbs_exist()/_describe_available_edge_verbs() in
data_source.py, called from both binding sites, mirroring the node-type
diagnostics: lists registered edge types, suggests a case-insensitive
near-match. Distinguishes "verb exists for no endpoint pair" (typo,
raises) from "verb exists but not for these endpoints" (silently pruned
from OR lists -- load-bearing for restricted edge_combinations schemas). - get_all_edge_schemas() added to SimpleSQLSchemaProvider.
Docs:
- New docs/limitations.md (mkdocs nav: User Guide > Limitations):
user-facing quick-reference table, CTE vs procedural feature matrix,
schema-binding error guide, parser notes (backticks, unaliased
projection aliases, != rejection), correctness/performance caveats.
user-guide.md's inline section now summarizes and links to it. - fix_transpiler_bugs skill: new step 7 requiring both limitations docs
(docs_help_dev/ internal, docs/ user-facing) to be updated whenever a
fix adds/removes a guardrail or NotSupported path.
Tests: tests/test_unknown_edge_type_errors.py (10 cases: single-hop, VLP,
OR-list typo detection, endpoint-pruning-not-a-typo regression). (554761c)
v0.12.1
v0.12.1 (2026-08-11)
Fix
- fix: escaped identifiers, unaliased projections, procedural BFS guardrails
Three parser-phase defects and a set of procedural-BFS correctness
guardrails, all found while diagnosing a "Failed to bind entity" error
on a real MATCH ... MATCH (root)-[*1..2]-(d) RETURN count(DISTINCT d)
query.
Parser (visitor.py):
- Backtick-escaped identifiers (
Label,prop,variable) leaked
their delimiters into entity_name/alias, so a correctly-registered
node type likecnpj_raizfailed to bind because the lookup used
the literal string with backticks still attached. Added
unescape_symbolic_name()/split_escaped_labels() and applied them at
every identifier extraction site: node labels, relationship types
(OR-joined), variables, property lookups, map literal keys, UNWIND/
path/REDUCE/list-comprehension variables. Table names (which come
from schema config, not the Cypher text) are untouched. - Unaliased projections (RETURN count(*), RETURN n.age + 1, ...) via
the live visit_oC_ProjectionItem path emitted an empty AS alias,
producing a SQL syntax error only at execution time. Added
sanitize_expression_alias() to derive a column name from the raw
Cypher expression text, matching the fallback the legacy (dead)
visit_oC_ReturnItem path already had.
Planner (data_source.py):
- Binding failure on an unknown node type now lists the registered
types and flags a case-insensitive near-match, instead of a bare
"Failed to bind entity" with no actionable next step.
Renderer (procedural_bfs_renderer.py, sql_renderer.py):
- Procedural BFS's single-source architecture (one global visited set
- CROSS JOIN frontier attribution) silently produced wrong results
for constructs CTE mode handles correctly: unfiltered/chained VLP
sources fabricated (start, end) pairs or returned empty, *0..N
omitted the start node, and length(p)/nodes(p) evaluated to NULL.
Added transpile-time NotSupportedException for all three (pointing
callers at vlp_rendering_mode="cte"), plus a runtime seed-cardinality
guard (RAISE_ERROR) since a start filter can still match >1 node.
relationships(p) and ALL(n IN nodes(p) WHERE ...) pushdown remain
supported in procedural mode.
- CROSS JOIN frontier attribution) silently produced wrong results
Verified: pyright/mypy 0 errors; 1738 passed (was 1694) + 7 skipped +
1 xfailed on the non-PySpark suite; 220 passed + 5 xfailed on PySpark.
Dual-mode (CTE vs procedural) A/B executed on real PySpark data across
~30 query variations with hand-computed expected results — no
divergence or wrong result survives after these fixes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> (3b948d9)
v0.12.0
v0.11.0
v0.10.5
v0.10.5 (2026-07-09)
Fix
- fix(resolver): derive entity id/src/dst columns from schema instead of hardcoding
Returning an entity (RETURN r/n, collect(r/n), WITH-propagated nodes after
VLP) built column references from the literals "src"/"dst"/"node_id" instead
of the schema's real column names. With a custom schema (e.g. edge_src_col=
"source_node_id", node_id_col="vertex_key") this emitted references to
columns never materialized in the inner projection, failing at runtime with
UNRESOLVED_COLUMN. Default schemas coincided with the literals, hiding the bug.
Derive keys/columns from rel_source_join_field / rel_sink_join_field /
node_join_field.field_name across all four sites (expression_renderer struct
collector, column_resolver bare-entity fallback, sql_renderer required-column
collection x2). Handle the post-WITH case where field_name arrives already
prefixed to avoid double-prefixing.
Default-schema output is byte-identical; struct keys for custom schemas now
use the real column names (the old keys pointed at nonexistent columns). (b4a0220)
v0.10.4
v0.10.4 (2026-07-08)
Documentation
- docs: link new repo org (
773c55d)
Fix
- fix(renderer): derive edge src/dst struct columns from schema instead of hardcoding
RETURN r / collect(r) hardcoded NAMED_STRUCT keys and required-columns to
"src"/"dst", which only worked by coincidence on the default schema. Custom
edge schemas (e.g. edge_src_col="source_node_id") produced dangling column
references, failing at runtime with UNRESOLVED_COLUMN. (bef9f39)