Skip to content

Releases: SAY-5/tradegraph

v5.1.1

Choose a tag to compare

@SAY-5 SAY-5 released this 28 Sep 18:53
66475bb

A browser demo fix release: nothing outside web/ changes apart from the version numbers and the documentation. The neighbourhood graph now places its labels instead of drawing each one on its spoke regardless of what is already there. For every node, in the order the graph receives them, web/src/components/ForceGraph.tsx tries the spot along the spoke, then level with the node, above or below it, and on the side facing the centre, and keeps the first whose box, 7.4 units per character in the page's monospace face, stays inside the view and clears every label already placed and every other node. A label with no such spot is left out, the centre's excepted, and the node keeps its full name in its accessible name and its tooltip.

Because a label can now be left out, the lists under the graph name every node it draws: one list for the selected entity and one for each expansion, each in the order neighbors.rq returns, lineage first and then holdings by value. In 5.1.0 there was a single list, built from every drawn edge as if it touched the selected entity, so the nodes an expansion added were mostly missing from it: expanding Price T ROWE Global Select Fund from Apple draws 49 nodes, and the 5.1.0 list left 16 of them unnamed, while five of the first ten single expansions from Apple left nodes unnamed and 5.1.1 names every node in all ten. The 5.1.0 notes called that list a parallel list of buttons for the graph; it was one only until the first expansion.

npm run selfcheck reads the demo block that README.md pastes and compares each figure it shares with web/src/data/demo-summary.json, the file the same run wrote: the store and endpoint, the dataset counts, the stats query time, the number of exposure queries, the three top pairs with their totals and latencies, the pairs with exposure and the latency p50 and max, the cached repeat, the operations store line, and the commit and host the sentence above the block names. Dollar totals are rounded the way scripts/demo_queries.py prints them, whose f-string takes a total ending in exactly .5 to the even dollar where Math.round would round it up. The same figures are also checked where the README's browser demo paragraph, web/README.md, etl/sample/README.md and ARCHITECTURE.md quote them. The self check runs 135 assertions, against 104 in 5.1.0.

The payload measures 844,919 bytes on disk and 241,422 gzipped under npm run weight, against 843,746 and 240,873 at 5.1.0, and web/README.md quotes the new figures; the gzip total is the one CI measured under Node 22.23.2 with zlib 1.3.1, and the same three files gave the same bytes under Node 26.3.0 with that zlib. The ETL package declares 5.1.1 in etl/pyproject.toml and in tradegraph_etl.version, which read 5.0.0 and 1.0.0 in 5.1.0, beside the API's pom.xml and the explorer's package.json. Tests at this commit: 47 ETL pytest, 60 API unit and 34 Testcontainers integration, 9 explorer specs, and 135 browser demo self check assertions, all passing in CI run 36466280184.

v5.1.0

Choose a tag to compare

@SAY-5 SAY-5 released this 27 Sep 00:27
6c8c8aa

A correctness and provenance pass over the whole repository. No new endpoints: the changes make the claims in the documents follow from the code and the data.

neighbors.rq ranks lineage above holdings, so the row limit truncates holdings rather than the parent and subsidiary edges the explorer draws the corporate tree from. Ordering by the relation name had sorted HELD_BY and HOLDS ahead of PARENT and SUBSIDIARY, so a well held issuer such as Apple, which has 27 distinct holders and five direct subsidiaries in the sample, kept three of those five lineage edges at the explorer's initial thirty row request and none at its fifteen row expansion; NeighborLimitIT covers it, and two of its three cases fail against the previous ORDER BY. ExplorerGraphIT now runs the endpoints the explorer calls against the full sample rather than the six entity fixture, which is where row limits, ordering and truncation behave differently at all. The query cost guard runs over every template as QueryTemplates loads it instead of over every rendered query, where the pattern also matched inside a caller's quoted literal and answered 422 to ordinary search terms such as tg:x+. /exposure/concentration is bounded in the store: a ranked page with HAVING and LIMIT plus one row each for the family total and the number of issuers above the threshold, where it previously returned a row per issuer the family held, up to 276 rows in the committed sample. A position id comes from its own namespace constant rather than a string rewrite of the entity namespace.

The live ETL path states what it reads: there is no Exhibit 21 reader in --live, so it produces holdings and no corporate tree, and a test asserts the transform emits no subsidiaryOf triple from live data. scripts/demo_queries.py --summary PATH writes the dataset counts, the exposure latencies and the commit, host and timestamp that produced them, and both the README demo block and the browser demo quote that file instead of transcribing numbers. ExposurePerformanceIT writes its latencies to target/benchmarks/exposure-latency.txt and fails rather than skipping when TRADEGRAPH_REQUIRE_SAMPLE=1, which CI now sets, so a missing artifact cannot become a silent pass; docs/benchmarks/2026-09-15-exposure-latency.txt records two runs with the dataset, JVM, host and commit behind them, in place of the two exposure latencies the README used to quote with no artifact behind either. etl/tests/test_manifest_parity.py asserts the ontology, triple, entity and position counts in the slice manifest against what rdflib counts over the same sample.

The property path lab renders the lineage templates again: the committed slice predated the directOnly placeholder that lineage_up.rq and lineage_down.rq gained, so the page quoted stale SPARQL and regenerating the slice would have thrown on the unresolved placeholder. The browser renderers now pass that parameter the way LineageService renders it with reasoning off, the concentration tab renders the floor and the limit the API sends, the self check renders all five tabs and asserts no placeholder survives, and every section sits behind an error boundary. The hero leads with the figures the page computes and names the full sample beside them, the interactive SVGs are no longer labelled as images, Space activates a node without scrolling the page, the neighbourhood graph has a parallel list of buttons, count up animations announce the settled number once, the dimmest text token measures 6.25:1 on the page background, and the payload dropped the full d3 package and the five templates the page never shows.

CI gained a web job: type check, bundle, self check, payload weight against ceilings of 1,100,000 bytes on disk and 300,000 gzipped, and a slice drift gate that regenerates the slice and fails on any diff, all of which make web runs locally. Documentation: the cache list, the --funds limit, the CLI command list and the contributor lint command now match the code, and the explorer's test fixture records the dataset the sample actually holds. Tests at this commit: 47 ETL pytest, 60 API unit and 34 Testcontainers integration, 9 explorer specs, and 104 browser demo self check assertions.

v5.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 10 Sep 10:14

What the API is doing and what it will refuse to do are both visible now. GET /ops/overview reports the store and its triple count, the Caffeine hit ratio across the caches, the slowest queries still in a 200 entry ring buffer, both depth limits and the headline of the last quality run, and every query is also a Micrometer timer tagged with the template it came from at /actuator/metrics/tradegraph.sparql. A cost guard answers 422 instead of running an unbounded property path or a walk deeper than the configuration allows, and the two limits are separate, four for exposure and five for lineage. deploy/fuseki adds a read only /ds-inf service backed by a Jena generic rule reasoner that materialises the transitive closure of subsidiaryOf and the inverse hasSubsidiary, which is what Stardog gives with reasoning=true for this vocabulary; with tradegraph.store.reasoning=true the exposure clauses ask for one hop instead of the bounded alternation and the lineage queries filter to direct parents. ReasoningParityIT loads the fixture into that service and asserts the exposure total is identical either way, and that the two hop edges exist only where the reasoner is running. The demo script prints the operations overview and no longer hardcodes the API version, and the README demo block is the output of the run that shipped this tag. Tests are 40 ETL pytest, 56 API unit, 27 Testcontainers integration and 9 explorer specs, all green against Fuseki.

v4.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 10 Sep 09:49

The ETL now says whether what it produced is well formed and the API serves that verdict. ontology/shapes.ttl holds SHACL shapes for entities, positions, instruments and filings, covering cardinality, datatypes, the ten digit CIK pattern, ownership fractions between 0 and 1 and the closed instrument class vocabulary. tradegraph-etl validate runs them with pyshacl and adds the three checks that belong to the dataset rather than to a single node, dangling references, subsidiaryOf cycles and issuers carrying neither a CIK nor a ticker, writing the result to etl/build/quality.json; --fail-on-violation makes it exit non zero, which is what CI now runs. The shapes earned their place on the first run by catching a real defect: a listed asset manager's Exhibit 21 filing shared an accession number with its own first 13F, so 13F filings in the sample are now numbered from 1000. GET /quality serves the report of the last load and answers 404 when there is none, and tradegraph-etl load --since only pushes the graph files that moved, so a rebuild that touched one graph replaces one graph instead of three. Tests are 40 ETL pytest, 45 API unit, 22 Testcontainers integration and 9 explorer specs, all green against Fuseki.

v3.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 10 Sep 09:33

Lineage edges now carry how much of a subsidiary its parent owns, so exposure can be read through that ownership instead of counting every position at face value. The ETL writes tg:ownershipFraction on every entity that has a parent, taking the percentage from the Exhibit 21 style list where one is stated and defaulting to whole ownership flagged with tg:ownershipAssumed where it is not; the committed sample states a percentage for 1,552 of its 2,148 subsidiaries. Passing weighted=true to exposure multiplies each line by the product of the fractions along its lineage path and reorders the answer by what is left, so debt issued by an entity owned 75 percent by a parent that is itself owned 80 percent counts 0.6 towards the top of the tree. Unweighted answers are byte for byte what they were, and weight and weightedValue are simply absent from them. A new endpoint at /exposure/concentration lists the issuers that make up at least a configurable share of what a fund family holds, largest first, with the default threshold in tradegraph.exposure.min-share. Tests are 32 ETL pytest, 43 API unit, 21 Testcontainers integration and 9 explorer specs, all green against Fuseki.

v2.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 10 Sep 09:18

Positions now carry the reporting period of the filing they came from, so an answer covers exactly one quarter instead of summing the same holding once per filed period. The ETL normalises whatever shape EDGAR reports the period in (ISO, compact, US and quarter labels) to the quarter end, and drops a filing whose report date cannot be read; the committed sample gains a prior quarter, so every fund now reports for 2024-03-31 and 2024-06-30 across 820 filings and 24,336 positions. Exposure and the positions listing take an as_of parameter that selects the latest period on or before the given date, and without it the latest period of all is used. A new endpoint at /positions/delta reports the lines a holder opened, closed and moved between two periods, matched on CUSIP, and /periods lists what the store holds. The explorer exposure panel gains a period selector and prints the period each answer was computed over. Tests are 30 ETL pytest, 35 API unit, 19 Testcontainers integration and 8 explorer specs, all green against Fuseki.

v1.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 10 Sep 09:00

TradeGraph turns SEC EDGAR data into a counterparty knowledge graph and answers lineage and exposure questions over it with SPARQL 1.1. The Python ETL maps company tickers, 13F-HR holdings and Exhibit 21 style subsidiary lists into three named graphs under a FIBO-inspired OWL ontology, and loads them through the Graph Store Protocol. The Spring Boot API exposes entity search, entity detail, lineage, exposure with a path explanation per contributing position, trades, a neighbourhood graph and store statistics. The Angular explorer renders the neighbour graph with d3, the corporate tree and the exposure breakdown. Stardog is the intended store and ships with a compose stack and a profile, but it needs a licence, so the tests, CI and the demo all run against Apache Jena Fuseki. This baseline passes 20 ETL pytest tests, 30 API unit tests, 12 Testcontainers integration tests and 7 explorer specs.