feat(pipeline): add file-backed FileProvenanceStore#636
Merged
Conversation
- Persist processing records in a single JSON file keyed by dataset URI, for pipelines that run without a triplestore - Write atomically (temp file + rename) so a run killed mid-write cannot corrupt later skip decisions - Document the single-writer scope and durable-volume requirement in JSDoc and README
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fix #634
Ships a file-backed
ProvenanceStorein@lde/pipeline, alongside the interface and the existingFileLoadedSparqlProvenanceStore, so a consumer that runs the pipeline without a triplestore (e.g. a single CronJob persisting to a mounted volume) no longer has to stand one up purely to remember processing records.Changes
FileProvenanceStorepersists all records to a single JSON file, keyed by dataset URI, with atomic writes (temp file + rename) so a run killed mid-write cannot corrupt the next run’s skip decisions. A missing file is the empty store; any other read error (including corruption) propagates instead of silently starting over.{ path }) rather than a positional argument, matchingFileLoadedSparqlProvenanceStore.fileProvenanceStore.tsis at 100% coverage and the package coverage thresholds auto-updated upward.