Ensure data provenance for raw mass spectrometry data #9
adamscharlotte
announced in
Hackathon Proposals
Replies: 1 comment
|
This letter in JPR gives a bit more background, as does this 2023 white paper. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Title
Ensure data provenance for raw mass spectrometry data
Abstract
Data provenance mechanisms already exist in regulated and clinical workflows. Unfortunately, these rely on centralized authorities, vendor PKI, and cumbersome infrastructure that is impractical for the day-to-day needs of the broader proteomics community. Meanwhile, with data simulation tools now capable of producing highly convincing raw mass spectrometry files, the ability to verify the provenance and integrity of deposited datasets is becoming a prerequisite for trust in publicly deposited datasets.
We propose a lightweight attestation framework that computes a canonical hash of the analytical data and stores an Ed25519 signature over that hash together with a stable key identifier directly inside the raw file. Integrity and identity can then be verified by anyone with access to the signer’s public key, retrieved from a trusted keystore.The framework is format-aware, with current support for Bruker .d (via a reserved table in analysis.tdf) and mzML (via a in ), and is designed to be extensible to further formats including Thermo RAW and mzPeak. Implementations are available in both Python and Rust. By embedding provenance directly into the formats used for downstream analysis, the system ensures that a verifiable chain of custody travels with the data through the full proteomics workflow.
While embedding signatures into files is technically straightforward, several broader ecosystem questions remain unresolved: how signing keys should be generated, distributed, and trusted within the proteomics community; how provenance information should propagate through downstream processing and format conversion; and how repositories and analysis software should expose and validate provenance metadata. The hackathon will therefore focus not only on implementation, but also on defining practical interoperability and trust models suitable for open scientific workflows.
Project Plan
The project comprises four major tasks:
After the hackathon we aim to have an initial community-endorsed specification, extended format coverage, and a robust cross-language test suite. Additional expected outcomes are:
Technical Details
Contact information
Charlotte Adams
University of Antwerp
charlotte.adams@uantwerpen.be
David Teschner
Johannes-Gutenberg University Mainz
dateschn@uni-mainz.de
Richard Scheltema
University of Liverpool
richard.scheltema@liverpool.ac.uk
Stefan Tenzer
University Medical Center of the Johannes-Gutenberg University Mainz
Helmholtz Institute for Translational Oncology (HI-TRON) Mainz
tenzer@uni-mainz.de
Magnus Palmblad
Center for Proteomics and Metabolomics
Leiden University Medical Center
n.m.palmblad@lumc.nl
All reactions