From 7542d98566c20731a69149d48e1db16c5e154f45 Mon Sep 17 00:00:00 2001 From: Birdmachine Date: Mon, 20 Jul 2026 16:45:09 -0400 Subject: [PATCH 1/2] simplify scriptwork to utilize HRA for allowable value data & drop need for api keys, generate Harmonized-style metadata documentation pages drectly from github issue number Add Paired Tag & fix required values for DNA Methylation --- .gitignore | 1 - docs/assays/metadata/dna-methylation.md | 28 +- docs/assays/metadata/index.md | 3 +- docs/assays/metadata/paired-tag.md | 26 + scripts/metaBuilder/README.md | 53 -- scripts/metaBuilder/assays.txt | 5 - scripts/metaBuilder/fetchMeta.py | 137 ---- scripts/metaBuilder/generate_all_md.py | 268 -------- scripts/metaBuilder/json_to_md.py | 57 -- scripts/metaBuilder/pageLayout.md | 12 - scripts/metadata-generator/generate.py | 875 ++++++++++++++++++++++++ 11 files changed, 916 insertions(+), 549 deletions(-) create mode 100644 docs/assays/metadata/paired-tag.md delete mode 100644 scripts/metaBuilder/README.md delete mode 100644 scripts/metaBuilder/assays.txt delete mode 100644 scripts/metaBuilder/fetchMeta.py delete mode 100644 scripts/metaBuilder/generate_all_md.py delete mode 100644 scripts/metaBuilder/json_to_md.py delete mode 100644 scripts/metaBuilder/pageLayout.md create mode 100755 scripts/metadata-generator/generate.py diff --git a/.gitignore b/.gitignore index 8278e665..de2b58f5 100644 --- a/.gitignore +++ b/.gitignore @@ -5,5 +5,4 @@ node_modules /.vscode/* .directory # scripts -scripts/* docs/metaTemplate/.gitignore \ No newline at end of file diff --git a/docs/assays/metadata/dna-methylation.md b/docs/assays/metadata/dna-methylation.md index dd0ce528..5a5cff0e 100644 --- a/docs/assays/metadata/dna-methylation.md +++ b/docs/assays/metadata/dna-methylation.md @@ -11,20 +11,18 @@ Fields that are collected for DNA Methylation data, available at ```dataset.meta | Attribute | Type | Description | Allowable Values | |------|------|-------------|-------------------| +| parent_sample_id * | | The unique identifier from HuBMAP or SenNet for the sample (such as a block, section, or suspension) used to perform the assay. For instance, in an RNAseq assay, the parent sample would be the suspension, while in imaging assays, it would be the tissue section. If the assay is derived from multiple parent samples, this field should contain a comma-separated list of identifiers. Example: HBM386.ZGKG.235, HBM672.MKPK.442 | | | lab_id | | A locally assigned identifier provided by the data provider for the dataset. It is used to reference an external metadata record that may be maintained independently, enabling traceability and supporting provenance tracking. Example: Visium_9OLC_A4_S1 | | -| dataset_type * | | The specific type of dataset being produced. Example: RNAseq | ```Visium HD``` ```4i``` ```Illumina Spatial ver0``` ```LC-MS``` ```Thick section Multiphoton MxIF``` ```Light Sheet``` ```ATACseq``` ```Resolve``` ```HiFi-Slide``` ```COMET``` ```DNA Methylation``` ```MPLEx``` ```10X Multiome``` ```MALDI``` ```MACSima``` ```Raman Imaging``` ```Histology``` ```Cell DIVE``` ```FACS``` ```MS Lipidomics``` ```Visium (no probes)``` ```MUSIC``` ```RNAseq``` ```GeoMx (NGS)``` ```GeoMx (nCounter)``` ```RNAseq (with probes)``` ```Singular Genomics G4X``` ```Molecular Cartography``` ```CosMx Transcriptomics``` ```MERFISH``` ```Pixel-seqV2``` ```2D Imaging Mass Cytometry``` ```Confocal``` ```seqFISH``` ```DART-FISH``` ```MIBI``` ```Olink``` ```Enhanced Stimulated Raman Spectroscopy (SRS)``` ```DESI``` ```Xenium``` ```iCLAP``` ```CyCIF``` ```SNARE-seq2``` ```nanoSPLITS``` ```STARmap``` ```Stereo-seq``` ```Visium (with probes)``` ```SIMS``` ```Auto-fluorescence``` ```CyTOF``` | -| analyte_class | | The analyte class which is the target molecule that the assay is measuring. Example: DNA | ```Nucleic acid + protein``` ```Lipid + metabolite``` ```Collagen``` ```RNA``` ```Fluorochrome``` ```DNA``` ```Metabolite``` ```DNA + RNA``` ```Saturated lipid``` ```Lipid``` ```Lipid + metabolite + protein``` ```RNA + protein``` ```Peptide``` ```Protein``` ```Unsaturated lipid``` ```Endogenous fluorophore``` ```Chromatin``` ```Polysaccharide``` | -| acquisition_instrument_vendor | | The company that manufactures or supplies the acquisition instrument. An acquisition instrument is a device equipped with signal detection hardware and signal processing software. It captures signals produced by assays, such as variations in light intensity or color, or signals corresponding to molecular mass. If the instrument was custom-built or developed internally, enter "In-House". Example: Illumina | ```Complete Genomics``` ```Cytek Biosciences``` ```Thermo Fisher Scientific``` ```Sciex``` ```Vizgen``` ```Leica Microsystems``` ```Akoya Biosciences``` ```Keyence``` ```Andor``` ```Standard BioTools (Fluidigm)``` ```Leica Biosystems``` ```Zeiss Microscopy``` ```Ionpath``` ```Motic``` ```In-House``` ```Miltenyi Biotec``` ```Revvity``` ```Evident Scientific (Olympus)``` ```GE Healthcare``` ```Element Biosciences``` ```Hamamatsu``` ```Waters``` ```Bruker``` ```Illumina``` ```3DHISTECH``` ```Singular Genomics``` ```Huron Digital Pathology``` ```Resolve Biosciences``` ```NanoString``` ```Cytiva``` ```10x Genomics``` ```Microscopes International``` ```BGI Genomics``` | -| acquisition_instrument_model | | The specific model of the acquisition instrument, as manufacturers often offer various versions with differing features or sensitivities. These differences may be relevant to the processing or interpretation of the data. If the instrument was custom-built or developed internally, enter "In-House". If the model is unknown, enter "Unknown". Example: HiSeq 4000 | ```NovaSeq X``` ```NovaSeq X Plus``` ```Cytek Northern Lights``` ```Lightsheet 7``` ```Resolve Biosciences Molecular Cartography``` ```timsTOF HT``` ```timsTOF Pro 2``` ```timsTOF Pro``` ```timsTOF Ultra``` ```timsTOF Ultra 2``` ```timsTOF SCP``` ```Axio Scan.Z1``` ```MALDI timsTOF Flex Prototype``` ```LSM 710 Confocal Microscope``` ```CosMx Spatial Molecular Imager``` ```Unknown``` ```MERSCOPE Ultra``` ```Juno System``` ```timsTOF FleX``` ```Custom: Multiphoton``` ```CyTOF XT``` ```Helios``` ```EVOS M7000``` ```Aperio AT2``` ```Phenocycler-Fusion 2.0``` ```Axio Observer 5``` ```Axio Observer 7``` ```Axio Observer 3``` ```NanoZoomer-SQ``` ```NanoZoomer S210``` ```NanoZoomer S60``` ```NanoZoomer S360``` ```DM6 B``` ```MoticEasyScan One``` ```In-House``` ```NextSeq 500``` ```BZ-X710``` ```MACSima System``` ```QTRAP 5500``` ```DMi8``` ```NextSeq 550``` ```HiSeq 2500``` ```HiSeq 4000``` ```NovaSeq 6000``` ```Opera Phenix Plus HCS``` ```SYNAPT G2-Si``` ```Q Exactive HF``` ```Orbitrap Fusion Tribrid``` ```Orbitrap Fusion Lumos Tribrid``` ```Q Exactive HF-X``` | -| source_storage_duration_value | | The length of time the sample was stored prior to processing it. For assays performed on tissue sections, this refers to how long the tissue section (e.g., slide) was stored before the assay began (e.g., imaging). For assays performed on suspensions, such as sequencing, it refers to how long the suspension was stored before library construction started. Example: 12 | | -| source_storage_duration_unit | | The unit of measurement used to specify the source storage duration value. Example: hour | ```hour``` ```month``` ```day``` ```minute``` ```year``` | -| time_since_acquisition_instrument_calibration_value | | The length of time since the acquisition instrument was last serviced or calibrated. This provides a metric for assessing drift in data capture. Example: 10 | | -| time_since_acquisition_instrument_calibration_unit | | The unit of measurement used to specify the time since acquisition instrument calibration value. Example: month | ```month``` ```day``` ```year``` | | preparation_protocol_doi * | | The DOI for the protocols.io page that details the assay or the procedures used for sample procurement and preparation. For example, in the case of an imaging assay, the protocol may start with tissue section staining and end with the generation of an OME-TIFF file. The documented protocol should also include any image processing steps involved in producing the final OME-TIFF. Example: https://dx.doi.org/10.17504/protocols.io.eq2lyno9qvx9/v1 | | -| is_targeted | | Indicates whether a specific molecule or set of molecules is targeted for detection or measurement by the assay. Example: Yes | | -| contributors_path | | The name of the file containing the ORCID IDs for all contributors to this dataset. Example: ./contributors.csv | | -| data_path | | The top-level directory containing the raw and/or processed data. For a single dataset upload, this might be represented as ".", whereas for a data upload containing multiple datasets, this would be the directory name for the respective dataset. For example, if the data is within a directory named "TEST001-RK", use the syntax "./TEST001-RK" for this field. If there are multiple directory levels, use the format "./TEST001-RK/Run1/Pass2", where "Pass2" is the subdirectory where the single dataset's data is stored. This is an internal metadata field used solely for data ingestion. Example: ./TEST001-RK | | -| parent_sample_id | | The unique identifier from HuBMAP or SenNet for the sample (such as a block, section, or suspension) used to perform the assay. For instance, in an RNAseq assay, the parent sample would be the suspension, while in imaging assays, it would be the tissue section. If the assay is derived from multiple parent samples, this field should contain a comma-separated list of identifiers. Example: HBM386.ZGKG.235, HBM672.MKPK.442 | | -| metadata_schema_id | | The unique string identifier for the metadata specification version, which is easily interpretable by computers for purposes of data validation and processing. Example: d70bfe24-e82a-46cb-9369-28ae03660d97 | | - - +| dataset_type * | | The specific type of dataset being produced. Example: RNAseq | ```10X Multiome``` ```2D Imaging Mass Cytometry``` ```4i``` ```ATACseq``` ```Auto-fluorescence``` ```Cell DIVE``` ```CODEX``` ```COMET``` ```Confocal``` ```CosMx Proteomics``` ```CosMx Transcriptomics``` ```CyCIF``` ```CyTOF``` ```DART-FISH``` ```DBiT-seq``` ```DESI``` ```DNA Methylation``` ```Enhanced Stimulated Raman Spectroscopy (SRS)``` ```FACS``` ```GeoMx (nCounter)``` ```GeoMx (NGS)``` ```HiFi-Slide``` ```Histology``` ```iCLAP``` ```Illumina Spatial ver0``` ```LC-MS``` ```Light Sheet``` ```MACSima``` ```MALDI``` ```MERFISH``` ```MIBI``` ```Molecular Cartography``` ```MPLEx``` ```MS Lipidomics``` ```MUSIC``` ```nanoSPLITS``` ```Olink``` ```PhenoCycler``` ```Pixel-seqV2``` ```Raman Imaging``` ```Resolve``` ```RNAseq``` ```RNAseq (with probes)``` ```Second Harmonic Generation (SHG)``` ```Seq-Scope``` ```seqFISH``` ```SIMS``` ```Singular Genomics G4X``` ```SNARE-seq2``` ```STARmap``` ```Stereo-seq``` ```Thick section Multiphoton MxIF``` ```Virtual Histology``` ```Visium (no probes)``` ```Visium (with probes)``` ```Visium HD``` ```Xenium``` | +| analyte_class * | | The analyte class which is the target molecule that the assay is measuring. Example: DNA | ```Chromatin``` ```Collagen``` ```DNA``` ```DNA + RNA``` ```Endogenous fluorophore``` ```Fluorochrome``` ```Lipid``` ```Lipid + metabolite``` ```Lipid + metabolite + protein``` ```Metabolite``` ```Nucleic acid + protein``` ```Peptide``` ```Polysaccharide``` ```Protein``` ```RNA``` ```RNA + protein``` ```Saturated lipid``` ```Unsaturated lipid``` | +| is_targeted * | | Indicates whether a specific molecule or set of molecules is targeted for detection or measurement by the assay. Example: Yes | ```Yes``` ```No``` | +| acquisition_instrument_vendor * | | The company that manufactures or supplies the acquisition instrument. An acquisition instrument is a device equipped with signal detection hardware and signal processing software. It captures signals produced by assays, such as variations in light intensity or color, or signals corresponding to molecular mass. If the instrument was custom-built or developed internally, enter "In-House". Example: Illumina | ```10x Genomics``` ```3DHISTECH``` ```Akoya Biosciences``` ```Andor``` ```BGI Genomics``` ```Bruker``` ```Complete Genomics``` ```Cytek Biosciences``` ```Cytiva``` ```Element Biosciences``` ```Evident Scientific (Olympus)``` ```GE Healthcare``` ```Hamamatsu``` ```Huron Digital Pathology``` ```Illumina``` ```In-House``` ```Ionpath``` ```Keyence``` ```Leica Biosystems``` ```Leica Microsystems``` ```Microscopes International``` ```Miltenyi Biotec``` ```Motic``` ```NanoString``` ```Resolve Biosciences``` ```Revvity``` ```Sciex``` ```Singular Genomics``` ```Standard BioTools (Fluidigm)``` ```Thermo Fisher Scientific``` ```Vizgen``` ```Waters``` ```Zeiss Microscopy``` | +| acquisition_instrument_model * | | The specific model of the acquisition instrument, as manufacturers often offer various versions with differing features or sensitivities. These differences may be relevant to the processing or interpretation of the data. If the instrument was custom-built or developed internally, enter "In-House". If the model is unknown, enter "Unknown". Example: HiSeq 4000 | ```Aperio AT2``` ```Aperio CS2``` ```AVITI``` ```Axio Observer 3``` ```Axio Observer 5``` ```Axio Observer 7``` ```Axio Scan.Z1``` ```Axio Zoom.V16``` ```Biomark HD``` ```BZ-X710``` ```BZ-X800``` ```BZ-X810``` ```Cell DIVE``` ```CosMx Spatial Molecular Imager``` ```Custom: Multiphoton``` ```Cytek Northern Lights``` ```CyTOF 2``` ```CyTOF XT``` ```Digital Spatial Profiler``` ```DM6 B``` ```DMi8``` ```DNBSEQ-T7``` ```EVOS M7000``` ```G4X Spatial Sequencer``` ```Helios``` ```HiSeq 2500``` ```HiSeq 4000``` ```Hyperion Imaging System``` ```IN Cell Analyzer 2200``` ```In-House``` ```Juno System``` ```Lightsheet 7``` ```LSM 710 Confocal Microscope``` ```MACSima System``` ```MALDI timsTOF Flex Prototype``` ```MERSCOPE``` ```MERSCOPE Ultra``` ```MIBIscope``` ```MoticEasyScan One``` ```NanoZoomer 2.0-HT``` ```NanoZoomer 2.0-RS``` ```NanoZoomer S210``` ```NanoZoomer S360``` ```NanoZoomer S60``` ```NanoZoomer-SQ``` ```NextSeq 2000``` ```NextSeq 500``` ```NextSeq 550``` ```Not applicable``` ```NovaSeq 6000``` ```NovaSeq X``` ```NovaSeq X Plus``` ```Opera Phenix HCS``` ```Opera Phenix Plus HCS``` ```Orbitrap Eclipse Tribrid``` ```Orbitrap Fusion Lumos Tribrid``` ```Orbitrap Fusion Tribrid``` ```Pannoramic MIDI II Digital Scanner``` ```Panoramic 150 Digital Scanner``` ```Phenocycler-Fusion 1.0``` ```Phenocycler-Fusion 2.0``` ```PhenoImager Fusion``` ```Q Exactive``` ```Q Exactive HF``` ```Q Exactive HF-X``` ```Q Exactive UHMR``` ```QTRAP 5500``` ```Resolve Biosciences Molecular Cartography``` ```SCN400``` ```solariX``` ```STELLARIS 5``` ```SYNAPT G2-Si``` ```timsTOF FleX``` ```timsTOF FleX MALDI-2``` ```timsTOF HT``` ```timsTOF Pro``` ```timsTOF Pro 2``` ```timsTOF SCP``` ```timsTOF Ultra``` ```timsTOF Ultra 2``` ```TissueScope LE Slide Scanner``` ```Unknown``` ```uScopeHXII-20``` ```VS200 Slide Scanner``` ```Xenium Analyzer``` ```Zeiss LightSheet Z.1``` ```Zyla 4.2 sCMOS``` | +| source_storage_duration_value * | | The length of time the sample was stored prior to processing it. For assays performed on tissue sections, this refers to how long the tissue section (e.g., slide) was stored before the assay began (e.g., imaging). For assays performed on suspensions, such as sequencing, it refers to how long the suspension was stored before library construction started. Example: 12 | | +| source_storage_duration_unit * | | The unit of measurement used to specify the source storage duration value. Example: hour | ```day``` ```hour``` ```minute``` ```month``` ```year``` | +| time_since_acquisition_instrument_calibration_value | | The length of time since the acquisition instrument was last serviced or calibrated. This provides a metric for assessing drift in data capture. Example: 10 | | +| time_since_acquisition_instrument_calibration_unit | | The unit of measurement used to specify the time since acquisition instrument calibration value. Example: month | ```day``` ```month``` ```year``` | +| contributors_path * | | The name of the file containing the ORCID IDs for all contributors to this dataset. Example: ./contributors.csv | | +| data_path * | | The top-level directory containing the raw and/or processed data. For a single dataset upload, this might be represented as ".", whereas for a data upload containing multiple datasets, this would be the directory name for the respective dataset. For example, if the data is within a directory named "TEST001-RK", use the syntax "./TEST001-RK" for this field. If there are multiple directory levels, use the format "./TEST001-RK/Run1/Pass2", where "Pass2" is the subdirectory where the single dataset's data is stored. This is an internal metadata field used solely for data ingestion. Example: ./TEST001-RK | | +| metadata_schema_id * | | The unique string identifier for the metadata specification version, which is easily interpretable by computers for purposes of data validation and processing. Example: d70bfe24-e82a-46cb-9369-28ae03660d97 | | diff --git a/docs/assays/metadata/index.md b/docs/assays/metadata/index.md index 3a910a63..30ffd3a3 100644 --- a/docs/assays/metadata/index.md +++ b/docs/assays/metadata/index.md @@ -15,7 +15,7 @@ A list of available dataset types (data types from multiple supported assays), w | [CosMx Proteomics](https://hubmapconsortium.github.io/ingest-validation-tools/cosmx-proteomics/current/) [](cosmx-proteomics "Attribute description")| CosMx Proteomics is a technology that enables the high-resolution, spatial analysis of proteins within their native tissue environment. It is part of the CosMx Spatial Molecular Imager (SMI) platform, which provides single-cell and subcellular resolution to map protein expression, cell states, and cell-cell interactions in FFPE and fresh frozen tissue samples. | | CosMx Transcriptomics [](cosmx-transcriptomics "Attribute description")| Dataset generated from performing the CosMx Transcriptomics assay. | | [DESI](https://pmc.ncbi.nlm.nih.gov/articles/PMC6053038/) [](DESI "Attribute description")| Desorption Electrospray Ionization (DESI), an ambient ionization technique that can be coupled to mass spectrometry (MS) for chemical analysis of samples at atmospheric conditions. Link to [DESI directory schema](https://hubmapconsortium.github.io/ingest-validation-tools/desi/current/). | -| DNA-Methylation [](dna-methylation "Attribute description") | DNA methylation is a critical, reversible epigenetic mechanism in biomedical research involving the addition of a methyl group to DNA (usually cytosine in CpG islands), altering gene expression without changing the underlying sequence. | +| [DNA Methylation](https://hubmapconsortium.github.io/ingest-validation-tools/dna-methylation/current/) [](dna-methylation "Attribute description") | DNA methylation is a critical, reversible epigenetic mechanism in biomedical research involving the addition of a methyl group to DNA (usually cytosine in CpG islands), altering gene expression without changing the underlying sequence. | | [Enhanced SRS](https://www.nature.com/articles/s41467-019-13230-1) [](enhancedsrs "Attribute description")| Refers to improvements made to Stimulated Raman Scattering (SRS), a technique used in microscopy and spectroscopy for chemical imaging and analysis. These enhancements aim to improve sensitivity, spatial resolution, and other capabilities of SRS. Link to [Enhanced SRS directory schema](https://hubmapconsortium.github.io/ingest-validation-tools/enhanced-srs/current/). | | [GeoMx](https://nanostring.com/products/geomx-digital-spatial-profiler/spatial-multiomics-enabled-with-geomx-dsp/) [](geomx "Attribute description")| A platform for spatial biology that analyzes RNA and protein expression within tissue sections which allows for non-destructive, in situ profiling of gene expression and protein levels from specific regions of interest (ROIs) within a tissue. Link to [GeoMx directory schema](https://hubmapconsortium.github.io/ingest-validation-tools/geomx-ngs/current/). | | [HiFi](https://www.researchgate.net/publication/370676672_HiFi-Slide_spatial_RNA-Sequencing_v2) [](hifi "Attribute description")| High-Fidelity Spatial Transcriptomic Slide (HiFi-Slide) sequencing, a super-resolution spatial transcriptomics sequencing technology, captures and spatially resolves genome-wide RNA expression in a submicron resolution for fresh-frozen tissue. Link to [HiFi directory schema](https://hubmapconsortium.github.io/ingest-validation-tools/hifi-slide/current/). | @@ -30,6 +30,7 @@ A list of available dataset types (data types from multiple supported assays), w | [MERFISH](https://pubmed.ncbi.nlm.nih.gov/27241748/) [](merfish "Attribute description") | A spatial transcriptomics technology that allows for the simultaneous imaging of hundreds to thousands of RNA species within single cells, providing both copy number and spatial distribution information. Link to [MERFISH directory schema](https://hubmapconsortium.github.io/ingest-validation-tools/merfish/current/). | | [MUSIC](https://www.nature.com/articles/s41586-024-07239-w) [](MUSIC "Attribute description") | A sequencing assay that allows profiling of gene expression, co-complexed DNA sequences, and RNA-chromatin interactions from the same single-cell nucleus. Both RNA and fragmented DNA within a nucleus are labelled with a unique cell barcode, enabling identification and matching of RNA and DNA sequences. Link to [MUSIC directory schema](https://hubmapconsortium.github.io/ingest-validation-tools/music/current/). | | [MxIF](https://pmc.ncbi.nlm.nih.gov/articles/PMC9959383/#) [](thicksectionmultiphotonmxif "Attribute description") | One version of MXIF (multiplexed fluorescence microscopy), an imaging platform whereby a large number of cellular and histological markers can be investigated on a single tissue section. Link to [TSM MxIF directory schema](https://hubmapconsortium.github.io/ingest-validation-tools/thick-section-multiphoton-mxif/current/).| +| [Paired Tag](https://hubmapconsortium.github.io/ingest-validation-tools/paired-tag/current/) [](paired-tag "Attribute description") | Paired-Tag (parallel analysis of individual cells for RNA expression and DNA from targeted tagmentation by sequencing) is a high-throughput, single-cell multiomics method that simultaneously maps histone modifications and gene expression (transcriptome) in the same single cell. | | Pixel-seqV2 [](pixel-seqv2 "Attribute description") | Pixel-seqV2 is a spatial transcriptomics method that utilizes polony gels to capture and sequence RNA, proteins or other molecules in tissues with high resolution. These polony gels are arrays of micron-scale DNA clusters, each containing a unique barcode, allowing for the mapping of molecules within their original spatial context in a tissue, thereby allowing researchers to study the spatial organization of cells and their gene expression profiles within tissues.| | [RNAseq](https://docs.hubmapconsortium.org/assays/rnaseq) [](RNAseq "Attribute description") | While bulk RNAseq elucidates the average gene expression profile in cells comprising a tissue sample, single-cell RNAseq employs per-cell and per-molecule barcoding to enable single-cell resolution of the gene expression profile. Link to [RNAseq directory schema](https://hubmapconsortium.github.io/ingest-validation-tools/rnaseq/current/).| | [RNAseq with Probes](https://pmc.ncbi.nlm.nih.gov/articles/PMC5717752/#) [](RNAseq-(with-probes) "Attribute description") | Uses probes to capture and enrich specific regions of the RNA for targeted sequencing, allowing for in-depth analysis of those regions. Link to [RNAseq with Probes directory schema](https://hubmapconsortium.github.io/ingest-validation-tools/rnaseq-with-probes/current/).| diff --git a/docs/assays/metadata/paired-tag.md b/docs/assays/metadata/paired-tag.md new file mode 100644 index 00000000..1b9d21ae --- /dev/null +++ b/docs/assays/metadata/paired-tag.md @@ -0,0 +1,26 @@ +--- +layout: page-triary +--- + +# Paired Tag Metadata Attributes + +Fields that are collected for Paired Tag data, available at ```dataset.metadata.``` +  + +* indicates a required field + +| Attribute | Type | Description | Allowable Values | +|------|------|-------------|-------------------| +| parent_sample_id * | | The unique identifier from HuBMAP or SenNet for the sample (such as a block, section, or suspension) used to perform the assay. For instance, in an RNAseq assay, the parent sample would be the suspension, while in imaging assays, it would be the tissue section. If the assay is derived from multiple parent samples, this field should contain a comma-separated list of identifiers. Example: HBM386.ZGKG.235, HBM672.MKPK.442 | | +| lab_id | | A locally assigned identifier provided by the data provider for the dataset. It is used to reference an external metadata record that may be maintained independently, enabling traceability and supporting provenance tracking. Example: Visium_9OLC_A4_S1 | | +| preparation_protocol_doi * | | The DOI for the protocols.io page that details the assay or the procedures used for sample procurement and preparation. For example, in the case of an imaging assay, the protocol may start with tissue section staining and end with the generation of an OME-TIFF file. The documented protocol should also include any image processing steps involved in producing the final OME-TIFF. Example: https://dx.doi.org/10.17504/protocols.io.eq2lyno9qvx9/v1 | | +| dataset_type * | | The specific type of dataset being produced. Example: RNAseq | ```10X Multiome``` ```2D Imaging Mass Cytometry``` ```4i``` ```ATACseq``` ```Auto-fluorescence``` ```Cell DIVE``` ```CODEX``` ```COMET``` ```Confocal``` ```CosMx Proteomics``` ```CosMx Transcriptomics``` ```CyCIF``` ```CyTOF``` ```DART-FISH``` ```DBiT-seq``` ```DESI``` ```DNA Methylation``` ```Enhanced Stimulated Raman Spectroscopy (SRS)``` ```FACS``` ```GeoMx (nCounter)``` ```GeoMx (NGS)``` ```HiFi-Slide``` ```Histology``` ```iCLAP``` ```Illumina Spatial ver0``` ```LC-MS``` ```Light Sheet``` ```MACSima``` ```MALDI``` ```MERFISH``` ```MIBI``` ```Molecular Cartography``` ```MPLEx``` ```MS Lipidomics``` ```MUSIC``` ```nanoSPLITS``` ```Olink``` ```PhenoCycler``` ```Pixel-seqV2``` ```Raman Imaging``` ```Resolve``` ```RNAseq``` ```RNAseq (with probes)``` ```Second Harmonic Generation (SHG)``` ```Seq-Scope``` ```seqFISH``` ```SIMS``` ```Singular Genomics G4X``` ```SNARE-seq2``` ```STARmap``` ```Stereo-seq``` ```Thick section Multiphoton MxIF``` ```Virtual Histology``` ```Visium (no probes)``` ```Visium (with probes)``` ```Visium HD``` ```Xenium``` | +| contributors_path * | | The name of the file containing the ORCID IDs for all contributors to this dataset. Example: ./contributors.csv | | +| data_path * | | The top-level directory containing the raw and/or processed data. For a single dataset upload, this might be represented as ".", whereas for a data upload containing multiple datasets, this would be the directory name for the respective dataset. For example, if the data is within a directory named "TEST001-RK", use the syntax "./TEST001-RK" for this field. If there are multiple directory levels, use the format "./TEST001-RK/Run1/Pass2", where "Pass2" is the subdirectory where the single dataset's data is stored. This is an internal metadata field used solely for data ingestion. Example: ./TEST001-RK | | +| capture_batch_id * | | A lab-assigned identifier used to indicate which cells or nuclei were captured together in the same run. For example, in a 10X Genomics Chromium Controller workflow, this could be the chip ID used to trace datasets derived from a single capture event. This ID helps users identify samples processed together on the same device. To avoid conflicts across institutions, it is recommended to prefix the ID with the name of the sequencing center. Example: Broad_Batch1234 | | +| preparation_instrument_vendor * | | The company that manufactures the instrument used to prepare the sample (e.g., for staining or other processing steps) prior to the assay. If the instrument was custom-built or developed internally, enter "In-House". If no sample preparation occurred, enter "Not applicable". Example: 10X Genomics | ```10x Genomics``` ```Akoya Biosciences``` ```Hamamatsu``` ```HTX Technologies``` ```In-House``` ```Ionpath``` ```Leica Biosystems``` ```Not applicable``` ```Roche Diagnostics``` ```SunChrom``` ```Thermo Fisher Scientific``` | +| preparation_instrument_model * | | The specific model of the instrument used for sample preparation, such as staining. Manufacturers may offer multiple models with varying features or sensitivities, which can influence how the sample is processed and how the resulting data is interpreted. If no sample preparation occurred, enter "Not applicable". Example: Chromium X | ```AutoStainer XL``` ```Chromium Connect``` ```Chromium Controller``` ```Chromium iX``` ```Chromium X``` ```Custom``` ```Discovery Ultra``` ```EVOS M7000``` ```M3+ Sprayer``` ```M5 Sprayer``` ```NanoZoomer S210``` ```NanoZoomer S360``` ```NanoZoomer S60``` ```Not applicable``` ```ST5020 Multistainer``` ```Sublimator``` ```SunCollect Sprayer``` ```TM-Sprayer``` ```Visium CytAssist``` | +| preparation_instrument_kit * | | The reagent kit used in conjunction with the preparation instrument for sample processing (e.g., staining or labeling). If the reagent kit was prepared using a custom protocol, enter “Custom”. Example: 10X Genomics; Chromium Next GEM Chip G Single Cell Kit, 48 rxns; PN 1000120 | ```10x Genomics; Chromium Chip E Single Cell ATAC Kit, 48 rxns; PN 1000155``` ```10x Genomics; Chromium GEM-X Single Cell 3' Kit v4, 16 rxns; PN 1000691``` ```10X Genomics; Chromium Next GEM Chip G Single Cell Kit, 16 rxns; PN 1000127``` ```10X Genomics; Chromium Next GEM Chip G Single Cell Kit, 48 rxns; PN 1000120``` ```10X Genomics; Chromium Next GEM Chip K Automated Single Cell Kit, 48 rxns; PN 1000289``` ```10X Genomics; Chromium Next GEM Chip K Single Cell Kit, 16 rxns; PN 1000287``` ```10X Genomics; Chromium Next GEM Chip K Single Cell Kit, 48 rxns; PN 1000286``` ```10X Genomics; Chromium Next GEM Chip Q Single Cell Kit, 16 rxns; PN 1000422``` ```10X Genomics; Chromium NextGem Single Cell Multiome ATAC + Gene Expression Reagent Bundle, 16 rxn; PN 1000283``` ```10X Genomics; Chromium NextGem Single Cell Multiome ATAC + Gene Expression Reagent Bundle, 4 rxn; PN 1000285``` ```10x Genomics; Chromium Single Cell B Chip Kit, 16 rxns; PN 1000074``` ```10X Genomics; Visium FFPE Reagent Kit v2-Small, PN 1000436``` ```Custom``` | +| number_of_pre_amplification_pcr_cycles * | | The number of PCR cycles performed following the Chromium Controller step and before the suspension is separated and library construction begins. Example: 7 | | +| non_global_files | | A semicolon separated list of non-shared files to be included in the dataset. The path assumes the files are located in the "TOP/non-global/" directory. For example, for the file is TOP/non-global/lab_processed/images/1-tissue-boundary.geojson the value of this field would be "./lab_processed/images/1-tissue-boundary.geojson". After ingest, these files will be copied to the appropriate locations within the respective dataset directory tree. | | +| metadata_schema_id * | | The unique string identifier for the metadata specification version, which is easily interpretable by computers for purposes of data validation and processing. Example: 22bc762a-5020-419d-b170-24253ed9e8d9 | | diff --git a/scripts/metaBuilder/README.md b/scripts/metaBuilder/README.md deleted file mode 100644 index 2b07f37c..00000000 --- a/scripts/metaBuilder/README.md +++ /dev/null @@ -1,53 +0,0 @@ -# Assay Metadata Markdown Generator - -This directory contains scripts to fetch, process, and convert HuBMAP assay metadata templates into human-readable Markdown documentation. - -## Workflow Overview -1. **Edit assays.txt** - - List the URLs of the HuBMAP ingest-validation-tools assay metadata schema pages you want to process, one per line. Under each URL, include the intended description of the assay (also on its own line; no line breaks) - - Example: - ``` - https://hubmapconsortium.github.io/ingest-validation-tools/iclap/current/ - iCLAP (individual-nucleotide resolution UV-crosslinking and affinity purification) is a specialized...[etc] - - https://hubmapconsortium.github.io/ingest-validation-tools/comet/current/ - COMET is a technique used to measure DNA damage in individual cells...[etc] - ``` - -2. **Run the Pipeline** - - Use the provided script to generate all Markdown documentation in one step: - ```sh - python3 generate_all_md.py - ``` - - This will: - 1. Run `fetchMeta.py` to fetch the associated metadata information from CEDAR and generate JSON files in `metaJSON/`. - 2. Convert each JSON file in `metaJSON/` to a Markdown file in `toMD/` using `json_to_md.py`. - -3. **Output** - - JSON files: `metaJSON/` - - Markdown files: `toMD/` - - After Generation, processed files are moved to `old/` subdirectories within each folder. (This folder's included in .gitignore) - -## Script Descriptions - -- **fetchMeta.py**: Fetches and processes metadata templates from the CEDAR API using the Template IDs. Outputs a JSON file for each assay. -- **json_to_md.py**: Converts a single JSON file to a Markdown table. Used by the batch script. -- **generate_all_md.py**: Runs the full pipeline: fetches all templates and generates Markdown for each. - -## Requirements -- Python 3 -- `requests` library (`pip install requests`) - -## Notes -- The URLs used in the assays.txt file can be found [here](https://hubmapconsortium.github.io/ingest-validation-tools/current/). -- The scripts will create the `metaJSON/` and `toMD/` folders if they do not exist. -- Re-running the Script will add additional rows to each index table, but overwrite the existing assay metadata page file (vs creating a duplicate) - -## Example Usage - -```sh -python3 generate_all_md.py -``` - -This will produce Markdown documentation for all specified templates. diff --git a/scripts/metaBuilder/assays.txt b/scripts/metaBuilder/assays.txt deleted file mode 100644 index 3011e94a..00000000 --- a/scripts/metaBuilder/assays.txt +++ /dev/null @@ -1,5 +0,0 @@ -https://hubmapconsortium.github.io/ingest-validation-tools/iclap/current/ -iCLAP (individual-nucleotide resolution UV-crosslinking and affinity purification) is a specialized, high-stringency technique designed to map the specific RNA binding sites of RNA-binding proteins (RBPs) at the single-nucleotide level. - -https://hubmapconsortium.github.io/ingest-validation-tools/comet/current/ -COMET is a technique used to measure DNA damage in individual cells. The name comes from the shape that damaged DNA fragments form when they migrate out of a cell's nucleus under an electric field, resembling a comet with a head and a tail. This assay is widely used in genetics research to study DNA damage from factors like radiation, chemicals, and environmental exposure. \ No newline at end of file diff --git a/scripts/metaBuilder/fetchMeta.py b/scripts/metaBuilder/fetchMeta.py deleted file mode 100644 index 57fa9eb2..00000000 --- a/scripts/metaBuilder/fetchMeta.py +++ /dev/null @@ -1,137 +0,0 @@ -import os -import json -import re -import requests -import argparse - -BASE_TEMPLATE_URL = "https://open.metadatacenter.org/templates/https%3A%2F%2Frepo.metadatacenter.org%2Ftemplates%2F{}" - -strip = [ - "@id", - "@type", - "schema:name", - "schema:description", -] - -def fetch_data(url): - response = requests.get(url) - response.raise_for_status() - return response.json() - -def process_assay(template_id): - url = BASE_TEMPLATE_URL.format(template_id) - assay = fetch_data(url) - props = assay.get("properties", {}) - assay_details = { - "assayName": assay.get("schema:name", ""), - "properties": [] - } - print(f" Captured {assay_details['assayName']} details...") - for property, prop_val in props.items(): - if property in strip: - continue - literal_values = [] - name = prop_val.get("skos:prefLabel", "") - shortcode = prop_val.get("schema:name", "") - type_ = prop_val.get("_ui", {}).get("inputType", "") - constraints = [] - literals = [] - value = "" - - value_constraints = prop_val.get("_valueConstraints", {}) - # Assigned Values - if value_constraints.get("branches"): - branches = value_constraints.get("branches") - - if isinstance(branches, list): - for branch in branches: - name = branch.get("name") - uri = branch.get("uri") - elif isinstance(branches, dict): - name = branches.get("name") - uri = branches.get("uri") - - constraints = branches - if isinstance(constraints, list) and constraints and isinstance(constraints[0], dict) and 'uri' in constraints[0]: - value_HRAVS = str(constraints[0]['uri']) - value = fetch_by_HRAVS(constraints) - type_ = "Assigned Value" - - # YES / NO - elif value_constraints.get("literals"): - literals = value_constraints["literals"] - for literal in literals: - if isinstance(literal, dict): - for key in ("label", "value", "prefLabel"): - if key in literal: - literal_values.append(str(literal[key])) - break - else: - literal_values.append(str(literal)) - else: - literal_values.append(str(literal)) - type_ = "Radio" - # value = ",".join(literal_values) - value =literal_values - - else: - type_ = prop_val.get("_ui", {}).get("inputType", "") - value = "" - - if name: - assay_details["properties"].append({ - "attribute": shortcode, - "label": name, - "type": type_.title() if isinstance(type_, str) else type_, - "description": prop_val.get("schema:description", ""), - "value": value, - "required": value_constraints.get("requiredValue") - }) - - title = f"{assay.get('schema:name', '')}.json" - title = re.sub(r"[\s;]+", "-", title) - os.makedirs("metaJSON", exist_ok=True) - with open(os.path.join("metaJSON", title), "w") as f: - json.dump(assay_details["properties"], f) - -def fetch_by_HRAVS(details): - detail = details[0]; - if detail.get("uri"): - value_HRAVS = str(detail.get("uri")) - - match = re.search(r'#HRAVS_(\d+)', value_HRAVS) - if match: - hravs_id = match.group(1) - valueSet = fetch_details(hravs_id) - return valueSet - -def fetch_details(url_id): - BASE_URL ="https://data.bioontology.org/ontologies/HRAVS/classes/https%3A%2F%2Fpurl.humanatlas.io%2Fvocab%2Fhravs%23HRAVS_{}" - url = BASE_URL.format(url_id)+"/children?apikey=ad1d9ae5-3781-48d2-a61b-ab243bea22ee&display=prefLabel&no_context=true" - headers = { - "User-Agent": ( - "Mozilla/5.0 (X11; Linux x86_64) " - "AppleWebKit/537.36 (KHTML, like Gecko) " - "Chrome/115.0.0.0 Safari/537.36" - ) - } - response = requests.get(url, headers=headers) - response.raise_for_status() - try: - data = json.loads(response.text) - collection = data.get("collection", []) - values = [item.get("prefLabel") for item in collection if "prefLabel" in item] - except Exception as e: - print(f"Error parsing JSON for {url_id}: {e}") - - return values - -def main(): - parser = argparse.ArgumentParser(description="Fetch and process multiple HuBMAP assay metadata templates.") - parser.add_argument('--templateID', action='append', required=True, help='Template ID to process (can be used multiple times)') - args = parser.parse_args() - for tid in args.templateID: - process_assay(tid) - -if __name__ == "__main__": - main() \ No newline at end of file diff --git a/scripts/metaBuilder/generate_all_md.py b/scripts/metaBuilder/generate_all_md.py deleted file mode 100644 index f3961e56..00000000 --- a/scripts/metaBuilder/generate_all_md.py +++ /dev/null @@ -1,268 +0,0 @@ -import os -import subprocess -import glob -import threading -import itertools -import sys -import time -import requests -from bs4 import BeautifulSoup -import re - -def spinner(msg, stop_event): - for c in itertools.cycle('⠷⠯⠟⠻⠽⠾'): - if stop_event.is_set(): - break - sys.stdout.write(f'\r{msg} {c}') - sys.stdout.flush() - time.sleep(0.1) - sys.stdout.write('\r' + ' ' * (len(msg) + 2) + '\r') - -def get_assay_info(assay_url): - response = requests.get(assay_url) - response.raise_for_status() - soup = BeautifulSoup(response.text, 'html.parser') - - # 1. Assay name (try h1 without 'category' class, then fallback) - h1s = soup.find_all('h1') - assay_name = None - for h1 in h1s: - if 'category' not in (h1.get('class') or []): - assay_name = h1.text.strip() - break - if not assay_name: - assay_name = soup.title.text.strip() if soup.title else "Unknown" - - # 2. Versions (found between tags within tags) - versions = [] - current_version = None - for summary in soup.find_all('summary'): - b_tag = summary.find('b') - if b_tag: - version = b_tag.text.strip() - versions.append(version) - if "use this one" in summary.text.lower(): - current_version = version - - # 3. Metadata schema links - meta_links = [] - template_ids = [] - for link in soup.find_all('a', href=True): - href = link['href'] - if href.startswith('https://openview.metadatacenter.org/templates/https:%2F%2Frepo.metadatacenter.org%2Ftemplates%2F'): - meta_links.append(href) - # Extract template ID - tid = href.split('templates%2F')[-1] - tid = tid.split('?')[0] # Remove any query params - template_ids.append(tid) - - return { - "assay_name": assay_name, - "versions": versions, - "current_version": current_version, - "metadata_schema_links": meta_links, - "template_ids": template_ids - } - -# --------- MAIN FINALIZATION --------- - -# Paths -script_dir = os.path.dirname(__file__) -toMD_dir = os.path.join(script_dir, "toMD") -layout_template_path = os.path.join(script_dir, "pageLayout.md") -output_dir = os.path.abspath(os.path.join(script_dir, "../../docs/assays/metadata/")) -index_md_path = os.path.join(output_dir, "index.md") -new_index_md_path = os.path.join(output_dir, "new-metadata-test.md") -assays_txt_path = os.path.join(script_dir, "assays.txt") -meta_json_dir = os.path.join(script_dir, 'metaJSON') - - -# Step 1: Process assays.txt as link-description pairs -assay_infos = [] -all_template_ids = [] -assay_descriptions = {} -with open(assays_txt_path, "r") as f: - lines = [l.strip() for l in f if l.strip()] - for i in range(0, len(lines), 2): - url = lines[i] - description = lines[i+1] if i+1 < len(lines) else "" - info = get_assay_info(url) - assay_infos.append(info) - print(f"Processed: {info['assay_name']}") - all_template_ids.extend(info['template_ids']) - assay_descriptions[info['assay_name']] = description - -# Step 2: Call fetchMeta.py with all collected template IDs -if all_template_ids: - template_id_args = [] - for tid in all_template_ids: - template_id_args.extend(["--templateID", tid]) - stop_event = threading.Event() - spin_thread = threading.Thread(target=spinner, args=('Fetching metadata', stop_event)) - spin_thread.start() - subprocess.run(['python3', 'fetchMeta.py'] + template_id_args, check=True) - stop_event.set() - spin_thread.join() - print('Fetching... done.') -else: - print("No template IDs found to pass to fetchMeta.py") - -# Step 3: Find all JSON files in metaJSON and run json_to_md.py on each -json_files = glob.glob(os.path.join(meta_json_dir, '*.json')) -for json_file in json_files: - stop_event = threading.Event() - spin_thread = threading.Thread(target=spinner, args=(f'Converting {os.path.basename(json_file)}', stop_event)) - spin_thread.start() - subprocess.run(['python3', 'json_to_md.py', json_file], check=True) - stop_event.set() - spin_thread.join() - -print('All markdown files generated into ./toMD') - -# --------- FINAL STEP: Build new markdown file from pageLayout.md --------- - -def build_final_md(assay_name, current_version, table_md_path, layout_template_path, output_dir): - # Read the layout template - with open(layout_template_path, "r") as f: - layout = f.read() - # Read the generated table - with open(table_md_path, "r") as f: - table = f.read() - # Fill in the template - content = layout.replace("{AssayName}", assay_name) - content = content.replace("{Version NUMBER (current)}", current_version) - content = content.replace("{Table Here}", table) - # Write to the output directory - out_path = os.path.join(output_dir, f"{assay_name}.md") - with open(out_path, "w") as f: - f.write(content) - print(f"Saved: {out_path}") - return out_path - -def update_index_md(index_md_path, assay_name, description=None): - # Add a row to the table in index.md for the new assay, keeping alphabetical order - with open(index_md_path, "r") as f: - lines = f.readlines() - # Find all table rows (lines starting with "|", but skip the header and separator) - table_rows = [(i, line) for i, line in enumerate(lines) if line.strip().startswith("|")] - header_idx = table_rows[0][0] - separator_idx = table_rows[1][0] - data_rows = table_rows[2:] # actual data rows - - # Prepare the new row - link = f"[attributes]({assay_name})" - desc = description or "" - new_row = f"| {assay_name} | {link} | {desc} |\n" - - # Extract assay names from data rows for sorting - def get_row_assay_name(row): - # Assay name is between first and second '|' - parts = row.split('|') - return parts[1].strip().lower() - - # Insert the new row into the correct alphabetical position - inserted = False - for idx, (line_idx, row) in enumerate(data_rows): - existing_name = get_row_assay_name(row) - if assay_name.lower() < existing_name: - insert_at = line_idx - lines.insert(insert_at, new_row) - inserted = True - break - if not inserted: - # If not inserted, append at the end of the table - last_table_row = table_rows[-1][0] - lines.insert(last_table_row + 1, new_row) - - with open(index_md_path, "w") as f: - f.writelines(lines) - print(f"Updated: {index_md_path}") - -# --- NEW: Update new-metadata-test.md in alphabetical order --- -def update_new_index_md(new_index_md_path, assay_name, description=None, schema_url=None): - """ - Add a row to the table in new-metadata-test.md for the new assay, keeping alphabetical order. - The row format is: - | Dataset Type | Description | - | [AssayName](schema_url) [](AssayName "Attribute description") | description | - """ - with open(new_index_md_path, "r") as f: - lines = f.readlines() - # Find all table rows (lines starting with '|', but skip the header and separator) - table_rows = [(i, line) for i, line in enumerate(lines) if line.strip().startswith("|")] - if len(table_rows) < 2: - print(f"Table not found in {new_index_md_path}") - return - header_idx = table_rows[0][0] - separator_idx = table_rows[1][0] - data_rows = table_rows[2:] # actual data rows - - # Prepare the new row - # Use schema_url if provided, else just the assay name - if schema_url: - assay_link = f"[{assay_name}]({schema_url}) []({assay_name} \"Attribute description\")" - else: - assay_link = f"{assay_name} []({assay_name} \"Attribute description\")" - desc = description or "" - new_row = f"| {assay_link} | {desc} |\n" - - # Extract assay names from data rows for sorting - def get_row_assay_name(row): - # Assay name is between first and second '|', but may contain markdown links - parts = row.split('|') - # Remove markdown link if present - name = parts[1].strip() - if name.startswith('['): - name = name.split(']')[0][1:] - return name.lower() - - # Insert the new row into the correct alphabetical position - inserted = False - for idx, (line_idx, row) in enumerate(data_rows): - existing_name = get_row_assay_name(row) - if assay_name.lower() < existing_name: - insert_at = line_idx - lines.insert(insert_at, new_row) - inserted = True - break - if not inserted: - # If not inserted, append at the end of the table - last_table_row = table_rows[-1][0] - lines.insert(last_table_row + 1, new_row) - - with open(new_index_md_path, "w") as f: - f.writelines(lines) - print(f"Updated: {new_index_md_path}") - - -assay_version_map = {info['assay_name']: info['current_version'] for info in assay_infos} - -# Find all generated markdown tables in toMD -md_tables = glob.glob(os.path.join(toMD_dir, "*.md")) -for table_md_path in md_tables: - assay_name = os.path.splitext(os.path.basename(table_md_path))[0] - # Use the actual current_version if available - current_version = assay_version_map.get(assay_name, "Version 2 (current)") - out_path = build_final_md(assay_name, current_version, table_md_path, layout_template_path, output_dir) - # Use the description from assays.txt for both index.md and new-metadata-test.md - description = assay_descriptions.get(assay_name, "") - update_index_md(index_md_path, assay_name, description=description) - update_new_index_md(new_index_md_path, assay_name, description=description) - -# After processing, move parsed JSON and Markdown files to /old - -# Ensure the /old directories exist -meta_json_old_dir = os.path.join(meta_json_dir, "old") -toMD_old_dir = os.path.join(toMD_dir, "old") -os.makedirs(meta_json_old_dir, exist_ok=True) -os.makedirs(toMD_old_dir, exist_ok=True) - -# Move JSON files -for json_file in json_files: - dest = os.path.join(meta_json_old_dir, os.path.basename(json_file)) - os.rename(json_file, dest) - -# Move Markdown files -for md_file in md_tables: - dest = os.path.join(toMD_old_dir, os.path.basename(md_file)) - os.rename(md_file, dest) \ No newline at end of file diff --git a/scripts/metaBuilder/json_to_md.py b/scripts/metaBuilder/json_to_md.py deleted file mode 100644 index 5e4f6a72..00000000 --- a/scripts/metaBuilder/json_to_md.py +++ /dev/null @@ -1,57 +0,0 @@ -import os -import json -import sys - -def json_to_markdown(json_path, md_path=None): - with open(json_path, 'r') as f: - data = json.load(f) - # Use the filename (without extension) as the header - base = os.path.basename(json_path) - header = os.path.splitext(base)[0] - # If the JSON has an 'assayName' field, prefer that for the header - if isinstance(data, dict) and 'assayName' in data: - header = data['assayName'] - properties = data.get('properties', []) - elif isinstance(data, dict) and 'properties' in data: - properties = data['properties'] - elif isinstance(data, list): - properties = data - else: - raise ValueError('Unrecognized JSON structure') - - md_lines = [] - md_lines.append("| Attribute Name | Type | Description | Allowable Values | Required |") - md_lines.append("|---------------|------|-------------|------------------|----------|") - for prop in properties: - attr = prop.get('attribute', '') - typ = prop.get('type', '') - desc = prop.get('description', '').replace('\n', ' ') - # Allowable Values: join list or show as string - val = prop.get('value', '') - if isinstance(val, list): - val = ', '.join(f'```{str(v)}```' for v in val) - elif val is None: - val = '' - else: - val = f'```{val}```' if val else '' - req = str(prop.get('required', '')) - md_lines.append(f"| {attr} | {typ} | {desc} | {val} | {req} |") - md_content = '\n'.join(md_lines) - if not md_path: - base = os.path.basename(json_path) - md_name = os.path.splitext(base)[0] + ".md" - md_dir = os.path.join(os.path.dirname(json_path), "../toMD") - md_dir = os.path.abspath(md_dir) - os.makedirs(md_dir, exist_ok=True) - md_path = os.path.join(md_dir, md_name) - with open(md_path, 'w') as f: - f.write(md_content) - # print(f"{header} markdown exported to ./toMD") - -if __name__ == "__main__": - if len(sys.argv) < 2: - print("Usage: python json_to_md.py [output_md_file]") - sys.exit(1) - json_file = sys.argv[1] - md_file = sys.argv[2] if len(sys.argv) > 2 else None - json_to_markdown(json_file, md_file) diff --git a/scripts/metaBuilder/pageLayout.md b/scripts/metaBuilder/pageLayout.md deleted file mode 100644 index 33cd771e..00000000 --- a/scripts/metaBuilder/pageLayout.md +++ /dev/null @@ -1,12 +0,0 @@ ---- -layout: page ---- -# {AssayName} - -
{Version NUMBER (current)} - -## {Version NUMBER (current)} - -{Table Here} - -
\ No newline at end of file diff --git a/scripts/metadata-generator/generate.py b/scripts/metadata-generator/generate.py new file mode 100755 index 00000000..b8cea36f --- /dev/null +++ b/scripts/metadata-generator/generate.py @@ -0,0 +1,875 @@ +#!/usr/bin/env python3 +"""Generate HuBMAP assay metadata documentation from GitHub issues. + +Issue format: + * Title: assay name + * Body: assay description and exactly one ingest-validation-tools URL + +Usage: + python3 scripts/metadata-generator/generate.py 70 71 + python3 scripts/metadata-generator/generate.py https://github.com/OWNER/REPO/issues/70 + python3 scripts/metadata-generator/generate.py --dry-run 70 + python3 scripts/metadata-generator/generate.py --self-test + +The generator uses the public HRA Knowledge Graph for HRAVS allowable values. +It has no third-party Python dependencies and requires no API key. +""" + +from __future__ import annotations + +import argparse +import json +import re +import subprocess +import sys +from dataclasses import dataclass +from html.parser import HTMLParser +from pathlib import Path +from typing import Any, Callable +from urllib.error import HTTPError, URLError +from urllib.parse import quote, unquote, urlencode, urlparse +from urllib.request import Request, urlopen + + +IVT_HOST = "hubmapconsortium.github.io" +IVT_PATH_PREFIX = "/ingest-validation-tools/" +CEDAR_OPEN_HOST = "open.metadatacenter.org" +OPENVIEW_HOST = "openview.metadatacenter.org" +HRA_SPARQL_URL = "https://lod.humanatlas.io/sparql" +HRAVS_GRAPH = "https://purl.humanatlas.io/vocab/hravs" +HRAVS_CONCEPT_PATTERN = re.compile( + r"https://purl\.humanatlas\.io/vocab/hravs#[A-Za-z0-9_]+" +) +METADATA_DIR = Path("docs/assays/metadata") +INDEX_PATH = METADATA_DIR / "index.md" + +TEXT_ICON = '' +RADIO_ICON = '' +VALUE_ICON = ( + '' +) + + +class GeneratorError(RuntimeError): + """An expected, user-actionable generator failure.""" + + +@dataclass(frozen=True) +class Issue: + number: int + title: str + body: str + url: str + repo: str + + +@dataclass(frozen=True) +class Assay: + name: str + description: str + schema_index_url: str + slug: str + + +@dataclass(frozen=True) +class SchemaLink: + url: str + version: str + + +@dataclass(frozen=True) +class Field: + attribute: str + description: str + required: bool + kind: str + allowable_values: tuple[str, ...] = () + + +def http_get(url: str, *, headers: dict[str, str] | None = None) -> bytes: + request_headers = {"User-Agent": "HuBMAP-documentation-metadata-generator/1.0"} + request_headers.update(headers or {}) + request = Request(url, headers=request_headers) + try: + with urlopen(request, timeout=45) as response: + return response.read() + except HTTPError as exc: + raise GeneratorError(f"HTTP {exc.code} while fetching {url}") from exc + except URLError as exc: + raise GeneratorError(f"Could not fetch {url}: {exc.reason}") from exc + + +def detect_repo(root: Path) -> str: + result = subprocess.run( + ["git", "remote", "get-url", "origin"], + cwd=root, + check=False, + capture_output=True, + text=True, + ) + if result.returncode: + raise GeneratorError("Could not determine the GitHub repository from git remote 'origin'.") + remote = result.stdout.strip() + match = re.search(r"github\.com[/:]([^/]+/[^/]+?)(?:\.git)?$", remote) + if not match: + raise GeneratorError(f"The origin remote is not a recognizable GitHub repository: {remote}") + return match.group(1) + + +def parse_issue_reference(reference: str, default_repo: str) -> tuple[str, int]: + if reference.isdigit(): + return default_repo, int(reference) + match = re.fullmatch( + r"https?://github\.com/([^/]+/[^/]+)/issues/(\d+)/?", reference + ) + if not match: + raise GeneratorError( + f"Invalid issue reference {reference!r}; use an issue number or full GitHub issue URL." + ) + return match.group(1), int(match.group(2)) + + +def fetch_issue(reference: str, default_repo: str, root: Path) -> Issue: + repo, number = parse_issue_reference(reference, default_repo) + result = subprocess.run( + [ + "gh", + "issue", + "view", + str(number), + "--repo", + repo, + "--json", + "number,title,body,url", + ], + cwd=root, + check=False, + capture_output=True, + text=True, + ) + if result.returncode: + detail = result.stderr.strip() or result.stdout.strip() + raise GeneratorError(f"Could not read {repo} issue #{number} with gh: {detail}") + return issue_from_gh_data(json.loads(result.stdout), repo, number) + + +def issue_from_gh_data(data: dict[str, Any], repo: str, number: int) -> Issue: + issue_url = str(data["url"]) + if re.search(r"/pull/\d+/?$", urlparse(issue_url).path): + raise GeneratorError( + f"{repo}#{number} is a pull request, not an issue; " + "check the assay issue number or pass its full issue URL." + ) + return Issue( + number=int(data["number"]), + title=data["title"].strip(), + body=data.get("body", ""), + url=issue_url, + repo=repo, + ) + + +def parse_assay(issue: Issue) -> Assay: + url_matches = list(re.finditer(r"https?://[^\s)>]+", issue.body)) + if not url_matches: + raise GeneratorError( + f"Issue #{issue.number} contains no URLs; check the issue number and " + "repository, then place its ingest-validation-tools URL at the bottom." + ) + + final_match = url_matches[-1] + schema_url = final_match.group().rstrip(".,") + parsed_schema_url = urlparse(schema_url) + if ( + parsed_schema_url.netloc != IVT_HOST + or not parsed_schema_url.path.startswith(IVT_PATH_PREFIX) + ): + raise GeneratorError( + f"The final link in issue #{issue.number} must be an " + "ingest-validation-tools URL." + ) + + parts = [part for part in urlparse(schema_url).path.split("/") if part] + if len(parts) < 2: + raise GeneratorError(f"Could not derive an assay slug from {schema_url}") + slug = parts[-2] if parts[-1] == "current" else parts[-1] + + body_without_url = issue.body[: final_match.start()] + issue.body[final_match.end() :] + raw_blocks = [ + block.strip() + for block in re.split(r"\n\s*\n", body_without_url) + if block.strip() + ] + description_blocks: list[str] = [] + if raw_blocks: + first_lines = [line.strip() for line in raw_blocks[0].splitlines() if line.strip()] + if first_lines and first_lines[0].casefold() == issue.title.casefold(): + first_lines.pop(0) + if first_lines: + description_blocks.append(" ".join(first_lines)) + description_blocks.extend(raw_blocks[1:]) + description = " ".join( + re.sub(r"\s+", " ", block).strip() + for block in description_blocks + if block.strip() + ).strip() + if not issue.title or not description: + raise GeneratorError( + f"Issue #{issue.number} must provide an assay name and description before the URL." + ) + return Assay( + name=issue.title, + description=description, + schema_index_url=schema_url, + slug=slug, + ) + + +class MetadataSchemaParser(HTMLParser): + def __init__(self) -> None: + super().__init__() + self.in_heading = False + self.heading_level = "" + self.heading_text: list[str] = [] + self.in_section = False + self.current_href: str | None = None + self.current_text: list[str] = [] + self.links: list[tuple[str, str]] = [] + + def handle_starttag(self, tag: str, attrs: list[tuple[str, str | None]]) -> None: + if tag in {"h1", "h2", "h3"}: + self.in_heading = True + self.heading_level = tag + self.heading_text = [] + elif self.in_section and tag == "a": + self.current_href = dict(attrs).get("href") + self.current_text = [] + + def handle_data(self, data: str) -> None: + if self.in_heading: + self.heading_text.append(data) + if self.current_href is not None: + self.current_text.append(data) + + def handle_endtag(self, tag: str) -> None: + if self.current_href is not None and tag == "a": + self.links.append((self.current_href, "".join(self.current_text).strip())) + self.current_href = None + self.current_text = [] + if self.in_heading and tag == self.heading_level: + heading = " ".join("".join(self.heading_text).split()).casefold() + self.in_section = heading == "metadata schema" + self.in_heading = False + + +def find_current_schema(page: bytes, source_url: str) -> SchemaLink: + parser = MetadataSchemaParser() + parser.feed(page.decode("utf-8", errors="replace")) + matches = [ + (href, text) + for href, text in parser.links + if "use this one" in text.casefold() + and urlparse(href).netloc == OPENVIEW_HOST + ] + if len(matches) != 1: + raise GeneratorError( + f"Expected one '(use this one)' OpenView link under Metadata schema at " + f"{source_url}; found {len(matches)}." + ) + href, text = matches[0] + version = re.sub( + r"\s*\(\s*use this one\s*\)\s*", "", text, flags=re.I + ).strip() + return SchemaLink(url=href, version=version or "Current version") + + +def cedar_api_url(openview_url: str) -> str: + parsed = urlparse(openview_url) + if parsed.netloc != OPENVIEW_HOST or not parsed.path.startswith("/templates/"): + raise GeneratorError(f"Unsupported OpenView template URL: {openview_url}") + template_uri = unquote(parsed.path.removeprefix("/templates/")) + if not re.fullmatch( + r"https://repo\.metadatacenter\.org/templates/[0-9a-f-]+", template_uri + ): + raise GeneratorError(f"Could not extract a CEDAR template ID from {openview_url}") + return f"https://{CEDAR_OPEN_HOST}/templates/{quote(template_uri, safe='')}" + + +class HraValueSource: + """Read and cache HRAVS narrower values from the public HRA SPARQL endpoint.""" + + def __init__( + self, + getter: Callable[..., bytes] = http_get, + ) -> None: + self.getter = getter + self.cache: dict[str, tuple[str, ...]] = {} + + def query(self, sparql: str) -> dict[str, Any]: + url = f"{HRA_SPARQL_URL}?{urlencode({'query': sparql})}" + data = self.getter( + url, headers={"Accept": "application/sparql-results+json"} + ) + return json.loads(data) + + def narrower_values(self, branch: dict[str, Any]) -> tuple[str, ...]: + acronym = branch.get("acronym") or "HRAVS" + uri = branch.get("uri") + if acronym != "HRAVS" or not isinstance(uri, str): + raise GeneratorError( + "Only HRAVS controlled-value branches are supported by the keyless HRA lookup." + ) + if not HRAVS_CONCEPT_PATTERN.fullmatch(uri): + raise GeneratorError(f"Unsupported HRAVS concept URI: {uri}") + if uri in self.cache: + return self.cache[uri] + + sparql = f""" +SELECT DISTINCT ?label +WHERE {{ + GRAPH <{HRAVS_GRAPH}> {{ + <{uri}> ?child . + ?child ?label . + FILTER(LANG(?label) = "" || LANGMATCHES(LANG(?label), "en")) + }} +}} +ORDER BY LCASE(STR(?label)) +""".strip() + data = self.query(sparql) + bindings = data.get("results", {}).get("bindings", []) + values = tuple( + str(binding["label"]["value"]).strip() + for binding in bindings + if isinstance(binding, dict) + and isinstance(binding.get("label"), dict) + and binding["label"].get("value") + ) + if not values: + raise GeneratorError(f"HRA returned no narrower values for {uri}") + self.cache[uri] = values + return values + + def version(self) -> str: + sparql = f""" +SELECT ?version +WHERE {{ + GRAPH <{HRAVS_GRAPH}> {{ + <{HRAVS_GRAPH}> ?version . + }} +}} +LIMIT 1 +""".strip() + data = self.query(sparql) + bindings = data.get("results", {}).get("bindings", []) + if not bindings or not bindings[0].get("version", {}).get("value"): + raise GeneratorError("HRA did not report its current HRAVS graph version.") + return str(bindings[0]["version"]["value"]) + + +def cedar_value_is_required(field_name: str, prop: dict[str, Any]) -> bool: + """Return the required-value flag OpenView displays for a CEDAR field.""" + constraints = prop.get("_valueConstraints") or {} + raw_value = constraints.get("requiredValue", False) + if isinstance(raw_value, bool): + return raw_value + if isinstance(raw_value, str) and raw_value.casefold() in {"true", "false"}: + return raw_value.casefold() == "true" + raise GeneratorError( + f"Field {field_name} has an invalid CEDAR requiredValue: {raw_value!r}" + ) + + +def fields_from_template( + template: dict[str, Any], + narrower_fetcher: Callable[[dict[str, Any]], tuple[str, ...]], +) -> list[Field]: + properties = template.get("properties", {}) + order = template.get("_ui", {}).get("order", []) + names = list(order) + [name for name in properties if name not in order] + value_required_names = { + name + for name, prop in properties.items() + if isinstance(prop, dict) and cedar_value_is_required(name, prop) + } + fields: list[Field] = [] + ignored = {"@context", "@id", "@type", "schema:name", "schema:description"} + for name in names: + if name in ignored: + continue + prop = properties.get(name) + if not isinstance(prop, dict) or "schema:name" not in prop: + continue + constraints = prop.get("_valueConstraints", {}) + branches = constraints.get("branches") or [] + literals = constraints.get("literals") or [] + kind = "text" + values: tuple[str, ...] = () + if branches: + if len(branches) != 1: + raise GeneratorError( + f"Field {name} has {len(branches)} allowable-value branches; expected one." + ) + kind = "allowable" + values = narrower_fetcher(branches[0]) + elif literals: + kind = "radio" + values = tuple( + str(item.get("label") or item.get("value") or item.get("prefLabel")) + if isinstance(item, dict) + else str(item) + for item in literals + ) + # Numeric and link controls intentionally use the documentation text-field icon. + fields.append( + Field( + attribute=str(prop.get("schema:name") or name), + description=str(prop.get("schema:description") or ""), + # CEDAR's template-level `required` array is structural: it says + # the field object must exist. OpenView's required-value marker is + # specifically `_valueConstraints.requiredValue` on the field. + required=name in value_required_names, + kind=kind, + allowable_values=values, + ) + ) + if not fields: + raise GeneratorError("The selected CEDAR template contains no metadata fields.") + return fields + + +def clean_cell(value: str) -> str: + return re.sub(r"\s+", " ", value).strip().replace("|", "\\|") + + +def render_page(assay: Assay, fields: list[Field]) -> str: + lines = [ + "---", + "layout: page-triary", + "---", + "", + f"# {assay.name} Metadata Attributes", + "", + f"Fields that are collected for {assay.name} data, available at ```dataset.metadata.```", + " ", + "", + '* indicates a required field', + "", + "| Attribute | Type | Description | Allowable Values |", + "|------|------|-------------|-------------------|", + ] + for field in fields: + attribute = clean_cell(field.attribute) + if field.required: + attribute += ' *' + icon = {"allowable": VALUE_ICON, "radio": RADIO_ICON}.get( + field.kind, TEXT_ICON + ) + values = " ".join( + f"```{clean_cell(value)}```" for value in field.allowable_values + ) + lines.append( + f"| {attribute} | {icon} | {clean_cell(field.description)} | {values} |" + ) + return "\n".join(lines) + "\n" + + +def index_row(assay: Assay) -> str: + name = clean_cell(assay.name) + description = clean_cell(assay.description) + return ( + f'| [{name}]({assay.schema_index_url}) ' + f'[]({assay.slug} "Attribute description")' + f" | {description} |\n" + ) + + +def row_sort_name(row: str) -> str: + cell = row.split("|", 2)[1].strip() + match = re.match(r"\[([^]]+)]", cell) + if match: + cell = match.group(1) + return re.sub(r"[^a-z0-9]+", "", cell.casefold()) + + +def row_metadata_slug(row: str) -> str | None: + match = re.search(r'\]\(([^ )]+)\s+"Attribute description"\)', row) + return match.group(1) if match else None + + +def update_index(content: str, assay: Assay) -> str: + lines = content.splitlines(keepends=True) + header = next( + (i for i, line in enumerate(lines) if line.startswith("| Dataset Type |")), + None, + ) + if header is None or header + 1 >= len(lines): + raise GeneratorError("Could not find the assay metadata index table.") + start = header + 2 + end = start + while end < len(lines) and lines[end].lstrip().startswith("|"): + end += 1 + rows = [row for row in lines[start:end] if row_metadata_slug(row) != assay.slug] + new_row = index_row(assay) + new_name = row_sort_name(new_row) + insert_at = next( + (index for index, row in enumerate(rows) if row_sort_name(row) > new_name), + len(rows), + ) + rows.insert(insert_at, new_row) + lines[start:end] = rows + return "".join(lines) + + +def repository_root() -> Path: + result = subprocess.run( + ["git", "rev-parse", "--show-toplevel"], + check=False, + capture_output=True, + text=True, + ) + if result.returncode: + raise GeneratorError("Run this command from inside the documentation git repository.") + return Path(result.stdout.strip()) + + +def prepare_issue( + issue: Issue, value_source: HraValueSource +) -> tuple[Assay, SchemaLink, list[Field]]: + assay = parse_assay(issue) + schema_link = find_current_schema( + http_get(assay.schema_index_url), assay.schema_index_url + ) + template = json.loads(http_get(cedar_api_url(schema_link.url))) + fields = fields_from_template(template, value_source.narrower_values) + return assay, schema_link, fields + + +def report_results( + prepared: list[tuple[Issue, Assay, SchemaLink, list[Field]]], + hravs_version: str, + *, + prefix: str, +) -> None: + for issue, assay, schema_link, fields in prepared: + print( + f"{prefix} docs/assays/metadata/{assay.slug}.md from " + f"{issue.repo}#{issue.number} " + f"({schema_link.version}, {len(fields)} fields, " + f"{sum(field.required for field in fields)} required, " + f"HRAVS {hravs_version})" + ) + + +def write_results( + prepared: list[tuple[Issue, Assay, SchemaLink, list[Field]]], + root: Path, + hravs_version: str, +) -> list[Path]: + index_path = root / INDEX_PATH + index_content = index_path.read_text(encoding="utf-8") + for _, assay, _, _ in prepared: + index_content = update_index(index_content, assay) + + output_paths = [] + for _, assay, _, fields in prepared: + output_path = root / METADATA_DIR / f"{assay.slug}.md" + output_path.write_text(render_page(assay, fields), encoding="utf-8") + output_paths.append(output_path) + index_path.write_text(index_content, encoding="utf-8") + report_results(prepared, hravs_version, prefix="Generated") + return output_paths + + +def run_self_tests() -> bool: + import unittest + + class GeneratorTests(unittest.TestCase): + def setUp(self) -> None: + self.issue = Issue( + number=70, + title="DNA Methylation", + body=( + "DNA Methylation\n" + "DNA methylation adds a methyl group to DNA.\n\n" + "https://hubmapconsortium.github.io/ingest-validation-tools/" + "dna-methylation/current/" + ), + url="https://github.com/hubmapconsortium/documentation/issues/70", + repo="hubmapconsortium/documentation", + ) + + def test_issue_body(self) -> None: + assay = parse_assay(self.issue) + self.assertEqual(assay.name, "DNA Methylation") + self.assertEqual( + assay.description, "DNA methylation adds a methyl group to DNA." + ) + self.assertEqual(assay.slug, "dna-methylation") + + def test_issue_body_without_repeated_title(self) -> None: + issue = Issue( + 71, + "New Assay", + "Description only.\n\nhttps://hubmapconsortium.github.io/" + "ingest-validation-tools/new-assay/current/", + "https://github.com/hubmapconsortium/documentation/issues/71", + "hubmapconsortium/documentation", + ) + self.assertEqual(parse_assay(issue).description, "Description only.") + + def test_final_link_selects_schema_with_other_links_present(self) -> None: + issue = Issue( + 72, + "Linked Assay", + "Linked Assay\nSee https://example.org/protocol for context.\n\n" + "https://hubmapconsortium.github.io/ingest-validation-tools/" + "old-assay/current/\n\n" + "https://hubmapconsortium.github.io/ingest-validation-tools/" + "linked-assay/current/", + "https://github.com/hubmapconsortium/documentation/issues/72", + "hubmapconsortium/documentation", + ) + assay = parse_assay(issue) + self.assertEqual(assay.slug, "linked-assay") + self.assertIn("https://example.org/protocol", assay.description) + + def test_issue_references(self) -> None: + self.assertEqual( + parse_issue_reference("70", "hubmapconsortium/documentation"), + ("hubmapconsortium/documentation", 70), + ) + self.assertEqual( + parse_issue_reference( + "https://github.com/somewhere/else/issues/12", "ignored/repo" + ), + ("somewhere/else", 12), + ) + + def test_pull_request_number_is_rejected_as_an_issue(self) -> None: + with self.assertRaisesRegex(GeneratorError, "is a pull request"): + issue_from_gh_data( + { + "number": 110, + "title": "package-lock update", + "body": "", + "url": "https://github.com/hubmapconsortium/documentation/pull/110", + }, + "hubmapconsortium/documentation", + 110, + ) + + def test_current_schema_section(self) -> None: + page = b""" +

Metadata schema

+ Version 2 (use this one) +

Directory schemas

+ Version 3 (use this one) + """ + link = find_current_schema(page, "https://example.test") + self.assertEqual(link.version, "Version 2") + self.assertIn("d70bfe24-e82a-46cb-9369-28ae03660d97", link.url) + + def test_hra_values_are_keyless_and_cached(self) -> None: + calls = [] + + def fake_get(url: str, *, headers: dict[str, str] | None = None) -> bytes: + calls.append((url, headers)) + return json.dumps( + { + "results": { + "bindings": [ + {"label": {"value": "Alpha"}}, + {"label": {"value": "Beta"}}, + ] + } + } + ).encode() + + source = HraValueSource(fake_get) + branch = { + "acronym": "HRAVS", + "uri": "https://purl.humanatlas.io/vocab/hravs#HRAVS_1000401", + } + self.assertEqual(source.narrower_values(branch), ("Alpha", "Beta")) + self.assertEqual(source.narrower_values(branch), ("Alpha", "Beta")) + self.assertEqual(len(calls), 1) + self.assertIn("lod.humanatlas.io", calls[0][0]) + self.assertEqual(calls[0][1]["Accept"], "application/sparql-results+json") + + def test_hra_rejects_non_hravs_branch(self) -> None: + source = HraValueSource(lambda *args, **kwargs: b"{}") + with self.assertRaisesRegex(GeneratorError, "Only HRAVS"): + source.narrower_values( + {"acronym": "OTHER", "uri": "https://example.test/value"} + ) + + def test_field_types(self) -> None: + template = { + "_ui": {"order": ["count", "doi", "choice", "controlled"]}, + "required": ["doi"], + "properties": { + "count": { + "schema:name": "count", + "schema:description": "A number", + "_ui": {"inputType": "numeric"}, + "_valueConstraints": {"requiredValue": True}, + }, + "doi": { + "schema:name": "doi", + "schema:description": "A link", + "_ui": {"inputType": "link"}, + "_valueConstraints": {}, + }, + "choice": { + "schema:name": "choice", + "schema:description": "Yes or no", + "_valueConstraints": { + "literals": [{"label": "Yes"}, {"label": "No"}] + }, + }, + "controlled": { + "schema:name": "controlled", + "schema:description": "Pick one", + "_valueConstraints": { + "branches": [ + { + "acronym": "HRAVS", + "uri": "https://purl.humanatlas.io/vocab/hravs#HRAVS_1000401", + } + ] + }, + }, + }, + } + fields = fields_from_template(template, lambda branch: ("A", "B")) + self.assertEqual( + [field.kind for field in fields], + ["text", "text", "radio", "allowable"], + ) + self.assertTrue(fields[0].required) + self.assertFalse(fields[1].required) + self.assertFalse(fields[2].required) + self.assertFalse(fields[3].required) + self.assertEqual(fields[2].allowable_values, ("Yes", "No")) + self.assertEqual(fields[3].allowable_values, ("A", "B")) + + def test_cedar_required_value_flag(self) -> None: + self.assertTrue( + cedar_value_is_required( + "parent_sample_id", + {"_valueConstraints": {"requiredValue": True}}, + ) + ) + self.assertFalse( + cedar_value_is_required( + "lab_id", {"_valueConstraints": {"requiredValue": False}} + ) + ) + self.assertFalse(cedar_value_is_required("optional", {})) + self.assertTrue( + cedar_value_is_required( + "legacy", {"_valueConstraints": {"requiredValue": "true"}} + ) + ) + with self.assertRaisesRegex(GeneratorError, "invalid CEDAR"): + cedar_value_is_required( + "broken", {"_valueConstraints": {"requiredValue": 1}} + ) + + def test_page_rendering(self) -> None: + assay = parse_assay(self.issue) + page = render_page( + assay, + [ + Field("count", "A number", True, "text"), + Field( + "choice", + "Pick one", + False, + "allowable", + ("A", "B"), + ), + ], + ) + self.assertIn("# DNA Methylation Metadata Attributes", page) + self.assertIn('count *', page) + self.assertIn(TEXT_ICON, page) + self.assertIn("```A``` ```B```", page) + self.assertNotIn("fa-arrow-up-right-from-square", page) + + def test_index_replacement_is_idempotent(self) -> None: + assay = parse_assay(self.issue) + index = """before +| Dataset Type | Description | +|--------------|-------------| +| Alpha [](alpha "Attribute description") | Alpha | +| Old DNA [](dna-methylation "Attribute description") | Old | +| Zebra [](zebra "Attribute description") | Zebra | +{: .assay-metadata-index } +after +""" + once = update_index(index, assay) + twice = update_index(once, assay) + self.assertEqual(once, twice) + self.assertEqual( + once.count('dna-methylation "Attribute description"'), 1 + ) + self.assertIn("DNA methylation adds a methyl group to DNA.", once) + + suite = unittest.defaultTestLoader.loadTestsFromTestCase(GeneratorTests) + result = unittest.TextTestRunner(verbosity=2).run(suite) + return result.wasSuccessful() + + +def main(argv: list[str] | None = None) -> int: + parser = argparse.ArgumentParser( + description="Generate assay metadata pages and index entries from GitHub issues." + ) + parser.add_argument( + "issues", + nargs="*", + help="Issue numbers from this repo and/or full GitHub issue URLs", + ) + parser.add_argument( + "--dry-run", + action="store_true", + help="Fetch and render everything without writing pages or the index", + ) + parser.add_argument( + "--self-test", + action="store_true", + help="Run the generator's offline test suite", + ) + args = parser.parse_args(argv) + + if args.self_test: + if args.issues: + parser.error("--self-test does not accept issue references") + return 0 if run_self_tests() else 1 + if not args.issues: + parser.error("provide at least one issue number or URL") + + try: + root = repository_root() + repo = detect_repo(root) + value_source = HraValueSource() + issues = [fetch_issue(reference, repo, root) for reference in args.issues] + # Complete every network read before writing so a failed item cannot leave a + # partially generated batch in the working tree. + prepared = [ + (issue, *prepare_issue(issue, value_source)) for issue in issues + ] + hravs_version = value_source.version() + if args.dry_run: + report_results(prepared, hravs_version, prefix="Would generate") + else: + write_results(prepared, root, hravs_version) + except (GeneratorError, json.JSONDecodeError) as exc: + print(f"error: {exc}", file=sys.stderr) + return 1 + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) From 626c56932eade6df0237c44b205b7db17567235e Mon Sep 17 00:00:00 2001 From: Birdmachine Date: Mon, 20 Jul 2026 17:00:36 -0400 Subject: [PATCH 2/2] adds MACSima --- docs/assays/metadata/index.md | 1 + docs/assays/metadata/macsima.md | 40 +++++++++++++++++++++++++++++++++ 2 files changed, 41 insertions(+) create mode 100644 docs/assays/metadata/macsima.md diff --git a/docs/assays/metadata/index.md b/docs/assays/metadata/index.md index 30ffd3a3..3f4c3077 100644 --- a/docs/assays/metadata/index.md +++ b/docs/assays/metadata/index.md @@ -25,6 +25,7 @@ A list of available dataset types (data types from multiple supported assays), w | [IMC](https://docs.hubmapconsortium.org/assays/imc) [](imc "Attribute description")| Combines standard immunohistochemistry with CyTOF mass cytometry to resolve the cellular localization of up to 40 proteins in a tissue sample. Link to [IMC directory schema](https://hubmapconsortium.github.io/ingest-validation-tools/imc-2d/current/). | | [LC-MS](https://docs.hubmapconsortium.org/assays/lcms) [](LC-MS "Attribute description")| Coupling of liquid chromatography (LC) to mass spectrometry (MS). Link to [LC-MS directory schema](https://hubmapconsortium.github.io/ingest-validation-tools/lcms/current/). | | [Light Sheet](https://en.wikipedia.org/wiki/Light_sheet_fluorescence_microscopy) [](Light-Sheet "Attribute description")| A fluorescence imaging technique that uses a thin sheet of laser light to illuminate a sample, allowing for high-resolution, 3D imaging with reduced photobleaching and phototoxicity; particularly useful for imaging large, thick, or delicate biological samples, like developing embryos or organoids. Link to [Light Sheet directory schema](https://hubmapconsortium.github.io/ingest-validation-tools/lightsheet/current/). | +| [MACSima](https://hubmapconsortium.github.io/ingest-validation-tools/macsima/current/) [](macsima "Attribute description") | MACSima is a fully automated, high-content spatial biology system designed for ultra-high-plex, cyclic immunofluorescence imaging. It enables researchers to map hundreds of protein markers, and some RNAs, on a single tissue sample using [MICS technology (MACSima Imaging Cyclic Staining)](https://www.nature.com/articles/s41598-022-05841-4), combining deep phenotyping with spatial context. | | [MALDI-IMS](https://docs.hubmapconsortium.org/assays/maldi-ims) [](MALDI "Attribute description") | Matrix-assisted laser desorption/ionization (MALDI) imaging mass spectrometry (IMS) combines the sensitivity and molecular specificity of MS with the spatial fidelity of classical microscopy. Link to [MALDI-IMS directory schema](https://hubmapconsortium.github.io/ingest-validation-tools/maldi/current/). | | [MIBI](https://www.researchgate.net/figure/Multiplexed-ion-beam-imaging-workflow-for-high-resolution-spatial-proteomics-Here_fig1_349770840) [](MIBI "Attribute description") | Preserved tissue sections, mounted on conductive substrates are incubated with unique isotopic transition metal-tagged antibody reporters. An oxygen primary ion beam rasters the sample surface, ejecting and ionizing the isotope reporters. Their masses are subsequently measured via a mass analyzer. Link to [MIBI directory schema](https://hubmapconsortium.github.io/ingest-validation-tools/mibi/current/). | | [MERFISH](https://pubmed.ncbi.nlm.nih.gov/27241748/) [](merfish "Attribute description") | A spatial transcriptomics technology that allows for the simultaneous imaging of hundreds to thousands of RNA species within single cells, providing both copy number and spatial distribution information. Link to [MERFISH directory schema](https://hubmapconsortium.github.io/ingest-validation-tools/merfish/current/). | diff --git a/docs/assays/metadata/macsima.md b/docs/assays/metadata/macsima.md new file mode 100644 index 00000000..8ef8df93 --- /dev/null +++ b/docs/assays/metadata/macsima.md @@ -0,0 +1,40 @@ +--- +layout: page-triary +--- + +# MACSima Metadata Attributes + +Fields that are collected for MACSima data, available at ```dataset.metadata.``` +  + +* indicates a required field + +| Attribute | Type | Description | Allowable Values | +|------|------|-------------|-------------------| +| parent_sample_id * | | The unique identifier from HuBMAP or SenNet for the sample (such as a block, section, or suspension) used to perform the assay. For instance, in an RNAseq assay, the parent sample would be the suspension, while in imaging assays, it would be the tissue section. If the assay is derived from multiple parent samples, this field should contain a comma-separated list of identifiers. Example: HBM386.ZGKG.235, HBM672.MKPK.442 | | +| lab_id | | A locally assigned identifier provided by the data provider for the dataset. It is used to reference an external metadata record that may be maintained independently, enabling traceability and supporting provenance tracking. Example: Visium_9OLC_A4_S1 | | +| preparation_protocol_doi * | | The DOI for the protocols.io page that details the assay or the procedures used for sample procurement and preparation. For example, in the case of an imaging assay, the protocol may start with tissue section staining and end with the generation of an OME-TIFF file. The documented protocol should also include any image processing steps involved in producing the final OME-TIFF. Example: https://dx.doi.org/10.17504/protocols.io.eq2lyno9qvx9/v1 | | +| dataset_type * | | The specific type of dataset being produced. Example: RNAseq | ```10X Multiome``` ```2D Imaging Mass Cytometry``` ```4i``` ```ATACseq``` ```Auto-fluorescence``` ```Cell DIVE``` ```CODEX``` ```COMET``` ```Confocal``` ```CosMx Proteomics``` ```CosMx Transcriptomics``` ```CyCIF``` ```CyTOF``` ```DART-FISH``` ```DBiT-seq``` ```DESI``` ```DNA Methylation``` ```Enhanced Stimulated Raman Spectroscopy (SRS)``` ```FACS``` ```GeoMx (nCounter)``` ```GeoMx (NGS)``` ```HiFi-Slide``` ```Histology``` ```iCLAP``` ```Illumina Spatial ver0``` ```LC-MS``` ```Light Sheet``` ```MACSima``` ```MALDI``` ```MERFISH``` ```MIBI``` ```Molecular Cartography``` ```MPLEx``` ```MS Lipidomics``` ```MUSIC``` ```nanoSPLITS``` ```Olink``` ```PhenoCycler``` ```Pixel-seqV2``` ```Raman Imaging``` ```Resolve``` ```RNAseq``` ```RNAseq (with probes)``` ```Second Harmonic Generation (SHG)``` ```Seq-Scope``` ```seqFISH``` ```SIMS``` ```Singular Genomics G4X``` ```SNARE-seq2``` ```STARmap``` ```Stereo-seq``` ```Thick section Multiphoton MxIF``` ```Virtual Histology``` ```Visium (no probes)``` ```Visium (with probes)``` ```Visium HD``` ```Xenium``` | +| analyte_class * | | The analyte class which is the target molecule that the assay is measuring. Example: DNA | ```Chromatin``` ```Collagen``` ```DNA``` ```DNA + RNA``` ```Endogenous fluorophore``` ```Fluorochrome``` ```Lipid``` ```Lipid + metabolite``` ```Lipid + metabolite + protein``` ```Metabolite``` ```Nucleic acid + protein``` ```Peptide``` ```Polysaccharide``` ```Protein``` ```RNA``` ```RNA + protein``` ```Saturated lipid``` ```Unsaturated lipid``` | +| is_targeted * | | Indicates whether a specific molecule or set of molecules is targeted for detection or measurement by the assay. Example: Yes | ```Yes``` ```No``` | +| acquisition_instrument_vendor * | | The company that manufactures or supplies the acquisition instrument. An acquisition instrument is a device equipped with signal detection hardware and signal processing software. It captures signals produced by assays, such as variations in light intensity or color, or signals corresponding to molecular mass. If the instrument was custom-built or developed internally, enter "In-House". Example: Illumina | ```10x Genomics``` ```3DHISTECH``` ```Akoya Biosciences``` ```Andor``` ```BGI Genomics``` ```Bruker``` ```Complete Genomics``` ```Cytek Biosciences``` ```Cytiva``` ```Element Biosciences``` ```Evident Scientific (Olympus)``` ```GE Healthcare``` ```Hamamatsu``` ```Huron Digital Pathology``` ```Illumina``` ```In-House``` ```Ionpath``` ```Keyence``` ```Leica Biosystems``` ```Leica Microsystems``` ```Microscopes International``` ```Miltenyi Biotec``` ```Motic``` ```NanoString``` ```Resolve Biosciences``` ```Revvity``` ```Sciex``` ```Singular Genomics``` ```Standard BioTools (Fluidigm)``` ```Thermo Fisher Scientific``` ```Vizgen``` ```Waters``` ```Zeiss Microscopy``` | +| acquisition_instrument_model * | | The specific model of the acquisition instrument, as manufacturers often offer various versions with differing features or sensitivities. These differences may be relevant to the processing or interpretation of the data. If the instrument was custom-built or developed internally, enter "In-House". If the model is unknown, enter "Unknown". Example: HiSeq 4000 | ```Aperio AT2``` ```Aperio CS2``` ```AVITI``` ```Axio Observer 3``` ```Axio Observer 5``` ```Axio Observer 7``` ```Axio Scan.Z1``` ```Axio Zoom.V16``` ```Biomark HD``` ```BZ-X710``` ```BZ-X800``` ```BZ-X810``` ```Cell DIVE``` ```CosMx Spatial Molecular Imager``` ```Custom: Multiphoton``` ```Cytek Northern Lights``` ```CyTOF 2``` ```CyTOF XT``` ```Digital Spatial Profiler``` ```DM6 B``` ```DMi8``` ```DNBSEQ-T7``` ```EVOS M7000``` ```G4X Spatial Sequencer``` ```Helios``` ```HiSeq 2500``` ```HiSeq 4000``` ```Hyperion Imaging System``` ```IN Cell Analyzer 2200``` ```In-House``` ```Juno System``` ```Lightsheet 7``` ```LSM 710 Confocal Microscope``` ```MACSima System``` ```MALDI timsTOF Flex Prototype``` ```MERSCOPE``` ```MERSCOPE Ultra``` ```MIBIscope``` ```MoticEasyScan One``` ```NanoZoomer 2.0-HT``` ```NanoZoomer 2.0-RS``` ```NanoZoomer S210``` ```NanoZoomer S360``` ```NanoZoomer S60``` ```NanoZoomer-SQ``` ```NextSeq 2000``` ```NextSeq 500``` ```NextSeq 550``` ```Not applicable``` ```NovaSeq 6000``` ```NovaSeq X``` ```NovaSeq X Plus``` ```Opera Phenix HCS``` ```Opera Phenix Plus HCS``` ```Orbitrap Eclipse Tribrid``` ```Orbitrap Fusion Lumos Tribrid``` ```Orbitrap Fusion Tribrid``` ```Pannoramic MIDI II Digital Scanner``` ```Panoramic 150 Digital Scanner``` ```Phenocycler-Fusion 1.0``` ```Phenocycler-Fusion 2.0``` ```PhenoImager Fusion``` ```Q Exactive``` ```Q Exactive HF``` ```Q Exactive HF-X``` ```Q Exactive UHMR``` ```QTRAP 5500``` ```Resolve Biosciences Molecular Cartography``` ```SCN400``` ```solariX``` ```STELLARIS 5``` ```SYNAPT G2-Si``` ```timsTOF FleX``` ```timsTOF FleX MALDI-2``` ```timsTOF HT``` ```timsTOF Pro``` ```timsTOF Pro 2``` ```timsTOF SCP``` ```timsTOF Ultra``` ```timsTOF Ultra 2``` ```TissueScope LE Slide Scanner``` ```Unknown``` ```uScopeHXII-20``` ```VS200 Slide Scanner``` ```Xenium Analyzer``` ```Zeiss LightSheet Z.1``` ```Zyla 4.2 sCMOS``` | +| source_storage_duration_value * | | The length of time the sample was stored prior to processing it. For assays performed on tissue sections, this refers to how long the tissue section (e.g., slide) was stored before the assay began (e.g., imaging). For assays performed on suspensions, such as sequencing, it refers to how long the suspension was stored before library construction started. Example: 12 | | +| source_storage_duration_unit * | | The unit of measurement used to specify the source storage duration value. Example: hour | ```day``` ```hour``` ```minute``` ```month``` ```year``` | +| time_since_acquisition_instrument_calibration_value | | The length of time since the acquisition instrument was last serviced or calibrated. This provides a metric for assessing drift in data capture. Example: 10 | | +| time_since_acquisition_instrument_calibration_unit | | The unit of measurement used to specify the time since acquisition instrument calibration value. Example: month | ```day``` ```month``` ```year``` | +| contributors_path * | | The name of the file containing the ORCID IDs for all contributors to this dataset. Example: ./contributors.csv | | +| data_path * | | The top-level directory containing the raw and/or processed data. For a single dataset upload, this might be represented as ".", whereas for a data upload containing multiple datasets, this would be the directory name for the respective dataset. For example, if the data is within a directory named "TEST001-RK", use the syntax "./TEST001-RK" for this field. If there are multiple directory levels, use the format "./TEST001-RK/Run1/Pass2", where "Pass2" is the subdirectory where the single dataset's data is stored. This is an internal metadata field used solely for data ingestion. Example: ./TEST001-RK | | +| antibodies_path * | | The path to the antibodies.tsv file relative to the root directory of the upload structure. This path should start with "." and is typically formatted as "./extras/antibodies.tsv". Example: ./extras/antibodies.tsv | | +| total_run_time_value * | | The total run time, which is the duration the instrument takes to fully complete all imaging rounds on the loaded slide. Example: 24 | | +| total_run_time_unit * | | The unit of measurement for the total run time value. If the total run time is not specified, this field may be left blank. Example: hour | ```hour``` ```minute``` | +| number_of_antibodies * | | The number of antibodies used in the assay. If no antibodies were utilized, enter 0. Example: 5 | | +| number_of_channels * | | The number of fluorescent channels that are imaged during each cycle. Example: 3 | | +| number_of_biomarker_imaging_rounds * | | The number of imaging rounds required to capture the tagged biomarkers. For CODEX, a biomarker imaging round includes steps such as (1) oligo application, (2) fluor application, and (3) washes. For Cell DIVE, it involves (1) the staining of a biomarker via secondary detection or direct conjugate, followed by (2) dye inactivation. Example: 3 | | +| number_of_total_imaging_rounds * | | The total number of imaging rounds performed using a microscope to collect either autofluorescence/background or stained signals, such as those used in histological analysis. Example: 5 | | +| slide_id | | The unique identifier assigned to each slide, enabling users to determine which tissue sections were processed together on the same slide. It is recommended that data providers prefix the ID with the center name to prevent overlapping values across different centers. Example: VAN0071-PA-1-1_AF | | +| cell_boundary_marker_or_stain * | | The name of the marker or stain used to identify all cell boundaries in the tissue. This name must exactly match the antibody-targeted molecule marker or non-antibody targeted molecule stain as found in the imaging data. For example, in the case of using the PhenoCycler, ensure the name corresponds to the value in the XPD output file. If multiple markers or stains are employed, list them in a comma-separated format. Example: Pan-Cytokeratin, E-Cadherin | | +| nuclear_marker_or_stain * | | The nuclear marker or stain used, which can be an antibody-targeted molecule present in or around the cell nucleus. For protein targets, use the protein or gene symbol that identifies the antibody target, ensuring it matches the antibody target from the panel used or custom panels. Preferably, if using a custom antibody marker, this symbol should be the HGNC symbol (https://www.genenames.org/). For non-protein targets, provide the stain name (e.g., DAPI) and, when applicable, include the associated staining kit and vendor. For the PhenoCycler, ensure the symbol matches the value found in the XPD output file. Example: DAPI | | +| non_global_files | | Specifies a semicolon-separated list of non-global files that are to be included in the dataset. The file paths assume that the files are located in the "TOP/non-global/" directory. For instance, if the file is located at TOP/non-global/lab_processed/images/1-tissue-boundary.geojson, the value for this field would be "./lab_processed/images/1-tissue-boundary.geojson". Once ingested, these files will be copied to their appropriate locations within the respective dataset directory tree. This field is intended for internal HuBMAP processing. Examples for GeoMx and PhenoCycler are provided in the File Locations documentation: https://docs.google.com/document/d/1n2McSs9geA9Eli4QWQaB3c9R3wo5d5U1Xd57DWQfN5Q/edit#heading=h.1u82i4axggee Example: ./lab_processed/images/1-tissue-boundary.geojson | | +| signal_quenching_method * | | An automated step in cyclic imaging workflows that removes fluorescent signals from a sample to enable subsequent rounds of staining. This process is critical for high-multiplex imaging, allowing the detection of numerous markers on a single sample while preserving tissue integrity, as demonstrated by technologies such as MACSima Imaging Cyclic Staining (MICS). Example: Photobleaching | ```Enzymatic cleavage``` ```Photobleaching``` | +| metadata_schema_id * | | The unique string identifier for the metadata specification version, which is easily interpretable by computers for purposes of data validation and processing. Example: 22bc762a-5020-419d-b170-24253ed9e8d9 | |