Dedup dataset metadata by dataset_id - #48
Merged
Merged
Conversation
The metadata channel kept one entry per dataset by filtering on rna_norm == "log_cp10k". That says "one per dataset" only as long as log_cp10k is one of the normalizations we run; otherwise the channel is empty and the output join has nothing to merge. Dedup on dataset_id, which is what was meant.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Describe your changes
The metadata channel wants one entry per dataset, but expresses that as "keep the log_cp10k ones":
That holds only while
log_cp10kis among the normalizations we materialise. The day a dataset arrives under a different one -- orlog_cp10kgets renamed -- the filter passes nothing,joinStatesnever fires, andoutput_dataset_info/output_method_configs/output_metric_configs/output_task_infoare simply absent from the final state. Silently, since an empty channel isn't an error.De-duplicating on
dataset_idinstead says the intended thing directly and doesn't care how many normalizations there are. Checked the Groovy under nextflow:(three states, two datasets,
normalization_idstripped as before.)No behavioural change today, since
log_cp10kis the only normalization in play.Part of a series of PRs coming out of a pre-run review of the benchmark.
Checklist before requesting a review
I have performed a self-review of my code
Check the correct box. Does this PR contain:
Proposed changes are described in the CHANGELOG.md
CI Tests succeed and look good!