If I have two record identifiers that share a common prefix, pipestat will return the value for the shorter one even when you request the longer one.
This may be the root of this problem I found earlier in looper: pepkit/looper#470
Watch this:
Create pipestatmanager object
psm3 = pipestat.PipestatManager(
record_identifier="sample1",
results_file_path="analysis/bug_test.yaml",
schema_path="analysis/seqcol_pipestat_schema.yaml",
)
Here's my seqcol_pipestat_schema.yaml
title: Refget henge back-end schema
description: Allows pipestat to be used as a simple key-value store.
type: object
properties:
pipeline_name: "refget"
samples:
type: object
properties:
value:
type: string
description: "Value of the object referred to by the key"
Report/retrieve results is broken if you use similar sample names:
psm3.report({"value": "abcdefg"}, record_identifier="sample1")
psm3.report({"value": "12345"}, record_identifier="sample")
psm3.retrieve_one("sample", "value")
# '12345'
psm3.retrieve_one("sample1", "value") # should return 'abcdefg'
# '12345'
And continuing:
psm3.report({"value": "This is a new value"}, record_identifier="sample1", force_overwrite=True)
psm3.retrieve_one("sample1", "value")
# '12345'
It seems that if anything matches the first prefix of the sample, it will return that somehow. This is a critical bug.
The problem is with retrieval, not with reporting, because the file itself is actually correct, it shows:
refget:
project: {}
sample:
sample1:
value: This is a new value
pipestat_created_time: '2024-02-21 07:49:04'
pipestat_modified_time: '2024-02-21 07:50:52'
sample:
value: '12345'
pipestat_created_time: '2024-02-21 07:49:13'
pipestat_modified_time: '2024-02-21 07:49:13'
If I have two record identifiers that share a common prefix, pipestat will return the value for the shorter one even when you request the longer one.
This may be the root of this problem I found earlier in looper: pepkit/looper#470
Watch this:
Create pipestatmanager object
Here's my
seqcol_pipestat_schema.yamlReport/retrieve results is broken if you use similar sample names:
And continuing:
It seems that if anything matches the first prefix of the sample, it will return that somehow. This is a critical bug.
The problem is with retrieval, not with reporting, because the file itself is actually correct, it shows: