2.22.1
2.22.1 (2026-10-01)
Fix
-
fix: count single-colour images and flag unanswerable queries (#5347)
-
feat: count single-colour images and flag unanswerable queries
ImageStatistics gains constant_images -- images that are one colour
everywhere. Such an image carries no visual signal; it is usually a failed
fetch silently replaced by a blank frame, which makes it missing data rather
than noise.
Auditing the image tasks available offline (2.2M images) found that what a
constant image costs depends entirely on which side it sits on:
- VisualNewsI2TRetrieval: 14 solid-colour query images, all in qrels. The
task is i2t, so the query is the image alone and those queries cannot be
answered. - WITT2IRetrieval: 21 constant corpus images, each the sole gold document
for its caption (the loader maps one image per caption), so 21 queries
have nothing retrievable. - VisualNewsT2IRetrieval: 174 constant corpus images, 14 of them a query's
only gold -- the same underlying files as the 14 above, on the other side
of the task pair. - WebQA (187), OVEN (420) and RP2k (1): constant documents no qrel
references, or one blank among 1069 positives. Inert; not flagged.
So the check is scoped by evaluation semantics rather than by column name.
A query or labelled sample is itself the evaluated unit, so a constant one
is always a defect. A document is only unretrievable when it is the whole
of some query's gold set, which qrels decide -- reported separately as
RelevantDocsStatistics.queries_with_all_gold_constant. Corpus columns are
no longer skipped; they are judged by the rule that fits them.
Counting is close to free: calculate_image_statistics already walks every
image for width and height, and the corpus flags are computed once and
shared between the image statistics and the qrels intersection.
Stats generated before these fields exist simply omit them and are skipped,
so both activate per task as descriptive statistics are regenerated. No
KNOWN_ISSUES entries are added.
- fix: narrow constant-image quality check to pure black/white images
Flagging any single-colour image treated template/placeholder art (a
solid red or brand-colour background) the same as a failed image fetch
silently replaced by a blank frame. Only pure black/white reliably
indicates the latter, so narrow detection to that and rename the
helpers/fields/messages so they no longer imply otherwise.
Two fixes were needed for the narrowed check to be correct rather than
just stricter: palette-mode images store indices, not colours, so an
index of 0 isn't black until resolved through the palette; and alpha
must be constant but excluded from the colour check itself, since an
ordinary opaque black/white pixel has alpha=255, not 0.
Qrels-aware semantics are unchanged: a query is only flagged when its
entire gold set is black/white, and corpus images no qrel references
stay unflagged.
-
fix: simplify black and white image checks following review
-
fix: convert to RGB for the black/white check and trim its test
Drop the per-mode handling (palette, CMYK, HSV, wide-integer white points)
in favour of a single RGB conversion, reduce the test to one small case,
and revert two formatting-only hunks.
- test: use the same relevance score for both gold documents
Only the sign of the score matters to the count, and the mixed 1/2 pair
read as if it meant something.
- test: report black/white image checks as warnings
A black or white image can be valid content, so both checks are
diagnostic heuristics rather than correctness invariants. Route them
through _WARNING_CHECK_KINDS, drop the failure language from the
docstrings, and name the gold-set statistic as the stronger signal for
manual review. The unit test now names its corpus ids instead of using
list indices.
- fix: match black/white doc ids to prefixed qrels in multilingual stats
The overall split of a multilingual retrieval task prefixes corpus and
qrel ids with split and subset, but the black/white doc id set was built
from the raw concatenated corpus ids, so
queries_with_all_gold_black_or_white was always 0 at the overall level.
Regenerate descriptive stats for VisualNewsI2TRetrieval (14 black/white
query images) and WITT2IRetrieval (21 black/white corpus images, 21
queries whose only gold document is black/white), matching the offline
audit.
- test: register WITT2IRetrieval duplicate images and captions as known
Regenerating its stats added num_documents and num_queries, which lets
the duplicate checks run on this task for the first time. The duplicates
are a property of WIT and are already registered for WITI2TRetrieval.
- optimize computation
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (78aa9a6)