Releases: mcp-tool-shop-org/plain-sight
Release list
v1.1.0 — the dataset lane can no longer mislabel silently
1.1.0 — 2026-08-20
Dogfood swarm: health pass (bugs, proactive, humanization) plus a pre-Phase-9
wave. 39 tests → 88. Two CRITICALs closed, both of the same shape — output a
downstream consumer would trust, wrong, with no signal.
Fixed — the dataset lane could silently mislabel training data
- Sidecar collisions are refused, not merged. Sidecar paths derived from the
image stem alone, soimg.pngandimg.jpgin one folder both claimed
img.txt. The second was reported asskip (exists)and the run exited 0 — a
green run in which one image carried a description of a different image. It
needed no flags. Colliding batches are now refused before the model loads,
naming the offenders. Sidecars are never renamed to dodge a clash: trainers
pair by exact stem, so a rename would orphan the caption. - OCR can no longer present invented text as extracted text. Florence-2
emits a decoded string for every image, including images with no text — a
photograph returns'2', lexically indistinguishable from a correct reading.
Results now always carryabsence_of_text_unreliable(MCP) or an
[OCR_CAVEAT]line (CLI). The text is never suppressed, emptied, or
length-thresholded, because a short reading may be genuine. - Sidecar writes are atomic — temp file plus
os.replace, so an interrupt
cannot leave a partial caption at the final path. An existing but empty
sidecar is treated as unfinished and re-captioned. - A mid-batch failure no longer aborts the run. The batch loop caught only
FileNotFoundError/ValueError, so a CUDA OOM killed the job and swallowed
the JSON summary. It now records the failure and continues; exit 3 on partial.
Fixed — broken promises
PLAIN_SIGHT_MODEL_IDwas documented in nine READMEs and honoured by the MCP
server, while the CLI silently ignored it. Resolved once, at module scope.PLAIN_SIGHT_LOG_LEVELwas named in the CLI's own error hint and had no
effect there — DEBUG was unreachable. Sharedconfigure_logging()now serves
both surfaces.PLAIN_SIGHT_EAGER_LOADis honoured on the MCP surface again, and a failure
during eager load no longer kills the server import: it surfaces via
sight_statusand as aToolErroron first use.- The CLI emitted a raw traceback at exit 1 when engine construction failed —
the construction sat outside the error boundary. Now a structured line, exit 2.
Added — provenance
- The model revision is pinned to
4271c66b88cdbc05735372ec13b2360108de5317.
Unpinned, HuggingFace resolves to whatever the default branch points at, so a
silent retag would change captions under unchanged inputs.
PLAIN_SIGHT_MODEL_REVISIONoverrides. - Every output payload names the weights —
model_idand the resolved
revision ondescribe_image,read_text,describe_batch,sight_selftest,
the CLI--jsonmodes and the batch summary.sight_statusreports requested
and resolved separately, so a mismatch is visible. --manifest PATHwrites an opt-in run record: versions, model, both
revisions, device, dtype, tier, prefix/suffix, per-image results. Never
inferred; a path colliding with a sidecar is refused.
Added — it now says what it is doing
- The load is announced before work begins, at default verbosity, with the
count of images that will actually be captioned — so the pause never appears
mid-run after a stretch of skips. - Progress heartbeat every 25 items or 30s: written / skipped / failed,
rate, ETA. Per-skip lines are gone; a re-run over a finished set is quiet.
Failures stay one line each. --dry-runprints the whole plan — model, revision, counts, collisions —
loading nothing and writing nothing.--helpcarries the exit-code table, the stderr/stdout split, and the
first-load cost. Every flag on every subcommand has help text.
Changed
- CI and
verify.shselect tests by marker (-m "not dogfood") rather than by
filename, so a new CI-safe test file is covered without touching CI. --helpand CLI error output are ASCII; the em-dash separator mojibaked when
stderr was a cp1252 pipe on Windows.pythonpath = ["."]so the console script andpython -m pytestagree.status()reports device-scoped VRAM, omitted on CPU.
Notes
Severities throughout this release were assigned by a cross-family panel of
pinned non-Claude model seats with authorship stripped, not by the authors of
the findings. Several moved in both directions.
plain-sight v1.0.0
An AI says what it sees. First shipped release: Florence-2 describer (MCP server + CLI), three detail tiers, OCR, LoRA-dataset caption sidecars with exact basename pairing and deterministic decoding. Model pinned to florence-community/Florence-2-large (MIT, native transformers — trust_remote_code never used). Shipcheck audit 100%; landing page + handbook: https://mcp-tool-shop-org.github.io/plain-sight/ — full details in CHANGELOG.md.