Releases: eko/qc
Releases · eko/qc
Release list
v1.0.0
First public release: a Go library and the qc CLI.
Added
- Technical analysis (
qc analyze): container and stream info, bitrate
over time with peaks, GOP structure and HDR metadata read from the
bitstream without decoding (--fast, under a second). Then one decode
fanned out to SI/TI (ITU-T P.910), shot detection, black and frozen
segments, letterbox/pillarbox detection and luma levels. - Audio quality control, alongside the frame analysis and without
adding to its wall time: every audio track (--audio-tracks) decoded by
ffmpeg to float samples and measured in pure Go (packagesaudio,
audio/loudness,audio/defect): integrated loudness, loudness range,
true peak (4× oversampling) and momentary/short-term series per ITU-R
BS.1770-5 and EBU Tech 3341/3342, checked against--loudness-target
(ebu,ebu-live,atsc,streaming,streaming-14or a LUFS value);
silence of the mix and of each channel (--silence-threshold,
--silence-duration), leading and trailing silence, muted channels,
empty LFE or centre, clipping, DC offset, out-of-phase segments,
inverted polarity, mono as stereo, sample rate, bit depth, layout and
stream start offsets. Terminal, HTML (loudness, levels and phase charts
per track) and JSON reports, and aloudnessitem for annotated videos;
--no-audioleaves it out,--fast --audioadds it to an inspection.
Validated against the synthetic EBU conformance cases (all within
tolerance), ffmpeg'sebur128on real content (within 0.01 LU / dB) and
synthetic defects (bench/audioval, docs/audio.md). - Camera motion (
analyze/motion, on by default,--no-motionto skip):
the global motion of every frame estimated on the shared thumbnails
(integral-projection predictor, block matching with a Lucas–Kanade
sub-pixel step, robust similarity fit), and each shot's camera work:
static, pan, tilt, zoom, tracking (parallax), handheld, mixed, with its
direction and a shake measure. JSON columns (motionPan,motionTilt,
motionZoom,motionRoll,motionShake,motionConfidence),
video.motionandshots[].camera; a camera section and a shot column in
the terminal report; a camera motion chart, card and shot column in the
HTML report; a note on shaky shots. About 0.17 ms of one core per frame;
validated on synthetic moves of known speed (bench/motionval). - Annotated videos (
--overlay annotated.mp4onanalyze,vmafand
run, and a question of the wizard): a copy of the video with the
analysis burnt in as a debug overlay: timecode, frame number, size and
keyframes, bitrate, shot and cut markers, camera work with a motion
vector, SI/TI, luma levels, HDR light levels, black/frozen/banded/
out-of-range badges, the VMAF and other metrics of each scored frame
(--exactfor every frame), and a timeline with a playhead
(--overlay-itemsto choose,--overlay-heightfor a smaller, faster
copy). Packageoverlaywrites an ASS script, frame-accurate at any
frame rate below 100 fps, that libass draws in the H.264 encode of the
copy (encode.FFmpeg.Burn); same frames, timestamps and audio as the
source. The copy is encoded by the media engine when there is one
(--overlay-encoder auto: VideoToolbox on macOS after a test encode,
NVENC with--gpu, x264 otherwise), in segments split at keyframes and
rendered concurrently (--overlay-workers), each with its slice of the
script, joined without re-encoding (a segment that does not render the
frames planned falls back to one pass): 59 minutes of 1080p25 annotated
in 217 s on an M2 Max (778 s with x264, 6371 s of CPU against 387 s).
ffmpeg's libass, and an explicit hardware encoder, are checked before any
work;qc version --checklists VideoToolbox; the Docker images ship
libass with DejaVu Sans Mono. - Fast frame analysis: on macOS the video is split at keyframes into
segments decoded concurrently by VideoToolbox (--hwaccel auto, the
default;videotoolboxandnonealso accepted), each fed to forks of
every analyzer (analyze.Forker,analyze.RunSegments) whose per-frame
series are merged in order: the report is identical to a single pass, to
the bit, and a segment that does not start where planned falls back to
one. Segments seek on the container's timeline, so a video starting after
its audio (some concatenations) no longer falls back. NEON loops on arm64 for SI, TI, luma statistics and thumbnails
(checked against the portable Go ones), and Unix sockets instead of pipes
for the raw frames. 59 minutes of 1080p25 H.264 analysed in 62–71 s at
26 Mbit/s (225 s before) and 49–52 s at 6 Mbit/s (230 s) on an M2 Max; the
single CPU pass (other systems,--hwaccel none) is 1.1× to 2.7× faster
too, depending on the share of the decode. See
analysis.md. - VMAF with a confidence interval (
qc vmaf): short clips sampled
across shots (stratified, two-stage) until the 95% interval is narrower
than--precision, with a real coverage validated by replaying thousands
of runs (bench/vmafsim). Fixed budgets with--sample(5%of the
frames,2/scene), and--exactto score every frame, bit-exact with
Netflix'svmaftool at 8 and 10 bits. VMAF v1 models (v1.0.16) picked
automatically, 4K and high frame rate aware. - Metrics beyond VMAF, measured on the same frames with their own
confidence intervals (--metrics): XPSNR (a pure-Go port matching
ffmpeg's filter, also usable on its own asquality/xpsnr), CAMBI banding
with banded segments, PSNR, PSNR-HVS, SSIM, MS-SSIM, CIEDE2000, and the
whole AOM AV2 common test conditions set with--av2-ctc. - VMAF per viewing device (
--devices phone,tv,4k), each scored with its
VMAF v1 model on the same frames. - Faster VMAF on macOS: both videos are decoded by concurrent
VideoToolbox sessions (--hwaccel auto, the default), sweeps are split
into concurrent runs, and exact measurements are scored in segments split
at keyframes, three libvmaf contexts at once, with two warm-up frames
before each so that every frame scores as in a single pass. XPSNR has
NEON loops (3.8× faster at 1080p). Results are identical to a CPU run,
frame by frame. On an M2 Max: 26.5 → 18.6 s for the ±0.5 measurement of
a 10:36 title, 58.7 → 42.1 s for a 59-minute one (184 → 108 s at
--sample 5%), with a third of the CPU; exact measurements, bound by
libvmaf's CPU, 149 → 131 s on the 10:36 title and 897 → 704 s on the
59-minute one. See vmaf.md. - Per-title ladders (
qc ladder) for H.264 (libx264), HEVC (libx265)
and AV1 (SVT-AV1): probe encodes of a representative digest, rate-quality
curves per resolution, their upper envelope, rungs one just-noticeable
difference apart, then a verification encode of every rung. Lands on the
exhaustive optimum (−0.04 VMAF, −0.7% bitrate on average) in a fraction
of the time (bench/ladderval). The curve of the resolution below the top
rung's is extended when it might reach the top quality at least 10%
cheaper, so easy content gets a cheaper lower-resolution top rung.- Imposed shapes: a rung count or the rung resolutions (
--rungs 1080,720,540,360), with the bitrates computed; top and minimum VMAF,
bitrate bounds, 8 or 10-bit encodes (--encode-bit-depth). - Adaptive probing (
--probing adaptive): extra probes where the rungs
are uncertain. - Probes at a faster preset (
--probe-preset): the rungs are planned on
them, then the top and bottom rungs encoded at--presetanchor the
probes onto it (bitrate ratio and CRF offset between the presets); for
slow delivery presets (SVT-AV1 preset 4 probed at 8: −24% time, same
rungs). - Per-shot rungs (
--per-shot): one CRF per shot at an equal
rate-quality slope, with the per-shot ladder view (every shot's CRF,
bitrate and VMAF per rung, and each rung's pooled bitrate over the
title) in the terminal, HTML and JSON reports. The chunks of an encode
and the per-shot verifications run concurrently, and the source's
shots come from the analysisqc runalready made, or from one run
alongside the probes (per-shot stage −20% on the drama). Experimental per-shot
resolution (--per-shot-resolution, implies--per-shot): shots also
pick their resolution among neighbouring rung resolutions, at equal
slope. - AV1 film grain synthesis (
--film-grain autoor a level), detected and
calibrated (the three calibration levels encoded concurrently),
fidelity scored against a denoised reference. It cannot be
combined with per-shot rungs: the combination is rejected before any
work. - Verified rungs checked for banding and for VMAF/XPSNR ranking
disagreements. - Renditions (
--encode-ladder <folder>, also asked by the wizard): the
ladder encoded on the whole title with the settings of each rung, and
each rung's per-shot version, with progress in the dashboard. Each
rendition is checked against the source (--no-rendition-checkskips
it): its VMAF with its confidence interval and its bitrate, shown next
to the predictions in the reports, with findings when they drift from
them (quality beyond the check's interval, bitrate beyond Apple's 10%
BANDWIDTHrule).
- Imposed shapes: a rung count or the rung resolutions (
- HDR (docs/hdr.md): HDR10, PQ and HLG detected from the
stream and its first frame (HDR10 metadata in SEI, HDR10+), Dolby Vision
and HDR10+ reported without processing their dynamic metadata; signalling
checks (BT.2020 primaries and matrix, 10 bits, narrow range, missing
mastering display or content light level). MaxCLL and MaxFALL measured
during the frame analysis (CTA-861.3, on a grid of 10-bit samples that
ffmpeg writes to a second pipe of the same decode, robust to 4:2:0 chroma
overshoots) and compared with the signalled
v...