Skip to content

Releases: eko/qc

Release list

v1.0.0

Choose a tag to compare

@github-actions github-actions released this 30 Sep 21:01
1d0183c

First public release: a Go library and the qc CLI.

Added

  • Technical analysis (qc analyze): container and stream info, bitrate
    over time with peaks, GOP structure and HDR metadata read from the
    bitstream without decoding (--fast, under a second). Then one decode
    fanned out to SI/TI (ITU-T P.910), shot detection, black and frozen
    segments, letterbox/pillarbox detection and luma levels.
  • Audio quality control, alongside the frame analysis and without
    adding to its wall time: every audio track (--audio-tracks) decoded by
    ffmpeg to float samples and measured in pure Go (packages audio,
    audio/loudness, audio/defect): integrated loudness, loudness range,
    true peak (4× oversampling) and momentary/short-term series per ITU-R
    BS.1770-5 and EBU Tech 3341/3342, checked against --loudness-target
    (ebu, ebu-live, atsc, streaming, streaming-14 or a LUFS value);
    silence of the mix and of each channel (--silence-threshold,
    --silence-duration), leading and trailing silence, muted channels,
    empty LFE or centre, clipping, DC offset, out-of-phase segments,
    inverted polarity, mono as stereo, sample rate, bit depth, layout and
    stream start offsets. Terminal, HTML (loudness, levels and phase charts
    per track) and JSON reports, and a loudness item for annotated videos;
    --no-audio leaves it out, --fast --audio adds it to an inspection.
    Validated against the synthetic EBU conformance cases (all within
    tolerance), ffmpeg's ebur128 on real content (within 0.01 LU / dB) and
    synthetic defects (bench/audioval, docs/audio.md).
  • Camera motion (analyze/motion, on by default, --no-motion to skip):
    the global motion of every frame estimated on the shared thumbnails
    (integral-projection predictor, block matching with a Lucas–Kanade
    sub-pixel step, robust similarity fit), and each shot's camera work:
    static, pan, tilt, zoom, tracking (parallax), handheld, mixed, with its
    direction and a shake measure. JSON columns (motionPan, motionTilt,
    motionZoom, motionRoll, motionShake, motionConfidence),
    video.motion and shots[].camera; a camera section and a shot column in
    the terminal report; a camera motion chart, card and shot column in the
    HTML report; a note on shaky shots. About 0.17 ms of one core per frame;
    validated on synthetic moves of known speed (bench/motionval).
  • Annotated videos (--overlay annotated.mp4 on analyze, vmaf and
    run, and a question of the wizard): a copy of the video with the
    analysis burnt in as a debug overlay: timecode, frame number, size and
    keyframes, bitrate, shot and cut markers, camera work with a motion
    vector, SI/TI, luma levels, HDR light levels, black/frozen/banded/
    out-of-range badges, the VMAF and other metrics of each scored frame
    (--exact for every frame), and a timeline with a playhead
    (--overlay-items to choose, --overlay-height for a smaller, faster
    copy). Package overlay writes an ASS script, frame-accurate at any
    frame rate below 100 fps, that libass draws in the H.264 encode of the
    copy (encode.FFmpeg.Burn); same frames, timestamps and audio as the
    source. The copy is encoded by the media engine when there is one
    (--overlay-encoder auto: VideoToolbox on macOS after a test encode,
    NVENC with --gpu, x264 otherwise), in segments split at keyframes and
    rendered concurrently (--overlay-workers), each with its slice of the
    script, joined without re-encoding (a segment that does not render the
    frames planned falls back to one pass): 59 minutes of 1080p25 annotated
    in 217 s on an M2 Max (778 s with x264, 6371 s of CPU against 387 s).
    ffmpeg's libass, and an explicit hardware encoder, are checked before any
    work; qc version --check lists VideoToolbox; the Docker images ship
    libass with DejaVu Sans Mono.
  • Fast frame analysis: on macOS the video is split at keyframes into
    segments decoded concurrently by VideoToolbox (--hwaccel auto, the
    default; videotoolbox and none also accepted), each fed to forks of
    every analyzer (analyze.Forker, analyze.RunSegments) whose per-frame
    series are merged in order: the report is identical to a single pass, to
    the bit, and a segment that does not start where planned falls back to
    one. Segments seek on the container's timeline, so a video starting after
    its audio (some concatenations) no longer falls back. NEON loops on arm64 for SI, TI, luma statistics and thumbnails
    (checked against the portable Go ones), and Unix sockets instead of pipes
    for the raw frames. 59 minutes of 1080p25 H.264 analysed in 62–71 s at
    26 Mbit/s (225 s before) and 49–52 s at 6 Mbit/s (230 s) on an M2 Max; the
    single CPU pass (other systems, --hwaccel none) is 1.1× to 2.7× faster
    too, depending on the share of the decode. See
    analysis.md.
  • VMAF with a confidence interval (qc vmaf): short clips sampled
    across shots (stratified, two-stage) until the 95% interval is narrower
    than --precision, with a real coverage validated by replaying thousands
    of runs (bench/vmafsim). Fixed budgets with --sample (5% of the
    frames, 2/scene), and --exact to score every frame, bit-exact with
    Netflix's vmaf tool at 8 and 10 bits. VMAF v1 models (v1.0.16) picked
    automatically, 4K and high frame rate aware.
  • Metrics beyond VMAF, measured on the same frames with their own
    confidence intervals (--metrics): XPSNR (a pure-Go port matching
    ffmpeg's filter, also usable on its own as quality/xpsnr), CAMBI banding
    with banded segments, PSNR, PSNR-HVS, SSIM, MS-SSIM, CIEDE2000, and the
    whole AOM AV2 common test conditions set with --av2-ctc.
  • VMAF per viewing device (--devices phone,tv,4k), each scored with its
    VMAF v1 model on the same frames.
  • Faster VMAF on macOS: both videos are decoded by concurrent
    VideoToolbox sessions (--hwaccel auto, the default), sweeps are split
    into concurrent runs, and exact measurements are scored in segments split
    at keyframes, three libvmaf contexts at once, with two warm-up frames
    before each so that every frame scores as in a single pass. XPSNR has
    NEON loops (3.8× faster at 1080p). Results are identical to a CPU run,
    frame by frame. On an M2 Max: 26.5 → 18.6 s for the ±0.5 measurement of
    a 10:36 title, 58.7 → 42.1 s for a 59-minute one (184 → 108 s at
    --sample 5%), with a third of the CPU; exact measurements, bound by
    libvmaf's CPU, 149 → 131 s on the 10:36 title and 897 → 704 s on the
    59-minute one. See vmaf.md.
  • Per-title ladders (qc ladder) for H.264 (libx264), HEVC (libx265)
    and AV1 (SVT-AV1): probe encodes of a representative digest, rate-quality
    curves per resolution, their upper envelope, rungs one just-noticeable
    difference apart, then a verification encode of every rung. Lands on the
    exhaustive optimum (−0.04 VMAF, −0.7% bitrate on average) in a fraction
    of the time (bench/ladderval). The curve of the resolution below the top
    rung's is extended when it might reach the top quality at least 10%
    cheaper, so easy content gets a cheaper lower-resolution top rung.
    • Imposed shapes: a rung count or the rung resolutions (--rungs 1080,720,540,360), with the bitrates computed; top and minimum VMAF,
      bitrate bounds, 8 or 10-bit encodes (--encode-bit-depth).
    • Adaptive probing (--probing adaptive): extra probes where the rungs
      are uncertain.
    • Probes at a faster preset (--probe-preset): the rungs are planned on
      them, then the top and bottom rungs encoded at --preset anchor the
      probes onto it (bitrate ratio and CRF offset between the presets); for
      slow delivery presets (SVT-AV1 preset 4 probed at 8: −24% time, same
      rungs).
    • Per-shot rungs (--per-shot): one CRF per shot at an equal
      rate-quality slope, with the per-shot ladder view (every shot's CRF,
      bitrate and VMAF per rung, and each rung's pooled bitrate over the
      title) in the terminal, HTML and JSON reports. The chunks of an encode
      and the per-shot verifications run concurrently, and the source's
      shots come from the analysis qc run already made, or from one run
      alongside the probes (per-shot stage −20% on the drama). Experimental per-shot
      resolution (--per-shot-resolution, implies --per-shot): shots also
      pick their resolution among neighbouring rung resolutions, at equal
      slope.
    • AV1 film grain synthesis (--film-grain auto or a level), detected and
      calibrated (the three calibration levels encoded concurrently),
      fidelity scored against a denoised reference. It cannot be
      combined with per-shot rungs: the combination is rejected before any
      work.
    • Verified rungs checked for banding and for VMAF/XPSNR ranking
      disagreements.
    • Renditions (--encode-ladder <folder>, also asked by the wizard): the
      ladder encoded on the whole title with the settings of each rung, and
      each rung's per-shot version, with progress in the dashboard. Each
      rendition is checked against the source (--no-rendition-check skips
      it): its VMAF with its confidence interval and its bitrate, shown next
      to the predictions in the reports, with findings when they drift from
      them (quality beyond the check's interval, bitrate beyond Apple's 10%
      BANDWIDTH rule).
  • HDR (docs/hdr.md): HDR10, PQ and HLG detected from the
    stream and its first frame (HDR10 metadata in SEI, HDR10+), Dolby Vision
    and HDR10+ reported without processing their dynamic metadata; signalling
    checks (BT.2020 primaries and matrix, 10 bits, narrow range, missing
    mastering display or content light level). MaxCLL and MaxFALL measured
    during the frame analysis (CTA-861.3, on a grid of 10-bit samples that
    ffmpeg writes to a second pipe of the same decode, robust to 4:2:0 chroma
    overshoots) and compared with the signalled
    v...
Read more