Skip to content

ASR and Video Analysis

Hermes Agent edited this page Oct 1, 2026 · 2 revisions

ASR and Video Analysis

English | 中文 | 日本語 | 한국어 | Español | Português | Русский

Attach a subtitle-less Bilibili / YouTube video and browsa fetches the audio in-page, runs cloud ASR, and produces a [mm:ss] transcript — or runs 视听精读 (audiovisual deep-read): a figure-and-text document of speech + visuals. The ASR provider today is Volcengine Ark.

Why "download → upload file" and not URL direct-pass

URL direct-pass is dead for BOTH platforms: Bilibili signs CDN URLs to the user's session IP (the Ark server can't fetch them); YouTube adds PO-token + session + cookie binding (server-side fetches 403). The browser — the client whose IP matches the signature — must fetch the bytes itself and upload them to the Files API. This is a root-cause conclusion, not one option among several.

The pipeline (eight steps)

  1. ATTACH_PAGE returns {mode:'asr-pending'} (deferred-storage handoff).
  2. buildAsrPendingCtx dispatches by platform: Bilibili re-injects its content script to read streams; YouTube calls MAIN-world __browsaFetchFreshYouTubeStreams (a fresh /player request minting a current PO token).
  3. Stream selection: lowest-bitrate pure audio (it gets transcoded anyway) but REJECT streams shorter than 90% of the true duration (real truncated-stream incident); prefer AAC-LC (decodeAudioData may reject HE-AAC).
  4. DNR header injection (lib/sidepanel/media-headers.js): Bilibili needs Referer; YouTube ALSO needs Origin rewrite + cookies (googlevideo 403s header-less extension fetches — verified live).
  5. Transcode to 16kHz mono WAV: a Bilibili m4s is a video-less fMP4 MP4; Ark classifies files BY CONTENT as "video" → Invalid video_url → status failed. WAV is the empirically proven format. Web Audio decode + hand-written linear-interpolation resample (pure, Node-testable).
  6. Upload to the Files API (≤512MB multipart): expire_at = 30 days, semantic filenames; browsaArkFileCache stores file_ids (key = base|key-fingerprint|assetId) so the same video skips re-upload within 30 days (liveness-probed before reuse).
  7. Poll file status → streaming transcription via /responses (input_audio.file_id), the prompt forcing [mm:ss] per line.
  8. ATTACH_ASR_CONFIRM stores the transcript + videoSrc stamp (the precondition for timestamp pills / the timeline drawer).

Truncation defenses — never silently store a partial transcript (four layers, incident-driven)

  1. Explicit max_output_tokens: 65536 on /responses (the model's default output cap silently cuts mid-audio).
  2. Parse the SSE incomplete / finish_reason:'length' signals → truncated:true.
  3. Completeness gate: last timestamp < 90% of video duration → throw, fail open to plain text + toast (NOT stored).
  4. Hole gate: largestTranscriptGapSec interval-merges the timeline; a gap > 300s rejects (an 81-minute video once lost an hour in the middle).

Video deep-read (the video itself goes to the model)

input_video (+ optional separate input_audio), max_output_tokens 65536, streaming figure-and-text output. Prompt discipline (calibrated against real runs): label speakers whenever ANY second voice appears, append identifiable names, compress ad reads to one line, keyframe screenshots ONLY when the argument depends on seeing them (soft density hint + explicit do-not-capture list; the only hard ceiling is the client-side SAFETY_KEYFRAME_CAP 24 — a pathological-output guard, not a quality knob). [图N] markers share one anchoring protocol with page images and the timeline drawer.

Platform notes

  • Bilibili: /x/player/wbi/v2 track list + CDN subtitle JSON fetched actively — independent of the CC toggle.
  • YouTube: timedtext serves real content ONLY to the player's own POT-bearing requests (direct fetches of captionTracks.baseUrl return empty bodies); the interceptor captures material only if CC was on at least once. Fallback stage 2c drives the player's caption module to mint one authorized request (the brief on-screen caption flash is expected); a tracklist-empty verdict is negative-cached per videoId.
  • "No subtitles" is judged from the structured noTranscript flag, NOT the ## 字幕 text marker (auto mode's silent Jina fallback swallows text markers — real bug).

The seam for a second ASR provider (reserved)

lib/asr-providers.js (metadata registry; zero UI changes) + ASR_ADAPTERS in attach-asr.js (same-signature {transcribeAudio, analyzeVideo}). The hard part is always the endpoint's file-upload capability (long audio can't ride inline base64). A full Qwen/DashScope adapter was built and removed after one day (browser→OSS unstable; preserved in git history); the tabCapture record-while-playing fallback was rejected (paused/skipped segments are silence — an unplayed video can't be transcribed).


Source of truth: AGENTS.md "Video ASR subtitles", "YouTube transcript acquisition vs POT"; the attach-asr.js header comment (full history). Synced 2026-10-01.

Clone this wiki locally