-
Notifications
You must be signed in to change notification settings - Fork 0
ASR and Video Analysis
English | 中文 | 日本語 | 한국어 | Español | Português | Русский
Attach a subtitle-less Bilibili / YouTube video and browsa fetches the audio in-page, runs cloud ASR, and produces a [mm:ss] transcript — or runs 视听精读 (audiovisual deep-read): a figure-and-text document of speech + visuals. The ASR provider today is Volcengine Ark.
URL direct-pass is dead for BOTH platforms: Bilibili signs CDN URLs to the user's session IP (the Ark server can't fetch them); YouTube adds PO-token + session + cookie binding (server-side fetches 403). The browser — the client whose IP matches the signature — must fetch the bytes itself and upload them to the Files API. This is a root-cause conclusion, not one option among several.
-
ATTACH_PAGEreturns{mode:'asr-pending'}(deferred-storage handoff). -
buildAsrPendingCtxdispatches by platform: Bilibili re-injects its content script to read streams; YouTube calls MAIN-world__browsaFetchFreshYouTubeStreams(a fresh/playerrequest minting a current PO token). - Stream selection: lowest-bitrate pure audio (it gets transcoded anyway) but REJECT streams shorter than 90% of the true duration (real truncated-stream incident); prefer AAC-LC (
decodeAudioDatamay reject HE-AAC). - DNR header injection (
lib/sidepanel/media-headers.js): Bilibili needsReferer; YouTube ALSO needsOriginrewrite + cookies (googlevideo 403s header-less extension fetches — verified live). -
Transcode to 16kHz mono WAV: a Bilibili m4s is a video-less fMP4 MP4; Ark classifies files BY CONTENT as "video" →
Invalid video_url→ status failed. WAV is the empirically proven format. Web Audio decode + hand-written linear-interpolation resample (pure, Node-testable). - Upload to the Files API (≤512MB multipart):
expire_at= 30 days, semantic filenames;browsaArkFileCachestores file_ids (key = base|key-fingerprint|assetId) so the same video skips re-upload within 30 days (liveness-probed before reuse). - Poll file status → streaming transcription via
/responses(input_audio.file_id), the prompt forcing[mm:ss]per line. -
ATTACH_ASR_CONFIRMstores the transcript +videoSrcstamp (the precondition for timestamp pills / the timeline drawer).
- Explicit
max_output_tokens: 65536on/responses(the model's default output cap silently cuts mid-audio). - Parse the SSE
incomplete/finish_reason:'length'signals →truncated:true. - Completeness gate: last timestamp < 90% of video duration → throw, fail open to plain text + toast (NOT stored).
- Hole gate:
largestTranscriptGapSecinterval-merges the timeline; a gap > 300s rejects (an 81-minute video once lost an hour in the middle).
input_video (+ optional separate input_audio), max_output_tokens 65536, streaming figure-and-text output. Prompt discipline (calibrated against real runs): label speakers whenever ANY second voice appears, append identifiable names, compress ad reads to one line, keyframe screenshots ONLY when the argument depends on seeing them (soft density hint + explicit do-not-capture list; the only hard ceiling is the client-side SAFETY_KEYFRAME_CAP 24 — a pathological-output guard, not a quality knob). [图N] markers share one anchoring protocol with page images and the timeline drawer.
-
Bilibili:
/x/player/wbi/v2track list + CDN subtitle JSON fetched actively — independent of the CC toggle. -
YouTube: timedtext serves real content ONLY to the player's own POT-bearing requests (direct fetches of
captionTracks.baseUrlreturn empty bodies); the interceptor captures material only if CC was on at least once. Fallback stage 2c drives the player's caption module to mint one authorized request (the brief on-screen caption flash is expected); a tracklist-empty verdict is negative-cached per videoId. - "No subtitles" is judged from the structured
noTranscriptflag, NOT the## 字幕text marker (auto mode's silent Jina fallback swallows text markers — real bug).
lib/asr-providers.js (metadata registry; zero UI changes) + ASR_ADAPTERS in attach-asr.js (same-signature {transcribeAudio, analyzeVideo}). The hard part is always the endpoint's file-upload capability (long audio can't ride inline base64). A full Qwen/DashScope adapter was built and removed after one day (browser→OSS unstable; preserved in git history); the tabCapture record-while-playing fallback was rejected (paused/skipped segments are silence — an unplayed video can't be transcribed).
Source of truth: AGENTS.md "Video ASR subtitles", "YouTube transcript acquisition vs POT"; the attach-asr.js header comment (full history). Synced 2026-10-01.
English
- Home
- Architecture
- Rendering Pipeline
- Storage Model
- Providers and Agents
- ASR and Video Analysis
- Security Model
- Design Decisions
- Contributing
中文
相关 / Related
日本語
한국어
Español
- Inicio
- Arquitectura
- Pipeline de renderizado
- Modelo de almacenamiento
- Proveedores y agentes
- ASR y análisis de vídeo
- Modelo de seguridad
- Decisiones de diseño
- Contribuir
Português
- Início
- Arquitetura
- Pipeline de renderização
- Modelo de armazenamento
- Provedores e agentes
- ASR e análise de vídeo
- Modelo de segurança
- Decisões de design
- Contribuindo
Русский