Skip to content

Releases: speechlab0210/tutorial-video-skill

Tutorial Video Skill v1.2.0

Choose a tag to compare

@speechlab0210 speechlab0210 released this 03 Oct 09:57

Tutorial Video Skill v1.2.0

Checking synthesized or recorded narration with a speech recognizer, made explicit. Full list: CHANGELOG.md.

New helper project.py speech. Give it a recognizer's plain-text transcript of a take or of the exported audio, and the scene script. It reports similarity, the passages that are missing, extra, or replaced, the numbers on each side with written and spoken forms treated as equal (1600 = 一千六百 = one thousand six hundred; 5萬7千; 6:09 = 六點零九分; August eleventh = August 11th), percent and minus signs, and negation words. It exits 1 when numbers, signs, or negations differ, the transcript is empty, or similarity falls below a threshold you calibrated. It runs no recognizer and does not listen: every flag is a place to hear.

Calibrated before release on 1,397 accepted Chinese and 103 English narration segments from finished videos (three recognizers). 4.4% of Chinese and 2.9% of English segments are flagged for a number difference, nearly all recognizer homophones or dropped words, and 0.9% of Chinese segments for a negation. Every deliberately planted digit, numeral, dropped number, and dropped or replaced negation was caught. A first draft flagged 9% of correct Chinese segments; reading those false alarms drove the fixes (for example, the 一 in 這一題 is not a quantity, 兩三 is a range, and percent signs are checked instead of dropped).

New guidance: a score ranks takes but does not say what the chosen take says, so read its transcript; compare numbers, signs, and negations as sets; re-check the exported audio scene by scene, because noise reduction, speed changes, and assembly can change speech that passed as a take; re-baseline accepted takes when the recognizer changes; convert Traditional/Simplified before judging differences.

Install: download tutorial-video-skill-v1.2.0.zip, unpack it, and copy its tutorial-video/ subfolder into your agent's skills directory (for example ~/.claude/skills/tutorial-video/ or ~/.codex/skills/tutorial-video/). English and Traditional Chinese guides are included. MIT licensed.

Validation: 39 automated tests pass with FFmpeg available. Two independently built ZIPs are byte-identical (SHA-256 in SHA256SUMS.txt). These tests and the calibration establish helper behavior only; no controlled learner study has been done.

This is an AI-authored resource, not an official conference resource or an instructor endorsement.

Tutorial Video Skill v1.1.0

Choose a tag to compare

@speechlab0210 speechlab0210 released this 03 Oct 06:43

Tutorial Video Skill v1.1.0

A review of v1.0.0 found a real bug in the bundled assembler and a set of smaller tool and guidance gaps. This release fixes them. Full list: CHANGELOG.md.

Most important fix — narration no longer drifts behind the slides. v1.0.0 compressed each scene's audio separately and joined the clips by stream copy, so every join added about 29 ms of delay. In a 40-scene synthetic test at 30 fps, measured in the exported MP4, the last scene's speech started 1.13 s after its slide. v1.1.0 places every scene on the frame grid and encodes the narration once as one continuous track; the same test measures 0.1 ms at the start, middle, and end (30, 25, and 7 fps). The assembler now checks decoded audio length against the video before writing the file. If you built videos with the v1.0.0 assembler, check their later scenes.

Also fixed: VBR MP3 endings, % in slide names, transparent slides rendering black, GIF slides, crashes on Chinese text when output is piped under a legacy Windows code page, extra blank lines in SRT files, Windows junctions in manifests, and error messages that now name the failing scene. New guidance covers measuring sync in the exported file, disclosing synthetic narration, checking heteronyms by ear, authorization scope, contrast, and Claude Code / Codex install paths.

Install: download tutorial-video-skill-v1.1.0.zip, unpack it, and copy its tutorial-video/ subfolder into your agent's skills directory (for example ~/.claude/skills/tutorial-video/ or ~/.codex/skills/tutorial-video/). English and Traditional Chinese guides are included. MIT licensed.

Validation: 28 automated tests pass with FFmpeg available, including media tests that measure the shipped MP4; on Windows, each new test fails against the v1.0.0 scripts. Two independently built ZIPs are byte-identical (SHA-256 in SHA256SUMS.txt). These tests establish helper behavior only. No controlled learner study has been done.

This is an AI-authored resource, not an official conference resource or an instructor endorsement.

Tutorial Video Skill v1.0.0

Choose a tag to compare

@speechlab0210 speechlab0210 released this 03 Oct 04:27

Tutorial Video Skill v1.0.0

A general AI agent skill for creating coherent, accurate, editable course videos across subjects and talk formats.

Includes teaching design, evidence checking, visual explanations, narration and pacing, media production, revision synchronization, and three-part quality review. The skill distills existing teaching craft, production engineering, and lessons reinforced through instructor collaboration.

Download tutorial-video-skill-v1.0.0.zip and copy its complete tutorial-video/ folder into your agent's configured skills directory. English and Traditional Chinese getting-started guides are included. Original package content is MIT licensed.

Optional Python helpers check plans, captions, and file integrity. The FFmpeg assembler builds narrated still scenes from your own images and authorized audio. It does not generate speech or slides, or establish audience comprehension.

Validation: official skill format check passed; 17 helper/integration tests passed with no skips, including a real synthetic two-scene movie, decoded scene order, audible test tones, retained ending silence, and no-overwrite behavior. Internal documentation links and public-source/history checks passed. No independent agent evaluation or learner-effectiveness study is claimed.

This is an AI-authored resource, not an official conference resource or an instructor endorsement. No private tutorial correspondence, original conference media, or voice model is included.

Known issue, fixed in v1.1.0: this version's assembler lets narration drift later than its slides at every scene join (about 29 ms per join; 1.13 s late by scene 40 in a synthetic test). Please use v1.1.0.