Tutorial Video Skill v1.2.0
Checking synthesized or recorded narration with a speech recognizer, made explicit. Full list: CHANGELOG.md.
New helper project.py speech. Give it a recognizer's plain-text transcript of a take or of the exported audio, and the scene script. It reports similarity, the passages that are missing, extra, or replaced, the numbers on each side with written and spoken forms treated as equal (1600 = 一千六百 = one thousand six hundred; 5萬7千; 6:09 = 六點零九分; August eleventh = August 11th), percent and minus signs, and negation words. It exits 1 when numbers, signs, or negations differ, the transcript is empty, or similarity falls below a threshold you calibrated. It runs no recognizer and does not listen: every flag is a place to hear.
Calibrated before release on 1,397 accepted Chinese and 103 English narration segments from finished videos (three recognizers). 4.4% of Chinese and 2.9% of English segments are flagged for a number difference, nearly all recognizer homophones or dropped words, and 0.9% of Chinese segments for a negation. Every deliberately planted digit, numeral, dropped number, and dropped or replaced negation was caught. A first draft flagged 9% of correct Chinese segments; reading those false alarms drove the fixes (for example, the 一 in 這一題 is not a quantity, 兩三 is a range, and percent signs are checked instead of dropped).
New guidance: a score ranks takes but does not say what the chosen take says, so read its transcript; compare numbers, signs, and negations as sets; re-check the exported audio scene by scene, because noise reduction, speed changes, and assembly can change speech that passed as a take; re-baseline accepted takes when the recognizer changes; convert Traditional/Simplified before judging differences.
Install: download tutorial-video-skill-v1.2.0.zip, unpack it, and copy its tutorial-video/ subfolder into your agent's skills directory (for example ~/.claude/skills/tutorial-video/ or ~/.codex/skills/tutorial-video/). English and Traditional Chinese guides are included. MIT licensed.
Validation: 39 automated tests pass with FFmpeg available. Two independently built ZIPs are byte-identical (SHA-256 in SHA256SUMS.txt). These tests and the calibration establish helper behavior only; no controlled learner study has been done.
This is an AI-authored resource, not an official conference resource or an instructor endorsement.