Skip to content

ArthurFDLR/yt-cc

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

7 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

YoutubeCC (yt-cc) - A Youtube Closed Captions (.json3) parser

yt-cc is a Python library that parses Youtube Closed Captions (.json3) files. It is a simple and easy-to-use library that can be used to query precise parts of the video transcript or iterate over the entire transcript.

%load_ext autoreload
%autoreload 2
from pathlib import Path

from yt_dlp import YoutubeDL
from yt_cc import YoutubeCC
DATA_PATH = Path.cwd() / ".tmp"

Downloading the Closed Captions with yt-dlp

To download the closed captions of a Youtube video, you can use the yt-dlp command-line tool. You can install yt-dlp using pip:

pip install yt-dlp

To download the closed captions of a Youtube video in the .json3 format, you can use the following command:

yt-dlp --write-auto-sub --skip-download --sub-format json3 <video-url>

You can also download the closed captions with the yt-dlp Python API:

YoutubeDL(
    params = dict(
        paths=dict(home=str(DATA_PATH)),
        skip_download=True,
        outtmpl="%(id)s.%(ext)s",
        subtitlesformat="json3",
        writeautomaticsub=True,
    )
).download(["https://www.youtube.com/watch?v=oHWuv1Aqrzk"])

yt_cc_file_path = next(DATA_PATH.glob("*.json3"))
yt_cc_file_path
[youtube] Extracting URL: https://www.youtube.com/watch?v=oHWuv1Aqrzk
[youtube] oHWuv1Aqrzk: Downloading webpage
[youtube] oHWuv1Aqrzk: Downloading ios player API JSON
[youtube] oHWuv1Aqrzk: Downloading android player API JSON


WARNING: [youtube] Skipping player responses from android clients (got player responses for video "aQvGIIdgFDM" instead of "oHWuv1Aqrzk")


[youtube] oHWuv1Aqrzk: Downloading m3u8 information
[info] oHWuv1Aqrzk: Downloading subtitles: en
[info] oHWuv1Aqrzk: Downloading 1 format(s): 248+251
Deleting existing file /home/arthur/Documents/02.workspace/02.active/clips-analytics/yt-cc/.tmp/oHWuv1Aqrzk.en.json3
[info] Writing video subtitles to: /home/arthur/Documents/02.workspace/02.active/clips-analytics/yt-cc/.tmp/oHWuv1Aqrzk.en.json3
[download] Destination: /home/arthur/Documents/02.workspace/02.active/clips-analytics/yt-cc/.tmp/oHWuv1Aqrzk.en.json3
[download] 100% of   71.66KiB in 00:00:00 at 1.28MiB/s





PosixPath('/home/arthur/Documents/02.workspace/02.active/clips-analytics/yt-cc/.tmp/oHWuv1Aqrzk.en.json3')

Parse the Closed Captions with yt-cc

youtube_caption = YoutubeCC(yt_cc_file_path)
youtube_caption

YoutubeCC

  • File: /home/arthur/Documents/02.workspace/02.active/clips-analytics/yt-cc/.tmp/oHWuv1Aqrzk.en.json3
  • lines: 202
  • Segments: 768
for i, line_cc in enumerate(youtube_caption):
    print(line_cc)
    if i > 5:
        break
LineCC(event_id=1.0, start_time_ms=0.0, duration_ms=218599.0, window_id=nan, window_style_id=1.0, window_position_id=1.0, append=nan, segments=[])
LineCC(event_id=nan, start_time_ms=2820.0, duration_ms=5279.0, window_id=1.0, window_style_id=nan, window_position_id=nan, append=nan, segments=['is', ' there', ' cool', ' small', ' projects', ' like', ' uh'])
LineCC(event_id=nan, start_time_ms=5510.0, duration_ms=2589.0, window_id=1.0, window_style_id=nan, window_position_id=nan, append=1.0, segments=['\n'])
LineCC(event_id=nan, start_time_ms=5520.0, duration_ms=6180.0, window_id=1.0, window_style_id=nan, window_position_id=nan, append=nan, segments=['archive', ' sanity', ' and', ' and', ' so', ' on', ' that', " you're"])
LineCC(event_id=nan, start_time_ms=8089.0, duration_ms=3611.0, window_id=1.0, window_style_id=nan, window_position_id=nan, append=1.0, segments=['\n'])
LineCC(event_id=nan, start_time_ms=8099.0, duration_ms=5580.0, window_id=1.0, window_style_id=nan, window_position_id=nan, append=nan, segments=['thinking', ' about', ' the', ' the', ' the', ' the', ' world', ' the'])
LineCC(event_id=nan, start_time_ms=11690.0, duration_ms=1989.0, window_id=1.0, window_style_id=nan, window_position_id=nan, append=1.0, segments=['\n'])
youtube_caption.lines
<style scoped> .dataframe tbody tr th:only-of-type { vertical-align: middle; }
.dataframe tbody tr th {
    vertical-align: top;
}

.dataframe thead th {
    text-align: right;
}
</style>
event_id start_time_ms duration_ms window_id window_style_id window_position_id append
0 1.0 0 218599.0 NaN 1.0 1.0 NaN
1 NaN 2820 5279.0 1.0 NaN NaN NaN
2 NaN 5510 2589.0 1.0 NaN NaN 1.0
3 NaN 5520 6180.0 1.0 NaN NaN NaN
4 NaN 8089 3611.0 1.0 NaN NaN 1.0
... ... ... ... ... ... ... ...
197 NaN 212580 3900.0 1.0 NaN NaN NaN
198 NaN 214850 1630.0 1.0 NaN NaN 1.0
199 NaN 214860 3739.0 1.0 NaN NaN NaN
200 NaN 216470 2129.0 1.0 NaN NaN 1.0
201 NaN 216480 2119.0 1.0 NaN NaN NaN

202 rows × 7 columns

youtube_caption.segments
<style scoped> .dataframe tbody tr th:only-of-type { vertical-align: middle; }
.dataframe tbody tr th {
    vertical-align: top;
}

.dataframe thead th {
    text-align: right;
}
</style>
text asr_confidence offset_ms pen_id start_time_ms line_id
0 is 248.0 NaN None 2820 1
1 there 248.0 599.0 None 3419 1
2 cool 248.0 839.0 None 3659 1
3 small 248.0 1020.0 None 3840 1
4 projects 248.0 1380.0 None 4200 1
... ... ... ... ... ... ...
763 sounds 248.0 1260.0 None 216120 199
764 kind 240.0 1320.0 None 216180 199
765 of 248.0 1500.0 None 216360 199
766 \n NaN NaN None 216470 200
767 crazy 248.0 NaN None 216480 201

768 rows × 6 columns

print(youtube_caption.get_text(start_time_ms=0, end_time_ms=10000))
is there cool small projects like uh
archive sanity and and so on that you're
thinking about the the the

About

💬 YoutubeCC - Parse JSON3 Youtube Closed Captions

Resources

Stars

Watchers

Forks

Used by

Contributors

Languages