[extractor/substack] Return canonical URLs #8219
Merged
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Description of your pull request and other information
The URL passed to _real_extract is of format "{username}.substack.com", but for blogs with custom domains the canonical URL would use the custom domain. Because of the wrong URL, the cookies in the resulting info dict come from the wrong domain, which breaks subscriber content extraction.
This isn't really testable because for whatever reason yt-dlp itself doesn't have any trouble downloading such content; however,
yt-dlp -j
consumers are broken without this change because of how thecookies
field is populated.Template
Before submitting a pull request make sure you have:
In order to be accepted and merged into yt-dlp each piece of code must be in public domain or released under Unlicense. Check all of the following options that apply:
What is the purpose of your pull request?
Copilot Summary
🤖 Generated by Copilot at fa51465
Summary
🎥🛠️🔗
Enhanced the Substack extractor to support custom domains. Used
canonical_url
instead ofusername
for video extraction and added more metadata fields.Walkthrough
_extract_video_formats
function to useurl
instead ofusername
andurllib.parse.urljoin
to construct video URL (link)_real_extract
function to getdomain
value fromwebpage_info
and use it to setcanonical_url
with custom domain if applicable (link)canonical_url
instead ofusername
to_extract_video_formats
function in_real_extract
function (link)webpage_url
andoriginal_url
fields to metadata dictionary in_real_extract
function, usingcanonical_url
as the value (link)