[extractor/substack] Fix embed URL extraction #8218
Merged
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Description of your pull request and other information
The format of the JSON payload being extracted has changed to a JSON.parse("...")-style format. This PR fixes the extraction process so that it ignores the added slashes.
There's a theoretical possibility of the extracted string being too short if it contains an escape sequence (which would contain a slash thus not matching the
[^\"]+
part). In practice, a valid DNS name is unlikely to get escaped by a sensible JSON encoder. I have searched for uses of_search_json
and_parse_json
and found none, so I opted for a fix to the regular expression.Regarding tests, there are two in
generic.py
which should have actually been failing. If I try them manually usingyt-dlp -j
, the result seems to match, but I don't know how to run specific download tests in a generic way.Template
Before submitting a pull request make sure you have:
In order to be accepted and merged into yt-dlp each piece of code must be in public domain or released under Unlicense. Check all of the following options that apply:
What is the purpose of your pull request?
Copilot Summary
馃 Generated by Copilot at 6cfb6a0
Summary
馃悰馃敡馃И
Fixed substack extractor to handle backslashes in subdomain names. Modified regex pattern in
yt_dlp/extractor/substack.py
.Walkthrough
yt_dlp/extractor/substack.py