Conversation
Signed-off-by: Evan Lloyd New-Schmidt <evan@new-schmidt.com>
Signed-off-by: Evan Lloyd New-Schmidt <evan@new-schmidt.com>
|
I've been using the Pomona College article ( The original file is 1.5mB. The output looks like this: pomona_college.html.txt |
|
Need to remove whole first paragraph. Remove section, span. |
Do you mean: or: ? |
|
Can empty sections like Notes be detected and deleted? Or maybe remove all Notes completely? Is there any value in them? |
|
The first one. |
Yes, I'm working on that - I think it should handle most of the galleries too. |
|
With empty sections removed, sections expanded, and doctype/html/body/whitespace removed: pomona_college2.html.txt I think it looks pretty good! Need to try it on more articles. |
|
The main issues I've seen, which are present in the old scraper extracts too:
|
|
Cool, so it works! Let's finish this PR! |
- Article contents are from the 2023-04-01 Wikipedia Enterprise Dump - Add benchmark for HTML processing Signed-off-by: Evan Lloyd New-Schmidt <evan@new-schmidt.com>
Signed-off-by: Evan Lloyd New-Schmidt <evan@new-schmidt.com>
Signed-off-by: Evan Lloyd New-Schmidt <evan@new-schmidt.com>
Signed-off-by: Evan Lloyd New-Schmidt <evan@new-schmidt.com>
Signed-off-by: Evan Lloyd New-Schmidt <evan@new-schmidt.com>
Signed-off-by: Evan Lloyd New-Schmidt <evan@new-schmidt.com>
Signed-off-by: Evan Lloyd New-Schmidt <evan@new-schmidt.com>
Signed-off-by: Evan Lloyd New-Schmidt <evan@new-schmidt.com>
Signed-off-by: Evan Lloyd New-Schmidt <evan@new-schmidt.com>
Signed-off-by: Evan Lloyd New-Schmidt <evan@new-schmidt.com>
- Combine expansion steps - Pull original steps into functions - Use parent sections for removing specific headers - Remove head in initial stage Signed-off-by: Evan Lloyd New-Schmidt <evan@new-schmidt.com>
Signed-off-by: Evan Lloyd New-Schmidt <evan@new-schmidt.com>
Additional work to bring HTML simplification in line with the Wikipedia Extracts API.
Closes #4.
Remaining work:
ids