GitHub - hrbrmstr/jericho: :notebook_with_decorative_cover: Extract plain or structured text from HTML content in R

jericho : Break Down the Walls of ‘HTML’ Tags into Usable Text

Structured ‘HTML’ content can be useful when you need to parse data tables or other tagged data from within a document. However, it is also useful to obtain “just the text” from a document free from the walls of tags that surround it. Tools are provied that wrap methods in the ‘Jericho HTML Parser’ Java library by Martin Jericho http://jericho.htmlparser.net/docs/index.html. Martin’s library is used in many at-scale projects, icluding the ‘The Internet Archive’.

As a result of using a Java library, this package requires rJava.

The following functions are implemented:

html_to_text: Convert HTML to Text
render_html_to_text: Render HTML to Text

Installation

If you do use devtools, then it should pickup the Remotes: section in DESCRIPTION. Until the package is on CRAN, you might want to also invoke the installation of jerichojars as shown below:

install.packages(c("jerichojars", "jericho"), repos = "https://cinc.rud.is/")

Usage

Let’s use this NASA blog post as an example.

library(jericho)

# current verison
packageVersion("jericho")

## [1] '0.2.0'

URL <- "https://blogs.nasa.gov/spacestation/2017/09/02/touchdown-expedition-52-back-on-earth/"
  
doc <- paste0(readr::read_lines(URL), collapse = "\n")

This is pure text extraction:

html_to_text(doc)

This provides a human readable version of the segment content that is modelled on the way Mozilla Thunderbird and other email clients provide an automatic conversion of HTML content to text in their alternative MIME encoding of emails.

render_html_to_text(doc)

You should run each to see and compare the output (GitHub markdown documents aren’t the best viewing medium).

`jericho` Metrics

Lang	# Files	(%)	LoC	(%)	Blank lines	(%)	# Lines	(%)
Java	2	0.18	49	0.38	9	0.19	14	0.13
R	6	0.55	40	0.31	10	0.21	62	0.56
Maven	1	0.09	23	0.18	1	0.02	1	0.01
Rmd	1	0.09	9	0.07	24	0.50	33	0.30
make	1	0.09	8	0.06	4	0.08	0	0.00

Name		Name	Last commit message	Last commit date
Latest commit History 13 Commits
R		R
inst		inst
java/jericho		java/jericho
man		man
tests		tests
.Rbuildignore		.Rbuildignore
.codecov.yml		.codecov.yml
.gitignore		.gitignore
.travis.yml		.travis.yml
DESCRIPTION		DESCRIPTION
LICENSE		LICENSE
NAMESPACE		NAMESPACE
NEWS.md		NEWS.md
README.Rmd		README.Rmd
README.md		README.md
appveyor.yml		appveyor.yml
jericho.Rproj		jericho.Rproj

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Installation

Usage

`jericho` Metrics

About

Releases

Packages

Languages

License

hrbrmstr/jericho

Folders and files

Latest commit

History

Repository files navigation

Installation

Usage

jericho Metrics

About

Topics

Resources

License

Stars

Watchers

Forks

Releases

Packages 0

Languages

`jericho` Metrics

Packages