read-art

Readability reference to Arc90's.
Scrape article from any page (automatically).
Make any web page readable, no matter Chinese or English.

快速抓取网页文章标题和内容，适合node.js爬虫使用，服务于ElasticSearch。

Guide

How it works

In my case, the speed of spider is about 1500k documents per day, and the maximize crawling speed is 1.2k /minute, avg 1k /minute, the memory cost are about 200 MB on each spider kernel, and the accuracy is about 90%, the rest 10% can be fixed by customizing Score Rules or Selectors. it's better than any other readability modules.

(4) Server infos:

20M bandwidth of fibre-optical

8 Intel(R) Xeon(R) CPU E5-2650 v2 @ 2.60GHz cpus

32G memory

Name		Name	Last commit message	Last commit date
Latest commit History 261 Commits
docs		docs
examples		examples
lib		lib
test		test
.gitignore		.gitignore
.npmignore		.npmignore
.overcommit.yml		.overcommit.yml
.travis.yml		.travis.yml
HISTORY.md		HISTORY.md
README.md		README.md
index.js		index.js
package.json		package.json

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

read-art

Guide

How it works

About

Releases 1

Packages

Contributors 3

Languages

Tjatse/node-readability

Folders and files

Latest commit

History

Repository files navigation

read-art

Guide

How it works

About

Resources

Stars

Watchers

Forks

Releases 1

Packages 0

Contributors 3

Languages

Packages