Skip to content
simplified scrapy, A Simple Web Crawle
Python
Branch: master
Clone or download
Latest commit 6b1354a Jan 20, 2020
Permalink
Type Name Latest commit message Commit time
Failed to load latest commit information.
simplified_html 0.7.90 Jan 4, 2020
simplified_scrapy Update regex_dic.py Jan 20, 2020
spiders 0.7.90 Jan 4, 2020
.gitignore 0.7.90 Jan 4, 2020
LICENSE Initial commit Aug 7, 2019
README.md Update README.md Jan 2, 2020
setting.json 0.7.90 Jan 4, 2020
setting.py 0.7.90 Jan 4, 2020
setup.py select Jan 18, 2020
spider_resource.py 0.7.90 Jan 4, 2020
start.py 0.2.35 Dec 7, 2019

README.md

simplified-scrapy

simplified scrapy, A Simple Web Crawle

Requirements

  • Python 2.7, 3.0+
  • Works on Linux, Windows, Mac OSX, BSD

run

from simplified_scrapy.simplified_main import SimplifiedMain
SimplifiedMain.startThread()

Demo

Custom crawler class needs to extend Spider class

from core.spider import Spider 
class DemoSpider(Spider):

Here is an example of collecting data

from simplified_scrapy.spider import Spider, SimplifiedDoc
class DemoSpider(Spider):
  name = 'demo-spider'
  start_urls = ['http://quotes.toscrape.com/']
  allowed_domains = ['quotes.toscrape.com']
  def extract(self, url, html, models, modelNames):
    doc = SimplifiedDoc(html)
    lstA = doc.listA(url=url["url"])
    return [{"Urls": lstA, "Data": None}]

from simplified_scrapy.simplified_main import SimplifiedMain
SimplifiedMain.startThread(DemoSpider())

pip install

pip install simplified-scrapy

Examples

Legal Issues

In particular, please be aware that

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.

Translated to human words:

In case your use of the software forms the basis of copyright infringement, or you use the software for any other illegal purposes, the authors cannot take any responsibility for you.

We only ship the code here, and how you are going to use it is left to your own discretion.

You can’t perform that action at this time.