Skip to content

2 Submitting multiple seeds to the crawler

jackpay edited this page Jan 29, 2020 · 10 revisions

This uses the CrawlerSubmissionService to submit a potentially large number of seed urls to crawl using JQM

This is achieved by running the CrawlerSubmissionService java program. To run this program navigate to the m52BLAHJ jar with the necessary input parameters.

The parameters required to submit seeds to JQM are as follows.

Preliminary steps for submitting multiple seeds:

  • Generate the csv from which the seed list json file can be created. This will typically be some variation on the source list provided by the client (i.e. ACLED).
  • The .csv file has three required fields in order for the crawling/scraping architecture to operate properly.
  1. SOURCE - this field specifies the root URL or domain which is crawled and must be identical to those sources named in the acled_source table in the acled_camunda database.
  2. LINK - The is the precise starting seed URL presented to a crawler.
  3. SCRAPER - The individual name of the scraper json file, which defines the scraping rules for the specific domain expressed in LINK.
    NB: The directory containing all scrapers is specified in the crawler JobDef xml (an example of which is given in this project).
  • The python file build_seed_list.py (provided in this repo) can be used to generate the seed json file. This script also allows the specification of a black-list (seed exclusion) or white-list (exclusive seed inclusion).

  • Once generated, the seed json file is ready for the submission service. The exact details of how to submit these seeds is given below.

The java program and main class:

  • Typically it is useful to set this process up as a separate screen on the server. e.g. screen -S SEEDSUBMISSION.
  • The program(jar)/project containing the submission service is m52-norconex-crawler.
  • The main class of the submission service is uk.ac.susx.tag.norconex.jobqueuemanager.CrawlerSubmissionService.
  • The two main parameters of the submission service given below.
    e.g. java -cp m52-norconex-crawler-x.x.x.jar uk.ac.susx.tag.norconex.jobqueuemanager.CrawlerSubmissionService '<SEED_JSON_LOCATION>/sample-source.json' '<PROPS_LOCATION>/crawlmanager.properties

Parameter 1: The seeds to be crawled.

Preparing the seed urls for submission.

The process expects a json file containing at minimum a collection of seed urls, with potentially any desired metadata. An example of what form this json file should take can be found in the file 'seed.json'

Parameter 2: The crawler manager properties file.

This file contains the properties used to configure JQM. An example of this file can be found in crawlmanager.properties. The most important parameter for the submission service is the property jqm url which represents the url and port used to send job requests to JQM.

Clone this wiki locally