Repository navigation
2 Submitting multiple seeds to the crawler
This uses the CrawlerSubmissionService to submit a potentially large number of seed urls to crawl using JQM
This is achieved by running the CrawlerSubmissionService java program.
To run this program navigate to the m52BLAHJ jar with the necessary input parameters.
The parameters required to submit seeds to JQM are as follows.
- Generate the csv from which the seed list json file can be created. This will typically be some variation on the source list provided by the client (i.e. ACLED).
- The .csv file has three required fields in order for the crawling/scraping architecture to operate properly.
-
SOURCE- this field specifies the root URL or domain which is crawled and must be identical to those sources named in theacled_sourcetable in theacled_camundadatabase. -
LINK- The is the precise starting seed URL presented to a crawler. -
SCRAPER- The individual name of the scraper json file, which defines the scraping rules for the specific domain expressed inLINK.
NB: The directory containing all scrapers is specified in the crawler JobDefxml(an example of which is given in this project).
-
The python file
build_seed_list.py(provided in this repo) can be used to generate the seed json file. This script also allows the specification of a black-list (seed exclusion) or white-list (exclusive seed inclusion). -
Once generated, the seed json file is ready for the submission service. The exact details of how to submit these seeds is given below.
- Typically it is useful to set this process up as a separate screen on the server. e.g.
screen -S SEEDSUBMISSION. - The program(jar)/project containing the submission service is
m52-norconex-crawler. - The main class of the submission service is
uk.ac.susx.tag.norconex.jobqueuemanager.CrawlerSubmissionService. - The two main parameters of the submission service given below.
e.g.java -cp m52-norconex-crawler-x.x.x.jar uk.ac.susx.tag.norconex.jobqueuemanager.CrawlerSubmissionService '<SEED_JSON_LOCATION>/sample-source.json' '<PROPS_LOCATION>/crawlmanager.properties
The process expects a json file containing at minimum a collection of seed urls, with potentially any desired metadata. An example of what form this json file should take can be found in the file 'seed.json'
This file contains the properties used to configure JQM.
An example of this file can be found in crawlmanager.properties.
The most important parameter for the submission service is the property jqm url which represents the url and port used to send job requests to JQM.