Skip to content

1.1 The crawler job definition file a.k.a JobDef

jackpay edited this page Dec 10, 2019 · 3 revisions

Creating and using the JobDef

Any task or program run in JQM needs a job definition. An example of this file can be found in general-scraper-def.xml. The required properties that must be set in this file are as follows:

<path> = This is an absolute path to the program (i.e. jar file) that is run by this job. In the example general-scraper-def.xml, when this JobDef is posted to JQM the jar file acled-spring-crawler-1.2.6.jar is run.

<javaClassName>= The program's main class.

<name> = The application name is the identifier used to communicate what job/program should be run by a job request. For example, to run the program specified in the example JobDef, given in this project, it would post a Job Request for the application named SpringCollector.

Submitting JobDef and program parameters

The remaining parameters can be set in this JobDef, but can also be submitted programmatically when the Job Request is made. Which method is chosen depends on the task. For example, the example JobDef given in this project specifies the majority of parameters needed by a crawler instance. In does not however include the seed url, which is supplied when the program is called. This allows for uniform submission of multiple seeds and also a reduction in the number of parameters needing to be supplied each time a crawler is submitted.

JQM documentation

For more information refer to the JQM documentation found here https://jqm.readthedocs.io/en/jqm-all-2.1.0/.

Clone this wiki locally