Web crawler



A general purpose web crawler.

Usage: php index.php -u [url] [options]...

      Print debug messages.

      Print this help message.

      Ignore robots.txt and rel="nofollow" on links

      Enable spanning across hosts when doing recursive retrieving.

  -l <depth>
  --level <depth>
      Specify maximum recursion depth level.

  -o <directory>
  --output-directory <directory>
      Log retrieved data to files in a directory.

      Turn off regular output.

      Turn on recursive retrieving. The default maximum depth is 5.

  -T <seconds>
  --timeout <seconds>
      Set the network timeout to <seconds> seconds.

  --connect-timeout <seconds>
      Set the connect timeout to <seconds> seconds.

  -u <url>
  --url <url>
      Retrieve a URL.

      Turn on verbose output.

  -w <seconds>
  --wait <seconds>
      Wait the specified number of seconds between the retrievals.

Example output:

$ php index.php -u https://mozilla.org -r
200 https://www.mozilla.org/en-US/
200 https://www.mozilla.org/en-US/mission/
200 https://www.mozilla.org/en-US/about/
200 https://www.mozilla.org/en-US/products/
200 https://www.mozilla.org/en-US/contribute/

Using Docker

$ docker run --rm aliasio/aranea -u https://mozilla.org -r