Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

5 Commits
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Given an input file of newline separated urls, this will output a list of domains linked to by the given urls. Inspired by Xenu's Link Sleuth.

This is the first non-trivial concurrent program I've written, it uses a simple pub/sub model with a queue of url's being processed by a number of worker threads. It's really meant to handle gigantic inputs. The input file is slurped into the queue, and the list of processed urls and linked domains stays in memory the entire time it runs. A lot of the code is heavily influenced by dakrone's Itsy, a more general web crawler.

The number of threads used for crawling the given links is configured in the config map at the top of the source. So is delay for the snapshot thread, which runs every x milliseconds and creates two files, linked-domains and urls-processed, which contain all the domains linked to by the urls processed so far, and the actual urls that have been processed so far.

####Example If sample.com links to google.com and yahoo.com, and another domain, example.com, links to google.com and mit.edu, given an input file containing

http://sample.com
http://example.com

Then running lein run input-file will produce a filed called linked-domains containing

google.com
yahoo.com
mit.edu

and urls-processed containing

sample.com
example.com

Note: The lines in the output files are in no particular order.

About

Find all the domains that are linked to by a set of urls. Written in Clojure.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages