-
Notifications
You must be signed in to change notification settings - Fork 0
Part 2
In our first installment of this series we showed how to create the simplest possible Cascading 2.0 application. If you haven't read that yet, it's probably best to start there.
Today's lesson takes the same app and stretches it a bit further. Undoubtedy you've seen Word Count before. We'd feel remiss at Cascading if we did not provide a Word Count example. It's the "Hello World" of MapReduce apps. Fortunately, this code is one of the basic steps toward developing a TF-IDF implementation. How convenient. We'll also show how to use Cascading to generate a visualization of your MapReduce app.
Download source for this example on GitHub. For quick reference, the source code and a log for this example is listed in a gist. The input data stays the same as in the earlier code
Note that the names of the taps have changed. Instead of inTap and outTap, we're now using docTap and wcTap. We'll be adding more taps, so this will help to have more descriptive names. Makes it simpler to follow all the plumbing.
uses a regex to split the input text lines into a token stream
generates a DOT file, to show the Cascading flow graphically
physical plan: 1 Mapper, 1 Reducer
...
Place those source lines all into a Main method, then build a JAR file. You should be good to go.
The build for this example is based on using Gradle. The script is in build.gradle and to generate an IntelliJ project use:
gradle ideaModule
To build the sample app from the command line use:
gradle clean jar
What you should have at this point is a JAR file which is nearly ready to drop into your Maven repo -- almost. Actually, we provide a community jar repository for Cascading libraries and extensions at http://conjars.org
Before running this sample app, you'll need to have a supported release of Apache Hadoop installed. Here's what was used to develop and test our example code:
$ hadoop version
Hadoop 1.0.3
Be sure to set your HADOOP_HOME environment variable. Then clear the output directory (Apache Hadoop insists, if you're running in standalone mode) and run the app:
rm -rf output
hadoop jar ./build/libs/impatient.jar data/rain.txt output/wc
Output text gets stored in the partition file output/wc which you can then verify:
more output/wc/part-00000
Again, here's a log file from our run of the sample app, part 2. If your run looks terribly different, something is probably not set up correctly. Drop us a line on the cascading-user email forum. Or visit one of our user group meetings. [Coming up real soon...]
Stay tuned for the next installments of our Cascading for the Impatient series.