Skip to content
ceteri edited this page Jul 5, 2012 · 22 revisions

Cascading for the Impatient, Part 3

In our second installment of this series we showed how to implement Word Count as a Cascading 2.0 application. If you haven't read that yet, it's probably best to start there.

Today's lesson takes the same app and stretches it even more. We'll show how to write a custom Operation. Again, this code is leading toward an implementation of TF-IDF in Cascading. We'll show best practices for workflow orchestration and test-driven development (TDD) at scale.

Source

Download source for this example on GitHub. For quick reference, the source code and a log for this example is listed in a gist. The input data stays the same as in the earlier code

TBD

Uses a custom Function to scrub the token stream Shows how to sort the output (ascending, based on token counts) Physical plan: 1 Mapper, 1 Reducer

Place those source lines all into a Main method, then build a JAR file. You should be good to go.

If you want to read in more detail about the classes in the Cascading API which were used, see the Cascading 2.0 User Guide and JavaDoc.

Build

The build for this example is based on using Gradle. The script is in build.gradle and to generate an IntelliJ project use:

gradle ideaModule

To build the sample app from the command line use:

gradle clean jar

What you should have at this point is a JAR file which is nearly ready to drop into your Maven repo -- almost. Actually, we provide a community jar repository for Cascading libraries and extensions at http://conjars.org

Run

Before running this sample app, you'll need to have a supported release of Apache Hadoop installed. Here's what was used to develop and test our example code:

$ hadoop version
Hadoop 1.0.3

Be sure to set your HADOOP_HOME environment variable. Then clear the output directory (Apache Hadoop insists, if you're running in standalone mode) and run the app:

rm -rf output
hadoop jar ./build/libs/impatient.jar data/rain.txt output/wc

Output text gets stored in the partition file output/wc which you can then verify:

more output/wc/part-00000

Here's a log file from our run of the sample app, part 2. If your run looks terribly different, something is probably not set up correctly. Drop us a line on the cascading-user email forum. Or visit one of our user group meetings. [Coming up real soon...]

Stay tuned for the next installments of our Cascading for the Impatient series.

Clone this wiki locally