A Scala API for Cascading
Scala Java Ruby Shell Thrift
Latest commit 34cc46c Feb 22, 2017 @dieu dieu committed on GitHub Merge pull request #1644 from dieu/apanasenko/reduce_estimator_vkvs
Use JobConf input files list for input size computation used by ReducersEstimators
Failed to load latest commit information.
docs/src/main Generate Scalding microsite via sbt-microsites (#1623) Nov 19, 2016
logo Add scalding logo Oct 25, 2013
maple/src/main/java/com Maple fix for HBaseTap, Issue 781 Mar 2, 2016
project Upgrade build to sbt 0.13.13 (#1629) Dec 10, 2016
scalding-args/src Convert procedures to methods Jun 5, 2016
scalding-avro Move Scalding to 2.10/2.11 Dec 5, 2014
scalding-benchmarks/src/test/scala/com/twitter/scalding import hygiene: remove unused imports Apr 1, 2015
scalding-commons Call side-effecting functions with () Jun 5, 2016
scalding-core Use JobConf input files list for input size computation used by Reduc… Feb 23, 2017
scalding-date/src Convert procedures to methods Jun 5, 2016
scalding-db Convert procedures to methods Jun 5, 2016
scalding-hadoop-test/src Added Execution API for work with distributed cache files Feb 6, 2017
scalding-hraven refactor toFlowStepHistory Jan 14, 2016
scalding-jdbc/src Fixes Dec 5, 2014
scalding-json/src Making side-effecting functions non-nullary May 13, 2016
scalding-parquet-fixtures/src/test/resources Add partitioned parquet thrift source Aug 30, 2016
scalding-parquet-scrooge-fixtures/src/test/resources Move the scrooge fixtures into their own build targets so the local b… Aug 3, 2015
scalding-parquet-scrooge Replace Manifests with ClassTags Sep 15, 2016
scalding-parquet Clean up implicit classtags Sep 15, 2016
scalding-repl Make side-effecting functions non-nullary and call side-effecting fun… Jun 5, 2016
scalding-serialization/src Little cleanup Jun 10, 2016
scalding-thrift-macros-fixtures/src/test/resources Because, because... fun, the scala compiler has special naming rules … Mar 4, 2016
scalding-thrift-macros Kill dead code Jun 10, 2016
scripts Upgrade build to sbt 0.13.13 (#1629) Dec 10, 2016
tutorial Moving the repl in wonderland to a dedicated md file (#1614) Nov 15, 2016
.gitignore Generate Scalding microsite via sbt-microsites (#1623) Nov 19, 2016
.travis.blacklist Add scalding-parquet-fixtures to travis blacklist Feb 1, 2016
.travis.yml Generate Scalding microsite via sbt-microsites (#1623) Nov 19, 2016
CHANGES.md Remove two extra changes May 3, 2016
COMMITTERS.md Create COMMITTERS.md Aug 26, 2016
CONTRIBUTING.md Generate Scalding microsite via sbt-microsites (#1623) Nov 19, 2016
LICENSE Initial Import Jan 10, 2012
NOTICE Initial Import Jan 10, 2012
README.md Generate Scalding microsite via sbt-microsites (#1623) Nov 19, 2016
build.sbt Pick up Algebird 0.13.0 Feb 13, 2017
sbt Update Scala and sbt version (#1610) Oct 13, 2016
version.sbt Setting version to 0.16.1-SNAPSHOT Jun 6, 2016



Build status Coverage Status Latest version Chat

Scalding is a Scala library that makes it easy to specify Hadoop MapReduce jobs. Scalding is built on top of Cascading, a Java library that abstracts away low-level Hadoop details. Scalding is comparable to Pig, but offers tight integration with Scala, bringing advantages of Scala to your MapReduce jobs.

Scalding Logo

Word Count

Hadoop is a distributed system for counting words. Here is how it's done in Scalding.

package com.twitter.scalding.examples

import com.twitter.scalding._
import com.twitter.scalding.source.TypedText

class WordCountJob(args: Args) extends Job(args) {
    .flatMap { line => tokenize(line) }
    .groupBy { word => word } // use each word for a key
    .size // in each group, get the size
    .write(TypedText.tsv[(String, Long)](args("output")))

  // Split a piece of text into individual words.
  def tokenize(text: String): Array[String] = {
    // Lowercase each word and remove punctuation.
    text.toLowerCase.replaceAll("[^a-zA-Z0-9\\s]", "").split("\\s+")

Notice that the tokenize function, which is standard Scala, integrates naturally with the rest of the MapReduce job. This is a very powerful feature of Scalding. (Compare it to the use of UDFs in Pig.)

You can find more example code under examples/. If you're interested in comparing Scalding to other languages, see our Rosetta Code page, which has several MapReduce tasks in Scalding and other frameworks (e.g., Pig and Hadoop Streaming).

Documentation and Getting Started

Please feel free to use the beautiful Scalding logo artwork anywhere.


For user questions or scalding development (internals, extending, release planning): https://groups.google.com/forum/#!forum/scalding-dev (Google search also works as a first step)

In the remote possibility that there exist bugs in this code, please report them to: https://github.com/twitter/scalding/issues

Follow @Scalding on Twitter for updates.

Chat: Gitter

Get Involved + Code of Conduct

Pull requests and bug reports are always welcome!

We use a lightweight form of project governence inspired by the one used by Apache projects. Please see Contributing and Committership for our code of conduct and our pull request review process. The TL;DR is send us a pull request, iterate on the feedback + discussion, and get a +1 from a Committer in order to get your PR accepted.

The current list of active committers (who can +1 a pull request) can be found here: Committers

A list of contributors to the project can be found here: Contributors


There is a script (called sbt) in the root that loads the correct sbt version to build:

  1. ./sbt update (takes 2 minutes or more)
  2. ./sbt test
  3. ./sbt assembly (needed to make the jar used by the scald.rb script)

The test suite takes a while to run. When you're in sbt, here's a shortcut to run just one test:

> test-only com.twitter.scalding.FileSourceTest

Please refer to FAQ page if you encounter problems when using sbt.

We use Travis CI to verify the build: Build Status

We use Coveralls for code coverage results: Coverage Status

Scalding modules are available from maven central.

The current groupid and version for all modules is, respectively, "com.twitter" and 0.16.0-RC1.

Current published artifacts are

  • scalding-core_2.10
  • scalding-args_2.10
  • scalding-date_2.10
  • scalding-commons_2.10
  • scalding-avro_2.10
  • scalding-parquet_2.10
  • scalding-repl_2.10

The suffix denotes the scala version.


  • Ebay
  • Etsy
  • Sharethrough
  • Snowplow Analytics
  • Soundcloud
  • Twitter

To see a full list of users or to add yourself, see the wiki


Thanks for assistance and contributions:

A full list of contributors can be found on GitHub.


Copyright 2016 Twitter, Inc.

Licensed under the Apache License, Version 2.0