Python clone of Spark, a MapReduce alike framework in Python
Switch branches/tags
Nothing to show
Pull request Compare This branch is 1071 commits behind douban:master.
Fetching latest commit…
Cannot retrieve the latest commit at this time.
Permalink
Failed to load latest commit information.
dpark
examples
tests
tools
.gitignore
AUTHORS
CONTRIBUTORS
LICENSE
README
TODO
setup.py

README

Dpark is a Python clone of Spark, MapReduce computing 
framework supporting regression computation.

Word count example wc.py:

 from dpark import DparkContext
 ctx = DparkContext()
 file = ctx.textFile("/tmp/words.txt")
 words = file.flatMap(lambda x:x.split()).map(lambda x:(x,1))
 wc = words.reduceByKey(lambda x,y:x+y).collectAsMap()
 print wc

This scripts can run locally or on Mesos cluster without
any modification, just with different command arguments:

$ python wc.py
$ python wc.py -m process
$ python wc.py -m mesos

See examples/ for more examples.

Some Chinese docs: https://github.com/jackfengji/test_pro/wiki

DPark run on Mesos -r 1292597@trunk or 112ea04 of github mirror

Mailing list: dpark-users@googlegroups.com (http://groups.google.com/group/dpark-users)