Skip to content
pwendell edited this page Oct 13, 2012 · 85 revisions

Shark (Hive on Spark)

Shark is a large-scale data warehouse system for Spark designed to be compatible with Apache Hive. It can answer Hive QL queries up to 30 times faster than Hive without modification to the existing data nor queries. Shark supports Hive's query language, metastore, serialization formats, and user-defined functions.

Documentation

There is a small corpus of helpful documents on Shark. The shark-users mailing list is also very active and will be a helpful resource for beginners.

For Users:

Shark User Guide - an introduction to running Shark and its API.

Building and Deploying Shark - getting Shark up and running locally, or in a cluster.

Compatibility with Apache Hive - for those deploying Shark in existing Hive Warehouses

For Developers:

Developer Guide - for people who are interested in contributing. We don't have much here yet. Will slowly add content to it.

Startup Tasks for New Contributors

Hive Patches - documents patches we made to Hive.

Related Projects

Spark

Clone this wiki locally