This repository has been archived by the owner. It is now read-only.
Warcbase is an open-source platform for managing analyzing web archives
Java Scala HTML JavaScript Python CSS
Clone or download
Fetching latest commit…
Cannot retrieve the latest commit at this time.
Permalink
Failed to load latest commit information.
vis
warcbase-core better io (#261) Feb 11, 2017
warcbase-hbase Cleaned up and simplified dependencies, etc. Jun 17, 2016
.gitignore Copy dependencies into solr home. Jul 13, 2015
.travis.yml Moved HBase and non-core code out of warcbase-core into warcbase-hbase. Jun 16, 2016
CONTRIBUTING.md draft of contributing file Apr 18, 2016
README.md Goodbye to Warcbase, hello to AUT! Sep 14, 2017
pom.xml Upgraded to CDH 5.7.1. Jun 24, 2016

README.md

Warcbase

Warcbase is an open-source platform for managing web archives built on Hadoop and HBase. The platform provides a flexible data model for storing and managing raw content as well as metadata and extracted knowledge. Tight integration with Hadoop provides powerful tools for analytics and data processing via Spark.

Bad news: Warcbase is defunct and no longer under active development!

Good news: In June 2017, the University of Waterloo and York University were awarded a grant from the Andrew W. Mellon Foundation to build the next generation of tools that will make historical internet content accessible to scholars. Warcbase serves as the foundation for the ArchivesUnleashed Toolkit!

If you're interested in reading about the development of Warcbase, check out this article:

Jimmy Lin, Ian Milligan, Jeremy Wiebe, and Alice Zhou. Warcbase: Scalable Analytics Infrastructure for Exploring Web Archives. ACM Journal on Computing and Cultural Heritage, 10(4), Article 22, 2017.

License

Licensed under the Apache License, Version 2.0.

Acknowledgments

This work has been supported in part by the U.S. National Science Foundation, the Natural Sciences and Engineering Research Council of Canada, the Social Sciences and Humanities Research Council of Canada, the Ontario Ministry of Research and Innovation's Early Researcher Award program, and the Mellon Foundation (via Columbia University). Any opinions, findings, and conclusions or recommendations expressed are those of the researchers and do not necessarily reflect the views of the sponsors.