GitHub - wavelets/go: The Open Source Data Science Masters

The Open-Source Data Science Masters

The open-source curriculum for learning Data Science. Foundational in both theory and technologies, the OSDSM breaks down the core competencies necessary to make data useful.

The Internet is Your Oyster

With Coursera, ebooks, Stack Overflow, and GitHub -- all free and open -- how can you afford not to take advantage of an open source education?

The Motivation

We need more Data Scientists.

...by 2018 the United States will experience a shortage of 190,000 skilled data scientists, and 1.5 million managers and analysts capable of reaping actionable insights from the big data deluge.

-- McKinsey Report Highlights the Impending Data Scientist Shortage 23 July 2013

There are little to no Data Scientists with 5 years experience, because the job simply did not exist.

-- David Hardtke How To Hire A Data Scientist 13 Nov 2012

An Academic Shortfall

Classic academic conduits aren't providing Data Scientists -- this talent gap will be closed differently.

Academic credentials are important but not necessary for high-quality data science. The core aptitudes – curiosity, intellectual agility, statistical fluency, research stamina, scientific rigor, skeptical nature – that distinguish the best data scientists are widely distributed throughout the population.

We’re likely to see more uncredentialed, inexperienced individuals try their hands at data science, bootstrapping their skills on the open-source ecosystem and using the diversity of modeling tools available. Just as data-science platforms and tools are proliferating through the magic of open source, big data’s data-scientist pool will as well.

And there’s yet another trend that will alleviate any talent gap: the democratization of data science. While I agree wholeheartedly with Raden’s statement that “the crème-de-la-crème of data scientists will fill roles in academia, technology vendors, Wall Street, research and government,” I think he’s understating the extent to which autodidacts – the self-taught, uncredentialed, data-passionate people – will come to play a significant role in many organizations’ data science initiatives.

-- James Kobielus, Closing the Talent Gap 17 Jan 2013

Ready?

The Open Source Data Science Curriculum

Start here. Intro to Data Science UW / Coursera

Topics: Python NLP on Twitter API, Distributed Computing Paradigm, MapReduce/Hadoop & Pig Script, SQL/NoSQL, Relational Algebra, Experiment design, Statistics, Graphs, Amazon EC2, Visualization.

Data Science / Harvard Video Archive & Course

Topics: Data wrangling, data management, exploratory data analysis to generate hypotheses and intuition, prediction based on statistical methods such as regression and classification, communication of results through visualization, stories, and summaries.

Data Science with Open Source Tools Book $27

Topics: Visualizing Data, Estimation, Models from Scaling Arguments, Arguments from Probability Models, What you Really Need to Know about Classical Statistics, Data Mining, Clustering, PCA, Map/Reduce, Predictive Analytics
Example Code in: R, Python, Sage, C, Gnu Scientific Library

A Note About Direction

This is an introduction geared toward those with at least a minimum understanding of programming, and (perhaps obviously) an interest in the components of Data Science (like statistics and distributed computing). Out of personal preference and need for focus, I geared the original curriculum toward Python tools and resources. R resources can be found here.

Math

[★ What are some good resources for learning about numerical analysis? / Quora ] (http://www.quora.com/What-are-some-good-resources-for-learning-about-numerical-analysis)

Linear Algebra & Programming
Linear Algebra / Levandosky Stanford / Book $10
Linear Programming (Math 407) University of Washington / Course
Statistics
Statistics I Princeton / Coursera
Stats in a Nutshell Book $29
Think Stats: Probability and Statistics for Programmers Digital & Book $25
Think Bayes Digital & Book $25
Differential Equations & Calculus
Differential Equations in Data Science Python Tutorial
Problem Solving
Problem-Solving Heuristics "How To Solve It" Polya / Book $10

Computing

Algorithms
Algorithms Design & Analysis I Stanford / Coursera
Algorithm Design, Kleinberg & Tardos Book $125
Distributed Computing Paradigms
*See Intro to Data Science UW / Lectures on MapReduce
Intro to Hadoop and MapReduce Cloudera / Udacity Course *includes select free excerpts of Hadoop: The Definitive Guide Book $29
Databases
Introduction to Databases Stanford / Online Course
SQL School Mode Analytics / Tutorials
SQL Tutorials SQLZOO / Tutorials
Data Mining
Mining Massive Data Sets Stanford / Digital & Book $58
Mining The Social Web Book $30
Introduction to Information Retrieval / Stanford Digital & Book $56

OSDSM Specialization: Web Scraping & Crawling

Machine Learning
Machine Learning Ng Stanford / Coursera
A Course in Machine Learning UMD / Digital Book
Machine Learning Caltech / Edx
Programming Collective Intelligence Book $27
The Elements of Statistical Learning / Stanford Digital^ & Book $80
Machine Learning for Hackers - Python port ipynb / digital book
Statistical Network Analysis & Modeling
Probabilistic Programming and Bayesian Methods for Hackers Github / Tutorials
Probabalistic Graphical Models Stanford / Coursera
Neural Networks U Toronto / Coursera
Network & Graph Analysis
Social and Economic Networks: Models and Analysis / Stanford / Coursera
Social Network Analysis for Startups Book $22
Natural Language Processing
NLP with Python (NLTK library) Digital, Book $36
Analysis
Python for Data Analysis Paper Book $24
Big Data Analysis with Twitter UC Berkeley / Lectures
Exploratory Data Analysis Tukey / Book $81
Visualization
Envisioning Information Tufte / Book $36
The Visual Display of Quantitative Information Tufte / Book $27
Data Visualization, CS 171 Harvard / Lectures
Data Visualization, CSE512 University of Washington / Slides
Scott Murray's Tutorial on D3 Blog / Tutorials
Berkely's Viz Class UC Berkeley / Course Docs
Rice University's Data Viz class Rice University

OSDSM Specialization: Data Journalism

Python (Learning)

Learn Python the Hard Way Digital & Book $23
Python Class / Google
Think Python Digital & Book $34
Introduction to Computer Science and Programming MIT OpenCourseWare / Lectures

Python (Libraries)

Installing Basic Packages Python, virtualenv, NumPy, SciPy, matplotlib and IPython & Using Python Scientifically

More Libraries can be found in related specialiaztions

Data Structures & Analysis Packages
- Flexible and powerful data analysis / manipulation library with labeled data structures objects, statistical functions, etc pandas & Tutorials Python for Data Analysis / Book
Machine Learning Packages
- scikit-learn - Tools for Data Mining & Analysis
Networks Packages
- networkx - Network Modeling & Viz
Statistical Packages
- PyMC - Bayesian Inference & Markov Chain Monte Carlo sampling toolkit
- Statsmodels - Python module that allows users to explore data, estimate statistical models, and perform statistical tests
- PyMVPA - Multivariate Pattern Analysis in Python
Natural Language Processing & Understanding
- NLTK - Natural Language Toolkit
- Gensim - Python library for topic modelling, document indexing and similarity retrieval with large corpora. Target audience is the natural language processing (NLP) and information retrieval (IR) community.
Live Data Packages
- twython - Python wrapper for the Twitter API
Visualization Packages
- Orange - Open source data visualization and analysis for novice and experts. Data mining through visual programming or Python scripting. Components for machine learning. Add-ons for bioinformatics and text mining
iPython Data Science Notebooks
Data Science in IPython Notebooks (Linear Regression, Logistic Regression, Random Forests, K-Means Clustering)
A Gallery of Interesting IPython Notebooks - Pandas for Data Analysis

Datasets are now here

R resources are now here

Data Science as a Profession

Doing Data Science: Straight Talk from the Frontline O'Reilly / Book $25

Capstone Project

Capstone Analysis of Your Own Design; Quora's Idea Compendium
Healthcare Twitter Analysis Coursolve & UW Data Science

Resources

DataTau - The "Hacker News" of Data Science
Metacademy - Search for a concept you want to learn
Coursera - Online university courses
Wolfram Alpha - The smart number and info cruncher
Khan Academy - High quality, free learning videos
Wikipedia - The free encyclopedia
The Signal and The Noise - Nate Silver Pop-Sci Data Analysis $15
Zipfian Academy's List of Resources
A Software Engineer's Guide to Getting Started with Data Science
Data Scientist Interviews Metamarkets
/r/MachineLearning Reddit

Notation

Paid books, courses, and resources are noted with $.

Contribute

Please Contribute Your Ideas -- this is Open Source!

Please showcase your own specialization & transcript by submitting a markdown file pull request in the /transcripts directory with your name! eg clare-corthell-2014.md

Follow me on Twitter @clarecorthell

Name		Name	Last commit message	Last commit date
Latest commit History 174 Commits
transcripts		transcripts
LICENSE.md		LICENSE.md
README.md		README.md
analysis-technologies.md		analysis-technologies.md
basic-programming.md		basic-programming.md
blogs-n-media.md		blogs-n-media.md
database-tech.md		database-tech.md
datasets.md		datasets.md
machine-learning.md		machine-learning.md
r-resources.md		r-resources.md
specializations.md		specializations.md

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Repository files navigation

The Open-Source Data Science Masters

The Internet is Your Oyster

The Motivation

An Academic Shortfall

Ready?

The Open Source Data Science Curriculum

A Note About Direction

Math

Computing

Python (Learning)

Python (Libraries)

Datasets are now here

R resources are now here

Data Science as a Profession

Capstone Project

Resources

Notation

Contribute

About

Uh oh!

Releases

Packages

License

wavelets/go

Folders and files

Latest commit

History

Repository files navigation

The Open-Source Data Science Masters

The Internet is Your Oyster

The Motivation

An Academic Shortfall

Ready?

The Open Source Data Science Curriculum

A Note About Direction

Math

Computing

Python (Learning)

Python (Libraries)

Datasets are now here

R resources are now here

Data Science as a Profession

Capstone Project

Resources

Notation

Contribute

About

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Packages