A complete, hands-on implementation of every topic from a university-level data analytics course — notes, exercises, quizzes, and mini-projects in Python, R, and Excel.
| Section | Purpose | Start Here |
|---|---|---|
| 01 Notes | Chapter-by-chapter study notes | Browse notes |
| 02 Exercises | Practice questions with worked answers, adapted from the course assignments | Open exercises |
| 03 Quiz | 310-question interactive bank with chapter and mixed practice attempts | Launch quiz hub |
| 04 Projects | 33 mini-projects and capstones in Python and R | View projects |
| 05 Toolkit | Setup guidance, dataset preparation, references, and project checklist | Open toolkit |
Read notes -> Practice exercises -> Take quiz -> Build project -> Review weak areas
- Start with the chapter notes in
01-notes. - Complete the matching set in
02-exercises. - Test recall in the Quiz Hub.
- Apply the concepts through a mini-project or capstone in
04-projects.
data-analytics/
├── 01-notes/ # 8 chapter notes files (Markdown)
├── 02-exercises/ # 8 practice sets, worked Q&A
├── 03-quiz/ # Static quiz app (index.html + quiz.html)
├── 04-projects/ # 33 project folders across 8 modules
└── 05-resources/ # Public learning toolkit; source materials stay local-only
Key concepts: Stevens' 4 measurement scales (Nominal, Ordinal, Interval, Ratio), dataset types (Record, Matrix, Transaction, Graph, Ordered), Discrete vs Continuous.
| Project | What it does | Languages |
|---|---|---|
| Attribute Auditor | Infers the scale type of each column from data characteristics | Python |
| Dataset Type Explorer | Classifies a dataset as Record/Matrix/Transaction/Graph/Ordered | Python |
| Scale Exercises | Solves the chapter exercise — identifies scale for 5 real-world attributes | Python, R |
Key concepts: Frequency tables, mode/median/mean/quartiles, IQR, MAD, standard deviation, skewness, box plots, Pearson r, Spearman ρ, covariance.
| Project | What it does | Languages |
|---|---|---|
| Frequency Table Builder | Absolute, relative, and cumulative frequency tables | Python |
| Statistics Explorer | Full location & dispersion stats with ASCII box plot | Python |
| Distribution Visualizer | Histograms, bar charts, frequency distributions | Python |
| Bivariate Quantitative | Covariance, Pearson r, Spearman ρ with step-by-step computation | Python |
| Categorical Analyzer | Contingency tables, grouped box plots, chi-square hint | Python |
| Exercise Solutions | Verified solutions for all chapter exercises (frequency table, quartiles, Pearson/Spearman) | Python, R |
Key concepts: Multivariate statistics matrices, 3D scatter, bubble charts, parallel coordinates, star/radar plots, Chernoff faces, scatter plot matrix, covariance matrix, Pearson correlation matrix, correlogram, heatmap with dendrograms, mosaic plots.
| Project | What it does | Languages |
|---|---|---|
| Multivariate Statistics | Location matrix, dispersion matrix, covariance matrix, correlation matrix | Python, R |
| Multivariate Visualization | Parallel coords, star plots, bubble charts, 3D scatter, Chernoff faces, box plots | Python, R |
| Correlation & Heatmap | Scatter plot matrix, correlogram, Pearson/Spearman comparison, heatmap | Python, R |
| Joint Frequency & Mosaic | Joint frequency tables, mosaic plots for qualitative attributes | Python, R |
Key concepts: Missing values (per-class imputation), duplicates, outlier detection (IQR), discretization (equal-width / equal-depth), one-hot encoding, gray code, thermometer code, min-max normalization, z-score standardization, Euclidean distance, log transformation, stratified sampling.
| Project | What it does | Languages |
|---|---|---|
| Data Quality Auditor | Detects and fixes missing values, duplicates, and outliers | Python, R |
| Discretization & Encoding | Equal-width/depth bins, one-hot, gray code, thermometer code | Python, R |
| Normalization & Distance | Min-max, z-score, Euclidean distance, stratified sampling | Python, R |
| Data Transformation | Log transform, absolute value, skewness reduction | Python, R |
Key concepts: Distance measures (Minkowski L1/L2, Hamming, Levenshtein), K-means, K-means++, SSE, elbow curve, hierarchical clustering (single/complete/average/Ward linkage), dendrograms, DBSCAN (core/border/noise points).
| Project | What it does | Languages |
|---|---|---|
| Distance Measures | All distance types from scratch: Minkowski, Hamming, edit distance | Python, R |
| K-Means Clustering | K-means + K-means++ + elbow curve; solves chapter exercise step-by-step | Python, R |
| Hierarchical & DBSCAN | All 4 linkage methods with ASCII dendrogram, BFS-based DBSCAN | Python, R |
Key concepts: Itemsets, support, confidence, lift, Apriori algorithm (anti-monotone pruning), FP-tree, FP-growth, maximal and closed frequent itemsets, association rules.
| Project | What it does | Languages |
|---|---|---|
| Association Rules | Apriori from scratch + rule generation sorted by lift | Python, R |
| FP-Growth | FP-tree construction, recursive mining, maximal/closed detection | Python, R |
| Pattern Mining Libraries | Practical mining with mlxtend (Python) and arules (R) | Python, R |
Key concepts: Decision trees (entropy, information gain, gain ratio, Gini), k-NN (majority vote, normalization), Naive Bayes (Gaussian, Laplace smoothing), confusion matrix, accuracy, precision, recall, F1 score, k-fold cross-validation.
| Project | What it does | Languages |
|---|---|---|
| Decision Trees | Build trees using entropy/IG/Gini; ASCII tree visualization | Python, R |
| k-NN & Naive Bayes | Distance-based k-NN, step-by-step Gaussian Naive Bayes | Python, R |
| Model Evaluation | Confusion matrix, precision/recall/F1, 5-fold cross-validation | Python, R |
Multi-chapter capstone projects that combine techniques from across the course on real-world scenarios.
| Project | Theme | Dataset | Chapters Combined |
|---|---|---|---|
| End-to-End Pipeline | Full pipeline on Friends data | 14 friends | Ch4 → Ch2/3 → Ch5 → Ch7 |
| Market Basket & Segmentation | Purchase patterns + customer clusters | 20 grocery customers | Ch2/3 + Ch4 + Ch6 + Ch5 |
| Classification Study | Compare 3 classifiers with evaluation | 20 friends | Ch2/3 + Ch4 + Ch7 |
| Iris Complete Walkthrough | Every technique applied to Iris | 150 Iris samples | Ch1 through Ch7 |
| Student Performance Predictor | Grade patterns → early warning system | 25 students | Ch1+Ch2+Ch3+Ch4+Ch5+Ch6+Ch7 |
| Retail Churn Analytics | Identify customers likely to churn | 30 retail customers | Ch1+Ch2+Ch3+Ch4+Ch5+Ch6+Ch7 |
| Medical Heart Risk Capstone | "Know Your Data" — 8-phase clinical analysis | 25 patients | All Ch1–Ch7 |
See the full dashboard (all 33 projects) in 04-projects/README.md.
A self-contained, dependency-free quiz app (03-quiz/) with 310 questions across seven chapter banks — random subsets per attempt, instant scoring, answer explanations, and a full post-quiz review.
| Feature | Details |
|---|---|
| Chapter coverage | 7 chapter quizzes + 3 mixed practices |
| Question pool | 310 total questions: 40 or 50 per chapter |
| Chapter attempts | 20 questions from 40-question banks; 25 from 50-question banks |
| Mixed attempts | 50, 75, or 100 questions drawn from the combined Chapter 1–7 bank |
| Attempt style | Randomized question subset and shuffled options |
| Access | Open the Quiz Hub |
| Language | Usage | Key Libraries |
|---|---|---|
| Python | Readable implementations of the course algorithms | Standard library for most labs; optional pandas/mlxtend comparison |
| R | Visualization, statistics, libraries | ggplot2, corrplot, pheatmap, arules, rpart, e1071, class, caret |
| Excel | Formulas and walkthroughs | COUNTIF, QUARTILE, CORREL, RANK.AVG, COVARIANCE.S |
Python files follow a consistent 8-section structure:
Section 1: Imports
Section 2: Constants
Section 3: Demo Dataset
Section 4: Helper Functions
Section 5: Core Analysis
Section 6: Printing / Reporting
Section 7: File I/O (includes standard load_csv)
Section 8: Main
- Most labs use no external libraries; the pattern-mining library comparison optionally uses pandas and mlxtend
- No lambda expressions — explicit for-loops throughout
- ASCII-only output (runs on any terminal/encoding)
- Each script can run standalone as a demo, or accept a CSV file as input
| Dataset | Description | Used in |
|---|---|---|
| Friends | 14 friends with Max_temp, Weight, Height, Years, Gender, Company | Ch02–Ch05, Ch08 |
| Iris | 150 flowers, 4 measurements, 3 species | Ch03 exercises, Ch08 |
| Friends Cuisine | 10 friends × 5 food categories (transaction data) | Ch06 |
| Friends Classification | 9 friends × Food/Distance/Company (target) | Ch07 |
| Grocery Customers | 20 customers × items purchased | Ch08 |
| Students | 25 students × 4 subject scores + demographics | Ch08 |
| Retail Customers | 30 customers × purchase behavior + churn label | Ch08 |
| Medical Patients | 25 patients × cardiovascular risk factors | Ch08 |
Run any Python project:
python path/to/script.py # runs built-in demo
python path/to/script.py data.csv # loads your own CSVRun any R project:
source("path/to/script.R") # runs in RStudio or RscriptRequired R packages (install once):
install.packages(c("ggplot2", "GGally", "corrplot", "pheatmap",
"scatterplot3d", "aplpack", "MASS",
"arules", "arulesViz",
"rpart", "rpart.plot",
"e1071", "class", "caret", "dbscan"))Study the notes. Drill the exercises. Prove it in the labs.