# Join Statements - Lab

## Introduction

In this lab, you'll practice your knowledge of `JOIN` statements, using various types of joins and various methods for specifying the links between them.

## Objectives

You will be able to:

* Write SQL queries that make use of various types of joins
* Compare and contrast the various types of joins
* Discuss how primary and foreign keys are used in SQL
* Decide and perform whichever type of join is best for retrieving desired data

## CRM ERD

In this lab, you'll use the same customer relationship management (CRM) database that you saw from the previous lesson.
<img src='https://curriculum-content.s3.amazonaws.com/data-science/images/Database-Schema.png' width="600">

## Connecting to the Database
Import the necessary packages and connect to the database `'data.sqlite'`.

In [None]:
# Your code here

import pandas as pd
import sqlite3
conn = sqlite3.connect('data.sqlite')

## Select the names of all employees in Boston 

Hint: join the employees and offices tables. Select the first and last name.

In [None]:
# Your code here

q = '''
SELECT employees.firstName, employees.lastName
FROM employees 
JOIN offices
USING (officeCode)
WHERE city = 'Boston';
 '''
pd.read_sql(q, conn)

## Are there any offices that have zero employees?
Hint: Combine the employees and offices tables and use a group by. Select the office code, city, and number of employees.

In [None]:
# Your code here

q = '''
SELECT offices.officeCode, city, COUNT(employees.employeeNumber) AS num_employees
FROM offices
LEFT JOIN employees ON offices.officeCode = employees.officeCode
GROUP BY offices.officeCode, offices.city
HAVING COUNT(employees.employeeNumber) = 0;
'''
pd.read_sql(q, conn)

## Write 3 questions of your own and answer them

In [None]:
# Answers will vary

# Example question: 
"""
How many customers are there per office?
"""

In [None]:
"""
How many customers made orders from Boston?
"""

# Your code here

q = """
SELECT COUNT(DISTINCT orders.customerNumber) AS num_customers_who_made_orders
FROM customers
RIGHT JOIN orders ON customers.customerNumber = orders.customerNumber;
"""
pd.read_sql(q, conn)

In [None]:
"""
How many Customers made payments in Boston?
"""

# Your code here

q = """
SELECT COUNT(DISTINCT payments.customerNumber) AS num_customers_who_paid
FROM customers
LEFT JOIN payments ON customers.customerNumber = payments.customerNumber
WHERE customers.city = 'Boston';
"""
pd.read_sql(q, conn)

## Level Up 1: Display the names of every individual product that each employee has sold

Hint: You will need to use multiple `JOIN` clauses to connect all the way from employee names to product names.

In [None]:
# Your code here

q = """
SELECT e.firstName, e.lastName, p.productName
FROM employees AS e
JOIN customers AS c ON e.firstName = c.contactFirstName
JOIN orders AS o ON c.customerNumber = o.customerNumber
JOIN orderdetails AS od ON o.orderNumber = od.orderNumber
JOIN products AS p ON od.productCode = p.productCode
ORDER BY e.firstName, p.productName;
"""
pd.read_sql(q, conn)

## Level Up 2: Display the number of products each employee has sold

Alphabetize the results by employee last name.

Hint: Use the `quantityOrdered` column from `orderDetails`. Also, think about how to group the data when some employees might have the same first or last name.

In [None]:
# Your code here

q = """
SELECT e.firstName, e.lastName, SUM(od.quantityOrdered) AS num_of_products_sold
FROM employees AS e
JOIN customers AS c ON e.firstName = c.contactFirstName
JOIN orders AS o ON c.customerNumber = o.customerNumber
JOIN orderdetails AS od ON o.orderNumber = od.orderNumber
JOIN products AS p ON od.productCode = p.productCode
GROUP BY e.firstName, e.lastName
ORDER BY e.firstName, num_of_products_sold;
"""
pd.read_sql(q, conn)

## Level Up 3: Display the names employees who have sold more than 200 different products

Hint: this is different from the previous question because the quantity sold doesn't matter, only the number of different products

In [None]:
# Your code here

q = """
SELECT e.firstName, e.lastName, COUNT(DISTINCT p.productCode) AS unique_products_sold
FROM employees AS e
JOIN customers AS c ON e.firstName = c.contactFirstName
JOIN orders AS o ON c.customerNumber = o.customerNumber
JOIN orderdetails AS od ON o.orderNumber = od.orderNumber
JOIN products AS p ON od.productCode = p.productCode
GROUP BY e.firstName, e.lastName
HAVING COUNT (DISTINCT p.productCode) > 200
ORDER BY e.lastName;
"""
pd.read_sql(q, conn)

## Summary

Congrats! You practiced using join statements and leveraged your foreign keys knowledge!