<a href="https://colab.research.google.com/github/paulteeter/DS-Unit-1-Sprint-2-Statistics/blob/master/Paul_Teeter_DSPT7_LS_DS_121_Statistics_Probability_Assignment.ipynb" target="_parent"><img src="https://colab.research.google.com/assets/colab-badge.svg" alt="Open In Colab"/></a>

<img align="left" src="https://lever-client-logos.s3.amazonaws.com/864372b1-534c-480e-acd5-9711f850815c-1524247202159.png" width=200>
<br></br>
<br></br>

## *Data Science Unit 1 Sprint 2 Assignment 1*

# Apply the t-test to real data

Your assignment is to determine which issues have "statistically significant" differences between political parties in this [1980s congressional voting data](https://archive.ics.uci.edu/ml/datasets/Congressional+Voting+Records). The data consists of 435 instances (one for each congressperson), a class (democrat or republican), and 16 binary attributes (yes or no for voting for or against certain issues). Be aware - there are missing values!

Your goals:

1. Load and clean the data (or determine the best method to drop observations when running tests)
2. Using hypothesis testing, find an issue that democrats support more than republicans with p < 0.01
3. Using hypothesis testing, find an issue that republicans support more than democrats with p < 0.01
4. Using hypothesis testing, find an issue where the difference between republicans and democrats has p > 0.1 (i.e. there may not be much of a difference)

Note that this data will involve *2 sample* t-tests, because you're comparing averages across two groups (republicans and democrats) rather than a single group against a null hypothesis.

In [3]:
# import and read in data
import pandas as pd
from scipy.stats import ttest_ind

column_headers = ['party','handicapped-infants','water-project',
                          'budget','physician-fee-freeze', 'el-salvador-aid',
                          'religious-groups','anti-satellite-ban',
                          'aid-to-contras','mx-missile','immigration',
                          'synfuels', 'education', 'right-to-sue','crime','duty-free',
                          'south-africa']

votes = pd.read_csv('https://archive.ics.uci.edu/ml/machine-learning-databases/voting-records/house-votes-84.data', header=None,
                    names=column_headers, na_values='?')

votes.head()


Unnamed: 0,party,handicapped-infants,water-project,budget,physician-fee-freeze,el-salvador-aid,religious-groups,anti-satellite-ban,aid-to-contras,mx-missile,immigration,synfuels,education,right-to-sue,crime,duty-free,south-africa
0,republican,n,y,n,y,y,y,n,n,n,y,,y,y,y,n,y
1,republican,n,y,n,y,y,y,n,n,n,n,n,y,y,y,n,
2,democrat,,y,y,,y,y,n,n,n,n,y,n,y,y,n,n
3,democrat,n,y,y,n,,y,n,n,n,n,y,n,y,n,n,y
4,democrat,y,y,y,n,y,y,n,n,n,n,y,,y,y,y,y


In [6]:
# Recode votes as numeric
votes = votes.replace({'y':1, 'n':0})
votes.sample(5)

Unnamed: 0,party,handicapped-infants,water-project,budget,physician-fee-freeze,el-salvador-aid,religious-groups,anti-satellite-ban,aid-to-contras,mx-missile,immigration,synfuels,education,right-to-sue,crime,duty-free,south-africa
243,democrat,0.0,1.0,1.0,0.0,0.0,0.0,1.0,1.0,,1.0,0.0,0.0,1.0,0.0,1.0,1.0
317,democrat,1.0,0.0,0.0,0.0,0.0,0.0,1.0,1.0,1.0,1.0,0.0,0.0,0.0,0.0,1.0,1.0
94,democrat,1.0,0.0,1.0,0.0,1.0,1.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,1.0
299,democrat,0.0,1.0,1.0,0.0,0.0,1.0,1.0,1.0,1.0,1.0,0.0,0.0,1.0,1.0,1.0,1.0
386,democrat,0.0,,1.0,0.0,1.0,1.0,0.0,1.0,0.0,0.0,1.0,0.0,0.0,0.0,0.0,


In [0]:
# Filter dataframes for easy calling

dems = votes[votes['party']=='democrat']
reps = votes[votes['party']=='republican']


In [35]:
# BUDGET
#check counts between dems/reps to see support levels 
print(reps['budget'].value_counts())
dems['budget'].value_counts()

# Dems-   231 - Yes   |   29 - No
# Reps-    22 - Yes   |  142 - No



0.0    142
1.0     22
Name: budget, dtype: int64


1.0    231
0.0     29
Name: budget, dtype: int64

In [11]:
# run the t-test
ttest_ind(reps['budget'], dems['budget'], nan_policy='omit')

Ttest_indResult(statistic=-23.21277691701378, pvalue=2.0703402795404463e-77)

**Budget:
The pvalue is far below the threshold of 0.01**

**_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _**

In [36]:
# IMMIGRATION
# Check support between reps/dems
print(reps['immigration'].sum()/len(reps))
print(dems['immigration'].sum()/len(dems))

# It appears rep support for this bill is higher

0.5476190476190477
0.46441947565543074


In [37]:
# run the t-test
ttest_ind(reps['immigration'], dems['immigration'], nan_policy='omit')

Ttest_indResult(statistic=1.7359117329695164, pvalue=0.08330248490425066)

**This does not meet the threshold of p-value < 0.01**

In [39]:
# RELIGIOUS GROUPS
# Check for support levels

print(reps['religious-groups'].sum()/len(reps))
print(dems['religious-groups'].sum()/len(reps))

# It also appears there is a higher republican support for this bill

0.8869047619047619
0.7321428571428571


In [40]:
# Ttest
ttest_ind(reps['religious-groups'], dems['religious-groups'], nan_policy='omit')

Ttest_indResult(statistic=9.737575825219457, pvalue=2.3936722520597287e-20)

**The p-value < 0.01**

In [8]:
# water-project
print(dems['water-project'].mean())
reps['water-project'].mean()

0.502092050209205


0.5067567567567568

In [10]:
ttest_ind(reps['water-project'], dems['water-project'], nan_policy='omit')

Ttest_indResult(statistic=0.08896538137868286, pvalue=0.9291556823993485)

In [0]:
# t stat = .08896538137868286
# pvalue = 0.9291556823993485
# There is no significant difference between Democrats and Repbublicans on this issue

# This satisfies prompt #4:
# Using hypothesis testing, find an issue where the difference between 
# republicans and democrats has p > 0.1 (i.e. there may not be much of a difference)


## Stretch Goals:

1. Refactor your code into functions so it's easy to rerun with arbitrary variables
2. Work on Performing a T-test without using Scipy in order to get "under the hood" and learn more thoroughly about this topic.
### Start with a 1-sample t-test
 - Establish the conditions for your test 
 - [Calculate the T Statistic](https://blog.minitab.com/hs-fs/hubfs/Imported_Blog_Media/701f9c0efa98a38fb397f3c3ec459b66.png?width=247&height=172&name=701f9c0efa98a38fb397f3c3ec459b66.png) (You'll need to omit NaN values from your sample).
 - Translate that t-statistic into a P-value. You can use a [table](https://www.google.com/search?q=t+statistic+table) or the [University of Iowa Applet](https://homepage.divms.uiowa.edu/~mbognar/applets/t.html)

 ### Then try a 2-sample t-test
 - Establish the conditions for your test 
 - [Calculate the T Statistic](https://lh3.googleusercontent.com/proxy/rJJ5ZOL9ZDvKOOeBihXoZDgfk7uv1YsRzSQ1Tc10RX-r2HrRpRLVqlE9CWX23csYQXcTniFwlBg3H-qR8MKJPBGnjwndqlhDX3JxoDE5Yg) (You'll need to omit NaN values from your sample).
 - Translate that t-statistic into a P-value. You can use a [table](https://www.google.com/search?q=t+statistic+table) or the [University of Iowa Applet](https://homepage.divms.uiowa.edu/~mbognar/applets/t.html)

 ### Then check your Answers using Scipy!