### Fit a logistic regression model that uses income and balance to predict default.

In [1]:
df = read.csv("Default.csv")

In [2]:
set.seed(7)
fit.glm = glm(default ~ income + balance, data = df, family = "binomial")
summary(fit.glm)


Call:
glm(formula = default ~ income + balance, family = "binomial", 
    data = df)

Deviance Residuals: 
    Min       1Q   Median       3Q      Max  
-2.4725  -0.1444  -0.0574  -0.0211   3.7245  

Coefficients:
              Estimate Std. Error z value Pr(>|z|)    
(Intercept) -1.154e+01  4.348e-01 -26.545  < 2e-16 ***
income       2.081e-05  4.985e-06   4.174 2.99e-05 ***
balance      5.647e-03  2.274e-04  24.836  < 2e-16 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

(Dispersion parameter for binomial family taken to be 1)

    Null deviance: 2920.6  on 9999  degrees of freedom
Residual deviance: 1579.0  on 9997  degrees of freedom
AIC: 1585

Number of Fisher Scoring iterations: 8


### Split the sample set into a training set and a validation set.

In [3]:
train = sample(dim(df)[1], (dim(df)[1])*.5)
test = df[-train,]

### Fit a multiple logistic regression model using only the training observations.

In [4]:
fit.glm = glm(default ~ income + balance, data = test, family = "binomial")

In [5]:
summary(fit.glm)


Call:
glm(formula = default ~ income + balance, family = "binomial", 
    data = test)

Deviance Residuals: 
    Min       1Q   Median       3Q      Max  
-2.1919  -0.1349  -0.0531  -0.0188   3.7829  

Coefficients:
              Estimate Std. Error z value Pr(>|z|)    
(Intercept) -1.194e+01  6.478e-01 -18.432  < 2e-16 ***
income       2.388e-05  7.335e-06   3.256  0.00113 ** 
balance      5.766e-03  3.372e-04  17.102  < 2e-16 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

(Dispersion parameter for binomial family taken to be 1)

    Null deviance: 1382.01  on 4999  degrees of freedom
Residual deviance:  736.93  on 4997  degrees of freedom
AIC: 742.93

Number of Fisher Scoring iterations: 8


### Obtain a prediction of default status for each individual in the validation set by computing the posterior probability of default for that individual, and classifying the individual to the default category if the posterior probability is greater than 0.5.

In [6]:
glm.probs = predict(fit.glm, newdata = test, type="response")
glm.pred=rep("No",5000)
glm.pred[glm.probs>0.5] = "Yes"
mean(glm.pred != test$default)

### Repeat the process in (b) three times, using three different splits of the observations into a training set and a validation set. Comment on the results obtained.

In [7]:
for(i in 1:3) {
    train = sample(dim(df)[1], (dim(df)[1])*.5)
    test = df[-train,]
    fit.glm = glm(default ~ income + balance, data = test, family = "binomial")
    glm.probs = predict(fit.glm, newdata = test, type="response")
    glm.pred=rep("No",length(glm.probs))
    glm.pred[glm.probs>0.5] = "Yes"
    print(mean(glm.pred != test$default))
}

[1] 0.0266
[1] 0.0276
[1] 0.027


### Now consider a logistic regression model that predicts the prob ability of default using income , balance , and a dummy variable for student . Estimate the test error for this model using the validation set approach. Comment on whether or not including a dummy variable for student leads to a reduction in the test error rate.

In [8]:
for(i in 1:3) {
    train = sample(dim(df)[1], (dim(df)[1])*.5)
    test = df[-train,]
    fit.glm = glm(default ~ income + balance + student, data = test, family = "binomial")
    glm.probs = predict(fit.glm, newdata = test, type="response")
    glm.pred=rep("No",length(glm.probs))
    glm.pred[glm.probs>0.5] = "Yes"
    print(mean(glm.pred != test$default))
}

[1] 0.0262
[1] 0.0292
[1] 0.0258
