# Heart Disease Prediction Using Logistic Regression
The code implements logistic regression for predicting the presence of heart disease based on various medical parameters. After importing necessary libraries, the heart disease dataset is loaded and explored, revealing details such as data shape, missing values, and target value distribution. The data is split into input features (X) and target labels (Y), followed by train-test splitting using the train_test_split() method. A logistic regression model is trained on the training data, and its accuracy is evaluated on both training and testing sets using the accuracy_score() metric. Finally, the model is used to predict the presence of heart disease for a sample input data point, demonstrating its potential for medical diagnosis.

Importing libraries

In [49]:
import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score

Data Collection and Processing

[Dataset Link](https://www.kaggle.com/datasets/johnsmith88/heart-disease-dataset/data)

In [50]:
# Loading csv
heart_data = pd.read_csv('/content/heart.csv')

In [51]:
# First rows of dataset
heart_data.head()

Unnamed: 0,age,sex,cp,trestbps,chol,fbs,restecg,thalach,exang,oldpeak,slope,ca,thal,target
0,52,1,0,125,212,0,1,168,0,1.0,2,2,3,0
1,53,1,0,140,203,1,0,155,1,3.1,0,0,3,0
2,70,1,0,145,174,0,1,125,1,2.6,0,0,3,0
3,61,1,0,148,203,0,1,161,0,0.0,2,1,3,0
4,62,0,0,138,294,1,1,106,0,1.9,1,3,2,0


In [52]:
# Last rows of dataset
heart_data.tail()

Unnamed: 0,age,sex,cp,trestbps,chol,fbs,restecg,thalach,exang,oldpeak,slope,ca,thal,target
1020,59,1,1,140,221,0,1,164,1,0.0,2,0,2,1
1021,60,1,0,125,258,0,0,141,1,2.8,1,1,3,0
1022,47,1,0,110,275,0,0,118,1,1.0,1,1,2,0
1023,50,0,0,110,254,0,0,159,0,0.0,2,0,2,1
1024,54,1,0,120,188,0,1,113,0,1.4,1,1,3,0


In [53]:
# shape of data
heart_data.shape

(1025, 14)

In [54]:
# Getting info of data
heart_data.info()

<class 'pandas.core.frame.DataFrame'>
RangeIndex: 1025 entries, 0 to 1024
Data columns (total 14 columns):
 #   Column    Non-Null Count  Dtype  
---  ------    --------------  -----  
 0   age       1025 non-null   int64  
 1   sex       1025 non-null   int64  
 2   cp        1025 non-null   int64  
 3   trestbps  1025 non-null   int64  
 4   chol      1025 non-null   int64  
 5   fbs       1025 non-null   int64  
 6   restecg   1025 non-null   int64  
 7   thalach   1025 non-null   int64  
 8   exang     1025 non-null   int64  
 9   oldpeak   1025 non-null   float64
 10  slope     1025 non-null   int64  
 11  ca        1025 non-null   int64  
 12  thal      1025 non-null   int64  
 13  target    1025 non-null   int64  
dtypes: float64(1), int64(13)
memory usage: 112.2 KB


In [55]:
# Checkng for missing values
heart_data.isnull().sum()

age         0
sex         0
cp          0
trestbps    0
chol        0
fbs         0
restecg     0
thalach     0
exang       0
oldpeak     0
slope       0
ca          0
thal        0
target      0
dtype: int64

In [56]:
# Description of data
heart_data.describe()

Unnamed: 0,age,sex,cp,trestbps,chol,fbs,restecg,thalach,exang,oldpeak,slope,ca,thal,target
count,1025.0,1025.0,1025.0,1025.0,1025.0,1025.0,1025.0,1025.0,1025.0,1025.0,1025.0,1025.0,1025.0,1025.0
mean,54.434146,0.69561,0.942439,131.611707,246.0,0.149268,0.529756,149.114146,0.336585,1.071512,1.385366,0.754146,2.323902,0.513171
std,9.07229,0.460373,1.029641,17.516718,51.59251,0.356527,0.527878,23.005724,0.472772,1.175053,0.617755,1.030798,0.62066,0.50007
min,29.0,0.0,0.0,94.0,126.0,0.0,0.0,71.0,0.0,0.0,0.0,0.0,0.0,0.0
25%,48.0,0.0,0.0,120.0,211.0,0.0,0.0,132.0,0.0,0.0,1.0,0.0,2.0,0.0
50%,56.0,1.0,1.0,130.0,240.0,0.0,1.0,152.0,0.0,0.8,1.0,0.0,2.0,1.0
75%,61.0,1.0,2.0,140.0,275.0,0.0,1.0,166.0,1.0,1.8,2.0,1.0,3.0,1.0
max,77.0,1.0,3.0,200.0,564.0,1.0,2.0,202.0,1.0,6.2,2.0,4.0,3.0,1.0


In [57]:
heart_data['target'].value_counts()

target
1    526
0    499
Name: count, dtype: int64

Target value <br>
 0 = no disease <br>
 1 = disease

In [58]:
# Splitting of data into target and input columns
X = heart_data.drop(columns='target', axis=1)
Y = heart_data['target']

In [59]:
print(X)

      age  sex  cp  trestbps  chol  fbs  restecg  thalach  exang  oldpeak  \
0      52    1   0       125   212    0        1      168      0      1.0   
1      53    1   0       140   203    1        0      155      1      3.1   
2      70    1   0       145   174    0        1      125      1      2.6   
3      61    1   0       148   203    0        1      161      0      0.0   
4      62    0   0       138   294    1        1      106      0      1.9   
...   ...  ...  ..       ...   ...  ...      ...      ...    ...      ...   
1020   59    1   1       140   221    0        1      164      1      0.0   
1021   60    1   0       125   258    0        0      141      1      2.8   
1022   47    1   0       110   275    0        0      118      1      1.0   
1023   50    0   0       110   254    0        0      159      0      0.0   
1024   54    1   0       120   188    0        1      113      0      1.4   

      slope  ca  thal  
0         2   2     3  
1         0   0     3  
2  

In [60]:
print(Y)

0       0
1       0
2       0
3       0
4       0
       ..
1020    1
1021    0
1022    0
1023    1
1024    0
Name: target, Length: 1025, dtype: int64


In [61]:
print("Shape of Training data : ", X.shape, "\nShape of testing data : ",Y.shape)

Shape of Training data :  (1025, 13) 
Shape of testing data :  (1025,)


Train Test Split

In [62]:
X_train, X_test, Y_train, Y_test = train_test_split(X,Y, test_size=0.2, stratify=Y, random_state=2)

In [63]:
print(X_train.shape, X_test.shape, Y_train.shape, Y_test.shape)

(820, 13) (205, 13) (820,) (205,)


### Model Training

In [64]:
# Logistic Regression
model = LogisticRegression()

In [65]:
model.fit(X_train, Y_train)

STOP: TOTAL NO. of ITERATIONS REACHED LIMIT.

Increase the number of iterations (max_iter) or scale the data as shown in:
    https://scikit-learn.org/stable/modules/preprocessing.html
Please also refer to the documentation for alternative solver options:
    https://scikit-learn.org/stable/modules/linear_model.html#logistic-regression
  n_iter_i = _check_optimize_result(


Model Evaluation

In [66]:
# Accuracy on training data
X_train_prediction = model.predict(X_train)
X_train_prediction_score = accuracy_score(X_train_prediction, Y_train)
print("The accuracy of model on traing data is : ", X_train_prediction_score)

The accuracy of model on traing data is :  0.8524390243902439


In [67]:
# Accuracy on testing data
X_test_prediction = model.predict(X_test)
X_test_prediction_score = accuracy_score(X_test_prediction, Y_test)
print("The accuracy of model on test data is : ", X_test_prediction_score)

The accuracy of model on test data is :  0.8048780487804879


Building a prediction system

In [71]:
input_data = (54,1,0,122,286,0,0,116,1,3.2,1,2,2)

# Converting into nparray
input_data_nparray = np.asarray(input_data)

# Reshapeing the array
input_data_reshape = input_data_nparray.reshape(1,-1)

input_prediction = model.predict(input_data_reshape)

print(input_prediction)

if input_prediction:
  print("Heart Disease")
else:
  print("No Heart Disease")

[0]
No Heart Disease


