# üè• Capstone Project - Healthcare Domain
## Diabetes Readmission Prediction

### Student: Vishnu  
### Domain: Healthcare  
### Project Objective:
To predict hospital readmission in diabetic patients using machine learning techniques by analyzing patient demographics, treatment history, and admission details.

In [None]:
import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import LabelEncoder
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import classification_report, confusion_matrix
import matplotlib.pyplot as plt
import seaborn as sns
import graphviz
import warnings
warnings.filterwarnings('ignore')

### üìä Step 1: Data Loading
We load the cleaned dataset containing records of diabetic patients including their demographics, diagnoses, treatments, and readmission status.

In [None]:
df = pd.read_csv("diabetic_data_cleaned.csv")
df.head()

### üìà Dataset Summary Diagram
- Total Rows: `df.shape[0]`
- Total Columns: `df.shape[1]`
- Sample Columns: 'race', 'age', 'gender', 'admission_type_id', 'diag_1', 'readmitted'

### üî§ Step 2: Data Preprocessing - Encoding
We encode categorical variables to numeric form using Label Encoding.

In [None]:
label_encoders = {}
for col in df.select_dtypes(include='object').columns:
    le = LabelEncoder()
    df[col] = le.fit_transform(df[col].astype(str))
    label_encoders[col] = le
df.head()

### üß± Block Diagram of Preprocessing:
```
Raw Data ‚Üí Handle Missing Values ‚Üí Encode Categorical Data ‚Üí Cleaned Numeric Dataset
```

### üß™ Step 3: Train-Test Split
Splitting the data into training and test sets for model evaluation.

In [None]:
X = df.drop(columns=['readmitted', 'readmitted_binary'])
y = df['readmitted_binary']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

### üß≠ Workflow Diagram:
```
Cleaned Dataset ‚Üí Features (X), Target (y) ‚Üí Train-Test Split
```

### üå≤ Step 4: Random Forest Classifier
We train a Random Forest model to classify whether a patient will be readmitted or not.

In [None]:
rf_model = RandomForestClassifier(n_estimators=100, random_state=42)
rf_model.fit(X_train, y_train)

### üìä Training Process Overview
```
Training Data ‚Üí Random Forest Model ‚Üí Trained Model
```

### ‚úÖ Step 5: Evaluation and Results
We evaluate the model using classification metrics and visualize the confusion matrix.

In [None]:
y_pred = rf_model.predict(X_test)
print("Classification Report:\n", classification_report(y_test, y_pred))
print("Confusion Matrix:\n", confusion_matrix(y_test, y_pred))

In [None]:
cm = confusion_matrix(y_test, y_pred)
plt.figure(figsize=(6,4))
sns.heatmap(cm, annot=True, fmt='d', cmap='Blues')
plt.title("Confusion Matrix")
plt.xlabel("Predicted")
plt.ylabel("Actual")
plt.show()

### üìå Summary Diagram:
```
Test Data ‚Üí Predict ‚Üí Compare with Actual ‚Üí Evaluation Metrics
```

### üìò Conclusion:
- Random Forest is effective in predicting diabetic patient readmissions.
- The model can support hospitals in reducing unnecessary readmissions.

**Future Enhancements:**
- Try other models (Gradient Boosting, Logistic Regression).
- Tune hyperparameters.
- Explore deeper feature engineering.