<a href="https://colab.research.google.com/github/AradhanaSingh22/CodeAlphaProject/blob/Tasks/Titanic_Classification.ipynb" target="_parent"><img src="https://colab.research.google.com/assets/colab-badge.svg" alt="Open In Colab"/></a>

# **TASK 1**

# **Titanic Classification**

We'll have separate files for data loading, preprocessing, model building, evaluation, and main script.

**data_loader.py:** Load the Titanic dataset.

**data_preprocessing.py:** Preprocess the data (handle missing values, encode categorical variables).

**model.py:** Build and train the classification model.

**evaluation.py:** Evaluate the model (calculate accuracy, classification report).

**main.py:** Main script to orchestrate the entire process. italicized text

**Step 1 - data_loader.py:**

In [None]:
import pandas as pd

def load_data(filepath="/content/archive (11).zip"):
    return pd.read_csv(filepath)

**Step 2 - data_preprocessing.py:**

In [None]:
from sklearn.preprocessing import LabelEncoder

def preprocess_data(data):
    data.drop(['PassengerId', 'Name', 'Ticket', 'Cabin'], axis=1, inplace=True)
    data['Age'].fillna(data['Age'].median(), inplace=True)
    data['Embarked'].fillna(data['Embarked'].mode()[0], inplace=True)
    label_encoder = LabelEncoder()
    data['Sex'] = label_encoder.fit_transform(data['Sex'])
    data['Embarked'] = label_encoder.fit_transform(data['Embarked'])
    return data


**Step 3 - model.py:**

In [None]:
from sklearn.ensemble import RandomForestClassifier

def build_model():
    return RandomForestClassifier(n_estimators=100, random_state=42)

def train_model(model, X_train, y_train):
    model.fit(X_train, y_train)



**Step 4 - evaluation.py:**

In [None]:
from sklearn.metrics import accuracy_score, classification_report

def evaluate_model(model, X_test, y_test):
    y_pred = model.predict(X_test)
    accuracy = accuracy_score(y_test, y_pred)
    report = classification_report(y_test, y_pred)
    return accuracy, report


**Step 5 - main.py:**

In [None]:
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score, classification_report
from sklearn.preprocessing import LabelEncoder

# Load data
titanic_data = pd.read_csv("/content/archive (11).zip")

# Display the first few rows of the dataset
print(titanic_data.head())

# Preprocess data
titanic_data.drop(['PassengerId', 'Name', 'Ticket', 'Cabin'], axis=1, inplace=True)
titanic_data['Age'].fillna(titanic_data['Age'].median(), inplace=True)
titanic_data['Embarked'].fillna(titanic_data['Embarked'].mode()[0], inplace=True)
label_encoder = LabelEncoder()
titanic_data['Sex'] = label_encoder.fit_transform(titanic_data['Sex'])
titanic_data['Embarked'] = label_encoder.fit_transform(titanic_data['Embarked'])

# Split data into features and target variable
X = titanic_data.drop('Survived', axis=1)
y = titanic_data['Survived']

# Split data into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# Build and train the model
clf = RandomForestClassifier(n_estimators=100, random_state=42)
clf.fit(X_train, y_train)

# Evaluate the model
y_pred = clf.predict(X_test)
accuracy = accuracy_score(y_test, y_pred)
report = classification_report(y_test, y_pred)
print("Accuracy:", accuracy)
print("Classification Report:\n", report)

# Feature importance
feature_importance = pd.Series(clf.feature_importances_, index=X.columns)
print("Feature Importance:\n", feature_importance)



   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   
3            4         1       1   
4            5         0       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   
3       Futrelle, Mrs. Jacques Heath (Lily May Peel)  female  35.0      1   
4                           Allen, Mr. William Henry    male  35.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  
3      0            113803  53.1000  C123        S  
4      0            373450   8.0500   NaN        S  
Ac