# Pandas for Machine Learning Preparation

Before building a Machine Learning (ML) model, raw data must be transformed into a **model-friendly format**. Pandas plays a crucial role in this phase, providing tools for **cleaning, transforming, and organizing datasets** to maximize model accuracy. In the Titanic dataset, preparation means ensuring that features are numerical (or properly encoded), scaled, and free from inconsistencies.

**1. Feature Selection**

Not all columns are useful for prediction. For example, `PassengerId`, `Name`, and `Ticket` may not directly help predict survival, so we drop them to reduce noise.

**2. Encoding Categorical Variables**

ML algorithms can’t process text directly. Pandas makes encoding easy via `pd.get_dummies()` or `LabelEncoder` to transform columns like `Sex` and `Embarked` into numeric form.

**3. Handling Missing Data**

We fill or drop missing values in columns like `Age` and `Cabin`. Filling with the median or a placeholder avoids model bias.

**4. Feature Scaling**

Some algorithms (like Logistic Regression, SVM, and KNN) perform better when features are on a similar scale. Pandas works with Scikit-learn scalers to standardize or normalize columns like `Fare` and `Age`.

**5. Train-Test Split Preparation**

The cleaned and transformed DataFrame is split into `X` (features) and `y` (target variable, `Survived`), ready for training.

### Example with Titanic Dataset

In [1]:
import pandas as pd
from sklearn.preprocessing import MinMaxScaler

# Load dataset
df = pd.read_csv("data/train.csv")
print("Initial Data Shape:", df.shape)
print("\nFirst 5 rows of the original dataset:")
print(df.head())

# Drop irrelevant columns
df = df.drop(['PassengerId', 'Name', 'Ticket'], axis=1)
print("\nShape after dropping irrelevant columns:", df.shape)
print(df.head())

# Encode categorical variables
df = pd.get_dummies(df, columns=['Sex', 'Embarked'], drop_first=True)
print("\nColumns after one-hot encoding:")
print(df.columns)

# Fill missing values
df['Age'] = df['Age'].fillna(df['Age'].median())
df['Cabin'] = df['Cabin'].fillna('Unknown')
print("\nMissing values after imputation:")
print(df.isnull().sum())

# Feature scaling
scaler = MinMaxScaler()
df[['Age', 'Fare']] = scaler.fit_transform(df[['Age', 'Fare']])
print("\nFirst 5 rows after scaling Age and Fare:")
print(df[['Age', 'Fare']].head())

# Split into features and target
X = df.drop('Survived', axis=1)
y = df['Survived']

print("\nFeature matrix shape:", X.shape)
print("Target vector shape:", y.shape)
print("\nFirst 5 rows of features:")
print(X.head())
print("\nFirst 5 target values:")
print(y.head())


Initial Data Shape: (891, 12)

First 5 rows of the original dataset:
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   
3            4         1       1   
4            5         0       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   
3       Futrelle, Mrs. Jacques Heath (Lily May Peel)  female  35.0      1   
4                           Allen, Mr. William Henry    male  35.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  
3      0            113803  53.1000  C1

### Summary

Pandas bridges the gap between raw datasets and ML-ready data. By **selecting relevant features, encoding categorical variables, handling missing values, and scaling numerical data**, you ensure your ML algorithm receives clean, consistent input. This preparation step directly impacts model accuracy, stability, and training speed. Without it, even the most advanced algorithm can underperform.