# LAB | Feature Engineering

**Load the data**

In this challenge, we will be working with the same Spaceship Titanic data, like the previous Lab. The data can be found here:

https://raw.githubusercontent.com/data-bootcamp-v4/data/main/spaceship_titanic.csv

Metadata

https://github.com/data-bootcamp-v4/data/blob/main/spaceship_titanic.md

In [1]:
#Libraries
import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split

In [2]:
spaceship = pd.read_csv("https://raw.githubusercontent.com/data-bootcamp-v4/data/main/spaceship_titanic.csv")
spaceship.head()

Unnamed: 0,PassengerId,HomePlanet,CryoSleep,Cabin,Destination,Age,VIP,RoomService,FoodCourt,ShoppingMall,Spa,VRDeck,Name,Transported
0,0001_01,Europa,False,B/0/P,TRAPPIST-1e,39.0,False,0.0,0.0,0.0,0.0,0.0,Maham Ofracculy,False
1,0002_01,Earth,False,F/0/S,TRAPPIST-1e,24.0,False,109.0,9.0,25.0,549.0,44.0,Juanna Vines,True
2,0003_01,Europa,False,A/0/S,TRAPPIST-1e,58.0,True,43.0,3576.0,0.0,6715.0,49.0,Altark Susent,False
3,0003_02,Europa,False,A/0/S,TRAPPIST-1e,33.0,False,0.0,1283.0,371.0,3329.0,193.0,Solam Susent,False
4,0004_01,Earth,False,F/1/S,TRAPPIST-1e,16.0,False,303.0,70.0,151.0,565.0,2.0,Willy Santantines,True


**Check the shape of your data**

In [3]:
spaceship.shape

(8693, 14)

**Check for data types**

In [5]:
spaceship.dtypes

PassengerId      object
HomePlanet       object
CryoSleep        object
Cabin            object
Destination      object
Age             float64
VIP              object
RoomService     float64
FoodCourt       float64
ShoppingMall    float64
Spa             float64
VRDeck          float64
Name             object
Transported        bool
dtype: object

**Check for missing values**

In [7]:
spaceship.isnull().sum()

PassengerId       0
HomePlanet      201
CryoSleep       217
Cabin           199
Destination     182
Age             179
VIP             203
RoomService     181
FoodCourt       183
ShoppingMall    208
Spa             183
VRDeck          188
Name            200
Transported       0
dtype: int64

There are multiple strategies to handle missing data

- Removing all rows or all columns containing missing data.
- Filling all missing values with a value (mean in continouos or mode in categorical for example).
- Filling all missing values with an algorithm.

For this exercise, because we have such low amount of null values, we will drop rows containing any missing value. 

In [10]:
# Handling null values
spaceship_cleaned=spaceship.dropna()


In [11]:
spaceship_cleaned.shape

(6606, 14)

- **Cabin** is too granular - transform it in order to obtain {'A', 'B', 'C', 'D', 'E', 'F', 'G', 'T'}

In [12]:
# Display the first few rows to understand the structure
print("Original Dataset:")
print(spaceship.head())

# Extract the first letter of the Cabin value and handle missing values
spaceship['Cabin'] = spaceship['Cabin'].str[0]

# Drop rows with null values
spaceship_cleaned = spaceship.dropna()

# Verify the transformation
print("Transformed Cabin Column:")
print(spaceship['Cabin'].unique())
print(spaceship.head())

Original Dataset:
  PassengerId HomePlanet CryoSleep  Cabin  Destination   Age    VIP  \
0     0001_01     Europa     False  B/0/P  TRAPPIST-1e  39.0  False   
1     0002_01      Earth     False  F/0/S  TRAPPIST-1e  24.0  False   
2     0003_01     Europa     False  A/0/S  TRAPPIST-1e  58.0   True   
3     0003_02     Europa     False  A/0/S  TRAPPIST-1e  33.0  False   
4     0004_01      Earth     False  F/1/S  TRAPPIST-1e  16.0  False   

   RoomService  FoodCourt  ShoppingMall     Spa  VRDeck               Name  \
0          0.0        0.0           0.0     0.0     0.0    Maham Ofracculy   
1        109.0        9.0          25.0   549.0    44.0       Juanna Vines   
2         43.0     3576.0           0.0  6715.0    49.0      Altark Susent   
3          0.0     1283.0         371.0  3329.0   193.0       Solam Susent   
4        303.0       70.0         151.0   565.0     2.0  Willy Santantines   

   Transported  
0        False  
1         True  
2        False  
3        False  
4

- Drop PassengerId and Name

In [13]:
# Drop PassengerId and Name columns
spaceship_cleaned = spaceship_cleaned.drop(columns=['PassengerId', 'Name'])

# Verify the transformation
print("Transformed Dataset:")
print(spaceship_cleaned.head())
print("Unique values in Cabin column:")
print(spaceship_cleaned['Cabin'].unique())

Transformed Dataset:
  HomePlanet CryoSleep Cabin  Destination   Age    VIP  RoomService  \
0     Europa     False     B  TRAPPIST-1e  39.0  False          0.0   
1      Earth     False     F  TRAPPIST-1e  24.0  False        109.0   
2     Europa     False     A  TRAPPIST-1e  58.0   True         43.0   
3     Europa     False     A  TRAPPIST-1e  33.0  False          0.0   
4      Earth     False     F  TRAPPIST-1e  16.0  False        303.0   

   FoodCourt  ShoppingMall     Spa  VRDeck  Transported  
0        0.0           0.0     0.0     0.0        False  
1        9.0          25.0   549.0    44.0         True  
2     3576.0           0.0  6715.0    49.0        False  
3     1283.0         371.0  3329.0   193.0        False  
4       70.0         151.0   565.0     2.0         True  
Unique values in Cabin column:
['B' 'F' 'A' 'G' 'E' 'C' 'D' 'T']


- For non-numerical columns, do dummies.

In [14]:
# Convert non-numerical columns to dummy variables
spaceship_dummies = pd.get_dummies(spaceship_cleaned)

# Verify the transformation
print("Transformed Dataset with Dummies:")
print(spaceship_dummies.head())
print("Columns in the transformed dataset:")
print(spaceship_dummies.columns)

Transformed Dataset with Dummies:
    Age  RoomService  FoodCourt  ShoppingMall     Spa  VRDeck  Transported  \
0  39.0          0.0        0.0           0.0     0.0     0.0        False   
1  24.0        109.0        9.0          25.0   549.0    44.0         True   
2  58.0         43.0     3576.0           0.0  6715.0    49.0        False   
3  33.0          0.0     1283.0         371.0  3329.0   193.0        False   
4  16.0        303.0       70.0         151.0   565.0     2.0         True   

   HomePlanet_Earth  HomePlanet_Europa  HomePlanet_Mars  ...  Cabin_D  \
0             False               True            False  ...    False   
1              True              False            False  ...    False   
2             False               True            False  ...    False   
3             False               True            False  ...    False   
4              True              False            False  ...    False   

   Cabin_E  Cabin_F  Cabin_G  Cabin_T  Destination_55 Canc

**Perform Train Test Split**

In [15]:
# Define features (X) and target (y)
X = spaceship_dummies.drop(columns=['Transported'])
y = spaceship_dummies['Transported']

# Perform train-test split with 80% training data and 20% testing data
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

In [18]:
X_train.head()

Unnamed: 0,Age,RoomService,FoodCourt,ShoppingMall,Spa,VRDeck,HomePlanet_Earth,HomePlanet_Europa,HomePlanet_Mars,CryoSleep_False,...,Cabin_D,Cabin_E,Cabin_F,Cabin_G,Cabin_T,Destination_55 Cancri e,Destination_PSO J318.5-22,Destination_TRAPPIST-1e,VIP_False,VIP_True
7832,25.0,0.0,1673.0,0.0,642.0,612.0,False,True,False,True,...,False,False,False,False,False,True,False,False,True,False
5842,36.0,0.0,2624.0,1657.0,2799.0,1.0,False,True,False,True,...,False,False,False,False,False,False,False,True,True,False
3928,34.0,0.0,0.0,0.0,0.0,0.0,False,False,True,False,...,False,False,True,False,False,True,False,False,True,False
4091,37.0,0.0,0.0,0.0,0.0,0.0,True,False,False,False,...,False,False,False,True,False,True,False,False,True,False
7679,22.0,0.0,0.0,0.0,0.0,0.0,True,False,False,False,...,False,False,False,True,False,False,True,False,True,False


In [19]:
y_train.head()

7832    False
5842    False
3928     True
4091     True
7679     True
Name: Transported, dtype: bool

In [20]:
X_test.head()

Unnamed: 0,Age,RoomService,FoodCourt,ShoppingMall,Spa,VRDeck,HomePlanet_Earth,HomePlanet_Europa,HomePlanet_Mars,CryoSleep_False,...,Cabin_D,Cabin_E,Cabin_F,Cabin_G,Cabin_T,Destination_55 Cancri e,Destination_PSO J318.5-22,Destination_TRAPPIST-1e,VIP_False,VIP_True
8441,29.0,0.0,0.0,0.0,0.0,0.0,False,True,False,False,...,False,False,False,False,False,False,False,True,True,False
8058,13.0,0.0,0.0,0.0,0.0,0.0,True,False,False,False,...,False,False,False,True,False,True,False,False,True,False
320,50.0,0.0,0.0,0.0,0.0,0.0,True,False,False,False,...,False,False,False,True,False,False,False,True,True,False
2548,6.0,0.0,0.0,0.0,0.0,0.0,True,False,False,True,...,False,False,False,True,False,True,False,False,True,False
8027,37.0,210.0,3.0,601.0,1.0,68.0,True,False,False,True,...,False,False,False,True,False,False,False,True,True,False


In [21]:
y_test.head()

8441     True
8058     True
320      True
2548    False
8027    False
Name: Transported, dtype: bool

**Model Selection**

In this exercise we will be using **KNN** as our predictive model.

In [25]:
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score

# Initialize the KNN model
knn_classifier = KNeighborsClassifier()

# Fit the model on the training data
knn_classifier.fit(X_train, y_train)

# Make predictions on the test data
y_pred = knn_classifier.predict(X_test)



- Evaluate your model's performance. Comment it

In [26]:
# Evaluating the model
accuracy = accuracy_score(y_test, y_pred)
precision = precision_score(y_test, y_pred)
recall = recall_score(y_test, y_pred)
f1 = f1_score(y_test, y_pred)

# Print evaluation metrics
print(f"Accuracy: {accuracy:.2f}")
print(f"Precision: {precision:.2f}")
print(f"Recall: {recall:.2f}")
print(f"F1 Score: {f1:.2f}")

Accuracy: 0.79
Precision: 0.79
Recall: 0.78
F1 Score: 0.79


In [27]:
print(f"The accuracy of the model is {knn_classifier.score(X_test, y_test)*100: .2f}%")

The accuracy of the model is  78.59%
