dbfs:/FileStore/shared_uploads/danielgarciav010@gmail.com/caso_1/orders_data.parquet

- Preparar el fichero `orders_data.parquet` de modo que pueda ser usado para contruir un 'forecasting model'.  
- Limpiar la dataset para que cumpla los requerimientos del equipo de Data y Machine Learning.  
- Guardar el archivo actualizado (limpio) como `orders_data_clean.parquet`

  
![](https://community.cloud.databricks.com/files/EP_3/1.png)

Como ingeniero de datos de una empresa de comercio electrónico llamada Voltmart, un equipo de aprendizaje automático le ha pedido que limpie los datos que contienen información sobre los pedidos realizados el año pasado. Tienen previsto utilizar estos datos depurados para crear un modelo de previsión de la demanda (Forecasting Model). Para ello, han compartido sus requisitos sobre el formato de tabla de salida deseado.

Un analista ha compartido un archivo parquet llamado `orders_data.parquet` para que usted los limpie y los preprocese.

A continuación puede ver el esquema del conjunto de datos junto con los requisitos de limpieza de los perezosos analistas de datos:

## `orders_data.parquet`

| column | data type | description | cleaning requirements | 
|--------|-----------|-------------|-----------------------|
| `order_date` | `timestamp` | Date and time when the order was made | _Modify: Remove orders placed between 12am and 5am (inclusive); convert from timestamp to date_ |
| `time_of_day` | `string` | Period of the day when the order was made | _New column containing (lower bound inclusive, upper bound exclusive): "morning" for orders placed 5-12am, "afternoon" for orders placed 12-6pm, and "evening" for 6-12pm_ |
| `order_id` | `long` | Order ID | _N/A_ |
| `product` | `string` | Name of a product ordered | _Remove rows containing "TV" as the company has stopped selling this product; ensure all values are lowercase_ |
| `product_ean` | `double` | Product ID | _N/A_ |
| `category` | `string` | Broader category of a product | _Ensure all values are lowercase_ |
| `purchase_address` | `string` | Address line where the order was made ("House Street, City, State Zipcode") | _N/A_ |
| `purchase_state` | `string` | US State of the purchase address | _New column containing: the State that the purchase was ordered from_ |
| `quantity_ordered` | `long` | Number of product units ordered | _N/A_ |
| `price_each` | `double` | Price of a product unit | _N/A_ |
| `cost_price` | `double` | Cost of production per product unit | _N/A_ |
| `turnover` | `double` | Total amount paid for a product (quantity x price) | _N/A_ |
| `margin` | `double` | Profit made by selling a product (turnover - cost) | _N/A_ |

<br>

In [0]:
from pyspark.sql import (
    SparkSession,
    types,
    functions as F,
)

spark = (
    SparkSession
    .builder
    .appName('cleaning_orders_dataset_with_pyspark')
    .getOrCreate()
)

In [0]:
df = spark.read.parquet('dbfs:/FileStore/shared_uploads/danielgarciav010@gmail.com/caso_1/orders_data.parquet')
df.toPandas().head()

Unnamed: 0,order_date,order_id,product,product_id,category,purchase_address,quantity_ordered,price_each,cost_price,turnover,margin
0,2023-01-22 21:25:00,141234,iPhone,5638009000000.0,Vêtements,"944 Walnut St, Boston, MA 02215",1,700.0,231.0,700.0,469.0
1,2023-01-28 14:15:00,141235,Lightning Charging Cable,5563320000000.0,Alimentation,"185 Maple St, Portland, OR 97035",1,14.95,7.475,14.95,7.475
2,2023-01-17 13:33:00,141236,Wired Headphones,2113973000000.0,Vêtements,"538 Adams St, San Francisco, CA 94016",2,11.99,5.995,23.98,11.99
3,2023-01-05 20:33:00,141237,27in FHD Monitor,3069157000000.0,Sports,"738 10th St, Los Angeles, CA 90001",1,149.99,97.4935,149.99,52.4965
4,2023-01-25 11:59:00,141238,Wired Headphones,9692681000000.0,Électronique,"387 10th St, Austin, TX 73301",1,11.99,5.995,11.99,5.995


### Respuestas:

In [0]:
from pyspark.sql.functions import *

##### 1. Modify: Remove orders placed between 12am and 5am (inclusive); convert from timestamp to date 

In [0]:
dataframe = df.withColumn("time", concat(hour(df.order_date), minute(df.order_date)))
dataframe = dataframe.filter(dataframe.time > 500)
dataframe.toPandas().head()

Unnamed: 0,order_date,order_id,product,product_id,category,purchase_address,quantity_ordered,price_each,cost_price,turnover,margin,time
0,2023-01-22 21:25:00,141234,iPhone,5638009000000.0,Vêtements,"944 Walnut St, Boston, MA 02215",1,700.0,231.0,700.0,469.0,2125
1,2023-01-28 14:15:00,141235,Lightning Charging Cable,5563320000000.0,Alimentation,"185 Maple St, Portland, OR 97035",1,14.95,7.475,14.95,7.475,1415
2,2023-01-17 13:33:00,141236,Wired Headphones,2113973000000.0,Vêtements,"538 Adams St, San Francisco, CA 94016",2,11.99,5.995,23.98,11.99,1333
3,2023-01-05 20:33:00,141237,27in FHD Monitor,3069157000000.0,Sports,"738 10th St, Los Angeles, CA 90001",1,149.99,97.4935,149.99,52.4965,2033
4,2023-01-25 11:59:00,141238,Wired Headphones,9692681000000.0,Électronique,"387 10th St, Austin, TX 73301",1,11.99,5.995,11.99,5.995,1159


##### 2. New column containing (lower bound inclusive, upper bound exclusive): "morning" for orders placed 5-12am, "afternoon" for orders placed 12-6pm, and "evening" for 6-12pm

In [0]:
dataframe_2 = dataframe.withColumn("morning", ((dataframe.time > 5) & (dataframe.time < 1200)))
dataframe_2 = dataframe_2.withColumn("afternoon", ((dataframe.time > 1200) & (dataframe.time < 1800)))
dataframe_2 = dataframe_2.withColumn("evening", (dataframe.time > 1800))
dataframe_2.toPandas().head()

Unnamed: 0,order_date,order_id,product,product_id,category,purchase_address,quantity_ordered,price_each,cost_price,turnover,margin,time,morning,afternoon,evening
0,2023-01-22 21:25:00,141234,iPhone,5638009000000.0,Vêtements,"944 Walnut St, Boston, MA 02215",1,700.0,231.0,700.0,469.0,2125,False,False,True
1,2023-01-28 14:15:00,141235,Lightning Charging Cable,5563320000000.0,Alimentation,"185 Maple St, Portland, OR 97035",1,14.95,7.475,14.95,7.475,1415,False,True,False
2,2023-01-17 13:33:00,141236,Wired Headphones,2113973000000.0,Vêtements,"538 Adams St, San Francisco, CA 94016",2,11.99,5.995,23.98,11.99,1333,False,True,False
3,2023-01-05 20:33:00,141237,27in FHD Monitor,3069157000000.0,Sports,"738 10th St, Los Angeles, CA 90001",1,149.99,97.4935,149.99,52.4965,2033,False,False,True
4,2023-01-25 11:59:00,141238,Wired Headphones,9692681000000.0,Électronique,"387 10th St, Austin, TX 73301",1,11.99,5.995,11.99,5.995,1159,True,False,False


##### 3. Remove rows containing "TV" as the company has stopped selling this product; ensure all values are lowercase

In [0]:
dataframe_3 = dataframe_2.filter(~ lower(dataframe_2.product).contains("tv"))

dataframe_3.toPandas().head()

Unnamed: 0,order_date,order_id,product,product_id,category,purchase_address,quantity_ordered,price_each,cost_price,turnover,margin,time,morning,afternoon,evening
0,2023-01-22 21:25:00,141234,iPhone,5638009000000.0,Vêtements,"944 Walnut St, Boston, MA 02215",1,700.0,231.0,700.0,469.0,2125,False,False,True
1,2023-01-28 14:15:00,141235,Lightning Charging Cable,5563320000000.0,Alimentation,"185 Maple St, Portland, OR 97035",1,14.95,7.475,14.95,7.475,1415,False,True,False
2,2023-01-17 13:33:00,141236,Wired Headphones,2113973000000.0,Vêtements,"538 Adams St, San Francisco, CA 94016",2,11.99,5.995,23.98,11.99,1333,False,True,False
3,2023-01-05 20:33:00,141237,27in FHD Monitor,3069157000000.0,Sports,"738 10th St, Los Angeles, CA 90001",1,149.99,97.4935,149.99,52.4965,2033,False,False,True
4,2023-01-25 11:59:00,141238,Wired Headphones,9692681000000.0,Électronique,"387 10th St, Austin, TX 73301",1,11.99,5.995,11.99,5.995,1159,True,False,False


##### 4. Ensure all values are lowercase

In [0]:
dataframe_4 = dataframe_3.withColumn("product", lower(dataframe_3.product))
dataframe_4 = dataframe_4.withColumn("category", lower(dataframe_3.category))
dataframe_4 = dataframe_4.withColumn("purchase_address", lower(dataframe_3.purchase_address))
dataframe_4.toPandas().head()

Unnamed: 0,order_date,order_id,product,product_id,category,purchase_address,quantity_ordered,price_each,cost_price,turnover,margin,time,morning,afternoon,evening
0,2023-01-22 21:25:00,141234,iphone,5638009000000.0,vêtements,"944 walnut st, boston, ma 02215",1,700.0,231.0,700.0,469.0,2125,False,False,True
1,2023-01-28 14:15:00,141235,lightning charging cable,5563320000000.0,alimentation,"185 maple st, portland, or 97035",1,14.95,7.475,14.95,7.475,1415,False,True,False
2,2023-01-17 13:33:00,141236,wired headphones,2113973000000.0,vêtements,"538 adams st, san francisco, ca 94016",2,11.99,5.995,23.98,11.99,1333,False,True,False
3,2023-01-05 20:33:00,141237,27in fhd monitor,3069157000000.0,sports,"738 10th st, los angeles, ca 90001",1,149.99,97.4935,149.99,52.4965,2033,False,False,True
4,2023-01-25 11:59:00,141238,wired headphones,9692681000000.0,électronique,"387 10th st, austin, tx 73301",1,11.99,5.995,11.99,5.995,1159,True,False,False


##### 5. New column containing: the State that the purchase was ordered from

In [0]:
dataframe_5 = dataframe_4.withColumn("state", split(split(dataframe_4.purchase_address, ",")[2], " ")[1])
dataframe_5.toPandas().head()

Unnamed: 0,order_date,order_id,product,product_id,category,purchase_address,quantity_ordered,price_each,cost_price,turnover,margin,time,morning,afternoon,evening,state
0,2023-01-22 21:25:00,141234,iphone,5638009000000.0,vêtements,"944 walnut st, boston, ma 02215",1,700.0,231.0,700.0,469.0,2125,False,False,True,ma
1,2023-01-28 14:15:00,141235,lightning charging cable,5563320000000.0,alimentation,"185 maple st, portland, or 97035",1,14.95,7.475,14.95,7.475,1415,False,True,False,or
2,2023-01-17 13:33:00,141236,wired headphones,2113973000000.0,vêtements,"538 adams st, san francisco, ca 94016",2,11.99,5.995,23.98,11.99,1333,False,True,False,ca
3,2023-01-05 20:33:00,141237,27in fhd monitor,3069157000000.0,sports,"738 10th st, los angeles, ca 90001",1,149.99,97.4935,149.99,52.4965,2033,False,False,True,ca
4,2023-01-25 11:59:00,141238,wired headphones,9692681000000.0,électronique,"387 10th st, austin, tx 73301",1,11.99,5.995,11.99,5.995,1159,True,False,False,tx


##### 6. Guardar archivo final limpio con nombre `orders_data_clean.parquet` 

In [0]:
dataframe_5.write.parquet("dbfs:/FileStore/shared_uploads/danielgarciav010@gmail.com/caso_1/orders_data_clean.parquet")

##### 7. Exportar archivo limpio en formato CSV 

In [0]:
spark.read.parquet("dbfs:/FileStore/shared_uploads/danielgarciav010@gmail.com/caso_1/orders_data_clean.parquet").write.csv("dbfs:/FileStore/shared_uploads/danielgarciav010@gmail.com/caso_1/orders_data_clean.csv")

In [0]:
spark.read.csv("dbfs:/FileStore/shared_uploads/danielgarciav010@gmail.com/caso_1/orders_data_clean.csv").display()

_c0,_c1,_c2,_c3,_c4,_c5,_c6,_c7,_c8,_c9,_c10,_c11,_c12,_c13,_c14,_c15
2023-01-22T21:25:00.000Z,141234,iphone,5638008983335.0,vêtements,"944 walnut st, boston, ma 02215",1,700.0,231.0,700.0,469.0,2125,False,False,True,ma
2023-01-28T14:15:00.000Z,141235,lightning charging cable,5563319511488.0,alimentation,"185 maple st, portland, or 97035",1,14.95,7.475,14.95,7.475,1415,False,True,False,or
2023-01-17T13:33:00.000Z,141236,wired headphones,2113973395220.0,vêtements,"538 adams st, san francisco, ca 94016",2,11.99,5.995,23.98,11.99,1333,False,True,False,ca
2023-01-05T20:33:00.000Z,141237,27in fhd monitor,3069156759167.0,sports,"738 10th st, los angeles, ca 90001",1,149.99,97.4935,149.99,52.4965,2033,False,False,True,ca
2023-01-25T11:59:00.000Z,141238,wired headphones,9692680938163.0,électronique,"387 10th st, austin, tx 73301",1,11.99,5.995,11.99,5.995,1159,True,False,False,tx
2023-01-29T20:22:00.000Z,141239,aaa batteries (4-pack),2953868554188.0,alimentation,"775 willow st, san francisco, ca 94016",1,2.99,1.495,2.99,1.495,2022,False,False,True,ca
2023-01-26T12:16:00.000Z,141240,27in 4k gaming monitor,5173670800988.0,vêtements,"979 park st, los angeles, ca 90001",1,389.99,128.69670000000002,389.99,261.2933,1216,False,True,False,ca
2023-01-01T10:30:00.000Z,141242,bose soundsport headphones,1508418177978.0,électronique,"867 willow st, los angeles, ca 90001",1,99.99,49.995,99.99,49.995,1030,True,False,False,ca
2023-01-22T21:20:00.000Z,141243,apple airpods headphones,1386344211590.0,électronique,"657 johnson st, san francisco, ca 94016",1,150.0,97.5,150.0,52.5,2120,False,False,True,ca
2023-01-07T11:29:00.000Z,141244,apple airpods headphones,4332898830865.0,vêtements,"492 walnut st, san francisco, ca 94016",1,150.0,97.5,150.0,52.5,1129,True,False,False,ca
