<h1>Creating, Reading and Writing</h1>

<h2>Introduction</h2>

<p>In this micro-course, you'll learn all about pandas, the most popular Python library for data analysis.</p>

<p>Along the way, you'll complete several hands-on exercises with real-world data. We recommend that you work on the exercises while reading the corresponding tutorials.</p>

<p>In this tutorial, you will learn how to create your own data, along with how to work with data that already exists.</p>

<h2>Getting started</h2>

<p>To use pandas, you'll typically start with the following line of code.</p>

In [1]:
import pandas as pd

<h2>Creating data</h2>

<p>There are two core objects in pandas: the <strong>DataFrame</strong> and the <strong>Series</strong>.</p>

<h3>DataFrame</h3>

<p>A DataFrame is a table. It contains an array of individual <em>entries</em>, each of which has a certain <em>value</em>. Each entry corresponds to a row (or <em>record</em>) and a <em>column</em>.</p>

<p>For example, consider the following simple DataFrame:</p>

In [2]:
pd.DataFrame({'Yes': [50, 21], 'No': [131, 2]})

Unnamed: 0,Yes,No
0,50,131
1,21,2


<p>In this example, the "0, No" entry has the value of 131. The "0, Yes" entry has a value of 50, and so on.</p>

<p>DataFrame entries are not limited to integers. For instance, here's a DataFrame whose values are strings:</p>

In [3]:
pd.DataFrame({'Bob': ['I liked it.', 'It was awful.'], 'Sue': ['Pretty good.', 'Bland.']})

Unnamed: 0,Bob,Sue
0,I liked it.,Pretty good.
1,It was awful.,Bland.


<p>We are using the <code>pd.DataFrame()</code> constructor to generate these DataFrame objects. The syntax for declaring a new one is a dictionary whose keys are the column names (<code>Bob</code> and <code>Sue</code> in this example), and whose values are a list of entries. This is the standard way of constructing a new DataFrame, and the one you are most likely to encounter.</p>

<p>The dictionary-list constructor assigns values to the <em>column labels</em>, but just uses an ascending count from 0 (0, 1, 2, 3, ...) for the <em>row labels</em>. Sometimes this is OK, but oftentimes we will want to assign these labels ourselves.</p>

<p>The list of row labels used in a DataFrame is known as an <strong>Index</strong>. We can assign values to it by using an <code>index</code> parameter in our constructor:</p>

In [4]:
pd.DataFrame({'Bob': ['I liked it.', 'It was awful.'],
              'Sue': ['Pretty good.', 'Bland.']},
             index=['Product A', 'Product B'])

Unnamed: 0,Bob,Sue
Product A,I liked it.,Pretty good.
Product B,It was awful.,Bland.


<h3>Series</h3>

<p>A Series, by contrast, is a sequence of data values. If a DataFrame is a table, a Series is a list. And in fact you can create one with nothing more than a list:</p>

In [5]:
pd.Series([1, 2, 3, 4, 5])

Unnamed: 0,0
0,1
1,2
2,3
3,4
4,5


<p>A Series is, in essence, a single column of a DataFrame. So you can assign row labels to the Series the same way as before, using an <code>index</code> parameter. However, a Series does not have a column name, it only has one overall <code>name</code>:</p>

In [6]:
pd.Series([30, 35, 40], index=['2015 Sales', '2016 Sales', '2017 Sales'], name='Product A')

Unnamed: 0,Product A
2015 Sales,30
2016 Sales,35
2017 Sales,40


<p>The Series and the DataFrame are intimately related. It's helpful to think of a DataFrame as actually being just a bunch of Series "glued together". We'll see more of this in the next section of this tutorial.</p>

<h1>Reading data files</h1>

<p>Being able to create a DataFrame or Series by hand is handy. But, most of the time, we won't actually be creating our own data by hand. Instead, we'll be working with data that already exists.</p>

<p>Data can be stored in any of a number of different forms and formats. By far the most basic of these is the humble CSV file. When you open a CSV file you get something that looks like this:</p>

<code>Product A,Product B,Product C,
30,21,9,
35,34,1,
41,11,11</code>

<p>So a CSV file is a table of values separated by commas. Hence the name: "Comma-Separated Values", or CSV.</p>

<p>Let's now set aside our toy datasets and see what a real dataset looks like when we read it into a DataFrame. We'll use the <code>pd.read_csv()</code> function to read the data into a DataFrame. This goes thusly:</p>

In [33]:
!curl -L -o wine-reviews.zip https://www.kaggle.com/api/v1/datasets/download/zynicide/wine-reviews

  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                                 Dload  Upload   Total   Spent    Left  Speed
  0     0    0     0    0     0      0      0 --:--:-- --:--:-- --:--:--     0
100 50.8M  100 50.8M    0     0  72.2M      0 --:--:-- --:--:-- --:--:--  165M


In [38]:
!unzip wine-reviews.zip

Archive:  wine-reviews.zip
  inflating: winemag-data-130k-v2.csv  
  inflating: winemag-data-130k-v2.json  
  inflating: winemag-data_first150k.csv  


In [39]:
wine_reviews = pd.read_csv("winemag-data-130k-v2.csv")

<p>We can use the <code>shape</code> attribute to check how large the resulting DataFrame is:</p>

In [40]:
wine_reviews.shape

(129971, 14)

<p>So our new DataFrame has 130,000 records split across 14 different columns. That's almost 2 million entries!</p>

<p>We can examine the contents of the resultant DataFrame using the <code>head()</code> command, which grabs the first five rows:</p>

In [41]:
wine_reviews.head()

Unnamed: 0.1,Unnamed: 0,country,description,designation,points,price,province,region_1,region_2,taster_name,taster_twitter_handle,title,variety,winery
0,0,Italy,"Aromas include tropical fruit, broom, brimston...",Vulkà Bianco,87,,Sicily & Sardinia,Etna,,Kerin O’Keefe,@kerinokeefe,Nicosia 2013 Vulkà Bianco (Etna),White Blend,Nicosia
1,1,Portugal,"This is ripe and fruity, a wine that is smooth...",Avidagos,87,15.0,Douro,,,Roger Voss,@vossroger,Quinta dos Avidagos 2011 Avidagos Red (Douro),Portuguese Red,Quinta dos Avidagos
2,2,US,"Tart and snappy, the flavors of lime flesh and...",,87,14.0,Oregon,Willamette Valley,Willamette Valley,Paul Gregutt,@paulgwine,Rainstorm 2013 Pinot Gris (Willamette Valley),Pinot Gris,Rainstorm
3,3,US,"Pineapple rind, lemon pith and orange blossom ...",Reserve Late Harvest,87,13.0,Michigan,Lake Michigan Shore,,Alexander Peartree,,St. Julian 2013 Reserve Late Harvest Riesling ...,Riesling,St. Julian
4,4,US,"Much like the regular bottling from 2012, this...",Vintner's Reserve Wild Child Block,87,65.0,Oregon,Willamette Valley,Willamette Valley,Paul Gregutt,@paulgwine,Sweet Cheeks 2012 Vintner's Reserve Wild Child...,Pinot Noir,Sweet Cheeks


<p>The <code>pd.read_csv()</code> function is well-endowed, with over 30 optional parameters you can specify. For example, you can see in this dataset that the CSV file has a built-in index, which pandas did not pick up on automatically. To make pandas use that column for the index (instead of creating a new one from scratch), we can specify an <code>index_col</code>.</p>

In [43]:
wine_reviews = pd.read_csv("winemag-data-130k-v2.csv", index_col=0)
wine_reviews.head()

Unnamed: 0,country,description,designation,points,price,province,region_1,region_2,taster_name,taster_twitter_handle,title,variety,winery
0,Italy,"Aromas include tropical fruit, broom, brimston...",Vulkà Bianco,87,,Sicily & Sardinia,Etna,,Kerin O’Keefe,@kerinokeefe,Nicosia 2013 Vulkà Bianco (Etna),White Blend,Nicosia
1,Portugal,"This is ripe and fruity, a wine that is smooth...",Avidagos,87,15.0,Douro,,,Roger Voss,@vossroger,Quinta dos Avidagos 2011 Avidagos Red (Douro),Portuguese Red,Quinta dos Avidagos
2,US,"Tart and snappy, the flavors of lime flesh and...",,87,14.0,Oregon,Willamette Valley,Willamette Valley,Paul Gregutt,@paulgwine,Rainstorm 2013 Pinot Gris (Willamette Valley),Pinot Gris,Rainstorm
3,US,"Pineapple rind, lemon pith and orange blossom ...",Reserve Late Harvest,87,13.0,Michigan,Lake Michigan Shore,,Alexander Peartree,,St. Julian 2013 Reserve Late Harvest Riesling ...,Riesling,St. Julian
4,US,"Much like the regular bottling from 2012, this...",Vintner's Reserve Wild Child Block,87,65.0,Oregon,Willamette Valley,Willamette Valley,Paul Gregutt,@paulgwine,Sweet Cheeks 2012 Vintner's Reserve Wild Child...,Pinot Noir,Sweet Cheeks
