# Intro to R and Jupyter Notebook

The goal of this lab is to introduce you to R and Jupyter Notebook, which you’ll be using throughout the course both to learn the statistical concepts discussed in class and also to analyze real data and come to informed conclusions. To straighten out which is which: R is the name of the programming language itself and Jupyter Notebook is a convenient interface.

As the labs progress, you are encouraged to explore beyond what the labs dictate; a willingness to experiment will make you a much better programmer. Before we get to that stage, however, you need to build some basic fluency in R. Today we begin with the fundamental building blocks of R and Jupyter Notebook: the interface, reading in data, and basic commands.



## The Notebook

Right now you are working in the Jupyter Notebook. The notebooks communicates with R through the 'kernel'. Look at the blue "R" at the top right corner of the page, this means the language of this partircular notebook is R. Jupyter supports several other languages beside R including Python and Julia. 

As you can see, the notebook is browser based (it opens a window in your browser) and works a lot like a web server. This notebook is running an R kernel, but we could choose a Python kernel, a bash kernel (unix shell) or from a long list that is currently expanding:

https://github.com/jupyter/jupyter/wiki/Jupyter-kernels

Be careful, though. Much of this is under development and considered ‘beta’ (or even alpha) - the tools can be buggy. We will be careful to use only the better developed parts of the Jupyter universe.

The notebook is comprised of ‘cells’. Cells can either be ‘code’ or ‘markdown’. Code is for writing R commands. Markdown is for text and is an extension of html. This cell is a markdown cell.

More about markdown here:

https://en.wikipedia.org/wiki/Markdown

You can type anything you want in markdown (though there are some special characters that will be interpreted as commands). Markdown cells recognize LaTeX syntax to write mathematical equations, for example

$$x^2 + y^2 = z^2$$

We are not going to cover LaTeX in this course, if you are interested in learning how to typeset documents in LaTeX you may want to start with "The not so Short Introduction to LaTeX" by Tobias Oetiker.

Code cells require R syntax. The following cell is an R cell:

In [None]:
# This is an R cell. The '#' tells R this is a comment
3+1

When I run the code cell, it executes the R code (3+1) and returns the output in an output cell. To run a code cell, you can type shift-enter, or press the run button at the top of the screen.

## R

R is a programming environment created specifically for statistics. It is a scripting language (if you don’t know what that means, don’t worry for now). R can be used interactively (as we will see in this notebook), or it can be told to execute a list of commands stored in a plain text file (called a ‘script’).

To get you started, run the command below:

In [None]:
source("datasets/arbuthnot.R")

This command instructs R to load some data from a script: the Arbuthnot baptism counts for boys and girls. As you interact with R, you will create a series of objects. Sometimes you load them as we have done here, and sometimes you create them yourself as the byproduct of a computation or some analysis you have performed. Note that because you are accessing data from a data file, this command (and the entire assignment) will work as long as such file exists. To list the objects that exist within this enviroment you can use the function `ls()`. Try it now.

In [None]:
ls()

## The Data: Dr. Arbuthnot’s Baptism Records

The Arbuthnot data set refers to Dr. John Arbuthnot, an 18th century physician, writer, and mathematician. He was interested in the ratio of newborn boys to newborn girls, so he gathered the baptism records for children born in London for every year from 1629 to 1710. We can take a look at the data by typing its name into the console.

In [None]:
arbuthnot

What you should see are four columns of numbers, each row representing a different year: the first entry in each row is simply the row number (an index we can use to access the data from individual years if we want), the second is the year, and the third and fourth are the numbers of boys and girls baptized that year, respectively. Use the scrollbar on the right side of the console window to examine the complete data set.

Note that the row numbers in the first column are not part of Arbuthnot’s data. R adds them as part of its printout to help you make visual comparisons. You can think of them as the index that you see on the left side of a spreadsheet. In fact, the comparison to a spreadsheet will generally be helpful. R has stored Arbuthnot’s data in a kind of spreadsheet or table called a data frame.

You can see the dimensions of this data frame by typing:

In [None]:
dim(arbuthnot)

This command should output 82 3, indicating that there are 82 rows and 3 columns. You can see the names of these columns (or variables) by typing:

In [None]:
names(arbuthnot)

You should see that the data frame contains the columns year, boys, and girls. At this point, you might notice that many of the commands in R look a lot like functions from math class; that is, invoking R commands means supplying a function with some number of arguments. The  dim and names commands, for example, each took a single argument, the name of a data frame.

## Some Exploration

Let’s start to examine the data a little more closely. We can access the data in a single column of a data frame separately using a command like

In [None]:
arbuthnot$boys

This command will only show the number of boys baptized each year.

**Excercise 1:** What command would you use to extract just the counts of girls baptized? Try it!

Notice that the way R has printed these data is different. When we looked at the complete data frame, we saw 82 rows, one on each line of the display. These data are no longer structured in a table with other variables, so they are displayed one right after another. Objects that print out in this way are called vectors; they represent a set of numbers.

R has some powerful functions for making graphics. We can create a simple plot of the number of girls baptized per year with the command

In [None]:
plot(x = arbuthnot$year, y = arbuthnot$girls)

By default, R creates a scatterplot with each x,y pair indicated by an open circle. Notice that the command above again looks like a function, this time with two arguments separated by a comma. The first argument in the plot function specifies the variable for the x-axis and the second for the y-axis. If we wanted to connect the data points with lines, we could add a third argument, the letter l for line.

In [None]:
plot(x = arbuthnot$year, y = arbuthnot$girls, type = "l")

You might wonder how you are supposed to know that it was possible to add that third argument. Thankfully, R documents all of its functions extensively. To read what a function does and learn the arguments that are available to you, just type in a question mark followed by the name of the function that you’re interested in. Try the following.

In [None]:
?plot

**Exercise 2:** Is there an apparent trend in the number of girls baptized over the years?
How would you describe it?

Now, suppose we want to plot the total number of baptisms. To compute this, we could use the fact that R is really just a big calculator. We can type in mathematical expressions like

In [None]:
5218 + 4683

to see the total number of baptisms in 1629. We could repeat this once for each year, but there is a faster way. If we add the vector for baptisms for boys and girls, R will compute all sums simultaneously.

In [None]:
arbuthnot$boys + arbuthnot$girls

What you will see are 82 numbers (in that packed display, because we aren’t looking at a data frame here), each one representing the sum we’re after. Take a look at a few of them and verify that they are right. Therefore, we can make a plot of the total number of baptisms per year with the command

In [None]:
plot(arbuthnot$year, arbuthnot$boys + arbuthnot$girls, type = "l")

This time, note that we left out the names of the first two arguments. We can do this because the help file shows that the default for plot is for the first argument to be the x-variable and the second argument to be the y-variable.

Similarly to how we computed the proportion of boys, we can compute the ratio of the number of boys to the number of girls baptized in 1629 with

In [None]:
5218 / 4683

or we can act on the complete vectors with the expression

In [None]:
arbuthnot$boys / arbuthnot$girls

The proportion of newborns that are boys

In [None]:
5218 / (5218 + 4683)

or this may also be computed for all years simultaneously:

In [None]:
arbuthnot$boys / (arbuthnot$boys + arbuthnot$girls)

Note that with R as with your calculator, you need to be conscious of the order of operations. Here, we want to divide the number of boys by the total number of newborns, so we have to use parentheses. Without them, R will first do the division, then the addition, giving you something that is not a proportion.

**Exercise 3:** Now, make a plot of the proportion of boys over time. What do you see?

Finally, in addition to simple mathematical operators like subtraction and division, you can ask R to make comparisons like greater than, >, less than, <, and equality, ==. For example, we can ask if boys outnumber girls in each year with the expression

In [None]:
arbuthnot$boys > arbuthnot$girls

This command returns 82 values of either TRUE if that year had more boys than girls, or FALSE if that year did not (the answer may surprise you). This output shows a different kind of data than we have considered so far. In the arbuthnot data frame our values are numerical (the year, the number of boys and girls). Here, we’ve asked R to create logical data, data where the values are either TRUE or FALSE. In general, data analysis will involve many different kinds of data types, and one reason for using R is that it is able to represent and compute with many of them.

This seems like a fair bit for your first lab, so let’s stop here. Remember to save your work either with `CTRL-s` or using the menu at the top.

## On your own

In the previous few pages, you recreated some of the displays and preliminary analysis of Arbuthnot’s baptism data. Your assignment involves repeating these steps, but for present day birth records in the United States. Load up the present day data with the following command.

In [39]:
source("datasets/present.R")

The data are stored in a data frame called present.

1. What years are included in this data set? What are the dimensions of the data frame and what are the variable or column names?

2. How do these counts compare to Arbuthnot’s? Are they on a similar scale?

3. Make a plot that displays the boy-to-girl ratio for every year in the data set. What do you see? Does Arbuthnot’s observation about boys being born in greater proportion than girls hold up in the U.S.? Include the plot in your response.

4. In what year did we see the most total number of births in the U.S.? You can refer to the help files or the R reference card [Short-refcard.pdf](Short-refcard.pdf) to find helpful commands.

These data come from a report by the Centers for Disease Control http://www.cdc.gov/nchs/data/nvsr/nvsr53/nvsr53_20.pdf. Check it out if you would like to read more about an analysis of sex ratios at birth in the United States.

That was a short introduction to R and RStudio, but we will provide you with more functions and a more complete sense of the language as the course progresses.

*This notebook is released under a Creative Commons Attribution-ShareAlike 3.0 Unported. This notebook was adapted from (1) an OpenIntro lab by Andrew Bray and Mine Çetinkaya-Rundel, which in turn was adapted from a lab written by Mark Hansen of UCLA Statistics; and (2) a notebook from the Duke NGS workshop. This version was made by Gonzalo G. Peraza Mues.*