Recent developments by the R community have revolutionized the data analysis pipeline in R, from manipulating and visualizing data to communicating results. Our workshop will provide hands-on training in tools from the tidyverse ecosystem, using real epidemiologic data. In the first section, we will teach data manipulation with dplyr, a package that makes data cleaning easy, flexible, and enjoyable. In the next section, we will teach data visualization with ggplot2, the most popular plotting package in R, with a focus on creating publication-quality plots. We will then put these tools together to make reproducible documents. Using R Markdown, we will weave code and text together and learn to write papers and reports, exported to PDF, Word, or HTML, entirely in R. This workflow easily propagates upstream changes to data or analyses throughout a document and eliminates copy and paste errors. Together, these tools form a data analysis pipeline for reproducible, publication-ready work.
- Make sure you have R and RStudio installed (see the introductory PDF). Additionally, make sure you have the tidyverse and gapminder packages installed:
- Set up an account at RStudio Cloud. We'll have a free cloud version R and RStudio available with all of the workshop materials included in case you have any issues with your local installation!
- Download these materials locally using the usethis package:
usethis will open the files right in RStudio. To open it in the future, open the
SER-tidyverse-workshop.Rproj file by double-clicking it or, in RStudio, navigate to File > Open Project.
- 8:50-9:00: Zoom room open
- 9:00-10:00: Data manipulation using dplyr (Malcolm), Slides
- 10:00-11:00: Data visualization using ggplot2 (Corinne), Slides
- 11:00-11:15: Break
- 11:15-12:15: Data visualization team exercise (Corinne and Malcolm), Slides
- 12:15-1:00: Reproducible reports and manuscripts using R markdown (Malcolm), Slides