SER Workshop: Data manipulation, visualization, and reproducible documents with R and the Tidyverse
8:30pm -12:30pm on June 18, 2019
Recent developments by the R community have revolutionized the data analysis pipeline in R, from manipulating and visualizing data to communicating results. Our workshop will provide hands-on training in tools from the tidyverse ecosystem, using real epidemiologic data. In the first section, we will teach data manipulation with dplyr, a package that makes data cleaning easy, flexible, and enjoyable. In the next section, we will teach data visualization with ggplot2, the most popular plotting package in R, with a focus on creating publication-quality plots. We will then put these tools together to make reproducible documents. Using R Markdown, we will weave code and text together and learn to write papers and reports, exported to PDF, Word, or HTML, entirely in R. This workflow easily propagates upstream changes to data or analyses throughout a document and eliminates copy and paste errors. Together, these tools form a data analysis pipeline for reproducible, publication-ready work.
- 8:30-9:30: Data manipulation using dplyr (Malcolm), Slides
- 9:30-10:30: Data visualization using ggplot2 (Corinne), Slides
- 10:30-10:45: Break
- 10:45-11:45: Data visualization team exercise (Corinne and Malcolm), Slides
- 11:45-12:30: Reproducible reports and manuscripts using R markdown (Malcolm), Slides