Skip to content

Creating epmGrid objects

ptitle edited this page May 12, 2021 · 31 revisions

What is an epmGrid object?

The core R object in the epm package is an object of class epmGrid.

An epmGrid object is a list that contains the following components:

  • a grid made up of either hexagonal or square cells
  • a list of the unique species communities found in these gridcells
  • an indexing vector that links each unique species community to the appropriate gridcells
  • a vector of the unique species found throughout the cell communities
  • the counts of cells that make up each species' geographic range (for calculating range-weighted metrics)
  • ecological / morphological data
  • a phylogeny

epmGrid objects can be generated either from species geographic range polygons or from point occurrence records. We will demonstrate how this is done, and what decisions need to be made in choosing from the various options.


Creating an epmGrid object from range polygons

We will present the following demonstration based on mammal range polygons downloaded from IUCN, but there are of course a variety of ways you might generate or acquire geographic range polygons for a set of species.

More specifically, we will focus our examples on North American squirrels.

Preparing the range polygons for the epm package

We will handle spatial data in R using the sf package. However, if you are not as familiar with sf and are more comfortable working SpatialPolygons objects via the sp package, that's ok! An alternative script will be linked below that uses SpatialPolygons rather than sf objects.

> library(sf)
> library(epm)

> # filename and location of IUCN mammals shapefile
> IUCNfile <- "MAMMALS_TERRESTRIAL_ONLY/MAMMALS_TERRESTRIAL_ONLY.shp"

> # load the shapefile as a simple features object
> mammals <- st_read(IUCNfile, stringsAsFactors = FALSE)

Reading layer `MAMMALS_TERRESTRIAL_ONLY' from data source `MAMMALS_TERRESTRIAL_ONLY/MAMMALS_TERRESTRIAL_ONLY.shp' using driver `ESRI Shapefile'
Simple feature collection with 12483 features and 28 fields
Geometry type: MULTIPOLYGON
Dimension:     XY
Bounding box:  xmin: -179.999 ymin: -55.97946 xmax: 179.999 ymax: 83.62744
Geodetic CRS:  WGS 84

The binomial field contains the taxon names. Make sure the taxon field is of mode character and not factor:

> class(mammals$binomial)
[1] "character"

As we are focusing on North American squirrels, we will identify all the relevant species using a bit of regular expression. Here, any taxon names that contains any of the listed terms will be returned (the | implies "or").

> allsp <- unique(mammals$binomial)
> squirrelSp <- grep("Spermophilus|Citellus|Tamias|Sciurus|Glaucomys|Marmota|Cynomys\\s", allsp, value = TRUE, ignore.case = TRUE)
> head(squirrelSp)
[1] "Ammospermophilus nelsoni" "Callosciurus adamsi"      "Callosciurus albescens"  
[4] "Callosciurus baluensis"   "Callosciurus caniceps"    "Callosciurus erythraeus" 

The mammal ranges are currently all combined into a multipolygon object. We will now extract the squirrel species and store them as a list of separate species range polygons.

> spList <- vector('list', length(squirrelSp))
> names(spList) <- squirrelSp
> 
> for (i in 1:length(squirrelSp)) {
+ 	ind <- which(mammals$binomial == squirrelSp[i])
+ 	spList[[i]] <- mammals[ind,]
+ }

Although we won't go into detail here as to the how and why, polygons can potentially have geometry issues -- that is, there may be issues with the topology of the polygons, and this can have undesirable effects downstream. We will demonstrate one way to address this potential problem using the lwgeom package. Here, we will test each range polygon for problems, and attempt to repair the polygon if need be.

library(lwgeom)

> for (i in 1:length(spList)) {
+ 	if (!any(st_is_valid(spList[[i]]))) {
+ 		message('\trepairing poly ', i)
+ 		spList[[i]] <- st_make_valid(spList[[i]])
+ 	}
+ }

In this case, no issues were detected. Good!

These range polygons are unprojected, in longitude/latitude. We would like to work with equal area grid cells, so we will transform these range polygons to an equal area projection. The North America Albers Equal Area projection works well for North America, so we will use the following coordinate system definition:

> EAproj <- '+proj=aea +lat_1=20 +lat_2=60 +lat_0=40 +lon_0=-96 +x_0=0 +y_0=0 +ellps=GRS80 +datum=NAD83 +units=m +no_defs'

And we will now transform our range polygons.

> spListEA <- lapply(spList, function(x) st_transform(x, st_crs(EAproj)))

We will also make some changes to some of the taxonomy so that it matches with taxonomic and phylogenetic data later on.

names(spList) <- gsub('Neotamias', 'Tamias', names(spList))
names(spListEA) <- gsub('Neotamias', 'Tamias', names(spListEA))

At this stage, we have range polygons for 187 squirrel species and we are ready to use the epm package. Click here to see R code that does the same thing as above, but with SpatialPolygons and the sp package, rather than the sf package.


The function createEPMgrid() will take range polygons and create an epmGrid object. This involves a number of options to consider. We will list the arguments in createEPMgrid() and briefly explain what options there are to choose from. You will find further demonstrations of the different options elsewhere on this wiki.

resolution: We need to choose the size of the gridcells. The nature of the input data and the purpose of the research will dictate what a reasonable resolution should be. Importantly, resolution is in the spatial units of the input data. As we projected our range polygons to an equal area projection, units are meters. We will set the resolution to 50km (= 50000 m).

method: There are two options for determining whether a species range polygon should register in a given grid cell. With method 'centroid', a range polygon registers if it intersects the centroid coordinates of the grid cell. With method 'areaCutoff', a range polygon registers for a given cell if it covers xx% of that cell, where xx is supplied via the 'coverCutoff' input argument. More details can be found here.

cellType: We can choose between hexagonal and square grid cells. Hexagonal grid cells are preferable for several reasons:

  1. Each cell has a clear set of 6 neighboring cells, which is good for turnover metrics,
  2. hexagonal grids more can more naturally follow contours and irregular shapes and provide nice map aesthetics, and
  3. if working with unprojected data, hexagonal grid cells can be of different sizes.

On the other hand, square grid cell might be preferable if you are working at a very high spatial resolution, as raster data storage is more efficient. Square grid cells might also be preferable if you intend to work with other raster datasets, such that the grid cells align.

Some of these differences can be seen in the graphic associated with the first extent example.

retainSmallRanges: an option where if a species has a range that is smaller than a single grid cell, it might get dropped because it registers nowhere (either because it does not overlap any cell midpoint coordinates or because it occupies too little any cell's area). If this option is set to TRUE, then the species will still register in the cell that it most occupies. More details can be found here.

extent: The extent of the epmGrid object can be specified a number of ways:

  • If extent = 'auto', then the full extent of the input data will be used.
  • If extent is a polygon, then the returned object will be cropped and masked to that polygon.
  • You can also supply c(minX, maxX, minY, maxY) as the x and y ranges for cropping.
  • Finally, if extent = 'interactive', then a map will be displayed where you can define a polygon by hand. The selected polygon will be returned by the function so that you can copy/paste it as a polygon supplied to the extent argument in a subsequent call to the function. More details can be found here.

template: As an alternative to the 'resolution' and 'extent' arguments, you can provide a raster as a template, and resolution and extent will be derived from that raster. Useful if you are trying to match the spatial configuration of another dataset.

percentWithin: This is an optional filter that will calculate the percent of each species' range that is within the analysis extent, and will exclude species if the percent area is below this threshold. More details can be found here.


Link back to the table of contents.

Clone this wiki locally