Skip to content

v4.0.0

Latest

Choose a tag to compare

@1beb 1beb released this 19 Aug 18:50
· 4 commits to main since this release
2a793b6

What's Changed

  • Warn that party has no effect on predictions, and correct the docs by @soodoku in #174
  • format_legacy_data() now returns an object predict_race() can consume, and its output changes shape as a result (#175). It is keyed by state and carries state/age/sex/year alongside the county, tract, block_group and block tables, matching get_census_data(). Three defects are fixed: the population counts were labelled with 2010 SF1 names (P005003 and friends) rather than the 2020 redistricting names the 2020 code path reads; the block-group table used a blockGroup id column that no merge looks for; and the tract, block-group and block ids were cut from the wrong end of the GEOID, so no block-level row ever matched a voter file. Wrapping the result in list(ST = ...) is no longer needed.
  • A census.data object that carries legacy column names is now recognised by those names rather than by year, so an object built for the 2020 redistricting tables or by an older format_legacy_data() works whichever year is requested (#175).
  • model = "fBISG" now respects skip_bad_geos = TRUE: voter rows whose geography is absent from the census data are dropped (with a message) so the model can initialize and sample on a consistent set of rows, instead of erroring with "Some initial race values are NA" (#163).
  • Fixed use.counties = TRUE raising year must be one of "2020" or "2010", not "3": the county-filtered Census download passed arguments positionally with year omitted, so retry was read as year (#161).
  • Provided census.data built with age = TRUE and/or sex = TRUE is now used instead of being silently re-downloaded; the cache check compared the stored flags against FALSE rather than the requested age/sex (#161).
  • Surname-only predictions (surname.only = TRUE or names.to.use = "surname") no longer require the first- and middle-name dictionaries to be present; those files are read only when first/middle names are actually used (#160).
  • Added a wru_progress option to suppress the Census-download progress bar. Set options(wru_progress = FALSE) to silence the furrr progress bar (which bloats batch/non-interactive log files) while keeping the other progress messages (#150).
  • predict_race() gains a return.unmatched argument (default FALSE). When TRUE, it appends boolean last_matched / first_matched / middle_matched columns reporting whether each name was found in the dictionary. The flags use a single dictionary match for every model (BISG, fBISG, eBISG), so they are comparable across models and name selections, and are computed before imputation (#105). Re-implements the feature proposed by @mdblocker in #165, keeping the original argument and column names.
  • Fixed unmatched non-hyphenated names being silently matched to the genuine surname "Na" after the name-cleaning cascade coerced their match key to the string "NA". Such names are now correctly treated as unmatched, so impute.missing = FALSE leaves them NA and impute.missing = TRUE imputes them rather than assigning the unrelated "Na" probabilities (#162).

Full Changelog: v3.1.1...v4.0.0