Skip to content

v0.3.3

Choose a tag to compare

@tsostarics tsostarics released this 12 Sep 14:03
· 68 commits to main since this release

A fair amount of refactoring done with a few breaking changes to use_contrasts, but most people won't use this function directly anyways. Overall, speed should be greatly improved since there is no longer any regular expression checking of the deparsed formulas. Details below.

  • .reset_comparison_labels no longer does manual comparisons for some cherry picked schemes. .process_contrasts will now pass the symbol passed to code_by to use_contrasts, which uses the symbol in the use_contrasts.function method since the function name is lost when doing eval(code_by). This allows .reset_comparison_labels to straightforwardly check the function name as a single string. However, this means that the function must be used to benefit from the automatically generated labels (see below). I think this is fine because the use of the function names should be encouraged, but this does mean that aliased functions can't benefit:
enlist_contrasts(mtcars, carb ~ helmert_code)
# Colnames: <2 <3 <4 <6 <8

cmat <- helmert_code(6)
enlist_contrasts(mtcars, carb ~ cmat)
# Colnames: 2 3 4 6 8


helmert <- helmert_code
enlist_contrasts(mtcars, carb ~ helmert)
# Colnames: 2 3 4 6 8
  • contr.poly and contr.helmert are the only stats-exported contrast functions that are handled by the new .reset_comparison_labels. The former gets the usual polynomial column names as before, but the latter will now denote the scaling factors in the name:
enlist_contrasts(mtcars, carb ~ contr.helmert)
# Colnames: (<2)/2  (<3)/3  (<4)/4  (<6)/5  (<8)/6
  • cumulative_split_code gains special labels now following the ordinal package's notation for threshold coefficients (since that's what inspired this function anyways)
enlist_contrasts(mtcars, carb ~ contr.helmert)
# Colnames:  1|2  2|3  3|4  4|6  6|8
  • .make_parameters has been refactored to use less helper functions. Many of the individual processing steps could be rewritten to follow the same pattern with iterative unnesting of the RHS of the expressions. As a result, many of the formula processing and validation steps could be avoided altogether or handled more elegantly while making the parameters. So, no more regular expression verification of formulas.
  • parse_drop_sequence, .omit_function_calls, .parse_formula, .check_XXX_coding, and .reinstate_dropped_trends (#34) are all obsolete and have been deleted.
  • Because there's no more regex checking of user formulas, the check for an invalid formula where the variable/function name is not the first term on the RHS will now throw an informative error suggesting to check the string (using CLI to give examples). Note that only a few cases of this error are implemented, since depending on the type or class of the object the actual R error might differ. The warnings vignette gives examples of this.
  • The handling of how the - operator is ignored when used with set_contrasts has been redone entirely so avoid regex checking of the deparsed formulas. Now, set_contrasts will apply an attribute that is read by enlist_contrasts, which (if present) will tell .process_contrasts to ignore the - operator.
  • Related to the above, but enlist_contrasts will also return the model data with any factor coercions applied if called within set_contrasts. This changes the potential return value to be either a named list of matrices to a named list containing a list of matrices and a list with a data frame. The standalone use of enlist_contrasts will only ever return the former, and the latter is only used within set_contrasts. This avoids needing to coerce columns twice, which speeds up set_contrasts a bit for large datasets.