Repository navigation
precrec 0.16.0
This release is on GitHub only. CRAN still carries 0.14.5, and this version
is not being submitted there yet.
remotes::install_github("evalclass/precrec")It also carries the 0.15.0 changes, which were prepared but never released;
see NEWS.md for those.
Changes in 0.16.0
-
Add
metric_curve(), which takes the name of a measure for the x axis and
the name of a measure for the y axis and draws one against the other, the
wayROCR::performance()does. Every measureevalmod()can calculate is
available on both axes, under its own name or under theROCRidentifier.
It returns anssxycurves,msxycurves,smxycurvesormmxycurves
object, chosen the wayevalmod()chooses between its own, and the object
works withprint(),as.data.frame(),fortify(),plot()and
autoplot().Only two pairs are joined by a line: false positive rate against
sensitivity, which is the ROC curve, and sensitivity against precision,
which is the precision-recall curve. Those two have a defined
interpolation, and for themmetric_curve()hands the work to the same
codeevalmod(mode = "rocprc")uses, so the two cannot disagree. Every
other pair is drawn as points, because joining raw per-cutoff points with
straight lines is the error this package was written to avoid; pass
type = "l"to join them anyway.metric_curve()draws one curve per test dataset and does not average
over them. An average needs a rule for interpolating between the points of
each curve, which is what an unregistered pair does not have. -
Add the evaluation measures
ROCRprovides thatprecrecdid not:
fpr,fnr,false_discovery_rate,false_omission_rate,
predicted_positive_rate,predicted_negative_rate,lift,odds,
mi,chisqandcost. Each also answers to the identifierROCRuses
for it -fall,miss,pcfall,pcmiss,rpp,rnp,
mutual_information- and to its standard abbreviation where it has one,
so a call written againstROCRkeeps working.evalmod()gains
cost_fpandcost_fnfor the two weights thecostmeasure takes;
with the default weights of1it is the error rate.They are not calculated unless asked for.
evalmod()gains ametrics
argument that names the measures to add, or takes"all"; the default
NULLis the fourteen measures the function has always returned, so an
existing call gets the same object with the same fourteenplot()and
autoplot()panels. A measure that was not calculated cannot be plotted,
and the error says which argument asks for it.The odds ratio and the chi-square statistic are
NAat the top and the
bottom of every dataset, where the 2x2 table has an empty cell and neither
is defined.ROCRreports an infinity or aNaNthere.NAis what
precisionandnpvalready do with their own undefined end, and it
keeps an infinity off a shared axis. The mutual information is0at
those two points rather thanNA, because a cutoff that predicts one
class for everything carries no information about the labels - that value
is defined, and it is zero. -
Add
prbe(), which finds the points of a precision-recall curve at which
precision and recall are equal. It takes the objectevalmod()returns and
gives back a data frame with one row per break-even point, in the manner of
auc().ROCR::performance(pred, "prbe")interpolates linearly between adjacent
raw precision-recall points to find the crossing, which is not correct and
is the reason this package exists;prbe()reads the crossing off the
curveevalmod()has already interpolated properly. -
Add the
sarmeasure, the mean of accuracy, the AUC of the ROC curve, and
one minus the root mean squared error. Like the other added measures it is
opt-in throughevalmod(metrics = ). The RMSE reads the values of the
scores rather than their ranks, sosarwarns and returnsNAwhen the
scores are not probabilities between 0 and 1; every other measure asked for
in the same call is still returned. -
Add the root mean squared error to
prob_metrics()and
prob_metrics_ci(), as the"rmse"metric. It is the square root of the
Brier score, so each model and dataset now takes up three rows rather than
two. -
Declare
statsinImports. It was used but not listed. -
Fix the y axis of
plot()for informedness and markedness. Both run from
-1 to 1, and both were drawn on a 0 to 1 axis, which cut off the negative
half of the curve.autoplot()was never affected. The axis range of a
measure now comes from one table rather than from a list of names that had
the internal short names missing from it.