Skip to content

Release v0.3.0: Classification-only metrics

Latest

Choose a tag to compare

@dfuchss dfuchss released this 22 Sep 13:41
· 1 commit to main since this release

ARDoCo Metrics is now a classification-only metrics library. This is a deliberately breaking release with no deprecated shims — downstream ArDoCo code needs updating, and the ArDoCo parent still pins metrics.version=0.2.1, so that wants a bump now that this is out.

💥 Removed: rank metrics

RankMetricsCalculator, its calculation functions and result types, the rank/aggRnk CLI subcommands, the /rank-metrics endpoints and the wiki page are all gone.

They were not merely untested — they were wrong. calculateWeightedAverage never divided auc by sumOfWeights, so the aggregated AUC was a weighted sum, not an average; the same loop branched on all { it.auc == null } while summing per element, so a list where only some results had an AUC produced a partial sum over a full denominator. Repairing that math with zero test coverage and no consumers was worse than deleting it.

🧱 Aggregations now keep their data

The old calculateMicroAverage summed TP/FP/FN/TN and then threw the counts away into a SingleClassificationResult<Nothing> with empty element sets. AggregatedClassificationResult had no confusion matrix at all, weights was null for micro but not otherwise, and all three aggregations each carried their own copy of the inputs — so an /average response repeated its inputs three times.

The types are now split by who owns what:

  • ClassificationResult carries a ConfusionMatrix. For a single result and the micro average, the metrics are exactly the metrics of that matrix; for macro and weighted average, which are means over the single results, it is the pooled matrix and describes the underlying data rather than the origin of the values. A test pins which of the two holds where.
  • AggregatedClassificationResult is reduced to the aggregated values plus the pooled matrix.
  • ClassificationAggregationResult (new) holds the single results and the weights once, and derives the rest on demand: the pooled confusionMatrix, the element unions (truePositives() / falsePositives() / falseNegatives()) and the distribution of a metric across inputs (spread(metric), fbetaSpread(beta)).

asList() and the union/spread accessors are functions rather than properties on purpose, so they don't end up duplicated in the serialized JSON.

🔀 calculateAverages returns an object

// before — order and cardinality were an undocumented contract
val macro = calculator.calculateAverages(results).first { it.type == MACRO_AVERAGE }

// after
val aggregation = calculator.calculateAverages(results)
aggregation.macroAverage.f1
aggregation[AggregationType.MICRO_AVERAGE].recall

The REST /average response is keyed by macroAverage / weightedAverage / microAverage instead of being a classificationResults array.

📐 F-beta everywhere

Every result carries fbetaScores: Map<Double, Double> next to f1, with the betas selectable per call — library overloads, CLI -b 0.5,2, REST "betas": [...]. Beta 1.0 is always included, duplicates are dropped, keys ascend. calculateF1 now delegates to calculateFBeta, removing the formula SingleClassificationResult.fBeta duplicated.

The aggregation rule is explicit and pinned by a test: macro and weighted F-beta are the (weighted) mean of the per-result scores, while micro is recalculated from the pooled matrix. Neither is the F-beta of the averaged precision and recall — on the test fixture that wrong computation gives 0.62 where the correct macro F1 is 0.49.

🐛 Fixed

  • Wrong-length weights walked off the end with IndexOutOfBoundsException, reachable straight from REST /average with user-supplied weights. Now a require, answered with 400.
  • A confusionMatrixSum smaller than the classified and expected elements silently produced a negative true-negative count and nonsense accuracy/phi. Now rejected.
  • HomeController registered a @Primary bare ObjectMapper, discarding Spring Boot's configured one. Under Spring Boot 4 jackson-module-kotlin is auto-registered, so the bean was redundant and harmful.
  • Handler mapped everything non-NPE to 500, swallowing Spring's own exceptions — an unknown path answered 500 instead of 404, an unsupported method 500 instead of 405, a malformed body 500 instead of 400. It now extends ResponseEntityExceptionHandler, and IllegalArgumentException maps to 400.
  • The CLI always exited 0. main() discarded the status execute() returned, so a missing input file, an invalid beta or an unreadable result file printed a message and still reported success — unusable for a pipeline branching on exit status. Pinned by a test that runs the CLI in its own process.
  • aggCl crashed with a raw stack trace on any file in -d that isn't a classification result — a stray note, a .DS_Store, or its own -o output written back into the input directory on a re-run. Hidden files are skipped; anything else that fails to parse is reported by name with exit status 1 rather than silently dropped from the aggregate.
  • calculateAccuracy(0, 0, 0, 0) returned NaN, which Jackson writes as the JSON string "NaN", violating the number type the REST schema declares for accuracy. It now returns the same 1.0 sentinel as its siblings, and a test pins that no metric of an empty confusion matrix is non-finite.

⚠️ Two Jackson traps worth knowing

Guarded by tests, since calculator deliberately depends on neither Jackson nor Swagger and so cannot carry annotations:

  1. The property is fbetaScores, not fScores. Jackson's legacy bean naming mangles getFScores() to fscores while jackson-module-kotlin reports the constructor parameter as fScores. They never match, so the mismatch is treated as read-only and silently dropped — every non-F1 score would vanish on read.
  2. Getter-only properties serialize but do not deserialize. confusionMatrix and ConfusionMatrix.total started out that way, so a consumer with a default ObjectMapper could not read back this library's own output. confusionMatrix became a creator property and total() became a function, guarded by a round-trip test with FAIL_ON_UNKNOWN_PROPERTIES enabled.

🔧 Build

The parent is pinned to the released io.github.ardoco:parent:2.0.1 rather than 2.1.0-SNAPSHOT. flattenMode is resolveCiFriendliesOnly, which keeps the <parent> element in the published pom, so a snapshot parent would have shipped artifacts no consumer could resolve. flatten-maven-plugin is pinned to 1.7.3 locally because 2.0.1 does not pin it; both can go once parent 2.1.0 is released.