Skip to content

Releases: Nicholas022400701/ulpwise

ulpwise 0.3.4

Choose a tag to compare

@github-actions github-actions released this 09 Oct 03:21
  • Regression corpus case for pytorch/rl #4504, merged upstream: MultiCategorical.to_one_hot iterated over self.nvec,
    a 2-D tensor for a spec with a batch shape, so one_hot got a tensor as num_classes and raised TypeError, and the
    0-D sample of a shape [1] spec raised IndexError on val[..., i]. The case one-hot encodes a sample of
    MultiCategorical([3, 2], shape=(4, 2)) and the scalar sample of MultiCategorical([5]).
  • ulpwise scan: new dropout-never-applied rule (medium), the first that reads a whole file rather than one
    function: a Dropout, DropPath or a ModuleDict or ModuleList under a dropout name is assigned to self
    and then never called, never passed on, never returned or iterated and its rate never read, anywhere in the file.
    The option is accepted and does nothing, so the model trains without the regularisation it reports. peft #3830 is
    the example: OFTLayer built module_dropout into self.oft_dropout, filled it in update_layer and never
    called it. Reading the rate counts as use (scaled_dot_product_attention(dropout_p=self.dropout.p)), as does
    passing the module to another one; filling it does not, and subclasses of Sequential, which run every
    attribute, are not read. On kornia, ultralytics, torchvision, diffusers, torchrl, vllm and detectron2 main the
    rule reports nothing; on peft main it reports the #3830 line, on timm main one attn_drop in coat.py that a
    comment already calls unused, and on transformers main nine lines, four of them modular files whose forward
    lives in another file.

ulpwise 0.3.3

Choose a tag to compare

@github-actions github-actions released this 08 Oct 03:14
  • Two regression corpus cases for peft, both open upstream with a fix PR: peft #3769, add_weighted_adapter with
    combination_type='svd' crashed for Conv1d and Conv3d LoRA layers because only Conv2d deltas were flattened before
    the SVD (the case combines two rank 2 adapters at svd_rank=4 and expects the output of both adapters active, which
    the fix matches to 2e-6), and peft #3830, OFT module_dropout was never applied, so ten training forwards of the
    same input were identical where the fix spreads them by 0.56.
  • ulpwise scan: new where-nan-gradient rule (medium) for where(d > eps, f(d), other) with f a division by d
    or a sqrt, log, acos or asin of it. where evaluates both branches and hands the discarded one a zero
    gradient, and the backward of f at the singularity turns that zero into 0 / 0 = nan, so the guard protects the
    value and not the gradient. kornia #5579, found by the kornia conventions audit, is the example: the mutual
    information losses divided by the range of the signal inside where(diff > eps, ...) and a constant input or
    target got a nan gradient. The condition may be a comparison, a name bound to one, or ~, &, | and a subscript
    of those, and the method form a.where(cond, b) counts. Quiet for numpy.where, which has no autograd, for a
    comparison against a number above 1 (strength < 50 in kornia's JPEG scale is a piecewise definition), for a
    floor named eps, tol, floor, thresh, tiny, bound or limit on the other side of the comparison, and
    once d is re-bound to a where, clamp or maximum of itself, which is the recommended fix and the shape of
    kornia's fit_line and rotation_matrix_to_quaternion. On kornia main d15741e2 the rule reports four lines, on
    ultralytics main 8df3534 two, both in metric code, and nothing in torchvision.
  • ulpwise scan: new clamp-at-singularity rule (medium) for clamp(x, min=0).sqrt(), sqrt(clamp(x, min=0)),
    clamp(c, -1, 1).acos() and the clip, clamp_min, clamp_max, asin and log forms: the bound is the point
    where the next function has an infinite derivative, and clamp's derivative at its own bound is 1 on torch 2.5.1 and
    2.9.1 and 0 on 2.14, so the gradient there is inf or nan on the older half of a supported range. kornia #4229,
    found by the kornia conventions audit, is the example, _cdist's clamp(min=0.0).sqrt() with identical float16
    descriptors, and kornia #5500's clamp(-1, 1).acos() the second. A bound strictly inside the domain, min=1e-8,
    is a floor and is not reported, a bound given by a name is not read, and numpy is quiet. On kornia main d15741e2
    the rule reports the three lines of _solve_cubic_real, on ultralytics main 8df3534 one line of DepthMetrics.
  • ulpwise scan: a finding inside a nested function is reported once, under the nested function, instead of once
    more under the enclosing one, and the enclosing function's call and division counts no longer include it.
  • ulpwise scan: new eps-floor rule (medium) for x * (1 - eps) + eps, in either order and with self.eps too: the
    floor moves every value, 0 becomes eps and 1 stays 1, so the entries of a one-hot sum to 1 + (C - 1) eps and a
    perfect prediction scores a loss that grows with the image. kornia #5538, found by the kornia conventions audit, is
    the example: one_hot floored its zeros with 1e-6 and a perfect prediction of a 256 × 384 image with a one-pixel
    class scored a macro Dice loss of 0.0234. On kornia main 9e188e08 the rule reports that line and nothing else, and
    nothing in ultralytics or torchvision; (p + eps) / (q + eps) is not a floor and stays quiet.
  • ulpwise scan: the hypot-by-hand advice no longer offers norm as the replacement. torch.linalg.vector_norm and
    numpy.linalg.norm square the entries first too, so in float32 both return inf for a vector with two entries of
    1.8e19 while hypot returns 2.5e19; the advice now says to divide a vector by its largest absolute entry when the
    point is a direction. The kornia maintainer caught the case in review of kornia #5505, where v / v.norm() turned
    such vectors into zeros and the angle between them read 0.
  • ulpwise scan: acos-for-angle moves from info to medium and names the mechanism: next to 1 the cosine has no
    digits left for a small angle, so acos recovers the angle with an absolute error of sqrt(eps), 0.02 degrees in
    float32, whatever the angle. The kornia scan listed kornia/metrics/pose.py under this rule, and both functions
    there, angle_error_mat and angle_error_vec, return exactly 0 for every rotation below 0.03 degrees in float32
    and 180.0 for 179.99 degrees (kornia #5500); the rule cites it.
  • Regression corpus case for kornia #5500: a 0.01 degree rotation scores 0.0 in both metrics on kornia main 050ac77f.
  • The kornia #5500 corpus case records its fix, kornia #5505, merged 2026-10-06: angle_error_mat reads the sine from
    the skew part of the relative rotation and angle_error_vec from the cross product of the scaled vectors, and both
    return atan2(sin, cos), so a 0.01 degree rotation scores 0.0099990 in float32. The case stays an expected failure
    until a kornia release carries the fix.
  • ulpwise scan: logsumexp-by-hand no longer fires on the stable form, where every exp argument has its
    maximum subtracted first (log(exp(x - x.max()).sum()), or s = x - m with m bound to a max, amax,
    maximum or torch.max(...).values). The lm-evaluation-harness scan listed the numerically stable _log_softmax
    of lm_eval/models/_onnx_base.py next to a real log(exp(a) + exp(b)); it now reports only the real one.

ulpwise 0.3.2

Choose a tag to compare

@github-actions github-actions released this 06 Oct 01:57
  • Two regression corpus cases for pytorch #199850: torch.erf in bfloat16 and float16 on CPU returns 0 at and
    below 1.8e-7 and loses relative accuracy below 1e-3 (13404 bfloat16 ulps, 5 float16 ulps at worst), found with
    the half precision rows of ulpwise survey --functions erf --backends torch --dtypes f16,bf16.
  • Two regression corpus cases for pytorch #199867: torch.special.logit in float16 and bfloat16 on CPU rounds 1 - x
    and x / (1 - x) to the input dtype before the log, so logit(0.499756) in float16 is -0.000488 for an exact -0.000977
    (512 ulp, 64 bfloat16 ulp at worst), found with ulpwise survey --functions logit --backends torch --dtypes f16,bf16.
  • studies/accuracy-survey-2026-10-half: the second accuracy survey, torch 2.14.0+cpu in float16 and bfloat16 for the
    45 functions with half precision CPU kernels. 33 float16 and 34 bfloat16 rows of 45 are correctly rounded; erf
    (pytorch #199850) and logit (pytorch #199867) lose their digits before the rounding step, polygamma(2, x) in
    float16 is off at every half integer in (-1024, -256) because the Hurwitz zeta sum accumulates in float for the
    reduced types, and the activation tails and the rsqrt and i0e vector versus scalar disagreements of September
    show again.
  • ulpwise survey: the polygamma_1 and polygamma_2 references at a non positive integer are now the signed
    infinity of the pole, +inf for an odd order (the limit from both sides) and -inf for an even one (the sign of
    (-1) ** (n + 1) * n! * zeta(n + 1, x), which is what scipy, torch and jax return), instead of a domain error
    that expected nan. The f16 and bf16 grids reach the integers from 2048 and 256 up, so every negative point
    past there was a pole and counted as a nonfinite mismatch: on torch 2.14 the two functions had 354 of 3546
    float16 points and 492 of 4016 bfloat16 points each, polygamma_2 has 0 now in every dtype and polygamma_1
    keeps 491 bfloat16, 4 float32 and 2 float16 points where torch returns a large finite value instead of inf
    at a negative integer (pytorch #198663).

ulpwise 0.3.1

Choose a tag to compare

@github-actions github-actions released this 05 Oct 22:44
  • ulpwise survey takes f16 and bf16 in --dtypes (the default stays f32,f64): torch and jax get half
    precision rows, measured in float16 or bfloat16 ulps against the same mpmath reference; numpy has no
    bfloat16 and scipy.special computes a float16 input in float32, so those two have no bf16 rows and scipy
    has no f16 rows. A backend without a kernel for a function in a dtype (torch's Bessel and Airy functions,
    erfcx, ndtri, log_ndtr and zeta in half precision) is logged and skipped before the reference is
    computed. The torch default tolerances of the two dtypes are (1.6e-2, 1e-5) and (1e-3, 1e-5), as in
    torch.testing. A log spaced domain wider than the dtype is now clipped to the dtype's finite range before
    the points are spaced: the float16 grid of sqrt and log over (1e-300, 1e300) kept 32 of 600 points and
    acosh 24, they have about 570 and 490 now; the float32 grids of the same functions go from about 100 points
    to about 600, the float64 grids are unchanged. The (0, 1) and (-1, 1) grids of ndtri, logit, asin,
    acos, atanh and erfinv stop at the float below 1 in half precision too, where 1 - 1e-4 and 1 - 1e-3
    rounded to 1. Two corrections to the ulp metric for every dtype: the threshold from which an exact result
    counts as a correct overflow to infinity was fmax * (1 + eps / 2), about one ulp above the largest finite
    value, and is now fmax + ulp / 2, the rounding threshold; and the spacing at the largest finite value was
    numpy's inf, which turned every finite result there into a 0 ulp error, it is ulpwise.spacing now.
    An unknown dtype raises ValueError instead of a KeyError from the grid.
  • next_up, next_down, spacing, neighbours, binade_edges and all_floats accept f16 and bf16 (and the
    float16, half, bfloat16, torch.float16, torch.bfloat16 aliases) like ulp_distance and special already
    did; they raised unsupported dtype before. The 16 bit versions are pure Python on the bit patterns, with the
    conventions of the Rust ones: the largest float steps to infinity, both zeros to the smallest subnormal, spacing
    is the distance from |x| to the next float and stays finite at the largest float, neighbours returns finite
    values only. A test walks every finite float16 and bfloat16 and checks the six functions against each other, against
    numpy's float16 nextafter and spacing and against torch's bfloat16 nextafter.
  • ulpwise scan read the first positional argument of a method call as the value the method is applied
    to, so log(exp(x).sum(-1)) was not a logsumexp-by-hand finding while log(exp(x).sum()) and
    log(exp(x).sum(dim=-1)) were. A method call whose first argument is a dimension (an integer, None
    or a tuple of them) is now read as applying to its receiver, like the argument-less form.
  • ulpwise scan printed a SyntaxWarning (a DeprecationWarning before Python 3.12) for every invalid
    escape sequence in the files it scanned, four lines on kornia. The parse now ignores those warnings;
    a file that does not parse is skipped as before.
  • ulpwise scan reported Windows paths with backslashes (kornia\geometry\conversions.py), so the same
    report read differently from the one made on Linux or macOS. Paths are now written with forward
    slashes on every platform; --run still imports the module from either form. The CI pytest job on
    ubuntu, macOS and windows runs tests/test_scan.py now, and that file also covers the reasons
    --run gives for a function it does not run (a class, a constant, an instance method, a builtin
    without a signature, a nested name, a module that fails to import) and the ulp comparison of the
    results across infinities, NaNs and an all zero reference.
  • ulpwise survey without scipy installed died at ndtri with a ModuleNotFoundError and lost every
    row computed before it: the reference started its Newton iteration from scipy.special.ndtri. It now
    starts from a bisection on math.erfc, returns an exact zero at p = 0.5 (a Newton step there leaves a
    rounding residual, which the ulp metric would read as 1e16 ulps against the exact 0.0 of torch and scipy)
    and gives the same rows as before on the 600 point float64 and float32 grids. The survey also evaluates
    the reference only when some requested backend implements the function; a numpy only run no longer
    spends its time on references for the torch and scipy only entries. The pytest job on ubuntu, macOS and
    windows (Python 3.12 and 3.9) runs tests/test_survey.py now, with numpy as the only backend.
  • ulpwise survey printed a numpy RuntimeWarning for every backend call that divided by zero or
    overflowed at the edge values of the grid (reciprocal and log at 0, reciprocal at the overflow
    edge), hundreds of lines on a full run. Those points are counted in the nonfinite column already, so
    the survey now evaluates the backends under np.errstate(all="ignore"), and a test runs reciprocal and
    log with warnings turned into errors.
  • The README said that two kornia fixes in the corpus were merged but not released; the corpus has 31
    kornia cases now, and the pytorch/rl, peft and pytorch (#198006) fixes are in the same state (kornia
    0.8.3 and 0.9.0rc1, torchrl 0.14.0, peft 0.21.2 and torch 2.14.1 predate them or cherry-pick other
    changes), while the timm and ultralytics fixes are released. The sentence says so now, and
    tests/test_corpus.py checks that the README table has one row per case with the merge date from
    cases.json, so a case added without its row, or a row with a stale date, fails the suite.
  • tests/test_survey.py checks every reference of the survey registry against the libraries it measures:
    each entry is surveyed on a 24 point float64 grid against torch, numpy and scipy and its median error must
    stay under 32 ulp, which a reference that is another function fails by fifteen orders of magnitude (torch's
    polygamma_1 is the worst true median at 9 ulp), and the piecewise references (selu, elu, entr,
    gelu_tanh, log_ndtr, lgamma, digamma, erfinv, ndtri, zeta, sinc, spherical_bessel_j0)
    are pinned once per branch, since a slip in one branch moves a third of the grid and keeps the median.
    survey.py goes from 77 to 94 percent covered.
  • With torch installed but expecttest missing, torch.testing._internal does not import and
    ulpwise survey silently used the dtype default as the OpInfo tolerance, so the op tol columns
    looked like an override that was never read. The survey now logs one line with the import error and
    the survey extra installs expecttest. The regression corpus CI job, which has torch but had no
    expecttest, failed on the two op_db tests for the same reason; it installs expecttest now and the
    tests skip with the reason where op_db is not importable.
  • ulpwise survey looked up torch's OpInfo tolerance by op name and took the first op_db entry
    with that name. polygamma has one entry per order, and the first, polygamma_n_0, has no
    override, so the op fail column of polygamma_1 and polygamma_2 in float32 was computed with
    the dtype default (rtol 1.3e-6, atol 1e-5) instead of torch's polygamma_n_1 and _n_2
    override (rtol 0.01, atol 1e-4). Entry gains opinfo_variant, torch_opinfo_tolerance
    takes a variant and otherwise prefers the base variant, and the two registry entries name
    theirs. In studies/accuracy-survey-2026-09 the polygamma_1 f32 row goes from 86 to 7 inputs
    failing the op's tolerance; the measured errors are unchanged (erratum in its README).
  • tests/test_survey.py covers the survey itself now: survey() on sqrt against numpy and torch
    (numpy within half an ulp, torch under one), an inexact backend, a backend that raises or returns
    the wrong shape, the backend resolution, the op_db tolerance lookup including the variant case,
    the csv and markdown writers, versions(), main() and the survey subcommand. CI installs
    mpmath for the Python job so these tests run there instead of being skipped.
  • Six more kornia cases, each present on kornia 0.8.3 and fixed on main: the Hessian of So3.exp
    at the identity (#4972, nine nan for -I / 4 and zeros), sampson_epipolar_distance of a point on
    its epiline with squared=False (#5116, sqrt(eps) = 1e-4 for 0) and of the same F scaled by
    1e-4 (#5116, a third lower), RandomHue on a float64 image (#5131, 8.7e-8 rad past a half turn,
    the float32 rounding error of pi), MS_SSIMLoss(sigmas=(0.5, 1.3), reduction="none") (#5143, an
    even 6-pixel window and a (1, 15, 19) map for a (1, 3, 16, 20) input) and MS_SSIMLoss on a
    uint8 pair (#5353, expected scalar type Byte but found Float; with data_range=255 it now scores
    the pair divided by 255 at the default to 1e-6). Six more: So3.right_jacobian of a float16
    45 rad rotation (#4967, the identity for a matrix of 0.2 to 0.4), average_quaternions with a member
    stored as 3 q (#4980, 41.8 degrees for the 22.5 degree bisector), Hyperplane.through a float16
    triangle with legs of 300 (#5104, the normal (0, 0, 1) for (0, 0, -1)), filter2d with a
    per-sample kernel on a channels-last image (#5301, torch's view error), lovasz_hinge_loss of a
    float16 prediction (#5303, a float32 loss) and get_box_kernel1d (#5357, a stride-0 view where
    one write zeroed all three taps). The kornia block is 31 cases, all present on 0.8.3 and all
    fixed on main. One more ultralytics case, scale_masks with the dataloader's ratio_pad and an
    odd letterbox padding (#26379, a padded row survived the crop and the bottom rows of every mask
    faded to 0) and verify_image_label on two triangles that tile one square (#26377, the second
    polygon was dropped as a duplicate of the first because only the class and the box were compared);
    both carry `f...
Read more

ulpwise 0.3.0

Choose a tag to compare

@github-actions github-actions released this 26 Sep 06:09
  • ulpwise scan: static scan of a repository (a directory, a GitHub URL or owner/repo, cloned
    with depth 1) for the floating point patterns behind the corpus bugs. Twelve rules with severity,
    reason, replacement and upstream example: exp-of-square, sin-of-pi-times, softplus-by-hand,
    logsumexp-by-hand, hypot-by-hand, sqrt-of-difference, one-minus-cos, log1p-by-hand,
    expm1-by-hand, atan-of-quotient, small-angle-division, acos-for-angle. Findings carry
    file, line, function and the source line; the report ends with the functions that do the most
    elementary math. --report writes Markdown, --rules filters, --fail-on gates CI,
    --include-tests widens the walk. On kornia main the four small-angle-division and
    one-minus-cos lines are the ones kornia #4897 fixes.
  • ulpwise scan --run: imports the math-heavy module level functions and calls them on the same
    grid in float32 and float64, reporting the largest error in ulps of the largest output and the
    worst elementwise ulp distance with its input. Static methods run, instance methods and other
    skips carry their reason. --run-limit bounds it, --run-installed imports the installed package
    instead of the scanned tree.
  • ulp_distance, ulp_distances, ordered, max_ulp and assert_max_ulp take f16 and
    bf16, and flatten reads the dtype off float16 and bfloat16 numpy arrays and torch tensors.
    Inputs that are not representable are rounded to nearest even first, so a bfloat16 tensor can be
    measured against a float64 reference; without an explicit dtype the less precise of the two
    inputs decides. Cross-checked against the numpy int16 view for float16 and the torch view for
    bfloat16.
  • special("f16") and special("bf16"): the same 29 named edge values as f32 and f64, so
    edge_values, the edge_f16 and edge_bf16 pytest fixtures and ulpwise special f16 work.
    Every value satisfies the same checks as the f32 and f64 tables, run through numpy for float16
    and torch for bfloat16.
  • ulpwise corpus: runs the regression corpus against the installed packages without pytest and
    prints one line per case, present, fixed or skipped, with the installed version and the upstream
    reference. --repo filters by repository or case id, --fail-if-present makes a present bug exit 1.
    Any exception from a repro counts as present, like the pytest run: several corpus bugs are crashes.