Releases: aughtone/aughtone-normalize
Releases · aughtone/aughtone-normalize
Release list
Release 0.0.5
Full Changelog: v0.0.4...v0.0.5
Release 0.0.4
Full Changelog: v0.0.3...v0.0.4
Release 0.0.3
[0.0.3] - 2026-09-13
Breaking
- The subaddress-keeping email policy is no longer called lenient.
EmailPolicy.ByteStableV1LenientbecomesEmailPolicy.ByteStableV1Subaddressed, and its id moves fromemail.byte-stable+lenienttoemail.byte-stable+subaddressed. It keeps the+-subaddress rather than accepting more input, so it was never a lenient policy. Its bytes are unchanged, andemail.byte-stable+lenientno longer resolves. - The four Unicode normalization forms moved to text policy ids.
nfc.u17,nfd.u17,nfkc.u17andnfkd.u17are removed; the same forms aretext.u17+nfc,text.u17+nfd,text.u17+nfkcandtext.u17+nfkd, andTextPolicy.NfcU17and its siblings remain as presets under the new ids. Their output bytes are unchanged. Taken during alpha, before any known consumer stored a text id. PublishedPolicies.resolveis final. A module that rebuilds policies from their ids overridesresolveBaseinstead, so opted-in forms are handled once for every module.NormalizationStepcontributes a group of links.link: PolicyLinkbecamelinks: List<PolicyLink>, so a configured policy can compose into another module's chain whole.
Added
TextPolicyRuleOrderlint check. The:unicodeAndroid artifact bundles a lint check that warns in the editor when aTextPolicy { }builder lists its rules in a different order from the one they run in, naming the order they run. Its rank table is tested against real policies so it cannot drift from the runtime.- Email subaddress as its own piece.
normalizeEmailWithSubaddresswithEmailSubaddressPolicy.ByteStableV1reads an address once and returns the mailbox, identical tonormalizeEmailunderByteStableV1, and the RFC 5233 subaddress under its own identityemail.subaddress(NormalizedEmailWithSubaddress,NormalizedEmailSubaddress), so the two can be tokenized separately without re-implementing the email rules. - Domain display conversion.
toUnicodeDomainruns UTS-46 ToUnicode under aDomainPolicy, validating every label with the policy's Unicode release and flags and reporting per label its U-label, its ASCII form and the first check it failed (UnicodeDomain,UnicodeLabel). It is display conversion with no policy identity;normalizeDomainis unchanged. The UCD generator now also emits the ToUnicode columns ofIdnaTestV2.txt, and both domain policies are tested against them. normalizeTextis a configurable text normalizer.TextPolicy { ascii { … } }andTextPolicy(UnicodeRelease.U17) { unicode { … }; ascii { … }; nonEmpty() }configure control stripping, trimming, space collapsing or removal, lowercase, uppercase, full case folding, the four normalization forms and a non-empty check. Rules run in a fixed order regardless of how they are written, and the id renders the configuration:text+trim+lower,text.u17+trim+casefold+nfc. Rules written out of order produce aTextPolicyWarning, reported through a replaceableTextPolicy.warningHandler.UnicodePoliciesresolves any text id by rebuilding it and refuses non-canonical spellings with the newPolicyIdentityError.NotCanonical. Convenience presetsTrimLowercaseandCaselessU17are provided.- Comparable forms. A policy can declare the canonical forms it writes, and offer forms a caller may opt into by naming them in the id after every other link (
ipv4.inet-aton+form.ipv4.address).PolicyResolver.comparabilityreports whether two stored identities are the same policy, comparable in a named form, or not comparable.LinkKind.Form,ComparableForm,Comparability,OptedInPolicyandPolicyIdentityError.FormNotOfferedare new. Forms are additive on a published policy; composed policies and skeletons declare none. - IP network normalization.
normalizeIpv4BlockandnormalizeIpv6Blockderive the network an address falls in at a prefix named in the policy id (ipv4.dotted-quad+block-24),normalizeIpv4BlocksandnormalizeIpv6Blocksderive several prefixes from one reading of an address, andnormalizeIpv4CidrandnormalizeIpv6Cidrcanonicalize CIDR input, refusing host bits (…+cidr) or clearing them (…+cidr+masked). Block derivation and CIDR input share one network formatter. - UUID normalization.
normalizeUuidwithUuidPolicy.Hex(any 128-bit value) andUuidPolicy.Rfc9562(variant and version checked, Nil and Max accepted) accepts any case, braces,urn:uuid:and the bare form, and writes lowercase hyphenated text under the comparable formuuid. TheguidBytes()mode reads Windows GUID byte dumps, refuses string forms, and offers theuuidform for opt-in.formatUuidrenders braces, URN, uppercase, bare and GUID byte order for display. - MAC address normalization.
normalizeMacwithMacPolicy.Eui48andMacPolicy.Eui64accepts colon, hyphen, Cisco dotted and bare spellings in any case, with leading zeros omitted, and writes lowercase colon pairs. The policies declare the comparable formsmac.eui48andmac.eui64.formatMacrenders a normalized address in IEEE, Cisco dotted or bare notation for display. - Strict and lenient pairs share a comparable form where leniency only widens what is accepted:
pan.digits(PanForms.Digits),iban.compact(IbanForms.Compact),domain.ascii.u17andurl.rfc3986.u17(DomainForms). Email has no lenient policy; its subaddress-keeping variant is a parameter,email.byte-stable+subaddressed, and is not comparable withemail.byte-stable. - Phone comparable form. Every phone policy, with or without a region and strict or lenient, declares the comparable form
phone.e164(PhoneForms.E164), so numbers read under different phone policies are comparable explicitly throughPolicyResolver.comparability. - IP comparable forms.
IpFormsnamesipv4.address,ipv6.address,ipv4.networkandipv6.network.ipv4.dotted-quadand the IPv6 policies declare their address forms, withunmapandnat64also declaringipv4.address; block derivation and CIDR input declare the network forms.ipv4.inet-atonand its networks only offer theirs, soipv4.inet-aton+form.ipv4.addressrecords a caller's choice to compare it withdotted-quad. - IPv6 modes.
Ipv6Policy.unmap()writes IPv4-mapped addresses as IPv4,nat64()does the same for64:ff9b::/96, andzone()keeps zone identifiers; each is named in the id, and a block underunmapornat64carries an IPv4 and an IPv6 prefix. IP block, CIDR and mode policies resolve fromQuodlibetPoliciesby rebuilding them from their ids. - Composed chains resolve. A combined resolver splits an id such as
username.basic+skeleton.u17ortext.u17+trim+casefold+skeleton.u17at each group opener, resolves every group with the module that owns it, and returns aComposedPolicywhosebaseandstepsre-derive the same bytes. A group no module owns fails the whole chain, and a group that cannot run as a step fails with the newPolicyIdentityError.NotAStep. - Unicode 17
CaseFolding.txtandPropList.txtare pinned, and the generator emits White_Space, control, case folding and simple case mapping tables for the text rules. - A portable spelling for policy ids.
PolicyId.toPortableandPolicyId.fromPortableconvert an id to and from a form that uses_in place of+, andPolicyId.portablerenders a parsed chain that way. It is for places that reject+: form-encoded query strings, Kubernetes label values, container image tags. The canonical+spelling remains the stored identity. StepPhase.rank. Chain order is validated against an explicit, frozen rank rather than enum declaration order. Every id parses exactly as before.
Changed
- Examples and internals use the
OutcomeAPI instead of hand-writtenwhenblocks. The README and KDoc consume results withonSuccess { }/onFailure { },dataOrNull()anddataOrThrow(), and thenormalizeEmailOrNull,normalizeDomainOrNullandnormalizeTextOrNullhelpers are nowdataOrNull()calls. No canonical output, policy id or signature changed.
Fixed
- Combining resolvers no longer loses what a module resolves alone.
a + bused to search a merged list of policies, so ids a module rebuilds on demand, such asphone.e164+region-ca, failed once combined. Each module's ownresolveis now asked, andplusflattens soa + b + cis a single composite.
Release 0.0.2
[0.0.2] - 2026-09-12
Added
- A policy identity can be resolved back into its policy.
PolicyResolverin:commonturns a stored(id, version)pair into the frozen policy that produced it, so an application adding rows to an existing store normalizes them under the policy those rows were derived with, and a policy can be named in configuration rather than hardcoded. Each module exposes one resolver —QuodlibetPoliciestoday — and callers combine the ones they depend on with+. There is no global registry and nothing registers at startup. - The policy id grammar is published API.
PolicyIdrenders and parses a chain of links, and every policy in the suite renders its own id through it, so the written and parsed forms cannot drift apart. Parsing is strict: a chain out of canonical order, an unknown link, or a version this build does not carry is a typedPolicyIdentityError, never a near match. Policyin:common— the base type every policy now implements, carryingidandversion.- The Unicode table generator. A JVM-only build module, never published, turns a pinned copy of the Unicode Character Database into the suite's frozen tables. The data files are checked in with their checksums and verified before parsing, so a build needs no network and a regeneration is reproducible. Tables are emitted as Kotlin source that compiles on every target, with no resource loading.
- A material-change check, wired into
check. Regenerating classifies every difference from the previous baseline as an addition, a modification or a removal; a modification in a normalization table fails the build, because the suite's byte-stability promise and its delta packaging both assume additions only. IDNA and confusable data, whose upstream guarantees are weaker, report changes without failing.verifyUnicodeTablesalso fails when a generated file is edited by hand or the data moves without a regeneration. :unicode: NFC, NFD, NFKC and NFKD, vianormalizeText(value, policy)under four frozen policies —TextPolicy.NfcU17,NfdU17,NfkcU17andNfkdU17, with idsnfc.u17,nfd.u17,nfkc.u17andnfkd.u17. The tables are frozen against Unicode 17.0.0 and ship with the library: nothing reads the platform's Unicode data, which is what would otherwise make the same input normalize differently on an old Android build and a current iOS one. A later Unicode release mints a new policy rather than changing one.- The Unicode Consortium's conformance suite runs on every target. All ~19,000 cases of
NormalizationTest.txtfor Unicode 17.0.0, on jvm, js, wasmJs and iOS, because a normalizer verified only on JVM has not been tested for the one property this suite sells. :phone: phone numbers to E.164, vianormalizePhone(value, policy)withPhonePolicy.E164,E164Lenient,e164ForRegion(region)ande164ForRegionLenient(region)— idsphone.e164,phone.e164+lenient,phone.e164+region-caandphone.e164+region-ca+lenient. The region is part of the policy identity and is never inferred (ADR-0001). Digits from any script are converted through the frozen tableio.github.aughtone:phonenumberpublishes, so this module ships no table of its own; letters are refused rather than dialled.- A URL whose host is an address is accepted only in canonical dotted-quad form, and every other spelling — leading zeros, hexadecimal parts, fewer than four parts, a bare integer — is refused rather than rewritten. Rewriting would mean choosing between readings that disagree and hiding that choice inside a URL;
normalizeIpv4is where that choice belongs, and its policy identity records it. :ubilibet: URL normalization, vianormalizeUrl(value, policy)withUrlPolicy.Rfc3986U17andRfc3986U17Lenient— idsurl.rfc3986+domain.ascii.u17andurl.rfc3986+domain.ascii.u17+lenient, which name the host policy they use. Limited to transforms that cannot change which resource is addressed: the query and fragment survive byte for byte, and sorting query parameters is deliberately not done.:confusables: UTS-39 skeletons, vianormalizeSkeleton(value, policy)withConfusablePolicy.SkeletonU17(idskeleton.u17). The fullskeletonis implemented, not the simplerinternalSkeleton: the standard defines it asbidiSkeleton(LTR, X), so the text is laid out by the bidirectional algorithm before it is reduced. A skeleton is documented throughout as a check rather than an identity — it is many-to-one by design, and storing one as an account key merges different users.- The Unicode Bidirectional Algorithm (UAX #9) through rule L2, with the paired-bracket rule, isolates and overrides. Its conformance suite runs in full on JVM — all 94,000 cases of
BidiCharacterTest.txt— and a deterministic sample of it on every other target, because the whole corpus is several megabytes once compiled into a test binary. :quodlibetgains five more normalizers, all table-free and all in the same bundle: credit-card/PAN (pan.digits,pan.digits+lenient), IBAN (iban.compact,iban.compact+lenient), IPv4 (ipv4.dotted-quad,ipv4.inet-aton), IPv6 (ipv6.rfc5952) and usernames (username.basic).- IPv4 ships two policies rather than one interpretation. Stacks disagree about the shorthand spellings —
192.168.0.010is 8 underinet_aton, 10 under a plain decimal reading, and refused outright by Go and Python — soipv4.dotted-quadaccepts only the unambiguous form whileipv4.inet-atonapplies the classic rules deliberately and records in its id that it did. Values under the two never match, which is correct: they are different claims. The financial pair gate on their check digits by default and relax that under a lenient policy, matching how:phonetreats an implausible number; IPv6 follows RFC 5952 and has no lenient variant, because every candidate relaxation changes which address is meant rather than how it is spelled; the username base mirrors email's ASCII-only rules and encodes no platform behaviour. :ubilibet: hostname and domain normalization under UTS-46, vianormalizeDomain(value, policy)withDomainPolicy.AsciiU17(iddomain.ascii.u17) andDomainPolicy.AsciiU17Lenient(iddomain.ascii.u17+lenient). Output is the A-label form. Every hostname is normalized here, ASCII included, so an ASCII name can never carry two policy identities for the same bytes. Nontransitional processing only; transitional processing is deprecated upstream and is not implemented.- The UTS-46 conformance suite runs on every target. All of
IdnaTestV2.txtfor Unicode 17.0.0, checked against both policies, with the file's own mapping from status codes to flags deciding what each policy must refuse. - Punycode (RFC 3492), including the bootstring overflow checks, which are part of the specification rather than defensive padding.
NormalizationStepin:common, the interface a table-free module accepts so a caller can compose a transform from another module without either module depending on the other. EachTextPolicyis one, soemail.byte-stable+nfc.u17names its transform in the identity.
Changed
- Breaking: the email normalizer moved to a new coordinate. It is published as
io.github.aughtone.normalize:quodlibetinstead ofio.github.aughtone.normalize:email. Change the dependency coordinate and nothing else: the packageio.github.aughtone.normalize.email, every type and function name, the canonical output and the policy versions are all unchanged, so no stored value or derived token is affected.io.github.aughtone.normalize:email:0.0.1remains on Maven Central and is not republished.:quodlibetis the bundle for normalizers that need no lookup table and no external dependency, so PAN, IBAN, IPv6 and the username base will join it rather than arriving as separate artifacts. - Breaking: the relaxed email policy is renamed.
EmailPolicy.LenientbecomesEmailPolicy.ByteStableV1Lenient, and itsidmoves fromemail.lenienttoemail.byte-stable+lenient. Canonical output, policy version and error types are unchanged — an address normalized under it produces exactly the bytes it did in0.0.1. - Policy ids are now chains. An id is an ordered list of links joined by
+: the base rule set, then anything qualifying it, then any steps from other modules. The id therefore describes the policy rather than merely labelling it, which is what lets a stored id be resolved back to the policy that produced it. - Taken deliberately in alpha, before any consumer had derived a token under
email.lenient. A published id does not move once anyone holds a value derived under it;email.byte-stableis unchanged and never will move.
Release 0.0.1
[0.0.1] - 2026-09-07
Added
- Byte-stable email normalizer (
:email):normalizeEmail(value, policy)produces a deterministic, byte-level canonical string for hashing and blind tokenization. The canonical form applies only ASCII-level, Unicode-version-independent operations (trim ASCII whitespace, ASCII-lowercase, RFC 5233+-subaddress handling), so its output is byte-identical on every platform and every build and never drifts when Unicode ships a new version. - Named, frozen policies (
EmailPolicy):EmailPolicy.ByteStableV1is the shared canonical form for blind tokenization (strips the+-subaddress);EmailPolicy.Lenientis a loose trim + ASCII-lowercase key for display and dedupe. Each policy carries anid+versionepoch that travels with any derived token, and a published policy's output is never edited in place — a rules change mints a new version. Outcome-based result with typed, value-free errors: normalization returns an aughtone-typesOutcome<NormalizedEmail>; malformed input yieldsOutcome.Failurecarrying a typedEmailNormalizationError(MissingAtSign,EmptyLocalPart,EmptyDomain,UnpairedSurrogate) whose messages never echo the input, so a rejected address cannot leak into a log. AString.normalizeEmailOrNull(policy)convenience is also provided.- Depends on
io.github.aughtone:types3.4.0, exposed transitively viaapi, forOutcome.Success/Outcome.Failureand therunOutcome { }builder. - Multiplatform targets: published for JVM, Android, iOS (
arm64,x64,simulatorArm64), JS (browser), wasmJs (browser) and Linux x64. The same test suite runs on each, so the canonical form is verified identical across them rather than assumed. - Shared
Normalizedcontract (:common): theNormalizedinterface (canonical,policyId,policyVersion) is the common result shape every normalizer in the suite reports, so a derived hash can always be stored beside the policy identity that produced it.