ktecma262 0.2.0
The library stops being only a regular expression engine.
Four more corners of ECMA-262 arrive, each one something no Kotlin target
reproduces on its own: JavaScript's number formatting and parsing, the four URI
escaping functions, Unicode normalisation, and identifier validation — plus the
handful of places Kotlin returns a different answer rather than an error.
Nothing in the regular expression engine changed behaviour, so upgrading from
0.1.4 is safe; the version is 0.2.0 because the surface grew rather than
because anything moved. The JVM artifact grows from 176 KiB to about 269 KiB,
almost all of it the normalisation tables. Shrinkers remove what you do not
call — the tables are a single leaf class with one inbound reference — and on
plain JVM an unused class is never even loaded.
Added
-
String.isEcmaIdentifierName(),isEcmaIdentifier()and
isEcmaReservedWord()inio.github.mgilbir.ecma262.text- ECMA-262 12.7,
reusing the Unicode tables already compiled in.The two questions are different and both are useful: a property key needs an
IdentifierName, a binding needs anIdentifier, and keywords are the first
but not the second. Words reserved only in some contexts -letand the
strict-mode set,awaitin a module,yieldin a generator - are parameters
rather than assumptions.Verified at two levels: every code point against the composed rule from 12.7,
and a sample against what node's parser actually accepts. The second caught a
mistake in the first oracle:o.a-bparses as(o.a) - b, so testing member
access reporteda-bas a valid name. An object literal key is the correct
probe.Five planted bugs are caught. One more showed that an explicit lone-surrogate
branch was dead code, since a surrogate is in neitherID_Startnor
ID_Continue; it has been removed rather than left as decoration. -
String.ecmaTrim(),ecmaTrimStart(),ecmaTrimEnd()andEcmaMathin
io.github.mgilbir.ecma262.textand.number. These are the places Kotlin
disagrees with JavaScript by returning a different answer rather than an
error, on every target including Kotlin/JS.Trimming differs on five characters, measured rather than assumed: Kotlin
strips U+001C to U+001F where JavaScript keeps them, and keeps U+FEFF where
JavaScript strips it. U+00A0 is stripped by both, contrary to a first guess.
Which characters count is verified over the whole BMP.EcmaMathcovers only the exactly specified functions —round,trunc,
sign,clz32,imul,fround— since the rest ofMathis
implementation-approximated and there is nothing to be correct against.
kotlin.math.roundrounds ties to even; JavaScript rounds them toward
positive infinity.Two implementations that look obvious are wrong:
floor(x + 0.5)answers 1
for0.49999999999999994, andtoFloat().toDouble()is a no-op on
Kotlin/JS, where aFloatis a JavaScript number. The second was caught by
the JS target failing while JVM and native passed.The whitespace predicate is shared with the number parser, so the two cannot
drift apart. Five planted bugs are caught. -
String.normalize()inio.github.mgilbir.ecma262.text— ECMA-262 22.1.3.15,
which defers to UAX #15, in all four forms.java.text.Normalizeris JVM-only
and Kotlin/Native has nothing, so multiplatform code comparing user-entered
text has been comparing sequences that look identical and are not.Tables are generated from the UCD, with decompositions stored fully expanded
so normalising is a lookup rather than a recursion, and Hangul left out
because its mappings are arithmetic.Checked against Unicode's own
NormalizationTest.txt: 20,034 rows whose
expectations no implementation produced, each verified against the invariants
the file states — all five columns must agree under each form. Generation
fails if node disagrees with any row. Every one of the 1,112,064 code points
is then checked individually in all four forms against node.Six planted bugs are caught: unstable canonical ordering, ignoring the
composition blocking rule, dropping a Hangul trailing jamo, NFKC using
canonical decomposition, and — in the generator — ignoring
Full_Composition_Exclusionand storing decompositions one step deep instead
of fully expanded. -
String.encodeUriComponent(),encodeUri(),decodeUriComponent()and
decodeUri()inio.github.mgilbir.ecma262.uri— ECMA-262 19.2.6. Common
Kotlin has no equivalent, andjava.net.URLEncoderis a different algorithm
(application/x-www-form-urlencoded) that is wrong for URIs.Encoding is verified over every one of the 1,114,112 code points against
node, not a sample: escaping is what stops untrusted text from changing a
URI's structure. Unpaired surrogates throwUriError.Decoding rejects what UTF-8 forbids — overlong forms, encoded surrogates,
code points past U+10FFFF, truncated escapes — since accepting them turns a
decoder into a filter bypass. Its input is text and cannot be enumerated, so
./gradlew uriFuzzcovers it: 600,000 cases clean, about half of them
rejections, and it runs nightly on three seeds.Five planted bugs are caught: a wrong unescaped set, accepting overlong
encodings, accepting encoded surrogates, decoding reserved escapes in
decodeUri, and substituting U+FFFD for a lone surrogate instead of throwing. -
Double.toEcmaString()inio.github.mgilbir.ecma262.number— JavaScript's
Number::toString, ECMA-262 6.1.6.1.20. No Kotlin target produces it:
Kotlin/JS does because it is JavaScript, but Kotlin/JVM and Kotlin/Native
print1.0,1.0E21and4.9E-324where JavaScript prints1,1e+21and
5e-324. Over 200,000 random doubles the JVM's string differs 98.4% of the
time.Almost all of that is layout — JavaScript stays positional out to 10^21 and in
to 10^-6 — but not all of it: a JDK 21Double.toStringis not shortest for
the smallest subnormals, so reusing the platform digits and re-laying them out
would be wrong in exactly the cases hardest to notice.The specification defines the result rather than an algorithm, which makes two
properties equivalent to it: the output round-trips, and no shorter decimal
does. Both are checked over random doubles, every power of two, the subnormal
range and short decimals, so the tests hold any implementation to the
specification rather than to this one. Four planted bugs — an extra digit, a
missing round-up, ignoring round-half-to-even, and dropping the asymmetric gap
below powers of two — are each caught by the property that should catch them.Ties escape both properties, since both candidates round-trip and are equally
short. Rounding them down instead of to the even significand costs 48 values
in 231,948; the differential fixture and an explicit test carry that rule.Implemented with the exact rational method of Steele & White as presented by
Burger & Dybvig: big integers, no lookup tables. Correctness first — it is 63x
slower thanjava.lang.Double.toStringat the extremes of the exponent range
and 6.8x slower for everyday values. A table-driven method would close that. -
String.toEcmaDouble()—StringToNumber, ECMA-262 7.1.4.1.1, the
conversionNumber("…")performs. The whole string must be a numeric
literal, so this is notparseFloat;0x/0o/0bliterals are accepted but
take no sign, and an empty or all-whitespace string is+0. All 68 grammar
and rounding cases are taken from node.Correctly rounded, which can turn on the 767th significant digit. Bounded
against the inputs that have historically hung decimal parsers: significant
digits are capped with the rest folded into a sticky flag, the exponent is
clamped as it is read, and out-of-range magnitudes resolve before any big
integer is built.Three planted bugs are caught — ties rounding up rather than to even,
accepting trailing garbage, and ignoring the sticky flag. The third initially
was not: no test exercised a value truncated at the cap that was also an
exact tie, because a genuine tie never needs more than 767 digits.
digitsBeyondTheCapBreakTiesconstructs one — the exact midpoint between 1
and the next double, padded past the cap, with a single digit beyond it that
decides the result. -
Formatting and parsing are now checked as inverses against each other, so the
round-trip property no longer leans on the host's decimal parser. -
Double.toEcmaString(radix)for radices 2 to 36. Unlike everything else
here this is compatibility rather than conformance: ECMA-262 calls the result
for a radix other than 10 implementation-approximated and defines nothing
further, so V8's behaviour is what is implemented and the 21,280 recorded
strings are the contract rather than a check on one. Radix 10 delegates to the
specified path. Four planted bugs are caught — dropping the round-half-to-even
step, removing the floor on the error term, skipping the zero fill above 2^53,
and stopping the fraction a digit early. -
Double.toEcmaFixed(),Double.toEcmaExponential()and
Double.toEcmaPrecision()—Number.prototype.toFixed,toExponentialand
toPrecision(21.1.3.3, 21.1.3.2, 21.1.3.5). Checked against node over a grid
of 4,054 values crossed with arguments from 0 to 100: 97,296 strings.All three round ties up, where
toStringrounds them to even. The
specification asks for exactly that difference, and a planted swap to
round-half-even is caught. So is losing the rounding carry —(99.995)
.toFixed(2)is"100.00", and taking the decimal exponent from a rounded
first digit instead of an unrounded one gave"999.95"until the grid caught
it — and getting thetoPrecisionexponential threshold off by one. -
Grisu3 as the fast path for
Double.toEcmaString(), with the exact method
kept as the fallback. It works in 64-bit arithmetic against a generated table
of cached powers of ten, tracks the rounding error it accumulates, and
declines when that error could change the answer — measured at 0.51% of
random doubles and 0% of subnormals and simple decimals.Correctness does not depend on it being right about which values are hard,
only on it refusing to answer when unsure.Grisu3AgreesWithExactTest
compares the two implementations directly, and asserts the fallback is still
reached so it cannot rot into dead code. Three planted bugs — never
declining, truncating instead of rounding in the 128-bit multiply, and
dropping the closer-lower-boundary rule at powers of two — are each caught by
both that test and the node differential.The cached powers are generated with
BigIntand checked to within half a
unit in the last place during generation, rather than transcribed.Against
java.lang.Double.toString: 63.5x slower becomes 3.5x for values at
the extremes of the exponent range, and 6.8x becomes 2.0x for everyday ones. -
A differential fixture for numbers, checked against node over 231,948 values:
a deterministic sweep of the bit space, every subnormal up to 20,000, every
power of two, and short decimals. Walked from an index on both sides, so the
fixture is 3 KB rather than a megabyte.
Testing
-
A nightly differential fuzz for the number functions, three seeds of 500,000
cases each, alongside the existing regex fuzz. Fresh cases rather than a
recorded set: random doubles throughtoString,toFixed,toExponential,
toPrecisionandtoString(radix), and random strings — signed, spaced,
radix-prefixed, corrupted, up to 1,800 digits long — through the parser.
900,000 cases run clean locally.Building it found a gap in itself. A planted bug that accepted a sign before a
radix prefix went undetected, because the generator almost never produced one;
it now emits signed and mixed-case radix literals, and the same plant is
caught within 435 cases. -
Test262's Number formatting tests, 205 cases across
toFixed,
toExponential,toPrecisionandtoString. Every expectation is the
literal from the suite's own assertion rather than an engine's answer, which
makes this the one source here that no implementation produced. The value is
really the inputs: corner cases chosen by people who knew where
implementations go wrong. The nightly regenerates the fixture and fails if
upstream has changed. -
A fourth V8 defect is recognised and skipped. A non-multiline
$lets V8
start its scan near the end of the input, using a minimum match length
counted in code points but applied to a UTF-16 index, so an astral tail
pushes the scan past a position that matches:/[^\w]$/ufinds nothing in
"\u{1F600}"while/^[^\w]$/u,/[^\w]$/uy,/[^\w]$/umand
/[^\w]$/vall match the same character. This engine matches over code
points throughout;EndAnchorAstralTestpins the behaviour down.The detector is behavioural rather than syntactic: it scans code-point
boundaries with a sticky copy of the pattern and reports a defect only when
sticky finds a match plainexecskipped, so it tests V8's actual
self-contradiction instead of a guess about which patterns trigger it. Over
500,000 fuzz cases it excluded exactly one, and no recorded corpus case.Found by a local fuzz run on
/([^\w]??){1,2}$\B(?!(?<=\b))/sug.
Available from Maven Central as io.github.mgilbir:ktecma262:0.2.0.