Skip to content

v0.4.0

Choose a tag to compare

@github-actions github-actions released this 04 Sep 00:23
· 56 commits to main since this release
v0.4.0
ac3fef1

M3 is complete. Text and numbers, which is the two conversions between them, the elementary functions, the bit level primitives underneath everything numeric, complex numbers, random numbers, and arbitrary precision integers, rationals and floats. Nine pull requests, four issues, and the parity count goes from 7.8 percent to 13.6 percent, which is 1,208 of Go's symbols across eighteen packages.

The thread that runs through all of it is that accuracy is a contract rather than a preference, and that the contract is what decided wrap against port every time the question came up. core.strconv prints the shortest decimal that reads back as the same float and reads back the nearest float to what was written. core.math holds every function to the tolerance Go's own test tables hold it to. Every operation in core.math.big's Float is correctly rounded, which means the true answer rounded once rather than approached through a sequence of roundings. Measured against those contracts, std.bit is exact and the thirty counting and reversing functions are one call each, while std.math.erf is off by a hundred and seventy million parts in the last place over Go's own inputs and is written out instead. The library calls Mojo's where Mojo's is right and ports Go's where it is not, and a measurement is what says which, not a preference for one or the other.

The other thread is that a port is worth what it is checked against, and that checking it against itself is worth nothing. Go's own test tables are now harvested out of the Go tree by tools/testgen rather than typed, and the tool fails rather than skipping when a table it was told to take is not there, since a table quietly disappearing after a Go upgrade would take its coverage with it and say nothing. Above that are the differential runs against a real Go process, which now include writing out and reading back all 4,294,967,296 float32 there are. And inside the suite are the independent oracles: the Float arithmetic is checked against a number held as a list of exponents, where adding is concatenating two lists, and the square root is checked by squaring the answer and measuring the distance back.

core.math.big, 166 of Go's 166 symbols with six waived. Arbitrary precision integers, rationals and floating point numbers, and the unsigned limb layer that all three are built on.

An Int is a sign and a list of sixty four bit limbs, a Rat is two Int values kept in lowest terms with the sign on the numerator, and a Float is a sign, a mantissa of as many bits as you ask for, and an exponent, with six rounding modes and an accuracy reported after every operation saying which way the answer was rounded. Every arithmetic result is correctly rounded, meaning it is the true answer rounded once rather than approached through a sequence of roundings.

The limb layer is Go's nat, split the way Go splits it. Multiplication is schoolbook below forty digits and Karatsuba above it with a squaring path of its own. Division is Knuth's algorithm D, with the divisor scaled so the two digit guess is never more than one too large and a reciprocal cached across the digits. Base conversion is a digit at a time for short numbers and a table of divisor powers for long ones. Exponentiation is a sliding window, with Montgomery multiplication for an odd modulus and Montgomery plus the Chinese remainder split for an even one. The greatest common divisor is Lehmer's, which does most of its work in single digit arithmetic on the leading digits. probably_prime is Baillie-PSW: the Miller-Rabin rounds the caller asked for, plus a base two round and an extra strong Lucas test that no known composite survives.

Two of Go's algorithms are not here, divRecursive and karatsubaSqr. Both are written entirely as overlapping subsections of one vector being read and written by the same call, which this port cannot express. Neither changes an answer, so a very large division and a very large squaring are slower here and correct.

There is no destination argument anywhere in the package, and that is the difference a caller meets first. z.Add(x, y) is x.add(y) returning a new value, because Mojo will not let one value arrive as both the mutable receiver and a borrowed argument, and z.Add(z, y) is exactly that. quo_rem, div_mod and gcd_ext have a second answer and write it through a mut argument, because a tuple would need Int to be implicitly copyable and a type holding a list cannot be. For a Float the same rule decides where the precision and the rounding mode come from: x.add(y) works at the wider of the two precisions and rounds by x's mode, which is what Go's new(Float).Add(x, y) does, and x.add(y, p) names the precision. copy keeps the source's precision and mode where set rounds to the receiver's, which is Go's own distinction between Copy and Set.

And, Or and Not are Mojo keywords, so they are __and__, __or__ and __invert__. Xor follows them to __xor__ and AndNot stays and_not. The type is named Int, so a file that imports it by name loses Mojo's own Int, and every test file and the package docstring say import core.math.big as big and spell it big.Int, which is what a Go caller writes anyway.

Every panic is a raise and so is every nil. That is division by zero, a negative bit index, a bit that is not zero or one, the square root of a negative number, a base outside two to MaxBase, an even second argument to jacobi, a negative round count to probably_prime, a fill_bytes buffer too small to hold the number, a zero denominator, the inverse of zero, a NaN or an infinity handed to Rat.set_float64, and the three places Go returns a nil Int for no inverse, no modular square root and a negative power of a number with no inverse. Go panics with an ErrNaN value in the six places reachable through a Float, which are two infinities of opposite signs added, two of the same sign subtracted, a zero times an infinity, a zero over a zero or an infinity over an infinity, the square root of a negative number, and a NaN handed to set_float64 or new_float, and all six raise with the ErrNaN code instead. Go's own documentation calls that panic a value carrying an error, which is a raise written the long way round.

A Rat's denominator is a one from the moment the value exists. Go leaves the zero value's empty and reads an empty one as a one, which costs it mulDenom, scaleDenom and half of norm. None of the three has a counterpart here and no number behaves differently, because Go normalises at the first assignment. The only visible trace is the gob encoding of a Rat nobody has assigned to, which Go writes with an empty denominator and this writes with a one. Both sides read both forms, so a value still crosses in either direction, and there is a test for reading Go's form.

Four more rows are in docs/deviations.md beside those. bits and set_bits copy, where Go documents them as raw access to the number's own array, and Rat.num and Rat.denom copy for the same reason, since a view of storage a value owns is the thing this library does not hand out anywhere. Rat.float64 and Rat.float32 report a number below the smallest subnormal as inexact, where Go reports 1/(1<<2000) as an exact zero, because a zero does not represent that number and the flag exists to say whether the float is the rational. Float.int and Float.rat raise for an infinity where Go returns a nil, since an infinity is not an integer and not a rational number and there is nothing to hand back. And Int.rand takes anything implementing core.math.rand.Source rather than a pointer to a Rand, so its digits differ from Go's for the same seed, because Go builds each sixty four bit digit from two thirty two bit draws left over from its v1 generator and core.math.rand is v2, whose sources are already sixty four bits wide.

The six waivers are Int.Format and Float.Format, which are fmt.Formatter, Int.Scan, Rat.Scan and Float.Scan, which are fmt.Scanner, and ErrNaN.Error. core.fmt takes its format string as a comptime parameter and typechecks the verbs when the program is compiled, so there is no run time State to hand the first five and no verb for them to look at; text and write_to do what Format did and set_string does what Scan did. ErrNaN.Error has no method to sit behind, because ErrNaN is an error code here rather than a struct.

One addition, Int.must_set_string. Go writes a constant modulus at package level and drops the boolean; a constant here is built inside a function and set_string raises, so without this a file holding a written down curve order would have to be a raising one. It aborts, so it is for a literal, and the linter refuses it on anything else. Writing the first must_ function in the tree found a hole in that check, since the line declaring one is not a call to one and its parameters are named rather than quoted, so the declaration reported itself.

191 tests, ported from int_test.go, intconv_test.go, intmarsh_test.go, nat_test.go, prime_test.go, rat_test.go, ratconv_test.go, ratmarsh_test.go, float_test.go, floatconv_test.go, floatmarsh_test.go, decimal_test.go, sqrt_test.go and bits_test.go. The arithmetic ones are the ones worth naming. They check the four Float operations against a second implementation that holds a number as a list of exponents, where adding is concatenating two lists and multiplying is every pairwise sum. It shares no code at all with the package under test, so agreement is evidence of correctness rather than of self consistency, and it is 3840 add and subtract combinations and 3840 multiply and divide combinations across eight numbers and twenty precisions. _bits.mojo is that second implementation and test_bits.mojo tests it first, since a broken oracle would agree with a broken package. The square root is checked three ways, against Go's 350 digit expansions, by squaring the answer and asking that the distance back to the number it came from is under one unit in the last place of the precision asked for, and against the machine square root at 53 bits on ten thousand random numbers. The text tests cross check against core.strconv on every row where the two have a syntax in common, so the formatting is confirmed by a completely different code path as well as by Go's table. Separately from all of that, Float was run against Go directly on about 34,800 comparison lines with no differences other than the documented ones.

Every table Go writes down was compared against Go's own source rather than trusted after typing, which caught two things no test could have: an extra digit in the 337 digit Arnault composite, where a mistyped composite is still a composite and the assertion still passes, and a wrong row in prodZZ.

The arithmetic is correct and it is slower than Go's. A modular exponentiation with a 2048 bit modulus takes 30.5 ms here against 9.44 ms for Go built with -tags math_big_pure_go and 5.47 ms for Go as it ships with its assembly, measured on a quiet Linux x86_64 machine. The ratio against portable Go is flat as the numbers grow, 3.05 at six words and 3.23 at thirty two, which rules out per call allocation from the no destination design and puts the whole gap inside the one multiply and add loop that Montgomery multiplication spends its time in. That loop is eight lines of ordinary Mojo over a Span[UInt64] and the code generated for it is what is behind, so it is filed as a toolchain issue rather than worked around here.

core.math.rand, all 59 of Go's symbols, no waivers. PCG and ChaCha8 as the two generators, Rand over either of them for everything derived from a uniform draw, Zipf for the fourth distribution, and the nineteen top level functions that need no generator at all.

This is math/rand/v2, named without the version because there is no version one here to tell it apart from. Go's original math/rand shares one locked generator between its top level functions, pays for a division to correct the bias in Int31n, and lets Seed make the whole program's randomness a global. None of that was worth carrying, and Go excluded it from this library on the same reasoning.

The top level functions are the reason the package declares unsafe, and only globals.mojo is. They need a generator that outlives the call, is not shared between threads, and is seeded by the operating system, and Mojo has nowhere to put one. So the C shim in core/errors/shim/slot.c grew a second slot beside the error record, the generator is allocated with malloc because a pthread key destructor is what frees it, and the seed comes from a raw getentropy. That is the whole of it, it is three functions long, and nothing a caller of Rand, PCG, ChaCha8 or Zipf touches goes near any of it. The dependency it adds is on core.errors, which every package in the library already has.

Six things differ from Go and each has a row in docs/deviations.md. A Rand owns its Source and a Zipf owns its Rand, where Go holds pointers and lets the caller go on drawing from the same generator alongside, which is the mistake neither language can check for. shuffle takes its swap as a compile time parameter, so the exchange is inlined rather than called through a pointer, the choice core.sort already made. Every panic is a raise, which is a bound that is not positive, a zero to uint64_n, a negative count to shuffle or perm, and a new_zipf with parameters that describe no distribution, where Go hands back a nil pointer for the caller to find by dereferencing it. n raises on a DType that is not integral, because Go constrains its type parameter and section 10 of docs/design.md is why that check cannot be a compile error here; it is a comptime if and folds away for every integer type. ChaCha8.read has io.Reader's signature exactly and declares nothing, because declaring the conformance would put core.io underneath this package and Go's own does not import io either.

The sixth is worth more than a line. ChaCha8.UnmarshalBinary in Go reads the readbuf: length byte and the block position straight out of the encoding and uses both without looking at them, so three shapes of untrusted input panic there with a slice bounds failure and a fourth restores a generator that silently skips forward in the stream. All four are refused here with ErrInvalidEncoding, and a refusal leaves the receiver alone. A marshalled generator is exactly the sort of thing that arrives from somewhere else, so the decoder is the wrong place to trust a length.

52 tests, ported from pcg_test.go, chacha8_test.go, rand_test.go, regress_test.go, auto_test.go and race_test.go. TestRegress is the most valuable of them and is three hundred and forty golden values, twenty from each of seventeen methods of a Rand on new_pcg(1, 2), with a fresh generator per method. Go builds that list by walking the methods with reflect. There is no reflection here, so the seventeen calls are written out in the order reflect walks them, which is alphabetical by Go's name, and each one reads the harvested list its return type went into at the offset its group starts at.

Go checks the ChaCha8.Read transcript by taking a sha256 of 2976 bytes and comparing it against a constant. There is no sha256 in this library yet and there did not need to be one for this: 2976 is 372 times 8, and the transcript is exactly the little endian expansion of chacha8output, which was confirmed against Go before it was relied on. So the read tests compare against the expansion, which pins the same bytes and also says which one is wrong when one is. The hash is harvested and waiting for core.crypto.sha256.

Harvesting those tables changed tools/testgen. Go's marshalled forms are byte strings rather than text, and a Mojo string literal is UTF-8, so a byte like \xb6 in a Go table would arrive as the two bytes 194 and 182 and a sixteen byte golden would come out eighteen bytes long. Byte strings are now emitted as hex through a generated _hex helper with the Go spelling above them as a comment. The two harvests that already existed, math and cmplx, came back byte identical afterwards, so the change reaches only tables that contain bytes.

Zipf has no test in Go, in either version of the package, and has had none since it landed in 2011. Its goldens here were generated from go1.26.7 instead: three parameter sets of twenty values each, all from new_pcg(3, 4) so that the same uniform draws feed all three and a difference between them is the parameters and nothing else. On top of those is a shape test, a hundred thousand draws asked whether the first three values come out in the ratio one to a quarter to a ninth, because a golden list taken from the implementation's own ancestor cannot notice a plausible looking wrong distribution.

The concurrency test found something about the sanitiser rather than about the package, and it is written down in tools/mojotest/run.py where the next person writing a threaded test will read it. The Mojo runtime serves every List out of its own allocator, the thread sanitiser does not intercept that allocator, so it never learns that a block one thread freed is the same block another thread was later handed, and a file containing nothing but a List built inside a thread reports a data race. The threaded test therefore leaves out perm, which is the one top level function that allocates, and perm draws through the same per thread generator as everything else. Everything else in the package runs from eight threads at once with the sanitiser clean.

core.math.cmplx, all 27 of Go's symbols, no waivers. The elementary functions on complex numbers, on std.complex.ComplexFloat64, which is Mojo's own type, so a value from this package is a value the rest of the language already works with.

The manifest said this would be a wrap over std.complex and it is a port of all twenty seven. The type and the functions turned out to be different questions. ComplexFloat64 is a good number type and carries the four operators, a conjugate and a norm, and none of the twenty seven, so there was nothing to wrap. Two of the three it does carry are not the ones Go uses either: norm() squares both parts and adds them, so it overflows at (1e308, 1e308) where the answer is 1.4e308 and underflows to zero at (1e-320, 1e-320) where the answer is 1.4e-320, and abs here is a hypotenuse and is right at both. Division is the naive formula and overflows the same way, which never comes up only because Go's math/cmplx never divides one complex number by another. The arithmetic is the type's, the functions are Go's.

The bodies are Cephes by way of Go and are short. The C99 Annex G special case switches in front of them are most of the source and all of the difficulty, because the answers there turn on the sign of a zero, on an infinity beside a not a number counting as an infinity rather than as a not a number, and on which side of a branch cut an argument arrived from. Go has one unreachable panic, in Pow, and it is reachable: IsNaN is false for a not a number sitting beside an infinity, so an exponent of (nan, inf) arrives with none of the three comparisons true. A not a number is the answer everywhere else a not a number exponent turns up and it is the answer there. polar returns Go's two values as a tuple that unpacks, the way frexp does.

The package depends on core.math and on core.math.bits, which puts it in tier 2. The dependency on bits is tan at a huge argument: reducing modulo pi in float64 has no correct digits left around 1e10, so the reduction runs in three hundred and twenty bits of an integer pi through mul64, add64 and leading_zeros64, and Go's tanHuge table holds it to three parts in the last place afterwards.

Building this changed core.math.hypot from a wrap into a port, and the reason is one worth keeping. It was the system routine on the grounds that sqrt(p*p + q*q) has one right answer and nothing to disagree about, and the system routine is not required to give it. On macOS hypot(-16.688001947990161, -24.14182405800559) comes back as 29.348203997243527 where the correctly rounded answer is 29.348203997243523, and Go's two lines give the correctly rounded one. A part in the last place would not matter if hypot were only ever the length of something, but sqrt here divides by it and then subtracts two nearly equal numbers, and at asinh(-5.010603618271075 + 9.636293707198417i) that one bit arrives fifty bits wide, which is thirty times the tolerance Go holds the function to. Go's version is reproducible and the system's is whatever the platform ships, so this package now gives the same answer everywhere.

24 tests, ported from cmath_test.go and huge_test.go. Go tests this package from the inside rather than from a _test package, so a NaN() in one of its tables is cmplx.NaN and not math.NaN, and tools/testgen grew the piece that harvests tables out of an in-package test file and dot imports the package under test so that the names in them keep meaning what they meant. Each function gets Go's general table at Go's own tolerance for it, then the special cases, then the two symmetries Go checks on every special case, f(conj(z)) == conj(f(z)) and f(-z) == -f(z) for the odd ones, and then the branch point table, which is pairs of points either side of a cut asked to give the same answer at the cut.

core.math, all 97 of Go's symbols, no waivers. Thirty constants and sixty seven functions, every one of them on float64, as Go's are.

The manifest said this would be a wrap over std.math and about half of it is a port. The reason is that Mojo's is a throughput first library and Go's tests are an accuracy contract. Measured against libm over Go's own inputs, std.math.erf is off by a hundred and seventy million parts in the last place, pow by ten million and log by two million, while the tolerances Go holds its math to are forty five parts and two. So exp, exp2, log, log1p, pow, erf, erfc, cosh and gamma are transcriptions of Go's, and Go's own test tables are what says they are close enough now. The functions that were already accurate enough are still one call each, because what is inside libm is FDLIBM and what Go wrote down is a transcription of the same code, so a third copy would add a place for a bug to live and take nothing away from either of the other two.

Accuracy was the easy half. The harder half is that a system library and Go disagree at the special values, and the special values are exactly what a caller reaches for math rather than for an arithmetic operator to get right. std.math.sinh is a not a number at negative infinity and std.math.tanh is one at a not a number, so both are ported whole. std.math.j1 hands back a negative zero at negative infinity, correctly for an odd function and not what Go returns, so j0, j1, y0 and y1 keep the system library for the arithmetic and answer Go's special cases in front of it. lgamma in the system library is positive infinity at both infinities where Go's is the argument. And gamma had to be ported outright rather than patched at the edges, because std.math.gamma(-170.5) is out by four parts in fourteen digits, four times the tolerance Go allows that function and about two hundred places in the last bit, and -170.5 is not a special value at all, just an ordinary argument in the range where the reflection formula runs.

Nine more are here because there was nothing to call: erfinv and erfcinv, which are a rational approximation out of a statistics paper, jn and yn, which are Bessel recurrences, ilogb, pow10, round, round_to_even and the sign half of lgamma, which in C is a global variable that the call sets as a side effect and which Go turned into a second return value.

Two of the constants are written as bit patterns rather than as the decimals Go writes, because a float literal is flushed to zero on its way into a Float32 or a Float64 when it is subnormal, and the smallest nonzero float of a width is subnormal by definition. MAX_FLOAT64 and MAX_FLOAT32 are written out too, and not taken from Float64.MAX, which in Mojo is positive infinity rather than the largest finite value. The fifteen integer limits have the type their name says instead of being untyped as Go's are. All four of those have a row in docs/deviations.md.

The eleven irrational constants needed none of that, and finding out why corrected something this project had written down wrongly. A comptime expression in Mojo is evaluated at arbitrary precision and rounded once, exactly as a Go untyped constant is, so LOG2E could have been written as 1 / LN2 here just as Go writes it. Between two float64s the same division rounds twice, and for LOG10E that costs a bit in the last place. tests/math/test_const.mojo asserts both spellings so the fact stays pinned. It matters beyond the constants: Go's large angle trig tests add 100000 * Pi to every input, and writing that as a product of two float64 values lands an ulp low, which is 6e-11 there and enough to move the fourth digit of the sine.

The package depends on nothing in core. Its manifest recorded core.math.bits from the day the roadmap was written, and the lint caught that nothing imports it: Go's math reaches for bits.Mul64 and bits.LeadingZeros64 in fma.go and trig_reduce.go, and here the fused multiply add is a hardware instruction and the argument reduction is inside libm, so neither call site exists. Dropping the dependency puts core.math in tier 0 and moved eleven other manifests down by one, core.math.cmplx by two, since the lint computes a tier from the graph rather than trusting what a manifest records. None of them changed in any other way. That also found a bug in pixi run pkg, which builds every package against only what it declares: it copied a package's whole subtree into the scratch build, so core.math was being asked to build core.math.bits as well and to declare its dependencies. A package is its own source files, and the copy now stops at a directory with a PACKAGE.toml in it.

104 tests, ported from all_test.go and huge_test.go. This is the first package whose tables are harvested rather than typed, so tools/testgen grew the piece that does it: extract.go parses the Go test source with go/ast and prints the literal tables as Mojo, plan.toml names what to harvest, and tests/generated/math.mojo is the checked in result, regenerated by pixi run -e oracle testgen math against the Go release that pixi.lock pins rather than whichever one is on PATH. The tool fails rather than skipping when a table named in the plan is not in the Go tree, because a table quietly disappearing after a Go upgrade would take its test coverage with it and say nothing. third_party/go/LICENSE carries the BSD terms the harvested data comes under. Go's own tolerance for each function is copied along with its table, so sin is checked to two parts in the last place and pow to forty five, and where Go compares with != these use a helper that also catches a zero of the wrong sign.

core.math.bits, all 50 of Go's symbols, no waivers. Bit counting and the wide arithmetic that a multiple precision integer is built out of.

This is the shortest package in the library and the reason is worth writing down. Go implements every one of these by hand, out of generated byte tables, de Bruijn constants and Knuth's algorithm D, and it does that because the Go compiler recognises some of these functions on some architectures and the written version is what the rest fall back to. Mojo has std.bit for the counting and the reversing and UInt128 for the arithmetic, so the twenty counting functions and the ten reversing functions are one call each, and the wide sum, difference, product and quotient are the ordinary operation done one width up and split in half. That is not a shortcut past Go's work, it is Go's work already done by the compiler. The tests take Go's own tables to say the answers are the same, including the byte table Go builds by shifting one bit at a time, which is exactly the independent implementation you want on the other side of a comparison.

The five rotations are the exception and are written out, because std.bit.rotate_bits_left wants the distance at compile time and Go's takes it at run time. Both shifts are masked, not just the first, since a shift by the full width is undefined below Mojo and a rotation by zero would otherwise land on it.

The six functions that can fail raise rather than panic. Go's zero divisor and its too wide quotient both come out as the runtime's own integer divide by zero and integer overflow, which read as though the hardware trapped when in fact the package checked. Here they are ErrDivideByZero and ErrOverflow, two new codes. A package in tier 1 is not the one that gets to end the process.

The package depends on core.errors and nothing else, which puts it under core.math and under core.math.big, the caller the arithmetic half exists for. That moved twelve other manifests down by one, since the lint computes a tier from the graph rather than trusting what a manifest records. None of them changed in any other way.

20 tests, ported from bits_test.go. UintSize is written down as 64 rather than derived from ^uint(0) >> 63, because every platform in docs/platforms.md is 64 bit, and it needed a line in parity/renames.toml to become UINT_SIZE on the same rule strconv.IntSize took.

core.strconv, all 43 of Go's symbols, no waivers. Numbers to text and back, and Go's quoted string syntax. Every float this library prints is the shortest decimal that reads back as the same float, and every float it reads is the nearest one to what was written, which are the two properties the package exists for and the two that are easy to get almost right.

The floating point half is a port of what Go's internal/strconv does as of go1.26.7, down to the tables. Parsing is Eisel-Lemire with an exact big decimal underneath it for the inputs the fast path declines, formatting is Dragonbox for the shortest form, a fixed digit path for up to eighteen digits, and the same big decimal for anything wider. The tables are generated by tools/gen/strconv.py from exact rational arithmetic, the way Go generates its own with math/big, and they are checked in and rerun by pixi run generated-check. Go's own test corpus passes row for row, and three new differential areas run against Go every night. strconv-floats writes generated floats out in eighteen ways and reads the shortest form back, strconv-parse feeds both parsers the same generated strings across eight shapes from plain digits to four hundred digit monsters, and strconv-float32 writes out and reads back all 4,294,967,296 float32 there are and compares a hash of every string produced. All three agree byte for byte.

The package depends on core.errors, core.unicode and core.unicode.utf8 and nothing else, which puts it in tier 1. That means it does not use core.math: the handful of routines it needs, a 64 by 64 to 128 bit multiply and a few bit counts, are in its own math.mojo, which is exactly what Go does and for the same reason, since core.math will want to format a float long before this package could be moved above it. Lowering the tier moved fourteen other manifests down by one, since the lint computes a tier from the graph rather than trusting what a manifest records. None of them changed in any other way.

A failure is a raise, so the clamped value Go hands back next to ErrRange has nowhere to go. errors.matches(e, strconv.ErrRange) says which limit was hit, the sign is in the input, and NumError.of(e) recovers the call, the input and the code, which is Go's NumError with the parts it had. Two codes are additions: ErrBase and ErrBitSize. Go folds both into a message with no sentinel behind it, so a caller cannot tell an impossible base from malformed digits, and the two are different bugs to go and fix, because the base came from the program and the digits came from its input. The two places Go panics, a bad bit size in FormatFloat and in FormatComplex, raise here.

parse_complex and format_complex work on std.complex.ComplexFloat64, because Go's complex128 is a language type and Mojo's complex is a library one. It is the standard library's type rather than core.math.cmplx's, since that package sits above this one and could not be depended on from here. The append forms take mut dst: List[UInt8] and return the byte count instead of returning the grown list, which is the rule the rest of this library already follows.

54 tests, ported from atob_test.go, atof_test.go, atoi_test.go, atoc_test.go, ftoa_test.go, ctoa_test.go, itoa_test.go and quote_test.go. The tables that Go writes with hexadecimal float literals are written as exact decimals with the Go spelling in a comment beside them, since Mojo has no such literal, and powers of two are built from bits rather than through ldexp, which writes the exponent field directly and so gives a wrong answer as soon as the result is subnormal. The tests Go writes against a hook it exports only for testing are not here; the differential run covers the same ground.

core.strings, 82 of Go's 82 symbols with one waived. The waiver is Title, the same one core.bytes took, deprecated in Go itself because the word boundary rule turns "they're" into "They'Re". to_title is here and is the rune wise mapping, which was never the broken part.

The whole package is built on one decision: nothing here reimplements a search. Every function that looks for something calls core.bytes over the string's own bytes and turns the byte offset it gets back into a slice. Go has two packages with the same twenty functions in them because []byte(s) copies the string, so calling one from the other would cost an allocation per call. A Mojo String already holds UTF-8 and as_bytes() lends those bytes without copying, so the reason for the second copy of the code is gone. index, split, trim, fields and the rest are thin, and the Rabin-Karp and the cutset bitmap that they run on live in one place and are tested in one place.

Turning a byte offset back into a string is where the safety comes from. s[byte=a:b] aborts if either bound is not a code point boundary, so every slice this package hands out is one that a search found, and a bug in an offset is a stopped program rather than a String holding half a rune. That is a language guarantee this library did not have to build.

There is no len. count_bytes, count_runes and count_graphemes are three names for the three answers, and which one a caller wants is a question that has a different answer for every one of them: "a👩‍👩‍👧‍👦é" is 28 bytes, 9 runes and 3 graphemes. Mojo makes len(s) a compile error and it is right to.

Every signature binds its origin as ImmOrigin rather than as a plain Origin. That is not a detail. Two spans over the same mutable origin cannot be passed as two arguments of one call, so with the plain bound trim(s, s) and has_prefix(s, s) would not compile, and a caller would meet the exclusivity checker while doing something obviously harmless. Declaring the bound immutable removes the restriction, and nothing in the package writes through a span anyway.

Builder is not copyable, so the mistake Go catches with a runtime panic does not compile here. string() raises when what was written is not valid UTF-8 and bytes() is the accessor that never refuses, the same split bufio.Scanner uses. reset() keeps the allocation, unlike Go's, which drops it, because the builder is not copyable and hands out nothing that views its storage, so there is no way to read the old bytes and no reason to give them back. What it costs is that string() copies: Go's reinterprets the accumulated bytes in place through unsafe, and here the String has to own what it holds.

Replacer is one trie rather than Go's four implementations. Go picks between a single byte table, a byte to string table, a single string search and a generic trie by inspecting the pairs, which is four bodies of code to keep in agreement about a subtle rule: replacements happen in the order they appear in the target, matches do not overlap, and at the same position the pair that came first in the argument list wins, which is priority and not longest match. One trie answers all of that once. The nodes are Int indices into a list rather than pointers, since a recursive struct cannot be written. Two visible differences from Go: the constructor takes a list of pairs, so the odd argument count Go panics on cannot be written down, and the trie is built by the constructor rather than on first use behind a sync.Once, because replace takes self by borrow and there is no interior mutability to build through. write_string builds the whole result and writes it once instead of streaming each piece, so a writer that fails sees one call rather than a half written result.

Reader carries its origin in the type, so reset takes a slice from the same string the reader was built over, exactly as bytes.Reader and bufio.Reader do. All eleven of Go's methods are there.

78 tests, ported from strings_test.go, builder_test.go, replace_test.go and reader_test.go. The tables that Go writes with invisible characters in them are built with chr(0x85), chr(0xA0), chr(0x3000) and chr(0x200B) instead of pasted literals, which keeps a source file that no editor and no tool can silently rewrite. test_trimming_the_same_string_twice exists only to prove the ImmOrigin bound does what the paragraph above says.

Adding core.bytes to this package's dependencies moved it from tier 2 to tier 3, and the lint computes a package's tier from the graph rather than trusting what it records, so 49 other manifests move up by one behind it. None of them changed in any other way.

What's Changed

  • core.strings: the function set over text, sharing bytes' search code (#20) by @tamnd in #122
  • core.strconv: numbers to text and back, correctly rounded both ways (#21) by @tamnd in #123
  • core.math.bits: bit counting and the wide arithmetic big is built on by @tamnd in #124
  • core.math: the constants and the functions, at Go's tolerances (#22) by @tamnd in #125
  • core.math.cmplx: twenty seven functions on Mojo's complex type by @tamnd in #126
  • core.math.rand: PCG, ChaCha8 and the distributions over them by @tamnd in #127
  • core.math.big: Int, over a limb layer ported from Go's nat by @tamnd in #128
  • core.math.big: Rat, Go's rational number type (#22) by @tamnd in #129
  • core.math.big: Float, Go's arbitrary precision floating point type by @tamnd in #130
  • changelog: core.math.big, and cut v0.4.0, the end of M3 by @tamnd in #132

Full Changelog: v0.3.0...v0.4.0