Skip to content

Releases: tamnd/mojo.core

v0.7.0

Choose a tag to compare

@tamnd tamnd released this 10 Sep 04:07
v0.7.0
27df258

What's Changed

  • core.encoding: the six marshalling traits by @tamnd in #190
  • core.encoding.base64: the whole of Go's package, ported by @tamnd in #191
  • core.encoding.base32: the whole of Go's package, ported by @tamnd in #192
  • core.encoding.hex and core.encoding.ascii85, both ported whole by @tamnd in #193
  • encoding/pem: port the block format by @tamnd in #194
  • encoding/csv: port the reader and the writer by @tamnd in #195
  • encoding/xml: port the tokenizer and the token encoder by @tamnd in #196
  • encoding/json: port the scanner, compact, indent and Number by @tamnd in #197
  • encoding/json: run JSONTestSuite in full by @tamnd in #198
  • encoding/json: read a stream one token at a time by @tamnd in #199
  • encoding/json: read a whole document into an arena by @tamnd in #200
  • encoding/binary: the byte orders and the varints by @tamnd in #201
  • encoding/json: one runtime for every generated codec by @tamnd in #202
  • encoding/asn1: the DER reader by @tamnd in #203
  • encoding/json: a value read later by @tamnd in #204
  • encoding/asn1: the DER writer by @tamnd in #205
  • encoding/asn1: read and write Mozilla's root store again by @tamnd in #206
  • encoding/binary: read and write whole values by @tamnd in #207
  • encoding/asn1: one value in and one value out by @tamnd in #208
  • encoding/json: the writer half by @tamnd in #210
  • encoding/json: the error types by @tamnd in #211
  • encoding/xml: the value half by @tamnd in #212
  • changelog: cut v0.7.0, the end of M6 by @tamnd in #213

Full Changelog: v0.6.0...v0.7.0

What's Changed

  • core.encoding: the six marshalling traits by @tamnd in #190
  • core.encoding.base64: the whole of Go's package, ported by @tamnd in #191
  • core.encoding.base32: the whole of Go's package, ported by @tamnd in #192
  • core.encoding.hex and core.encoding.ascii85, both ported whole by @tamnd in #193
  • encoding/pem: port the block format by @tamnd in #194
  • encoding/csv: port the reader and the writer by @tamnd in #195
  • encoding/xml: port the tokenizer and the token encoder by @tamnd in #196
  • encoding/json: port the scanner, compact, indent and Number by @tamnd in #197
  • encoding/json: run JSONTestSuite in full by @tamnd in #198
  • encoding/json: read a stream one token at a time by @tamnd in #199
  • encoding/json: read a whole document into an arena by @tamnd in #200
  • encoding/binary: the byte orders and the varints by @tamnd in #201
  • encoding/json: one runtime for every generated codec by @tamnd in #202
  • encoding/asn1: the DER reader by @tamnd in #203
  • encoding/json: a value read later by @tamnd in #204
  • encoding/asn1: the DER writer by @tamnd in #205
  • encoding/asn1: read and write Mozilla's root store again by @tamnd in #206
  • encoding/binary: read and write whole values by @tamnd in #207
  • encoding/asn1: one value in and one value out by @tamnd in #208
  • encoding/json: the writer half by @tamnd in #210
  • encoding/json: the error types by @tamnd in #211
  • encoding/xml: the value half by @tamnd in #212
  • changelog: cut v0.7.0, the end of M6 by @tamnd in #213

Full Changelog: v0.6.0...v0.7.0

v0.6.0

Choose a tag to compare

@tamnd tamnd released this 06 Sep 13:24
v0.6.0
3d2d828

What's Changed

  • core.syscall: bindings generated from the host headers (#27) by @tamnd in #140
  • core.syscall: open takes a mode, and fcntl is bound (#139) by @tamnd in #141
  • core.syscall: reading a directory (#142) by @tamnd in #143
  • core.syscall: reading a clock by @tamnd in #148
  • core.path: the nine lexical path functions (#146) by @tamnd in #147
  • core.time: instants, durations and the calendar (#149) by @tamnd in #150
  • core.time: Location and the compiled zone file reader (#151) by @tamnd in #152
  • core.time: a Time that carries its location (#153) by @tamnd in #154
  • core.time: finding a zone file on the host (#155) by @tamnd in #156
  • core.time: the layout language and format by @tamnd in #159
  • core.time: reading a layout back, parse and parse_duration by @tamnd in #161
  • core.time: the marshalling methods, go_string, is_dst and local by @tamnd in #163
  • core.time.tzdata: the zone database compiled in by @tamnd in #165
  • core.path.filepath: the lexical half, and fs.valid_path by @tamnd in #167
  • core.io.fs, core.os: FileMode, FileInfo and the path errors by @tamnd in #170
  • core.os: File, the open flags and the standard streams by @tamnd in #172
  • core.os, core.io.fs: DirEntry, read_dir and the directory methods on File by @tamnd in #175
  • core.os: the path calls, read_file and write_file by @tamnd in #179
  • core.os: the environment, expand and the user directories by @tamnd in #180
  • core.os: the process ids, args, hostname, executable and exit by @tamnd in #181
  • core.os: create_temp, mkdir_temp and chtimes by @tamnd in #182
  • filepath: abs, eval_symlinks, glob, walk and walk_dir by @tamnd in #183
  • io/fs: the FS trait, and os.dir_fs over a real disk by @tamnd in #184
  • core.os.exec and core.os.signal by @tamnd in #185
  • core.os.Root, copy_fs, and the rest of core.os by @tamnd in #186
  • core.time.sleep, and the timers deferred to M9 by @tamnd in #187
  • changelog: cut v0.6.0, the end of M5 by @tamnd in #189

Full Changelog: v0.5.0...v0.6.0

What's Changed

  • core.syscall: bindings generated from the host headers (#27) by @tamnd in #140
  • core.syscall: open takes a mode, and fcntl is bound (#139) by @tamnd in #141
  • core.syscall: reading a directory (#142) by @tamnd in #143
  • core.syscall: reading a clock by @tamnd in #148
  • core.path: the nine lexical path functions (#146) by @tamnd in #147
  • core.time: instants, durations and the calendar (#149) by @tamnd in #150
  • core.time: Location and the compiled zone file reader (#151) by @tamnd in #152
  • core.time: a Time that carries its location (#153) by @tamnd in #154
  • core.time: finding a zone file on the host (#155) by @tamnd in #156
  • core.time: the layout language and format by @tamnd in #159
  • core.time: reading a layout back, parse and parse_duration by @tamnd in #161
  • core.time: the marshalling methods, go_string, is_dst and local by @tamnd in #163
  • core.time.tzdata: the zone database compiled in by @tamnd in #165
  • core.path.filepath: the lexical half, and fs.valid_path by @tamnd in #167
  • core.io.fs, core.os: FileMode, FileInfo and the path errors by @tamnd in #170
  • core.os: File, the open flags and the standard streams by @tamnd in #172
  • core.os, core.io.fs: DirEntry, read_dir and the directory methods on File by @tamnd in #175
  • core.os: the path calls, read_file and write_file by @tamnd in #179
  • core.os: the environment, expand and the user directories by @tamnd in #180
  • core.os: the process ids, args, hostname, executable and exit by @tamnd in #181
  • core.os: create_temp, mkdir_temp and chtimes by @tamnd in #182
  • filepath: abs, eval_symlinks, glob, walk and walk_dir by @tamnd in #183
  • io/fs: the FS trait, and os.dir_fs over a real disk by @tamnd in #184
  • core.os.exec and core.os.signal by @tamnd in #185
  • core.os.Root, copy_fs, and the rest of core.os by @tamnd in #186
  • core.time.sleep, and the timers deferred to M9 by @tamnd in #187
  • changelog: cut v0.6.0, the end of M5 by @tamnd in #189

Full Changelog: v0.5.0...v0.6.0

v0.5.0

Choose a tag to compare

@github-actions github-actions released this 04 Sep 07:51
v0.5.0
fe89517

What's Changed

  • tools/docjson: read mojo doc JSON as the reflection substitute by @tamnd in #134
  • tools/codec: generate JSON encoders and decoders (#24) by @tamnd in #135
  • core.fmt: the compile time half of Printf (#25) by @tamnd in #136
  • core.fmt: the runtime path, and %v without reflection (#26) by @tamnd in #137
  • changelog: the generators, and cut v0.5.0, the end of M4 by @tamnd in #138

Full Changelog: v0.4.0...v0.5.0

v0.4.0

Choose a tag to compare

@github-actions github-actions released this 04 Sep 00:23
v0.4.0
ac3fef1

M3 is complete. Text and numbers, which is the two conversions between them, the elementary functions, the bit level primitives underneath everything numeric, complex numbers, random numbers, and arbitrary precision integers, rationals and floats. Nine pull requests, four issues, and the parity count goes from 7.8 percent to 13.6 percent, which is 1,208 of Go's symbols across eighteen packages.

The thread that runs through all of it is that accuracy is a contract rather than a preference, and that the contract is what decided wrap against port every time the question came up. core.strconv prints the shortest decimal that reads back as the same float and reads back the nearest float to what was written. core.math holds every function to the tolerance Go's own test tables hold it to. Every operation in core.math.big's Float is correctly rounded, which means the true answer rounded once rather than approached through a sequence of roundings. Measured against those contracts, std.bit is exact and the thirty counting and reversing functions are one call each, while std.math.erf is off by a hundred and seventy million parts in the last place over Go's own inputs and is written out instead. The library calls Mojo's where Mojo's is right and ports Go's where it is not, and a measurement is what says which, not a preference for one or the other.

The other thread is that a port is worth what it is checked against, and that checking it against itself is worth nothing. Go's own test tables are now harvested out of the Go tree by tools/testgen rather than typed, and the tool fails rather than skipping when a table it was told to take is not there, since a table quietly disappearing after a Go upgrade would take its coverage with it and say nothing. Above that are the differential runs against a real Go process, which now include writing out and reading back all 4,294,967,296 float32 there are. And inside the suite are the independent oracles: the Float arithmetic is checked against a number held as a list of exponents, where adding is concatenating two lists, and the square root is checked by squaring the answer and measuring the distance back.

core.math.big, 166 of Go's 166 symbols with six waived. Arbitrary precision integers, rationals and floating point numbers, and the unsigned limb layer that all three are built on.

An Int is a sign and a list of sixty four bit limbs, a Rat is two Int values kept in lowest terms with the sign on the numerator, and a Float is a sign, a mantissa of as many bits as you ask for, and an exponent, with six rounding modes and an accuracy reported after every operation saying which way the answer was rounded. Every arithmetic result is correctly rounded, meaning it is the true answer rounded once rather than approached through a sequence of roundings.

The limb layer is Go's nat, split the way Go splits it. Multiplication is schoolbook below forty digits and Karatsuba above it with a squaring path of its own. Division is Knuth's algorithm D, with the divisor scaled so the two digit guess is never more than one too large and a reciprocal cached across the digits. Base conversion is a digit at a time for short numbers and a table of divisor powers for long ones. Exponentiation is a sliding window, with Montgomery multiplication for an odd modulus and Montgomery plus the Chinese remainder split for an even one. The greatest common divisor is Lehmer's, which does most of its work in single digit arithmetic on the leading digits. probably_prime is Baillie-PSW: the Miller-Rabin rounds the caller asked for, plus a base two round and an extra strong Lucas test that no known composite survives.

Two of Go's algorithms are not here, divRecursive and karatsubaSqr. Both are written entirely as overlapping subsections of one vector being read and written by the same call, which this port cannot express. Neither changes an answer, so a very large division and a very large squaring are slower here and correct.

There is no destination argument anywhere in the package, and that is the difference a caller meets first. z.Add(x, y) is x.add(y) returning a new value, because Mojo will not let one value arrive as both the mutable receiver and a borrowed argument, and z.Add(z, y) is exactly that. quo_rem, div_mod and gcd_ext have a second answer and write it through a mut argument, because a tuple would need Int to be implicitly copyable and a type holding a list cannot be. For a Float the same rule decides where the precision and the rounding mode come from: x.add(y) works at the wider of the two precisions and rounds by x's mode, which is what Go's new(Float).Add(x, y) does, and x.add(y, p) names the precision. copy keeps the source's precision and mode where set rounds to the receiver's, which is Go's own distinction between Copy and Set.

And, Or and Not are Mojo keywords, so they are __and__, __or__ and __invert__. Xor follows them to __xor__ and AndNot stays and_not. The type is named Int, so a file that imports it by name loses Mojo's own Int, and every test file and the package docstring say import core.math.big as big and spell it big.Int, which is what a Go caller writes anyway.

Every panic is a raise and so is every nil. That is division by zero, a negative bit index, a bit that is not zero or one, the square root of a negative number, a base outside two to MaxBase, an even second argument to jacobi, a negative round count to probably_prime, a fill_bytes buffer too small to hold the number, a zero denominator, the inverse of zero, a NaN or an infinity handed to Rat.set_float64, and the three places Go returns a nil Int for no inverse, no modular square root and a negative power of a number with no inverse. Go panics with an ErrNaN value in the six places reachable through a Float, which are two infinities of opposite signs added, two of the same sign subtracted, a zero times an infinity, a zero over a zero or an infinity over an infinity, the square root of a negative number, and a NaN handed to set_float64 or new_float, and all six raise with the ErrNaN code instead. Go's own documentation calls that panic a value carrying an error, which is a raise written the long way round.

A Rat's denominator is a one from the moment the value exists. Go leaves the zero value's empty and reads an empty one as a one, which costs it mulDenom, scaleDenom and half of norm. None of the three has a counterpart here and no number behaves differently, because Go normalises at the first assignment. The only visible trace is the gob encoding of a Rat nobody has assigned to, which Go writes with an empty denominator and this writes with a one. Both sides read both forms, so a value still crosses in either direction, and there is a test for reading Go's form.

Four more rows are in docs/deviations.md beside those. bits and set_bits copy, where Go documents them as raw access to the number's own array, and Rat.num and Rat.denom copy for the same reason, since a view of storage a value owns is the thing this library does not hand out anywhere. Rat.float64 and Rat.float32 report a number below the smallest subnormal as inexact, where Go reports 1/(1<<2000) as an exact zero, because a zero does not represent that number and the flag exists to say whether the float is the rational. Float.int and Float.rat raise for an infinity where Go returns a nil, since an infinity is not an integer and not a rational number and there is nothing to hand back. And Int.rand takes anything implementing core.math.rand.Source rather than a pointer to a Rand, so its digits differ from Go's for the same seed, because Go builds each sixty four bit digit from two thirty two bit draws left over from its v1 generator and core.math.rand is v2, whose sources are already sixty four bits wide.

The six waivers are Int.Format and Float.Format, which are fmt.Formatter, Int.Scan, Rat.Scan and Float.Scan, which are fmt.Scanner, and ErrNaN.Error. core.fmt takes its format string as a comptime parameter and typechecks the verbs when the program is compiled, so there is no run time State to hand the first five and no verb for them to look at; text and write_to do what Format did and set_string does what Scan did. ErrNaN.Error has no method to sit behind, because ErrNaN is an error code here rather than a struct.

One addition, Int.must_set_string. Go writes a constant modulus at package level and drops the boolean; a constant here is built inside a function and set_string raises, so without this a file holding a written down curve order would have to be a raising one. It aborts, so it is for a literal, and the linter refuses it on anything else. Writing the first must_ function in the tree found a hole in that check, since the line declaring one is not a call to one and its parameters are named rather than quoted, so the declaration reported itself.

191 tests, ported from int_test.go, intconv_test.go, intmarsh_test.go, nat_test.go, prime_test.go, rat_test.go, ratconv_test.go, ratmarsh_test.go, float_test.go, floatconv_test.go, floatmarsh_test.go, decimal_test.go, sqrt_test.go and bits_test.go. The arithmetic ones are the ones worth naming. They check the four Float operations against a second implementation that holds a number as a list of exponents, where adding is concatenating two lists and multiplying is every pairwise sum. It shares no code at all with the package under test, so agreement is evidence of correctness rather than of self consistency, and it is 3840 add and subtract combinations and 3840 multiply and divide combinations across eight numbers and twenty precisions. _bits.mojo is that second implementation and test_bits.mojo tests it first, since a broken oracle would agree with a broken package. The square root ...

Read more

v0.3.0

Choose a tag to compare

@tamnd tamnd released this 03 Sep 02:06
v0.3.0
99c2560

M2 is complete. Reading and writing, buffering, ordering, and the two packages that everything textual is built on. Eight pull requests, seven issues, and the parity count goes from 6.7 percent to 7.8 percent, which is 691 of Go's symbols across eleven packages.

The thread that runs through all of it is that Go's interfaces are checked at run time and these are not. io.copy in Go type asserts its argument to io.WriterTo and then to io.ReaderFrom, and a caller cannot find out ahead of time which path it will take. Here every reader and writer answers capabilities() with a set of bits, the fast paths are selected from those bits, and a caller can ask the same question the library asks. The erased forms exist for when the type genuinely is not known until run time, and they carry the same bits, so nothing has to be discovered by trying.

The other thread is that a for loop in Mojo swallows an error raised out of __next__, which was found with a probe before anything depended on it. That one language fact is why core.iter.Cursor exists and why fallible iteration in this library is written as an explicit has_next and next pair rather than as a loop. The lint enforces it.

core.bytes, 97 of Go's 99 symbols with two waived and one renamed. That closes M2. The waivers are Title, which Go's own documentation deprecates because its word boundary rule turns "they're" into "They'Re", and Buffer.AvailableBuffer, which hands out the buffer's spare capacity to be appended into and has nowhere to land here. The rename is MinRead to MIN_READ, on the constants rule the rest of the library already follows.

The rule the package is built around is one sentence: a method never hands out a view into the buffer's own storage. Go documents Buffer.Bytes, Next and Peek as valid only until the next write, and a Go program that breaks that rule reads stale bytes out of a live allocation. Here the next write can reallocate, so the same mistake reads freed memory, and probes/span_outlives_its_owner.mojo pins that the compiler does not stop it. So all three return owned bytes and there is no view returning accessor beside them under any name. The copy is real and Buffer is a hot type, so there are four ways not to pay for it, and the docstring leads with them: len(), write_to, read into a span the caller already owns, and string(), which builds the String directly rather than through a List[Byte] first. The functions that carve a slice up, trim and cut and split and fields, do return spans, exactly as Go returns subslices: those borrow the caller's bytes rather than the package's own.

Everything Go panics about raises instead. There are four: Repeat on a negative count or an overflowing length, Join on an overflowing length, Buffer.Grow on a negative one, and Buffer.Truncate past what is buffered. Each is a check the caller could have made. Running out of memory raises ErrTooLarge rather than panicking with it, because a buffer that cannot grow is a condition only the caller can report on, since they are the one who knows whether the input that caused it was theirs.

Two more places answer differently and both follow the rule the rest of this library uses: bytes now, the reason next call. peek short of what was asked returns what there is and does not raise, and only raises EOF when the buffer is empty and the count is positive; read_bytes and read_string with no delimiter left return everything remaining and let the end arrive on the following call. Go returns the bytes and io.EOF together in both cases, which puts the answer in the value and the reason in the error at the same time and is the shape that produces callers who check the error first and drop the data. deviations.md has the row.

Reader carries its origin in the type, and that costs one thing: reset takes a span from the same place the reader was built over, because Reader[o] cannot be re-pointed at a different o without becoming a different type. Pointing a reader at another part of the same buffer works, which is what Reset is mostly for; pointing it at a different buffer means new_reader, which is two words and no allocation. Same shape as bufio.Reader.Reset, and the row is next to it.

search.mojo is the answer to the second half of issue #15. All fourteen functions are written against Span[Byte, o] and none of them allocates, so core.strings will call these rather than having its own copy: a Mojo String already holds UTF-8 and lends its bytes through as_bytes() without copying. Go has two of everything precisely because []byte(s) copies. index is Rabin-Karp over a rolling hash, which is Go's fallback rather than its amd64 assembly, neither of which is available here, with the naive loop below the crossover Go puts it at and the same FNV prime as Go so that a divergence in behaviour cannot be a divergence in the hash.

121 tests, ported from bytes_test.go, buffer_test.go and reader_test.go. The thing that made them portable is four functions in tests/bytes/_fixtures.mojo. Go's tables are full of byte string literals like "a\xffb" that a Mojo String cannot hold, so enc reads that notation into bytes, quote prints bytes back into it, joined joins pieces with a bar, and expect runs a table's want column through both so the two sides meet. Converting the expectation rather than the result is the direction that keeps the tables readable, because printable ASCII and the bar survive the round trip unchanged, so a row about characters still reads as characters and a row about bytes still reads as \xNN. A failure prints a\xffb instead of a column of numbers.

Two of those tests are worth naming. test_the_dead_prefix_is_reclaimed is Go's issue 5154: a buffer read from and written to forever must reuse the space in front of the read cursor rather than growing without bound, and nothing else in the file would catch it. And test_fields_splits_on_the_space_table_and_not_the_ascii_six is where a wrong assumption of mine got caught by the library rather than the other way round: unicode.IsSpace includes U+0085 NEL and U+00A0 NBSP, which most people would not guess, and does not include U+200B ZERO WIDTH SPACE, which most people would.

core.unicode, all 309 symbols, no waivers: 247 range tables, six maps, the case and fold arithmetic, the thirteen predicates and the Turkish and Azerbaijani special cases. It is scheduled for M3 and lands early for the same reason utf8 did — core.bytes and core.strings are next and both want case mapping — and the manifest says Adapt because the tables cannot be a Go map here and one of Go's types cannot be constructed at all.

Everything is generated by tools/gen/unicode.py from five files of the Unicode character database, vendored under tests/data/ucd at edition 15.0.0 and pinned by sha256 with Unicode-3.0 recorded against each. The generated source is checked in, so a build never reaches the network, and pixi run generated-check reruns the generator and fails on a diff. The generator emits code the formatter would not change, which took two rounds to get right: mojo format wraps at 80 columns and leaves docstrings alone, so an 82 character map entry and a missing blank line between two functions were enough to make format and generated-check disagree about the same file, each of them correct.

data.mojo is nine thousand lines and five arrays, and the shape of it is the reason the package is fast. A range is packed into one 64-bit word as lo | hi << 21 | stride << 42, because a code point needs 21 bits and three of them fit in a word with room over, and every table is laid out end to end in one array rather than one array each. What that buys is that a RangeTable is four integers naming a slice — no pointer, no origin, no lifetime — so it can be a comptime constant, unicode.Greek costs nothing to pass, and reading a range is materialize[_RANGES]()[i], a load from read only memory rather than a copy of the array. docs/design.md section 6 has the measurement that makes that true, and if it stops being true this package gets slow rather than wrong.

The price is the first deviation and it is a real loss: a caller cannot build a RangeTable, because there is nowhere to put the ranges. Go's rangetable.New has no equivalent here. is_one_of over a List[RangeTable] of the tables that do exist covers most of what people build one for, and the rest is a predicate the caller writes. The second deviation is that Go's six maps — Categories, Scripts, Properties, FoldCategory, FoldScript, CategoryAliases — are functions returning a fresh Dict, since a Dict cannot be a compile time value, and so are GraphicRanges, PrintRanges and CaseRanges for the same reason about List. Calling Categories() in a loop builds a 38 entry dictionary every time round; reaching for unicode.Lu instead is free, and that is the advice the module docstring leads with. docs/deviations.md has all four rows, the fourth being CaseRange.Delta, which is a four lane SIMD[DType.int32, 4] because Mojo's fixed size array is not implicitly copyable and its vector type only comes in powers of two. The fourth lane is always zero and nothing reads it.

Is is is_in and In is in_any, because both of Go's names are Mojo keywords. Those two are the questions the package exists to answer, so they had to keep reading as questions rather than becoming is_ and in_. The ten constants are capitals — MAX_RUNE, UPPER_CASE, REPLACEMENT_CHAR — on the same rule that gave io its seek constants, and tools/parity/renames.toml carries all twelve renames with the reasoning.

The check that the tables are Go's is exhaustive rather than sampled, which is what #19 asked for. Two new differ areas do it. unicode-tables dumps all 245 tables the five maps name, range by range, the 328...

Read more

v0.2.0

Choose a tag to compare

@tamnd tamnd released this 02 Sep 18:44
v0.2.0
337feb7

M1 is complete. core.errors is the first package in this library with code in it, and the first at full parity. It is the mechanism every fallible function written here from now on is written against, which is why it comes before anything that could use it.

The problem it solves is that Mojo's Error carries a string and nothing else, while half of Go's standard library returns a struct with a path, a syscall number or a byte count in it. The answer is a record in thread local storage, matched to the error by the message it was raised with, and capture for when the error has to outlive the raise. All three of the ways that can fail silently now have a test that was watched failing.

Two language probes and the design facts they pin, ahead of core.errors in M1.

Mojo has no global mutable state. A module level var is refused outright and the message tells you to move it into a function or make it a comptime constant, so there is nowhere in the language to put a package level counter, cache, registry or default. Go's standard library has one in a dozen places. docs/design.md now records that, and that every one of those becomes a value the caller owns and passes.

The one exception is the thread local error record in section 4, which has to outlive the call that wrote it and cannot be passed. It gets its slot from a small C object that core.errors links, described below. The cost is stated in the design rather than discovered later, and the alternative it was weighed against was threading an explicit context through every fallible call in the library, which puts the error mechanism in the signature of every function in it.

tools/probe/probes/thread_local.mojo pins that pthread's per thread storage really is per thread. Four threads each claim a slot, hand its address to pthread_setspecific, wait at a barrier until all four have written, then read the pointer back and write through it, while the main thread holds a different value in the same key across all of it. Without the barrier a shared slot would still look correct, because each thread would set and read before the next arrived. The failure this rules out is one thread reading another thread's error fields, which is a wrong answer rather than a crash.

Both probes were checked against the two ways they can fail: compiling when they should not, and still being refused for a different reason than the one recorded.

The first library code in the tree: the thread local error record, which is the mechanism every fallible function in this library will be written against.

Report(message).with_field("path", name).error() writes a record into this thread's slot and hands back the Error to raise. field(e, "path"), code(e) and partial(e) read it back at the catch site, so the fields Go would have put in a struct survive a raise that carries only a string, and so does the n from Go's (n, err) that a raise would otherwise drop.

A record is matched to an error by the message it was raised with and by nothing else. That is what makes the two silent failures safe. An error raised by std, or by anything that has never heard of this mechanism, finds a record whose message is not its own and is correctly reported as carrying nothing. An error held past the next raise on the same thread finds the newer record, sees a different message, and reports nothing rather than the newer error's fields. Both of those would otherwise run, print something plausible and be wrong, so both have a test, and both tests were watched failing with the identity check removed. Nothing is appended to the message to make this work: a message with a token in it is a message that cannot be printed, and matching on text that somebody also reads would make every wording change a breaking change to a lookup.

The slot is C, in core/errors/shim/slot.c, and that directory's README says why at length. It is a pthread key rather than a _Thread_local pointer because a key has a destructor, so a thread that exits still holding a record frees it. The destructor is a Mojo function handed to the shim rather than exported for it to find, because a function only C calls is a function a dead code pass removes, and the exported version linked on Linux and not on macOS. core.errors therefore declares unsafe = true, taking the linter's count from fifteen packages to sixteen, and running the tests now needs a C compiler on the host.

core.errors is tier zero, so every binary built on this library links that object. The cost is stated in docs/design.md section 4 rather than left to be found in a link line, and the alternative it was weighed against, threading an explicit context through every fallible call, is recorded there too.

Two more language facts, each with a probe. A struct valued field cannot be moved out of an owned self, which is why errors.Report is both the builder and the record rather than the two structs that would read better. And the pthread key destructor really does call back into Mojo on a worker thread's exit, which was proved before the record depended on it.

tools/lib/native.py is now the one place that goes looking for a C compiler, shared by pixi run baseline and the test runner, and it explains why neither takes one from the lockfile.

The rest of Go's errors package: wrap, matches, join, unwrap, causes and new. core.errors is the first package in this library at full parity, five symbols present and two waived.

The record is a tree now, because wrapping and joining both make one. It is an arena of links with integer indices, which is the technique the design already committed to for every recursive type here, used for the first time. A chain survives any number of levels with each level keeping its own fields, and wrap copies the cause out of the thread's slot before the new record replaces it, which is the only order that works when there is one slot per thread. field and code answer about the error you hand them and not about what it wraps, because two links in a chain can each carry a path and the wrong one is worse than nothing. matches is the one that walks, and it walks every cause of a joined error rather than the first, which is the part of Go's contract that is easy to get wrong and which now fails a test when it is.

A sentinel is a code, and nobody picks the number. Go's errors.Is(err, io.EOF) works because io.EOF is a value you can hold, and there is none here, so a sentinel is an integer on the record. Integers chosen by hand collide, and a collision makes matches(e, io.EOF) quietly true for an os error, so core/errors/codes.toml lists every sentinel in the library and tools/gen/codes.py numbers them. That makes a collision impossible rather than unlikely, at the cost that the numbers move when a line is inserted, so a code is meaningless outside the process that produced it and the type says so. Code is a struct rather than an Int, which is what stops with_code(300) compiling where with_count(300) was meant.

errors.As and errors.AsType are waived. Both are reflection over a type hierarchy and this library has neither, so there is no honest partial version; the replacement is the field lookup, the sentinel comparison, and a per package helper such as os.PathError.of(e). errors.Is is renamed to matches, because is is a Mojo keyword.

join is weaker than Go's and the deviations page says so rather than leaving it to be found. Go holds error values and every field with them. At most one of the errors passed here still owns the thread's record, so the others contribute their message and their place in the tree and nothing more. capture is what will close it.

Sixteen new tests, and the suite was checked against three ways of getting this wrong: a matches that does not walk the chain, a join that reports only its first cause, and a wrap that refers to the record instead of copying it. Each one fails the tests that name it and nothing else.

errors.capture(e) and ErrorValue, which is what makes an error a value again. It copies the error's whole subtree out of the thread and owns every message, field and cause it can reach, so a failure can go in a list, sit in a struct field, or be read by a thread that never saw it raised. ErrorValue.error() is the way back: it installs the record on whatever thread calls it, so a task's error can be re-raised by the thread that collected it and every function in this package then works on it as though it had just happened.

The cross thread case is the one that decides whether any of this is real, and it now has a test. The reading thread first raises an error of its own, so its slot holds something else entirely, and then reads the captured value. Routing ErrorValue.field through the thread's slot instead of the value makes that test fail with left: /the wrong one, which is the exact wrong answer it exists to rule out.

join over captured errors keeps every field of every cause. Over live errors it cannot, because a record is written at raise time and the next raise replaces it, so all but one arrive as a message. Both halves have a test, so the code and the deviations page cannot drift apart.

With that, the three ways design.md section 4 can fail silently are all covered: a foreign error mistaken for ours, an error held past the next raise, and a capture read on a thread whose own storage holds something else.

What's Changed

  • probe: pin no global mutable state and per thread storage by @tamnd in #102
  • core.errors: the thread local error record by @tamnd in #103
  • core.errors: wrap, matches, join and the code registry by @tamnd in #104
  • core.errors: capture and ErrorValue, including across a thread by @tamnd in #10...
Read more

v0.1.0

Choose a tag to compare

@github-actions github-actions released this 02 Sep 17:05
v0.1.0
213e1ee

What's Changed

  • tools/mojotest: the test runner by @tamnd in #98
  • tools/baseline: the platform tables by @tamnd in #99
  • Record what branch protection does and does not require by @tamnd in #100
  • release: v0.1.0 by @tamnd in #101

Full Changelog: v0.0.1...v0.1.0

v0.0.1

Choose a tag to compare

@github-actions github-actions released this 02 Sep 16:08
v0.0.1
1fda99b

What's Changed

  • core: lay out the package tree and the manifests by @tamnd in #77
  • Write the ten language probes by @tamnd in #78
  • lint: fail the build on our own compile time diagnostics by @tamnd in #94
  • parity: measure the surface against Go's own manifests by @tamnd in #95
  • release: v0.0.1 by @tamnd in #96
  • Hold the Go API generator to the Go pixi.lock pins by @tamnd in #97

New Contributors

  • @tamnd made their first contribution in #77

Full Changelog: https://github.com/tamnd/mojo.core/commits/v0.0.1