Repository navigation
v0.3.0
M2 is complete. Reading and writing, buffering, ordering, and the two packages that everything textual is built on. Eight pull requests, seven issues, and the parity count goes from 6.7 percent to 7.8 percent, which is 691 of Go's symbols across eleven packages.
The thread that runs through all of it is that Go's interfaces are checked at run time and these are not. io.copy in Go type asserts its argument to io.WriterTo and then to io.ReaderFrom, and a caller cannot find out ahead of time which path it will take. Here every reader and writer answers capabilities() with a set of bits, the fast paths are selected from those bits, and a caller can ask the same question the library asks. The erased forms exist for when the type genuinely is not known until run time, and they carry the same bits, so nothing has to be discovered by trying.
The other thread is that a for loop in Mojo swallows an error raised out of __next__, which was found with a probe before anything depended on it. That one language fact is why core.iter.Cursor exists and why fallible iteration in this library is written as an explicit has_next and next pair rather than as a loop. The lint enforces it.
core.bytes, 97 of Go's 99 symbols with two waived and one renamed. That closes M2. The waivers are Title, which Go's own documentation deprecates because its word boundary rule turns "they're" into "They'Re", and Buffer.AvailableBuffer, which hands out the buffer's spare capacity to be appended into and has nowhere to land here. The rename is MinRead to MIN_READ, on the constants rule the rest of the library already follows.
The rule the package is built around is one sentence: a method never hands out a view into the buffer's own storage. Go documents Buffer.Bytes, Next and Peek as valid only until the next write, and a Go program that breaks that rule reads stale bytes out of a live allocation. Here the next write can reallocate, so the same mistake reads freed memory, and probes/span_outlives_its_owner.mojo pins that the compiler does not stop it. So all three return owned bytes and there is no view returning accessor beside them under any name. The copy is real and Buffer is a hot type, so there are four ways not to pay for it, and the docstring leads with them: len(), write_to, read into a span the caller already owns, and string(), which builds the String directly rather than through a List[Byte] first. The functions that carve a slice up, trim and cut and split and fields, do return spans, exactly as Go returns subslices: those borrow the caller's bytes rather than the package's own.
Everything Go panics about raises instead. There are four: Repeat on a negative count or an overflowing length, Join on an overflowing length, Buffer.Grow on a negative one, and Buffer.Truncate past what is buffered. Each is a check the caller could have made. Running out of memory raises ErrTooLarge rather than panicking with it, because a buffer that cannot grow is a condition only the caller can report on, since they are the one who knows whether the input that caused it was theirs.
Two more places answer differently and both follow the rule the rest of this library uses: bytes now, the reason next call. peek short of what was asked returns what there is and does not raise, and only raises EOF when the buffer is empty and the count is positive; read_bytes and read_string with no delimiter left return everything remaining and let the end arrive on the following call. Go returns the bytes and io.EOF together in both cases, which puts the answer in the value and the reason in the error at the same time and is the shape that produces callers who check the error first and drop the data. deviations.md has the row.
Reader carries its origin in the type, and that costs one thing: reset takes a span from the same place the reader was built over, because Reader[o] cannot be re-pointed at a different o without becoming a different type. Pointing a reader at another part of the same buffer works, which is what Reset is mostly for; pointing it at a different buffer means new_reader, which is two words and no allocation. Same shape as bufio.Reader.Reset, and the row is next to it.
search.mojo is the answer to the second half of issue #15. All fourteen functions are written against Span[Byte, o] and none of them allocates, so core.strings will call these rather than having its own copy: a Mojo String already holds UTF-8 and lends its bytes through as_bytes() without copying. Go has two of everything precisely because []byte(s) copies. index is Rabin-Karp over a rolling hash, which is Go's fallback rather than its amd64 assembly, neither of which is available here, with the naive loop below the crossover Go puts it at and the same FNV prime as Go so that a divergence in behaviour cannot be a divergence in the hash.
121 tests, ported from bytes_test.go, buffer_test.go and reader_test.go. The thing that made them portable is four functions in tests/bytes/_fixtures.mojo. Go's tables are full of byte string literals like "a\xffb" that a Mojo String cannot hold, so enc reads that notation into bytes, quote prints bytes back into it, joined joins pieces with a bar, and expect runs a table's want column through both so the two sides meet. Converting the expectation rather than the result is the direction that keeps the tables readable, because printable ASCII and the bar survive the round trip unchanged, so a row about characters still reads as characters and a row about bytes still reads as \xNN. A failure prints a\xffb instead of a column of numbers.
Two of those tests are worth naming. test_the_dead_prefix_is_reclaimed is Go's issue 5154: a buffer read from and written to forever must reuse the space in front of the read cursor rather than growing without bound, and nothing else in the file would catch it. And test_fields_splits_on_the_space_table_and_not_the_ascii_six is where a wrong assumption of mine got caught by the library rather than the other way round: unicode.IsSpace includes U+0085 NEL and U+00A0 NBSP, which most people would not guess, and does not include U+200B ZERO WIDTH SPACE, which most people would.
core.unicode, all 309 symbols, no waivers: 247 range tables, six maps, the case and fold arithmetic, the thirteen predicates and the Turkish and Azerbaijani special cases. It is scheduled for M3 and lands early for the same reason utf8 did — core.bytes and core.strings are next and both want case mapping — and the manifest says Adapt because the tables cannot be a Go map here and one of Go's types cannot be constructed at all.
Everything is generated by tools/gen/unicode.py from five files of the Unicode character database, vendored under tests/data/ucd at edition 15.0.0 and pinned by sha256 with Unicode-3.0 recorded against each. The generated source is checked in, so a build never reaches the network, and pixi run generated-check reruns the generator and fails on a diff. The generator emits code the formatter would not change, which took two rounds to get right: mojo format wraps at 80 columns and leaves docstrings alone, so an 82 character map entry and a missing blank line between two functions were enough to make format and generated-check disagree about the same file, each of them correct.
data.mojo is nine thousand lines and five arrays, and the shape of it is the reason the package is fast. A range is packed into one 64-bit word as lo | hi << 21 | stride << 42, because a code point needs 21 bits and three of them fit in a word with room over, and every table is laid out end to end in one array rather than one array each. What that buys is that a RangeTable is four integers naming a slice — no pointer, no origin, no lifetime — so it can be a comptime constant, unicode.Greek costs nothing to pass, and reading a range is materialize[_RANGES]()[i], a load from read only memory rather than a copy of the array. docs/design.md section 6 has the measurement that makes that true, and if it stops being true this package gets slow rather than wrong.
The price is the first deviation and it is a real loss: a caller cannot build a RangeTable, because there is nowhere to put the ranges. Go's rangetable.New has no equivalent here. is_one_of over a List[RangeTable] of the tables that do exist covers most of what people build one for, and the rest is a predicate the caller writes. The second deviation is that Go's six maps — Categories, Scripts, Properties, FoldCategory, FoldScript, CategoryAliases — are functions returning a fresh Dict, since a Dict cannot be a compile time value, and so are GraphicRanges, PrintRanges and CaseRanges for the same reason about List. Calling Categories() in a loop builds a 38 entry dictionary every time round; reaching for unicode.Lu instead is free, and that is the advice the module docstring leads with. docs/deviations.md has all four rows, the fourth being CaseRange.Delta, which is a four lane SIMD[DType.int32, 4] because Mojo's fixed size array is not implicitly copyable and its vector type only comes in powers of two. The fourth lane is always zero and nothing reads it.
Is is is_in and In is in_any, because both of Go's names are Mojo keywords. Those two are the questions the package exists to answer, so they had to keep reading as questions rather than becoming is_ and in_. The ten constants are capitals — MAX_RUNE, UPPER_CASE, REPLACEMENT_CHAR — on the same rule that gave io its seek constants, and tools/parity/renames.toml carries all twelve renames with the reasoning.
The check that the tables are Go's is exhaustive rather than sampled, which is what #19 asked for. Two new differ areas do it. unicode-tables dumps all 245 tables the five maps name, range by range, the 328 case ranges and the 42 category aliases, and the 7,571 lines are byte identical to the same dump out of Go. unicode-runes prints one line per code point holding a thirteen bit predicate mask and five mappings, and all 1,114,112 of them are byte identical too. Together that is 1,121,683 compared inputs in about eighteen seconds. Both drivers accept --seed and ignore it: there is nothing random about a character database, and always starting at zero is what makes the differ's case index the code point that diverged, so a failure names itself. The nightly passes the run number as a seed, and without that rule it would have started the sweep at an arbitrary code point and skipped everything below it.
tools/differ grew the harness to go with them. Areas are declared in tools/differ/cases.toml with the command for each side, what it needs on PATH, and how many cases it wants, and differential.yml now derives its matrix from that file rather than from a list in the workflow. The list that was there named eight areas that do not exist and passed flags the tool does not have, so the first nightly after this would have failed on all eight; an area now reaches the nightly by being declared, and the areas docs/testing.md describes and nobody has written are simply absent, which is the difference between a plan and a job that fails every night. The runner also caps how many divergences it prints and says how many it swallowed, because an area that is wrong everywhere should not produce a million line log.
Eighty tests, ported from Go's letter_test.go, graphic_test.go, digit_test.go and script_test.go with the fixtures transcribed from Go's sources rather than from memory — those code points were chosen by somebody reading the Unicode files, one per interesting block and one either side of every boundary, so a table derived from this implementation would agree with it by construction and prove nothing. Go's category and property tests carry one line that is easy to miss and worth more than the rest of the file: after every fixture row has been checked, every category the fixture did not mention is a failure. That assertion is here too, so a table cannot arrive untested. The other two that earn their place are TestLetterOptimizations, which asks the 256 byte Latin-1 table and the range tables the same questions and is the only thing that would catch one bit set wrong in the byte table, and TestNegativeRune, which exists because those shortcuts narrow a rune to a byte and a negative rune narrowed to a byte looks like a perfectly ordinary letter.
Go's quirks are reproduced rather than fixed, and the tests say so by name. FoldCategory has no LC key even though LC is a category with fold exceptions, because Go's generator only walks the categories that appear in the file it reads; test_fold_category_has_no_cased_letter asserts the absence so that nobody quietly corrects it. to(-1, r) answers REPLACEMENT_CHAR rather than failing, which is one of the few places in either library where a programming mistake produces a code point. And the case mappings are simple mappings only: ß upper cases to ß and not to SS, the Turkish tables are four code points rather than a language, and full case mapping is a string operation that belongs to core.strings.
simple_fold is the function in here worth reading the docstring of. It is not a fold — it walks the cycle of code points equal to its argument under case folding, so calling it repeatedly enumerates an equivalence class. K, k and U+212A KELVIN SIGN are one such class and S, s and U+017F LONG S are another, and the case mappings only go one way through each: U+212A lower cases to k and nothing upper cases to U+212A, while U+017F upper cases to S and nothing lower cases to U+017F. A program comparing case insensitively by mapping both sides has to map towards lower case to get Kelvin right and towards upper case to get long s right, and no single direction gets both. That is the whole reason the function exists, and there is a test that says it in five lines.
core.maps, all ten symbols, no waivers. The manifest says Adapt, because two of Go's five shapes do not survive the crossing and the docstrings argue for both rather than presenting them as translations.
The three producers return a List where Go returns an iter.Seq. A Dict already has keys(), values() and items(), they borrow rather than copy, and a keys here that handed one back would be a second name for something that exists. What does not exist is the thing a Go programmer writes slices.Collect(maps.Keys(m)) for, so that is what these three are, each with an _into sibling that appends to a list the caller already owns. That is thirteen functions for ten symbols and it is the cheaper half of the deviation.
The order is the Dict's and nothing here promises what that is, with one exception that is worth having and costs nothing: with no mutation in between, keys, values and all walk the dict the same way, so keys(d)[i] and values(d)[i] are one entry. Go promises the opposite on purpose — each of its iterators randomises separately — and the effect there is that anyone wanting parallel lists has to go through All and unpack. test_keys_and_values_line_up is the test that makes it a guarantee rather than an accident.
The two consumers take a Span[Tuple[K, V], o] rather than the core.iter.Cursor that slices.collect takes, and that one is forced. A cursor's element type is an associated comptime Element, and there is no way to say "and that element is a Tuple[K, V]" that the compiler will take apart again: where C.Element == Tuple[K, V] parses, is checked at the call site, and leaves C.Element opaque in the body, so pair[0] on it is still an error. probes/pair_cursor.mojo pins that and design.md records it next to the weaker form slices.sorted ran into. A fallible source is therefore drained with slices.collect first and the list handed over, which costs a line and a copy of the pairs; two tests do exactly that against a fixture cursor that fails partway, because a documented workaround nobody compiled is a guess.
delete_func gathers the keys to delete before deleting any of them, so the predicate sees every entry exactly once. Go deletes as it iterates, which its map specification allows and which nothing here promises about a Dict; the cost is one list of keys and the gain is that the rule is about this function rather than about the implementation underneath it. equal compares values with == and so says a dict holding a NaN is not equal to itself, which is core.cmp.compare's total order deliberately not applying outside sorting, and there is a test named after it.
Every bound is spelled K: KeyElement & Copyable & Deinitable. KeyElement is Movable & Hashable & Equatable and implies neither of the other two, so the first version of this package compiled until it tried to copy a key and then said it lacked evidence to prove the copy correct. Forty-three tests, sorting both sides of every comparison through one shared helper rather than repeating the sort, because a test that sorts one side and forgets the other passes for the wrong reason.
Four packages now have a tests/<pkg>/_fixtures.mojo and two of them imported it as a bare from _fixtures import. The suite is one generated main built with -I at the root, so the fourth one made the name ambiguous and the build failed inside tests/slices, which had not changed. All of them are now spelled from tests.<pkg>._fixtures import, which is what the other two already did.
core.cmp, core.sort and core.slices, eighty-four symbols across the three, no waivers anywhere in them. They landed together because they are one stack: cmp is the ordering, sort is the algorithm over it, slices is the collection of loops people actually call. Splitting the work would have meant writing slices.sort twice.
core.cmp is four symbols and one of them is a decision. Ordered is a comptime alias for Comparable rather than a new trait, because Go writes cmp.Ordered as a type set only for want of a way to say "supports <", and Mojo has one. Inventing a trait here would have produced a worse Comparable that Int did not conform to. The name survives so that a bound written T: Ordered reads the same in both languages.
The NaN rules in compare are Go's and they are deliberately not IEEE's: a NaN sorts before every number and two NaNs are equal. IEEE makes a < b, a > b and a == b all false for the same pair, and a partition loop given that answer three times has no invariant left and walks off the end of the array. That is not a hypothetical — it is the same failure the next paragraph is about. The cost is that compare and == disagree about NaN on purpose, so slices.equal and slices.compare give different answers on the same two spans, and there is a test named after the disagreement rather than a comment apologising for it.
core.sort is a port of Go's pdqsort and SymMerge rather than a wrapper around std.builtin.sort, and the manifest says Port where it used to say Wrap. The reason is memory safety, not fidelity. Handed an inconsistent comparator, the standard library's sort passes that comparator elements from outside the span it was given: McIlroy's antiquicksort adversary extracts 645 such comparisons from a 1,000 element range, and an index outside the data is a dead process. Go's sort does not do this at any size, stable or unstable. A caller who gets the ordering wrong deserves a wrong order, and that guarantee is most of the reason to want Go's sort in the first place. tests/sort/test_adversary.mojo runs the adversary at 1,000 and at 10,000, and the stable sort is run through it too, because "that one has nothing to attack" is a claim and not a fact.
A comparator here is a compile time parameter, and tools/probe/probes/comparator_capture_list.mojo pins the one subtlety that makes the whole approach work. The parameter has to be declared capturing [_]. Spelled capturing thin, which is how std.builtin.sort prints its own signature, it compiles, drops the capture silently, and sorts by whatever the closure would have read had it been able to read anything — and with a capture owning heap memory the same shape aborts instead. No error and no warning at that spelling. The probe is the working half, forwarded through one wrapper and through two, which is exactly what core.slices calling core.sort is. The broken half is not written down, because a dropped capture is undefined behaviour and a probe should not be.
Reverse holds a pointer at the data rather than a value, because a Reverse holding a T would sort a copy and hand the caller back an unchanged slice with nothing to report. StringSlice.swap goes through a temporary and two copies, because borrowing both elements of one span mutably is refused; sorting an index list with slice is the way out when the strings are long enough for that to matter. Sixty-two tests, including Bentley and McIlroy's grid of distributions and both adversaries.
core.slices is forty symbols over four files, and the split is by what the function does to the length rather than by Go's file layout. find.mojo and order.mojo take a span and leave it the length it was. edit.mojo is everything that changes the length, and every function in it takes mut s: List[T] and returns nothing. Go returns the new slice because a slice header is a value and returning one is how a length change gets published; here the length lives in the list the caller already owns, and handing back a second List over the same buffer is the thing ownership exists to prevent. insert rotates the appended block into place with three reversals rather than shifting, clip gives the spare allocation back rather than capping a capacity that a span does not have, and concat is variadic over List[T] rather than over spans because a Mojo variadic binds one origin parameter for every argument and cannot take spans into two different lists.
seq.mojo is where Go's symmetry breaks. Go's producers return an iter.Seq and its consumers take one, so the two halves meet in the middle. Here the four producers return ordinary Mojo iterators and a for loop over one is correct, while the five consumers take a core.iter.Cursor and are written as a while. Walking a span cannot fail, and core/iter/cursor.mojo says outright that a list iterator is not a Cursor and should not be made into one. But csv records, sql rows and directory walks can all fail, and a collect that only accepted infallible iterators would be a collect with nothing to collect from. The producer tests are written as for loops on purpose: if one of them stops compiling, the producers have quietly become fallible and the rule needs looking at again.
append_seq takes the cursor first and the destination second, which is backwards from Go and from every other length-changing function in the package. That is not taste. C is inferred from the cursor and the destination is a List[C.Element], so with the destination first the compiler reaches it before C is resolved and refuses the call. The three sorted functions have a related scar: where conforms_to(C.Element, Ordered) does narrow the associated type where the type stands alone, and does not let T be inferred through a Span[C.Element, o], so they go through two private helpers whose T is explicit. probes/refined_associated_type.mojo has the reduced case and design.md records it, so when the compiler learns to infer through the refinement the probe starts compiling and the helpers go away.
min_func and max_func both return the first of a tie. That is Go's documented rule for both, it was implemented here as first and last before anyone checked, and the obvious wrong max_func — >= where it should be > — passes every test that does not have three records comparing equal on the key and differing in the payload. There is now such a test. One hundred and twenty tests over the package, with a Tracked element type that counts live objects rather than destructor calls, because counting destructors made the assertions depend on how many copies compact happened to make.
The linter's iteration rule is narrower by one case. It rejected any __next__ that raises, which is design.md section 7 made mechanical, and Mojo 1.0's Iterator trait requires the signature def __next__(mut self) raises StopIteration. That raise is not an error being swallowed — it is how the language spells end of input and the loop catching it is the loop working — so StopIteration alone is now exempt, and the check reads what follows raises rather than stopping at the word. A bare raises is still rejected and so is raises Error or any other named type, which is the shape a lazier exemption would have waved through: a check that stopped at the word raises and accepted anything typed after it would let a read failure arrive by the route the loop uses to learn that the data ended. tests/lint/typed_fallible_next.mojo is the third fixture and it keeps that half of the rule provably alive.
core.unicode.utf8, all nineteen symbols, no waivers. It is scheduled for M3 with the rest of Unicode and it landed in M2 instead, because core.bufio needed a decoder and core.bytes and core.strings need one next, and the milestone order had them all depending on something that did not exist. Splitting it was cheap: the nineteen symbols are arithmetic, and the tables, the character classes and the case mapping that make unicode a week of generated code are still in #19. Issue #115 has the reasoning and #19 lost these from its scope.
The package works on Span[UInt8] rather than on String, and that is the point rather than an implementation detail. A String here is already valid UTF-8, so a decoder over one has nothing to decide; the input worth having a decoder for is bytes off a socket. The four ...InString functions are kept anyway, one line each, because Go has them for a reason that does not apply — []byte(s) copies and String.as_bytes borrows — and a Go programmer should still find the name they went looking for.
Invalid input decodes to RUNE_ERROR with a size of exactly one, never zero and never the width of the malformed sequence. That is Go's rule and it is not the obvious one: a decoder meeting a byte it dislikes could reasonably skip the whole sequence and resynchronise. Advancing one byte is what makes a decode loop terminate, what makes rune_count agree with the loop a caller writes, and what leaves resynchronisation to the caller, who is the only one who knows whether the right answer is to show the damage or reject the message.
The corollary is the trap in the package. RUNE_ERROR is an ordinary code point that valid input may contain, so valid is written against the size rather than the value — EF BF BD is three bytes of perfectly good text, and a check on the value would have this library calling its own error output malformed. rune_len is the other one: it answers -1 for a value that is not a code point where encode_rune substitutes and writes three bytes, so sizing a buffer from it is a bug that only shows up on input somebody else chose. bufio's private predecessor made those two agree by returning 3, which was the wrong fix — it made the question unanswerable. There is a test named after the disagreement.
AppendRune is the one signature that changed. Go returns the grown slice because that is what append does; append_rune takes mut dst and returns the byte count, since two names for the same list is what ownership forbids. deviations.md has the row.
Forty tests. The tables are Go's utf8map, surrogateMap and the invalid sequence list, transcribed by hand as byte lists rather than string literals — half of those sequences cannot be held in a Mojo String at all, which is itself the reason the package takes spans. Two of them are exhaustive rather than sampled: every code point from zero to MAX_RUNE round trips through encode_rune and decode_rune with rune_len checked in the same loop, and all 256 single bytes are decoded against the stated rule with full_rune checked alongside so the two answers cannot drift.
Forty two mutations, one per claim, forty one caught on the first pass. The survivor was the landing check in decode_last_rune, and the reason it survived is worth more than the fix. Every truncated input in the table answers (RUNE_ERROR, 1) because the forward decode from the recovered start byte fails on its own; the check that the decode landed exactly on the end never decides anything for them. It decides for 41 C2 80 80 — an A, a complete two byte rune, and one loose continuation byte — where the backwards scan finds the C2, decodes a real rune from it, and is still wrong. That input is now a test, and the table test's docstring says which half of the work it is doing.
core/bufio/_rune.mojo is deleted. Its first line promised that it would go when core.unicode.utf8 arrived and that the replacement would be an import change and a deletion, and that is what it was: three imports repointed, _rune_len dropped because nothing ever called it, and 124 bufio tests passing unchanged. core.bufio gains the dependency and the PACKAGE.toml comment explaining the absence is gone.
core.bufio, all of it: the buffered reader, the buffered writer, the scanner, the four split functions and the pair of halves Go calls ReadWriter. pixi run parity core.bufio reports nothing missing, with six waivers each carrying a reason about Go rather than about effort. It landed at tier 2 rather than the tier 3 the graph had planned for it, because it turned out to need core.bytes for nothing that a loop over a span does not do, and core.unicode.utf8 is tier 0, so neither dependency costs a tier. Four packages downstream moved with it.
The decision that shaped every file is that nothing here returns a view into a buffer. Go's Peek, ReadSlice, ReadLine and Scanner.Bytes all hand back a slice of the reader's own memory and document that the next call invalidates it, and design.md used to claim that origins turn that documented hazard into a compile error here. That claim was wrong and it had no probe, which is why every claim in that file has one now. A Span built over a List stays usable across a mutation of that list, including one that reallocates and leaves it pointing at freed memory, and nothing in the compiler notices; probes/span_outlives_its_owner.mojo pins it. So all four return owned bytes, the copy is on deviations.md as a cost rather than a win, and that page's better column is one row shorter than it was.
Scanner is a core.iter.Cursor, which is what Cursor was written for: it names bufio.Scanner as the thing it exists to prevent. for s.Scan() {} ends early and silently on a read failure and the check that would have caught it is s.Err(), somewhere after the loop, easy to forget and impossible for a reader of the code to notice is missing. Here a clean end of input is False out of has_next and everything else is a raise, so there is no Err and nothing to forget. The deviations page has the row and the module docstring has the loop written out, deliberately as a while rather than a for.
The split function is a trait rather than a function pointer, and it is the first use of a shape the rest of the library will lean on, so design.md section 3 gained a paragraph about it. Go's SplitFunc takes a []byte, and a function type here has to name concrete types in every position, origins included, so the argument has no spelling that does not launder an origin through core.runtime.box — in a package that otherwise needs nothing unsafe. A trait method may be parametric over the origin exactly as io.Reader.read is, so the hole does not exist, and the receiver is where a closure's captures go. mut self makes a stateful splitter ordinary rather than clever. The consequence is that the splitter is a type parameter fixed when the scanner is built, so Go's Split setter and the panic it carries for being called after the first Scan are both gone rather than reimplemented.
A token is a start and a stop into the data the splitter was given rather than a subslice, for the same reason peek returns a copy. That moves a bounds check from the compiler to us and buys a check Go cannot make: a split function whose range is outside the window is caught rather than handing out neighbouring bytes. Split also carries final where Go has ErrFinalToken, which is Go's one place where an error is not a failure; the error channel here only ever carries failures.
read_string and Scanner.text raise on input that is not valid UTF-8, and that is the one place this package is deliberately stricter than Go rather than merely different. A Go string is bytes that are usually text, so ReadString hands back whatever it read; a Mojo String says it is UTF-8. The alternatives were to substitute U+FFFD, which loses bytes silently, or to assert the encoding without checking, which is the unsafe_ constructor and would have made this package unsafe on behalf of a caller who never asked. read_bytes and Scanner.bytes are the versions that never refuse, and a protocol carrying both text and binary wants them anyway.
ReadWriter is two named fields and twenty forwarders rather than Go's two embedded pointers, and it comes out ahead. Go's version has two Buffered methods, two Size methods and two Reset methods, and promotion resolves none of them: rw.Buffered does not compile and a Go caller writes rw.Reader.Buffered() anyway, one mistake later. The three ambiguous names are simply not forwarded here and the fields are public, so that spelling is what everybody writes from the start. Everything with one honest meaning is forwarded.
Two things about Go's scanner had to be checked against Go's own loop rather than assumed, and both were assumed wrong first. Once the input has ended, a decision carrying no token ends the scan even if bytes are left — a stateful splitter that skips on alternate calls loses its last token, and that is Go's rule, not a bug in the scanner. And the token ceiling counts the delimiter: a sixteen byte line needs a seventeen byte buffer, because the delimiter is held before the splitter drops it. Both are now named tests rather than quiet fixes.
The tests are four files over six shared fixtures, and the fixtures are where the value is. A buffered reader is a loop around a source that does not cooperate, so Half, OneByte and a reader that fails after handing over its bytes are what the whole suite is run through, alongside a sink that accepts half of every write and one that fails after a fixed number of bytes. Nine hand written splitters go with them, five misbehaving on purpose: one that goes backwards, one that goes past the end, one whose token range is outside the data it was given, one that never advances, one that raises.
Twenty seven mutations, one per claim. Twenty three were caught on the first pass and four were not, which is the reason for doing it rather than a footnote to it. Refusing a peek larger than the buffer can ever hold was checked for the error but not for the fact that it is refused before the source is touched. The straight-through read path was tested with a span more than twice the buffer, so a version that took it half as often passed. A cut line ending in a carriage return was tested only where the newline did arrive, where dropping the return and putting it back are indistinguishable. And both ScanWords inputs ended in whitespace, so the branch that produces the last word because the input ended was never reached. Four tests later all twenty seven are caught.
_rune.mojo was a private decoder written because core.unicode.utf8 was scheduled for M3 and read_rune, write_rune and ScanRunes cannot be written without one. It is gone, deleted by #115 above, which is what its first line promised. ScanRunes deviated while it was here and keeps deviating after, for a reason that has nothing to do with the decoder: Go substitutes the encoding of U+FFFD for a byte that is not valid UTF-8, and a token that is a range of the input has nowhere for bytes that are not in the input to come from, so the offending byte is the token. One per byte, same advance, so the token stream still lines up with Go's.
The rest of core.io. Nineteen more traits, seven functions and five types, which is Go's io complete apart from the pipe, and pixi run parity core.io now reports every symbol present or waived with a reason.
The small traits are one required method each and none of them needed a decision, with one exception. ReaderFrom and WriterTo already existed as methods on Writer and Reader with capability bits behind them, so declaring them again as traits looked like two mechanisms for one thing. It is not, and the reason is that a single method satisfies both: a struct that lists Writer, ReaderFrom and writes read_from once conforms to both, which was checked before any of this was written. Which name to use is a question about what a function needs rather than about what a type is. A function that only makes sense over a writer with a fast path takes [W: ReaderFrom] and gets a compile error otherwise; copy cannot, because it takes any writer and finds out at run time, and that is what the bit is for. The bit stays the only thing copy reads.
read_full and read_at_least are the two functions that justify the two new sentinels. A truncated stream raises ErrUnexpectedEOF and an empty one raises EOF, and the distinction is the entire reason those are separate numbers: a decoder reading records in a loop wants to stop cleanly at a record boundary and complain loudly anywhere else, and only this layer can tell it which happened. ErrShortBuffer is raised before anything is read, because a buffer smaller than the minimum asked for is a caller mistake that no amount of input would fix.
Both loop, and a version of either written as one read passes every test over a reader that fills the span it is given. So they are tested through a half reader as well, which is Go's iotest.HalfReader and the type most of the value in Go's io suite is carried by. The same reader is what proves read_all keeps its offsets right across the three list growths a 1,500 byte stream needs.
copy_n and copy_buffer are the two copies with something invisible to get wrong. copy_n has to clip its last read to the count, because a version that read a whole buffer and wrote only part of it produces exactly the right output and silently eats the rest of a stream somebody else was going to read. copy_buffer has to use the buffer it was handed and still take a fast path over it, which surprises people and is Go's behaviour. Neither claim is visible in the bytes, so both are checked with counters.
copy_n differs from Go and says so. Go reaches the destination's read_from by wrapping the source in a LimitReader and calling Copy; the wrapper here would have to hold a borrowed reader in a field, which erased.mojo forbids because a view does not keep its target alive. A caller that owns its source gets Go's behaviour exactly by writing copy(dst, limit_reader(src^, n)), and the docstring says that rather than leaving the difference to be discovered.
LimitedReader, SectionReader and OffsetWriter are generic over what they wrap instead of holding an erased value. Go has to hold an interface, so every read through io.LimitReader is an indirect call; here LimitedReader[Fixed] calls Fixed.read directly and can inline into it. The cost is that they own their argument, because a parameter has to be a type and the alternative is a view in a field.
SectionReader.outer is the one place a Go signature would not translate. Go hands back the underlying ReaderAt along with the offset and the size, and a method here cannot return a borrow of a field alongside two values, so the source is the public field r and outer returns the pair. Two section readers over one source still read at the same time without disturbing each other, which is the property ReaderAt exists to promise and there is a test for it.
WriterAt takes mut self where ReaderAt takes self, and the asymmetry is deliberate. Go documents parallel non overlapping WriteAt calls as safe, which is a property of the destination rather than of the value, and Go can express it because its implementations write through a pointer. A sink that keeps its bytes in a field cannot change them through an immutable borrow without interior mutability, which this library does not have, so the promise survives as a comment and the borrow checker serialises the calls. Nothing was lost that could have been said.
multi_reader, multi_writer, tee_reader, nop_closer and Discard are the combinators, and three of them behave differently from Go on purpose. A multi reader whose middle source is empty keeps going rather than handing back a zero length read, because a zero without an error is the thing ErrNoProgress exists to catch and producing one deliberately is worse than the extra loop. A tee reader raises when its sink fails, because by then the bytes are out of the source and a swallowed failure loses them with nobody told. And it does not flatten a nested multi reader the way Go does, because the erased value's type is gone by the time it is seen and there is nothing to recognise.
Discard is a type rather than a variable, since there are no package level variables and there was nothing to construct anyway, so it is Discard() at the use site. It sets READER_FROM and implements read_from, which is what makes draining a reader into it allocate one 8 kB buffer instead of thirty two.
write_string does not type assert. Go's io.WriteString looks for an io.StringWriter and the only thing that buys is skipping the copy []byte(s) makes; String.as_bytes is a borrow, so there is no copy and the assertion would have nothing to gain. The trait is still declared, because a writer that wants the string rather than its bytes is a real thing and a function can ask for one by name.
io.Pipe is waived to M4 with the milestone and the issue named on the waiver line. Every one of its methods blocks until the other side arrives, which needs core.sync, and a pipe written against nothing would deadlock on first use. ErrClosedPipe is numbered now regardless, so the sentinel sits with the rest of the io ones and does not move when the pipe lands.
Twenty mutations, one per claim, each reverted after the suite caught it. Deleting the short buffer check, reporting EOF where a truncation should report ErrUnexpectedEOF, running read_at_least once instead of looping, dropping LimitedReader's clip, reading a section's read_at offset as absolute, letting copy_n read past its count, having copy_buffer allocate its own buffer or skip the fast path, swallowing a tee's write failure, clearing Discard's bit. All twenty were caught.
One more Mojo fact, with a probe. A struct's own parameter has to be written Self.R in a field declaration; the bare name is refused with "unqualified access to struct parameter" while being perfectly legal in a method signature three lines below and in the return type of a free function next to it. Every wrapper type in this package hit it.
The parity tool was under-reporting again, and this time the bug had the shape worth recording: it read re-exports with a regular expression that only matched a single line from ... import, so a package measured worse the more it exported, because mojo format rewrites a long import into the parenthesised multi-line form the moment the names stop fitting. It matches both spellings now. The seek constants are in renames.toml, because the rule turns SeekStart into seek_start and this library spells a constant in capitals.
core.iter.Cursor, which is the shape every fallible iterator in this library has to be written in. A Mojo for loop drops an error raised out of __next__, so a csv reader that hits a malformed row returns the rows before it and the program carries on believing it read the whole file. The replacement is an explicit has_next and next pair with a trait to hang it on, and the loop written out three lines longer, which is the entire point: the failure comes out of the while or out of the next and lands on the caller, because there is nowhere for it to go quietly.
It is deliberately not called Iterator. The hazard being defended against is somebody assuming for x in it works, and a trait with that name invites the assumption at every use site. Infallible iteration keeps Mojo's own __iter__ and __next__, which are safe for exactly the reason this is not.
The element type is an associated comptime Element rather than a parameter, because trait Cursor[T] does not compile: trait declarations do not take parameters. Its bound is Deinitable & Movable, the least a caller needs to own what it is handed and let it go again, so a cursor may yield a buffer or a file handle. Requiring Copyable would have ruled out most of the interesting ones. Both methods may raise and the implementation chooses which side does the work, because a reader cannot know whether there is another record without parsing one, and a reader sitting on a buffer answers has_next cheaply and fails in next.
The swallowing turns out to be specific to __next__, which nobody had checked. A raising __has_next__ makes the for statement itself a raising call, so the loop cannot be written in a function that is not raises and the error comes out normally. There is a second probe for that half, because the two behaviours being different is what makes the linter rule correct as narrowly as it is written, and a release that made them consistent could go either way.
The linter rule had a hole of the same kind the unsafety check had. Its pattern required the parenthesis straight after the name, so def __next__[o: Origin](mut self) raises went through unmatched, and that is not an exotic spelling: it is what an iterator borrowing its input looks like. The pattern now skips an optional parameter list, and a second fixture goes with it, because the first one names __next__ the old way and would still be rejected with the new half deleted.
Go's four symbols in iter are waived. Seq is func(yield func(V) bool), so the sequence is a closure and iterating it means storing another one, which design.md section 3 says is not available; Pull gets to a next and stop pair, which is the shape used here, but it gets there by taking a Seq. There is no Cursor2, because Go needs Seq2 only to give a range over func a second yield argument and a Cursor whose Element is a tuple covers it.
core.io has its first three symbols: Reader, Writer and copy. It is the package the rest of the library is shaped around, so the interesting part is not the code but which of Go's io decisions survived contact with Mojo.
Go's io.Copy opens with two type assertions, src.(WriterTo) and dst.(ReaderFrom), and neither is available here. A trait is a compile time constraint rather than a value, so there is nothing to assert on, and a generic function also cannot ask whether its R happens to implement something, so both halves of Go's mechanism are missing rather than one. They are replaced by one thing: capabilities, an Int of bits declared on the trait with a default body returning zero, alongside the optional methods declared with default bodies that raise. A type that implements write_to overrides two methods, the one that does the work and the one that admits to it. What makes this worth the redundancy is that it reads identically on both paths. A static reader answers from a constant the compiler sees through; an erased one answers from a field it copied off its target at construction. Go needs an interface table lookup for the erased case and cannot do the static case at all.
A type can therefore lie, and setting a bit without writing the method gets the inherited raising stub, which is a clear failure at the first call rather than a wrong answer. There is a test for that, and for the path choice itself: the four combinations of a static and an erased side, with the bit set and with it cleared, proven by a counter on the test types rather than by reading the code. A suite that only compared the output bytes would pass with the whole dispatch deleted.
Read is stricter than Go's. A read that moved bytes returns them and does not raise, so EOF always arrives with a count of zero. Go allows both at once and then asks every caller to handle the n before the err, which is a rule a caller can forget and a familiar source of silently truncated data. Go's own bufio already behaves the way this does.
Two erased types rather than one, and the split is the safety story. AnyReader owns what it wraps and is what a function returning "some reader" returns. ReaderView borrows: it is an address and a table and it keeps nothing alive, and it exists because read_from is handed a mut src it has no right to take and must still call it through a function pointer, which is the one thing a trait bound cannot do. A view is an argument and never a field, nothing checks that, and it is confined to two constructions inside the package that both hand it straight to a call.
Two Mojo findings came out of building it. A pair of structs whose function pointer fields name each other compile on their own and stop compiling once the slots are filled by parametric thunks, reported as struct has recursive reference to itself against whichever came second, which is not where the problem is; the slots carry the other view's address as an Int instead, so only the thunk bodies name it. That one is recorded next to the code and not in design.md, because every claim there comes with a probe and no small program reproduces it. The second one does: a value is destroyed after its last use rather than at the end of its scope, so taking a view's address and never mentioning it again means the view is already gone when the callee reads it. _ = view^ after the call is the fix, tools/probe/probes/last_use_destroys.mojo pins that it is still needed, and design.md now says so.
The parity tool was measuring this package at zero and then at half. It ran mojo doc without an include path, so the first package in the tree with a dependency failed to parse and reported no symbols at all while looking like a package nobody had started; it now passes the root and prints the compiler's own message when a package fails, rather than leaving the reader to reproduce it. It also counted only declarations, so a name a package republishes for its users was invisible, which is EOF and its two siblings here: they are declared in core.errors.codes because they come out of one generated table, and they are importable exactly as Go's io.EOF is. It now reads the from ... import lines of each package's __init__.mojo, and only that file, because an import anywhere else is a private dependency rather than a promise.
A correction to design.md section 5, and a probe so it stays corrected. The section said a struct field cannot carry an unbound origin, and left the reader to conclude that the erased box stores an integer because nothing else was available. That conclusion is wrong: AnyOrigin in a field is rejected, but Pointer[T, UntrackedOrigin[mut=True]] in a field is accepted and erases just as thoroughly. There are now two probes, one on each side of that line. The integer is a choice, and the reason is that the box has forgotten its pointee type as well as its origin, so a pointer field would have to name a placeholder T and reinterpret at every use.
The bottom of the erasure design: core.runtime.box, a refcounted heap box holding a value whose type has been forgotten. io.Reader, io.Writer, net.Conn, the nine database/sql driver interfaces and every value in a JSON document sit on this, because Mojo has no trait objects and each of those has to be built by hand out of a box and a table of function pointers.
It is its own package rather than forty lines inside core.io, which is the decision worth explaining. All four of the consumers already depend on core.io, so that would have cost nothing in dependency edges. It was rejected because core.io would have to declare unsafe = true, and that permission then covers every line of the reader, the writer and copy rather than the forty that need it, and because erasure is not an IO idea: anything that ever wants a heterogeneous collection would have to depend on core.io to get one. That makes 138 packages and 17 unsafe, with core.io still safe, which was the point.
The count and the value share one allocation, because erasure sits on the read path of every buffered reader in this library and two allocations per box is a cost that shows up. That means the block is aligned by posix_memalign rather than malloc, since malloc promises sixteen bytes and a SIMD type wants sixty four, and there is a test that boxes one.
The destructor is the part that was not obvious. A parametric _release[T] can be materialized into a def (Int) thin -> None and stored in a struct field, which is what lets a box that has forgotten its type still run that type's destructor. That is design.md section 2 used for something other than a vtable.
Copying is an atomic increment and it is spelled out at the call site, because Go's interface copy is free and this one is not. errors.ErrorValue made the same call for the same reason.
pixi run race builds the whole suite under the thread sanitiser. Refcounts and locks are the two things in this library whose bugs do not arrive as a failing assertion, so a green pixi run test is not evidence about either of them. The sanitiser reports and then leaves the exit code alone, which would have meant a suite with a data race passing with the report sitting in the log, so the runner reads its output and promotes the report. Making the box's count non atomic produces four race reports and a failing count, and both halves are printed rather than the first of the two.
One trap found on the way and worth knowing about: two packages in the same binary declaring the same foreign symbol with different argument types is a build that fails, and nothing says so until both packages land in one link. core.errors frees a pointer, this box was freeing an integer, and the test suite is the first build that contains both.
The linter's unsafety check had a hole big enough to drive the box through, and now does not. It was a list of eight names, all of them types or free functions, and it never looked at the unsafe operations that hang off a safe type as methods. span.unsafe_ptr().as_unsafe_any_origin() is a raw pointer with its origin erased, it is the entire trick core.runtime.box exists to contain, and the word Pointer does not appear in it. The check now matches the whole unsafe_ prefix as a family, which covers the fifty odd names that exist today and the ones that arrive with the next release. A second fixture goes with it: the old one names the raw types outright, so it would still be rejected with the family match deleted and therefore cannot prove the family match is alive.
What's Changed
- core.runtime.box: the refcounted type erased heap box by @tamnd in #107
- lint: match the whole unsafe_ prefix, not a list of eight names by @tamnd in #108
- design: a field can hold UntrackedOrigin, and now a probe says so by @tamnd in #109
- core.io: Reader, Writer and copy, on capability bits rather than assertions by @tamnd in #110
- core.iter: Cursor, the shape fallible iteration has to be written in by @tamnd in #111
- core.io: the rest of the surface, on the same capability bits by @tamnd in #113
- core.bufio: the buffered reader, writer and scanner, over core.io by @tamnd in #114
- core.unicode.utf8, and the deletion it was written to allow by @tamnd in #116
- core.cmp, core.sort and core.slices: the ordering stack by @tamnd in #117
- core.maps: all ten symbols over Dict by @tamnd in #118
- core.unicode: the character database, generated and checked exhaustively by @tamnd in #119
- core.bytes: Buffer, Reader and the search half by @tamnd in #120
- changelog: cut v0.3.0, the end of M2 by @tamnd in #121
Full Changelog: v0.2.0...v0.3.0