v1.1.7
Unsafe slice to string conversation apparently could cause an allocation -- it was in tests, but only when converting a slice that was actually an 8192 byte array in the benchmark itself. We can just duplicate the logic to drop the allocation. This was tested by deleting the 128 bit assembly functions, and as it turns out, hand rolled assembly is still a bit faster.