Skip to content

Bug fixes for ToCamelCase and Translate, plus a Scrub doc clarification

Latest

Choose a tag to compare

@huandu huandu released this 30 Sep 04:50
96b69c7

This release documents everything since v1.6.0, including the two changes which were parked on the v1.6.1 tag. That tag never got a release note of its own, so its changes are listed here as well and the changelog link at the bottom compares against v1.6.0.

Both Translate bugs silently returned a wrong string instead of reporting an error, so please read the Behaviour changes section if you pass patterns which contain a literal replacement character.

Bug fixes

Translate / Delete / Count: a literal U+FFFD in a pattern

The internal marker for "no rune" was utf8.RuneError, which is the very same value as the rune U+FFFD. A literal replacement character in a from or to pattern was therefore mistaken for the end of the pattern: the rune was dropped from the pattern and every pattern rune behind it was mismapped.

  • Translate("\uFFFDa", "\uFFFDa", "xy") returned "\uFFFDx" instead of "xy".
  • Delete("\uFFFDa\uFFFDb", "\uFFFDa") returned "\uFFFD\uFFFDb" instead of "b", and Count("\uFFFDa\uFFFDb", "\uFFFDa") returned 1 instead of 3.
  • In a to pattern the documented "repeat the last rune in to" rule broke as well: Translate("abc", "abc", "x\uFFFD") returned "xxx" instead of "x\uFFFD\uFFFD".

The internal marker is noRune = -1 now, a value which decoding a string can never produce, so U+FFFD is an ordinary rune in a pattern. The historical fallback for a pattern which parses to no rune at all ("-", "---", "\\", and the same shapes behind a range such as "0-9-") is kept, and is pinned by a test now.

Translate: the last rune of a from range was dropped when to contains literals

When a from range was mapped onto the literal runes of a to pattern, the last rune of the range never got a mapping: it was left unchanged in the result, and as soon as more source runes followed, their replacement runes shifted.

  • Translate("ab", "a-b", "xy") returned "xb" instead of "xy".
  • Translate("abc", "a-c", "xyz") returned "xyc" instead of "xyz".
  • Translate("abc", "a-bc", "xyz") returned "xby" instead of "xyz", i.e. 'c' was mapped to the rune 'b' should have used.

Translator.TranslateRune was affected too: NewTranslator("a-b", "xy").TranslateRune('b') returned ('b', false) although 'b' is part of the pattern. It returns ('y', true) now.

ToCamelCase: a single uppercase rune behind a separator was lowercased

ToCamelCase lowercased the last uppercase initial of an input, so it disagreed with the lowercase spelling of the same word:

  • ToCamelCase("option_A") returned "optiona" instead of "optionA"; the same applied to "option-A" and "option A".
  • ToCamelCase("HTTP_X") returned "httpx" instead of "httpX", and ToCamelCase("é_Ö") returned "éö" instead of "éÖ".

The final lowercasing step was redundant — the loop already normalizes the uppercase runes inside each word — so it was removed. ToPascalCase is not affected, and neither is input which starts with an uppercase word: ToCamelCase("URL_Parser") is still "urlParser".

Scrub: documentation only

Scrub replaces invalid UTF-8 bytes and valid U+FFFD runes with repl, and a consecutive run of them is replaced only once. The doc comment now says so and has samples: Scrub("a\uFFFDb", "?") => "a?b". Runtime behaviour is unchanged.

Behaviour changes

  • ToCamelCase keeps a single uppercase rune behind a separator: "option_A" → "optionA" (it was "optiona"), "HTTP_X" → "httpX" (it was "httpx").
  • A pattern which contains a literal U+FFFD behaves differently now, because the rune is part of the pattern instead of being removed from it. The results listed above are the only ones which change.
  • Delete, Count and Squeeze are not affected by the range fix, and ^-reverted patterns are not affected by either of the two Translate fixes.

Tests

  • 16 new table driven cases for the range fix: ascending and descending source ranges, shorter and longer replacement patterns, runes before and after a range, several ranges in one pattern, Unicode, U+0000 as a replacement rune, and an escaped - in the to pattern.
  • Two focused tests for the replacement character fix: literal positions, range endpoints, escaped and reverted patterns, the "repeat the last rune" rule, and the shared Delete, Count and TranslateRune paths.
  • Six new ToCamelCase and six new ToPascalCase cases for the camel case fix, covering underscore, hyphen and space separators, repeated separators, an uppercase prefix and non-ASCII runes.
  • Statement coverage stays at 99.9%.
  • The two Translate fixes were cross checked with a differential test against an independent implementation of the documented pattern semantics; no mismatch was found in tens of thousands of random from/to pattern pairs. The camel case change was compared against the previous implementation over 22,621 short strings, and only the intended results differ.

Thanks

Thank you for the four pull requests in this release:

  • #68 Clarify Scrub handling of U+FFFD by @x0Lazarus
  • #69 Preserve trailing uppercase initials in camel case by @jakezwang
  • #70 Preserve replacement characters in translation patterns by @x0Lazarus
  • #71 Preserve the last character when translating a range to literals by @x0Lazarus

Full Changelog: v1.6.0...v1.6.2