APIs to read and write code points. #145

swankjesse · 2015-04-25T16:02:55Z

The String APIs transcode UTF-16 to UTF-8 and back.

These APIs avoid the UTF-16 intermediate form altogether, and
go right from UTF-8 to a codepoint and back.

swankjesse · 2015-04-25T16:04:45Z

okio/src/main/java/okio/Buffer.java

+
+    } else if (codePoint < 0x10000) {
+      if (codePoint >= 0xd800 && codePoint <= 0xdfff) {
+        throw new IllegalArgumentException(


Interesting design decision here. What to do when the input's not quite right:

we could throw

we could encode it (it's still recoverable on the other end, though encoders might return the replacement char)

we could encode the replacement character ourselves

Thoughts?

Tough. Option 2 doesn't seem reasonable to me. I'm inclined to lean towards 1 (as implemented), but I see the value in 3. I think it's hard to be too forgiving at this level. A higher-level API could handle the replacement character writing, but when you're at this level I think throwing is correct.

swankjesse · 2015-04-25T16:06:50Z

HttpUrl wants this. It's defined in terms of code points, and wants to parse codepoint-by-codepoint. I'm currently doing it Java char-by-char, but it's awkward.
square/okhttp#1486

JakeWharton · 2015-04-26T05:31:34Z

okio/src/main/java/okio/BufferedSource.java

+   * Removes and returns a single UTF-8 code point, reading between 1 and 4 bytes as necessary.
+   *
+   * <p>If this source is exhausted before a complete code point can be read, this throws an {@link
+   * java.io.EOFException} and consumes no input.


Nice doc. I don't think we're very good about documenting the behavior when methods throw (related to the contents of the underlying buffer).

JakeWharton · 2015-04-26T06:00:07Z

LGTM

The String APIs transcode UTF-16 to UTF-8 and back. These APIs avoid the UTF-16 intermediate form altogether, and go right from UTF-8 to a codepoint and back.

APIs to read and write code points.

swankjesse reviewed Apr 25, 2015
View reviewed changes

JakeWharton reviewed Apr 26, 2015
View reviewed changes

APIs to read and write code points.

85e9f94

The String APIs transcode UTF-16 to UTF-8 and back. These APIs avoid the UTF-16 intermediate form altogether, and go right from UTF-8 to a codepoint and back.

swankjesse force-pushed the jwilson_0425_utf8_code_points branch from 6529430 to 85e9f94 Compare April 26, 2015 12:30

swankjesse added a commit that referenced this pull request Apr 26, 2015

Merge pull request #145 from square/jwilson_0425_utf8_code_points

3c61fdb

APIs to read and write code points.

swankjesse merged commit 3c61fdb into master Apr 26, 2015

swankjesse deleted the jwilson_0425_utf8_code_points branch May 16, 2015 15:55

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

APIs to read and write code points. #145

APIs to read and write code points. #145

swankjesse commented Apr 25, 2015

swankjesse Apr 25, 2015

JakeWharton Apr 26, 2015

swankjesse commented Apr 25, 2015

JakeWharton Apr 26, 2015

JakeWharton commented Apr 26, 2015

APIs to read and write code points. #145

APIs to read and write code points. #145

Conversation

swankjesse commented Apr 25, 2015

swankjesse Apr 25, 2015

Choose a reason for hiding this comment

JakeWharton Apr 26, 2015

Choose a reason for hiding this comment

swankjesse commented Apr 25, 2015

JakeWharton Apr 26, 2015

Choose a reason for hiding this comment

JakeWharton commented Apr 26, 2015