[SPARK-48148][CORE] JSON objects should not be modified when read as STRING #46408

eric-maynard · 2024-05-06T22:21:02Z

What changes were proposed in this pull request?

Currently, when reading a JSON like this:

{"a": {"b": -999.99999999999999999999999999999999995}}

With the schema:

a STRING

Spark will yield a result like this:

{"b": -1000.0}

Other changes such as changes to the input string's whitespace may also occur. In some cases, we apply scientific notation to an input floating-point number when reading it as STRING.

This applies to reading JSON files (as with spark.read.json) as well as the SQL expression from_json.

Why are the changes needed?

Correctness issues may occur if a field is read as a STRING and then later parsed (e.g. with from_json) after the contents have been modified.

Does this PR introduce any user-facing change?

Yes, when reading non-string fields from a JSON object using the STRING type, we will now extract the field exactly as it appears.

How was this patch tested?

Added a test in JsonSuite.scala

Was this patch authored or co-authored using generative AI tooling?

No

sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/json/JacksonParser.scala

sql/core/src/test/scala/org/apache/spark/sql/execution/datasources/json/JsonSuite.scala

sadikovi · 2024-05-08T20:18:29Z

sql/core/src/test/scala/org/apache/spark/sql/execution/datasources/json/JsonSuite.scala

+
+      val df = spark.read.schema("data STRING").json(path.getAbsolutePath)
+
+      val expected = s"""{"v": ${granularFloat}}"""


Can you add more test cases for the following?

{"data": {"v": "abc"}}, expected: "{"v": "abc"}"

{"data":{"v": "0.999"}}, expected: "{"v": "0.999"}"

{"data": [1, 2, 3]}, expected: "[1, 2, 3]"

{"data": }, expected the object as string.

Added more tests -- can you clarify the last example and what we expect that to do? It seems like invalid JSON

sadikovi · 2024-05-08T20:19:32Z

sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/json/JacksonParser.scala

-          Utils.tryWithResource(factory.createGenerator(writer, JsonEncoding.UTF8)) {
-            generator => generator.copyCurrentStructure(parser)
+          val startLocation = parser.getTokenLocation
+          startLocation.contentReference().getRawContent match {


Is there an existing API to get the remaining content as string? Also, would it work with multi-line JSON?

I was not able to find such an existing API -- there is JacksonParser.getText but that appears to simply get the current value if it's a string value.

wrt. multiline JSON, I have added a test to cover this. It seems that the content reference is not a byte array when using multiline mode.

sadikovi · 2024-05-08T20:30:08Z

cc @dongjoon-hyun @HyukjinKwon

HyukjinKwon · 2024-05-09T02:56:30Z

SPARK-48148: values are unchanged when read as string *** FAILED *** (134 milliseconds)

seems it fails

eric-maynard · 2024-05-09T16:33:50Z

sql/core/src/test/scala/org/apache/spark/sql/execution/datasources/json/JsonSuite.scala

        expectedExactData = Seq(s"""{"v": ${granularFloat}}""")
      )
      // In multiLine, we fall back to the inexact method:
      extractData(
-        s"""{"data": {"white":\n"space"}}""",


This makes \n no longer function as a newline here

eric-maynard · 2024-05-09T18:10:30Z

Hey @HyukjinKwon, can you take another look and possible re-trigger tests? I believe multiline should be working now.

HyukjinKwon · 2024-05-09T23:41:15Z

Merged to master.

HyukjinKwon · 2024-05-09T23:41:45Z

btw you can trigger on your own https://github.com/eric-maynard/spark/runs/24789350525 I can't trigger :-).

…STRING ### What changes were proposed in this pull request? Currently, when reading a JSON like this: ``` {"a": {"b": -999.99999999999999999999999999999999995}} ``` With the schema: ``` a STRING ``` Spark will yield a result like this: ``` {"b": -1000.0} ``` Other changes such as changes to the input string's whitespace may also occur. In some cases, we apply scientific notation to an input floating-point number when reading it as STRING. This applies to reading JSON files (as with `spark.read.json`) as well as the SQL expression `from_json`. ### Why are the changes needed? Correctness issues may occur if a field is read as a STRING and then later parsed (e.g. with `from_json`) after the contents have been modified. ### Does this PR introduce _any_ user-facing change? Yes, when reading non-string fields from a JSON object using the STRING type, we will now extract the field exactly as it appears. ### How was this patch tested? Added a test in `JsonSuite.scala` ### Was this patch authored or co-authored using generative AI tooling? No Closes apache#46408 from eric-maynard/SPARK-48148. Lead-authored-by: Eric Maynard <eric.maynard@databricks.com> Co-authored-by: Hyukjin Kwon <gurwls223@gmail.com> Signed-off-by: Hyukjin Kwon <gurwls223@apache.org>

initial commit

46bdf7d

github-actions bot added the SQL label May 6, 2024

eric-maynard commented May 6, 2024

View reviewed changes

sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/json/JacksonParser.scala Show resolved Hide resolved

sadikovi reviewed May 8, 2024

View reviewed changes

eric-maynard and others added 4 commits May 8, 2024 14:18

add flag

79e8457

improve tests

07083fd

should be stable

2f697da

Apply suggestions from code review

a5c3761

HyukjinKwon approved these changes May 9, 2024

View reviewed changes

sadikovi approved these changes May 9, 2024

View reviewed changes

eric-maynard commented May 9, 2024

View reviewed changes

stable; fixed multiline

9a78a8d

eric-maynard added 3 commits May 9, 2024 11:10

pull master

114d8a3

resolve conflicts

e9cb9d2

polish

fc77ed0

HyukjinKwon closed this in b47d785 May 9, 2024

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[SPARK-48148][CORE] JSON objects should not be modified when read as STRING #46408

[SPARK-48148][CORE] JSON objects should not be modified when read as STRING #46408

eric-maynard commented May 6, 2024 •

edited

sadikovi May 8, 2024

eric-maynard May 8, 2024

sadikovi May 8, 2024

eric-maynard May 8, 2024

eric-maynard May 8, 2024 •

edited

sadikovi commented May 8, 2024

HyukjinKwon commented May 9, 2024

eric-maynard May 9, 2024

eric-maynard commented May 9, 2024

HyukjinKwon commented May 9, 2024

HyukjinKwon commented May 9, 2024


		val df = spark.read.schema("data STRING").json(path.getAbsolutePath)

		val expected = s"""{"v": ${granularFloat}}"""

[SPARK-48148][CORE] JSON objects should not be modified when read as STRING #46408

[SPARK-48148][CORE] JSON objects should not be modified when read as STRING #46408

Conversation

eric-maynard commented May 6, 2024 • edited

What changes were proposed in this pull request?

Why are the changes needed?

Does this PR introduce any user-facing change?

How was this patch tested?

Was this patch authored or co-authored using generative AI tooling?

sadikovi May 8, 2024

Choose a reason for hiding this comment

eric-maynard May 8, 2024

Choose a reason for hiding this comment

sadikovi May 8, 2024

Choose a reason for hiding this comment

eric-maynard May 8, 2024

Choose a reason for hiding this comment

eric-maynard May 8, 2024 • edited

Choose a reason for hiding this comment

sadikovi commented May 8, 2024

HyukjinKwon commented May 9, 2024

eric-maynard May 9, 2024

Choose a reason for hiding this comment

eric-maynard commented May 9, 2024

HyukjinKwon commented May 9, 2024

HyukjinKwon commented May 9, 2024

eric-maynard commented May 6, 2024 •

edited

eric-maynard May 8, 2024 •

edited