[SPARK-3308][SQL] Ability to read JSON Arrays as tables #2400

yhuai · 2014-09-15T19:51:23Z

This PR aims to support reading top level JSON arrays and take every element in such an array as a row (an empty array will not generate a row).

JIRA: https://issues.apache.org/jira/browse/SPARK-3308

SparkQA · 2014-09-15T19:54:17Z

QA tests have started for PR 2400 at commit 990077a.

This patch merges cleanly.

SparkQA · 2014-09-15T21:21:24Z

QA tests have finished for PR 2400 at commit 990077a.

This patch passes unit tests.
This patch merges cleanly.
This patch adds no public classes.

marmbrus · 2014-09-16T18:17:26Z

Oh wow, this was easier/cleaner than I expected. LGTM

…type in the same fieild ## What changes were proposed in this pull request? This #2400 added the support to parse JSON rows wrapped with an array. However, this throws an exception when the given data contains array data and struct data in the same field as below: ```json {"a": {"b": 1}} {"a": []} ``` and the schema is given as below: ```scala val schema = StructType( StructField("a", StructType( StructField("b", StringType) :: Nil )) :: Nil) ``` - **Before** ```scala sqlContext.read.schema(schema).json(path).show() ``` ```scala Exception in thread "main" org.apache.spark.SparkException: Job aborted due to stage failure: Task 7 in stage 0.0 failed 4 times, most recent failure: Lost task 7.3 in stage 0.0 (TID 10, 192.168.1.170): java.lang.ClassCastException: org.apache.spark.sql.types.GenericArrayData cannot be cast to org.apache.spark.sql.catalyst.InternalRow at org.apache.spark.sql.catalyst.expressions.BaseGenericInternalRow$class.getStruct(rows.scala:50) at org.apache.spark.sql.catalyst.expressions.GenericMutableRow.getStruct(rows.scala:247) at org.apache.spark.sql.catalyst.expressions.GeneratedClass$SpecificPredicate.eval(Unknown Source) ... ``` - **After** ```scala sqlContext.read.schema(schema).json(path).show() ``` ```bash +----+ | a| +----+ | [1]| |null| +----+ ``` For other data types, in this case it converts the given values are `null` but only this case emits an exception. This PR makes the support for wrapped rows applied only at the top level. ## How was this patch tested? Unit tests were used and `./dev/run_tests` for code style tests. Author: hyukjinkwon <gurwls223@gmail.com> Closes #11752 from HyukjinKwon/SPARK-3308-follow-up.

…type in the same fieild ## What changes were proposed in this pull request? This apache#2400 added the support to parse JSON rows wrapped with an array. However, this throws an exception when the given data contains array data and struct data in the same field as below: ```json {"a": {"b": 1}} {"a": []} ``` and the schema is given as below: ```scala val schema = StructType( StructField("a", StructType( StructField("b", StringType) :: Nil )) :: Nil) ``` - **Before** ```scala sqlContext.read.schema(schema).json(path).show() ``` ```scala Exception in thread "main" org.apache.spark.SparkException: Job aborted due to stage failure: Task 7 in stage 0.0 failed 4 times, most recent failure: Lost task 7.3 in stage 0.0 (TID 10, 192.168.1.170): java.lang.ClassCastException: org.apache.spark.sql.types.GenericArrayData cannot be cast to org.apache.spark.sql.catalyst.InternalRow at org.apache.spark.sql.catalyst.expressions.BaseGenericInternalRow$class.getStruct(rows.scala:50) at org.apache.spark.sql.catalyst.expressions.GenericMutableRow.getStruct(rows.scala:247) at org.apache.spark.sql.catalyst.expressions.GeneratedClass$SpecificPredicate.eval(Unknown Source) ... ``` - **After** ```scala sqlContext.read.schema(schema).json(path).show() ``` ```bash +----+ | a| +----+ | [1]| |null| +----+ ``` For other data types, in this case it converts the given values are `null` but only this case emits an exception. This PR makes the support for wrapped rows applied only at the top level. ## How was this patch tested? Unit tests were used and `./dev/run_tests` for code style tests. Author: hyukjinkwon <gurwls223@gmail.com> Closes apache#11752 from HyukjinKwon/SPARK-3308-follow-up.

Handle top level JSON arrays.

990077a

asfgit closed this in 7583699 Sep 16, 2014

yhuai deleted the SPARK-3308 branch October 6, 2014 16:23

HyukjinKwon mentioned this pull request Mar 16, 2016

[SPARK-13719][SQL] Parse JSON rows having an array type and a struct type in the same fieild #11752

Closed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[SPARK-3308][SQL] Ability to read JSON Arrays as tables #2400

[SPARK-3308][SQL] Ability to read JSON Arrays as tables #2400

yhuai commented Sep 15, 2014

SparkQA commented Sep 15, 2014

SparkQA commented Sep 15, 2014

marmbrus commented Sep 16, 2014

[SPARK-3308][SQL] Ability to read JSON Arrays as tables #2400

[SPARK-3308][SQL] Ability to read JSON Arrays as tables #2400

Conversation

yhuai commented Sep 15, 2014

SparkQA commented Sep 15, 2014

SparkQA commented Sep 15, 2014

marmbrus commented Sep 16, 2014