[SPARK-36797][SQL] Union should resolve nested columns as top-level columns #34038

viirya · 2021-09-18T07:54:11Z

What changes were proposed in this pull request?

This patch proposes to generalize the resolving-by-position behavior to nested columns for Union.

Why are the changes needed?

Union, by the API definition, resolves columns by position. Currently we only follow this behavior at top-level columns, but not nested columns.

As we are making nested columns as first-class citizen, the nested-column-only limitation and the difference between top-level column and nested column do not make sense. We should also resolve nested columns like top-level columns for Union.

Does this PR introduce any user-facing change?

Yes. After this change, Union also resolves nested columns by position.

How was this patch tested?

Added tests.

viirya · 2021-09-18T07:55:42Z

sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/analysis/CheckAnalysis.scala

-                      |column types. ${dt1.catalogString} <> ${dt2.catalogString} at the
-                      |${ordinalNumber(ci)} column of the ${ordinalNumber(ti + 1)} table
-                    """.stripMargin.replace("\n", " ").trim())
+              if (!isUnion) {


Not sure if we should also generalize to all set operations? Although it looks reasonable, but by their API definition seems we don't have the by-position definition as Union.

How are top-level columns handled for other set operations? In general I feel Spark SQL is built around column names by default, not positions, so I would expect it to be by-name. I was surprised to realize recently that union is by-position.

I think these set operations work basically the same. But at the API doc, we don't have document it for all set operations except for union. The by-position resolution for union, I think, is to follow SQL. It only requires the columns to union have the same data types in same order, but not column names.

I think it's better to make top-level and nested columns consistent in other set operations as well, which is, do by-position resolution. We can't go with the other direction as that will be a breaking change.

Ok. I think it makes more sense. I will make other set operations as by-position too at nested column level.

BTW I will make the change for other set operations in another PR (JIRA). It might require more change (doc, test..).

SparkQA · 2021-09-18T08:37:30Z

Kubernetes integration test starting
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/47942/

SparkQA · 2021-09-18T08:46:02Z

Kubernetes integration test status failure
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/47942/

SparkQA · 2021-09-18T18:45:46Z

Kubernetes integration test starting
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/47944/

SparkQA · 2021-09-18T18:54:40Z

Kubernetes integration test status failure
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/47944/

SparkQA · 2021-09-18T21:42:41Z

Kubernetes integration test starting
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/47946/

SparkQA · 2021-09-18T21:51:35Z

Kubernetes integration test status failure
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/47946/

SparkQA · 2021-09-19T01:49:32Z

Test build #143438 has finished for PR 34038 at commit 47bec5d.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

viirya · 2021-09-20T22:03:33Z

cc @cloud-fan @dongjoon-hyun @xkrogen @HyukjinKwon

dongjoon-hyun · 2021-09-20T22:07:52Z

Thank you for pinging me, @viirya .

xkrogen · 2021-09-20T22:28:26Z

sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/analysis/CheckAnalysis.scala

-                      |column types. ${dt1.catalogString} <> ${dt2.catalogString} at the
-                      |${ordinalNumber(ci)} column of the ${ordinalNumber(ti + 1)} table
-                    """.stripMargin.replace("\n", " ").trim())
+              if (!isUnion) {


How are top-level columns handled for other set operations? In general I feel Spark SQL is built around column names by default, not positions, so I would expect it to be by-name. I was surprised to realize recently that union is by-position.

xkrogen · 2021-09-20T22:38:09Z

sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/analysis/CheckAnalysis.scala

+              if (!isUnion) {
+                dataTypes(child).zip(ref).zipWithIndex.foreach { case ((dt1, dt2), ci) =>
+                  // SPARK-18058: we shall not care about the nullability of columns
+                  if (TypeCoercion.findWiderTypeForTwo(dt1.asNullable, dt2.asNullable).isEmpty) {
+                    failAnalysis(
+                      s"""
+                         |${operator.nodeName} can only be performed on tables with the compatible
+                         |column types. ${dt1.catalogString} <> ${dt2.catalogString} at the
+                         |${ordinalNumber(ci)} column of the ${ordinalNumber(ti + 1)} table
+                      """.stripMargin.replace("\n", " ").trim())
+                  }
+                }
+              } else {
+                // `TypeCoercion` takes care of type coercion already. If any columns or nested
+                // columns are not compatible, we detect it here and throw analysis exception.
+                val typeChecker = (dt1: DataType, dt2: DataType) => {
+                  !TypeCoercion.findWiderTypeForTwo(dt1.asNullable, dt2.asNullable).isEmpty
+                }
+                dataTypes(child).zip(ref).zipWithIndex.foreach { case ((dt1, dt2), ci) =>
+                  if (!DataType.equalsStructurally(dt1, dt2, true, typeChecker)) {
+                    failAnalysis(
+                      s"""
+                         |${operator.nodeName} can only be performed on tables with the compatible
+                         |column types. ${dt1.catalogString} <> ${dt2.catalogString} at the
+                         |${ordinalNumber(ci)} column of the ${ordinalNumber(ti + 1)} table
+                      """.stripMargin.replace("\n", " ").trim())
+                  }


Maybe we can simplify like:

val dataTypesAreCompatibleFn = if (isUnion) { // `TypeCoercion` takes care of type coercion already. If any columns or nested // columns are not compatible, we detect it here and throw analysis exception. val typeChecker = (dt1: DataType, dt2: DataType) => { !TypeCoercion.findWiderTypeForTwo(dt1.asNullable, dt2.asNullable).isEmpty } (dt1: DataType, dt2: DataType) => !DataType.equalsStructurally(dt1, dt2, true, typeChecker) } else { // SPARK-18058: we shall not care about the nullability of columns (dt1: DataType, dt2: DataType) => TypeCoercion.findWiderTypeForTwo(dt1.asNullable, dt2.asNullable).isEmpty } dataTypes(child).zip(ref).zipWithIndex.foreach { case ((dt1, dt2), ci) => if (dataTypesAreCompatibleFn(dt1, dt2)) { failAnalysis( s""" |${operator.nodeName} can only be performed on tables with the compatible |column types. ${dt1.catalogString} <> ${dt2.catalogString} at the |${ordinalNumber(ci)} column of the ${ordinalNumber(ti + 1)} table """.stripMargin.replace("\n", " ").trim()) } }

...alyst/src/main/scala/org/apache/spark/sql/catalyst/plans/logical/basicLogicalOperators.scala

SparkQA · 2021-09-22T19:50:51Z

Kubernetes integration test starting
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48022/

SparkQA · 2021-09-22T20:52:20Z

Kubernetes integration test status failure
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48022/

viirya · 2021-09-22T21:35:42Z

retest this please

SparkQA · 2021-09-22T22:53:24Z

Kubernetes integration test starting
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48029/

SparkQA · 2021-09-22T23:41:39Z

Kubernetes integration test status failure
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48029/

SparkQA · 2021-09-23T01:43:33Z

Test build #143519 has finished for PR 34038 at commit 8db8b50.

This patch fails Spark unit tests.
This patch merges cleanly.
This patch adds no public classes.

HyukjinKwon · 2021-09-23T02:07:01Z

retest this please

SparkQA · 2021-09-23T02:47:55Z

Test build #143520 has finished for PR 34038 at commit 8db8b50.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

SparkQA · 2021-09-23T02:59:59Z

Kubernetes integration test starting
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48034/

SparkQA · 2021-09-24T03:01:13Z

Kubernetes integration test starting
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48089/

SparkQA · 2021-09-24T03:46:09Z

Kubernetes integration test status failure
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48089/

SparkQA · 2021-09-24T06:08:01Z

Kubernetes integration test starting
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48100/

SparkQA · 2021-09-24T07:12:11Z

Kubernetes integration test status failure
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48100/

cloud-fan · 2021-09-24T09:51:58Z

sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/analysis/TypeCoercion.scala

        // Otherwise, record the result in the queue and find the type for the next column
        case Some(widenType) =>
-          castedTypes.enqueue(widenType)
+          castedTypes.enqueue(Some(widenType))
          getWidestTypes(children, attrIndex + 1, castedTypes)


the code can be simplified

findWiderCommonType(children.map(_.output(attrIndex).dataType)).map { widenTypeOpt => castedTypes.enqueue(widenTypeOpt) getWidestTypes(children, attrIndex + 1, castedTypes) }

findWiderCommonType returns Opion[DataType]. map can iterate over the DataType if any, but we still need to enqueue the None.

Just simplified to

val widenTypeOpt = findWiderCommonType(children.map(_.output(attrIndex).dataType)) castedTypes.enqueue(widenTypeOpt) getWidestTypes(children, attrIndex + 1, castedTypes)

cloud-fan · 2021-09-24T09:53:05Z

sql/core/src/test/resources/sql-tests/results/postgreSQL/union.sql.out

-org.apache.spark.SparkException
-Failed to merge incompatible data types decimal(38,18) and string
+org.apache.spark.sql.AnalysisException
+Union can only be performed on tables with the compatible column types. string <> decimal(38,18) at the first column of the second table


string <> decimal(38,18) at the first column of the second table is it valid English syntax?

Ha, this comes from CheckAnalysis's original error message. We can improve it, although there are some more tests relying on the error message.

Updated the error message. Please let me know if it looks good to you.

SparkQA · 2021-09-24T10:03:53Z

Test build #143589 has finished for PR 34038 at commit be31929.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

SparkQA · 2021-09-24T17:18:58Z

Kubernetes integration test starting
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48119/

SparkQA · 2021-09-24T18:17:43Z

Kubernetes integration test status failure
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48119/

SparkQA · 2021-09-24T21:22:52Z

Test build #143607 has finished for PR 34038 at commit 80bb6e1.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

SparkQA · 2021-09-25T08:54:54Z

Kubernetes integration test starting
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48132/

SparkQA · 2021-09-25T09:38:16Z

Kubernetes integration test status failure
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48132/

SparkQA · 2021-09-25T10:06:15Z

Test build #143620 has finished for PR 34038 at commit 5e17567.

This patch fails Spark unit tests.
This patch merges cleanly.
This patch adds no public classes.

SparkQA · 2021-09-25T19:09:26Z

Kubernetes integration test starting
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48136/

SparkQA · 2021-09-25T20:08:47Z

Kubernetes integration test status failure
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48136/

SparkQA · 2021-09-25T22:49:25Z

Test build #143624 has finished for PR 34038 at commit 87cf3e1.

This patch fails Spark unit tests.
This patch merges cleanly.
This patch adds no public classes.

SparkQA · 2021-09-26T00:55:52Z

Kubernetes integration test starting
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48139/

SparkQA · 2021-09-26T01:39:41Z

Kubernetes integration test status failure
URL: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder-K8s/48139/

SparkQA · 2021-09-26T05:05:57Z

Test build #143628 has finished for PR 34038 at commit f382cf2.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

cloud-fan · 2021-09-27T07:51:55Z

thanks, merging to master!

Union should resolve nested columns as top-level columns.

f9c133c

github-actions bot added the SQL label Sep 18, 2021

viirya commented Sep 18, 2021

View reviewed changes

This comment has been minimized.

Sign in to view

Keep original check logic.

b91c08e

This comment has been minimized.

Sign in to view

Fix test.

47bec5d

viirya mentioned this pull request Sep 19, 2021

[SPARK-36673][SQL] Fix incorrect schema of nested types of union #34025

Closed

xkrogen reviewed Sep 20, 2021

View reviewed changes

cloud-fan reviewed Sep 22, 2021

View reviewed changes

...alyst/src/main/scala/org/apache/spark/sql/catalyst/plans/logical/basicLogicalOperators.scala Show resolved Hide resolved

For review comment.

8db8b50

This comment has been minimized.

Sign in to view

Update test result.

a0af93c

This comment has been minimized.

Sign in to view

Fix test.

be31929

cloud-fan reviewed Sep 24, 2021

View reviewed changes

Simplify code.

80bb6e1

Update exception message.

5e17567

Update test for updated error message.

87cf3e1

Update more test.

f382cf2

cloud-fan approved these changes Sep 27, 2021

View reviewed changes

cloud-fan closed this in 44070e0 Sep 27, 2021

viirya deleted the SPARK-36797 branch December 27, 2023 18:25

[SPARK-36797][SQL] Union should resolve nested columns as top-level columns #34038

[SPARK-36797][SQL] Union should resolve nested columns as top-level columns #34038

Conversation

viirya commented Sep 18, 2021

What changes were proposed in this pull request?

Why are the changes needed?

Does this PR introduce any user-facing change?

How was this patch tested?

Choose a reason for hiding this comment

Choose a reason for hiding this comment

viirya Sep 20, 2021 • edited Loading

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

SparkQA commented Sep 18, 2021

SparkQA commented Sep 18, 2021

This comment has been minimized.

SparkQA commented Sep 18, 2021

SparkQA commented Sep 18, 2021

This comment has been minimized.

SparkQA commented Sep 18, 2021

SparkQA commented Sep 18, 2021

SparkQA commented Sep 19, 2021

viirya commented Sep 20, 2021

dongjoon-hyun commented Sep 20, 2021

Choose a reason for hiding this comment

Choose a reason for hiding this comment

SparkQA commented Sep 22, 2021

SparkQA commented Sep 22, 2021

This comment has been minimized.

viirya commented Sep 22, 2021

SparkQA commented Sep 22, 2021

SparkQA commented Sep 22, 2021

SparkQA commented Sep 23, 2021

HyukjinKwon commented Sep 23, 2021

SparkQA commented Sep 23, 2021

SparkQA commented Sep 23, 2021

SparkQA commented Sep 24, 2021

SparkQA commented Sep 24, 2021

This comment has been minimized.

SparkQA commented Sep 24, 2021

SparkQA commented Sep 24, 2021

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

SparkQA commented Sep 24, 2021

SparkQA commented Sep 24, 2021

SparkQA commented Sep 24, 2021

SparkQA commented Sep 24, 2021

SparkQA commented Sep 25, 2021

SparkQA commented Sep 25, 2021

SparkQA commented Sep 25, 2021

SparkQA commented Sep 25, 2021

SparkQA commented Sep 25, 2021

SparkQA commented Sep 25, 2021

SparkQA commented Sep 26, 2021

SparkQA commented Sep 26, 2021

SparkQA commented Sep 26, 2021

cloud-fan commented Sep 27, 2021

viirya Sep 20, 2021 •

edited

Loading