SOLR-16810: Under certain situations Solr produces managed schema XML with duplicate fields #1654

thiru-mg · 2023-05-21T13:57:39Z

https://issues.apache.org/jira/browse/SOLR-16810

Description

While persisting the ManagedIndexSchema as XML, non-printable characters in field names get escaped as #nn;, where nn is the decimal representation of the non-printable character. For example, if the field name has the byte 0x14, it gets escaped as #20;. This in indistinguishable from the literal #20; in the field name. If we have two fields - one with the non-printable character and the other with the literal string, two fields get generated with the same name. Loading the resulting XML, naturally, causes an exception. To fix this, any occurrence of literal # in the field name should be escaped, with say ##.
A second problem is that while escaping happens when generating XML, the corresponding unescaping does not happen on loading it.

Solution

Both the suggested fixes are here. There are tests to expose the bugs and the fixes that pass the tests.

Tests

Added two sets of tests - one for IndexSchema.java and the other for XML.java

Checklist

Please review the following and check all that apply:

I have reviewed the guidelines for How to Contribute and my code conforms to the standards described there to the best of my ability.
I have created a Jira issue and added the issue ID to my pull request title.
I have given Solr maintainers access to contribute to my PR branch. (optional but recommended)
I have developed this patch against the main branch.
I have run ./gradlew check.
I have added tests for my changes.
I have added documentation for the Reference Guide

sonatype-lift · 2023-05-22T04:32:12Z

solr/solrj/src/java/org/apache/solr/common/util/XML.java

+          sb.append(g3);
+        } else {
+          sb.append(m.group(2));
+          sb.append(escapes[m.group(5).charAt(0)]);


NULL_DEREFERENCE: object returned by m.group(5) could be null and is dereferenced at line 247.

ℹ️ Expand to see all @sonatype-lift commands

You can reply with the following commands. For example, reply with @sonatype-lift ignoreall to leave out all findings.

Command Usage

@sonatype-lift ignore Leave out the above finding from this PR

@sonatype-lift ignoreall Leave out all the existing findings from this PR

@sonatype-lift exclude <file|issue|path|tool> Exclude specified file|issue|path|tool from Lift findings by updating your config.toml file

Note: When talking to LiftBot, you need to refresh the page to see its response.
_{Click here to add LiftBot to another repo.}

github-actions · 2024-02-15T12:21:04Z

This PR had no visible activity in the past 60 days, labeling it as stale. Any new activity will remove the stale label. To attract more reviewers, please tag someone or notify the dev@solr.apache.org mailing list. Thank you for your contribution!

epugh · 2024-02-15T13:21:39Z

Thank you StaleBot..... I just checked the JIRA and I was last to chime in, so I'll take this and try and get it over the finish line.

github-actions · 2024-05-01T00:00:28Z

This PR had no visible activity in the past 60 days, labeling it as stale. Any new activity will remove the stale label. To attract more reviewers, please tag someone or notify the dev@solr.apache.org mailing list. Thank you for your contribution!

thiru-mg force-pushed the SOLR-16810 branch from d0b1986 to 6ed0ca5 Compare May 21, 2023 14:12

Fixed field-name duplicates in XML

84c43d7

thiru-mg force-pushed the SOLR-16810 branch from 6ed0ca5 to 84c43d7 Compare May 22, 2023 03:45

sonatype-lift bot reviewed May 22, 2023

View reviewed changes

Fixed an issue reported by sonatype-lift

114d3ec

noblepaul changed the title ~~SOLR-16810: Under certain situations Solr produces managed schema XML that cannot be loaded~~ SOLR-16810: Under certain situations Solr produces managed schema XML with duplicate fields May 22, 2023

Thiruvalluvan M g added 4 commits May 23, 2023 17:54

Fixed field-name duplicates in XML

c5fb4f5

Fixed field-name duplicates in XML

7c0691f

Fixed field-name duplicates in XML

e80a66c

Fixed field-name duplicates in XML

ef646ba

github-actions bot added the stale PR not updated in 60 days label Feb 15, 2024

Merge remote-tracking branch 'upstream/main' into pr/1654

e46a45b

github-actions bot added the client:solrj label Feb 15, 2024

github-actions bot removed the stale PR not updated in 60 days label Mar 1, 2024

github-actions bot added the stale PR not updated in 60 days label May 1, 2024

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

SOLR-16810: Under certain situations Solr produces managed schema XML with duplicate fields #1654

SOLR-16810: Under certain situations Solr produces managed schema XML with duplicate fields #1654

thiru-mg commented May 21, 2023 •

edited

sonatype-lift bot May 22, 2023

github-actions bot commented Feb 15, 2024

epugh commented Feb 15, 2024

github-actions bot commented May 1, 2024

Command	Usage
`@sonatype-lift ignore`	Leave out the above finding from this PR
`@sonatype-lift ignoreall`	Leave out all the existing findings from this PR
`@sonatype-lift exclude <file\|issue\|path\|tool>`	Exclude specified `file\|issue\|path\|tool` from Lift findings by updating your config.toml file

SOLR-16810: Under certain situations Solr produces managed schema XML with duplicate fields #1654

Are you sure you want to change the base?

SOLR-16810: Under certain situations Solr produces managed schema XML with duplicate fields #1654

Conversation

thiru-mg commented May 21, 2023 • edited

Description

Solution

Tests

Checklist

sonatype-lift bot May 22, 2023

Choose a reason for hiding this comment

github-actions bot commented Feb 15, 2024

epugh commented Feb 15, 2024

github-actions bot commented May 1, 2024

thiru-mg commented May 21, 2023 •

edited