Improve handling of unsupported Matcher attributes #4070

adrianeboyd · 2019-08-02T08:11:06Z

Feature description

I decided to separate this from #4069 (see also #4063 ) since it's not quite the same issue.

With Matcher, it's very confusing that "bad" parts of patterns (e.g., with attributes that aren't supported) are silently discarded and end up matching every token rather than no token.

{'ASDF': True} shouldn't be equivalent to {}.

How slow is validation? Would it make sense to make validate=True the default? Are there simpler checks that could be added for when attributes are discarded that don't require full validation so that you could provide warnings or errors in these cases?

I think that less confusing default behavior would be:

{'ASDF': True} matches nothing
there are always warnings or errors when it is known that patterns will match nothing due to this kind of problem

The text was updated successfully, but these errors were encountered:

ines · 2019-08-02T08:21:51Z

Yes, I agree! 👍

How slow is validation? Would it make sense to make validate=True the default? Are there simpler checks that could be added for when attributes are discarded that don't require full validation so that you could provide warnings or errors in these cases?

Yes, we should probably just consider implementing simpler checks for just the key names and whether they map to a valid attribute or not.

The current validation uses jsonschema, which we've made an optional dependency (we didn't want it to drag in too many other dependencies down the tree). JSON schema validation is nice because it lets you write pretty complex constraints and gives you helpful and detailed error messages. It's also cross-Python compatible. (Once we drop 2.7 and 3.5, we could probably implement something similar using native Python type hints and make that enabled by default.)

Add more detailed token pattern checks without full JSON pattern validation and provide more detailed error messages. Addresses explosion#4070 (also related: explosion#4063, explosion#4100). * Check whether top-level attributes in patterns and attr for PhraseMatcher are in token pattern schema * Check whether attribute value types are supported in general (as opposed to per attribute with full validation) * Report various internal error types (OverflowError, AttributeError, KeyError) as ValueError with standard error messages * Check for tagger/parser in PhraseMatcher pipeline for attributes TAG, POS, LEMMA, and DEP * Add error messages with relevant details on how to use validate=True or nlp() instead of nlp.make_doc() * Support attr=TEXT for PhraseMatcher * Add NORM to schema * Expand tests for pattern validation, Matcher, PhraseMatcher, and EntityRuler

* Fix typo in rule-based matching docs * Improve token pattern checking without validation Add more detailed token pattern checks without full JSON pattern validation and provide more detailed error messages. Addresses #4070 (also related: #4063, #4100). * Check whether top-level attributes in patterns and attr for PhraseMatcher are in token pattern schema * Check whether attribute value types are supported in general (as opposed to per attribute with full validation) * Report various internal error types (OverflowError, AttributeError, KeyError) as ValueError with standard error messages * Check for tagger/parser in PhraseMatcher pipeline for attributes TAG, POS, LEMMA, and DEP * Add error messages with relevant details on how to use validate=True or nlp() instead of nlp.make_doc() * Support attr=TEXT for PhraseMatcher * Add NORM to schema * Expand tests for pattern validation, Matcher, PhraseMatcher, and EntityRuler * Remove unnecessary .keys() * Rephrase error messages * Add another type check to Matcher Add another type check to Matcher for more understandable error messages in some rare cases. * Support phrase_matcher_attr=TEXT for EntityRuler * Don't use spacy.errors in examples and bin scripts * Fix error code * Auto-format Also try get Azure pipelines to finally start a build :( * Update errors.py Co-authored-by: Ines Montani <ines@ines.io> Co-authored-by: Matthew Honnibal <honnibal+gh@gmail.com>

lock · 2019-09-20T21:42:53Z

This thread has been automatically locked since there has not been any recent activity after it was closed. Please open a new issue for related bugs.

ines added enhancement Feature requests and improvements feat / matcher Feature: Token, phrase and dependency matcher labels Aug 2, 2019

ines added the help wanted Contributions welcome! label Aug 2, 2019

adrianeboyd mentioned this issue Aug 12, 2019

Improve token pattern checking without validation #4105

Merged

3 tasks

ines added feat / ux Feature: User experience, error messages etc. and removed help wanted Contributions welcome! labels Aug 15, 2019

ines closed this as completed Aug 21, 2019

lock bot locked as resolved and limited conversation to collaborators Sep 20, 2019

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Improve handling of unsupported Matcher attributes #4070

Improve handling of unsupported Matcher attributes #4070

adrianeboyd commented Aug 2, 2019

ines commented Aug 2, 2019

lock bot commented Sep 20, 2019

Improve handling of unsupported Matcher attributes #4070

Improve handling of unsupported Matcher attributes #4070

Comments

adrianeboyd commented Aug 2, 2019

Feature description

ines commented Aug 2, 2019

lock bot commented Sep 20, 2019