You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Improving STIX indicator parsing and validation without making feeds brittle
#940
STIX indicator parsing: handling edge cases without making detection logic more complex
Hey everyone,
While working on #939, I ran into an interesting edge case in MVT's STIX indicator parsing.
The immediate issue is an indicator value containing =. For example:
https://example.com/track?id=1
Splitting the pattern on every = causes the value to be truncated, even though the URL itself is completely valid. I've addressed that specific case in #939 by limiting the split to the first =.
While looking at this, I started wondering whether the parsing boundary itself is worth discussing.
MVT doesn't need to become a full STIX parser, but indicator feeds can contain values with characters that have meaning in the STIX syntax itself. Some examples include:
= inside URLs
quotes or apostrophes inside values
additional whitespace
malformed or partially malformed patterns
properties or indicator forms that MVT doesn't support
This raises a broader question:
Should MVT keep the current lightweight parsing approach and handle edge cases incrementally, or would it be useful to introduce a small explicit parsing/validation boundary for the STIX patterns that MVT supports?
I'm thinking of something deliberately small:
Parse the expected indicator structure explicitly.
Preserve the complete value rather than relying on unrestricted string splitting.
Validate the parts that MVT actually supports.
Skip or log malformed individual indicators where possible instead of making the entire collection fail.
Leave the existing matching and detection semantics unchanged.
The goal would be input robustness rather than expanding MVT's STIX implementation.
I'd be interested in hearing how maintainers currently think about this boundary, especially whether there are existing feeds or indicator formats that would make a more explicit parser undesirable.
The concrete = regression is covered by #939; I'm opening this discussion separately because the broader parsing question seems better suited for design discussion than for the bug fix itself.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
STIX indicator parsing: handling edge cases without making detection logic more complex
Hey everyone,
While working on #939, I ran into an interesting edge case in MVT's STIX indicator parsing.
The immediate issue is an indicator value containing
=. For example:Splitting the pattern on every
=causes the value to be truncated, even though the URL itself is completely valid. I've addressed that specific case in #939 by limiting the split to the first=.While looking at this, I started wondering whether the parsing boundary itself is worth discussing.
MVT doesn't need to become a full STIX parser, but indicator feeds can contain values with characters that have meaning in the STIX syntax itself. Some examples include:
=inside URLsThis raises a broader question:
Should MVT keep the current lightweight parsing approach and handle edge cases incrementally, or would it be useful to introduce a small explicit parsing/validation boundary for the STIX patterns that MVT supports?
I'm thinking of something deliberately small:
The goal would be input robustness rather than expanding MVT's STIX implementation.
I'd be interested in hearing how maintainers currently think about this boundary, especially whether there are existing feeds or indicator formats that would make a more explicit parser undesirable.
The concrete
=regression is covered by #939; I'm opening this discussion separately because the broader parsing question seems better suited for design discussion than for the bug fix itself.All reactions