Repository navigation
Lexer
The lexerfile defines the information required for lexical analysis and thus tokenization.
The simplest form of tokenization is based on the help of regex. The symbol groups and the corresponding patterns are mapped together and define the alphabet and the vocabulary of the grammar.
To go into the notation of the lexerinput, the operators are described below.
The usual form of a lexer definition rule is as follows:
[TOKENNAME] := "[REGEX]"
The tokenname should start on a new line. It is good practice to define this name capslock for clarity.
Nonterminal elements should be mapped with the keyword NONTERMINAL to tell the generator which elements are nonterminal. The mapped regex element can be combined individually or as usual with a regex pipe operator.
Symbols (such as the empty symbol) may appear in the grammar which have no influence on the parsing. Well-known examples of this are spaces or line breaks. These are called ignorables in Parssist.
To define these elements, you can write a lot with the respective symbols. The notation is as follows:
%"[SYMBOL1]", "[SYMBOL2]" [...]
Make sure that the quantities are each defined on a new line. These symbols are then automatically given the label "IGNORE".
In the lexer file there are only line comments that must begin with a hashtag. these are not taken into account during the lexical analysis.