Skip to content

Data format

Marius Jøhndal edited this page Jun 6, 2016 · 11 revisions

This document describes some points of divergence between PROIEL XML and the internal data format used by proiel-webapp.

Hierachical organisation

Precedence refers to the linear order of objects in the underlying text.

Precedence is encoded on objects using an integer index. If an index i is less than j, the object with index i precedes the object with index j. The index is not guaranteed to be zero-based nor does it have to be contiguous. The following, for example, is a valid representation of precedence in a sentence with three tokens:

Token form Token precedence index
Marcus 5
puellam 100
amat 12345678

For objects that do not correspond to anything in the underlying text (e.g. pro subject tokens), the index is undefined. Undefined in this context means that the index may have any integer value or a null value; in either case it carries no meaning and must be ignored. The following three representations, for example, are equivalent. Although the index corresponding to the pro subject has different values, the pro subject in each representation has no position in the linear order in the underlying text.

Token form Token precedence index
pro 0
puellam 100
amat 12345678
Token form Token precedence index
pro 1000
puellam 100
amat 12345678
Token form Token precedence index
pro null
puellam 100
amat 12345678

The core attributes for tokens

In the following, an attribute that may be NULL can have the value NULL in the SQL database, have an empty value in the XML representation or be absent altogether from the XML representation.

All tokens must have all the following attributes, which may never be NULL:

Attribute                     | Description
---------                     | -----------
`id`                          | A unique ID for the token
`sentence_id`                 | The sentence that the token belongs to
`token_number`                | The position of the token in the linear string

The pair (sentence_id, token_number) must be unique.

Any tokens can have any of the following attributes, which may be NULL:

Attribute                       | Description
---------                       | -----------
`citation_part`                 |
`information_status_tag`        |
`antecedent_id`                 |
`contrast_group`                |
`token_alignment_id`            |
`automatic_token_alignment`     |
`dependency_alignment_id`       |
`source_morphology_tag`         | (deprecated)
`source_lemma`                  | (deprecated)
`created_at`                    |
`updated_at`                    |
`lemma_id`                      |
`morphology_tag`                |
`head_id`                       |
`relation_tag`                  |

If form is NULL, empty_token_sort may not be NULL, while presentation_before, presentation_after and foreign_ids must be NULL.

If form is not NULL, empty_token_sortmust be NULL.

Part of speech and morphology

The part_of_speech_tag is a two-character positional tag whose first character is the major part of speech and second character is the minor part of speech. Valid part-of-speech tags are defined in the configuration file config/tagsets/parts_of_speech.yml.

The morphology_tag is a positional tag with 10 characters. Valid morphology tags are defined in the configuration file config/tagsets/morphology.yml.

The inflection field

The inflection field indicates whether the annotated token is morphologically inflected or uninflected. The field is part of the morphology of each token rather than associated with the lemma assigned to the token to allow for sporadic lack of inflection of lemmata that otherwise usually appear with inflection.

Presentation text

Presentation text is textual material that should be presented when a token, sentence, source division or source is rendered but should remain unavailable for annotation. Typical examples include punctuation, verse numbering and section headings.

Presentation text is stored in the columns presentation_before and presentation_after in the tables tokens, sentences and source_divisions.

The columns are intended for different purposes. The columns in the source_divisions table can contain any amount of text. They are therefore suitable for introductory material like prologues or chapter headings.

The columns in the sentences and tokens tables are intended for short strings that are exempted from annotation. These columns are therefore in practice restricted to strings of a certain length (see the database schema for the exact value).

The columns in the sentences table should only be used for text that unambiguously indicates a fixed break between sentences. An example would be stage directions in drama. The columns in the tokens table, on the other hand, can contain any presentation text that intervenes in the running text. This is typically punctuation and inter-word spacing (including line breaks in poetry or drama).

Because of their different purposes, the columns have different semantics with respect to the merging of sentences. Two sentences can be merged when there is presentation text intervening between them but only when this is represented in the relevant tokens columns. If there is material in any of the sentences columns, it is interpreted as a fixed sentence boundary, which prevents merging of the two sentences.

Sentences from different source divisions cannot be merged for independent reasons so any presentation text in the source_divisions table is irrelevant when merging sentences.

The effect on foreign_ids of splitting and merging tokens

When a token is split, the original value of foreign_ids is parenthesised and _1 is appended to the result and used as the new value for the first token. Similarly, _2 is appended and used as the new value for the second token. Example:

Before split                        After split
------------                        -----------
token.foreign_ids = "note=foobar"   token1.foreign_ids = "(note=foobar)_1"
token.token_number = 10             token1.token_number = 10
                                    token2.foreign_ids = "(note=foobar)_2"
                                    token2.token_number = 11

Repeated splits will add further parentheses and suffixes:

note=foobar
(note=foobar)_1
((note=foobar)_1)_1

When two tokens are merged, the original values of foreign_ids are parenthesised and concatenated with . as separator. Example:

Before merge                        After merge
------------                        -----------
token1.foreign_ids = "note=foobar"  token.foreign_ids = "(note=foobar).(note=barfoo)"
token1.token_number = 10            token.token_number = 10
token2.foreign_ids = "note=barfoo"
token2.token_number = 11

Repeated mergers will add further parentheses and concatenation symbols:

note=foobar
(note=foobar).(note=barfoo)
((note=foobar).(note=barfoo)).(note=foobaz)

Alignment data

For all alignments the directionality of the alignment relations is from the secondary source, i.e. the assumed translation, to the primary source, i.e. the assumed original.

Sentence alignment

Column Description
source_divisions.aligned_source_division_id The source division this source division should be aligned to. This should be set on secondary sources only. The relation is used to determine the scope of automatic sentence alignment.
sentences.unalignable If true, the sentence is, for the purposes of sentence alignment, not to be considered as an independent unit, but rather as part of the previous sentence in the linear ordering of sentences. This is, in other words, an indication that the sentence has been 'black- listed' from sentence alignment.
sentences.automatic_alignment If true, the sentence alignment indicated by sentences.sentence_alignment_id has been generated automatically and is therefore more likely to be wrong and thus more likely to be a candidate for deletion should alignment need to be adjusted at a later stage. The flag is set when automatic sentence alignment is committed and unset when it is uncommitted.
sentences.sentence_alignment_id The sentence this sentence is aligned with. This sentence alignment has been provided manually unless sentences.automatic_alignment is set.

Token alignment

Column Description
tokens.token_alignment_id The token this token is aligned with for the purposes of token alignment. This token alignment has been provided manually unless tokens.automatic_token_alignment is set.
tokens.automatic_token_alignment If true, the token alignment indicated by tokens.token_alignment_id has been generated automatically.

Dependency alignment

Column Description
tokens.dependency_alignment_id The token that is the head of the dependency subgraph this token is aligned with for the purposes of dependency alignment.
dependency_alignment_terminations.token_id The token whose dependency subgraph has a termination, i.e. which is not part of its heads dependency subgraph alignment.
dependency_alignment_terminations.source_id The target source for the termination in dependency_alignment_terminations.token_id.

Clone this wiki locally