-
Notifications
You must be signed in to change notification settings - Fork 4
Data format
This document describes some points of divergence between PROIEL XML and the internal data format used by proiel-webapp.
Precedence refers to the linear order of objects in the underlying text.
Precedence is encoded on objects using an integer index. If an index i is less than j, the object with index i precedes the object with index j. The index is not guaranteed to be zero-based nor does it have to be contiguous. The following, for example, is a valid representation of precedence in a sentence with three tokens:
| Token form | Token precedence index |
|---|---|
| Marcus | 5 |
| puellam | 100 |
| amat | 12345678 |
For objects that do not correspond to anything in the underlying text (e.g. pro subject tokens), the index is undefined. Undefined in this context means that the index may have any integer value or a null value; in either case it carries no meaning and must be ignored. The following three representations, for example, are equivalent. Although the index corresponding to the pro subject has different values, the pro subject in each representation has no position in the linear order in the underlying text.
| Token form | Token precedence index |
|---|---|
| pro | 0 |
| puellam | 100 |
| amat | 12345678 |
| Token form | Token precedence index |
|---|---|
| pro | 1000 |
| puellam | 100 |
| amat | 12345678 |
| Token form | Token precedence index |
|---|---|
| pro | null |
| puellam | 100 |
| amat | 12345678 |
In the following, an attribute that may be NULL can have the value NULL in the SQL database, have an empty value in the XML representation or be absent altogether from the XML representation.
All tokens must have all the following attributes, which may never be NULL:
Attribute | Description
--------- | -----------
`id` | A unique ID for the token
`sentence_id` | The sentence that the token belongs to
`token_number` | The position of the token in the linear string
The pair (sentence_id, token_number) must be unique.
Any tokens can have any of the following attributes, which may be NULL:
Attribute | Description
--------- | -----------
`citation_part` |
`information_status_tag` |
`antecedent_id` |
`contrast_group` |
`token_alignment_id` |
`automatic_token_alignment` |
`dependency_alignment_id` |
`source_morphology_tag` | (deprecated)
`source_lemma` | (deprecated)
`created_at` |
`updated_at` |
`lemma_id` |
`morphology_tag` |
`head_id` |
`relation_tag` |
If form is NULL, empty_token_sort may not be NULL, while presentation_before, presentation_after and foreign_ids must be NULL.
If form is not NULL, empty_token_sortmust be NULL.
The part_of_speech_tag is a two-character positional tag whose first
character is the major part of speech and second character is the minor
part of speech. Valid part-of-speech tags are defined in the configuration file
config/tagsets/parts_of_speech.yml.
The morphology_tag is a positional tag with 10 characters. Valid morphology
tags are defined in the configuration file config/tagsets/morphology.yml.
The inflection field indicates whether the annotated token is morphologically
inflected or uninflected. The field is part of the morphology of each token
rather than associated with the lemma assigned to the token to allow for
sporadic lack of inflection of lemmata that otherwise usually appear with
inflection.
Presentation text is textual material that should be presented when a token, sentence, source division or source is rendered but should remain unavailable for annotation. Typical examples include punctuation, verse numbering and section headings.
Presentation text is stored in the columns presentation_before and
presentation_after in the tables tokens, sentences and
source_divisions.
The columns are intended for different purposes. The columns in the
source_divisions table can contain any amount of text. They are therefore
suitable for introductory material like prologues or chapter headings.
The columns in the sentences and tokens tables are intended for short
strings that are exempted from annotation. These columns are therefore in
practice restricted to strings of a certain length (see the database schema for
the exact value).
The columns in the sentences table should only be used for text that
unambiguously indicates a fixed break between sentences. An example would be
stage directions in drama. The columns in the tokens table, on the other
hand, can contain any presentation text that intervenes in the running text.
This is typically punctuation and inter-word spacing (including line breaks in
poetry or drama).
Because of their different purposes, the columns have different semantics with
respect to the merging of sentences. Two sentences can be merged when there is
presentation text intervening between them but only when this is represented in
the relevant tokens columns. If there is material in any of the sentences
columns, it is interpreted as a fixed sentence boundary, which prevents merging
of the two sentences.
Sentences from different source divisions cannot be merged for independent
reasons so any presentation text in the source_divisions table is irrelevant
when merging sentences.
When a token is split, the original value of foreign_ids is parenthesised and _1 is appended to the result and used as the new value for the first token. Similarly, _2 is appended and used as the new value for the second token. Example:
Before split After split
------------ -----------
token.foreign_ids = "note=foobar" token1.foreign_ids = "(note=foobar)_1"
token.token_number = 10 token1.token_number = 10
token2.foreign_ids = "(note=foobar)_2"
token2.token_number = 11
Repeated splits will add further parentheses and suffixes:
note=foobar
(note=foobar)_1
((note=foobar)_1)_1
When two tokens are merged, the original values of foreign_ids are parenthesised and concatenated with . as separator. Example:
Before merge After merge
------------ -----------
token1.foreign_ids = "note=foobar" token.foreign_ids = "(note=foobar).(note=barfoo)"
token1.token_number = 10 token.token_number = 10
token2.foreign_ids = "note=barfoo"
token2.token_number = 11
Repeated mergers will add further parentheses and concatenation symbols:
note=foobar
(note=foobar).(note=barfoo)
((note=foobar).(note=barfoo)).(note=foobaz)
For all alignments the directionality of the alignment relations is from the secondary source, i.e. the assumed translation, to the primary source, i.e. the assumed original.
| Column | Description |
|---|---|
source_divisions.aligned_source_division_id |
The source division this source division should be aligned to. This should be set on secondary sources only. The relation is used to determine the scope of automatic sentence alignment. |
sentences.unalignable |
If true, the sentence is, for the purposes of sentence alignment, not to be considered as an independent unit, but rather as part of the previous sentence in the linear ordering of sentences. This is, in other words, an indication that the sentence has been 'black- listed' from sentence alignment. |
sentences.automatic_alignment |
If true, the sentence alignment indicated by sentences.sentence_alignment_id has been generated automatically and is therefore more likely to be wrong and thus more likely to be a candidate for deletion should alignment need to be adjusted at a later stage. The flag is set when automatic sentence alignment is committed and unset when it is uncommitted. |
sentences.sentence_alignment_id |
The sentence this sentence is aligned with. This sentence alignment has been provided manually unless sentences.automatic_alignment is set. |
| Column | Description |
|---|---|
tokens.token_alignment_id |
The token this token is aligned with for the purposes of token alignment. This token alignment has been provided manually unless tokens.automatic_token_alignment is set. |
tokens.automatic_token_alignment |
If true, the token alignment indicated by tokens.token_alignment_id has been generated automatically. |
| Column | Description |
|---|---|
tokens.dependency_alignment_id |
The token that is the head of the dependency subgraph this token is aligned with for the purposes of dependency alignment. |
dependency_alignment_terminations.token_id |
The token whose dependency subgraph has a termination, i.e. which is not part of its heads dependency subgraph alignment. |
dependency_alignment_terminations.source_id |
The target source for the termination in dependency_alignment_terminations.token_id. |