v0.9.0
Recent changes:
- Implemented a script that fixes the DBNL FoLiA v2 documents as delivered in 2019. This scripts fixes the IDs and adds the necessary (sub)metadata. Discussed in Jira ticket: https://jira.socialhistoryservices.org/browse/TT-709
Older changes (2019):
- enable ignore option for wikiente
- implement support for language constrain in modernisation
- added resources (migrated from inl/nederlab-linguistic-enrichment)
- only do language identification on sentences!
- simplifying the pipeline, do not run frog in batches anymore but one frog per file (at cost of init time and extra memory, but easier to handle potential errors)
- replacing folialangid with colibri-lang, use --subcodes for colibri-lang
- do language detection before tokenization