The kaikki.org website has some links to "postprocessed" wiktextract data, and in the near future links to these files will be removed because they're causing problems when users mistakenly download them instead of the raw wiktextract dump data.
Some of these issues are related to the fact that the data is used specifically in the generation of the kaikki.org website, which means modifying it by frex. sorting words alphabetically instead of in the order of generation (which is generally the same as the order they have on the wiktionary page they were extracted from). Other issues are old projects by Tatu trying to create identifier numbers, and trying to create a system to disambiguate words from certain senses or meanings; these things will still remain in the output of the kaikki.org website, but they are unlikely to progress any further and they're probably not useful to users wanting to do their own processing of the data.
The raw extract file is the more complete of the datasets. Most of the things done to the postprocessed data is specific to generating the kaikki website, like the sorting, and some of it will lose information in the process.
The kaikki.org website has some links to "postprocessed" wiktextract data, and in the near future links to these files will be removed because they're causing problems when users mistakenly download them instead of the raw wiktextract dump data.
Some of these issues are related to the fact that the data is used specifically in the generation of the kaikki.org website, which means modifying it by frex. sorting words alphabetically instead of in the order of generation (which is generally the same as the order they have on the wiktionary page they were extracted from). Other issues are old projects by Tatu trying to create identifier numbers, and trying to create a system to disambiguate words from certain senses or meanings; these things will still remain in the output of the kaikki.org website, but they are unlikely to progress any further and they're probably not useful to users wanting to do their own processing of the data.
language_codeentry to get the entries that would appear in the old file.The raw extract file is the more complete of the datasets. Most of the things done to the postprocessed data is specific to generating the kaikki website, like the sorting, and some of it will lose information in the process.