This record provides the reference for the Modeling Languages with Their Own Parameters: A Response to subs2vec manuscript. The project investigated the use of large-scale film and television subtitle corpora to train distributional word embeddings across multiple languages.
The word2vec approach was designed to capture semantic similarity from naturalistic language use in subtitles, and to enable cross-linguistic comparisons of embedding spaces. We systematically explored a range of hyperparameters—including vector dimensionality and context window size—and reported their impact across 45+ languages.
Key contributions of the paper include:
• Introducing subtitle-based corpora as a multilingual resource for word embeddings.
• Providing embeddings for dozens of languages under varied training configurations.
• Critically assessing the limitations of the original subs2vec paper, and outlining how future work might better adapt embeddings to each language’s properties.
This release serves as a canonical reference for the manuscript, complementing the per-language embedding datasets archived on Zenodo. Together, the paper and datasets provide a reproducible record of the project and a resource for further computational and cross-linguistic research.
Citation