Skip to content

Corpus is 書面語, so colloquial (口語) tone behaviour is guessed, not measured #2

Description

@candpixie

Mainstream Cantopop lyrics are written largely in 書面語. The consequence shows
up immediately: of 103 unique characters in a colloquial test lyric, the corpus
could resolve 95. The eight it could not were

冇 刪 咗 哋 唔 嘅 嘢 為

which is almost exactly the set of high-frequency 口語 characters.

Right now these are covered by cantojam/data/colloquial.json,
a hand-written supplement. That gives them a reading, so lookups work, but it
gives them no measured tone-melody behaviour, because they appear in too
few corpus songs to contribute statistics.

Two ways to help

Transcribe colloquial songs (the real fix). Any song written in 口語 rather
than 書面語 puts these characters into the model with actual pitch data behind
them. See #1.

Audit the supplement (quick, still useful). The file is hand-authored and
unreviewed by a second pair of eyes. Wrong readings are plausible. One line per
character, jyutping with tone digit, most common reading first, both listed if a
character genuinely has two live readings.

Metadata

Metadata

Assignees

No one assigned

    Labels

    corpusTranscription work, the highest-value contributionhelp wantedExtra attention is needed

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions