The Unified Benchmark for Metaphor Identification (UBMI) gathers a wide range of metaphor and idiomaticity datasets into a single, harmonised format. The collected datasets were not all created for the automatic identification of metaphors: some come from psycholinguistic studies designed to collect human ratings or neurological measurements (e.g. CARD, JANK), some were originally created in other languages and later manually translated by their authors (e.g. WANG, BAMB), some are collections of metaphors obtained with concordancers (e.g. CCM), and others are full documents annotated to provide a standard annotation scheme (e.g. VUAC). Even within NLP, the datasets serve varied purposes — studying compositionality (GUT), evaluating metaphor generation (CHAK), analysing metaphoricity (DUNN), or disambiguating potentially idiomatic expressions.
Because of these differences, UBMI is not designed to merge all labelled phrases into one large dataset. What the collected datasets have in common is that they encode the metaphoric or idiomatic properties of expressions in context — or, conversely, their absence. UBMI re-encodes them with unified field names to facilitate experiments and comparisons across this diversity. In the open benchmark, only datasets with open licences are included.
HuggingFace repository : https://huggingface.co/datasets/Joanne/UBMI
Every example is stored with a unique tagged Potentially Metaphoric Expression (PME) per
sentence, together with the minimal information needed for a binary classification task
(the position of the expression in its context and its label), plus optional unified and
dataset-specific fields kept under additional_information.
An example UBMI entry from the CARD_N dataset (Cardillo et al., 2010). The core fields
(ref_id, dataset_id, id, context, expression/pme, position, label, split)
locate and label the PME, while additional_information preserves unified optional fields
(e.g. lemma, pos, mscore, tenor, tenor_position, set_id, five-fold splits) and
the original dataset-specific ratings.
| Field | Description |
|---|---|
| Core — dataset level | |
ref_id |
Name of the original dataset (first author's abbreviated name if no other name exists). Several UBMI datasets may come from the same reference dataset. |
dataset_id |
Unique id for a UBMI dataset, often composed of the reference id and a split (random, lexical, etc.). |
data_split |
Random, lexical, original split, etc. |
| Core — example level | |
id |
Entry id for a single example. |
context |
Text containing the PME (most often a sentence). |
expression |
The PME: the labelled words in the context. |
position |
Offset of the labelled expression within the context field. |
label |
Only metaphorical (m), literal (l) and other (o) labels appear here; finer labels are kept in original_label. |
split |
Whether the example belongs to the train, validation or test set of the main split. |
| Added for all examples | |
five_folds |
Train/validation/test membership across 5 folds (for cross-validation). |
lemma |
Lemma(s) of the PME, obtained with spaCy. |
pos |
A single PoS, or a list of PoS for PMEs longer than one word. |
| Common optional unified fields | |
long_context |
Additional context when the original dataset provides more than one sentence. |
mscore |
Metaphoricity or confidence score scaled to [0, 1]. |
tenor / tenor_lemma / tenor_pos / tenor_position |
The topic word(s) of the metaphor and their lemma, PoS and position. |
source_domain / target_domain |
Conceptual source-to-target domain mapping. |
original_label |
The original (usually more fine-grained) label when it differs from the binary label. |
set_id |
Present when the sentence is grouped with others (usually pairs or triples). |
| other | Non-unified fields are stored as original dataset information in additional_information. |
All datasets containing both literal and metaphorical labels are split into training (70%), development (10%) and test (20%) sets. Collections of metaphors without literal examples are not split. Splits come in two main sampling methods:
- Random splits — simple random shuffles of the data.
- Lexical splits — ensure a PME present in one split does not appear in the others, probing model generalisation.
Some datasets contain duplicated sentences (e.g. MIPVU datasets) or duplicated contexts where only the tagged expression changes (e.g. CHAK, GUT, TONG); for these, additional context-based splits keep the same context out of more than one split. Fixed 5-fold cross-validation splits are provided for both random and lexical sampling, and any original split shipped with a dataset is preserved.
The table below lists all datasets in the open benchmark, with the number of entries,
number/percentage of metaphors, number of distinct expressions and contexts, the context
span, and the sizes of the random-split train/dev/test sets.
The In Full? column indicates whether the dataset is already available in the Full/ directory (✅ Added) or will be added shortly (⏳ Soon).
| Dataset Name | In Full? | N Entry | N Met (%) | N dist. Expr. | N dist. Ctxt. | Ctxt. Span | Train | Dev | Test |
|---|---|---|---|---|---|---|---|---|---|
| Nominal PME | |||||||||
| CARDN | ✅ Added | 512 | 256 (50) | 256 | 511 | sentence | 358 | 51 | 103 |
| JANK | ✅ Added | 360 | 120 (33) | 120 | 360 | sentence | 252 | 36 | 72 |
| BAMB | ⏳ Soon | 115 | 115 (100) | 91 | 84 | noun | 0 | 0 | 115 |
| WANG | ⏳ Soon | 240 | 120 (50) | 120 | 240 | sentence | 128 | 24 | 48 |
| Fig-QA | ✅ Added | 8922 | 8922 (100) | 6845 | 4454 | sentence | 0 | 0 | 8922 |
| 2x2Meta | ⏳ Soon | 691 | 329 (48) | 547 | 203 | paragraph | 484 | 69 | 138 |
| Adjectival PME | |||||||||
| GUT | ✅ Added | 8591 | 4601 (54) | 23 | 3479 | noun | 6014 | 859 | 1718 |
| NEU | ✅ Added | 100 | 56 (56) | 5 | 93 | noun | 70 | 10 | 20 |
| TSVA | ✅ Added | 1945 | 979 (50) | 687 | 1072 | noun | 1362 | 195 | 388 |
| Verbal PME | |||||||||
| CARDV | ✅ Added | 280 | 140 (50) | 140 | 280 | sentence | 196 | 28 | 56 |
| CHAK | ✅ Added | 468 | 312 (67) | 355 | 155 | sentence | 328 | 47 | 93 |
| MOH | ✅ Added | 1632 | 407 (25) | 438 | 1630 | sentence | 1142 | 163 | 327 |
| DUNN | ✅ Added | 60 | 40 (67) | 20 | 60 | sentence | 0 | 0 | 100 |
| TSVV | ⏳ Soon | 222 | 111 (50) | 119 | 222 | sentence | 0 | 0 | 222 |
| TroFi | ✅ Added | 3642 | 2098 (58) | 50 | 3642 | sentence | 2549 | 364 | 729 |
| NewsMet | ⏳ Soon | 1205 | 594 (49) | 680 | 1205 | sentence | 844 | 120 | 241 |
| Various | |||||||||
| TONG | ⏳ Soon | 1428 | 655 (46) | 1031 | 739 | sentence | 1000 | 142 | 286 |
| CCM | ✅ Added | 8492 | 8492 (100) | 590 | 8490 | sentence | 0 | 0 | 8492 |
| GORD | ⏳ Soon | 1771 | 1771 (100) | 574 | 1771 | sentence | 0 | 0 | 1771 |
| ATT-META-S | ⏳ Soon | 500 | 500 (100) | 500 | 500 | sentence | 0 | 0 | 500 |
| PIEs | |||||||||
| VNC | ✅ Added | 2568 | 2017 (79) | 53 | 2534 | 3 sentences | 1798 | 257 | 513 |
| PVC | ✅ Added | 1348 | 878 (65) | 23 | 1348 | 3 sentences | 944 | 135 | 269 |
| SE2013ALL | ✅ Added | 1969 | 1172 (60) | 55 | 1939 | 3 sentences | 1378 | 197 | 394 |
| SE2013LEX | ✅ Added | 2371 | 1199 (51) | 10 | 2339 | 3 sentences | 1660 | 237 | 474 |
| PIE | ✅ Added | 3025 | 1434 (47) | 1608 | 2204 | 3 sentences | 2118 | 303 | 604 |
| MAD | ✅ Added | 4558 | 2190 (48) | 443 | 4554 | 3 sentences | 3191 | 456 | 911 |
| MAGPIE | ✅ Added | 48395 | 36328 (75) | 9462 | 47280 | 5 sentences | 33876 | 4839 | 9680 |
| PARSEME | ⏳ Soon | 1114 | 1114 (100) | 530 | 958 | sentence | 0 | 0 | 1114 |
| MIPVU | |||||||||
| VUACBO | ✅ Added | 39223 | 20350 (52) | 8722 | 39108 | 3 sentences | 27456 | 3922 | 7845 |
| VUACST1 | ✅ Added | 23113 | 6554 (28) | 2382 | 22894 | 3 sentences | 16179 | 2311 | 4623 |
| VUACST2 | ✅ Added | 94807 | 15026 (16) | 13509 | 93382 | 3 sentences | 66365 | 9481 | 18961 |
| NACEY | ⏳ Soon | 498 | 249 (50) | 244 | 498 | 3 sentences | 349 | 49 | 100 |
| JUL | ⏳ Soon | 4268 | 2134 (50) | 1649 | 4268 | 3 sentences | 2988 | 426 | 854 |
The datasets are grouped along two dimensions: the syntactic form of the labelled expression, and the granularity of the annotation. The first three groups are organised by syntactic form (nominal; adjectival and verbal; mixed), and the last two by the type of expression and annotation procedure (multi-word expressions; MIPVU). Each table below lists the UBMI datasets of that category with their original reference and original licence.
Datasets designed for the study of the metaphorical usage of nouns or noun phrases (NPs). The figurative examples are direct metaphors, where the tenor NP and the vehicle NP are explicitly related within the sentence (e.g. Man is a wolf, an ocean of happiness). Most were created for psycholinguistic or neurolinguistic studies; two come from NLP projects.
| Dataset | Reference | Original licence |
|---|---|---|
| CARDN | Cardillo et al. (2010); Cardillo et al. (2017) | CC BY-NC |
| JANK | Jankowiak (2020) | CC BY 4.0 |
| BAMB | Bambini et al. (2014) | CC BY 4.0 |
| WANG | Wang et al. — Chinese metaphor norms | Open |
| Fig-QA | Liu et al. (2022) | MIT License — copy, modification and redistribution allowed; must always include the original license: https://github.com/nightingal3/Fig-QA/blob/master/LICENSE |
| 2x2Meta | Boisson et al. (2025) | CC BY-NC |
Indirect metaphors — the source concept is suggested by the words used metaphorically but not explicitly named (e.g. tasty metaphor, pour money, economy flourishes). They are the most frequent type of metaphor in natural text, and the datasets were created in majority for NLP studies. Adjectival datasets annotate the literal/metaphorical use of an adjective in the context of a noun (often adjective-noun pairs); verbal datasets are collections of full sentences.
Adjectival PME
| Dataset | Reference | Original licence |
|---|---|---|
| GUT | Gutiérrez et al. (2016) | AFL-3.0 |
| NEU | Turney et al. (2011); Assaf et al. (2013) | CC BY-NC 4.0 |
| TSVA | Tsvetkov et al. (2014) | Redistribution allowed; license at https://github.com/ytsvetko/metaphor/blob/master/LICENSE.md |
Verbal PME
| Dataset | Reference | Original licence |
|---|---|---|
| CARDV | Cardillo et al. (2010); Cardillo et al. (2017) | CC BY-NC |
| CHAK | Chakrabarty et al. (2021) | No explicit licence; original data shared on GitHub: https://github.com/tuhinjubcse/MetaphorGenNAACL2021 |
| MOH | Mohammad et al. (2016) | Redistribution allowed; license at https://saifmohammad.com/WebPages/metaphor.html |
| DUNN | Dunn (2014) | CC BY-SA 3.0 |
| TSVV | Tsvetkov et al. (2014) | Redistribution allowed; license at https://github.com/ytsvetko/metaphor/blob/master/LICENSE.md |
| TroFi | Birke & Sarkar (2006) | CC BY-NC 4.0 |
| NewsMet | Joseph et al. (2023) | Apache-2.0 |
Datasets labelled with various types of metaphoric expressions, with (mostly) no syntactic constraint on the context or the PME, and expressions that may span more than one word. They include collections of metaphors focused on mapping analysis and datasets built for metaphor identification. (In the UBMI table these appear under the Various group.)
| Dataset | Reference | Original licence |
|---|---|---|
| TONG | Tong et al. (2024) | CC BY 4.0 |
| CCM | MacWhinney & Fromm (2014); Levin et al. (2014) | CC BY-NC 4.0 |
| GORD | Gordon et al. (2015) | Redistribution allowed with required attribution. Any product, report, publication, presentation or document including or referencing the data must contain the text: "This effort contains or makes use of the IARPA-funded Metaphor Program USC/ISI annotated metaphorical language collection, release iarpa_metaphor_isi.edu_metaphor_corpus_20150403" |
| ATT-META-S | ATT-Meta Databank — Barnden et al. | Non-commercial research/instructional use; most pages of the databank may be freely copied for that purpose: https://www.cs.nmsu.edu/atmet/Databank/root.html.OLD |
Datasets where all labelled expressions are multi-word expressions. Most are Potentially Idiomatic Expressions (PIEs) — the same surface form occurs both literally and idiomatically — compiled and annotated for binary classification. Some target MWEs with specific patterns (noun compounds, verb-preposition compounds); others impose no restriction on the compound form. PARSEME 1.3 is an exception: it was created for verbal-MWE identification (negative cases are not annotated by default). (In the UBMI table these appear under the PIEs group.)
| Dataset | Reference | Original licence |
|---|---|---|
| VNC | Cook et al. (2008) | No specific licence stated (no entry in the licence appendix) |
| PVC | Tu & Roth (2012) | No explicit licence; original dataset: https://cogcomp.seas.upenn.edu/page/resource_view/26 |
| SE2013ALL | Korkontzelos et al. (2013) | CC BY-SA 4.0; redistribution authorised directly by the authors. Original download: https://www.inf.uni-hamburg.de/en/inst/ab/lt/resources/data.html |
| SE2013LEX | Korkontzelos et al. (2013) | CC BY-SA 4.0; redistribution authorised directly by the authors. Original download: https://www.inf.uni-hamburg.de/en/inst/ab/lt/resources/data.html |
| PIE | Haagsma et al. (2019) | CC BY 4.0 |
| MAD | Tayyar Madabushi et al. (2021, 2022) | GPL-3.0 |
| MAGPIE | Haagsma et al. (2020) | CC BY 4.0 |
| PARSEME | Savary et al. (2023) | CC BY-SA 4.0 / CC BY-SA 3.0 / CC BY 4.0 |
Datasets annotated with the MIPVU procedure (and its extensions): richly annotated documents in which each lexical unit is labelled metaphoric or literal. The benchmark retains three open-source MIPVU datasets (VUAC, NACEY, JUL); the VUAC is additionally sampled and reformatted for binary classification — balanced multi-word sampling (VUACBO) and the shared-task samplings (VUACST1, VUACST2). The licence shown for the VUAC samplings is that of the underlying VUAC corpus.
| Dataset | Reference | Original licence |
|---|---|---|
| VUACBO | Boisson et al. (2023) sampling of Steen et al. (2010) — VUAC | CC BY-SA 3.0 |
| VUACST1 | Leong et al. (2020) sampling of Steen et al. (2010) — VUAC | CC BY-SA 3.0 |
| VUACST2 | Leong et al. (2020) sampling of Steen et al. (2010) — VUAC | CC BY-SA 3.0 |
| NACEY | Nacey (2019) | CC0 1.0 |
| JUL | Julich (2022) | CC BY-SA 4.0 |
