비교적 대부분의 사람들이 접근할 수 있는 오픈 데이터를 정리하였다. 구할 수 있는 모든 데이터를 쏟아 부어서 end to end로 모델을 만들어 보겠다는 포부를 가진 분들의 진입을 쉽게하기 위한 목적이고, 정교한 데이터 구축을 위해서는 이후에 어떠한 데이터가 필요한지를 살펴보기 위한 과정이다.
12월 15일 첫번째 버전을 만들었고 이후 박조은님의 코멘트 그리고 2020년 8월 21일 @warnikchow님의 다양한 기여와 의견을 반영하여 수정하였고 2020년 10월 18일 메인 레포 를 이동하였다. 메인 레포와 함께 데이터 링크와 약간의 설명을 보강하여 운영 중이다.
Natural language processing의 각 분야에 대한 자료 정리는 다음 링크를 참고 Awesome-Korean-NLP
다양한 전처리 및 다운로더를 포함한 데이터 링크는 다음을 참조 https://ratsgo.github.io/embedding/preprocess.html
코퍼스 패키지에 많은 관심이 필요합니다! 웹에 공개되어 있는 한국어 텍스트 데이터들을 손쉽게 로딩하고, 이를 이용하여 모델링한 후 evaluation 까지 편하게 수행하는 / 즉 한국어 텍스트 데이터를 위한 huggingface.nlp 작업 중인 페이지는 다음을 참조, ko-nlp
24년 10월 9일 한글날 기준 허깅페이스에서 2번 이상 다운받은 데이터를 다음 링크 참조 https://github.com/songys/huggingface_KoreanDataset
Commercially available(com), academic use only(Academia), unknown(unk)
Redistribution is possible with or without modification, if neither, or unknown (rd, rd/mod-x, no, unk)
Internationally available publication(inter)
아래 목록은 Open-korean-corpora 의 2026년 6월 업데이트 기준을 따라 나누었다. 각 표의 열은 Dataset, Typical Usage, Provider, Docu., License, Redist, mod-x, Volume 순서이다.
2026년 6월 업데이트에서 새로 추가한 데이터셋 (2020–2025)
이번 업데이트에서는 2020년부터 2025년까지 공개된 데이터셋을 분야별로 추가하였다.
Benchmark studies: Open-Ko-LLM, HAE-RAE Bench, KMMLU, KULTURE Bench, KMMLU-Redux/Pro, KoBALT
Entailment, similarity, paraphrase: KoSEnd
Intention and sentiment: KMRE, ToM-Diary, KoCoSa, KOTE, KPoEM
Offensive language, fairness and bias (추가된 분야): KoMultiText, K-HATERS, KoSBi, SQuARe, KoBBQ, KCDD, LifeTox
QA and dialogue: KorWikiTQ, CLIcK, KoDialogBench, KorNAT, K-MMBench, K-Viscuit, KoSimpleQA, KoPIQA
Summarization, translation, transliteration: OPUS-MT ko-en, Naver News Summary, KoreaScience Summary, SSL, KPC, KNOTICED
Korean in multilingual corpora: XL-Sum, MASSIVE
Speech corpora: OLKAVS, KMSAV
Other topics: KorMedMCQA, KBMC, ESG-Kor, KBL, FunctionChat-Bench, KCL, KorMedLawQA
연도별 수록 데이터셋 수는 다음과 같다 (총 96개).
연도
2015
2016
2018
2019
2020
2021
2022
2023
2024
2025
데이터셋 수
1
3
8
6
11
5
21
9
23
9
Dataset
Typical Usage
Provider
Docu.
License
Redist
mod-x
Volume
KLUE
Benchmark studies
Industry
int'l
all
rd
0
DP - 15K (s), DST - 10K (d), MRC - 29K (p), NER - 31K (s), NLI - 31K (p), RE - 48K (s), STS - 12K (p), TC - 63K (s)
KoBEST
Benchmark studies
Industry
int'l
all
rd
0
BoolQ - 5.8K (p), COPA - 4.6K, KB-WiC - 5.2K (p), KB-HellaSwag - 3K, SentiNeg - 4.8K
Ko-H5/Open-Ko-LLM
LLM benchmark
Industry
int'l
all
rd
0
Ko-ARC, Ko-HellaSwag, Ko-MMLU, Ko-TruthfulQA, Ko-CommonGen v2; Season 2 tasks
HAE-RAE Bench
LLM benchmark
Academia
int'l
all
rd
0
1.5K questions, 6 tasks
KMMLU
LLM benchmark
Academia
int'l
all
rd
0
35K MCQs, 45 subjects
KULTURE Bench
Cultural benchmark
Academia
int'l
academic
unk
0
-
KMMLU-Redux / KMMLU-Pro
LLM benchmark
Industry
int'l
all
rd
1
2.6K / 2.8K problems
KoBALT
Linguistic benchmark
Academia
int'l
academic
rd
1
700 MCQs
Dataset
Typical Usage
Provider
Docu.
License
Redist
mod-x
Volume
UD Korean KAIST
Dependency parsing
Academia
int'l
academic
none
0
27K (s)
PKT-UD
Dependency parsing
Academia
int'l
academic
none
0
5K (s)
AIR x NAVER NER/SRL
NER, SRL
Competition
dom.
academic
none
0
NER - 90K (s), SRL - 35K (s)
KMOU NER
NER
Academia
dom.
academic
rd
0
24K (s)
OpenKorPOS
POS tagging
Academia
int'l
all
rd
0
55M (w)
KoNEC & KoNNEC
NER
Academia
dom.
all
rd
0
26K (s)
3. Entailment, sentence similarity, and paraphrase
Dataset
Typical Usage
Provider
Docu.
License
Redist
mod-x
Volume
Question Pair
Paraphrase detection
Academia
dom.
all
rd
0
10K (p)
KorNLI/KorSTS
NLI, STS
Industry
int'l
all
rd
0
KorNLI - 940K train (p), KorSTS - 5.7K train (p)
ParaKQC
Paraphrase detection
Academia
int'l
all
rd
0
540K (p)
StyleKQC
Style transfer, paraphrase detection
Academia
int'l
all
rd
0
30K (s)
Korean Smile Style Dataset
Style transfer
Industry
dom.
academic
rd
0
2.5K (d)
KoSEnd
Linguistic evaluation
Academia
int'l
academic
unk
0
-
4. Intention understanding and sentiment analysis
Dataset
Typical Usage
Provider
Docu.
License
Redist
mod-x
Volume
NSMC
Sentiment analysis
Academia
int'l
all
rd
0
200K (s)
3i4K
Speech act classification
Academia
int'l
all
rd
0
61K (s)
Kocasm
Sarcasm detection
Academia
dom.
all
rd
0
9K (s)
KMRE
Emotion classification
Academia
int'l
all
rd
0
-
ToM-Diary
Theory of mind analysis
Academia
dom.
academic
rd
1
18K diaries, 74K (s)
KoCoSa
Sarcasm detection
Academia
int'l
all
rd
0
12.8K (d)
KOTE
Emotion classification
Academia
int'l
academic
rd
1
50K comments, 250K annotations
KPoEM
Emotion detection
Academia
int'l
all
rd
0
-
5. Offensive language detection, fairness and bias
Dataset
Typical Usage
Provider
Docu.
License
Redist
mod-x
Volume
BEEP!
Hate speech detection
Academia
int'l
all
rd
0
9.4K (s)
APEACH
Hate speech detection
Academia
int'l
all
rd
0
4K (s)
Korean Unsmile Dataset
Hate speech detection
Industry
dom.
academic
rd
1
19K (s)
HateScore
Hate speech detection
Academia
int'l
academic
rd
0
35K (s)
KOLD
Offensive language detection
Academia
int'l
all
rd
0
40K (s)
K-MHaS
Hate speech detection
Academia
int'l
all
rd
0
109K (s)
DKTC
Threatening conversation detection
Industry
dom.
academic
rd
0
4.5K (d)
KODOLI
Offensive language detection
Academia
int'l
all
rd
0
38K (s)
KoMultiText
Bias and profanity detection
Academia
int'l
all
rd
0
150K comments
K-HATERS
Offensive language detection
Academia
int'l
all
rd
0
192K comments
KoSBi
Social bias detection
Industry
int'l
all
rd
0
34K (p)
SQuARe
Safe response generation
Industry
int'l
all
rd
0
49K questions, 88K responses
KoBBQ
Bias benchmark
Industry
int'l
all
rd
0
76K samples
KCDD
Violence classification
Academia
int'l
academic
none
0
22K (d)
LifeTox
Implicit toxicity detection
Academia
int'l
academic
unk
0
-
Dataset
Typical Usage
Provider
Docu.
License
Redist
mod-x
Volume
KorQuAD 1.0, 2.0
QA
Industry
int'l
all
rd
1
70K / 100K questions
KorWikiTQ
Table QA
Industry
int'l
all
rd
0
-
HuLiC
Dialog
Industry
dom.
academic
rd
0
115K turns
OPELA
Dialog
Industry
int'l
academic
rd
0
600 (d)
CareCall
Dialog
Industry
int'l
academic
rd
0
10K (d)
CLIcK
Cultural QA
Academia
int'l
all
rd
0
2K QA pairs
KoDialogBench
Dialogue benchmark
Academia
int'l
all
rd
0
83K examples
KorNAT
Social value QA
Industry
int'l
academic
rd
0
10K questions
K-MMBench
Vision-language QA
Industry
int'l
academic
rd
1
4.3K questions
K-Viscuit
VQA
Academia
int'l
academic
unk
0
-
KoSimpleQA
Factual QA
Academia
int'l
academic
unk
0
1K questions
KoPIQA
Physical commonsense QA
Academia
int'l
academic
unk
0
441 QA pairs
7. Summarization, Translation, and Transliteration
Dataset
Typical Usage
Provider
Docu.
License
Redist
mod-x
Volume
Sci-news-sum-kr
Summarization
Academia
dom.
academic
rd
0
50 (p)
Korean Parallel Corpora
MT
Academia
int'l
academic
rd
1
100K (p)
Transliteration Dataset
Transliteration
Academia
dom.
all
rd
0
35K (p)
sae4K
Summarization
Academia
int'l
all
rd
0
50K (p)
OPUS-MT ko-en
MT
Academia
int'l
all
rd
0
Tatoeba Challenge test sets
Naver News Summary
Summarization
Academia
none
all
rd
0
-
KoreaScience Summary
Summarization
Academia
int'l
all
rd
0
-
SSL
Sign language translation
Academia
int'l
all
unk
0
-
KPC
MT
Academia
int'l
academic
rd
1
131K (p)
KNOTICED
MT error detection
Academia
int'l
academic
unk
0
-
8. Korean in multilingual corpora
Dataset
Typical Usage
Provider
Docu.
License
Redist
mod-x
Volume
PAWS-X
Paraphrase detection
Industry
int'l
all
rd
0
5K / 2K / 2K (p)
SIGMORPHON G2P
G2P conversion
Competition
int'l
all
rd
0
3.6K / 0.45K / 0.45K (p)
TyDi-QA
QA
Industry
int'l
all
rd
0
11K / 1.7K / 1.7K (p)
XPersona
Dialog
Academia
int'l
all
rd
0
0.3K (d) / 4.7K (s)
XL-Sum
Summarization
Academia
int'l
all
rd
0
-
MultiCoNER
NER
Competition
int'l
all
rd
0
178K / 2.6K (s)
MINT
Sentiment analysis
Competition
int'l
unk
unk
0
2K train/test, 0.5K zero-shot
MASSIVE
NLU
Industry
int'l
all
rd
0
1M examples
IWSLT 2023
MT
Competition
int'l
all
rd
0
3K (p)
Dataset
Typical Usage
Provider
Docu.
License
Redist
mod-x
Volume
KSS
ASR, TTS
Academia
int'l
academic
rd
0
12K (u), 1 speaker
Zeroth
ASR
Industry
int'l
all
rd
0
50+ (h)
Pansori-TEDxKR
ASR
Academia
int'l
academic
rd
1
3+ (h)
ProSem
SLU
Academia
int'l
all
rd
0
7.1K (u), 2 speakers
ClovaCall
ASR
Industry
int'l
academic
none
0
80+ (h)
JIT/JSS
ASR, TTS
Industry
int'l
all
rd
0
10K (JSS), 170K (JIT)
kosp2e
Speech translation
Academia
int'l
academic
rd
0
30K (u)
OLKAVS
Audio-visual speech recognition
Academia
int'l
all
rd
0
1,150 (h) audio, 5,750 (h) video
KMSAV
Audio-visual speech recognition
Academia
int'l
academic
rd
1
150 (h) transcribed, 2,000+ (h) untranscribed
Dataset
Typical Usage
Provider
Docu.
License
Redist
mod-x
Volume
K2NLG
Data-to-text generation
Academia
dom.
academic
rd
0
4K (s)
KommonGen
Common sense generation
Academia
int'l
all
rd
0
43K train, 2K test
KoCHET
Cultural heritage entity tasks
Academia
int'l
academic
unk
0
NER 112K, RE 39K, ET 113K
LBox Open
Legal language understanding
Academia
int'l
academic
rd
0
150K precedents
Korean GEC dataset
GEC
Academia
int'l
academic
rd
0
155K (s pair)
Korean Ambiguity Dataset
Word sense disambiguation
Academia
int'l
all
rd
0
35K (s), 8.2K surface forms
KorMedMCQA
Medical QA
Academia
int'l
academic
rd
1
7.5K questions
KBMC
Medical NER
Academia
int'l
all
rd
0
-
ESG-Kor
ESG information extraction
Academia
int'l
academic
rd
0
119K (s)
KBL
Legal language understanding
Academia
int'l
academic
rd
0
3.3K exam examples, 150K precedents
FunctionChat-Bench
Function calling benchmark
Industry
int'l
all
rd
0
-
KCL
Legal reasoning
Academia
int'l
academic
rd
1
283 MCQA, 169 essay questions
KorMedLawQA
Medical law QA
Academia
int'l
academic
rd
0
-
번호
데이터 종류
데이터 설명
1
우리말샘
이 사전에 대한 설명 : 다양한 어휘와 유의어 정보 등을 얻을 수 있는 대사전 : 로그인 후 전체 사전 데이터 다운로드 가능
2
NIA 사전
묻지도 따지지도 않고 다음 링크에서 엑셀로 다운로드 가능
3
국립국어원 언어정보나눔터
로그인 후 세종2007 코퍼스나 낭독체 음성 파일 등도 다운로드 가능, 다운 받을 때 간단한 서약에 체크만 하면 되는데 자료의 크기를 작게 나누어 놓아서 여러번 체크해야 한다는 것이 단점
4
AIHub
텍스트와 음성 멀티모달까지 가장 광범위한 데이터, 로그인 및 사용 목적과 기간을 명시한 사용 신청서 작성 후 허가 메일이 오면(하루 정도 걸린다) 다운로드 가능, 개별 데이터 종류·분량은 자주 갱신되므로 위 사이트에서 최신 목록 확인
5
국립국어원 모두의 말뭉치
다양한 분석 말뭉치(형태소 분석과 구문 분석 말뭉치 등), 다양한 도메인의 말뭉치(문어, 신문, 구어, 웹), 자연어 추론을 위한 말뭉치(유사 문장) 등 다양한 데이터들이 체계적으로 구축되어 있다. 말뭉치 종류·분량은 자주 갱신되므로 위 신청 페이지에서 최신 목록 확인, 로그인·메일 인증을 거쳐 데이터를 신청할 수 있고 다운로드 받기 위해서는 연구과제명과 수행기관, 약정 기간 등이 필수 입력 요소이다.
저작권 : 국어원이 승인한 이용 범위 내에서만 저작물을 낼 수 있으며 저작물을 내는 경우 국어원의 정보 제공 사실을 명시 필요, 즉 말뭉치 신청시 명시하면 학습에 사용할 수 있음, 그런 경우에도 아이디등을 제외한 텍스트 자체를 재배포 하는 것은 금지됨, AIHUB의 경우 회원 가입 절차 후에 승인을 받는 절차가 간소화되어 있음.
이 자료를 인용할 때는 다음을 사용한다. ACL Anthology 판과 arXiv 판 중 하나를 쓰면 된다.
@inproceedings{cho-etal-2020-open,
title = "Open {K}orean Corpora: A Practical Report",
author = "Cho, Won Ik and
Moon, Sangwhan and
Song, Youngsook",
booktitle = "Proceedings of Second Workshop for NLP Open Source Software (NLP-OSS)",
month = nov,
year = "2020",
address = "Online",
publisher = "Association for Computational Linguistics",
url = "https://www.aclweb.org/anthology/2020.nlposs-1.12",
pages = "85--93",
}
@article{cho2026open,
title = "Open {K}orean Corpora: A Practical Report",
author = "Cho, Won Ik and Moon, Sangwhan and Song, Youngsook",
journal = "arXiv preprint arXiv:2012.15621",
year = "2026",
note = "v3, revised 9 Jun. 2026",
url = "https://arxiv.org/abs/2012.15621",
}