-
Notifications
You must be signed in to change notification settings - Fork 1
Concept vectors
** General Terms** ABSTRACT ALGORITHMS, EXPERIMENTATION
The availability of machine readable taxonomy has been demonstrated by various applications such as document clas- sication and information retrieval. One of the main top- Keywords ics of automated taxonomy extraction research is Web min- Wikipedia, Web mining, categorization, concept vector ing based statistical NLP and a signicant number of re- searches have been conducted. However, existing works on automatic dictionary building have accuracy problems due 1. INTRODUCTION to the technical limitation of statistical NLP (Natural Lan- In the research area of linguistics and taxonomy, it is an guage Processing) and noise data on the WWW. To solve important task to sort out concepts in the world. Actu- these problems, in this work, we focus on mining Wikipedia, ally, a large number of dictionaries have been built by man- a large scale Web encyclopedia. Wikipedia has high-quality power to encourages learning such as bilingual dictionary and huge-scale articles and a category system because many and encyclopedia. Especially, constructing a machine read- users in the world have edited and rened these articles and able dictionary has been needed as a fundamental technol- category system daily. Using Wikipedia, the decrease of ogy for Semantic Web, which considers and processes the accuracy deriving from NLP can be avoided. However, af- meanings of texts in contrast to the traditional Web. Tax- liation relations cannot be extracted by simply descending onomy, one of (machine readable) hierarchical dictionaries, the category system automatically since the category system describes the information about to what categories each con- in Wikipedia is not in a tree structure but a network struc- cept belongs as a tree structure or DAG (Directed Acyclic ture. We propose concept vectorization methods which are Graph) structure. There are a considerable number of auto- applicable to the category network structured in Wikipedia. matic methods to build taxonomy, which recently use Web text data with statistical NLP (Natural Language Process- Categories and Subject Descriptors ing). However these methods have an accuracy problem due to the technical limitation of statistical NLP and noise data H.3.6 [Information Storage and Retrieval]: Library Automa- of Web text data. To improve the accuracy, the coverage is tion; M.7 [Knowledge Retrieval] sacriced. In order to resolve the accuracy problem deriving from NLP, we focus on Wikipedia. Wikipedia is a collabora- tive Wiki[8]-based encyclopedia. Since Wikipedia is based Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are on Wiki, anyone can edit and rene the articles using Web not made or distributed for profi or commercial advantage and that copies browser, that makes Wikipedia high quality and huge scale. bear this notice and the full citation on the firs page. To copy otherwise, to As for high quality, according to the statistics of Nature[6], republish, to post on servers or to redistribute to lists, requires prior specifi Wikipedia is about as accurate in covering scientic topics as permission and/or a fee. the Encyclopedia Britannica. As for huge scale, Wikipedia ICUIMC-09, January 15-16, 2009, Suwon, S. Korea contains not only general terms but also a large amount of Copyright 2009 ACM 978-1-60558-405-8....00.