The role of word sense disambiguation in automated text categorization

Gómez Hidalgo, José María; De Buenaga Rodríguez, Manuel; Cortizo Pérez, José Carlos

The role of word sense disambiguation in automated text categorization

Identifiers

URI: http://hdl.handle.net/11268/5479

DOI: 10.1007/11428817_27

Publication date

2005

Authors

Gómez Hidalgo, José María

De Buenaga Rodríguez, Manuel

Cortizo Pérez, José Carlos

Metrics

Abstract

Automated Text Categorization has reached the levels of accuracy of human experts. Provided that enough training data is available, it is possible to learn accurate automatic classifiers by using Information Retrieval and Machine Learning Techniques. However, performance of this approach is damaged by the problems derived from language variation (specially polysemy and synonymy). We investigate how Word Sense Disambiguation can be used to alleviate these problems, by using two traditional methods for thesaurus usage in Information Retrieval, namely Query Expansion and Concept Indexing. These methods are evaluated on the problem of using the Lexical Database WordNet for text categorization, focusing on the Word Sense Disambiguation step involved. Our experiments demonstrate that rather simple dictionary methods, and baseline statistical approaches, can be used to disambiguate words and improve text representation and learning in both Query Expansion and Concept Indexing approaches.

UNESCO Subjects

Lenguajes controlados
Inteligencia artificial
Robótica

Bibliographic reference

Gómez Hidalgo, J. M., De Buenaga Rodríguez, M., & Cortizo Pérez, J. C. (2005). The Role of word sense disambiguation in automated text categorization. Lecture Notes in Computer Science, 3513, 298-309.

Type of document

journal article

Collections

Otras Áreas de STEAM

Full item page

The role of word sense disambiguation in automated text categorization

Identifiers

Publication date

Authors

Advisors

Editors

Journal Title

Journal ISSN

Volume Title

Publisher

Metrics

Research Projects

Organizational Units

Journal Issue

Abstract

Description

UNESCO Subjects

Keywords

Bibliographic reference

Type of document

Collections