Extracting Named Entity Translingual Equivalence with Limited Resources

Fei Huang, Stephan Vogel, Alex Waibel

Research output: Contribution to journalArticle

9 Citations (Scopus)

Abstract

In this article we present an automatic approach to extracting Hindi-English (H-E) Named Entity (NE) translingual equivalences from bilingual parallel corpora. In the absence of a Hindi NE tagger or H-E translation dictionary, this approach adapts a Chinese-English (C-E) surface string transliteration model for H-E NE extraction. The model is initially trained using automatically extracted C-E NE pairs, then iteratively updated based on newly extracted H-E NE pairs. For each English person and location NE in each sentence pair, this approach searches for its Hindi correspondence with minimum transliteration cost and constructs an H-E NE list from the bilingual corpus. Experiments show that this approach extracted 1000 H-E NE pairs with a precision of 91.8%.

Original languageEnglish
Pages (from-to)124-129
Number of pages6
JournalACM Transactions on Asian Language Information Processing
Volume2
Issue number2
DOIs
Publication statusPublished - 1 Jun 2003
Externally publishedYes

    Fingerprint

Keywords

  • Algorithms
  • information extraction
  • Language
  • machine translation
  • Named entity translation
  • Performance
  • transliteration

ASJC Scopus subject areas

  • Computer Science(all)

Cite this