Semi-Automatic Alignment of Multilingual Parts of Speech Tagsets

Authors

  • S. Yashothara University of Moratuwa
  • R.T. Uthayasanker University of Moratuwa
  • G.V. Dias University of Moratuwa
  • S. Jayasena University of Moratuwa

DOI:

https://doi.org/10.13053/cys-26-3-4357

Keywords:

Parts of speech, POS tagset mapping, POS tagset alignment, semi-automatic approach, BIS tagset, UOM tagset, tamil NLP, sinhala NLP

Abstract

We cast the problem of mapping a pair of Parts of Speech (POS) tagsets as a labelled tree mapping problem and present a general-purpose semi automatic POS tree alignment algorithm to solve the alignment. This algorithm can be used to align two POS tagsets of different languages or the same language. We evaluate its usefulness using POS tagsets of two languages: Tamil and Sinhala. The proposed approach shows that manual effort in prior approaches is drastically reduced due to the proposed algorithm and eliminates the need to create new POS tagsets.

Author Biographies

S. Yashothara, University of Moratuwa

Department of Computer Science and Engineering

R.T. Uthayasanker, University of Moratuwa

Department of Computer Science and Engineering

G.V. Dias, University of Moratuwa

Department of Computer Science and Engineering

S. Jayasena, University of Moratuwa

Department of Computer Science and Engineering

Downloads

Published

2022-08-31

Issue

Section

Articles