An Analysis of Variance Method to Detect Collocations in a Pedagogical Domain Corpus

Autores/as

  • Yuridiana Alemán Benemérita Universidad Autónoma de Puebla, Computer Science Faculty
  • María Somodevilla Benemérita Universidad Autónoma de Puebla, Computer Science Faculty
  • Darnes Vilariño Benemérita Universidad Autónoma de Puebla, Computer Science Faculty

DOI:

https://doi.org/10.13053/cys-24-2-3411

Palabras clave:

Pedagogical domain, variance, collocations, ontology, important concepts

Resumen

In this paper, an exploratory experiment, based on analysis of variance, was carried out in order to get collocations in a pedagogical domain corpus. A semi-automatic corpus containing learning styles papers in Spanish was built. After wards, the corpus was lemmatized and a bigrams representation was extracted. The proposed method consists on divide the list of bigrams in quartiles, and analyzing the variance on each one of them. A list of collocations, which was evaluated using a gold standard built by an expert in the domain, was retrieved from each experiment accordingto established thre sholds for the method. Results showed a retrieved list with important collocation in the selected domain.

Descargas

Publicado

2020-06-23