Evaluation of Croatian Word Embeddings

Svoboda, Lukáš; Beliga, Slobodan

Repository landing page

oai:www.unirepository.svkri.uniri.hr:infri_1036

Evaluation of Croatian Word Embeddings

Authors: Lukáš Svoboda
Slobodan Beliga
Publication date: 1 January 2018
Publisher: European Language Resources Association

Abstract

Croatian is poorly resourced and highly inflected language from Slavic language family. Nowadays, research is focusing mostly on English. We created a new word analogy dataset based on the original English Word2vec word analogy dataset and added some of the specific linguistic aspects from the Croatian language. Next, we created Croatian WordSim353 and RG65 datasets for a basic evaluation of word similarities. We compared created datasets on two popular word representation models, based on Word2Vec tool and fastText tool. Models have been trained on 1.37B tokens training data corpus and tested on a new robust Croatian word analogy dataset. Results show that models are able to create meaningful word representation. This research has shown that free word order and the higher morphological complexity of Croatian language influences the quality of resulting word embeddings

Similar works

Full text

Repository of the University of Rijeka

oai:www.unirepository.svkri.un...

Last time updated on 14/02/2023

This paper was published in Repository of the University of Rijeka.

Having an issue?

Is data on this page outdated, violates copyrights or anything else? Report the problem now and we will take corresponding actions after reviewing your request.