The Finno-Ugric Languages and The Internet Project

descriptionPublicationkeyboard_double_arrow_right Article , Other literature type , Conference object 17 Jun 2015 Finland Publisher:UiT The Arctic University of NorwayJournal:Septentrio Conference Series (eissn: 2387-3086,

Copyright policy )

Authors: Jauhiainen, Heidi; Jauhiainen, Tommi; Lindén, Krister;

doi: 10.7557/5.3471

handle: 10138/159402

The Finno-Ugric Languages and The Internet Project

- Summary
- Subjects
- Metrics

Abstract

This paper describes a Kone Foundation funded project called "The Finno-Ugric Languages and The Internet" together with some of the achieved results. The main activity of the project is to crawl the internet and gather texts written in small Uralic languages. The sentences and words of the found texts will be assembled into a freely available corpus. Crawling is done using the open source crawler Heritrix, which is developed by the Internet Archive. Heritrix crawls through the pages and passes the found texts to a language identifier. We are using a state of the art language identifier, which has been further developed within the project and has been evaluated using 285 languages. We describe the language identification evaluation results concerning the 34 Uralic languages known by the language identifier. We also describe the initial observations and results from the first five large crawls which were done in the national internet domains of Finland, Sweden, Norway, Russia, and Estonia.

Country

Finland

Related Organizations

University of Helsinki
Finland

Keywords

Computer and information sciences, Languages

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	3
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average