Welcome to the VLO!
Use the search bar below to start searching through hundreds of thousands of language resources, or continue to browse everything and use facets to narrow down to your area of interest or discover new resources.
See all records Learn more Take a quick tourUse the categories below to limit the search results to those matching the selected value(s).
These levels provide an indication of the degree to which resources and tools are publicly accessible. Please check the specific conditions on any resource or tool that you end up using.
The Kielipankki version of Fenno-Ugrica (http://urn.fi/urn:nbn:fi:lb-2014073056) is available in Kielipankki - the Langu…
The Kielipankki version of Fenno-Ugrica (http://urn.fi/urn:nbn:fi:lb-2014073056) is available in Kielipankki - the Language Bank of Finland at http://urn.fi/urn:nbn:fi:lb-2015103001 Change Log: 26.2.2019 - Missing Eastern Mari added to location-URN (urn:nbn:fi:lb-2015103001) - Name changed: - ugrica > Ugrica; Fenno-…
The Dictionaries for Hill Mari and Finnish are based on lexical material from a large array of dictionaries and written …
The Dictionaries for Hill Mari and Finnish are based on lexical material from a large array of dictionaries and written literature from the 20th and 21th centuries. A note-worthy dictionary is the Hill Mari - Russian dictionary by Anna A. Savatkova (1981). Translation work is being conducted at the University of Helsin…
A Mari-English dictionary with 42,560 headwords and 82,740 subentries, including 10,750 set phrases.
A Mari-English dictionary with 42,560 headwords and 82,740 subentries, including 10,750 set phrases.
HeLI off-the-shelf language identifier with language models for 200 languages. The program will read the <infile> and cl…
HeLI off-the-shelf language identifier with language models for 200 languages. The program will read the <infile> and classify the language of each line as one of the 200 languages it knows and writes the results, one ISO 639-3 code per line, into file <outfile>. It can identify c. 3000 sentences per second using one c…
HeLI off-the-shelf language identifier with language models for 200 languages. The program will read the <infile> and cl…
HeLI off-the-shelf language identifier with language models for 200 languages. The program will read the <infile> and classify the language of each line as one of the 200 languages it knows and writes the results, one ISO 639-3 code per line, into file <outfile>. It can identify c. 3000 sentences per second using one c…
Wanca 2016 is a collection of web corpora in small Uralic languages. The collection is composed of 29 sentence corpora i…
Wanca 2016 is a collection of web corpora in small Uralic languages. The collection is composed of 29 sentence corpora in different languages. The corpora have been collected from the Internet using the automated system developed in the Finno-Ugric Languages and the Internet project (SUKI) supported by the Kone foundat…
The VRT version of Wanca 2016 is a collection of web corpora in small Uralic languages. The collection is composed of 29…
The VRT version of Wanca 2016 is a collection of web corpora in small Uralic languages. The collection is composed of 29 sentence corpora in different languages. The corpora have been collected from the Internet using the automated system developed in the Finno-Ugric Languages and the Internet project (SUKI) supported …
The Korp version of Wanca 2016 is a collection of web corpora in small Uralic languages. The collection is composed of 2…
The Korp version of Wanca 2016 is a collection of web corpora in small Uralic languages. The collection is composed of 29 sentence corpora in different languages. The corpora have been collected from the Internet using the automated system developed in the Finno-Ugric Languages and the Internet project (SUKI) supported…
HeLI off-the-shelf language identifier with language models for 200 languages. The program will read the <infile> and cl…
HeLI off-the-shelf language identifier with language models for 200 languages. The program will read the <infile> and classify the language of each line as one of the 200 languages it knows and writes the results, one ISO 639-3 code per line, into file <outfile>. It can identify c. 3000 sentences per second using one c…
HeLI off-the-shelf language identifier with language models for 200 languages. The program will read the <infile> and cl…
HeLI off-the-shelf language identifier with language models for 200 languages. The program will read the <infile> and classify the language of each line as one of the 200 languages it knows and writes the results, one ISO 639-3 code per line, into file <outfile>. It can identify c. 3000 sentences per second using one c…