Neurotechnology contributed to the General Lithuanian Language Corpus by training and optimizing two vectorized Lithuanian language models designed to support the next generation of artificial intelligence solutions and digital services.
- 3.9 billion words licensed Lithuanian language corpus
- 2 vectorized language models pre-trained completely from scratch
- ~2000 hours of model training and optimization
- Multiple architectures extensively evaluated and benchmarked
words licensed Lithuanian
language corpus
language models pre-trained completely
from scratch
hours of model training
and optimization
architectures extensively evaluated
and benchmarked
Neurotechnology, together with Vytautas Magnus University, UAB Tilde IT and MB Krilas, completed a national project focused on creating the General Lithuanian Language Corpus and vectorized Lithuanian language models. The initiative resulted in a large-scale licensed Lithuanian language dataset and two vectorized Lithuanian language models designed to support future artificial intelligence development in Lithuanian.
As artificial intelligence adoption continues to grow, the availability of high-quality language resources has become essential for developing reliable AI systems. While major languages benefit from extensive datasets and pretrained models, smaller languages often lack the resources needed to develop advanced AI solutions.
Associate Professor at Vytautas Magnus University Faculty of Humanities
Within the project, Neurotechnology was responsible for training and optimizing the language models. The company's natural language processing team evaluated model architectures, optimized training methodologies and developed reusable technologies that can support future Lithuanian language model development.
Training the Lithuanian Language Models
Building an effective language model requires both high-quality linguistic resources paired with significant computational capabilities. Moving past initial experimental stages, the team dedicated heavy computational runtimes to finalize two distinct classes of models:
-
Large Language Model (LLM)
- Architecture: Generative AI model, built on the Llama3 framework.
- Model Size: Approximately 1.04 billion parameters.
- Data Utilization: Trained on the entire Lithuanian text dataset (3.9 billion words).
- Training Time: ~1,920 hours on powerful computers.
- Technical Highlights: Built completely from scratch with no prior knowledge. It includes a custom tool to understand Lithuanian words and can process massive files or long documents all at once.
-
Masked Language Model (MLM)
- Architecture: Text-analysis model (used for searching, sorting, and analyzing text) built on ModernBERT.
- Model Size: Approximately 0.2 Billion parameters.
- Data Utilization: Trained on about half of the Lithuanian text dataset.
- Training Time: 210 hours on powerful computers.
- Technical Highlights: Uses a custom tool specifically designed to break down complex Lithuanian grammar. It is highly accurate and can easily analyze long passages of text.
Technical Lead of the NLP Department at Neurotechnology
Evaluating Reliability and Results
The reliability and real-world suitability of the corpus and models were assessed by project partner MB Krilas, which carried out a comprehensive validation and testing process. The corpus was reviewed from multiple angles: scope and representativeness, linguistic quality, diversity of topics and styles, and whether it contained biased or discriminatory content or personal data.
Model capability was tested through a practical task – generating summaries of Lithuanian-language texts. The evaluation separately assessed how well the model had learned Lithuanian, whether it produced false or fabricated information, and whether it remained unbiased.
The project found the corpus to be high-quality, balanced and safe to use, and the model to have learned Lithuanian well – producing fluent and, in most cases, factually accurate text suitable for real-world application.
Optimizing Performance for the Lithuanian Language
Model quality was evaluated using the Perplexity metric, which assesses how accurately a language model can predict the next word in a text sequence. Lower Perplexity scores indicate better language understanding and prediction capabilities and a better command of grammatical structures.
By fine-tuning text segment limits and engineering structural configurations tailored specifically to Lithuanian syntax, the project achieved Perplexity results that comfortably exceeded previously available regional open-source models. In addition to the main models, the project also includes a reusable training methodology and a tokenizer specifically adapted for the Lithuanian language. A tokenizer is a method by which text is broken down into the smallest meaningful units.
NLP Team Lead at Neurotechnology
From Language Understanding to Real-World Applications
The General Lithuanian Language Corpus was compiled from fiction and non-fiction literature, media content, spoken language and documents, collected by Vytautas Magnus University and UAB Tilde IT. Numerous media and cultural partners contributed data, including DELFI, LRT.lt, Vakarų ekspresas, Lrytas.lt, Švenčionių kraštas, Gargždai.lt, Sena.lt, AutoGidas.lt, 15min, Bernardinai.lt, Humanitarų meka, ELTA, BNS, Kas vyksta Kaune, VDA and others.
Following pretraining, the models were adapted for Named Entity Recognition (NER), a task that involves identifying and classifying people, locations, dates and other entities within text.
Validation results demonstrated that the model could accurately identify and classify key information within Lithuanian-language documents. These capabilities enable practical applications such as intelligent search systems, virtual assistants, customer service automation, document search, information extraction, data anonymization and automated document processing.
All project outputs – the open-source AI models and the corpus – are freely available on the Hugging Face platform, and the corpus is also accessible via the CLARIN-LT platform.
The project was implemented by the State Digital Solutions Agency and carried out by a four-partner consortium led by Vytautas Magnus University, together with UAB Neurotechnology, UAB Tilde IT and MB Krilas. The project ran from December 2024 to April 2026.
help you create and deploy the NLP-based solutions.
Related Research
Our first open-source large language model in the Lithuanian language is available to promote AI solutions development in the region.
expanding LLM capacity
The new research demonstrates how the application of continual learning advances LLMs' linguistic capabilities.
