GigaChat Family: Efficient Russian Language Modeling Through Mixture of Experts Architecture

GigaChat team; Valentin, Mamedov; Kosarev, Evgenii; Leleytner, Gregory; Shchuckin, Ilya; Berezovskiy, Valeriy; Smirnov, Daniil; Kozlov, Dmitry; Averkiev, Sergei; Ivan, Lukyanenko; Proshunin, Aleksandr; Israfilova, Ainur; Baskov, Ivan; Chervyakov, Artem; Shakirov, Emil; Kolesov, Mikhail; Khomich, Daria; Latortseva, Darya; Porkhun, Sergei; Fedorov, Yury; Kutuzov, Oleg; Kudriavtseva, Polina; Soldatova, Sofiia; Egor, Kolodin; Pyatkin, Stanislav; Menshykh, Dzmitry; Sergei, Grafov; Damirov, Eldar; Vladimir, Karlov; Gaitukiev, Ruslan; Shatenov, Arkadiy; Fenogenova, Alena; Savushkin, Nikita; Minkin, Fedor

Computer Science > Computation and Language

arXiv:2506.09440 (cs)

[Submitted on 11 Jun 2025]

Title:GigaChat Family: Efficient Russian Language Modeling Through Mixture of Experts Architecture

Abstract:Generative large language models (LLMs) have become crucial for modern NLP research and applications across various languages. However, the development of foundational models specifically tailored to the Russian language has been limited, primarily due to the significant computational resources required. This paper introduces the GigaChat family of Russian LLMs, available in various sizes, including base models and instruction-tuned versions. We provide a detailed report on the model architecture, pre-training process, and experiments to guide design choices. In addition, we evaluate their performance on Russian and English benchmarks and compare GigaChat with multilingual analogs. The paper presents a system demonstration of the top-performing models accessible via an API, a Telegram bot, and a Web interface. Furthermore, we have released three open GigaChat models in open-source (this https URL), aiming to expand NLP research opportunities and support the development of industrial solutions for the Russian language.

Comments:	ACL-2025 System Demo
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2506.09440 [cs.CL]
	(or arXiv:2506.09440v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2506.09440

Submission history

From: Alena Fenogenova Ms [view email]
[v1] Wed, 11 Jun 2025 06:46:49 UTC (1,603 KB)

Computer Science > Computation and Language

Title:GigaChat Family: Efficient Russian Language Modeling Through Mixture of Experts Architecture

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:GigaChat Family: Efficient Russian Language Modeling Through Mixture of Experts Architecture

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators