Skip to main content

Showing 1–2 of 2 results for author: Pyatkin, S

Searching in archive cs. Search in all archives.
.
  1. arXiv:2506.09440  [pdf, ps, other

    cs.CL cs.AI

    GigaChat Family: Efficient Russian Language Modeling Through Mixture of Experts Architecture

    Authors: GigaChat team, Mamedov Valentin, Evgenii Kosarev, Gregory Leleytner, Ilya Shchuckin, Valeriy Berezovskiy, Daniil Smirnov, Dmitry Kozlov, Sergei Averkiev, Lukyanenko Ivan, Aleksandr Proshunin, Ainur Israfilova, Ivan Baskov, Artem Chervyakov, Emil Shakirov, Mikhail Kolesov, Daria Khomich, Darya Latortseva, Sergei Porkhun, Yury Fedorov, Oleg Kutuzov, Polina Kudriavtseva, Sofiia Soldatova, Kolodin Egor, Stanislav Pyatkin , et al. (9 additional authors not shown)

    Abstract: Generative large language models (LLMs) have become crucial for modern NLP research and applications across various languages. However, the development of foundational models specifically tailored to the Russian language has been limited, primarily due to the significant computational resources required. This paper introduces the GigaChat family of Russian LLMs, available in various sizes, includi… ▽ More

    Submitted 11 June, 2025; originally announced June 2025.

    Comments: ACL-2025 System Demo

  2. Probabilistically Robust Watermarking of Neural Networks

    Authors: Mikhail Pautov, Nikita Bogdanov, Stanislav Pyatkin, Oleg Rogov, Ivan Oseledets

    Abstract: As deep learning (DL) models are widely and effectively used in Machine Learning as a Service (MLaaS) platforms, there is a rapidly growing interest in DL watermarking techniques that can be used to confirm the ownership of a particular model. Unfortunately, these methods usually produce watermarks susceptible to model stealing attacks. In our research, we introduce a novel trigger set-based water… ▽ More

    Submitted 18 September, 2024; v1 submitted 16 January, 2024; originally announced January 2024.

    Journal ref: Proceedings of the International Joint Conferences on Artificial Intelligence, 33 (2024), 4778-4787