Kandinsky 3
By Sber
Kandinsky 3 is a text-to-image diffusion model developed by Russian technology company Sber, part of the Kandinsky model series, notable for its multilingual capability including strong native support for Russian-language prompts alongside…
Definition
Kandinsky 3 is a text-to-image diffusion model developed by Russian technology company Sber, part of the Kandinsky model series, notable for its multilingual capability including strong native support for Russian-language prompts alongside English. It follows a latent diffusion design similar to other diffusion generators and has been released with open or research-permissive licensing for parts of the series, giving developers and researchers a way to run the model outside Sber's own hosted applications.
Overview
Kandinsky 3 is a text-to-image diffusion model developed by the Russian technology company Sber, continuing the Kandinsky series, and built in part to address a gap left by most mainstream text-to-image models: reliable handling of Russian-language prompts alongside English, rather than treating non-English input as an afterthought translated internally or handled poorly. That focus reflects Sber's positioning of the Kandinsky series as infrastructure for Russian-market products, where the majority of end users are expected to write prompts in Russian rather than English, making native-language prompt fidelity a core requirement rather than a secondary feature. The model follows a latent diffusion approach broadly similar in structure to other diffusion-based generators, encoding an image into a compressed latent space and training a denoising network conditioned on a text encoder, but Sber trained and tuned that text conditioning specifically to represent Russian text well, alongside English, rather than fine-tuning an English-centric encoder after the fact. This multilingual training is the detail that differentiates it mechanically from many Western-developed models built primarily around English-language datasets and encoders. Among text-to-image systems, Kandinsky 3 occupies a niche similar to other diffusion generators from outside the dominant English-language ecosystem, such as Kolors from Kuaishou for Chinese-language support, positioned against the largely English-first defaults of Stable Diffusion, DALL-E, and Midjourney. Its differentiator is this native multilingual capability rather than a distinct architectural innovation. In practice, Kandinsky 3 is used through Sber's own applications and public demos, and by developers building products for Russian-speaking users who want prompts written naturally in Russian to produce accurate results without needing to first translate to English, a step that can lose nuance or introduce translation artifacts in less multilingual-aware models. Its limitations include a smaller surrounding ecosystem of fine-tunes, extensions, and community tooling compared to the dominant open models like Stable Diffusion, since most third-party development effort in the open text-to-image space has concentrated around the more widely adopted architectures. Teams outside its target linguistic use case, or those wanting the largest possible library of community checkpoints and plugins, generally choose a more widely supported base model instead, reserving Kandinsky 3 specifically for Russian-language-first applications. The Kandinsky series overall has been developed with support from Sber's AI research groups and released with public demos and, for parts of the series, open weights, giving it a visible presence within the open text-to-image research community even though its practical adoption outside Russian-language contexts has remained comparatively limited relative to the dominant English-first open models.
Key Concepts
- Provides strong native support for Russian-language text prompts
- Uses a latent diffusion architecture similar to Stable Diffusion
- Developed by Sber as part of the ongoing Kandinsky model series
- Released with open or research-permissive licensing
- Compatible with community diffusion tooling ecosystems
- Supports multilingual prompting beyond English and Russian
- Iterates on earlier Kandinsky model architecture and training data