Habr AI→ original

Sber released GigaChat 3.5 Ultra to open source with 432 billion parameters

The GigaChat team released GigaChat 3.5 Ultra to open source—a new model with 432 billion parameters. This is smaller than the 700 billion parameters of the previous GigaChat 3.1 Ultra flagship, but due to new data, an updated training recipe, and architectural changes, the model became stronger in code and agentic scenarios, as well as more efficient in memory and generation speed.

AI-processed from Habr AI; edited by Hamidun News
Sber released GigaChat 3.5 Ultra to open source with 432 billion parameters
Source: Habr AI. Collage: Hamidun News.
◐ Listen to article

Sberbank released GigaChat 3.5 Ultra with 432 billion parameters as open source

Sberbank released a new flagship model, GigaChat 3.5 Ultra, to the public domain (open source) — it contains 432 billion parameters, almost 40% fewer than the previous flagship GigaChat 3.1 Ultra with 700 billion parameters. Model weights are available to developers for independent deployment, without contacting the company's closed API.

How much more compact the model became

GigaChat 3.5 Ultra weighs 432 billion parameters versus 700 billion for GigaChat 3.1 Ultra — a reduction of almost 40%. Developers emphasize that this is not a trade-off of "smaller, but cheaper": according to them, new data, updated training recipe, and architectural changes made the model simultaneously more compact and stronger.

  • GigaChat 3.5 Ultra — 432 billion parameters, distributed as open source
  • GigaChat 3.1 Ultra (previous flagship of the line) — 700 billion parameters
  • Model size reduction — approximately 40%, or 268 billion parameters
  • First time for the GigaChat line: proprietary hybrid architecture scaled to hundreds of billions of parameters

What changed in the architecture

The Sber team scaled its proprietary hybrid GigaChat architecture to hundreds of billions of parameters for the first time within a single release. Hybrid architecture in such models typically refers to a combination of different types of layers and attention mechanisms, which allows scaling without proportional growth in computational costs. Developers say this made it possible to simultaneously speed up inference (answer generation) and expand the model's capabilities — primarily in code, agent scenarios (when the model itself calls tools and performs chains of actions), and complex subject areas.

"This is not a compromise of 'smaller, but cheaper': due to new data,

an updated training recipe, and architectural changes, the model became stronger, as well as more efficient in memory and generation speed," — according to Sber's official blog on Habr.

Why the company reduced the model size

Sber explains the model size reduction not as economy, but as engineering calculation: a smaller model is more efficient in memory and generation speed while maintaining or improving quality. The difference of 268 billion parameters between GigaChat 3.1 Ultra and GigaChat 3.5 Ultra directly reflects the infrastructure: the fewer parameters a model has, the less video memory is needed to load it and the higher the token generation speed with the same equipment.

What this means for developers

For those who deploy open source models independently, reducing the flagship from 700 billion to 432 billion parameters means a lower hardware entry threshold with comparable or higher answer quality. GigaChat 3.5 Ultra is oriented, among other things, at code and agent scenarios — exactly the tasks where developers most often compare open source models with competing closed APIs.

What this means

Sber continues to publish flagship GigaChat models in open access and simultaneously shifts toward more compact and faster architectures — a trend now characteristic of the entire large language model industry, where developers are seeking a balance between model size, infrastructure cost, and its practical utility in code and agent scenarios.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Want to stop reading about AI and start using it?

AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.

What do you think?
Loading comments…