IEEE Spectrum AI→ original

Prometheus: Majestic Labs server with 128 TB of memory for AI

Startup Majestic Labs is working on the Prometheus server, designed to solve the problem of limited memory when working with large language models. Unlike NVIDIA DGX B300 with 2 TB of memory, Prometheus can contain up to 128 TB. The company is preparing a DRAM-oriented architecture that will replace the traditional combined approach (HBM+DRAM).

AI-processed from IEEE Spectrum AI; edited by Hamidun News
Prometheus: Majestic Labs server with 128 TB of memory for AI
Source: IEEE Spectrum AI. Collage: Hamidun News.
◐ Listen to article

Artificial intelligence hardware startup Majestic Labs is developing a new AI server Prometheus with memory capacity up to 128 terabytes — more than 60 times the capacity of Nvidia's flagship DGX B300 server, one of the most powerful serial systems for processing AI workloads on the market. As IEEE Spectrum reports, the company is betting on an architecture built entirely around DRAM memory to solve the so-called "memory wall" problem — one of the main performance bottlenecks for inference in large language models.

Why memory became the main bottleneck for AI

  • Product — Prometheus server from Majestic Labs
  • Memory capacity — up to 128 terabytes
  • Comparison — more than 60 times more memory than Nvidia DGX B300
  • Architecture — unified, based entirely on DRAM (specifically LPDDR6), unlike Nvidia's HBM and DRAM combination
  • Key figure — Sha Rabii, co-founder and president of Majestic Labs

According to influential scientific work cited by IEEE Spectrum, token generation in large language models is fundamentally limited by memory speed: the speed at which a model outputs text is determined by how quickly data — primarily model weights — can be read from memory, not by how fast the computing units operate. The larger the model, the more acute this problem becomes, and eventually a "memory wall" emerges, which limits inference performance regardless of growth in processor computational power — no matter how many cores and tensor blocks engineers add, the final text output speed is constrained by memory bandwidth and available memory capacity.

How

Prometheus's DRAM architecture differs from Nvidia's approach

Majestic Labs co-founder and president Sha Rabii believes that memory scale gives the company a competitive advantage. According to him, Nvidia "has done phenomenal work creating a system capable of scaling," but as models grow, this approach becomes increasingly uneconomical: the system "ends up with excessive computational power and starving for memory." Nvidia servers today use fast, high-bandwidth memory (HBM) — usually for storing model weights — and a separate, typically larger, but slower DRAM pool, which handles other LLM tasks and server overhead.

Majestic Labs chose a fundamentally different path: a unified architecture built entirely on DRAM, specifically LPDDR6 memory, without division into fast and slow pools. According to Rabii, most memory interfaces are designed to operate at very short physical distances — sometimes just a few millimeters — and this is precisely what physically limits how much memory can be placed in a system while maintaining high bandwidth.

What this could mean for the AI infrastructure market

If Majestic Labs' bet pays off, the company will offer the market an alternative source of computing power for workloads that today lack memory — work with long context, mixed expert models with large total parameter volume under sparse activation, intensive KV-cache usage when serving multiple parallel requests. This challenges the common assumption that the race in AI infrastructure is primarily about the number of graphics processors and computational power in FLOPs: Prometheus shifts part of competitive focus to memory volume and architecture. Given Nvidia's dominant position in the AI server market, a startup's bet on memory-centric rather than computation-centric architecture stands out against most players who broadly follow the computationally-oriented roadmap set by Nvidia in recent years.

The very fact of such a project appearing shows that the "memory wall" has stopped being a narrow academic topic and has become a commercial incentive for launching a new hardware startup. If Majestic Labs brings Prometheus to serial production, data center operators and cloud providers will for the first time in a long time have a real alternative to the standard "GPU plus HBM" set, and competition in the AI infrastructure segment will cease to be exclusively a comparison of the number and generation of graphics accelerators.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Need AI working inside your business — not just in your newsfeed?

I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).

What do you think?
Loading comments…