Archetype Capital

Archetype Capital

News & Updates

The Fastest Frontier AI Inference System

Faster than Cerebras?

Archetype Capital's avatar
Archetype Capital
Oct 04, 2026
∙ Paid

They call it the Volantis-A1. And they make a pretty big claim:

Faster than Cerebras?

If we take Volantis’s disclosed A-1 targets at face value and rank the systems specifically for low-latency frontier-model inference, then A-1 would be #1 on paper.

Inference: Up to 10,000 tok/s/user on >20T models

Memory: 10 TB @ 240 TB/s

Volantis just receieved $88m in funding. Sounds small-fry, but what they are working on is game-changing. And they say they are shipping to customers next year.

What problem has Volantis solved?

In photonics, the “memory wall” refers to the same fundamental problem as in electronic computing: compute capability is improving faster than the system can move data to and from memory.

Let’s take a look at the NVIDIA Rubin for example. The chip is absurdly powerful, so much so, that the engineering challenge has become “how do we move the data from the Rubin to the High-bandwidth-memory (HBM) and back, fast enough to keep the GPU from idling and wasting precious compute?

The problem is visually illustrated in the image below- you simply don’t have space for more memory chiplets next to the GPU because you are limited by the reach of the current (copper) electrical links.

These electrical links are microscopic copper interconnects. Put simply, if you try make the copper links longer, you get worse performance. That forces the memory to sit extremely close to the compute die, limiting how many high-bandwidth memory devices can surround it.. Less than you would like.

Here’s an analogy: Lian Yunzhi is a 6 year old speedcuber from China who recently broke the women’s Rubik’s Cube speed record- at 3.549 seconds.

Her brain is a NVIDIA Rubin, and her brain can actually solve the cube in 2 seconds flat.

The process looks like this:

  • Lian’s brain — the Rubin GPU — can solve the cube in 2 seconds.

  • But the information it needs is stored in HBM — her working memory.

  • Her brain constantly has to pull data from HBM, process it, and send the results back.

  • Meanwhile, her hands are waiting for the next instruction.

So where’s the problem?

Her brain can think faster than information can move between her brain and her memory.

That gap is the memory wall.

NVIDIA is trying many different solutions to the memory wall too, by using a lot more copper, but not in the way Volantis is attacking the problem- by using mature VSCEL technology.

Coherent is also developing wide-and-slow VCSEL-array architectures for AI scale-up interconnects.

The Volantis solution: integrated micro-VSCEL arrays and an optical memory fabric. If you can move data between the GPU and memory fast enough, more of Rubin’s enormous compute capacity can actually be used. (I don’t have confirmation for whether Volantis is using a Rubin, just using it as an example here).

The image on the left, below, shows the eletrical-interconnect memory wall problem.

On the right- Volantis' solution- you can attach vastly more memory because optical links relax the short-reach limitation of the electrical memory interface.

Who are the public winners?

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 David · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture