Constellation
Nodes that together.
Constellation makes multiple think AI nodes operate as one coordinated system: an office MicroNode, a lab SuperNode, and a rack row of RackNodes in the same pool. This page carries the numbers behind the story.
Drag a job onto a node
SUPERNODE
healthy
edge office
VRAM in use0%
Booting first node
1 node online
Any mix of nodes.
One pool of compute.
One unified compute pool, wherever your nodes live. No rack, no raised floor, no industrial cooling solution needed.
Three engines, one pool. vLLM, SGLang and llama.cpp side by side.
ONE MODULAR SUPERCOMPUTE CLUSTER
2.0 TB
Pooled VRAM
21.0 PFLOPS FP8 DENSE
ILM ORCHESTRATED
MicroNode
edge office
192 GB · 2 PFLOPS
vLLM · Mistral 7B
SuperNode
lab bench
384 GB · 4 PFLOPS
SGLang · agents
UltraNode
server room
672 GB · 7 PFLOPS
vLLM · GPT-OSS 120B
RackNode
rack row A
768 GB · 8 PFLOPS
llama.cpp · R1 32B
Pooled VRAM allocation2.0 TB serving
Serving 2.0 TB
Training 0 TB
No rack required
Office, lab, or datacentre: nodes join one fabric wherever they live.
Unified VRAM + compute
Pooled capacity serves even larger models, faster.
Zero-downtime hot-reload
Atomic cutover keeps every session alive as the fabric reshapes.
Sovereign fabric
Every byte stays on premises. Air-gapped, always.
Start with one node. Keep the door open to a fleet.
Every think AI node is already Constellation-ready. Talk to us about where your estate should go next.
Contact Think AI