A typical AI server.
Four accelerators in one chassis. Each one is an independent piece of silicon with its own memory, given its own work. Two are carrying load. Across clusters deployed today they average 40 per cent activity.
Every square is one accelerator.
768 of them: a hundred and ninety-two chassis, about twenty racks, a real cluster. Across the majority of clusters and AI silicon deployed today, silicon like this runs at 30 to 50 per cent utilisation.
Most of it is idle.
Median silicon activity is 40 per cent. The rest is bought, powered, cooled and doing nothing. Orchestration sees which accelerator it holds and how hot it runs. It does not see what that accelerator could still deliver, so the gap never closes.
- 768accelerators on screen, one per square: a hundred and ninety-two four-unit chassis
- $21.5mof silicon, at observed street prices
- 40%median silicon activity, measured inside a data centre built only for LLM work
- $12.9mof this cluster not computing, at any moment
Cluster utilisation: published measurement across AI serving and training fleets. Cost and idle value: modelled at observed street prices. Register and method.
Work is issued one device at a time.
A workload is handed a whole accelerator, because that is the unit the tooling knows how to give out. One model, one device. Whatever memory that model does not fill cannot be reached by anything else, and a model larger than one device has to be split by hand.
Loaded. Hot. Doing nothing.
Requests do not arrive evenly, so a serving device stays resident between them: allocated, its program loaded, its clocks held up to catch the next one. It draws almost full power and does almost no work. Nothing reports it as idle, because the job is still running.
This idleness is everywhere.
About twenty-four million accelerators have shipped worldwide. Every new site adds more silicon, more power, more cooling. The fraction working does not move.
The stack is fragmented.
Six layers of hardware. Six of software. Every one of them built for compatibility and for scaling quickly, not one of them built for the efficiency of the layer above or below it. The compute goes into the gaps between them.
The only way to tackle that inefficiency was to engineer them as one cohesive, symbiotic system.
Hardware and software designed against each other rather than around each other, and shipped as one thing. No layer waiting on a layer that cannot see it, and no gaps left to fall into. That system is the Think AI Fabric.
The same silicon.
The same infrastructure.
Awake.
One unit, working. 92.3 per cent sustained utilisation across the stack, on the same open ecosystem and the same silicon, at 0 to 3.5 per cent orchestration overhead.
One model. Different silicon. At the same time.
AI compute is chosen one vendor at a time, and the choice compounds. Every accelerator after the first has to match the first. The Fabric runs a single model across silicon of different vendors, different generations and different classes, together.
Not everyone can build this. Everyone can reach it.
think Grid puts the Fabric within reach without a build programme, from a tier three data centre coming online in Riyadh now, and from sites opening after it.
Demand more from your AI compute.
State of the art AI compute,
powering state of the art intelligence.
Every figure comes from a validated run or from published list prices. The bare metal delta is measured against the same serving stack in a container, over a concurrency sweep from one to two hundred, with the launch configuration verified identical.
Who we work with
Pre-seed · July 2026
$8M+
The largest deeptech AI pre-seed raise in MENA history.
Backed by
and others
Three layers.
One ecosystem.
Intelligent software bonded with high-performance hardware. Each layer stands alone. Together, they are the fabric.
A datacentre condensed into one sealed, liquid-cooled enclosure. Rack-class compute without the rack.
Bare-metal, hardware-aware orchestration. The intelligence layer for the next generation of infrastructure.
Nodes that scale together and think together. Any mix of think nodes becomes one supercompute cluster.
One fabric.
Across the silicon.
ILM runs today on NVIDIA Blackwell and Intel Arc, with more coming. Generations and architectures pool as one, workstation beside rack.
A datacentre,
condensed.
think Node carries its own datacentre. A sealed loop circulates coolant through the cold plates and rejects the heat straight to ambient air. No chiller, no facility water, nothing plumbed into the building. 7.1x the cooling density of a liquid-cooled 5U server and the CDU, chiller and piping it needs.
Measured across the true system volume: chassis, radiators, pumps and every part of the cooling path.
Mixed silicon. Mixed AI workloads.
Zero idle compute.
Most deployed silicon delivers a fraction of what it could. ILM runs serving, fine-tuning and utilities side by side across NVIDIA and Intel today, and the accelerators still to come, placing every job by memory fit, thermal headroom and affinity.
Four small models, four accelerators, and most of the VRAM doing nothing.
Nodes that scale together
and think together.
Constellation links any mix of think nodes into one pool of compute. Start with one. Grow without limits.
AI infrastructure
that thinks.
AI Node, ILM, and Constellation, bonded into one fabric. One node or an entire constellation, on your silicon.
Talk to us