A typical AI server.

Eight accelerators in one chassis. Each one is an independent piece of silicon with its own memory, given its own work. This is one day of it. Across clusters deployed today they average 40 per cent activity.

Ahmed AlSharif. Founder and CEO.

A systems software engineer. The place I am happiest is the layer between the silicon and the software. This gets a little technical. It will be worth it.

Every square is one server in a cluster.

Ninety-six servers. Seven hundred and sixty-eight accelerators. About nineteen racks.

Most of them are idle.

This is not a diagram. It is one real fleet, read live: the National Research Platform. Every cell is one of its accelerators. The lit ones are working. The rest are bought, powered, cooled and doing nothing.

Live now

  • 1,292accelerators reporting across the fleet, this second
  • 46%of them reading exactly zero: powered, allocated, doing nothing at all
  • 7%what the tensor cores are really doing, averaged over the whole fleet
  • $2.0mof silicon doing nothing this second, priced model by model at street rates

Read live from the National Research Platform's public telemetry, which reports activity off the card rather than inferring it from the scheduler, and stamped above with the minute it was read. The valuation prices all thirty models in that fleet at street rates, September 2026; those are asking prices, not settled sales. Register and method.

Not one cluster. All of them.

Four public fleets, same treatment. Only the first reports what its silicon is actually doing. The other three report what has been handed out, which is the number the industry quotes and is not the same thing.

Live where marked

  • 43%what Nautilus's counter calls busy: a kernel is loaded on the chip
  • 7%what its tensor cores are really doing in that same time
  • 18%of a busy reading that is real arithmetic. That measured ratio is what makes the other three comparable
  • 95%of Polaris allocated at Argonne, so at most a fifth of it can be computing
64% 46% $2.0m

Device telemetry at the National Research Platform is read live in the browser. Argonne and Akash report allocation, counted as nodes holding a job or a reservation, and are read through a caching proxy so a busy page never reaches their infrastructure. Each column carries the minute it was read; a feed that cannot be reached falls back to its own last reading and is stamped with that date rather than called live. The comparable number applies Nautilus's measured ratio to each fleet's allocation, which makes it a ceiling rather than an estimate: allocation counts a whole node the moment a job lands on it. Aurora and Akash carry no price. Aurora's Intel Max 1550 never traded as a loose part, and Akash is a marketplace of mixed hardware with no published bill of materials. Register and method.

This is everywhere.

About twenty-four million accelerators have shipped worldwide. Every new site adds more silicon, more power, more cooling.

Idle is not one problem. It is three.

Three separate leaks, one result. Compute falls between the layers. Memory sits unreachable inside the device. Energy goes on the gaps between requests.

The same silicon.
The same infrastructure.
Awake.

One unit, working. 92.3 per cent sustained utilisation across the stack, on the same open ecosystem and the same silicon, at 0 to 3.5 per cent orchestration overhead.

One model. Different silicon. At the same time.

AI compute is chosen one vendor at a time, and the choice compounds. Every accelerator after the first has to match the first. The Fabric runs a single model across silicon of different vendors, different generations and different classes, together.

The Think AI Fabric. Own it, or reach it.

Efficiency is built into every node. Own the hardware and run it in your building: MicroNode, SuperNode, UltraNode, RackNode. Or reach those same nodes from anywhere on think Grid. Same silicon. Same efficiency.

See how it works

Demand more from your AI compute.

State of the art AI compute,
powering state of the art intelligence.

92.3%Sustained utilisationIndustry norm 30 to 50 per cent
2.35×Power densityOver standard architecture
10×Cheaper per million tokens$0.96 against a $9.90 cloud average
+20.7%Faster first token, bare metalAgainst a containerised deployment

Every figure comes from a validated run or from published list prices. The bare metal delta is measured against the same serving stack in a container, over a concurrency sweep from one to two hundred, with the launch configuration verified identical.

Who we work with

Pre-seed · July 2026

$8M+

The largest deeptech AI pre-seed raise in MENA history.

Backed by

  • WAED
  • RAED
  • DTV

and others

Three layers.
One ecosystem.

Intelligent software bonded with high-performance hardware. Each layer stands alone. Together, they are the fabric.

Fabric

One fabric.
Across the silicon.

ILM runs today on NVIDIA Blackwell and Intel Arc, with more coming. Generations and architectures pool as one, workstation beside rack.

Blackwell architecture CDNA architecture Battlemage architecture Hopper architecture Fabric
Supported todayIn enablement

A datacentre,
condensed.

think Node carries its own datacentre. A sealed loop circulates coolant through the cold plates and rejects the heat straight to ambient air. No chiller, no facility water, nothing plumbed into the building. 7.1x the cooling density of a liquid-cooled 5U server and the CDU, chiller and piping it needs.

Conventional 4 kW liquid-cooled datacentre plant room required
Outdoor dry cooler Chiller warm climates Facility loop valves, piping Pump skid N+1 5U server CDU 42U rack 42U rack + plant To scale To scale vs 42U rack think think think SuperNode 471 x 285 x 513 mm · 68.9 L · 3 kW 7.1x vs 5U + CDU
Component layout illustrative. Volumes measured across the true system, including every part of the cooling path.
True system volume 659 L 68.9 L
Cooling density 6.1 W/L 43.6 W/L

Measured across the true system volume: chassis, radiators, pumps and every part of the cooling path.

Mixed silicon. Mixed AI workloads.
Zero idle compute.

Most deployed silicon delivers a fraction of what it could. ILM runs serving, fine-tuning and utilities side by side across NVIDIA and Intel today, and the accelerators still to come, placing every job by memory fit, thermal headroom and affinity.

Four small models, four accelerators, and most of the VRAM doing nothing.

Beforeone model per chip
Afterthink AI Node with ILM
Inference ManagementCPU · single chip · multi chip · sharded
Fine-Tuning OrchestrationSFT · LoRA · QLoRA
Model ConversionHuggingFace to GGUF · quantization
Model Acquisitionsecure downloads · recovery
Real-Time Observabilitymetrics · logs · safe cancellation

Nodes that scale together
and think together.

Constellation links any mix of think nodes into one pool of compute. Start with one. Grow without limits.

Drag a job onto a node
Booting first node 1 node online

AI infrastructure
that thinks.

AI Node, ILM, and Constellation, bonded into one fabric. One node or an entire constellation, on your silicon.

Talk to us