Research notes.
Thermal density at the true system boundary
Density figures for accelerated compute are usually quoted at the chassis, even though much of the thermal system of a modern machine, the distribution units, pump skids and chillers, sits outside it. We define a boundary that follows the whole heat path from die to environment and recompute the density of nine systems against it. The ordering changes considerably: once its plant is counted, a 50 kW liquid-cooled rack comes out less dense than the 10 kW air-cooled rack it is normally compared against, while a sealed think node reaches 43.6 W per litre with no facility plant at all. We then show that at a fixed build height the annual cost of the floor a delivered watt occupies is inversely proportional to volumetric density, so the same ratios price real estate directly, and we propose a five-number reporting convention that would make this recomputation unnecessary.
Think AI Research22 August 20263 min read
-
Research
One model per accelerator
Most AI clusters run between 30 and 50 per cent utilised, and inference silicon sits idle more than half its life. We argue that a single convention is responsible for most of it, one model to one accelerator, and follow that convention from mechanism to money. The shortfall decomposes into four mechanisms sitting at four layers of the stack under four different owners, which is why fixing any one of them recovers so little. We treat the second mechanism, the memory ceiling, at length, because escaping it means splitting a model across devices that do not match and most stacks decline rather than degrade. We then measure what closing all four together produces on one node, 92.3 per cent sustained at 0 to 3.5 per cent orchestration overhead, and price the same twelve-model workload against on-demand rates from three providers, where the convention costs between $3,941 and $5,143 per model per month against $1,250. The conclusion we defend is the decomposition rather than either headline number.Think AI Research22 August 20263 min read
-
Research
Memory budgets for adapter fine-tuning at trillion-parameter scale
We work through a sizing exercise: whether a trillion parameter model, quantised to four bits and fine-tuned through adapters on a frozen base, fits inside the pooled memory of a single eight-accelerator RackNode. At the conservative end of the published footprint range it does, with 18 GB to spare, a margin of 2.3 per cent. Because that margin is smaller than the uncertainty in the multiplier it depends on, we present the result as a boundary case rather than a comfortable fit, and show how it moves with both the multiplier and the quantisation.Think AI Research22 August 20263 min read