ML Systems Software Engineer
Work at the seam between the model and the silicon: the layer where a scheduler decision, a memory choice or a kernel path is the difference between silicon that computes and silicon that waits. This is the core of ILM.
Work at the seam between the model and the silicon: the layer where a scheduler decision, a memory choice or a kernel path is the difference between silicon that computes and silicon that waits. This is the core of ILM.
What you will do
- Build and optimise the serving path: batching, memory management, cache behaviour and scheduling under real concurrency.
- Profile end to end and find the actual bottleneck rather than the assumed one.
- Make models share accelerators well, including across silicon of different capacities and generations.
- Work close to the runtime: drivers, kernels, memory allocators and the telemetry that makes behaviour visible.
- Build the measurement harness. If it is not measured on real hardware it is not a result.
- Take a change from a hypothesis to a benchmark to production.
What you need
- Four years or more in systems or ML infrastructure engineering.
- Strong Python and a systems language, typically C++ or Rust.
- You have optimised inference or training throughput and can explain exactly where the win came from.
- Understanding of accelerator memory hierarchies, kernel launch behaviour and where time actually goes.
- Comfort on bare metal rather than behind a managed service.
Useful, not required
- CUDA, ROCm, Triton or comparable kernel-level work.
- Contributions to vLLM, TensorRT-LLM, SGLang or similar.
- Distributed serving and multi-accelerator sharding.
- Published benchmark or systems work.
What we offer
- One of the first deep tech companies in the region, building foundational technology in house.
- Meaningful ownership and impact at an early stage.
- Competitive early-stage compensation.
- Close collaboration with a small, senior team.
- Problems that combine hardware, systems and AI at scale.