A constraint-aware GPU allocator was benchmarked against a FIFO scheduler across seven scenarios, showing up to 33 percentage points higher utilization and up to 105% higher priority-weighted value.
Aug 17, 2026
14d agoKey Details
- Evaluated across seven benchmark scenarios including mixed control, real-time contention, training-heavy, large mixed, oversubscribed, scale test, and uniform priority
- GPU utilization rose by as much as 33 percentage points (53.6% to 87.0% on 8 GPUs and 16 training-heavy jobs)
- Priority-weighted value rose between 24.6% and 105.1% across contended scenarios, averaging a 52% gain
- Achieved allocation decision latency of 1 to 2 milliseconds on contended workloads and 15 milliseconds on a 64-GPU 30-job scale test
- Operates a fast structural constraint heuristic on the hot path combined with a 24-hour horizon formal model re-optimized every 30 to 60 minutes