Why Your GPU Fleet Is Both Full and Idle: Building a Capacity Orchestration Layer
Managing GPU capacity manually produces fleets that are simultaneously fully allocated and poorly utilized. A capacity orchestration layer can fix the problem.
Join the DZone community and get the full member experience.
Join For FreeMost machine learning (ML) platforms still hand out accelerators through static quotas, spreadsheet-driven approvals, and hallway negotiation. That model was tolerable when capacity was cheap, and demand was steady. Neither holds for modern AI workloads. This article walks through why static allocation breaks down, the five components of a capacity orchestration layer (intent capture, policy, hardware brokering, continuous reallocation, and goodput accounting), and a 90-day rollout plan that won't destroy your users' trust in the scheduler.
The Contradiction You Have Probably Already Measured
If you run infrastructure for a large ML organization, you have likely seen a dashboard that says something uncomfortable. The fleet is 90-plus percent allocated. Teams are filing tickets about blocked capacity. And a sampled profile of the same fleet shows a meaningful fraction of devices doing no useful work.
Buying more accelerators does not resolve this. It scales the inefficiency. The contradiction comes from a specific architectural choice: treating an accelerator reservation as ownership rather than as a lease with conditions.
The industry has spent years optimizing acquisition: supply agreements, multi-year commitments, reserved cloud capacity. Comparatively little engineering has gone into what happens after the hardware lands in the fleet.
Why Static Quotas Fail for ML Specifically
Conventional service capacity planning works because service demand is autocorrelated. Today's traffic predicts tomorrow's. ML workload demand does not behave this way, for four reasons.
Demand is event-driven. An eval result at 2 p.m. can turn a 16-device research project into a 512-device training run by the next morning. No quarterly quota cycle accommodates that.
Jobs differ in shape, not just size. A gang-scheduled training job with all-reduce dependencies across 64 devices and a stateless inference replica are both "GPU requests" to a naive scheduler, but they have incompatible placement, topology, and preemption semantics. A fleet fragmented without topology awareness ends up with allocatable devices that no large job can actually use.
Reservation and consumption diverge. A job holding 32 devices while blocked on an input pipeline is occupying capacity and producing nothing. Device-level utilization metrics will not reliably catch this, because a device spinning on a synchronization barrier can still report activity.
Returning capacity is asymmetrically risky. This one dominates. If reacquiring capacity takes two weeks of negotiation, the rational move for every team is to hold everything they might need at peak, indefinitely. Static allocation rewards hoarding. Each over-request is a locally correct decision inside a badly designed system, which is why exhortations to "be good citizens" never fix it.
The Architecture
A capacity orchestration layer sits between your workload submission API and the underlying scheduler, whether that is Kubernetes with Kueue or Volcano, Slurm, Ray, or an internal system. It does not replace the scheduler. It feeds the scheduler better information and revisits its decisions after admission.
1. Intent Capture: A Device Count Is an Answer, Not a Requirement
When a user requests "32 A100s," they have already collapsed their actual requirement into one number. The scheduler needs the requirement itself: what the workload can substitute, what it can tolerate, and what it cannot give up. In practice that looks like a spec along these lines:
apiVersion: capacity.internal/v1
kind: WorkloadIntent
metadata:
name: ranking-hparam-sweep
spec:
class: research-batch # drives policy, not placement
preferred:
accelerator: a100-80gb
count: 32
substitutable:
- accelerator: h100-80gb
count: 20 # based on measured step-time on this model
- accelerator: tpu-v5e
count: 64
requiresRecompile: true # XLA path validated in CI
topology:
gangScheduled: true
minInterconnect: nvlink-domain
elasticity:
minReplicas: 8
maxReplicas: 32
resizeCost: 90s # checkpoint + reshard
preemption:
tolerated: true
gracePeriod: 120s
checkpointInterval: 600s
deadline: "2026-08-14T00:00:00Z"
Two fields carry most of the value. substitutable is what makes cross-architecture brokering possible. preemption.tolerated, paired with a real checkpointInterval, is what makes a job safely reclaimable.
Interruption tolerance should be verified, not self-declared. Run a preemption drill in CI: kill the job, confirm it resumes from checkpoint within the stated grace period. A self-reported tolerance that has never been exercised will eventually be exercised in production, on a bad day.
Be honest about the engineering cost. Portable workload definitions, container images built for multiple backends, and checkpoint formats that survive a device-count change are all real work. Elastic resharding in particular is hard for models with complex parallelism strategies. Treat the full intent spec as something teams grow into. A job that declares only preferred still schedules; it just gets fewer options.
2. Policy Engine: Encode the Argument Once
The point of a policy engine is not to remove human judgment. It moves judgment out of per-request negotiation and into rules that can be reviewed, versioned, and revised.
On Kubernetes, Kueue's cohort model already expresses the core mechanics through nominal quota, borrowing, and lending:
apiVersion: kueue.x-k8s.io/v1beta1
kind: ClusterQueue
metadata:
name: research-shared
spec:
cohort: fleet-accelerators
preemption:
reclaimWithinCohort: Any
withinClusterQueue: LowerPriority
resourceGroups:
- coveredResources: ["nvidia.com/gpu"]
flavors:
- name: h100-80gb
resources:
- name: "nvidia.com/gpu"
nominalQuota: 256
borrowingLimit: 512 # may burst into idle cohort capacity
lendingLimit: 192 # keeps a protected floor of 64 devices
Read those preemption lines together with the limits, because the pair is what changes team behavior. lendingLimit guarantees a floor that is never lent out, and reclaimWithinCohort: Any guarantees that lent capacity comes back the moment the owner needs it, by preempting borrowers if necessary. Once returning idle capacity carries no risk of losing it, the economic case for hoarding mostly evaporates.
Beyond quota mechanics, useful policy inputs include production criticality, launch dates, queue age, the historical efficiency of the requesting workload, and per-organization fairness floors. Dominant Resource Fairness (Ghodsi et al., NSDI 2011) is a sensible default for multi-resource sharing, with queue aging layered on top so low-priority work eventually runs instead of starving.
One requirement is non-negotiable: explainability. Every scheduling decision needs a retrievable answer to "why is my job still queued?" naming the binding policy, the preempting workload, and a projected start time. Without that, engineers conclude the system is arbitrary and route around it. An orchestration layer people distrust is worse than static quotas, because now you have a scheduler plus a shadow economy of private allocations.
3. Hardware Broker: One Portfolio, Not Several Fiefdoms
Many organizations run accelerator families as separate silos, each with its own team, toolchain, and queue. The fragmentation hides substitution opportunities and strands capacity on whichever side happens to be idle.
A broker matches substitutable declarations against live supply and makes the tradeoff explicit: 32 A100s in four hours, versus 20 H100s now, versus 64 TPU v5e chips now plus a 12-minute recompile. Comparing those options requires measured throughput per accelerator class per model family, which in turn means maintaining a benchmark corpus and refreshing it as software evolves. Compiler-level changes alone can shift the relative ranking of hardware classes over time; our work on compiler autotuning for TPU fleets (CATWILD, MLSys 2026) found that tuning decisions which were optimal at one point degrade as workloads and toolchains drift, and the same staleness problem applies to any performance model a broker relies on. A stale model will confidently route work to the wrong hardware.
Scope this realistically. Workloads with hand-tuned CUDA kernels, custom collectives, or architecture-specific memory layouts are not portable, and forcing portability on them wastes effort. The goal is narrower: preserve flexibility where it is technically sound, and make sure the jobs that genuinely are portable are visible to the scheduler as such.
4. Continuous Reallocation: Placement Does Not End at Admission
Traditional schedulers decide placement once, at admission, and then stop thinking. AI infrastructure needs a control loop that keeps asking whether each device is still attached to the highest-value feasible work.
Conditions the loop should detect and act on:
- Idle reservations. Allocated devices with sustained near-zero productive compute past a grace threshold get reclaimed.
- Over-provisioned elastic jobs. When scaling efficiency degrades past a threshold, the marginal device is buying almost nothing. Resize down and return the difference.
- Blocked jobs. A job stalled on data loading, or an external dependency for longer than its checkpoint interval, gets checkpointed and requeued rather than left squatting.
- Priority inversions. A launch-critical workload arriving after lower-priority jobs have claimed capacity triggers graceful preemption, per declared policy.
- Substitution openings. A portable queued job starts now on a different accelerator class instead of waiting for its preferred one.
The safety constraints matter as much as the actions. Rate-limit preemptions per team per hour. Never preempt a job whose checkpoint drill has not passed. Cap how many times a single job can be preempted before it gets promoted, or some unlucky workload will never finish. And write every reallocation event into the job's own log with a reason attached. Unexplained preemption is the fastest way to lose organizational support for the whole system.
5. Goodput Accounting: Allocated Is Not Productive
You cannot manage what your metrics obscure, and device-level utilization obscures a lot. A layered view separates the questions:
| Layer | Question | Rough definition |
|---|---|---|
| Allocation rate | Is the fleet handed out? | Allocated device-seconds / total device-seconds |
| Occupancy | Are allocated devices active? | Active device-seconds / allocated device-seconds |
| Goodput | Is the activity producing committed progress? | Device-seconds contributing to non-discarded work / allocated device-seconds |
| MFU | How close is the work to the hardware roofline? | Achieved model FLOPs / peak FLOPs |
Goodput is the layer that changes behavior. Computing it means joining hardware telemetry with job-level progress signals such as steps committed, checkpoints written, and requests served, so that restart replay, discarded partial work, and stalled collectives count as loss rather than as utilization. A training job that crashes and replays six hours from a stale checkpoint looks fine by occupancy and terrible by goodput, and goodput is telling the truth.
Model FLOPs utilization, introduced in the PaLM paper (Chowdhery et al., 2022), is worth tracking but answers a different question. MFU is a compiler and kernel efficiency signal, not an allocation signal. A job can be perfectly scheduled and still leave performance on the floor; that is a different team's bug.
The Fairness Constraint
A system that optimizes purely for near-term commercial value fails predictably. Exploratory research never runs, because its value is uncertain by definition, and small teams lose every contest against large ones.
The safeguards are concrete: a protected research pool that revenue workloads cannot preempt, queue aging that raises effective priority over time, per-team caps within each cohort, and reservation expiry so held capacity gets re-justified periodically. All of these cost measured efficiency, and that cost should be budgeted deliberately. The alternative is an infrastructure layer that quietly defunds the products you have not built yet.
A 90-Day Rollout
Days 0 to 15: instrument before changing anything. Inventory accelerator types, ownership boundaries, queue mechanisms, and approval paths. Separate reservation time from productive execution time in telemetry, and publish the gap. If you cannot produce a defensible goodput number, everything downstream is guesswork. Expect the first honest measurement to be politically uncomfortable.
Days 16 to 45: write the allocation contract. Define priority classes, fairness floors, preemption rules, reservation expiry, and substitution policies, then circulate them with engineering, infrastructure, and product leadership before automating anything. Decide which workloads carry hard guarantees and which are explicitly flexible. This phase is mostly negotiation rather than implementation, and skipping it is the most common way these projects die.
Days 46 to 90: automate one bounded pool. Choose a cohort with a contained blast radius. Batch research capacity is usually right, since it tolerates interruption. Enable reclamation, resizing, and substitution there only, and track queue time, end-to-end completion time, goodput, failed migrations, and policy exceptions.
Then hold. Do not expand scope until teams trust both the decisions and the explanations behind them. Trust, not scheduler throughput, is the binding constraint on adoption.
Closing
Accelerator demand is not going to flatten. Models grow, inference spreads into more product surfaces, and experiment volume rises with headcount. Newer hardware improves per-device throughput while doing nothing about contention; if anything, it raises the stakes, since a stranded H100 is a more expensive mistake than the stranded V100 it replaced.
Organizations that scale well will stop treating accelerators as assets assigned to teams and start managing them as a continuously rebalanced portfolio governed by explicit, explainable policy. Procurement still matters. Orchestration determines how much of what you procured turns into shipped product.
The most valuable accelerator in your fleet is not the next one you buy. It is the one you already own, currently attached to the wrong job.
Opinions expressed by DZone contributors are their own.
Comments