DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

Related

  • The AI Co-Pilot: How to Lead When Your Team's Best Player Is a Machine
  • How To Boost Your Software Engineer Career: Code and Life
  • Resume the Evaluation, Not the Entire Batch: Build a Checkpoint-Aware AI Job Controller With Temporal
  • An Enterprise AI Governance Checklist for Software Teams

Trending

  • Building an AI-Ready Data Layer Without Rebuilding the Enterprise
  • OpenSearch Heap Sizing: Swap, Page Cache, and the 50% Rule
  • Embabel vs LangGraph4j: Two Agentic Philosophies for Investment and Risk Analysis in BFSI
  • From Wild West to Context-Driven Engineering: How Culture Shapes Software Decisions
  1. DZone
  2. Data Engineering
  3. AI/ML
  4. Why Your GPU Fleet Is Both Full and Idle: Building a Capacity Orchestration Layer

Why Your GPU Fleet Is Both Full and Idle: Building a Capacity Orchestration Layer

Managing GPU capacity manually produces fleets that are simultaneously fully allocated and poorly utilized. A capacity orchestration layer can fix the problem.

By 
Ankit Sinha user avatar
Ankit Sinha
·
Oct. 09, 26 · Opinion
Likes (0)
Comment
Save
Tweet
Share
131 Views

Join the DZone community and get the full member experience.

Join For Free

Most machine learning (ML) platforms still hand out accelerators through static quotas, spreadsheet-driven approvals, and hallway negotiation. That model was tolerable when capacity was cheap, and demand was steady. Neither holds for modern AI workloads. This article walks through why static allocation breaks down, the five components of a capacity orchestration layer (intent capture, policy, hardware brokering, continuous reallocation, and goodput accounting), and a 90-day rollout plan that won't destroy your users' trust in the scheduler.

The Contradiction You Have Probably Already Measured

If you run infrastructure for a large ML organization, you have likely seen a dashboard that says something uncomfortable. The fleet is 90-plus percent allocated. Teams are filing tickets about blocked capacity. And a sampled profile of the same fleet shows a meaningful fraction of devices doing no useful work.

Buying more accelerators does not resolve this. It scales the inefficiency. The contradiction comes from a specific architectural choice: treating an accelerator reservation as ownership rather than as a lease with conditions.

The industry has spent years optimizing acquisition: supply agreements, multi-year commitments, reserved cloud capacity. Comparatively little engineering has gone into what happens after the hardware lands in the fleet.

Why Static Quotas Fail for ML Specifically

Conventional service capacity planning works because service demand is autocorrelated. Today's traffic predicts tomorrow's. ML workload demand does not behave this way, for four reasons.

Demand is event-driven. An eval result at 2 p.m. can turn a 16-device research project into a 512-device training run by the next morning. No quarterly quota cycle accommodates that.

Jobs differ in shape, not just size. A gang-scheduled training job with all-reduce dependencies across 64 devices and a stateless inference replica are both "GPU requests" to a naive scheduler, but they have incompatible placement, topology, and preemption semantics. A fleet fragmented without topology awareness ends up with allocatable devices that no large job can actually use.

Reservation and consumption diverge. A job holding 32 devices while blocked on an input pipeline is occupying capacity and producing nothing. Device-level utilization metrics will not reliably catch this, because a device spinning on a synchronization barrier can still report activity.

Returning capacity is asymmetrically risky. This one dominates. If reacquiring capacity takes two weeks of negotiation, the rational move for every team is to hold everything they might need at peak, indefinitely. Static allocation rewards hoarding. Each over-request is a locally correct decision inside a badly designed system, which is why exhortations to "be good citizens" never fix it.

The Architecture

A capacity orchestration layer sits between your workload submission API and the underlying scheduler, whether that is Kubernetes with Kueue or Volcano, Slurm, Ray, or an internal system. It does not replace the scheduler. It feeds the scheduler better information and revisits its decisions after admission.

1. Intent Capture: A Device Count Is an Answer, Not a Requirement

When a user requests "32 A100s," they have already collapsed their actual requirement into one number. The scheduler needs the requirement itself: what the workload can substitute, what it can tolerate, and what it cannot give up. In practice that looks like a spec along these lines:

YAML
 
apiVersion: capacity.internal/v1
kind: WorkloadIntent
metadata:
  name: ranking-hparam-sweep
spec:
  class: research-batch          # drives policy, not placement
  preferred:
    accelerator: a100-80gb
    count: 32
  substitutable:
    - accelerator: h100-80gb
      count: 20                  # based on measured step-time on this model
    - accelerator: tpu-v5e
      count: 64
      requiresRecompile: true    # XLA path validated in CI
  topology:
    gangScheduled: true
    minInterconnect: nvlink-domain
  elasticity:
    minReplicas: 8
    maxReplicas: 32
    resizeCost: 90s              # checkpoint + reshard
  preemption:
    tolerated: true
    gracePeriod: 120s
    checkpointInterval: 600s
  deadline: "2026-08-14T00:00:00Z"


Two fields carry most of the value. substitutable is what makes cross-architecture brokering possible. preemption.tolerated, paired with a real checkpointInterval, is what makes a job safely reclaimable.

Interruption tolerance should be verified, not self-declared. Run a preemption drill in CI: kill the job, confirm it resumes from checkpoint within the stated grace period. A self-reported tolerance that has never been exercised will eventually be exercised in production, on a bad day.

Be honest about the engineering cost. Portable workload definitions, container images built for multiple backends, and checkpoint formats that survive a device-count change are all real work. Elastic resharding in particular is hard for models with complex parallelism strategies. Treat the full intent spec as something teams grow into. A job that declares only preferred still schedules; it just gets fewer options.

2. Policy Engine: Encode the Argument Once

The point of a policy engine is not to remove human judgment. It moves judgment out of per-request negotiation and into rules that can be reviewed, versioned, and revised.

On Kubernetes, Kueue's cohort model already expresses the core mechanics through nominal quota, borrowing, and lending:

YAML
 
apiVersion: kueue.x-k8s.io/v1beta1
kind: ClusterQueue
metadata:
  name: research-shared
spec:
  cohort: fleet-accelerators
  preemption:
    reclaimWithinCohort: Any
    withinClusterQueue: LowerPriority
  resourceGroups:
    - coveredResources: ["nvidia.com/gpu"]
      flavors:
        - name: h100-80gb
          resources:
            - name: "nvidia.com/gpu"
              nominalQuota: 256
              borrowingLimit: 512   # may burst into idle cohort capacity
              lendingLimit: 192     # keeps a protected floor of 64 devices


Read those preemption lines together with the limits, because the pair is what changes team behavior. lendingLimit guarantees a floor that is never lent out, and reclaimWithinCohort: Any guarantees that lent capacity comes back the moment the owner needs it, by preempting borrowers if necessary. Once returning idle capacity carries no risk of losing it, the economic case for hoarding mostly evaporates.

Beyond quota mechanics, useful policy inputs include production criticality, launch dates, queue age, the historical efficiency of the requesting workload, and per-organization fairness floors. Dominant Resource Fairness (Ghodsi et al., NSDI 2011) is a sensible default for multi-resource sharing, with queue aging layered on top so low-priority work eventually runs instead of starving.

One requirement is non-negotiable: explainability. Every scheduling decision needs a retrievable answer to "why is my job still queued?" naming the binding policy, the preempting workload, and a projected start time. Without that, engineers conclude the system is arbitrary and route around it. An orchestration layer people distrust is worse than static quotas, because now you have a scheduler plus a shadow economy of private allocations.

3. Hardware Broker: One Portfolio, Not Several Fiefdoms

Many organizations run accelerator families as separate silos, each with its own team, toolchain, and queue. The fragmentation hides substitution opportunities and strands capacity on whichever side happens to be idle.

A broker matches substitutable declarations against live supply and makes the tradeoff explicit: 32 A100s in four hours, versus 20 H100s now, versus 64 TPU v5e chips now plus a 12-minute recompile. Comparing those options requires measured throughput per accelerator class per model family, which in turn means maintaining a benchmark corpus and refreshing it as software evolves. Compiler-level changes alone can shift the relative ranking of hardware classes over time; our work on compiler autotuning for TPU fleets (CATWILD, MLSys 2026) found that tuning decisions which were optimal at one point degrade as workloads and toolchains drift, and the same staleness problem applies to any performance model a broker relies on. A stale model will confidently route work to the wrong hardware.

Scope this realistically. Workloads with hand-tuned CUDA kernels, custom collectives, or architecture-specific memory layouts are not portable, and forcing portability on them wastes effort. The goal is narrower: preserve flexibility where it is technically sound, and make sure the jobs that genuinely are portable are visible to the scheduler as such.

4. Continuous Reallocation: Placement Does Not End at Admission

Traditional schedulers decide placement once, at admission, and then stop thinking. AI infrastructure needs a control loop that keeps asking whether each device is still attached to the highest-value feasible work.

Conditions the loop should detect and act on:

  • Idle reservations. Allocated devices with sustained near-zero productive compute past a grace threshold get reclaimed.
  • Over-provisioned elastic jobs. When scaling efficiency degrades past a threshold, the marginal device is buying almost nothing. Resize down and return the difference.
  • Blocked jobs. A job stalled on data loading, or an external dependency for longer than its checkpoint interval, gets checkpointed and requeued rather than left squatting.
  • Priority inversions. A launch-critical workload arriving after lower-priority jobs have claimed capacity triggers graceful preemption, per declared policy.
  • Substitution openings. A portable queued job starts now on a different accelerator class instead of waiting for its preferred one.

The safety constraints matter as much as the actions. Rate-limit preemptions per team per hour. Never preempt a job whose checkpoint drill has not passed. Cap how many times a single job can be preempted before it gets promoted, or some unlucky workload will never finish. And write every reallocation event into the job's own log with a reason attached. Unexplained preemption is the fastest way to lose organizational support for the whole system.

5. Goodput Accounting: Allocated Is Not Productive

You cannot manage what your metrics obscure, and device-level utilization obscures a lot. A layered view separates the questions:

Layer Question Rough definition
Allocation rate Is the fleet handed out? Allocated device-seconds / total device-seconds
Occupancy Are allocated devices active? Active device-seconds / allocated device-seconds
Goodput Is the activity producing committed progress? Device-seconds contributing to non-discarded work / allocated device-seconds
MFU How close is the work to the hardware roofline? Achieved model FLOPs / peak FLOPs

Goodput is the layer that changes behavior. Computing it means joining hardware telemetry with job-level progress signals such as steps committed, checkpoints written, and requests served, so that restart replay, discarded partial work, and stalled collectives count as loss rather than as utilization. A training job that crashes and replays six hours from a stale checkpoint looks fine by occupancy and terrible by goodput, and goodput is telling the truth.

Model FLOPs utilization, introduced in the PaLM paper (Chowdhery et al., 2022), is worth tracking but answers a different question. MFU is a compiler and kernel efficiency signal, not an allocation signal. A job can be perfectly scheduled and still leave performance on the floor; that is a different team's bug.

The Fairness Constraint

A system that optimizes purely for near-term commercial value fails predictably. Exploratory research never runs, because its value is uncertain by definition, and small teams lose every contest against large ones.

The safeguards are concrete: a protected research pool that revenue workloads cannot preempt, queue aging that raises effective priority over time, per-team caps within each cohort, and reservation expiry so held capacity gets re-justified periodically. All of these cost measured efficiency, and that cost should be budgeted deliberately. The alternative is an infrastructure layer that quietly defunds the products you have not built yet.

A 90-Day Rollout

Days 0 to 15: instrument before changing anything. Inventory accelerator types, ownership boundaries, queue mechanisms, and approval paths. Separate reservation time from productive execution time in telemetry, and publish the gap. If you cannot produce a defensible goodput number, everything downstream is guesswork. Expect the first honest measurement to be politically uncomfortable.

Days 16 to 45: write the allocation contract. Define priority classes, fairness floors, preemption rules, reservation expiry, and substitution policies, then circulate them with engineering, infrastructure, and product leadership before automating anything. Decide which workloads carry hard guarantees and which are explicitly flexible. This phase is mostly negotiation rather than implementation, and skipping it is the most common way these projects die.

Days 46 to 90: automate one bounded pool. Choose a cohort with a contained blast radius. Batch research capacity is usually right, since it tolerates interruption. Enable reclamation, resizing, and substitution there only, and track queue time, end-to-end completion time, goodput, failed migrations, and policy exceptions.

Then hold. Do not expand scope until teams trust both the decisions and the explanations behind them. Trust, not scheduler throughput, is the binding constraint on adoption.

Closing

Accelerator demand is not going to flatten. Models grow, inference spreads into more product surfaces, and experiment volume rises with headcount. Newer hardware improves per-device throughput while doing nothing about contention; if anything, it raises the stakes, since a stranded H100 is a more expensive mistake than the stranded V100 it replaced.

Organizations that scale well will stop treating accelerators as assets assigned to teams and start managing them as a continuously rebalanced portfolio governed by explicit, explainable policy. Procurement still matters. Orchestration determines how much of what you procured turns into shipped product.

The most valuable accelerator in your fleet is not the next one you buy. It is the one you already own, currently attached to the wrong job.

AI FLOPS Requirements engineering YAML Accelerator (software) Bad Day (viral video) career Checkpoint (pinball) IDLE (Python) teams

Opinions expressed by DZone contributors are their own.

Related

  • The AI Co-Pilot: How to Lead When Your Team's Best Player Is a Machine
  • How To Boost Your Software Engineer Career: Code and Life
  • Resume the Evaluation, Not the Entire Batch: Build a Checkpoint-Aware AI Job Controller With Temporal
  • An Enterprise AI Governance Checklist for Software Teams

Partner Resources

×

Comments

The likes didn't load as expected. Please refresh the page and try again.

  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook