Declarative Scheduling
The other pages in this section show how to hand-write Workload and PodGroup objects alongside a JobSet, and point pods at a PodGroup via schedulingGroup.podGroupName. JobSet also supports a declarative integration: you describe the scheduling behavior you want directly on the JobSet with spec.scheduling, and the JobSet controller creates, updates, and deletes the matching Workload/PodGroup objects for you.
This page walks through that spec.scheduling field using the example manifests in site/static/examples/scheduling, one use case at a time.
Prerequisites
This feature is alpha and off by default, gated by two independent switches:
The JobSet controller-manager feature gate
JobSetWorkloadAwareSchedulingAPI, enabled via the manager’sConfiguration(thejobset-manager-configConfigMap consumed through--config):apiVersion: config.jobset.x-k8s.io/v1alpha1 kind: Configuration featureGates: JobSetWorkloadAwareSchedulingAPI: trueThe Kubernetes cluster’s own WAS feature gates and runtime-config, same as the general Workload Aware Scheduling prerequisites, plus two additional gates needed by some of the use cases below:
GenericWorkloadWorkloadWithJobTopologyAwareWorkloadScheduling— required for the topology-constrained examplesDRAWorkloadResourceClaims— required for the shared DRA claim example- API server
--runtime-config=scheduling.k8s.io/v1alpha3=true,scheduling.k8s.io/v1beta1=true
The local and E2E Kind setup uses Kubernetes
v1.37.0release binaries, which include the requiredscheduling.k8s.ioAPIs. Seehack/kind-config-scheduling.yamlandhack/e2e-scheduling-cluster.shfor the cluster setup, or runmake kind-cluster-schedulingto create one locally.
How It Works
spec.scheduling is optional and immutable once a JobSet is created. If it is left unset, nothing changes: no Workload or PodGroup objects are created, and existing JobSets are unaffected. Setting it — even to an empty {} — tells the controller to compile exactly one Workload, owned by the JobSet, containing one or more PodGroupTemplates, and to keep matching PodGroup objects in sync with it.
spec.scheduling supports two mutually exclusive models — a JobSet must use exactly one, since composite Gang-of-Gangs PodGroup hierarchies linking a parent PodGroup to leaf PodGroups aren’t implemented in alpha:
schedulingPolicy,schedulingConstraints, anddisruptionModeat the top level ofspec.schedulingconfigure a single composite PodGroup (or, under sequenced startup, one PodGroup perReplicatedJob) covering the whole JobSet. LeavereplicatedJobsunset when using this model.replicatedJobslets you target one or moreReplicatedJobs by name with their own leaf-level policy, producing onePodGroupper policy entry. EveryReplicatedJobin the JobSet must be targeted by exactly one entry, since there’s no top-level policy for an untargetedReplicatedJobto fall back to. Leave the top-levelschedulingPolicy,schedulingConstraints, anddisruptionModeunset when using this model.jobnested inside areplicatedJobsentry goes one level deeper still, giving each Job replica of aReplicatedJobits ownPodGroup(“gang-of-gangs”).
Child Jobs are annotated with scheduling.k8s.io/group-template-name so you can trace which PodGroupTemplate each Job belongs to.
Use Case: Whole-JobSet Gang Scheduling
The most common pattern: every pod across every ReplicatedJob must be schedulable before any of them are admitted. Set a top-level gang policy and no replicatedJobs — the controller creates a single PodGroup sized to every pod in the JobSet.
# Gang Scheduling Example
#
# All pods across driver and workers are gang-scheduled: they must be admitted
# together atomically or none at all. This is the most common pattern for
# distributed ML training.
#
# Expected results:
# - One Workload is created with 1 PodGroupTemplate
# (top-level gang with no per-RJ overrides → single PodGroup)
# - A single PodGroup with Gang policy (computed minCount)
# - All child Jobs carry the scheduling.k8s.io/group-template-name annotation
# - The Workload has an OwnerReference pointing to the JobSet
#
# Verify (generated Workload and PodGroup names include identity hashes):
# kubectl get workloads,podgroups,jobs -n default
# kubectl describe workload -n default -l jobset.sigs.k8s.io/jobset-name=gang-training
# kubectl get jobs -n default -o jsonpath='{range .items[*]}{.metadata.name}: {.metadata.annotations}{"\n"}{end}'
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: gang-training
namespace: default
spec:
successPolicy:
operator: All
network:
enableDNSHostnames: true
scheduling:
schedulingPolicy:
gang: {}
disruptionMode:
all: {}
replicatedJobs:
- name: driver
replicas: 1
template:
spec:
parallelism: 1
completions: 1
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: driver
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "30s"]
resources:
requests:
cpu: 100m
memory: 64Mi
- name: workers
replicas: 2
template:
spec:
parallelism: 2
completions: 2
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: worker
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "30s"]
resources:
requests:
cpu: 100m
memory: 64Mi
A single-ReplicatedJob variant of the same idea:
# Gang Scheduling Example
#
# A single ReplicatedJob ("workers", 2 replicas x parallelism 4) is
# gang-scheduled: all 8 pods must be admitted together atomically or none at
# all (one PodGroup, computed minCount=8).
#
# Verify (generated Workload and PodGroup names include identity hashes):
# kubectl get workloads,podgroups,jobs -n default
# kubectl describe workload -n default -l jobset.sigs.k8s.io/jobset-name=single-gang
# kubectl get jobs -n default -o jsonpath='{range .items[*]}{.metadata.name}: {.metadata.annotations}{"\n"}{end}'
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: single-gang
namespace: default
spec:
successPolicy:
operator: All
network:
enableDNSHostnames: true
scheduling:
schedulingPolicy:
gang: {}
replicatedJobs:
- name: workers
replicas: 2
template:
spec:
parallelism: 4
completions: 4
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: worker
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "30s"]
resources:
requests:
cpu: 100m
memory: 64Mi
kubectl get workloads,podgroups,jobs -n default
kubectl describe workload -n default -l jobset.sigs.k8s.io/jobset-name=gang-training
Use Case: No Scheduling (Backward Compatibility)
Existing JobSets that don’t set spec.scheduling must keep working exactly as before: no Workload, no PodGroup, no scheduling annotations on child Jobs.
# No-Scheduling Baseline
#
# This JobSet does NOT set spec.scheduling. It verifies backward compatibility:
# existing JobSets without scheduling config must continue to work identically,
# producing zero Workload or PodGroup objects.
#
# Expected results:
# - Jobs are created normally
# - NO Workload objects exist for this JobSet
# - NO PodGroup objects exist for this JobSet
# - Child Jobs do NOT carry scheduling.k8s.io annotations
#
# Verify:
# kubectl get jobs -n default -l jobset.sigs.k8s.io/jobset-name=no-sched-baseline
# kubectl get workloads -n default # should NOT contain no-sched-baseline
# kubectl get podgroups -n default # should NOT contain no-sched-baseline
# kubectl get jobs -n default -o jsonpath='{range .items[*]}{.metadata.name}: {.metadata.annotations}{"\n"}{end}'
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: no-sched-baseline
namespace: default
spec:
successPolicy:
operator: All
network:
enableDNSHostnames: true
# No spec.scheduling — this is the baseline.
replicatedJobs:
- name: workers
replicas: 2
template:
spec:
parallelism: 1
completions: 1
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: worker
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "15s"]
resources:
requests:
cpu: 100m
memory: 64Mi
kubectl get workloads -n default # should not list no-sched-baseline
kubectl get podgroups -n default # should not list no-sched-baseline
Use Case: Independent PodGroups per ReplicatedJob
Sometimes different ReplicatedJobs should be scheduled independently of one another rather than as one giant gang — for example, a driver that can start on its own while workers are gang-scheduled among themselves. Give each ReplicatedJob its own entry in replicatedJobs, and each gets its own PodGroup, sized from its own parallelism × replicas.
# KEP-969 Representative API: "Schedule groups independently"
#
# scheduling:
# replicatedJobs:
# - targetReplicatedJobs: [launcher]
# schedulingPolicy: {gang: {}}
# - targetReplicatedJobs: [worker]
# schedulingPolicy: {gang: {}}
#
# Expected: one Workload containing one identity-hashed PodGroup per targeted
# ReplicatedJob: one for launcher and one for worker. Each PodGroup's minCount
# is independently computed from its own
# ReplicatedJob's parallelism x replicas (launcher: 1, worker: 4).
#
# Verify:
# kubectl get workloads,podgroups,jobs -n default
# kubectl get podgroups -n default -o custom-columns=NAME:.metadata.name,MINCOUNT:.spec.schedulingPolicy.gang.minCount
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: per-rj-independent
namespace: default
spec:
successPolicy:
operator: All
network:
enableDNSHostnames: true
scheduling:
replicatedJobs:
- targetReplicatedJobs: ["launcher"]
schedulingPolicy:
gang: {}
- targetReplicatedJobs: ["worker"]
schedulingPolicy:
gang: {}
replicatedJobs:
- name: launcher
replicas: 1
template:
spec:
parallelism: 1
completions: 1
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: launcher
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "30s"]
resources:
requests:
cpu: 100m
memory: 64Mi
- name: worker
replicas: 2
template:
spec:
parallelism: 2
completions: 2
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: worker
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "30s"]
resources:
requests:
cpu: 100m
memory: 64Mi
kubectl get podgroups -n default -o custom-columns=NAME:.metadata.name,MINCOUNT:.spec.schedulingPolicy.gang.minCount
Use Case: One Shared PodGroup Across Multiple ReplicatedJobs
The opposite grouping: multiple ReplicatedJobs that must be admitted together as a single gang. List them all in one targetReplicatedJobs entry, and the controller creates one PodGroup whose minCount is the sum of both ReplicatedJobs’ pod counts.
# KEP-969 Representative API: "Schedule groups together"
#
# scheduling:
# replicatedJobs:
# - targetReplicatedJobs: [launcher, worker]
# schedulingPolicy: {gang: {}}
#
# Expected: one Workload containing one PodGroup shared by launcher and worker.
# minCount is the sum of the pod counts of both targeted ReplicatedJobs
# (launcher: 1*1=1, worker: 2*2=4 -> total 5).
#
# Verify (generated Workload and PodGroup names include identity hashes):
# kubectl get workloads,podgroups,jobs -n default
# kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=per-rj-grouped -o jsonpath='{.items[0].spec.schedulingPolicy.gang.minCount}{"\n"}'
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: per-rj-grouped
namespace: default
spec:
successPolicy:
operator: All
network:
enableDNSHostnames: true
scheduling:
replicatedJobs:
- targetReplicatedJobs: ["launcher", "worker"]
schedulingPolicy:
gang: {}
disruptionMode:
all: {}
replicatedJobs:
- name: launcher
replicas: 1
template:
spec:
parallelism: 1
completions: 1
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: launcher
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "30s"]
resources:
requests:
cpu: 100m
memory: 64Mi
- name: worker
replicas: 2
template:
spec:
parallelism: 2
completions: 2
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: worker
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "30s"]
resources:
requests:
cpu: 100m
memory: 64Mi
kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=per-rj-grouped -o jsonpath='{.items[0].spec.schedulingPolicy.gang.minCount}{"\n"}'
Use Case: Gang-of-Gangs — Independent PodGroup per Job Replica
job goes one level finer than replicatedJobs: instead of one PodGroup covering every replica of a ReplicatedJob, each Job replica gets its own PodGroup, sized to just that Job’s own parallelism. This suits workloads where each replica is an independently-schedulable gang, such as per-replica launcher/worker sets.
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: per-ij-job-independent
namespace: default
spec:
successPolicy:
operator: All
network:
enableDNSHostnames: true
scheduling:
replicatedJobs:
- targetReplicatedJobs: ["launcher"]
schedulingPolicy:
gang: {}
- targetReplicatedJobs: ["worker"]
job:
schedulingPolicy:
gang: {}
disruptionMode:
all: {}
replicatedJobs:
- name: launcher
replicas: 1
template:
spec:
parallelism: 1
completions: 1
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: launcher
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "30s"]
resources:
requests:
cpu: 100m
memory: 64Mi
- name: worker
replicas: 2
template:
spec:
parallelism: 2
completions: 2
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: worker
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "30s"]
resources:
requests:
cpu: 100m
memory: 64Mi
kubectl get workloads,podgroups,jobs -n default
Use Case: Topology-Aware Placement
Combine gang scheduling with schedulingConstraints.topology to require that all pods in a gang land in the same topology domain — for example, co-locating pods on the same rack for RDMA-sensitive training, or coordinating placement across TPU slices. topology-constrained.yaml below uses replicatedJobs to give the driver an independent basic policy while workers get a gang policy with the rack constraint — every ReplicatedJob is targeted by its own entry since no top-level schedulingPolicy is set.
# Topology-Constrained Workers with Independent Driver
#
# The driver uses a Basic scheduling policy (schedules independently).
# Workers use an explicit Gang scheduling policy with topology
# constraints requiring co-location on the same rack. This is ideal for
# RDMA-based distributed training where network locality matters.
#
# The top-level schedulingPolicy/schedulingConstraints/disruptionMode API and
# replicatedJobs are mutually exclusive (there is no composite
# PodGroup linking leaf PodGroups together in alpha), so this uses
# replicatedJobs only: every ReplicatedJob is targeted by its own
# entry, each with an explicit leaf-level schedulingPolicy.
#
# Requires the TopologyAwareWorkloadScheduling feature gate on the apiserver;
# without it, the topology constraint is silently dropped on write and the
# JobSet controller will continuously delete/recreate the PodGroup trying to
# reconcile the drift (see hack/kind-config-scheduling.yaml).
#
# Expected results:
# - One Workload with 2 PodGroupTemplates
# - One driver PodGroup with Basic policy (no Gang)
# - One workers PodGroup with Gang policy (computed minCount=4) and a
# topology constraint on "topology.kubernetes.io/rack"
# - All child Jobs carry scheduling annotations
#
# Verify (generated Workload and PodGroup names include identity hashes):
# kubectl get workloads,podgroups -n default -l jobset.sigs.k8s.io/jobset-name=topo-training
# kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=topo-training -o yaml
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: topo-training
namespace: default
spec:
successPolicy:
operator: All
network:
enableDNSHostnames: true
scheduling:
# Per-ReplicatedJob policies. Every ReplicatedJob must be targeted by
# exactly one entry, since no top-level schedulingPolicy is set.
replicatedJobs:
- targetReplicatedJobs: ["driver"]
# Driver schedules independently; no gang requirement.
schedulingPolicy:
basic: {}
- targetReplicatedJobs: ["workers"]
# Workers are gang-scheduled with rack constraints.
schedulingPolicy:
gang: {}
schedulingConstraints:
topology:
- key: topology.kubernetes.io/rack
replicatedJobs:
- name: driver
replicas: 1
template:
spec:
parallelism: 1
completions: 1
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: driver
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "30s"]
resources:
requests:
cpu: 100m
memory: 64Mi
- name: workers
replicas: 2
template:
spec:
parallelism: 2
completions: 2
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: worker
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "30s"]
resources:
requests:
cpu: 100m
memory: 64Mi
# KEP-969 Representative API: "Run a TPU multi-slice workload"
#
# scheduling:
# schedulingPolicy:
# gang: {}
# schedulingConstraints:
# topology: ...
#
# Expected: one Workload and one JobSet-wide PodGroup with the requested
# topology constraint, coordinating placement across TPU slices. TPU-specific
# pod resources/topology values stay in the JobSet pod templates (not shown
# here in full since this example doesn't request real TPU resources).
#
# Requires the TopologyAwareWorkloadScheduling feature gate on the apiserver;
# see the note in topology-constrained.yaml for what happens without it.
#
# Verify (generated Workload and PodGroup names include identity hashes):
# kubectl get workloads,podgroups,jobs -n default
# kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=tpu-multislice -o jsonpath='{.items[0].spec.schedulingPolicy}{" "}{.items[0].spec.schedulingConstraints}{"\n"}'
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: tpu-multislice
namespace: default
spec:
successPolicy:
operator: All
network:
enableDNSHostnames: true
scheduling:
schedulingPolicy:
gang: {}
disruptionMode:
all: {}
schedulingConstraints:
topology:
- key: cloud.google.com/gke-tpu-slice
replicatedJobs:
- name: slice
replicas: 2
template:
spec:
parallelism: 2
completions: 2
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: worker
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "15s"]
resources:
requests:
cpu: 100m
memory: 64Mi
Both examples require the TopologyAwareWorkloadScheduling feature gate; without it the constraint is silently dropped on write and the controller will continuously try to reconcile the resulting drift.
kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=topo-training -o yaml
Use Case: Sequenced Startup
When ReplicatedJobs use dependsOn (or an InOrder StartupPolicy) to create Jobs sequentially, not all pods exist at the same time — so a single JobSet-wide gang could never be satisfied. An explicit top-level schedulingPolicy.gang is therefore rejected. Leave schedulingPolicy unset (for example, scheduling: {}) to use the default Gang policy independently for each ReplicatedJob, or configure Gang policies explicitly under scheduling.replicatedJobs, targeting each ReplicatedJob separately.
# KEP-969 Representative API: "Combine sequencing with Gang scheduling"
#
# scheduling: {}
# replicatedJobs:
# - name: leader
# dependsOn: []
# - name: worker
# dependsOn:
# - name: leader
# status: Ready
#
# Expected: because DependsOn makes Job creation sequential, the controller
# uses one identity-hashed Gang PodGroup per ReplicatedJob instead of a single
# JobSet-wide PodGroup, avoiding
# deadlock (a single PodGroup requiring every pod at once could never be
# satisfied while worker's Job doesn't exist yet). Each PodGroup's minCount is
# computed independently from its own ReplicatedJob (leader: 1, worker: 2);
# an explicit top-level schedulingPolicy.gang is rejected with DependsOn or
# InOrder startup. Leave schedulingPolicy unset to use per-ReplicatedJob defaults.
#
# Verify:
# kubectl get workloads,podgroups,jobs -n default
# kubectl get jobs -n default # worker's Job appears only after leader is Ready
# kubectl get podgroups -n default -o custom-columns=NAME:.metadata.name,MINCOUNT:.spec.schedulingPolicy.gang.minCount
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: sequenced-gang
namespace: default
spec:
successPolicy:
operator: All
network:
enableDNSHostnames: true
scheduling: {}
replicatedJobs:
- name: leader
replicas: 1
template:
spec:
parallelism: 1
completions: 1
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: leader
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "10s"]
resources:
requests:
cpu: 100m
memory: 64Mi
- name: worker
replicas: 1
dependsOn:
- name: leader
status: Ready
template:
spec:
parallelism: 2
completions: 2
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: worker
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "10s"]
resources:
requests:
cpu: 100m
memory: 64Mi
kubectl get jobs -n default # worker's Job appears only after leader is Ready
Use Case: Suspend and Resume
Suspending a JobSet (spec.suspend: true) shouldn’t hold a scheduling reservation for pods that don’t exist. While suspended, the controller deletes the Workload/PodGroup; resuming recreates them under the same names.
# KEP-969 Representative API: "Avoid reserving resources for suspended workloads"
#
# spec:
# suspend: true
# scheduling: ...
#
# Expected: the Workload/PodGroup are deleted while suspended, and recreated
# (same names/mapping) when spec.suspend is set back to false.
#
# Verify:
# kubectl apply -f site/static/examples/scheduling/suspend-resume.yaml
# kubectl get workloads,podgroups -n default # created
# kubectl patch jobset gang-suspend -n default --type=merge -p '{"spec":{"suspend":true}}'
# kubectl get workloads,podgroups -n default # gone
# kubectl patch jobset gang-suspend -n default --type=merge -p '{"spec":{"suspend":false}}'
# kubectl get workloads,podgroups -n default # recreated
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: gang-suspend
namespace: default
spec:
suspend: false
successPolicy:
operator: All
network:
enableDNSHostnames: true
scheduling:
schedulingPolicy:
gang: {}
replicatedJobs:
- name: workers
replicas: 1
template:
spec:
parallelism: 2
completions: 2
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: worker
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "120s"]
resources:
requests:
cpu: 100m
memory: 64Mi
kubectl patch jobset gang-suspend -n default --type=merge -p '{"spec":{"suspend":true}}'
kubectl get workloads,podgroups -n default # gone
kubectl patch jobset gang-suspend -n default --type=merge -p '{"spec":{"suspend":false}}'
kubectl get workloads,podgroups -n default # recreated
Use Case: Elastic Scaling
For an ElasticJobSet, scaling a ReplicatedJob’s parallelism/completions patches the corresponding PodGroup’s Gang.minCount in place rather than deleting and recreating the Workload.
# Elastic Gang Scheduling Example
#
# Mirrors the integration test "should patch Gang minCount in place when
# ElasticJobSet changes parallelism" (test/integration/scheduling/scheduling_test.go)
# and the KEP-969 "Scale an elastic workload" representative API.
#
# A single ReplicatedJob is gang-scheduled with parallelism/completions=2.
#
# Expected results (before scaling):
# - One Workload with 1 PodGroupTemplate
# - A single PodGroup with Gang policy, minCount=2
#
# The generated leaf GangSchedulingPolicy.MinCount field is required. Rather
# than persisting a derived value in the immutable spec.scheduling, the
# controller computes it from the live represented pod count whenever it
# compiles the scheduling objects. On every reconcile it recomputes minCount
# and patches the Workload/PodGroup in place when scaling changes the value.
#
# Because spec.scheduling is immutable, scale it by patching only
# spec.replicatedJobs (see elastic-gang-scale-patch.yaml) rather than
# re-applying the whole JobSet with a modified spec.scheduling block.
#
# Verify (generated Workload and PodGroup names include identity hashes):
# kubectl get workloads,podgroups,jobs -n default
# kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=gang-elastic -o jsonpath='{.items[0].spec.schedulingPolicy.gang.minCount}{"\n"}'
# kubectl get workloads -n default -l jobset.sigs.k8s.io/jobset-name=gang-elastic -o jsonpath='{.items[0].metadata.uid}{"\n"}'
#
# To reproduce the scale-up and observe minCount update from 2 to 4 in place:
# kubectl patch jobset gang-elastic -n default --type=json -p \
# '[{"op":"replace","path":"/spec/replicatedJobs/0/template/spec/parallelism","value":4},
# {"op":"replace","path":"/spec/replicatedJobs/0/template/spec/completions","value":4}]'
# (see elastic-gang-scale-patch.yaml for why a plain `kubectl apply` or
# `--type=merge` patch don't work for this.)
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: gang-elastic
namespace: default
spec:
successPolicy:
operator: All
network:
enableDNSHostnames: true
scheduling:
schedulingPolicy:
gang: {}
disruptionMode:
all: {}
replicatedJobs:
- name: workers
replicas: 1
template:
spec:
parallelism: 2
completions: 2
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: worker
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "120s"]
resources:
requests:
cpu: 100m
memory: 64Mi
# Scale patch for elastic-gang.yaml
#
# Apply this after elastic-gang.yaml to scale the "workers" ReplicatedJob's
# parallelism/completions from 2 to 4. See elastic-gang.yaml for how the
# controller keeps the PodGroup's minCount in sync with the live pod count
# across this kind of scale-up.
#
# NOTE: this manifest must keep spec.scheduling byte-for-byte identical to
# elastic-gang.yaml (including disruptionMode). spec.scheduling as a whole is
# immutable (enforced by a CEL rule), so a `kubectl apply` that changes it
# even incidentally (e.g. by omitting a field the live object has) is
# rejected with "spec.scheduling: Invalid value: Value is immutable" even
# though minCount itself is not what's being edited here.
#
# Usage:
# kubectl apply -f site/static/examples/scheduling/elastic-gang.yaml
# kubectl apply -f site/static/examples/scheduling/elastic-gang-scale-patch.yaml
#
# Or, without a full re-apply, patch just the scaled fields directly. Use a
# JSON patch (--type=json) rather than a merge patch (--type=merge): a merge
# patch replaces the whole replicatedJobs[0] list entry and therefore
# requires every required field (including the pod template) to be repeated,
# not just the two fields being changed.
# kubectl patch jobset gang-elastic -n default --type=json -p \
# '[{"op":"replace","path":"/spec/replicatedJobs/0/template/spec/parallelism","value":4},
# {"op":"replace","path":"/spec/replicatedJobs/0/template/spec/completions","value":4}]'
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: gang-elastic
namespace: default
spec:
successPolicy:
operator: All
network:
enableDNSHostnames: true
scheduling:
schedulingPolicy:
gang: {}
disruptionMode:
all: {}
replicatedJobs:
- name: workers
replicas: 1
template:
spec:
parallelism: 4
completions: 4
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
containers:
- name: worker
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "120s"]
resources:
requests:
cpu: 100m
memory: 64Mi
The generated leaf Gang.minCount field is required. Rather than persisting a derived value in the immutable spec.scheduling, the controller computes it from the live represented pod count whenever it compiles the scheduling objects. It recomputes that value on every reconcile and patches the Workload/PodGroup in place when scaling changes it. An explicit non-zero minCount under replicatedJobs[].job.schedulingPolicy.gang remains a user-selected fixed quorum instead.
Because spec.scheduling is immutable, scale by patching only spec.replicatedJobs rather than re-applying a JobSet whose spec.scheduling differs even incidentally (e.g. a missing field) from what’s already stored — that fails with spec.scheduling: Invalid value: Value is immutable, since the API server can’t tell that the rest of spec.scheduling was meant to stay the same.
kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=gang-elastic -o jsonpath='{.items[0].spec.schedulingPolicy.gang.minCount}{"\n"}'
# after applying elastic-gang-scale-patch.yaml, or an equivalent patch to
# spec.replicatedJobs, this reflects the new pod count (2 -> 4)
Use Case: Shared DRA Resource Claims
replicatedJobs[].job.resourceClaims attaches a PodGroup-level DRA ResourceClaimTemplate reference to each Job replica’s PodGroup, so all pods in that Job share a single allocated claim (for example, an NVLink/IMEX channel) instead of each pod claiming its own device.
# KEP-969 Representative API: "Share a DRA claim across a Job's pods"
#
# scheduling:
# replicatedJobs:
# - targetReplicatedJobs: [worker]
# job:
# schedulingPolicy: {gang: {}}
# resourceClaims:
# - name: imex-channel
# resourceClaimTemplateName: imex-channel-template
#
# Expected: the "worker" PodGroup carries the resourceClaims entry. Each pod
# in "worker" references the same claim name + resourceClaimTemplateName in
# its own pod.spec.resourceClaims, which (with the DRAWorkloadResourceClaims
# feature gate enabled) resolves to the single ResourceClaim generated for
# and owned by the PodGroup instead of creating one ResourceClaim per pod.
#
# Requires: DRAWorkloadResourceClaims feature gate enabled on the apiserver,
# and a DRA driver + DeviceClass so the ResourceClaimTemplate is schedulable.
# This example only exercises the JobSet -> PodGroup resourceClaims wiring;
# without a real DRA driver installed the pods may stay Pending on claim
# allocation, but the PodGroup/ResourceClaimTemplate wiring can still be
# verified directly.
#
# Verify:
# kubectl get workloads,podgroups,resourceclaimtemplates -n default
# kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=dra-shared-claim -o jsonpath='{.items[0].spec.resourceClaims}{"\n"}'
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
name: imex-channel-template
namespace: default
spec:
spec:
devices:
requests:
- name: channel
exactly:
deviceClassName: imex-channel-verify-class
---
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: dra-shared-claim
namespace: default
spec:
successPolicy:
operator: All
network:
enableDNSHostnames: true
scheduling:
replicatedJobs:
- targetReplicatedJobs: ["worker"]
job:
schedulingPolicy:
gang: {}
resourceClaims:
- name: imex-channel
resourceClaimTemplateName: imex-channel-template
replicatedJobs:
- name: worker
replicas: 1
template:
spec:
parallelism: 2
completions: 2
completionMode: Indexed
template:
spec:
restartPolicy: OnFailure
resourceClaims:
- name: imex-channel
resourceClaimTemplateName: imex-channel-template
containers:
- name: worker
image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
command: ["sleep", "10s"]
resources:
requests:
cpu: 100m
memory: 64Mi
claims:
- name: imex-channel
This requires the DRAWorkloadResourceClaims feature gate and a real DRA driver/DeviceClass to fully allocate; without one, the PodGroup-to-claim wiring can still be verified directly even if the pods stay Pending.
kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=dra-shared-claim -o jsonpath='{.items[0].spec.resourceClaims}{"\n"}'
Use Case: Preemption
disruptionMode: {all: {}} (set on several examples above, such as gang-scheduling.yaml) tells the scheduler to preempt the entire gang together rather than individual pods, so a higher-priority JobSet doesn’t leave a lower-priority one half-evicted. Pair it with a PriorityClass on the pod template.
Feedback
Was this page helpful?
Glad to hear it! Please tell us how we can improve.
Sorry to hear that. Please tell us how we can improve.