Declarative Scheduling

Configure Workload-Aware Scheduling declaratively through the JobSet spec, with JobSet managing the Workload and PodGroup objects for you

The other pages in this section show how to hand-write Workload and PodGroup objects alongside a JobSet, and point pods at a PodGroup via schedulingGroup.podGroupName. JobSet also supports a declarative integration: you describe the scheduling behavior you want directly on the JobSet with spec.scheduling, and the JobSet controller creates, updates, and deletes the matching Workload/PodGroup objects for you.

This page walks through that spec.scheduling field using the example manifests in site/static/examples/scheduling, one use case at a time.

Prerequisites

This feature is alpha and off by default, gated by two independent switches:

  1. The JobSet controller-manager feature gate JobSetWorkloadAwareSchedulingAPI, enabled via the manager’s Configuration (the jobset-manager-config ConfigMap consumed through --config):

    apiVersion: config.jobset.x-k8s.io/v1alpha1
    kind: Configuration
    featureGates:
      JobSetWorkloadAwareSchedulingAPI: true
    
  2. The Kubernetes cluster’s own WAS feature gates and runtime-config, same as the general Workload Aware Scheduling prerequisites, plus two additional gates needed by some of the use cases below:

    • GenericWorkload
    • WorkloadWithJob
    • TopologyAwareWorkloadScheduling — required for the topology-constrained examples
    • DRAWorkloadResourceClaims — required for the shared DRA claim example
    • API server --runtime-config=scheduling.k8s.io/v1alpha3=true,scheduling.k8s.io/v1beta1=true

    The local and E2E Kind setup uses Kubernetes v1.37.0 release binaries, which include the required scheduling.k8s.io APIs. See hack/kind-config-scheduling.yaml and hack/e2e-scheduling-cluster.sh for the cluster setup, or run make kind-cluster-scheduling to create one locally.

How It Works

spec.scheduling is optional and immutable once a JobSet is created. If it is left unset, nothing changes: no Workload or PodGroup objects are created, and existing JobSets are unaffected. Setting it — even to an empty {} — tells the controller to compile exactly one Workload, owned by the JobSet, containing one or more PodGroupTemplates, and to keep matching PodGroup objects in sync with it.

spec.scheduling supports two mutually exclusive models — a JobSet must use exactly one, since composite Gang-of-Gangs PodGroup hierarchies linking a parent PodGroup to leaf PodGroups aren’t implemented in alpha:

  • schedulingPolicy, schedulingConstraints, and disruptionMode at the top level of spec.scheduling configure a single composite PodGroup (or, under sequenced startup, one PodGroup per ReplicatedJob) covering the whole JobSet. Leave replicatedJobs unset when using this model.
  • replicatedJobs lets you target one or more ReplicatedJobs by name with their own leaf-level policy, producing one PodGroup per policy entry. Every ReplicatedJob in the JobSet must be targeted by exactly one entry, since there’s no top-level policy for an untargeted ReplicatedJob to fall back to. Leave the top-level schedulingPolicy, schedulingConstraints, and disruptionMode unset when using this model.
  • job nested inside a replicatedJobs entry goes one level deeper still, giving each Job replica of a ReplicatedJob its own PodGroup (“gang-of-gangs”).

Child Jobs are annotated with scheduling.k8s.io/group-template-name so you can trace which PodGroupTemplate each Job belongs to.

Use Case: Whole-JobSet Gang Scheduling

The most common pattern: every pod across every ReplicatedJob must be schedulable before any of them are admitted. Set a top-level gang policy and no replicatedJobs — the controller creates a single PodGroup sized to every pod in the JobSet.

# Gang Scheduling Example
#
# All pods across driver and workers are gang-scheduled: they must be admitted
# together atomically or none at all. This is the most common pattern for
# distributed ML training.
#
# Expected results:
#   - One Workload is created with 1 PodGroupTemplate
#     (top-level gang with no per-RJ overrides → single PodGroup)
#   - A single PodGroup with Gang policy (computed minCount)
#   - All child Jobs carry the scheduling.k8s.io/group-template-name annotation
#   - The Workload has an OwnerReference pointing to the JobSet
#
# Verify (generated Workload and PodGroup names include identity hashes):
#   kubectl get workloads,podgroups,jobs -n default
#   kubectl describe workload -n default -l jobset.sigs.k8s.io/jobset-name=gang-training
#   kubectl get jobs -n default -o jsonpath='{range .items[*]}{.metadata.name}: {.metadata.annotations}{"\n"}{end}'

apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
  name: gang-training
  namespace: default
spec:
  successPolicy:
    operator: All
  network:
    enableDNSHostnames: true
  scheduling:
    schedulingPolicy:
      gang: {}
    disruptionMode: 
      all: {}
  replicatedJobs:
    - name: driver
      replicas: 1
      template:
        spec:
          parallelism: 1
          completions: 1
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: driver
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "30s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi
    - name: workers
      replicas: 2
      template:
        spec:
          parallelism: 2
          completions: 2
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: worker
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "30s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi

A single-ReplicatedJob variant of the same idea:

# Gang Scheduling Example
#
# A single ReplicatedJob ("workers", 2 replicas x parallelism 4) is
# gang-scheduled: all 8 pods must be admitted together atomically or none at
# all (one PodGroup, computed minCount=8).
#
# Verify (generated Workload and PodGroup names include identity hashes):
#   kubectl get workloads,podgroups,jobs -n default
#   kubectl describe workload -n default -l jobset.sigs.k8s.io/jobset-name=single-gang
#   kubectl get jobs -n default -o jsonpath='{range .items[*]}{.metadata.name}: {.metadata.annotations}{"\n"}{end}'

apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
  name: single-gang
  namespace: default
spec:
  successPolicy:
    operator: All
  network:
    enableDNSHostnames: true
  scheduling:
    schedulingPolicy:
      gang: {}
  replicatedJobs:
    - name: workers
      replicas: 2
      template:
        spec:
          parallelism: 4
          completions: 4
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: worker
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "30s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi
kubectl get workloads,podgroups,jobs -n default
kubectl describe workload -n default -l jobset.sigs.k8s.io/jobset-name=gang-training

Use Case: No Scheduling (Backward Compatibility)

Existing JobSets that don’t set spec.scheduling must keep working exactly as before: no Workload, no PodGroup, no scheduling annotations on child Jobs.

# No-Scheduling Baseline
#
# This JobSet does NOT set spec.scheduling. It verifies backward compatibility:
# existing JobSets without scheduling config must continue to work identically,
# producing zero Workload or PodGroup objects.
#
# Expected results:
#   - Jobs are created normally
#   - NO Workload objects exist for this JobSet
#   - NO PodGroup objects exist for this JobSet
#   - Child Jobs do NOT carry scheduling.k8s.io annotations
#
# Verify:
#   kubectl get jobs -n default -l jobset.sigs.k8s.io/jobset-name=no-sched-baseline
#   kubectl get workloads -n default         # should NOT contain no-sched-baseline
#   kubectl get podgroups -n default         # should NOT contain no-sched-baseline
#   kubectl get jobs -n default -o jsonpath='{range .items[*]}{.metadata.name}: {.metadata.annotations}{"\n"}{end}'

apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
  name: no-sched-baseline
  namespace: default
spec:
  successPolicy:
    operator: All
  network:
    enableDNSHostnames: true
  # No spec.scheduling — this is the baseline.
  replicatedJobs:
    - name: workers
      replicas: 2
      template:
        spec:
          parallelism: 1
          completions: 1
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: worker
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "15s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi
kubectl get workloads -n default   # should not list no-sched-baseline
kubectl get podgroups -n default   # should not list no-sched-baseline

Use Case: Independent PodGroups per ReplicatedJob

Sometimes different ReplicatedJobs should be scheduled independently of one another rather than as one giant gang — for example, a driver that can start on its own while workers are gang-scheduled among themselves. Give each ReplicatedJob its own entry in replicatedJobs, and each gets its own PodGroup, sized from its own parallelism × replicas.

# KEP-969 Representative API: "Schedule groups independently"
#
# scheduling:
#   replicatedJobs:
#     - targetReplicatedJobs: [launcher]
#       schedulingPolicy: {gang: {}}
#     - targetReplicatedJobs: [worker]
#       schedulingPolicy: {gang: {}}
#
# Expected: one Workload containing one identity-hashed PodGroup per targeted
# ReplicatedJob: one for launcher and one for worker. Each PodGroup's minCount
# is independently computed from its own
# ReplicatedJob's parallelism x replicas (launcher: 1, worker: 4).
#
# Verify:
#   kubectl get workloads,podgroups,jobs -n default
#   kubectl get podgroups -n default -o custom-columns=NAME:.metadata.name,MINCOUNT:.spec.schedulingPolicy.gang.minCount
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
  name: per-rj-independent
  namespace: default
spec:
  successPolicy:
    operator: All
  network:
    enableDNSHostnames: true
  scheduling:
    replicatedJobs:
      - targetReplicatedJobs: ["launcher"]
        schedulingPolicy:
          gang: {}
      - targetReplicatedJobs: ["worker"]
        schedulingPolicy:
          gang: {}
  replicatedJobs:
    - name: launcher
      replicas: 1
      template:
        spec:
          parallelism: 1
          completions: 1
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: launcher
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "30s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi
    - name: worker
      replicas: 2
      template:
        spec:
          parallelism: 2
          completions: 2
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: worker
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "30s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi
kubectl get podgroups -n default -o custom-columns=NAME:.metadata.name,MINCOUNT:.spec.schedulingPolicy.gang.minCount

Use Case: One Shared PodGroup Across Multiple ReplicatedJobs

The opposite grouping: multiple ReplicatedJobs that must be admitted together as a single gang. List them all in one targetReplicatedJobs entry, and the controller creates one PodGroup whose minCount is the sum of both ReplicatedJobs’ pod counts.

# KEP-969 Representative API: "Schedule groups together"
#
# scheduling:
#   replicatedJobs:
#     - targetReplicatedJobs: [launcher, worker]
#       schedulingPolicy: {gang: {}}
#
# Expected: one Workload containing one PodGroup shared by launcher and worker.
# minCount is the sum of the pod counts of both targeted ReplicatedJobs
# (launcher: 1*1=1, worker: 2*2=4 -> total 5).
#
# Verify (generated Workload and PodGroup names include identity hashes):
#   kubectl get workloads,podgroups,jobs -n default
#   kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=per-rj-grouped -o jsonpath='{.items[0].spec.schedulingPolicy.gang.minCount}{"\n"}'
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
  name: per-rj-grouped
  namespace: default
spec:
  successPolicy:
    operator: All
  network:
    enableDNSHostnames: true
  scheduling:
    replicatedJobs:
      - targetReplicatedJobs: ["launcher", "worker"]
        schedulingPolicy:
          gang: {}
        disruptionMode:
          all: {}
  replicatedJobs:
    - name: launcher
      replicas: 1
      template:
        spec:
          parallelism: 1
          completions: 1
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: launcher
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "30s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi
    - name: worker
      replicas: 2
      template:
        spec:
          parallelism: 2
          completions: 2
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: worker
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "30s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi
kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=per-rj-grouped -o jsonpath='{.items[0].spec.schedulingPolicy.gang.minCount}{"\n"}'

Use Case: Gang-of-Gangs — Independent PodGroup per Job Replica

job goes one level finer than replicatedJobs: instead of one PodGroup covering every replica of a ReplicatedJob, each Job replica gets its own PodGroup, sized to just that Job’s own parallelism. This suits workloads where each replica is an independently-schedulable gang, such as per-replica launcher/worker sets.

apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
  name: per-ij-job-independent
  namespace: default
spec:
  successPolicy:
    operator: All
  network:
    enableDNSHostnames: true
  scheduling:
    replicatedJobs:
      - targetReplicatedJobs: ["launcher"]
        schedulingPolicy:
          gang: {}
      - targetReplicatedJobs: ["worker"]
        job:
          schedulingPolicy:
            gang: {}
          disruptionMode: 
            all: {}
  replicatedJobs:
    - name: launcher
      replicas: 1
      template:
        spec:
          parallelism: 1
          completions: 1
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: launcher
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "30s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi
    - name: worker
      replicas: 2
      template:
        spec:
          parallelism: 2
          completions: 2
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: worker
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "30s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi
kubectl get workloads,podgroups,jobs -n default

Use Case: Topology-Aware Placement

Combine gang scheduling with schedulingConstraints.topology to require that all pods in a gang land in the same topology domain — for example, co-locating pods on the same rack for RDMA-sensitive training, or coordinating placement across TPU slices. topology-constrained.yaml below uses replicatedJobs to give the driver an independent basic policy while workers get a gang policy with the rack constraint — every ReplicatedJob is targeted by its own entry since no top-level schedulingPolicy is set.

# Topology-Constrained Workers with Independent Driver
#
# The driver uses a Basic scheduling policy (schedules independently).
# Workers use an explicit Gang scheduling policy with topology
# constraints requiring co-location on the same rack. This is ideal for
# RDMA-based distributed training where network locality matters.
#
# The top-level schedulingPolicy/schedulingConstraints/disruptionMode API and
# replicatedJobs are mutually exclusive (there is no composite
# PodGroup linking leaf PodGroups together in alpha), so this uses
# replicatedJobs only: every ReplicatedJob is targeted by its own
# entry, each with an explicit leaf-level schedulingPolicy.
#
# Requires the TopologyAwareWorkloadScheduling feature gate on the apiserver;
# without it, the topology constraint is silently dropped on write and the
# JobSet controller will continuously delete/recreate the PodGroup trying to
# reconcile the drift (see hack/kind-config-scheduling.yaml).
#
# Expected results:
#   - One Workload with 2 PodGroupTemplates
#   - One driver PodGroup with Basic policy (no Gang)
#   - One workers PodGroup with Gang policy (computed minCount=4) and a
#     topology constraint on "topology.kubernetes.io/rack"
#   - All child Jobs carry scheduling annotations
#
# Verify (generated Workload and PodGroup names include identity hashes):
#   kubectl get workloads,podgroups -n default -l jobset.sigs.k8s.io/jobset-name=topo-training
#   kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=topo-training -o yaml

apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
  name: topo-training
  namespace: default
spec:
  successPolicy:
    operator: All
  network:
    enableDNSHostnames: true
  scheduling:
    # Per-ReplicatedJob policies. Every ReplicatedJob must be targeted by
    # exactly one entry, since no top-level schedulingPolicy is set.
    replicatedJobs:
      - targetReplicatedJobs: ["driver"]
        # Driver schedules independently; no gang requirement.
        schedulingPolicy:
          basic: {}
      - targetReplicatedJobs: ["workers"]
        # Workers are gang-scheduled with rack constraints.
        schedulingPolicy:
          gang: {}
        schedulingConstraints:
          topology:
            - key: topology.kubernetes.io/rack
  replicatedJobs:
    - name: driver
      replicas: 1
      template:
        spec:
          parallelism: 1
          completions: 1
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: driver
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "30s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi
    - name: workers
      replicas: 2
      template:
        spec:
          parallelism: 2
          completions: 2
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: worker
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "30s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi
# KEP-969 Representative API: "Run a TPU multi-slice workload"
#
# scheduling:
#   schedulingPolicy:
#     gang: {}
#   schedulingConstraints:
#     topology: ...
#
# Expected: one Workload and one JobSet-wide PodGroup with the requested
# topology constraint, coordinating placement across TPU slices. TPU-specific
# pod resources/topology values stay in the JobSet pod templates (not shown
# here in full since this example doesn't request real TPU resources).
#
# Requires the TopologyAwareWorkloadScheduling feature gate on the apiserver;
# see the note in topology-constrained.yaml for what happens without it.
#
# Verify (generated Workload and PodGroup names include identity hashes):
#   kubectl get workloads,podgroups,jobs -n default
#   kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=tpu-multislice -o jsonpath='{.items[0].spec.schedulingPolicy}{" "}{.items[0].spec.schedulingConstraints}{"\n"}'
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
  name: tpu-multislice
  namespace: default
spec:
  successPolicy:
    operator: All
  network:
    enableDNSHostnames: true
  scheduling:
    schedulingPolicy:
      gang: {}
    disruptionMode: 
      all: {}
    schedulingConstraints:
      topology:
        - key: cloud.google.com/gke-tpu-slice
  replicatedJobs:
    - name: slice
      replicas: 2
      template:
        spec:
          parallelism: 2
          completions: 2
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: worker
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "15s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi

Both examples require the TopologyAwareWorkloadScheduling feature gate; without it the constraint is silently dropped on write and the controller will continuously try to reconcile the resulting drift.

kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=topo-training -o yaml

Use Case: Sequenced Startup

When ReplicatedJobs use dependsOn (or an InOrder StartupPolicy) to create Jobs sequentially, not all pods exist at the same time — so a single JobSet-wide gang could never be satisfied. An explicit top-level schedulingPolicy.gang is therefore rejected. Leave schedulingPolicy unset (for example, scheduling: {}) to use the default Gang policy independently for each ReplicatedJob, or configure Gang policies explicitly under scheduling.replicatedJobs, targeting each ReplicatedJob separately.

# KEP-969 Representative API: "Combine sequencing with Gang scheduling"
#
# scheduling: {}
# replicatedJobs:
#   - name: leader
#     dependsOn: []
#   - name: worker
#     dependsOn:
#       - name: leader
#         status: Ready
#
# Expected: because DependsOn makes Job creation sequential, the controller
# uses one identity-hashed Gang PodGroup per ReplicatedJob instead of a single
# JobSet-wide PodGroup, avoiding
# deadlock (a single PodGroup requiring every pod at once could never be
# satisfied while worker's Job doesn't exist yet). Each PodGroup's minCount is
# computed independently from its own ReplicatedJob (leader: 1, worker: 2);
# an explicit top-level schedulingPolicy.gang is rejected with DependsOn or
# InOrder startup. Leave schedulingPolicy unset to use per-ReplicatedJob defaults.
#
# Verify:
#   kubectl get workloads,podgroups,jobs -n default
#   kubectl get jobs -n default   # worker's Job appears only after leader is Ready
#   kubectl get podgroups -n default -o custom-columns=NAME:.metadata.name,MINCOUNT:.spec.schedulingPolicy.gang.minCount
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
  name: sequenced-gang
  namespace: default
spec:
  successPolicy:
    operator: All
  network:
    enableDNSHostnames: true
  scheduling: {}
  replicatedJobs:
    - name: leader
      replicas: 1
      template:
        spec:
          parallelism: 1
          completions: 1
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: leader
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "10s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi
    - name: worker
      replicas: 1
      dependsOn:
        - name: leader
          status: Ready
      template:
        spec:
          parallelism: 2
          completions: 2
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: worker
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "10s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi
kubectl get jobs -n default   # worker's Job appears only after leader is Ready

Use Case: Suspend and Resume

Suspending a JobSet (spec.suspend: true) shouldn’t hold a scheduling reservation for pods that don’t exist. While suspended, the controller deletes the Workload/PodGroup; resuming recreates them under the same names.

# KEP-969 Representative API: "Avoid reserving resources for suspended workloads"
#
# spec:
#   suspend: true
#   scheduling: ...
#
# Expected: the Workload/PodGroup are deleted while suspended, and recreated
# (same names/mapping) when spec.suspend is set back to false.
#
# Verify:
#   kubectl apply -f site/static/examples/scheduling/suspend-resume.yaml
#   kubectl get workloads,podgroups -n default   # created
#   kubectl patch jobset gang-suspend -n default --type=merge -p '{"spec":{"suspend":true}}'
#   kubectl get workloads,podgroups -n default   # gone
#   kubectl patch jobset gang-suspend -n default --type=merge -p '{"spec":{"suspend":false}}'
#   kubectl get workloads,podgroups -n default   # recreated
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
  name: gang-suspend
  namespace: default
spec:
  suspend: false
  successPolicy:
    operator: All
  network:
    enableDNSHostnames: true
  scheduling:
    schedulingPolicy:
      gang: {}
  replicatedJobs:
    - name: workers
      replicas: 1
      template:
        spec:
          parallelism: 2
          completions: 2
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: worker
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "120s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi
kubectl patch jobset gang-suspend -n default --type=merge -p '{"spec":{"suspend":true}}'
kubectl get workloads,podgroups -n default   # gone
kubectl patch jobset gang-suspend -n default --type=merge -p '{"spec":{"suspend":false}}'
kubectl get workloads,podgroups -n default   # recreated

Use Case: Elastic Scaling

For an ElasticJobSet, scaling a ReplicatedJob’s parallelism/completions patches the corresponding PodGroup’s Gang.minCount in place rather than deleting and recreating the Workload.

# Elastic Gang Scheduling Example
#
# Mirrors the integration test "should patch Gang minCount in place when
# ElasticJobSet changes parallelism" (test/integration/scheduling/scheduling_test.go)
# and the KEP-969 "Scale an elastic workload" representative API.
#
# A single ReplicatedJob is gang-scheduled with parallelism/completions=2.
#
# Expected results (before scaling):
#   - One Workload with 1 PodGroupTemplate
#   - A single PodGroup with Gang policy, minCount=2
#
# The generated leaf GangSchedulingPolicy.MinCount field is required. Rather
# than persisting a derived value in the immutable spec.scheduling, the
# controller computes it from the live represented pod count whenever it
# compiles the scheduling objects. On every reconcile it recomputes minCount
# and patches the Workload/PodGroup in place when scaling changes the value.
#
# Because spec.scheduling is immutable, scale it by patching only
# spec.replicatedJobs (see elastic-gang-scale-patch.yaml) rather than
# re-applying the whole JobSet with a modified spec.scheduling block.
#
# Verify (generated Workload and PodGroup names include identity hashes):
#   kubectl get workloads,podgroups,jobs -n default
#   kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=gang-elastic -o jsonpath='{.items[0].spec.schedulingPolicy.gang.minCount}{"\n"}'
#   kubectl get workloads -n default -l jobset.sigs.k8s.io/jobset-name=gang-elastic -o jsonpath='{.items[0].metadata.uid}{"\n"}'
#
# To reproduce the scale-up and observe minCount update from 2 to 4 in place:
#   kubectl patch jobset gang-elastic -n default --type=json -p \
#     '[{"op":"replace","path":"/spec/replicatedJobs/0/template/spec/parallelism","value":4},
#       {"op":"replace","path":"/spec/replicatedJobs/0/template/spec/completions","value":4}]'
# (see elastic-gang-scale-patch.yaml for why a plain `kubectl apply` or
# `--type=merge` patch don't work for this.)

apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
  name: gang-elastic
  namespace: default
spec:
  successPolicy:
    operator: All
  network:
    enableDNSHostnames: true
  scheduling:
    schedulingPolicy:
      gang: {}
    disruptionMode:
      all: {}
  replicatedJobs:
    - name: workers
      replicas: 1
      template:
        spec:
          parallelism: 2
          completions: 2
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: worker
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "120s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi
# Scale patch for elastic-gang.yaml
#
# Apply this after elastic-gang.yaml to scale the "workers" ReplicatedJob's
# parallelism/completions from 2 to 4. See elastic-gang.yaml for how the
# controller keeps the PodGroup's minCount in sync with the live pod count
# across this kind of scale-up.
#
# NOTE: this manifest must keep spec.scheduling byte-for-byte identical to
# elastic-gang.yaml (including disruptionMode). spec.scheduling as a whole is
# immutable (enforced by a CEL rule), so a `kubectl apply` that changes it
# even incidentally (e.g. by omitting a field the live object has) is
# rejected with "spec.scheduling: Invalid value: Value is immutable" even
# though minCount itself is not what's being edited here.
#
# Usage:
#   kubectl apply -f site/static/examples/scheduling/elastic-gang.yaml
#   kubectl apply -f site/static/examples/scheduling/elastic-gang-scale-patch.yaml
#
# Or, without a full re-apply, patch just the scaled fields directly. Use a
# JSON patch (--type=json) rather than a merge patch (--type=merge): a merge
# patch replaces the whole replicatedJobs[0] list entry and therefore
# requires every required field (including the pod template) to be repeated,
# not just the two fields being changed.
#   kubectl patch jobset gang-elastic -n default --type=json -p \
#     '[{"op":"replace","path":"/spec/replicatedJobs/0/template/spec/parallelism","value":4},
#       {"op":"replace","path":"/spec/replicatedJobs/0/template/spec/completions","value":4}]'

apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
  name: gang-elastic
  namespace: default
spec:
  successPolicy:
    operator: All
  network:
    enableDNSHostnames: true
  scheduling:
    schedulingPolicy:
      gang: {}
    disruptionMode:
      all: {}
  replicatedJobs:
    - name: workers
      replicas: 1
      template:
        spec:
          parallelism: 4
          completions: 4
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              containers:
                - name: worker
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "120s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi

The generated leaf Gang.minCount field is required. Rather than persisting a derived value in the immutable spec.scheduling, the controller computes it from the live represented pod count whenever it compiles the scheduling objects. It recomputes that value on every reconcile and patches the Workload/PodGroup in place when scaling changes it. An explicit non-zero minCount under replicatedJobs[].job.schedulingPolicy.gang remains a user-selected fixed quorum instead.

Because spec.scheduling is immutable, scale by patching only spec.replicatedJobs rather than re-applying a JobSet whose spec.scheduling differs even incidentally (e.g. a missing field) from what’s already stored — that fails with spec.scheduling: Invalid value: Value is immutable, since the API server can’t tell that the rest of spec.scheduling was meant to stay the same.

kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=gang-elastic -o jsonpath='{.items[0].spec.schedulingPolicy.gang.minCount}{"\n"}'
# after applying elastic-gang-scale-patch.yaml, or an equivalent patch to
# spec.replicatedJobs, this reflects the new pod count (2 -> 4)

Use Case: Shared DRA Resource Claims

replicatedJobs[].job.resourceClaims attaches a PodGroup-level DRA ResourceClaimTemplate reference to each Job replica’s PodGroup, so all pods in that Job share a single allocated claim (for example, an NVLink/IMEX channel) instead of each pod claiming its own device.

# KEP-969 Representative API: "Share a DRA claim across a Job's pods"
#
# scheduling:
#   replicatedJobs:
#     - targetReplicatedJobs: [worker]
#       job:
#         schedulingPolicy: {gang: {}}
#         resourceClaims:
#           - name: imex-channel
#             resourceClaimTemplateName: imex-channel-template
#
# Expected: the "worker" PodGroup carries the resourceClaims entry. Each pod
# in "worker" references the same claim name + resourceClaimTemplateName in
# its own pod.spec.resourceClaims, which (with the DRAWorkloadResourceClaims
# feature gate enabled) resolves to the single ResourceClaim generated for
# and owned by the PodGroup instead of creating one ResourceClaim per pod.
#
# Requires: DRAWorkloadResourceClaims feature gate enabled on the apiserver,
# and a DRA driver + DeviceClass so the ResourceClaimTemplate is schedulable.
# This example only exercises the JobSet -> PodGroup resourceClaims wiring;
# without a real DRA driver installed the pods may stay Pending on claim
# allocation, but the PodGroup/ResourceClaimTemplate wiring can still be
# verified directly.
#
# Verify:
#   kubectl get workloads,podgroups,resourceclaimtemplates -n default
#   kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=dra-shared-claim -o jsonpath='{.items[0].spec.resourceClaims}{"\n"}'
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
  name: imex-channel-template
  namespace: default
spec:
  spec:
    devices:
      requests:
        - name: channel
          exactly:
            deviceClassName: imex-channel-verify-class
---
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
  name: dra-shared-claim
  namespace: default
spec:
  successPolicy:
    operator: All
  network:
    enableDNSHostnames: true
  scheduling:
    replicatedJobs:
      - targetReplicatedJobs: ["worker"]
        job:
          schedulingPolicy:
            gang: {}
          resourceClaims:
            - name: imex-channel
              resourceClaimTemplateName: imex-channel-template
  replicatedJobs:
    - name: worker
      replicas: 1
      template:
        spec:
          parallelism: 2
          completions: 2
          completionMode: Indexed
          template:
            spec:
              restartPolicy: OnFailure
              resourceClaims:
                - name: imex-channel
                  resourceClaimTemplateName: imex-channel-template
              containers:
                - name: worker
                  image: gcr.io/k8s-staging-perf-tests/sleep:v0.1.0
                  command: ["sleep", "10s"]
                  resources:
                    requests:
                      cpu: 100m
                      memory: 64Mi
                    claims:
                      - name: imex-channel

This requires the DRAWorkloadResourceClaims feature gate and a real DRA driver/DeviceClass to fully allocate; without one, the PodGroup-to-claim wiring can still be verified directly even if the pods stay Pending.

kubectl get podgroups -n default -l jobset.sigs.k8s.io/jobset-name=dra-shared-claim -o jsonpath='{.items[0].spec.resourceClaims}{"\n"}'

Use Case: Preemption

disruptionMode: {all: {}} (set on several examples above, such as gang-scheduling.yaml) tells the scheduler to preempt the entire gang together rather than individual pods, so a higher-priority JobSet doesn’t leave a lower-priority one half-evicted. Pair it with a PriorityClass on the pod template.