Upcoming Kubernetes 1.38 preview: Sneak Peek
Akın Özer
·
5 minute read
Kubernetes 1.38 is due on December 16. The release team is tracking 86 enhancements, and 71 of them are alpha or beta, which is where the new toys are. I went through the KEPs and the API types on the master branch and picked the ones I'd reach for first. Code freeze is November 17, and most of what follows still has a pull request to merge before then.
Headed for beta
These fields are already shipped as alpha, most of them in 1.37, so you can try them today with the feature gate turned on. The gate has to be on in kube-apiserver and in the component that acts on the field. With the API server's gate off, the pod and Job fields below are dropped without an error, so check that yours survived the kubectl apply.
On October 6 the gates were still off by default on master. Beta won't always mean on by default either: the gRPC TLS and Job KEPs both allow a beta that stays off, so check the defaults when 1.38 ships.
StatefulSet Recreate strategy
A StatefulSet rollout that hits a pod stuck in ImagePullBackOff stops until somebody deletes that pod by hand, even after you fix the template. With Recreate, the controller deletes all the old pods, waits for them to terminate, then creates the new revision. You get downtime, and the PVCs are kept. This is KEP-3541, behind the StatefulSetRecreateStrategy gate.
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: web
spec:
updateStrategy:
type: Recreate
# replicas, selector, and template as usual
Stop signals in the pod spec
If your app wants SIGUSR1 or SIGQUIT on shutdown, you no longer need to rebuild the image or write a preStop hook. Set lifecycle.stopSignal on the container. The API rejects the field unless spec.os.name is set, and an older container runtime will accept it and then ignore it. KEP-4960, gate ContainerStopSignals.
spec:
os:
name: linux
containers:
- name: app
image: registry.example/app:1.0
lifecycle:
stopSignal: SIGUSR1
Probes for HTTP/2 and TLS-only gRPC
Two small fields save you from opening a second port just for health checks. protocol: HTTP2 on an httpGet probe sends cleartext HTTP/2 (KEP-5999, gate H2CContainerProbe). mode: TLS on a grpc probe connects over TLS, without verifying the server certificate (KEP-4939, gate GRPCContainerProbeTLS).
containers:
- name: app
image: registry.example/app:1.0
readinessProbe:
httpGet:
path: /readyz
port: 8080
protocol: HTTP2
livenessProbe:
grpc:
port: 8443
mode: TLS
Volume modes and file owners
emptyDir.mode gives you a proper sticky /tmp without an init container running chmod (KEP-5502, gate EmptyDirVolumeMode). defaultUser on Secret, ConfigMap, downward API, and projected volumes writes the files as a UID you choose, which helps software that insists on owning its key file (KEP-5936, gate AtomicWriteVolumeUserFields). If the pod also sets fsGroup, the kubelet adds group permissions and the setgid bit on top of your mode, so check the result with stat.
volumes:
- name: tmp
emptyDir:
mode: 01777
- name: key
secret:
secretName: app-key
defaultUser: 1000
A third volume feature, bindMountOptions: [noexec, nosuid, nodev] on volumeMounts (KEP-5855), is on the beta list too. It needs support in containerd and CRI-O that neither project has released yet, so I'd wait on that one.
Default sysctls from the kubelet config
Node-wide tuning such as TCP keepalive usually gets copied into every pod spec or injected by a webhook. defaultPodSysctls in the kubelet config applies namespaced sysctls to every new pod on the node. A pod that sets the same sysctl itself keeps its own value, and net.* defaults are skipped for hostNetwork pods. KEP-5996, gate DefaultPodSysctls on the kubelet, which rejects the field at startup if the gate is off.
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
featureGates:
DefaultPodSysctls: true
defaultPodSysctls:
net.ipv4.tcp_keepalive_time: "600"
Gang scheduling for Jobs
Add spec.scheduling to a Job and the scheduler places all of its pods or none of them. The Job controller creates the Workload and PodGroup objects for you. With the gate on it does that for every Job that has no owner, whether or not the Job sets spec.scheduling, so start on a test cluster. This is KEP-5547. You need the GenericWorkload and WorkloadWithJob gates, and the API server has to serve scheduling.k8s.io/v1beta1, which is off by default.
apiVersion: batch/v1
kind: Job
metadata:
name: train
spec:
parallelism: 4
completions: 4
scheduling:
schedulingPolicy:
gang:
minCount: 4 # optional, defaults to parallelism
# template as usual
New in alpha
None of these had merged on October 6, and the gates named below don't exist in any build yet. Don't add them to --feature-gates early, because an unknown gate stops the component from starting. Expect a rename or two, and expect some of them to slip to 1.39.
Pod-level PID and disk limits
Two KEPs add entries to spec.resources, the pod-level resources block. With pids, a pod can set a tighter process limit than the node-wide podPidsLimit, though it can't raise it (KEP-6063, gate PerPodPIDLimit). With ephemeral-storage, you size local disk once for the whole pod and stop splitting a shared emptyDir between containers on paper (KEP-6386, gate PodLevelResourcesEphemeralStorage).
spec:
resources:
requests:
ephemeral-storage: 2Gi
limits:
ephemeral-storage: 5Gi
pids: "2048"
containers:
- name: app
image: nginx
A sync period per HPA
Every HPA in a cluster is reconciled on one kube-controller-manager flag, 15 seconds by default. syncPeriodSeconds lets a single HPA run faster or slower, anywhere from 3 to 3600 seconds. KEP-6003, gate HPAConfigurableSyncPeriod.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: queue-worker
spec:
syncPeriodSeconds: 5
# scaleTargetRef, minReplicas, maxReplicas, and metrics as usual
Also on the list
More alphas are tracked with no merged code yet: in-place probe updates (KEP-6318), per-container ulimits (KEP-5758), seccomp profiles pulled from OCI registries (KEP-6061), replaceable node liveness detection (KEP-6371), and etcd object compression (KEP-6264). Pods with no network at all (defaultNetwork: None, KEP-6313) belong here too. That field was renamed twice in September, and no released container runtime supports it.
For DRA, derived attributes (KEP-6080) and list-typed attributes (KEP-5491) target beta. The EvictionRequest API (KEP-4563) shipped its types in 1.37, and the controllers that make it do something are still in review.
What to check before you upgrade
Most of this comes from the v1.38.0-alpha.1 changelog, so the list will grow. The deprecated API migration guide has nothing for 1.38 yet.
Feature gates removed: AggregatedDiscoveryRemoveBetaType, CRDValidationRatcheting, CustomResourceFieldSelectors, DisableAllocatorDualWrite, DynamicResourceAllocation, JobManagedBy, KubeletTracing, MatchLabelKeysInPodAffinity, MultiCIDRServiceAllocator, NFTablesProxyMode, PodSchedulingReadiness, PreferSameTrafficDistribution, SeparateTaintEvictionController, ServiceAccountTokenPodNodeInfo, and WindowsHostNetwork. The changelog names nine of them. The other six show up when you compare the feature gate list on master with v1.37.0. Components refuse to start when --feature-gates names a gate they don't recognize, so search your flags and config files for these first.
Feature gates newly locked on: StrictIPCIDRValidation, DRADeviceTaints, DRADeviceTaintRules, EnvFiles, and HPAGeneration. Setting a locked gate to false is a startup error too.
Strict IP and CIDR validation can no longer be turned off. It has been on by default since 1.36, and KEP-4858 goes GA in 1.38 with the gate locked. IPv4 addresses with leading zeros and IPv4-mapped IPv6 addresses such as ::ffff:1.2.3.4 are rejected. Existing objects keep working through ratcheting validation, but templates that create new objects with those forms will fail.
LimitRanger compares quantities exactly. A request of 0.9999 against a minimum of 1 used to round onto the bound and pass, and now it is rejected. This applies to pod and PVC creation, and to PVC updates and pod resizes that change the value.
kube-apiserver returns a JSON Status object for unknown /apis/... paths. It used to return the plain-text body 404 page not found, so any client that matches on that string needs a look.
A DeviceTaintRule with no spec.deviceSelector now matches no devices. It used to match every device in the cluster. Set deviceSelector: {} if you want the old behavior.
The endpoint_slice_controller_changes metric is deprecated in favor of endpoint_slice_controller_changes_total. Both are emitted during the deprecation period.
Go packages removed: the metrics.k8s.io/v1alpha1 types and client in k8s.io/metrics, and k8s.io/api/apidiscovery/v2beta1. A default 1.37 cluster wasn't serving either one, so this only matters if you build against those modules.
Dependencies: Go 1.27.1, default etcd 3.7.1, CoreDNS v1.14.7. The kubelet is now built without CGO by default.
Dates
November 4: v1.38.0-beta.0
November 9 to 12: KubeCon + CloudNativeCon North America in Salt Lake City
November 17: code freeze
November 23: the release team's deprecations and removals post
December 3 and December 9: release candidates rc.0 and rc.1
December 16: v1.38.0
KubeCon ends on a Thursday and code freeze follows on Tuesday, so if you're in Salt Lake City, that's the week to ask a SIG lead whether their KEP is going to make it.