Contributing guidelines
I've found a bug and checked that ...
Description
With the kubernetes driver, the builder pod becomes Ready about 30s after its container starts, although buildkitd is serving within about 1s. buildx create --bootstrap waits for that, so bootstrap takes about 31s instead of a few seconds.
Expected behaviour
The builder pod becomes Ready, and buildx create --bootstrap returns, within a few seconds of buildkitd starting to serve.
Actual behaviour
The pod becomes Ready 30–32s after container start in every run. The first readiness probe that actually runs is the one triggered by the 30s periodSeconds timer, because initialDelaySeconds: 5 drops the probe kubelet attempts at container start. See Additional info for details and measurements.
Buildx version
github.com/docker/buildx v0.37.2 Homebrew
Docker info
Builders list
NAME/NODE DRIVER/ENDPOINT STATUS BUILDKIT PLATFORMS
probe-test kubernetes
\_ probe-test0 \_ kubernetes:///probe-test?deployment=buildkit-3d1526b0-b87c-49f7-8472-8cbaefdfe66d-862a5&kubeconfig=%2FUsers%2Fuser%2F.config%2Fkube%2Fconfig running v0.33.1 linux/amd64 (+4), linux/386
Configuration
No Dockerfile is involved. The delay is in bootstrap.
$ docker buildx create --name probe-test --driver kubernetes \
--driver-opt namespace=buildx-probe-test,replicas=1 \
--bootstrap
$ pod=$(kubectl -n buildx-probe-test get pods -l app=probe-test0 -o name)
$ kubectl -n buildx-probe-test get $pod -o jsonpath='{range .status.conditions[*]}{.type}={.lastTransitionTime}{"\n"}{end}containerStarted={.status.containerStatuses[0].state.running.startedAt}{"\n"}'
$ kubectl -n buildx-probe-test logs --timestamps $pod
$ docker buildx rm probe-test
Build logs
=== v0.37.2 (release) ===
#1 waiting for 1 pods to be ready, timeout: 2 minutes
#1 waiting for 1 pods to be ready, timeout: 2 minutes 32.0s done
#1 DONE 32.5s
--- pod status
PodScheduled=2026-10-02T20:05:32Z
Initialized=2026-10-02T20:05:32Z
PodReadyToStartContainers=2026-10-02T20:05:33Z
containerStarted=2026-10-02T20:05:33Z
ContainersReady=2026-10-02T20:06:03Z
Ready=2026-10-02T20:06:03Z
--- readinessProbe / livenessProbe
{"exec":{"command":["buildctl","debug","workers"]},"failureThreshold":3,"initialDelaySeconds":5,"periodSeconds":30,"successThreshold":1,"timeoutSeconds":60}
--- buildkitd logs
2026-10-02T20:05:33.352683496Z creating cgroup namespace
2026-10-02T20:05:33.394549109Z time="2026-10-02T20:05:33Z" level=info msg="auto snapshotter: using overlayfs"
2026-10-02T20:05:33.394908981Z time="2026-10-02T20:05:33Z" level=warning msg="using host network as the default"
2026-10-02T20:05:33.500291667Z time="2026-10-02T20:05:33Z" level=warning msg="failed check for fsverity support" error="enable fsverity failed: inappropriate ioctl for device" path=/var/lib/buildkit/runc-overlayfs/content
2026-10-02T20:05:33.510910020Z time="2026-10-02T20:05:33Z" level=info msg="found worker \"28mywzl2177xcdacpeetjxwbe\", labels=map[org.mobyproject.buildkit.worker.executor:oci org.mobyproject.buildkit.worker.hostname:probe-test0-5877ddc5c-nmhwh org.mobyproject.buildkit.worker.network:host org.mobyproject.buildkit.worker.oci.process-mode:sandbox org.mobyproject.buildkit.worker.selinux.enabled:false org.mobyproject.buildkit.worker.snapshotter:overlayfs], platforms=[linux/amd64 linux/amd64/v2 linux/amd64/v3 linux/amd64/v4 linux/386]"
2026-10-02T20:05:33.511627804Z time="2026-10-02T20:05:33Z" level=warning msg="skipping containerd worker, as \"/run/containerd/containerd.sock\" does not exist"
2026-10-02T20:05:33.511638233Z time="2026-10-02T20:05:33Z" level=info msg="found 1 workers, default=\"28mywzl2177xcdacpeetjxwbe\""
2026-10-02T20:05:33.511641389Z time="2026-10-02T20:05:33Z" level=warning msg="currently, only the default worker can be used."
2026-10-02T20:05:33.521255986Z time="2026-10-02T20:05:33Z" level=info msg="running server on /run/buildkit/buildkitd.sock"
=== #4124 (github.com/docker/buildx v0.0.0+unknown 685f46b511a6bd558223b120653df830d6ae5a2f) ===
#1 waiting for 1 pods to be ready, timeout: 2 minutes
#1 waiting for 1 pods to be ready, timeout: 2 minutes 6.2s done
#1 DONE 6.6s
--- pod status
PodScheduled=2026-10-02T20:06:15Z
Initialized=2026-10-02T20:06:15Z
PodReadyToStartContainers=2026-10-02T20:06:15Z
containerStarted=2026-10-02T20:06:15Z
ContainersReady=2026-10-02T20:06:21Z
Ready=2026-10-02T20:06:21Z
--- startupProbe
{"exec":{"command":["buildctl","debug","workers"]},"failureThreshold":24,"periodSeconds":5,"successThreshold":1,"timeoutSeconds":60}
--- readinessProbe / livenessProbe
{"exec":{"command":["buildctl","debug","workers"]},"failureThreshold":3,"periodSeconds":30,"successThreshold":1,"timeoutSeconds":60}
--- buildkitd logs
2026-10-02T20:06:15.690363682Z creating cgroup namespace
2026-10-02T20:06:15.737762999Z time="2026-10-02T20:06:15Z" level=info msg="auto snapshotter: using overlayfs"
2026-10-02T20:06:15.738142313Z time="2026-10-02T20:06:15Z" level=warning msg="using host network as the default"
2026-10-02T20:06:15.959975612Z time="2026-10-02T20:06:15Z" level=warning msg="failed check for fsverity support" error="enable fsverity failed: inappropriate ioctl for device" path=/var/lib/buildkit/runc-overlayfs/content
2026-10-02T20:06:15.970055747Z time="2026-10-02T20:06:15Z" level=info msg="found worker \"bwhe7ntguic6fs713ayooi1d4\", labels=map[org.mobyproject.buildkit.worker.executor:oci org.mobyproject.buildkit.worker.hostname:probe-test0-7795758b54-t65jj org.mobyproject.buildkit.worker.network:host org.mobyproject.buildkit.worker.oci.process-mode:sandbox org.mobyproject.buildkit.worker.selinux.enabled:false org.mobyproject.buildkit.worker.snapshotter:overlayfs], platforms=[linux/amd64 linux/amd64/v2 linux/amd64/v3 linux/amd64/v4 linux/386]"
2026-10-02T20:06:15.971938943Z time="2026-10-02T20:06:15Z" level=warning msg="skipping containerd worker, as \"/run/containerd/containerd.sock\" does not exist"
2026-10-02T20:06:15.971951871Z time="2026-10-02T20:06:15Z" level=info msg="found 1 workers, default=\"bwhe7ntguic6fs713ayooi1d4\""
2026-10-02T20:06:15.971955595Z time="2026-10-02T20:06:15Z" level=warning msg="currently, only the default worker can be used."
2026-10-02T20:06:15.981312133Z time="2026-10-02T20:06:15Z" level=info msg="running server on /run/buildkit/buildkitd.sock"
Additional info
Environment
Why the pod waits 30s
initialDelaySeconds drops probe attempts made before the delay ends. It doesn't schedule an attempt for when the delay ends. Probes run on the periodSeconds timer, plus extra attempts kubelet makes while a container isn't Ready:
- The container starts. Kubelet triggers an immediate readiness probe (prober_manager.go#L350-L357) and resets the readiness timer to 30s (worker.go#L194-L197).
- The container has been running for less than 5s, so the probe returns without running the command (worker.go#L330-L333). No result is recorded, so nothing triggers another attempt.
- The timer fires at about +30s. This is the first probe that runs, and the pod becomes Ready.
The same behaviour is reported in kubernetes/website#43259 (1.22, 1.27) and kubernetes/website#48519 (1.30): when initialDelaySeconds is lower than periodSeconds, the first probe runs at periodSeconds.
Removing initialDelaySeconds alone lets the immediate attempt run. Whether the pod becomes Ready then depends on buildkitd already serving at that moment; if it isn't, readiness falls back to the 30s timer. A startup probe doesn't depend on that timing. When the startup probe succeeds, kubelet triggers a readiness probe, and with no initial delay that probe runs right away.
Reproducer without buildx
Three pods:
Each probe command first writes a line to PID 1's stdout, so kubectl logs --timestamps shows when kubelet actually ran it. A probe skipped by initialDelaySeconds leaves no line.
repro.yaml
apiVersion: v1
kind: Pod
metadata:
name: probes-master
spec:
containers:
- name: buildkitd
image: moby/buildkit:buildx-stable-1
securityContext:
privileged: true
readinessProbe:
exec:
command: ["sh", "-c", "echo probe: readiness >/proc/1/fd/1; buildctl debug workers >/dev/null"]
initialDelaySeconds: 5
periodSeconds: 30
timeoutSeconds: 60
failureThreshold: 3
livenessProbe:
exec:
command: ["sh", "-c", "echo probe: liveness >/proc/1/fd/1; buildctl debug workers >/dev/null"]
initialDelaySeconds: 5
periodSeconds: 30
timeoutSeconds: 60
failureThreshold: 3
---
apiVersion: v1
kind: Pod
metadata:
name: probes-pr
spec:
containers:
- name: buildkitd
image: moby/buildkit:buildx-stable-1
securityContext:
privileged: true
startupProbe:
exec:
command: ["sh", "-c", "echo probe: startup >/proc/1/fd/1; buildctl debug workers >/dev/null"]
periodSeconds: 5
timeoutSeconds: 60
failureThreshold: 24
readinessProbe:
exec:
command: ["sh", "-c", "echo probe: readiness >/proc/1/fd/1; buildctl debug workers >/dev/null"]
periodSeconds: 30
timeoutSeconds: 60
failureThreshold: 3
livenessProbe:
exec:
command: ["sh", "-c", "echo probe: liveness >/proc/1/fd/1; buildctl debug workers >/dev/null"]
periodSeconds: 30
timeoutSeconds: 60
failureThreshold: 3
---
apiVersion: v1
kind: Pod
metadata:
name: probes-nodelay
spec:
containers:
- name: buildkitd
image: moby/buildkit:buildx-stable-1
securityContext:
privileged: true
readinessProbe:
exec:
command: ["sh", "-c", "echo probe: readiness >/proc/1/fd/1; buildctl debug workers >/dev/null"]
periodSeconds: 30
timeoutSeconds: 60
failureThreshold: 3
livenessProbe:
exec:
command: ["sh", "-c", "echo probe: liveness >/proc/1/fd/1; buildctl debug workers >/dev/null"]
periodSeconds: 30
timeoutSeconds: 60
failureThreshold: 3
kubectl apply -f repro.yaml
kubectl wait --for=condition=Ready pod/probes-master pod/probes-pr pod/probes-nodelay --timeout=180s
for p in master pr nodelay; do
echo "== probes-$p"
kubectl get pod probes-$p -o jsonpath='started={.status.containerStatuses[0].state.running.startedAt} ready={.status.conditions[?(@.type=="Ready")].lastTransitionTime}{"\n"}'
kubectl logs --timestamps probes-$p | grep -E 'probe:|running server'
done
Sample output (EKS):
== probes-master
started=2026-10-02T18:14:04Z ready=2026-10-02T18:14:35Z
2026-10-02T18:14:05.201179767Z time="2026-10-02T18:14:05Z" level=info msg="running server on /run/buildkit/buildkitd.sock"
2026-10-02T18:14:34.754522653Z probe: liveness
2026-10-02T18:14:35.043446260Z probe: readiness
2026-10-02T18:14:35.080853887Z probe: readiness
== probes-pr
started=2026-10-02T18:14:05Z ready=2026-10-02T18:14:10Z
2026-10-02T18:14:05.711110124Z time="2026-10-02T18:14:05Z" level=info msg="running server on /run/buildkit/buildkitd.sock"
2026-10-02T18:14:10.213776351Z probe: startup
2026-10-02T18:14:10.403317746Z probe: readiness
2026-10-02T18:14:35.216043687Z probe: liveness
2026-10-02T18:14:40.403790102Z probe: readiness
== probes-nodelay
started=2026-10-02T18:14:05Z ready=2026-10-02T18:14:06Z
2026-10-02T18:14:05.988668879Z probe: readiness
2026-10-02T18:14:06.180991662Z time="2026-10-02T18:14:06Z" level=info msg="running server on /run/buildkit/buildkitd.sock"
2026-10-02T18:14:06.394152984Z probe: readiness
2026-10-02T18:14:35.693876775Z probe: liveness
2026-10-02T18:14:36.404469637Z probe: readiness
Measurements
Definitions: container start is the container's startedAt, buildkitd serving is the timestamp of the running server on log line, and Ready is the Ready condition's lastTransitionTime (1s resolution).
buildx create --bootstrap on EKS, 10 runs each, alternating master and #4124:
|
bootstrap, median (range) |
container start → buildkitd serving |
container start → Ready |
| master |
32.7s (32.4–34.3s) |
0.2–1.1s |
30–32s |
| #4124 |
7.4s (7.1–8.2s) |
0.3–1.0s |
5–6s |
Manifest reproducer, container start → Ready:
|
EKS, 16 runs |
minikube, 10 runs |
| master probes |
30–31s |
30–31s |
| #4124 probes |
4–6s (one run 10s) |
5–6s |
master probes without initialDelaySeconds |
≤1s in 14 runs; 12s and 30s in the other 2 |
0–1s |
With #4124, the pod becomes Ready at the first startup probe tick after buildkitd is serving, which is within one 5s startup period.
Startup probe failure budget
- Healthy daemon:
buildctl debug workers takes about 10ms.
- No socket yet: it fails immediately.
- Socket open but daemon not responding:
buildctl gives up after 20s (gRPC-Go's default minConnectTimeout), before the 60s probe timeout.
Time from container start until kubelet restarts a hung buildkitd. The container command is sh -c 'buildkitd & sleep 2; kill -STOP $!; wait', run on EKS:
|
container start → restart |
| master (liveness probe) |
110s |
#4124 (startup timeoutSeconds: 60) |
485s (24 × ~20s) |
#4124 with startup timeoutSeconds: 5 |
125s |
failureThreshold: 24 × periodSeconds: 5 gives buildkitd 120s to start when probes fail fast. The 60s timeout was reused from the existing probes, and the startup probe doesn't need it.
Contributing guidelines
I've found a bug and checked that ...
Description
With the kubernetes driver, the builder pod becomes Ready about 30s after its container starts, although buildkitd is serving within about 1s.
buildx create --bootstrapwaits for that, so bootstrap takes about 31s instead of a few seconds.Expected behaviour
The builder pod becomes Ready, and
buildx create --bootstrapreturns, within a few seconds of buildkitd starting to serve.Actual behaviour
The pod becomes Ready 30–32s after container start in every run. The first readiness probe that actually runs is the one triggered by the 30s
periodSecondstimer, becauseinitialDelaySeconds: 5drops the probe kubelet attempts at container start. See Additional info for details and measurements.Buildx version
github.com/docker/buildx v0.37.2 Homebrew
Docker info
Builders list
Configuration
No Dockerfile is involved. The delay is in bootstrap.
Build logs
Additional info
Environment
moby/buildkit:buildx-stable-1)Why the pod waits 30s
initialDelaySecondsdrops probe attempts made before the delay ends. It doesn't schedule an attempt for when the delay ends. Probes run on theperiodSecondstimer, plus extra attempts kubelet makes while a container isn't Ready:The same behaviour is reported in kubernetes/website#43259 (1.22, 1.27) and kubernetes/website#48519 (1.30): when
initialDelaySecondsis lower thanperiodSeconds, the first probe runs atperiodSeconds.Removing
initialDelaySecondsalone lets the immediate attempt run. Whether the pod becomes Ready then depends on buildkitd already serving at that moment; if it isn't, readiness falls back to the 30s timer. A startup probe doesn't depend on that timing. When the startup probe succeeds, kubelet triggers a readiness probe, and with no initial delay that probe runs right away.Reproducer without buildx
Three pods:
probes-masteruses the probes buildx master generates.probes-pruses the probes from kubernetes: replace probe's InitialDelay with StartupProbe #4124.probes-nodelayuses master's probes withoutinitialDelaySeconds.Each probe command first writes a line to PID 1's stdout, so
kubectl logs --timestampsshows when kubelet actually ran it. A probe skipped byinitialDelaySecondsleaves no line.repro.yaml
Sample output (EKS):
Measurements
Definitions: container start is the container's
startedAt, buildkitd serving is the timestamp of therunning server onlog line, and Ready is the Ready condition'slastTransitionTime(1s resolution).buildx create --bootstrapon EKS, 10 runs each, alternating master and #4124:Manifest reproducer, container start → Ready:
initialDelaySecondsWith #4124, the pod becomes Ready at the first startup probe tick after buildkitd is serving, which is within one 5s startup period.
Startup probe failure budget
buildctl debug workerstakes about 10ms.buildctlgives up after 20s (gRPC-Go's defaultminConnectTimeout), before the 60s probe timeout.Time from container start until kubelet restarts a hung buildkitd. The container command is
sh -c 'buildkitd & sleep 2; kill -STOP $!; wait', run on EKS:timeoutSeconds: 60)timeoutSeconds: 5failureThreshold: 24×periodSeconds: 5gives buildkitd 120s to start when probes fail fast. The 60s timeout was reused from the existing probes, and the startup probe doesn't need it.