There is a useful kind of bad measurement.
It does not tell you what is wrong, exactly. It tells you that the shape of the problem is wrong.
On staging-build-1, normal Canopy Host Dockerfile builds had become that kind of problem. A build that should take a few minutes would run for an hour or more. Some did not complete inside an 80-minute window. We had already consolidated Dockerfiles and moved large, stable inputs later. That was sensible work. It reduced the number of opportunities to pay a bad cost. It did not explain the cost.
Then we noticed this:
#9 [build 2/6] WORKDIR /src DONE 297.4s
WORKDIR /src creates a directory. It took four minutes and 57 seconds.
That was the clue. The relevant cost was plainly not the number of bytes in the instruction.
The layer that was almost empty
The build was a normal Dockerfile build, run by BuildKit in a VM-isolated Kata pod on a build Grove. BuildKit state lived on an ext4 filesystem created on a dedicated block device. The checkout came in separately. Those details mattered, because there were several storage paths available to blame.
| What BuildKit was doing | Observed time |
|---|---|
| Extracting a roughly 80-byte layer | 318.6 s |
WORKDIR /src | 297.4 s |
| Extracting the roughly 820 KB distroless image | 1676.6 s |
| Extracting the golang base image | 1065.4 s |
A 67-byte layer and a 67.8 MB layer both took about 74 seconds in the same concurrent pair. Later pairs rose together: roughly 55 seconds, then 74, then 121, then 168.
That is not a bandwidth curve. It is accumulated state showing up as time.
We also captured a real dispatch through the actual egress proxy. About 313 MB of base-image blobs arrived in 12 seconds. A real checkout of the monorepo took 6.8 seconds. Neither can explain a WORKDIR that takes five minutes.
What the snapshotter was doing
BuildKit uses a snapshotter to prepare the filesystem seen by each Dockerfile instruction. At the time, our builder image explicitly selected containerd’s native snapshotter:
[worker.oci]
snapshotter = "native"
The important distinction is not that native is universally slow. For this workload, it did not provide a copy-on-write layer. Preparing a child snapshot recursively copied the parent filesystem before applying the instruction’s own change.
flowchart LR
layer1["Layer 1<br/>snapshot A"] --> layer2["Layer 2<br/>copy snapshot A<br/>apply tiny change"]
layer2 --> layer3["Layer 3<br/>copy snapshot A + B<br/>apply tiny change"]
layer3 --> layer4["Layer 4<br/>copy snapshot A + B + C<br/>apply tiny change"]
The copy happens inside the snapshotter’s metadata transaction. That explains the lockstep pairs: concurrent vertices can serialise behind it and each reports the wall-clock time.
With BuildKit’s default overlayfs snapshotter, a child layer initially refers to its parent and records only changed files. A small instruction can be small in the usual way.
flowchart LR
layer1["Layer 1<br/>lowerdir = A"] --> layer2["Layer 2<br/>lowerdir = A<br/>upperdir = tiny change B"]
layer2 --> layer3["Layer 3<br/>lowerdir = A + B<br/>upperdir = tiny change C"]
An almost-empty layer could therefore be expensive because it landed on a large filesystem. An 820 KB image could take 28 minutes to unpack because it was many successive snapshots, not because 820 KB is a large download.
Why it was there
The setting had a history. An earlier Kata build failed during a cross-stage COPY with an xattr-related ENOTSUP. Two changes landed close together: one selected the native snapshotter, and one enabled xattr support in virtiofsd.
virtio_fs_extra_args = ["--thread-pool-size=1", "--announce-submounts", "--xattr"]
The latter was the actual fix required for the xattr failure. Its investigation reproduced the failure regardless of snapshotter choice. Native was an understandable workaround that arrived beside the real fix, then survived as a historical assumption.
We had even reverted to overlayfs once before. It did not improve the build, but that was before the Grove had a real BuildKit block device. The experiment did not have the storage precondition the later re-test needed. We treated it as more conclusive than it was.
This is how fixes accumulate. The code looks like one answer. The history says it is two.
The fix that did not fix production
We removed the setting from the builder image source. The image used BuildKit’s overlayfs default on the ext4 BuildKit-state device. We kept virtiofsd --xattr, because removing it would have reintroduced a different failure.
Then production appeared unchanged.
We changed the code, built the fix image, updated FOREST_GROVE_BUILD_IMAGE_DOCKERFILE, restarted build workers, and observed another slow build. It looked as though the snapshotter theory was wrong.
The benchmark pass seemed to confirm that. A high-fidelity replica ran in about 31 seconds. buildctl debug workers reported overlayfs. The block device, ext4, Kata, BuildKit, cache flags, guest memory, swap, and contention all looked healthy. The benchmark was correct.
Its premise was wrong.
The benchmark used the fixed builder image. It was not the image staging-build-1 dispatched for real builds.
The configuration we changed was not the configuration in use
Forest resolves a worker image from a per-Grove database value and a global environment variable. The per-Grove value wins. The environment variable is a fallback.
flowchart TD
grove["Per-Grove build_dockerfile_image_ref"] -->|wins when set| resolved["Resolved builder image"]
global["FOREST_GROVE_BUILD_IMAGE_DOCKERFILE"] -->|fallback when no per-Grove value exists| resolved
staging-build-1 had a build_dockerfile_image_ref override. It still pointed at the old builder image, whose entrypoint wrote snapshotter = "native". Updating the global value and restarting workers did not change that Grove’s dispatches.
We had changed source and a real configuration value. Neither changed the image selected for production.
That is why the evidence contradicted itself. The slow production build was still native. The fast benchmark was already overlayfs. Both results were right about the software they actually ran. We were wrong about which software that was.
We confirmed this directly: the image requested by a fresh production dispatch, its resolved image ID, and the entrypoint inside a throwaway pod all matched the pre-fix source, including the native snapshotter line. The fixed image was in the registry. It simply was not selected.
The governing change was to update the Grove’s own image pin. The global setting was aligned too, but had no behavioural effect on this Grove. A per-Grove change needs no restart because resolution happens at dispatch time.
The actual before and after
We compared two real dispatches of the same workload on the same Grove. The successful run had no warm-cache escape hatch: it took about 170 seconds in buildctl, with 0 of 26 cache steps hit.
| Operation | Native snapshotter | Overlayfs default |
|---|---|---|
WORKDIR /src | 297.4 s | 1.6 s |
FROM distroless/static:nonroot extraction | 1676.6 s | 6.1 s |
FROM golang:1.26 extraction | 1065.4 s | 27.8 s |
| Byte-sized layers near 300 seconds | about 300 s | about 0.1 s |
Complete uncached buildctl | did not finish in the ordinary window | about 170 s |
Caching did not hide the bad path. There were no cache hits. The pathological filesystem behaviour was gone, and extraction time was proportional to the work again.
We did find a separate virtiofs performance problem during the benchmark pass. Its single-threaded configuration makes many-small-file work slower on that path. It was real, but not this incident. On the production build, checkout and context overhead bounded its contribution at roughly eight seconds. Virtiofs costs per file. Native snapshotting cost grew with the copied parent tree. Those are different signatures.
Making the next wrong pin visible
The immediate fix was small. The follow-up matters more.
Image resolution is now a shared, explicit Host module. Dispatch logs a warning when a per-Grove value shadows a different global value. A new read-only command shows the effective image, its source, and the shadowing condition:
canopy admin groves image-plan <grove-id>
Tests cover all five worker-image kinds, the per-Grove override, the global fallback, an unconfigured result, and the exact incident case: a global image set to one value while a Grove quietly selects another.
Per-Grove image references still exist for a reason, including registry reachability. The point is not to make stale configuration impossible. It is to make precedence visible before an operator has to reconstruct it from a live pod spec.
Source configuration is not evidence of runtime configuration. The system should be able to answer what it will run next.
Another way a build could wait an hour
While investigating hour-long builds, we found another, completely different way a build could legitimately wait an hour.
A VM-isolated BuildKit build receives an ephemeral, block-mode PVC for its state. On staging-build-1, that PVC was backed by the Grove’s one local BuildKit block-device PV. Kubernetes correctly keeps the PVC while the pod that owns it exists. The problem was the pod lifetime.
Finished Jobs were retained for TTLSecondsAfterFinished = 3600, keeping useful Job specifications and logs for debugging. A completed build pod therefore retained its ephemeral PVC. The local PV remained bound. The next build had no suitable persistent volume and could not schedule.
flowchart TD
complete["Build A completes"] --> retained["Pod remains for the Job TTL"]
retained --> pvc["Ephemeral BuildKit-state PVC remains Bound"]
pvc --> unavailable["Local BuildKit PV stays unavailable"]
unavailable --> blocked["Build B cannot schedule"]
The scheduler said so directly: no available persistent volumes to bind. Deleting the already-completed pod released the resource immediately. The local provisioner cleaned and re-offered the device, and the next build started about two and a half minutes later.
This was not the snapshotter incident. It was a resource-lifetime bug. An etcd and debugging retention decision had accidentally become the release deadline for a scarce physical device.
The controller now waits two minutes after a BuildKit Job becomes terminal, long enough to collect its exit code, artifact digest, workspace digest, and cache counters. It then deletes only finished build pods that hold the BuildKit-state volume. Other finished build pods keep their normal debugging TTL.
flowchart TD
complete["Build A completes"] --> facts["Status and artifact facts are collected"]
facts --> reclaim["BuildKit-state pod is reclaimed after a short grace"]
reclaim --> recycle["Local PV is recycled"]
recycle --> scheduled["Build B can schedule"]
Scheduler reasons now reach Host after a grace period, rather than making an unschedulable build look like an unexplained long wait. The reclamation code has focused tests for the grace window, failed and successful Jobs, orphaned pods, pods without a block volume, and volume identification.
What changed in how we look at this
The native snapshotter made a normal Dockerfile instruction pay for an increasingly large parent filesystem. That was the performance failure.
The more durable failure was ambiguity. We could change source, build an image, update a global value, restart workers, and still not know that a particular Grove had selected an older image. A benchmark could be accurate and still fail as a production reproduction because it had the wrong pin.
The PV incident made the same point from a different direction. Kubernetes was behaving correctly. The bug was our decision to let a debugging retention period own the lifetime of a one-device resource pool.
Forest puts Canopy Host’s build and runtime work through Kubernetes without asking application teams to operate Kubernetes objects. That only works if the abstraction preserves the facts needed to repair it. A scheduler reason is useful. The selected image digest is useful. The distinction between an xattr failure, a snapshotter, a shared filesystem, and a bound PVC is useful.
The build is fast again. More importantly, the next person debugging a build Grove has a better chance of proving what it is actually running before they explain its behaviour.