Skip to content

Fetch parallel pull layers from configured mirrors - #2095

Open
komapa wants to merge 4 commits into
awslabs:mainfrom
komapa:parallel-pull-mirrors
Open

komapa wants to merge 4 commits into
awslabs:mainfrom
komapa:parallel-pull-mirrors

Conversation

@komapa

@komapa komapa commented Sep 24, 2026 •

Copy link
Copy Markdown

Issue #, if available: Fixes #2094 (mirror part). Related to #1013 (same root cause, for SOCI artifacts instead of parallel pull layers)

Description of changes:

Parallel pull (parallel_pull_unpack) fetched every layer from the image registry, even when [resolver.host] mirrors were configured. preloadAllLayers created a single ORAS store for the image registry and only reused the HTTP client of the first resolved host.

This PR makes parallel pull use the configured mirrors:

  1. Resolver: wildcard hosts and mirror URL schemes (service/resolver)
    • [resolver.host."*"] applies to every registry without its own entry, like the "*" mirror in containerd's CRI registry config. Node-local mirrors such as Spegel are usually configured this way.
    • A mirror host with an http:// scheme uses plain HTTP; before, it was switched to https unless insecure = true was also set.
    • A mirror host without a scheme (mirror.example.com:5000) defaults to https; before, host:port without a scheme failed to parse and broke host resolution for the image.
  2. Parallel pull: fetch layers from mirrors (fs)
    • One blob source per resolved host, in resolver order (mirrors, then the image registry).
    • For each layer, mirrors are probed with a single HEAD request; the layer is fetched from the first mirror that has it, otherwise from the image registry. Probing runs inside each layer's premount goroutine, so it is parallel.
    • If the download from a host fails, it is retried from the next host (next mirror, then the image registry). The layer is written to the ingest file first and only then decompressed, unpacked, and verified, so the retry is limited to registry I/O; decompression and unpack errors still fail immediately. (This follows the retry scope discussed in Honor registry mirrors when fetching SOCI artifacts #1844.)
    • Requests to mirrors carry the ns=<registry> query parameter, as containerd sends it, so a mirror can tell which registry the image belongs to.
    • The image registry keeps its own scheme. Today isInsecureHost makes the image registry plain HTTP whenever any of its mirrors is insecure = true; on this path that only happens if the registry is configured as an insecure mirror of itself.
    • Mirrors with a custom path are skipped (blob URLs are built under /v2).
    • Behavior without mirrors is unchanged (no extra requests).

Docs: docs/config.md (host scheme handling, "*", and the insecure default, which is false in code) and docs/parallel-mode.md (Mirrors).

Testing performed:

  • New unit tests: TestAsRegistryHostsMirrors (wildcard, precedence, http scheme, scheme-less host); TestSelectBlobSource* (httptest mirror and origin: hit on mirror, miss falls back to origin, no probes without mirrors); TestNewBlobSourcesOriginScheme; TestParallelFetchFromMirrorWithFallback (mirror serves the layer; mirror GET fails and the image registry serves it; ns sent only to the mirror). The resolver tests fail without the resolver change, and the fallback test fails with the fallback disabled.
  • go test -race for all non-integration packages and golangci-lint run (v2.13.0): pass.
  • Parsed a Bottlerocket-rendered config (companion PR soci-snapshotter: render registry mirrors as resolver hosts bottlerocket-os/bottlerocket-core-kit#1061) with this branch: an ECR image resolves to http://<node>:30031 (Spegel) first, then the ECR endpoint.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.

Allow [resolver.host."*"] to configure mirrors for every registry without
its own entry, matching the "*" mirror in containerd's CRI registry config.
Registry mirrors running on the node, such as Spegel, are usually configured
this way.

Use plain HTTP when a mirror host is given with an http:// scheme, instead
of silently switching it to https unless insecure = true is set. A host
given without a scheme defaults to https; previously a host:port without a
scheme failed to parse and broke host resolution for the image.

Signed-off-by: Kiril Angov <kangov@seatgeek.com>
Parallel pull built a single remote store for the image registry and only
took the HTTP client from the first resolved host, so layers were always
fetched from the image registry even when [resolver.host] mirrors were
configured.

Build one blob source per resolved host (mirrors first, then the image
registry). For each layer, probe the mirrors in order with a single HEAD
request and fetch from the first one that has the blob, falling back to the
image registry. If the download fails, retry it from the next host; the
layer is only unpacked and verified after the download completes, so
decompression and unpack errors still fail immediately.

Requests to mirrors carry the ns=<registry> query parameter, as containerd
sends it. The image registry keeps its own scheme: an insecure mirror no
longer makes the image registry plain HTTP, unless the registry is
configured as an insecure mirror of itself. Mirrors with a custom path are
skipped, since blob URLs are always built under /v2.

This lets node-local registry mirrors such as Spegel serve layers to
parallel pull.

Signed-off-by: Kiril Angov <kangov@seatgeek.com>
@komapa
komapa force-pushed the parallel-pull-mirrors branch from 4b4b73b to 9fd00ac Compare September 24, 2026 19:15
@komapa

komapa commented Sep 24, 2026

Copy link
Copy Markdown
Author

For context on overlap with other work:

What this PR adds on top of both: [resolver.host."*"], http:// and scheme-less mirror hosts in the resolver, which is what Bottlerocket needs to pass settings.container-registry.mirrors through (bottlerocket-os/bottlerocket-core-kit#1061). #2096 is a separate fix for distribution source labels on stored layers.

@sondavidb sondavidb left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Doing a quick first pass, will look closer on Monday, couple of concerns but this will be great once it's merged. Thanks!

Comment thread docs/config.md Outdated
Comment thread fs/artifact_fetcher.go Outdated
Comment thread fs/fs.go
// download the target layer
s := src[0]
client := s.Hosts[0].Client
if len(s.Hosts) == 0 {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can probably also apply this logic in MountLocal?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

MountLocal now builds the mirror and origin stores with newBlobSources and picks one with selectBlobSource. One limitation I see is that MountLocal doesn't retry on download failure the way the parallel path does, so a mirror that has the layer but fails mid-download won't fall back to the next host.

I can add that if you want it in this PR?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

a mirror that has the layer but fails mid-download won't fall back to the next host

Hm, that's a good point. I think this would be a good followup PR then, since I think we should separate this behavioral change. Though I suppose two commits for this PR would also suffice.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let me know which way you prefer because this change is already getting somewhat big and I do not want it to get held up on unrelated change. I can still open a new PR for it, that's not the problem :)

Comment thread fs/fs.go Outdated
Comment thread fs/fs.go Outdated
Probe mirrors with GetHeader so registries that reject HEAD still work,
apply mirror selection in MountLocal, fold blobSource into orasBlobStore,
drop 'parallel pull' from the custom-path log, and leave the documented
mirror insecure default untouched.

Signed-off-by: Kiril Angov <kangov@seatgeek.com>
@sondavidb

Copy link
Copy Markdown
Contributor

CI looks to be failing on the new test

GetHeader makes three requests on a miss. A 404 to the HEAD probe is
final, so only fall back to GetHeader for other statuses, such as
registries that do not allow HEAD.

Signed-off-by: Kiril Angov <kangov@seatgeek.com>
@komapa

komapa commented Oct 8, 2026 •

Copy link
Copy Markdown
Author

CI looks to be failing on the new test

(Hopefully) Fixed the CI failure and hasBlob now treats a HEAD 404 as a definitive miss.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

go Pull requests that update Go code testing Unit and/or integration tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] Parallel pull ignores registry mirrors and stores layers without distribution source labels

2 participants