Skip to content

Instantly share code, notes, and snippets.

@jcpowermac
Created September 14, 2026 21:46
Show Gist options
  • Select an option

  • Save jcpowermac/f455c102ee2297907600a1f02b6183ff to your computer and use it in GitHub Desktop.

Select an option

Save jcpowermac/f455c102ee2297907600a1f02b6183ff to your computer and use it in GitHub Desktop.
OpenShift pull-through cache on a bastion (vSphere IPI, 4.20): registry:2 proxies, MachineConfig CRI-O mirrors, preload, verification

OpenShift pull-through cache on a bastion (vSphere IPI, 4.20)

Reduce cloud-provider egress for scale-out perf testing: every node image pull is rewritten to a registry:2 pull-through cache running on the bastion host. First pull of a unique image still goes upstream (bastion → provider); all subsequent pulls from any node are served from the bastion at L2 speed.

Tested on OpenShift 4.20 / vSphere IPI, CRI-O 1.36.5, ~50-80 churning autoscaled workers pulling quay-proxy.ci.openshift.org CI images.

Topology

Port Upstream Auth Notes
5000 https://quay.io none public images
5001 https://quay-proxy.ci.openshift.org pull-secret CI component images
5002 — (plain registry) none direct pushes only (caches reject pushes); used for a locally-built operator image
5003 https://registry.redhat.io pull-secret ubi/rhel images

BASTION_IP below = the bastion's IP on the node subnet (e.g. 10.22.13.252). Nodes reach it directly; no firewall changes needed on a vSphere cluster where the bastion is on the same network as workers.

1. Cache containers (on bastion)

Plain docker.io/library/registry:2 with the built-in proxy mode. One container per upstream so creds/failures are isolated.

# anonymous upstream
podman run -d --name cache-quayio \
  -p BASTION_IP:5000:5000 \
  -v mao-cache-quayio:/var/lib/registry \
  -e REGISTRY_HTTP_ADDR=0.0.0.0:5000 \
  -e REGISTRY_PROXY_REMOTEURL=https://quay.io \
  docker.io/library/registry:2

# authenticated upstream — creds from the cluster pull-secret
#   oc get secret pull-secret -n openshift-config -ojsonpath='{.data.\.dockerconfigjson}' | base64 -d | jq .
podman run -d --name cache-ciproxy \
  -p BASTION_IP:5001:5000 \
  -v mao-cache-ciproxy:/var/lib/registry \
  -e REGISTRY_HTTP_ADDR=0.0.0.0:5000 \
  -e REGISTRY_PROXY_REMOTEURL=https://quay-proxy.ci.openshift.org \
  -e REGISTRY_PROXY_USERNAME=<pull-secret-username> \
  -e REGISTRY_PROXY_PASSWORD=<pull-secret-password> \
  docker.io/library/registry:2

# same pattern for registry.redhat.io on port 5003 (creds: registry.redhat.io entry in pull-secret)

# plain registry for images you built yourself (caches reject pushes by design)
podman run -d --name mao-plain \
  -p BASTION_IP:5002:5000 \
  -v mao-plain-data:/var/lib/registry \
  -e REGISTRY_HTTP_ADDR=0.0.0.0:5000 \
  docker.io/library/registry:2

Container listens on 5000 internally; the host port does the mapping. Volumes are rootless podman volumes (~/.local/share/containers/storage/volumes/...).

Path semantics: the cache host is the prefix. A request for BASTION_IP:5001/openshift/ci@sha256:... makes the proxy fetch openshift/ci@sha256:... from the upstream — the origin host must NOT appear in the path.

Fallback: if the cache is down, CRI-O falls back to pulling from the origin — the bastion is not a hard SPOF.

2. Insecure registries (cluster-wide)

oc patch image.config.openshift.io/cluster --type=merge -p '{
  "spec":{"registrySources":{"insecureRegistries":[
    "BASTION_IP:5000","BASTION_IP:5001","BASTION_IP:5002","BASTION_IP:5003"
  ]}}}'

4.20 gotcha: spec.registrySources.rewritePolicies no longer exists in the API (patch fails with "unknown field"). Host rewrites must go through CRI-O registries.conf.d mirrors, delivered by MachineConfig — see below.

3. MachineConfig: CRI-O mirror rewrites

Drop /etc/containers/registries.conf.d/99-pull-through-cache.conf on every node. 99- prefix wins over distro defaults.

apiVersion: machineconfiguration.openshift.io/v1
kind: MachineConfig
metadata:
  labels:
    machineconfiguration.openshift.io/role: worker   # repeat for masters
  name: 99-pull-through-cache-worker
spec:
  config:
    ignition:
      version: 3.5.0
    storage:
      files:
      - contents:
          source: data:text/plain;base64,<base64 of the toml below>
        filesystem: root
        mode: 420
        path: /etc/containers/registries.conf.d/99-pull-through-cache.conf
[[registry]]
  prefix = "quay.io"
  location = "quay.io"
  [[registry.mirror]]
    prefix = "quay.io"
    location = "BASTION_IP:5000"
    insecure = true

[[registry]]
  prefix = "quay-proxy.ci.openshift.org"
  location = "quay-proxy.ci.openshift.org"
  [[registry.mirror]]
    prefix = "quay-proxy.ci.openshift.org"
    location = "BASTION_IP:5001"
    insecure = true

[[registry]]
  prefix = "registry.redhat.io"
  location = "registry.redhat.io"
  [[registry.mirror]]
    prefix = "registry.redhat.io"
    location = "BASTION_IP:5003"
    insecure = true

Generate the base64: base64 -w0 99-pull-through-cache.conf.

Notes:

  • [[registry.mirror]] with prefix is supported by CRI-O 1.36.x (4.18+).
  • Existing nodes get the file via MCO rollout (~1-2 min/node, CRI-O restart, no reboot). Nodes booted after the MachineConfig exists get it at bootstrap, so new autoscaled workers are wired from the start.
  • While the rollout is in flight, unserviced nodes can try HTTPS against the plain-HTTP cache and fail with http: server gave HTTP response to HTTPS client — harmless, they fall back to origin until rolled.

4. Preload the cache

The cache is lazy: it fills on the first pull per unique image, not per node. Preload = pull every image the fleet uses, once, from the bastion.

# collect unique images
oc get pods -A -ojson \
  | jq -r '.items[] | .spec.initContainers[]?.image, .spec.containers[]?.image' \
  | sort -u > images.txt

# warm each through the right cache, 4-parallel
warm() {
  img="$1"
  case "$img" in
    quay-proxy.ci.openshift.org/*) cache="BASTION_IP:5001/${img#quay-proxy.ci.openshift.org/}" ;;
    quay.io/*)                     cache="BASTION_IP:5000/${img#quay.io/}" ;;
    registry.redhat.io/*)          cache="BASTION_IP:5003/${img#registry.redhat.io/}" ;;
    *) echo "skip: $img"; return ;;
  esac
  podman pull --tls-verify=false "$cache" >/dev/null \
    && echo "OK   $img" || echo "FAIL $img"
}
export -f warm
xargs -a images.txt -P 4 -I{} bash -c 'warm "$1"' _ {}

YAGNI note: on vSphere machinesets, new workers are usually template clones (check Machine.spec.providerSpec.value.template) — the RHCOS rootfs is not downloaded per VM, so the only per-worker egress is the node's CRI-O component pulls, which is exactly what this covers.

5. Verification

# pod on a rolled node pulls via the cache
oc run cache-test --rm -it --overrides='{
  "spec":{"nodeSelector":{"kubernetes.io/hostname":"<rolled-worker>"},
          "tolerations":[{"operator":"Exists"}],
          "containers":[{"name":"t","image":"BASTION_IP:5001/openshift/ci@sha256:<digest>",
                         "command":["sleep","30"]}]}}' --restart=Never
# expect Running in a few seconds; cache logs show GETs from the node IP

# prove a second pull hits the cache, not upstream
podman logs cache-ciproxy | wc -l   # note N
# ... second pull ...
podman logs cache-ciproxy | tail -n +$((N+1)) | grep -c "proxy fetch"   # 0 = served locally

# node has the config?
oc debug node/<worker> -- test -f /host/etc/containers/registries.conf.d/99-pull-through-cache.conf

6. Teardown

oc delete mc 99-pull-through-cache-master 99-pull-through-cache-worker
oc patch image.config.openshift.io/cluster --type=merge \
  -p '{"spec":{"registrySources":{"insecureRegistries":[]}}}'
podman rm -f cache-quayio cache-ciproxy cache-redhat mao-plain
podman volume rm mao-cache-quayio mao-cache-ciproxy mao-plain-data mao-cache-redhat

(Skip on throwaway test clusters.)

Gotchas hit while building this

  • registry:2 proxy mode rejects pushes — keep a separate plain registry for locally-built images.
  • A xargs-fan-out preload script that writes to a log file must export LOG, or every child silently no-ops (>> $LOG → empty filename).
  • pkill -f <script> from an ssh one-liner matches the ssh shell's own command line (it contains the script text) and kills itself. Use pkill -xf "/bin/bash /path/to/script" (exact argv match).
  • CRI-O mirror + insecure = true requires the host:port in insecureRegistries (step 2) and the MachineConfig (step 3); either one alone is not enough.
  • If the CVO is running it will fight you on image pins, but it does not touch registries.conf — no CVO scale-down needed for the cache itself.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment