Repository navigation
Fix metrics-server: migrate from retired Bitnami chart to upstream - #320
arnoldcastro5000 wants to merge 3 commits into
Conversation
The metrics-server pod has been stuck in ImagePullBackOff for ~19 days because the pinned Bitnami image no longer exists: Bitnami retired its free public catalog in Aug 2025, removing the versioned tag public.ecr.aws/bitnami/metrics-server:0.4.2-debian-10-r30. Switch the Ansible role from the Bitnami chart (metrics-server-5.8.3, app v0.4.2) to the upstream kubernetes-sigs chart (3.13.1, app v0.8.1), which is actively maintained and supports Kubernetes 1.31+ (the cluster runs 1.33). The DigitalOcean-specific --kubelet-insecure-tls flag, the cloud-services-node-pool affinity, and the metrics-server namespace are kept; only the values schema is translated (extraArgs map -> args list). Deployment remains manual via: ansible-playbook monitoring/metrics-server-playbook.yml
|
Hi @dadiorchen , the metrics-server is down for 19 days. This PR updates the monitoring/metrics-server-playbook.yml file to pull a new image. Restoring the metrics-server will require a manual execution of ansible-playbook for the prod cluster to pull the image. Requesting your review. |
|
hi @dadiorchen , review please. |
|
Has this been tested in dev ? Reviewing the docs here https://kubernetes-sigs.github.io/metrics-server/ indicates that there are several other components that are created by the installer for the updated metrics server. So simply changing the image might not be enough. |
|
Thanks @ZavenArra. You are right that the metrics-server needs more than an image, so let me show exactly what changes between the two charts and what stays the same. Short version: this PR does not swap the image on the old release. It replaces the whole Bitnami chart with the upstream kubernetes-sigs chart, and that chart ships every component itself. 1. Lineage and versions
The upstream chart lives in the same repo as the metrics-server code, so the chart and the binary release together. The Bitnami chart was a third-party repackage. 2. The image, which is the real root cause
In August 2025 Bitnami retired its free catalog and removed that exact tag, so the pod went to 3. Components each chart createsBoth charts create the full stack. This is your point directly.
The upstream chart ships a few more templates ( 4. How flags are passed (the one real translation)Bitnami took a map Upstream takes a list Upstream runs with extra sane defaults Bitnami never set ( 5. Default differences that matter
6. Kubernetes compatibility
SummaryThe change is not an image swap on the same chart. It replaces a retired third-party repackage (old binary, deleted image, Not yet tested on dev. Our flow here is approve, then I run it on the dev cluster, then prod. So the next step after this merges is the dev run, and I will post the results in this thread ( One migration note I want to flag for that dev run: because we reuse the release name, this is a |
The upstream chart renders defaultArgs then args and appends duplicate flags, so keeping --kubelet-preferred-address-types in args left the container with InternalIP,ExternalIP,Hostname,InternalIP. Override defaultArgs to set it once to InternalIP. Only --kubelet-insecure-tls stays in args.
What this PR does
Fixes the metrics-server on the
prod-k8s-treetrackercluster, which hasbeen completely down for about 19 days. It updates our Ansible role so it
installs metrics-server from the official Kubernetes source instead of the old
Bitnami source that no longer works.
metrics-server is the component that collects CPU and memory usage for nodes and
pods. While it is down:
kubectl top nodesandkubectl top podsreturn "Metrics API not available".based on CPU/memory does not work.
Why metrics-server is currently down
very old version from 2021 (
metrics-server-5.8.3, app version v0.4.2).image our old version needs —
public.ecr.aws/bitnami/metrics-server:0.4.2-debian-10-r30— was removed.ImagePullBackOffand never starts. The cluster has been retrying thedownload tens of thousands of times over the last 19 days.
In short: this is not a config or network problem — the image the old version
depends on simply no longer exists.
The fix
Point the Ansible role at the upstream kubernetes-sigs metrics-server chart
instead of Bitnami:
charts.bitnami.com/bitnamikubernetes-sigs.github.io/metrics-serverbitnami/metrics-server5.8.3metrics-server/metrics-server3.13.1v0.4.2(2021)v0.8.1(current)Everything specific to our setup is kept unchanged:
--kubelet-insecure-tls(required on DigitalOcean — see below)--kubelet-preferred-address-types=InternalIPcloud-services-node-poolnode affinitymetrics-servernamespaceThe only translation needed is the values format: Bitnami used an
extraArgsmap, the upstream chart uses an
argslist. Same flags, different shape.Why version v0.8.1 (chart 3.13.1) was chosen
It is the current, actively maintained release of metrics-server. The old
Bitnami version was 5 years out of date.
The official compatibility matrix says the 0.8.x line supports Kubernetes
1.31 and newer, and our cluster runs 1.33 — so this is the release line
that officially matches our cluster:
The old v0.4.2 only ever targeted very old Kubernetes versions, so even if its
image still existed it would not be a safe match for a 1.33 cluster.
The chart-to-app mapping was verified directly: chart
3.13.1ships appv0.8.1(checked in the chart'sChart.yaml).Why
--kubelet-insecure-tlsis still neededOn DigitalOcean the kubelet's TLS certificate does not include an IP address,
so metrics-server cannot verify it and would otherwise fail to scrape metrics.
This flag tells metrics-server to skip that specific certificate check. This was
already in our old config; we are keeping it.
How to deploy (after merge)
This PR only changes the Ansible role. It does not deploy anything by itself.
Deployment is manual and, because we are switching charts, needs one extra step
the first time.
The release keeps the same name (
metrics-server) and namespace, so Helm wouldotherwise treat this as an
upgradeover the old Bitnami release. Combined withatomic: true, a failed rollout would then roll back to the broken Bitnamirevision. To avoid that, uninstall the old release first, then run the playbook:
# one-time, because we are moving from the Bitnami chart to the upstream chart helm uninstall metrics-server -n metrics-server ansible-playbook monitoring/metrics-server-playbook.ymlHow to verify it worked
Rollback
If anything goes wrong, revert this PR. Note that rolling back to the previous
Bitnami version will not work, because its image no longer exists — that is
the very reason for this change. Any rollback must stay on an image source that
still publishes images (i.e. the upstream chart).
Because this is a chart migration, do not rely on
helm rollback; the previousrevision is the broken Bitnami release. Roll back by reverting this PR and
re-running the playbook on the upstream chart.
Safer rollout
The Ansible task now runs the Helm release with
wait: trueandatomic: true.This is a direct lesson from this outage: previously the install could report
success even though the pod never actually started, so the failure went
unnoticed for ~19 days.
wait: trueblocks until the pods are actually Ready.atomic: trueautomatically rolls the release back if it does not converge,so we never leave a half-broken release behind.
Follow-up (not in this PR)
metrics-server currently runs as a single replica pinned to the
cloud-services-node-poolwith a hard scheduling requirement. That means if thatnode pool is full, cordoned, or unavailable, metrics-server cannot run at all —
a single point of failure. A separate PR should consider running 2 replicas
with a PodDisruptionBudget for high availability. That is a behavior change
beyond restoring the service, so it is intentionally kept out of this fix.
Risk
Low. metrics-server has already been non-functional for ~19 days, so this change
can only restore functionality, not remove something currently working. It runs
in its own namespace and is pinned to the
cloud-services-node-pool.