Skip to content

Fix metrics-server: migrate from retired Bitnami chart to upstream - #320

Open
arnoldcastro5000 wants to merge 3 commits into
Greenstand:masterfrom
arnoldcastro5000:fix/metrics-server-upstream-chart
Open

arnoldcastro5000 wants to merge 3 commits into
Greenstand:masterfrom
arnoldcastro5000:fix/metrics-server-upstream-chart

Conversation

@arnoldcastro5000

@arnoldcastro5000 arnoldcastro5000 commented Jul 2, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does

Fixes the metrics-server on the prod-k8s-treetracker cluster, which has
been completely down for about 19 days. It updates our Ansible role so it
installs metrics-server from the official Kubernetes source instead of the old
Bitnami source that no longer works.

metrics-server is the component that collects CPU and memory usage for nodes and
pods. While it is down:

  • kubectl top nodes and kubectl top pods return "Metrics API not available".
  • Any Horizontal Pod Autoscaler (HPA) has no metrics to act on, so autoscaling
    based on CPU/memory does not work.

Why metrics-server is currently down

  1. Our role installed metrics-server using the Bitnami Helm chart, pinned to a
    very old version from 2021 (metrics-server-5.8.3, app version v0.4.2).
  2. In August 2025, Bitnami retired its free public image catalog. The exact
    image our old version needs —
    public.ecr.aws/bitnami/metrics-server:0.4.2-debian-10-r30 — was removed.
  3. Because Kubernetes can no longer download that image, the pod is stuck in
    ImagePullBackOff and never starts. The cluster has been retrying the
    download tens of thousands of times over the last 19 days.

In short: this is not a config or network problem — the image the old version
depends on simply no longer exists.

The fix

Point the Ansible role at the upstream kubernetes-sigs metrics-server chart
instead of Bitnami:

Before After
Helm repo charts.bitnami.com/bitnami kubernetes-sigs.github.io/metrics-server
Chart bitnami/metrics-server 5.8.3 metrics-server/metrics-server 3.13.1
App version v0.4.2 (2021) v0.8.1 (current)

Everything specific to our setup is kept unchanged:

  • --kubelet-insecure-tls (required on DigitalOcean — see below)
  • --kubelet-preferred-address-types=InternalIP
  • the cloud-services-node-pool node affinity
  • the metrics-server namespace

The only translation needed is the values format: Bitnami used an extraArgs
map, the upstream chart uses an args list. Same flags, different shape.

Why version v0.8.1 (chart 3.13.1) was chosen

  • It is the current, actively maintained release of metrics-server. The old
    Bitnami version was 5 years out of date.

  • The official compatibility matrix says the 0.8.x line supports Kubernetes
    1.31 and newer
    , and our cluster runs 1.33 — so this is the release line
    that officially matches our cluster:

    metrics-server Supported Kubernetes
    0.8.x 1.31+
    0.7.x 1.27+
    0.6.x 1.25+
  • The old v0.4.2 only ever targeted very old Kubernetes versions, so even if its
    image still existed it would not be a safe match for a 1.33 cluster.

  • The chart-to-app mapping was verified directly: chart 3.13.1 ships app
    v0.8.1 (checked in the chart's Chart.yaml).

Why --kubelet-insecure-tls is still needed

On DigitalOcean the kubelet's TLS certificate does not include an IP address,
so metrics-server cannot verify it and would otherwise fail to scrape metrics.
This flag tells metrics-server to skip that specific certificate check. This was
already in our old config; we are keeping it.

How to deploy (after merge)

This PR only changes the Ansible role. It does not deploy anything by itself.
Deployment is manual and, because we are switching charts, needs one extra step
the first time.

The release keeps the same name (metrics-server) and namespace, so Helm would
otherwise treat this as an upgrade over the old Bitnami release. Combined with
atomic: true, a failed rollout would then roll back to the broken Bitnami
revision. To avoid that, uninstall the old release first, then run the playbook:

# one-time, because we are moving from the Bitnami chart to the upstream chart
helm uninstall metrics-server -n metrics-server

ansible-playbook monitoring/metrics-server-playbook.yml

How to verify it worked

# pod should become Running (1/1), not ImagePullBackOff
kubectl get pods -n metrics-server

# the metrics API should report AVAILABLE = True
kubectl get apiservice v1beta1.metrics.k8s.io

# these should return real numbers instead of "Metrics API not available"
kubectl top nodes
kubectl top pods -A

Rollback

If anything goes wrong, revert this PR. Note that rolling back to the previous
Bitnami version will not work, because its image no longer exists — that is
the very reason for this change. Any rollback must stay on an image source that
still publishes images (i.e. the upstream chart).

Because this is a chart migration, do not rely on helm rollback; the previous
revision is the broken Bitnami release. Roll back by reverting this PR and
re-running the playbook on the upstream chart.

Safer rollout

The Ansible task now runs the Helm release with wait: true and atomic: true.
This is a direct lesson from this outage: previously the install could report
success even though the pod never actually started, so the failure went
unnoticed for ~19 days.

  • wait: true blocks until the pods are actually Ready.
  • atomic: true automatically rolls the release back if it does not converge,
    so we never leave a half-broken release behind.

Follow-up (not in this PR)

metrics-server currently runs as a single replica pinned to the
cloud-services-node-pool with a hard scheduling requirement. That means if that
node pool is full, cordoned, or unavailable, metrics-server cannot run at all —
a single point of failure. A separate PR should consider running 2 replicas
with a PodDisruptionBudget for high availability. That is a behavior change
beyond restoring the service, so it is intentionally kept out of this fix.

Risk

Low. metrics-server has already been non-functional for ~19 days, so this change
can only restore functionality, not remove something currently working. It runs
in its own namespace and is pinned to the cloud-services-node-pool.

The metrics-server pod has been stuck in ImagePullBackOff for ~19 days
because the pinned Bitnami image no longer exists: Bitnami retired its
free public catalog in Aug 2025, removing the versioned tag
public.ecr.aws/bitnami/metrics-server:0.4.2-debian-10-r30.

Switch the Ansible role from the Bitnami chart (metrics-server-5.8.3,
app v0.4.2) to the upstream kubernetes-sigs chart (3.13.1, app v0.8.1),
which is actively maintained and supports Kubernetes 1.31+ (the cluster
runs 1.33). The DigitalOcean-specific --kubelet-insecure-tls flag, the
cloud-services-node-pool affinity, and the metrics-server namespace are kept;
only the values schema is translated (extraArgs map -> args list).

Deployment remains manual via:
  ansible-playbook monitoring/metrics-server-playbook.yml
@arnoldcastro5000

Copy link
Copy Markdown
Collaborator Author

Hi @dadiorchen , the metrics-server is down for 19 days. This PR updates the monitoring/metrics-server-playbook.yml file to pull a new image.

Restoring the metrics-server will require a manual execution of ansible-playbook for the prod cluster to pull the image.

Requesting your review.

@arnoldcastro5000

Copy link
Copy Markdown
Collaborator Author

hi @dadiorchen , review please.

ZavenArra
ZavenArra previously approved these changes Aug 31, 2026

@ZavenArra ZavenArra left a comment •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@ZavenArra
ZavenArra self-requested a review August 31, 2026 23:57
@ZavenArra

Copy link
Copy Markdown
Member

Has this been tested in dev ?

Reviewing the docs here https://kubernetes-sigs.github.io/metrics-server/ indicates that there are several other components that are created by the installer for the updated metrics server. So simply changing the image might not be enough.

@arnoldcastro5000

Copy link
Copy Markdown
Collaborator Author

Thanks @ZavenArra. You are right that the metrics-server needs more than an image, so let me show exactly what changes between the two charts and what stays the same. Short version: this PR does not swap the image on the old release. It replaces the whole Bitnami chart with the upstream kubernetes-sigs chart, and that chart ships every component itself.

1. Lineage and versions

Old (Bitnami) New (upstream)
Repo charts.bitnami.com/bitnami kubernetes-sigs.github.io/metrics-server
Chart bitnami/metrics-server 5.8.3 metrics-server/metrics-server 3.13.1
App (binary) v0.4.2 (2021) v0.8.1 (current)
Maintainer Bitnami, now retired kubernetes-sigs, the project itself

The upstream chart lives in the same repo as the metrics-server code, so the chart and the binary release together. The Bitnami chart was a third-party repackage.

2. The image, which is the real root cause

  • Bitnami default: docker.io/bitnami/metrics-server:0.4.2-debian-10-r30 (Debian rebuild with an -rNN suffix).
  • Upstream: registry.k8s.io/metrics-server/metrics-server:v0.8.1 (distroless, project-owned).

In August 2025 Bitnami retired its free catalog and removed that exact tag, so the pod went to ImagePullBackOff. The upstream image lives on registry.k8s.io, which the Kubernetes project runs and does not retire. That durability is the core reason to switch, not just the version bump.

3. Components each chart creates

Both charts create the full stack. This is your point directly.

Component Bitnami 5.8.3 Upstream 3.13.1
Deployment deployment.yaml deployment.yaml
Service svc.yaml service.yaml
ServiceAccount serviceaccount.yaml serviceaccount.yaml
APIService v1beta1.metrics.k8s.io metrics-api-service.yaml apiservice.yaml
ClusterRole cluster-role.yaml clusterrole.yaml + clusterrole-aggregated-reader.yaml
ClusterRoleBinding metrics-server-crb.yaml clusterrolebinding.yaml
auth-delegator binding auth-delegator-crb.yaml clusterrolebinding-auth-delegator.yaml
auth-reader RoleBinding (kube-system) role-binding.yaml rolebinding.yaml
PodDisruptionBudget pdb.yaml pdb.yaml

The upstream chart ships a few more templates (psp.yaml, servicemonitor.yaml, certificate.yaml, addon-resizer "nanny" templates), but they are opt-in and stay off by default. psp.yaml needs rbac.pspEnabled (PodSecurityPolicy was removed in Kubernetes 1.25, so it must stay off on 1.33). The nanny, ServiceMonitor, and Certificate templates need addon-resizer, Prometheus-Operator, or cert-manager, all off by default. So on our cluster the upstream chart renders the same functional set that Bitnami did.

4. How flags are passed (the one real translation)

Bitnami took a map extraArgs and rendered --secure-port plus each entry, with no other baked-in args. Effective args under Bitnami:

--secure-port=<port>
--kubelet-insecure-tls=true
--kubelet-preferred-address-types=InternalIP

Upstream takes a list args and renders --secure-port, then a defaultArgs list, then our args:

# defaultArgs (chart default)
--cert-dir=/tmp
--kubelet-preferred-address-types=InternalIP,ExternalIP,Hostname
--kubelet-use-node-status-port
--metric-resolution=15s
# our args
--kubelet-insecure-tls
--kubelet-preferred-address-types=InternalIP

Upstream runs with extra sane defaults Bitnami never set (--cert-dir, --metric-resolution=15s, --kubelet-use-node-status-port). One side effect: --kubelet-preferred-address-types now appears twice and resolves to InternalIP,ExternalIP,Hostname,InternalIP. This is harmless because the resolver takes the first match and InternalIP is first, but I will set it in defaultArgs instead of appending via args to keep it clean.

5. Default differences that matter

  • apiService.create: Bitnami default is false, so the role had to set it true. Upstream default is already true, so that line is now redundant but harmless.
  • Affinity: Bitnami exposed both nodeAffinityPreset and raw affinity. Upstream exposes only raw affinity. Our PR uses the raw block, which both charts accept, so it carried over cleanly.

6. Kubernetes compatibility

  • v0.4.2 targeted about Kubernetes 1.19 to 1.21. It was never validated against 1.33.
  • v0.8.x officially supports Kubernetes 1.31 and newer. Our cluster runs 1.33, so 0.8.1 is the supported line.

Summary

The change is not an image swap on the same chart. It replaces a retired third-party repackage (old binary, deleted image, apiService.create off by default) with the project's own chart (current binary, durable registry.k8s.io image, sane default args, apiService.create on). Every component flagged is created and owned by Helm.

Not yet tested on dev. Our flow here is approve, then I run it on the dev cluster, then prod. So the next step after this merges is the dev run, and I will post the results in this thread (kubectl get pods -n metrics-server, kubectl get apiservice v1beta1.metrics.k8s.io, kubectl top nodes) before anything reaches prod.

One migration note I want to flag for that dev run: because we reuse the release name, this is a helm upgrade over the old Bitnami release, and atomic: true would roll back to that broken revision on failure. To avoid that trap I plan to helm uninstall the old release first, then install the upstream chart clean. I will confirm that path in dev and report back.

The upstream chart renders defaultArgs then args and appends duplicate
flags, so keeping --kubelet-preferred-address-types in args left the
container with InternalIP,ExternalIP,Hostname,InternalIP. Override
defaultArgs to set it once to InternalIP. Only --kubelet-insecure-tls
stays in args.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants