Repository navigation
fix: pin airflow redis image to the bitnamilegacy archive - #322
Merged
arnoldcastro5000 merged 1 commit intoJul 19, 2026
Merged
arnoldcastro5000 merged 1 commit into
arnoldcastro5000 merged 1 commit into
Conversation
Airflow has been down since 2026-06-13 (Greenstand#321): the embedded Redis pod cannot pull public.ecr.aws/bitnami/redis:5.0.14-debian-10-r32 because Bitnami delisted its free public images. Without Redis (the Celery broker) the scheduler and workers crash loop and no DAG runs. This pins the Redis image explicitly to docker.io/bitnamilegacy/redis, Bitnami's read-only archive of the same images. The archive does not carry the old suffixed tag; it carries the bare 5.0.14 version alias (verified against the Docker Hub API on 2026-07-03), which points at the final rebuild of the exact Redis version that ran in production before the outage. The playbook previously set no image override, so the deployed release inherited a default that no longer exists anywhere pullable. Deploy note: do not roll this out before Greenstand/treetracker-tile-server#113 is merged and deployed. Restoring Airflow resumes the map pre-warm DAGs, and the tile server needs its query timeout in place first (see issue Greenstand#315 for the sequencing rationale).
dadiorchen
approved these changes
Jul 14, 2026
Collaborator
Author
|
Deployed to dev. Will monitor before deploying to prod. |
Collaborator
Author
|
Airflow redis restored in prod and running for 35h. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #321. Related to #315 (the tile server investigation) and #320 (metrics-server, which failed for the same root cause).
Background: why Airflow has been down since June 13
Airflow uses Redis as its message queue: the scheduler puts tasks into Redis and the workers take them out. No Redis means no tasks run at all.
On June 13 the cluster nodes were rebuilt and the Redis pod was rescheduled onto a fresh node, which had to pull the Redis container image again. The pull fails with ErrImagePull, because Bitnami (the image publisher) delisted its free public images in 2025. The image the cluster asks for,
public.ecr.aws/bitnami/redis:5.0.14-debian-10-r32, no longer exists at that address. Since then the scheduler has restarted 57+ times (its health check fails without a working queue) and the workers 550+ times. Every DAG has been stopped for three weeks.This is the same event that broke metrics-server (fixed in #320): two casualties of the one Bitnami delisting.
The change
One block added to the Airflow Helm values in the Ansible playbook, pinning the Redis image explicitly:
bitnamilegacyis Bitnami's official read-only archive on Docker Hub, created for exactly this situation: it holds the delisted images.5.0.14-debian-10-r32), but it carries the bare5.0.14version alias, which Bitnami used as the rolling name for the newest rebuild of that version. Same Redis 5.0.14 that production ran, so no version change and no upgrade risk (and the queue data is disposable regardless: the chart runs this Redis without persistence).The value structure (
redis.image.registry/repository/tag) was verified against the chart source: airflow chart 8.5.2 embeds thestable/redis10.5.7 chart, whose image settings are exactly these three fields.Why not migrate to an upstream image, like #320 did for metrics-server?
Deliberately different strategies for the two Bitnami casualties:
registry.k8s.ioimages), so Fix metrics-server: migrate from retired Bitnami chart to upstream #320 could move to the canonical source and be done with Bitnami entirely./opt/bitnami/redis/mounted-etc/, starts redis with--include /opt/bitnami/redis/etc/redis.conf, and reads its password from/opt/bitnami/redis/secrets/redis-password(verified in the chart templates). The officialredisimage on Docker Hub has none of those paths, so swapping it in would crashloop. Only a Bitnami-layout image works here, andbitnamilegacyis the archive of exactly those.The durable escape routes exist but are much bigger changes than a 3 week outage should wait for: either
externalRedis.*(run our own Redis outside the chart, on any image we choose) or an Airflow chart upgrade. Both are worth considering later; this PR is the minimal fix that brings Airflow back. Being frozen is acceptable for this image because Redis 5.0.14 was already frozen (it is past end of life); the real modernization path is the chart upgrade, tracked separately.IMPORTANT: do not deploy this before node-mapnik-1#113
Restoring Airflow restarts the map pre-warm DAGs, which send the heaviest database queries in the system to the tile server every 5 minutes. The tile server currently has no query time limit; that combination is what caused the original #315 incident. Two protections must be in place first:
productionbranch, so they are in force the moment Airflow wakes up.Merging THIS PR is safe at any time: the playbook change does nothing until someone runs the Ansible playbook against production.
Checks to run at deploy time
docker manifest inspect docker.io/bitnamilegacy/redis:5.0.14(the tag's existence was already confirmed via the Hub API on 2026-07-03; this re-confirms nothing moved). If the archive is ever removed, the fallback is to mirror the image into the greenstand Docker Hub organization and point this block there (a one line change).helm get values airflow -n airflow. The dead pod referenced an image this playbook never set, so the live release contains overrides the playbook does not know about; make sure running the playbook does not silently revert something else.Follow-up (not in this PR)
The Redis pod runs with no resource requests or limits, which puts it first in line for eviction under node memory pressure (the same weakness the tile server had, see #315). We deliberately did not guess values here: after restoration the pod will be measured under real load for a day or two, and a follow-up PR will add measured requests and limits.