Repository navigation
Raise work_mem to 64MB for readonlyuser on treetracker to cut temp-file spill - #325
Open
arnoldcastro5000 wants to merge 2 commits into
Open
arnoldcastro5000 wants to merge 2 commits into
arnoldcastro5000 wants to merge 2 commits into
Conversation
The tile-map read queries spill large volumes of temporary files to disk because work_mem is too small for their sorts and hashes. Two 12h monitor runs show about 65GB of temp writes per window while replica memory sits near 30 percent, the signature of an under-sized work_mem. Add an operator-applied SQL runbook that sets work_mem to 64MB for the readonlyuser role on the treetracker database only. The scope avoids OOM on the smaller primary; the value is sized to the measured peak concurrency of 30 to 33 heavy queries. This is a reversible first step; the spatial indexes remain the real fix. The cyrilgdn/postgresql provider (1.22.0) cannot express a per-role, per-database configuration parameter, so this lives under database-grants/tuning/ as reviewable SQL rather than Terraform.
…eline Review against the captured query plans showed the temp spill comes from single-process Sort nodes, so work_mem applies at 1x; hash_mem_multiplier and parallel workers do not apply. The earlier "about 46 percent" figure assumed a hash and parallel model and over-stated the gain, especially for the large case1 spatial queries. Reword to a meaningful but partial reduction, largest for the mid-size queries, with the exact number to be measured on the validation fork. Also state the current baseline (work_mem 13MB, the DigitalOcean default on the 16GB replica) and the temp-write range (about 55 to 65GB per 12h window).
arnoldcastro5000
marked this pull request as ready for review
August 30, 2026 00:13
Collaborator
Author
|
tested on dev. configured on prod today. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Give the
readonlyuserrole a largerwork_mem(64MB, up from the cluster default) on thetreetrackerdatabase only. This lets the map-tile read queries do more of their sorting and grouping in memory instead of spilling to temporary files on disk, which is a large, safe, and fully reversible reduction in disk load. It is a first step, not the complete fix (see Follow-ups).Background (what is going wrong)
The tile server draws the map by running heavy read queries against a read replica of the
treetrackerdatabase. Many of these queries sort or group large amounts of spatial data. PostgreSQL gives each such operation a memory budget calledwork_mem. When an operation needs more thanwork_mem, PostgreSQL does not fail; instead it writes the overflow to temporary files on disk and reads them back. That disk traffic is slow and it competes with every other query on the replica.Two 12-hour monitoring runs show this clearly:
work_memthat is set too small: the database trades spare RAM for disk spill.case1) and a recursive query over the organization tree (a recursive CTE).The replica has 16GB of RAM. The current
work_memis 13MB (the DigitalOcean default on a 16GB node), so even moderate sorts spill. This change raises it to 64MB, about 5x.The change
readonlyuserrole, and only inside thetreetrackerdatabase. No other role, database, or the primary writer is affected. This is deliberate:work_memis allocated per operation per connection, so a database-wide or cluster-wide change would multiply memory use across unrelated workloads and risk an out-of-memory (OOM) condition, especially on the smaller 8GB primary.Expected effect
The captured query plans show the spill comes from Sort nodes running in a single process, so
work_memapplies at 1x here (hash_mem_multiplierand parallel workers do not apply to these sorts). At 64MB this gives a meaningful but partial reduction in temp-file writes, largest for the mid-size queries that just cross the current 13MB limit. Thecase1spatial queries, whose per-call working set is about 1.33GB, are far above any safework_memand improve only marginally; they need the spatial index, not more memory (see Follow-ups). The exact percentage should be measured on the validation fork withEXPLAIN (ANALYZE, BUFFERS); an earlier estimate of about 46 percent was withdrawn because it assumed a hash and parallel model that the plans do not support. In short: this PR is a low-risk win, but it is a mitigation, not the cure.How it is applied (and why this is not Terraform)
The
cyrilgdn/postgresqlprovider (v1.22.0) used indatabase-grants/has no resource for a per-role, per-database configuration parameter (theALTER ROLE ... IN DATABASE ... SETform), so this change cannot be expressed as a Terraform resource. It is committed as an operator-applied SQL runbook underdatabase-grants/tuning/prod/, and applied by a database administrator using thedoadminaccount. The CI service account is intentionally not permitted to change production data.Apply steps:
database-grants/tuning/prod/work_mem-readonlyuser.sqlagainst the production cluster asdoadmin.pgpooldeployment). The new value only takes effect for new backend sessions at login time; pgpool holds long-lived pooled connections, so they must be recycled to pick up 64MB.Rollback
Fully reversible, no restart, no data change:
ALTER ROLE readonlyuser IN DATABASE treetracker RESET work_mem;Then recycle pgpool again.
Verification
readonlyusersession after the pgpool recycle:SHOW work_mem;returns64MB.pg_stat_statementsand the monitor report) should drop noticeably for the mid-size queries. If there is no drop at all, re-check that pgpool actually recycled.Risk and blast radius
RESET.work_memuse is roughly 4 to 6GB on top ofshared_buffers(about 3.2GB), well under 16GB, so headroom is comfortable. Monitoring the replica memory after the pgpool recycle is the safety check.Follow-ups (not in this PR)
case1family (the real fix for the largest spillers).organization_childrenCTE.wallet.token.capture_id::textcast that defeats an existing index.work_memtoward 256MB once the indexes ease the concurrent pile-up.Note on Terraform activation
The
database-grantsTerraform is applied manually (there is no CI apply), and whether prod state is currently reconciled has not been confirmed from the remote state. This does not affect the change: the provider does not modelwork_mem, so aterraform applywill neither create nor revert it, and there is no drift on this parameter. If the role turns out to be Terraform-managed, no change to this PR is needed beyond noting that the role identity and grants stay in Terraform while this GUC stays here by provider limitation.Checklist
doadminaccess ready