Repository navigation
[BUG] DLRM dictionary generation repeatedly hits GPU OOM during CSV read and aggregation #15671
Description
Activity
- addedbugSomething isn't workingSomething isn't working? - Needs TriageNeed team to review and classifyNeed team to review and classify
on Aug 14, 2026 - removed? - Needs TriageNeed team to review and classifyNeed team to review and classify
on Aug 18, 2026 It would be good to run this locally to see if we can repro this. If we can't repro it locally, then we could commit a test with this change, to see if the cloud repros it.
This is likely a leak. We should add this to spark's executor (and driver if running local mode) extraJavaOptions
-Dai.rapids.refcount.debug=true. This will complain loudly when a leak has happened, I can't imagine why this would start failing otherwise.- addedbot_watchSlack bot watched issue for LLM analyzerSlack bot watched issue for LLM analyzer
on Sep 16, 2026 I reran this job with
-Dai.rapids.refcount.debug=trueinspark.executor.extraJavaOptions. The job still hit an RMM memory limit while reading CSV, and executor 8 subsequently reported eight pinned host buffer leaks with reference-count histories.[2026-09-16T03:20:13Z] [INFO] DeviceMemoryEventHandler: Device allocation of 32000000 bytes failed. Device store spilled 0 bytes. First attempt. Total RMM allocated is 14817513585 bytes. [2026-09-16T03:20:13Z] [WARNING] DeviceMemoryEventHandler: Device store exhausted, unable to allocate 32000000 bytes. Total RMM allocated is 14817513585 bytes. [2026-09-16T03:20:13Z] [WARNING] GpuSemaphore: Dumping stack traces. The semaphore sees 4 tasks, 4 threads are holding onto the semaphore.The first captured
OutOfMemoryErroron executor 8 was at 03:20:13.110220414 UTC, inTable.readCSV:[2026-09-16T03:20:13.110220414Z] java.lang.OutOfMemoryError: Could not allocate native memory: std::bad_alloc: out_of_memory: RMM failure at:_deps/rmm-src/cpp/src/mr/detail/limiting_resource_adaptor_impl.cpp:61: Exceeded memory limit 14819262464; Allocated bytes 14817631232; Requested bytes 32000000 at ai.rapids.cudf.Table.readCSV(Native Method) ~[rapids-4-spark_2.12-26.10.0-SNAPSHOT-cuda12.jar:?] at ai.rapids.cudf.Table.readCSV(Table.java:855) ~[rapids-4-spark_2.12-26.10.0-SNAPSHOT-cuda12.jar:?] at com.nvidia.spark.rapids.CSVPartitionReader$.$anonfun$readToTable$2(GpuCSVScan.scala:454) ~[spark-shared/:?] at com.nvidia.spark.rapids.NvtxId.apply(NvtxRangeWithDoc.scala:84) ~[spark-shared/:?] at com.nvidia.spark.rapids.NvtxIdWithMetrics$.apply(NvtxWithMetrics.scala:67) ~[spark-shared/:?] at com.nvidia.spark.rapids.CSVPartitionReader$.$anonfun$readToTable$1(GpuCSVScan.scala:454) ~[spark-shared/:?] at com.nvidia.spark.rapids.RmmRapidsRetryIterator$AutoCloseableAttemptSpliterator.next(RmmRapidsRetryIterator.scala:586) ~[spark-shared/:?] at com.nvidia.spark.rapids.RmmRapidsRetryIterator$RmmRapidsRetryIterator.next(RmmRapidsRetryIterator.scala:749) ~[spark-shared/:?] at com.nvidia.spark.rapids.RmmRapidsRetryIterator$RmmRapidsRetryAutoCloseableIterator.next(RmmRapidsRetryIterator.scala:634) ~[spark-shared/:?] at com.nvidia.spark.rapids.RmmRapidsRetryIterator$.drainSingleWithVerification(RmmRapidsRetryIterator.scala:332) ~[spark353/:?] at com.nvidia.spark.rapids.RmmRapidsRetryIterator$.withRetryNoSplit(RmmRapidsRetryIterator.scala:135) ~[spark353/:?] at com.nvidia.spark.rapids.CSVPartitionReader$.readToTable(GpuCSVScan.scala:452) ~[spark-shared/:?] at com.nvidia.spark.rapids.CSVPartitionReader.readToTable(GpuCSVScan.scala:516) ~[spark-shared/:?] at com.nvidia.spark.rapids.CSVPartitionReader.readToTable(GpuCSVScan.scala:464) ~[spark-shared/:?] at com.nvidia.spark.rapids.GpuTextBasedPartitionReader.readToTable(GpuTextBasedPartitionReader.scala:561) ~[spark353/:?] at com.nvidia.spark.rapids.GpuTextBasedPartitionReader.$anonfun$readBatch$1(GpuTextBasedPartitionReader.scala:453) ~[spark353/:?] [... remaining stack frames omitted ...]A subsequent exception also reported
Requested bytes 16000004, with the same limit (14819262464) and allocated-byte count (14817631232).At 03:20:15 UTC, executor 8 reported these eight pinned host buffer leaks:
[2026-09-16T03:20:15Z] [ERROR] PinnedMemoryPool: A PINNED HOST BUFFER WAS LEAKED (ID: 780 7a11a8020000) [2026-09-16T03:20:15Z] [ERROR] PinnedMemoryPool: A PINNED HOST BUFFER WAS LEAKED (ID: 1416 7a12a80a0000) [2026-09-16T03:20:15Z] [ERROR] PinnedMemoryPool: A PINNED HOST BUFFER WAS LEAKED (ID: 1417 7a13a8130200) [2026-09-16T03:20:15Z] [ERROR] PinnedMemoryPool: A PINNED HOST BUFFER WAS LEAKED (ID: 1418 7a13280e0000) [2026-09-16T03:20:15Z] [ERROR] PinnedMemoryPool: A PINNED HOST BUFFER WAS LEAKED (ID: 1419 7a1168000000) [2026-09-16T03:20:15Z] [ERROR] PinnedMemoryPool: A PINNED HOST BUFFER WAS LEAKED (ID: 1420 7a1268080000) [2026-09-16T03:20:15Z] [ERROR] PinnedMemoryPool: A PINNED HOST BUFFER WAS LEAKED (ID: 1421 7a1228060000) [2026-09-16T03:20:15Z] [ERROR] PinnedMemoryPool: A PINNED HOST BUFFER WAS LEAKED (ID: 929 7a12e80c0000)Here is the allocation stack for buffer 780, through the CSV input buffering path (remaining frames omitted):
[2026-09-16T03:20:15Z] [ERROR] MemoryCleaner: Leaked pinned host buffer (ID: 780): 2026-09-16 03:20:10.0600 UTC: INC java.base/java.lang.Thread.getStackTrace(Thread.java:1602) ai.rapids.cudf.MemoryCleaner$RefCountDebugItem.<init>(MemoryCleaner.java:419) ai.rapids.cudf.MemoryCleaner$Cleaner.addRef(MemoryCleaner.java:79) ai.rapids.cudf.MemoryBuffer.incRefCount(MemoryBuffer.java:267) ai.rapids.cudf.MemoryBuffer.<init>(MemoryBuffer.java:106) ai.rapids.cudf.HostMemoryBuffer.<init>(HostMemoryBuffer.java:189) ai.rapids.cudf.PinnedMemoryPool.tryAllocateInternal(PinnedMemoryPool.java:269) ai.rapids.cudf.PinnedMemoryPool.tryAllocate(PinnedMemoryPool.java:194) com.nvidia.spark.rapids.HostAlloc.tryAllocPinned(HostAlloc.scala:113) com.nvidia.spark.rapids.HostAlloc.tryAllocInternal(HostAlloc.scala:208) com.nvidia.spark.rapids.HostAlloc.alloc(HostAlloc.scala:267) com.nvidia.spark.rapids.HostAlloc.allocate(HostAlloc.scala:281) ai.rapids.cudf.HostMemoryBuffer.allocate(HostMemoryBuffer.java:127) ai.rapids.cudf.HostMemoryBuffer.allocate(HostMemoryBuffer.java:138) com.nvidia.spark.rapids.HostLineBufferer.<init>(GpuTextBasedPartitionReader.scala:118) com.nvidia.spark.rapids.FilterCsvEmptyHostLineBuffererFactory$$anon$1.<init>(GpuTextBasedPartitionReader.scala:105) com.nvidia.spark.rapids.FilterCsvEmptyHostLineBuffererFactory$.createBufferer(GpuTextBasedPartitionReader.scala:105) com.nvidia.spark.rapids.FilterCsvEmptyHostLineBuffererFactory$.createBufferer(GpuTextBasedPartitionReader.scala:102) com.nvidia.spark.rapids.GpuTextBasedPartitionReader.$anonfun$readPartFile$1(GpuTextBasedPartitionReader.scala:428) com.nvidia.spark.rapids.NvtxId.apply(NvtxRangeWithDoc.scala:84) com.nvidia.spark.rapids.GpuTextBasedPartitionReader.readPartFile(GpuTextBasedPartitionReader.scala:422) com.nvidia.spark.rapids.GpuTextBasedPartitionReader.$anonfun$readToTable$1(GpuTextBasedPartitionReader.scala:539) com.nvidia.spark.rapids.GpuMetric.ns(GpuMetrics.scala:514) com.nvidia.spark.rapids.GpuMetric.ns(GpuMetrics.scala:507) com.nvidia.spark.rapids.GpuTextBasedPartitionReader.readToTable(GpuTextBasedPartitionReader.scala:539) [... remaining stack frames omitted ...]The approach we are thinking of is to change the call to
Table.readCSVto be able to split its input (HostLineBuffererproduces aHostMemoryBuffer.. we would need to do some special handling but essentially read less CSV lines at a time). This would make this code more reliable to any batch size and fragmentation. @sameerz fyi.- addedreliabilityFeatures to improve reliability or bugs that severly impact the reliability of the pluginFeatures to improve reliability or bugs that severly impact the reliability of the plugin
on Sep 16, 2026 Also we should make sure those allocations we leaked on OOM are guarded by withResource or closeOnExcept.
#16105 is adding split retry for CSV reader, to fix another OOM issue, which might also fix this issue.
- added a commit that references this issue
on Oct 9, 2026
Describe the bug
build: dlrm-etl-on-dataproc/800, 802, 803
Three independent nightly DLRM ETL runs failed during dictionary generation with the same GPU out-of-memory pattern.
The first failures occurred while reading CSV data on the GPU, followed by partial hash aggregation:
Table.readCSV → GpuCSVScan → RmmRapidsRetryIterator → DynamicGpuPartialAggregateIterator
RMM was already near its approximately 14.8 GB limit when relatively small allocations of 16–96 MB failed. Two runs subsequently lost executors with exit code 134 (Aborted (core dumped)). The remaining run continued reporting CUDA/RMM allocation failures until the 20-minute stage timeout was reached.
The same workload and configuration reproduced the issue in three independent runs using different cudf-spark snapshot revisions.
Error logs:
Another run reported the CUDA allocator variant:
Two runs later lost executors:
Environment details (please complete the following information)