Skip to content

[BUG] DLRM dictionary generation repeatedly hits GPU OOM during CSV read and aggregation #15671

Description

@yinqingh

Describe the bug
build: dlrm-etl-on-dataproc/800, 802, 803

Three independent nightly DLRM ETL runs failed during dictionary generation with the same GPU out-of-memory pattern.

The first failures occurred while reading CSV data on the GPU, followed by partial hash aggregation:
Table.readCSV → GpuCSVScan → RmmRapidsRetryIterator → DynamicGpuPartialAggregateIterator
RMM was already near its approximately 14.8 GB limit when relatively small allocations of 16–96 MB failed. Two runs subsequently lost executors with exit code 134 (Aborted (core dumped)). The remaining run continued reporting CUDA/RMM allocation failures until the 20-minute stage timeout was reached.

The same workload and configuration reproduced the issue in three independent runs using different cudf-spark snapshot revisions.

Error logs:

java.lang.OutOfMemoryError: Could not allocate native memory:
std::bad_alloc: out_of_memory: RMM failure:
Exceeded memory limit 14815068160;
Allocated bytes 14811834368;
Requested bytes 96126719

    at ai.rapids.cudf.Table.readCSV(Native Method)
    at ai.rapids.cudf.Table.readCSV(Table.java:949)
    at com.nvidia.spark.rapids.CSVPartitionReader.readToTable(GpuCSVScan.scala)
    at com.nvidia.spark.rapids.RmmRapidsRetryIterator.next(RmmRapidsRetryIterator.scala)
    at com.nvidia.spark.rapids.DynamicGpuPartialAggregateIterator.hasNext(GpuAggregateExec.scala)

Another run reported the CUDA allocator variant:

java.lang.OutOfMemoryError: Could not allocate native memory:
std::bad_alloc: out_of_memory:
CUDA error (failed to allocate 96126719 bytes):
cudaErrorMemoryAllocation: out of memory

Two runs later lost executors:

Container exited with a non-zero exit code 134.
Aborted (core dumped)

Environment details (please complete the following information)

  • Google Cloud Dataproc 2.3 on Ubuntu 22.04
  • Apache Spark 3.5.3
  • Scala 2.12
  • CUDA 12
  • NVIDIA T4 GPUs
  • cudf-spark and cuDF 26.08.0-SNAPSHOT

Activity

  1. abellina commented on Aug 18, 2026

    @abellina
    Collaborator

    It would be good to run this locally to see if we can repro this. If we can't repro it locally, then we could commit a test with this change, to see if the cloud repros it.

    This is likely a leak. We should add this to spark's executor (and driver if running local mode) extraJavaOptions -Dai.rapids.refcount.debug=true. This will complain loudly when a leak has happened, I can't imagine why this would start failing otherwise.

  2. yinqingh commented on Sep 16, 2026

    @yinqingh
    CollaboratorAuthor

    I reran this job with -Dai.rapids.refcount.debug=true in spark.executor.extraJavaOptions. The job still hit an RMM memory limit while reading CSV, and executor 8 subsequently reported eight pinned host buffer leaks with reference-count histories.

    [2026-09-16T03:20:13Z] [INFO] DeviceMemoryEventHandler: Device allocation of 32000000 bytes failed. Device store spilled 0 bytes. First attempt. Total RMM allocated is 14817513585 bytes.
    [2026-09-16T03:20:13Z] [WARNING] DeviceMemoryEventHandler: Device store exhausted, unable to allocate 32000000 bytes. Total RMM allocated is 14817513585 bytes.
    [2026-09-16T03:20:13Z] [WARNING] GpuSemaphore: Dumping stack traces. The semaphore sees 4 tasks, 4 threads are holding onto the semaphore.
    

    The first captured OutOfMemoryError on executor 8 was at 03:20:13.110220414 UTC, in Table.readCSV:

    [2026-09-16T03:20:13.110220414Z] java.lang.OutOfMemoryError: Could not allocate native memory: std::bad_alloc: out_of_memory: RMM failure at:_deps/rmm-src/cpp/src/mr/detail/limiting_resource_adaptor_impl.cpp:61: Exceeded memory limit 14819262464; Allocated bytes 14817631232; Requested bytes 32000000
    	at ai.rapids.cudf.Table.readCSV(Native Method) ~[rapids-4-spark_2.12-26.10.0-SNAPSHOT-cuda12.jar:?]
    	at ai.rapids.cudf.Table.readCSV(Table.java:855) ~[rapids-4-spark_2.12-26.10.0-SNAPSHOT-cuda12.jar:?]
    	at com.nvidia.spark.rapids.CSVPartitionReader$.$anonfun$readToTable$2(GpuCSVScan.scala:454) ~[spark-shared/:?]
    	at com.nvidia.spark.rapids.NvtxId.apply(NvtxRangeWithDoc.scala:84) ~[spark-shared/:?]
    	at com.nvidia.spark.rapids.NvtxIdWithMetrics$.apply(NvtxWithMetrics.scala:67) ~[spark-shared/:?]
    	at com.nvidia.spark.rapids.CSVPartitionReader$.$anonfun$readToTable$1(GpuCSVScan.scala:454) ~[spark-shared/:?]
    	at com.nvidia.spark.rapids.RmmRapidsRetryIterator$AutoCloseableAttemptSpliterator.next(RmmRapidsRetryIterator.scala:586) ~[spark-shared/:?]
    	at com.nvidia.spark.rapids.RmmRapidsRetryIterator$RmmRapidsRetryIterator.next(RmmRapidsRetryIterator.scala:749) ~[spark-shared/:?]
    	at com.nvidia.spark.rapids.RmmRapidsRetryIterator$RmmRapidsRetryAutoCloseableIterator.next(RmmRapidsRetryIterator.scala:634) ~[spark-shared/:?]
    	at com.nvidia.spark.rapids.RmmRapidsRetryIterator$.drainSingleWithVerification(RmmRapidsRetryIterator.scala:332) ~[spark353/:?]
    	at com.nvidia.spark.rapids.RmmRapidsRetryIterator$.withRetryNoSplit(RmmRapidsRetryIterator.scala:135) ~[spark353/:?]
    	at com.nvidia.spark.rapids.CSVPartitionReader$.readToTable(GpuCSVScan.scala:452) ~[spark-shared/:?]
    	at com.nvidia.spark.rapids.CSVPartitionReader.readToTable(GpuCSVScan.scala:516) ~[spark-shared/:?]
    	at com.nvidia.spark.rapids.CSVPartitionReader.readToTable(GpuCSVScan.scala:464) ~[spark-shared/:?]
    	at com.nvidia.spark.rapids.GpuTextBasedPartitionReader.readToTable(GpuTextBasedPartitionReader.scala:561) ~[spark353/:?]
    	at com.nvidia.spark.rapids.GpuTextBasedPartitionReader.$anonfun$readBatch$1(GpuTextBasedPartitionReader.scala:453) ~[spark353/:?]
    [... remaining stack frames omitted ...]
    

    A subsequent exception also reported Requested bytes 16000004, with the same limit (14819262464) and allocated-byte count (14817631232).

    At 03:20:15 UTC, executor 8 reported these eight pinned host buffer leaks:

    [2026-09-16T03:20:15Z] [ERROR] PinnedMemoryPool: A PINNED HOST BUFFER WAS LEAKED (ID: 780 7a11a8020000)
    [2026-09-16T03:20:15Z] [ERROR] PinnedMemoryPool: A PINNED HOST BUFFER WAS LEAKED (ID: 1416 7a12a80a0000)
    [2026-09-16T03:20:15Z] [ERROR] PinnedMemoryPool: A PINNED HOST BUFFER WAS LEAKED (ID: 1417 7a13a8130200)
    [2026-09-16T03:20:15Z] [ERROR] PinnedMemoryPool: A PINNED HOST BUFFER WAS LEAKED (ID: 1418 7a13280e0000)
    [2026-09-16T03:20:15Z] [ERROR] PinnedMemoryPool: A PINNED HOST BUFFER WAS LEAKED (ID: 1419 7a1168000000)
    [2026-09-16T03:20:15Z] [ERROR] PinnedMemoryPool: A PINNED HOST BUFFER WAS LEAKED (ID: 1420 7a1268080000)
    [2026-09-16T03:20:15Z] [ERROR] PinnedMemoryPool: A PINNED HOST BUFFER WAS LEAKED (ID: 1421 7a1228060000)
    [2026-09-16T03:20:15Z] [ERROR] PinnedMemoryPool: A PINNED HOST BUFFER WAS LEAKED (ID: 929 7a12e80c0000)
    

    Here is the allocation stack for buffer 780, through the CSV input buffering path (remaining frames omitted):

    [2026-09-16T03:20:15Z] [ERROR] MemoryCleaner: Leaked pinned host buffer (ID: 780): 2026-09-16 03:20:10.0600 UTC: INC
    java.base/java.lang.Thread.getStackTrace(Thread.java:1602)
    ai.rapids.cudf.MemoryCleaner$RefCountDebugItem.<init>(MemoryCleaner.java:419)
    ai.rapids.cudf.MemoryCleaner$Cleaner.addRef(MemoryCleaner.java:79)
    ai.rapids.cudf.MemoryBuffer.incRefCount(MemoryBuffer.java:267)
    ai.rapids.cudf.MemoryBuffer.<init>(MemoryBuffer.java:106)
    ai.rapids.cudf.HostMemoryBuffer.<init>(HostMemoryBuffer.java:189)
    ai.rapids.cudf.PinnedMemoryPool.tryAllocateInternal(PinnedMemoryPool.java:269)
    ai.rapids.cudf.PinnedMemoryPool.tryAllocate(PinnedMemoryPool.java:194)
    com.nvidia.spark.rapids.HostAlloc.tryAllocPinned(HostAlloc.scala:113)
    com.nvidia.spark.rapids.HostAlloc.tryAllocInternal(HostAlloc.scala:208)
    com.nvidia.spark.rapids.HostAlloc.alloc(HostAlloc.scala:267)
    com.nvidia.spark.rapids.HostAlloc.allocate(HostAlloc.scala:281)
    ai.rapids.cudf.HostMemoryBuffer.allocate(HostMemoryBuffer.java:127)
    ai.rapids.cudf.HostMemoryBuffer.allocate(HostMemoryBuffer.java:138)
    com.nvidia.spark.rapids.HostLineBufferer.<init>(GpuTextBasedPartitionReader.scala:118)
    com.nvidia.spark.rapids.FilterCsvEmptyHostLineBuffererFactory$$anon$1.<init>(GpuTextBasedPartitionReader.scala:105)
    com.nvidia.spark.rapids.FilterCsvEmptyHostLineBuffererFactory$.createBufferer(GpuTextBasedPartitionReader.scala:105)
    com.nvidia.spark.rapids.FilterCsvEmptyHostLineBuffererFactory$.createBufferer(GpuTextBasedPartitionReader.scala:102)
    com.nvidia.spark.rapids.GpuTextBasedPartitionReader.$anonfun$readPartFile$1(GpuTextBasedPartitionReader.scala:428)
    com.nvidia.spark.rapids.NvtxId.apply(NvtxRangeWithDoc.scala:84)
    com.nvidia.spark.rapids.GpuTextBasedPartitionReader.readPartFile(GpuTextBasedPartitionReader.scala:422)
    com.nvidia.spark.rapids.GpuTextBasedPartitionReader.$anonfun$readToTable$1(GpuTextBasedPartitionReader.scala:539)
    com.nvidia.spark.rapids.GpuMetric.ns(GpuMetrics.scala:514)
    com.nvidia.spark.rapids.GpuMetric.ns(GpuMetrics.scala:507)
    com.nvidia.spark.rapids.GpuTextBasedPartitionReader.readToTable(GpuTextBasedPartitionReader.scala:539)
    [... remaining stack frames omitted ...]
    
  3. abellina commented on Sep 16, 2026

    @abellina
    Collaborator

    The approach we are thinking of is to change the call to Table.readCSV to be able to split its input (HostLineBufferer produces a HostMemoryBuffer.. we would need to do some special handling but essentially read less CSV lines at a time). This would make this code more reliable to any batch size and fragmentation. @sameerz fyi.

  4. added
    reliabilityFeatures to improve reliability or bugs that severly impact the reliability of the plugin
    on Sep 16, 2026
  5. abellina commented on Sep 16, 2026

    @abellina
    Collaborator

    Also we should make sure those allocations we leaked on OOM are guarded by withResource or closeOnExcept.

  6. thirtiseven commented on Sep 29, 2026

    @thirtiseven
    Collaborator

    #16105 is adding split retry for CSV reader, to fix another OOM issue, which might also fix this issue.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bot_watchSlack bot watched issue for LLM analyzerbugSomething isn't workingreliabilityFeatures to improve reliability or bugs that severly impact the reliability of the plugin

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions