PIM-tree: A Skew-resistant Index for Processing-in-Memory

Hongbo Kang, Yiwei Zhao, Guy E. Blelloch, Laxman Dhulipala, Yan Gu, Charles McGuffey, Phillip B. Gibbons

Published: 2025, Last Modified: 06 Nov 2025VLDB J. 2025EveryoneRevisionsBibTeXCC BY-SA 4.0

Abstract: The performance of today’s in-memory indexes is bottlenecked by the memory latency/bandwidth wall. Processing-in-memory (PIM) is an emerging approach that potentially mitigates this bottleneck by enabling low-latency memory access whose aggregate memory bandwidth scales with the number of PIM nodes. There is an inherent tension, however, between minimizing inter-node communication and achieving load balance in PIM systems, in the presence of workload skew. This paper presents PIM-tree, an ordered index for PIM systems that achieves both low communication and high load balance, regardless of the degree of skew in data/queries. Our skew-resistant index is based on a novel division of labor between the multi-core host CPU and the PIM nodes, which leverages the strengths of each. We introduce push-pull search, which dynamically decides whether to push queries to a PIM-tree node (CPU \(\rightarrow \) PIM-node) or pull the node’s keys back to the CPU (PIM-node \(\rightarrow \) CPU) based on workload skew. Combined with other PIM-friendly optimizations (shadow subtrees and chunking), PIM-tree achieves high throughput, (guaranteed) low communication, and (guaranteed) high load balance, for batches of point queries, updates, and range scans. We implement the PIM-tree structure, in addition to prior proposed PIM indexes, on the latest PIM system from UPMEM, with 32 CPU cores and 2048 PIM nodes. On workloads with 500 million keys and batches of 1 million queries, the throughput using PIM-trees is up to \(69.7\times \) and \(59.1\times \) higher than the two best prior PIM-based methods. As far as we know these are the first implementations of ordered indexes on real PIM systems.

External IDs:dblp:journals/vldb/KangZBDGMG25