PARALLEL PROCESSING OF MASSIVE LIDAR DATASETS USING APACHE SPARK AND DISTRIBUTED SPATIAL INDEXING

Main Article Content

Nguyen Thi Huu Phuong
Pham Thi Hai Van
Nguyen Thi Thanh

Abstract

Efficiently processing terabyte-scale 3D Airborne Light Detection and Ranging (LiDAR) point clouds remains a critical challenge for smart city and Digital Twin infrastructures. While in-memory distributed engines like Apache Spark accelerate geospatial analytics, they suffer from severe performance bottlenecks caused by geometric-blind partitioning, data skewness workload imbalances, and excessive network shuffle latencies during 3D proximity queries. To resolve these systemic limitations, this paper introduces a terrain-aware, multi-tier distributed spatial indexing and parallel processing framework optimized for massive LiDAR datasets on Apache Spark. The proposed architecture operates through three integrated phases: first, a parallel binary stream parser maps raw data directly into unmanaged off-heap memory, establishing a high-throughput 3D Spatial DataFrame while neutralizing Java Virtual Machine serialization delays. Second, a dynamic, terrain-aware adaptive octree partitioning algorithm is executed at the Master tier, automatically optimizing global split boundaries based on localized point density and geomorphic ruggedness to enforce strict cluster load balancing. Finally, worker nodes autonomously synthesize in-memory local 3D R-Trees over their designated partitions. By coupling global pruning with sub-logarithmic local searches, the framework ensures that complex neighborhood queries execute entirely within local RAM footprints, reducing cross-node network communication substantially reducing (Shuffe cost = 0) Experimental evaluations on a high-performance cluster demonstrate that our framework yields exceptional computational throughput and near-linear speedup factors, establishing a robust computational infrastructure for large-scale geospatial big data analytics.

Article Details

Section

Articles

References

[1] J. Yu, J. Zhang, and Z. Liu, "Scalable and High-Performance Processing of Massive 3D Point Clouds on Apache Spark," IEEE Transactions on Big Data, vol. 10, no. 2, pp. 145–158, Apr. 2024.

[2] M. S. Khan, A. R. Mahmood, and T. Shahzad, "Distributed Spatial Indexing Techniques for Big Geospatial Data: A Review of Recent Advances (2020–2024)," ACM Computing Surveys, vol. 56, no. 8, pp. 201–225, Aug. 2024.

[3] X. Jia, J. Yu, and A. Shrestha, "Apache Sedona: A High-Performance Distributed Framework for Big Geospatial Data Processing," Proceedings of the VLDB Endowment, vol. 17, no. 12, pp. 3412–3425, Aug. 2024.

[4] L. Wang, Y. Chen, and X. Li, "An Adaptive Global-Local Indexing Structure for Terabyte-Scale LiDAR Point Clouds in Distributed Environments," Information Sciences, vol. 632, pp. 102–119, July 2023.

[5] R. Santos, P. Furtado, and J. Bernardino, "Mitigating Data Skewness in Parallel Spatial Joins on Cloud-Based Distributed Frameworks," Future Generation Computer Systems, vol. 151, pp. 42–56, Feb. 2024.

[6] H. Zhao, W. Xu, and Y. Liang, "Massive LiDAR Point Cloud Management and Querying for Large-Scale Digital Twins: An In-Memory Distributed Approach," Automation in Construction, vol. 160, p. 105312, Apr. 2024.

[7] Z. Zhang, J. Wang, and L. Qi, "Distributed Semantic and Spatial Processing of Terabyte-Scale 3D Point Clouds Using Cloud-Native Apache Spark Architectures," IEEE Transactions on Parallel and Distributed Systems, vol. 36, no. 3, pp. 512–527, Mar. 2025.

[8] H. Al-Mubaid, Y. Kim, and S. Prasad, "Next-Generation Distributed Spatial Indexing for Big Geospatial Data: A Comprehensive Evaluation of Hybrid Global-Local Models," ACM Computing Surveys, vol. 57, no. 2, pp. 114–139, Feb. 2025.

[9] T. George, D. Haynes, and E. Shook, "Extending Spatial DataFrame Engines for High-Performance 3D LiDAR Processing in Apache Sedona," Proceedings of the VLDB Endowment, vol. 18, no. 4, pp. 895–908, Dec. 2024.

[10] Y. Liu, X. Chang, and M. Zhou, "A Terrain-Aware Distributed Spatial Indexing Structure for Heterogeneous LiDAR Datasets on Cloud Frameworks," Information Sciences, vol. 689, p. 115420, Jan. 2025.

[11] K. Srinivasan and L. Gomez, "Dynamic Workload Balancing and Data Skewness Mitigation for Parallel Spatial Queries in In-Memory Distributed Engines," Future Generation Computer Systems, vol. 164, pp. 208–222, Mar. 2025.

[12] M. R. N. Asyraf et al., "In-Memory Distributed Management of Massive LiDAR Datasets for Large-Scale Digital Twins and Urban Flood Modeling," Automation in Construction, vol. 171, p. 106104, Mar. 2025.