Jianheng Liu

SpaRV: Sparse Radiance Voxels for Scalable Geometrically Consistent LiDAR-Visual Mapping and Rendering

Jianheng Liu, Yunfei Wan, Fangming Cheng, Xiangyu Hong, Zuhao Zou, Haotian Li, Longji Yin, Chunran Zheng, Jiarong Lin, and Fu Zhang

The University of Hong Kong

Under review

Color
Scene:
Compare color with:

SpaRV builds a sparse radiance voxel map from LiDAR and cameras. The map stays geometrically consistent away from the captured trajectory. SpaRV also optimizes the camera poses and focal lengths against the map, and reconstructs a university campus captured in 63,582 images on one 24 GB GPU.

Abstract


Robotic simulation, navigation, and inspection need photorealistic and geometrically accurate maps. Radiance fields achieve photorealism, but building them from robot data raises three challenges. A robot's trajectory serves its task, not reconstruction, so images from few viewpoints leave the geometry under-constrained. Residual errors in camera poses and intrinsics blur the map. The map's GPU memory grows with scene size, which limits its level of detail in large scenes. We present SpaRV, a Sparse Radiance Voxel map rendered by differentiable ray marching, which lets irregular LiDAR scans supervise it alongside images. LiDAR depth and occupancy anchor the under-constrained geometry, and normal and smoothness regularizers keep it smooth between LiDAR points. To correct camera errors, SpaRV alternates map training with photometric optimization of the camera poses and focal lengths. To bound GPU memory in training and rendering, SpaRV holds one submap at a time at full resolution and the rest at coarser octree levels. Experiments show that SpaRV keeps the geometry consistent away from the captured trajectory and sharpens textures by optimizing the camera parameters. On one 24 GB GPU, SpaRV reconstructs a university campus captured in 63,582 images and simulates other cameras and LiDARs. We will open-source our implementation.

Video


The overview video will be posted here.

Method


Overview of the SpaRV pipeline
Overview of SpaRV. (1) Inputs: RGB images, LiDAR scans, and initial camera parameters. The top panel shows one frame as its allocated voxels (left) and the view rendered from them (right). (2) A granularity-driven cut selects voxels from the octree, and a differentiable ray marcher renders them into color and depth. (3) Training runs over the submaps coarse-to-fine and alternates map training with camera optimization. The trained map is used for novel view synthesis, surface reconstruction, and LiDAR simulation.

1Ray-marching rendering of sparse radiance voxels

SpaRV stores the scene as sparse voxels in an octree, allocated only around surfaces and refined where more detail is needed.

A differentiable ray marcher traverses the voxels along each ray independently. Camera pixels, LiDAR rays, and rays of non-pinhole cameras therefore supervise the map through one renderer. The rasterizer of SVRaster, in contrast, batches a pinhole pixel lattice and does not directly apply to LiDAR rays or non-pinhole cameras.

Sparse voxels on a two-object scene
Sparse voxels on a two-object scene. (a) Voxels are allocated only around the surfaces. (b) Finer voxels (darker) concentrate where the detail is required. (c) Trilinear interpolation of the eight vertex densities.
Rasterization versus ray marching
Rasterization versus ray marching. (a) Rasterization batches a pinhole pixel lattice. (b) Ray marching evaluates each ray independently.

2Geometry-grounded map training

Color supervision fits every camera ray but leaves the geometry under-constrained, since different density profiles can render the same pixel color. LiDAR depth supervision anchors the field at the measured range, but only along measured rays. Between these anchors, normal and smoothness regularizers keep the surface locally smooth.

Loss terms and the regions they constrain
Loss terms and the regions they constrain. (a) Color: a floater (dashed) and a surface-aligned density profile (solid) render the same pixel color. (b) LiDAR depth: the target distribution (rose) pulls the rendered depth distribution (blue) onto the measured range. (c) Normal and smoothness: between the LiDAR anchors, a rippling surface (blue, dashed) fits the images as well as a smooth one, and the regularizers smooth it (rose).

A LiDAR scan is sparse, and its field of view does not coincide with the camera's. Floaters off the measured rays therefore get no depth supervision. The point cloud still marks which voxels contain a measured surface. Occupied voxel guidance turns that occupancy into a termination target for camera rays that hit an occupied voxel.

Occupied voxel guidance
Occupied voxel guidance. Teal dots are input LiDAR points. (a) LiDAR depth supervision leaves floaters off the measured rays unconstrained. (b) Guiding camera rays to terminate in the first occupied voxel suppresses floaters in front of it. (c) Voxel adaptation removes floater voxels whose blending weights fall below the pruning threshold.
Without occupied voxel guidance, color
With occupied voxel guidance, color
Without guidanceWith guidance

The same view on an Oxford Spires scene, from maps trained without and with occupied voxel guidance. The color images are nearly the same, but without the guidance the depth and normal maps show a band of floaters over the lawn.

3Map-driven camera optimization

The cameras of robot data are metrically plausible but carry residual errors in every frame, and these errors blur the textures of the map. SpaRV alternates map training with photometric optimization of the pose and focal length of each camera against the current map. On the six FAST-LIVO2 scenes, camera optimization raises the training-view PSNR of SpaRV from 26.55 dB with the odometry cameras to 29.87 dB.

Drive, odometry cameras
Drive, optimized cameras
Odometry camerasOptimized cameras
Show optimized cameras vs. ground truth
Drive, optimized cameras
Drive, ground truth
Optimized camerasGround truth

SpaRV on a FAST-LIVO2 training view, trained with the odometry cameras (left) and with the cameras it optimized (right).

FAST-LIVO2 Drive along the training trajectory. 3DGS, SVRaster, and SpaRV, each trained on the odometry cameras (top) and on the cameras optimized by SpaRV (bottom).
Training-view PSNR↑ on FAST-LIVO2.
MethodLandmarkCBDCultureDriveSculptureStreetAvg.
odometry cameras
InstantNGP31.70525.31223.13118.37925.27623.42924.539
3DGS32.26022.03321.57125.57024.29426.83625.427
SVRaster31.61626.41924.74926.03524.63626.02526.580
SpaRV31.22325.92324.68126.87124.48326.13026.552
SpaRV (cam. opt.)36.19028.11426.51229.81728.54830.00729.865
optimized cameras
InstantNGP31.69325.18224.06419.62225.29023.59624.908
3DGS36.25623.15921.23127.00627.93730.65727.708
SVRaster36.01128.68024.75527.58728.41129.44329.148
LetsGo32.06627.10724.21817.71223.17728.40825.448
GS-SDF35.74428.58427.14328.29828.06630.49329.721
SpaRV36.40328.15726.60129.91428.60530.01329.949

Odometry cameras: every method is trained on the FAST-LIVO2 odometry poses, and SpaRV (cam. opt.) also runs camera optimization. Optimized cameras: every method is trained on the cameras that SpaRV (cam. opt.) optimized.

4Large-scale reconstruction under bounded GPU memory

The GPU memory of a map grows with the extent of the scene. A single view, however, needs fine detail only near the camera, since distant regions project to few pixels. The octree of SpaRV keeps every coarser level that training passes through, and a coarser voxel needs no further optimization to be rendered as a level of detail. A granularity-driven cut selects voxels by the number of pixels they span on the image plane. It holds the submap in focus at the finest available level and the rest of the scene at coarser levels. The same cut rule serves training and rendering.

Large-scale training and deployment
Large-scale training and deployment. (1) Coarse training subdivides the octree over the whole scene up to the memory budget and keeps each coarser level as frozen nodes. (2) Submap training trains one submap at a time, with its relevant views, on a cut: the interior (blue) is subdivided, the exterior (violet) stays frozen, and grey nodes are not in the cut. (3) Deployment renders from the cut around the deploy box that contains the camera.
Octree level 10
Level 10
Octree level 11
Level 11
Octree level 12
Level 12
Octree level 16
Level 16
Cut at tau 1
τ = 1
Cut at tau 5
τ = 5
Cut at tau 9
τ = 9
Cut at tau 17
τ = 17

Levels of detail of one octree, seen from a single viewpoint. Top: increasingly fine levels. Bottom: cuts at an increasing granularity target τ, in pixels.

Results


We compare SpaRV with InstantNGP, 3DGS, SVRaster, LetsGo, and GS-SDF on synthetic indoor scenes and two outdoor LiDAR-visual datasets, with H3DGS on two large scenes we collected, and with EgoNeRF and OmniGS on an omnidirectional benchmark. In the tables, bold is the best result and underlined the second best.

Rendering away from the captured trajectory

A robot plans, simulates, and navigates beyond the views it captured, so the map is expected to stay geometrically consistent at those viewpoints. On FAST-LIVO2 Street, the methods are hard to tell apart at an interpolated view and separate once the camera moves off the captured trajectory. The image-only methods (InstantNGP, 3DGS, SVRaster) degrade. The Gaussian-based LetsGo and GS-SDF show popping on the billboard, where SpaRV stays consistent.

Baseline, extrapolated view of Street
SpaRV, extrapolated view of Street
3DGSSpaRV (Ours)
Compare with:

FAST-LIVO2 Street at an extrapolated view away from the captured trajectory. Pick a baseline and drag the divider.

All six methods on Street, with the viewpoint moved away from the captured trajectory.

Replica

Replica is a synthetic indoor dataset with ground-truth camera poses, which remove camera error from the comparison. We evaluate 8 scenes at interpolated viewpoints along the trajectory and at extrapolated viewpoints sampled uniformly in each scene. Pixels that no training view sees are excluded from the extrapolated evaluation.

Baseline, Replica room-0
SpaRV, Replica room-0
3DGSSpaRV (Ours)
Compare with:

Replica Room-0. Every method is close to the reference at the interpolated training view; at the extrapolated view of the same maps, the methods separate.

Rendering on Replica, averaged over the 8 scenes.
MethodExtrapolated (E)Interpolated (I)
PSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓
InstantNGP35.6430.9440.15033.7130.9330.154
3DGS28.7140.9220.16140.4390.9740.095
SVRaster33.3470.9460.13539.6000.9750.090
LetsGo25.0510.8860.19838.8920.9650.122
GS-SDF36.0530.9620.10340.5480.9730.078
SpaRV36.6900.9630.07741.0010.9820.047

FAST-LIVO2

FAST-LIVO2 is a real-world LiDAR-visual dataset captured with a solid-state Livox AVIA LiDAR. Its non-repetitive scan pattern makes the range measurements of each frame sparse and irregular. We use six outdoor scenes with three trajectory types: forward-facing (Landmark, CBD, Street), object-centric (Sculpture), and free-view (Culture, Drive). The initial cameras come from the FAST-LIVO2 odometry.

The six FAST-LIVO2 scenes
The six FAST-LIVO2 scenes: LiDAR map, trajectory, and a SpaRV render per scene.
SpaRV (Ours)
Compare with:

Fly-throughs along the test-view trajectory, with every method trained on the optimized cameras. Pick a scene and a baseline, and drag the divider.

On a test split that holds out one image in every eight, SpaRV has the lowest LPIPS on all six scenes and the highest SSIM on five of them. Its average PSNR is 0.13 dB below GS-SDF.

FAST-LIVO2 test split under the optimized cameras, averaged over the six scenes.
MethodPSNR↑SSIM↑LPIPS↓
InstantNGP25.3500.7610.325
3DGS27.1980.8450.224
SVRaster28.0420.8730.168
LetsGo28.3950.8810.164
GS-SDF28.8020.8920.149
SpaRV28.6760.9110.103

Oxford Spires

Oxford Spires has 12 outdoor sequences around Oxford landmarks, captured with a 64-channel Hesai QT64 LiDAR and three synchronized global-shutter cameras. Its ground-truth trajectories are registered against a millimeter-accurate laser scan, so this benchmark compares the methods with accurate cameras. SpaRV leads on all three averaged metrics, and its LPIPS is the lowest on all 12 sequences.

Baseline, Oxford Spires
SpaRV, Oxford Spires
3DGSSpaRV (Ours)
Compare with:

Oxford Spires, averaged over the 12 sequences.
MethodPSNR↑SSIM↑LPIPS↓
InstantNGP24.1020.7720.428
3DGS25.7590.7960.366
SVRaster25.2860.8140.302
LetsGo25.7140.8350.304
GS-SDF25.8800.8170.329
SpaRV26.8370.8570.232

Large scenes: Garden and Campus

We captured Garden and Campus with a handheld scanner that carries a solid-state Livox MID360 LiDAR and three fisheye cameras. Garden is an outdoor district densely covered with complex objects: 44,550 images and 146.7 M LiDAR points along a 2,239 m trajectory. Campus spans an entire university: 63,582 images and 151.6 M LiDAR points along a 3,534 m trajectory. We compare SpaRV at two stages with H3DGS. SpaRV (coarse) is the output of coarse training, and SpaRV (deploy) adds submap training and renders through a cut around the camera.

Garden and Campus reconstructed by SpaRV
Garden (top) and Campus (bottom), reconstructed by SpaRV on a single GPU. Each scene is drawn as the accumulated LiDAR point cloud of the capture, viewed from above. Beside each red dashed box, points is the point cloud of that region seen from a free viewpoint and render is the SpaRV rendering from the same viewpoint.

SpaRV (coarse), H3DGS, and SpaRV (deploy) along the whole fly-through of each scene. Campus plays at 2–3× speed.

Submap training improves SpaRV (coarse) on every metric, most on Campus (+2.5 dB PSNR). On Campus, SpaRV (deploy) renders at 68.8 FPS with a peak of 15.6 GiB of GPU memory, against 29.6 FPS and 22.5 GiB for H3DGS.

Rendering quality on Garden and Campus, evaluated on the training views.
SceneMethodPSNR↑SSIM↑LPIPS↓
GardenH3DGS18.020.6200.282
SpaRV (coarse)20.890.7120.346
SpaRV (deploy)21.120.7130.295
CampusH3DGS20.010.6600.295
SpaRV (coarse)20.620.6520.451
SpaRV (deploy)23.160.7600.283
Training and rendering cost on Campus (3,960-frame fly-through at 1280×960).
MethodTrain [h]↓FPS↑Peak VRAM [GiB]↓
TrainRender
H3DGS7.329.68.322.5
SpaRV (coarse)0.8248.412.54.3
SpaRV (deploy)8.768.811.715.6

Omnidirectional cameras: EgoNeRF

SpaRV renders equirectangular cameras with the same ray marcher, and uses the renderer and training schedule of the pinhole experiments unchanged. OmniGS, in contrast, specializes its rasterizer for the equirectangular camera. SpaRV outperforms OmniGS on all three metrics on both splits. The PSNR margin is small (0.34 dB on OmniBlender, 0.15 dB on Ricoh360), while LPIPS drops by 26% and 11%.

Novel view synthesis on the test frames of the EgoNeRF dataset.
MethodOmniBlender (11 scenes)Ricoh360 (11 scenes)
PSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓
EgoNeRF30.930.8920.20724.890.7520.353
OmniGS33.090.9180.17425.990.8250.262
SpaRV33.430.9260.12926.140.8380.233

Applications


Sensor simulation

The ray marcher accepts arbitrary rays, so one trained map can render other cameras and LiDARs. From the Garden map, SpaRV renders three camera models (pinhole, fisheye, and omnidirectional) and three LiDAR scan patterns (Livox AVIA, Livox MID360, and an OS128-style ring pattern).

Sensor simulation from one trained map: three camera models (top) and three LiDAR scan patterns (bottom).

Mesh extraction

Mesh of Street
Street
Mesh of CBD
CBD
Mesh of Landmark
Landmark
Mesh of Sculpture
Sculpture
Mesh of Culture
Culture
Mesh of Drive
Drive

Meshes extracted from the six FAST-LIVO2 maps by marching cubes.

BibTeX


@misc{liu2026sparv,
  title   = {SpaRV: Sparse Radiance Voxels for Scalable Geometrically Consistent LiDAR-Visual Mapping and Rendering},
  author  = {Liu, Jianheng and Wan, Yunfei and Cheng, Fangming and Hong, Xiangyu and Zou, Zuhao and
             Li, Haotian and Yin, Longji and Zheng, Chunran and Lin, Jiarong and Zhang, Fu},
  note    = {Under review},
  year    = {2026}
}