3D-USE logo3D-USE: From Image-Level to Scene-Level Underwater Enhancement

1College of Computer Science, Nankai University 2NKIARI, Shenzhen Futian
3School of Automation, Southeast University 4Advanced Ocean Institute of Southeast University, Nantong
*Corresponding author
TL;DR

We introduce 3D-USE, which shifts the underwater enhancement objective from physical radiance inversion to scene-level appearance learning. Without paired enhanced 3D supervision, it transfers 2D UIE knowledge into a persistent, view-consistent 3D scene.

2D3D

via Appearance Transition Consensus

00

Motivation

Why view-wise enhancement cannot define a consistent 3D scene.

Limitations of physical inversion, image pre-processing, and direct 3D supervision
Why scene-level enhancement is challenging. Degraded multi-view images do not uniquely determine a desired enhanced 3D appearance: physical inversion is under-constrained, image-wise enhancement is inconsistent across views, and paired enhanced 3D data are unavailable.
  • (a) Physical inversion. Different scene-radiance and water-medium decompositions can reproduce the captured views but leave medium-dependent artifacts in the recovered appearance.
  • (b) Image-based pre-processing. Independently enhanced views introduce incompatible color and contrast cues as conflicting RGB supervision for one shared 3D scene.
  • (c) Direct 3D supervision. Although conceptually ideal, it requires paired enhanced 3D ground truth that is unavailable for real underwater scenes.
3D-USE transfers a paired 2D transition prior into a consistent enhanced 3D scene
3D-USE learns a scene-consistent appearance transition.
  • (d) Scene-level appearance learning. 3D-USE learns global and local transition priors from external paired 2D data, reaches consensus across captured views, and embeds the resulting object and medium appearance in the 3D scene for persistent enhanced rendering.
01

Overview

Underwater enhancement should persist across viewpoints.

Abstract

Underwater 3D reconstruction faithfully reproduces the color shifts and visibility loss of captured views, while physical inversion may leave estimation errors in the recovered scene appearance. We formulate Underwater Scene-level Enhancement (USE) as learning a persistent, visibility-enhanced 3D scene representation from degraded multi-view underwater observations, enabling consistent enhanced rendering. Realizing USE requires both a reliable scene representation for enhancement and a consistent enhancement target without paired enhanced 3D data. Therefore, we present 3D-USE, a two-stage framework. First, MediumRBF establishes a medium-aware Gaussian scene by representing water effects with shared radial-basis anchors and explicitly decomposing object and medium contributions. Based on this fixed scene representation, Appearance Transition Consensus (ATC) transfers paired 2D UIE knowledge into scene-global and Gaussian-local targets, avoiding direct supervision from inconsistent enhanced views. An Underwater Bilateral Appearance Field (U-BAF) then realizes these targets in Gaussian radiance and medium appearance. The scene directly renders enhanced novel views without a 2D UIE model at inference. Experiments on real underwater scenes show improved visibility and cross-view consistency while preserving reconstruction quality.

01

Enhancement lives in 3D

A persistent enhanced Gaussian scene replaces independent post-processing of rendered views.

02

Medium-aware reconstruction

MediumRBF separates object appearance from water-induced effects over shared scene geometry.

03

Transition consensus

ATC consolidates 2D enhancement knowledge into scene-global and Gaussian-local guidance.

02

Method

A two-stage path from reconstruction to enhancement.

Overview of the 3D-USE pipeline
Method overview Reserved for the final pipeline PNG Save as media/figures/pipeline.png
Stage 1

Medium-aware reconstruction

MediumRBF models water-induced effects through shared radial-basis anchors, establishing explicit object and medium components over reliable Gaussian geometry.

Stage 2

Scene-level appearance enhancement

ATC forms fixed scene-global and Gaussian-local targets. U-BAF realizes their consensus in object and medium appearance before physics-based rasterization.

03

Results

Compare reconstruction fidelity and enhanced appearance.

Qualitative videos first show reconstruction and enhancement along matched camera paths, followed by quantitative comparisons on the evaluation benchmarks.

Quantitative Comparison

Reconstruction and enhancement benchmarks

Results are reported on held-out views, with scene-macro averaging for the multi-scene benchmarks.

Qualitative Comparison

Matched-view reconstruction and enhancement

Use the synchronized controls to inspect geometry, fine texture, color correction, and visibility.

Enhancement Quality

Scene-level enhancement across SeaThru-4 and D3

Two synchronized videos show the complete frame for judging color correction, visibility, and spatial uniformity.

Scene
Compare against
Stage 1IUI3-RedSea
Video pendingStage 1
3D-USE (Ours)IUI3-RedSea
Video pending3D-USE (Ours)
Synchronized playback

Quantitative enhancement

Enhanced novel-view quality and consistency

Higher is better, except wLPIPS.

Method SeaThru-4 DRUVA-20 D3 D5 wLPIPS
โ†“
UCIQEURankerMUSIQ UCIQEURankerMUSIQ UCIQEURankerMUSIQ UCIQEURankerMUSIQ
WaterSplatting0.50831.36158.5290.385โˆ’0.85633.1930.4700.54158.8330.4990.69055.5680.173
SeaSplat0.58791.55554.4700.518โˆ’0.18628.2520.5041.00758.2510.442โˆ’0.89540.6480.209
Plenodium0.50341.05858.3700.395โˆ’0.79234.7860.4830.70560.1580.4430.03352.4560.158
MarineSTD-GS0.52190.91855.8900.447โˆ’0.56739.0700.4430.41656.2460.413โˆ’0.11754.4240.216
3D-UIR0.60871.65854.5400.5370.17432.7520.5191.35659.3510.4680.13751.9910.208
3D-USE (Ours)0.60851.67959.8810.5750.45636.5980.5432.03462.9930.5131.02758.3240.147

SeaThru-4 and DRUVA-20 are scene-macro averages; wLPIPS is averaged over all 26 scenes. Bold and underline denote best and second-best results.

Reconstruction Quality

Matched-view reconstruction and detail inspection

Play the matched camera path to inspect structural stability, then move the detail region to examine boundaries and fine textures.

Scene
Compare against
IUI3-RedSea

Play or drag horizontally to compare.

Video pending3D-USE (Ours)
Video pending3DGS
3DGS 3D-USE (Ours)
Selected detail3DGS
Detail preview
Selected detail3D-USE (Ours)
Detail preview

All methods use matched viewpoints and their native underwater reconstruction outputs.

Quantitative reconstruction

Held-out reconstruction fidelity

Higher is better, except LPIPS.

Method SeaThru-4 DRUVA-20 D3 D5 FPS
โ†‘
PSNRSSIMLPIPS PSNRSSIMLPIPS PSNRSSIMLPIPS PSNRSSIMLPIPS
WaterSplatting29.3400.9190.12927.0070.8640.35727.3800.7760.25329.1130.8170.31399.78
SeaSplat27.1460.8780.18422.8660.8280.38827.8850.7850.28227.1130.8060.382100.77
MarineSTD-GS30.0340.9010.16026.4720.8560.36429.5100.7810.31630.2370.8180.32941.13
Plenodium30.3310.9220.12527.4130.8660.33528.6860.8030.21829.0880.8170.31491.60
3D-UIR28.2600.8890.17925.4320.8200.35428.0980.8300.28328.9730.8230.32442.74
3D-USE (Ours)30.3690.9240.12427.4950.8670.35229.7590.8580.18331.4000.9080.235105.36

SeaThru-4 and DRUVA-20 are scene-macro averages. Bold and underline denote best and second-best results.

Citation

BibTeX

@misc{3duse,
  title     = {3D-USE: From Image-Level to Scene-Level Underwater Enhancement},
  author    = {Yuan, Jieyu and Zhang, Yuanlin and Li, Jihong and Guo, Chun-Le and Lu, Huimin and Li, Chongyi},
  note      = {Project page}
}