[Paper Review, KR] FlashWorld: High-Quality 3D Scene Generation within Seconds
Paper Information
Title: FlashWorld: High-Quality 3D Scene Generation within Seconds
Authors: Xinyang Li, Tengfei Wang, Zixiao Gu, Shengchuan Zhang, Chunchao Guo, Liujuan Cao
Venue: ICLR 2026 Oral
Link: [Paper], [Project], [Github]
Teaser Image (Poster)
Introduction
3D Generation ๋ถ์ผ๋ ํฌ๊ฒ ์ฑ์ฅํ๊ณ ์๋ ๋ถ์ผ์ด์ง๋ง, scarcity of high-quality 3D scene data์ exponential complexity of modeling real-world scenes๋ผ๋ ๋ ๊ฐ์ ํฐ ์ฅ์ ๋ฌผ ๋๋ฌธ์ ์ด๋ ค์์ ๊ฒช๊ณ ์๋ค๊ณ ๋งํ๋ค.
์ฌ๊ธฐ์๋ ํฌ๊ฒ 2๊ฐ์ ํจ๋ฌ๋ค์์ด ์๋ค.
๋จผ์ multi-view-oriented(MV-oriented) ํ์ดํ๋ผ์ธ์ด๋ค. diffusion model์ด ํ ์คํธ๋ ์ฐธ์กฐ ์ด๋ฏธ์ง๋ก๋ถํฐ ์ฌ๋ฌ ์์ ์ ์ด๋ฏธ์ง๋ค์ ๋จผ์ ์์ฑํ ๋ค์, 3D reconstruction์ ์ํํ๋ ๋ฐฉ์์ด๋ค. ๊ทธ๋ฌ๋ ์์ ํฉ์ฑ ๊ณผ์ ์์ ๋ช ์์ ์ธ 3D ์ ์ฝ ์กฐ๊ฑด์ด ์์ด geometric ํน์ semantic inconsistencies๋ฅผ ๋ฐ์์ํฌ ์ ์๋ค. ๊ฒ๋ค๊ฐ ์ด๋ ์๋นํ computational overhead๊ฐ ๋ฐ์ํ๊ณ ์์ฑ ์๊ฐ๋ ์๋นํ๋ค๋ ๊ฒ์ ์ ์ ์๋ค.
diffusion model์ ํจ์จ์ฑ์ ๋์ด๊ธฐ ์ํด, post-training distillation ๊ธฐ์ ๋ค์ด ์์ฃผ ์ฌ์ฉ๋๋ค. ์ด๋ฌํ distillation ๊ธฐ๋ฒ์ ์ง์ ์ ์ฉํ๋ฉด ํ๋ ์์ํฌ๊ฐ ๋ณธ์ง์ ์ผ๋ก ๊ฐ์ง ํ๊ณ์ ์ ์คํ๋ ค ์ฆํญ์ํค๊ฒ ๋ ์ ์๋ค.
๋ค์์ผ๋ก 3D-oriented ํจ๋ฌ๋ค์์ด๋ค. ์ด ๋ฐฉ์์ diffusion model์ด๋ ๋ฏธ๋ถ๊ฐ๋ฅํ rendering์ combineํ๋ ๊ฒ์ด๋ค. ์ด ๋ฐฉ์์ ๋ฌผ์ฒด๋ ๋ฐฐ๊ฒฝ์ ๊ธฐํํ์ ํํ๊ฐ ์ด๊ธ๋์ง ์๊ณ ๋ฌผ๋ฆฌ์ ์ผ๊ด์ฑ์ ์ ์งํ๋ ํ์ง์ด ๋ค์ ํ๋ฆฟํด์ง๋ ๋ฌธ์ ๊ฐ ์๋ค. ๊ฒ๋ค๊ฐ refinement stage๋ฅผ ์ถ๊ฐ๋ก ํ์๋กํ๋ค.
Preliminary
FlashWorld์ ํต์ฌ์ธ cross-mode post-training์ ์ดํดํ๊ธฐ ์ํด์๋ ๋จผ์ Diffusion Model๊ณผ Distribution Matching Distillation (DMD)์ ๋ํ ์ดํด๊ฐ ํ์ํ๋ค.
Diffusion Model
Diffusion model์ ์ผ๋ฐ์ ์ผ๋ก Gaussian noise์์ ์์ํ์ฌ ์ ์ง์ ์ผ๋ก noise๋ฅผ ์ ๊ฑฐํ๋ฉด์ target data distribution์ sample์ ์์ฑํ๋ค.
์๋ณธ ๋ฐ์ดํฐ \(x\)์ timestep \(t\)์ ๋ฐ๋ฅธ Gaussian noise๋ฅผ ์ถ๊ฐํ๋ forward process๋ ๋ค์๊ณผ ๊ฐ์ด ์ ์๋๋ค:
\[x_t = F(x,t) = \alpha_t x + \sigma_t \epsilon, \qquad \epsilon \sim \mathcal{N}(0,I)\]์ฌ๊ธฐ์ \(\alpha_t\)์ \(\sigma_t\)๋ timestep \(t\)์ ๋ฐ๋ฅธ signal๊ณผ noise์ ๋น์จ์ ๊ฒฐ์ ํ๋ค.
์ฆ,
\[x_t = \underbrace{\alpha_t x}_{\text{signal}} + \underbrace{\sigma_t\epsilon}_{\text{noise}}\]๋ก ๋ณผ ์ ์๋ค.
Denoising network๋ noisy sample \(x_t\)์ timestep \(t\)๋ฅผ ์ ๋ ฅ์ผ๋ก ๋ฐ์ ์๋์ clean data \(x\)๋ฅผ ์์ธกํ๋๋ก ํ์ต๋๋ค.
\[\mathcal{L} = \mathbb{E}_{x,t,\epsilon} \left[ \left\| x-\hat{x}_\theta(x_t,t) \right\|^2 \right]\]์ ์์์์๋ clean data \(x\)๋ฅผ ์ง์ ์์ธกํ๋ \(x\)-prediction์ ์ฌ์ฉํ์ง๋ง, diffusion model์ noise \(\epsilon\)์ ์์ธกํ๊ฑฐ๋ \(x\)์ \(\epsilon\)์ ์ ํ ๊ฒฐํฉ์ธ \(v\)๋ฅผ ์์ธกํ๋ ๋ฐฉ์์ผ๋ก๋ ํ์ต๋ ์ ์๋ค.
์ด๋ฌํ prediction๋ค์ ๋ชจ๋ denoised estimate \(\mu(x_t,t)\)๋ก ๋ณํํ ์ ์์ผ๋ฉฐ, ์ด๋ฅผ ์ด์ฉํ๋ฉด distribution์ score๋ฅผ ๋ค์๊ณผ ๊ฐ์ด ํํํ ์ ์๋ค.
\[s(x_t,t) = \nabla_{x_t}\log p_t(x_t) = - \frac{x_t-\alpha_t\mu(x_t,t)} {\sigma_t^2}\]Score
\[s(x_t,t)=\nabla_{x_t}\log p_t(x_t)\]๋ ํ์ฌ sample \(x_t\)๊ฐ ํด๋น data distribution์์ probability๊ฐ ๋ ๋์ ์์ญ์ผ๋ก ์ด๋ํ๋ ค๋ฉด ์ด๋ ๋ฐฉํฅ์ผ๋ก ์์ง์ฌ์ผ ํ๋์ง๋ฅผ ๋ํ๋ด๋ gradient๋ผ๊ณ ์ดํดํ ์ ์๋ค.
์ฆ, diffusion model์ ๋จ์ํ denoising ๊ฒฐ๊ณผ๋ฅผ ์์ธกํ๋ ๊ฒ๋ฟ๋ง ์๋๋ผ, ํ์ฌ sample์ data distribution์ ๋ ๊ฐ๊น์ด ๋ฐฉํฅ์ผ๋ก ์ด๋์ํค๊ธฐ ์ํ score field๋ฅผ ์ ๊ณตํ ์ ์๋ค.
Distribution Matching Distillation (DMD)
Distribution Matching Distillation (DMD)์ ๋ง์ denoising step์ด ํ์ํ diffusion model์ ์ ์ step๋ง์ผ๋ก generation์ ์ํํ๋ generator๋ก distillationํ๊ธฐ ์ํ ๋ฐฉ๋ฒ์ด๋ค.
๊ธฐ์กด diffusion teacher๊ฐ
\[z \rightarrow x_{T-1} \rightarrow x_{T-2} \rightarrow \cdots \rightarrow x_0\]์ฒ๋ผ ์ฌ๋ฌ ๋ฒ์ denoising step์ ๊ฑฐ์ณ sample์ ์์ฑํ๋ค๋ฉด, DMD์ ๋ชฉ์ ์ few-step student generator \(G_\theta\)๊ฐ ์์ฑํ๋ distribution์ teacher์ target distribution๊ณผ ์ผ์น์ํค๋ ๊ฒ์ด๋ค.
์ฆ,
\[p_{\text{fake}} \rightarrow p_{\text{real}}\]์ด ๋๋๋ก student generator๋ฅผ ํ์ตํ๋ค.
์ฌ๊ธฐ์, \(p_{\text{real}}\)๋ teacher diffusion model์ด ํํํ๋ target distribution, $p_{\text{fake}}$๋ ํ์ฌ student generator \(G_\theta\)๊ฐ ์์ฑํ๋ distribution
์ ์๋ฏธํ๋ค.
DMD์์๋ randomly sampled noise \(z\)๋ฅผ student generator์ ์ ๋ ฅํ์ฌ
\[x_{\text{fake}} = G_\theta(z)\]๋ฅผ ์์ฑํ๊ณ , ์ฌ๊ธฐ์ ๋ค์ timestep \(t\)์ ํด๋นํ๋ noise๋ฅผ ์ถ๊ฐํ๋ค.
\[x_t = F(G_\theta(z),t)\]์ด noisy sample์ ๋ํด real distribution๊ณผ fake distribution ๊ฐ๊ฐ์ score๋ฅผ ๊ณ์ฐํ๋ค.
\[s_{\text{real}}(x_t,t) = \nabla_{x_t} \log p_{\text{real}}(x_t)\] \[s_{\text{fake}}(x_t,t) = \nabla_{x_t} \log p_{\text{fake}}(x_t)\]DMD์ ํต์ฌ gradient๋ ๋ค์๊ณผ ๊ฐ์ด ๋ score์ ์ฐจ์ด๋ฅผ ์ด์ฉํ๋ค.
\[\nabla \mathcal{L}_{\mathrm{DMD}} = - \mathbb{E}_{t} \left[ \int \left( s_{\mathrm{real}} \left( F(G_\theta(z),t),t \right) - s_{\mathrm{fake}} \left( F(G_\theta(z),t),t \right) \right) \frac{dG_\theta(z)}{d\theta} \,dz \right]\]์ฌ๊ธฐ์ ํต์ฌ์ ์ธ ๋ถ๋ถ์
\[s_{\text{real}} - s_{\text{fake}}\]์ด๋ค.
Score์ ์ ์๋ฅผ ์ด์ฉํ๋ฉด
\[s_{\text{real}} - s_{\text{fake}} = \nabla_x\log p_{\text{real}}(x) - \nabla_x\log p_{\text{fake}}(x)\]์ด๋ฏ๋ก,
\[s_{\text{real}} - s_{\text{fake}} = \nabla_x \log \frac{p_{\text{real}}(x)} {p_{\text{fake}}(x)}\]๋ก ๋ณผ ์ ์๋ค.
๋ฐ๋ผ์ \(s_{\text{real}}-s_{\text{fake}}\)๋ ๋จ์ํ student์๊ฒ loss๋ฅผ ์ ๋ฌํ๊ธฐ ์ํด ์ฌ์ฉํ๋ ๊ฒ์ด ์๋๋ผ, ํ์ฌ student์ output distribution์ด real distribution๊ณผ ๋น๊ตํ์ ๋ ์ด๋ ๋ฐฉํฅ์ผ๋ก ์์ ๋์ด์ผ ํ๋์ง๋ฅผ ๋ํ๋ด๋ gradient๋ผ๊ณ ์ดํดํ ์ ์๋ค.
์ด๋ฅผ ๋ค์
\[\frac{dG_\theta(z)}{d\theta}\]๋ฅผ ํตํด generator parameter \(\theta\)๊น์ง ์ ๋ฌํจ์ผ๋ก์จ
\[p_{\text{fake}} \rightarrow p_{\text{real}}\]์ด ๋๋๋ก student generator๋ฅผ ํ์ตํ๋ค.
Real Score Model and Fake Score Model
DMD์์๋ \(s_{\text{real}}\)๊ณผ \(s_{\text{fake}}\)๋ฅผ ์ง์ ์ ์ ์๊ธฐ ๋๋ฌธ์ ๊ฐ๊ฐ diffusion model์ ์ด์ฉํ์ฌ score๋ฅผ ์ถ์ ํ๋ค.
Real score์ ๊ฒฝ์ฐ pretrained diffusion model
\[\mu_{\text{real}}\]์ ์ฌ์ฉํ๋ค.
$\mu_{\text{real}}$์ target data distribution์ ๋ํด ์ด๋ฏธ ํ์ต๋์ด ์์ผ๋ฏ๋ก training ๊ณผ์ ์์ frozen ์ํ๋ก ์ ์ง๋๋ค.
๋ฐ๋ฉด fake distribution์ student generator๊ฐ ํ์ต๋ ๋๋ง๋ค ๊ณ์ ๋ณํํ๋ค.
\[p_{\text{fake}}^{(0)} \neq p_{\text{fake}}^{(1)} \neq p_{\text{fake}}^{(2)} \neq \cdots\]๋ฐ๋ผ์ fake distribution์ score๋ฅผ ์ถ์ ํ๋ ๋ณ๋์ diffusion model
\[\mu_{\text{fake}}\]๊ฐ ํ์ํ๋ค.
$\mu_{\text{fake}}$๋ ํ์ฌ student generator๊ฐ ์์ฑํ sample๋ค์ ์ด์ฉํ diffusion loss๋ฅผ ํตํด ์ง์์ ์ผ๋ก update๋๋ฉฐ, ํ์ฌ์
\[p_{\text{fake}}\]๋ฅผ ์ถ์ ํ๋๋ก ํ์ต๋๋ค.
์ ์ฒด์ ์ธ ๊ตฌ์กฐ๋ ๋ค์๊ณผ ๊ฐ์ด ์๊ฐํ ์ ์๋ค.
\[z \rightarrow G_\theta(z) \rightarrow F(G_\theta(z),t)\]์์ฑ๋ noisy sample์ ๋ score model์ ์ ๋ ฅ๋๋ค.
\[F(G_\theta(z),t) \rightarrow \begin{cases} \mu_{\text{real}} \rightarrow s_{\text{real}} \\ \mu_{\text{fake}} \rightarrow s_{\text{fake}} \end{cases}\]๊ทธ๋ฆฌ๊ณ
\[s_{\text{real}}-s_{\text{fake}}\]๋ฅผ ์ด์ฉํ์ฌ student generator \(G_\theta\)๋ฅผ updateํ๋ค.
Why DMD Accelerates Inference
DMD์ ์ฃผ๋ ๋ชฉ์ ์ training ์์ฒด๋ฅผ ๋น ๋ฅด๊ฒ ํ๋ ๊ฒ์ด ์๋๋ผ inference์ ํ์ํ denoising step์ ์ค์ด๋ ๊ฒ์ด๋ค.
์ผ๋ฐ์ ์ธ multi-step diffusion teacher๊ฐ
\[z \xrightarrow{\text{many denoising steps}} x\]๋ฅผ ํตํด target distribution์ sample์ ์์ฑํ๋ค๋ฉด, DMD๋ student๊ฐ
\[z \xrightarrow{\text{few steps}} \hat{x}\]๋ง์ผ๋ก๋
\[p(\hat{x}) \approx p(x)\]๊ฐ ๋๋๋ก ํ์ตํ๋ค.
์ฆ teacher์ ๊ธด denoising trajectory๋ฅผ ๊ทธ๋๋ก student๊ฐ ๋ฐ๋ผ๊ฐ๋ ๊ฒ์ด ์๋๋ผ, teacher๊ฐ ๋ง์ denoising step์ ํตํด ์ต์ข ์ ์ผ๋ก ํ์ฑํ๋ output distribution์ few-step generator๊ฐ ์ฌํํ๋๋ก ํ์ตํ๋ ๊ฒ์ด๋ค.
๋ฐ๋ผ์ distillation training์๋ real score model, fake score model, student generator ๋ฑ์ด ํ์ํ๊ธฐ ๋๋ฌธ์ ํ์ต ๊ณผ์ ์์ฒด๊ฐ ๋จ์ํด์ง๋ ๊ฒ์ ์๋์ง๋ง, ํ์ต์ด ์๋ฃ๋ ๋ค inference์์๋ teacher์ fake score model์ด ํ์ํ์ง ์์ผ๋ฉฐ few-step student๋ง ์ฌ์ฉํ๋ฉด ๋๋ค.
์ ๋ฆฌํ๋ฉด,
\[\boxed{ \text{DMD: Multi-step Teacher Distribution} \rightarrow \text{Few-step Student Generator} }\]์ด๋ฉฐ, FlashWorld์์๋ ์ด๋ฌํ DMD๋ฅผ ๊ธฐ๋ฐ์ผ๋ก ๋์ visual quality๋ฅผ ๊ฐ์ง๋ MV-oriented mode์ distribution์ 3D consistency๋ฅผ ๊ฐ์ง๋ 3D-oriented few-step generator์ ์ ๋ฌํ๋ค.
Method & Technical Details
ํด๋น ๋ ผ๋ฌธ์ ๊ธฐ๋ฒ์ ์ train ๋๊ณ high-quality multi-view ๋ฅผ ์์ฑํ ์ ์๋ โMV-oriented multi-view diffusion modelโ๊ณผ few-step ๋ง์ 3D consistency๋ฅผ ๋ถ์ฌํ๋ โ3D-oriented generatorโ๋ฅผ ์์ ๋งํ DMD ๊ธฐ๋ฒ์ ํตํด distillationํ๋ ๊ฒ์ด ๋ชฉํ๋ค.
๊ทธ๋ฌ๊ธฐ ์ํด์๋ ์ ์๋ค์ ๋ ๊ฐ์ง challenge๋ฅผ ํด๊ฒฐํด์ผ ํ๋ค๊ณ ํ๋ค:
- 3D-oriented few step generator๋ ์ถฉ๋ถํ robustํ prior์ ๊ฐํ generative ๋ฅ๋ ฅ์ด ํ์ํ๋ค.
- high-quality multi-view dataset์ ์ถฉ๋ถ์น ์๊ธฐ ๋๋ฌธ์ ๊ธฐ์กด์ style, object, camera ๊ถค์ ๊ณผ ๊ฐ์ ๋ณ์๋ค์ handling ํ๋ ์ ๋ต์ develop ํด์ผํ๋ค.
Dual-mode Pre-training
์ด challenge๋ค์ ํด๊ฒฐํ๊ธฐ ์ํ ๊ธฐ๋ฐ์ ๋ค์ง๊ธฐ ์ํด framework๋ฅผ ๋จผ์ ์ ์ํ๋ค.
์ฌ๊ธฐ์ dual-mode๋ ์๋์ ๋ ๊ฐ์ง ๋ชจ๋๋ฅผ ์๋ฏธํ๋ค:
- MV-oriented mode: multi-view image๋ฅผ ์์ธกํ๋ ๊ฒ
- 3D-oriented mode: ์ค๊ฐ feature์์ 3DGS๋ฅผ ์ง์ ๋ง๋ค๊ณ ๋ ๋๋งํ๋ ๊ฒ
๋จผ์ , training dataset์์ $X = {X_1, X_2, โฆ , X_V}$ multi-view image๋ค์ ๊ฐ์ ธ์จ๋ค. ๊ทธ๋ฆฌ๊ณ , $C={C_1, C_2, โฆ , C_V}$ ๊ฐ view์ ํด๋นํ๋ camera parameter๋ ๊ฐ์ ธ์จ๋ค. ์ถ๊ฐ๋ก, $y$๋ผ๋ condition(text prompt, single-view image ๋ฑ)์ด ๋ค์ด๊ฐ๋ค.
์ด๋ ๊ฒ ์ ๋ ฅ์ด ์ค๋น๊ฐ ๋๋ฉด,
\[Z = E(X)\]์ ๋ ฅ multi-view images $X$๋ฅผ VAE encoder $E$์ ๋ฃ์ด์ latent๋ก ๋ณํํ๋ค. ๊ทธ๋ฌ๋ฉด ์๋ ์ฒ๋ผ multi-view data ํ ๋ฐฐ์น์ ๋ํ latent ์งํฉ์ด ์์ฑ๋๋ค:
\[Z = \{Z_1, Z_2, Z_2, ...\}\]๊ทธ๋ฆฌ๊ณ ๋์, ์ผ๋ฐ์ ์ธ diffusion training์ฒ๋ผ random timestep $t$๋ฅผ ์ ํํ๊ณ noise๋ฅผ ๋ฃ๋๋ค.
\[Z_t = \alpha Z + \sigma_t \epsilon\]์์ํ ์ ์๊ฒ ์ง๋ง, ๊ทธ๋ฌ๋ฉด ํ์ต์ผ๋ก ์ฌ์ฉ๋๋ ์ ๋ ฅ multi-view image๊ฐ ์๋ noisy multi-view latent์ธ $Z_t$๋ฅผ ์ฌ์ฉํ๊ฒ ๋๋ค. ๊ทธ๋ฌ๋ฉด, denoising network์ ์ต์ข ์ ์ผ๋ก ๋ค์ด๊ฐ๋ ์ ๋ ฅ์:
\[(Z_t, C, y) \rightarrow \text{Denoising Network}\]์์๋๋ก, noisy latent, camera parameter, condition ์ด 3๊ฐ๊ฐ ๋ค์ด๊ฐ๊ฒ ๋๋ค. ์ฌ๊ธฐ์ camera parameter๋ฅผ ํํํ๋ ๋ฐฉ์์๋ Plรผcker Coordinates raymap๋ฅผ ์ฌ์ฉํ๋ฉฐ, ์ด๋ 3D generation ๋ถ์ผ์์ multi-view๋ฅผ ์์ฑํ ๋ camera parameter๋ฅผ ๋ง์ด ์ฌ์ฉ๋๋ ๋ฐฉ์์ด๋ค.
Denoising Network๋ Diffusion Transformer (DiT)๋ฅผ ๊ธฐ๋ฐ์ผ๋ก ํ๋ฉฐ, 3D attention block์ด ์ถ๊ฐ์ ์ผ๋ก ๋ค์ด๊ฐ๋ค. ์ด network๋ ๋ ๊ฐ์ง ๊ฒฐ๊ณผ๋ฅผ ์ถ๋ ฅํ๋ ค๊ณ ํ๋ค:
\[\hat{Z}_{MV}, F\]์์๋๋ก, clean multi-view latent(MV-oriented mode), multi-view scene information ์ ๋ด๊ณ ์๋ auxiliary feature ์ด๋ค. ์ด ์ค์์ ํ์์ ๊ฒฝ์ฐ์๋ ์ดํ์ 3DGS decoder๋ก ๋ณด๋ด 3D Gaussian์ ๋ง๋ ๋ค(3D-oriented mode). ์ฆ, ์ฌ๊ธฐ์๋ถํฐ ๋ ๊ฐ์ branch๋ก ๋๋์ด์ ์งํ๋๋ค.
๋จผ์ , MV-oriented mode์ ๊ฒฝ์ฐ, DiT๊ฐ $Z_t, C, y$๋ฅผ ๋ฐ์ $\hat{Z}_{MV}$๋ฅผ ์์ธกํ๋ค. ๊ทธ๋ฆฌ๊ณ ground-truth clean latent $Z$์ ๋น๊ตํ๋ค:
\[\mathcal{L}_{MV} = \mathbb{E}_{X, t, \epsilon, y, C} [\lVert Z - \hat{Z}_{MV} \rVert ^ 2]\]์ฆ, noisy multi-view latent ๋ฅผ clean multi-view latent๋ก ๋ณํํ๋ ๊ฒ์ ๋ฐฐ์ฐ๋ diffusion objective๋ค. ์ค์ํ ๊ฑด, MV-oriented mode๊ฐ ์ค์ ๋ก ๋ง๋๋ ๊ฑด 3D representation์ด ์๋๋ค. ๊ฐ camera view์ ํด๋นํ๋ ์ด๋ฏธ์ง๋ฅผ ์ง์ ์์ฑํ๋ ๊ฒ์ด๋ค. ๊ทธ๋ฌ๋ฏ๋ก, ๊ฐ view๊ฐ diffusion์ ์ํด ์ด๋ฏธ์ง ๊ณต๊ฐ์์ ์์ฑ๋๋ฏ๋ก, view 1๊ณผ view 2๊ฐ ๋์ผํ 3D geometry์์ ๋์จ๋ค๊ณ ๋ณด์ฅ๋์ง ์๋๋ค.
์ฆ, ์์ diffusion์ ์์ฑ ๋ฅ๋ ฅ์ ๋ฐ๋ผ high quality๋ฅผ ์ถ๋ ฅํ ์ ์์ง๋ง, multi-view inconsistency๋ผ๋ ๋ฌธ์ ๊ฐ ์๊ธธ ์ ์๋ค.
์ด๋ฅผ ํด๊ฒฐํ๊ธฐ ์ํด ๋๋ค๋ฅธ branch์ธ 3D-oriented mode๊ฐ ๋ฑ์ฅํ๋ค. DiT์ intermediate/output feature์ธ $F$๋ฅผ ๋ณ๋์ 3DGS decoder $D_G$์ ๋ฃ๋๋ค:
\[D_G(F) = \{\tau, q, s, \alpha, c\}\]decoder๊ฐ ๋ด๋๋ ๊ฒฐ๊ณผ๋ ์์๋๋ก, depth, rotation quaternion, scale, opacity, spherical harmonics coefficients๋ค. ์ข ๋ ๊ฐ๋จํ๊ฒ, ๊น์ด, ํ์ , ํฌ๊ธฐ, ๋ถํฌ๋ช ๋, ์์์ ๊ด๋ จํ parameter๋ฅผ ๋ด๋๋๋ค๊ณ ์๊ฐํ๋ฉด๋๋ค. ์ด๋ 3D Gaussian parameter๋ฅผ ์ ๊ณตํ๋ค๊ณ ์๊ฐํ๋ฉด ๋๋ค.
์ฐ๋ฆฌ๋ 3DGS์ ์ตํ ์๊ณ ์๋ค๋ฉด, ์๋ ๋ณดํต 3DGS primitives ๋ผ๊ณ ํ๋ค๋ฉด, gaussian์ position์ธ $\mu$๊ฐ ์์ด์ผ ํ๋ค๋ ๊ฒ์ ๋์น์ฑ ์ ์๋ค. ๊ทธ๋ฐ๋ฐ ์์ parameter ์ค์๋ position์ ๊ด๋ จํ parameter๊ฐ ์๊ณ , ๋์ ๊ทธ ์๋ฆฌ์ depth $\tau$๊ฐ ์๋ค. ์ ์๋ค์ ์ด position์ parameter๋ฅผ ๋ฐ๋ก ์ ํ๋ ๋ฐฉ์์ด ์๋๋ผ ์ด depth๋ผ๋ parameter๋ฅผ ์ฌ์ฉํด position parameter๋ฅผ ์์ธกํ๋ค:
\[\mu = o + \tau d\]์ฌ๊ธฐ์, $o$๋ camera origin(์นด๋ฉ๋ผ ์์น, ๋ณดํต 3D ๊ณต๊ฐ ์์์ ์์ ), $d$๋ ray direction(์นด๋ฉ๋ผ๊ฐ ๋ฐ๋ผ๋ณด๋ ๋ฐฉํฅ, ์ฆ ์ฐ๋ฆฌ๊ฐ ๋ฐ๋ผ๋ณด๋ ๋ฐฉํฅ), $\tau$๋ ์์ธก๋ depth๋ฅผ ์๋ฏธํ๋ค. ๊ฐ๋ ์ ์ผ๋ก, ํด๋น ์์์ด ๊ฐ pixel ๋จ์์์ ์ผ์ด๋๋ฉฐ, ๊ฐ ํฝ์ ์ ๋์ํ๋ 3D gaussian๋ค์ liftingํ๋ค.
๊ทธ๋ฌ๋ฉด, ์ต์ข ์ ์ผ๋ก ์ฐ๋ฆฌ๊ฐ ํ๋ Gaussian parameter๋ฅผ ์์ฑํ ์ ์๋ค:
\[G = \{\mu, q, s, \alpha, c\}\]๊ทธ๋ฌ๋ฉด, ์ฐ๋ฆฌ๋ 3DGS renderer $R$๋ฅผ ์ด์ฉํด์ ํด๋น novel camera view์ ํด๋นํ๋ ์ฅ๋ฉด์ rendering ํ ์ ์๋ค:
\[R(G, C_{novel})\]๊ทธ๋ฌ๋ฉด objective๊ฐ ๋ฌด์์ธ์ง๋ฅผ ์๊ฐํด์ผํ๋ค. ์ ํํ๊ฒ๋ 3D Gaussian์ ๊ด๋ จํ ground-truth๊ฐ ์กด์ฌํ๊ธฐ ํ๋ค๋ค. depth๋ฅผ ์ถ์ ํ๋ ๋ชจ๋ธ๋ ๊ฐ๊ฐ์ ๊ฐ ํฝ์ ๋จ์๋ก depth๋ฅผ ๋ค๋ฅด๊ฒ ์ถ์ ํ๋ค. ์ฌ์ฉํ๋ depth model์ด ๋ฌ๋ผ depth ์ถ์ ์ด ๋ค๋ฅด๋ค๋ฉด, ๊ทธ์ ๋ฐ๋ผ ๊ฐ์ฐ์์์ position์ ๋ฌ๋ผ์ง ๊ฒ์ด๋ค. ๊ทธ๋ฐ๋ฐ, ๊ฒฐ๊ตญ์๋ ์ฌ์ฉ์๊ฐ ์ฅ๋ฉด์ ๋ฃ๊ณ 3D์ฒ๋ผ ๋ณด์ธ๋ค๋ผ๋ ๋๋์ด ๋ค๋ฉด ๊ทธ๊ฒ ์ ๋ต์ด๋ค๋ผ๊ณ ๋ณผ ์ ์๋ ๊ฒ์ด๋ค. ์ฆ, ํด๋น ๊ด๋ จํ ํ๋ผ๋ฏธํฐ๋ค์ ground-truth๊ฐ ์๋ค:
\[\mu_{GT}, q_{GT}, s_{GT}\]๊ทธ๋์ ๋๋ถ๋ถ์ 3D Generation ๋ ผ๋ฌธ๋ค์ ์ด rendering(ํน์ view์ ๊ดํ ์ฅ๋ฉด์ ์บก์ฒํ๋ ๊ฒ)์ ์ฌ์ฉํด์ rendering supervision์ ์ฌ์ฉํ๋ค:
\[\mathcal{L}_{3D} = \mathbb{E} [\lVert X_{novel} - R(G, C_{novel}) \rVert ^2]\]์ฌ๊ธฐ์, ๋ ํน์ดํ ์ ์ ๋ณดํต์ ๋ ผ๋ฌธ๋ค์ camera view๋ค(pose 1 ~ pose n)์ ํ์ดํ๋ผ์ธ ์คํ ์ ์ ์ ํด๋๊ณ , ํด๋น view๋ค์ ๊ดํด์๋ง supervision์ ์งํํ๋๋ฐ ํด๋น ๋ ผ๋ฌธ์ ๊ฒฝ์ฐ์๋ ๋ค๋ฅธ ๋ฐฉ์์ผ๋ก ์ ๊ทผํ๋ค.
novel-view์์ renderingํ๊ณ novel-view๋ฅผ ๋์์ผ๋ก supervision์ ์งํํ๋ค.
์ด๋ ๊ฒ ํ๋ ์ด์ ๋ ์ ๋ ฅํ view์์๋ง reconstruction loss๋ฅผ ๊ฑธ๊ฒ ๋๋ฉด, ์ ๋ ฅ view ์์๋ง ๋ง๊ฒ ๋ณด์ด๊ณ , ๋ค๋ฅธ view์์๋ 3D ๊ตฌ์กฐ๊ฐ ์ด์ํ ์ ์๋ค. ๊ทธ๋ฌ๋ ์ ๋ ฅ view ์ด์ธ์ view์์๋ ๋ง๊ฒ ๋ณด์ด๊ฒ novel-view reconstruction loss๋ฅผ ๊ฑธ๊ฒ ๋๋ฉด, ์ฌ๋ฌ ๊ด์ ์์ ์ผ๊ด๋๊ฒ ๋ฐฐ์น๋์ด์ผ ํ๋ฏ๋ก, 3D consistency constraint๊ฐ ์๊ธฐ๊ฒ ๋๋ค. ์ฌ๊ธฐ์, ์ฐ๋ฆฌ๋ ์ด์ ์ ์ ํด๋์(training dataset์ด ์ ํด๋์, ๋์ผํ Camera parameter $C$) ๊ฐ view์์ 3D consistentํ rendering multi-view๋ฅผ ์ป๊ฒ ๋๋ค.
MV-oriented branch์์๋ ๊ฒฐ๊ณผ์ ์ผ๋ก clean multi-view latent $\hat{Z}_{MV}$๋ฅผ ๋ง๋ค์ด๋๋ค. ํด๋น 3D-oriented branch ์์๋ ๋น์ทํ๊ฒ, $\hat{Z}_{3D}$๋ฅผ ๋ง๋ค์ด๋ธ๋ค. ์ด๋ rendering multi-view์ ๋จ์ํ VAE Encoder $E$๋ฅผ ๊ฑฐ์ณ ๋์จ latent๋ค:
\[\hat{Z}_{3D} = E(R(G, C))\]cross-mode post-training
์ด์ ๋ค์์ผ๋ก ๋ ผ๋ฌธ์ distillation ๋ฐฉ์์ ํตํด few-step 3D Scene์ ์์ฑํ ์ ์๋๋ก ํ์ต์ ํ๋ ค๊ณ ํ๋ค.
์ฐ๋ฆฌ๋ ์์ preliminary ์น์ ์์ DMD์ ๋ํด ๋ฐฐ์ ๊ณ , ๊ทธ์ ๋ํด
\[\mu_{real}, \mu_{fake}, G_{\theta}\]๊ฐ ์กด์ฌํ๋ค๋ ๊ฒ์ ์ ์ ์์๋ค. ์ด ํ์ดํ๋ผ์ธ์์๋ ๊ฐ๊ฐ MV-oriented mode(teacher), 3D student distribution์ fake score๋ฅผ ์ถ์ ํ๋ model, 3D-oriented few-step generator(student)๋ผ๊ณ ์ดํดํ๋ฉด ๋๋ค.
์ฌ๊ธฐ์ $\mu_{real}$๋ frozen ๋ ์ํ๋ก ์ฌ์ฉ๋๋ค. ์ฐ๋ฆฌ๋ ์์ 3D-oriented mode์์ ๋ง์ step์ผ๋ก ๋๋ ธ๋๋ฐ, ์ด๋ฅผ ๊ทธ๋๋ก ์ฌ๊ธฐ์ ๋ง์ step์ผ๋ก ๋๋ฆฌ๋ ๊ฒ์ด ์๋๋ผ, ์์ 3D-oriented mode๋ฅผ ์ง๊ธ few-step student์ initialization์ผ๋ก ์ฌ์ฉํ๋ค. ์ฆ, ์ ์ dual-mode pretraining ์น์ ์์์ architecture์ ๋ค๋ฅธ architecture๊ฐ ์๋ก ์ถ๊ฐ๋๋๊ฒ ์๋๋ค.
๋ ผ๋ฌธ์ ์ด์ ์น์ ์ 3D-oriented mode์ ํ์ดํ๋ผ์ธ์ ๊ฑฐ์ ๊ทธ๋๋ก ๋ฌผ๋ ค ๋ฐ๋๋ค:
\[(\{Z_{t_i}, t_i, y, C\} \rightarrow DiT \rightarrow F_i \rightarrow D_G \rightarrow G_i) \rightarrow R(G_i,C) \rightarrow E(R(G_i,C)) \rightarrow \text{noise injection} \rightarrow Z_{t_{i+1}}\]์ฌ๊ธฐ์, ๊ดํธ ์์ ์๋ ํ์ดํ๋ผ์ธ์ด 3D-oriented generation process์ denoising $G_{\theta, 3D}$์ด๋ผ๊ณ ๋ณด๋ฉด ๋๊ณ , ์ด๊ฑธ 4๋ฒ์ step ๋ง์ ์ฒ๋ฆฌํ๋ ๊ฑธ ๋ชฉํ๋ก ํ๋ค. ์ฌ๊ธฐ์ ํท๊ฐ๋ฆฌ๋ฉด ์๋๋ ๊ฒ์ด DiT์ timestep์ด 4๋ผ๋๊ฒ ์๋๋ผ, ํด๋น ํ๋ก์ธ์ค๊ฐ 4๋ฒ์ ํ์๋ก ์งํ๋๋ ๊ฒ์ ๋งํ๋ค. ํด๋น ๋ ผ๋ฌธ์์๋ timestep์ด $t_i={1000, 900, 759, 500}$ ์ผ๋ก ์ด๋ฃจ์ด์ ธ ์๋ค๊ณ ํ๋ค. ํด๋น ํ๋ก์ธ์ค์ step ํ ๋ฒ์ ๋์๋ ๋งค๋ฒ noise injection์ด ์ํ๋๊ณ , ์ด๋ฅผ ๋ค์ step $t_{i+1}$์ ๋๊ฒจ์ค๋ค.
์ด๋ ๊ฒ ํ์ดํ๋ผ์ธ์ ํ๋ฆ์ ์ ํด๋๊ณ , DMD2 ๊ธฐ๋ฒ์ ์ฌ์ฉํ๋ค. ์๋ 3D-oriented branch ์์๋ many step์ ์ฌ์ฉํ๋๋ฐ, ๋ฐ๋ก 4-step branch๋ก ๋ฐ๊ฟ๋ฒ๋ฆฌ๊ฒ ๋๋ฉด, quality๊ฐ ๋ณด์ฅ๋์ง ์๋๋ค. ๊ทธ๋์ ์ด few-step student๋ฅผ ๋ค์ ํ์ตํด์ few-step์ผ๋ก๋ ์ข์ output distribution์ ๋๋ฌํ ์ ์๋๋ก ํ์ตํด์ผํ๋ค.
์ฌ๊ธฐ์ ์ด์ frozen๋ high-quality MV distribution์ score์ธ $s_{real}$์ ์ถ์ ํ๋ $\mu_{real}$์ ์ฌ์ฉํ์ฌ teacher model๋ก ์ฌ์ฉํ๋ค. ๊ทธ๋ฆฌ๊ณ , $\mu_{fake}$๋ student์ธ $G_{\theta, 3D}$๋ฅผ ๊ณ์ ์ถ์ ํ๋ฉด์ student๊ฐ ์์ฑํ๋ distribution์ธ $p_{fake}$์ score์ธ $s_{fake}$๋ฅผ ์ถ์ ํ๋ค. ๋น์ฐํ student๊ฐ ๋ฐ๋ ๋๋ง๋ค $p_{fake}$๋ ๋ฐ๋๋ฏ๋ก $\mu_{fake}$๋ ๊ณ์ update ๋๋ค.
์ฆ,
\[s_{MV} - s_{\text{current 3D student}} = s_{real} - s_{fake}\]์ด gradient๋ฅผ ์ฌ์ฉํด์ $G_{\theta, 3D}$๋ฅผ updateํ๋ค. ์ด๋ ๊ฒ ๋๋ฉด,
\[p_{\text{3D student}} \rightarrow p_{MV}\]๊ฐ ๋๋๋ก ํ๋ค. ์ถ๊ฐ์ ์ผ๋ก DMD2 loss๋ DMD loss์ GAN loss๋ฅผ ํฉํ ๋ฒ์ ์ด๋ค:
\[L_{DMD2} \approx L_{DMD}+\lambda_{GAN}L_{GAN}\]์ฌ๊ธฐ์, $\lambda$๋ R1 regularization์ด๋ค. ๊ฐ๋ ์ ์ผ๋ก discriminator๋
\[D(X_{real}) \rightarrow 1 D(X_{fake}) \rightarrow 0\]๊ฐ ๋๋๋ก ํ์ต๋๊ณ , generator $G_{\theta, 3D}$๋
\[D(X_{fake}) \rightarrow 1\]์ด ๋๋๋ก ํ์ต๋๋ค.
๊ทธ๋ฐ๋ฐ, ๋ฌธ์ ๋ ์ด๋ ๊ฒ ๋๋ฉด ๊ฒฐ๊ตญ์ 3D student๊ฐ 3d consistency์ ๋ํ ๊ฐ์ ์ ์์ด๋ฒ๋ฆฌ๊ณ MV-oriented ์ชฝ์ผ๋ก ๋๋ ค๊ฐ ์ํ์ฑ์ด ์๋ค๋ ๊ฒ์ด๋ค. ๋ฌผ๋ก , 3d-oriented branch๋ ํ๋์ share๋ 3D Scene์์ ๋ ๋๋ง๋์ด ์ด gradient๊ฐ update์ ์ฃผ์ฒด๊ฐ ๋์ด 3d consistency๋ฅผ ์ ์งํ ๊ฒ์ฒ๋ผ ๋ณด์ธ๋ค.
๋ ผ๋ฌธ์ ์ ์๋ค์ ์๋์ 3๊ฐ์ ์ด์ ๋ก 3d consistency๊ฐ ์ ์ง๋๋ค๊ณ ๋งํ์ง๋ง, ํด๋น ๋ฆฌ๋ทฐ๋ฅผ ์ ๋ ๋ณธ์ธ์ ํด๋น ์ด์ ๋ง์ผ๋ก ์ฆ๋ช ๋๊ธฐ๋ ํ๋ค๋ค๊ณ ๋ณธ๋ค:
- 3D supervision (์ด์ ์น์ ์์ novel-view์ ๋ํ rendering supervision์ ์ํํ๋ ๊ฒ)
- pretrained 3D-oriented weights๋ก ์์ํ๋ ๊ฒ
- ๋งค step ๋ง๋ค ๊ณ์ $G_i \rightarrow \text{Render}$ ํ๋ ๊ฒ
์ฌ๊ธฐ์ ์ด ๋ ผ๋ฌธ์ ๋ฆฌ๋ทฐ๋ฅผ ์ ๋ ์์ฑ์๊ฐ ์๊ฐํ๊ธฐ์ 2๋ฒ์ด ์ค์ํด ๋ณด์ด๋๋ฐ, ์ด weights๊ฐ ๊ฒฐ๊ตญ์๋ MV-oriented weights๋ก ๊ณ์ ๋ณํํ ํ ๋ฐ, ๊ทธ๋ฌ๋ฉด 3d consistency์ ๊ฐ์ ์ ์ด๋ป๊ฒ ์ ์์ด๋ฒ๋ฆฌ๊ณ ์ ์งํ ์ง ์๋ฌธ์ด๋ค. MV์ชฝ์ ๊ณ ์ ๋์ด ์๋ ์ํ์์ 3D์ชฝ๋ง updateํ์ฌ ์งํํ๋ ๋ฐฉ์์ด๊ธฐ ๋๋ฌธ์ด๋ค.
์ถ๊ฐ์ ์ผ๋ก ์ ์๋ค์ Cross-Mode Consistency Loss๋ฅผ ์ถ๊ฐํ๋ค. ์ฌ๊ธฐ์, 3D Student๊ฐ ์ฌ์ฉํ๋ DiT backbone์ ๊ณต์ ํ๋ MV_oriented student branch๋ฅผ ์ฌ์ฉํ๋ค.
\[\begin{aligned} \hat{Z}_{3D} &= E(R(G_{\theta, 3D}(Z_t, t_i, y, C), C)) \\ \hat{Z}_{MV} &= G_{\theta, MV}(Z_t, t_i, y, C) \end{aligned}\]๋ฅผ ์ป์ด์
\[\mathcal{L}_{CMC} = \lVert \hat{Z}_{3D} - \hat{Z}_{MV} \rVert ^2\]๋ก ๋ง์ถ๋ค. ์ด๋ DMD2๋ฅผ ์งํํ์ ๋ ์๊ธฐ๋ floating artifact์ ๋ถ์์ ํ 3D prediction์ ์ค์ด๊ธฐ ์ํ ๋ณด์กฐ loss๋ก MV-oriented student๋ ๋ฎ์ frequency๋ก ์ด๋ฐ์ดํธํ๊ณ , ๋ mode์ prediction์ ๋ง์ถ๋ค.
๊ฒฐ๋ก ์ ์ผ๋ก, 3D student๋ ํด๋น signal์ ๋ฐ๊ฒ ๋๋ค:
\[\underbrace{\mathcal{L}_{\mathrm{DMD}}}_{\text{MV teacher distribution์ผ๋ก ์ด๋}} + \underbrace{\mathcal{L}_{\mathrm{GAN}}}_{\text{rendering์ real/high-qualityํ๊ฒ}} + \underbrace{\lambda \mathcal{L}_{\mathrm{CMC}}}_{\text{3D branch ์์ ํ}}\]Out-of-Distribution Data Co-Training
์ด๋ ๊ฒ ํ์ต์ ์งํํ๊ฒ ๋๋ฉด, ํ๋์ ์ฌ์ํ ๋ฌธ์ ๊ฐ ๋ ๋ฐ์ํ๊ฒ ๋๋ค. Flashworld์ ๊ฒฝ์ฐ, MVImgNet, RealEstate10K, DL3DV10K๋ฅผ ๊ธฐ๋ฐ์ผ๋ก ํ์ต๋์๋ค. ๊ทธ๋ฐ๋ฐ, DiT์ ๊ฒฝ์ฐ์๋ ์ด ์ด์ธ์๋ ๋๊ท๋ชจ image์ video data๋ก ํ์ต๋์๊ธฐ ๋๋ฌธ์ ์์ 3๊ฐ์ multi-view datasets์ ์ ์ธํ ๋ค๋ฅธ ๋ฐ์ดํฐ๊ฐ ๋ค์ด์๋ robustํ๊ฒ ๋์ฒํ ์ ์๋ ๋ฐ๋ฉด, 3D branch์ ์๋ 3DGS Decoder๋ 3๊ฐ์ multi-view datasets์ ์ ์ธํ ๋ค๋ฅธ ๋ฐ์ดํฐ์ ๋ํด robust ํ์ง์๋ค.
์ฆ, 3DGS Decoder $D_G$๋ multi-view distribution๋ง ๊ฒฝํํ๋ค.
๊ทธ๋ ๋ค๋ฉด 3DGS Decoder๊ฐ ๋ฐ์ input์ distribution์ ๋ํ์ฃผ๋ฉด ํด๊ฒฐ๋๋ค. ๊ทธ๋ฌ๊ธฐ ์ํด์, ๋จผ์ ํด๋น ํ์ดํ๋ผ์ธ์์ ์ฒ์์ ์ ๋ ฅํ๋ ๋ฐ์ดํฐ๋ฅผ multi-view dataset์ด ์๋๋ผ single image๋ text๋ง ๋๊ฒจ์ฃผ๊ฒ ๋๋ค.
text ํน์ single image $y$๋ง ๋๊ฒจ์ฃผ๋ ๊ฒฝ์ฐ์๋ ์ด์ camera trajectory $C$๋ฅผ ๊ฐ์ด ๋๊ฒจ์ฃผ๊ณ , DMD ๋ฐฉ์์ผ๋ก๋ง ํ์ต์ ์งํํ๊ฒ ๋๋ค. ๋น์ฐํ ์ค์ Ground truth๊ฐ ์์ผ๋ฏ๋ก GAN Loss๋ ์ฌ์ฉํ์ง ์๊ณ DMD์ CMC loss ๋ง์ผ๋ก ํ์ต์ ์งํํ๋ค:
\[(y, C_{random}) \rightarrow DiT \rightarrow F \rightarrow D_G \rightarrow G\]์ฌ๊ธฐ์ Camera trajectory๋ RealEstate10K, WorldScore ๊ฐ์ ์ด๋ฏธ ์๋ multi-view dataset์ trajectory๋ฅผ ์ฌ์ฉํ๋ค. ๊ทธ๋ฆฌ๊ณ , ํด๋น ํ์ต์ pre-training ๋จ๊ณ๊ฐ ์๋ post-training ๋จ๊ณ์์ multi-view data์ ood data๋ฅผ 2:1์ ๋น์จ๋ก ์์ด์ ์ฌ์ฉํ๋ค.
Experiments
figure 4 ์์๋ baseline๋ค์ MV-oriented ๋ฐฉ์๋ค๋ก ๊ตฌ์ฑํ๊ณ , 3D oriented ํ์ดํ๋ผ์ธ์ธ flashworld์ ๊ฐ์ ์ ๋ณด์ฌ์ค๋ค. ํด๋น ๋ฒ ์ด์ค๋ผ์ธ๋ค์ ์ฝ๋๊ฐ ์คํ๋์ด ์์ง ์์ง๋ง, ๊ฐ ํ๋ก์ ํธ ํ์ด์ง์ ์ ๊ณต๋ ๋น๋์ค ๊ฒฐ๊ณผ๋ฌผ์ ํ์ฉํ๊ณ , ViPE๋ผ๋ ๋ชจ๋ธ์ ์ฌ์ฉํด์ ์นด๋ฉ๋ผ ํฌ์ฆ์ ๋ด์ฌ ํ๋ผ๋ฏธํฐ๋ฅผ ์ถ์ ํ์ฌ ์ต๋ํ ๋น์ทํ ๊ฐ๋์์ ์์ฑํ๋๋ก ํ์๋ค.
ํ ์คํธ ๊ธฐ๋ฐ ์์ฑ ๋น๊ต์ ๋ํด์๋ ์ ์ฑํ๊ฐ์ ์ ๋ํ๊ฐ๋ฅผ ๋ชจ๋ ์งํํ์๋ค. ์ ์ฑํ๊ฐ์์ Prometheus๋ MV-oriented ํ์ดํ๋ผ์ธ์ ๋ณธ์ง์ ์ธ ๋ถ์ผ์น๋ก ์ธํด ์์ฑ๋ ์ฅ๋ฉด์ด ์์ฃผ ํ๋ฆฟํด์ง๊ณ ๊ธฐํ ๊ตฌ์กฐ๊ฐ ์๋ชป ํํ๋๊ธฐ๋ ํ๋ค. ๊ทธ๋ฆฌ๊ณ , SplatFlow์ VideoRFSplat ์ญ์ ํ๋ฆฟํ ์๊ณก์ผ๋ก ์ด๋ ค์์ ๊ฒช์ผ๋ฉฐ ๋ฐ๋ฅ์ด๋ ์๋ ๋ฑ์์ ๋ฐ๊ฒฌ๋๋ ์ธ๋ถ์ ์ธ ๋ํ ์ผ์ ์ฌํํ๋๋ฐ ํ๊ณ๋ฅผ ๋ณด์ธ๋ค.
์ ๋ํ๊ฐ์ ๊ฒฝ์ฐ์๋ T3Bench, DL3DV, WorldScore์์ 600๊ฐ์ ํ ์คํธ ํ๋กฌํฌํธ๋ฅผ ์ํ๋งํ์๋ค. ํด๋น table์์ ๋น๊ต ๋์์ด ๋๋ ๋ฐฉ๋ฒ๋ค์ด 3D Gaussian representation์ ๊ธฐ๋ฐ์ผ๋ก ํ๊ธฐ ๋๋ฌธ์, ์นด๋ฉ๋ผ ์ ์ด ๋ฐ 3d consistency๊ณผ ๊ด๋ จ๋ ์งํ๋ค์ ๋ณธ ์คํ ์ค์ ์์ ์ ์ฉํ๊ธฐ ์ ํฉํ์ง ์์์ CLIP IQA+, CLIP Aesthetic, CLIP Score, Q-Align์ ํฌํจํ์ฌ ํ๊ฐ ์งํ์ ์ง์คํ๋ค. ํนํ, CLIP-Aesthetic ์งํ์ ๊ฒฝ์ฐ, ๋๋๋ก smooth ์ถ๋ ฅ๋ฌผ์ ์ ํธํ๋ ๊ฒฝํฅ์ด ์์ด ๋ณธ ์ฐ๊ตฌ ๋ฐฉ๋ฒ์ด ๋ง๋ค์ด๋ด๋ ์ ๊ตํ๊ณ ์ฌ์ค์ ์ธ ๊ฒฐ๊ณผ์ ํญ์ ๋ถํฉํ์ง ์์ ์ ์์์ ์ ์ ์๋ค.
๋ ผ๋ฌธ์ ์ ์๋ค์ worldscore benchmark์ ๋ํด์๋ ํ๊ฐ๋ฅผ ์งํํ๋ค. Flashworld๋ WonderJourney, LucidDreamer, WonderWorld 3D ์์ฑ ๋ฐฉ๋ฒ๋ก ๋ค๊ณผ ๋น๊ตํ๋ค. ์ฌ๊ธฐ์ ํด๋น ์ฐ๊ตฌ์์ 3D ์์ฑ ๋ฐฉ๋ฒ๋ก ์๋ง ์ง์คํ๊ณ ์์ด, Camera Control ์ด๋ผ๋ ์งํ๋ ์ฃผ๋ก ๊ฐ ๋ฐฉ๋ฒ๋ก ์ ํ๊ฐ ํ๋กํ ์ฝ์ ๋ํ ๊ฐ๊ฑด์ฑ๋ง์ ๋ฐ์ํ ๋ฟ์ด์ด์ ๋ณธ ์คํ ์ธํ ์์๋ informative๊ฐ ๋จ์ด์ ธ ํด๋น ์งํ๋ ํฌํจํ์ง ์์๋ค. ๋ํ, ๊ธฐ์กด WorldScore ๋ฒค์น๋งํฌ๋ ๋๋ถ๋ถ์ ์งํ๋ฅผ anchor frames์์๋ง ํ๊ฐํ๋๋ฐ, ์ด๋ novel view synthesis๊ฐ ์๊ตฌ๋๋ 3D ์๋ ์์ฑ ๊ณผ์ ์ suboptimal ์ผ ์ ์๋ค.
๊ทธ๋์ ๋์ฑ ๊ณต์ ํ ๋น๊ต๋ฅผ ์ํด ํน์ interval ์์ ์๋ ํ๋ ์๋ค ์ค์ ๋ฌด์์๋ก ํ๋ ์๋ค์ ๋ฝ์๋ด์ด์ ์ฌํ๊ฐํ๋ค๊ณ ํ๋ค. ๋ชจ๋ ์ ๊ทผ ๋ฐฉ์ ์ค์์ ๊ฐ์ฅ ๋์ ํ๊ท ์ ์์ ๊ฐ์ฅ ๋น ๋ฅธ ์ถ๋ก ์๋๋ฅผ ๋ฌ์ฑํ์๋ค.
๊ทธ๋ฆฌ๊ณ ๋ ผ๋ฌธ์ ์ ์๋ค์ ๋ค์ํ ablation study๋ฅผ ์งํํ์๋ค. w/ MV-Diff์ ๊ฒฝ์ฐ MV-oriented diffusion model์ ์๋ฏธํ๊ณ , w/ 3D-Diff์ ๊ฒฝ์ฐ 3D oriented diffusion model์ ์๋ฏธํ๊ณ , w/ MV-Dist์ ๊ฒฝ์ฐ MV-oriented model์ few-step์ผ๋ก distillํ ๊ฒฝ์ฐ๋ฅผ ์๋ฏธํ๊ณ , w/o CMC์ ๊ฒฝ์ฐ์๋ 3D-oriented model์ few-step์ผ๋ก distillํ์ง๋ง CMC loss๊ฐ ์ ๊ฑฐ๋ ๊ฒฝ์ฐ, ๋ง์ง๋ง์ผ๋ก w/o OOD์ ๊ฒฝ์ฐ์๋ Full cross-mode model์์ OOD co-training์ ์ ๊ฑฐํ ๊ฒฝ์ฐ๋ฅผ ๋ณด์ฌ์ค๋ค.
์ ๋ง ์ ๊ธฐํ๊ฒ๋, full model์ ๋นํด์ w/o CMC์ ๊ฒฝ์ฐ์์ ๋ง์ metric์ด ๋ ์ฐ์ํ ๊ฒฝ์ฐ๋ค์ ๋ณด์ฌ์ค๋ค. ์ด๋ ๋จ์ํ CMC๊ฐ ์๊น ๋งํ๋ฏ์ด 3D student model์ด 3D consistency์ ๋ํ distribution์ ์๋ ๊ฒ์ ๋์ด์ ์ฆ๋ช ๋ ํ๋ฌ์ ๋ณด์ฌ์ฃผ๋ ๊ฒ ๊ฐ๋ค.
Contributions
- MV-oriented ๋ฐฉ์๊ณผ 3D-oriented ๋ ๋ฐฉ์ ๋ชจ๋์์ ์๋ํ๋ multi-view diffusion model์ ํ์ตํ๋ pretraining strategy๋ฅผ ์๊ฐํ๋ค.
- visual quality, 3d consistency ๋ชจ๋์์ robustํ ์ฑ๋ฅ์ ๋ณด์ฌ์ฃผ๋ cross-mode post-training strategy๋ฅผ ์ ์ํ๋ค.
- out-of-distribution์์ generalization ability๋ฅผ ๋์ด๋ novel strategy๋ฅผ ์๊ฐํ๋ค.
Limitations & Future works
view์ ์๋ฅผ ๋๋ ธ์์๋, ์์ฑ๋๋ 3D ์ฅ๋ฉด์ ๋ค์์ฑ๊ณผ ๊ท๋ชจ๋ ์ฌ์ ํ ๊ธฐ์กด์ ์กด์ฌํ๋ ๋ฐ์ดํฐ์ ์ ์ปค๋ฒ๋ฆฌ์ง ๋ฒ์์ ์ํด ์ ํ๋๋ค. ์ ๋ฐํ ๊ธฐํ ๊ตฌ์กฐ๋ ๊ฑฐ์ธ ๋ฐ์ฌ, ๊ธ์จ์ ๊ฐ์ ์ธ๋ฐํ ํํ์๋ ํ๊ณ๋ฅผ ์ง๋๋ค. ์๊ธฐํ๊ท ์์ฑ๊ธฐ๋ฒ์ ๋์ ๊ณผ ๋์ 4D ์ฅ๋ฉด ์์ฑ์ผ๋ก์ ํ์ฅ์ ํฅํ ์ฐ๊ตฌ ๋ฐฉํฅ์ผ๋ก ๋ด๋์๋ค.
๊ทธ๋ฆฌ๊ณ ์๋๋ ๋ฆฌ๋ทฐ ์์ฑ์๊ฐ ์๊ฐํ๋ limitation ์ด๋ค:
distillation๋ Student model์ด 3D consistency distribution์ ์ ์งํ ์ ์๋ ์ง ์ฆ๋ช ๋์ง ์์๋ค.