Single-Stage 3D Generation · Self-Conditioning

Position Forcing:
Self-Conditioning 3D Generation

1VCIP, Nankai University 2Tencent Hunyuan 3MMLab, CUHK 4Fudan University 5Shanghai Innovation Institute 6HKU
†Project lead  ·  ‡Corresponding author
TL;DR — Position Forcing improves single-stage 3D generation by recovering token positions from geometry latents and using them as coarse-to-fine spatial self-conditioning during denoising.

Abstract

Recent single-stage 3D generative models commonly adopt VecSet representations, encoding 3D shapes as unordered sets of latent tokens. However, compared with two-stage methods that provide explicit positional guidance, these models must implicitly infer token positions throughout denoising, limiting their generation quality. We observe that, despite the absence of explicit positional conditioning, VecSet tokens retain recoverable spatial correspondences. Building on this observation, we propose Position Forcing, a position-based self-conditioning framework. During denoising, Position Forcing recovers token positions from the current clean latent estimate, quantizes them at progressively finer resolutions according to the denoising stage, and feeds the resulting positional encodings back into the diffusion Transformer. This progressively refined positional feedback provides spatial guidance at a granularity appropriate to each denoising stage, guiding shape generation along a coarse-to-fine trajectory and substantially improving generation quality without a separate position generation stage. Experiments demonstrate that Position Forcing achieves strong performance among single-stage 3D generative methods and outperforms several competitive multi-stage approaches.

The generation process, step by step

One 50-step sampling trajectory on a real case. Press play and watch the spatial guidance sharpen from a single cell to full resolution.

Live trajectory Positional guidance at R = 1
Conditioning image for the demonstrated case
input image
show
step0 / 50
timestep t1.000
grid R(t)1
Clean-state conditioning — one step
z₁ zt ztβˆ’1 zβ‚€ αΊ‘β‚€ P0|t RoPE
  1. Predict velocity at zt using the previous step's position
  2. Invert the flow path to a clean estimate αΊ‘0
  3. Decode positions P0|t from that clean estimate
  4. Quantize at the next step's resolution R(t−1)
  5. Inject through RoPE and take the Euler step to zt−1
↻ the recovered positions condition the next step
speed
t = 1  (pure noise)t = 0  (clean geometry)

Interactive results

Every mesh below is the real generated .glb, shaded by surface normal so the geometry reads without texture hiding it. Page through all 48 results — the conditioning image is inset in each corner.

Loading cases…

The observation that makes it possible

In a VecSet VAE with point queries, every latent token is built from a spatial query — and that correspondence survives encoding.

Our observation and positional conditioning designs
(I) Latents decode into both shape geometry and the spatial positions associated with each token. (II) (a) VecSet uses no explicit positional conditioning; (b) VoxSet and two-stage pipelines rely on positions established beforehand; (c) a naive design recovers positions from the current noisy state; (d) Position Forcing recovers them from the predicted clean state and applies progressive quantization for coarse-to-fine guidance.
99.9%
position recovery on clean latents, after joint finetuning
fraction of tokens landing in the correct cell of a $128^3$ grid
92.5%
still correct under 20% noise perturbation
vs. 23.0% for a frozen VAE β€” this robustness is what makes mid-denoising feedback usable
1 stage
no separate structure-generation model
spatial guidance is produced inside the single denoising trajectory

Method

Position Forcing method overview: Position VAE, Position-Forced DiT, progressive quantization
(a) Position VAE learns latent tokens that preserve geometry–position correspondences. (b) Position-Forced DiT recovers token positions from the predicted clean latents and injects them into subsequent denoising through RoPE. (c) Recovered positions are progressively quantized from coarse to fine as denoising proceeds.

Position VAE

A VoxSet VAE augmented with a 4-layer Transformer position decoder $D_{\mathrm{pos}}$, trained in two stages:

  1. Stage I β€” position learning on frozen latents. The VAE is frozen; only $D_{\mathrm{pos}}$ trains, supervised by a token-wise L1 loss against the jittered query coordinates. $$\mathcal{L}_{\mathrm{pos}}=\tfrac{1}{3N}\textstyle\sum_i\|\widehat{\mathbf{p}}_i-\mathbf{p}_i\|_1$$ This leaves the geometry latent space untouched, but accuracy is limited and collapses under noise.
  2. Stage II β€” joint position-aware training. Encoder, geometry decoder and position decoder train together, so the latents support reconstruction and position recovery. $$\mathcal{L}_{\mathrm{joint}}=\mathcal{L}_{\mathrm{rec}}+\lambda_{\mathrm{KL}}\mathcal{L}_{\mathrm{KL}}+\lambda_{\mathrm{pos}}\mathcal{L}_{\mathrm{pos}}$$ Positions now survive heavy noise β€” the prerequisite for reading them off an imperfect $\widehat{\mathbf{Z}}_0$.

Position-Forced DiT

A FLUX-style rectified-flow Transformer (12 double-stream + 24 single-stream blocks, width 1536) with DINOv2-Giant image conditioning and 3D RoPE carrying the positional guidance.

Why the clean state? Position predictions from the noisy latent $\mathbf{Z}_t$ are unreliable at high noise, and full-resolution encodings then lock the model onto wrong fine-grained relationships. Position Forcing instead inverts the flow path to a clean estimate first:

$$\widehat{\mathbf{Z}}_0=\mathbf{Z}_t-t\widehat{\mathbf{V}}_t,\qquad P_{0\mid t}=D_{\mathrm{pos}}(\widehat{\mathbf{Z}}_0)$$

Why quantize? Coarse cells at high noise keep the global layout while staying insensitive to positional error; the grid then refines toward local geometry.

Training vs. inference. Reproducing the feedback during training would require unrolling the trajectory, so training keeps single-timestep sampling and conditions on ground-truth query positions under the same schedule $R(t)$, plus random $\{-1,0,1\}$ cell perturbations for robustness:

$$\mathcal{L}_{\mathrm{FM}}=\mathbb{E}\!\left[\left\|v_\theta\!\left(\mathbf{Z}_t,t,\mathbf{c}_I,\operatorname{PE}\!\left(Q_{R(t)}(\mathcal{P})+\boldsymbol{\delta}\right)\right)-(\boldsymbol{\epsilon}-\mathbf{Z}_0)\right\|_F^2\right]$$

Results

Image-conditioned generation

Position Forcing reaches the best or tied-best score on all four metrics — ahead of both single-stage and multi-stage baselines.

ModelSingle stageULIP-T ↑ULIP-I ↑Uni3D-T ↑Uni3D-I ↑
TRELLISβœ—0.0760.1260.2490.311
TRELLIS 2βœ—0.0770.1240.2450.317
Direct3D-S2βœ—0.0740.1220.2470.314
Hi3DGenβœ—0.0650.1130.2520.301
CraftsMan 1.5βœ“0.0740.1290.2370.298
UniLat3Dβœ“0.0710.1190.2520.307
Hunyuan3D 2.1βœ“0.0750.1250.2500.320
Position Forcing (Ours)βœ“0.0770.1300.2570.321

Geometry reconstruction

MethodLatent sizeCD ↓F1 ↑
TripoSG64 Γ— 409617.9692.33
64 Γ— 819216.6194.19
Hunyuan3D-2.164 Γ— 40969.0889.43
64 Γ— 81928.2890.99
64 Γ— 204807.6292.06
Position Forcing64 Γ— 40968.2890.74
64 Γ— 81926.3793.83
64 Γ— 204805.3995.38

Joint training → robust position recovery

Accuracy (%) = fraction of predicted positions falling in the same cell as their query position on a $128^3$ grid.

Training strategy0%10%20%30%
Frozen VAE67.451.323.06.7
Joint finetuning99.999.792.567.3
Positional consistency along the sampling trajectory
Positions from the predicted clean latent ($P_{0\mid t}$) agree with the final positions earlier and stay consistent under progressive quantization — unlike those read from the noisy latent ($P_t$).

For qualitative results, see the interactive meshes above — every case is the real generated asset, rotatable from any angle rather than a fixed render.

Ablations

Same input, seed and sampling settings throughout. (a) is the VecSet baseline with no positional conditioning; (f) is the full model.

Configuration(a)(b)(c)(d)(e)(f) Ours(g)
Train pos.β€”PredPredPredQueryQueryQuery
Infer pos.β€”PredPred$x_0$Pred$x_0$$x_0$
Progressive quant.βœ—βœ—βœ“βœ“βœ“βœ“βœ“
Resolutionβ€”128Prog.Prog.Prog.Prog.Prog.
VAEVecSetJointJointJointJointJointStage-I
ULIP-T ↑0.0750.0760.0740.0750.0580.0770.076
ULIP-I ↑0.1250.1270.1240.1260.0950.1300.128
Uni3D-T ↑0.2500.2470.2560.2570.2250.2570.248
Uni3D-I ↑0.3200.3050.3210.3230.2650.3210.308

Query / Pred / $x_0$ = ground-truth query positions, predicted positions, and positions recovered from the predicted clean latent. Bold = best, underline = second best.

Algorithm

Training conditions on ground-truth query positions; inference recovers them from the predicted clean latent and re-quantizes each step.

# Training: use query positions from the frozen encoder
def training_loss(point_cloud, image):
    with no_grad():
        z0, query_pos = vae.encode(point_cloud)
        image_cond    = image_encoder(image)

    t, noise = sample_time_and_noise(z0)
    zt = (1 - t) * z0 + t * noise

    # Quantize query positions at the current resolution
    pos = quantize(query_pos, R(t))
    v   = dit(zt, t, image_cond, rope_idx=pos)

    return mse(v, noise - z0)


# Inference: recover positions from predicted clean latents
@no_grad()
def sample(noise, image, times):
    z = noise
    image_cond = image_encoder(image)

    # Initially R = 1: all tokens share the same coordinates
    pos = zeros_like_positions(z)

    for t, t_next in consecutive_pairs(times):
        # Predict velocity using current position
        v = dit_cfg(z, t, image_cond, rope_idx=pos)

        # Predict clean latents and recover token positions
        z0_pred       = z - t * v
        predicted_pos = pos_decoder(z0_pred)

        # Prepare position conditions for the next step
        pos = quantize(predicted_pos, R(t_next))

        # Advance sampling with an Euler update
        z = z + (t_next - t) * v

    # Decode the generated geometry
    return mesh_decoder(z)

BibTeX

@misc{ouyang2026positionforcingselfconditioning3d,
      title={Position Forcing: Self-Conditioning 3D Generation}, 
      author={Ziheng Ouyang and Zeqiang Lai and Jiarui Chen and Jiangshan Wang and Yuhao Wan and Jingbo Gong and Xiangyu Yue and Hengshuang Zhao and Qibin Hou and Chunchao Guo},
      year={2026},
      eprint={2610.10342},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.10342}, 
}