SpaceFlow
Locally Controllable 3D Generation
Training-free 3D generation with local geometric control
and part-specific text or image appearance conditioning.
1 ETH Zürich2 Stanford University* Equal contribution† Equal supervision
Abstract
Current 3D generation methods lack explicit local control: geometric adherence is often defined by a global control strength, and appearance cannot be specified locally. We present SpaceFlow, a training-free pipeline for locally controllable 3D generation from text descriptions and a collection of geometric primitives. Each primitive serves as a proxy for an object part and is assigned a local control level, enabling users to specify whether regions should strictly follow the input shape or allow generative completion. During structure generation, we enforce these spatial constraints within the generative flow process. For appearance synthesis, the generated structure is segmented and matched to the primitives. Each generated part is conditioned only on its assigned text or image cue, thereby limiting cross-part leakage. Regional geometry metrics demonstrate that SpaceFlow preserves the specified geometry in high-control regions and enables plausible shape variation in low-control areas. A user study further indicates that the resulting balance between geometric fidelity and generative freedom remains competitive in overall quality. When evaluating appearance on fixed geometry, text-conditioned routing achieves state-of-the-art prompt faithfulness and color/material accuracy. Qualitative results additionally show localized routing of image cues.
01 / OVERVIEW
Local geometric and
appearance control
SpaceFlow assigns geometric control strengths and appearance conditions to individual parts of a generated 3D asset.
Each input primitive represents an object part. Strong geometric guidance encourages adherence to the input shape, while weak guidance allows prompt-conditioned completion. During appearance generation, text or image cues are routed to the corresponding regions of the generated structure.
The pretrained generator remains frozen throughout both stages.
Geometric and appearance control
LOCAL GEOMETRY + LOCAL APPEARANCEINTERACTIVE EXAMPLES
Explore the generated assets
03 / METHOD
Structure generation and
appearance conditioning
SpaceFlow applies local guidance during sampling. Input primitives specify geometric constraints and associate text or image cues with object parts.
The input consists of editable geometric parts, each with a local control level and an optional text or image appearance cue. Deformable superquadrics provide a compact representation for specifying these parts.

Primitive-to-part appearance routing
Input primitives are matched to regions of the generated structure. Corresponding colors identify the same appearance condition; gray regions receive the global condition. These colors indicate cue assignments, not the final materials.


“red metal truck with a white cabin and blue glass”


“blue wooden toy elephant with pink wooden ears”


“white enamel mug with a blue handle”
Selected examples from supplementary Figure S9. Routing is represented on a 32³ voxel grid using 30 PartField clusters.
Experimental results
SpaceFlow win rate
Overall preference balances fidelity to the input and realism. These are VLM-judge results over 83 assets, aggregated by majority vote across three passes; intervals are 95% Wilson confidence intervals.
05 / RESOURCES
Paper and resources
The paper, the workflow, and the generated assets.
Citation
@misc{delafuente2026spaceflow,
title = {SpaceFlow: Locally Controllable 3D Generation},
author = {De La Fuente, Neil and Lafuente Baeza, Joan and
Sayfiddinov, Mukhammadali and Scharitzer, Felicia and
Pollefeys, Marc and Çelen, Ata and
Deb Sarkar, Sayan and Fedele, Elisabetta},
year = {2026},
note = {Preprint}
}Preprint citation. Publication metadata will be updated with the release.


