WACV 2027 ยท Accepted in Round 1

Goal Images Are Control Targets

Spatially-Grounded Goal-Image Synthesis for Robotic Manipulation

Riddhi Chatterjee Ritish Shrirao Kushagra Mehrotra Viswanath Gopalakrishnan
International Institute of Information Technology Bangalore (IIIT-B)
Left: general-purpose editors introduce unmentioned changes, ignore the instruction, or hallucinate objects. Right: SGE-Goal edits the correct object with spatial grounding.
Goal images are control targets, not artworks. A goal image must move the correct object to the correct final state and avoid spurious changes that mislead the policy. SGE-Goal achieves this through explicit spatial grounding, while strong general-purpose editors introduce unmentioned changes, miss the instruction, or hallucinate objects.
Abstract

A better edit is not a better goal

Goal-conditioned manipulation policies are steered by a goal image, a picture of the desired post-manipulation scene synthesized from the current observation and a language instruction. The prevailing assumption is that better image-editing quality yields better goals, so goal generators are built and compared using general image-editing metrics. We show this assumption is false.

Across two independent goal-conditioned policies on CALVIN, the image editors that top aggregate editing-quality leaderboards (Gemini 3 Pro Image and GPT Image 1.5) produce worse control targets than our far smaller fine-tuned open models, and the gap widens sharply for the multi-step tasks robots actually face. A per-dimension analysis explains why: downstream policy success is governed almost entirely by manipulation adherence (selecting the right object and the right final state; Spearman ฯ โ‰ˆ 0.9 with success), while photorealism and aesthetic fidelity โ€” precisely where large proprietary editors excel โ€” are uncorrelated with success (ฯ โ‰ˆ 0). This dissociation holds on both CALVIN and the real-world-aligned SIMPLER-Bridge benchmark.

We then present SGE-Goal, a spatially-grounded decomposed editing framework that predicts initial and final 2D edit-region masks before goal-image synthesis. SGE-Goal turns open base models (FLUX.1 Kontext, FLUX.2 klein) into the strongest goal generators we evaluate on CALVIN, and second only to Gemini 3 Pro Image on SIMPLER-Bridge. Fine-tuned with SGE-Goal, they also hold up under chaining far better than the best proprietary editor, sustaining roughly seven times its success rate at five-step chains. Goal generators must be selected by downstream utility, not by editing-quality leaderboards.

Key findings

What actually makes a good goal image

Editing-quality rankings and downstream success rankings disagree, on all three datasets and under both editing evaluators.

ฯ 0.92vs 0.25
Spearman |ฯ| with GR-MG success on CALVIN: source adherence (R14) vs photorealism (R13).
โ‰ˆ7ร—
Success rate of SGE-Goal (FLUX.1 Kontext) over Gemini 3 Pro Image at chain length 5 under GR-MG (30.9% vs 4.4%).
10.6 s
End-to-end goal-image inference on a single A10G, entirely with open models โ€” no proprietary API.

Editing quality is a misaligned proxy

Aggregate editing score correlates only weakly with policy success (ฯ โ‰ˆ 0.67) and mis-orders the top methods.

Adherence predicts success; photorealism does not

An exact permutation test on CALVIN gives ฯ = 0.98 (p < 10โปยณ) for source adherence and ฯ = 0.03 (p = 0.95) for photorealism.

A property of the goal images, not of one policy

GR-MG and GHIL-Glue rank the nine generators near-identically (ฯ = 0.96), and the dissociation reproduces on SIMPLER-Bridge.

Visual goals beat text goals at every horizon

Image + text conditioning beats text-only conditioning at every chain length: 73.1% vs 59.0% at CL 1, widening to 30.9% vs 0.2% at CL 5.

Method

SGE-Goal: ground the edit, then synthesize

Given the current observation Io and an instruction T, SGE-Goal predicts where the edit happens before it decides what the goal looks like. Grounding here means 2D localization of the edit as binary edit-region masks, not metric 3D reasoning.

Training of the final-mask predictor and the conditional goal-image generator, and the three-step inference pipeline.
Training and inference. LoRA adapters on a frozen base editor predict the final edit-region mask (rectified-flow + channel-consistency loss) and synthesize the goal image (rectified-flow + static-scene-consistency loss). At inference, the initial mask comes from a training-free grounding module.
1

Initial mask Mi

Training-free. Target objects parsed from the instruction are detected with Florence-2 and Grounding DINO, verified by BLIP-2 VQA, and segmented by SAM.

2

Final mask Mf

A LoRA on the base editor predicts where the objects end up after manipulation, with a channel-consistency loss (ฮป = 25, gated to ฯƒ < 0.35).

3

Goal image รŽe

A second LoRA synthesizes the goal from (Io, T, Mi, Mf); a static-scene-consistency loss (ฮป = 7) preserves everything outside the edit region.

SGE-Goal operates on atomic instructions: composite instructions are decomposed into their atomic constituents and processed in sequence, with each generated goal becoming the next step's input.

Data

Training data from raw robot video, without manual annotation

Each sample is a tuple (Io, Ie, Mi, Mf, T) extracted automatically by chaining vision foundation models.

Data-extraction pipeline: EVF-SAM and BLIP-2 for arm detection, SAM and XMem for object tracking, LoFTR and KMeans for movement quantification, mask and edited-frame generation, and VideoChat-Flash for instruction generation.
Automated data extraction. Robotic-arm masking and reference-frame selection (EVF-SAM, BLIP-2), object proposals and tracking (SAM, XMem), movement quantification (LoFTR, 2-means), and mask / edited-frame generation. For datasets without instructions, VideoChat-Flash writes one from the video.

Released datasets

DatasetTrainValTestHugging Face
BridgeDataV2 (teleoperated)40,8511,0002,000RiddhiCh/LGRM_BridgeDataV2_Teleoperated
CALVIN-ABCD22,0009661,087RiddhiCh/Calvin-Dataset-Final
LIBERO-10/90 (combined)6,464358376RiddhiCh/LGRM_Libero_10_90_Combined

Splits are assigned by MD5 hashing (seed 42). Rows carry original_image, edited_image, initial_edit_region_mask, final_edit_region_mask and edit_instruction.

Violin plots of human quality ratings for each extracted component across the three datasets.
Extraction quality. 50 annotators rated 500 tuples per dataset on each extracted component (1โ€“5). The lowest per-component mean across all three datasets is 4.44.
Results

Downstream policy success

Only the goal generator varies; the goal-conditioned policies are pre-trained and frozen. FK = FLUX.1 Kontext [Dev], F2K = FLUX.2 [klein], base = untuned editor.

Success rate vs. chain length on CALVIN
SR % (โ†‘) for five-step chained tasks

CALVIN (GR-MG and GHIL-Glue, chain lengths 1โ€“5)

GR-MGGHIL-Glue
Method1234512345
GT goal (oracle)73.6โ€“โ€“โ€“โ€“39.0โ€“โ€“โ€“โ€“
SGE-Goal (FK)73.159.747.336.830.929.212.03.70.80.0
SGE-Goal (F2K)71.853.834.821.417.322.88.11.50.10.0
Gemini 3 Pro Image63.336.918.18.14.413.22.40.00.00.0
Clover63.234.212.83.61.917.33.20.20.00.0
GPT Image 1.560.236.717.16.23.48.00.70.00.00.0
Qwen Image Edit59.433.713.44.62.27.21.10.00.00.0
SuSIE55.828.49.92.61.06.10.10.00.00.0
FLUX.2 klein (base)52.029.613.45.22.92.90.00.00.00.0
FLUX.1 Kontext (base)50.328.612.84.82.62.60.00.00.00.0
A five-step CALVIN chain: pull the switch down, put blue block on the table, put pink block on the table, close the drawer and move the blue block to the left.
Chained prediction. SGE-Goal goals across a long-horizon CALVIN chain, one atomic step at a time.

SIMPLER-Bridge (GHIL-Glue)

MethodSpoonCarrotStackEggplantMean
Gemini 3 Pro Image16.729.20.04.212.5
SGE-Goal (F2K)16.716.70.08.310.4
SGE-Goal (FK)8.320.80.04.28.3
SuSIE16.78.34.20.07.3
Qwen Image Edit8.34.20.00.03.1
GPT Image 1.54.20.00.04.22.1
FLUX.2 klein (base)4.24.20.00.02.1
FLUX.1 Kontext (base)4.24.20.00.02.1
Clover4.20.00.00.01.0

SR % on the WidowX / BridgeData-V2 simulator, four atomic tasks ร— 24 episodes. GHIL-Glue is SuSIE's own low-level policy, yet both SGE-Goal variants surpass SuSIE.

Adherence vs. photorealism

CALVINSIMPLER
Dimension (|ฯ| with SR)GR-MGGHILBridge
R14 ยท source adherence0.920.890.85
R15 ยท destination adherence0.900.790.66
R13 ยท photorealism0.250.070.02

Absolute Spearman correlation between each dimension of our 15-dimension evaluator and downstream success (CALVIN averaged over chain lengths 1โ€“5). Adherence predicts success on every benchmark; photorealism does not.

The inversion: editing quality vs. control utility

Ranked by editing quality, the field's usual ordering returns โ€” and it is wrong for manipulation. Gemini tops our human-aligned evaluator and GPT tops I2I-Bench, yet both are only mid-pack downstream.

Our evaluator ยท rank โ†“I2I-Bench ยท score โ†‘
MethodBridgeCALVINLIBEROBridgeCALVINLIBERO
Gemini 3 Pro Image2.072.651.810.7850.7070.675
GPT Image 1.55.155.093.040.8010.7160.774
Qwen Image Edit4.144.172.060.6460.6420.680
SGE-Goal (F2K)2.993.414.980.6550.6260.484
SGE-Goal (FK)3.813.185.540.6490.6160.479
Cloverโ€“5.48โ€“โ€“0.606โ€“
SuSIE6.167.865.390.4860.4070.377
FLUX.1 Kontext (base)5.487.336.510.5330.4490.378
FLUX.2 klein (base)6.205.836.660.5680.4700.324

Our evaluator agrees with human judgment on 91.7% of 278 annotated triplets (I2I-Bench: 71.2%). It is a diagnostic that explains downstream behaviour, not a substitute for downstream evaluation.

Ablations: every spatial component helps

CALVIN ยท GR-MGCALVIN ยท GHIL-GlueSIMPLER
Variant1234512345Bridge
FLUX.1 Kontext
SGE-Goal (full)73.159.747.336.830.929.212.03.70.80.08.3
w/o final mask60.128.412.95.73.513.80.30.00.00.05.2
w/o masks50.721.37.52.81.17.30.00.00.00.05.2
base50.328.612.84.82.62.60.00.00.00.02.1
FLUX.2 klein
SGE-Goal (full)71.853.834.821.417.322.88.11.50.10.010.4
w/o final mask63.535.015.25.72.913.01.50.00.00.07.3
w/o masks60.032.112.54.11.912.61.20.00.00.02.1
base52.029.613.45.22.92.90.00.00.00.02.1

โ€œfullโ€ uses both masks and the static-scene loss. The final edit-region mask is the dominant ingredient, and the masks matter most under long-horizon chaining.

Runtime

Inference stage (A10G)Time (s)
Initial mask Mi2.47
Final mask Mf5.16
Goal image รŽe2.92
Total10.55
Data generation (A6000, one-time)Time (s)
EVF-SAM + BLIP-24.46
SAM (first frame)1.07
XMem tracking7.76
LoFTR + KMeans8.74
Qualitative

Goal images side by side

SGE-Goal moves the instructed object to the instructed place and leaves the rest of the scene alone. General editors often change unmentioned objects, ignore the instruction or hallucinate content.

Click an image to open it at full resolution.

Citation

BibTeX

@inproceedings{chatterjee2027goalimages,
  title     = {Goal Images Are Control Targets: Spatially-Grounded
               Goal-Image Synthesis for Robotic Manipulation},
  author    = {Chatterjee, Riddhi and Shrirao, Ritish and Mehrotra, Kushagra
               and Gopalakrishnan, Viswanath},
  booktitle = {Proceedings of the IEEE/CVF Winter Conference on
               Applications of Computer Vision (WACV)},
  year      = {2027}
}
This entry is temporary: the paper has been accepted at WACV 2027 but is not yet published. It will be updated with the proceedings details.