No human-edited targets
Training uses unlabeled source images. Every target is one of the editor's own samples that passed verification.
Self-Evolving Image Editing via Reward-Verified Self-Distillation
Overview
Improving instruction-guided editors usually needs human-edited training pairs or an external reward model, and a single scalar score can reward plausible failures such as an unchanged image or a lost background. Rubric-CEPR checks each candidate edit with the editor's own features, rejects any candidate that fails a check, and distills only the verified ones into a lightweight adapter.
Training uses unlabeled source images. Every target is one of the editor's own samples that passed verification.
The Critic reuses features already exposed by the editor. External judges are used only to evaluate on benchmarks.
The adapted editor is evaluated with one generation per instruction: ImgEdit 4.36 → 4.60 on Qwen-Image-Edit, and a gain on Step1X-Edit too.
Why it is hard
Self-training an editor is risky because an output that passes a lenient score becomes supervision for the next model.
On a fixed 210-pair probe, a scalar reward still assigns 0.39–0.60 to no-op, corrupted and wrong-edit candidates. Our fix: separate checks for edit realization, removal of the old state and preservation, combined with non-compensatory gates.
Training on the editor's own candidates without gates improves ImgEdit by only +2.0%, versus +5.5% for Rubric-CEPR, because rejected candidates outnumber usable targets. Our fix: distill only gate-passing, best-scoring candidates.
Framework
A frozen Qwen-Image-Edit backbone provides understanding features and VAE latents. Lightweight Planner and Editor adapters are updated; the Critic stays fixed.
Emits an edit instruction and a structured specification: edit type, entities, target region, required and forbidden after-states, and content to preserve.
Samples K candidate edits for each proposal and is the model that is finally improved, through a LoRA adapter trained on verified targets.
A fixed internal Critic scores each candidate with the Contrastive Edit-Preservation Reward and rubric checks. Any failed gate makes a candidate ineligible.
Key ideas
CEPR contrasts support for the requested edit with incorrect alternative instructions, and rewards preservation of unrelated content. The rubric adds explicit checks for required states, removal of forbidden old states, and preservation constraints.
The reward is R = G · √(Edit × Preserve). A candidate that fails any applicable check gets zero reward, so a high score on one term cannot hide a failed edit or lost content.
Best-of-K sampling already finds better edits than the base editor produces in one shot. Distilling the best gate-passing candidate into an adapter turns that headroom into a single-shot checkpoint.
Results
Every adapted editor is compared with its own base under the same protocol, with one generation per instruction. The ImgEdit overall score is the mean over three training seeds (4.58, 4.60, 4.62; 4.60 ± 0.02 SD).
| Model | ImgEdit | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Add | Remove | Replace | Adjust | Action | Style | Background | Extract | Compose | Overall | |
| Image editing models | ||||||||||
| GPT-4o-Image | 4.65 | 3.81 | 4.49 | 4.26 | 4.76 | 4.75 | 4.62 | 2.96 | 4.54 | 4.32 |
| Step1X-Edit | 3.90 | 2.61 | 3.45 | 3.13 | 3.43 | 4.44 | 3.19 | 1.87 | 2.52 | 3.17 |
| ImgEdit-E1 | 3.82 | 2.40 | 2.80 | 4.04 | 3.21 | 4.38 | 3.38 | 2.55 | 2.87 | 3.27 |
| UltraEdit | 3.63 | 1.71 | 3.13 | 3.01 | 3.57 | 3.69 | 3.31 | 2.02 | 2.33 | 2.93 |
| MagicBrush | – | – | – | – | – | – | – | – | – | 1.90 |
| InstructPix2Pix | 2.29 | 1.49 | 1.93 | 1.79 | 1.51 | 3.54 | 1.67 | 1.33 | 1.48 | 1.89 |
| Reward post-trained methods | ||||||||||
| UniWorld-Qwen-Edit† | 4.34 | 4.65 | 4.62 | 4.40 | 4.80 | 4.89 | 4.31 | 4.37 | 3.90 | 4.48 |
| UniWorld-V2† | 4.29 | 4.72 | 4.69 | 4.44 | 4.83 | 4.91 | 4.41 | 4.32 | 3.83 | 4.49 |
| Qwen-Image-Edit-2509 | 4.51 | 4.36 | 4.72 | 4.40 | 4.69 | 4.70 | 4.40 | 3.41 | 4.05 | 4.36 |
| Rubric-CEPR (ours) | 4.65 | 4.50 | 4.86 | 4.54 | 4.82 | 4.83 | 4.59 | 4.26 | 4.39 | 4.60 |
| Δ Improvement | +3.1% | +3.2% | +3.0% | +3.2% | +2.8% | +2.8% | +4.3% | +24.9% | +8.4% | +5.5% |
| Model | GEdit-Bench | Complex-Edit | |||||
|---|---|---|---|---|---|---|---|
| Semantic Consistency | Perceptual Quality | Overall | Instruction Following | Identity Preservation | Perceptual Quality | Overall | |
| Image editing models | |||||||
| GPT-4o | 7.74 | 8.13 | 7.49 | 9.29 | 7.51 | 9.47 | 8.76 |
| Imagen3 | – | – | – | 7.56 | 6.55 | 7.67 | 7.26 |
| SeedEdit | 7.22 | 7.89 | 6.98 | 8.49 | 6.91 | 8.74 | 8.04 |
| Step1X-Edit | 7.13 | 7.00 | 6.44 | – | – | – | – |
| Step1X-Edit v1.1 | 7.66 | 7.35 | 6.97 | – | – | – | – |
| Gemini 2.0 Flash | 6.87 | 7.44 | 6.51 | – | – | – | – |
| OmniGen | 5.88 | 5.87 | 5.01 | 6.25 | 6.42 | 7.54 | 6.74 |
| AnyEdit | 3.05 | 5.88 | 2.85 | 1.60 | 8.15 | 7.25 | 5.67 |
| UltraEdit | – | – | – | 6.56 | 5.93 | 7.29 | 6.59 |
| MagicBrush | 4.52 | 6.37 | 4.19 | – | – | – | – |
| InstructPix2Pix | 3.30 | 6.19 | 3.22 | – | – | – | – |
| Reward post-trained methods | |||||||
| UniWorld-Qwen-Edit† | 8.36 | 7.87 | 7.76 | 9.73 | 9.08 | 7.61 | 8.81 |
| UniWorld-V2† | 8.39 | 8.02 | 7.83 | – | – | – | – |
| Qwen-Image-Edit-2509 | 8.11 | 7.10 | 7.39 | 9.69 | 9.02 | 7.60 | 8.77 |
| Rubric-CEPR (ours) | 9.05 | 7.92 | 8.31 | 9.75 | 9.18 | 7.97 | 8.97 |
| Δ Improvement | +11.6% | +11.5% | +12.4% | +0.7% | +1.8% | +4.9% | +2.3% |
| Model | GEdit-Bench | ImgEdit | ||
|---|---|---|---|---|
| Semantic Consistency | Perceptual Quality | Overall | Overall | |
| Step1X-Edit | 7.07 | 7.58 | 6.69 | 3.86 |
| Rubric-CEPR (ours) | 7.84 | 8.05 | 7.24 | 4.16 |
| Δ Improvement | +10.9% | +6.2% | +8.2% | +7.8% |
Mechanism evidence
Qualitative results
The baseline often returns a realistic image that leaves the requested change undone. The adapted editor completes the edit and keeps the content the instruction leaves unchanged.
Citation
Copy the BibTeX entry below.
@article{thawkar2026rubriccepr,
title={Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Distillation},
author={Thawkar, Ritesh and Patle, Shubham and Venkatraman, Shravan and Anwer, Rao Muhammad},
journal={arXiv preprint arXiv:2610.12469},
year={2026}
}