Rubric-CEPR

Self-Evolving Image Editing via Reward-Verified Self-Distillation

Ritesh Thawkar1 Shubham Patle1 Shravan Venkatraman1 Rao Muhammad Anwer1,2
1Mohamed bin Zayed University of Artificial Intelligence 2Aalto University
Overview of Rubric-CEPR: a Planner proposes edits, an Editor samples candidates, a frozen Critic verifies them, and the best verified candidate is distilled into the Editor and Planner.
A Planner proposes structured edits on unlabeled images, the Editor samples K candidates, and a frozen internal Critic keeps only the best gate-passing candidate as a training target.

Overview

Can an image editor improve itself by verifying its own edits?

Improving instruction-guided editors usually needs human-edited training pairs or an external reward model, and a single scalar score can reward plausible failures such as an unchanged image or a lost background. Rubric-CEPR checks each candidate edit with the editor's own features, rejects any candidate that fails a check, and distills only the verified ones into a lightweight adapter.

No human-edited targets

Training uses unlabeled source images. Every target is one of the editor's own samples that passed verification.

No external reward model in training

The Critic reuses features already exposed by the editor. External judges are used only to evaluate on benchmarks.

Single-shot at test time

The adapted editor is evaluated with one generation per instruction: ImgEdit 4.36 → 4.60 on Qwen-Image-Edit, and a gain on Step1X-Edit too.

Why it is hard

Plausible failures look like successes.

Self-training an editor is risky because an output that passes a lenient score becomes supervision for the next model.

01

Scalar scores are not enough

On a fixed 210-pair probe, a scalar reward still assigns 0.39–0.60 to no-op, corrupted and wrong-edit candidates. Our fix: separate checks for edit realization, removal of the old state and preservation, combined with non-compensatory gates.

02

Naive self-training gains less

Training on the editor's own candidates without gates improves ImgEdit by only +2.0%, versus +5.5% for Rubric-CEPR, because rejected candidates outnumber usable targets. Our fix: distill only gate-passing, best-scoring candidates.

Framework

Propose, edit, verify, distill.

A frozen Qwen-Image-Edit backbone provides understanding features and VAE latents. Lightweight Planner and Editor adapters are updated; the Critic stays fixed.

Rubric-CEPR framework: frozen Qwen-Image-Edit backbone, Planner, Editor, internal rubric Critic and the self-evolution loop.
01

Propose Planner

Emits an edit instruction and a structured specification: edit type, entities, target region, required and forbidden after-states, and content to preserve.

02

Edit Editor

Samples K candidate edits for each proposal and is the model that is finally improved, through a LoRA adapter trained on verified targets.

03

Verify Critic

A fixed internal Critic scores each candidate with the Contrastive Edit-Preservation Reward and rubric checks. Any failed gate makes a candidate ineligible.

Key ideas

What makes verified self-distillation work.

1

Contrastive Edit-Preservation Reward (CEPR)

CEPR contrasts support for the requested edit with incorrect alternative instructions, and rewards preservation of unrelated content. The rubric adds explicit checks for required states, removal of forbidden old states, and preservation constraints.

2

Non-compensatory gates

The reward is R = G · √(Edit × Preserve). A candidate that fails any applicable check gets zero reward, so a high score on one term cannot hide a failed edit or lost content.

3

Distill the verified best-of-K

Best-of-K sampling already finds better edits than the base editor produces in one shot. Distilling the best gate-passing candidate into an adapter turns that headroom into a single-shot checkpoint.

Results

Matched base-versus-ours gains.

Every adapted editor is compared with its own base under the same protocol, with one generation per instruction. The ImgEdit overall score is the mean over three training seeds (4.58, 4.60, 4.62; 4.60 ± 0.02 SD).

Table 1 · ImgEdit · Qwen-Image-Edit-2509

Nine edit families · scores 0–5 · Δ Improvement over base
ModelImgEdit
AddRemoveReplaceAdjustActionStyleBackgroundExtractComposeOverall
Image editing models
GPT-4o-Image4.653.814.494.264.764.754.622.964.544.32
Step1X-Edit3.902.613.453.133.434.443.191.872.523.17
ImgEdit-E13.822.402.804.043.214.383.382.552.873.27
UltraEdit3.631.713.133.013.573.693.312.022.332.93
MagicBrush–––––––––1.90
InstructPix2Pix2.291.491.931.791.513.541.671.331.481.89
Reward post-trained methods
UniWorld-Qwen-Edit†4.344.654.624.404.804.894.314.373.904.48
UniWorld-V2†4.294.724.694.444.834.914.414.323.834.49
Qwen-Image-Edit-25094.514.364.724.404.694.704.403.414.054.36
Rubric-CEPR (ours)4.654.504.864.544.824.834.594.264.394.60
Δ Improvement+3.1%+3.2%+3.0%+3.2%+2.8%+2.8%+4.3%+24.9%+8.4%+5.5%
Comparison with image editing methods on ImgEdit. We compare published editors, UniWorld (external reward), and our Rubric-CEPR, which self-distills Qwen-Image-Edit with only internal rewards. Rubric-CEPR improves every family and raises the overall score from 4.36→4.60 (+5.5%; +0.24±0.02 SD over three seeds), with the largest gain on extraction (3.41→4.26, +24.9%), the weakest base family. Scores are on the 0–5 scale and Δ Improvement is the relative improvement over the base, 100×(adapted-base)/base. †UniWorld post-trains Qwen-Image-Edit with an external multimodal-LLM reward model (training-free, logit-based) and reports 4.35→4.48 (+3.0%) over its own base. Rows from other papers use their own judges and are not directly comparable.

Table 2 · GEdit-Bench and Complex-Edit

Scores 0–10 · Δ Improvement over base
ModelGEdit-BenchComplex-Edit
Semantic
Consistency
Perceptual
Quality
OverallInstruction
Following
Identity
Preservation
Perceptual
Quality
Overall
Image editing models
GPT-4o7.748.137.499.297.519.478.76
Imagen3–––7.566.557.677.26
SeedEdit7.227.896.988.496.918.748.04
Step1X-Edit7.137.006.44––––
Step1X-Edit v1.17.667.356.97––––
Gemini 2.0 Flash6.877.446.51––––
OmniGen5.885.875.016.256.427.546.74
AnyEdit3.055.882.851.608.157.255.67
UltraEdit–––6.565.937.296.59
MagicBrush4.526.374.19––––
InstructPix2Pix3.306.193.22––––
Reward post-trained methods
UniWorld-Qwen-Edit†8.367.877.769.739.087.618.81
UniWorld-V2†8.398.027.83––––
Qwen-Image-Edit-25098.117.107.399.699.027.608.77
Rubric-CEPR (ours)9.057.928.319.759.187.978.97
Δ Improvement+11.6%+11.5%+12.4%+0.7%+1.8%+4.9%+2.3%
Transfer of Rubric-CEPR to GEdit-Bench and Complex-Edit. We evaluate the same checkpoint as in Table 1 without any further adaptation. Rubric-CEPR improves every reported metric over its Qwen-Image-Edit base, raising the GEdit-Bench overall score from 7.39→8.31 (+12.4%) and the Complex-Edit overall score from 8.77→8.97 (+2.3%), with the largest Complex-Edit gains on perceptual quality (+4.9%) and identity preservation (+1.8%). This suggests that the verified targets improve editing quality beyond the edit types emphasized during training. Scores are on the 0–10 scale; GEdit-Bench reports Semantic Consistency, Perceptual Quality, and Overall VIEScore; Complex-Edit reports Instruction Following, Identity Preservation, Perceptual Quality, and Overall, the mean of the three metrics. Results of other methods are reported under their own protocols. Δ Improvement is the relative improvement over the base. †UniWorld uses an external multimodal-LLM reward model during post-training and reports 7.54→7.76 (+2.9%) over its own Qwen-Image-Edit base on GEdit-Bench.

Table 3 · Second editor · Step1X-Edit

GEdit-Bench 0–10 · ImgEdit 0–5 · Δ Improvement over base
ModelGEdit-BenchImgEdit
Semantic
Consistency
Perceptual
Quality
OverallOverall
Step1X-Edit7.077.586.693.86
Rubric-CEPR (ours)7.848.057.244.16
Δ Improvement+10.9%+6.2%+8.2%+7.8%
Rubric-CEPR generalizes to a second editor, Step1X-Edit. We apply the same procedure to Step1X-Edit using its own VLM and VAE features, without any Qwen-Image-Edit weights. Self-distillation raises the overall GEdit-Bench score from 6.69→7.24 (+8.2%) and the overall ImgEdit score from 3.86→4.16 (+7.8%). This shows that internal verification is not specific to a single editor. Δ Improvement is the relative improvement over the base. The overall gains are +0.55±0.05 SD over three training seeds for GEdit-Bench and +0.30±0.06 s.e. for ImgEdit.

Mechanism evidence

Why verification, headroom and base quality matter.

Three panels: gate validation against a scalar reward, best-of-K headroom on the extraction family, and the share of headroom closed per ImgEdit family.
(a) Gate validation. A scalar reward scores invalid candidates 0.39–0.60, while the feasibility gates accept 51% of valid edits and none of the invalid ones. (b) Best-of-K headroom. On extraction, reward-best selection rises from 3.41 (K=1) to 4.45 (K=8), and the single-shot distilled checkpoint reaches 4.26, on par with best-of-4. (c) Headroom closed. The share of the gap to the 5.0 judge ceiling that Rubric-CEPR closes is 22–53% on every family, so the large extraction gain largely reflects its large headroom.

Qualitative results

Completing edits while preserving the rest.

The baseline often returns a realistic image that leaves the requested change undone. The adapted editor completes the edit and keeps the content the instruction leaves unchanged.

Qualitative comparison between Qwen-Image-Edit and Rubric-CEPR on ImgEdit.
ImgEdit. Input, baseline and adapted outputs for extraction, removal, background, replacement and action edits.
Additional high-headroom qualitative results.
High-headroom edits. Object extraction, removal and background edits where the baseline fails to realize the operation.
Qualitative results on other edit families.
Other edit families. Replacement, addition, adjustment, action and style edits, where the baseline is often already strong.

Citation

Cite this work.

Copy the BibTeX entry below.

@article{thawkar2026rubriccepr,
  title={Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Distillation},
  author={Thawkar, Ritesh and Patle, Shubham and Venkatraman, Shravan and Anwer, Rao Muhammad},
  journal={arXiv preprint arXiv:2610.12469},
  year={2026}
}