Under Review

ReGain: Restoring Subject Fidelity in Personalization on Synthetic Images

University of Illinois at Urbana-Champaign

TL;DR

DreamBooth on generated images of a subject produces oversaturated, over-textured subjects. We trace this to classifier-free guidance: the guidance term is inflated because the angle between the conditional and unconditional predictions widens. ReGain corrects it at sampling time, per frequency band and denoising step, with no real photos or retraining.

One of the real photos of the subject, a corgi puppy Real photo of the subject
“a photo of a [V] dog in front of the Acropolis”

Both images are sampled from the same Stable Diffusion v1.5 model, fine-tuned on synthetic images of this subject, with the same prompt and seed. Left: classifier-free guidance (CFG) alone. Right: CFG with ReGain.

51 to 64%
of the lost subject fidelity recovered (SD1.5)
Training-free
applied at sampling time, with no real photos or retraining
≤ 0.004
change in prompt fidelity (CLIP-T) in every setting
1.1×
the sampling time of plain CFG (SD1.5)

Abstract

Text-to-image diffusion models are personalized to a subject by DreamBooth fine-tuning on a handful of its images. Increasingly, these images come from a diffusion model rather than a camera. We show that fine-tuning on such synthetic images degrades subject fidelity, producing oversaturated color and excess high-frequency detail. To isolate the cause, we fine-tune two models from the same base model with the same DreamBooth recipe, one on real photos of a subject and one on synthetic images of that subject generated by the first. We trace the degradation to classifier-free guidance (CFG). For the model personalized on synthetic images, the angle between the conditional and unconditional noise predictions, and with it the norm of their difference, is much larger than for the model personalized on real photos. This inflation grows toward high frequencies and also appears at other prompts semantically close to the subject, such as its class noun, but not at unrelated ones. We propose ReGain, a training-free correction applied at sampling time that measures how much each frequency band of the guidance is inflated relative to the base model and scales that band down accordingly. ReGain needs no real photos. On Stable Diffusion v1.5, ReGain closes 51–64% of the subject-fidelity gap to the model personalized on real photos, as measured by DINO, DINOv2 and CLIP-I. It also improves subject fidelity on SDXL and SD 3.5 and preserves text alignment on all three backbones.

Problem

Personalizing a diffusion model on synthetic images of a subject degrades its fidelity. Why?

To study this, we train diffusion models Mreal and Msyn with the same recipe, the only difference being their training data: real photos and synthetic images respectively.

Real photo of the subject Real photo of the subject Real photo of the subject
Real photos of the subject
DreamBooth from Mbase
Mreal
sample, same prompt and seed
Clean sample generated by the model trained on real photos
Generated: clean
Mreal generates the five synthetic training images of the subject
Synthetic image of the subject Synthetic image of the subject Synthetic image of the subject
Synthetic images of the subject
DreamBooth from Mbase, same recipe
Msyn
sample, same prompt and seed
Oversaturated sample with excess texture generated by the model trained on synthetic images
Generated: oversaturated, excess texture

Prompt for both samples: “a photo of a [V] dog in front of the Acropolis”, where [V] is the identifier that DreamBooth binds to the subject.

Both models are fine-tuned with DreamBooth from Mbase and sampled with the same prompt and seed.

Findings

Notation

Classifier-free guidance (CFG) with weight w:Δ = εc − ε∅ε̂ = ε∅ + w Δ

εc and ε∅ are the noise predictions with and without the prompt.

The outputs of Msyn carry more power at every frequency.

Power spectrum of generated images in the latent space.

The guidance term Δ of Msyn is inflated throughout denoising.

The plots in Findings are for this subject on SD1.5. Step t = 0 is the noisiest. The two plots above show the mean ± s.e. over 10 seeds.

The Δ inflation, the excess in ‖Δ‖ of Msyn over Mreal, has four properties.

1Angle-driven.

The norms of εc and ε∅ do not grow. The angle θ between them is wider for Msyn at every step, so Δ inflates.

Mreal trained on real photos

Msyn trained on synthetic images

Angles are drawn 6× wider so that they are visible. Mean over 10 seeds.

2Semantic.

The class name (“a dog”) shows nearly the full inflation and a related class (“a cat”) about half. Unrelated prompts show little to none.

Angle excess θMsyn(t) − θMreal(t) for the subject prompt and increasingly semantically distant prompts.

3Persists on the same input.

When Mreal and Msyn denoise the same noisy copies of the training images of Msyn (as opposed to their own sampling trajectories in the plots above), the Δ of Msyn is still larger. The inflation is a property of the model.

4High-frequency and time-varying.

High-frequency bands inflate several-fold, and each band follows its own time profile. A lower guidance weight w acts uniformly across bands and steps, so it cannot undo this.

Energy of Δ in each radial frequency band k and step t, Msyn divided by Mreal; red marks inflation.

low k (1-8) mid k (9-24) high k (25-45)

Excess of the angle θ of Msyn over Mreal, in percent, within the low, mid and high bands.

Method

ReGain uses the base model Mbase as the reference for how large Δ should be, because Mbase is freely available and its Δ for the subject’s class is uninflated.

ReGain overview: (A) personalization on synthetic images, (B) the wider angle inflates the guidance term, (C) one-time calibration against the base model and per-step band-wise correction inside the subject mask
Overview.
AMsyn is DreamBooth fine-tuned on synthetic subject images generated by Mreal.
BThe angle θ between the unconditional and conditional predictions is wider for Msyn, so its guidance term Δ = εc − ε∅ is inflated.
C
ReGain.
(i)Once per model, Mbase and Msyn are evaluated at forward-noised training images zt; the square root of the ratio of their masked band energies e gives the gain ĝ(k, t), compressed into the schedule gb(t) (darker blue attenuates more; the gain schedule below).
(ii)At each step, Δ is rescaled per band group by gb(t) inside the subject mask m (green) and left unchanged outside it (grey), then scaled by w and added to ε∅, the ReGain update below.

csubj is the prompt with the identifier [V] and cclass is the same prompt without it.

C (i)Gain schedule gb(t)
g^(k,t)=eMbase(k,t)eMsyn(k,t)compressed intogb(t)
e(k, t)energy of Δ in frequency band k at step t, inside the subject mask
ĝ(k, t)gain for band k at step t
gb(t)gain schedule, ĝ compressed into a few cells of constant gain, one row per band group b

For different subjects, measured using our setup.

C (ii)ReGain update:
ε̂ = ε∅+w m ⊙ ∑b gb(t) Bb(Δ)inside the subject mask+w (1 − m) ⊙ Δoutside the mask
ε̂corrected noise prediction
mbinary mask of the subject
bband group, a contiguous range of frequency bands that share one gain
gb(t)gain schedule, the factor that scales band group b at step t
Bb(Δ)component of Δ in band group b
⊙elementwise product

Results

Visual results

Each card is one subject, prompt and seed. Drag the slider to compare plain CFG with ReGain at the same guidance weight w.

Quantitative improvements in subject fidelity on the DreamBooth dataset

Each track runs from Msyn with plain CFG to Mreal. The number is the share of that gap that ReGain recovers.

Show Table 1
SD1.5, DreamBoothSD1.5, DreamBooth-LoRA
MethodCLIP-IDINODINOv2CLIP-TCLIP-IDINODINOv2CLIP-T
Mreal + CFG0.8140.6810.6610.3000.7910.6300.6060.306
Msyn + CFG0.7880.6110.6110.2890.7500.5400.5100.300
Msyn + S-CFG0.7810.6050.6040.2920.7440.5360.5030.302
Msyn + CFG++0.7880.6060.6080.2890.7480.5350.5080.299
Msyn + FDG0.7920.6080.6140.2840.7550.5380.5170.297
Msyn + CFG + ReGain0.8020.6470.6430.2900.7750.5940.5680.304
  ↪ Gap recovered54%51%64%+0.00161%60%60%+0.004
SDXL, DreamBooth-LoRASD3.5, DreamBooth-LoRA
Mreal + CFG0.7570.5580.5350.2890.8110.6820.6500.311
Msyn + CFG0.7240.4770.4670.2780.7760.5860.5670.307
Msyn + CFG + ReGain0.7360.5020.4940.2790.7920.6050.5960.307
  ↪ Gap recovered36%31%40%+0.00146%20%35%0.000

Table 1: Subject and prompt fidelity on DreamBooth, averaged over 30 subjects, 25 prompts and 3 seeds (higher is better). The first row is the reference model Mreal, and every other row samples from Msyn. Baselines, on SD1.5 only: S-CFG (Shen et al., 2024), CFG++ (Chung et al., 2025) and FDG (Sadat et al., 2025b). Bold marks the best Msyn result. Gap recovered is the part of the drop from Mreal to Msyn, both with CFG, that ReGain wins back. For CLIP-T we report the change from Msyn + CFG instead.

Over-guidance artifacts

ReGain moves saturation, contrast and high-band power inside the subject region back toward Mreal.

Show Table 2
MethodSaturationContrastHigh band
Mreal + CFG0.3010.2110.061
Msyn + CFG0.3550.2580.069
Msyn + CFG + ReGain0.3090.2340.066
↪ Gap recovered85%51%38%

Table 2: Over-guidance artifacts inside the subject region (SD1.5, DreamBooth). Bold marks the Msyn row closer to Mreal.

BibTeX

If you find our work useful, please consider citing:

@misc{bhatnagar2026regainrestoringsubjectfidelity,
      title={ReGain: Restoring Subject Fidelity in Personalization on Synthetic Images},
      author={Shubhang Bhatnagar and Ishan Bhatnagar and Viraj Shah and Narendra Ahuja},
      year={2026},
      eprint={2609.38680},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.38680},
}