Skip to content

Introducing Photoroom’s Fidelity Layer for Image Editing

Product fidelity is one of the biggest barriers to using generative imagery in commerce: an image can look convincing while quietly changing the product being sold. In our benchmark, even the strongest base model preserved the complete product in only 29.0% of cases, with the other leading frontier models performing similarly. That is why we are introducing the Photoroom Fidelity Layer, a reasoning system built specifically for product imagery that combines product understanding, visual verification, and targeted correction around image-generation models.

Image-editing models have improved quickly. They can now turn a simple product photo into a convincing lifestyle image or place a garment on a virtual model. The resulting images are often realistic enough to look ready for product catalogs.

Example catalog listing generated from a single product image, including a flat lay view and front and back virtual model shots.

For brands, however, realism is not enough. The generated image must also preserve the product being sold, a requirement we refer to as product fidelity.

Sometimes the failure is easy to see. The chest pocket may have a different construction, an additional button may appear on a cuff, color may slightly drift, or an element may disappear altogether. The image can still be attractive, but it is no longer an accurate representation of the product. This can break consistency across a marketing campaign, or the trust of a user buying the product online.

Generated catalog detail closeups highlighting key product features, including the collar, chest pocket, and cuff construction.

A person can usually catch these errors and try again. They can adjust the prompt, generate another image, or manually edit the affected region. This is workable when producing a small number of images, although it already requires time, attention, and some understanding of why the model failed. At scale, however, this becomes much harder to manage.

The more difficult failures are much less obvious. A button may still be present but have a distorted shape or construction, or the shape of a small embroidered emblem may change. Unless someone compares the generated image carefully with the source product, a small discrepancy can pass through review and reach a product page, catalogue, or campaign.

Generated campaign hero image featuring the virtual model wearing the product in a studio setting.

At Photoroom, product fidelity has been one of the main barriers to using generative image models at scale in commerce. A system that creates beautiful images but changes the product still requires careful supervision, and the most subtle changes are precisely the ones most likely to escape it. Addressing this required two things: a reliable way to measure the problem, and a system capable of improving it.

Additional generated images from the same campaign, showing consistent identity, styling, and product presentation across multiple compositions.

How to measure product fidelity?

Before trying to improve the problem, we needed to measure it.

We built the Photoroom Product Fidelity Benchmark to evaluate how well leading image-editing models preserve real products. It covers 850 products across clothing, footwear, bags, jewelry, and accessories, with each generation failing if any product discrepancy is detected. We describe the full methodology and results in our benchmark article.

Product Fidelity Leaderboard, July 2026. Four image-editing models evaluated out of the box, plus Nano Banana 2 with Photoroom’s Fidelity Layer applied as a correction system.

Even the strongest base model preserved the complete product in only 29.0% of cases. The three leading frontier models, including Google’s Nano Banana 2 and OpenAI’s GPT Image 2, performed similarly, suggesting that simply switching from one strong model to another would not be enough to make product generation reliable.

Applying the Photoroom Fidelity Layer to the strongest base model increased the pass rate from 29.0% to 38.2%, an improvement of 9.2 percentage points. We will explain how the system achieves this in the following sections.

From one-shot generation to a reasoning system

Most image-generation systems still work as one-shot pipelines. A model receives an image and an instruction, produces a result, and returns it. The system does not usually compare the output with the original product, determine whether important details were preserved, or decide what to do when they were not.

We think reliable product imagery requires a broader system. Language models have shown the value of inspecting intermediate results, reasoning about errors, using tools, and revising an answer. The same general principle can be applied to image generation, with the reasoning grounded in visual comparison.

The Photoroom Fidelity Layer is our approach to this problem. It operates around the complete generation process. It builds an understanding of the source product before the first image is created, gives generation models the context they need to preserve it, and verifies each result against the original. When it detects a discrepancy, it can use that diagnosis to guide another generation or a more targeted correction.

Left: original product detail from the reference image provided as input. Right: virtual-model generation with fidelity errors, where the sleeve placket incorrectly includes an extra button and the shirt color shifts slightly.

To do this, we train in-house visual language models to analyze product fidelity. When they detect a problem, they are trained not only to flag it, but also to describe it and localize the affected region. A useful diagnosis should be specific: a logo has become unreadable, a button is missing from the left sleeve, or the color of a bag has shifted from dark green to black.

Left: original product detail from the reference image provided as input. Right: virtual-model generation with a fidelity error, where the embroidered fox emblem is significantly altered and no longer matches the reference.

This structured diagnosis is used in two ways. It prevents an inaccurate image from being accepted, while also providing generation and editing tools with the information needed to correct it.

Left: original product detail from the reference image provided as input. Right: virtual-model generation with a fidelity error, where the front button is rendered incorrectly as an artifact.

The Fidelity Layer connects its reasoning models with generation and editing tools, including in-house technologies related to Product Fixer, which are designed to correct localized regions while remaining grounded in the original product. Depending on the scope of the issue, the system can choose between a broader regeneration and a targeted correction of the affected region.

Conceptually, the process is straightforward:

Understand → generate → inspect → diagnose → correct or regenerate → verify

The process can repeat, with each new result verified again before it is accepted or discarded.

The difficult part is making every step reliable enough to be useful. The system must identify the details that matter, detect subtle changes without flagging harmless differences, localize problems accurately, preserve the parts of the image that are already correct, and verify that the corrected result is actually better, while still aligning to the user’s intent.

Building reliable image systems

The 9.2 percentage-point improvement shown in the benchmark does not mean that product fidelity is solved. A 38.2% pass rate is still far from the level of reliability required for fully automated production. It does, however, show that improving the system around a generation model can produce meaningful gains over using the model by itself.

Better generation models will remain essential. They will produce stronger images, preserve more details, and give correction systems fewer problems to solve. But for product imagery, model quality alone is unlikely to be the complete answer.

A reliable system also needs to understand the source product, give the generator the right context, inspect each result against the source of truth, explain what went wrong, and use the appropriate tools to correct it. Generation needs to be treated as a process rather than a single inference call.

The Fidelity Layer is designed around this principle. Generation is one part of a broader system that also includes product understanding, visual reasoning, verification, and correction. Because these capabilities operate around the generator, the system can benefit from stronger base models as they become available without making reliability entirely dependent on any one of them.

There is still a great deal of work ahead. We are continuing to improve the reasoning models, the correction strategies, and our own specialized image-generation models.

The goal is simple: brands should be able to generate product imagery at scale without losing confidence that the image still represents the product they are selling.

This work was led by Louis Roussel, Aharon Azulay, Marco Forte, Matthieu Toulemont, and Jon Almazán, with contributions from the wider Machine Learning team at Photoroom.

Jon AlmazánResearch scientist
Introducing Photoroom’s Fidelity Layer for Image Editing

Frequently asked questions

What is the Photoroom Fidelity Layer?

Which Photoroom tools use the Fidelity Layer?

Does the Photoroom Fidelity Layer slow down generation?

How does the Fidelity Layer decide between regeneration and local correction?

Keep reading

How to make AI product images look real
Closing the fidelity gap in AI product photography
How often do top editing image models maintain product details? Only 29% of the time