The model improved; the approval problem did not disappear
OpenAI launched ChatGPT Images 2.5 on 8 September with a practical promise: better preservation of reference subjects, more precise changes and less degradation across multiple editing turns. The API has two versions. Flare is positioned as the faster default for high-volume work, while Sunburst is intended for premium visual workflows that need tighter control and accepts longer latency. OpenAI says Flare delivers higher quality than GPT-Image-2 at 50% lower latency; that is a vendor comparison, not a guarantee for a particular campaign.
For brands, agencies and studios, the meaningful shift is not simply a prettier result. A product, performer, location or layout can now remain recognisable while the team asks for background, wardrobe, copy or compositional changes inside the same working thread. That makes the system feel closer to an editable master than a slot machine.
But consistency is not identity. A face can remain recognisable while its age, expression or skin texture moves. A pack can retain its silhouette while a label, legal line or colour value changes. A layout can feel on-brand while hierarchy, accessibility or claim meaning drifts. The supplier has improved the probability of preservation; the production team still owns the definition and acceptance of preservation.

Create a reference lock before the first edit
A reference lock is a small contract attached to an approved source asset. Give the source a stable ID and hash, then divide its visible properties into three groups: locked, editable and negotiable. Locked elements may include a person's identity and skin tone, a product's geometry and label, a logo, an approved claim, a colour standard or a camera viewpoint. Editable elements are the purpose of the task. Negotiable elements may move within a named tolerance, such as crop, shadow density or prop position.
Write the edit as a delta: ‘replace the neutral background with the approved retail interior; do not change the bottle, label, liquid colour, reflections, camera angle or crop’. Attach the exact source files rather than relying on an earlier message, a descriptive prompt or a thumbnail in a board. If multiple references are supplied, state the role and priority of each one: product truth, talent likeness, lighting, composition, wardrobe or brand palette.
A sketch or annotated frame can make spatial intent clearer than prose. It should not become an unlicensed source by accident. Record who made it, what it controls and whether it is merely directional or must be matched. Reference-led production gets safer when every input has a role, an owner and a permitted use—not when the prompt contains more adjectives.
Treat multi-turn editing as version control
Do not overwrite the accepted image with the latest conversational result. Save every candidate with a parent version, instruction, model snapshot, quality and size, reference IDs, timestamp, operator and cost. A simple sequence might read A001 source, A002 background change, A003 copy correction and A004 approved delivery. Branch when two stakeholders request different directions; do not blend both into an ambiguous next turn.
Reset to the last approved version when an edit introduces unrelated drift. Continuing from a compromised frame compounds the error because the new image becomes the reference for the next instruction. The correct question is not ‘can the model fix it?’ but ‘which version still contains the approved truth?’ That version should remain available even if a chat preview changes or a supplier interface moves on.
OpenAI demonstrates multi-turn consistency with sequences such as a rotating cube. The production lesson is broader: every new instruction creates both the requested change and a new opportunity for cumulative drift. The review system needs to compare the current candidate with the locked source and with its immediate parent, because those two comparisons reveal different failures.

Build the acceptance test around commercial failure
Before generating, decide what would make the asset unusable. For a product visual, measure silhouette, logo and label text, cap and closure, material finish, colour against an approved value, required legal copy and safe areas for every destination. For talent, check consented likeness, age presentation, anatomy, wardrobe, expression, distinguishing features and whether the treatment could imply an endorsement that was never agreed.
Use an overlay, side-by-side review and targeted crops at delivery resolution. OCR can flag copy changes, perceptual comparison can identify large pixel shifts and a colour sample can test a pack value, but none of them replaces a named creative, brand, legal or rights approver. Automation should route attention to likely failures; it should not quietly approve an image because a single similarity score passed.
Set a change budget. For example: zero tolerance for logo, claim and pack geometry; a small agreed tolerance for crop and shadow; subjective approval for background dressing. Then run a representative pilot of perhaps 20 edits, not the easiest demo image. Track first-pass acceptance, hidden-drift rate, human review minutes, correction turns, cost per accepted asset and the percentage that survive final resize and compression without reopening.
Separate platform safety from production clearance
OpenAI's system card describes checks on prompts, image inputs and outputs, and reports results from a fixed adversarial test set. OpenAI explicitly says those prompts are not representative of production traffic and notes that automated labels can contain errors. Those safeguards matter, but they answer whether the service permits or blocks content under its policies. They do not clear a campaign's contracts, trademarks, claims, likeness rights, territories or client rules.
OpenAI's current service terms say visual capabilities may not reproduce a person's likeness without express consent and all necessary rights. Make that consent asset-specific. Record the performer or contributor, permitted transformations, media, territory, duration, exclusivity, sensitive treatments, approval rights and revocation route. Confirm that the team also has rights to every product shot, location image, font, artwork, pack design and reference placed into the workflow.
Keep the model policy and the production policy as two separate gates. A system refusal is not proof that the brief was unlawful; a successful generation is not proof that the output is cleared. Procurement should ask which account terms govern the job, how customer inputs are handled, what evidence the supplier retains and which responsibilities remain with the user before accepting any broad ‘commercially safe’ label.
Keep provenance outside the rendered pixels
OpenAI says Images 2.5 output continues to use C2PA metadata and invisible watermarking through SynthID. The system card also recognises that no single provenance method is sufficient. That is the right boundary for production: embedded signals are useful evidence, but they can be lost or altered when an image is cropped, recompressed, placed in a layout, exported through another tool or captured from a screen.
Keep an external manifest beside the campaign files. It should link the locked source and its hash, every ingredient, consent and licence references, prompts or change instructions, model and snapshot, operator, candidate versions, human edits, approval decisions, exported derivatives and destinations. C2PA's ingredient model is useful here because it describes source assets as parents, components or inputs to a computational process; the same relationships should remain understandable even when the final JPEG carries no readable manifest.
The final handoff should contain the approved master, delivery derivatives, the source lock, provenance manifest and a contact for corrections or withdrawal. A screenshot of a green credential badge is not the rights record. The practical objective is that a producer can answer where the image came from, what changed, who approved it and where it ran without reopening a vendor conversation.
Choose the model on accepted cost, not render price
OpenAI publishes the same token rates for Flare and Sunburst: $8 per million image input tokens, $2 per million cached image input tokens, $30 per million image output tokens, $5 per million text input tokens and $1.25 per million cached text input tokens. The guide warns that equal rates do not mean equal cost per image because token use varies by model and quality. Responses API work can also add mainline-model token charges, and streamed partial images add output tokens.
Therefore log actual usage from representative requests at explicit sizes and quality settings. Compare Flare and Sunburst on cost per accepted result, not one generation. A slower, more precise model can be cheaper if it prevents three corrections; a faster model can win when the reference lock is simple and review is automated. Include operator time, approver time, rejected candidates, storage, delivery and recovery from a bad public asset.
Set a job ceiling and a stop condition before scaling. If first-pass acceptance, drift or review time misses the pilot target, change the reference package, mask strategy, model or production route. Do not respond by silently generating a larger pile. The useful promise of Images 2.5 is controlled iteration. The buyer's job is to make ‘controlled’ measurable enough that speed does not outrun approval.

Run one controlled edit before approving the workflow
Choose one real asset with meaningful constraints: a product and label, a consented performer, a retail environment or a campaign layout. Freeze the source, define the permitted delta and score ten candidates against the same lock. Take the strongest candidate through copy correction, format adaptation, compression, approval and delivery rather than stopping at the attractive preview.
Measure accepted assets per hour, human review minutes, correction turns, hidden-drift escapes, rights exceptions and total cost. Preserve the losing candidates long enough to diagnose recurring failure, then apply the agreed retention policy. If the final derivative passes, package the source lock and provenance record with it. If it fails, return to the last approved parent instead of repairing an already compromised version.
That is the production opportunity in stronger multi-turn editing: fewer rebuilds and more useful continuity between decisions. The competitive advantage is not a prompt that produces a convincing first image. It is a system that lets a team change one thing, prove that only one thing changed and deliver the result with its rights and history intact.
Build
Need a repeatable AI production workflow?
Mike designs the tools, review loops, and publishing systems that make it usable.
Launching a business of your own? Founder Launch OS connects the brand, offer, website and visual campaign in one guided Codex or Claude Code workspace.



