Practical guide
How to Combine Two Images with AI
A practical two-reference workflow for choosing inputs, writing a clear prompt, selecting an available model, troubleshooting failures, and reviewing the result.
Provider capability statements link to the official sources listed below. Demonstrations and untested comparisons are labeled; this article does not imply a live multi-model benchmark.
Combining two images with AI is different from placing one file on top of another. A conventional editor can mask, crop, and blend existing pixels. An image-generation model can interpret the relationship between two references and create a new composition.
If you came here to combine 2 images, the practical goal is simple: give Image 1 and Image 2 separate jobs, then review the new output before publishing it.
This guide follows the same sequence as the AI Image Combiner workspace: choose two useful inputs, assign each one a role, review the suggested direction, select an Available model and output settings, generate, and inspect the result.
1. Give each reference one clear job
Start by deciding what Image 1 and Image 2 contribute. Strong pairs usually have distinct roles:
- Subject + environment: an adult portrait plus a landscape.
- Product + scene: a clean product image plus a compatible lifestyle setting.
- Pet + concept: a pet photo plus an original costume or character direction.
- Identity + visual direction: a portrait plus a palette, lighting, wardrobe, or illustration reference.
- Character + composition: an original character plus a background that controls framing and atmosphere.
Avoid asking both images to control everything. If each contains several competing subjects, backgrounds, text blocks, and lighting directions, the model must guess what to retain.
2. Use technically clean inputs
The workspace accepts JPG, PNG, and WebP files up to 10MB each. A useful source has enough resolution to show the features you want preserved, but it does not need to be an enormous camera original.
Check these points before upload:
- The important subject is in focus and not heavily compressed.
- Hands, paws, product edges, labels, and other required details are visible.
- The two camera angles can plausibly fit one composition.
- The lighting directions do not contradict each other unnecessarily.
- You have permission to use both images.
For identity-sensitive work, use an adult subject and avoid deceptive or harmful requests.
3. Write a relationship-first prompt
A useful prompt says more than "merge these images." Name the role of each input and describe the final composition:
Keep the adult subject from Image 1 recognizable and place them naturally in the mountain environment from Image 2. Match the golden-hour lighting, preserve realistic anatomy, and create one coherent cinematic portrait.
This structure makes four decisions explicit: what must be preserved, what may change, what the second reference contributes, and what the final image should look like.
4. Select a model that is actually Available
The catalog includes three upstream model families routed through the configured DMX, KIE, or Cloudflare REST transport:
- GPT Image 2, provided by OpenAI, whose documentation covers image generation and editing with image input and output.
- Nano Banana 2, Google's Gemini 3.1 Flash Image, whose documentation covers native image generation, editing, and multiple-reference processing.
- Gemini 3 Pro Image, provided by Google, with Standard, HD, and Ultra HD two-reference requests mapped to native 1K, 2K, and 4K output settings.
Only a model marked Available in the live workspace can be submitted. A visible but unavailable card is not a working integration.
Review the current OpenAI image generation guide and Google Gemini image generation guide before implementing against either provider. Model identifiers, limits, and lifecycle states can change after this article's source-check date.
5. Choose ratio and quality before submission
Choose the canvas for the intended placement:
- 1:1 for profile images and square product cards.
- 3:2 for flexible landscape compositions.
- 2:3 for portraits and posters.
- 16:9 for wide placements.
- 9:16 for vertical mobile placements.
The workspace displays the exact credit quote for the current model and quality before submission. A higher-quality request is a new generation, not an unlock for an existing result.
6. Review the output before download
Inspect the details that matter to the intended use:
- Is the adult subject, pet, product, or character still recognizable?
- Are hands, paws, edges, labels, and small objects coherent?
- Do scale, perspective, light direction, and shadows agree?
- Is any visible text accurate?
- Did the output invent another subject or remove a required object?
- Does it contain unintended trademarks or private information?
If the result misses the brief, make one correction at a time.
Troubleshooting common failures
| Symptom | Likely cause | Better next step |
|---|---|---|
| The subject is replaced | The prompt did not state which identity or object to preserve. | Name Image 1 as the identity anchor and list two or three required traits. |
| The scene looks pasted together | Camera angle, scale, or lighting conflicts between inputs. | Use a more compatible pair or explicitly request one perspective and light direction. |
| Product geometry changes | Style language overpowered structural constraints. | Put shape, material, color, and label preservation before mood words. |
| Extra limbs or duplicate objects appear | The composition contains competing subjects or an ambiguous count. | Crop distractions and state the exact number of subjects or objects. |
| Visible text is wrong | Generated text is not guaranteed to reproduce a label accurately. | Request a clean label area, then verify or add critical copy in a conventional editor. |
Reusable prompt template
Use this starting point:
Preserve [subject or object] from Image 1, especially [required traits]. Use [scene, lighting, palette, wardrobe, or composition] from Image 2. Create [final image type] with [ratio or framing]. Keep [important constraints], natural perspective, and consistent lighting.
The prompt does not need to be long. It needs to resolve the relationship between the two references. When your inputs and instructions are ready, open the AI Image Combiner workspace.
Official sources · Checked 2026-08-03
Provider documentation can change. Recheck these pages before making an implementation or purchasing decision.
- GPT Image 2 model documentation — OpenAI
- Image generation guide — OpenAI
- Gemini image generation guide — Google
- Gemini 3.1 Flash Image model documentation — Google
- Gemini 3 Pro Image model documentation — Google
Ready to combine two references?
Open the shared workspace, review the suggested direction, and choose quality before generating.
Open workspace