Franklin AI Explainer

Ideogram 4.5 targets precise image edits, with masks and reference controls

Key Takeaways

  • Ideogram 4.5 focuses on preserving image details through repeated edits.
  • Its official API references describe copied unchanged pixels, masks, reference images, and asynchronous res Ideogram 4.5 focuses on preserving image details through repeated edits.
  • Its official API references describe copied unchanged pixels, masks, reference images, and asynchronous results, while also documenting size limits that qualify the model page's high-resolution claims.
  • Changing a product's color should not require accepting a different product shape in the next image.
  • That is the problem behind Ideogram 4.5's focus on precise editing.

Changing a product's color should not require accepting a different product shape in the next image. That is the problem behind Ideogram 4.5's focus on precise editing. The company says Ideogram 4.5 reduces the pixel shifts, color changes, and texture artifacts that can accumulate across multiple edits.

Its model page illustrates color and lighting changes, text modification, product photography, interior design, and other editing tasks. The accompanying API documentation makes the workflow more concrete: there is a dedicated precise-edit endpoint, along with a generation endpoint that can also accept source images. Their controls and size rules are not identical.

Limit the change with a mask

The precise-edit reference describes uploading one image with an instruction, optional reference images, and an optional mask. Ideogram says pixels the edit does not meaningfully change are copied exactly from the original. That is a specific preservation mechanism, rather than simply asking the model to recreate a similar composition.

A mask limits the edit to part of the image. Black marks the area to change; white marks the area to keep. Intermediate values are rounded to the nearer of those two values, so the documented mask is not a graded transparency control. It must match the input image's width and height and contain both black and white areas.

Up to four reference images can guide an unmasked edit. They are not edited themselves. Adding a mask reduces the reference allowance to three because the mask occupies one reference slot. The input image and each reference can be JPEG, PNG, or WEBP, with a maximum file size of 25MB each.

Read the high-resolution promise alongside the limits

The model page shows zoom editing on a 4,016-by-6,016-pixel source image, described as 24.2 megapixels. Its explanation is to edit a crop while preserving its edges, then stitch that crop back into the original. That is different from promising that every API request will process an unrestricted full-resolution image.

The precise-edit reference says results match the input's dimensions, but also says images too large for the model are scaled down proportionally and returned as rendered. It rejects aspect ratios outside 1:6 to 6:1. The captured documentation does not give a numerical threshold for that downscaling, so there is no basis for declaring a universal maximum resolution here.

The generation endpoint's size controls add another distinction. With source images, its default auto setting selects a supported 2K size; source requests the first image's dimensions, subject to downscaling if necessary. An exact requested size must use multiples of 32, have sides of at least 256 pixels, stay within 2048-by-2048 total pixels, and satisfy the aspect-ratio limit. Choosing an exact size reshapes the source to it.

Prompt preparation affects the request

Both references accept natural-language or structured prompts. For precise editing, natural language is automatically converted into a structured prompt, while valid structured JSON is used as supplied.

Generation exposes an additional magic_prompt control. Its auto and on modes rewrite and expand natural-language prompts before generation. The off mode keeps the wording while converting it into a structured prompt. When source images are supplied, every mode converts the instruction into a structured edit prompt.

Precise editing defaults to medium quality. Its four choices run from very_low through high; the documentation says higher quality takes longer and costs more. Generation defaults to medium with source images and high without them, and its very_low setting requires source images. These are documented trade-offs, not measured timings or a published price comparison.

Quote a request before running it

Both endpoints support a dry_run option that validates and prices the actual request without generating, storing, or billing an image. No safety review is performed during that quote. This lets a developer inspect the charge for selected options without treating a successful quote as approval of the eventual output.

Requests normally wait for completed images. Setting async, or supplying a webhook URL, instead returns a generation identifier for polling or later delivery. The references describe signed webhook results and reject private or loopback webhook destinations.

Ideogram also presents side-by-side multi-turn comparisons against other image models. Those are publisher-supplied demonstrations, not independent proof that competitors consistently fail or that a particular edit will preserve every required detail. The practical distinction to inspect is whether the edited region changes as intended while the surrounding image stays usable.

Our read

Franklin AI Take

Preserving the surrounding image is valuable when a small revision would otherwise undo approved work. The documented mask and copied-pixel behavior make that goal more concrete than a comparison gallery alone. We would check the original and edited result side by side, especially near the mask boundary, and confirm dimensions before using the result in a larger design. The API's downscaling caveat matters even when the model page demonstrates edits on a high-resolution crop.