mrkeyoor.com_
Sat 03 Oct 05:40 UTC
AI7 min read

FLUX 3 Image Puts Bounding Boxes Inside the Prompt

Black Forest Labs gives FLUX 3 Image a 0-to-1000 layout grid, one API for generation and editing, and a steep price jump at 4K.

FLUX 3 Image reached 295 points on Hacker News after arriving with an unusually concrete developer interface: a 0-to-1000 layout grid embedded directly in the prompt. That grid turns a request such as "put the product near the bottom" into coordinates an application can save, revise, and send again. Black Forest Labs released the model on October 1 as an API for generating and editing images at up to 4K resolution, according to its release notes.

The more consequential part is the contract around the model. The same endpoint accepts a text prompt, as many as ten reference images, or a prompt containing element IDs and bounding boxes. Black Forest Labs' overview says the prompt itself determines whether the request creates a new image or edits an existing one. For developers building design tools or agent-driven media pipelines, FLUX 3 Image is asking to be treated as a programmable renderer rather than a chat box with an image attached.

The prompt is now a layout contract

Each FLUX 3 box uses four integers in the order top, left, bottom, right. The values run from 0 to 1000 regardless of the final image size, so [0, 0, 500, 500] always means the upper-left quarter. The bounding-box guide tells developers to give every object an ID, describe the whole scene, then append a JSON array containing the position and description of each element. There is no separate layout parameter.

A minimal request can carry a scene description followed by data in BFL's documented element-table format:

[
  {"id": "headline_1", "bbox": [80, 100, 240, 900], "desc": "A short blue headline on a cream background"},
  {"id": "product_1", "bbox": [300, 200, 900, 800], "desc": "A front-facing red desk lamp"}
]

That format connects IDs in the scene caption to rows in the array. Because the coordinates are normalized, a layout editor can draw boxes on a canvas without first converting everything to the requested output resolution. The application still has to preserve the aspect ratio for which the boxes were designed, since the grid stretches with the frame, as the bounding-box guide explains.

This makes the prompt partly machine-readable, even though BFL sends the boxes as text. An agent can draft a caption and element table, a person can move one box, and the model can render the revised layout. BFL demonstrates that sequence on its FLUX 3 Image page: an LLM plans the elements, the boxes remain editable, and FLUX generates each element inside its assigned region. The useful distinction is that a user can change geometry without rewriting the scene in increasingly fussy prose.

One endpoint handles generation and edits

The API surface is small. Developers send POST https://api.bfl.ai/v1/flux-3-image with an API key in the x-key header, and only prompt is required. Optional fields include images, aspect_ratio, resolution, safety_tolerance, and grounding. The API overview lists 15 fixed aspect ratios plus auto, with resolutions from a 768-pixel square preview through 1K, 1.5K, 2K, and 4K.

Adding one image turns the request into an edit. Adding two to ten lets the prompt combine references. The order matters because the prompt refers to "image 1," "image 2," and so on, while auto takes its aspect ratio from the first reference. BFL's editing documentation uses that scheme to put an object from one photo into another scene. For product software, that is simpler than maintaining different request shapes for generation, inpainting, and multi-reference composition. It also means input ordering becomes part of the application's state and should not be treated as an incidental array detail.

The response is asynchronous. Submission returns a polling_url. Clients poll until the status becomes Ready, then download the file from result.sample. BFL's text-to-image guide says that signed result URL expires after one hour and warns clients not to send the API key header when fetching it. The endpoint also rejects unknown fields with HTTP 422, including familiar parameters from other BFL APIs such as seed, width, and input_image. Existing FLUX clients will need an explicit adapter rather than a renamed model string.

Editing still needs a leakage test

Bounding boxes create a testable promise: change the boxed object and leave the rest alone. BFL says pixels outside the selected boxes usually remain the same, but its editing guide also says shadows, reflections, and nearby lighting can change. That caveat matters for catalog images and brand assets, where a small color shift outside the target can invalidate the result even when the requested object looks right.

A box is optional for a simple edit. When a developer names one element precisely, the model may locate it and add a box while expanding the instruction into a fuller prompt. The returned result.prompt exposes that expanded version, according to the same guide. Logging it gives teams something concrete to compare when two near-identical requests produce different edits, although BFL does not claim that the expanded prompt makes a request deterministic.

Small targets are another boundary. In BFL's tests, a newly requested element inside a box around 40 by 25 pixels often failed to appear. The docs advise using larger boxes and giving text its own region. That is more useful than a vague warning about complex prompts because an editor can enforce a minimum selectable area, flag tiny objects before spending a request, and record failure rates by box size. The bounding-box tutorial also specifies separate source and target boxes for moving or resizing an existing element.

Grounding changes the default request

FLUX 3 Image enables grounding by default. BFL says the model searches the web and images before generation to improve its handling of real things. Setting the field to false produces a faster request that relies only on the prompt. The official overview illustrates the difference with an Oktoberfest poster, comparing an invented result with one based on the current official design.

That default deserves attention in production. A team evaluating prompt consistency should record whether grounding was enabled, because a retrieved source can change while the stored prompt does not. Applications that handle factual visuals also need their own provenance rules. BFL exposes a boolean control, but its documentation does not describe a citation field for the web or image results used during grounding. Turning search on may improve a reference to a current object while leaving the application responsible for showing where that information came from.

The 4K price jump is steep

BFL charges per completed image class rather than per token. Its October 1 release notes list $0.041 for 768sq, $0.048 for roughly one-megapixel 1K output, $0.100 for roughly four-megapixel 2K output, and $0.607 for roughly 16-megapixel 4K output. The submission response includes the cost, which lets a client attach actual spend to each job instead of estimating it later.

The jump from 1K to 4K is about 12.6 times, while the approximate pixel count rises by about 16 times. A batch of 1,000 accepted 1K generations would cost $48 at the listed rate. The same count at 4K would cost $607. The pricing table gives no separate discount for using references, boxes, or edits. Teams can use the 768-square or 1K modes for layout trials and reserve 4K for the final pass, but they still need to measure whether a low-resolution approval reliably survives regeneration at the larger size. BFL does not promise that two separate calls will preserve every detail.

API access arrives before public weights

FLUX 3 Image is currently documented as a hosted API, with commercial weights available to companies that contact BFL for deployment at scale. The model page does not offer a public checkpoint download for this image product. That limits the immediate story for local inference developers, even though the API has more structured control than a plain prompt field.

BFL's broader FLUX 3 launch plan separates image synthesis through APIs and private weights from a planned FLUX 3 Dev multimodal backbone with open-weight access. The company said in July that those capabilities would arrive over the following weeks and months, but it did not give a date for FLUX 3 Dev. Until public weights, licensing terms, and hardware requirements appear, developers should treat the hosted image endpoint and the promised open model as separate products.

The next evidence should come from repeated, boring production tests: whether boxes hold their positions across aspect ratios, how often an edit changes pixels outside its target, what grounding adds to latency, and how many paid attempts survive review. The 0-to-1000 grid gives teams a stable object to measure. If it holds up under messy references and real brand layouts, FLUX 3 Image can sit behind an editor or an agent without asking users to negotiate every pixel in prose. If it drifts, the coordinates will at least make the failure visible.

We reviewed this

  1. editor — our honest review
  2. editor — our honest review
  3. bottom — our honest review

Sources

  1. FLUX 3 Image
  2. FLUX 3 Image documentation
  3. Bounding boxes with FLUX 3 Image
  4. FLUX 3 Image editing
  5. FLUX 3 text-to-image guide
  6. FLUX 3 Image release notes
  7. Black Forest Labs API pricing
  8. FLUX 3 launch plan