Studio
Queue work for the fleet — text to image, editing, interleaved generation, and visual question answering.
Queue work for the fleet — text to image, editing, interleaved generation, and visual question answering.
Native 2048² buckets, optional reasoning pass.
Four-panel comic with Chinese dialogue
Dense in-image text is the model's standout — all four speech bubbles render correctly.
2048×2048 · 50 steps
Instruction editing with single or multi-image reference.
Recolour without disturbing the text
Editing regenerates the whole frame, so naming what must not change is the whole skill.
Drop an image here, or
Multiple images act as references — say what each one contributes.
Add at least one image
Alternating text and images in one pass.
Illustrated step-by-step tutorial
Returns a planning trace plus one composed instructional image — the MLX backend does not emit a multi-image sequence.
Emits several images — expect this to take longer than a single render.
Image-grounded question answering.
Read a comic and retell it
Understanding only — no image is rendered, so this comes back in well under a minute.
Drop an image here, or
The question is answered against these.
Understanding only — no image is rendered, so this returns fast.