Click to Play Episode
Editing replaced generation as the core task. How GPT Image 2.5, Google's Nano Banana line, Midjourney V8.2 and Flux 2 differ, what open weights and LoRAs buy you, ControlNet vs instruction editing, and how licensing and C2PA provenance work now.
Companion show. This episode is the overview. For weekly, hands-on coverage of the full image and video pipeline, from a first usable clip to scenes that cut together, listen to AI Video Generation.
First of three episodes on AI media generation (this one is images and editing; then video, then the pipeline). A decision guide rather than a leaderboard: where GPT Image, Nano Banana, Midjourney and Flux each fit in late 2026, why instruction-driven editing replaced generation as the core task, what open weights buy you, and how licensing and provenance work.
The 2025 "artist vs collaborator" split is over and the collaborators won. Every frontier image model now sits behind a language model that reads the prompt with world knowledge and accepts images as input, so the unit of work became "here is an image, change this one thing and keep everything else." Text rendering, precise instruction following and identity-preserving reference images all landed at once for one reason: the image model became, or was paired with, a multimodal language model. Generation from scratch is now the special case where the input image is empty.
The lineage runs gpt-image-1 (2025), gpt-image-2, then ChatGPT Images 2.5 in September 2026, with a precision variant (Sunburst) and a fast default (Flare); the same models power ChatGPT and the Images API. As of this recording it holds the top slots on both the LMArena text-to-image and image-edit boards and on Artificial Analysis. For editing it does mask inpainting, up to four reference images, and multi-turn editing via the Responses API. Its weaknesses are latency at high quality and occasional text and consistency slips. Billing is per token (text in, image in, image out) with a cached-input discount; see the gpt-image-2.5-sunburst model page.
Nano Banana was gemini-2.5-flash-image (August 2025, now legacy). It was succeeded by Nano Banana Pro (gemini-3-pro-image) and Nano Banana 2 (gemini-3.1-flash-image, February 2026), plus a Lite tier. Per the Gemini image generation docs, Nano Banana 2 accepts up to fourteen reference images (objects plus characters), outputs up to 4K, does multi-turn sequential editing, and can ground generation in Google Search; the Pro model adds style references and identity preservation across up to five subjects. Every output carries a SynthID watermark with no opt-out. Imagen appears to be superseded for new work, though no formal retirement notice was found.
Midjourney moved from V7 (2025) to V8 alpha in March 2026 and V8.2 as the default in July 2026 (version history). V8.2's Edit Model replaces Omni Reference, Character Reference, Retexture and the Editor with one instruction-driven model taking up to four references, the same convergence OpenAI and Google made. Aesthetics and draft-mode ideation remain its strengths. The hard limits: still no official API and terms that bar automation; generations public by default below the Stealth tier; and it is the defendant in Disney Enterprises v. Midjourney (filed June 2025, joined by a separate Warner Bros. Discovery suit), currently in discovery with Midjourney demanding the studios' own AI records.
Black Forest Labs, founded by the original Stable Diffusion authors, shipped FLUX.1 Kontext (in-context editing without masks) in 2025 and FLUX.2 in November 2025: a Mistral vision-language model paired with a rectified-flow transformer, up to ten references, 4MP editing. Tiers: Pro and Flex (API only); FLUX.2 dev (32B, open weights, non-commercial license); and FLUX.2 klein (January 2026), where the 4B model is Apache 2.0 and the 9B is non-commercial. The open-weights editing sub-board puts FLUX.2 and HunyuanImage roughly 130-150 Elo behind the closed frontier. Open weights earn their place through fine-tuning, on-prem privacy and composability with ControlNets and node graphs rather than raw quality.
Change this thing is instruction editing, now standard everywhere; masked inpainting and outpainting (Photoshop Generative Fill, GPT Image masks, FLUX.1 Fill) remain the hard constraint when the instruction is not enough. Keep this subject is reference conditioning, whose open-world mechanism is IP-Adapter (decoupled image cross-attention on a frozen base) and whose closed equivalents are Google's fourteen references and Midjourney's edit-model references; single references drift on fine detail. When drift is unacceptable, train a LoRA: roughly ten to twenty images and about a thousand steps on Flux (fal guide, FLUX.2 LoRA guide, Replicate trainer), open weights only. Keep this structure is ControlNet depth/edge/pose conditioning, still required for geometry fidelity per Autodesk's testing, with first-party support in Qwen-Image-Edit and FLUX.2 ComfyUI nodes. Multi-turn is for exploration; for repeatability, reproduce the winning edit as one instruction from the original.
Ownership is a human-authorship question: the US Copyright Office's Copyrightability report requires human authorship, treats prompts as unprotectable instructions, and protects AI-assisted work to the extent of the human contribution; the Supreme Court declined Thaler v. Perlmutter in March 2026, leaving the "how much human is enough" line undrawn. Commercial use is a vendor question: OpenAI (you own outputs), Midjourney (paid plans, no indemnity), Flux (tiered), Firefly (indemnified). Training-data fair use is unresolved in the US; Andersen v. Stability AI is the image bellwether, while the UK High Court largely rejected Getty's claims against Stability in November 2025. Provenance now has two layers: C2PA 2.3 manifests plus pixel watermarks; OpenAI now embeds both C2PA and SynthID (API guide), as Google already did. Manifests are detailed but stripped on re-save, screenshot and most platform uploads; SynthID survives those but carries little information. Labeling is now law: EU AI Act Article 50 and California SB 942 both enforceable from August 2026, China's rules from September 2025.
A note on how this one was made. I generated this episode with Gnothi, my tool that takes a topic, researches it, writes the chapters, and reads them aloud. This episode is the overview. The dedicated Gnothi show on AI video generation goes further, from a first usable clip to scenes that cut together. O C devel dot com slash video. That's O C D E V E L dot com, slash video.
This is the first of three episodes on AI media generation. This one is images, and specifically image editing, because that's what the job turned into. The next episode is AI Video Generation in twenty twenty six, and the one after that is the AI Media Pipeline, which covers voice, music, ComfyUI, APIs and finishing. This is a decision guide, not a leaderboard. Which image models matter, why "edit this" replaced "generate this" as the core task, what open weights actually buy you, and how the licensing works.
Start with what changed. If you last looked at this space in twenty twenty four or even mid twenty twenty five, the mental model you have is wrong in a specific way.
The field used to be described as a split between artists and collaborators. The artists, Midjourney above all, gave you gorgeous, opinionated images and ignored half your instructions. The collaborators, the image models bolted onto ChatGPT and Gemini, followed instructions, rendered text, and produced images that were competent and a little dull. You picked a camp based on whether you wanted beauty or obedience.
That split is gone, and the collaborators won. Not because Midjourney got worse, but because every serious model now sits behind a language model that reads your prompt the way a person would, and every serious model now takes images as input, not just text. The unit of work stopped being "generate an image from a description" and became "here is an image, or four images, change this one thing and keep everything else." Instruction-following editing is the core task. Generation from scratch is the degenerate case where the input image happens to be empty.
Three specific things got solved along the way. First, text rendering. For years the giveaway of an AI image was the mangled lettering on a sign. That's over for the frontier models. They render paragraphs, multilingual text, logos with the right letters in the right order. Second, instruction following. You can say "move the lamp to the left side of the desk and make it brass" and get exactly that, with the rest of the room untouched. Third, the one that changed workflows most, reference images. You hand the model a photo of a person, a product, a style, and it keeps that identity across generations. Character consistency used to require training your own adapter. Now it's a parameter.
All three fell at once for the same reason. The image model became a multimodal language model, or a diffusion transformer paired with one, so the thing reading your prompt understands language and images together, and it understands the world. Ask for a diagram of how a bicycle derailleur works and it knows what a derailleur is. That world knowledge is what turned these from paint programs into something closer to a junior designer who has read everything.
So that's the shift. Now the players. I'll spend real time on four of them, because those four set the shape of the market, and then run through the second tier faster.
First, OpenAI. The lineage is easy to follow because they stopped calling it GPT four o image generation and gave it a name. It was GPT Image one in twenty twenty five, then GPT Image two, and as of this recording the current family is what OpenAI calls ChatGPT Images two point five, which shipped in early September of twenty twenty six. There are two variants. One is tuned for maximum quality and editing precision, the other is the fast default that OpenAI says beats the previous generation at about half the latency. The same models power ChatGPT and the developer API, so what you get in the chat window is what you get programmatically.
Why it matters. As of this recording this family is at the top of both major human-preference leaderboards, for text to image and for image editing, and it isn't close. The two crowd-sourced arenas disagree on which of the two variants is number one, which tells you the top is noisy, but they agree OpenAI holds the top three slots. If you want the best image from a plain English prompt and you don't care about running it yourself, this is the default answer right now.
What it does for editing. Mask-based inpainting, where you paint the region to change. Up to four reference images the model treats as context, so you can say "put the product from image one on the table from image two in the lighting of image three." And multi-turn editing through the Responses API, meaning you carry on a conversation about an image and each turn edits the previous result, which is how ChatGPT itself does it. The docs admit that fine text and consistency still slip sometimes, and that complex prompts at the highest quality setting can take a couple of minutes. It's slow. Everyone who builds on it complains about that first.
For developers, the API bills in tokens, text in, image in, image out, with a cache discount on repeated inputs. I won't say prices on air because they'll be wrong by the time you hear this. Structurally, a high-quality large image costs real money per call, which is why the fast variant exists and why people route the cheap variant for drafts and the expensive one for finals.
Second, Google. This is the one whose naming will confuse you, so here's the map. In twenty twenty five Google shipped an image model inside Gemini that the internet knew as Nano Banana before Google admitted it was theirs. Its real identifier was Gemini two point five Flash Image. It became a phenomenon for editing, specifically for keeping a person's face identical across edits, and it made "nano banana" a verb for a while. Since then the lineage moved to Nano Banana Pro, which is the Gemini three Pro image model, and then Nano Banana two, which is the Gemini three point one Flash image model, plus a Lite tier for cheap high-volume work. The original Nano Banana is now marked legacy in Google's docs. Imagen, the older Google image brand, is no longer where the new work lands, though I couldn't find a formal retirement notice, so I'll just say the Gemini-branded models are the ones to use.
Why it matters. Google's differentiator is references and resolution. Nano Banana two takes up to fourteen reference images in one request, split between object references and character references, and the Pro model adds style references on top. It outputs up to four K. And it does something nobody else does at this level: it grounds generation in Google Search. Ask for an infographic about a current event and it will look the event up before drawing it. For editing, it does multi-turn sequential edits, localized changes, and camera and lighting transformations on an existing photo, and the identity preservation across up to five subjects is why people use it for anything involving real faces.
One more Google fact worth a full sentence. Every image these models produce carries SynthID, Google's invisible watermark, with no opt-out. I'll come back to that in the provenance section, but if you need to know whether your pipeline will emit a detectable watermark, with Google the answer is always yes.
Third, Midjourney. Midjourney stayed a product for artists rather than a platform for developers, and that choice defines both what it's good at and what it can't do for you.
Versions. Version seven came in spring twenty twenty five. Version eight went to alpha in March twenty twenty six, and as of this recording the default is version eight point two, from late July. The headline of eight point two is that Midjourney collapsed all its scattered control features into one thing it calls the Edit Model. Before, you had omni reference for putting a subject into a scene, character reference for faces, retexture for restyling, and a separate editor for inpainting and outpainting, each with its own flag. Now one model takes up to four reference images plus a plain-language instruction and does all of it. That's the convergence I described at the top. Midjourney watched Google and OpenAI ship instruction editing and rebuilt itself around the same shape, but kept its own aesthetic.
And the aesthetic is still the reason to use it. If you want an image that looks composed, lit, and art-directed with no effort, Midjourney's default output is still the one that makes people stop scrolling. Draft mode gives you a grid of cheap low-resolution candidates you promote and refine, which is the right workflow for ideation. Personalization learns your taste from your ratings.
Now the limits, and they're hard limits. There is still no API. None. Midjourney's terms of service bar automated access and reselling, and every third-party "Midjourney API" you see advertised is someone scripting the web app and risking a ban. If your job is a pipeline, Midjourney isn't in it. That has been true for two years. Second, generations are public by default on the standard plans. Stealth is a higher tier. If you're working on unreleased products, know that.
Third, Midjourney is the defendant in the industry's biggest copyright suit. Disney, Universal and DreamWorks sued in June twenty twenty five over outputs of characters like Darth Vader and the Minions, and Warner Brothers followed with a separate suit. The case is in discovery as of this recording, with a trial plausibly in twenty twenty seven. Midjourney is arguing fair use and, interestingly, demanding the studios disclose their own internal AI use. I'm not going to predict the outcome. I am going to say that a client-facing agency should think about it before standardizing on the tool.
Fourth, Flux, from Black Forest Labs. This is the open-weight story, and it's the one where I want to be precise, because the license tiers are easy to misread.
Black Forest Labs is a German lab founded by the researchers who wrote the original Stable Diffusion papers before they left Stability. Flux one, in twenty twenty four, became the model the open community actually moved to when Stability faltered. Flux one Kontext, in mid twenty twenty five, was the open world's answer to Nano Banana. Kontext means in-context editing: give it an image and an instruction, no mask, and it keeps character and style consistent across iterations. Kontext is what runs under the hood of several Adobe and Meta integrations.
Flux two shipped in late November twenty twenty five and is the current generation. The architecture pairs a Mistral vision-language model with a rectified-flow transformer, which is the same "put a language model in front of the image model" move everyone made. It takes up to ten reference images and edits at four megapixels. The lab raised a large round in December twenty twenty five, so it isn't going anywhere.
The tiers. Flux two comes in four. Pro and Flex are API-only. Dev is the thirty-two billion parameter open-weight model, and it's under a non-commercial license. You can download it, run it, build on it, but the moment you generate revenue from it you need Black Forest Labs' paid self-hosted commercial license or you need to use their API.
Then there's Klein, which shipped in January twenty twenty six. Klein comes in a four billion and a nine billion parameter size, four-step distilled so it runs in under a second on a consumer GPU, and the four billion fits in around thirteen gigabytes of video memory. The four billion Klein is Apache two, meaning free for commercial use. The nine billion is non-commercial like Dev. So the pattern from Flux one, where Schnell was free and Dev was not, carried forward, just with the free tier now being the smallest model. If your startup's plan is "we'll self-host Flux and never pay anyone," what that plan actually means is "we'll self-host the four billion Klein," and you should test whether its quality is enough for your use before you architect around it.
Why open weights matter at all, if the closed models lead the leaderboards by a hundred and thirty or so Elo points. Three reasons. You can fine-tune them, which is the LoRA story I'll get to. You can run them on your own hardware with no rate limits, no content policy that refuses your medical illustration, and no images leaving your network. And you can compose them with the whole toolkit of ControlNets, adapters, and node graphs that only exists for open models. For a lot of production work, "eighty-five percent of frontier quality with total control" beats "frontier quality through someone else's API."
So those are the four. Now the second tier, quickly, because each one does a job the big four don't.
Stable Diffusion. Stability AI is still alive. It raised money again in late summer twenty twenty six with entertainment industry backers, and Stable Diffusion three point five is still its flagship image family. There's no Stable Diffusion four as of this recording. The community license is free under a revenue threshold of about a million dollars a year, then you need an enterprise deal. Stable Diffusion stopped being the frontier open model, but SDXL from twenty twenty three still has the deepest collection of fine-tunes and LoRAs in existence, thousands of LoRAs on Civitai, and a lot of style work still happens there because the adapters exist.
Qwen Image, from Alibaba. Twenty billion parameters, Apache two, and it's the model to reach for when text in the image is the point, especially Chinese text, which it renders better than anyone. Qwen Image Edit, with several point releases through late twenty twenty five, does multi-image editing and identity preservation, all open. The catch is that the newer Qwen Image two line is closed and API-only. So the open Qwen you can download is the twenty twenty five generation.
Seedream, from ByteDance. Seedream four unified generation and editing in one model, does four K, and takes up to six references, with four point five raising that to ten and dropping the four K surcharge. It's available through fal, through ByteDance's cloud, and inside CapCut and Dreamina. It had a real moment in twenty twenty five as the fastest high-resolution model, and as of this recording it's no longer in the top five of either arena. Good model, particularly on cost, no longer the headline.
Ideogram. For years Ideogram was the text-rendering specialist, and its version three added character reference from a single image. In June twenty twenty six Ideogram four did something surprising: it went open-weight, nine billion parameters, with structured JSON prompting where you place elements by bounding box. Read the license, though. The code is Apache two but the weights are non-commercial, so it's Flux Dev's situation, not Flux Klein's.
Recraft. The design-system model. Recraft version four and four point one, from early and mid twenty twenty six, output native vector, meaning real SVG with editable paths you can open in Figma or Illustrator and commit to a repository. It also does exact brand-palette control and clean white-background product shots. If you're generating icons, logos, or anything a designer will edit afterward, this is the one that produces the right file type.
Adobe Firefly. Firefly's pitch hasn't changed. It's trained on licensed Adobe Stock and public domain content, Adobe doesn't train on your work, and paid Creative Cloud plans carry intellectual property indemnification. That last part is what matters to an enterprise, and nobody else offers it. What changed is that Firefly stopped being only Adobe's model. The Firefly app and Photoshop now let you pick partner models, GPT Image, Nano Banana Pro, Flux, Ideogram, inside Adobe's tools, with Adobe's own Image Model five as the default. Content Credentials get attached automatically when the whole image is Firefly-generated, and partner-model output is labeled differently. Firefly became the enterprise front door to everyone else's models.
Leonardo, owned by Canva since twenty twenty four, runs its own Phoenix model alongside Flux and sells a developer API, and it's the practical pick if your organization already lives in Canva. Krea is the real-time one. Its own model streams the image onto the canvas as you type, and the business is really aggregation, sixty-plus models behind one interface. Krea is the one I'd point people at when they want to try everything without ten accounts.
That's the field. Now I want to slow down and talk about control, because control is what separates people who get usable output from people who get lucky.
I think about control as three different problems that happen to use similar words. Keep this structure. Change this thing. Keep this subject. Each has its own tool, and only one of the three got commoditized.
Change this thing is instruction editing, and it's the one that got commoditized. Every headline model does it now. You provide an image and a sentence, the model changes what you asked and preserves the rest. This replaced masked inpainting as the default way to edit. Inpainting, where you paint a mask over a region and regenerate only that region, still exists in every product, Photoshop's generative fill, GPT Image's mask parameter, Flux Fill in the open world, and outpainting is the same operation aimed outside the frame. But you reach for the mask now only when the instruction isn't precise enough, or when you need to guarantee that pixels outside the mask don't move at all. Instruction editing is a request. A mask is a constraint.
Keep this subject is reference conditioning, and it's where the most progress happened. The mechanism, in the open world, is a twenty twenty three paper called IP-Adapter. It adds a small adapter that feeds image features into the model through a separate cross-attention path from the text, so the base model stays frozen and the image reference composes with everything else. The closed models do a more integrated version of the same idea, which is what Google's fourteen references and Midjourney's edit model references are. The practical rule is that a single reference gets you the gist of a subject, the face shape, the outfit, the vibe, but fine details like freckles, a specific logo, or exact jewelry will drift. Midjourney's own docs say that outright. When drift is unacceptable you go to the next step, which is training.
That step is LoRA, low-rank adaptation. It was a large language model paper from twenty twenty one, and the image community adopted it as the way to teach a model a specific character, product, or style without retraining the whole thing. You freeze the model and train small low-rank matrices that get added to the attention weights. The output is a file of a few dozen megabytes you load on top of the base model. For Flux one, a character LoRA takes roughly ten to twenty images, a thousand or so steps, and somewhere between five minutes on a hosted trainer like fal and an hour on rented hardware. Flux two guidance widens the dataset range considerably, and these numbers are vendor blog numbers, not benchmarks, so treat them as order of magnitude.
The craft is in the dataset. Vary the setting, expression and background so the model learns the person and not the room. Keep the things you want fixed, like a haircut, constant. Don't let hands cover the face. A LoRA is the only way to get a product rendered with pixel-exact fidelity across hundreds of images, and it only works on open-weight models, which is the single strongest argument for open weights in a business setting.
Keep this structure is ControlNet, and it didn't get commoditized either. ControlNet is the twenty twenty three paper that lets you condition generation on a spatial map, an edge detection of a sketch, a depth map, a stick-figure pose, and the output follows that geometry exactly. Instruction editing can't do this. If you tell a model "same room but from three feet to the left" you get an approximation. If you give it the depth map of the room from three feet to the left you get the room. Qwen Image Edit ships depth, edge and keypoint conditioning natively. Flux two has first-party ControlNet nodes in ComfyUI. Autodesk tested whether next-generation models still need ControlNet for architecture and product renders and concluded yes, for geometry fidelity, you still need it.
So the practical stack is instruction editing for the everyday change, references or a LoRA for identity, ControlNet for geometry, and a mask when you need a hard guarantee. Most jobs need one of these. Serious production work needs two or three composed, which is why the node-graph tools exist, and that's the episode after next.
One more control concept, because it comes up constantly. Multi-turn. Both OpenAI and Google let you carry a conversation where each turn edits the last result. This is wonderful for exploration and terrible for reproducibility. Each edit is a new sample, and small drift compounds. My practice is to explore in multi-turn, then take the winning image and reproduce the final edit as a single instruction from the original, so I have a prompt I can rerun.
So that's control. Now licensing and provenance, which programmers skip and then regret.
Who owns the output. The US Copyright Office's position, from its twenty twenty five report on copyrightability, is that human authorship is required, prompts alone are instructions and aren't protectable no matter how elaborate, and AI-assisted work is protectable to the extent of the human contribution. The Supreme Court declined in March twenty twenty six to hear Thaler versus Perlmutter, which leaves that standing. Thaler settles almost nothing about normal practice, because Thaler disclaimed all human input. Where the line is for a person who prompted, edited, composed and retouched is still undrawn. Practically, a purely generated image is probably not something you can copyright, and an image you meaningfully edited and composed probably contains protectable human work. Plan for that. If a brand asset needs to be defensible, a human needs to have done real work on it, and you should keep the record.
Who can use the output commercially. That's a vendor question, not a law question. OpenAI, you own your outputs subject to the usage policies. Midjourney, paid subscribers own their assets and can use them commercially, free users get non-commercial terms, and there's no indemnity. Flux, Dev and the nine billion Klein are non-commercial, the four billion Klein and the API are commercial. Adobe Firefly, paid plans carry IP indemnification, which is the enterprise reason it exists. Read the actual terms the week you sign, because they move.
Then the training-data question, which is the one that could reshape the field. No US court has yet ruled on whether training an image model on unlicensed images is fair use. The bellwether is Andersen versus Stability AI, the artists' class action, with trial pushed to spring twenty twenty seven as of this recording. In the UK, Getty lost most of its case against Stability in November twenty twenty five, because the training happened outside the UK and the court held that importing a model is not importing an infringing copy, with Getty winning only narrow trademark points over watermarks that surfaced in early outputs. And there's the Disney and Universal suit against Midjourney I mentioned. None of this changes what you should do this week, but it's why Firefly's indemnity and the "trained on licensed data" pitch have customers.
Provenance is the more actionable half. The standard is C two P A, Content Credentials, a signed manifest embedded in the file that records what made the image and what was done to it. Adobe, Microsoft and Google were founding members. What changed in twenty twenty six is that OpenAI joined fully. In May twenty twenty six OpenAI announced that images from ChatGPT and the API carry both a C two P A manifest and Google's SynthID invisible watermark, with a public verification tool. Google already did both.
So the two biggest generators now do two layers, and the layers matter for different reasons. The manifest is detailed and durable while the file stays intact, and it dies the moment someone screenshots, re-saves, or uploads to a platform that strips metadata, which most do. SynthID is the opposite. It lives in the pixels, survives screenshots and recompression, carries almost no information, and only the vendor's detector can read it. An attacker now has to defeat both. A casual re-share defeats the first automatically.
On the display side, LinkedIn shows a credentials icon, Instagram shows them read-only, Google Search's about-this-image and YouTube can surface them. The caveat is that display isn't preservation. Credentials are generally not preserved through upload, so the second hop loses them. Absence of a credential proves nothing. And a credential attests provenance, not truth. A signed manifest on a fake tells you honestly that it's a fake, which is the entire point.
The reason this stopped being optional is regulation. The EU AI Act's transparency obligations became enforceable in August twenty twenty six, requiring machine-readable marking of synthetic content. California's AI Transparency Act was aligned to the same date and requires large providers to embed watermarks and publish a detection tool. China's labeling rules took effect in September twenty twenty five and require both visible labels and embedded metadata. If you ship generated images to users, you now have a compliance surface, and the practical answer is to keep the manifests your provider gives you and not strip them in your own pipeline.
So that's the legal and provenance layer. Let me close with the decision guide, by job.
Marketing asset, meaning a social image, a hero banner, an ad variant. Use GPT Image or Nano Banana. You want instruction following, correct text, and fast iteration, and you don't care about running it yourself. If the brand has a real face or product in it, use Nano Banana's references first, and if the drift is unacceptable, train a Flux LoRA.
Concept art and mood boards. Midjourney, still. Draft mode for volume, then the edit model to push the keeper. Accept that it's a manual tool and keep it out of the pipeline. If the concept art is for a client with a legal department, ask them about the Disney suit before you deliver.
Product photography. This is a consistency and geometry problem. A Flux LoRA of the product plus ControlNet depth for the placement, or Nano Banana with the product as an object reference for the quick version. Recraft if the deliverable is a vector. Firefly if the deliverable needs indemnification.
Developer pipeline, meaning you're calling this from code at volume. Three questions. Do images leave your network? If not, open weights, meaning Flux Klein four billion for commercial or Flux two Dev with a license. Do you need fine-tuning? If so, open weights again. Otherwise, OpenAI or Google's API, cheap variant for drafts and expensive variant for finals, and never Midjourney. We cover the aggregator APIs, fal and Replicate, and running open models locally in ComfyUI, two episodes from now.
The recap. Editing replaced generation as the core task, and every serious model converged on the same shape, a language model reading an instruction plus reference images. OpenAI leads the quality boards as of this recording, Google leads on references, resolution and search grounding and watermarks everything, Midjourney leads on aesthetics and refuses to be a platform, and Flux is the open-weight default with a license you must actually read. Control is three separate problems, change a thing, keep a subject, keep a structure, and only the first got easy. Output ownership depends on human contribution, commercial use depends on the vendor, and provenance is now two layers and a compliance requirement.
Next episode is video, where every one of these ideas, references, consistency, native multimodal models, provenance, shows up again with a time axis and a much higher price per usable output. That is AI Video Generation in twenty twenty six.