The moment this problem surfaces: someone needs a specific image for a presentation, a landing page, or a document — they know exactly what they want in their head — and they type something reasonable into Firefly and get something that's close but wrong in ways they can't articulate or fix. They tweak the prompt. It gets different but not better. They spend twenty minutes on what felt like a thirty-second job. The complaint 'it can be difficult to get the image in my head to be what is on the screen' is the cleanest expression of this — the gap between intent and output is real and it's not obvious how to close it.
The reason this gap persists is structural: prompt engineering is a skill that takes time to learn, and the feedback loop in image generation tools is slow and opaque. You don't know why a prompt didn't work, so you don't know how to fix it. Adobe has built Firefly's interface around people who already have a vocabulary for visual description — art direction language like 'golden hour lighting' or 'shallow depth of field' — but the users coming to it from productivity contexts (marketers, HR teams, ops managers making a quick graphic) don't have that vocabulary and have no way to acquire it in-context.
What's missing is a structured intake that asks non-visual questions — 'What's the mood?', 'What's it for?', 'Who's in it?', 'What should the background feel like?' — and translates those answers into a prompt that actually produces something close to correct on the first try. Not a prompt editor. A structured form that asks for intent and handles the translation layer so the user never has to learn prompt syntax.
This is a business because the organizations deploying Firefly to non-creative employees — marketing teams, internal comms, HR, product — keep running into the same adoption problem: people try it, get inconsistent results, and fall back on stock photo libraries or asking a designer. Every failed Firefly session is a recurring cost. A structured intake layer that dramatically improves first-try success rates would be bought at the team or department level by anyone trying to get actual ROI out of a Firefly or Creative Cloud for Teams license.
What to build
Build a web-based structured form that walks users through five to seven intent questions about their image (purpose, tone, setting, subjects, color feel), assembles a Firefly-optimized prompt from those answers using a prompt template library, and submits it via the Firefly API — presenting results alongside the generated prompt so users can see the translation and learn from it over time.
Where to start
Target HR and internal communications teams specifically, who regularly need images for internal newsletters, culture decks, and training materials and who are the least likely to have any design support but are also early adopters of AI tools for productivity — a single enterprise HR team is a repeatable sales motion into a company's broader comms function.
The hard part
The quality of the output is dependent on Firefly's underlying model and API access terms, which means you're building on infrastructure you don't control — if Adobe tightens API access or changes the model in ways that break your prompt templates, your core value proposition degrades overnight without warning.
How it makes money
Per-seat monthly subscription sold at the team level ($12-18/seat/month), with a usage-based tier for smaller teams based on number of image generation sessions per month.
See the evidence. The complaints behind this idea, the products they came from, and similar ideas in Photo Editing.
More ideas in Photo Editing