A creative director or senior retoucher is evaluating whether to standardize their team on Adobe Firefly generations, or route certain jobs to a different generation model based on subject matter. Right now, the only way to do this is to manually run the same prompt through multiple tools, screenshot the results, and make a gut call. There's no structured way to test, and no institutional memory of which model handles which subject matter better.
The complaint that 'the gen AI results are really behind competitors' and the wish that 'AI could select the best model' both point to the same structural problem: users know models differ by task, but there's no tooling to help them understand those differences systematically. Adobe has no reason to build a comparison tool that highlights where their model loses. The model vendors themselves only show their best cases.
What doesn't exist is a prompt-anchored benchmark runner where a user submits their actual working prompts — not synthetic test prompts — and gets back side-by-side generations with a structured evaluation interface. Not a generic leaderboard. A testing tool calibrated to your specific use case: your subjects, your lighting scenarios, your typical fill requests. The results accumulate into a team-level knowledge base: 'for product backgrounds, model X; for portrait retouching, model Y.'
Without this, teams either default to whatever's bundled in their existing software (often Adobe Firefly, which multiple users call behind competitors) or waste hours in informal testing that nobody documents. The cost is invisible but recurring: wrong model choices slow retouching workflows and produce more retries.
What to build
Build a web-based prompt testing tool where users paste their real generative fill or image generation prompts, run them against multiple models simultaneously, rate outputs on a structured rubric (realism, prompt adherence, edge quality, text accuracy), and export a model recommendation report for their specific use case.
How it makes money
Charged per testing session (a bundle of N prompt runs across M models), with team accounts on a monthly subscription for ongoing access to accumulated benchmark history — free for the first session to drive inbound from people actively evaluating tools.
See the evidence. The complaints behind this idea, the products they came from, and similar ideas in Photo Editing.
More ideas in Photo Editing