The problem that never gets named clearly: when an AI-generated draft lands in an editor's queue, they have no way of knowing how much work it needs before they open it. 'The generated texts are sometimes really bad' and 'takes time to edit' are two ends of a spectrum — but right now the editor discovers where a draft falls on that spectrum only after they've already invested time reading it. The triage is invisible, so every draft gets the same full editorial pass regardless of how much it actually needs.

This persists because AI writing tools are optimized for generation volume, not for output quality transparency. Telling users 'this draft is bad' would suppress usage — so there's a direct commercial incentive not to surface quality signals. The tools show a word count. They don't show an editing estimate.

What editors actually need is a fast signal before they open a draft: is this 10 minutes of light polish or 45 minutes of structural reconstruction? That signal can be built from patterns that are measurable without a human reading the full piece — sentence-level tone consistency, deviation from the brand's established vocabulary, structural coherence across sections, reading level drift, and whether the draft mirrors the factual claims in the original brief. Users say 'long-form content generation still needs occasional manual polishing' and 'editing was required as the company was messed up after the content was generated' — the second complaint is a structural coherence failure that's detectable programmatically.

The business case is simple: an editorial team processing 50 AI drafts a week and mis-triaging half of them is burning hours on the wrong pieces and letting under-edited content through. A reliable quality signal lets the editor spend their time where it matters. The need recurs every publishing cycle, and it scales directly with how much AI content the team produces — which is going up, not down.

What to build

Build a Slack and Google Docs-integrated draft scoring service that analyzes submitted AI-generated documents against a team's defined quality criteria — tone consistency, structural completeness, brief adherence — and posts a score with a breakdown of specific problem areas before the editor opens the file.

Where to start

Start with content agencies that produce AI-assisted content at scale for multiple clients and already have a QA step before delivery — they have an existing workflow slot for a scoring tool and a direct financial cost when under-edited content reaches the client.

The hard part

Calibrating the scoring model to a specific team's quality bar requires enough of their historical content and editorial feedback to be meaningful — early customers will need a hands-on onboarding process that doesn't scale easily and requires domain judgment calls that can't be fully automated at launch.

How it makes money

Monthly subscription priced per editor seat, with a per-document overage charge above a volume threshold; free trial capped at 20 documents to let teams feel the triage improvement before committing.

See the evidence. The complaints behind this idea, the products they came from, and similar ideas in AI Writing Assistant.

More ideas in AI Writing Assistant