A list of complaint topics isn't a build list. Five named clusters from a plugin's (or a theme's) reviews tells you where that product's problems are, not which one is actually worth your time to go build.
Turning a topic list into an actual short list takes three separate judgments, weighed against each other on purpose instead of by gut feeling. Here's the framework.
The core of it: To prioritize which WordPress plugin complaints are worth fixing first, weigh each named cluster's frequency, severity, and buildability against each other on purpose - a topic list of complaints on its own isn't a build list.
Why a Topic List Isn't a Priority List
It's tempting to just chase whichever complaint topic has the most reviews attached to it. I hold that to be a reasonable tiebreaker, but a bad primary rule — a complaint that shows up in five reviews but describes a genuinely broken workflow is worth more than one that shows up in fifty reviews and amounts to "I wish the button were a different color."
Frequency alone also rewards whichever issue is easiest to complain about in one sentence, which isn't the same thing as whichever issue represents the biggest real opportunity.
The Three Questions That Actually Decide Order
Three separate questions, asked about every complaint topic on the list — remember, this list came from a plugin you're researching, not necessarily one you already own:
- Frequency — how often does it come up? Count how many distinct reviewers raised this complaint, not how many total sentences mention it. Ten reviewers each writing a paragraph outweighs two reviewers who each left three separate reviews.
- Severity — how bad is it for the person hitting it? Review Gap Finder already gives you an AI-ranked severity score for each complaint topic, which is a fast starting read — a complaint that breaks a core workflow or loses data outranks one about a missing nice-to-have, even at a lower frequency.
- Buildability — could you actually build a good solution for this? A severe, frequent complaint that would need a genuinely hard technical solution competes on a different timeline than an equally frequent complaint that turns out to be a straightforward feature to add well.
No single one of these should decide the order alone. A complaint that's frequent and severe but hard to solve well might still rank below a moderately frequent, moderately severe complaint you could build a clean, well-executed feature for quickly — the second one is the better use of your next few weeks.
A Simple Way to Score Each Topic
You don't need a complicated formula for this — a rough high/medium/low rating on each of the three questions, side by side, is usually enough to see the order clearly.
- List every complaint topic down the left side.
- Rate each one High, Medium, or Low on frequency, severity, and buildability (buildability read as "how realistic," so High buildability is good news).
- Anything High on frequency and severity goes to the top of the list regardless of buildability — those are worth scheduling as a serious project even if they're hard, just not necessarily the first thing you ship.
- Among the rest, favor whatever combines decent frequency or severity with High buildability — that's where a realistic amount of work addresses a real, validated gap.
- Anything Low across all three drops to the bottom, or off the list — not every named pattern is worth building for, especially a rare, minor one.
A Hypothetical Example, Scored
This example is illustrative only — not real data from any specific plugin. Say you've run Review Gap Finder on a plugin you're researching and it turns up four complaint topics:
| Complaint topic | Frequency | Severity | Buildability | Priority |
|---|---|---|---|---|
| Fatal error with a specific caching plugin | Medium | High | High | Worth building first |
| Settings page confusing on first setup | High | Low | Medium | Worth building soon |
| Missing a feature only power users want | Low | Low | Low | Backlog, maybe never |
| No export option for a key report, chronic complaint | Medium | High | Low | Bigger project, worth scoping |
In my read, the highest-frequency complaint here — the confusing settings page — isn't the top priority. It's real and worth building around, but the caching conflict is both severe and something you could realistically build a clean compatibility fix for, which makes it the better first project even though fewer people mentioned it.
Where Timing Fits Into the Score
Whether a complaint is brand new or has been sitting for years changes how it should be scored, on top of these three questions. I've come to see a chronic complaint that's been unaddressed for a long time as stronger evidence of a real, standing gap than a fresh spike that the plugin's own team might patch before you'd ever ship anything. See how to tell if a complaint pattern is new or ongoing for how to make that call before you finalize the order.
Where the Topics and Severity Scores Come From

Review Gap Finder's summary view — this is where the severity read for each topic starts.
StatWP's Review Gap Finder does the first part of this for you automatically — it groups a plugin's (or theme's) reviews into complaint topics using AI and ranks them by severity, which gives you the raw list of topics, roughly how often each one comes up, and a starting severity read. What I keep finding: buildability isn't something a review-grouping tool can judge for you; that comes from your own read of what each complaint actually describes and your own sense of what you could realistically ship. Before scoring anything, it's also worth a quick check that each topic is really a product gap and not a setup issue being misattributed to the plugin — no sense scoring a complaint you'd never actually fix by building anything. The framework above is what turns the tool's output into an order.
Rank a Plugin's Complaints by Severity
When to Skip the Scoring and Just Move
The three-question framework is worth running formally when you're looking at a real list of several competing complaint topics and need to decide what to work on next. It's overkill for the obvious cases.
If a complaint topic is severe, clearly affects more than a handful of people, is chronic rather than fresh, and turns out on inspection to be a straightforward feature to build well, don't wait for a full scoring pass to justify starting — I use this framework for genuinely close calls, not to slow down a decision that's already obvious once you've read the complaint.
Save the formal scoring for the harder cases: a handful of topics that all seem to deserve attention, competing for the same limited time, where it's genuinely unclear which one to tackle first without laying them out side by side.
Frequently Asked Questions
What if two complaint topics score identically?
Use timing as the tiebreaker — a chronic topic that's been unaddressed for a long time outranks a fresh spike, even at an identical frequency/severity/buildability score, since the chronic one is stronger evidence the plugin's own team isn't going to close the gap.
Should I ever start with a Low-priority complaint before a higher one?
Occasionally, if it happens to be a genuinely trivial, well-scoped feature you could ship quickly without displacing anything else — but don't let easy wins crowd out the actual top of the list session after session.
How do I score severity for something like "confusing," which isn't a crash?
I ask what the reviewer actually did next — did they abandon setup entirely, or work around it and keep using the plugin? Abandonment is high severity even without an error message; a workaround that got someone to a working state is lower severity. Review Gap Finder's own severity ranking is a fast starting point for this same judgment.
Does a high review count always mean high frequency?
A high review count doesn't always mean high frequency — I count distinct reviewers, not total review text, when scoring a complaint topic this way. A handful of people who each left several angry follow-up reviews about the same issue can inflate a raw review count without actually reflecting how many separate users hit the problem.
Is this framework any different if I'm scoring my own plugin's complaints instead of a competitor's?
The three questions work the same either way. The main difference is buildability — on your own plugin you already know the codebase, while on a plugin you don't own, buildability really means "how hard would it be to build this well from scratch," which is a bigger judgment call.
Frequency, severity, and buildability, weighed against each other on purpose — that's the difference between a topic list and an actual plan worth acting on. For the step before this one, see how to turn a plugin's bad reviews into your next feature idea, and for confirming a pattern is real before you score it, see how to find the real complaint hiding in a 1-star review.