# How to Identify Grading Errors From Return Patterns Before They Hit Your Seller Performance

_By GradeThread Team · Published October 1, 2026_

> A healthy overall return rate can hide one badly graded category. Here's how to tag returns, slice them by category and defect, and catch the pattern early.

To identify grading errors from return patterns, tag every return with a reason code and a defect note, then group them by category, grade tier, and defect type. A grading problem looks like a cluster, not a spike: the same defect, in the same category, at the same tier, showing up again and again.

Here is the scenario. You sell 200 items a month and take 12 returns. That is 6%, which feels fine. But 7 of those 12 are knits sold as Very Good, and all 7 say some version of "pilling not shown." Your overall rate looks healthy. Your knit grading is broken. Nothing on a standard dashboard will tell you that.

## Why the overall return rate is a blind spot

eBay's seller performance metrics roll returns up into one account-level number. Item-not-as-described (INAD) cases are tracked as their own signal, and too many of them can cost you Top Rated status or search visibility. We won't quote thresholds here because eBay changes them, so check your current Seller Dashboard for the exact numbers.

The problem is averaging. Say 150 of your 200 monthly items are denim and tees that almost never come back. They dilute a knit category that returns at 25%. By the time the blended number moves, you have already shipped dozens of mis-graded items and collected the cases to prove it.

The failure mode is simple: you see the symptom (a rising rate) weeks after the cause (a grader who keeps missing the same defect). The fix is to look at returns the way you look at sourcing: by category.

## What data you need to capture on every return

You cannot analyze what you did not record. Most resellers log the refund amount and nothing else. For pattern work, capture five fields per return:

- **SKU**, so you can pull the original grade and listing.
- **Category** (denim, knits, outerwear, shoes, and so on).
- **Grade tier at listing**: NWT, NWOT, Excellent, Very Good, Good, Fair, or Poor.
- **Reason for return code**: INAD, changed mind, wrong size, damaged in transit, other.
- **Defect description**: the buyer's words, then your normalized tag (pilling, odor, seam wear, stain, fading, missing button).

The reason code and the defect tag do different jobs. A reason code of INAD tells you the listing missed. The defect tag tells you what it missed. Track reason for return code by category and you find where the problem lives. Add the defect tag and you find what to fix.

[Screenshot placeholder: FlipDesk return intake form with reason code, defect tag, and linked SKU grade]

## How to identify grading errors from return patterns: a step-by-step workflow

Run this monthly, or weekly if you ship more than 500 items a month. It takes about 20 minutes once your data is clean.

1. **Export last 90 days of returns.** Thirty days is too thin for most categories. Ninety gives you enough volume to see clusters.
2. **Join each return to its original sale.** You need the grade tier and category from listing time, not from memory.
3. **Filter to INAD and condition-related reasons.** Remove wrong size, changed mind, and transit damage. Those are real costs, but they are not grading errors.
4. **Group by category and calculate each category's return rate.** Returns divided by items sold in that category, not by total sales.
5. **Within any category above your baseline, group by defect tag.** One tag that dominates is your signal.
6. **Cross-check against grade tier.** If the cluster sits in Very Good and Excellent, you are likely over-grading. If it sits in Good and Fair, check whether your photos or description left something out.

That last step is where most diagnoses get sharper. A category returning at 20% across all tiers suggests a sourcing or description issue. A category returning at 20% only in Excellent suggests the line between Excellent and Very Good is drifting.

## Reading the clusters: what the pattern is telling you

Different shapes point to different causes. Use this table as a starting diagnosis, then verify against the actual garments if you still have return stock.

| Pattern | Likely cause | First fix |
| --- | --- | --- |
| One defect tag dominates in one category | A systematic miss in that category's inspection | Add that defect to your intake checklist for the category |
| Returns concentrated in Excellent or Very Good | Over-grading at the top tiers | Re-grade a sample of current inventory at those tiers |
| Odor complaints across categories | Odor & Cleanliness not checked or not disclosed | Add a smell check step and a disclosure line |
| Defects buyers cite are visible in your photos | Buyers not reading, or description vague | Name the defect in the title or first line of the description |
| Defects are not visible in your photos | Photography gap | Add a defect close-up shot for that category |
| Spread evenly, no dominant tag | Probably noise or fit-related | Keep watching; do not change grading yet |

Map each defect tag back to the five grading factors: Fabric Condition, Structural Integrity, Cosmetic Appearance, Functional Elements, and Odor & Cleanliness. Pilling and fading are Fabric Condition. Seam wear and loose hems are Structural Integrity. Stains are Cosmetic Appearance. Broken zippers and missing buttons are Functional Elements. If your clusters keep landing on one factor, that is the factor your inspection routine under-weights.

## Catching undisclosed defects early, before the cases stack up

Waiting for 90 days of data is the safety net. You want a faster trigger too. To catch undisclosed defects in returns early, set a simple tripwire: when any category-plus-defect combination produces two returns inside 30 days, stop and inspect.

Two is a small number, but it matters because the cause is usually upstream. If two knits came back for pilling, the other knits you listed that month were graded by the same eye under the same light. Pull the active listings in that category and re-inspect five of them. If two of the five have the same problem, revise those listings now. Editing a listing costs minutes. An INAD case costs the sale, return shipping, the relist, and a mark on your record.

Show the math on one cluster:

- Sale price: $42.00
- Final value fee and per-order fee (example, around 13.6% plus a fixed fee): about $6.20
- Outbound shipping: $7.50
- Return shipping you cover on INAD: $7.50
- Refund: $42.00 (fees may be credited back in part, depending on the case)

On a single INAD return you may be out the original shipping, the return shipping, and a month of dead stock. Seven of them in one category in one quarter is a few hundred dollars, plus the performance risk. Fee rates vary by category and change over time, so treat these as example numbers and plug in your own.

## Fixing the grade, not just the listing

Once you have found the pattern, resist the urge to only soften your descriptions. Better wording helps, but the lasting fix is in the grading process.

1. **Write a category-specific defect checklist.** Knits get a pilling check under angled light. Denim gets a crotch and hem wear check. Outerwear gets a lining and zipper check.
2. **Define the tier boundary in writing.** What exactly separates Excellent from Very Good for knits? One sentence, with a photo reference.
3. **Re-grade the at-risk inventory.** Downgrade items that would not pass the new boundary. A Very Good that should be Good sells at a lower price but stays sold.
4. **Watch the next 30 days.** If the defect tag stops clustering, the fix worked. If it persists, the problem is deeper than the checklist.

Honest grading does not guarantee zero returns. Buyers change their minds and sizes run small. What it does is remove the returns you caused, which are the only ones you can control.

## Make the review a habit, not a rescue

The resellers who get hurt by this are rarely careless. They are busy, and their dashboard showed a decent number. A standing monthly review fixes that. Put it on the calendar next to your sales tax and payout reconciliation, because it belongs in the same conversation: every return is a payout line that does not match a sale.

If you track returns in FlipDesk, the reason code, defect tag, and original grade tier sit on the same record as the SKU, so the category slice above is a filter instead of an afternoon of spreadsheet joins. To test it, log your last 20 returns with a defect tag each and see whether a cluster appears. If you want a consistent starting grade on the garments themselves, run one through GradeThread and compare its factor breakdown to your own call.

---

Canonical: [https://gradethread.com/blog/identify-grading-errors-from-return-patterns](https://gradethread.com/blog/identify-grading-errors-from-return-patterns)
