GradeThread

How GradeThread grades

A condition grade is only trustworthy if the method behind it is open to inspection. This page documents what actually produces a GradeThread grade, how each rubric version is evaluated before it can serve, exactly what a grade does and doesn't claim, how we handle errors, and where human judgment sits in the loop — the same standard we'd want from anyone grading our own items.

What does the grading

Every garment is scored against one fixed rubric — five weighted factors (Fabric Condition 30%, Structural Integrity 25%, Cosmetic Appearance 20%, Functional Elements 15%, Odor & Cleanliness 10%) combined into a single 1.0–10.0 grade. See the full rubric on the grading standard.

We do not train a model of our own, and we think saying so plainly is worth more than the alternative. Your photos are read by a general-purpose vision model built by Anthropic — the Claude family — which we direct with our own versioned rubric and scoring instructions. Nothing about your garment adjusts that model's weights. The judgment being applied is the rubric; the model is what reads the photographs.

Expert reviewer corrections and real post-sale outcomes do feed a continuous accuracy loop — as scored reference examples and as cases in the golden set used to test each new rubric version, not as training data. That distinction is why the exact model behind a grade is recorded alongside it, and why a rubric version that passed its evaluation on one model is refused if a different one would serve it. An evaluation result does not transfer across models, so we do not let it.

The one thing we do not publish is the wording of the scoring instructions themselves. Everything they are measured against is public: the rubric and its tolerances, the confidence threshold below, the error and agreement bars a new version has to beat, and the accuracy we actually achieve. You can check the output against the standard without reading the prompt, which is the part that matters.

What a grade claims — and doesn't

A grade is

  • An objective assessment of physical condition
  • Reproducible against a published rubric
  • Independently verifiable via a certificate

A grade is not

Error handling & human review

Every grade carries a confidence score between 0 and 1. When confidence falls below 0.75, the submission is routed to a human reviewer before the grade is finalized — low-confidence cases never ship unchecked. Additional triggers force review regardless of the score: two or more contradictions found when the photos are re-checked against the assessment, an authenticity or tampering flag, or an incomplete photo set. A single contradiction lowers confidence rather than forcing review, which can route the grade anyway once it falls below the threshold. Buyers can dispute a grade, and reviewer corrections feed back into the accuracy loop. A grading prompt promoted through our admin flow cannot serve live traffic unless its most recent run against a golden set of expert-graded garments cleared fixed error and agreement thresholds, on the same model that will run it. That refusal is enforced in code. The prompt that ships inside the service is the fallback when no promoted version is active, and changing it is a deploy rather than a promotion, so it is held to the same shadow, eval and canary sequence by policy rather than by that automatic refusal. We publish the resulting platform-wide accuracy on the transparency report.

Methodology FAQ

Does GradeThread train its own grading model?
No, and the distinction is worth stating. Photos are read by a general-purpose vision model built by Anthropic, directed by GradeThread's own versioned rubric — five weighted factors (fabric, structure, cosmetics, function, odor) combined into a 1.0–10.0 score. Nothing about a graded garment adjusts that model's weights. Expert reviewer corrections and post-sale outcomes feed a continuous accuracy loop as scored reference examples and as cases in the golden set that tests each new rubric version. A version promoted through our admin flow cannot serve live traffic unless its most recent run against that golden set cleared fixed error and agreement thresholds, measured on the same model that will run it. The rubric shipped inside the service is the fallback when no promoted version is active; it follows the same shadow, eval and canary sequence as policy rather than as an automatic refusal.
What does a GradeThread grade claim — and not claim?
A grade is an assessment of a garment's physical CONDITION against a published rubric: wear, damage, structural soundness, function, and cleanliness. It is not an authentication (it doesn't verify a brand or that an item is genuine — see grading vs. authentication), not an appraisal of monetary value, and not a guarantee of fit. It grades condition, objectively and reproducibly, and nothing more.
How does GradeThread handle grading errors?
Every grade carries a confidence score. When confidence falls below threshold, the submission is routed to human review before the grade is finalized, so low-confidence cases never ship unchecked. Buyers can dispute a grade, reviewer corrections feed back into the accuracy loop, and platform-wide agreement and error rates are published on the transparency report.
Are humans involved in grading?
Yes. Human reviewers correct low-confidence grades before they finalize, adjudicate disputes, and maintain the golden set that gates every promoted rubric version. The AI does the volume; humans hold the standard.

Ready to Grade Smarter?

Join resellers who trust GradeThread to standardize condition grading, build buyer confidence, and sell faster.