Most data quality systems rate people: this annotator is good, that one is not. It feels intuitive, but it is the wrong unit. Quality does not live in a person, it lives in the work they produced on a given task at a given time. The right unit to score is the batch.
Why per-person scoring misleads
A strong contributor has a bad day, misreads a new guideline, or hits a task type outside their strength, and produces a weak batch. A weaker contributor nails a simple batch. Per-person reputation smooths over both, so you keep shipping the strong contributor's bad batch and needlessly discount the weak contributor's good one.
Per-person scores also react slowly. By the time someone's average drops enough to notice, many batches have already gone into training.
What per-batch trust gives you
Attach a trust signal to each batch, from live inter-rater agreement, gold-standard items, and review outcomes, and you get a live, actionable metric. A batch below threshold can be paused, re-reviewed, re-clarified or re-routed before it contaminates a dataset. You act on the actual unit of risk, not a lagging average.
It also lets you quarantine, not distrust
The biggest practical win is containment. When a problem appears, per-batch lineage lets you quarantine the specific affected batches rather than throwing out a whole dataset or blacklisting a contributor. That is cheaper, fairer, and keeps good people doing good work.
Spot checks are the weakest version of this
After-the-fact spot checks are per-batch scoring done too late and too sparsely. They sample a little, look backwards, and miss clustered failures. Live per-batch trust is the same idea done continuously and completely, so drift is caught while you can still fix it.
Score the right unit
Pathwize scores trust per batch with live agreement and recorded oversight, so you catch drift early and quarantine precisely. Book a demo to see it on a sample of your data.