Preference compares two answers. Evaluation scores one. But sometimes you do not have a good answer to compare or grade at all, you need someone to write it. That is a domain judgment task: an expert produces the ideal answer, or the authoritative verdict, for each item. It is supervised fine-tuning data at its source, the gold response the model will learn from.
Gold data starts with a qualified author
The value of an SFT example is bounded by whoever wrote it. A gold answer in oncology, contract law, or frontier maths is only gold if the author is genuinely qualified to produce it. This is the task type where credentials matter most, because there is no second answer to check against; the expert's output is the standard, so the standard is only as good as the expert.
Two flavours: ideal answer and expert verdict
Some domain judgment tasks ask the expert to write the ideal answer from scratch, the response you wish the model would give. Others ask for a verdict on an existing answer: is this AI-drafted response correct, incorrect, or an edge case, and why? Both produce authoritative expert judgment. The first builds gold responses; the second builds a labelled record of what correct looks like, which is just as useful for training and evaluation.
Why edge cases are worth capturing
The most valuable domain judgments are often the ones that are not clean pass or fail. An expert flagging that an answer is technically correct but misleading, or right in general but wrong for this patient, teaches the model the nuance a binary label would erase. Give experts room to record the edge case rather than forcing a verdict, and capture the reasoning alongside it.
Where domain judgment fits
This is the foundation stage of a data programme. Domain judgment builds the gold set the model is fine-tuned on and the reference answers your evaluations grade against. Preference then tunes behaviour, evaluation measures it, and red teaming stress-tests it, but all of that rests on the quality of the expert answers underneath. Get this layer right and the rest has something solid to stand on.
On Pathwize
A domain judgment task on Pathwize routes each item to credential-verified experts in the relevant field and records the answer or verdict, the reasoning, and the author, as a signed trail. You get gold data whose provenance you can prove, which is exactly what Annex IV expects you to be able to show. Create one from the AI tasks console and set the required expertise.