Research
Latest research
7Latest research7
Verifiable human judgment: how to prove your AI training data is trustworthy
Verifiable human judgment: how to prove your AI training data is trustworthy
Frontier labs train and evaluate on human judgment they cannot verify. This guide explains how data provenance, expert and batch trust scores, and live inter-rater agreement make human-labeled data auditable.
Dr. Helena Vogt
June 18, 2026
Synthetic data vs human data: where generated data hits a wall
Synthetic data vs human data: where generated data hits a wall
Synthetic data scales until the task gets hard. This guide maps where generated data works, where it breaks down on frontier maths, medicine and security, and how to decide between synthetic and expert data.
Mareike Hoffmann
December 17, 2025
How to evaluate LLMs with expert graders and resist benchmark gaming
How to evaluate LLMs with expert graders and resist benchmark gaming
Most LLM benchmarks grade from final text and reward the appearance of correctness. This guide shows how to design expert-graded evaluations that resist gaming, including undisclosed-AI submissions.
Dr. Helena Vogt
March 16, 2026
The expert data economy: where demand for frontier AI experts is heading
The expert data economy: where demand for frontier AI experts is heading
AI data work is moving from cheap crowd labelling to credentialed physicians, lawyers and scientists. This guide maps where demand is concentrating and what it means for experts and for Europe.
Pathwize Research
January 2, 2026
Human in the loop: how human expertise shapes modern AI
Human in the loop: how human expertise shapes modern AI
Every leap in model capability is underwritten by human judgment. This guide explains how human feedback and domain expertise shape modern AI, and what an honest human-in-the-loop workflow looks like.
Jonas Albrecht
January 5, 2026






