Skip to content

AI disguised as human feedback is poisoning RLHF

Pathwize ResearchData quality and provenance2 min read
ProvenancePathwize AI

When contractors quietly route tasks through an LLM, the human signal you paid for disappears. How to detect and design against it.

There is a failure mode in RLHF that almost no one puts on a slide: some of the human feedback you paid for was produced by a machine. A contributor pastes the task into an LLM, lightly edits the output, and submits it as their own judgment. It looks human, passes a quick review, and quietly poisons the reward model you are training.

Why it happens

It is not usually malice. Contributors are paid per task, deadlines are tight, and a capable model is one tab away. When the incentive is speed and the oversight is thin, routing tasks through an LLM is the path of least resistance. The more a platform pushes volume over quality, the worse it gets.

Why it is so damaging

The whole point of RLHF is to teach a model what humans prefer. If the 'human' preferences are actually a model's preferences, you are training your model to imitate another model, and inheriting its blind spots and errors, while believing you captured genuine human judgment.

It is also self-concealing. The contaminated answers are fluent and confident, so they sail through light review. The damage only surfaces later, as unexplained quality drift in evaluations, by which point it is baked into the training set.

Why the usual defences fail

Spot checks sample randomly and assume the rest is fine, but contamination clusters around specific contributors, task types and deadlines. Attestations ('I did not use AI') are unenforceable. And blunt AI-text detectors are unreliable on short, edited answers. None of these give you something you could show a customer.

Design against it, do not just detect it

The durable fix is structural: verify contributors, capture the human step as work happens, score trust per batch rather than per person, and monitor live agreement so drift shows up immediately. Combined with behavioural signals (implausible timing, boilerplate phrasing, characteristic model errors), this lets you quarantine a suspect batch instead of distrusting a whole dataset.

The goal is not to win an arms race with detectors. It is to make the authentic human signal the path of least resistance, and to be able to prove it.

Keep the human signal you paid for

Pathwize pairs credential-verified experts with recorded human oversight and per-batch trust, so the feedback is genuinely human and you can evidence it. Book a demo to pressure-test it on a sample of your pipeline.

Frequently asked questions

What does 'AI disguised as human feedback' mean?+

It is when contributors route RLHF or evaluation tasks through an LLM and submit the output as their own human judgment. The feedback looks human but reflects a model's preferences, which contaminates the reward model you train.

Why is fake human feedback so harmful to RLHF?+

RLHF teaches a model what humans prefer. If the preferences are actually a model's, you train your model to imitate another model and inherit its errors, while believing you captured human judgment. It also hides in fluent answers and only surfaces later as quality drift.

Can AI-text detectors solve this?+

Not reliably. Detectors struggle on short, edited answers, and attestations are unenforceable. The durable fix is structural: verify contributors, capture the human step, score trust per batch, and monitor live agreement, so you can prove and quarantine rather than guess.

Related reading

See Pathwize on your own data

Source verifiable expert data with provenance built in, EU-native and audit-ready.

Book a demo
← All posts

Related stories

ProvenancePathwize AI

How to prove your RLHF data was actually labelled by humans

AI quietly routed through a contractor looks like human feedback until it poisons your reward model. Here is how to detect it and design a pipeline that can prove the human signal.

InsightPathwize AI
Provenance

Data provenance for AI, explained

Provenance is the record of where your data came from and how it was produced. Here is why it is becoming the difference between a defensible model and a liability.

ProvenancePathwize AI

A provenance trail every AI auditor will accept

Auditors do not want a summary, they want to follow one output back to the person and process that produced it. Here is what a trail they trust looks like.