Collecting preference data at scale means turning "which output is better" into a dataset large enough and clean enough to train a reward model that doesn't get gamed. This guide covers the exact steps, the tools, and the mistakes that quietly wreck RLHF pipelines in 2026.
- Preference data quality beats volume: 1,000 clean pairwise comparisons outperform 10,000 noisy ones.
- Reward hacking starts with sloppy annotation guidelines, not with the model — fix the rubric first.
- Rapidata's API delivers crowd-sourced preference data and RLHF labeling without building an annotator pipeline from scratch.
- Inter-annotator agreement below 70% means your reward model is training on noise, not signal.
Why this matters
A reward model is only as good as the preference data behind it. If your comparisons are inconsistent, ambiguous, or gathered from too few raters, the model learns to exploit the noise instead of learning what "good" actually means.
This is the failure mode teams hit in 2026 when they scale RLHF fast: annotation throughput goes up, agreement quality goes down, and the reward model starts rewarding outputs that look right to a rubric loophole instead of outputs that are actually right. Preference data collected without a repeatable process is the single biggest reason reward models get hacked by the policy they're supposed to be training.
What you'll need
- A clear task definition: what two outputs are being compared, and on what axis (helpfulness, safety, factuality, style)
- A pairwise or ranked-choice comparison format — avoid single-item scoring for RLHF, it doesn't transfer well to reward modeling
- A pool of human raters large enough to get 3-5 independent judgments per comparison
- An annotation interface or API that logs rater identity, timestamp, and confidence
- A quality control step: gold-standard questions with known answers mixed into the batch
- Access to a crowd-sourced labeling platform if you don't have in-house annotators — Rapidata provides this as an API/SDK layer specifically for RLHF preference data and model evaluation
The steps
1. Define the comparison axis before you write a single prompt
Decide exactly what "better" means for this batch — helpfulness, honesty, harmlessness, tone, format compliance. One axis per batch, not a blended "overall quality" judgment.
Mixed-axis comparisons are the top cause of low inter-annotator agreement. If raters don't know whether they're judging accuracy or friendliness, they'll disagree on grounds that have nothing to do with the model's actual weakness.
Common mistake: writing a rubric that says "pick the better response" with no criteria. That instruction alone can drop agreement to near-chance levels on subjective content.
2. Generate response pairs from diverse prompts and diverse model checkpoints
Pull prompts from your real usage distribution, not a curated demo set. Pair outputs from different checkpoints, temperatures, or even different base models so the reward model sees genuine quality variance, not near-identical text.
Aim for at least 1,000 comparison pairs per task category before you trust a reward model trained on the results — fewer than that and the model overfits to quirks of a small sample.
3. Write an annotation guideline with 3-5 concrete criteria and examples
Each criterion needs a one-sentence definition and at least one worked example showing a clear win and a clear loss. Vague rubrics ("choose the more helpful answer") get inconsistent labels; specific ones ("prefer the response that includes a working code example when the prompt asks for code") don't.
Include a tie-breaker rule. Raters need permission to say "equally good" — forcing a false choice on genuine ties injects noise straight into your preference data.
4. Route each comparison to 3-5 independent raters, not one
Single-rater labels can't detect disagreement, and disagreement is the signal that tells you a prompt is ambiguous or a rubric is broken. Three to five independent judgments per pair lets you compute inter-annotator agreement and flag low-agreement items for review instead of silently baking them into training data.
This is where crowd-sourced platforms earn their cost over in-house annotation: getting five independent judgments per item at scale requires a rater pool you don't build overnight.
5. Insert gold-standard control questions at a fixed rate
Mix in comparisons with a known correct answer — roughly 1 in every 10 items — and track each rater's accuracy against them. Raters scoring below your agreement floor on gold items get their batches re-reviewed or excluded entirely.
Common mistake: running gold checks only at the start of a session. Rater attention drops over long sessions; spacing controls throughout catches fatigue-driven errors mid-batch.
6. Compute inter-annotator agreement before you touch the reward model
Calculate agreement (Cohen's kappa or simple majority-vote consistency) per task category. An agreement rate above 70-80% signals a usable batch; anything lower means the rubric, the prompt set, or the rater pool needs fixing before you spend compute training on it.
Don't skip this step to hit a deadline. A reward model trained on preference data with poor agreement will optimize for whatever spurious pattern the noise contains, and that shows up later as reward hacking during RL fine-tuning.
7. Audit for reward-hacking-friendly patterns before training
Check whether raters are systematically preferring longer responses, more confident-sounding phrasing, or specific formatting regardless of actual correctness. These biases get amplified during RL and produce a policy that games the reward model instead of improving the underlying task.
A reward model built on preference data that resists these exploitable patterns is the difference between RLHF that improves model behavior and RLHF that just teaches the model to write longer, more confident-sounding wrong answers.
8. Version and log every batch of preference data
Store the prompt set, rubric version, rater pool composition, and agreement scores alongside the labels themselves. When a reward model behaves unexpectedly six months later, you need to trace it back to the exact preference data batch and rubric version that produced it.
Scale RLHF preference data collection
API and SDK access to crowd-sourced human feedback for reward model training.
Troubleshooting
- Agreement scores stuck below 60%: the rubric is ambiguous or the axis is mixed. Split the comparison into two separate single-axis batches and re-run.
- Raters converging too fast, near-perfect agreement on subjective tasks: check for anchoring — are they seeing each other's labels, or defaulting to the first option every time? Randomize response order per rater.
- Reward model rewards length over quality: audit your preference data for a length bias in the human labels themselves; the model is often just reflecting what raters actually preferred.
- Gold-standard accuracy dropping over a session: rater fatigue. Cap session length and rotate gold items so raters can't memorize the answer key.
- Reward model scores diverge sharply from held-out human judgment: the reward model has likely learned a proxy signal from the preference data rather than the intended axis — re-check rubric specificity before retraining.
- Preference data volume looks fine but reward model quality is flat: volume without diversity doesn't help; check whether your comparison pairs are pulled from a narrow prompt distribution.
Tools and resources
- A pairwise comparison interface with randomized response order and mandatory tie option
- Inter-annotator agreement calculators (Cohen's kappa, Fleiss' kappa for 3+ raters)
- A crowd-sourced labeling platform for rater diversity and scale, built around a reward model designed to resist gaming
- Version control for rubrics and prompt sets, not just for model checkpoints
- A dashboard tracking agreement rate and gold-standard accuracy per batch, updated continuously through 2026 as your rater pool changes
What to do next
Once a batch clears your agreement threshold, train a reward model on it and validate against a held-out human-judged test set before running any RL fine-tuning. Skipping that validation step is how teams end up discovering reward hacking only after the policy model has already been deployed.
FAQ
What is preference data in RLHF?
Preference data is a set of paired model outputs where a human rater indicates which one is better on a defined axis like helpfulness or accuracy. It's the training signal used to build a reward model in RLHF pipelines.
How much preference data do you need to train a reward model?
Most teams need at least 1,000 comparison pairs per task category before a reward model generalizes reliably. Smaller samples tend to overfit to quirks in the specific prompts used.
What is a good inter-annotator agreement rate for preference data?
An agreement rate of 70-80% or higher signals usable preference data in 2026. Lower agreement means the rubric or task definition needs revision before training on the labels.
Is crowd-sourced preference data better than in-house annotation?
Crowd-sourced preference data gives you rater diversity and scale that in-house teams struggle to match quickly. In-house annotation can work for narrow, specialized tasks but doesn't scale as fast for broad RLHF training runs.
How do you prevent reward hacking in RLHF?
Reward hacking is prevented at the preference data stage by writing specific, single-axis rubrics and auditing for biases like length or confident-sounding phrasing before training. A reward model built on biased preference data will get exploited by the policy during reinforcement learning.
What's the difference between pairwise and ranked-choice preference data?
Pairwise comparisons ask raters to pick the better of two outputs, while ranked-choice orders three or more outputs at once. Pairwise is simpler and produces cleaner reward model training signal for most RLHF use cases.
How many raters should judge each comparison?
Route each comparison to 3-5 independent raters so you can measure agreement and catch ambiguous items. A single rater per comparison gives you no way to detect noisy or biased labels.
Can preference data be collected via API?
Yes, platforms like Rapidata offer API and SDK access to crowd-sourced human feedback, letting teams request preference data and model evaluation at scale without building an annotation pipeline from scratch.
One last thing
The reward model that gets hacked almost never fails because of the model architecture — it fails because the preference data behind it let a spurious pattern (length, tone, confidence) slip through undetected. Audit the rubric before you audit the model.




