For AI teams

Types of AI training data

Choose between demonstrations, preferences, rubrics, trajectories and adversarial examples.

4 min read

The right data format depends on the behaviour you are trying to create or measure. Selecting a familiar format without defining that behaviour first creates expensive, low-signal data.

Demonstrations

High-quality input-output examples show the model what a good response looks like. They work best when contributors can explain the standard and cover meaningful variation.

Preferences and critiques

Pairwise rankings reveal which of two outputs better meets the goal. Critiques add diagnostic value by explaining the error and what should change.

Evaluations and trajectories

Rubric evaluations measure behaviour without necessarily training on the answer. For agents, step-by-step trajectories reveal planning, tool selection, recovery and completion failures.

  • Gold and calibration items
  • Safety red-team prompts
  • Tool-use traces
  • Outcome verification
  • Expert adjudication

Corrections and structured error labels

A corrected response shows both the failure and a better alternative. Structured error labels make the dataset easier to analyse: teams can distinguish factual errors, missing constraints, unsafe advice, weak tool use and style problems instead of treating every rejection as equivalent.

Synthetic data and human verification

Synthetic examples can expand coverage and create rare scenarios quickly, but they inherit the generating model's assumptions and mistakes. Use humans to validate realism, remove duplicates, check specialist facts and confirm that generated cases represent the target population.

Match the format to the learning objective

Demonstrations are useful when a clear ideal output exists. Preferences help optimise relative quality. Critiques explain failure modes, while evaluations measure a system without necessarily training it. Agent trajectories are needed when the path—planning, tools and recovery—matters alongside the final answer.

  • Define the behaviour before choosing a format
  • Collect the minimum fields needed for that behaviour
  • Keep evaluation data separate from training data
  • Version schemas and instructions together
  • Measure downstream effect before scaling collection

Plan a balanced dataset

A balanced set reflects meaningful variation in users, tasks, difficulty, language and risk—not necessarily equal counts in every category. Oversample rare, consequential failures for learning while preserving a representative evaluation set for honest measurement.

Putting “Types of AI training data” into practice

The useful next step is to turn the concept into an observable workflow with a defined standard. These checks help contributors and AI teams avoid collecting activity without a clear learning or evaluation purpose.

Define the intended behaviour

Write down who the system serves, what a successful demonstrations looks like and which mistakes matter most. Concrete examples should show both acceptable variation and failures that require correction or escalation.

Pilot before scaling

Run a small batch with representative contributors, compare disagreements and revise unclear instructions. A pilot reveals missing context and inconsistent labels before those problems are multiplied across a larger dataset or evaluation run.

Preserve evidence and feedback

Keep the instruction version, source material, reviewer rationale and final decision together. When teams change a model, prompt, tool or policy, those records make it possible to rerun difficult cases and measure whether the change genuinely improved quality.

Continue learning

Start a campaign