On this page

For AI agents: a documentation index is available at /docs/llms.txt. Append .md to any page URL for markdown, or send Accept: text/markdown.

Create and calibrate evaluators

Evaluators are checks you define for criteria specific to your product, such as whether your agent quoted the right refund policy or picked the right tool. They complement the built-in signals, which run on every session without setup. Each active evaluator writes one [Agent] Evaluator Result event per closed session, so you can chart, filter, and build cohorts on its results.

LLM evaluators make model calls with a model provider key you add. Refer to Set up custom evaluators to add one. Code evaluators don't need a key.

Choose an evaluator type

Each evaluator has an output type:

  • Binary: True or False. Binary questions calibrate most reliably, so start here.
  • Classification: One label from a set you define, with a description and examples for each label. Use three or more labels. For a two-outcome question, use a binary evaluator, which always emits True or False.

Code evaluators make no model calls, run fast, and give the same answer every time. LLM evaluators spend tokens on your key for each session they score.

Create an evaluator

  1. Open Evaluators in the Agent Analytics left navigation.
  2. Start from a suggested evaluator, such as "When did the agent select the wrong tool?", or create one from a blank Binary or Classification template.
  3. Write the prompt or the code rule. For an LLM evaluator, describe each label and add examples of sessions that fit it.
  4. For an LLM evaluator, pick the judge model. A small, fast model is usually enough for a well-scoped binary question.
  5. Choose which sessions the evaluator runs on. Run it on all in-scope sessions, set a sample rate, or filter by conditions such as tool errors, user feedback, latency, turn count, or a signal result.
  6. Choose what the prompt includes. Besides the conversation turns, you can send session context keys and named spans to the model.
  7. Choose the agents the evaluator applies to. By default, an evaluator applies to every agent in the project.
  8. Save the evaluator as a draft.

A draft doesn't score live sessions. Test and calibrate it first.

Calibrate against your reviews

Calibration measures how often the evaluator agrees with your own judgment on the same sessions.

  1. Review sessions in the Sessions view and tag the outcome you expect on each one.
  2. Run the draft evaluator on those manually tagged sessions. Amplitude saves them as a dataset named after the evaluator, such as "Refund Policy Calibration Set".
  3. Open the evaluator's Calibration panel. It shows the agreement score between the evaluator and your tags, and labels it Looks aligned or Needs work.
  4. If agreement is low, refine the prompt or the labels and run again. Around 90% agreement is a good bar to activate.

The calibration score covers manually reviewed sessions only, updates as you add or change tags, and shows per version after your first revision. To test on a larger sample, run the evaluator on a dataset from the Runs page. Refer to Datasets and Runs.

Activate an evaluator

When you activate an evaluator, it scores every new closed session it applies to.

  • Model required: An LLM evaluator needs a judge model and a model provider key before you can activate it.
  • Locked while active: You can't edit an active evaluator. Editing creates a new draft version, and earlier versions stay in the version history.
  • One version at a time: Activating a new version deactivates and replaces the active one.
  • All agents by default: If the evaluator applies to every agent, Amplitude asks you to confirm. To scope it, go to Edit Eval > Agents > Select Agent.

Active evaluators score sessions that close after activation. To score past sessions, start a run from the Runs page over specific sessions, a date range, or a dataset.

Creating and activating evaluators depends on your role-based access control permissions. Evaluator Result events count against your Amplitude event volume.

Reuse an evaluator in another project

To use the same evaluator in other projects, duplicate it to one or more projects. Calibrate each copy on that project's sessions before you rely on its results.

Analyze evaluator results

Each [Agent] Evaluator Result carries the result, the judge model, the evaluator version, and the evaluation cost, plus session dimensions such as agent, cost, turns, and topic. Filter the Sessions view by an evaluator result, or chart the event in any Amplitude chart. For the event's properties, refer to [Agent] Evaluator Result.

Last verified on October 6, 2026.

Was this helpful?