SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
How-to Guide 21 mins

Build a Private Agent Evaluation Set: Sampling and Holdout Design

A practical protocol for sampling real agent work, separating development from holdout cases, and redacting sensitive data without erasing task difficulty.

The PADISO Team ·

Prerequisites

Before collecting examples, agree on the agent behavior being evaluated and the decision the evaluation should inform. A set for deciding whether to expand a support workflow needs different cases from one used to compare prompt revisions. Write down the intended decision in one sentence, such as: “We will use this set to identify whether the revised agent can resolve routine account-access requests without increasing unsafe actions.” If the decision is vague, sampling will be vague too.

You also need access to a bounded source of real work, a person who understands that work, and a safe place to store a minimized dataset. The source could be resolved support tickets, completed operations tasks, or internal requests. You do not need to begin with every historical record. A carefully described sample of a defined period is easier to audit than a large, opaque export.

Decide who can review examples and who can adjudicate ambiguous labels. A domain reviewer should be able to distinguish a correct outcome from a plausible-sounding but incomplete one. Include a technical owner who can map examples to the agent’s available tools and record what the agent was expected to do. One person may fill more than one role, but the responsibilities should still be explicit.

Finally, establish a retention and access plan before copying source records. Remove fields that are irrelevant to the task, restrict access to the working set, and define when temporary copies will be deleted. This is a dataset-design practice, not a claim that any particular handling plan satisfies a legal or regulatory requirement. Ask the organization’s privacy and security specialists to review the plan where appropriate.

Pro tip: Freeze the intended evaluation question before inspecting model results. If the team repeatedly changes the question to fit what the current agent does well, the set becomes a moving target rather than evidence for a decision.

Step 1: Define what one case represents

A case is the smallest unit that can be scored consistently against the evaluation question. For a support agent, it might be one customer request together with the information the agent was legitimately allowed to use and the expected business outcome. For an operations agent, it might be a request, a relevant system state, and the action or explanation that would count as an acceptable resolution.

Do not treat a transcript as a complete case by default. The transcript may omit crucial context, such as whether the customer had already tried a recovery step, whether an account was locked, or whether an agent was allowed to take an irreversible action. Conversely, including an entire conversation history can expose irrelevant personal information and make the case harder to understand. Include only the context needed to interpret the task and judge the outcome.

Write a case contract before sampling. It should specify the input boundary, any permitted reference context, the expected output or outcome, the scoring method, and the conditions that make the case unscorable. “Answer the user correctly” is not a sufficient expected outcome. A more useful contract might say that the agent must determine whether the request is eligible for a self-service reset, provide the approved next step, and avoid changing account state when required identity evidence is absent.

Agent evaluations should assess outcomes as well as transcripts; the agent harness and evaluation harness are distinct components. (Anthropic, “Demystifying evals for AI agents”) That distinction matters when designing a case: a convincing final message is not proof that the intended action happened, and a correct system state is not proof that the agent communicated it clearly. Decide which evidence belongs in the case and which belongs in the scoring record.

A case contract also prevents label drift. If one reviewer considers “ask for more information” a successful result and another considers it a failure, their disagreement may reflect an underspecified target rather than a reviewer problem. Define acceptable alternatives and necessary stopping conditions before building a large sample.

Step 2: Build task strata before drawing examples

A task stratum is a group of cases that share a meaningful feature of the work. Strata help prevent the common mistake of sampling only the most frequent, clean, easy-to-label requests. Start from the workflow and the evaluation decision, not from whatever categories happen to be available in the ticketing system.

For a hypothetical account-support agent, useful strata might include routine eligible requests, missing-information requests, requests with conflicting account context, requests that require escalation, and requests where the requested action is not permitted. These categories describe different decisions the agent must make. “Email” or “chat” may be useful metadata, but should not replace decision-relevant categories unless channel changes the task in a way that affects success.

Use a small, stable taxonomy. Each case should have one primary stratum so counts are interpretable, with optional secondary tags for overlapping properties such as language, urgency, or tool dependency. If a case can fit several primary strata, define a precedence rule or revise the categories. A taxonomy that reviewers cannot apply reliably will produce counts that look precise but are not comparable.

StratumWhat it testsTypical sampling concernExample evidence to retain
Routine and eligibleCorrect completion of a common, bounded requestOverrepresentation can make the set too easyRequest, eligibility context, expected completion state
Missing informationRecognition that the agent cannot safely proceed yetSource records may not show what was absent at decision timeAvailable fields and the specific missing prerequisite
Conflicting contextHandling of inconsistent or stale informationConflicts may disappear when records are normalizedRelevant conflicting values and their provenance
Escalation requiredAppropriate handoff or safe stoppingEscalated cases may lack a clear final outcomeReason for escalation and acceptable next step
Disallowed or out of scopeRefusal or redirection without an unintended actionRare cases can be missed in frequency-based samplingRequest boundary and action that must not occur

This table is a starting point, not a universal taxonomy. Merge categories if the agent makes the same decision and the same scoring rule applies. Split them if one category combines materially different decisions. For example, “missing information” may need separate strata if missing identity evidence and missing preference details lead to different permitted actions.

Keep a written inclusion rule for each stratum. “Escalation cases” could mean cases ultimately routed to a human, or cases that should have been routed to a human regardless of what happened historically. For evaluation, the latter is often the relevant definition. Historical handling is evidence about the work, not automatically the standard for correct handling.

Step 3: Choose a sampling plan that matches the decision

First define the population: which work, from what source, over what date range, and under what inclusion rules. Record any known gaps, such as a queue not represented in the source or records that cannot be safely de-identified. Without a bounded population statement, later readers cannot tell whether a result describes the intended workflow or only the records that happened to be easiest to retrieve.

Then choose how to sample. A proportionate sample preserves the approximate mix of the defined population and can help describe performance on typical incoming work. A stratified sample deliberately selects cases from each decision category, which can make rare but important failure modes visible. A targeted challenge set can probe a specific risk, but its case mix should not be presented as representative of routine work.

For a first private set, stratified sampling is often useful because the purpose is usually to expose weaknesses across distinct task decisions, not to estimate a population-wide rate precisely. Keep the selection method consistent within each stratum. For example, randomly select from a time-bounded eligible pool, or use a reproducible interval after sorting records by a documented stable key. Avoid hand-picking only memorable failures unless the set is explicitly labeled as a challenge collection.

Set a target count per stratum based on review capacity and the importance of the decision. A small sample can reveal obvious defects, but cannot support fine-grained claims about rare events. Record the intended use of each stratum’s count: broad coverage, diagnosis, or a specific decision threshold. Do not imply statistical confidence merely because the dataset has a numeric size.

Illustrative calculation: suppose a team’s initial review capacity is 120 cases. It reserves 60 for development and 60 for a private holdout. Within each portion, it allocates cases across five strata, giving additional slots to routine requests while preserving minimum coverage for escalation and disallowed requests. The numbers are planning assumptions, not a recommended universal sample size. If a rare stratum contains only a few eligible historical cases, the team should report that limitation rather than manufacture a balanced-looking sample by duplicating records.

Record the sampling seed or deterministic selection rule, source period, strata counts, exclusions, and the person or process that performed extraction. If the selection includes judgment, describe it. Reproducibility means another authorized reviewer can understand how the set was assembled; it does not require exposing sensitive source records to everyone who reads the protocol.

Warning: A balanced sample answers “Can the agent handle these kinds of cases?” more readily than “How often will the agent fail in production?” Keep those questions separate. A deliberately enriched set changes the case mix.

Step 4: Separate development cases from the private holdout

Split cases before using them to tune prompts, routing, instructions, examples, or decision rules. Development cases are available for inspection and iteration. Holdout cases are reserved for evaluation after a candidate design has been fixed. The point is not that holdout examples are mysterious; it is that repeated exposure can turn them into informal training material, weakening their value as an independent check.

Make the split at the level that prevents leakage. If one customer, incident, or underlying workflow appears in multiple records, putting closely related records on both sides can make a holdout deceptively easy. Group related records first, then assign groups to development or holdout. Where grouping is uncertain, record the known relationship and treat a later overlap as a limitation rather than silently assuming independence.

Choose a split approach that fits the source. A time-based split can test whether a system handles later work, but it may confound the result with policy or workflow changes over time. A randomized group split may preserve a more similar mix, but it can allow near-duplicate cases across the boundary if grouping is weak. Neither method is automatically superior. State the risk each method leaves behind.

Keep the holdout manifest separate from the development workspace. The manifest should contain stable case identifiers, stratum, source cohort, split assignment, and version information—not unnecessary raw content. Restrict the raw holdout to the small group responsible for assembling and adjudicating it. The product team can receive aggregate results and case-level findings selected for diagnosis after an evaluation, subject to the organization’s handling rules.

A useful operating rule is to treat each full inspection of holdout cases as a consumption event. Record the candidate version, evaluation date, reason for opening the cases, and whether any examples were discussed in detail. If the team repeatedly adjusts the system in response to specific holdout examples, retire that holdout from future confirmatory use and assemble a fresh one from a later or untouched source pool.

The holdout is not a substitute for a broader evaluation framework. It is one controlled dataset component that can support a specific decision. For how the dataset fits into test harnesses, scoring, and ongoing evaluation, see AI Agents in Production: Agent Evaluation Frameworks. Keep this article’s focus on what cases enter the set and how they remain useful.

Step 5: Redact without deleting the task

Redaction should reduce exposure while preserving the information required to judge the task. Start by listing every field in the proposed case and asking what decision or score it supports. Remove fields with no stated purpose. Then classify the remaining content into direct identifiers, indirect identifiers, sensitive details, operational context, and task-relevant facts. This is a practical review aid, not a formal legal classification system.

Replace direct identifiers with stable, dataset-scoped tokens when linkage is genuinely necessary. If the reviewer must know that two messages concern the same account, a token can preserve that relationship without retaining the original account identifier in the evaluation copy. Store any mapping separately with stricter access, or avoid retaining a mapping when no future linkage is needed. Do not reuse a token across unrelated datasets by default; broad reuse can make linkage easier than the task requires.

Review free text separately from structured fields. Names, email addresses, phone numbers, account numbers, and addresses can appear in unexpected places: quoted replies, pasted signatures, logs, or tool output. Automated detection can help locate common patterns, but a reviewer should inspect transformed examples for residual identifiers and for damage to task meaning. A redaction pass that hides a reference to the very condition the agent must detect has made the case unusable.

Preserve decision-relevant relationships rather than original values wherever possible. If the task depends on whether a request came from the account holder, represent the relevant evidence as “verified” or “not verified” only if that accurately reflects the source and is sufficient for the scoring rule. If timing matters, preserve the relative ordering or a coarse time interval when exact timestamps are unnecessary. If a value’s format matters, use a synthetic value that retains the format while not resembling a real identifier.

Keep a redaction log with the case ID, fields transformed or removed, transformation category, reviewer status, and any known impact on interpretation. Do not copy the sensitive value into the log. A case that cannot be transformed without losing its meaning should be excluded or held for a separate approved review, not forced into the dataset with ambiguous labels.

Redaction can create its own sampling bias. For example, if complex cases contain more sensitive free text and are removed more often, the final set may overrepresent simple requests. Count exclusions by stratum and reason. If a category has a disproportionate exclusion rate, revise the collection method or state that coverage gap plainly.

Step 6: Create a case record and scoring contract

Give each case a stable, non-meaningful identifier such as case_0042. The ID should not encode a customer name, date of birth, region, or other source detail. Store content and metadata in a consistent schema so reviewers can filter and compare without reconstructing the dataset from scattered notes.

A practical record can contain the following fields:

case_id: case_0042
case_version: 1
source_cohort: support_2026_q2
split: holdout
primary_stratum: missing_information
secondary_tags: [account_access]
input: [minimized task content]
permitted_context: [facts the agent may use]
expected_outcome: [observable acceptable result]
prohibited_outcomes: [actions or claims that fail]
score_dimensions: [outcome, policy_boundary, communication]
label_rationale: [brief explanation]
redaction_status: reviewed
review_status: adjudicated

These are proposed fields, not a claim about a required platform format. Adapt names to the tools used by the team, but preserve the distinctions between the user’s input, permitted context, expected outcome, and scoring evidence. If the agent can call tools, describe the relevant initial state and the state change that would count as success. Do not make a transcript alone stand in for an external effect.

Write scoring criteria so a reviewer can apply them without guessing the author’s intent. For each dimension, specify what counts as pass, fail, or not scorable. Separate outcome correctness from communication quality where they can differ. For instance, an agent may give a clear explanation but fail to complete the permitted action; another may reach the correct state while making an unsupported claim about what it did. A single undifferentiated “good answer” label hides that distinction.

Include an acceptable-variation field for cases with more than one safe resolution. It should identify what can vary and what must remain true. Do not enumerate every possible wording. The purpose is to protect legitimate alternatives from inconsistent scoring while retaining a stable standard for the consequential decision.

Step 7: Label, adjudicate, and check the dataset

Have a domain reviewer label the expected outcome and explain the evidence that supports it. When practical, ask a second reviewer to independently assess a subset that includes ambiguous and consequential cases. Compare disagreements by cause: unclear case context, unclear scoring rule, reviewer interpretation, or genuinely different acceptable outcomes. Each cause calls for a different correction.

Do not resolve disagreement by averaging labels or choosing whichever answer favors the current system. Revise the case contract or scoring guide where needed, then document the adjudication. If the source record cannot establish the expected outcome, mark the case unscorable or exclude it. A confidently stated label is not more trustworthy merely because it is decisive.

Check coverage and quality before sealing the holdout. Confirm that every case has a primary stratum, split, source cohort, usable input, expected outcome, scoring rule, and redaction status. Review duplicates and near-duplicates, check that linked records remain on one split where required, and compare inclusion and exclusion counts across strata. Inspect a sample of transformed cases for both residual sensitive content and accidental loss of task meaning.

A flow for deciding what to do with a candidate record:

flowchart TD
  accDescr: Workflow stages and decisions: Candidate record, Fits task contract?, Exclude and log reason, Can meaning be minimized?, Assign stratum and group, Split before tuning, Label and review. The adjacent text explains the conditions and exceptions.
  accTitle: Build a Private Agent Evaluation Set — Sampling and Holdout Design workflow
    A["Candidate record"] --> B["Fits task contract?"]
    B -->|"No"| C["Exclude and log reason"]
    B -->|"Yes"| D["Can meaning be minimized?"]
    D -->|"No"| C
    D -->|"Yes"| E["Assign stratum and group"]
    E --> F["Split before tuning"]
    F --> G["Label and review"]

The first decision prevents unrelated work from entering the set simply because it is available. The second prevents a privacy transformation from erasing the evidence required for scoring. Records that fail either test are excluded with a reason, allowing the team to see whether exclusions cluster in an important category. Grouping related records before the split reduces leakage; labeling and review come after the split so case selection is not quietly shaped by a model’s performance on those examples.

Step 8: Version the set and control its use

Version the dataset whenever a change could affect comparability: new cases, changed labels, revised redactions, altered strata, a different sampling window, or a rewritten scoring contract. Keep the previous version’s manifest and change log so an apparent performance shift can be traced to the system, the data, or the scoring rules. Do not overwrite a case in place when the change alters its meaning.

Give each evaluation run a dataset version and candidate identifier. Record the harness and relevant configuration alongside the result, including which cases were skipped and why. This article does not prescribe a particular harness; the point is to make the evidence interpretable later. A score with no dataset version, candidate version, or missing-case record cannot reliably answer whether a change improved the intended behavior.

Retire a holdout when its examples have been repeatedly exposed or when a material workflow change makes its labels stale. Preserve the historical record according to the organization’s retention plan, but do not keep calling a known development resource an untouched holdout. Plan the next sample from a new eligible source period or a previously unused pool. If the workflow changes substantially, first revisit the case contract and strata instead of merely adding a few new examples to an outdated set.

The dataset can also become misleading through operational drift. A policy update may change which action is permitted; a new tool may change the evidence available to the agent; a source system may alter how it records outcomes. Assign an owner to review these triggers and mark affected cases as needing revalidation. Revalidation does not mean every historical case must be discarded, but it does mean the team should not assume old labels remain correct.

Do not use this set to answer questions it was not designed for. A stratified holdout can help uncover failure patterns across task types, but it does not automatically establish population failure rates, generalize to new regions, or predict performance on unrelated tasks. For public model-release comparisons and the limits of external benchmark evidence, see Benchmarks That Actually Matter for New Model Releases. If the team is also choosing how much reasoning effort to permit or whether to use repeated attempts, those are separate design questions; see Choosing Reasoning Effort: When More Thinking Costs More Than It Saves and Pass@1 vs Best-of-N: Which AI Benchmark Matches Your Workflow?. Task-specific conclusions should not be inferred from a general benchmark; the limits of common software and computer-use benchmarks are discussed in What SWE-bench, Terminal-Bench and OSWorld Can—and Cannot—Tell You.

Worked example: account-access requests

Consider a hypothetical mid-market software company preparing a private set for an agent that handles account-access requests. The intended decision is narrow: whether a revised agent can identify when a self-service recovery step is appropriate, request missing prerequisites when needed, and avoid taking an account action when the evidence does not support it. The example is illustrative; its categories and counts are assumptions, not reported results.

The team defines one case as a minimized request, the relevant account-state facts available at decision time, and an expected outcome. It identifies five strata: eligible routine recovery, missing verification, conflicting account context, escalation required, and a request outside the permitted workflow. Each receives a written inclusion rule. “Conflicting account context” means the case contains two materially inconsistent relevant facts, not merely that the conversation is long.

For an initial 120-case plan, the team sets aside 60 development cases and 60 holdout cases. It targets coverage in all five strata in each split, with more cases for routine work because that category dominates the source queue. It does not claim the resulting proportions represent the production mix. A separate frequency-weighted sample would be needed for a different question about typical incoming work.

Before splitting, the team groups records that refer to the same incident or account recovery episode. It then assigns groups using a recorded selection rule, so a sequence of follow-up messages cannot appear on both sides as near-duplicates. If a group is too large to fit the intended allocation, the team records the exception and makes the choice before inspecting agent outputs.

The source includes names, email addresses, account identifiers, and timestamps. The evaluation copy replaces account identifiers with set-scoped tokens where continuity matters, removes names and email addresses, and retains only a coarse ordering of events. A case involving an unverified requester keeps that task-relevant fact, but not the original identity evidence. Reviewers check whether the transformation preserves the distinction between verified and unverified context.

For each case, the expected outcome states what the agent may do, what it must not do, and what counts as an acceptable handoff. Reviewers separately score whether the outcome is correct, whether the agent respected the action boundary, and whether its explanation is accurate. If the agent claims that it completed a recovery step, the evaluation record must distinguish the claim from evidence that the relevant state actually changed.

Suppose a later candidate handles routine cases well but repeatedly treats a missing verification field as if verification had succeeded. The holdout has surfaced a decision-boundary weakness, but it does not tell the team how common that failure is in the live queue because the strata were deliberately balanced. The team can investigate the failure pattern, revise the candidate using development cases and other evidence, and then decide whether the exposed holdout should be retired before a confirmatory evaluation.

Now consider the counterexample: the team samples only resolved cases with complete transcripts and copies the historical support agent’s final answer as the expected result. This looks efficient, but it can exclude abandoned requests, missing-context cases, and requests that were mishandled historically. It may also reward reproducing a past response rather than meeting the intended workflow standard. The dataset would be easy to assemble and difficult to trust. A smaller, deliberately stratified sample with explicit exclusions is more informative for the stated decision.

Printable dataset worksheet

Use this worksheet as a record to complete before assembling cases. “Unknown” is a valid entry when the team cannot establish a fact; it is more useful than an unmarked assumption. Keep sensitive source content out of this planning document unless there is a specific, approved need for it.

  • Decision: What concrete product or operational decision will this set inform?
  • Task boundary: What work is in scope, and what neighboring work is excluded?
  • Case unit: What must one record contain to be independently scorable?
  • Population: Which source, date range, workflow, and inclusion rules define eligible work?
  • Known gaps: Which teams, languages, channels, or task types are absent or underrepresented?
  • Strata: What primary categories correspond to distinct decisions the agent must make?
  • Assignment rules: Can reviewers assign cases to strata consistently? What precedence rule resolves overlap?
  • Sampling method: Is the sample proportionate, stratified, or a labeled challenge set? Why does that suit the decision?
  • Counts: What target is assigned to each stratum and split, and what limitations follow from those counts?
  • Grouping: Which records are related closely enough that they must remain on one side of the split?
  • Holdout boundary: Who can inspect holdout content, and what event will trigger retirement or replacement?
  • Redaction: Which fields are removed, tokenized, generalized, or retained, and why is each retained field necessary?
  • Redaction review: Who checks for residual identifiers and for loss of task meaning?
  • Expected outcomes: What observable result counts as acceptable, and what alternatives are allowed?
  • Scoring: Which dimensions are scored separately, and what qualifies as fail or not scorable?
  • Adjudication: How are ambiguous labels resolved, and how are disagreements recorded?
  • Versioning: What changes create a new dataset version, and where are manifests and change logs kept?
  • Evaluation record: Which candidate, harness configuration, dataset version, skipped cases, and run date are recorded?
  • Review trigger: Which workflow, policy, source-system, or tool changes require label revalidation?
  • Limit statement: What conclusions is this set explicitly not intended to support?

A completed worksheet should leave a reviewer able to explain why each case was eligible, what decision it tests, how its holdout status was protected, and what the score can reasonably mean. If those answers depend on undocumented knowledge held by one person, the set is not yet reproducible.

Summary and next step

Build the set around a specific decision, define the case contract, and create strata that reflect distinct agent choices. Sample with a documented method, group related records before splitting, and keep development cases separate from a private holdout before tuning begins. Minimize records without removing the facts needed to score them; record exclusions and redaction effects rather than hiding gaps.

Treat labels, scoring rules, dataset versions, and holdout exposure as part of the artifact—not afterthoughts. Use the set to diagnose the behaviors it covers, and do not turn a deliberately enriched sample into a claim about typical production frequency. When workflow changes or holdout examples have been repeatedly exposed, review labels or retire the set for confirmatory use.

Teams that need help turning a workflow decision into a sampling plan, case schema, and review process can explore AI evaluation and strategy. The useful deliverable is not simply a folder of examples; it is a reproducible account of what those examples test, how they were prepared, and where their conclusions stop.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call