A model switch is not a single prompt comparison. It changes a component inside a working system: prompts, tool definitions, application logic, output parsing, retries and the business process that consumes the result. A regression suite makes those dependencies visible before a new model becomes the default.
This checklist focuses on migration gates for prompts, tools and representative tasks. It is for teams moving an existing workload to Sonnet 5.5, not for declaring a universal model winner. Anthropic announced Sonnet 5.5 on September 28, 2026, positioning it as a lower-cost, faster model for scoped work; that positioning is a reason to evaluate a fit, not evidence of a production result for your workload (announcement).
Use the checklist to build a controlled decision: freeze a baseline, replay representative cases, inspect behavior as well as final outcomes, and set explicit conditions for a limited release. A green result should mean the candidate passed criteria you set in advance—not that a sample looked persuasive or that a model label sounds appropriate.
1. Set the migration boundary
Start by naming the application path under review. A team may use one model across several workflows, but a migration decision should have a bounded unit: for example, “draft a support reply from a ticket and account record,” rather than “customer service.” Boundaries make it possible to assemble relevant cases and identify who can judge the result.
-
Name the workflow and its business output. Write one sentence describing the input, the model-assisted action and the output consumed downstream. “Summarize a ticket” is incomplete if the summary is later used to issue a refund; describe the actual path far enough to identify consequential decisions.
Record where the model’s work ends and where ordinary application logic or a person takes over. This boundary determines what the regression suite can assess. A model that produces a plausible summary has not necessarily produced a valid refund decision, and a valid structured response does not prove that a downstream action was appropriate.
-
List the production path, not just the prompt. Record the system and user instructions, inserted context, tool schemas, tool execution code, output validation, retry rules and any transformation before the result reaches a user or another service.
For each component, note its owner and current version or revision. A prompt snapshot without its tool schema and parser is not a reproducible baseline. If a request passes through several stages, sketch the actual order, including conditions that cause another model call or a fallback.
-
Choose the migration unit. Decide whether the candidate replaces a model for one workflow, one application route or a broader deployment. Keep the initial boundary as narrow as is useful for a meaningful evaluation.
A narrow migration unit reduces the number of changes whose effects must be interpreted. It also limits what a rollback must restore. Do not quietly include unrelated prompt cleanup or a new tool in the model-switch change; first establish whether the candidate works with the current behavior, then consider other changes separately.
-
Name an accountable decision owner and reviewers. Identify who can accept changes in task quality, who understands the application’s technical path, and who can stop a rollout if an operational failure appears.
Reviewers should know what evidence they are assessing. A domain reviewer can judge whether a proposed answer is useful; an engineer can determine whether the tool payload was valid; an operator can tell whether latency or failures make the workflow unreliable. These are different judgments, even when the same person holds more than one role.
-
Set the evaluation window and evidence date. Put a date on the suite and record when cases, prompts, tools and model identifiers were frozen. If the evaluation is rerun after any of them changes, mark it as a new run rather than silently replacing the previous result.
A date helps distinguish an actual comparison from an old approval reused after the application has changed. The public model name is not enough for reproducibility: capture the exact identifier selected in the deployment configuration and preserve it with the run record.
flowchart TD
accTitle: Evaluate a Sonnet migration before switching
accDescr: Use the same cases and scoring rules for both runs. A better average cannot compensate for a newly unacceptable failure in a critical case.
N0["Freeze task and grader versions"] --> N1["Run current model baseline"]
N1["Run current model baseline"] --> N2["Run Sonnet 5.5 candidate"]
N2["Run Sonnet 5.5 candidate"] --> N3["Review paired regressions"]
N3["Review paired regressions"] --> N4["Check budget and latency"]
N4["Check budget and latency"] --> N5["Approve bounded rollout"]
Use the same cases and scoring rules for both runs. A better average cannot compensate for a newly unacceptable failure in a critical case.
2. Freeze a trustworthy baseline
A baseline is the version of the existing system against which the candidate is assessed. It should be a known, repeatable configuration, not a recollection of how the incumbent model “usually” behaves. Without it, differences can be attributed to the wrong cause.
-
Capture the incumbent model identifier and request settings. Copy the exact configured model identifier from the environment being evaluated, along with the settings that affect its responses. Record which settings are fixed and which vary by request.
Do not substitute a marketing name for the identifier actually used by the application. If staging and production use different identifiers or settings, state which configuration the comparison represents. Keep secrets out of the evaluation record; identifiers and non-sensitive settings are generally sufficient to reproduce the decision context.
-
Version the prompts and relevant application code. Store the complete prompt material and a code revision or release reference for formatting, parsing and post-processing. Preserve the context assembly rules, including which fields can be absent or truncated.
Text that appears minor can change meaning: an instruction reordered near the end, a new field inserted into context, or a parser that normalizes empty values can affect output. A frozen baseline lets the team attribute observed changes to the model candidate rather than an unnoticed application edit.
-
Capture tool contracts and side-effect boundaries. Save the tool names, descriptions, schemas, required fields, validation rules, execution conditions and responses made available to the model. Document whether a tool reads information, proposes an action or changes external state.
A tool-call regression can be more consequential than a wording difference. The suite should make visible whether the candidate selects the correct tool, supplies a valid payload, handles a rejected call and stops when the requested action is outside the workflow’s authority. Never treat a valid-looking tool request as proof that the external action succeeded.
-
Save representative inputs and expected evaluation evidence. Preserve privacy-appropriate input snapshots and the context required to interpret them. For each case, state what a satisfactory outcome must contain or avoid.
If the input includes changing records, preserve a stable test fixture or record the values used at run time. Otherwise, an apparent model difference may actually be a difference in the source data. Avoid storing unnecessary sensitive information in the suite; use redacted or synthetic fixtures when they still test the behavior at issue.
-
Run the incumbent against the frozen cases. Record its outputs, tool requests, errors and relevant timing or retry observations using the same harness planned for the candidate.
This run is not an endorsement of every baseline output. It gives the reviewers a reference and can expose cases where the current system already fails. Mark known baseline defects rather than quietly treating them as acceptable. A candidate should not receive credit for a supposed improvement if the old result was never captured or the scoring rule changed afterward.
-
Keep a change log for suite edits. Record additions, removals and scoring changes, including why each was made and whether results from prior runs remain comparable.
Editing cases after seeing candidate outputs can bias the decision, even when the editor intends to be fair. Where possible, define the case and acceptance rule before running the candidate. When a new failure reveals a missing case, add it for the next run and preserve the distinction between the original and expanded suite.
For guidance on choosing useful measures before a release evaluation, see which model benchmarks matter in practice. The migration suite below has a narrower job: test whether this application’s current path remains acceptable after a specific change.
3. Assemble cases that represent actual work
A regression suite is a compact, deliberate sample of the work the application really receives. It does not need to represent every possible request, but it should include ordinary cases, difficult cases and cases where the right behavior is to stop, clarify or decline an action. A set of easy demonstrations is not a migration gate.
-
Sample routine, high-volume tasks. Choose common requests that exercise the primary workflow with realistic variation in length, wording and supplied context.
These cases establish whether ordinary service remains useful. Do not let frequency alone determine importance: a routine case can still have a critical requirement, such as preserving a date, not inventing a policy condition or returning a field that downstream software requires.
-
Include edge cases that alter the correct response. Add missing fields, contradictory details, ambiguous requests, unusual formats and inputs near any known application limits.
The goal is not to collect oddities for their own sake. Each edge case should test a decision the workflow must make. For example, if an account identifier is absent, the expected behavior may be to request clarification rather than infer one from an unrelated field.
-
Include explicit non-action cases. Add examples where the safe or correct outcome is to ask for more information, state a limitation, or leave a tool unused.
Teams often measure whether the model completes requested actions and overlook whether it refrains when the prerequisites are missing. Record the desired stopping condition plainly. “No unsafe action” is too vague unless reviewers can identify what counts as an unsafe action in the particular workflow.
-
Represent tool paths separately. Include cases that should use each important tool, cases that should not call it, and cases where tool output is incomplete or reports an error.
For an action path, include the information necessary to assess both the request sent to the tool and the result returned to the application. A prompt-only answer cannot establish that a tool interaction is correct. If the system has several branches, ensure the suite covers the branches that would change the downstream outcome.
-
Preserve difficult cases rather than replacing them with averages. Keep examples that expose recurring ambiguity, important business rules or costly failure modes even if they are uncommon.
Summary scores can hide a small number of severe errors inside a larger set of acceptable responses. Label critical cases and define their pass rule separately, so the team can see whether the candidate regressed on them instead of letting a broad average wash out the result.
-
Ask a domain reviewer to check the case set. Have a person who understands the intended work confirm that the inputs are plausible and the expected outcomes reflect the real task.
Engineers can identify malformed fixtures, but a technically valid example may still misrepresent how the work is done. Conversely, a domain expert may miss a schema dependency. Review the case set from both perspectives before treating it as representative.
Case count is a design decision, not a badge of rigor. Expand cases where tasks vary meaningfully or failures carry different consequences; remove duplicates that test the same behavior without adding coverage. For a separate discussion of sizing comparisons, see how to decide how many test cases a model comparison needs.
4. Define gates before running the candidate
A gate converts a set of observations into a decision rule. It should be specific enough that two reviewers can apply it consistently, but it should not pretend that every useful output has one mechanically exact answer. Use rules suited to the workflow: exact checks for required fields, bounded human judgment for content quality, and explicit stop conditions for critical failures.
-
Separate hard failures from graded quality. List failures that automatically block a migration, such as malformed required output, an unauthorized tool request or omission of a mandatory business constraint.
Then define quality dimensions where a reviewer can distinguish strong, acceptable and unacceptable work. Keeping the two categories separate prevents a high score on style or completeness from compensating for a disqualifying failure.
-
Write an acceptance rule for each critical case. State the observable behavior that passes: required fields present, prohibited claims absent, correct tool selected, or a request for clarification when a necessary input is missing.
Prefer an outcome rule over a preferred phrasing. If multiple answers can be correct, identify the essential facts or decisions rather than requiring one reference sentence. That makes the evaluation more faithful to the task and less sensitive to harmless wording differences.
-
Define how human review will be conducted. Give reviewers the input, relevant context, candidate output and scoring guide. Where feasible, hide which model produced an output until scoring is complete.
Ask reviewers to record the reason for a low score or disagreement, not just a number. A score without an explanation can show that reviewers differed but cannot help the team tell whether the rubric is unclear, the case is ambiguous or the model output has a real defect.
-
Choose a minimum passing condition for each gate. State what must pass, what can be accepted with remediation, and what blocks rollout. Keep rules for critical cases distinct from aggregate measures.
Avoid inventing a threshold after seeing results to make a preferred candidate pass. If the appropriate threshold is uncertain, record that uncertainty and have the decision owner resolve it before the candidate run. A defensible threshold follows the workflow’s tolerance for failure, not a universal score borrowed from another task.
-
Set an allowed-change rule. Define whether improvements on one case can compensate for a small degradation elsewhere, and prohibit compensation for specified hard failures.
This is especially important when a migration changes response style or behavior in ways reviewers value differently. Make the trade-off visible in advance. If the team cannot agree whether a behavior is an acceptable change, that is a decision to resolve—not a reason to hide the disagreement inside an average.
-
Record what the gate does not prove. State the suite’s boundaries: which workflows, languages, data shapes, operating conditions or tools it does not cover.
This is not boilerplate. It prevents a pass on one bounded workflow from being used as approval for unrelated routes. A narrow, well-understood decision is more useful than a broad claim that the model is “regression-free.”
Agent evaluation should inspect task outcomes as well as transcripts; the agent harness that performs work and the harness that evaluates it serve different roles (evaluation design for AI agents). For a migration, that means reviewing what the application accomplished, not only whether a response sounded convincing.
5. Run the same path for both models
The comparison is meaningful only if the candidate and incumbent receive equivalent cases through equivalent application paths. If one run gets different context, a revised prompt or a different tool response, the observed difference cannot be cleanly attributed to the model change.
-
Freeze one request fixture for each case. Use the same input, assembled context and tool results for the incumbent and candidate wherever the workflow permits.
If a case depends on external data, fix that data for the test or record the exact snapshot. Do not rerun a changing live query for one model and assume its response is comparable to the earlier one. Where fixed responses would conceal behavior the team needs to test, design a separate controlled case for that interaction.
-
Hold prompts, tools and application code constant for the first pass. Change the model selection only, unless a necessary compatibility adjustment is documented as a separate change.
This isolates the question “does the candidate work with the current implementation?” If the prompt must change to support the candidate, save that version and run an additional comparison that labels both changes. Otherwise, a good result cannot tell whether the model or the prompt rewrite produced it.
-
Record complete interaction traces needed for review. Preserve the request, relevant response, tool-call proposal, tool result, validation outcome and final application output, subject to the organization’s data-handling rules.
A transcript may omit a failed parser, a rejected tool payload or an application fallback. Capture enough of the path to explain the final result, but do not collect unrelated data merely because it is available. The trace should let a reviewer reconstruct the decision without pretending that one visible assistant message represents the whole system.
-
Repeat cases when variability affects the decision. Identify cases where different valid outputs or intermittent behavior could change the pass/fail result, and run enough repeat observations to understand that risk.
Repeats are most valuable for cases near a gate or involving a sensitive branch. They consume review effort, so do not repeat every case automatically without a reason. Record the repeat policy in advance and avoid selecting only favorable runs for the final summary.
-
Review candidate failures against the baseline. Classify each difference as a candidate regression, an existing baseline defect, an acceptable change or an unresolved judgment.
Keep the original outputs beside the decision. A candidate’s different wording may be harmless, while a seemingly small field change may break a parser. Classify by effect on the workflow, not by whether the output resembles the incumbent’s phrasing.
-
Track execution failures separately from task failures. Record timeouts, request errors, retries, malformed responses, tool failures and harness problems as distinct observations.
A broken test harness must not be scored as a model failure, and a retry that eventually returns an answer must not erase the initial operational event. Make the failure timeline visible: first request, error, retry or fallback, and final disposition. This helps distinguish quality from reliability without collapsing either into a misleading success count.
6. Evaluate prompts and tools as migration gates
Prompts and tools are where a model change often meets application assumptions. A regression can occur even when the final text looks plausible: a model may choose a different route, populate a field differently or infer that an action is ready when required information is missing.
-
Check instruction coverage, not prompt similarity. For each material instruction, identify a test case that demonstrates whether the intended behavior occurred.
A candidate does not need to follow the incumbent’s phrasing pattern. The question is whether it respects the operational constraint. If the prompt requires a specific ordering, a refusal boundary or a distinction between known facts and assumptions, test that behavior directly instead of judging the output by stylistic resemblance.
-
Validate structured outputs at the application boundary. Test required keys, types, allowed values, empty fields, extra fields and malformed outputs using the same parser and validator as the application.
Record both the model response and whether the application accepted it. If the application silently repairs a response, note the repair. A migration can appear successful in a notebook but fail in production code that expects a stricter structure.
-
Test tool selection and payload semantics. For each action case, assess whether the right tool was selected and whether the payload expresses the user’s intent accurately.
Validate meaning, not just schema. A payload can be syntactically valid and still contain the wrong account, amount, date or requested operation. Where a tool changes external state, test against a controlled environment or a safe non-executing harness appropriate to the application; do not infer a real-world outcome from the model’s claim that it performed the action.
-
Require a clear boundary between proposal and execution. Confirm that application logic—not unreviewed model language—determines whether an external action is allowed to proceed.
When human approval is part of the workflow, the approval must apply to the exact action payload and expire rather than remain reusable after that payload changes. Re-checking or replacing a payload should require a fresh decision. This design reduces the risk that an apparently small model change alters what gets executed after an earlier review.
-
Test tool errors and partial results. Feed cases with rejected requests, missing fields, stale information or an unavailable result, and verify that the application does not present an unsupported success claim.
The desired recovery may be to ask for help, retry under a defined policy or return a limited result. Specify which behavior is expected. Tool failure should be visible to the user or operator in the way the workflow requires, not hidden behind confident prose.
-
Check that retries do not duplicate external effects. Inspect how the application behaves if a request fails after a tool may have acted, or if the response is lost before the application can confirm the outcome.
Do not assume that repeating an action is harmless. The application needs a way to verify business state before deciding what to do next. A migration gate should reject any design in which a retry can accidentally repeat a consequential action without checking whether it already occurred.
7. Work through one hypothetical migration
Consider a hypothetical mid-market operations team that uses a model to prepare a service-credit recommendation from a case record. The application can look up the account, retrieve the relevant service event and submit a recommendation for a staff member to review. The migration question is not whether a response sounds polished; it is whether the new configuration preserves the correct recommendation, uses the tools appropriately and leaves the final action under the team’s established process.
The team freezes the current prompt, account-lookup and event-retrieval tool contracts, parser, and a set of case fixtures. The cases include a normal record, a missing event date, two conflicting records and a case outside the policy range. Those inputs are illustrative. They are not PADISO client data or a claim about measured model behavior.
Before running the candidate, the team writes down the expected outcome for each case. The normal case must preserve the verified event date and produce the required recommendation fields. The missing-date case must not infer a date. Conflicting records must be surfaced rather than silently resolved. The out-of-range case must not prepare an unsupported credit recommendation. The team also says what a valid tool payload looks like and what the application must do when a lookup returns no result.
Reviewers then run both configurations through the same fixtures and record the complete path. Suppose the candidate produces well-formed output in the normal case but selects a tool for the missing-date case that cannot establish the date. That is a failure of the migration gate even if the final prose says “date unavailable.” The request and application trace reveal whether the system actually respected the missing information.
Suppose a second case produces a sensible recommendation but omits a field required by the application parser. The team records an application-boundary failure rather than accepting the answer on content alone. It can then choose to reject the migration, or make a separately versioned compatibility change and rerun the affected suite. It must not quietly modify the parser, count the repaired result as an unchanged-model pass and lose track of the cause.
Now consider a counterexample: a reviewer prefers the candidate’s concise wording in several routine cases and proposes accepting it despite a failed conflicting-record case. The gate should block that trade if the policy requires conflicts to be surfaced. A style improvement cannot compensate for a failure on a case the team declared critical. If the team believes that case no longer reflects the real task, it must revise the case and rule with a recorded reason before using the result to decide.
The decision record should contain the candidate’s exact configured identifier, incumbent identifier, run date, fixture revision, prompt and tool versions, observations and unresolved issues. This makes the outcome reviewable later and prevents a passing result from being casually applied to a different workflow.
8. Decide whether the migration passes
A migration decision should be an explicit disposition of evidence: proceed to a bounded release, remediate and rerun, or stop. “Promising” is not a release state. The decision should identify what passed, what failed, what remains unknown and what the next action is.
-
Apply critical gates first. Check all predeclared hard failures before considering aggregate quality or reviewer preference.
If a critical case failed, record the failure and block progression under the stated rule. Do not conceal it in an average or remove the case after the result. If the failure exposes a flawed case or rule, document that analysis and create a new run against the corrected suite.
-
Summarize results by behavior and workflow branch. Report outcomes for routine tasks, edge cases, tool paths, non-action cases and operational failures separately where those distinctions matter.
A single pass percentage can hide exactly the failure the team needs to see. Include counts and denominators if you calculate rates, explain how repeats are handled and preserve notable examples. Do not report more precision than the suite supports.
-
Resolve reviewer disagreements openly. Keep disagreements attached to the relevant case and ask whether they arise from the rubric, task interpretation or a real trade-off.
Where no shared rule exists, do not manufacture consensus by averaging scores. Assign the decision to the appropriate accountable owner, document the rationale and consider adding a clearer case or rubric for future runs. Unresolved disagreement on a critical behavior is a reason to pause, not a cosmetic note.
-
Choose one disposition and name its conditions. Mark the candidate as ready for a limited release, blocked pending remediation, or not selected for this workflow.
A limited release should state its scope, what observations will be monitored and what event triggers a pause or rollback. A blocked disposition should name the defect and the evidence required to return for review. A not-selected result can still be useful: preserve why it failed so the same question is not reopened without new evidence.
-
Record the rollback target and method. Identify the configuration to restore, the person or process authorized to initiate that change and how the application will confirm that the rollback took effect.
Rollback should restore a known model configuration and its compatible prompt, tools and application behavior—not just switch a label while leaving a changed parser in place. If the release involved several coordinated changes, record their dependency order and how to reverse them safely.
-
Set a review date for the decision. State when the team will revisit the result, especially if the candidate’s prompt, model identifier, tools, task mix or application path changes.
A regression result is evidence for a particular configuration and scope. It does not automatically transfer to a new route or a changed task. A review date prevents an old decision from becoming an informal permanent approval.
If the team needs help defining model selection criteria or connecting evaluation results to a deployment decision, AI strategy and model selection is an appropriate next step. For separate release-level comparison and routing questions, consult a practical comparison of Astra, Sonnet 5.5 and Opus 5.5 and how teams can divide work between Sonnet 5.5 and Opus 5.5; those are distinct decisions from whether this migration passes its own gates.
9. Roll out only within the tested boundary
A passing suite supports a controlled next step; it does not justify an unbounded replacement. Keep the release aligned with what the team tested, and watch for evidence that the live path differs from the fixtures.
-
Start with an explicitly bounded deployment. State which workflow, route or user group is included and which remains on the incumbent configuration.
A limited rollout is useful only if the application can identify which configuration handled a request and the team can respond to a failure. If routing is managed outside the model host, document where the decision is made and how the selected configuration is recorded. Do not assume a hosted endpoint itself provides traffic splitting.
-
Define live signals that correspond to suite gates. Monitor the operational and business outcomes that matter to this workflow, such as structured-output rejection, tool errors, retry patterns or review corrections.
A signal is useful when someone knows what it means and what action follows. Define the observation window and response owner before release. Avoid treating raw model response time or a general satisfaction measure as a substitute for a critical task outcome.
-
Set a stop condition and rehearse the response. State which failure, pattern or operational condition pauses the candidate route, and confirm that the team can restore the previous configuration.
The stop condition should be actionable: identify the signal, decision owner and rollback action. For a consequential tool path, include how the team checks the underlying business state before retrying or reverting an operation that may already have occurred.
-
Compare production cases with the evaluation assumptions. Check whether live inputs, missing fields, tool responses and user behavior remain within the conditions represented by the suite.
A suite can be well designed and still miss a change in the incoming work. Add material new patterns to the next evaluation revision. Do not use a smooth initial rollout as proof that unrepresented cases are safe.
-
Close the loop on incidents and near misses. Preserve the relevant trace, classify the failure, decide whether it changes the migration decision and add a useful test case when appropriate.
Update the suite with an explanation rather than simply accumulating cases. A new case should test a distinct behavior or a newly important failure mode. Keep the previous run intact so the team can tell whether the candidate’s status changed because the software changed, the evidence changed or the decision rule changed.
10. Versioned decision worksheet and printable summary
Use the following worksheet inside the change record. It is deliberately a decision artifact, not a claim that any model has passed a test. Complete identifiers from the actual deployment configuration before running the suite, and preserve the finished record with the evaluation traces.
| Field | Record for this migration |
|---|---|
| Worksheet version | 1.0 |
| Prepared / evidence checked | 2026-09-30 |
| Test execution date | Not run in this article; enter the actual run date |
| Workflow and migration boundary | Enter the specific application path and excluded routes |
| Incumbent model name and exact configured ID | Enter both from the environment under test |
| Candidate model name | Claude Sonnet 5.5 |
| Candidate exact configured ID | Enter the exact identifier selected in the deployment configuration; not specified here |
| Prompt, tools and application revision | Enter immutable versions or commit references |
| Case-set revision and case count | Enter revision, total cases and relevant category counts |
| Hard gates | List each blocking condition and its observable pass rule |
| Outcome and operational observations | Record by task category; include failures and denominator |
| Decision | Limited release / remediate and rerun / not selected |
| Decision owner and reviewers | Enter names or role identifiers in the internal record |
| Scope, stop condition and rollback target | Record route, trigger, owner and restoration method |
| Next review date | Enter date or change event that reopens the decision |
The candidate model name above is a public-facing name, not a substitute for the configured identifier. The worksheet is prepared as of September 30, 2026, and does not state that a test was run on that date. Enter the actual run date after execution; if the exact identifier cannot be captured, do not describe the comparison as reproducible.
Printable pre-switch summary
- The workflow and migration boundary are specific.
- Incumbent and candidate identifiers, prompt versions, tool contracts and application revision are recorded.
- The frozen case set includes routine tasks, edge cases, non-action cases and material tool paths.
- Expected outcomes and hard gates were written before the candidate run.
- Both configurations received equivalent fixtures through the same evaluation path.
- Outputs, tool requests, tool results, parser outcomes and operational failures are reviewable.
- Critical failures and disagreements have an explicit disposition; aggregate scores do not mask them.
- A bounded release scope, stop condition, rollback target and review date are recorded.
A checked list is not the decision by itself. The decision is the completed record: what was tested, what the candidate did, which rules applied and why the chosen next step follows from that evidence. If one of those parts is missing, pause before switching the default.