SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
How-to Guide 18 mins

Astra vs Sonnet 5.5 vs Opus 5.5: How to Run Your Own Comparison

A reproducible protocol for comparing Astra, Sonnet 5.5 and Opus 5.5 using your own tasks, measured results and a versioned decision record.

The PADISO Team ·

Prerequisites

Before comparing Astra, Sonnet 5.5 and Opus 5.5, make sure the comparison will answer a real decision. A model ranking without a defined use case is hard to interpret: the best result on a synthetic puzzle may say little about customer support, code review or document extraction in your business.

You will need access to each model through the environment you expect to use, a set of representative tasks, an agreed scoring method and a way to record model identifiers and run settings. Confirm access and permitted use with your provider before scheduling the work. Do not assume that similarly named models, endpoints or settings behave identically across environments.

Choose a decision owner who can resolve tradeoffs among quality, latency, cost and operational effort. Engineering can measure response time and build the harness; the business owner should define what counts as a useful outcome. Keep a test operator separate from the person judging ambiguous outputs where practical, so expectations about a favored model do not quietly influence scores.

Prepare a place to store prompts, input versions, outputs, annotations, failures and run metadata. If the test uses sensitive material, use data approved for that environment and minimize what you retain. This guide is about comparison design, not a claim that any model or platform is suitable for a particular data classification.

The model labels refer to distinct offerings, not a pre-existing ranking. OpenAI identifies GPT-6 Astra as an OpenAI model and distinguishes model capability from Codex permissions and execution controls (Astra). Anthropic announced Sonnet 5.5 on September 28, 2026, positioning it as a lower-cost, faster option for scoped work (Sonnet 5.5). Anthropic presents Opus 5.5 for complex work; an architecture assistant cannot make business tradeoffs on your behalf (Opus 5.5). These are vendor descriptions, not results from this protocol.

Warning: Do not publish a winner before running the comparison. A plan, a vendor description or an unexecuted worksheet is not a benchmark.

1. Define the decision before choosing test cases

Write one sentence that states what the comparison will help you decide. For example: “Which model should draft a support response from an approved account history, subject to a human review?” This is more useful than “Which model is best?” because it names the work and the intended role of the output.

Then specify the decision boundary. Are you selecting one model for a workflow, deciding whether to route different task types to different models, or checking whether a proposed change is safe enough to evaluate further? Those are different decisions. A three-model comparison can inform them, but it should not quietly become a claim that one model is universally superior.

Describe the workflow around the model, not just the prompt. Record what information the model receives, the output format, who reviews the result, what actions follow, and where a human can stop the process. This prevents a candidate from receiving credit for a polished answer that cannot be used in the real process.

Choose the outcomes that matter before seeing results. For a support draft, that might include factual correctness, policy adherence, completeness, edit effort and time to a reviewable draft. For an extraction task, the outcomes might instead be field accuracy, missing-value handling and downstream validation. Keep the primary outcome narrow enough to drive a decision, then retain secondary measures to explain tradeoffs.

Set failure thresholds as well as targets. A response that invents a refund approval may be unacceptable even if it is otherwise fluent. Define such failures in task terms: what is wrong, why it matters and whether a human reviewer is expected to catch it. A score that averages away a severe failure can mislead a team into accepting a result it would reject in practice.

This is also the point to decide what will not be tested. If the question is model behavior on a fixed task, do not mix in a new user interface, revised retrieval pipeline and different human-review process in the same comparison. If the question is the whole workflow, state that broader scope and measure its components separately where possible.

For a broader view of how model choice fits an operating architecture, see Multi-Model Production Stacks in 2026: Routing Between Claude Opus 4.7 and GPT-5.5. Keep the present test focused on the comparison decision rather than turning it into a general routing design.

flowchart TD
  accTitle: Compare three candidate models under one protocol
  accDescr: The three candidates use the same workload definition, with supported settings and resource budgets recorded separately. The diagram describes a proposed experiment and does not declare a measured winner.
  A["Same held-out cases"] --> B["Astra candidate"]
  A --> C["Sonnet 5.5 candidate"]
  A --> D["Opus 5.5 candidate"]
  B --> E["Blind outcome grading"]
  C --> E
  D --> E
  E --> F["Workload-specific decision"]

The three candidates use the same workload definition, with supported settings and resource budgets recorded separately. The diagram describes a proposed experiment and does not declare a measured winner.

2. Freeze the comparison protocol

Write down the protocol before running candidate models. Include the task definition, input set, prompt text, any permitted context, output requirements, scoring rubric, exclusion rules, run date and the method for handling errors. Freeze a copy with a version identifier. If you alter a prompt or rubric after seeing results, treat the changed setup as a new test version.

Make the conditions as comparable as your environment allows. Use the same task inputs and the same intended instructions for each model. If a model requires a different representation or configuration, document the difference instead of pretending the conditions are identical. The goal is a fair, interpretable comparison, not artificial uniformity that prevents a candidate from being used correctly.

Record all settings that can affect the result and are available to you, including the exact model identifier returned by the environment, prompt version, any configured generation controls, input version, timestamp and relevant execution context. Do not substitute a friendly label for the exact identifier. A future reader needs to know what was actually tested, not just what the team called it.

Decide how repeated runs will work. If a task can produce variable outputs, determine in advance whether you will run it once or repeat it, how you will summarize variation and what counts as a failed attempt. Avoid rerunning only the outputs you dislike. That creates a selection bias that makes one candidate appear more consistent than it is.

Set a stopping rule. A practical rule might be to complete the preselected case set once, resolve only documented infrastructure failures, then score all valid outputs. If a provider outage or a harness defect interrupts the test, mark the affected run and resume according to the written rule. Do not silently replace an awkward output with a more favorable rerun.

Pro tip: Keep a read-only copy of the protocol and input manifest before the first run. If your team changes the evaluation after seeing an answer, preserve the original and label the new run rather than overwriting history.

For principles on selecting meaningful evaluation measures, read Benchmarks That Actually Matter for New Model Releases. This comparison still needs measures specific to your workflow; a general benchmark does not settle your operational decision.

3. Assemble representative cases and reference material

Build the test set from real work patterns, while respecting data-handling rules. A case can be anonymized or reconstructed if raw production content is unsuitable. Preserve the features that make the task difficult: conflicting facts, missing fields, unusual formatting, ambiguous requests or the need to decline an unsupported conclusion.

For each case, record the expected properties of a useful answer. Avoid writing a single “gold” response if several formulations could be correct. Instead, define required facts, forbidden claims, required actions or fields, and acceptable ways to express uncertainty. This gives reviewers room to judge meaning without rewarding one model merely for matching a preferred sentence.

Include ordinary cases and meaningful edge cases. An evaluation made only of rare disasters can overstate risk; one made only of clean, easy examples can conceal it. The mix should reflect the decision’s intended use and the consequences of mistakes. Document why each case is included so you can explain what the test set represents and what it does not.

Use the same case identifiers across candidates. Store a stable input hash or another version reference if the source material may change. If you revise a case because the original was ambiguous or mislabeled, note the revision and rerun the affected candidates under the revised protocol. Do not change a reference answer only for the model whose output exposed the problem.

Blind reviewers to candidate identity where feasible. Replace model names with neutral labels and randomize output order. Reviewers may still infer a candidate from style, so blinding is not perfect; it is a way to reduce expectation effects, not a guarantee of impartiality. Ask reviewers to score against the rubric and add a short rationale for borderline judgments.

Keep test-set size in proportion to the decision. A small exploratory set can reveal obvious workflow mismatches, but it should not support a broad claim. A higher-stakes migration or broad deployment decision usually needs more coverage, repeated assessment and separate failure analysis. See How Many Test Cases Are Enough for an AI Model Comparison? for a deeper treatment of sample-size choices; this guide focuses on how to run and report the comparison.

4. Run candidates under controlled conditions

Run the frozen cases through each candidate using the intended environment and workflow. Use the same input version and prompt version, and capture the complete output needed for review. Store outputs against a case identifier and candidate identifier, not in a spreadsheet whose rows can be detached from their settings.

Separate model output from orchestration behavior. If a wrapper truncates a response, transforms the prompt or fails to deliver an output, record that as a harness or execution event. The comparison may need to evaluate the full path, but the team should be able to tell whether a result came from the model, the surrounding software or the test operator.

Measure elapsed time consistently. Define the start and end points: for example, request submission through receipt of a complete response. If you also care about human review time, measure it separately. A quick response that requires extensive correction may not shorten the workflow, while a slower response that is immediately usable may be preferable in a low-volume process.

Track failures rather than dropping them. Distinguish a malformed answer, an incomplete answer, a timeout, an infrastructure error and a task the model could not complete. The distinction matters when interpreting a result: an infrastructure error may call for a rerun under a predeclared rule, while an incorrect answer is part of the candidate’s observed behavior in that test.

Avoid accidental prompt coaching. If an operator edits instructions for one candidate after seeing an answer, that candidate has received a different test. If adaptation is part of the real workflow, design it explicitly, give every candidate equivalent opportunities and count the additional human effort. Otherwise, use the frozen prompt and record where it fails.

A comparison run should be traceable from a reported score back to the individual output and the rubric judgment. Keep the mapping among case, candidate, prompt version, run timestamp, output and reviewer decision intact. Without that chain, a surprising aggregate cannot be investigated and a published result cannot be reproduced.

5. Score quality and operational fit separately

Use a rubric with observable criteria. For a draft response, “good” is too vague. A reviewer might assess whether the answer uses only supported account facts, answers the customer’s request, follows a specified policy, avoids unauthorized commitments and is ready for review. Define what earns each score and what constitutes a critical failure.

Keep factual correctness distinct from style. A polished but unsupported statement should not outscore a plain, accurate answer when correctness is the business requirement. Likewise, if tone is important, score it separately rather than allowing it to obscure whether the requested work was completed.

Report the distribution of results, not just an average. Show how many cases passed, needed material correction or failed a critical criterion. Note repeated failure patterns and provide examples with sensitive details removed. An overall average can conceal a small group of unacceptable errors or large differences among task types.

Treat human review as part of the cost and quality picture. Record whether the reviewer accepted, lightly edited, substantially rewrote or rejected the output. If reviewers disagree, preserve the disagreement and adjudicate according to your prewritten rule. A difference of opinion may signal an unclear rubric or a genuinely ambiguous task rather than a simple scoring mistake.

Measure latency and operational effort using consistent definitions. Where relevant, distinguish time to first response from time to complete response, and model waiting time from human work. Capture configuration and setup effort as observations, not as universal product properties. Do not turn a brief test into unsupported claims about future capacity, total cost or service behavior.

A useful results table can look like this:

MeasureCandidate ACandidate BCandidate CInterpretation rule
Cases meeting required qualityRecord resultRecord resultRecord resultState denominator and pass definition
Critical failuresRecord count and examplesRecord count and examplesRecord count and examplesAny disqualifying pattern?
Median review effortRecord measured valueRecord measured valueRecord measured valueDefine what review time includes
Completion timeRecord measured valueRecord measured valueRecord measured valueUse the same timing boundary
Run failuresClassify each eventClassify each eventClassify each eventSeparate infrastructure from answer failures

The labels in this table are placeholders, not measured results. Replace them only after the test is complete. For every figure, state the number of cases and runs behind it, the date and the measurement definition. If a measure is unavailable or unreliable, say so instead of filling the cell with an estimate that looks precise.

6. Interpret tradeoffs against the decision

Compare candidates against the thresholds established before the run. A candidate that meets the quality floor but is slower may fit a low-volume, high-consequence task. One with faster drafts but more correction may fit only if review capacity is available and errors remain within the stated limits. The right answer depends on the workflow and its risk tolerance.

Do not combine unlike outcomes into a single score unless the weighting is explicit and agreed in advance. If a team assigns quality 60 percent, latency 20 percent and editing effort 20 percent, that arithmetic expresses a business preference; it is not an objective property of the models. Show component measures alongside any weighted score and explain how changing the weights could change the ranking.

Inspect disagreements and outliers before making a recommendation. Re-read the original input, the output and the rubric. A surprising result can expose a weak case, an unclear instruction, a reviewer inconsistency or a genuine model difference. Record the cause where you can establish it; label it unresolved where you cannot.

Check whether the result holds across important segments. If a workflow includes short and long requests, or routine and exceptional cases, an aggregate may conceal a meaningful difference. Segment only on categories defined before inspecting the result or clearly mark later analysis as exploratory. Avoid turning a handful of examples into a categorical product claim.

A result is not a deployment decision by itself. It is evidence for a bounded choice under a particular prompt, input set, configuration and date. If you intend to switch a production workflow, build a regression suite and plan a separate migration evaluation. The distinct Sonnet 5.5 migration guide covers that transition problem rather than this three-candidate comparison.

7. Worked example: a hypothetical support-draft comparison

Assume a mid-market software company is considering model assistance for drafting replies to billing questions. This is a hypothetical scenario, not a PADISO test or result. The company has a human agent review every draft. It wants to compare the three candidates on factual grounding, policy adherence, material editing effort and completion time before deciding whether to run a limited operational evaluation.

The team defines a useful draft as one that answers the question using the supplied account summary, does not invent account events, follows a provided refund policy and flags missing information rather than guessing. A fabricated refund approval is a critical failure. Tone is scored separately. The decision is not “which model is best?” but “which candidate, if any, meets the minimum quality threshold for a human-reviewed drafting workflow?”

The team freezes a prompt, a case manifest, the rubric and a run plan. It uses the same set of anonymized cases for each candidate, including routine questions, incomplete account histories, a policy boundary and a request that cannot be answered from the supplied facts. Exact model identifiers, test date, prompt version and run conditions are recorded in the results file. No scores are assumed in this example.

Suppose, for illustration only, the team sets a minimum of 18 acceptable drafts out of 20 and zero critical failures before considering a candidate for the next evaluation stage. Those figures are a hypothetical internal decision rule, not a recommended universal threshold. The team should choose its own thresholds based on consequences, review practice and the cost of an incorrect draft.

After scoring, the team might find that one candidate produces fewer complete drafts but needs little editing on those it completes; another might be more consistent on routine cases but mishandle the policy boundary; a third might require more reviewer correction while preserving factual restraint. These are possible patterns, not claims about Astra, Sonnet 5.5 or Opus 5.5. The point is to report the observed pattern and its case-level evidence rather than invent a simple ranking.

If a candidate misses the threshold because reviewers disagree about tone, the team should not present that as the same kind of failure as an invented account fact. It can clarify the tone rubric and rerun a revised protocol, preserving the first result. If the candidate instead repeatedly asserts a refund not present in the input, the team can reject it for this use case even if its average score is otherwise high.

The next decision might be to run a bounded, human-reviewed pilot, revise the prompt and test again, or select no candidate. The hypothetical team should not infer that the comparison establishes suitability for autonomous customer communication. That would be a different workflow with different consequences and acceptance criteria.

8. Use this decision artifact to preserve the result

Create a versioned comparison record that lets another team member understand what was tested and what decision it supports. Keep the artifact with the test data or a stable reference to it. Do not publish a summary table without the definitions and context needed to interpret its numbers.

FieldRecord
Decision being informedSpecific workflow and intended next decision
Test date and protocol versionDate, version and any revised run identifier
Candidate model IDsExact identifiers returned by the environment for each candidate
Environment and run conditionsProvider path, relevant settings and known constraints
Prompt and input versionsStable references or hashes where appropriate
Case setCount, task categories and stated limitations
Scoring rubricCriteria, pass definition and critical-failure rule
ResultsPer-measure values, denominators and case-level findings
ExceptionsReruns, outages, exclusions and reasons
Decision and ownerRecommendation, unresolved risks and accountable decision-maker
Reassessment triggerPrompt, workflow, model-ID or input-distribution change that requires retesting

A concise decision statement should connect evidence to action: “Under protocol version ___, tested on ___, candidate ___ met/did not meet the stated thresholds for ___; proceed to ___ because ___; reassess if ___ changes.” Fill this in only after the run. If results are inconclusive, say that explicitly and identify the next experiment that would resolve the uncertainty.

Publish enough methodology for a reader to judge the scope. Include the test date, model IDs, case count, measurement definitions, failure handling and known limits. Protect confidential inputs and outputs; a public report can describe the test set and redacted examples without exposing customer material. Never imply that a small internal evaluation represents all tasks or all configurations.

If your team needs help translating a model-selection question into a test plan and operating decision, AI strategy and model selection is a relevant next step. The comparison remains your organization’s evidence: the service link is not a claim that PADISO ran or endorsed these tests.

9. Diagnose common comparison failures

A candidate wins on an average but fails a critical case. Keep the failure visible and apply the predeclared threshold. Do not let a high score on easy cases cancel an unacceptable outcome. Revisit the use case only if the business owner is willing to change the workflow or risk boundary, then document that as a new decision.

One model receives more prompt tuning. The outputs are not directly comparable. Preserve the run as exploratory, freeze a revised prompt protocol and give candidates equivalent treatment in a new comparison. If model-specific prompt adaptation is part of the intended system, make adaptation effort part of the evaluation rather than concealing it.

Reviewers disagree sharply. Check whether the rubric defines the disputed criterion, whether reviewers saw the same input context and whether the disagreement changes the decision. Adjudicate using a written method. If ambiguity remains, report a range or unresolved judgment rather than forcing false precision into one score.

Infrastructure problems produce missing outputs. Preserve timestamps and error details, classify the incident and follow the predeclared retry rule. Do not count a delivery failure as a content error without explanation, and do not erase it if operational reliability is part of the question. Separate model behavior from the surrounding service while retaining both if the actual decision concerns the whole workflow.

A run is interrupted or settings drift. Stop, record the affected cases and establish whether the protocol can be resumed consistently. If the environment or model identifier changed, do not merge the results as though nothing changed. Start a new run version or clearly separate the periods and their conditions.

The report contains a winner but no denominator. Add case counts, run counts and definitions. “Higher accuracy” means little without knowing what was counted as correct, how many cases were evaluated and whether excluded cases differed systematically. A table of scores is not self-explanatory evidence.

The team treats a test result as a permanent property. Attach the conclusion to its date, model identifier, prompt and workflow. Retest when a change could affect the decision. This does not require rerunning every test on an arbitrary schedule; it requires a defined trigger and an owner who notices when the tested conditions no longer match reality.

10. Prioritize the next action

Start with the decision sentence and acceptance thresholds, not with a race to collect model outputs. Freeze the protocol, run candidates on the same representative cases, preserve failures, score against observable criteria and examine case-level differences before interpreting aggregates.

Then make a bounded recommendation. Name the tested model identifiers and date, state what the results support, describe what remains unknown and identify the next step. If no candidate clears the threshold, the useful result may be to revise the workflow or test design rather than declare a winner.

Finally, keep the artifact versioned and publish only completed measurements. A reproducible comparison is not a contest between labels; it is a traceable decision record that another team can inspect, challenge and rerun when the conditions change.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call