Anthropic launched Claude Sonnet 5.5 on September 28, 2026, bringing cyber fallbacks like those used on its more capable models to the Sonnet line. A flagged request may be answered by Sonnet 5 or remain blocked, depending on the surface and configuration. Security teams upgrading a review pipeline need to decide what counts as a completed assessment when that happens.
Consider an automated code review that receives an empty response and records "no vulnerabilities found." If that response was a refusal, the pipeline has turned missing analysis into a clean verdict. A successful fallback introduces another question: does the replacement model meet the requirements for this review?
The procedure below is our engineering recommendation based on Anthropic's documentation. We have not measured the safeguards' accuracy or tested a live fallback integration.
Separate interactive switching from API routing#
In Claude's interactive experience, automatic switching is enabled by default. After a fallback, the model picker stays on Sonnet 5 for the remaining chat until the user changes it. The control is Settings > Capabilities, or Config > MODEL & OUTPUT in Claude Code: turn off Switch models when a message is flagged to pause instead. Interactive switching reference.
The API needs its own configuration. Anthropic documents server-side fallback as a Claude API beta using fallbacks: "default" and the server-side-fallback-2026-07-01 header. That parameter is unavailable on Bedrock, Google Cloud, and Microsoft Foundry, and unsupported in Message Batches. The documentation describes SDK middleware for client-side routing. Choose one routing mechanism; do not stack middleware and server-side fallback on the same request. Fallback setup and platform scope.
For Sonnet 5.5, the documented default server route sends cyber and frontier_llm declines to Sonnet 5. It does not retry bio, reasoning_extraction, or general_harms. This is category-specific routing, not a general retry for outages or rate limits. Sonnet 5.5 behavior.
For a review that requires one approved model, leave automatic fallback out of the API request and check that no client middleware adds it. If the team accepts a second model, approve that route deliberately and evaluate its reports separately. An instruction in the prompt is insufficient to enforce either choice.
See how your AI systems hold up under real attacks
General Analysis maps AI applications and agents, red teams prompts, retrieval, tools, MCP servers, browser actions, permissions, and business workflows, then turns findings into evidence your team can reproduce and retest.
Give an incomplete review its own outcome#
Anthropic's refusal response returns HTTP 200 with stop_reason: "refusal". Its category can be null; a classifier label is not required for the refusal to count. Branch on the stop reason rather than parsing the explanation text.
We recommend the following policy for a workflow that expects a completed report. The decisions are application policy; Anthropic does not enforce them for your release gate.
| Response or observation | Application decision |
|---|---|
refusal, including an unknown or null category | Mark the review refused. Do not emit a clean verdict. |
max_tokens, pause_turn, or tool_use | Keep the review incomplete; hand it to the appropriate continuation or tool handler. |
| A fallback ran, or the serving model differs from the approved model | Require the workflow's model-change review before accepting its report. |
end_turn from the approved model | Validate the report's schema, assessed revision, and required coverage before accepting it. |
| Missing or unfamiliar response fields | Preserve an unresolved outcome and investigate the integration. |
The stop-reason reference distinguishes natural completion, token exhaustion, pauses, and tool requests. end_turn only establishes that generation ended. It cannot establish that the requested files were reviewed or that the findings are correct.
In the code-review example, require the report to identify the exact commit and the files it assessed. Reject a report for a different revision, and keep skipped files visible. A valid report with no findings can then remain distinct from a review that never produced a valid report.
Put a small gate before the verdict parser#
This illustrative Python function accepts a fully assembled, non-streaming JSON response. It makes no API calls and never executes tools. The state names are application-defined. It deliberately sends every detected model change to review; deployments that permit fallback can replace that branch with their own evaluated policy.
Code source: illustrative; response fields follow Anthropic's refusal and fallback reference.
def review_state(message, expected_model):
stop = message.get("stop_reason")
if stop == "refusal":
return "refused"
if stop != "end_turn":
return "incomplete"
usage = message.get("usage") or {}
iterations = usage.get("iterations") or []
blocks = message.get("content") or []
fallback_seen = any(
item.get("type") == "fallback_message"
for item in iterations
) or any(block.get("type") == "fallback" for block in blocks)
served_model = message.get("model")
if not served_model:
return "incomplete"
if fallback_seen or served_model != expected_model:
return "model_review_required"
has_text = any(
block.get("type") == "text"
and bool(block.get("text", "").strip())
for block in blocks
)
return "validate_report" if has_text else "incomplete"
Validate the incoming JSON shape before calling this function. Pass the expected resolved model ID from application configuration. Only validate_report reaches the report parser, and parser failure must leave the review incomplete. That state is not permission to approve a pull request or deploy a change.
Test the handler with synthetic responses: a refusal with a null category; an empty end_turn; a truncated report; a tool request; a fallback followed by another refusal; a fallback that answers; and a report from the expected model. Then test the report validator with a stale commit ID and missing file coverage. These fixtures exercise your application's decisions without trying to trigger a safeguard or reproduce an exploit.
Record who answered before accepting the result#
Save the requested and serving models with the assessment's final stop reason, request or message ID, input revision, and report-validation result. Include fallback boundaries and usage iterations when present. These recommended audit fields let an investigator identify a model change without putting raw source code or full prompts in a general-purpose event log.
Streaming needs a separate implementation. A fallback can happen after output has begun, so the initial message_start model is insufficient. Anthropic documents the handoff in a fallback content block and the final usage iterations. Buffer any release decision until the stream has finished and the report has passed validation. A terminal refusal leaves partial output unsuitable as a completed assessment. Streaming fallback behavior.
A model switch also changes the available reasoning context. Sonnet 5 cannot read Sonnet 5.5 thinking blocks; the API drops incompatible blocks. Do not describe the fallback as the same evaluator simply continuing unchanged. Check request compatibility and conversation handling against the Sonnet 5.5 migration guide.
Keep tool authority independent of this admission check. A model change must not grant new credentials or skip approval for a write. The coding-agent security baseline covers those boundaries; Managed Agents tool approvals covers that product's separate confirmation and recovery flow. For supported enterprise app sessions, the Compliance API guide explains what retained transcripts can and cannot establish.
Set expectations for legitimate security work#
Anthropic says Sonnet 5.5's checks inspect context as well as the latest message, including files, connector content, and web results. A refusal therefore does not establish that the user submitted a malicious request. Preserve the task's authorization and route a blocked assessment to an operator or an approved alternative workflow; do not turn repeated rewording into an automatic retry strategy. Safeguard scope.
Existing Cyber Verification Program approval is not a launch-day exception for Sonnet 5.5. Anthropic's current program guidance excludes Sonnet 5.5 and Opus 5.5 while describing a planned expansion. Keep that distinction in the rollout plan instead of promising access the program does not yet provide.
Track refused, incomplete, model-changed, and validated reviews separately. Before making Sonnet 5.5 a required review step, prove that each of those outcomes reaches the intended owner and that only a validated report can satisfy the gate.
Frequently asked questions
- Does HTTP 200 mean Sonnet 5.5 completed the review?
No. A refusal can return HTTP 200 with stop_reason set to refusal. Other stop reasons can indicate unfinished work, and even end_turn requires application-level validation of the report. Keep refused and incomplete reviews separate from completed reviews with no findings.
- Does Sonnet 5.5 automatically fall back on every API request?
No. API fallback must be configured. Anthropic documents beta server-side fallback on the Claude API, while its interactive automatic model switching is enabled by default. Do not infer API configuration from a Claude app setting.
- Does existing Cyber Verification Program approval cover Sonnet 5.5?
Anthropic's launch documentation says Sonnet 5.5 is not available through the Cyber Verification Program at launch. Expansion is planned; existing approval should not be treated as current access for this model.

