TopxAI

Loading...

Blog

Moving to Claude Sonnet 5.5 on TopxAI

· anthropic, models, api, topxai

Update the Sonnet model ID, keep the same token prices, and check thinking and tool settings when moving your client to Claude Sonnet 5.5.

Use claude-sonnet-5-5 for Claude Sonnet 5.5 on TopxAI's Claude shared and official routes. The upstream ID is the same. Sonnet 5's input, output and cache rates carry over unchanged, but the old claude-sonnet-5 ID is retired and does not forward.

Update the client and key

Replace the saved model ID in your SDK or coding client. Keep the same base URL and key if the key has no model restriction. For a restricted key, edit its allowed-model list separately; changing the client's model does not expand the key's permissions.

For Claude Code, update ANTHROPIC_MODEL and any ANTHROPIC_DEFAULT_SONNET_MODEL value in your shell or ~/.claude/settings.json:

export ANTHROPIC_BASE_URL=https://ai.topxea.com
export ANTHROPIC_AUTH_TOKEN=$TOPXAI_API_KEY
export ANTHROPIC_MODEL=claude-sonnet-5-5
export ANTHROPIC_DEFAULT_SONNET_MODEL=claude-sonnet-5-5
claude

The Claude Code guide covers persistent settings. For a direct Messages request:

curl https://ai.topxea.com/v1/messages \
  -H "x-api-key: $TOPXAI_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{"model":"claude-sonnet-5-5","max_tokens":4096,"output_config":{"effort":"medium"},"messages":[{"role":"user","content":"Review the trade-offs in this API design."}]}'

The five prices stay the same

Default USD prices per million tokens, checked against Anthropic's model page on October 2, 2026:

Category Shared Official Anthropic list
Input $1 $1.80 $2
Output $5 $9 $10
Cache read $0.10 $0.18 $0.20
Cache write, 5 minutes $1.25 $2.25 $2.50
Cache write, 1 hour $2 $3.60 $4

The shared route is 50% of list and the official route is 90%, including both cache-write durations. There is no long-context premium. Sonnet 5.5 supports a 1M-token context and up to 128K output tokens. Actual charges use the upstream-reported token categories and the request's captured rates; a cache miss is not automatically a cache write. The model page and usage log show the applicable prices and counts.

Check thinking settings

With no thinking field, Sonnet 5.5 uses adaptive thinking at high effort. Adaptive mode accepts low, medium, high, xhigh and max. Thinking tokens count toward the output limit and are billed as output.

TopxAI translates an older thinking: {"type":"disabled"} request, or OpenAI reasoning.effort: "none", to between_tools with low effort. This turns off up-front thinking while retaining short progress updates between tool calls. Manual budget_tokens settings are translated into effort levels.

If you send native thinking: {"type":"between_tools"}, use only low, medium or high effort. That mode accepts no display, budget_tokens or other thinking subfields. Use adaptive mode for xhigh or max, or for display: "summarized" / "updates"; those adaptive display settings are preserved.

Read returned content blocks by type. Native thinking blocks and signatures need to travel back unchanged in tool conversations. Existing thinking state belongs to its model; start a fresh conversation when switching an agent from Sonnet 5 to Sonnet 5.5.

Check sampling and tools

Sonnet 5.5 does not support temperature, top_p or top_k; TopxAI removes those settings before forwarding. Forced tool selection (any / a named tool, or OpenAI required / a named function) returns 400 without a charge. Use auto; native tool_use / tool_result conversations and thinking signatures remain intact.

Clients using Anthropic's built-in computer-use tools must migrate to its new official toolset format. Ordinary custom tools keep their definitions. Anthropic's migration guide explains the upstream changes; the mappings above describe TopxAI's compatibility behavior.

During the October 2 supplier checks, some between_tools requests returned 502, including medium on both routes and OpenAI none on the official route. Adaptive mode with explicit low effort passed on both routes: use thinking: {"type":"adaptive"} with output_config: {"effort":"low"} for Messages, or reasoning_effort: "low" for Chat Completions, when it fits your task. Adaptive mode can still think before answering, and those tokens are billed as output. TopxAI preserves the requested mode rather than silently substituting another one.

Models: claude-sonnet-5-5 · Models & pricing

Related

Other languages: 简体中文