Course
In this tutorial, I’ll be testing a scenario: A few minutes after a software update, HarborCart — a store I'll use for this scenario — starts seeing checkout failures. Some customers wait more than 30 seconds; others see a server error and cannot pay. The payment provider is also having a short outage, so it looks like the obvious cause.
But the provider outage does not explain why the cart and order pages fail, too. Finding the missing link requires application logs, charts, request records, and the recent code change. This tutorial tests whether Claude Opus 5.5 can follow that evidence, test its explanation under controlled conditions, and report only what the evidence supports.
Some background: Claude Opus 5.5 arrived earlier that week, just before I started this project. Our Claude Opus 5.5 overview covers the launch and benchmarks, so this tutorial stays on the API and builds one investigation agent from the first request to a verified report.
We'll cover how to:
- Make your first Claude Opus 5.5 call and read its content blocks by type
- Give the agent read-only tools with strict schemas
- Let Claude's own code filter logs and traces with programmatic tool calling
- Treat screenshots as hypotheses and check them against metrics
- Test a root cause with a counterfactual replay
- Compare effort levels against the same evidence
- Return a structured report that is allowed to say "inconclusive"
- Calculate the investigation cost from API usage records
TL;DR
HarborCart's investigator separated the payment gateway burst from the retry policy that amplified it, then tested that explanation before returning a report.
- The gateway failure is the trigger, not the complete root cause. Retried charges hold database connections long enough to take down endpoints that never call the gateway.
- Investigation and reporting use separate requests. Web search and citations are available during investigation; a second request formats verified evidence as JSON.
- Programmatic tool calling reduced serialized evidence by 98.8%. Across the three investigations, 142.8 KB of tool results became 1.7 KB of summaries returned to the model.
- Higher effort didn't change the core replay plan. Medium and high selected the same hypothesis and core causal tests.
- The three measured full investigations averaged $0.2737 and about two minutes.
What Is the Claude Opus 5.5 API?
You access Claude Opus 5.5 through Anthropic's Messages API with the model ID claude-opus-5-5. Per the model overview, it takes text and images, with a 1M-token context window and 128K max output. Adaptive thinking is always on, and the default effort is medium.
Standard pricing is $4 per million input tokens and $20 per million output tokens. Five-minute cache writes cost $5 per million and cache reads $0.20. Once prompt caching is active, matching prefixes are billed at that lower cache-read rate.

What changed from Claude Opus 5?
Four points from the migration guide show up directly in this project.
-
Forced
tool_choicewithanyor a named tool returns a 400 error. -
The default effort dropped from
highon Claude Opus 5 tomedium. -
Thinking can't be switched off, and
thinkingblocks must go back unchanged inside a tool loop. -
Notes the model writes between tool calls arrive inside
thinkingblocks, which are empty by default.
What Will We Build With Claude Opus 5.5?
The agent only investigates. It gets read-only evidence tools and no production credentials. After evidence gathering, separate planning requests propose counterfactual tests, and Python validates and runs the medium-effort plan.
The complete code, evidence generator and web app included, is in this GitHub repository.
What happened to HarborCart checkout?
HarborCart is a fictional store. Its checkout-api serves the cart pages, order status, and POST /checkout, which charges a third-party payment gateway. All of those endpoints share one PostgreSQL pool of 15 connections per instance.
A deploy ships, and five minutes later the gateway returns 503s for about 90 seconds. Checkout latency climbs past 30 seconds while the pool sits at 15 of 15. Blaming the payment provider is the easy call, and the gateway really did fail.
The hidden cause sits one step further in. The deploy let failed POST charges retry up to three times, for four total charge attempts, with no pause while the handler still holds its database connection. Slow failing charges now hold connections for 30 seconds or more, until the pool runs out and cart pages that never call the gateway fail too.
I use three terms consistently from here. Here, the trigger is the temporary gateway failure. Retrying checkout POSTs while holding scarce database connections is the amplification mechanism; shared connection-pool exhaustion is the system failure.
What evidence can the agent inspect?
The agent starts with the alert, a monitoring screenshot, and an architecture diagram. Everything else comes through tools: logs, traces, five metrics, deployment metadata, the Git diff, and a runbook. Three competing explanations are seeded into the evidence: an inventory warning, a frontend warning, and possible CPU saturation.

HarborCart checkout path and shared pool. Image by Author.
The diagram says the connection is held for the whole request. It doesn't say that's a problem; the investigation has to work that out.
How will we know the diagnosis is correct?
Define success before building the agent. A correct report must:
- Name the retry change that allowed
POSTretries - State that the database connection stays held through the gateway call
- Explain how longer holds exhaust the pool
- Treat the gateway burst as the trigger, not the amplification mechanism
- Reject at least two of the three alternative explanations
- Cite concrete evidence, including the diff and a metric
- Include a counterfactual replay whose result matches the verdict
How to Use the Claude Opus 5.5 API in Python
You need Python 3.10 or newer and an Anthropic API key with access to claude-opus-5-5. These PowerShell commands clone the project and install its pinned dependencies, including anthropic 1.8.0. If you're on Amazon Bedrock, read the FAQs first, because several features won't port.
git clone https://github.com/KhalidAbdelaty/opus-5-5-api-tutorial.git
cd opus-5-5-api-tutorial
python -m venv .venv
.venv\Scripts\Activate.ps1
pip install -r requirements.txt
Copy-Item .env.example .env
On macOS or Linux, activate with source .venv/bin/activate and copy with cp .env.example .env. Add your key to .env, and python-dotenv loads it for the SDK; our environment variables guide explains the pattern. If you've called Claude from Python before, skip the next subsection, since it only confirms setup.
Make your first Claude Opus 5.5 API call
The smallest useful request confirms the key and shows what comes back.
import anthropic
from dotenv import load_dotenv
load_dotenv()
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=2048,
messages=[{"role": "user", "content": "A checkout API returns HTTP 503 right after a deploy. Name the first two things to check."}],
)
print([block.type for block in response.content])
text = "".join(block.text for block in response.content if block.type == "text")
In this request, the response contains thinking and text blocks. Select blocks by type instead of reading response.content[0].
How to Build a Claude Opus 5.5 Tool-Calling Agent
Tool-calling agents pair Claude's Messages API with Python functions that control data access. The application follows one rule: Claude decides what evidence it needs, and Python decides what it may access.
Our agent harness engineering guide covers broader tool boundaries and loops; HarborCart keeps its tools read-only and limited to this incident.
Define read-only incident tools
Every tool reads a fixed set of evidence and returns a limited JSON result. Log and trace queries return at most 200 rows plus a count, and metric queries return at most 60 points.
Programmatic tool calling doesn't support strict: true, so split the tools in two. Keep the evidence controls and investigation stop tool strict and direct-only. Logs, traces, and metrics use code execution only, which gives Claude one clear path for large evidence queries.
{"name": "finish_investigation", "strict": True,
"allowed_callers": ["direct"],
"input_schema": {"type": "object",
"properties": {"summary": {"type": "string"}},
"required": ["summary"],
"additionalProperties": False}},
{"name": "query_traces",
"allowed_callers": ["code_execution_20260120"],
"input_schema": {...}},
allowed_callers guides the model but is not a security boundary. Python checks the caller before executing each tool and rejects a direct query call. Programmatic calls also skip strict validation, so query functions still validate their own arguments.
The application tags every accepted tool result, rejects findings that cite missing evidence, and accepts documentation URLs only when web search returned them. Python, not the model, records replay outputs.
Use strict schemas instead of forced tool choice
As noted in the migration section, keep tool_choice at auto. Say in the prompt when a tool applies, and use strict schemas where arguments must be exact.
Build the multi-turn investigation loop
The loop sends the conversation, runs any tool_use blocks, appends the results, and repeats. Append the assistant's blocks unchanged, thinking included, and while programmatic code is paused, pass the container ID back with only tool_result blocks.
The investigation request includes vision, tools, web search, effort, and the task budget, but no output schema. This keeps citation-bearing search results away from structured JSON output while the stable request prefix keeps prompt caching active:
request = dict(
model="claude-opus-5-5",
max_tokens=16_000,
system=[{"type": "text", "text": SYSTEM_PROMPT, "cache_control": {"type": "ephemeral"}}],
tools=investigation_tools,
cache_control={"type": "ephemeral"},
thinking={"type": "adaptive", "display": "updates"},
output_config={
"effort": "medium",
"task_budget": {"type": "tokens", "total": 20_000},
},
betas=["task-budgets-2026-03-13", "thinking-display-updates-2026-08-18"],
)
How to Send Images to the Claude Opus 5.5 API
Attach the dashboard and architecture diagram to the first user message as base64 PNGs. Tell Claude to treat anything it reads from an image as a hypothesis and confirm it with query_metrics.

Dashboard shows pool saturation, flat CPU. Image by Author.
Both the dashboard and metric queries use the same source data. CPU stays near 30% while the pool is full, which argues against "the host is overloaded" before any query runs.
Cross-check visual observations against raw metrics
The screenshot suggests where to look, but the numerical series decides whether the observation holds. Vision generates a hypothesis; metrics test it.
For image-first workflows, see our agentic vision tutorial. HarborCart uses vision only to choose the next metric.
How Does Claude Opus 5.5 Programmatic Tool Calling Work?
Programmatic tool calling lets Claude write Python that runs in a code execution container and calls your tools as functions. Raw results stay in the sandbox, and only the code's printed output reaches the model.
Fan out across logs and traces
The agent writes short scripts that pull failing traces and print only the counts by endpoint. In one complete investigation, programmatic tool calling reduced serialized evidence returned to the model by 98.8%. The tool results were 42.9 KB and the summaries were 0.5 KB, a byte measure rather than billed input-token savings.

Tool calls narrow the incident evidence. Image by Author.
Add documentation search for uncertain dependency behavior
The application exposes restricted web search for retry-library semantics. Claude did not call it during the final evaluation, so the measured diagnosis rests on the diff, metrics, logs, and traces. The urllib3 reference independently confirms that allowed_methods=None retries any verb and backoff_factor=0 removes the wait, but that page is not part of the measured evidence.
How to Verify a Root Cause With a Counterfactual Replay
A counterfactual replay reruns the incident's traffic with one suspected cause removed and checks whether the failure goes away. It turns "these lines rise together" into a test.
Keep the replay honest
The replay reuses the same traffic pattern. For the comparison below, each scenario changes one condition, and the application controls which changes are allowed.
The summary separates gateway 503s from pool timeouts, and checkout 503s from cart and order reads. That split is what lets the model tell the trigger from the amplifier.

Each replay changes exactly one thing. Image by Author.
The baseline replay produced 124 503s: 105 pool timeouts, including 68 failures on read endpoints, and 19 gateway errors. Reverting the retry policy removed every pool timeout and read failure but surfaced 93 gateway 503s on checkout. Releasing the connection before the gateway call also removed pool failures while leaving 33 gateway 503s, and removing the gateway burst produced no errors.
The replay exposes the trade-off: a rollback protects the shared pool but allows more checkout failures through. Use it as a stopgap. Then add an idempotency key so a repeated charge cannot bill twice, and stop holding the connection during the gateway call.
Make verification a rule in code
The system prompt asks for a replay, but a prompt is not an enforcement mechanism. The loop checks whether replay evidence exists and rejects an untested diagnosis.
Keep this check in Python. A sharper prompt may improve compliance, but it can't guarantee it.
How to Use Effort and Task Budgets With Claude Opus 5.5
Effort sets how much Claude reasons per step, and a task budget sets how much work the whole loop should take. Our Claude Opus 5 API tutorial compares all five effort levels; here, medium and high receive the same pre-replay evidence.
Compare medium and high on the same evidence
Production stays at medium. Before replay, the application asks medium and high effort to design a causal test from the same evidence. It executes only the medium recommendation; the high response is used only for comparison.
The high request uses a per-message output_config.effort change behind mid-conversation-output-config-2026-07-01. It does not see the medium answer.
Both effort levels selected the same hypothesis and the same three core replay scenarios. High used 2,631 output tokens on average versus 2,307 at medium, and cost about 11% more without changing the causal test.
Set a task budget for the whole loop
Choose a task budget from observed use rather than guessing. HarborCart's largest unbounded investigation consumed 13,322 counted tokens, including model output and tool-result text Claude saw. Adding a 25% margin gives 16,653, below Anthropic's 20,000-token minimum, so the configured budget is 20,000.
Keep turns and elapsed time as application limits. The experiment runner stopped starting new work after recorded spend reached $2.50. This is not a hard cap because a request already in progress can finish above it.
How to Use Claude Opus 5.5 Structured Outputs
The final answer uses structured outputs. Its flat schema covers the verdict, cause, rejected hypotheses, evidence, and fix. Cost and latency stay out because the application measures them.
Separate investigation from reporting
Web search citations and output_config.format cannot share a request: citations need interleaved content blocks, while the schema requires JSON. HarborCart therefore investigates without an output schema. It stores findings tied to their sources and the replay results, then sends only that verified evidence to a second request with no tools or web search.
import json
report_response = client.messages.create(
model="claude-opus-5-5",
max_tokens=16_000,
system=report_instructions,
messages=[{"role": "user", "content": json.dumps(verified_evidence)}],
output_config={
"effort": "medium",
"format": {"type": "json_schema", "schema": report_schema},
},
)
The second request needs only the verified evidence, so preserving the full investigation cache is unnecessary.
Allow "inconclusive" in the verdict field. A report should not be forced into a verified diagnosis when the replay contradicts its explanation.
Schema-valid does not mean correct
The schema validates the report's shape, while the replay validates the diagnosis. A refusal also returns HTTP 200 with stop_reason: "refusal" and may not match your schema, so check the stop reason before parsing.
Did Claude Opus 5.5 Find the Real Root Cause?
All three final reports found the core causal mechanism and ruled out the three alternative explanations. Two met all eight checks; the third scored 6/8 because it omitted the explicit POST-retry configuration change and did not cite the deployment diff. That is why offline scoring stays separate from schema validation: valid JSON and the right diagnosis can still produce an incomplete report.
The reports also identify a second risk: retrying a charge can bill a customer twice. RFC 9110 does not define POST as inherently idempotent and advises against automatic retries unless the client knows the operation is safe to repeat. An idempotency key supported by the payment provider is one common way to make those retries safer.
Our Streamlit tutorial covers the interface setup. HarborCart's interface shows investigation events, the medium and high replay plans, replay results, the final report, and cost. For status between tool calls, the Claude Opus 5.5 prompting guide describes display: "updates"; the app also renders tool events when an update block is empty.
How Much Did the Claude Opus 5.5 Investigation Cost?
A complete investigation cost $0.2582 to $0.2838 and took 108.8 to 129.7 seconds. The average cost was $0.2737, including the optional high-effort comparison. Output averaged $0.2043, about three quarters of the total.
Count cache tokens the way the API reports them
input_tokens already excludes cached tokens, so total input is the sum of three fields. Don't subtract cache reads from it. If your cost tracking already handles this, skip the snippet.
cost = (
usage.input_tokens * 4.00 # uncached input only
+ usage.cache_read_input_tokens * 0.20
+ cache_creation.ephemeral_5m_input_tokens * 5.00
+ cache_creation.ephemeral_1h_input_tokens * 8.00
+ usage.output_tokens * 20.00
) / 1_000_000 + web_search_requests * 0.01 # from usage.server_tool_use
Read the search count from usage.server_tool_use. With response_inclusion: "excluded", counting search blocks in the response can undercount.
Every investigation request includes web_search_20260318, so Anthropic does not add a separate code-execution container charge beyond token and search costs. If you remove the qualifying web tool, track code-execution time separately.
Prompt caching on Claude Opus 5.5 needs at least 512 tokens. During investigation, top-level cache_control moves the breakpoint as the history grows. The report receives only compact verified evidence and intentionally starts without the full investigation cache.
What Would Need to Change Before Production?
A real on-call tool needs more controls than this demo, all of them in application code:
-
Scope observability credentials to the data the tools read, keep remediation in a separate permission tier, and enforce caller permissions in Python rather than trusting prompts or
allowed_callers. -
Treat logs, tickets, web pages, and tool results as untrusted data. Validate their shape and never execute text copied from them.
-
Classify and redact production logs before sending them to code execution. Anthropic's data-retention table marks code execution and programmatic tool calling as ineligible for ZDR and HIPAA readiness, with container data retained for up to 30 days. Web-search filtering through code execution is also outside ZDR and HIPAA eligibility.
-
Branch on
stop_reasonbefore parsing, count refusals separately from HTTP errors, and routeinconclusivereports to a human. -
Save tool calls, replays, hypotheses, token usage, and timing as the evidence log. Do not store hidden reasoning.
When Should You Use Claude Opus 5.5 for Agentic Work?
Use Claude Opus 5.5 when a mistaken diagnosis would cost more than the API call. Root-cause analysis, repository-wide debugging, migration planning, and investigations that combine logs, images, documentation, and several tools fit that test.
Skip it for formatting, classification, extraction, and short questions that do not need a tool loop. A smaller model will usually finish those tasks faster and at a lower cost.
For important agent work, prefer tasks where conclusions can be checked against tests, metrics, source evidence, or human review. Keep production at medium unless paired evaluations show that higher effort improves the plan on your workload.
Final Thoughts
We built an incident investigator that reads mixed evidence, calls bounded tools, tests its own diagnosis, and returns a structured report. All three final reports preserved the trigger and root-cause split described earlier, but Python still had to require the replay.
I would not generalize that result to every incident or codebase. What carries over is the method: limit data access, filter large tool results before they reach the model, allow an "inconclusive" verdict, and verify the explanation outside the model. That replay is the part I would keep even in a smaller version of this project.
Changing the evidence tools and verification step lets the same pattern support a CI failure investigator, a pull-request reviewer, or a migration checker. My first extension would be a router that sends simple incidents to a cheaper model and reserves Claude Opus 5.5 for cases that need several evidence sources. For the model-level picture, see the Claude Opus 5.5 overview linked in the introduction.
I’m a data engineer and community builder who works across data pipelines, cloud, and AI tooling while writing practical, high-impact tutorials for DataCamp and emerging developers.
FAQs
Can you turn thinking off in Claude Opus 5.5?
No. A request with thinking: {"type": "disabled"} returns a 400 error at every effort level, so lower effort when you want less reasoning and lower cost.
Does the API tell you how much task budget is left?
No. The countdown is visible only to the model, and usage has no budget field. Sum usage in the application if you need to track spend.
Is Claude Opus 5.5 better than Claude Opus 5?
Not for every task. Claude Opus 5.5 changes the price, default effort, and several API behaviors, but model quality still needs an evaluation on your own workload.
Can I run this agent on Amazon Bedrock?
Not unchanged. The basic Messages and client-side tool loop can move to Amazon Bedrock with the model ID anthropic.claude-opus-5-5. Bedrock currently lacks the structured outputs, server-side code execution, web search, and programmatic tool calling used here. Claude Platform on AWS is a separate service with broader feature support.
Can Claude Opus 5.5 run Python code?
Yes. The code execution tool lets Claude run Python in a managed container. Programmatic tool calling also lets that code call tools you allow, but your application still runs client-side tools and controls their permissions.

