Course
Every month, a finance team has to confirm that its records match the money that actually reached the bank. Sales, minus refunds and the fees the card processor keeps, should equal the deposits. This check is called a reconciliation, and when the numbers don't agree, someone has to dig through the records to find out why.
In this tutorial, we'll hand that job to Claude Sonnet 5.5 and build an AI agent around it in Python. Here, an agent is a program that Claude can call tools, such as a function that looks up refunds, and use the results to decide what to check next. The test case is Rivermark, a fictional subscription company whose September numbers don't add up.
The hard part is trust. Claude should see every record, but it shouldn't change the books until its explanation holds up. So Claude starts with tools that can only read. When it proposes a correction, Python checks the evidence first. Only then does Claude get a tool that records that one correction in a separate list, while the original data stays untouched. A final Python check compares the result with bank records kept outside Claude's tools.
What interested me was whether this setup could catch a mistake that looked reasonable. We'll cover how to:
- Make a first Claude Sonnet 5.5 API call in Python
- Give Claude tools that can read records but not change them
- Check Claude's proposed correction in Python before it can write anything
- Give Claude a new tool partway through the conversation with a mid-conversation system message
- Change Claude's effort on the later steps
- Check the final numbers in Python and work out what each API call costs
TL;DR
At medium effort, Claude Sonnet 5.5 found a $149.00 refund counted in the wrong month but missed a separate $15.00 fee kept by the card processor. Python's final check showed the totals still didn't match, so Claude kept going in the same conversation, found the fee, and fixed it.
-
Claude had already seen the fee it missed. It opened both records for a disputed payment but decided the $15.00 fee was already counted.
-
Python decided when Claude could write. The tool for recording corrections stayed hidden until Claude's proposal passed Python's checks, which rejected 2 of 4 proposals.
-
Changing tools and effort didn't reset the conversation. Because nothing earlier was rewritten, 89.3% of the 118,308 input tokens came from the prompt cache, which is billed at a lower rate.
-
Higher effort was not necessary in the matched replay. A separate replay from the same failure point stayed at
mediumand also found the fee after the same message from Python. -
The main reconciliation moved from
mediumtohigh. It took 15 API calls and cost $0.1190. The matched replay is separate.
Those numbers describe one fictional dataset. Treat them as behavior to test in your own application, not a benchmark.
Introduction to Claude Models
What Is Claude Sonnet 5.5?
Claude Sonnet 5.5 is part of Anthropic's Claude 5.5 family. It had just been released when I started this project, and its API model ID is claude-sonnet-5-5. Per the model overview, it has a 1M-token context window, up to 128K output tokens, adaptive thinking on by default, and a default effort of high on the API. Standard pricing is $2 per million input tokens and $10 per million output tokens.
Our Claude Sonnet 5.5 overview covers benchmarks, pricing comparisons, and access. Three of its API features are new in this release, and Rivermark uses all of them.
What's new in the Claude Sonnet 5.5 API?
Claude Sonnet 5.5 adds three ways to change a conversation while it runs. According to What's new in Claude Sonnet 5.5, none of them is available on Claude Sonnet 5:
- Per-message effort: change how much Claude reasons on later turns.
- Mid-conversation system messages: add system instructions partway through.
- Mid-conversation tool changes: show or hide declared tools partway through.
What Will We Build With the Claude Sonnet 5.5 API?
The Rivermark agent is a Python application built around one Messages API conversation with two permission levels. During the investigation, Claude can read orders, refunds, processor transactions, the close policy, and Rivermark's reconciliation check. After approval, it can record only the approved adjustments.
Rivermark uses a custom Messages API loop rather than the Claude Agent SDK because the approval gate has to sit between Claude's tool calls and their execution.
The complete code and sample data are in the Rivermark GitHub repository.

Claude proposes, Python grants write access. Image by Author.
What is the Rivermark reconciliation problem?
Rivermark's check reports $3,400.14 as the expected payout and $3,251.14 as its calculated processor total, a $149.00 difference. Claude has to explain the difference between the records without seeing either of the hidden causes.
Rivermark sells three monthly plans: Starter at $29, Team at $79, and Business at $149. The sample contains 58 September orders, 7 refund records, and 65 September processor transactions. Each processor record has an amount, a fee, and a net value.
How does Python define a successful reconciliation?
Python, not Claude, decides whether the reconciliation is complete:
-
The month is September 2026, by the processor settlement date.
-
The total of September bank deposits is the experiment's independent settlement target.
-
Balanced means expected payout plus adjustments equals those deposits to the cent.
-
Every adjustment cites the processor
txn_idsthat Claude retrieved, and its amount equals their net. -
Claude can only add approved adjustments and submit the final report.
-
Raw exports are hashed before processing and must match afterward.
Claude cannot inspect the bank records or target total during its initial investigation. After a failed check, Python reveals only the expected payout, aggregate deposit total, and remaining difference, not the bank records themselves.
How to Set Up the Claude Sonnet 5.5 API in Python
You need Python 3.10 or newer, which the Python SDK requires, an Anthropic API key, and anthropic 1.9.0. These PowerShell commands clone the project, create the environment, and build the sample data:
git clone https://github.com/KhalidAbdelaty/sonnet-5-5.git
cd sonnet-5-5
python -m venv .venv
.venv\Scripts\Activate.ps1
pip install -r requirements.txt
Copy-Item .env.example .env
python build_data.py
On macOS or Linux, use source .venv/bin/activate and cp .env.example .env, then put your key in .env. Our environment variables guide explains the pattern.
streamlit run app_streamlit.py opens a web interface that shows each step of the reconciliation as it happens, and our Streamlit tutorial covers the setup.
If your API key already works, skip the next request and move to adaptive thinking.
How to make your first Claude Sonnet 5.5 API call
If API request and response objects are new to you, our Python API guide covers the basics. One refund question is enough to confirm the key and inspect the returned content blocks:
import anthropic
from dotenv import load_dotenv
load_dotenv()
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=4096,
messages=[{"role": "user", "content": "A refund was requested on August 31 and settled on "
"September 2. Which month's payout should it reduce, and why?"}],
)
print([block.type for block in response.content])
print("".join(block.text for block in response.content if block.type == "text"))
print(response.usage)
In my run, the response started with a thinking block. Select blocks by type instead of reading response.content[0]; thinking tokens are billed as output.

First response separates thinking from text. Image by Author.
How to configure adaptive thinking and effort
Every request sends the same top-level settings, and only messages grows:
response = client.beta.messages.create(
model=MODEL, max_tokens=MAX_TOKENS, system=SYSTEM_PROMPT, tools=TOOLS,
cache_control={"type": "ephemeral"}, # automatic caching, breakpoint moves forward
thinking={"type": "adaptive", "display": "updates"},
output_config={"effort": START_EFFORT}, # never changes: per-message changes do that
messages=messages, betas=BETAS,
)
Despite the API's high default, this workflow starts at medium. Anthropic's effort guide says: "For agentic coding and multistep tool use, start with medium for well-specified tasks and move to high for harder or longer ones."
Thinking stays adaptive because the effort change later in the build depends on it. display: "updates" (beta, thinking-display-updates-2026-08-18) returns the notes Claude writes between tool calls. Without that setting, the thinking blocks are empty.
The top-level cache_control turns on automatic prompt caching, with a breakpoint that moves forward as the conversation grows. The first request wrote 2,080 tokens to cache, well above Claude Sonnet 5.5's 512-token minimum.
How to Build a Read-Only Reconciliation Agent
A read-only investigation agent lets Claude request evidence but exposes no write tool. Rivermark also rejects unapproved write calls in Python.
Which read-only tools does Claude use?
Claude gets five read tools and one proposal tool, all with strict: true. The descriptions say what each tool returns and nothing about where to look:
-
list_sourcesreturns sources, columns, and row counts. -
query_recordsreturns up to 40 rows from one source, with an optional filter and date range. -
aggregate_recordscounts rows and totals amount_cents by any column. -
read_policyreturns the close policy. -
run_reconciliation_checkruns Rivermark's existing internal logic, bugs included. -
submit_plansends a diagnosis and proposed adjustments to Python for validation, and writes nothing.
Two more tools sit in the same tools array, but defer_loading: true keeps them out of Claude's view for now. We'll get to how they appear later:
{"name": "run_reconciliation_check", "strict": True,
"description": "Run Rivermark's current internal reconciliation logic for September 2026, "
"including adjustments recorded so far.",
"input_schema": _schema({}, [])},
{"name": "create_adjustment", "strict": True, "defer_loading": True,
"description": "Record one approved adjustment in the close adjustments ledger. Never edits source files.",
"input_schema": _schema({...}, ["evidence_txn_ids", "rule", "amount_cents", "memo"])},
The write-tool schema is known at the first request, so the tool is declared upfront. Named or any tool choice returns a 400 error, so the prompt states when submit_plan applies.
How does the Claude tool-use loop work?
Our agent harness engineering guide explains how Python can manage longer agent loops. Rivermark's loop sends the conversation, runs any tool_use blocks in Python, and appends the results. Every record ID a read tool returns goes into an observed set that the plan gate checks later:
messages.append({"role": "assistant", "content": response.content}) # thinking blocks go back unchanged
if response.stop_reason == "tool_use":
results = []
for block in response.content:
if block.type != "tool_use":
continue
if block.name in READ_TOOLS:
out = reads.run(block.name, block.input) # adds returned IDs to gate.observed
results.append({"type": "tool_result", "tool_use_id": block.id, "content": dumps(out)})
... # submit_plan goes to the gate; create_adjustment to the executor
messages.append({"role": "user", "content": results})
The assistant's turn goes back exactly as received, empty thinking blocks included. The migration guide explains that Claude Sonnet 5.5 binds thinking blocks to the earlier messages, so editing that history can return a 400 error.
What did Claude find at medium effort?
At medium, the investigation took six API calls and nine read-tool calls. Claude pulled the refunds and grouped the processor lines by reporting_category. It found RF-1043, a $149.00 refund for an August 31 order that settled on September 2. Policy rule POL-3 puts it in September.
Then it opened both dispute rows. TXN-50036 contains a -$149.00 principal amount, a $15.00 fee, and a -$164.00 net cash effect. TXN-50052 returns the $149.00 principal with no fee. Claude wrote: "DSP-0077 nets to zero and its $15 fee is already correctly booked, so RF-1043 fully explains the variance."
Claude had confused the returned principal with the cash effect after fees:
- The principal does net to zero: -$149.00 + $149.00 = $0.00.
- The transaction nets do not: -$164.00 + $149.00 = -$15.00.
The gate rejected Claude's first plan because it cited order ORD-20813 without retrieving it. Claude fetched the order, resubmitted, and PLAN-1 passed with one adjustment.
Gate Write Access Behind an Approved Reconciliation Plan
Before exposing the write tool, the gate checks where the evidence came from and what the plan would change.
How does the plan gate check evidence?
Each adjustment in a plan cites processor txn_ids. The gate accepts it only if every cited line came back from a read tool in this conversation and the lines net to the proposed amount:
def evidence_problems(self, item: dict) -> list[str]:
"""Provenance: every cited line was retrieved, and the lines net to the adjustment."""
ids = item["evidence_txn_ids"]
problems = [f"{t} was never returned by a read tool in this conversation."
for t in ids if t not in self.observed]
unknown = [t for t in ids if t not in self.lines]
if unknown or not ids:
problems.append(f"Evidence must be processor txn_ids; not found: {', '.join(unknown) or 'none given'}.")
elif sum(self.lines[t]["net_cents"] for t in ids) != item["amount_cents"]:
problems.append(f"amount_cents {item['amount_cents']} is not the net_cents total of {', '.join(ids)}.")
return problems
A $15.00 adjustment that cites only the dispute debit fails because that line's net is -$164.00. The plan must cite the reversal too.
When does the plan gate reject a correction?
The gate also checks policy rules and duplicate transactions. A plan is rejected, and write access stays locked, if any item does one of these:
- Cites a supporting order or refund that Claude never retrieved
- Uses a policy rule other than POL-2, POL-3, or POL-4
- Covers transactions that another adjustment already covers
Rejections return as the submit_plan tool result, so Claude can investigate further and resubmit. The gate rejected 2 of 4 submissions, and Claude fixed each on its next call. Even after approval, create_adjustment accepts only entries that match an approved item exactly.
Add the write tool mid-conversation
Once the gate approves a plan, Python appends a role: "system" message with a tool_addition block. The change requires the inline-tools-2026-09-15 beta header. The tools array and every earlier message stay unchanged, so the cached prefix still matches. The instruction text comes from Python, not from Claude:
text = UNLOCK_TEXT.format(plan_id=approved_plan)
append_system([{"type": "text", "text": text},
{"type": "tool_addition", "tool": {"type": "tool_reference",
"name": "create_adjustment"}}])
gate.write_unlocked = True
A system message with content has to follow a user turn, including one with tool_result blocks. It cannot sit between a tool_use block and its result. System messages have higher priority, so never insert Claude's plan text, tool output, or data into one. The tool_addition block names create_adjustment by reference, and the tool becomes visible only after the plan passes.
Caching continued after the tool change. The request processed 231 uncached input tokens and read 6,883 from cache.
Why Was the First Reconciliation Adjustment Incomplete?
The first adjustment was correct and still left the job unfinished. Claude recorded ADJ-001, -$149.00 under POL-3, and reported it done. Rivermark's internal check would have agreed, showing a $0.00 variance. That sounds finished, but it isn't.
The independent Python check compares against bank deposits instead. Expected payout after adjustments was $3,251.14, deposits were $3,236.14, and $15.00 remained.
That gap is why the completion check lives in Python, not in Claude's final message.

Matched replay branches from failed verification. Image by Author.
Escalate Effort After Verification Fails
Changing effort mid-conversation on Claude Sonnet 5.5 means appending a system message with empty content and a new output_config.effort. The new level applies from the next user turn, and everything before it stays cached.
How to change effort without restarting the conversation
Per-message effort is in beta and needs the mid-conversation-output-config-2026-07-01 header. It also needs adaptive thinking: with between_tools, the same change returns a 400 error. When the independent check fails, Python appends the new effort setting before the next user message:
if escalate:
append_system([], output_config={"effort": ESCALATED_EFFORT}) # effort-only: accepted anywhere
messages.append({"role": "user", "content": (
f"The harness's independent check failed. Expected payout after adjustments: "
f"{_cents(result['expected_after_adjustments_cents'])}. Processor deposits for September (bank "
f"record): {_cents(result['processor_deposits_cents'])}. Residual: {_cents(result['residual_cents'])}. "
f"Recorded adjustments ({ids}) stay in the ledger. Investigate what the residual is, using the same "
f"tools, and submit an amended plan that contains only new adjustments.")})
A top-level effort change would restart the cache, since top-level effort is part of the cached prompt. The per-message form didn't: the first high-effort request read 8,012 tokens from cache and processed 4 uncached ones.
The $15.00 difference gives Claude a target, but not evidence for a correction. The gate still requires transaction IDs that Claude retrieved, and their net_cents must total -$15.00. A proposed -$15.00 adjustment that cites only TXN-50036 still fails because that line's net is -$164.00.
What did Claude find at high effort?
At high, Claude grouped the processor lines by payout and by fee_cents, then re-ran the internal check. Its next note added the fee lines to 12,586 cents. The dispute fee raised that total to 14,086 cents. Rivermark's check had left it out.
Its first amended plan hit that rule because it cited only the debit. The next one cited both dispute lines, PLAN-2 passed, and ADJ-002 recorded -$15.00 under POL-4.
Did the matched replay need high effort?
This experiment does not show that high was necessary. A separate replay continued from the same failure point with the same conversation history and Python message, but stayed at medium; it found the fee too.
The main run's six high-effort calls produced 2,763 output tokens (607 thinking) and cost $0.0484. The separate control's six medium-effort investigation calls produced 2,713 output tokens (628 thinking) and cost $0.0464, including the same gate rejection.
One final report call brought the control to 7 calls and $0.0615 in total. None of those calls or costs are included in the main run's 15 calls and $0.1190.
Both paths received the same failed-check message; only their effort differed. One replay cannot measure the size of any effort effect, but it does show that high was not necessary for this instance. The same effort guide reserves xhigh and max for cases where "your evals show a quality gain." Test high the same way before selecting it.
How to Verify the Final Reconciliation in Python
Final verification repeats 2 gate checks on purpose: evidence and write scope. The gate reviews a proposal before writing; final verification inspects what Python actually wrote, then adds the number and raw-file checks.
After ADJ-002, Python recomputed everything from the raw records, the approved adjustments, and the bank total:
checks = {
"numbers": adjusted == deposits,
"provenance": not provenance,
"raw_unchanged": hash_dir(self.raw) == self.hashes_before,
"write_scope": set(created) <= ALLOWED_OUTPUTS,
}
All four passed. Expected payout after adjustments was $3,236.14, matching the deposits. Both adjustments traced to retrieved lines, the raw source data stayed unchanged, and Python wrote only the approved entries.
Only then does the report step start. The application appends a message that sets effort back to medium, a short user turn, and a system message that swaps the tools:
append_system([{"type": "text", "text": REPORT_TEXT},
{"type": "tool_removal", "tool": {"type": "tool_reference", "name": "create_adjustment"}},
{"type": "tool_addition", "tool": {"type": "tool_reference", "name": "submit_report"}}])
The report is the last output, not the proof. Its follow-up suggestions still require human review. The recording below follows permissions, effort, checks, and cost through one Streamlit session.
Streamlit follows the reconciliation from start. Video by Author.
How Much Did the Claude Sonnet 5.5 Agent Cost?
The main reconciliation moved from medium to high, cost $0.1190 across 15 API calls, and took 70.0 seconds, including 69.0 seconds waiting on the API. The separate matched replay is not included. Every figure comes from response usage and the Claude Sonnet 5.5 rates.
For a broader cost breakdown, our Claude API guide covers prompt caching and batch processing.
How do you calculate Claude Sonnet 5.5 cache costs?
input_tokens counts only what came after the cache breakpoint, so total input is the sum of three fields, as the prompt caching docs linked earlier explain. Cache writes and reads have their own rates, and thinking tokens are already inside output_tokens:
cost = (
usage.input_tokens * 2.00 # uncached input only
+ cache_creation.ephemeral_5m_input_tokens * 2.50
+ cache_creation.ephemeral_1h_input_tokens * 4.00
+ usage.cache_read_input_tokens * 0.20
+ usage.output_tokens * 10.00 # includes thinking
) / 1_000_000
Across the reconciliation, Claude read 105,614 of 118,308 input tokens from cache (about 89%), and only 636 were billed as uncached input. The chart applies the four token rates to the measured usage.

Output tokens dominate the measured cost. Image by Author.
API Limitations and Production Considerations
Rivermark writes local adjustment records, so a production finance system still needs:
-
Local, fictional data. A real close needs authentication, audit logs, human approval of postings, and a data-retention review.
-
Beta features. The headers for per-message effort, tool changes, and thinking updates can change, so test them again before deployment.
-
Varying results. Claude Sonnet 5.5 rejects non-default temperature, so repeated attempts can differ. Test the pattern on your own data before relying on it.
Final Thoughts
We built a reconciliation agent that investigates with read-only tools, gets one write tool only after Python approves its plan, and finishes only when an independent check against bank deposits passes. Claude Sonnet 5.5 found the misplaced refund on its own, but it took that failed check to send it back to the $15.00 fee it had already read.
I would not generalize one fictional month to every close. What carries over is the method: hide the write tool until a plan passes, keep the bank records outside the model, require transaction evidence for every correction, and append tool or effort changes so the cache survives.
The independent check is the part I would keep even in a smaller version of this project. The effort change is the part I would test before trusting, for the reason covered in the effort section.
Swapping the read tools and the final check lets the same pattern handle data cleanup fixes, support refunds, or controlled document updates. My first extension would be a human approval step before each adjustment is written, since a real close needs one.
To practice the Anthropic API basics this build relies on, I recommend our Introduction to Claude Models course.
FAQs
Does this workflow work on Amazon Bedrock or Google Cloud?
Not unchanged. Claude Sonnet 5.5 and mid-conversation system messages are available on the Claude API, Amazon Bedrock, and Google Cloud. This build also uses per-message effort, which Anthropic currently documents on the Claude API and Google Cloud, not Bedrock. It sends the Claude API's inline-tools-2026-09-15 header; reference-based tool changes on Bedrock and Google Cloud use mid-conversation-tool-changes-2026-07-01.
When should tool_addition define a tool inline?
Define the tool inline when it was unknown at the first request, or when its schema changes later. Keep at least one tool visible from the start, or the first inline definition causes a full cache miss.
Does changing Claude Sonnet 5.5 effort reset the prompt cache?
A top-level effort change starts the cache over because it changes the request's prompt prefix. The per-message output_config used here leaves earlier messages unchanged, so the cached prefix remains available.
What happens if the independent check fails twice?
The first failure sends Claude the remaining difference and opens 1 more investigation step. A second failure stops the process instead of allowing more writes or accepting a final report.
Should every Claude Sonnet 5.5 agent start at medium effort?
No. Anthropic suggests medium for clearly defined tool tasks, medium or low for chat that needs fast replies, and high otherwise. The levels changed from Claude Sonnet 5, so evaluate them again for your workload.
I’m a data engineer and community builder who works across data pipelines, cloud, and AI tooling while writing practical, high-impact tutorials for DataCamp and emerging developers.


