Skip to main content

OpenAI Agents API Computer Use Tutorial: Build a Browser QA Agent

Follow this OpenAI Agents API Computer Use tutorial to build a Python QA agent that tests a checkout in a hosted browser and retests a fix in the same session.
Oct 7, 2026  · 11 min read

Explore with AI

ChatGPTClaudePerplexity

Imagine a checkout that shows $48 in the cart and $24 on the review page. The customer sees two totals in the same checkout.

Teams usually run quality assurance (QA) tests on this flow with a browser script: click this button, open that page, check this value. A scripted test checks only the states its author wrote down.

An AI agent is a model that can take actions toward a goal. OpenAI's Agents API manages the agent loop and keeps its work in a session. In this tutorial, Computer Use also provides the hosted browser.

Northstar Checkout is a fictional test store with a hidden subtotal bug.

The agent receives the correct checkout result, but not the bug's location or a list of buttons to press. A small Python program, called the harness, compares the values the agent reports and then asks the same session to test the fixed store.

In this tutorial, I'll cover how to:

  • Create an Agents API session with Computer Use that can reach only the test site
  • Approve the browser's request to open that site, and refuse any other
  • Have your own code decide whether the test passed
  • Retest the fixed site in the same session and work out what the experiment cost

The code and measurements use version 3.22.1 of the openai Python package.

In a Nutshell

If you only have a minute, here are the key takeaways.

  • The buggy build failed only the review subtotal; quantity stayed correct.
  • The fixed build passed in the same session with no second origin approval.
  • The token counters produced a $0.9469 standard-rate estimate. Cache-write charges and hosted sandbox compute are not included, and Agents API usage is best-effort rather than a final invoice.
  • In each test, the API returned 2 screenshots, from 7 and 5 computer_use_call items respectively.

This is one test store with one planted bug, not a reliability benchmark.

What Is Computer Use in the OpenAI Agents API?

Computer Use is a tool in the OpenAI Agents API that lets an agent operate a browser running on OpenAI's servers. Your code follows the session's events and answers its requests. OpenAI lists website testing as one use.

OpenAI manages the agent loop, the session, and recovery. Our OpenAI Agents API tutorial covers those basics.

Older computer use setups, like the one in our GPT-5.4 computer use tutorial, make developer code run the screenshot-and-action loop instead.

Cover graphic for the OpenAI Agents API Computer Use tutorial: a QA task flows through a hosted browser to Northstar Checkout and a record_qa_result call, producing FAIL on build ns-1041 and PASS on build ns-1042 in the same session

Why use Computer Use for browser QA testing?

In browser QA, the page itself is the thing under test.

Calling a checkout API directly would skip the page where Northstar's bug lives, so the agent follows the same path a customer would, from the product page to the cart, checkout, and review.

Architecture diagram showing a Python harness sending a QA task to an Agents API session running GPT-6 Astra, which operates an OpenAI-hosted browser against Northstar Checkout while events, approvals, screenshots, and function calls return to the harness

Harness, session, hosted browser, staging site. Image by Author.

OpenAI manages the session and browser inside the gray zone; the harness and Northstar remain outside it.

What Will We Build With Agents API Computer Use?

The project is a fictional staging store, a Python harness, and one Agents API session.

The complete code is in this GitHub repository.

The Northstar Checkout test case

Northstar sells one $24 Trail Bottle. The test moves from product to cart, checkout, and review; there is no shipping, tax, login, or working purchase button.

Northstar Checkout staging product page showing the Trail Bottle at $24 with an Add to cart button and an empty cart

Northstar product page before the test. Image by Author.

Build ns-1041 contains the bug, while ns-1042 contains the fix. Adding ?reset=1 to a build's start URL empties the cart before either test.

The QA request is written as a goal. Its acceptance criteria ask the agent to:

  • Find the Trail Bottle and put 2 in the cart
  • Check that the cart subtotal is $48.00
  • Continue to the order review page and check that the quantity and subtotal still match
  • Report only values visible in the browser

A separate safety constraint says never to place, submit, or pay for an order. The request defines the outcome, not the clicks.

The planted checkout bug

The buggy build adds up unit prices on the review page and forgets the quantity. Both pages show quantity 2, but the cart subtotal is $48.00 and the review subtotal is $24.00.

The answer key lives in the application code. Neither the instructions nor the task message mentions the bug.

How application code decides pass or fail

The agent reports the build ID and 4 observed values through one function tool, record_qa_result.

The harness first checks that the reported build is the one under test, since both builds share a hostname, then compares the values with the answer key.

A function tool only runs if the agent calls it. A missing record, missing value, or wrong build makes the result incomplete, which never counts as a pass.

Flow diagram showing a QA objective leading to a browser test, a build check, four observed values, the record_qa_result function call, and a pass, fail, or incomplete verdict from application code

From QA objective to application verdict. Image by Author.

How to Set Up OpenAI Agents API Browser Testing

You need Python, a scoped API key, GPT-6 Astra access, and one session with Computer Use.

Prerequisites for Agents API Computer Use

  • Python 3.10 or newer and openai==3.22.1 (the SDK sends the OpenAI-Beta: agents=v1 header for you)
  • An API key with the api.agents.read, api.agents.write, and api.responses.write scopes, on a project that can use gpt-6-astra

The Agents API is in public beta, so field names and behavior can change between SDK releases. The repository pins version 3.22.1 in requirements.txt.

The hosted browser needs a reachable URL, so the code uses a Vercel deployment of Northstar.

git clone https://github.com/KhalidAbdelaty/OpenAI-Agents-API.git
cd OpenAI-Agents-API
python -m venv .venv
.venv\Scripts\Activate.ps1
pip install -r requirements.txt
Copy-Item .env.example .env   # then add your OPENAI_API_KEY
python run_qa.py

For more on isolated dependencies, see our virtual environment guide. On macOS or Linux, activate with source .venv/bin/activate and copy the file with cp. Keep the key in .env, never in code.

The experiment uses GPT-6 Astra, the model in OpenAI's Computer Use examples. Our GPT-6 Astra overview covers the model itself.

The code uses the Agents API (client.beta.agents), not the Agents SDK or the Responses API computer tool used in our GPT-6 Astra API tutorial.

Configure a Computer Use session

Create one session with the computer_use tool and an OpenAI-hosted desktop, then reuse it for both tests:

session = client.beta.agents.sessions.create(
    agent={"model": MODEL, "instructions": INSTRUCTIONS,
           "reasoning": {"effort": REASONING_EFFORT},  # "medium", set explicitly
           "tools": [{"type": "computer_use", "include_screenshots": True}, RECORD_QA_RESULT]},
    environment={"type": "openai_hosted", "desktop": {"enabled": True},
                 "network": {"access": "restricted", "allowed_domains": [host]}},
    metadata={"experiment": "northstar-browser-qa"},
)

include_screenshots: True exposes any screenshots the API returns, while restricted network access limits the browser to Northstar.

The environment uses the default medium size (2 vCPU, 4 GB RAM).

Add a function tool for QA results

The function records what the agent observed. If the agent cannot read one of the 4 checked quantity or subtotal values, it must report that field as null.

Listing every property under required tells the model to answer all of them, using null for anything it did not see. The harness still treats a missing field as incomplete:

"properties": {
    "build_id": {"type": "string", "description": "Build id shown on the page."},
    "stage_reached": {"type": "string", "enum": ["product", "cart", "checkout_details", "review"]},
    "cart_quantity": {"type": ["integer", "null"]},
    "cart_subtotal": {"type": ["string", "null"], "description": "Exactly as displayed, e.g. $10.00"},
    "review_quantity": {"type": ["integer", "null"]},
    "review_subtotal": {"type": ["string", "null"], "description": "Exactly as displayed"},
    "purchase_control": {"type": "string", "enum": ["disabled", "absent", "enabled", "not_seen"]},
    "evidence_note": {"type": "string", "description": "One or two sentences on what you saw."},
},
"required": ["build_id", "stage_reached", "cart_quantity", "cart_subtotal",
             "review_quantity", "review_subtotal", "purchase_control", "evidence_note"],
"additionalProperties": False,

The harness converts each displayed price to cents, verifies the build ID, and compares the values with the answer key:

EXPECTED = {"cart_quantity": 2, "cart_subtotal_cents": 4800,
            "review_quantity": 2, "review_subtotal_cents": 4800}

def judge(record, expected_build):
    observed = {
        "cart_quantity": record.get("cart_quantity"),
        "cart_subtotal_cents": to_cents(record.get("cart_subtotal")),
        "review_quantity": record.get("review_quantity"),
        "review_subtotal_cents": to_cents(record.get("review_subtotal")),
    }
    missing = [field for field, value in observed.items() if value is None]
    if record.get("build_id") != expected_build:
        return {"verdict": "incomplete", "observed": observed, "failed_checks": [],
                "missing": [f"build_id={expected_build}", *missing]}
    if record.get("stage_reached") != "review":
        missing.append("stage_reached=review")
    failed = [{"field": field, "expected": EXPECTED[field], "observed": value}
              for field, value in observed.items()
              if value is not None and value != EXPECTED[field]]
    verdict = "fail" if failed else "incomplete" if missing else "pass"
    return {"verdict": verdict, "observed": observed, "failed_checks": failed, "missing": missing}

An unreadable or missing value produces an incomplete verdict, never a pass.

A report from the wrong build returns incomplete before its values can affect the verdict.

Write the QA instructions

The same instructions govern both tests:

INSTRUCTIONS = (
    "You are a QA tester for the Northstar Checkout staging site. "
    "Use the browser to run the test you are given. "
    "Stay on the approved staging origin and do not visit any other website. "
    "Inspect what is visible on a page before you make any claim about it. "
    "Stop before any purchase: never place, submit, or pay for an order. "
    "Never invent an observed value. If you could not see a value, report null. "
    "Call record_qa_result once, only after the browser test is finished, then give a short summary."
)

Only the website build changes between tests.

How to Run a Browser QA Test With Computer Use

Open the event stream, send the QA objective once, then handle approvals and function calls until the turn completes.

Send a QA task to the Agents API session

Open the event stream first, then send the task exactly once:

with self.client.beta.agents.sessions.events.stream(self.session_id) as events:
    if not sent:  # open the stream first, then send the task exactly once
        self.client.beta.agents.sessions.events.create(self.session_id, events=[message(text)])
        sent = True
    else:  # reconnected: act on what is still pending, never resend the task
        yield from self.handle_required_actions()
    for event in events:
        yield from self.handle(event)

Streams do not replay missed events. If the stream drops, open a new one, then retrieve the session and its saved items while it stays connected.

The task message names the build, the acceptance criteria, and the safety constraint, but nothing about the bug:

QA objective for Northstar Checkout staging build ns-1041. Start at https://northstar-checkout-staging.vercel.app/b/ns-1041/?reset=1
Scenario: a customer adds 2 Trail Bottles to the cart and continues through checkout to the order review page.
Acceptance criteria:
- The cart shows quantity 2 and a subtotal of $48.00 (unit price $24.00, no shipping or taxes).
- The order review page shows the same quantity and subtotal as the cart.
Safety constraint: never place, submit, or pay for an order.
Record the cart values and the review values as separate fields.

Keep the session ID for the retest.

Handle browser origin approval

The hosted browser asks for approval before opening each new website origin.

The stream emits agent.session.requires_action; retrieve the session and read required_actions for the request.

def answer_approval(self, action):
    request = action.request
    if request.type == "browser_origin_access":
        decision = "approve" if request.origin.rstrip("/") == self.origin else "deny"
        response = {"type": "browser_origin_access", "decision": decision}
    else:  # browser_authentication: Northstar has no login, so sign-in is refused
        response = {"type": "browser_authentication", "action": "cancel"}
    self.client.beta.agents.sessions.events.create(self.session_id, events=[{
        "type": "agent.session.input.computer_use_approval_request_result",
        "request_id": action.request_id, "response": response}])

Track browser activity with session events

Browser work shows up as computer_use_call items, each with a short title and a status. The event stream for the first test showed:

   12.4s  turn     sent       build=ns-1041
   59.4s  browser  completed  Connecting to the staging test browser
   63.6s  browser  completed  Connecting to the staging test browser
   68.6s  approval approve    https://northstar-checkout-staging.vercel.app
   70.8s  browser  completed  Inspecting the Trail Bottle product
   73.5s  browser  completed  Adding the first Trail Bottle
   78.2s  browser  completed  Checking cart quantity and subtotal
   85.7s  browser  completed  Continuing to checkout details
   89.2s  browser  completed  Checking order review values
   95.9s  record              cart 2 $48.00, review 2 $24.00, purchase disabled

About 47 seconds passed before the first browser activity.

All 7 computer_use_call items completed, but an item status is not the QA verdict; the function result is.

Did the Agent Catch the Checkout Bug?

Yes. More importantly, the function call isolated the failure to one field: the review subtotal.

What GPT-6 Astra reported

The record_qa_result call contained:

{
  "build_id": "ns-1041",
  "cart_quantity": 2,
  "cart_subtotal": "$48.00",
  "review_quantity": 2,
  "review_subtotal": "$24.00",
  "stage_reached": "review",
  "purchase_control": "disabled"
}

Every value matches the buggy page. Quantity stayed 2 on the review page, which rules out a visible quantity mismatch.

How the harness turned the report into a fail

judge() confirmed build ns-1041, compared the 4 values with the expected values, and found only the review subtotal wrong.

This is the only verdict the experiment uses:

{
  "verdict": "fail",
  "failed_checks": [{"field": "review_subtotal_cents", "expected": 4800, "observed": 2400}],
  "missing": []
}

Retest a Fix in the Same Agents API Session

After the fix is live, send one more message to the same session.

This small regression test uses the same instructions and verdict function.

Ship the fix without changing the test

The fix in build ns-1042 is one line of Northstar's JavaScript:

-const reviewSubtotal = (cart) => cart.reduce((sum, line) => sum + line.unitCents, 0);
+const reviewSubtotal = (cart) => cart.reduce((sum, line) => sum + line.unitCents * line.qty, 0);

Send the follow-up in the same session

The start link includes ?reset=1, so the retest begins with an empty cart. Then the follow-up goes to the same session:

A fix is deployed as staging build ns-1042 at https://northstar-checkout-staging.vercel.app/b/ns-1042/?reset=1
That link starts from an empty cart. Run the same QA objective and acceptance criteria against this build from the start of the journey, and record a new result.

The retest kept the hosted environment and required no new origin approval. Do not depend on browser state, because cookies can expire and recycling the environment clears it.

Diagram of one Agents API session and one hosted environment spanning two turns, with build ns-1041 failing in turn 1, a one-line fix shipping as ns-1042, and turn 2 passing without a new origin approval

One session carried both QA tests. Image by Author.

A hosted sandbox can be deleted if activity and keep-alives stop for 1 hour. Watch for agent.session.environment.reset and start each retest from a known state.

Did the retest pass?

Yes. The retest reported cart quantity 2 and $48.00, then review quantity 2 and $48.00, and judge() returned a pass with no failed checks.

It took 38.9 seconds with 5 browser activity items, against 96.5 seconds and 7 items for the first test, which included a 47-second wait before browser activity began.

Terminal output of the retest on build ns-1042 showing five completed browser activity items, no origin approval line, the record_qa_result values, and a PASS verdict

Retest passed without a new approval. Image by Author.

Does Computer Use Return a Screenshot for Every Activity?

Not necessarily. Even with include_screenshots set, the first test returned 2 screenshots from 7 browser activity items, and the retest returned 2 from 5.

Some items return output: null, so reports cannot assume a picture for every activity.

The event stream is not a continuous video feed of the hosted browser; it returns browser activity items and screenshots when available.

Northstar uses rrweb to capture Document Object Model (DOM) changes and interactions, send them to the same host, and replay both journeys below.

The agent's browser on both staging builds. Video by Author.

The replay shows quantity 2 and $24.00 on ns-1041, then $48.00 on ns-1042; the disabled purchase button remains untouched.

The repository also includes a small Streamlit viewer for the saved verdict, browser evidence, session details, cost, and event log.

How Much Did the Agents API Computer Use Test Cost?

The best-effort usage counters produced a $0.9469 standard-rate token estimate for both tests.

Token usage for the 2 tests

Metric Test 1 (ns-1041) Retest (ns-1042)
Input tokens 255,550 223,533
Cached input tokens 217,041 (84.9%) 219,449 (98.2%)
Output tokens 982 708
Estimated token cost $0.6512 $0.2957
Turn time 96.5 seconds 38.9 seconds
Browser activity items 7 5

The retest used fewer input tokens, and 98.2% of them came from the prompt cache. Together, the 2 tests cost $0.9469.

The observability guide says usage can be null when unknown and that recorded counts may change, so check them again before deleting the session.

What Agents API usage numbers leave out

When I ran the tests, these were the standard GPT-6 Astra rates on OpenAI's pricing page:

Token type Rate per 1M tokens
Input $10.00
Cached input $1.00
Cache writes $12.50
Output $50.00

The 272K long-context threshold applies per request. Both turns' combined input stayed below it, so no single request could have triggered the higher long-context rates.

The estimate still cannot reproduce the final invoice because Agents API usage is best-effort and does not expose separate cache-write counts.

The hosted sandbox is billed separately at standard container rates. The pricing page lists the 4 GB medium container at $0.12 per 20-minute session, with eligible container sessions billed by the minute and a 5-minute minimum.

How to Keep Agents API Computer Use Tests Safe

Safety depends on what the browser can reach and what the page lets it do.

Diagram of three safety layers between the agent and a purchase: a restricted network policy with exact hostnames, origin approval handled by the harness, and a disabled Place order button in the staging page

Three layers between agent and checkout. Image by Author

What origin approval covers in Computer Use

Network policy controls which hosts the browser can reach, and origin approval decides whether it may open each new origin. Neither one confirms individual browser actions.

Approving northstar-checkout-staging.vercel.app therefore does not approve each click separately.

The no-purchase rule is a safety constraint, and purchase_control is saved as evidence rather than judged as an acceptance criterion. Northstar's disabled "Place order" button is the control that enforces it.

How network policy limits the hosted browser

Under restricted, the browser can reach only the hostnames you list.

OpenAI's sandbox guide accepts 1 to 100 exact hostnames, with no wildcards, protocols, paths, or ports. Content delivery networks (CDNs), subdomains, and redirect targets need separate entries.

How to handle screenshots and session data

Screenshots and rrweb recordings contain whatever the page shows, so Northstar uses fictional data, has no login, and discloses recording in its footer.

The recorder masks inputs, but a production deployment would still need a data policy and masking suited to the page.

The Agents API supports data residency only in the United States and is not eligible for Zero Data Retention (ZDR), even with a self-hosted sandbox.

Save the results and screenshots you need, then delete the session instead of leaving a staging checkout in retained session state.

Deleting the Agents API session does not delete rrweb recordings stored by the site. Remove those separately according to the recording policy.

Final Thoughts

Northstar failed when the cart and review subtotals split, then passed after the fix in the same session. The harness, not the model's summary, decided both verdicts.

I would keep scripted regression tests for known invariants and use goal-based browser agents for exploratory journeys that are harder to express as an assertion. The agent explores; application code decides.

For API basics, I recommend our Working with the OpenAI API course.

FAQs

Is Computer Use in the Agents API generally available?

No, it ships as part of the Agents API public beta, and every request carries the OpenAI-Beta: agents=v1 header. Pin the SDK version you test with, because event names and fields can still change before general availability.

Does a high cached-input share mean the retest saved money?

Not by itself. The observability guide says a high cached-input percentage doesn't measure savings on the total task cost, since cached input is still billed and repeated calls can reprocess a large history.

Does one origin approval cover later session turns?

It did here: the retest raised no new request. Keep the approval handler running on every turn and never assume a site is still approved.

Why does your listener never see agent.session.action_required?

That name belongs to the webhook. On the event stream, the pause arrives as agent.session.requires_action. Handle it through the same required-action flow used for origin approval.

What if the agent calls record_qa_result twice in one turn?

The harness keeps the last call, which is fine for a read-only check. If your function writes anywhere, store each result by session, turn, and call id, and check for an earlier result before acting twice.


Khalid Abdelaty's photo
Author
Khalid Abdelaty
LinkedIn

I’m a data engineer and community builder who works across data pipelines, cloud, and AI tooling while writing practical, high-impact tutorials for DataCamp and emerging developers.

Topics
Artificial Intelligence
Large Language Models
OpenAI

Top DataCamp Courses

Course

Working with the OpenAI API

3 hr
179.5K
Start your journey developing AI-powered applications with the OpenAI API. Learn about the functionality that underpins popular AI applications like ChatGPT.
See DetailsRight Arrow
Start Course
See MoreRight Arrow
Related

Tutorial

OpenAI AgentKit Tutorial With Demo Project: Build an AI Agent

Learn how to use OpenAI AgentKit to automate GitHub issue reviews with AI agents that classify, detect duplicates, and summarize reports.
Aashi Dutt's photo

Aashi Dutt

6 min

Tutorial

OpenAI Agents API Tutorial: Build an Agent That Writes and Runs Code in the Cloud

Build and run a cloud agent with the OpenAI Agents API that can analyze files, execute code, verify results, and return finished artifacts from a single request.
Abid Ali Awan's photo

Abid Ali Awan

8 min

Tutorial

OpenAI Agents SDK: How to Run Agents in Modal Sandboxes

Learn how to build an OpenAI agent app that runs inside Modal Sandboxes, works with files, executes code, and returns results in this hands-on Python tutorial.
Abid Ali Awan's photo

Abid Ali Awan

10 min

Tutorial

OpenAI Codex CLI Tutorial

Learn to use OpenAI Codex CLI to build a website and deploy a machine learning model with a custom user interface using a single command.
Abid Ali Awan's photo

Abid Ali Awan

9 min

Tutorial

OpenAI Codex: A Step-by-Step Guide With 3 Practical Examples

Learn how to use OpenAI's Codex software engineering agent to fix bugs, explain code, and generate pull requests directly from ChatGPT.
Aashi Dutt's photo

Aashi Dutt

7 min

Tutorial

Kimi Browser Extension Tutorial: Automate Web Browsing With AI Agents

Learn how to set up the Kimi Browser Extension using both Kimi Work and Kimi Code CLI, and run a product search that extracts structured results from a website.
Aashi Dutt's photo

Aashi Dutt

8 min

See MoreSee More