Test a Computer Use API Controller Before Connecting a Model
Run 19 offline tests for computer-action validation, screenshot call IDs and form completion. Includes Python source and an unvalidated live adapter.
A Computer Use integration has two independent questions: can the model return supported actions, and can your application execute them without mistaking a success message for a completed task? This tutorial addresses the second question first. The download contains a Python controller and 19 offline tests you can run without an API key, browser session or model charge.
Download the controller kit. Its tests passed with Python 3.9.6, including a repeat from a freshly extracted archive. The included run_live.py adapter has not been executed or validated against a real provider. There are no real API screenshots, usage measurements or endpoint-compatibility results in this article.
Run the tests before choosing an endpoint
Extract the ZIP, enter api-kit, and run:
python3 -B -m unittest discover -s . -p 'test_*.py' -v
Python 3.9 or later and the standard library are sufficient for these offline tests. There are 15 controller cases and four receiver cases. The fake runtime and synthetic model responses do not open a browser or send a request. Their screenshot placeholders are deliberately not real image files; never send them to a provider.
| File | Purpose |
|---|---|
controller.py | Validate action batches, return observations and stop on verified completion |
test_controller.py | Exercise controller behavior with synthetic responses and a fake runtime |
test_receiver.py | Check exact receiver-state changes using in-memory records |
run_live.py | Reference adapter for later Playwright and HTTPS Responses integration; unvalidated live |
README.md | Offline commands and requirements for a separate live integration test |
Separate the model, runtime and verifier
The model chooses an action from an observation. The runtime owns the browser and executes supported actions. The verifier inspects an application result independently of the model. In this lab, that result is exactly one new form record containing a unique synthetic address ending in @example.test.
Use the form from the landing-page QA lab. Read /submissions before starting and again during execution. The old record prefix must be unchanged and exactly one matching record must have been appended. An old record, a duplicate or an unrelated address does not pass. A green success banner also does not pass by itself.
The receiver verifier is specific to this fixture. A production workflow needs a result check appropriate to its own system, such as a saved draft ID. Do not generalize this form rule to arbitrary websites.
Follow one computer-tool contract
The reference code follows the structured-action path described in the OpenAI Computer Use guide. The controller sends an initial screenshot, accepts an ordered computer_call action batch, executes supported operations, and returns a screenshot as computer_call_output with the matching call_id. The next request also carries previous_response_id.
Those identifiers connect an observation to the action request it answers. A screenshot without the correct call ID is not a substitute for tool output. The source is a protocol reference, not proof that the unexecuted adapter is compatible with a specific model today.
Before live use, verify the exact endpoint, model and tool schema together. A text request succeeding on an OpenAI-compatible endpoint does not establish computer-tool support. This kit makes no claim of Ofox endpoint compatibility and selects no default provider or model.
Validate actions before execution
The controller supports a deliberately small action set: left click, bounded text input, selected single-key presses, bounded scroll and screenshot. Coordinates must fit the runtime’s screenshot dimensions; booleans, non-finite numbers and malformed payloads are rejected. Text is limited to 200 characters, a response to 12 actions, and the default run to four model turns. Unknown actions stop execution.
The controller accepts one computer call per response. It rejects repeated call IDs, incomplete responses and pending safety checks instead of silently accepting them. This is a teaching implementation; it does not implement every tool action or a complete approval system.
Completion is checked before and after each action. That matters when a submit action finishes asynchronously: the next action must not blindly click again after the record already exists. The reference browser adapter also polls the receiver briefly after click and keypress actions. None of these local checks proves that a real browser integration works; that requires a separate test.
Follow one action through the controller
The controller is the application code between a model response and a browser executor. Its job is to reject unsupported actions, preserve call correlation and stop when application evidence confirms completion. It is useful to developers building that boundary; it is not a ready-made production agent.
This simplified response is a synthetic teaching fixture, matching the local controller’s expected shape. The coordinates describe an imaginary 100 × 100 viewport, not a button location in the QA page. Do not paste them into a live task.
{
"id": "response_demo_1",
"status": "completed",
"output": [{
"type": "computer_call",
"call_id": "call_demo_1",
"actions": [{"type": "click", "button": "left", "x": 10, "y": 20}]
}]
}
The outer status indicates response generation finished. It does not mean the click ran or the form saved. The controller validates the action batch, asks the runtime to perform permitted actions in order, checks receiver state, and obtains another observation when needed. For a continuing call, the screenshot output uses call_demo_1; previous_response_id points to response_demo_1. Mixing those identifiers breaks the relationship between the model’s request and the resulting observation.
Run a small example without a model or browser
Save the following as walkthrough.py beside controller.py and run python3 -B walkthrough.py. It uses the actual controller with a deliberately tiny fake runtime. The fake marks the task complete after one action so that you can see the second queued click being skipped.
from controller import run
class DemoRuntime:
width, height = 100, 100
def __init__(self):
self.actions = []
self.done = False
def screenshot(self):
return b"offline-placeholder-not-a-real-png"
def assert_allowed(self):
pass # Fake only: a real runtime must enforce its allowed surface.
def complete(self):
return self.done
def perform(self, action):
self.actions.append(action)
self.done = True # Simulated outcome, not receiver verification.
runtime = DemoRuntime()
def transport(payload):
return {
"id": "response_demo_1",
"status": "completed",
"output": [{
"type": "computer_call",
"call_id": "call_demo_1",
"actions": [
{"type": "click", "x": 10, "y": 20},
{"type": "click", "x": 30, "y": 40}
]
}]
}
result = run(transport, runtime, "Synthetic controller walkthrough")
assert len(runtime.actions) == 1
print(result["status"], result["turns"], len(runtime.actions))
The locally verified output is verified 1 1: the controller returned its success status after one simulated turn and one performed action. Here, verified trusts DemoRuntime.complete(), which we intentionally made trivial. It does not prove a form was saved. Replacing that method with a real, independent acceptance check is essential before connecting a model. The placeholder bytes also cannot be sent as a real API screenshot.
Replace the fake completion rule with receiver evidence
In the companion browser adapter, the form’s acceptance rule is stricter: the previous records must remain unchanged, and exactly one additional record must contain the unique address for this run. Consider these illustrative cases:
| Before → after | Decision | Reason |
|---|---|---|
| Old record → same old record + current address | Accept | One matching addition |
| Old record → same old record | Continue checking or stop inconclusive | No persisted result yet |
| Old record → same old record + two copies | Reject | Duplicate submission |
| Old record → different old record + current address | Reject | Baseline was changed |
| Old record → same old record + another address | Reject | Result belongs to a different input |
The four local receiver tests cover a matching addition, an unrelated address, extra records and changed history using in-memory data. The no-change row describes the adapter’s continuation rule; it is not a fifth receiver test. A real integration also needs to handle time, concurrency and failures in the receiver read. This fixture is suitable for an isolated exercise, not a concurrent production submission system. The QA walkthrough explains how a person collects the same before/after evidence.
What the offline tests establish
The tests cover malformed and out-of-bounds actions, unsupported input, call/output correlation, turn limits, repeated calls, early stopping and exact receiver changes. Regression cases check that a verified completion stops later actions in the same batch and that duplicate or wrong-address records fail verification.
For completed responses with a valid response ID, the audit retains reported usage, including a final response without actions, and records attempted actions if execution fails. Synthetic usage is test data, not observed billing. Exceptions carry the accumulated controller audit where implemented; the code does not promise an artifact for every possible startup failure.
You can inspect each test input and expected result in the source. These checks establish local controller behavior for the listed cases. They do not establish model accuracy, visual quality, browser reliability or a real-world success rate.
Diagnose a stopped run from the first failed boundary
| Controller outcome | Likely boundary to inspect | Safe next step |
|---|---|---|
Coordinates outside current viewport | Action validation | Compare the action with the actual current screenshot dimensions |
Unsupported action type | Supported action subset | Inspect the requested type; do not silently execute an unimplemented action |
Missing or repeated call id | Response correlation | Inspect response/call IDs and retry history |
Safety check requires human review | Pending approval | Stop; this example does not implement an approval UI |
Model stopped before receiver verification | Completion | Inspect receiver evidence; model text is insufficient |
Turn limit reached | Bounded loop | Check whether a submission already happened before retrying |
The batch is validated before execution, so a later unsupported action stops the entire batch before its first action. Once execution begins, an earlier action can already have had an effect when a later runtime failure occurs. Audit entries distinguish an attempted action from one that returned successfully, but they cannot undo side effects. Read the receiver before restarting.
For browser connection problems, use the Computer Use permissions guide. For tasks without a form, define a different outcome: a research task might require a parseable evidence file with source URLs, as in the competitor-table tutorial.
What is required for a real run
Passing the offline suite establishes only the tested controller behavior. Choose the provider protocol and an independently checkable outcome before adding a live executor.
The unvalidated run_live.py adapter uses a fresh Playwright Chromium context, a loopback-only fixture origin and an explicitly configured HTTPS Responses endpoint. It does not attach your normal browser profile. It routes ordinary page requests through a fixture-origin allowlist and stops on an unexpected page origin or tab. This is a controlled-fixture safeguard, not a general browser security sandbox. It does not automatically retry failed model requests.
Live dependencies have not been installed and pinned as tested. Follow the Playwright installation documentation in a dedicated environment, then record the actual package and browser versions. Supply COMPUTER_RESPONSES_URL, COMPUTER_MODEL and COMPUTER_API_KEY through your environment; do not put a secret in the archive or repository.
Real execution needs permission for the browser environment and any API spending. Do not use this adapter to bypass an existing browser access denial. Use a new synthetic address and a new output directory for every attempt. The README shows the invocation shape with an explicit acknowledgement flag.
Four calls and an output-token cap are not a dollar budget. Confirm current provider rates and an account-side spending limit separately. A timeout can still correspond to billed work; inspect provider usage and receiver state before deciding whether to retry.
Accept a live integration only with its own evidence
For an eventual live test, preserve the model ID, endpoint host, runtime versions, task, redacted raw responses, call IDs, screenshots and receiver result. Review logs before sharing. Record failures as failures, including unsupported tools and incomplete tasks; do not silently switch to text-only output and call it Computer Use.
Run a second attempt from a clean environment with a different synthetic address. Report both outcomes and observed usage separately. Until those runs exist, the defensible result is narrow but useful: the downloadable controller passed 19 offline tests, and the provider/browser connection remains unverified.
Frequently Asked Questions
- Does this include a completed browser test?
- No. The controller and receiver checks passed 19 offline tests. The live adapter has not been executed or validated against a provider, and the tests establish neither browser operation nor endpoint compatibility.
- Can I run the controller tests without an API key?
- Yes. The included standard-library tests use synthetic model responses and fake runtimes. They do not call a provider or launch a browser.
- What does the verified status mean?
- It means the supplied runtime reported completion. In the walkthrough that result is simulated; in a real integration completion must be checked against independent application evidence.
- Does limiting turns guarantee a spending budget?
- No. A turn limit bounds loop iterations, not money. Live use also requires verified rates, usage accounting and account-level spending controls.


