What the interface tax costs you.
Put in your volume. Get tokens, dollars, energy and time — computed from a benchmark we ran, with its provenance and its limits on the same page. Nothing here is a point estimate pretending to be a fact.
Net of what you would pay us on the Operator Pro tier. 255,000 tasks stay on your current path and keep costing what they cost.
On Operator Pro this only pays for itself above 178,661 tasks a year at your coverage. Below that, do not buy it.
| Path | Tokens / yr | Cost / yr | kWh / yr |
|---|---|---|---|
| Screenshot agent (computer use) | 58.91B | $41,050 | 1,920 |
| Machine surface (intent-native, over MCP) | 1.94B | $3,357 | 176 |
This is one synthetic task on one deliberately lean site, scaled linearly to your volume. Real tasks vary, real sites are heavier than our test site (which makes the gap larger), and real agents retry (which makes it larger again). We have not measured any of that, so we have not added it.
Read every limitation →Prove it on your own traffic
The calculator is a model. The pilot is a measurement. Thirty days, capped, with kill criteria you set.
Book a 30-day pilotEvery assumption, moved one at a time.
Including the two cases where the machine surface loses. A surface only wins if it stays lean; ship a bloated tool list and the advantage disappears.
| Case | Screenshot ÷ machine ($ / Wh) | DOM agent ÷ machine ($ / Wh) |
|---|---|---|
| Baseline: oracle agents, image token = text token energy | 6.0× / 5.3× | 3.2× / 2.9× |
| Vision energy weight x1.5 (screenshot only) | 6.0× / 6.5× | — |
| Vision energy weight x3 (screenshot only) | 6.0× / 10.0× | — |
| GUI agents take 1.4x the steps (OSWorld-Human low) | 8.5× / 7.6× | 4.3× / 3.9× |
| GUI agents take 2.7x the steps (OSWorld-Human high) | 18.7× / 17.2× | 8.0× / 7.5× |
| Both agents take 1.3x the steps | 6.4× / 5.6× | 3.4× / 3.0× |
| Machine surface agent takes 1.5x the steps, GUI agents perfect | 4.3× / 3.7× | 2.3× / 2.0× |
| GUI-favourable toolsets (no zoom; Playwright core tools only) | 5.9× / 5.2× | 2.4× / 2.2× |
| against usMachine surface ships 10k tokens of tool definitions | 1.9× / 1.9× | 1.0× / 1.0× |
| against usWorst case for us: GUI-favourable toolsets + 10k-token tools + 1.5x steps | 1.5× / 1.5× | 0.6× / 0.6× |
What this benchmark does not show.
- One synthetic task on one deliberately lean fictional site. Not a sample of the real web.
- The machine surface was designed for this task (3 tools, small responses). That bias is in our favour and is stated.
- Dollars, watt-hours and model latency are modelled, not measured. Only tokens and interface time are measured.
- Text tokens use o200k_base as a proxy; production tokenizers differ, typically within about 20%.
- Oracle policies favour the GUI agents wherever a choice existed.
- Not measured: real LLM runs, real energy at the wall, real public websites, failure and retry rates.
- Loopback interface time excludes model time and is not production latency.
machine-surface/bench@0.2 · measured 2026-09-16T02:19:51.194Z · 3 repeats · commit ec9c4ff · Chromium 141.0.7390.37 · Playwright MCP 0.0.81 · Node v22.22.2