Calculator

What the interface tax costs you.

Put in your volume. Get tokens, dollars, energy and time — computed from a benchmark we ran, with its provenance and its limits on the same page. Nothing here is a point estimate pretending to be a fact.

Your numbers
Tasks where the agent actually transacts — quotes, bookings, orders, comparisons that end in a commitment.
How your agents work todayLooks at the website the way a person does. Renders pixels, reads them back with a vision model, clicks and types.
Energy conversion band
Watt-hours per token is not a measured constant. We publish all three bands rather than pick the flattering one. The band moves only the energy figure; the dollars are identical in all three, which we would rather say than let you assume otherwise.

Hypothetical pricing scenarioThis calculator used to leave our own invoice out of it, which made a number called “net” a gross figure. It is subtracted now.
Our own sweep answered 59 of 120 intent slots, so 49% is the default rather than 100%. Raise it if your counterparties are better than the median.
Planning model only: paid quotas and billing are not available. One discovery, one read and one bounded query count as three requests in this scenario.
Sensitivity caseThe table below used to be decoration: it published cases down to 1.5× that the calculator could not compute. Now it can.
Annual saving, moving to a machine surface
$10,770Modelled net after hypothetical fee$16,758 saved, less $5,988 for Operator Pro
764 kWhEnergy avoided0.28 UK homes for a year · 0.153 t CO₂e (modelled)
26.93BInput tokens avoided14.9× fewer
3.2 daysInterface time avoidedLoopback only — excludes model time
Read this before you quote the number above

Net of what you would pay us on the Operator Pro tier. 255,000 tasks stay on your current path and keep costing what they cost.

On Operator Pro this only pays for itself above 178,661 tasks a year at your coverage. Below that, do not buy it.


PathTokens / yrCost / yrkWh / yr
Screenshot agent (computer use)58.91B$41,0501,920
Machine surface (intent-native, over MCP)1.94B$3,357176
Honest reading

This is one synthetic task on one deliberately lean site, scaled linearly to your volume. Real tasks vary, real sites are heavier than our test site (which makes the gap larger), and real agents retry (which makes it larger again). We have not measured any of that, so we have not added it.

Read every limitation →

Prove it on your own traffic

The calculator is a model. The pilot is a measurement. Thirty days, capped, with kill criteria you set.

Book a 30-day pilot
Sensitivity

Every assumption, moved one at a time.

Including the two cases where the machine surface loses. A surface only wins if it stays lean; ship a bloated tool list and the advantage disappears.

CaseScreenshot ÷ machine ($ / Wh)DOM agent ÷ machine ($ / Wh)
Baseline: oracle agents, image token = text token energy6.0× / 5.3×3.2× / 2.9×
Vision energy weight x1.5 (screenshot only)6.0× / 6.5×
Vision energy weight x3 (screenshot only)6.0× / 10.0×
GUI agents take 1.4x the steps (OSWorld-Human low)8.5× / 7.6×4.3× / 3.9×
GUI agents take 2.7x the steps (OSWorld-Human high)18.7× / 17.2×8.0× / 7.5×
Both agents take 1.3x the steps6.4× / 5.6×3.4× / 3.0×
Machine surface agent takes 1.5x the steps, GUI agents perfect4.3× / 3.7×2.3× / 2.0×
GUI-favourable toolsets (no zoom; Playwright core tools only)5.9× / 5.2×2.4× / 2.2×
against usMachine surface ships 10k tokens of tool definitions1.9× / 1.9×1.0× / 1.0×
against usWorst case for us: GUI-favourable toolsets + 10k-token tools + 1.5x steps1.5× / 1.5×0.6× / 0.6×
Limits

What this benchmark does not show.

  • One synthetic task on one deliberately lean fictional site. Not a sample of the real web.
  • The machine surface was designed for this task (3 tools, small responses). That bias is in our favour and is stated.
  • Dollars, watt-hours and model latency are modelled, not measured. Only tokens and interface time are measured.
  • Text tokens use o200k_base as a proxy; production tokenizers differ, typically within about 20%.
  • Oracle policies favour the GUI agents wherever a choice existed.
  • Not measured: real LLM runs, real energy at the wall, real public websites, failure and retry rates.
  • Loopback interface time excludes model time and is not production latency.

machine-surface/bench@0.2 · measured 2026-09-16T02:19:51.194Z · 3 repeats · commit ec9c4ff · Chromium 141.0.7390.37 · Playwright MCP 0.0.81 · Node v22.22.2