Evidence · swept 2026-09-16

What the real freight web costs an agent.

Our benchmark measured one synthetic task on one deliberately lean site we built ourselves. That is the most attackable thing we own, so we measured 35 real, named freight and logistics sites instead — twice, by two different methods, because the first method turned out to be measuring the wrong thing.

49.5×Cheaper, measured head to head91,455 tokens of screenshots against 1,846 of bounded queries, for the same forty questions
22 vs 17Questions answered, of 40Structure against vision, same ten sites. Cheaper is the smaller half of it
4/30Fully operableQuote, track, contact and book all reachable structurally
0Genuinely closed to an agentTen looked closed at first. Nine of them opened; the tenth we never knocked on, because its robots.txt says not to

Correction · the headline was 148×, and it was wrong

We charged one side for reading everything and the other for asking one question.

The old vision baseline was not measured. It was assumed: page height divided by viewport height, on the premise that an agent screenshots a whole page before it can find a link. Our own side was charged for one targeted query. That is not a comparison, and we had already written the same criticism about someone else’s number.

So we measured it. Ten independent vision agents were given the actual captured screenshots of ten of these sites, one at a time in scroll order, and asked to name the visible element for each of four intents. They read 67 screens — a median of 7 per site — and answered 17 of 40.

The real ratio is 49.5×, not 148×. The assumption was doing roughly fortyfold of the work, and it was ours.

There is a second, harsher number. If you insist the whole structured read must pass through the model — it does not, the transport builds it and answers from it — the ratio is 3.7×. Both are in the table below, because which one is right depends on an architectural boundary and not on us.

The result we did not expect is the one we would now lead with anyway: structure answered 22 of 40 against vision’s 17, and there was not a single question vision could answer that structure could not.



The vision arm

Ten agents, 67 screenshots, no HTML.

Each trial was run by a separate vision agent that received ONLY 1280x800 viewport screenshots of the site, one at a time in scroll order, and was asked after each whether it could yet name a specific visible clickable element for each of four intents. It had no access to the HTML, the accessibility tree, or the network. It was instructed that an unsure answer counts as not found, that a cookie banner or a 'Book a Demo' button satisfies nothing, and that a collapsed menu it cannot see inside does not count. Agents stopped early once all four intents were found.

SiteScreens readVision tokensVision answeredQuery tokensStructure answered
Loadlink Technologies
loadlink.ca
45,4600/41091/4
Shiply
shiply.com
912,2852/41722/4
Zencargo
zencargo.com
79,5550/4520/4
DSV
dsv.com
11,3652/42684/4
UPS
ups.com
22,7303/42953/4
Pall-Ex
pallex.co.uk
45,4603/42033/4
XPO
xpo.com
912,2853/42054/4
MSC
msc.com
79,5552/41362/4
Wincanton
wincanton.co.uk
1115,0151/41451/4
project44
project44.com
1317,7451/42612/4
Total6791,45517/401,84622/40
Where they disagreed

Of 40 questions, both paths found the affordance 17 times and neither found it 18 times. Structure found 5 that vision missed. Vision found 0 that structure missed.

That last number started at three. Vision was the thing that caught them, and fixing them is how the matcher lost its third round of false positives — a prose headline counted as a tracking affordance, and a “Pricing” link whose target was #.

Every way to build the ratio
  • Measured, charging the model for the whole structured read3.7×
  • Measured, same but the median site rather than the total3.65×
  • Measured, queries only — what the model actually receives49.5×
  • Assumed full-page vision coverage (what we used to publish)183.4×
  • Assumed coverage, and counting queries that found nothing945×
  • One screenshot only, against a bounded query20.4×

We publish the one marked measured, queries only. The two largest numbers on this list are the ones we used to publish, and they rest on the assumption this arm replaced.


Correction to our own published result

We said a third of the freight web was closed to agents. Most of that was our fault.

The first sweep ran a headless browser advertising itself as headless, and did not dismiss cookie walls. On that basis ten of 35 sites gave us nothing, and we published that as a finding about the industry.

Re-running with a normal Chrome identity and consent handling, 6 of those sites served us. Maersk, MSC, DHL, FedEx, UPS and Forto were not closed to agents. They were closed to the specific way we were knocking.

One qualifier on that sentence, because a reviewer caught us overstating it: FedEx served a country-selector interstitial rather than its homepage. It let us in; it did not let us in to anything. Five of the six is the honest claim.

Of 35 rows, 30 were read. 2 are excluded by us — a duplicate host after an acquisition, and a site whose robots.txt says not to knock. 3 refused a datacentre crawler outright, and every one of those later served an entitled browser. So the number of sites that are closed to an agent, as opposed to closed to an anonymous crawler, is zero. That is a much less dramatic number than the one we first published and it is the correct one.

We did not add stealth. No navigator.webdriver masking, no anti-detection, no evasion — that is a losing arms race and we are not playing it. We stopped announcing ourselves as a robot and started handling cookie banners like any browser does. Both identities were run so the difference is published rather than quietly absorbed.

Served us only after we stopped presenting as headless
Maersk
maersk.com
now readable
MSC
msc.com
now readable
DHL
dhl.com
now readable
FedEx
fedex.com
now readable
UPS
ups.com
now readable
Forto
forto.com
now readable

These observations do not establish why an edge refused access or whether a caller was authorised. Browser access and permission to automate are separate questions.


Third arm · the entitled browser

Refused a crawler; served a browser.

Some sites still refused a datacentre crawler even with a normal browser identity. So we ran the identical extractor inside an ordinary signed-in browser on a home connection. All 4 of them served it. That is a statement about what one signed-in visit saw on one day, not about those companies’ security posture, and one of the four is reported in aggregate rather than by name because its edge had presented a challenge page to the crawler and we do not publish a named row that reads as “we got past your challenge”. We did not: a person’s own browser is not a bypass.

Four manually recorded trials reported access from the operator’s own browser after datacentre refusals. This is a limited historical observation, not proof that every site permits agents or that authenticated access is authorised for automation. The original MachineViews were not archived, so their intent classifications cannot be independently replayed.

SiteFrom a datacentreHistorical browser observationAffordancesIntentsQuery costWhat it found
Hapag-Lloyd
hapag-lloyd.com
HTTP 502served283/4~43Get a Quote · Track
One ocean carrier
named on request, see /contact
HTTP 403 / CAPTCHA wallserved684/4~98Price · Shipment Tracking
Palletways
palletways.com
ERR_TUNNEL_CONNECTION_FAILEDserved523/4~102Get a quote now · Track pallet
sennder
sennder.com
timeout after 30sserved662/4~82Logistics Tech article · Contact us
What we did not do

No stealth, no navigator.webdriver masking, no CAPTCHA solving, and nothing was clicked — not even a cookie banner, because accepting terms on someone's behalf is not ours to do. The extractor is read-only. The only variable that changed was where the browser was and whose session it carried.

What this costs us to admit

It retires the most dramatic number we had. There is no wall of closed doors in freight. What there is, is a wall against anonymous automated traffic — which is reasonable of them, and which means the answer is to run the agent where it is already entitled to be rather than to argue about it.

Token counts in this arm are estimated from measured character counts at 3.688 characters per token. The structural results are direct measurement.

Width of this arm, stated

This arm is 4 sites and one run: the sites that refused the datacentre sweep, visited once each. It is not the whole sample re-run from a home connection, and it does not show that a browser sees more than the crawler on sites that served both — only that the refusals were about provenance, not capability. The runner that produced it is committed as operator/local.mjs(CDP, isolated world, read-only, --signed-in opt-in) and its README gives the one command that re-runs it across the full sample; we have not published that run because it means driving a person’s own browser across thirty-five third-party sites, and we treat that as their decision, not ours. The structured views from the four-site run were not archived, so unlike the main sweep its per-intent classifications cannot be re-derived against the current matcher; the counts are a record of that run, not something this repository can replay.


Method, corrected

The first sweep measured the wrong quantity.

It compared raw HTML, the accessibility tree and a screenshot — three ways of swallowing an entire page. Raw HTML came out 56.4× heavier than our lean site, which sounds decisive, and the accessibility tree came out only 5.3× heavier, which undercut the whole argument.

Both were answering a question no competent agent asks. An agent does not swallow the page. It asks a bounded question: where do I get a quote here. Measuring that is the second arm.

The method that replaced it

Discover first. Read a structure. Ask a bounded question.

Before any browser starts, probe for a sealed surface, an OpenAPI document, a GraphQL endpoint, a feed or a sitemap. If a structured API exists, wrapping it in a page read is pure overhead — route to it and launch nothing.

Where none exists, the browser becomes a dumb transport: no pixel ever reaches the agent. What reaches it is a compact structure — landmarks, heading outline, affordances with stable references, form field schemas including native validation rules, and media flagged honestly as unreadable when it has no alt text. Then the agent queries that structure for the one thing it needs.


Discovery

Where discovery sends you before a browser starts.

9Read the feed
21DOM as transport
3Blocked, inconclusive
1Call the API
1unknown

We used to write this section as “ten of 35 never needed a browser at all”, and that was too strong by nine. Nine of those ten publish an ordinary blog feed. A feed is genuinely cheaper for a content question, and our own discovery code says in the same breath that it will not support transacting — which is what all four intents on this page are. Those nine are measured through a browser further down this very document.

The honest count is one. Zencargo answers on a GraphQL endpoint, and for that one the honest answer is that we add nothing: call the API. Even that detection is a substring match on a probe response rather than a schema we retrieved, so treat it as a lead rather than a finding.


What each path costs

Median across the 30 sites we could read.

PathMedian tokensWorstWhat the agent gets
Raw HTML dump102,4071,519,280Everything, including the tracking and the styling. Blows a context window on one page.
Vision, measured9,14617,745What ten agents actually consumed per site. No refs, nothing actionable, no validation rules — and six of the ten had content hidden behind a cookie panel they could not dismiss.
Accessibility tree3,225Cheaper, and still the whole page rather than the answer.
Structured read (whole page)2,4139,255Landmarks, outline, every affordance with a usable reference, form schemas, honest media flags.
Bounded query67556The affordance for the one intent, with a reference the agent can act on. Median of queries that returned something; a query that finds nothing is cheap and useless and averaging those in flatters us.

The whole-page read is not the product. The bounded query is, because it returns the answer rather than the page — and unlike the screenshot it returns something the agent can act on. That is the entire argument, and it is why the accessibility-tree comparison in our first sweep was answering the wrong question.


Site by site

What an anonymous first visit to each homepage exposes, structurally.

Get a quote, track a shipment, reach a human, book or post a load. A tick means the affordance was found structurally on the homepage, on one visit, on 2026-09-16, unauthenticated, and carries a reference an agent could act on. A dash means it was not found there, on that visit. It is not a measure of these companies’ APIs, EDI links, customer portals or products, several of which are excellent and none of which we tested. Each row is hashed; the archived structure it was derived from is at /api/archive/<host>. If we have named you and a row is wrong, the correction route is on /contact.

15/30Get a quote
10/30Track a shipment
24/30Reach a human
10/30Book or post a load
SiteRouteQuoteTrackContactBookFour queriesScreensStructureArchive
DHL
dhl.com
Blocked, inconclusive20182,7330b25fab9
XPO
xpo.com
DOM as transport20591,3460e72a6f7
Maersk
maersk.com
DOM as transport25293,28173d330eb
DSV
dsv.com
DOM as transport26812,89587e4ead7
Kuehne+Nagel
kuehne-nagel.com
DOM as transport19291,960ec02ce9c
Pall-Ex
pallex.co.uk
DOM as transport20342,024bba9b198
GEODIS
geodis.com
Read the feed22563,30395e6806b
UPS
ups.com
DOM as transport29551,2475b49b543
DAT Freight & Analytics
dat.com
Read the feed32984,2506e610b24
Flexport
flexport.com
DOM as transport109141,163f4bd4d6f
MSC
msc.com
Blocked, inconclusive13673,3397c60d6cb
Forto
forto.com
Read the feed169112,591b0a54ee6
Shiply
shiply.com
DOM as transport17292,1893a6994f6
Returnloads
returnloads.net
DOM as transport18982,42002b75a70
123Loadboard
123loadboard.com
Read the feed234113,237e854ddf7
Freightos
freightos.com
Read the feed251122,274df7b9abe
Truckstop
truckstop.com
Read the feed25872,520ac425699
project44
project44.com
Read the feed261135,16109388a4b
uShip
uship.com
DOM as transport669123,075b720e051
GXO Logistics
gxo.com
Read the feed71101,71814f65d12
Expeditors
expeditors.com
DOM as transport7952,771ed7df332
Gregory Distribution
gregory.co.uk
DOM as transport8292,0493ec88374
FourKites
fourkites.com
DOM as transport89143,0103aabf7e9
sennder
sennder.com
DOM as transport108111,816b04f21c8
Loadlink Technologies
loadlink.ca
Read the feed10941,2586b116343
Loadsmart
loadsmart.com
DOM as transport11881,711f98a09a9
Wincanton
wincanton.co.uk
DOM as transport145112,405b3f8b8a7
FedEx
fedex.com
DOM as transport5269,255069aeb05
Zencargo
zencargo.com
Call the API5271,183235dda56
Transfix
transfix.io
DOM as transport52105850975630f

Tracking is the weakest column, and honestly so: most tracking sits behind a login, which an unauthenticated sweep cannot reach and should not try to.


Refusals and exclusions

The 3 that refused a crawler, and the 2 we excluded ourselves.

SiteSegmentWhat happened
Hapag-Lloyd
hapag-lloyd.com
Ocean carrierHTTP 502 — served an entitled browser, see the third arm
CMA CGM
cma-cgm.com
Ocean carrierHTTP 403 — served an entitled browser, see the third arm
Palletways
palletways.com
Pallet networkError: page.goto: net::ERR_TUNNEL_CONNECTION_FAILED at https://palletways.com/ — served an entitled browser, see the third arm
C.H. Robinson
chrobinson.com
3PLrobots.txt disallows /
DB Schenker
dbschenker.com
3PLredirects to dsv.com; de-duplicated

First sweep, kept for the record

Raw page weight, against the lean control.

Superseded as an argument, not deleted. The control is our own lean test site, measured the same day with the same instrument: 1,856 HTML tokens, 0.0 MB, 1 request.

SiteHTMLa11y treeTransferReqsScreens
Wincanton
wincanton.co.uk
What is served to a fresh visitor is a cookie-consent wall naming 1,075 partners, containing 2,006 base64 vendor logos. Only 6,232 characters are visible text. This is the real front door, not the site content.
1,521,6324,03818.9 MB7310.27
Pall-Ex
pallex.co.uk
588,2922,9405.4 MB2713.47
Freightos
freightos.com
470,7853,8504.8 MB10111.85
Loadlink Technologies
loadlink.ca
448,6411,69323.1 MB853.69
DAT Freight & Analytics
dat.com
425,1326,3569.2 MB5227.28
Flexport
flexport.com
374,7255,9504.4 MB6810.58
project44
project44.com
293,5055,9895.0 MB6612.1
FourKites
fourkites.com
Redirects to fourkites.ai.
222,7766,9396.6 MB7114.8
Kuehne+Nagel
kuehne-nagel.com
211,1152,5752.1 MB788.15
Transfix
transfix.io
183,8882,7363.8 MB1179.37
DSV
dsv.com
131,4814,8134.1 MB381
Gregory Distribution
gregory.co.uk
111,6633,2252.8 MB338.9

Wincanton is the clearest case of why weight was the wrong measure. Its front door is 1.5 million HTML tokens of cookie-consent wall. Dismiss the wall and read the structure and the same page is 2,405 tokens.


Limits of the structured arm

What this does not show.

  • The vision comparator is now MEASURED, on ten sites, by ten independent agents that saw only screenshots. Ten is a small sample and it is the limiting number in this arm. The previous comparator was modelled as full-page coverage and inflated the headline roughly fortyfold; that is recorded as a finding rather than quietly replaced.
  • The headline does not charge the model for the structured read, because the structured read is never sent to a model — the transport builds it and answers a bounded question from it. If you disagree with that boundary, the number that charges it is published beside the headline and is 3.7x.
  • The intent matcher has been wrong three times: unanchored terms, correctly anchored terms meaning something else in freight, and prose or dead links counted as affordances. It is re-derivable from the committed archive, so you can disagree with it and recompute rather than take our word.
  • A bounded query answers "where is the affordance for this intent", not "what is the price". Completing a quote means filling and submitting a form, which this sweep deliberately does not do to sites we do not own.
  • We accept nothing. The sweep used to click accept inside a cookie banner on every site, which contradicted our own entitlement arm and had never been checked; measured, it moves the answerable-intent total by one in 124, in the wrong direction. The comparison arm is still run so that claim stays checkable. Vision has no such option either way, and cookie panels obscured content in six of the ten vision trials.
  • We present as a normal Chrome build. We do NOT mask navigator.webdriver or run any other anti-detection. Both identities are run so the difference is published rather than absorbed.
  • The vision outcomes are manually recorded observations; the original model responses and screenshot bytes are not committed. The reported token costs use a per-screen formula, not provider usage receipts. Arithmetic reproduces; the original observations cannot be independently replayed from this repository.
  • The entitlement arm contains four historical, manually recorded sites. Its original structured views are absent and its older intent classifications have not been revalidated against the current matcher. Its counts are not evidence of current product accuracy or universal site access.
  • One page view per site, homepage only, one day, one location.
Limits of the weight arm

And what the first one did not.

  • One page view per site, of the homepage only. A real transaction would traverse several pages, so per-task totals are larger than one observation.
  • Measured from a datacentre IP with headless Chromium. Some of the blocks and timeouts would not happen from a consumer browser on a home connection — but a datacentre headless browser is exactly what an AI agent is, so this is the right population to measure, not a confound to hide.
  • The accessibility-tree figure is Playwright's ariaSnapshot, a proxy for what a real browser-automation agent receives. A real one also carries element references, so the true figure is somewhat higher.
  • Text tokens use o200k_base as a vendor-neutral proxy; production tokenizers differ, typically within about 20%.
  • Cookie-consent overlays are counted as part of what is served, because that is what an arriving agent actually receives.
  • Measured once, on one day, from one location. Sites change; this is a snapshot with a date on it, not a standing claim.
  • Dollar and energy projections derived from these observations remain modelled, and inherit the step-count and price assumptions of the original benchmark.

Everything we have found wrong in our own work, including both corrections on this page, is in the findings register.


Measure your own site.

The pilot begins by running this against your actual surfaces rather than a median. If the number comes back small, that is the answer and you have bought nothing.