Research · September 2026
What 598 contact pages look like from the outside
We read 598 small-business contact pages the way software reads them — one fetch, no browser. Then we opened the same pages in a real browser to find out what the first pass had missed.
It missed 87 of the 440 forms a browser could find, one in five. Which fifth depends almost entirely on what built the site: a plain fetch found 94.3% of forms on WordPress and 21.9% on Squarespace.
- 598
- 87
- 72 pts
That is a finding about measurement, and it comes first because every other number here — and in every report like this one — is produced by a method with that problem in it. Most of them do not say so. We can, because we measured ours.
The sample
Five samples, drawn five different ways
None of them is a random draw of small businesses, and saying otherwise would be the flaw that sinks the whole thing. So they were drawn to tilt in different directions.
| How it was drawn | Selects for | Sites |
|---|---|---|
| Commercial-intent search across five US trades, 65 metros | Businesses that invest in being found | 263 |
| Town chamber and local trade association member lists — UK, DE, NL, IE | Businesses that pay to belong locally | 70 |
| Utility contractor lists and professional body registers, 27 US states | Businesses that gave a website to a register | 86 |
| US chamber and Main Street directories, 25 metros | The same, plus a long dated tail | 70 |
| The same method again, 64 localities across five countries | The same | 109 |
A finding that holds across all five is standing on something. One that appears in a single sample is reported with the sample named beside it.
The worst skews, stated rather than discovered: membership in most of these bodies costs money, so all five lean toward established businesses. The website field is optional in most registers, so all five are conditioned on a business having supplied one. And a business with no website at all is invisible to every sample here by construction.
The domains are not published. The sources are, so the sample can be redrawn and the figures rechecked. No business that was measured is named here.
How they were read
Two passes over the same list
Pass one — what software sees
A single fetch per candidate path, stopping at the first page carrying a form, reading the markup as the server delivers it, without running any of the page's JavaScript. This is the production code, not a copy of it.
Pass two — what a visitor sees
The same page opened in a browser, allowed to finish loading, read from the rendered DOM. This is the ground truth the first pass is scored against.
Both passes send the same user agent and say what they are. Both read one public page per site. Nothing was typed into a form, no field was touched, nothing was submitted, and neither pass stored a page.
The browser uses our user agent rather than Chrome's, deliberately. Chrome's own string would get past sites that refuse us, and the comparison would then contain two variables — whether a page renders and whether we are let in. Sending the same string on both passes isolates the one being measured.
Both passes are published, along with the pre-registration and the sampling sources. They need no key, no account and no paid service, so every figure below can be checked by running it again — which is the point, and is not true of the numbers this study was written against.
The finding
What a single fetch can see
Of the contact forms a browser found, the share a static fetch also found.
| Platform | Forms found by browser | Also found statically | Rate |
|---|---|---|---|
| WordPress | 193 | 182 | 94.3% |
Wix
|
26 | 24 | 92.3% |
| Shopify | 23 | 21 | 91.3% |
| Webflow | 13 | 11 | 84.6% |
| Unrecognised | 141 | 102 | 72.3% |
| Showit | 11 | 5 | 45.5% |
| Squarespace | 32 | 7 | 21.9% |
| All | 440 | 353 | 80.2% |
The obvious reading — site builders hide their forms from software — is wrong, and that is the interesting part. Wix is at 92.3%; Squarespace is at 21.9%. Both are hosted builders, both host the whole site, and they are seventy points apart. It is not the category. It is how one platform assembles a page, and no general rule about builders would have predicted it.
On Squarespace the protection is as invisible as the form it protects: of the CAPTCHAs a browser could see on those sites, none were in the markup. A tool reading markup reports those sites as having no form and no CAPTCHA. Both statements are false, and they fail together — a page we discard takes its CAPTCHA with it.
If you are wondering what this says about your own page, the same static pass runs on our free check — one page, no account, nothing stored. It carries the same blind spot everything above describes, and now says so.
What changed when the sample grew
The headline moved against us, and we said in advance that it might
The first four samples gave 489 sites and the headline rested on 39 Squarespace ones — the thinnest denominator under the most interesting number here. So a fifth sample was drawn, registered in advance, and deliberately not selected on platform.
WordPress
Squarespace
Every other platform held steady. Squarespace rose eight points on eleven more readable pages, which is what a thin denominator does when you feed it — and it is why that number is still the most attackable thing here. The finding survives: a static read misses roughly four in five Squarespace forms and one in seventeen WordPress ones.
But anybody quoting 14.3% is quoting a figure a larger sample already corrected, and the larger figure may correct again. It is printed because the rule saying it would be was written down before the data existed.
The pages we can see
What is true of them
Everything in this section is measured on the subset a static fetch could read, which the table above shows is not the same population on every platform. Read it as a floor, not a census.
Nobody on Wix adds a CAPTCHA
Neither is a blind spot — we can see over 90% of forms on both. Two platforms decide for the owner, in opposite directions — and the one that adds it decides on behalf of every visitor the W3C has been writing about since 2003. On WordPress, where nothing is decided for anybody, 42% of owners decided for themselves.
Half print an address in the markup
Anything that reads pages can copy it. That is the mechanism by which an address reaches a list — not a guess about why anybody gets spam.
Most publish no DMARC record
Nothing tells a receiving mailbox what to do with mail forged in their name — the whole job of a DMARC record, and something Google has required of bulk senders since 2024. It is the one finding here with a one-line fix, and the one we have nothing to sell against.
Two questions we asked before looking, whose answer was nothing
No business in the sample has an SPF record past the ten-lookup limit. Zero, on every platform, across all 598. The question was worth asking and the answer is that this is not a problem small businesses have. It is printed because it was asked.
The trade a business is in predicts nothing once platform is accounted for. An earlier pass comparing wedding vendors against restaurants found CAPTCHA rates of 23.9% and 26.3% — no difference, and the gap that did appear pointed the wrong way and sat inside one standard error. The platform was doing the work the whole time.
What this cannot tell you
Including the question that matters most
Nothing in a public page can say whether anybody's filter is holding back real enquiries.
Three vendors have said as much in their own support channels. Gravity Forms staff: marking an entry as spam “simply moves the entry from the all view to the spam view, so no training is performed”. CleanTalk: “there is currently no mechanism to automatically resend the request after marking it as ‘Not Spam’”. Wix: an automation that would show a submission in the inbox “currently won't run if the submission is considered spam”.
And the correction goes nowhere even when a person makes it: marking an entry “Not Spam” “will not have any impact on future spam assessments. IPs and emails are not whitelisted”. So the businesses it happens to do not know, and the one thing they can do about it does not carry forward. That is worth finding out, and no amount of reading contact pages will find it out.
It also cannot say how much spam anybody receives — no figure in a page's markup carries it. Every circulating "N% of form submissions are spam" number we could trace ends at a vendor blog with no stated method, and we are not adding another.
Limitations, in one place
- None of the five samples is random. All five lean toward businesses established enough to belong to something, be listed somewhere, or rank for something.
- A business with no website is invisible here, as is one whose only presence is a social page. Between 30% and 75% of the members of the directories we read had no website to measure. That is the largest filter applied to this study, and it was applied before we touched anything.
- 36% of sites matched no platform signature and are reported as unrecognised rather than distributed among the others. Most are likely hand-built.
- 17 of 598 refused one or both passes. They are reported as refused rather than dropped — a site refusing a scripted request is itself evidence about that site.
- Several figures rest on fewer than twenty readable pages and are printed as fractions so they cannot be mistaken for rates.
- The instrument changed mid-study. Sample A was first read by an earlier version of the fetcher that found forms on 149 of 263 sites; after the path list and the embed signatures were fixed, the same 263 gave 175. Every figure here is from the later version, and sample A was re-read from scratch so nothing is compared across two instruments.
Who made this
Humainbox filters contact-form submissions: junk is held back, real enquiries are forwarded to whoever should answer them, and nothing is ever deleted. We built the measurement first, to find out whether we could see what we claimed to see, and published what it got wrong because a report with no error bars is an advertisement. The free check runs the same static pass against one page and tells whoever asks what it found — no account, no email field, nothing stored.
Where the non-measured claims come from
- Broken Gates: commercial CAPTCHA solvers against hCaptcha, reCAPTCHA and Turnstile — arXiv
- The W3C on why CAPTCHAs exclude people, first published 2003 — W3C
- RFC 7489 — DMARC, and what a receiving mailbox does when a check fails — IETF
- RFC 7208 §4.6.4 — the ten-lookup limit an SPF record may not pass — IETF
- The authentication Google has required of bulk senders since 2024 — Google
- Gravity Forms: marking an entry as spam moves it and trains nothing — Gravity Forms community
- CleanTalk: no mechanism to resend a request after marking it Not Spam — WordPress.org support
- Wix: a submission judged spam does not trigger the show-in-inbox automation — Wix Studio forum
- Gravity Wiz: marking Not Spam does not affect future assessments, and nothing is whitelisted — Gravity Forms community