0 cookiesnothing loads before a clickunlimited teammatesreply from slack or discord<4kb launcherno per-seat tax

/blog / receipts

how we test chat widgets

· by kuba · 8 min read

A long ruler along the bottom of a dark background measuring a small block and a framed picture of the Spookat cat, with n=5 written above.

We sell a chat widget, and we’re about to publish a series of posts comparing it with six others. You should not trust a vendor’s benchmark by default, so this post is the lab behind every number: what we measure, how, on which setup, what we couldn’t measure, and where the raw data lives so you can check it.

the seven widgets

Spookat, Tawk.to, Crisp, Intercom, Tidio, LiveChat and Chatwoot: the same six rivals we put next to ourselves on this site’s comparison pages.

For each rival we use the install snippet from the vendor’s own documentation and a widget id the vendor itself runs on its own homepage. That second part matters: we made no accounts anywhere, and we didn’t borrow some customer’s id. What we load is exactly what the vendor serves to visitors of its own site, with the vendor’s own configuration.

For Spookat we load the production launcher from cdn.spookat.com, the same file every customer’s page gets, with a key that belongs to no site. Before the click nothing but that one file is involved, so the key makes no difference.

what we measure, and why each one

A chat widget sits on every page of a site, and most visitors never open it. So the first questions are about what it costs the people who don’t click. The last ones are about the people who do.

test what we record why it matters
weigh-in bytes (compressed and uncompressed), requests by type and by domain, before any interaction every visitor downloads this, most of them for nothing
cookies cookies, localStorage, sessionStorage, IndexedDB and the domains contacted before any interaction storage before consent is what decides whether you need a cookie banner for the widget
lighthouse performance score, Total Blocking Time, Largest Contentful Paint, main-thread time: the page alone vs the page with each widget the number your SEO person looks at, measured the way Google’s tools measure it
click-to-open time from the click until the chat can be used, cold and warm cache, on a throttled phone-class setup the one moment the widget exists for
phone launcher size and position on a 375×812 screen, what it covers, tap target size, what opening it does, what happens with the keyboard up most chat happens on a phone
accessibility axe-core on the closed and the open widget, then keyboard only: reach it, open it, type, close it, see where focus goes a widget that can’t be used without a mouse locks people out of your support

Each test gets its own post with its own numbers. This one is the method they share.

the setup

  • Browser: Google Chrome 153.0.8010.53, the installed stable build, driven by Playwright 1.63. Headed, a real window, because some widgets refuse to boot when the browser says “HeadlessChrome” in its user agent, and a real window is what visitors have.
  • Clean profile, every run: each run starts Chrome with a brand-new profile in a fresh temp directory. Empty cache, no cookies, no storage, Chrome’s default settings. When a test needs a warm cache, it says so and reuses the profile on purpose.
  • Test page: a plain page for a made-up mug shop: a heading, a few paragraphs, a price grid, a footer, system fonts, no images, no scripts of its own. Each widget’s snippet goes at the end of <body>, where the vendors tell you to put it. The page alone is the baseline for every comparison.
  • Machine: an Apple silicon Mac. Lighthouse reports a CPU benchmark index with every run, and we publish it next to the scores.
  • Throttling: Lighthouse runs with its default mobile settings: a Moto G Power screen, simulated slow 4G (150 ms round trip, 1.6 Mbps down) and a 4x CPU slowdown. The click-to-open test uses the same values as DevTools throttling, switched on right before the click. The weigh-in and cookie tests don’t throttle, because bytes and cookies don’t depend on speed.
  • Runs: at least 5 per widget per test, 7 for the weigh-in and Lighthouse. We report medians and publish every single run.
  • Time window: before-click tests watch the page for 25 seconds without touching it. Several widgets keep loading well after the page’s load event, and a short window would flatter them.

four widgets won’t run on a test page

This is the part that changed the plan. Crisp, Intercom, LiveChat and Chatwoot check which domain they’re running on, and the ids we’re allowed to use belong to the vendors’ own sites. On our test page:

  • Crisp shows a red “Invalid website” label next to its launcher.
  • Intercom’s server answers 403 and the console says the domain is not allowed.
  • LiveChat logs that the domain is “not added to the allowed domains” and disables itself.
  • Chatwoot’s chat window is an iframe that its own security policy only lets a few domains embed, so the browser blocks it.

Tawk.to and Spookat boot anywhere. Tidio boots too, but the vendor set its own widget to stay hidden, so our page calls Tidio’s documented tidioChatApi.show() before any click to get the launcher a default install shows.

We could have tricked the browser into believing our test page lives on each vendor’s domain. We didn’t. So those four get measured twice:

  1. On our test page, which shows what they load before they refuse. That part is real and every visitor on a wrong domain pays it, but it’s a floor, not the full weight.
  2. On the vendor’s own page, as any visitor sees it. There we count only requests to the widget’s own addresses (for example client.crisp.chat or js.intercomcdn.com), and we list every counted URL in the raw data. For cookies and storage we load the same page again with the widget blocked, and a name counts as the widget’s only if it shows up with the widget and never without it. For Lighthouse we compare the vendor’s page as it is against the same page with the widget’s loader script blocked.

The vendor’s own page comes with the vendor’s own settings. Intercom’s site pushes a message to new visitors, LiveChat’s opens with a greeting card, Tawk.to’s shows an AI greeting. A default install on your site may load less, or more. Where that changes a number, the post says so.

One more limit: Chatwoot’s widget didn’t load on chatwoot.com’s homepage in any of our three tries, so we use the help center on the same domain, where it did.

what we never do

  • No accounts, no trials, no ids that aren’t the vendor’s own.
  • No messages. We open chats, we type into the box in the accessibility test, and we never press send. Nobody’s support team gets a fake visitor from us.
  • No tuning. Every widget runs with the configuration its owner chose (Tidio’s show call aside, see above). That includes ours: Spookat runs with no look published, so it shows its default style.
  • No numbers from marketing pages. If a number is in a post, it came out of a run.

where ours gets an easier ride, and where it doesn’t

Spookat’s launcher and chat panel load from the production CDN. The chat socket is the one exception. It normally goes to api.spookat.com, and it would need a real site to answer, so in the lab it goes to a small mock on the same machine that sends the same first frame our server sends. Before a click that makes no difference, because the launcher opens no socket. After a click, DevTools throttling still applies to that socket: we checked, and its first frame took 587 ms throttled against 1.3 ms without. What the mock skips is a real TLS handshake and the real server’s work, and the click-to-open post says what that could be worth.

Spookat’s rows were measured twice on 2026-09-27. The first run found three things we didn’t like in our own widget: a first open that took three round trips in a row, a launcher that sat on a shop’s bottom bar on a phone, and three accessibility findings in the open panel. We fixed all three that morning, shipped the fix to production, and measured our rows again with the same setup. Later that morning we made the launcher lighter and switched off two error-reporting headers our CDN added to every response, and measured the weigh-in, Lighthouse and click-to-open rows once more. Every Spookat number in the series is from its latest run, and each research file marks Spookat’s rows with when they were measured and why. The rivals’ rows are from the first run. Since the fix, Spookat’s panel doesn’t wait for the socket before you can type, so the mock socket doesn’t change its click-to-open time.

how to check us

Every test writes one JSON file with the method, the versions, the medians and every run:

Each file also lists the widgets with the exact snippet and id we used. Each post links its own file, and the posts go live one by one over the next weeks.

The scripts are public too: the widget lab serves every file exactly as it ran, with one command per test (bun run weigh-in, bun run lighthouse, and so on). You don’t need them to check a number. Open the vendor’s page or your own staging site, open DevTools, tick “Disable cache”, reload, and filter the Network tab by the widget’s domains. That’s the whole weigh-in. The guide on adding live chat without slowing your site down walks through it step by step. Lighthouse is in every Chrome.

If you get a different number, tell us. Widgets change every week, and so do we. We’d rather fix a post than defend it.

why we’re doing this at all

We think a chat widget should cost the page almost nothing until someone wants to talk, and we built Spookat around that. Saying so is cheap. Measuring it next to everyone else, publishing the raw runs and showing where our own setup gets help is the only version of that claim worth reading.

The launcher you’d be testing is under 4 kb and makes 0 requests to our API before a click. Put it on a staging page and run the weigh-in yourself.

liked it? get in before launch.

invites go out in small batches. then 14 days free, no card.