Running a model in the browser, and when that is a good idea
Transformers.js and WebGPU put real models on the client. The demo is easy; knowing which features belong there is the part worth thinking about.
Contents
The first time I ran a speech model entirely in a browser tab, with the network disconnected, I sat there for a minute feeling slightly silly about every architecture diagram I had drawn in the previous five years.
It is genuinely easy now. It is also genuinely the wrong choice for most features. Both things are worth knowing before you spend a sprint on it.
The easy part
Transformers.js runs a large slice of the Hugging Face model catalogue in JavaScript, on WebGPU where available and WebAssembly where not. A working sentiment classifier is about this much code:
import { pipeline } from '@huggingface/transformers';
const classify = await pipeline(
'sentiment-analysis',
'Xenova/distilbert-base-uncased-finetuned-sst-2-english',
{ device: 'webgpu' }
);
const out = await classify('the checkout flow is a nightmare');
// [{ label: 'NEGATIVE', score: 0.9996 }]
No server. No API key. No per-request cost. The model downloads once and sits in the cache.
Swap the pipeline name and you get feature extraction (embeddings), summarisation, translation, object detection, or speech-to-text. The API is deliberately boring, which is the best compliment I can pay a library.
The part that decides whether you ship it
Three costs, and you should know all three before you write a line.
Download size. A small classification model is 30–80 MB. A useful embedding model is 30–100 MB. Whisper’s small variants start around 40 MB and get large quickly. That is not a JavaScript bundle, it is a video download, and on a phone on a metered connection in Karachi it is a decision you are making on your user’s behalf.
First-run time. Download, then compile shaders, then warm up. On a laptop with a decent GPU this is a couple of seconds after the download. On a mid-range Android phone it can be fifteen, and the tab is not exactly lively while it happens.
Memory. These models sit in memory as long as the tab lives. Mobile Safari will discard your tab under pressure and the user will blame your site, not the browser.
When local genuinely wins
I now use a fairly short checklist, and a feature has to tick at least one line hard.
The data must not leave the device. This is the strongest reason and it is not a performance argument at all. Medical notes, legal documents, private recordings, anything a user would hesitate to paste into a website. Scribe, the meeting tool I build, transcribes locally for exactly this reason — the promise that nothing leaves the machine is the product.
It runs constantly. Per-keystroke, per-frame, per-scroll. A server round trip per keystroke is untenable at any price; local inference at 5 ms is free after the first load. Smart autocomplete, live camera filters, real-time moderation on a text box.
It must work offline. Field apps, in-flight, patchy connectivity. If the feature is useless without a network, ship it on the server.
The per-request cost would sink you. A free tool with a million users and no revenue cannot pay for a million API calls. It can happily pay for zero.
If none of those apply — and for most business features none of them do — put the model on a server. You get a better model, one place to update it, no download, and instant first use. That is not a compromise, it is usually the right answer.
The hybrid that I actually ship
The pattern that has held up best: a small local model decides, and the server handles the hard cases.
// In a worker. Cheap, private, instant.
const { score } = await localModerate(draftComment);
if (score > 0.9) return reject(); // obviously fine to block locally
if (score < 0.1) return accept(); // obviously fine to allow
return serverReview(draftComment); // the 8% that need judgement
Ninety-odd percent of traffic never touches the network, the expensive model only sees genuinely ambiguous input, and the feature stays responsive when the connection does not.
Practical notes from doing this for real
Quantise. Most models ship in 8-bit or 4-bit variants that are a quarter of the size and, for classification and embeddings, near-identical in quality. Try the small one first; only reach for full precision if you can measure the difference.
Cache deliberately. Model weights go in the Cache API via a service worker, not in whatever the browser decides. You want to control when a 60 MB download happens, and you want it to survive a reload.
Ask before downloading. A button that says “Enable on-device transcription (42 MB)” respects the user. An automatic 42 MB download on page load does not, and it will show up in your bandwidth bill too.
Detect WebGPU and fall back. if (!navigator.gpu) means WASM, which works but is several times slower. Decide in advance whether the WASM path is good enough to ship or whether the feature should simply not appear.
Test on a real mid-range phone. Not a simulator, not your phone. The gap between a flagship and a three-year-old Android is enormous here, far larger than it is for ordinary JavaScript.
Where this is going
The direction of travel is obvious: models keep getting smaller for the same quality, WebGPU keeps getting faster and more widely available, and browsers are starting to expose built-in models of their own so the download disappears entirely.
What will not change is the checklist. Privacy, frequency, offline, cost. Run down those four lines, and if a feature does not hit at least one of them convincingly, the server is still the boring correct answer — and boring correct answers are most of the job.