~/free-browser-based-ai Support me apps about me 🌐 中文
██████╗ ██████╗ ██████╗ ██╗ ██╗███████╗███████╗██████╗ ██████╗ █████╗ ███████╗███████╗██████╗ █████╗ ██╗ ██╗ ██╗ ███╗ ███╗ ██╔══██╗██╔══██╗██╔═══██╗██║ ██║██╔════╝██╔════╝██╔══██╗ ██╔══██╗██╔══██╗██╔════╝██╔════╝██╔══██╗ ██╔══██╗██║ ██║ ██║ ████╗ ████║ ██████╔╝██████╔╝██║ ██║██║ █╗ ██║███████╗█████╗ ██████╔╝ ██████╔╝███████║███████╗█████╗ ██║ ██║ ███████║██║ ██║ ██║ ██╔████╔██║ ██╔══██╗██╔══██╗██║ ██║██║███╗██║╚════██║██╔══╝ ██╔══██╗ ██╔══██╗██╔══██║╚════██║██╔══╝ ██║ ██║ ██╔══██║██║ ██║ ██║ ██║╚██╔╝██║ ██████╔╝██║ ██║╚██████╔╝╚███╔███╔╝███████║███████╗██║ ██║ ██████╔╝██║ ██║███████║███████╗██████╔╝ ██║ ██║██║ ███████╗███████╗██║ ╚═╝ ██║ ╚═════╝ ╚═╝ ╚═╝ ╚═════╝ ╚══╝╚══╝ ╚══════╝╚══════╝╚═╝ ╚═╝ ╚═════╝ ╚═╝ ╚═╝╚══════╝╚══════╝╚═════╝ ╚═╝ ╚═╝╚═╝ ╚══════╝╚══════╝╚═╝ ╚═╝

Free Browser Based AI LLM_

A real LLM running 100% in your browser · no login · no server · no data leaves your device
Built for demonstration purposes — not tested in heavy-use scenarios. This is a showcase of local small-model capabilities running entirely on your device.
Model not loaded
System prompt (persona)
First load downloads the weights once (cached in your browser). Subsequent visits start instantly.
Chat
sys
Pick a model and hit "Load model" to begin. Once loaded, this page works fully offline.
0.70 tokens: 0 · speed: 0 tok/s
💡 Writing scripts or long code? Increase max tokens to 2048, 4096, or 8192 above — otherwise the model will cut off mid-response.

How this works

Runtime
WebLLM + WebGPU
Where the model runs
Your GPU. Your tab.
Data leaving your device
Zero
After first download
Works offline
Terminal
jasper@macbook:~/free-browser-based-ai $ inspect --privacy
# checking data flows...
✓ no telemetry
✓ no API calls to inference servers
✓ no cookies, no tracking, no accounts
→ model weights fetched from CDN, then cached locally
→ inference runs on your GPU via WebGPU
⚠ small models = limited reasoning. Don't trust it with facts.

Model comparison

ModelMakerSizeModeBest for
Qwen 2.5 0.5BAlibaba~350 MBWebGPUFastest load, low-end GPUs
TinyLlama 1.1BStatNLP~650 MBWebGPUVery light, basic rewrites
Llama 3.2 1BMeta~880 MBWebGPUBest all-round default
Qwen 2.5 Coder 1.5BAlibaba~950 MBWebGPUCoding help
Gemma 2 2BGoogle~1.5 GBWebGPUNatural, fluent prose
Phi 3.5 MiniMicrosoft~2.2 GBWebGPUHighest quality, reasoning & code
Mistral 7B v0.3Mistral AI~4.3 GBWebGPUPowerful GPUs only
SmolLM2 135M / 360MHuggingFace~70–200 MBWASM (mobile)Phones, universal CPU
Sizes are approximate quantised (q4) download sizes. WebGPU models need Chrome/Edge/Brave on desktop; WASM models run on any device including iOS and Android.

Why run an LLM in your browser?

Most free AI chatbots send every prompt to a cloud server, require an account, and log your conversations. This tool works differently: it downloads an open-weights language model — Llama, Qwen, Phi, Gemma or Mistral — straight into your browser tab and runs it on your own hardware via WebGPU, or on any CPU through WebAssembly. No signup, no API key, no subscription, and no data collection of any kind.

Private by design — your prompts never leave your device, which makes this safe for drafts, client notes, or anything you wouldn't paste into a cloud chatbot. Works offline — after the one-time model download you can keep chatting on a plane or behind a strict firewall. Free without limits — no trial, no message caps, no upsell; the models are open-source and the compute is yours.

Typical uses: rewriting and summarizing text, drafting emails, explaining code, generating KQL or SQL queries, extracting JSON from messy text, brainstorming, and language practice — a lightweight, private alternative to cloud AI assistants, running as a local AI chatbot in your browser.

More local-AI tools on this site: IBM Granite AI in your browser · Whisper speech-to-text (local) · LLM token counter · AI-generated text detector

01. Getting started

What is this, exactly?

A real language model running inside your browser tab — not a front-end for someone else's API. You pick an open-weights model, it downloads once, and from then on every prompt is processed on your own GPU or CPU. There is no account, no server doing the thinking, and no per-message cost.

How do I run an LLM in my browser?

Open the page, pick a model and press Load model. The weights download once, are cached in your browser, and then run locally — on your GPU through WebGPU, or on your CPU in the Mobile / WASM tab. There is no install, no command line and no server involved.

Do I need to download or install anything?

No app, no runtime, no admin rights. The only thing that downloads is the model file itself, straight into the browser cache. Nothing is written to your Downloads folder and nothing is registered with the operating system — clearing site data removes every trace.

Do I need an account, an API key or an OpenAI subscription?

None of the three. Because the model runs on your own hardware there is nothing to authenticate against and no per-token bill. You never enter an email address, a key or a card number.

Why does the first load take so long?

You are downloading the model itself — anything from ~70 MB for the smallest mobile model to several gigabytes for a 7B one. It is stored with the browser's Cache API, so the second visit starts in seconds and needs no network at all. A cached model is marked with a badge in the picker.

Does it work offline?

Yes, once the weights are cached. You can switch off Wi-Fi entirely and keep chatting — a useful test in itself, because a tool that still answers with the network disconnected clearly is not calling a server. Only the first download of each model needs a connection.

What does “Best for my device” do?

It asks your GPU what the largest buffer it can allocate is, and picks a model that comfortably fits: a 1B model on modest hardware, Phi 3.5 Mini on a mid-range GPU, a 7B model on a card with several gigabytes to spare. It is a starting point rather than a verdict — you can always pick something bigger or smaller by hand.

Is it free? Are there any limits?

Completely free, with no account, no subscription, no message cap and no watermark. There is no rate limit either, because there is no server to rate-limit — the only ceiling is how fast your own hardware can generate tokens.

02. Choosing a model

Which model should I pick?

Llama 3.2 1B is the sensible default: around 880 MB, quick to load, and reliable at rewriting, summarising and everyday instructions. Qwen 2.5 0.5B (~350 MB) is the one to choose when speed matters more than depth. Phi 3.5 Mini (~2.2 GB) gives the best reasoning, coding and structured output if your GPU can hold it, and Gemma 2 2B writes the most natural prose. Start at 1B and move up only if the answers disappoint you.

What do 0.5B, 1B, 3B and 7B actually mean?

The number of parameters, in billions — roughly, how much the model knows and how well it can reason. Each step up costs download size, memory and speed: a 1B model is around 880 MB and feels instant, a 3B is about 2 GB and noticeably more coherent, a 7B is over 4 GB and needs a proper graphics card. Quality rises with size, but so does the wait.

Which model is best for coding?

Qwen 2.5 Coder 1.5B is trained specifically on code and is remarkably good for its size — the best value if you want completions, small functions and shell one-liners. Phi 3.5 Mini is the stronger all-rounder when you need the model to reason about the code rather than just produce it. Neither will replace a frontier model on a large codebase.

What are the DeepSeek R1 distills?

Small models trained on the output of a much larger reasoning model, so they work through a problem step by step before answering. That makes them better at puzzles and multi-step logic than their size suggests — and slower, because they spend tokens thinking out loud. Raise max tokens when you use them or they will be cut off mid-thought.

Which model is best for languages other than English?

The Qwen 2.5 family is the strongest multilingual option here, with 1.5B and 3B versions covering a wide range of languages. Phi 3.5 Mini also handles 20+ languages well. Small models drift more in languages other than English, so keep prompts explicit — setting the system prompt to “always answer in Dutch” works better than asking mid-conversation.

Can I run a 7B model?

Only with a discrete GPU that has roughly 5 GB or more of free video memory — the weights alone are around 4.3 GB before the conversation cache. Mistral 7B and the DeepSeek R1 7B distill are in the picker for machines that can take them. On integrated graphics they will either fail to allocate or crawl, and a 3B model is the better trade.

Which models does the Mobile / WASM tab use?

Much smaller ones, because it runs on the CPU: SmolLM2 135M (~70 MB), SmolLM2 360M (~200 MB, the recommended default), Qwen 1.5 0.5B and SmolLM2 1.7B for devices with memory to spare. They are noticeably simpler than the desktop models, but they run anywhere — no WebGPU required.

If I switch models, does it download again?

Each model is cached separately, so the first switch to a new model downloads it and every later switch is instant. Models you have already cached are flagged in the picker, which makes it easy to hop between a fast one and a smarter one without paying the download twice.

What is a context window, and which model has the largest?

It is how much text the model can hold in mind at once — your system prompt, the conversation so far and its own reply all count against it. Phi 3.5 Mini is the roomiest here at 128K tokens; older models such as Phi 3 Mini 4K are limited to a few thousand. When a long chat starts to lose the plot, reset it and paste in only what matters.

What does “q4f16” in the model name mean?

It is the quantisation: the weights are compressed to 4 bits with 16-bit floating point used during computation. That is what makes a multi-billion-parameter model small enough to download and fit in browser memory. The cost is a small amount of accuracy, which is a bargain compared with not being able to run the model at all.

03. Privacy and data

Does my data get sent anywhere?

No. Prompts and replies never leave the tab. The single network request the tool makes is the one-time download of the model weights from the Hugging Face CDN; after that you can disconnect entirely and it keeps working. There is no inference server, because your device is the inference server.

Are my chats stored or logged?

The conversation lives in the page and disappears when you close or reset the tab. Nothing is uploaded, nothing is written to a database, and the only things kept in browser storage are your theme choice and a note of which models you have already cached — never the messages. Use copy log or ↓ .md if you want to keep a conversation.

Is browser-based AI genuinely private, or just marketed that way?

Genuinely, and you can check it yourself rather than take anyone's word for it. There is no API key in the page to send anything with, no endpoint to send it to, and the model answers with the network switched off — which is impossible for a cloud chatbot.

How can I verify nothing is uploaded?

Two ways. Open your browser's developer tools, go to the Network tab, and watch it while you send a message: after the model has loaded there are no requests at all. Or simply disable Wi-Fi and keep chatting — if the answers still come, they are being generated locally.

Can I paste confidential or GDPR-sensitive text into it?

Technically nothing leaves your device, so there is no transfer to a third party and no processor to sign an agreement with — which is exactly why local models are attractive for sensitive text. That said, your own organisation's policy still applies, so check it before pasting client data anywhere, including here.

Does the model download tell anyone what I am doing?

Downloading the weights is an ordinary file request to a CDN, so that CDN sees a file being fetched, as it would for any image or script. It carries no prompt, no conversation and no identity beyond a normal web request — and it happens once per model, not per message.

Does it work in private or incognito mode?

It works, but private windows discard storage when you close them, so the model downloads again the next time. If you plan to use it regularly, a normal window keeps the cache and saves you the wait.

04. Speed and hardware

What hardware do I need?

For the desktop tab, a GPU that supports WebGPU and has enough free memory for the model: roughly 1 GB for a 1B model, 2–3 GB for a 3B, and 5 GB or more for a 7B. Integrated graphics handle the 0.5B–1B models comfortably. If none of that applies to your machine, the Mobile / WASM tab runs on the CPU on practically anything.

How fast should it be?

On a modern discrete GPU a 1B model typically streams faster than you can read, and a 3B model still feels conversational. On integrated graphics expect it to be readable but unhurried, and in CPU / WASM mode expect a few words per second. The live tok/s counter under the chat box shows exactly what your device is doing.

Why is generation slow or stuttering?

Usually the model is too large for the hardware, so memory is being shuffled instead of used. Drop a size class, close other GPU-heavy tabs — video calls and 3D pages compete for the same memory — and check that your laptop is not in a battery-saver profile that caps the GPU. A long conversation also slows things down, because the whole history is reprocessed; reset the chat to get the speed back.

Why does the answer stop in the middle?

It hit the max tokens limit, which defaults to 1024. For scripts, long code or detailed documents raise it to 2048, 4096 or 8192 next to the Send button. It is a deliberate cap rather than a failure — without it a small model can ramble for a very long time.

What does the temperature slider do?

It controls how adventurous the sampling is. Low values (around 0.2) make the model repetitive but predictable, which is what you want for extraction, JSON and code. Higher values (0.8 and above) produce more varied writing at the cost of accuracy. The default of 0.7 is a reasonable middle for chat.

How much faster is GPU mode than WASM mode?

Typically several times faster, sometimes an order of magnitude, which is why the WebGPU tab is the default on desktop. WASM mode exists so that the tool still works on hardware and browsers without WebGPU — it trades speed for running absolutely anywhere.

It says out of memory, or the GPU device was lost.

The model did not fit. Choose a smaller one, close other tabs using the GPU, and try again — the error message names the limit it hit. On laptops that switch between integrated and discrete graphics, forcing the browser onto the discrete GPU in the system settings often solves it outright.

Will it drain my battery or heat the laptop?

While it is generating, yes — you are running a neural network on your own silicon, so the fans may spin up much as they would during a game. It only draws power while a reply is being produced; the model sitting in memory idle costs nothing.

05. Browsers, storage and troubleshooting

Which browsers support WebGPU today?

Chrome and Edge have shipped it by default since version 113 on Windows, macOS and ChromeOS, and on Android 12+ since Chrome 121. Safari enables it by default in macOS Tahoe 26, iOS 26 and iPadOS 26. Firefox ships it on Windows from version 141 and on Apple Silicon Macs from 147, with Linux and Android still in Nightly. Anything not on that list can still use the Mobile / WASM tab.

Why does it say “WebGPU not available”?

Your browser, or that machine's graphics driver, is not exposing WebGPU — common on older browser versions, on Linux with certain drivers, in some virtual machines, and where enterprise policy disables it. The page detects this and moves you to the Mobile / WASM tab automatically, so you can still use the tool; updating the browser is the fix if you want GPU speed.

Does it work on iPhone or iPad?

Yes, in the Mobile / WASM tab, with one caveat: iOS Safari enforces a hard memory ceiling of roughly 1–1.5 GB per tab, so only the smallest model (SmolLM2 135M) loads reliably. Larger ones crash the tab with “a problem repeatedly occurred”. For anything more capable, open the page on a desktop.

Does it work on Android?

Yes. Chrome on Android 12 and later supports WebGPU on most recent chipsets, and where it does not, the WASM tab runs on the CPU. Stick to the small models — phone memory limits bite long before the model quality does.

It will not download the model on my work network.

Corporate proxies and content filters frequently block the Hugging Face CDN, which is where the weights come from, and some inspect and break large downloads. That is a network policy question rather than a browser one — the same page usually works immediately on a home connection or a phone hotspot.

Where are the weights stored, and how do I delete them?

In the browser's Cache Storage for this site, not in your file system. Clearing site data (or “Cookies and other site data” for this domain) removes every cached model and frees the space; the tool then behaves like a first visit. Nothing outside the browser profile is touched.

Can the browser throw the cached model away?

Yes. Browsers evict cached data when disk space runs low or a profile is cleaned, and private windows discard it on close. If a model that loaded instantly yesterday starts downloading again, that is what happened — nothing is broken.

How much disk space will this use?

Only what the models you actually load take up: from about 70 MB for the smallest to roughly 4.5 GB for a 7B. Every model you try is kept until you clear site data, so trying several does add up — worth remembering on a machine that is short on space.

06. What it is good at — and what it is not

Why is it less capable than ChatGPT or Claude?

Because it is smaller by three orders of magnitude. Frontier models run on server racks with hundreds of billions of parameters; a 1B model in a browser tab is a fraction of a percent of that, and it is remarkable that it works at all. Judge it as a fast local assistant for small jobs, not as a replacement for a frontier model.

So what is it actually good at?

Short, well-defined text work: rewriting a paragraph formally or casually, summarising something you paste in, drafting a reply, extracting fields into JSON, explaining a concept simply, and small code and query snippets. The preset buttons above the chat box are shortcuts to exactly those tasks. Anything that needs current facts or long reasoning chains is the wrong job for it.

Can it search the web or read my files?

No. It has no internet access and no access to your files — it only sees what you type into the box. That is the same property that makes it private: there is no channel in either direction. If you want it to work on a document, paste the relevant part in.

Does it make things up?

More often than a large model, yes. Small models are confident and frequently wrong about names, dates, numbers and anything specialised, and they have no way to look anything up. Treat factual claims as drafts to verify, and lean on the tasks where the source material is in your prompt — rewriting, summarising, extracting.

Can I change its personality or behaviour?

Yes. Open System prompt (persona) under the model picker and describe how it should respond — “you are a terse senior engineer”, “always answer in Dutch”, “reply only with JSON”. It applies to every message in the conversation, and it is by far the most effective way to improve the output of a small model.

Can it write code?

For small, self-contained things, yes — a regex, a shell one-liner, a function, a KQL query from the preset. Use Qwen 2.5 Coder or Phi 3.5 Mini, keep the request narrow, raise max tokens so it is not truncated, and read what it produces before running it. It has no idea what is in your codebase.

Can I use the output commercially?

Generally yes, but the licence belongs to the model rather than to this page, and they differ: Llama has Meta's community licence, Gemma has Google's terms of use, most Qwen and SmolLM2 releases are Apache 2.0 and Phi is MIT. If the output is going into a product, read the licence for the specific model you used.

07. Under the hood

What technology powers this?

The desktop tab uses WebLLM from MLC AI, which compiles open-weights models to run on WebGPU. The mobile tab uses Hugging Face's transformers.js on top of ONNX Runtime Web, executing on the CPU through WebAssembly. Everything around them is plain JavaScript and CSS in a single HTML file.

Where do the model weights come from?

From the Hugging Face CDN, in the pre-compiled MLC format for the WebGPU tab and ONNX for the WASM tab. They are open-weights models published by Meta, Alibaba, Microsoft, Google DeepMind, Hugging Face and DeepSeek — the same files anyone can download and run locally.

Is this the same as Ollama or LM Studio?

Same idea, different delivery. Those are desktop applications you install, and they will always be faster and support far larger models because they use your hardware directly. This runs the model in a browser tab with nothing installed at all, which makes it the easier way to try a local model — on a locked-down machine, on someone else's computer, or simply to see what the fuss is about.

When should I use a cloud model instead?

When the task needs breadth, accuracy on facts, long documents, up-to-date information or serious reasoning — a frontier model will do in one attempt what a 1B model cannot do at all. Use this when the text is sensitive, when you are offline, when you want an instant answer without opening an account, or when the job is small enough that a small model is simply the quicker tool.