~/free-browser-based-ai ☕ Support me apps ← about me 🌐 中文
██████╗ ██████╗ ██████╗ ██╗ ██╗███████╗███████╗██████╗ ██████╗ █████╗ ███████╗███████╗██████╗ █████╗ ██╗ ██╗ ██╗ ███╗ ███╗ ██╔══██╗██╔══██╗██╔═══██╗██║ ██║██╔════╝██╔════╝██╔══██╗ ██╔══██╗██╔══██╗██╔════╝██╔════╝██╔══██╗ ██╔══██╗██║ ██║ ██║ ████╗ ████║ ██████╔╝██████╔╝██║ ██║██║ █╗ ██║███████╗█████╗ ██████╔╝ ██████╔╝███████║███████╗█████╗ ██║ ██║ ███████║██║ ██║ ██║ ██╔████╔██║ ██╔══██╗██╔══██╗██║ ██║██║███╗██║╚════██║██╔══╝ ██╔══██╗ ██╔══██╗██╔══██║╚════██║██╔══╝ ██║ ██║ ██╔══██║██║ ██║ ██║ ██║╚██╔╝██║ ██████╔╝██║ ██║╚██████╔╝╚███╔███╔╝███████║███████╗██║ ██║ ██████╔╝██║ ██║███████║███████╗██████╔╝ ██║ ██║██║ ███████╗███████╗██║ ╚═╝ ██║ ╚═════╝ ╚═╝ ╚═╝ ╚═════╝ ╚══╝╚══╝ ╚══════╝╚══════╝╚═╝ ╚═╝ ╚═════╝ ╚═╝ ╚═╝╚══════╝╚══════╝╚═════╝ ╚═╝ ╚═╝╚═╝ ╚══════╝╚══════╝╚═╝ ╚═╝

Free Browser Based AI LLM_

A real LLM running 100% in your browser · no login · no server · no data leaves your device
Built for demonstration purposes — not tested in heavy-use scenarios. This is a showcase of local small-model capabilities running entirely on your device.
▸ Model not loaded
⚙ System prompt (persona)
First load downloads the weights once (cached in your browser). Subsequent visits start instantly.
▸ Chat
sys
Pick a model and hit "Load model" to begin. Once loaded, this page works fully offline. Drop a PDF or text file here to chat with it.
0.70 tokens: 0 · 0 tok/s · first token – · context 0%
💡 Writing scripts or long code? Increase max tokens to 2048, 4096, or 8192 above — otherwise the model will cut off mid-response.

What it can do

📎 Chat with documents
Attach a PDF, text, CSV, JSON or code file. Long files are indexed locally and only the relevant parts are sent to the model.
🖼 Ask about images
Phi 3.5 Vision reads photos, screenshots, receipts and charts — paste a screenshot straight into the box.
🧠 Watch it think
Qwen 3, DeepSeek R1 and Ministral Reasoning show their chain of thought in a folding block. Think can be switched off for speed.
{ } Strict JSON
JSON mode constrains generation with a grammar, so the output always parses — handy for extracting fields from messy text.
💾 History that stays local
Recent chats are kept in this browser only. Regenerate an answer, read it aloud, copy it or export the chat as Markdown.
◻ Simple & full screen
Hide everything but the prompt, or take the chat full screen. Open with ?mode=simple to start there.

How this works

Runtime
WebLLM + WebGPU
Where the model runs
Your GPU. Your tab.
Data leaving your device
Zero
After first download
Works offline
▸ Terminal
jasper@macbook:~/free-browser-based-ai $ inspect --privacy
# checking data flows...
✓ no telemetry
✓ no API calls to inference servers
✓ no cookies, no tracking, no accounts
→ model weights fetched from CDN, then cached locally
→ attached files are parsed and searched in the tab
→ chat history lives in localStorage, never on a server
→ inference runs on your GPU via WebGPU
⚠ small models = limited reasoning. Don't trust it with facts.

Model comparison

ModelMakerSizeModeBest for
Qwen 3 0.6BAlibaba335 MBWebGPUSmallest thinking model, low-end GPUs
Llama 3.2 1BMeta695 MBWebGPUBest all-round default
Gemma 3 1BGoogle563 MBWebGPUFluent, multilingual, light
Qwen 3 1.7BAlibaba968 MBWebGPUReasoning on modest GPUs
Qwen 3 4BAlibaba2.26 GBWebGPUBest quality per gigabyte
Phi 4 MiniMicrosoft2.16 GBWebGPUReasoning, maths & code
Ministral 3 3BMistral AI1.93 GBWebGPUEuropean languages, instructions
Phi 3.5 VisionMicrosoft2.77 GBWebGPUReads images & screenshots
Qwen 2.5 Coder 1.5B / 3B / 7BAlibaba0.87–4.3 GBWebGPUCode
Qwen 3 8BAlibaba4.61 GBWebGPUStrongest on the page; powerful GPUs
DeepSeek R1 Distill 1.5B / 7B / 8BDeepSeek1.0–4.5 GBWebGPULong step-by-step reasoning
SmolLM2 135M / 360MHugging Face137–365 MBWASM (CPU)Phones, any browser
Qwen 3 0.6B · Llama 3.2 1BAlibaba · Meta0.6–1.2 GBWASM (CPU)Better CPU quality, slower
Sizes are the exact download sizes of the quantised weights (4-bit for WebGPU, 8-bit for WASM). The full list of 30+ models is in the pickers above. WebGPU models need a browser with WebGPU; WASM models run on any device including iOS and Android.

Why run an LLM in your browser?

Most free AI chatbots send every prompt to a cloud server, require an account, and log your conversations. This tool works differently: it downloads an open-weights language model — Llama, Qwen, Phi, Gemma or Mistral — straight into your browser tab and runs it on your own hardware via WebGPU, or on any CPU through WebAssembly. No signup, no API key, no subscription, and no data collection of any kind.

Private by design — your prompts never leave your device, which makes this safe for drafts, client notes, or anything you wouldn't paste into a cloud chatbot. Works offline — after the one-time model download you can keep chatting on a plane or behind a strict firewall. Free without limits — no trial, no message caps, no upsell; the models are open-source and the compute is yours.

Typical uses: rewriting and summarizing text, drafting emails, explaining code, generating KQL or SQL queries, extracting JSON from messy text, brainstorming, and language practice — a lightweight, private alternative to cloud AI assistants, running as a local AI chatbot in your browser.

More local-AI tools on this site: IBM Granite AI in your browser · Whisper speech-to-text (local) · LLM token counter · AI-generated text detector

01. Getting started

What is this, exactly?

A real language model running inside your browser tab — not a front-end for someone else's API. You pick an open-weights model, it downloads once, and from then on every prompt is processed on your own GPU or CPU. There is no account, no server doing the thinking, and no per-message cost.

How do I run an LLM in my browser?

Open the page, pick a model and press Load model. The weights download once, are cached in your browser, and then run locally — on your GPU through WebGPU, or on your CPU in the Mobile / WASM tab. There is no install, no command line and no server involved.

Do I need to download or install anything?

No app, no runtime, no admin rights. The only thing that downloads is the model file itself, straight into the browser cache. Nothing is written to your Downloads folder and nothing is registered with the operating system — clearing site data removes every trace.

Do I need an account, an API key or an OpenAI subscription?

None of the three. Because the model runs on your own hardware there is nothing to authenticate against and no per-token bill. You never enter an email address, a key or a card number.

Why does the first load take so long?

You are downloading the model itself — anything from ~70 MB for the smallest mobile model to several gigabytes for a 7B one. It is stored with the browser's Cache API, so the second visit starts in seconds and needs no network at all. A cached model is marked with a badge in the picker.

Does it work offline?

Yes, once the weights are cached. You can switch off Wi-Fi entirely and keep chatting — a useful test in itself, because a tool that still answers with the network disconnected clearly is not calling a server. Only the first download of each model needs a connection.

What does “Best for my device” do?

It asks your GPU for the largest buffer it can allocate and picks a model that fits: Llama 3.2 1B on modest hardware, Qwen 3 4B on a mid-range GPU, and Qwen 3 8B on a card with several gigabytes to spare. It is a starting point rather than a verdict — you can always pick something bigger or smaller by hand, and the note under the picker shows each model's memory needs.

Is it free? Are there any limits?

Completely free, with no account, no subscription, no message cap and no watermark. There is no rate limit either, because there is no server to rate-limit — the only ceiling is how fast your own hardware can generate tokens.

02. Documents, images, thinking & modes

What does 🧠 Think do?

Qwen 3 models can reason before they answer. With Think on, they write out their chain of thought first — shown in a collapsible 💭 block with how long it took — and then answer. It makes arithmetic, logic and multi-step instructions noticeably more accurate, at the cost of time: Qwen 3 0.6B wrote about 550 thinking tokens before answering whether 91 is prime. Turn it off for quick replies. DeepSeek R1 and Ministral Reasoning always think; the toggle only appears for models that support switching.

Can it read images or screenshots?

Yes, with Phi 3.5 Vision. Pick it, load it, and attach an image with 🖼 — or paste a screenshot straight into the message box, or drop a file onto the chat. The image is resized to at most 1024 px inside your browser before the model sees it, and it never leaves your device. It is a 2.77 GB download and needs about 4 GB of GPU memory.

Can I chat with a PDF or a document?

Yes. Press 📎 or drop a file on the chat: PDF, TXT, Markdown, CSV, JSON, HTML or source code, up to 25 MB. Short files go into your message in full. Longer ones are split into overlapping parts and indexed locally with IBM's Granite Embedding Multilingual R2 (a 98 MB one-time download); for every question the four most relevant parts are handed to the chat model, labelled with the file and part number so you can check the source. Because the models here have a 4K-token window, this is what makes a 40-page PDF usable at all.

What is JSON mode?

Tick { } JSON and the answer is constrained to valid JSON: WebLLM applies a grammar while the model generates, so it cannot produce anything else. It is ideal with the Extract JSON template — “Sophie lives in Gent” comes back as {"name": "Sophie", "city": "Gent"} every time, ready to paste into code.

Is there a distraction-free or full-screen mode?

Yes. Simple hides everything on the page except the chat and lets it fill the window; the page remembers it, and ?mode=simple opens straight into it. Full screen takes the chat over your whole screen — press Esc to come back. The model stays loaded either way.

What happens when a conversation gets too long?

These models see about 4,000 tokens at a time. The small meter under the input shows how full that window is. When a conversation outgrows it, the oldest turns are quietly left out of what the model sees (they stay on your screen), so the chat keeps working instead of failing with an error.

03. Choosing a model

Which model should I pick?

Llama 3.2 1B is still the sensible default: 695 MB, quick to load and reliable at everyday instructions. For visible step-by-step reasoning pick Qwen 3 (0.6B, 1.7B, 4B or 8B) and switch 🧠 Think on. Qwen 3 4B is the best quality-per-gigabyte here if your GPU has about 3.5 GB free, and Qwen 3 8B the strongest overall. To ask about pictures, Phi 3.5 Vision is the only model that reads images. The new Qwen 3.5 models are very fast but answer each message on its own in this runtime. Press 🎯 Best for my device to let the page choose from your GPU's limits.

What do 0.5B, 1B, 3B and 7B actually mean?

The number of parameters, in billions — roughly, how much the model knows and how well it can reason. Each step up costs download size, memory and speed: a 1B model is around 880 MB and feels instant, a 3B is about 2 GB and noticeably more coherent, a 7B is over 4 GB and needs a proper graphics card. Quality rises with size, but so does the wait.

Which model is best for coding?

Qwen 2.5 Coder is trained specifically on code and comes in three sizes here: 1.5B for quick completions and shell one-liners, 3B for most everyday scripting, and 7B — the best coding model on the page — if your GPU has about 5 GB free. Qwen 3 4B with Think on is the better choice when you need the model to reason about why code fails rather than just write it. None of them will replace a frontier model on a large codebase.

What are the DeepSeek R1 distills?

Small models trained to imitate DeepSeek R1's long chain of thought. They write their reasoning before the answer; the page now folds that reasoning into a collapsible 💭 Thought for N s block so the answer stays readable, and keeps it out of the conversation history so later turns stay within the context window.

Which model is best for languages other than English?

The Qwen families are the strongest multilingual options here: Qwen 3 officially covers more than 100 languages and dialects, and Qwen 2.5 handles Chinese, Dutch, French, German, Spanish and many more well for its size. Gemma 3 and Phi 4 Mini are good alternatives. Small models drift more outside English, so keep prompts explicit — setting the system prompt to “always answer in Dutch” works better than asking mid-conversation. The Translate templates help too.

Can I run a 7B model?

Only with a GPU that has roughly 5–6 GB of free video memory — the weights alone are 4–5 GB before the conversation cache. The Large group holds them: Qwen 3 8B (the strongest model on the page), Llama 3.1 8B, Qwen 2.5 Coder 7B, the two DeepSeek R1 distills, Qwen 3.5 9B and Mistral 7B. The note under the picker shows how much GPU memory each one needs. On integrated graphics they either fail to allocate or crawl, and a 3–4B model is the better trade.

Which models does the Mobile / WASM tab use?

Small ones that run on the CPU: SmolLM2 135M (137 MB), SmolLM2 360M (365 MB, the default), Qwen 2.5 0.5B and Qwen 3 0.6B for other languages, and Llama 3.2 1B and SmolLM2 1.7B when quality matters more than speed. All run in 8-bit, which measured about 7× faster than 4-bit on the CPU (13.9 vs 1.9 tokens/s for SmolLM2 360M), and in a background thread so the page never freezes while it writes.

If I switch models, does it download again?

Each model is cached separately, so the first switch to a new model downloads it and every later switch is instant. Models you have already cached are flagged in the picker, which makes it easy to hop between a fast one and a smarter one without paying the download twice.

What is a context window, and which model has the largest?

It is how much text the model can hold in mind at once — your system prompt, the conversation so far, any document parts and its own reply all count against it. Many of these models support 32K tokens or more on paper, but WebLLM runs its builds with a window of about 4,000 tokens to keep GPU memory in check. The meter under the input shows how full it is; when a chat outgrows it, the oldest turns are left out of what the model sees. For long documents, attach them with 📎 instead of pasting — only the relevant parts are sent.

What does “q4f16” in the model name mean?

It is the quantisation: the weights are compressed to 4 bits with 16-bit floating point used during computation. That is what makes a multi-billion-parameter model small enough to download and fit in browser memory. The cost is a small amount of accuracy, which is a bargain compared with not being able to run the model at all.

04. Privacy and data

Does my data get sent anywhere?

No. Prompts and replies never leave the tab. The single network request the tool makes is the one-time download of the model weights from the Hugging Face CDN; after that you can disconnect entirely and it keeps working. There is no inference server, because your device is the inference server.

Are my chats stored or logged?

Nothing is ever uploaded. By default the page also keeps your recent conversations in this browser's local storage so you can reopen them from History — they never leave your device, and you can switch 💾 Save chats off or delete all history at any time. Attached documents and images are never stored; only the text of the conversation is.

Is browser-based AI genuinely private, or just marketed that way?

Genuinely, and you can check it yourself rather than take anyone's word for it. There is no API key in the page to send anything with, no endpoint to send it to, and the model answers with the network switched off — which is impossible for a cloud chatbot.

How can I verify nothing is uploaded?

Two ways. Open your browser's developer tools, go to the Network tab, and watch it while you send a message: after the model has loaded there are no requests at all. Or simply disable Wi-Fi and keep chatting — if the answers still come, they are being generated locally.

Can I paste confidential or GDPR-sensitive text into it?

Technically nothing leaves your device, so there is no transfer to a third party and no processor to sign an agreement with — which is exactly why local models are attractive for sensitive text. That said, your own organisation's policy still applies, so check it before pasting client data anywhere, including here.

Does the model download tell anyone what I am doing?

Downloading the weights is an ordinary file request to a CDN, so that CDN sees a file being fetched, as it would for any image or script. It carries no prompt, no conversation and no identity beyond a normal web request — and it happens once per model, not per message.

Does it work in private or incognito mode?

It works, but private windows discard storage when you close them, so the model downloads again the next time. If you plan to use it regularly, a normal window keeps the cache and saves you the wait.

05. Speed and hardware

What hardware do I need?

For the desktop tab, a GPU that supports WebGPU and has enough free memory for the model: roughly 1 GB for a 1B model, 2–3 GB for a 3B, and 5 GB or more for a 7B. Integrated graphics handle the 0.5B–1B models comfortably. If none of that applies to your machine, the Mobile / WASM tab runs on the CPU on practically anything.

How fast should it be?

On a modern discrete GPU a 1B model typically streams faster than you can read, and a 3B model still feels conversational. On integrated graphics expect it to be readable but unhurried, and in CPU / WASM mode expect a few words per second. The live tok/s counter under the chat box shows exactly what your device is doing.

Why is generation slow or stuttering?

Usually the model is too large for the hardware, so memory is being shuffled instead of used. Drop a size class, close other GPU-heavy tabs — video calls and 3D pages compete for the same memory — and check that your laptop is not in a battery-saver profile that caps the GPU. A long conversation also slows things down, because the whole history is reprocessed; reset the chat to get the speed back.

Why does the answer stop in the middle?

It hit the max tokens limit, which defaults to 1024. For scripts, long code or detailed documents raise it to 2048, 4096 or 8192 next to the Send button. It is a deliberate cap rather than a failure — without it a small model can ramble for a very long time.

What does the temperature slider do?

It controls how adventurous the sampling is. Low values (around 0.2) make the model repetitive but predictable, which is what you want for extraction, JSON and code. Higher values (0.8 and above) produce more varied writing at the cost of accuracy. The default of 0.7 is a reasonable middle for chat.

How much faster is GPU mode than WASM mode?

Measured on the same Apple-silicon laptop: Qwen 3 0.6B wrote about 60–80 tokens a second on WebGPU, while the CPU models manage 3–35 depending on size (SmolLM2 135M 35/s, 360M 14/s, Llama 3.2 1B 3.5/s). WASM mode exists so the tool still works on hardware and browsers without WebGPU.

It says out of memory, or the GPU device was lost.

The model did not fit. Choose a smaller one, close other tabs using the GPU, and try again — the error message names the limit it hit. On laptops that switch between integrated and discrete graphics, forcing the browser onto the discrete GPU in the system settings often solves it outright.

Will it drain my battery or heat the laptop?

While it is generating, yes — you are running a neural network on your own silicon, so the fans may spin up much as they would during a game. It only draws power while a reply is being produced; the model sitting in memory idle costs nothing.

06. Browsers, storage and troubleshooting

Which browsers support WebGPU today?

Chrome and Edge have shipped it by default since version 113 on Windows, macOS and ChromeOS, and on Android 12+ since Chrome 121. Safari enables it by default in macOS Tahoe 26, iOS 26 and iPadOS 26. Firefox ships it on Windows from version 141 and on Apple Silicon Macs from 147, with Linux and Android still in Nightly. Anything not on that list can still use the Mobile / WASM tab.

Why does it say “WebGPU not available”?

Your browser, or that machine's graphics driver, is not exposing WebGPU — common on older browser versions, on Linux with certain drivers, in some virtual machines, and where enterprise policy disables it. The page detects this and moves you to the Mobile / WASM tab automatically, so you can still use the tool; updating the browser is the fix if you want GPU speed.

Does it work on iPhone or iPad?

Yes, in the Mobile / WASM tab, with one caveat: iOS Safari enforces a hard memory ceiling of roughly 1–1.5 GB per tab, so only the smallest model (SmolLM2 135M) loads reliably. Larger ones crash the tab with “a problem repeatedly occurred”. For anything more capable, open the page on a desktop.

Does it work on Android?

Yes. Chrome on Android 12 and later supports WebGPU on most recent chipsets, and where it does not, the WASM tab runs on the CPU. Stick to the small models — phone memory limits bite long before the model quality does.

It will not download the model on my work network.

Corporate proxies and content filters frequently block the Hugging Face CDN, which is where the weights come from, and some inspect and break large downloads. That is a network policy question rather than a browser one — the same page usually works immediately on a home connection or a phone hotspot.

Where are the weights stored, and how do I delete them?

In the browser's Cache Storage for this site, not in your file system. Clearing site data (or “Cookies and other site data” for this domain) removes every cached model and frees the space; the tool then behaves like a first visit. Nothing outside the browser profile is touched.

Can the browser throw the cached model away?

Yes. Browsers evict cached data when disk space runs low or a profile is cleaned, and private windows discard it on close. If a model that loaded instantly yesterday starts downloading again, that is what happened — nothing is broken.

How much disk space will this use?

Only what the models you actually load take up: from 137 MB for the smallest CPU model and 204 MB for the smallest GPU model to about 5 GB for Qwen 3.5 9B. Attaching a large document also downloads the 100 MB embedding model once. Every model you try is kept until you clear site data, so trying several does add up — worth remembering on a machine that is short on space.

07. What it is good at — and what it is not

Why is it less capable than ChatGPT or Claude?

Because it is smaller by three orders of magnitude. Frontier models run on server racks with hundreds of billions of parameters; a 1B model in a browser tab is a fraction of a percent of that, and it is remarkable that it works at all. Judge it as a fast local assistant for small jobs, not as a replacement for a frontier model.

So what is it actually good at?

Short, well-defined text work: rewriting a paragraph formally or casually, summarising something you paste in, drafting a reply, extracting fields into JSON, explaining a concept simply, and small code and query snippets. The preset buttons above the chat box are shortcuts to exactly those tasks. Anything that needs current facts or long reasoning chains is the wrong job for it.

Can it search the web or read my files?

It has no internet access. It can read files you attach: text, Markdown, CSV, JSON, code and PDF. Small files are added to your message whole; larger ones are split into parts and searched locally with a multilingual embedding model, so only the most relevant passages are shown to the chat model. Everything happens in your browser — the file content is never uploaded.

Does it make things up?

More often than a large model, yes. Small models are confident and frequently wrong about names, dates, numbers and anything specialised, and they have no way to look anything up. Treat factual claims as drafts to verify, and lean on the tasks where the source material is in your prompt — rewriting, summarising, extracting.

Can I change its personality or behaviour?

Yes. Open System prompt (persona) under the model picker and describe how it should respond — “you are a terse senior engineer”, “always answer in Dutch”, “reply only with JSON”. It applies to every message in the conversation, and it is by far the most effective way to improve the output of a small model.

Can it write code?

For small, self-contained things, yes — a regex, a shell one-liner, a function, a KQL query from the preset. Use Qwen 2.5 Coder or Phi 3.5 Mini, keep the request narrow, raise max tokens so it is not truncated, and read what it produces before running it. It has no idea what is in your codebase.

Can I use the output commercially?

Generally yes, but the licence belongs to the model rather than to this page, and they differ: Llama has Meta's community licence, Gemma has Google's terms of use, most Qwen and SmolLM2 releases are Apache 2.0 and Phi is MIT. If the output is going into a product, read the licence for the specific model you used.

08. Under the hood

What technology powers this?

The desktop tab uses WebLLM 0.2.85 from MLC AI, which compiles open-weights models to run on WebGPU. The mobile tab uses Hugging Face transformers.js 4.3 on ONNX Runtime Web, executing 8-bit models on the CPU through WebAssembly in a Web Worker. Document search uses IBM's Granite Embedding Multilingual R2, and PDFs are read with Mozilla's pdf.js. Everything around them is plain JavaScript and CSS in a single HTML file.

Where do the model weights come from?

From the Hugging Face CDN, in the pre-compiled MLC format for the WebGPU tab and ONNX for the WASM tab. They are open-weights models published by Meta, Alibaba (Qwen), Microsoft (Phi), Google DeepMind (Gemma), Mistral AI, Hugging Face (SmolLM2), DeepSeek and IBM (the Granite embedding model used for document search) — the same files anyone can download and run locally.

Is this the same as Ollama or LM Studio?

Same idea, different delivery. Those are desktop applications you install, and they will always be faster and support far larger models because they use your hardware directly. This runs the model in a browser tab with nothing installed at all, which makes it the easier way to try a local model — on a locked-down machine, on someone else's computer, or simply to see what the fuss is about.

When should I use a cloud model instead?

When the task needs breadth, accuracy on facts, long documents, up-to-date information or serious reasoning — a frontier model will do in one attempt what a 1B model cannot do at all. Use this when the text is sensitive, when you are offline, when you want an instant answer without opening an account, or when the job is small enough that a small model is simply the quicker tool.