~/granite-ai ☀ LIGHT ☕ Support me apps ← about me
// Apache 2.0 · runs on your GPU · nothing uploaded

IBM Granite AI in Your Browser — Speech, Chat, Documents & Search

Granite Speech 4.1, the hybrid Granite 4.0-H chat models, Docling and multilingual Granite Embedding R2 — running entirely on your own machine via WebGPU. Transcribe and translate audio, ask questions, turn a scanned page into Markdown, or search text by meaning in 200+ languages, with no account, no server and nothing ever uploaded.

What loading looks like: the download bar moves quickly, then there is a second phase where the model is compiled for your GPU. That step reports no progress anywhere — it is the browser building a shader for every operator in the network — so expect a pause of a few seconds for the small models and up to a couple of minutes for Micro 3.4B. It only happens on first load; after that the model is cached and starts in seconds.
New: Granite Speech 4.1 now runs here — punctuated transcripts, translation into six languages and keyword biasing — alongside the multilingual Granite Embedding R2 (search across 200+ languages) and the hybrid Mamba-2 Granite 4.0-H chat models, from 238 MB. Granite 4.1 and 4.2 chat models still have no browser build; what runs and what doesn’t →
System prompt & sampling
Granite Docling 258M — turns a photo or scan into Markdown, tables included
Drop an image or screenshot here, or click to choose
PNG, JPG or WebP · a photo of a page, a receipt, a screenshot, a form
The extracted text will appear here.
English, French, German, Spanish, Portuguese, Japanese
Drop an audio file here, or click to choose
MP3, WAV, M4A, OGG · a voice note, an interview, a meeting — long files are split at pauses
Granite Speech does speech translation, not just transcription — French audio in, English text out, in one step.
The transcript will appear here.
Paste a list — support tickets, notes, product lines, FAQ entries — one per line. Then ask for something in your own words. It ranks by meaning, so “my card was declined” finds “payment failed at checkout” with no shared keywords.
Examples:
Results will appear here, best match first.
💾 Models stored in this browser
Reading the browser cache…

Why run a model on your own machine

Every mainstream AI assistant sends your text to someone else's server. Usually that is fine. Sometimes it is not: an unpublished manuscript, a client contract, patient notes, an internal incident report, source code under NDA. The moment that content leaves your device it is covered by somebody's retention policy rather than your own judgement.

Local inference removes the question entirely. The model weights download once, then everything — your prompt, the document you drop in, the audio you record — is processed by your own GPU inside this browser tab. There is no account to create, no request to a server, and no log anywhere. You can prove it: load a model, disconnect from the internet, and keep working.

Every Granite model here, measured

Each model below was loaded and run in this page’s own runtime (transformers.js 4.3 on WebGPU) before it was listed. Download sizes are the exact bytes fetched at the precision the page uses.

ToolModelDownloadGood for
ChatGranite 4.0-H 350M · hybrid Mamba-2238 MBThe fastest first answer on this page — rewriting, short summaries
ChatGranite 4.0-H 1B · hybrid Mamba-2927 MBThe best quality per megabyte; noticeably sharper than the 350M
ChatGranite 4.0 350M / 1B709 MB / 1.78 GBThe dense transformer versions, kept for comparison
ChatGranite 4.0-H Micro / 4.0 Micro 3B1.95 GB / 2.30 GBStrongest instruction following, for a machine with a decent GPU
Read a documentGranite Docling 258M1.15 GBA photo, scan, receipt or screenshot to Markdown — tables included
SpeechGranite Speech 4.1 2B1.49 GBPunctuated transcripts, translation, keyword biasing — 6 languages
SpeechGranite Speech 4.0 1B1.49 GBThe previous release, lowercase transcripts, kept for comparison
SearchGranite Embedding Multilingual R2 97M98 MBFinding text by meaning across 200+ languages — a Dutch query finds an English answer
SearchGranite Embedding English R2 47M / English 30M52 MB / 30 MBEnglish-only lists; the smallest downloads on the page

Start with Search by meaning if you want to see the idea work in seconds, and try the mixed languages example: type “waar blijft mijn pakket” and the multilingual model ranks “The parcel never arrived” first, with a similarity of 0.82. The English-only models cannot do that — switch and see them fail. The chat tab defaults to the 238 MB hybrid model for the same reason: a fast first answer beats a better one you never wait for.

What the “H” means

Granite 4.0-H models are hybrids: most layers are Mamba-2 state-space layers, with a few attention layers kept for precision. A state-space layer carries a fixed-size summary of everything so far instead of a key/value cache that grows with every token, which is why IBM ships them for long documents and long agent sessions on modest hardware. In this browser the practical effect is smaller downloads — the 350M hybrid is 238 MB against 709 MB for the dense 350M — and memory that stays flatter as a conversation grows. They need the Mamba kernels added in transformers.js 4, which is why they were not on this page before.

Granite 4.1 and 4.2: what runs in a browser today

IBM shipped Granite 4.1 on 29 April 2026 and Granite 4.2 — with a native thinking / non-thinking switch — on 25 August 2026. Not every release has a build a browser can run. The honest state, checked on 21 September 2026:

ModelBrowser build?On this page
Granite Speech 4.1 2BYes — ONNX for transformers.jsRuns — Speech tab
Granite Embedding Multilingual R2 (97M)Yes — ONNX for transformers.jsRuns — Search tab
Granite 4.0-H 350M / 1B / MicroYes — ONNX, needs transformers.js 4Runs — Chat tab
Granite 4.1 3B / 8B / 30BOnly for onnxruntime-genai; the 3B is a single 6.8 GB full-precision fileNot browser-viable
Granite 4.2 3B / 8B / 30B (reasoning)No ONNX build yetNot available
Granite Vision 4.1, Guardian 4.1 8BNo ONNX buildNot available

So the Speech and Search tabs are on the newest Granite generation, while the chat models are still 4.0 — not by choice, but because the newer language models have no web build. When one appears it will be added here. If you want Granite 4.1 or 4.2 chat today, run them locally with Ollama or LM Studio.

Frequently asked questions

Is anything I type or upload sent to a server?

No. The only network requests are for the model weights themselves, fetched once from Hugging Face's CDN and then cached by your browser. After that your prompts, documents and audio are processed entirely on your device. Disconnect from the internet once a model has loaded and everything still works — which is the simplest proof available.

Why is the first load so large?

Because you are downloading an actual neural network, not calling an API. A 350M-parameter model quantised to 4-bit is roughly 350 MB of weights. That is the price of the model living on your machine instead of someone else's. It happens once — the browser caches it, and returning visits load from disk in seconds. The loader tells you when that has happened.

The download finished but it sat on "preparing" for ages. Is that normal?

Yes, and it is unavoidable. Loading happens in two phases. First the weights download, which is network-bound and usually quick. Then ONNX Runtime has to parse the graph, allocate GPU buffers and compile a WebGPU shader for every operator in the network. That second phase is compute-bound, happens entirely on your machine, and reports no progress of any kind — there is no event to hook into, which is why you get an elapsed timer rather than a percentage.

Expect a few seconds for the 350M and embedding models, and up to a couple of minutes for Micro 3.4B on a modest GPU. It only happens on the first load of each model; afterwards both the weights and much of the compiled result are cached, and startup drops to seconds.

What is the .onnx_data file being downloaded?

It holds the actual model weights. The ONNX format stores a network as a protobuf file, which has a 2 GB size limit, so anything larger keeps the graph structure in model_q4f16.onnx and moves the weight tensors into one or more external .onnx_data files beside it. Seeing several of them — _data, _data_1 and so on — is completely normal for the bigger models and simply means the weights were split across files.

Do I need to download it again next time?

No. Weights are stored in the browser’s Cache Storage, so the second visit skips the download and the model list marks it ✓ on this device. The Models stored in this browser panel under the tools shows every cached model with its size and lets you delete one or all of them. Clearing site data does the same. Each model and each precision is cached separately.

What is WebGPU and do I need it?

WebGPU is the browser API that lets a page use your graphics card for computation. With it, generation runs at a usable speed. Without it the page falls back to WASM on your CPU, which still works but is several times slower — fine for search, painful for the larger chat models. WebGPU ships in recent Chrome, Edge and Brave on desktop, and in Safari 26 and later, which this page’s runtime now enables. Firefox support is newer and may still be behind a setting.

Which chat model should I pick?

Start with Granite 4.0-H 350M: at 238 MB it is the quickest to load and fine for rewriting, summarising and short questions. Once you know you want more, 4.0-H 1B (927 MB) is the best quality per megabyte on the page. The Micro models are the strongest and slowest, worth their two gigabytes only on a machine with a decent GPU. The dense 4.0 350M and 1B are kept for comparison; the hybrids are smaller at the same size class. Every model you load stays cached, so sizing up later costs nothing twice.

How good are these compared to ChatGPT or Claude?

Not close, and it would be dishonest to suggest otherwise. Frontier models have hundreds of billions of parameters; a 1B model is roughly 0.3% of that. What these are genuinely good at is bounded, mechanical work — rewriting a paragraph, summarising something you paste in, extracting structured data, drafting boilerplate. Do not rely on them for facts, current events or multi-step reasoning. The trade you are making is capability for privacy and cost.

The model output just "QQQQQQQ" or one repeated character. What went wrong?

That is a numerical failure, not a prompting problem. Granite 4.0 uses a hybrid Mamba/SSM architecture where state accumulates across every token generated. If the weights are quantised too aggressively for that maths to hold up on your hardware, the values collapse, every token scores identically, and the model emits the same character forever.

The fix is precision. Each model needs a different one — the 350M runs at fp16, the 1B at q4, and Micro 3.4B at q4f16, matching what IBM ships in their own reference demo. Those are the defaults here. If you still hit it, change Precision in the bar above to fp32, which is the most numerically robust option, and load the model again. The page detects a run of repeated characters and stops rather than filling the screen.

What do fp32, fp16, q4 and q4f16 mean?

They describe how precisely the model's numbers are stored. fp32 is full 32-bit precision — largest download, most reliable. fp16 halves that. q4 compresses the weights to 4 bits while computing in higher precision, and q4f16 does both, giving the smallest file.

Smaller is not simply worse-but-workable: below a certain point the arithmetic stops being stable, especially for SSM layers. Bigger models tolerate heavy quantisation better because they have more redundancy, which is why Micro is fine at q4f16 while the 350M is not.

Will it make things up?

Yes, more readily than a large model. Small models hallucinate confidently, particularly on dates, names, numbers and anything requiring knowledge rather than manipulation. Treat every factual claim as unverified. Where they are reliable is when you give them the material — paste the text and ask for a summary or a rewrite, rather than asking what they know.

What is Granite Docling and what can it read?

It is IBM's small document-understanding model — it looks at an image and produces text. It handles photographed pages, scans, receipts, forms, screenshots and simple tables. It is the most downloaded Granite browser model by a wide margin, which is why it is here. Handwriting is unreliable, very low-resolution images produce poor results, and complex multi-column layouts can come out in an unexpected order.

Why does the document output contain tags like <loc_42> and <otsl>?

Because Docling does not produce plain prose — it produces DocTags, a structured markup that records not just the words but where they sit on the page. The <loc_…> values are coordinates, and tags like <otsl>, <section_header> and <caption> record document structure. That is the whole point of the model: it understands layout rather than just running OCR.

The panel strips all of that for readability by default. Tick show raw DocTags to see the original, which is what you want if you intend to rebuild the document's structure rather than just lift the text out of it.

What are the different document prompts for?

Docling is prompt-directed rather than one-size-fits-all, and these are the phrasings IBM uses in their own demo. Convert this page to docling handles a full page. Convert this table to OTSL targets a table specifically — OTSL is a compact table markup that preserves cells and spans. Convert chart to OTSL pulls the underlying values out of a chart image. Convert code to text is tuned for code screenshots. Using the matching prompt gives noticeably better results than a generic instruction.

Can it read PDFs?

Not directly — it takes images. For a PDF, convert the page to an image first with the PDF to JPG converter and drop that in. If your PDF already contains a real text layer you do not need a model at all: copy the text out, which is faster and perfectly accurate.

Which languages does the speech model handle?

Both speech models cover English, French, German, Spanish, Portuguese and Japanese. Granite Speech 4.1 adds punctuation and capitalisation in all six, translation into Japanese, and keyword biasing; IBM reports a mean word error rate of 5.33% on the Open ASR leaderboard for it. Accuracy is best on clear single-speaker audio. It transcribes rather than diarises, so it will not tell you who said what. Dutch is not a supported language.

Why was the speech tab disabled before?

It is not any more. It used to be: Granite Speech needs a class called GraniteSpeechForConditionalGeneration, which no transformers.js 3.x release shipped. The page now runs transformers.js 4.3, which has it, and both Granite Speech 4.0 and 4.1 were run end to end in the browser before the tab was switched on. If you ever see the tab greyed out again, the page detected at startup that your browser loaded a build without that class, and disabled it rather than letting you download 1.5 GB for nothing.

Can it translate speech, or only transcribe it?

Both. Granite Speech does bidirectional speech translation as well as transcription, so it can take French, German, Spanish or Portuguese audio and return English text directly — no separate translation step. Pick the target from the dropdown above the Run button. Transcribing as spoken is the default, and if the model gets a prompt it does not understand it falls back to plain transcription rather than failing.

Can I record straight from my microphone?

Yes — press Record from microphone, allow access, and press it again to stop. The recording stays in the page and is transcribed locally, exactly like a dropped file. Nothing is uploaded, which makes it usable for the kind of meeting note you would not send to a transcription service.

How long can the audio be?

There is no fixed limit. Audio longer than about 30 seconds is cut into sections — at the quietest moment in the last few seconds of each window, so the cut falls between words — and the transcript streams in section by section. On an Apple-silicon laptop a 60-second talk transcribed in 12.6 seconds, roughly five times faster than real time. Memory, not a limit, is what eventually stops a very long recording: the whole file is decoded into memory first, so an hour-long meeting is better split before you drop it in.

What does "search by meaning" actually do?

It converts every line and your query into a vector — a list of numbers representing meaning — then ranks lines by how close they sit to the query. That is why “I can't pay for my order” finds “My card was declined at checkout” despite sharing no words, and why the multilingual model finds it from “Mijn betaling is mislukt” too. It is the retrieval half of what people call RAG, running in your browser on a 30–98 MB model.

What would I use semantic search for?

Sorting support tickets by topic, finding the closest existing FAQ entry to a new question, deduplicating a list of feature requests, or matching a description against a product catalogue. Anything where people describe the same thing in different words and keyword search consequently fails.

What is IBM Granite, and what licence is it under?

Granite is IBM's family of open models aimed at enterprise use — trained with documented data provenance, which matters to organisations that need to answer questions about what a model was trained on. They are released under Apache 2.0, a genuinely permissive licence: commercial use, modification and redistribution are all allowed, with no user-count restrictions of the kind some “open” model licences include.

Can I use the output commercially?

Yes. Apache 2.0 places no restriction on what you do with what the model produces, and this page adds none. Check the licence text yourself if you are making a decision that matters — but for ordinary commercial use, Granite is about as unencumbered as open models get.

Does this work offline?

Once a model is cached, yes, completely. It is genuinely useful on a plane, behind a restrictive corporate firewall, or anywhere the network is unreliable. The one caveat is that the page itself and the transformers.js library must load first, so visit once while connected before relying on it.

Will this work on my phone?

Partly. Phones have far less memory and mobile WebGPU is still uneven, so the CPU fallback often kicks in. Search by meaning is fine — its models run on the CPU anyway. The 238 MB chat model is usable on a recent phone; the gigabyte models will be slow or fail. In practice this is a desktop tool.

Why is generation slow, and what is tokens per second?

A token is roughly three-quarters of a word, and the counter shows how many the model produces each second along with how long the first one took. Speed depends almost entirely on your GPU and the model size — a discrete GPU running the 350M model feels instant, while an integrated one running Micro can drop to a few tokens per second. If it is slower than you want, pick a smaller model.

Can I change how it behaves?

Yes. Open System prompt & sampling under the chat and rewrite the system prompt — it is applied to every message. Temperature controls randomness: near 0 makes it repeatable and literal, which suits extraction and code, while 0.7–1.0 gives more varied prose. Max new tokens caps the response length; raise it if answers are cut off mid-sentence.

Are my conversations saved?

No. The conversation lives in the page and is gone when you close the tab — nothing is written to storage and nothing is uploaded. Use Export .md if you want to keep it. Deliberately nothing is persisted, since the point of a local tool is that sensitive material does not linger.

Can I use this in a corporate environment?

Technically it is well suited to it — Apache 2.0, no data egress, no vendor account. Two practical notes: the one-time weight download comes from Hugging Face's CDN, which some corporate proxies block, and you should confirm your own policy on running models locally. Nothing here transmits company data anywhere, which is usually the hard part of that conversation.

Why does the model sometimes stop mid-sentence?

It hit the max new tokens limit. Raise it in the sampling panel — 1024 or 2048 is sensible for longer answers. The trade-off is time: every extra token is more work for your GPU, and the default of 512 keeps responses quick.

Is it free, and are there ads or limits?

Free, no account, no ads, no message cap and no subscription. There is nothing to meter — the compute is yours, not a server's, so it costs nothing to run regardless of how much you use it. Analytics on this site are privacy-first and cookie-free and record only that a page was viewed.

What are the Granite 4.0-H models, and why are they smaller?

They are hybrid models: mostly Mamba-2 state-space layers with a few attention layers. A state-space layer keeps a fixed-size running summary instead of a key/value cache that grows with every token, so memory stays flatter over a long conversation. Here they also ship at 4-bit with fp16 activations (q4f16) without degrading, which makes the 350M hybrid 238 MB against 709 MB for the dense 350M. They need transformers.js 4 to run at all.

Can it search across languages?

Yes, with Granite Embedding Multilingual R2, which IBM trained on 200+ languages. Load the mixed languages example: “waar blijft mijn pakket” ranks “The parcel never arrived” first, “mot de passe oublié” finds “How do I reset my password?”, and “cancel my subscription” finds “Ik wil mijn abonnement opzeggen”. The English models rank the wrong line first on all three, which is the clearest way to see what multilingual training buys you.

What is keyword biasing in Granite Speech 4.1?

You give the model a short list of words it is likely to mishear — product names, surnames, acronyms — and it favours those spellings when the audio is ambiguous. Type them comma-separated in the Keywords box; the page appends them to the transcription prompt exactly as IBM’s model card specifies. It works with the 4.1 model only.

Can the document reader give me Markdown and tables?

Yes. Docling answers in DocTags, a markup that keeps every block’s position and type. The page rebuilds that into Markdown by default: headings stay headings, list items stay lists, and tables — which Docling emits in a compact format called OTSL — become real Markdown tables with every cell in its column. Switch to plain text for tab-separated tables, or to raw DocTags if you want the layout coordinates, and download any of the three.

How do I free up the disk space the models use?

Open Models stored in this browser below the tools. It lists every downloaded model, its size and how much of your browser’s storage allowance this site is using, with a Delete button per model and a Delete all. The list reads the browser’s own cache directly, so what it shows is what is actually on your disk.

Is Granite 4.2 with reasoning available here?

Not yet. Granite 4.2 (3B, 8B and 30B, released 25 August 2026) adds a thinking / non-thinking switch, but there is no ONNX build a browser can load, and its smallest model is 3B. Granite 4.1’s language models have ONNX files only for onnxruntime-genai, a native runtime — the 3B is one 6.8 GB full-precision file. When a transformers.js build of either appears, it will be added here.

Can I use the chat full screen, without everything else on the page?

Yes, two ways. Simple hides the whole page except the chat — no header, explanations, notices or other tools — and lets the conversation fill the window; the page remembers it, and ?mode=simple opens straight into it. Full screen takes the chat over your entire screen; press Esc or Exit full screen to come back. Both keep the model loaded and the conversation intact, and New chat clears the conversation without reloading the model.

Which version of transformers.js does this use, and why does it matter?

Version 4.3. The 4.x line replaced the WebGPU runtime with one rewritten in C++, added the Mamba-2 and Granite Speech support this page depends on, made BERT-style embedding models several times faster, and enabled WebGPU on Safari 26. Two Granite Speech quirks in it are worked around here: audio of certain exact lengths makes the encoder produce three more audio tokens than the processor expected, or crashes the GPU session outright, so every section is padded with 20 ms of silence to a safe length before it is encoded.

How do I run Granite outside the browser?

For the larger models — Granite 4.1 and 4.2 at 8B and 30B — use Ollama or LM Studio on your desktop, which take the GGUF builds directly. On Apple Silicon the MLX builds are faster still. Granite 4.2’s thinking switch is set in its chat template, so check that your runner passes it through. This page exists for the case where installing anything is not an option.