A private, client-side interface for running LLMs directly in your browser.
No Backend Required · WebGPU Accelerated · Local Privacy
WebLLM Pro is a lightweight, single-file HTML application that brings the power of Large Language Models (like Llama-3, Gemma, and Qwen) to your web browser.
Leveraging WebGPU technology, all model inference happens locally on your machine's GPU. This means your conversation data never leaves your device, ensuring 100% privacy and zero server costs.
- 🔒 Privacy First: No data is sent to the cloud. Everything runs inside your browser.
- 💾 Model Manager: Built-in interface to download, cache, and delete models (Llama-3, etc.).
- 💬 Persistent Chat: Conversations are automatically saved to your browser's local storage.
- 🎨 Rich UI Experience:
- Markdown rendering (tables, lists, math).
- Code syntax highlighting (Highlight.js).
- Bilingual support (English / Chinese).
index.html file. You must serve it via a local HTTP server.
- Install the Live Server extension in VS Code.
- Right-click
index.htmlin the file explorer. - Select "Open with Live Server".
If you have Python installed, run the following command in the project directory:
# Python 3
python -m http.server 8000Then open http://localhost:8000 in your browser.
The first time you load a model, the application will download model weights (approx. 2GB - 5GB). Please ensure you have a stable internet connection.
👇 Click to view Hardware & Browser Requirements
- Browser:
- Chrome 113+ or Edge 113+ (Chromium-based browsers are recommended).
- Must support WebGPU.
- GPU (Graphics Card):
- Integrated (Minimum): Apple M1/M2/M3 or Intel Iris Xe. Capable of running smaller models like
Qwen-1.5-1.8B. - Discrete (Recommended): NVIDIA RTX 3060+ or equivalent with 6GB+ VRAM. Required for larger models like
Llama-3-8B.
- Integrated (Minimum): Apple M1/M2/M3 or Intel Iris Xe. Capable of running smaller models like
| Inference Engine | Core Logic | Rendering | Styling |
| WebLLM | Vanilla JS (ES Modules) | Marked.js | Native CSS3 |
This project is open-source and available under the MIT License.