Skip to content
tongwentaoPublic

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

WebLLM Pro 🤖

A private, client-side interface for running LLMs directly in your browser.

No Backend Required · WebGPU Accelerated · Local Privacy

WebLLM WebGPU License


📖 Introduction

WebLLM Pro is a lightweight, single-file HTML application that brings the power of Large Language Models (like Llama-3, Gemma, and Qwen) to your web browser.

Leveraging WebGPU technology, all model inference happens locally on your machine's GPU. This means your conversation data never leaves your device, ensuring 100% privacy and zero server costs.

✨ Key Features

  • 🔒 Privacy First: No data is sent to the cloud. Everything runs inside your browser.
  • 💾 Model Manager: Built-in interface to download, cache, and delete models (Llama-3, etc.).
  • 💬 Persistent Chat: Conversations are automatically saved to your browser's local storage.
  • 🎨 Rich UI Experience:
    • Markdown rendering (tables, lists, math).
    • Code syntax highlighting (Highlight.js).
    • Bilingual support (English / Chinese).

🚀 Quick Start

⚠️ Note: Due to browser security policies (CORS), you cannot run this app by simply double-clicking the index.html file. You must serve it via a local HTTP server.

Option 1: Using VS Code (Recommended)

  1. Install the Live Server extension in VS Code.
  2. Right-click index.html in the file explorer.
  3. Select "Open with Live Server".

Option 2: Using Python

If you have Python installed, run the following command in the project directory:

# Python 3
python -m http.server 8000

Then open http://localhost:8000 in your browser.


⚠️ System Requirements

The first time you load a model, the application will download model weights (approx. 2GB - 5GB). Please ensure you have a stable internet connection.

👇 Click to view Hardware & Browser Requirements
  • Browser:
    • Chrome 113+ or Edge 113+ (Chromium-based browsers are recommended).
    • Must support WebGPU.
  • GPU (Graphics Card):
    • Integrated (Minimum): Apple M1/M2/M3 or Intel Iris Xe. Capable of running smaller models like Qwen-1.5-1.8B.
    • Discrete (Recommended): NVIDIA RTX 3060+ or equivalent with 6GB+ VRAM. Required for larger models like Llama-3-8B.

🛠️ Tech Stack

Inference Engine Core Logic Rendering Styling
WebLLM Vanilla JS (ES Modules) Marked.js Native CSS3

📄 License

This project is open-source and available under the MIT License.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages