PaddleOCR-VL Local Web UI is a small FastAPI application that queues OCR jobs and runs the official PaddleOCR-VL command inside Docker. Documents stay on your machine: uploads, logs, generated Markdown, JSON, text, and previews are stored below the local data/ directory.
Important
This project is designed for a trusted local machine. It has no authentication. Keep the default 127.0.0.1 binding and do not expose it directly to the internet.
- Drag-and-drop PDF and image uploads.
- CPU or NVIDIA GPU inference.
- Background job queue with live logs and status history.
- Preview for Markdown, plain text, JSON, and images.
- One-click ZIP download for each job.
- Persistent PaddleOCR model caches between runs.
- No cloud upload or external database required.
flowchart LR
A["Browser on localhost"] -->|"PDF / image"| B["FastAPI web server"]
B --> C["Local job queue"]
C --> D["PaddleOCR-VL Docker container"]
D --> E["Markdown, text, JSON, images"]
E --> F["Local data/jobs directory"]
F --> A
The Python server owns the queue and filesystem layout. Each job starts a short-lived container from the configured PaddleOCR-VL image. Only the job input, output directory, and local model caches are mounted into that container.
- Python 3.10 or newer
- Docker Desktop or Docker Engine
- Around 10 GB of free disk space for the OCR image and model cache
- For GPU mode: an NVIDIA GPU, current drivers, and NVIDIA Container Toolkit support
The default inference image is:
ccr-2vdh3abv-pub.cnc.bj.baidubce.com/paddlepaddle/paddleocr-vl:latest-nvidia-gpu
git clone https://github.com/egore4606/paddle-ocr-ui.git
cd paddle-ocr-ui
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
pip install -r server\requirements.txt
docker pull ccr-2vdh3abv-pub.cnc.bj.baidubce.com/paddlepaddle/paddleocr-vl:latest-nvidia-gpu
uvicorn server.app:app --host 127.0.0.1 --port 8000git clone https://github.com/egore4606/paddle-ocr-ui.git
cd paddle-ocr-ui
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
pip install -r server/requirements.txt
docker pull ccr-2vdh3abv-pub.cnc.bj.baidubce.com/paddlepaddle/paddleocr-vl:latest-nvidia-gpu
uvicorn server.app:app --host 127.0.0.1 --port 8000Open http://127.0.0.1:8000, choose CPU or GPU, upload one or more supported files, and start the job. The first run can take a while because Docker and PaddleOCR download several gigabytes.
| Type | Extensions |
|---|---|
| Documents | .pdf |
| Images | .png, .jpg, .jpeg, .bmp, .tif, .tiff, .webp |
Results are written to data/jobs/<job-id>/output/. The entire data/ directory is ignored by Git and should be treated as private because OCR results can contain sensitive document content.
The current release intentionally keeps configuration small:
- the PaddleOCR-VL v1 pipeline is used for every job;
- GPU is the default device, with CPU available in the UI;
- model caches live in
~/.paddleocr-vl-cache; - application data lives in the repository's
data/directory.
See Architecture for implementation details and Troubleshooting for Docker, GPU, startup, and cache problems.
python -m venv .venv
source .venv/bin/activate # Windows: .\.venv\Scripts\Activate.ps1
pip install -r requirements-dev.txt
ruff check .
pip-audit -r server/requirements.txt
python -m compileall server
node --check web/app.js
pytest --cov=server --cov-report=term-missingDocker is not required for the unit test suite. See CONTRIBUTING.md before opening a pull request.
Uploads and OCR output may contain confidential information. Do not commit data/, .env, logs, or model cache files. If you discover a vulnerability, use GitHub's private vulnerability reporting instead of opening a public issue. See SECURITY.md.
- Job cancellation and cleanup controls
- Configurable storage and upload limits
- More export formats
- Better progress reporting from the inference container
- Optional authentication for deliberate network deployments
Ideas and questions are welcome in Discussions.
Released under the MIT License. PaddleOCR and the referenced container image are separate projects with their own licenses and terms.