Skip to content

Latest commit

 

History

850 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

English | 简体中文 | 한국어 | 日本語

MangaTranslator

Gradio-based web application for automating the translation of manga/comic page images using AI. Targets speech bubbles and text outside of speech bubbles. Supports 60 languages and custom font pack usage.

Original Translated (w/ a single click)

Table of Contents

Features

  • Detection: Speech bubble detection & segmentation (YOLO, SAM 2.1/3)
  • Cleaning: Inpaint speech bubbles and OSB text (FLUX.2 Klein, FLUX.1 Kontext, or OpenCV)
  • Translation: LLM-powered OCR & translation (60 languages)
  • Rendering: Custom text rendering engine with alignment and custom font packs
  • Upscaling: Text region and full page artwork upscaling (2x-AnimeSharpV4)
  • Processing: Single/batch processing with directory preservation and ZIP support
  • Configuration: Flexible controls to adapt to diverse page layouts and fine-tune output quality
  • Interfaces: Web UI (Gradio) and CLI
  • Automation: One-click translation; no intervention required

Requirements

  • Python 3.10+
  • PyTorch (CPU, CUDA, ROCm, XPU, MPS)
  • Font pack with .ttf/.otf files; included with portable package
  • LLM for Japanese source text; VLM for other languages (API or local)

Install

Portable Package (Recommended)

Download the standalone zip from the releases page: Portable Build

Requirements:

  • Windows: Bundled Python/Git included; no additional requirements
  • Linux/macOS: Python 3.10+ and Git must be installed on your system

Tip

In the event that you need to transfer to a fresh portable package:

  • You can safely move the fonts, models, and output directories to the new portable package
  • You might be able to move the runtime directory over, assuming the same setup configuration is wanted

Manual install

  1. Clone and enter the repo
git clone https://github.com/meangrinch/MangaTranslator.git
cd MangaTranslator
  1. Create and activate a virtual environment (recommended)
python -m venv venv
# Windows PowerShell/CMD
.\venv\Scripts\activate
# Linux/macOS
source venv/bin/activate
  1. Install PyTorch (see: PyTorch Install)
# Example (CUDA 13.0)
pip install torch==2.11.0+cu130 torchvision==0.26.0+cu130 --extra-index-url https://download.pytorch.org/whl/cu130
# Example (ROCm 7.1)
pip install torch==2.11.0+rocm7.1 torchvision==0.26.0+rocm7.1 --extra-index-url https://download.pytorch.org/whl/rocm7.1
# Example (XPU)
pip install torch==2.11.0+xpu torchvision==0.26.0+xpu --extra-index-url https://download.pytorch.org/whl/xpu
# Example (MPS/CPU)
pip install torch==2.11.0 torchvision==0.26.0
  1. Install Nunchaku (optional, for FLUX.1 Kontext via Nunchaku backend)
  • Nunchaku wheels are not on PyPI. Install directly from the v1.3.0dev20260213 GitHub release URL, matching your OS and Python version. CUDA only, and requires a 2000-series card or newer.
# Example (Windows, Python 3.13, PyTorch 2.11.0, CUDA 13.0)
pip install https://github.com/nunchaku-ai/nunchaku/releases/download/v1.3.0dev20260213/nunchaku-1.3.0.dev20260213+cu13.0torch2.11-cp313-cp313-win_amd64.whl

# Example (Linux, Python 3.13, PyTorch 2.11.0, CUDA 13.0)
pip install https://github.com/nunchaku-ai/nunchaku/releases/download/v1.3.0dev20260213/nunchaku-1.3.0.dev20260213+cu13.0torch2.11-cp313-cp313-linux_x86_64.whl

Note

Nunchaku is not necessary for the use of Flux models via the sd.cpp/SDNQ backends.

  1. Install dependencies
pip install -r requirements.txt

Post-Install Setup

Models

  • The application will automatically download and use all required models

Fonts

  • Put font packs as subfolders in fonts/ with .otf/.ttf files
  • Prefer filenames that include italic/bold or both so variants are detected
  • Example structure:
fonts/
├─ CC Wild Words/
│  ├─ CCWildWords-Regular.otf
│  ├─ CCWildWords-Italic.otf
│  ├─ CCWildWords-Bold.otf
│  └─ CCWildWords-BoldItalic.otf
└─ Komika Hand/
   ├─ KOMIKA-HAND.ttf
   └─ KOMIKA-HANDBOLD.ttf

LLM setup

  • Providers: Google, OpenAI, Anthropic, SpaceXAI, Meta Model, DeepSeek, Z.ai, Moonshot AI, Xiaomi MiMo, QwenCloud, OpenCode, OpenRouter, OpenAI-Compatible
  • Web UI: configure provider/model/key in the Config tab (stored locally)
  • CLI: pass keys/URLs as flags or via env vars
  • Env vars: GOOGLE_API_KEY / GEMINI_API_KEY, OPENAI_API_KEY, ANTHROPIC_API_KEY, SPACEXAI_API_KEY / XAI_API_KEY, META_MODEL_API_KEY / META_API_KEY, DEEPSEEK_API_KEY, ZAI_API_KEY, MOONSHOT_API_KEY, MIMO_API_KEY, QWENCLOUD_API_KEY / QWEN_API_KEY, OPENCODE_API_KEY / OPENCODE_ZEN_API_KEY / OPENCODE_GO_API_KEY, OPENROUTER_API_KEY, OPENAI_COMPATIBLE_API_KEY
  • OpenAI-Compatible provider supports local endpoints (e.g., http://localhost:8080/v1), and Azure OpenAI endpoints (e.g., https://<resource>.openai.azure.com)

Note

The following models are automatically detected when used via the OpenAI-Compatible provider and receive optimized prompting. They are text-only and require two-step translation + local OCR. The special_instructions field maps to their corresponding glossary/terminology (one entry per line, e.g., term -> translation).

  • YanoljaNEXT-Rosetta (e.g., yanolja/YanoljaNEXT-Rosetta-4B-2511-GGUF)
  • Hy-MT2 (e.g., tencent/Hy-MT2-7B). Also pre-fills the model's recommended sampling parameters

OSB text setup (optional)

If you want to use the OSB text pipeline, you need a Hugging Face token with access to the following repositories:

  • deepghs/AnimeText_yolo

Steps to create a token:

  1. Sign in or create a Hugging Face account
  2. Visit and accept the terms on:
  3. Create a new access token in your Hugging Face settings with read access to gated repos ("Read access to contents of public gated repos")
  4. Add the token to the app:
    • Web UI: set hf_token in Config
    • Env var (alternative): set HF_TOKEN
  5. Save config to preserve the token across sessions

Run

Web UI (Gradio)

  • Portable package:
    • Run start-webui.bat (Windows) or ./start-webui.sh (Linux/macOS), located in MangaTranslator/
  • Manual install:
    • Run python app.py --open-browser

Run python app.py --help for launch options. First launch can take ~1–2 minutes.

Once launched, configure your LLM provider in the Config tab, then upload images and click Translate.

CLI

Examples:

# Single image, Japanese → English, Google provider, OSB text pipeline, custom OSB text font
python main.py --input <image_path> \
  --font-dir "fonts/Komika Hand" --provider Google --google-api-key <...> \
  --osb-enable --osb-font-dir "fonts/Comicka"

# Batch folder, Japanese → Chinese (Simplified), OpenAI-Compatible provider (llama.cpp), OSB text pipeline, custom OSB text font
python main.py --input <folder_path> --batch \
  --font-dir "fonts/Noto Sans SC" --output-language "Chinese (Simplified)" \
  --provider OpenAI-Compatible --openai-compatible-url http://localhost:8080/v1 \
  --output ./output --osb-enable --osb-font-dir "fonts/Noto Sans SC"

# Cleaning-only mode (no translation)
python main.py --input <image_path> --cleaning-only

# Upscaling-only mode (no translation)
python main.py --input <image_path> --upscaling-only --image-upscale-mode final --image-upscale-factor 2.0

# Full options
python main.py --help

Documentation

Updating

Portable Package

  • Run update.bat (Windows) or ./update.sh (Linux/macOS) from the portable package root

Manual Install

From the repo root:

git pull
pip install -r requirements.txt  # Or activate venv first if present

License & credits

ML Models & Libraries

Releases

Packages

Contributors

Languages