TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
-
Updated
Jul 20, 2026 - C++
TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
ComfyUI-QwenVL custom node: Integrates the Qwen-VL series, including Qwen2.5-VL and the latest Qwen3-VL, with GGUF support for advanced multimodal AI in text generation, image understanding, and video analysis.
A most Frontend Collection and survey of vision-language model papers, and models GitHub repository. Continuous updates.
A minimal codebase for finetuning large multimodal models, supporting llava-1.5/1.6, llava-interleave, llava-next-video, llava-onevision, llama-3.2-vision, qwen-vl, qwen2-vl, phi3-v etc.
Reinforcement Learning of Vision Language Models with Self Visual Perception Reward
Mark web pages for use with vision-language models
Local Video RAG Engine. A FastAPI microservice for video understanding: Scene Detection + Whisper ASR + Qwen3-VL. Optimized for Apple Silicon (MLX) & Windows/Linux (Llama.cpp).
给 DeepSeek 装上眼睛 — MCP Server + 通义千问VL, 剪贴板图片→视觉模型→文字描述 / Give DeepSeek the ability to see images via clipboard + Qwen-VL
An AI Agent that is able to control your screen to complste any task
Give non-multimodal Claude Code main models the ability to see pasted screenshots — a ~200-line UserPromptSubmit hook.
Self-evolving agentic reward framework for image-editing evaluation — 47.4% on EditReward-Bench from only 100 preference demos, no reward-model training. arXiv 2605.08703.
基于Qwen2.5-VL-3B + QLoRA的智能驾驶场景结构化理解系统
🎬 Extract AI prompts from video using Vision LLM (llama.cpp API) — Gradio WebUI + CLI
基于 Qwen3-VL-Flash 视觉语言模型的 Web 端自动化图像标注平台,支持 2D 物体检测、3D 空间定位(9-DOF)与视角遮挡分析,可导出 COCO / Pascal VOC 格式。
PriorTR (ECCV 2026): training-free, prior-corrected visual token reduction for accelerating multimodal LLMs — image & video.
A robotic sequential grasping system integrating YOLO detection and Qwen-VLM fine-tuning, enabling a full loop from manual teaching to LLM-based logical manipulation.
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
Run Qwen3-8B at ~18 tok/s on a Core Ultra laptop iGPU. Local-only LLM + VLM + ReAct agent stack with $0 token cost. Drop-in Claude Code backend.
Qwen-VL base model for use with Autodistill.
Add a description, image, and links to the qwen-vl topic page so that developers can more easily learn about it.
To associate your repository with the qwen-vl topic, visit your repo's landing page and select "manage topics."