This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Jacob is building a self-hosted local LLM inference stack across his Proxmox homelab — no cloud dependency. The stack targets NVIDIA consumer GPUs (8–16 GB VRAM per node) and serves coding assistance, general chat/reasoning, and an OpenAI-compatible API for apps and agents.
Use cases: Coding assistant (Python, Rust, Go, C#), general chat/reasoning, API serving for apps/agents, batch inference.
Current state: Infrastructure files are in place and tested locally on Windows (Podman + RTX 3080, ~107 tok/s confirmed). Not yet deployed to Proxmox.
Proxmox Host
└── Ubuntu Server VM (UEFI/q35, GPU PCI passthrough via vfio-pci)
└── Docker + NVIDIA Container Toolkit
├── Ollama :11434 ← Primary inference engine
└── Open WebUI :3000 ← Chat frontend
Multi-node expansion: Run docker-compose.yml on each node. On a coordinator node, set OLLAMA_NODES in .env (semicolon-separated endpoints) and run docker-compose.multi.yml. Open WebUI uses OLLAMA_BASE_URLS (plural) to merge model lists and randomly distribute requests across backends. No orchestration layer needed.
Why Ollama over TabbyAPI or vLLM:
- Widest model selection (GGUF, 135K+ models on HuggingFace)
- Works on any consumer GPU VRAM size — flexible for mixed hardware
- Multi-node expansion is additive: add endpoints to
OLLAMA_NODESand restart - No Python orchestration layer to maintain
- TabbyAPI is pre-1.0 and explicitly not production-grade
- vLLM's multi-node story (Ray) is heavier to operate and lacks GGUF support
infrastructure/
├── .env.example # Configuration template — copy to .env
├── docker-compose.yml # Single-node: Ollama + Open WebUI
├── docker-compose.multi.yml # Multi-node: Open WebUI routing to multiple Ollama nodes
└── scripts/
└── pull-models.sh # Pull recommended models by VRAM profile (8gb|12gb|16gb)
Ollama has no config file — all tuning is via environment variables in .env.
Key .env settings:
MODELS_PATH— leave empty for a named Docker volume, or set an absolute host path (e.g./mnt/models) for a dedicated diskOLLAMA_KEEP_ALIVE— how long a model stays in VRAM when idle (-1to never unload,5mfor shared use)OLLAMA_NODES— semicolon-separated Ollama endpoints for multi-node (used bydocker-compose.multi.ymlonly, maps to Open WebUI'sOLLAMA_BASE_URLS)
Common commands (use docker or podman interchangeably — pull-models.sh auto-detects):
cd infrastructure
cp .env.example .env
docker compose up -d # or: podman compose up -d
docker compose down
./scripts/pull-models.sh 8gb # or 12gb, 16gb
docker exec ollama ollama list
docker logs -f ollamaPodman on Windows: Requires one-time GPU setup before podman compose up -d will use the GPU. See docs/podman-gpu-windows.md.
| Role | 8 GB VRAM | 12 GB VRAM | 16 GB VRAM |
|---|---|---|---|
| Coding / autocomplete (FIM) | Qwen 2.5.1 Coder 7B @ Q5_K_M (~5.4 GB) | Qwen 2.5 Coder 14B @ Q5_K_M (~10.5 GB) | Qwen 2.5 Coder 14B @ Q6_K (~12.1 GB) |
| Daily driver / all-rounder | Qwen 3.5 9B @ Q4_K_M (~5.7 GB) | Qwen 3.5 9B @ Q6_K (~7.7 GB) | Qwen 3.5 9B @ Q8_0 (~9.6 GB) |
| Deep reasoning | DeepSeek R1 0528 Qwen3 8B @ Q5_K_M (~5.9 GB) | DeepSeek R1 0528 Qwen3 8B @ Q6_K (~6.7 GB) | DeepSeek R1 14B @ Q6_K (~12.1 GB) |
| Chat alternative | — | Gemma 3 12B @ Q5_K_M (~8.4 GB) | Phi-4 14B @ Q6_K (~12.0 GB) |
Qwen 3.5 9B is the recommended daily driver — multimodal, 262K context, toggleable thinking mode.
Quantization strategy: Each VRAM tier uses the highest quant that fits with room for KV cache. Ollama's library only ships Q4_K_M and Q8_0 — the pull script sources Q5_K_M and Q6_K from HuggingFace (bartowski, unsloth) via hf.co/ pulls. Bartowski uses imatrix quantization which preserves critical weight quality better than standard quants.
VRAM rule of thumb: VRAM ≈ (params_B × 0.56) + 1 GB overhead + KV cache. Keep context ≤ 8K tokens on 8 GB cards, ≤ 16K on 12 GB cards, ≤ 32K on 16 GB cards.
- PCI passthrough over LXC — full VM isolation via
vfio-pci(UEFI/q35 machine type,hostCPU). LXC is viable for GPU sharing between containers but not the primary target here. - Ollama as sole inference engine — no TabbyAPI or vLLM in the stack. Keep it simple.
- Multi-node via load balancer, not distributed inference — each node runs its own model independently. When a model needs to span multiple GPUs, use llama.cpp RPC (pipeline parallelism over Ethernet) or GPUStack.
- IDE integration via Continue extension — points at
http://<vm-ip>:11434. Qwen 2.5 Coder for FIM/autocomplete, Qwen 3.5 9B for chat.
architecture/plan.md— Full deployment guide (Proxmox setup, Docker config, model selection, IDE integration)architecture/high_level_context.md— 2026 inference engine landscape and benchmarksdocs/running.md— Step-by-step guide to starting the stack, pulling models, and using the prompt scriptdocs/podman-gpu-windows.md— GPU passthrough setup for Podman on Windows/WSL2 (one-time setup, CDI-based)