ElBruno.LocalLLMs supports 35 models across 6 tiers. This guide details each model, its capabilities, and how to use it.
| Tier | Model | Params | HuggingFace ID | ONNX Status | Chat Template | Tool Calling | Recommended RAM | Speed |
|---|---|---|---|---|---|---|---|---|
| ⚪ Tiny | TinyLlama-1.1B-Chat | 1.1B | TinyLlama/TinyLlama-1.1B-Chat-v1.0 | ✅ Native | ChatML | — | 2–4 GB | ⚡⚡⚡ |
| ⚪ Tiny | SmolLM2-1.7B-Instruct | 1.7B | HuggingFaceTB/SmolLM2-1.7B-Instruct | ✅ Native | ChatML | — | 2–4 GB | ⚡⚡⚡ |
| ⚪ Tiny | Qwen2.5-0.5B-Instruct | 0.5B | Qwen/Qwen2.5-0.5B-Instruct | ✅ Native | Qwen | ✅ | 1–2 GB | ⚡⚡⚡ |
| ⚪ Tiny | Qwen2.5-1.5B-Instruct | 1.5B | Qwen/Qwen2.5-1.5B-Instruct | ✅ Native | Qwen | ✅ | 2–4 GB | ⚡⚡⚡ |
| ⚪ Tiny | Gemma-2B-IT | 2B | elbruno/Gemma-2B-IT-onnx | ✅ Native | Gemma | — | 4 GB | ⚡⚡⚡ |
| ⚪ Tiny | Gemma-4-E2B-IT | 5.1B (2B active) | elbruno/Gemma-4-E2B-IT-onnx | ✅ Native | Gemma | ✅ | 4–6 GB | ⚡⚡⚡ |
| ⚪ Tiny | StableLM-2-1.6B-Chat | 1.6B | elbruno/StableLM-2-1.6B-Chat-onnx | ✅ Native | ChatML | — | 3–4 GB | ⚡⚡⚡ |
| 🟢 Small | Phi-3.5-mini-instruct | 3.8B | microsoft/Phi-3.5-mini-instruct-onnx | ✅ Native | Phi3 | ✅ | 6–8 GB | ⚡⚡ |
| 🟢 Small | Qwen2.5-3B-Instruct | 3B | Qwen/Qwen2.5-3B-Instruct | ✅ Native | Qwen | ✅ | 6–8 GB | ⚡⚡ |
| 🟢 Small | Llama-3.2-3B-Instruct | 3B | elbruno/Llama-3.2-3B-Instruct-onnx | ✅ Native | Llama3 | — | 6–8 GB | ⚡⚡ |
| 🟢 Small | Gemma-2-2B-IT | 2.6B | elbruno/Gemma-2-2B-IT-onnx | ✅ Native | Gemma | — | 6 GB | ⚡⚡ |
| 🟢 Small | Gemma-4-E4B-IT | 8B (4B active) | elbruno/Gemma-4-E4B-IT-onnx | ✅ Native | Gemma | ✅ | 8–10 GB | ⚡⚡ |
| 🟡 Medium | Qwen2.5-7B-Instruct | 7B | Qwen/Qwen2.5-7B-Instruct | ✅ Native | Qwen | ✅ | 8–12 GB | ⚡ |
| 🟡 Medium | Qwen2.5-Coder-7B-Instruct | 7B | elbruno/Qwen2.5-Coder-7B-Instruct-onnx | ✅ Native | Qwen | ✅ | 8–12 GB | ⚡ |
| 🟡 Medium | Llama-3.1-8B-Instruct | 8B | meta-llama/Llama-3.1-8B-Instruct | ✅ Native | Llama3 | — | 8–12 GB | ⚡ |
| 🟡 Medium | Mistral-7B-Instruct-v0.3 | 7B | mistralai/Mistral-7B-Instruct-v0.3 | ✅ Native | Mistral | — | 8–12 GB | ⚡ |
| 🟡 Medium | Gemma-2-9B-IT | 9B | elbruno/Gemma-2-9B-IT-onnx | ✅ Native | Gemma | — | 12 GB | ⚡ |
| 🟡 Medium | Gemma-4-12B-IT | 12B | elbruno/Gemma-4-12B-IT-onnx | ✅ Native | Gemma | ✅ | 12–16 GB | ⚡ |
| 🟡 Medium | Phi-4 | 14B | microsoft/phi-4 | ✅ Native | Phi3 | ✅ | 12–16 GB | ⚡ |
| 🟡 Medium | DeepSeek-R1-Distill-Qwen-14B | 14B | deepseek-ai/DeepSeek-R1-Distill-Qwen-14B | ✅ Native | ChatML | — | 12–16 GB | ⚡ |
| 🟡 Medium | Mistral-Small-24B-Instruct | 24B | mistralai/Mistral-Small-24B-Instruct-2501 | ✅ Native | Mistral | — | 16–20 GB | ⚡ |
| 🔴 Large | Qwen2.5-14B-Instruct | 14B | Qwen/Qwen2.5-14B-Instruct | ✅ Native | Qwen | ✅ | 16–24 GB | 🐢 |
| 🔴 Large | Qwen2.5-32B-Instruct | 32B | Qwen/Qwen2.5-32B-Instruct | ✅ Native | Qwen | ✅ | 24–32 GB | 🐢 |
| 🔴 Large | Llama-3.3-70B-Instruct | 70B | elbruno/Llama-3.3-70B-Instruct-onnx | ✅ Native | Llama3 | — | 40+ GB | 🐢 |
| 🔴 Large | Mixtral-8x7B-Instruct-v0.1 | 46.7B (MoE) | elbruno/Mixtral-8x7B-Instruct-v0.1-onnx | ✅ Native | Mistral | — | 24–32 GB | 🐢 |
| 🔴 Large | DeepSeek-R1-Distill-Llama-70B | 70B | elbruno/DeepSeek-R1-Distill-Llama-70B-onnx | ✅ Native | Llama3 | — | 40+ GB | 🐢 |
| 🔴 Large | Command-R (35B) | 35B | elbruno/Command-R-35B-onnx | ✅ Native | ChatML | — | 24–32 GB | 🐢 |
| 🔴 Large | Gemma-4-26B-A4B-IT | 25.2B (3.8B active) | elbruno/Gemma-4-26B-A4B-IT-onnx | ✅ Native | Gemma | ✅ | 20–28 GB | ⚡ |
| 🔴 Large | Gemma-4-31B-IT | 30.7B | elbruno/Gemma-4-31B-IT-onnx | ✅ Native | Gemma | ✅ | 24–32 GB | 🐢 |
| 🟣 Next-Gen | Llama-4-Scout | ~17B (MoE) | meta-llama/Llama-4-Scout-17B-16E-Instruct | 🔄 Convert | Llama3 | — | 24–32 GB | ⚡ |
| 🟣 Next-Gen | Llama-4-Maverick | ~17B (MoE) | meta-llama/Llama-4-Maverick-17B-128E-Instruct | 🔄 Convert | Llama3 | — | 64+ GB | 🐢 |
| 🟣 Next-Gen | Qwen3-8B | 8B | Qwen/Qwen3-8B | 🔄 Convert | Qwen | ✅ | 8–12 GB | ⚡ |
| 🟣 Next-Gen | Qwen3-32B | 32B | Qwen/Qwen3-32B | 🔄 Convert | Qwen | ✅ | 24–32 GB | 🐢 |
| 🟣 Next-Gen | Gemma-3-12B-IT | 12B | google/gemma-3-12b-it | 🔄 Convert | ChatML | — | 12–16 GB | ⚡ |
| 🟣 Next-Gen | DeepSeek-V3 | 671B (MoE) | deepseek-ai/DeepSeek-V3 | 🔄 Convert | ChatML | — | 128+ GB | 🐢 |
| 🟣 Next-Gen | Qwen3-14B-Instruct | 14.77B | onnx-community/Qwen3-14B-ONNX | ✅ Native | Qwen3 | ✅ | 16–24 GB | ⚡ |
These models are purpose-built for multi-agent orchestration, tool calling, and agentic loops. They follow the MagenticUI protocol with submit as a terminal tool signal.
| Model | Params | HuggingFace ID | ONNX Status | Chat Template | Tool Calling | Recommended RAM | Notes |
|---|---|---|---|---|---|---|---|
| MagenticBrain | ~14.77B | elbruno/MagenticBrain-onnx | ✅ Native | Qwen3 | ✅ | 16–24 GB | INT4, ~11 GB; set EnsureModelDownloaded = true |
| Fara1.5-9B | ~9.4B | elbruno/Fara1.5-9B-onnx | ✅ Native | Fara | — | 12–16 GB | Full multimodal package published and validated with LocalVisionChatClient |
Auto-download: MagenticBrain and Fara are both ready today. Fara uses the published multimodal package generated by
scripts/convert_fara_multimodal.py.MagenticBrain is a fine-tune of Qwen3-14B for agentic tasks. Converted from
microsoft/MagenticBrainusing ORT-GenAI built-in builder (INT4 CPU, ~11 GB).Fara1.5-9B is a computer-use agent fine-tuned from Qwen3.5-9B-VL. The built-in builder still maps
Qwen3_5ForConditionalGenerationtoQwen35TextModelwithexclude_embeds=true; the publishedelbruno/Fara1.5-9B-onnxpackage is completed via the custom multimodal export pipeline.Recommended sampling:
temperature=0.7, top_p=0.8, presence_penalty=1.0— greedy decoding causes infinite loops.Submit protocol: MagenticBrain signals task completion by calling
submit(a declared tool with no parameters). TheMagenticBrainAgentsample demonstrates the full OmniAgent loop.
See src/samples/MagenticBrainAgent/ for a ready-to-run demo.
BitNet models use 1.58-bit ternary weights {-1, 0, 1} and run via bitnet.cpp (a llama.cpp fork with custom kernels). These models require the ElBruno.LocalLLMs.BitNet NuGet package and a pre-built bitnet.cpp native library.
| Model | Params | HuggingFace ID | Format | Kernel | Chat Template | Approx. Size | RAM | Speed |
|---|---|---|---|---|---|---|---|---|
| BitNet b1.58 0.7B | 0.7B | 1bitLLM/bitnet_b1_58-large | GGUF | I2_S | ChatML | ~150 MB | 1–2 GB | ⚡⚡⚡ |
| Falcon3-1B-1.58bit | 1B | tiiuae/Falcon3-1B-Instruct-1.58bit | GGUF | TL2 | Falcon | ~200 MB | 2–3 GB | ⚡⚡⚡ |
| BitNet b1.58 2B-4T ⭐ | 2.4B | microsoft/BitNet-b1.58-2B-4T-gguf | GGUF | TL2 | Llama3 | ~400 MB | 3–4 GB | ⚡⚡ |
| BitNet b1.58 3B | 3B | 1bitLLM/bitnet_b1_58-3B | GGUF | I2_S | ChatML | ~650 MB | 4–6 GB | ⚡⚡ |
| Falcon3-3B-1.58bit | 3B | tiiuae/Falcon3-3B-Instruct-1.58bit | GGUF | TL2 | Falcon | ~600 MB | 4–6 GB | ⚡⚡ |
⭐ Default model. BitNet b1.58 2B-4T is the recommended starting point — MIT licensed, good quality, and tiny footprint.
Setup: BitNet models require a pre-built
bitnet.cppnative library. See the BitNet Guide for build instructions and configuration.
Fine-tuned variants of Qwen2.5-0.5B, optimized for specific tasks with ElBruno.LocalLLMs' chat template format. These models are ready-to-use ONNX INT4 and download automatically from HuggingFace.
| Model | Params | HuggingFace ID | Task | Chat Template | Tool Calling | RAM | Speed |
|---|---|---|---|---|---|---|---|
| Qwen2.5-0.5B-LocalLLMs-ToolCalling | 0.5B | elbruno/Qwen2.5-0.5B-LocalLLMs-ToolCalling | Tool/function calling | Qwen | ✅ | ~1 GB | ⚡⚡⚡ |
| Qwen2.5-0.5B-LocalLLMs-RAG | 0.5B | elbruno/Qwen2.5-0.5B-LocalLLMs-RAG | RAG with citations | Qwen | — | ~1 GB | ⚡⚡⚡ |
| Qwen2.5-0.5B-LocalLLMs-Instruct | 0.5B | elbruno/Qwen2.5-0.5B-LocalLLMs-Instruct | General (tools + RAG) | Qwen | ✅ | ~1 GB | ⚡⚡⚡ |
Tip: A fine-tuned 0.5B model often matches or exceeds a base 1.5B model on its specialized task. See the Fine-Tuning Guide for details.
- ✅ Native ONNX — ONNX weights are published on HuggingFace. Download and use immediately, no conversion needed.
- 🔄 Convert — Only PyTorch weights are available. Requires ONNX conversion using the Python scripts in
/scripts/. See Conversion below.
Best for:
- Edge devices and IoT hardware
- Fast prototyping and testing
- Learning the library
- Real-time applications with strict latency budgets (<100ms)
- Limited-memory environments (Raspberry Pi, old laptops)
Trade-offs:
- ✅ Super fast (100–500ms per response)
- ✅ Tiny memory footprint (1–4 GB)
- ❌ Limited reasoning ability
- ❌ Poor long-context understanding
- ❌ Lower code quality in code tasks
Examples:
var options = new LocalLLMsOptions
{
Model = KnownModels.Qwen25_05BInstruct // 0.5B — fastest
};Realistic outputs:
- Simple Q&A: ✅ Excellent
- Creative writing:
⚠️ Basic - Code generation: ❌ Poor
- Math reasoning: ❌ Poor
Best for:
- Your first local LLM project — start here
- Most chatbot and Q&A applications
- Content generation (summaries, emails, social posts)
- Edge deployment with decent quality
- Prototyping before scaling up
Trade-offs:
- ✅ Fast (2–5 seconds per response)
- ✅ Reasonable memory (6–8 GB)
- ✅ Best quality-to-size ratio
⚠️ Moderate reasoning ability⚠️ Limited multi-step logic
Examples:
// Recommended starting model
var client = await LocalChatClient.CreateAsync();
// Or explicitly:
var options = new LocalLLMsOptions
{
Model = KnownModels.Phi35MiniInstruct // 3.8B — best default
};Realistic outputs:
- Simple Q&A: ✅✅ Excellent
- Creative writing: ✅ Very good
- Code generation: ✅ Good (simple functions)
- Math reasoning:
⚠️ Basic - Multi-turn conversations: ✅ Good
Best for:
- Production-grade local LLM deployments
- Complex reasoning tasks
- Code generation and explanation
- Advanced content creation
- Systems with 12+ GB RAM/VRAM
Trade-offs:
- ✅ Excellent quality (comparable to GPT-3.5)
- ✅ Strong reasoning and coding ability
⚠️ Slower (5–15 seconds per response on CPU)⚠️ Needs 12–20 GB memory⚠️ Slower first token (KV cache is larger)
Popular choices:
Phi-4(14B) — best reasoning, Microsoft-published, native ONNXQwen2.5-7B-Instruct— excellent instruction-followingDeepSeek-R1-Distill-Qwen-14B— exceptional at reasoning/math
Example:
var options = new LocalLLMsOptions
{
Model = KnownModels.Phi4, // 14B — production-grade
ExecutionProvider = ExecutionProvider.Cuda // Use GPU
};
using var client = await LocalChatClient.CreateAsync(options);Realistic outputs:
- Simple Q&A: ✅✅ Near-perfect
- Creative writing: ✅✅ Excellent
- Code generation: ✅✅ Excellent (complex functions, patterns)
- Math reasoning: ✅ Very good
- Multi-step logic: ✅✅ Excellent
Best for:
- Multi-GPU systems or very high-end GPUs (RTX 4090, H100)
- Heavy research and advanced reasoning
- Production systems with massive context windows
- Organizations with dedicated ML infrastructure
Trade-offs:
- ✅ State-of-the-art quality
- ✅ Exceptional reasoning, coding, and analysis
- ❌ Requires 40+ GB memory or multi-GPU
- ❌ Very slow on CPU (minutes per response)
- ❌ High power consumption
Note on MoE models:
- Mixtral-8x7B (46.7B params but 2× speedup) — uses Mixture of Experts; only 2 of 8 experts active per token
- DeepSeek-R1-Distill-Llama-70B — exceptional reasoning but slower
Realistic outputs:
- Research-grade writing: ✅✅✅
- Complex code: ✅✅✅
- Advanced math/physics: ✅✅
- Reasoning chains: ✅✅✅
Best for:
- Cutting-edge capabilities
- Future-proofing your application
- Research and experimentation
- Trying the latest architectures (Llama 4, Qwen 3, DeepSeek-V3)
Status:
- Recently released by Meta, Qwen (Alibaba), DeepSeek
- May require architecture-specific ONNX conversion work
- Not all are well-tested in the ElBruno.LocalLLMs ecosystem yet
⚠️ Some frontier models are not viable for local inference. Very large Mixture-of-Experts and/or multimodal models — e.g. DeepSeek-V3 (671B MoE) and Inkling (975B MoE, multimodal text/image/audio) — cannot be converted to ONNX or run locally with this library (MoE routing is unsupported, and weights run to hundreds of GB even at INT4). See Blocked Models for the full analysis and recommended alternatives.
Example:
// Create custom ModelDefinition for Qwen3-8B
var qwen3 = new ModelDefinition
{
Id = "qwen3-8b",
DisplayName = "Qwen3-8B",
HuggingFaceRepoId = "Qwen/Qwen3-8B",
RequiredFiles = ["model.onnx"], // After conversion
ModelType = OnnxModelType.GenAI,
ChatTemplate = ChatTemplateFormat.Qwen,
Tier = ModelTier.Medium
};
var options = new LocalLLMsOptions { Model = qwen3 };Each model family uses a different prompt format. The library handles this automatically:
| Format | Models | Example | Notes |
|---|---|---|---|
| ChatML | Qwen, Mistral (old) | <|im_start|>user\nQuestion<|im_end|> |
Standard multi-turn format |
| Gemma | Gemma, Gemma 2, Gemma 4 | <start_of_turn>user\nQuestion<end_of_turn> |
Google's format |
| Phi3 | Phi-3, Phi-3.5, Phi-4 | <|user|>\nQuestion<|end|> |
Microsoft's format |
| Llama3 | Llama-3.x, Llama-4 | <|start_header_id|>user<|end_header_id|> |
Meta's modern format |
| Qwen | Qwen series | <|im_start|>user\nQuestion<|im_end|> |
Alibaba's format |
| Mistral | Mistral-7B+ | [INST] Question [/INST] |
Mistral's format |
You don't need to worry about these — the library applies the correct format automatically when you pass messages through GetResponseAsync().
Tool calling enables models to call functions/tools you define, enabling agent-like behavior. Not all models support this — it depends on the model's training data and chat template.
| Priority | Model | Why |
|---|---|---|
| 🥇 Best Starting Point | Phi-3.5-mini-instruct (3.8B) | Native ONNX, no conversion needed, good quality |
| 🥈 Smallest Option | Qwen2.5-0.5B-Instruct (0.5B) | Only ~1-2 GB RAM, great for testing/demos |
| 🥉 Best Quality/Size | Qwen2.5-7B-Instruct (7B) | Best tool calling accuracy among supported models |
| 🏆 Best Overall | Phi-4 (14B) | Highest accuracy, needs more RAM (12-16 GB) |
| 🎯 Best Small + Reliable | Qwen2.5-0.5B-LocalLLMs-ToolCalling (0.5B) | Fine-tuned for tool calling — cleaner JSON than base 0.5B |
Note: Tool calling quality scales with model size. Smaller models may hallucinate tool calls or miss them. For production use, prefer 3B+ models or a fine-tuned variant.
Tool calling in local models is prompt-based — tools are described as JSON schemas in the system prompt, and the model responds with JSON tool call objects. The library handles all formatting and parsing automatically through IChatClient.
See Getting Started for usage examples.
These models have ONNX weights pre-published on HuggingFace. Just use them:
// No conversion needed — these are ready
var client1 = await LocalChatClient.CreateAsync(new LocalLLMsOptions
{
Model = KnownModels.Phi35MiniInstruct // ✅ Native ONNX
});
var client2 = await LocalChatClient.CreateAsync(new LocalLLMsOptions
{
Model = KnownModels.Phi4 // ✅ Native ONNX
});All other models require conversion. See below.
Prerequisites:
- Python 3.10+ installed
- Git installed
- ~50 GB free disk space (for large models)
Steps:
-
Navigate to the scripts directory:
cd scripts/ -
Install Python dependencies:
pip install -r requirements.txt
This installs:
transformers,optimum,onnx,onnxruntime, etc. -
Run the conversion script:
# Example: Convert Qwen2.5-7B to ONNX python convert_to_onnx.py \ --model-id Qwen/Qwen2.5-7B-Instruct \ --output-dir ./onnx-models/qwen2.5-7b -
Wait for completion:
- Small models (3B): ~10–15 minutes on modern CPU
- Large models (70B): ~1–2 hours
- Output will be in
./onnx-models/qwen2.5-7b/
-
Use in your app:
var options = new LocalLLMsOptions { ModelPath = @"./onnx-models/qwen2.5-7b" }; using var client = await LocalChatClient.CreateAsync(options);
Detailed conversion guide: See scripts/README.md
To use a model not in KnownModels, create a custom ModelDefinition:
using ElBruno.LocalLLMs;
using Microsoft.Extensions.AI;
// Define a custom model
var customModel = new ModelDefinition
{
Id = "custom-qwen-7b",
DisplayName = "Custom Qwen2.5-7B",
HuggingFaceRepoId = "Qwen/Qwen2.5-7B-Instruct",
RequiredFiles = ["onnx/model.onnx", "onnx/model.onnx_data"],
ModelType = OnnxModelType.GenAI,
ChatTemplate = ChatTemplateFormat.Qwen,
Tier = ModelTier.Medium,
HasNativeOnnx = false // You'll need to convert it
};
var options = new LocalLLMsOptions { Model = customModel };
using var client = await LocalChatClient.CreateAsync(options);Adding to KnownModels permanently:
- Edit
src/ElBruno.LocalLLMs/Models/KnownModels.cs - Add a new static
readonlyfield - Add it to the
Allcollection - Submit a PR!
Here's how models compare in real-world scenarios (on NVIDIA RTX 4080, 8GB VRAM):
| Model | Size | CPU (tokens/sec) | GPU (tokens/sec) | Memory | Quality |
|---|---|---|---|---|---|
| Qwen2.5-0.5B | 0.5B | 120 | 450 | 2 GB | ⭐ |
| Phi-3.5-mini | 3.8B | 8 | 180 | 6 GB | ⭐⭐⭐ |
| Qwen2.5-7B | 7B | 2 | 85 | 10 GB | ⭐⭐⭐⭐ |
| Phi-4 | 14B | 0.5 | 40 | 14 GB | ⭐⭐⭐⭐ |
| Llama-3.1-70B | 70B | <0.1 | 8 | 45 GB | ⭐⭐⭐⭐⭐ |
Key observations:
- GPU provides 5–50x speedup depending on model size
- Small models are surprisingly capable for most tasks
- Token generation speed ≈ 1–2 words per second on consumer GPU
- First token is slowest (KV cache initialization)
START: Choosing a model?
│
├─ "I just want to try this library"
│ └─> Use Phi-3.5-mini-instruct (default, native ONNX, solid quality)
│
├─ "I have <6 GB RAM"
│ └─> Use Qwen2.5-0.5B-Instruct (Tiny, needs conversion)
│
├─ "I need fast responses and decent quality"
│ └─> Use Phi-3.5-mini-instruct (Small, native ONNX, 3–5 sec on CPU)
│
├─ "I want production-grade quality"
│ ├─ "I have GPU"
│ │ └─> Use Phi-4 or Qwen2.5-7B (Medium, native or 🔄 convert)
│ └─ "CPU only"
│ └─> Use Phi-3.5-mini-instruct or Qwen2.5-3B
│
├─ "I need advanced reasoning (math, code, logic)"
│ ├─ "Speed matters"
│ │ └─> Use Phi-4 (14B, native ONNX)
│ └─ "Quality over speed"
│ └─> Use Qwen2.5-7B or DeepSeek-R1-Distill (need conversion)
│
├─ "I need a code assistant / local Copilot replacement"
│ ├─ "ONNX Runtime (this library)"
│ │ └─> Use Qwen2.5-Coder-7B-Instruct (Medium, needs conversion)
│ └─ "Any runtime (llama.cpp, vLLM)"
│ └─> Use Devstral-Small-2 (24B, Apache 2.0, GGUF format)
│
├─ "I have multi-GPU / high-end hardware"
│ └─> Use Large models (70B+) for max quality
│
└─ "I want the latest cutting-edge models"
└─> Use Next-Gen (Llama-4, Qwen3, DeepSeek-V3) — needs ONNX conversion
| Use Case | Model | ExecutionProvider | RAM | Notes |
|---|---|---|---|---|
| Learning / Testing | Phi-3.5-mini | CPU | 6 GB | Default, native ONNX, instant download |
| Simple Chatbot | Qwen2.5-3B | CPU | 8 GB | Great instruction-following, needs conversion |
| Production API | Phi-4 | CUDA | 14 GB | Best reasoning, native ONNX, fast on GPU |
| RAG Pipeline | Phi-3.5-mini | CUDA | 6 GB | Native ONNX, excellent quality-to-size ratio, proven for RAG |
| Real-time App | Qwen2.5-0.5B | CUDA | 2 GB | Tiny, ultra-fast, weak quality |
| Content Gen | Qwen2.5-7B | CUDA | 12 GB | Excellent writing, powerful instruction-follow |
| Code Assistant | Qwen2.5-Coder-7B | CUDA | 10 GB | Code-specialized, Qwen architecture, needs conversion |
| Edge (RPi, IoT) | Qwen2.5-0.5B | CPU | 2 GB | Minimal, but usable for simple tasks |
| Advanced Reasoning | Llama-3.1-70B | CUDA | 45 GB | State-of-the-art, multi-GPU |
- 📚 Getting Started Guide — detailed setup and examples
- 🏗️ Architecture — internal design
- 📝 Contributing — add a new model
- 🐍 ONNX Conversion Scripts — convert models manually
- 🤗 HuggingFace Hub — browse all models
Happy experimenting! 🚀