Skip to content

Repository files navigation

ElBruno.LocalLLMs

NuGet NuGet Downloads Build Status License: MIT HuggingFace .NET GitHub stars Twitter Follow

Run local LLMs in .NET through IChatClient 🧠

Run local LLMs in .NET through IChatClient — the same interface you'd use for Azure OpenAI, Ollama, or any other provider. Powered by ONNX Runtime GenAI and BitNet.

What's New

  • ⬆️ v0.20.9 — Upgraded onnxruntime-genai to 0.15.1 and Microsoft.Extensions.AI.Abstractions to 10.8.3 across all projects. No API changes.
  • 🧩 ElBruno.LocalLLMs.BlazorComponents — new Razor Class Library with 7 ready-to-use Blazor components: ModelStatusCard (download progress bar + actions), ModelGallery (filterable grid), ModelSelector (two-way-bindable dropdown), ChatBox (streaming token display), EnvironmentDashboard (CPU/CUDA/DirectML badges), LocalLLMHealthBadge (nav-bar status dot), and RagPlayground. Call services.AddLocalLLMsBlazorComponents() to register. See the Blazor Components Guide and the BlazorDemo sample.
  • 📦 9 more models now support auto-download (v0.20.4) — StableLM-2-1.6B-Chat, Gemma-4-E2B-IT, Gemma-4-E4B-IT, Gemma-4-12B-IT, Gemma-4-26B-A4B-IT, Gemma-4-31B-IT, Mixtral-8x7B-Instruct-v0.1, DeepSeek-R1-Distill-Llama-70B, and Command-R (35B) are now HasNativeOnnx=true with ONNX weights hosted at elbruno/*-onnx on HuggingFace. Set EnsureModelDownloaded = true to auto-download.
  • 🔍 ModelDefinition.IsVisionCapable — new computed property. Consumer apps can now check model.IsVisionCapable instead of comparing ModelType == OnnxModelType.VisionGenAI. Fara1.5-9B returns true; all text models return false.
  • 🛑 Fail-fast for non-downloadable modelsOptionsValidator now throws an InvalidOperationException with actionable text when EnsureModelDownloaded = true is paired with a model that has HasNativeOnnx = false and no ModelPath. Eliminates the confusing "auto-download enabled" UX that previously gave no guidance.
  • 📄 Auto-download guide — new doc covering the auto-download flow, cache path, vision model usage, and first-run examples for MagenticBrain and Fara.
  • 📋 ListCachedModels() + GetModelCacheSize(model) — new cache inspection APIs on both LocalChatClient and LocalVisionChatClient. List all cached model directories with sizes, or get the byte count for a specific model. Delegates to ElBruno.HuggingFace.Downloader 1.4.4.
  • 🗑️ DeleteModelFromCacheAsync — now delegates to HuggingFaceDownloader.DeleteCachedFilesAsync (was a direct Directory.Delete). Available on both LocalChatClient and LocalVisionChatClient.
  • 🤖 MagenticBrain + Fara native ONNX ready — Both published ONNX repos now auto-download via EnsureModelDownloaded = true. Fara uses the validated multimodal package produced by scripts/convert_fara_multimodal.py.
  • 👁️ LocalVisionChatClient auto-download — Vision models with HasNativeOnnx=true now download automatically, just like text models.
  • 🧪 E2E lifecycle tests for all 35 models — 3-phase lifecycle (download → cache hit → delete) with automated markdown reports written to docs/tests/ after each run. Phases 1 and 3 also assert cache size and list membership.
  • 📦 Transitive native runtime fix (v0.20.1) — buildTransitive packaging so onnxruntime-genai.dll is copied for downstream consumers (Issue #24).
  • 🤖 Qwen3 & MagenticBrain support (v0.20.0) — New ChatTemplateFormat.Qwen3 formatter and KnownModels.Qwen3_14BInstruct for agentic multi-agent orchestration loops.
  • 👁️ Fara 1.5-9B vision-language model (v0.20.0) — Run Microsoft's Fara VLM locally via the new LocalVisionChatClient, IVisionGenerationModel, and VisionChatOptions with image paths.
  • 🌐 ElBruno.MagenticUI reference app — Full Blazor Server multi-agent app (FileSurfer, WebFetcher, Coder, UserProxy) powered by this library.
  • 📡 OpenTelemetry diagnostics (v0.19.0) — Generation lifecycle activities and metrics (gen_ai.client.*) via ActivitySource + Meter both named ElBruno.LocalLLMs.
  • Gemma 4 family activeE2B, E4B, 12B Unified, 26B-A4B, 31B all supported.
  • ⬆️ ONNX Runtime GenAI 0.14.1 — upgraded across library, tests, samples, and benchmarks.

Features

  • 🧩 Blazor componentsModelStatusCard, ChatBox, ModelGallery, ModelSelector, EnvironmentDashboard, LocalLLMHealthBadge, RagPlayground via ElBruno.LocalLLMs.BlazorComponents (guide)
  • 🔌 IChatClient implementation — seamless integration with Microsoft.Extensions.AI
  • 📦 Automatic model download — models are fetched from HuggingFace on first use
  • 🚀 Zero friction — works out of the box with sensible defaults (Phi-3.5 mini)
  • 🖥️ Multi-hardware — CPU, CUDA, and DirectML execution providers
  • 💉 DI-friendly — register with AddLocalLLMs() or AddBitNetChatClient() in ASP.NET Core
  • 🔄 Streaming — token-by-token streaming via GetStreamingResponseAsync
  • 📊 Multi-model — switch between Phi-3.5, Phi-4, Qwen2.5, Qwen3, Llama 3.2, MagenticBrain, and more
  • 👁️ Vision-language models — run Fara 1.5-9B image+text models via LocalVisionChatClient
  • 🤖 Agentic models — Qwen3 / MagenticBrain support for multi-agent orchestration loops
  • 🎯 Fine-tuned models — pre-trained Qwen2.5 variants for tool calling and RAG (guide)
  • BitNet support — run 1.58-bit ternary models via bitnet.cpp with extreme efficiency (guide)
  • 📈 OpenTelemetry diagnostics — lifecycle activities and metrics for queued, first-token, completion, cancellation, and failure (guide)

Packages

Package NuGet Downloads Description
ElBruno.LocalLLMs NuGet Downloads Core library — ONNX Runtime GenAI models via IChatClient
ElBruno.LocalLLMs.Rag NuGet Downloads RAG pipeline — document chunking, indexing, retrieval
ElBruno.LocalLLMs.BitNet NuGet Downloads BitNet 1.58-bit models via bitnet.cpp + IChatClient
ElBruno.LocalLLMs.BlazorComponents NuGet Downloads Blazor components — ModelStatusCard, ChatBox, ModelGallery, and more

Installation

dotnet add package ElBruno.LocalLLMs

For CPU scenarios, no extra package is required — the transitive buildTransitive shim copies onnxruntime-genai.dll automatically on Windows.

Add a runtime package only when you want a specific GPU provider:

# 🟢 NVIDIA GPU (CUDA):
dotnet add package Microsoft.ML.OnnxRuntimeGenAI.Cuda

# 🔵 Any Windows GPU — AMD, Intel, NVIDIA (DirectML):
dotnet add package Microsoft.ML.OnnxRuntimeGenAI.DirectML

⚠️ Add at most one GPU runtime package. Do not reference both Microsoft.ML.OnnxRuntimeGenAI.Cuda and Microsoft.ML.OnnxRuntimeGenAI.DirectML simultaneously.

If you use a GPU runtime package and want to disable the transitive CPU copy shim, set: <ElBrunoLocalLLMsDisableCpuNativeCopy>true</ElBrunoLocalLLMsDisableCpuNativeCopy> in your application .csproj.

🚀 The library defaults to ExecutionProvider.Auto — it tries GPU first and falls back to CPU automatically. No code changes needed.

Quick Start

using ElBruno.LocalLLMs;
using Microsoft.Extensions.AI;

// Create a local chat client (downloads Phi-3.5 mini on first run)
using var client = await LocalChatClient.CreateAsync();

var response = await client.GetResponseAsync([
    new(ChatRole.User, "What is the capital of France?")
]);

Console.WriteLine(response.Text);

First Run

The first time you create a LocalChatClient, the model is downloaded from HuggingFace to your local cache directory (~2-4 GB). This typically takes 30-60 seconds depending on your internet connection.

Track download progress:

using var client = await LocalChatClient.CreateAsync(
    new LocalLLMsOptions { Model = KnownModels.Phi35MiniInstruct },
    progress: new Progress<ModelDownloadProgress>(p =>
    {
        var percent = (p.BytesDownloaded * 100) / p.TotalBytes;
        Console.WriteLine($"{p.FileName}: {percent:F1}%");
    })
);

Subsequent runs load instantly from cache (%LOCALAPPDATA%/ElBruno/LocalLLMs/models).

Skip auto-download if using a pre-downloaded model:

var options = new LocalLLMsOptions
{
    Model = KnownModels.Phi35MiniInstruct,
    ModelPath = "/path/to/local/model",
    EnsureModelDownloaded = false
};
using var client = await LocalChatClient.CreateAsync(options);

Streaming

using ElBruno.LocalLLMs;
using Microsoft.Extensions.AI;

using var client = await LocalChatClient.CreateAsync(new LocalLLMsOptions
{
    Model = KnownModels.Phi35MiniInstruct
});

await foreach (var update in client.GetStreamingResponseAsync([
    new(ChatRole.System, "You are a helpful assistant."),
    new(ChatRole.User, "Explain quantum computing in simple terms.")
]))
{
    Console.Write(update.Text);
}

GPU Acceleration

By default, ExecutionProvider.Auto tries GPU first (CUDA → DirectML) and falls back to CPU automatically:

// Use explicit GPU provider (fails if CUDA not installed; use Auto to fallback to CPU)
var options = new LocalLLMsOptions
{
    ExecutionProvider = ExecutionProvider.Cuda
};

// Multi-GPU systems: select device ID
var options2 = new LocalLLMsOptions
{
    ExecutionProvider = ExecutionProvider.Cuda,
    GpuDeviceId = 1  // Use second GPU
};

Auto fallback behavior:

  • CUDA available → uses NVIDIA GPU
  • CUDA unavailable, DirectML available → uses AMD/Intel Arc GPU
  • GPU unavailable → falls back to CPU (no errors, just slower)

See Troubleshooting: GPU Setup for debugging GPU issues.

Model Metadata

Inspect model capabilities at runtime — context window size, model name, and vocabulary:

using var client = await LocalChatClient.CreateAsync();

var metadata = client.ModelInfo;
Console.WriteLine($"Model:          {metadata?.ModelName}");
Console.WriteLine($"Context window: {metadata?.MaxSequenceLength}");
Console.WriteLine($"Vocab size:     {metadata?.VocabSize}");

This is useful for prompt-length validation, adaptive chunking, and model selection logic.

Dependency Injection

builder.Services.AddLocalLLMs(options =>
{
    options.Model = KnownModels.Phi35MiniInstruct;
    options.ExecutionProvider = ExecutionProvider.DirectML;
});

// Inject IChatClient anywhere
public class MyService(IChatClient chatClient) { ... }

Error Handling

The library provides structured exception types for graceful error handling:

using ElBruno.LocalLLMs;
using Microsoft.Extensions.AI;

try
{
    using var client = await LocalChatClient.CreateAsync();
    var response = await client.GetResponseAsync([
        new(ChatRole.User, "Your question here")
    ]);
}
catch (ExecutionProviderException ex)
{
    // GPU/provider-specific error (no CUDA, DirectML not available, etc.)
    Console.WriteLine($"Provider error: {ex.Message}");
}
catch (ModelCapacityExceededException ex)
{
    // Prompt/response too long for model's context window
    Console.WriteLine($"Capacity error: {ex.Message}");
    // Solution: use a larger model or truncate the prompt
}
catch (InvalidOperationException ex)
{
    // General operation error (model not found, download failed, etc.)
    Console.WriteLine($"Operation error: {ex.Message}");
}

Observability

LocalChatClient emits generation lifecycle diagnostics through ActivitySource and Meter, both named ElBruno.LocalLLMs.

using ElBruno.LocalLLMs.Diagnostics;

builder.Services.AddOpenTelemetry()
    .WithTracing(tracing => tracing.AddSource(LocalLLMsInstrumentation.ActivitySourceName))
    .WithMetrics(metrics => metrics.AddMeter(LocalLLMsInstrumentation.MeterName));

By default, telemetry excludes prompt and completion text. Opt in only when you want content attached:

var options = new LocalLLMsOptions
{
    CaptureTelemetryContent = true
};

See docs/observability.md for the lifecycle event contract, metric names, and Aspire wiring notes, and docs/cancellation.md for voice barge-in cancellation behavior.

Cache Management

Inspect and manage the local model cache programmatically:

// Remove a model from the cache (no-op if not cached)
await LocalChatClient.DeleteModelFromCacheAsync(KnownModels.Phi35MiniInstruct);

// Or use a custom cache directory
await LocalChatClient.DeleteModelFromCacheAsync(
    KnownModels.Phi35MiniInstruct,
    cacheDirectory: @"D:\my-models");

// Get cached size in bytes for one model (0 if not downloaded)
long bytes = LocalChatClient.GetModelCacheSize(KnownModels.Phi35MiniInstruct);
Console.WriteLine($"Cached: {bytes / 1024 / 1024:N0} MB");

// List all cached models with size and last-modified date
var cached = LocalChatClient.ListCachedModels();
foreach (var repo in cached)
    Console.WriteLine($"{repo.LocalDirectory}  {repo.TotalSizeBytes / 1024 / 1024:N0} MB  {repo.LastModified:yyyy-MM-dd}");

// Same APIs available on LocalVisionChatClient for vision models
await LocalVisionChatClient.DeleteModelFromCacheAsync(KnownModels.Fara15_9B);
long visionBytes = LocalVisionChatClient.GetModelCacheSize(KnownModels.Fara15_9B);

The default cache directory is %LOCALAPPDATA%/ElBruno/LocalLLMs/models (Windows) or ~/.local/share/ElBruno/LocalLLMs/models (Linux/macOS).

These operations delegate to ElBruno.HuggingFace.Downloader which provides the underlying DeleteCachedFilesAsync, GetCachedSize, and ListCachedRepos implementation.

Troubleshooting

GPU not working? Use ExecutionProvider.Cpu explicitly. See GPU Setup Validation.

Out of memory? Try a smaller model:

var options = new LocalLLMsOptions
{
    Model = KnownModels.Qwen25_05BInstruct  // 0.5B instead of 3.8B
};

Model download fails?

  • Check your internet connection
  • For private HuggingFace models, set the HF_TOKEN environment variable

For detailed troubleshooting, see docs/troubleshooting-guide.md.

Supported Models

Tier Model Parameters ONNX ID
⚪ Tiny TinyLlama-1.1B-Chat 1.1B ✅ Native tinyllama-1.1b-chat
⚪ Tiny SmolLM2-1.7B-Instruct 1.7B ✅ Native smollm2-1.7b-instruct
⚪ Tiny Qwen2.5-0.5B-Instruct 0.5B ✅ Native qwen2.5-0.5b-instruct
⚪ Tiny Qwen2.5-1.5B-Instruct 1.5B ✅ Native qwen2.5-1.5b-instruct
⚪ Tiny Gemma-2B-IT 2B ✅ Native gemma-2b-it
⚪ Tiny Gemma-4-E2B-IT 5.1B (2B active) ✅ Native gemma-4-e2b-it
⚪ Tiny StableLM-2-1.6B-Chat 1.6B ✅ Native stablelm-2-1.6b-chat
🟢 Small Phi-3.5 mini instruct 3.8B ✅ Native phi-3.5-mini-instruct
🟢 Small Qwen2.5-3B-Instruct 3B ✅ Native qwen2.5-3b-instruct
🟢 Small Llama-3.2-3B-Instruct 3B ✅ Native llama-3.2-3b-instruct
🟢 Small Gemma-2-2B-IT 2B ✅ Native gemma-2-2b-it
🟢 Small Gemma-4-E4B-IT 8B (4B active) ✅ Native gemma-4-e4b-it
🟡 Medium Qwen2.5-7B-Instruct 7B ✅ Native qwen2.5-7b-instruct
🟡 Medium Qwen2.5-Coder-7B-Instruct 7B ✅ Native qwen2.5-coder-7b-instruct
🟡 Medium Llama-3.1-8B-Instruct 8B ✅ Native llama-3.1-8b-instruct
🟡 Medium Mistral-7B-Instruct-v0.3 7B ✅ Native mistral-7b-instruct-v0.3
🟡 Medium Gemma-2-9B-IT 9B ✅ Native gemma-2-9b-it
🟡 Medium Gemma-4-12B-IT 12B ✅ Native gemma-4-12b-it
🟡 Medium Phi-4 14B ✅ Native phi-4
🟡 Medium DeepSeek-R1-Distill-Qwen-14B 14B ✅ Native deepseek-r1-distill-qwen-14b
🟡 Medium Mistral-Small-24B-Instruct 24B ✅ Native mistral-small-24b-instruct
🔴 Large Qwen2.5-14B-Instruct 14B ✅ Native qwen2.5-14b-instruct
🔴 Large Qwen2.5-32B-Instruct 32B ✅ Native qwen2.5-32b-instruct
🔴 Large Llama-3.3-70B-Instruct 70B ✅ ONNX llama-3.3-70b-instruct
🔴 Large Mixtral-8x7B-Instruct-v0.1 8x7B ✅ Native mixtral-8x7b-instruct-v0.1
🔴 Large DeepSeek-R1-Distill-Llama-70B 70B ✅ Native deepseek-r1-distill-llama-70b
🔴 Large Command-R (35B) 35B ✅ Native command-r-35b
🔴 Large Gemma-4-26B-A4B-IT 25.2B (3.8B active) ✅ Native gemma-4-26b-a4b-it
🔴 Large Gemma-4-31B-IT 30.7B ✅ Native gemma-4-31b-it
🟣 Next-Gen Qwen3-14B-Instruct 14.77B ✅ Native qwen3-14b-instruct
🤖 Agentic MagenticBrain ~14.77B ✅ Native magentic-brain
👁️ VLM Fara 1.5-9B ~9.4B ✅ Native fara-1.5-9b

🔄 Convert = Use the conversion scripts in scripts/ to export ONNX locally before running the model.

¹ MagenticBrain ONNX: Native ONNX hosted at elbruno/MagenticBrain-onnx (INT4 quantized). Auto-downloads when EnsureModelDownloaded=true.

² Fara 1.5-9B ONNX: elbruno/Fara1.5-9B-onnx now includes the validated multimodal package (qwen3vl-vision.onnx, qwen3vl-embedding.onnx, patched genai_config.json, and ORT-compatible processor_config.json). See ONNX Conversion — Fara.

Fine-Tuned Models

Pre-trained variants optimized for specific tasks. A fine-tuned 0.5B model often matches or exceeds a base 1.5B on its specialized task.

Model Size Task HuggingFace ID
Qwen2.5-0.5B-ToolCalling ~1 GB Tool/function calling elbruno/Qwen2.5-0.5B-LocalLLMs-ToolCalling
Qwen2.5-0.5B-RAG ~1 GB RAG with citations elbruno/Qwen2.5-0.5B-LocalLLMs-RAG
Qwen2.5-0.5B-Instruct ~1 GB General-purpose elbruno/Qwen2.5-0.5B-LocalLLMs-Instruct

See the Supported Models Guide for detailed model cards, performance benchmarks, and selection guidance.

Samples

Sample Description
HelloChat Minimal console chat
StreamingChat Token-by-token streaming
MultiModelChat Switch models at runtime
DependencyInjection ASP.NET Core DI registration
ToolCallingAgent Function calling and tool use
FineTunedToolCalling Fine-tuned model for improved tool calling
RagChatbot RAG pipeline with document retrieval
ZeroCloudRag Zero-cloud RAG pipeline with real local embeddings and LLM inference
BitNetChat BitNet 1.58-bit model chat completion
BitNetPerformance Performance benchmark: BitNet vs ONNX models
MagenticBrainAgent Multi-agent orchestration loop using Qwen3/MagenticBrain
FaraVisionAgent Vision-language model (Fara 1.5-9B) image+text inference
MagenticUIServer ASP.NET Core + SignalR multi-agent server (FileSurfer, WebFetcher, Coder)
ConsoleAppDemo Interactive console application

🌐 Reference App: ElBruno.MagenticUI — full Blazor Server port of microsoft/magentic-ui running locally with this library.

Requirements

  • .NET 8.0 or .NET 10.0
  • CPU (default), NVIDIA GPU (CUDA), or Windows GPU (DirectML)
  • ~2-8 GB disk space per model (depending on size and quantization)

Building from Source

git clone https://github.com/elbruno/ElBruno.LocalLLMs.git
cd ElBruno.LocalLLMs
dotnet restore ElBruno.LocalLLMs.slnx
dotnet build ElBruno.LocalLLMs.slnx
dotnet test ElBruno.LocalLLMs.slnx --framework net8.0

Run integration tests (downloads real models — requires internet):

RUN_INTEGRATION_TESTS=true dotnet test ElBruno.LocalLLMs.slnx --framework net8.0

Integration tests validate the full lifecycle (download → infer → cache hit → delete) for all 35 supported models. See docs/tests/README.md for details.

Documentation

🤝 Contributing

Contributions are welcome! Please:

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

📄 License

This project is licensed under the MIT License — see the LICENSE file for details.

👋 About the Author

Hi! I'm ElBruno 🧡, a passionate developer and content creator exploring AI, .NET, and modern development practices.

Made with ❤️ by ElBruno

If you like this project, consider following my work across platforms:

  • 📻 Podcast: No Tienen Nombre — Spanish-language episodes on AI, development, and tech culture
  • 💻 Blog: ElBruno.com — Deep dives on embeddings, RAG, .NET, and local AI
  • 📺 YouTube: youtube.com/elbruno — Demos, tutorials, and live coding
  • 🔗 LinkedIn: @elbruno — Professional updates and insights
  • 𝕏 Twitter: @elbruno — Quick tips, releases, and tech news

🙏 Acknowledgments

About

C# local LLM chat completions library using ONNX Runtime, compatible with Microsoft.Extensions.AI

Topics

Resources

Contributing

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages