Description
On Linux x64, .NET 10.0.12 leaks native memory each time that a cached TypeInitializationException is rethrown from native runtime code (InitClassSlow) and caught in managed code.
The leaked blocks are 3400 B posix_memalign allocations. They contain a CONTEXT and an EXCEPTION_RECORD (AllocateExceptionRecords, pal/src/exception/seh-unwind.cpp).
In our Kubernetes production deployment, the trigger is Npgsql (GSS encryption mode Prefer, SSL on) in a container without libgssapi_krb5.so.2.
Each new physical connection calls NegotiateAuthentication.GetOutgoingBlob. This rethrows the TypeInitializationException of Interop+NetSecurityNative. Npgsql catches it.
A service grew about 28 MiB/day. A heap dump showed about 44,600 leaked 3400 B blocks after 5 days. The managed heap was stable at about 55 MB.
Analysis
- The rethrow starts in
InitClassSlow, which throws the cached TypeInitializationException from native code.
- In the dump, the leaked CONTEXT records were captured in
DispatchManagedException(PAL_SEHException& ex, bool isHardwareException) (vm/exceptionhandling.cpp). This function calls RtlCaptureContext(ex.GetContextRecord()).
UNINSTALL_MANAGED_EXCEPTION_DISPATCHER_EX (vm/exceptmacros.h) moves the PAL_SEHException into the local exCopy and calls DispatchManagedException(exCopy, false). This call does not return. Managed EH resumes in the managed catch, so ~PAL_SEHException() (FreeRecords) does not run.
- As far as I can see, the
ExInfo that is created in DispatchManagedException(OBJECTREF, CONTEXT*, EXCEPTION_RECORD*) does not take ownership of these records on this path. Nothing frees them.
Related issues
Reproduction Steps
Use mcr.microsoft.com/dotnet/runtime:10.0-noble-chiseled (it has no libgssapi):
using System.Net.Security;
static long Rss() => long.Parse(File.ReadAllLines("/proc/self/status").First(l => l.StartsWith("VmRSS")).Split(' ', StringSplitOptions.RemoveEmptyEntries)[1]);
void Once()
{
try { new NegotiateAuthentication(new NegotiateAuthenticationClientOptions { TargetName = "postgres/db" }).GetOutgoingBlob(ReadOnlySpan<byte>.Empty, out _); }
catch (TypeInitializationException) { }
}
for (int i = 0; i < 2000; i++) Once();
GC.Collect(); GC.WaitForPendingFinalizers(); GC.Collect();
long r0 = Rss();
for (int i = 0; i < 40000; i++) Once();
GC.Collect(); GC.WaitForPendingFinalizers(); GC.Collect();
Console.WriteLine($"RSS delta {Rss() - r0} kB");
Results (.NET 10.0.12, linux-x64)
| Image |
Calls |
RSS delta |
Per call |
| runtime:10.0-noble-chiseled |
20,000 |
82 MB |
~4.2 KB |
| runtime:10.0-noble-chiseled |
40,000 |
152 MB |
~3.9 KB |
Same loop with a managed throw/catch |
40,000 |
15 MB (not linear) |
- |
End-to-end with Npgsql 10.0.3 (Pooling=false, SslMode=Require, PostgreSQL 17 with SSL, same image):
GssEncryptionMode |
Connections |
RSS delta |
Per connection |
First-chance TIE |
| Prefer |
10,000 |
37 MB |
~3.8 KB |
10,000 |
| Prefer |
30,000 |
103 MB |
~3.5 KB |
30,000 |
| Disable |
30,000 |
5 MB |
- |
0 |
The runs used amd64 emulation (Podman) on an arm64 host. The production service runs natively on linux-x64 and shows the same leak.
Expected behavior
Memory RSS stays at a predictable level for long-running dotnet containers, where the runtime throughs TypeInitializationException.
Actual behavior
A steady, uncontrollable growth of native memory (not managed heap), which eventually causes applications deployed in containerized environments to reach environment limits and become unstable or get killed.
Regression?
Yes
Known Workarounds
In our case: Set GssEncryptionMode=Disable in the Npgsql connection string.
Configuration
No response
Other information
Note: Claude Opus 5.5 (GitHub Copilot CLI) assisted in this analysis. A human examined and verified the results and the repro.
Description
On Linux x64, .NET 10.0.12 leaks native memory each time that a cached
TypeInitializationExceptionis rethrown from native runtime code (InitClassSlow) and caught in managed code.The leaked blocks are 3400 B
posix_memalignallocations. They contain aCONTEXTand anEXCEPTION_RECORD(AllocateExceptionRecords,pal/src/exception/seh-unwind.cpp).In our Kubernetes production deployment, the trigger is Npgsql (GSS encryption mode
Prefer, SSL on) in a container withoutlibgssapi_krb5.so.2.Each new physical connection calls
NegotiateAuthentication.GetOutgoingBlob. This rethrows theTypeInitializationExceptionofInterop+NetSecurityNative. Npgsql catches it.A service grew about 28 MiB/day. A heap dump showed about 44,600 leaked 3400 B blocks after 5 days. The managed heap was stable at about 55 MB.
Analysis
InitClassSlow, which throws the cachedTypeInitializationExceptionfrom native code.DispatchManagedException(PAL_SEHException& ex, bool isHardwareException)(vm/exceptionhandling.cpp). This function callsRtlCaptureContext(ex.GetContextRecord()).UNINSTALL_MANAGED_EXCEPTION_DISPATCHER_EX(vm/exceptmacros.h) moves thePAL_SEHExceptioninto the localexCopyand callsDispatchManagedException(exCopy, false). This call does not return. Managed EH resumes in the managed catch, so~PAL_SEHException()(FreeRecords) does not run.ExInfothat is created inDispatchManagedException(OBJECTREF, CONTEXT*, EXCEPTION_RECORD*)does not take ownership of these records on this path. Nothing frees them.Related issues
GssEncryptionMode=Disablefixes it.TypeInitializationExceptionand continues without GSS. But the exception still occurs for each new physical connection (see the Npgsql 10.0.3 results above).TypeInitializationException.Reproduction Steps
Use
mcr.microsoft.com/dotnet/runtime:10.0-noble-chiseled(it has no libgssapi):Results (.NET 10.0.12, linux-x64)
throw/catchEnd-to-end with Npgsql 10.0.3 (
Pooling=false,SslMode=Require, PostgreSQL 17 with SSL, same image):GssEncryptionModeThe runs used amd64 emulation (Podman) on an arm64 host. The production service runs natively on linux-x64 and shows the same leak.
Expected behavior
Memory RSS stays at a predictable level for long-running dotnet containers, where the runtime throughs TypeInitializationException.
Actual behavior
A steady, uncontrollable growth of native memory (not managed heap), which eventually causes applications deployed in containerized environments to reach environment limits and become unstable or get killed.
Regression?
Yes
Known Workarounds
In our case: Set
GssEncryptionMode=Disablein the Npgsql connection string.Configuration
No response
Other information
Note: Claude Opus 5.5 (GitHub Copilot CLI) assisted in this analysis. A human examined and verified the results and the repro.