Skip to content

Support for JPEG XL (JXL) images - #3153

Draft
winscripter wants to merge 181 commits into
SixLabors:mainfrom
winscripter:jxl-support
Draft

winscripter wants to merge 181 commits into
SixLabors:mainfrom
winscripter:jxl-support

Conversation

@winscripter

@winscripter winscripter commented Jul 15, 2026

Copy link
Copy Markdown

Prerequisites

  • I have written a descriptive pull-request title
  • I have verified that there are no overlapping pull-requests open
  • I have verified that I am following the existing coding patterns and practice as demonstrated in the repository. These follow strict Stylecop rules 👮.
  • I have provided test coverage for my change (where applicable)

Description

This is a work-in-progress PR whose goal is to introduce decoding and encoding of JPEG XL (*.jxl) images.

Reference software
I use libjxl as reference. See https://github.com/libjxl/libjxl.

Performance
I will begin by applying light optimizations as I implement parts of the JPEG XL codec. Once the codec seems complete enough to handle decoding and encoding of JPEG XL images, I will apply heavier optimizations. Examples include but are not limited to stack allocation, array pooling, and SIMD.

Implementations
The JPEG XL codec lives under src/ImageSharp/Formats/Jxl.

Testing
I will start adding tests whenever the codec is complete enough to handle decoding of JPEG XL images.

Additionally, JPEG XL reference software, libjxl, contains its own tests too, which I might also implement without modification.

Progress

  • 🔴 - work didn't start yet
  • 🟠 - exists, but incomplete
  • 🟡 - either decoding side, encoding side, or tests are complete, but not the other
  • 🟢 - complete

  • 🟡 AC strategy

  • 🟢 AC strategy image/row

  • 🟢 AC context

  • 🔴 AC strategy tests

  • 🟠 Adaptive Quantization (encoder)

  • 🟠 ANS Entropy

  • 🟢 ANS Entropy: Common

  • 🟢 ANS Entropy: Common (Tests)

  • 🟠 ANS Entropy: Decoder (symbol reader is incomplete)

  • 🟠 ANS Entropy: Encoder (SIMD bit-cost calculation only)

  • 🔴 ANS Entropy: Tests

  • 🟢 Alpha Blending

  • 🟡 Bit I/O

  • 🟢 Bit I/O: Bit reader

  • 🟢 Bit I/O: Bit writer

  • 🔴 Bit I/O: tests

  • 🟢 Box Content Decoder

  • 🟢 Box Content Decoder: Uncompressed boxes

  • 🟢 Box Content Decoder: Brotli-compressed boxes

  • 🟡 Butteraugli

  • 🟢 Butteraugli: Shared methods

  • 🟢 Butteraugli: Abstractions

  • 🟢 Butteraugli: Default comparator

  • 🔴 Butteraugli: Encoder comparator

  • 🟡 Cache

  • 🟠 Cache: Decoder

  • 🔴 Cache: Encoder

  • 🟠 Chroma From Luma

  • 🟢 Chroma From Luma Abstractions

  • 🔴 Chroma From Luma Encoder

  • 🟡 Coefficient Order

  • 🟢 Coefficient Order: Forward

  • 🟢 Coefficient Order: Main

  • 🟢 Coefficient Order: Encoder

  • 🔴 Coefficient Order: Tests

  • 🔴 Compressed DC

  • 🟡 Context Map

  • 🟢 Context Map: Abstractions

  • 🟢 Context Map: Decoder

  • 🔴 Context Map: Encoder

  • 🟡 Convolution

  • 🟢 Convolution: Symmetric

  • 🟢 Convolution: Separable

  • 🟢 Convolution: Slow

  • 🟢 Convolution: SIMD

  • 🔴 Convolution: Separable5 Encoder

  • 🔴 Convolution: Tests

  • 🟠 Decoder: Frame

  • 🟢 Decoder: Group

  • 🟢 Decoder: Group Border

  • 🟠 Decoder: Main

  • 🟢 Decoder: Main: Codestream parser

  • 🟢 Decoder: Main: Container format parser

  • 🟡 Discrete Cosine Transform

  • 🟢 Discrete Cosine Transform: DCT scales

  • 🟢 Discrete Cosine Transform: Block data wrapper

  • 🟠 Discrete Cosine Transform: Block-based

  • 🟢 Discrete Cosine Transform: Slow DCT for reference in tests

  • 🔴 Discrete Cosine Transform: Tests

  • 🔴 Encoder dot detection

  • 🔴 Encoder dot dictionary

  • 🟠 Entropy coding

  • 🔴 Encoder entropy coding

  • 🟠 Fast Lossless Encoder

  • 🟢 Fields

  • 🟢 Fields: Visitor abstractions

  • 🟢 Fields: Parser

  • 🟢 Fields: Writer

  • 🔴 Frame Encoder

  • 🟢 Gaborish

  • 🟢 Gaborish Encoder

  • 🟢 Gaborish Tests

  • 🔴 Group Encoder

  • 🔴 Heuristics Encoder

  • 🟡 Image Bundle

  • 🟠 Image Bundle: Decoder/Common

  • 🔴 Image Bundle: Encoder

  • 🟡 LZ77 compression

  • 🟢 LZ77: Fast Lossless Encoder

  • 🔴 LZ77: Standard Encoder

  • 🟢 Huffman compression

  • 🟢 Huffman compression: Shared

  • 🟢 Huffman compression: Decoder

  • 🟢 Huffman compression: Encoder

  • 🟡 Modular

  • 🟢 Modular: Transforms

  • 🟢 Modular: Transforms: Palette (Inverse)

  • 🟢 Modular: Transforms: Palette (Forward)

  • 🟢 Modular: Transforms: RCT (Inverse)

  • 🟢 Modular: Transforms: RCT (Forward)

  • 🟢 Modular: Transforms: Squeeze (Inverse)

  • 🟢 Modular: Transforms: Squeeze (Forward)

  • 🟢 Modular: Encoding

  • 🟢 Modular: Encoding: Context Prediction

  • 🟢 Modular: Encoding: MA decoder

  • 🟢 Modular: Encoding: MA encoder

  • 🟢 Modular: Encoding: Tree Samples

  • 🟢 Modular: Encoding: Encoding decoder

  • 🟢 Modular: Encoding: Encoding encoder

  • 🔴 Modular: Decoder

  • 🔴 Modular: Encoder

  • 🔴 Modular: Encoder SIMD

  • 🔴 Modular: Tests

  • 🟠 Patch Dictionary: Decoder

  • 🔴 Patch Dictionary: Encoder

  • 🟠 Passes State: Decoder

  • 🔴 Passes State: Encoder

  • 🟢 Passes State: Shared

  • 🔴 Encoder Main

  • 🔴 Encoder Main

  • 🔴 Encoder Internal

  • 🔴 Encoder Tests

  • 🟢 Encoder: Linear Algebra

  • 🟢 Encoder: Linear Algebra Tests

  • 🟢 Image Operations

  • 🟢 Image Operations

  • 🟢 Image Operations: Tests

  • 🟢 Common I/O: Frame Header

  • 🟢 Common I/O: Metadata

  • 🟢 Common I/O: Container format

  • 🟠 JPEG to JPEG XL lossless compression

  • 🟢 JPEG to JPEG XL lossless compression: JPEG parser/writer

  • 🟠 JPEG to JPEG XL lossless compression (decoder)

  • 🔴 JPEG to JPEG XL lossless compression (encoder)

  • 🟡 Splines

  • 🟢 Splines

  • 🔴 Splines Tests

  • 🟢 Quantizer

  • 🟢 Dequantizer matrices

  • 🟢 Quantizer encoding

  • 🟢 Quantizer weights

  • 🟡 Noise

  • 🟢 Noise: Shared

  • 🟢 Noise: Decoder

  • 🟠 Noise: Encoder

  • 🟢 Noise: Simulation of Photon Noise

  • 🟢 Noise: Simulation of Photon Noise (Tests)

  • 🔴 Noise Tests

  • 🔴 JPEG XL Testing Tools

  • 🟠 Render Pipeline

  • 🔴 Render Pipeline: Main

  • 🔴 Render Pipeline: Low Memory Render Pipeline

  • 🟠 Render Pipeline: Stages Abstractions

  • 🔴 Render Pipeline: Stages: Blending

  • 🔴 Render Pipeline: Stages: Chroma Upsampling

  • 🔴 Render Pipeline: Stages: CMS

  • 🟢 Render Pipeline: Stages: EPF

  • 🔴 Render Pipeline: Stages: From Linear

  • 🟢 Render Pipeline: Stages: Gaborish

  • 🔴 Render Pipeline: Stages: Noise

  • 🟢 Render Pipeline: Stages: Patches

  • 🔴 Render Pipeline: Stages: Splines

  • 🟢 Render Pipeline: Stages: Spot color

  • 🔴 Render Pipeline: Stages: To Linear

  • 🔴 Render Pipeline: Stages: Tone mapping

  • 🔴 Render Pipeline: Stages: Upsampling

  • 🟠 Render Pipeline: Stages: Write to Output

  • 🔴 Render Pipeline: Stages: XYB

  • 🟢 Render Pipeline: Stages: Y'Cb'Cr -> RGB

  • 🟡 Color Management System (CMS)

  • 🟢 CMS: Transfer Functions

  • 🟢 CMS: Abstractions/Color Encoding

  • 🔴 CMS: Tone Mapping

  • 🔴 CMS: Interface

Other completed things:

  • 🟢 Frame Dimensions
  • 🟢 Field Encodings
  • 🟢 Gamma correction (tests, encoder)
  • 🟢 Shared math functions
  • 🟢 Lehmer Codes
  • 🟢 Lehmer Codes: Tests
  • 🟢 Luminance helper
  • 🟢 Opsin parameters
  • 🟢 Passes shared state
  • 🟢 SIMD utilities
  • 🟢 TOC
  • 🟢 XorShift128Plus
  • 🟢 XorShift128Plus Tests

@CLAassistant

CLAassistant commented Jul 15, 2026

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

Implementation of ac_strategy.h and ac_strategy.c
For now JxlMemoryManager will be a wrapper around MemoryPool<T>.
Implementation of image.h and image.c; AC strategy implementation was slightly adjusted to reduce errors.
This is an implementation of field_encodings.h.

Note that I avoided implementing EnumValid() and Values() functions, as we have dedicated methods in .NET to do exactly that (Enum.IsDefined, Enum.GetValues)
Implementation of spline.h
Implemented ANS constants
@winscripter

Copy link
Copy Markdown
Author

While I'm working on this, I'd like to note something important.

Libjxl is licensed under the BSD 3-Clause license, and since I'm using libjxl code as reference, that means the license must be included.

I'm not really sure what would be the proper way to include the license. I might place the LICENSE.txt file in the Jxl folder or add a README linking to the libjxl repo.

Comment thread src/ImageSharp/Formats/Jxl/Metadata/JxlExifOrientation.cs Outdated
Comment thread src/ImageSharp/Formats/Jxl/Metadata/JxlExtraChannel.cs Outdated
Comment thread src/ImageSharp/Formats/Jxl/Splines/JxlSplineEntropyContext.cs Outdated
Comment thread src/ImageSharp/Formats/Jxl/JxlFrameDimensions.cs Outdated
It is too large for a struct.
Add JxlAnsEntry and JxlAnsSymbol.

See ans_common.h. These correspond to the Entry and Symbol structures within AliasTable.
Comment thread src/ImageSharp/Formats/Jxl/IO/JxlAnsHelper.cs Outdated
Comment thread src/ImageSharp/Formats/Jxl/IO/JxlAnsHelper.cs Outdated
Comment thread src/ImageSharp/Formats/Jxl/JxlThrowHelper.cs Outdated
Currently, there's a VarLenUint8/VarLenUint16 as well as histogram parsing implementation.

I will additionally have to implement parsing of ANS codes, uint config and LZ77 parameters.
Comment thread src/ImageSharp/Formats/Jxl/IO/Jpeg/JpegBitWriter.cs Outdated
{
[Conditional("DEBUG")]
public static void LogWarning(string message) =>
// TODO: do we allow Debug.WriteLine in this codebase?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment thread src/ImageSharp/Formats/Jxl/IO/Jpeg/JpegSerializationStage.cs Outdated
Comment thread src/ImageSharp/Formats/Jxl/IO/Jpeg/JpegWriter.cs Outdated
Comment thread src/ImageSharp/Formats/Jxl/IO/Jpeg/JpegWriter.cs Outdated
Comment thread src/ImageSharp/Formats/Jxl/IO/Jpeg/JpegWriter.cs Outdated
Comment thread src/ImageSharp/Formats/Jxl/IO/Jpeg/JpegWriter.cs Outdated

namespace SixLabors.ImageSharp.Formats.Jxl.Processing.Decoder.Modular;

internal sealed class JxlModularFrameDecoder

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just as reminder: this class is WIP

Comment thread src/ImageSharp/Formats/Jxl/Processing/Dct/JxlDctScales.cs Outdated
Comment thread src/ImageSharp/Formats/Jxl/Processing/Decoder/JxlTransformsDecoder.cs Outdated
Comment thread src/ImageSharp/Formats/Jxl/Processing/Decoder/JxlTransformsDecoder.cs Outdated
Comment on lines +12 to +15
Vector<short> res = Vector.Create<short>(residuals);
Vector<short> token = FjxlSimdUtils.ValueToToken(res);
Vector<short> nbits = FjxlSimdUtils.SaturateSubtract(token, Vector<short>.One);
Vector<short> bits = FjxlSimdUtils.SaturateSubtract(res, FjxlSimdUtils.Pow2(nbits));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These are SIMD vectors, we can't stackalloc-slice them.

Comment thread src/ImageSharp/Formats/Jxl/Processing/Encoder/FastLossless/FjxlPrefixCode.cs Outdated
InlineArray3<int> result = default;
result[0] = 1;
result[2] = 2;
return result;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This way the inline array gets created on the stack, the copied to caller.
So maybe pass the result as out parameter? So the caller provides the memory, and this method just fills it.

But note it's only 12 bytes (maybe with padding 16), so I'd do this only when the code remains clean, i.e. the usage doesn't become awkward.

return false;
}

foreach (int c in new[] { 1, 0, 2 })

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

foreach (int c in (ReadOnlySpan<int>)[1, 0, 2])

to avoid the allocation of the array.

Note: starting with .NET 10 the JIT will very likely stack-allocate that array, because it sees that the reference to the array doesn't escape that scope. But

  • older runtimes (< .NET 10) can't do this
  • it's still more code that just referencing at as ROS


public void StartBox(bool boxUntilEof, long contentsSize)
{
this.buffer.Dispose();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For the memory stream this has no effect, so can be omitted.

public void StartBox(bool boxUntilEof, long contentsSize)
{
this.buffer.Dispose();
this.buffer = new MemoryStream();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

And just reset the buffer.Position = 0, instead of creating a new object?


// TODO: unrolling this would lead to big performance benefits
for (int iy = 0; iy < 4; iy++)
for (int iy = 0, iy4 = 0, iy28 = 0; iy < 4; iy++, iy4 += 4, iy28 += 16)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the JIT won't unroll that loop for us. See bottom of #3153 (comment) comment.

So write it as

int iy28 = 0;
for (int iy4 = 0; iy4 < 16; iy4 += 4)
{
    for (int ix = 0; ix < 4; ix++)
    {
        if (ix == 0 && iy4 == 0)
        {
            // iy4 == 0 is correct, as iy4 := iy * 4, and when iy == 0 <=> iy4 == 0.
        }
    }

    // ...

    iy28 += 16;
}

// IDCT4x4 in (odd, even) positions.
block[0] = dcs[1];
for (int iy = 0; iy < 4; iy++)
for (int iy = 0, iy4 = 0, iy28 = 0; iy < 4; iy++, iy4 += 4, iy28 += 16)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Similar, only one loop variable should be incremented as it seems.

block[0] = dcs[2];

for (int iy = 0; iy < 4; iy++)
for (int iy = 0, iy8 = 0; iy < 4; iy++, iy8 += 8)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also.

Comment on lines +376 to +379
float r00 = c00 + c01 + c10 + c11;
float r01 = c00 + c01 - c10 - c11;
float r10 = c00 - c01 + c10 - c11;
float r11 = c00 - c01 - c10 + c11;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please add braces to achieve instruction level parallelism at CPU level.

Look e.g. at godbolt example. In Foo all sums go to xmm0 register, creating a dependency on that. In Bar the intermediate sums go to xmm0 and xmm1, so a + b and c + d can be done in parallel.

Well, of course not a huge gain here, but the braces don't complicate the code.

Comment on lines +656 to +659
coefficients[0] = (block00 + block01 + block10 + block11) * 0.25f;
coefficients[1] = (block00 + block01 - block10 - block11) * 0.25f;
coefficients[8] = (block00 - block01 + block10 - block11) * 0.25f;
coefficients[9] = (block00 - block01 - block10 + block11) * 0.25f;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's the same as above? Maybe extract it to a helper method?
And the * 0.25f on the 4 floats could be done vectorized in one pass then more easily.

Comment on lines +19 to +24
fixed (float* pRowTop = rowTop)
{
fixed (float* pRow = row)
{
fixed (float* pRowBottom = rowBottom)
{

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

More or less a matter of taste, but to avoid some horizontal indentation:

fixed (float* pRowTop = rowTop)
fixed (float* pRow = row)
fixed (float* pRowBottom = rowBottom)
{
    // ...
}

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants