Support for JPEG XL (JXL) images - #3153
winscripter wants to merge 181 commits into
Conversation
Implementation of ac_strategy.h and ac_strategy.c
For now JxlMemoryManager will be a wrapper around MemoryPool<T>.
Implementation of image.h and image.c; AC strategy implementation was slightly adjusted to reduce errors.
This is an implementation of field_encodings.h. Note that I avoided implementing EnumValid() and Values() functions, as we have dedicated methods in .NET to do exactly that (Enum.IsDefined, Enum.GetValues)
Implementation of spline.h
Implemented ANS constants
|
While I'm working on this, I'd like to note something important. Libjxl is licensed under the BSD 3-Clause license, and since I'm using libjxl code as reference, that means the license must be included. I'm not really sure what would be the proper way to include the license. I might place the LICENSE.txt file in the Jxl folder or add a README linking to the libjxl repo. |
See ans_common.h
It is too large for a struct.
See ans_common.h
Add JxlAnsEntry and JxlAnsSymbol. See ans_common.h. These correspond to the Entry and Symbol structures within AliasTable.
Currently, there's a VarLenUint8/VarLenUint16 as well as histogram parsing implementation. I will additionally have to implement parsing of ANS codes, uint config and LZ77 parameters.
| { | ||
| [Conditional("DEBUG")] | ||
| public static void LogWarning(string message) => | ||
| // TODO: do we allow Debug.WriteLine in this codebase? |
|
|
||
| namespace SixLabors.ImageSharp.Formats.Jxl.Processing.Decoder.Modular; | ||
|
|
||
| internal sealed class JxlModularFrameDecoder |
There was a problem hiding this comment.
Just as reminder: this class is WIP
| Vector<short> res = Vector.Create<short>(residuals); | ||
| Vector<short> token = FjxlSimdUtils.ValueToToken(res); | ||
| Vector<short> nbits = FjxlSimdUtils.SaturateSubtract(token, Vector<short>.One); | ||
| Vector<short> bits = FjxlSimdUtils.SaturateSubtract(res, FjxlSimdUtils.Pow2(nbits)); |
There was a problem hiding this comment.
These are SIMD vectors, we can't stackalloc-slice them.
…e), add tests (incomplete), add compressed DC (incomplete), update folder structure
| InlineArray3<int> result = default; | ||
| result[0] = 1; | ||
| result[2] = 2; | ||
| return result; |
There was a problem hiding this comment.
This way the inline array gets created on the stack, the copied to caller.
So maybe pass the result as out parameter? So the caller provides the memory, and this method just fills it.
But note it's only 12 bytes (maybe with padding 16), so I'd do this only when the code remains clean, i.e. the usage doesn't become awkward.
| return false; | ||
| } | ||
|
|
||
| foreach (int c in new[] { 1, 0, 2 }) |
There was a problem hiding this comment.
foreach (int c in (ReadOnlySpan<int>)[1, 0, 2])to avoid the allocation of the array.
Note: starting with .NET 10 the JIT will very likely stack-allocate that array, because it sees that the reference to the array doesn't escape that scope. But
- older runtimes (< .NET 10) can't do this
- it's still more code that just referencing at as ROS
|
|
||
| public void StartBox(bool boxUntilEof, long contentsSize) | ||
| { | ||
| this.buffer.Dispose(); |
There was a problem hiding this comment.
For the memory stream this has no effect, so can be omitted.
| public void StartBox(bool boxUntilEof, long contentsSize) | ||
| { | ||
| this.buffer.Dispose(); | ||
| this.buffer = new MemoryStream(); |
There was a problem hiding this comment.
And just reset the buffer.Position = 0, instead of creating a new object?
|
|
||
| // TODO: unrolling this would lead to big performance benefits | ||
| for (int iy = 0; iy < 4; iy++) | ||
| for (int iy = 0, iy4 = 0, iy28 = 0; iy < 4; iy++, iy4 += 4, iy28 += 16) |
There was a problem hiding this comment.
I think the JIT won't unroll that loop for us. See bottom of #3153 (comment) comment.
So write it as
int iy28 = 0;
for (int iy4 = 0; iy4 < 16; iy4 += 4)
{
for (int ix = 0; ix < 4; ix++)
{
if (ix == 0 && iy4 == 0)
{
// iy4 == 0 is correct, as iy4 := iy * 4, and when iy == 0 <=> iy4 == 0.
}
}
// ...
iy28 += 16;
}| // IDCT4x4 in (odd, even) positions. | ||
| block[0] = dcs[1]; | ||
| for (int iy = 0; iy < 4; iy++) | ||
| for (int iy = 0, iy4 = 0, iy28 = 0; iy < 4; iy++, iy4 += 4, iy28 += 16) |
There was a problem hiding this comment.
Similar, only one loop variable should be incremented as it seems.
| block[0] = dcs[2]; | ||
|
|
||
| for (int iy = 0; iy < 4; iy++) | ||
| for (int iy = 0, iy8 = 0; iy < 4; iy++, iy8 += 8) |
| float r00 = c00 + c01 + c10 + c11; | ||
| float r01 = c00 + c01 - c10 - c11; | ||
| float r10 = c00 - c01 + c10 - c11; | ||
| float r11 = c00 - c01 - c10 + c11; |
There was a problem hiding this comment.
Please add braces to achieve instruction level parallelism at CPU level.
Look e.g. at godbolt example. In Foo all sums go to xmm0 register, creating a dependency on that. In Bar the intermediate sums go to xmm0 and xmm1, so a + b and c + d can be done in parallel.
Well, of course not a huge gain here, but the braces don't complicate the code.
| coefficients[0] = (block00 + block01 + block10 + block11) * 0.25f; | ||
| coefficients[1] = (block00 + block01 - block10 - block11) * 0.25f; | ||
| coefficients[8] = (block00 - block01 + block10 - block11) * 0.25f; | ||
| coefficients[9] = (block00 - block01 - block10 + block11) * 0.25f; |
There was a problem hiding this comment.
That's the same as above? Maybe extract it to a helper method?
And the * 0.25f on the 4 floats could be done vectorized in one pass then more easily.
| fixed (float* pRowTop = rowTop) | ||
| { | ||
| fixed (float* pRow = row) | ||
| { | ||
| fixed (float* pRowBottom = rowBottom) | ||
| { |
There was a problem hiding this comment.
More or less a matter of taste, but to avoid some horizontal indentation:
fixed (float* pRowTop = rowTop)
fixed (float* pRow = row)
fixed (float* pRowBottom = rowBottom)
{
// ...
}
Prerequisites
Description
This is a work-in-progress PR whose goal is to introduce decoding and encoding of JPEG XL (*.jxl) images.
Reference software
I use libjxl as reference. See https://github.com/libjxl/libjxl.
Performance
I will begin by applying light optimizations as I implement parts of the JPEG XL codec. Once the codec seems complete enough to handle decoding and encoding of JPEG XL images, I will apply heavier optimizations. Examples include but are not limited to stack allocation, array pooling, and SIMD.
Implementations
The JPEG XL codec lives under
src/ImageSharp/Formats/Jxl.Testing
I will start adding tests whenever the codec is complete enough to handle decoding of JPEG XL images.
Additionally, JPEG XL reference software, libjxl, contains its own tests too, which I might also implement without modification.
Progress
🟡 AC strategy
🟢 AC strategy image/row
🟢 AC context
🔴 AC strategy tests
🟠 Adaptive Quantization (encoder)
🟠 ANS Entropy
🟢 ANS Entropy: Common
🟢 ANS Entropy: Common (Tests)
🟠 ANS Entropy: Decoder (symbol reader is incomplete)
🟠 ANS Entropy: Encoder (SIMD bit-cost calculation only)
🔴 ANS Entropy: Tests
🟢 Alpha Blending
🟡 Bit I/O
🟢 Bit I/O: Bit reader
🟢 Bit I/O: Bit writer
🔴 Bit I/O: tests
🟢 Box Content Decoder
🟢 Box Content Decoder: Uncompressed boxes
🟢 Box Content Decoder: Brotli-compressed boxes
🟡 Butteraugli
🟢 Butteraugli: Shared methods
🟢 Butteraugli: Abstractions
🟢 Butteraugli: Default comparator
🔴 Butteraugli: Encoder comparator
🟡 Cache
🟠 Cache: Decoder
🔴 Cache: Encoder
🟠 Chroma From Luma
🟢 Chroma From Luma Abstractions
🔴 Chroma From Luma Encoder
🟡 Coefficient Order
🟢 Coefficient Order: Forward
🟢 Coefficient Order: Main
🟢 Coefficient Order: Encoder
🔴 Coefficient Order: Tests
🔴 Compressed DC
🟡 Context Map
🟢 Context Map: Abstractions
🟢 Context Map: Decoder
🔴 Context Map: Encoder
🟡 Convolution
🟢 Convolution: Symmetric
🟢 Convolution: Separable
🟢 Convolution: Slow
🟢 Convolution: SIMD
🔴 Convolution: Separable5 Encoder
🔴 Convolution: Tests
🟠 Decoder: Frame
🟢 Decoder: Group
🟢 Decoder: Group Border
🟠 Decoder: Main
🟢 Decoder: Main: Codestream parser
🟢 Decoder: Main: Container format parser
🟡 Discrete Cosine Transform
🟢 Discrete Cosine Transform: DCT scales
🟢 Discrete Cosine Transform: Block data wrapper
🟠 Discrete Cosine Transform: Block-based
🟢 Discrete Cosine Transform: Slow DCT for reference in tests
🔴 Discrete Cosine Transform: Tests
🔴 Encoder dot detection
🔴 Encoder dot dictionary
🟠 Entropy coding
🔴 Encoder entropy coding
🟠 Fast Lossless Encoder
🟢 Fields
🟢 Fields: Visitor abstractions
🟢 Fields: Parser
🟢 Fields: Writer
🔴 Frame Encoder
🟢 Gaborish
🟢 Gaborish Encoder
🟢 Gaborish Tests
🔴 Group Encoder
🔴 Heuristics Encoder
🟡 Image Bundle
🟠 Image Bundle: Decoder/Common
🔴 Image Bundle: Encoder
🟡 LZ77 compression
🟢 LZ77: Fast Lossless Encoder
🔴 LZ77: Standard Encoder
🟢 Huffman compression
🟢 Huffman compression: Shared
🟢 Huffman compression: Decoder
🟢 Huffman compression: Encoder
🟡 Modular
🟢 Modular: Transforms
🟢 Modular: Transforms: Palette (Inverse)
🟢 Modular: Transforms: Palette (Forward)
🟢 Modular: Transforms: RCT (Inverse)
🟢 Modular: Transforms: RCT (Forward)
🟢 Modular: Transforms: Squeeze (Inverse)
🟢 Modular: Transforms: Squeeze (Forward)
🟢 Modular: Encoding
🟢 Modular: Encoding: Context Prediction
🟢 Modular: Encoding: MA decoder
🟢 Modular: Encoding: MA encoder
🟢 Modular: Encoding: Tree Samples
🟢 Modular: Encoding: Encoding decoder
🟢 Modular: Encoding: Encoding encoder
🔴 Modular: Decoder
🔴 Modular: Encoder
🔴 Modular: Encoder SIMD
🔴 Modular: Tests
🟠 Patch Dictionary: Decoder
🔴 Patch Dictionary: Encoder
🟠 Passes State: Decoder
🔴 Passes State: Encoder
🟢 Passes State: Shared
🔴 Encoder Main
🔴 Encoder Main
🔴 Encoder Internal
🔴 Encoder Tests
🟢 Encoder: Linear Algebra
🟢 Encoder: Linear Algebra Tests
🟢 Image Operations
🟢 Image Operations
🟢 Image Operations: Tests
🟢 Common I/O: Frame Header
🟢 Common I/O: Metadata
🟢 Common I/O: Container format
🟠 JPEG to JPEG XL lossless compression
🟢 JPEG to JPEG XL lossless compression: JPEG parser/writer
🟠 JPEG to JPEG XL lossless compression (decoder)
🔴 JPEG to JPEG XL lossless compression (encoder)
🟡 Splines
🟢 Splines
🔴 Splines Tests
🟢 Quantizer
🟢 Dequantizer matrices
🟢 Quantizer encoding
🟢 Quantizer weights
🟡 Noise
🟢 Noise: Shared
🟢 Noise: Decoder
🟠 Noise: Encoder
🟢 Noise: Simulation of Photon Noise
🟢 Noise: Simulation of Photon Noise (Tests)
🔴 Noise Tests
🔴 JPEG XL Testing Tools
🟠 Render Pipeline
🔴 Render Pipeline: Main
🔴 Render Pipeline: Low Memory Render Pipeline
🟠 Render Pipeline: Stages Abstractions
🔴 Render Pipeline: Stages: Blending
🔴 Render Pipeline: Stages: Chroma Upsampling
🔴 Render Pipeline: Stages: CMS
🟢 Render Pipeline: Stages: EPF
🔴 Render Pipeline: Stages: From Linear
🟢 Render Pipeline: Stages: Gaborish
🔴 Render Pipeline: Stages: Noise
🟢 Render Pipeline: Stages: Patches
🔴 Render Pipeline: Stages: Splines
🟢 Render Pipeline: Stages: Spot color
🔴 Render Pipeline: Stages: To Linear
🔴 Render Pipeline: Stages: Tone mapping
🔴 Render Pipeline: Stages: Upsampling
🟠 Render Pipeline: Stages: Write to Output
🔴 Render Pipeline: Stages: XYB
🟢 Render Pipeline: Stages: Y'Cb'Cr -> RGB
🟡 Color Management System (CMS)
🟢 CMS: Transfer Functions
🟢 CMS: Abstractions/Color Encoding
🔴 CMS: Tone Mapping
🔴 CMS: Interface
Other completed things: