diff --git a/docs/cross_encoder/usage/usage.rst b/docs/cross_encoder/usage/usage.rst index 682e47d19..780bd991c 100644 --- a/docs/cross_encoder/usage/usage.rst +++ b/docs/cross_encoder/usage/usage.rst @@ -16,6 +16,7 @@ Once you have `installed <../../installation.html>`_ Sentence Transformers, you 1. :class:`~sentence_transformers.cross_encoder.model.CrossEncoder` 2. :meth:`CrossEncoder.predict ` 3. :meth:`CrossEncoder.rank ` + 4. :doc:`Speeding up Inference ` .. note:: MS Marco models return logits rather than scores between 0 and 1. Load the :class:`~sentence_transformers.cross_encoder.model.CrossEncoder` with ``activation_fn=torch.nn.Sigmoid()`` to get scores between 0 and 1. This does not affect the ranking. @@ -140,4 +141,3 @@ In this example, the multimodal CrossEncoder uses the same modular architecture Cross-Encoder vs Bi-Encoder <../../../examples/cross_encoder/applications/README> ../../../examples/sentence_transformer/applications/retrieve_rerank/README custom_models - efficiency diff --git a/docs/img/backends_benchmark_cpu.png b/docs/img/backends_benchmark_cpu.png index 614803cca..861b2be13 100644 Binary files a/docs/img/backends_benchmark_cpu.png and b/docs/img/backends_benchmark_cpu.png differ diff --git a/docs/img/backends_benchmark_cpu_cloud.png b/docs/img/backends_benchmark_cpu_cloud.png new file mode 100644 index 000000000..0c1fe951b Binary files /dev/null and b/docs/img/backends_benchmark_cpu_cloud.png differ diff --git a/docs/img/backends_benchmark_gpu.png b/docs/img/backends_benchmark_gpu.png index e3b2839eb..571c8146b 100644 Binary files a/docs/img/backends_benchmark_gpu.png and b/docs/img/backends_benchmark_gpu.png differ diff --git a/docs/img/llamacpp_gpu_imdb.svg b/docs/img/llamacpp_gpu_imdb.svg new file mode 100644 index 000000000..395de024f --- /dev/null +++ b/docs/img/llamacpp_gpu_imdb.svg @@ -0,0 +1,7798 @@ + + + + + + + + 2026-09-08T12:25:52.438646 + image/svg+xml + + + Matplotlib v3.11.1, https://matplotlib.org/ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + diff --git a/docs/img/llamacpp_gpu_nq.svg b/docs/img/llamacpp_gpu_nq.svg new file mode 100644 index 000000000..c54689347 --- /dev/null +++ b/docs/img/llamacpp_gpu_nq.svg @@ -0,0 +1,7900 @@ + + + + + + + + 2026-09-08T12:25:51.422977 + image/svg+xml + + + Matplotlib v3.11.1, https://matplotlib.org/ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + diff --git a/docs/img/llamacpp_gpu_stsb.svg b/docs/img/llamacpp_gpu_stsb.svg new file mode 100644 index 000000000..7a7f753e3 --- /dev/null +++ b/docs/img/llamacpp_gpu_stsb.svg @@ -0,0 +1,7966 @@ + + + + + + + + 2026-09-08T12:25:50.405213 + image/svg+xml + + + Matplotlib v3.11.1, https://matplotlib.org/ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + diff --git a/docs/multi_vector_encoder/usage/usage.rst b/docs/multi_vector_encoder/usage/usage.rst index 5117d92ed..b961f8222 100644 --- a/docs/multi_vector_encoder/usage/usage.rst +++ b/docs/multi_vector_encoder/usage/usage.rst @@ -18,6 +18,7 @@ Once you have `installed <../../installation.html>`_ Sentence Transformers, you 4. :meth:`MultiVectorEncoder.encode_document ` 5. :meth:`MultiVectorEncoder.similarity ` 6. :meth:`MultiVectorEncoder.similarity_pairwise ` + 7. :doc:`Speeding up Inference ` :: @@ -206,4 +207,3 @@ Multi-vector models can be loaded from any of the following sources, transparent ../../../examples/multi_vector_encoder/applications/README ../../../examples/multi_vector_encoder/evaluation/README custom_models - efficiency diff --git a/docs/quickstart.rst b/docs/quickstart.rst index 98b1c1bbe..578d180bc 100644 --- a/docs/quickstart.rst +++ b/docs/quickstart.rst @@ -25,7 +25,7 @@ Once you have `installed `_ Sentence Transformers, you can ea - :meth:`SentenceTransformer.similarity_pairwise ` - `SentenceTransformer > Usage <./sentence_transformer/usage/usage.html>`_ - - `SentenceTransformer > Usage > Speeding up Inference <./sentence_transformer/usage/efficiency.html>`_ + - `SentenceTransformer > Speeding up Inference <./sentence_transformer/usage/efficiency.html>`_ - `SentenceTransformer > Pretrained Models <./sentence_transformer/pretrained_models.html>`_ - `SentenceTransformer > Training Overview <./sentence_transformer/training_overview.html>`_ - `SentenceTransformer > Dataset Overview <./sentence_transformer/dataset_overview.html>`_ @@ -101,7 +101,7 @@ Finetuning Sentence Transformer models is easy and requires only a few lines of .. tip:: - Read `Sentence Transformer > Usage > Speeding up Inference `_ for tips on how to speed up inference of models by up to 2x-3x. + Read `Sentence Transformer > Speeding up Inference `_ for tips on how to speed up inference of models by up to 2x-3x. Cross Encoder ------------- @@ -228,7 +228,7 @@ Finetuning CrossEncoder models is easy and requires only a few lines of code. Fo .. tip:: - Read `CrossEncoder > Usage > Speeding up Inference `_ for tips on how to speed up inference of models by up to 2x-3x. + Read `CrossEncoder > Speeding up Inference `_ for tips on how to speed up inference of models by up to 2x-3x. Sparse Encoder -------------- @@ -294,7 +294,7 @@ Finetuning Sparse Encoder models is easy and requires only a few lines of code. .. tip:: - Read `Sparse Encoder > Usage > Speeding up Inference `_ for tips on how to speed up inference of models by up to 2x-3x. + Read `Sparse Encoder > Speeding up Inference `_ for tips on how to speed up inference of models by up to 2x-3x. Multi-Vector Encoder -------------------- @@ -391,7 +391,7 @@ Finetuning Multi-Vector Encoder models is easy and requires only a few lines of .. tip:: - Read `Multi-Vector Encoder > Usage > Speeding up Inference `_ for benchmarks and tips on how to speed up inference with the ONNX and OpenVINO backends. + Read `Multi-Vector Encoder > Speeding up Inference `_ for benchmarks and tips on how to speed up inference with the ONNX and OpenVINO backends. Next Steps ---------- diff --git a/docs/sentence_transformer/pretrained_models.md b/docs/sentence_transformer/pretrained_models.md index 3522643df..b2242dbf1 100644 --- a/docs/sentence_transformer/pretrained_models.md +++ b/docs/sentence_transformer/pretrained_models.md @@ -38,7 +38,7 @@ similarities = model.similarity(embeddings, embeddings) .. tip:: - Read `Sentence Transformer > Usage > Speeding up Inference <./usage/efficiency.html>`_ for tips on how to speed up inference of models by up to 2x-3x. + Read `Sentence Transformer > Speeding up Inference <./usage/efficiency.html>`_ for tips on how to speed up inference of models by up to 2x-3x. ``` ## Original Models diff --git a/docs/sentence_transformer/usage/efficiency.rst b/docs/sentence_transformer/usage/efficiency.rst index c1a0e6629..c7e127ad9 100644 --- a/docs/sentence_transformer/usage/efficiency.rst +++ b/docs/sentence_transformer/usage/efficiency.rst @@ -122,9 +122,13 @@ If you're using a GPU, then you can use the following options to speed up your i :alt: Flash Attention 2 Input Flattening Benchmark :width: 100% - Flash Attention 2 with input flattening always outperforms standard Flash Attention 2, while using considerably less VRAM. The gains grow with the variance in input length, with the mixed dataset with wildly varying lengths (10-500 tokens) benefitting the most. + Input flattening improves throughput and reduces VRAM use in this benchmark, with the largest gains on the + mixed dataset, where input lengths range from 10 to 500 tokens. This makes it particularly useful for batches + of texts with varying lengths, although the benefit depends on the model and batch size. - In the full backend benchmark below, fp16 with Flash Attention and input unpadding was the fastest configuration measured (3.87x over fp32, comparing each backend at its best batch size) at no loss of quality, making it the recommended GPU configuration when the model architecture supports it. + The `backend benchmark <#benchmarks>`_ below also shows the benefit of combining half precision with Flash Attention and + input unpadding: FP16 with this setup achieves the highest median speedup across the tested models (3.87x + over FP32), without reducing average task quality. Input flattening also speeds up training. When training with a gradient-cached loss such as :class:`~sentence_transformers.sentence_transformer.losses.CachedMultipleNegativesRankingLoss`, you can additionally set @@ -539,7 +543,7 @@ See this example for quantizing a model to ``int8`` with `static quantization Expand the benchmark details
- Speedup ratio: - - Quality ratio: The same models and hardware were used. We compare the evaluation quality against that of PyTorch with fp32, i.e. the default backend and precision. -
    -
  • - Evaluation: -
      -
    • - Semantic Textual Similarity: Spearman rank correlation based on cosine similarity on the sentence-transformers/stsb test set, computed via the EmbeddingSimilarityEvaluator. -
    • -
    • - Information Retrieval: NDCG@10 based on cosine similarity on the entire NanoBEIR collection of datasets, computed via the InformationRetrievalEvaluator. -
    • -
    -
  • -
- -
    -
  • - Backends: -
      -
    • - torch-fp32: PyTorch with float32 precision (default). -
    • -
    • - torch-fp16: PyTorch with float16 precision, via model_kwargs={"torch_dtype": "float16"}. -
    • -
    • - torch-bf16: PyTorch with bfloat16 precision, via model_kwargs={"torch_dtype": "bfloat16"}. -
    • -
    • - torch-fp16-fa2: PyTorch with float16 precision and FlashAttention-2 with automatic input unpadding, via model_kwargs={"torch_dtype": "float16", "attn_implementation": "flash_attention_2"}. -
    • -
    • - torch-bf16-fa2: the same with bfloat16 precision. -
    • -
    • - onnx: ONNX with float32 precision, via backend="onnx". -
    • -
    • - onnx-O1: ONNX with float32 precision and O1 optimization, via export_optimized_onnx_model(..., optimization_config="O1", ...) and backend="onnx". -
    • -
    • - onnx-O2: ONNX with float32 precision and O2 optimization, via export_optimized_onnx_model(..., optimization_config="O2", ...) and backend="onnx". -
    • -
    • - onnx-O3: ONNX with float32 precision and O3 optimization, via export_optimized_onnx_model(..., optimization_config="O3", ...) and backend="onnx". -
    • -
    • - onnx-O4: ONNX with float16 precision and O4 optimization, via export_optimized_onnx_model(..., optimization_config="O4", ...) and backend="onnx". -
    • -
    • - onnx-qint8: ONNX quantized to int8 with "avx512_vnni", via export_dynamic_quantized_onnx_model(..., quantization_config="avx512_vnni", ...) and backend="onnx". The different quantization configurations resulted in roughly equivalent speedups. -
    • -
    • - openvino: OpenVINO, via backend="openvino". -
    • -
    • - openvino-qint8: OpenVINO quantized to int8 via export_static_quantized_openvino_model(..., quantization_config=OVQuantizationConfig(), ...) and backend="openvino". -
    • -
    -
  • -
- - On GPU, fp16 with FlashAttention-2 and input unpadding is the strongest configuration everywhere (3.87x median at no loss of quality), making it the recommended setup when the model architecture supports it. Plain fp16 (2.92x) also beats every ONNX level on every dataset: compared to the original benchmark, the ONNX ratios dropped to roughly 1.2x because exported graphs are static and do not inherit the torch runtime improvements that made the fresh fp32 baseline much faster. The original recommendation of ONNX for short texts on GPU no longer holds: even on the short-text stsb dataset, plain fp16 (2.80x) outperforms onnx-O4 (2.57x), with the Flash Attention backends at 3.9x. The one niche ONNX keeps on GPU is latency-critical workloads pinned to small batches of short texts, where onnx-O4 still leads at the smallest measured batch sizes. Note that checkpoints stored in half precision (e.g. mxbai-embed-large-v1) load in that dtype by default since transformers v5, so they already run at fp16 speed without any configuration. The benchmark forces true fp32 for its baseline.
-
- For CPU, ONNX is stronger than OpenVINO on the short-text stsb dataset (1.35x against 1.24x), while int8 quantization pays off most: 3.23x for onnx-qint8 and 5.29x for openvino-qint8 overall, at a quality cost of less than half a percent. For longer texts, ONNX and OpenVINO can perform slightly worse than PyTorch, so we recommend testing the different backends with your specific model and data to find the best one for your use case. Half precision must be avoided on CPU altogether: torch-fp16 and torch-bf16 run largely emulated there and collapse to roughly 0.01x. + +I measured GPU throughput on an RTX 3090 and CPU throughput on an i7-13700K. Each speedup compares a backend with the matching model and workload in PyTorch FP32, and the bars summarize these ratios across the tested combinations. The whiskers show variation between combinations, rather than confidence intervals. The GPU llama.cpp measurements use their own matching FP32 baseline. + +**Datasets:** the workloads range from short sentences to long reviews: + +- `sentence-transformers/stsb `_: 38.9 characters on average (SD=13.9) + +- `sentence-transformers/natural-questions `_: answers only, 619.6 characters on average (SD=345.3) + +- `stanfordnlp/imdb `_: texts repeated 4 times, 9589.3 characters on average (SD=633.4) + +**Models:** + +- `sentence-transformers/all-MiniLM-L6-v2 `_: 22.7M parameters. + +- `BAAI/bge-base-en-v1.5 `_: 109M parameters. + +- `mixedbread-ai/mxbai-embed-large-v1 `_: 335M parameters. + +- `BAAI/bge-m3 `_: 567M parameters (GPU only). + +The GPU Sentence Transformers and ONNX tests use 2,000 samples per dataset. The CPU tests use 1,000 samples for MiniLM and BGE-base and 512 for mxbai-large. Mxbai-large is tested only on short sentences and NQ answers, giving eight CPU model and workload combinations in total. + +The throughput comparison uses each backend's best tested batch size, with 8 or 20 threads selected on CPU. After warmup, CPU timings use either the median of five passes or the mean of two passes. The CPU ``torch-fp16`` and ``torch-bf16`` results come from smaller checks on 128 samples, using the FP32-selected settings and two timed passes against matching FP32 controls. + +The ranking changes on Hugging Face Jobs ``cpu-upgrade`` instances, where llama.cpp leads instead of OpenVINO INT8. The cloud figure covers six model and workload combinations, with speedups relative to PyTorch FP32 on that CPU: + +.. image:: ../../img/backends_benchmark_cpu_cloud.png + :alt: Backend speedups on Hugging Face Jobs cpu-upgrade instances + :width: 75% + +For quality, each configuration's task scores are compared with PyTorch FP32 and the resulting ratios are averaged. The evaluation covers both sentence similarity and retrieval to capture different uses of the embeddings: + +- **Semantic Textual Similarity:** Spearman rank correlation based on cosine similarity on the `sentence-transformers/stsb `_ test set, computed via the EmbeddingSimilarityEvaluator. + +- **Information Retrieval:** NDCG@10 based on cosine similarity on the entire `NanoBEIR `_ collection of datasets, computed via the InformationRetrievalEvaluator. + +The CPU quality bars cover both tasks for MiniLM, BGE-base and mxbai-large, giving six ratios per backend. Quality evaluations use matching model artifacts on CPU or GPU, with OpenVINO and ONNX INT8 evaluated on CPU. The CPU ``torch-fp16`` and ``torch-bf16`` quality bars use GPU scores relative to matching GPU FP32 scores. + +For llama.cpp in the GPU figure, quality is evaluated on MiniLM and BGE-base, while throughput also covers mxbai-large and BGE-M3. BGE-M3 Q4 is excluded because its embeddings failed the agreement check against FP32. + +**Backends:** + +- ``torch-fp32``: PyTorch with float32 precision (default). + +- ``torch-fp16``: PyTorch with float16 precision, via ``model_kwargs={"torch_dtype": "float16"}``. + +- ``torch-bf16``: PyTorch with bfloat16 precision, via ``model_kwargs={"torch_dtype": "bfloat16"}``. + +- ``torch-fp16-fa2``: PyTorch with float16 precision and FlashAttention-2 with automatic input unpadding, via ``model_kwargs={"torch_dtype": "float16", "attn_implementation": "flash_attention_2"}``. + +- ``torch-bf16-fa2``: the same with bfloat16 precision. + +- ``onnx``: ONNX with float32 precision, via ``backend="onnx"``. + +- ``onnx-O1``: ONNX with float32 precision and O1 optimization, via ``export_optimized_onnx_model(..., optimization_config="O1", ...)`` and ``backend="onnx"``. + +- ``onnx-O2``: ONNX with float32 precision and O2 optimization, via ``export_optimized_onnx_model(..., optimization_config="O2", ...)`` and ``backend="onnx"``. + +- ``onnx-O3``: ONNX with float32 precision and O3 optimization, via ``export_optimized_onnx_model(..., optimization_config="O3", ...)`` and ``backend="onnx"``. + +- ``onnx-O4``: ONNX with float16 precision and O4 optimization, via ``export_optimized_onnx_model(..., optimization_config="O4", ...)`` and ``backend="onnx"``. + +- ``onnx-qint8``: ONNX quantized to int8 with "avx512_vnni", via ``export_dynamic_quantized_onnx_model(..., quantization_config="avx512_vnni", ...)`` and ``backend="onnx"``. The different quantization configurations resulted in roughly equivalent speedups. + +- ``openvino``: OpenVINO, via ``backend="openvino"``. + +- ``openvino-qint8``: OpenVINO quantized to int8 via ``export_static_quantized_openvino_model(..., quantization_config=OVQuantizationConfig(), ...)`` and ``backend="openvino"``. + +- ``llamacpp-*``: native llama.cpp with GGUF models in F16, BF16, Q8_0 or Q4_K_M format. + +.. raw:: html
+.. tab:: GPU -.. image:: ../../img/backends_benchmark_gpu.png - :alt: Benchmark for GPUs - :width: 45% + Half precision provides a substantial speedup on these models: plain FP16 reaches a median 2.92x the throughput of FP32. Combining it with Flash Attention 2 and input unpadding raises that to 3.87x, with BF16 performing similarly at 3.84x. This makes half precision with Flash Attention a useful starting point when your model supports it. -.. image:: ../../img/backends_benchmark_cpu.png - :alt: Benchmark for CPUs - :width: 45% + .. image:: ../../img/backends_benchmark_gpu.png + :alt: Benchmark for GPUs + :width: 75% + +.. tab:: CPU + + OpenVINO INT8 performs best on the local i7-13700K shown here, but llama.cpp leads on the cloud instance I tested. It's worth testing backends on your deployment hardware, with inputs representative of your application. + + .. image:: ../../img/backends_benchmark_cpu.png + :alt: Benchmark for CPUs + :width: 75% + +.. _llama-cpp-gpu-comparison: + +Sentence Transformers performed best on the smaller GPU models I tested, while llama.cpp became competitive on 8B models. CPU rankings differed between my local machine and the cloud instance, so benchmark on your deployment hardware. + +.. raw:: html + +
+ Compare with llama.cpp + + +The aggregate results above cover models up to BGE-M3. To explore how the comparison changes with larger models and different input lengths, I also compared Sentence Transformers with native llama.cpp on models up to Qwen3-Embedding-8B. These measurements use an RTX 3090 with 24 GB of VRAM under WSL2, with batch sizes tuned for each backend. + +The charts show median throughput, with whiskers indicating the interquartile range across repeated measurements. For Sentence Transformers, they compare the default unpadded FA2 configuration with padding enabled. Unpadding is automatic for text-only inputs when the Flash Attention implementation and model architecture support it. For llama.cpp, the GPU-table bars move the input embedding table from its default placement in CPU memory to CUDA. + +.. tab:: Short sentences + + .. image:: ../../img/llamacpp_gpu_stsb.svg + :alt: Sentence Transformers and native llama.cpp throughput on short sentences + :width: 100% + +.. tab:: NQ answers + + .. image:: ../../img/llamacpp_gpu_nq.svg + :alt: Sentence Transformers and native llama.cpp throughput on nq answers + :width: 100% + +.. tab:: Long reviews + + .. image:: ../../img/llamacpp_gpu_imdb.svg + :alt: Sentence Transformers and native llama.cpp throughput on long reviews + :width: 100% + +Sentence Transformers with BF16, Flash Attention 2 and unpadding leads on the small models, running 2.6 to 6.7 times faster than the fastest tested llama.cpp configuration on MiniLM and 2.4 to 3.0 times faster on BGE-base. The gap is much smaller on Qwen3-Embedding-4B, where the same configuration leads by about 5 to 11 percent. + +On Qwen3-Embedding-8B, the comparison shifts in favor of llama.cpp: Q8_0 with default embedding-table placement is about 12%, 2% and 12% faster than Sentence Transformers with BF16, FA2 and unpadding on short sentences, NQ answers and long reviews, respectively. More aggressive quantization does not help here, as Q4 is slower than Q8 on all three workloads. These results make llama.cpp worth considering for larger models, particularly when quantization helps them fit in memory, although models above 8B were not tested. + +.. note:: + + These results measure throughput with tuned batches, so the best configuration for single-request latency may differ. The comparison also checks embedding agreement rather than fully evaluating retrieval quality. Test both speed and task quality on representative inputs before choosing a backend. + +Timings for Sentence Transformers and native llama.cpp include tokenization, model execution, pooling, normalization and returning embeddings to CPU memory. The benchmark calls llama.cpp directly, so its timings do not include HTTP transport. + +To compare each model on the same workload, both backends use identical inputs and matching truncation limits: 256 tokens for MiniLM, 384 for MPNet, 512 for BGE-base and mxbai-large, 8,192 for BGE-M3, 32,768 for Qwen 0.6B and 40,960 for Qwen 4B and 8B. The workloads contain 2,000 short sentences, 1,000 NQ answers and 256 long reviews (repeated text). Qwen 4B and 8B use prefixes of 1,024, 256 and 64 texts. + +I checked embedding agreement on 552 texts against FP32, using BF16 as the reference for 8B. Configurations that fail this check are marked as withheld, including BGE-M3 Q4. MPNet has no supported FA2 or llama.cpp implementation in the tested versions, so those configurations are marked as unsupported. + +Sentence Transformers uses the same checkpoints as the llama.cpp GGUF conversions, which are tested in F16, BF16, Q8_0 and Q4_K_M. Decoder models use ``model[0].config.use_cache = False`` to avoid retaining a generation cache during embedding inference. The llama.cpp GPU-table variants reuse the default F16 token budgets. + +The software versions are Sentence Transformers 6.1.0.dev0 (084d9f7183b7), PyTorch 2.11.0+cu128, Transformers 5.14.1 and kernels 0.15.2 with kernels-community/flash-attn2. llama.cpp uses revision 4d91760, built with CUDA. + +.. raw:: html + +
+
Recommendations ^^^^^^^^^^^^^^^ @@ -680,7 +713,7 @@ Based on the benchmarks, this flowchart should help you decide which backend to }}%% graph TD A("What is your hardware?") -->|GPU| B("Does your model support
Flash Attention?") - A -->|CPU| C("Is a 0.4% accuracy loss
acceptable?") + A -->|CPU| C("Is a small accuracy loss
acceptable?") B -->|yes| K["float16 + Flash Attention"] B -->|no| F[float16] C -->|yes| G[openvino-qint8] diff --git a/docs/sentence_transformer/usage/usage.rst b/docs/sentence_transformer/usage/usage.rst index 22fb8eb96..b05261fcc 100644 --- a/docs/sentence_transformer/usage/usage.rst +++ b/docs/sentence_transformer/usage/usage.rst @@ -18,6 +18,7 @@ Once you have `installed <../../installation.html>`_ Sentence Transformers, you 3. :meth:`SentenceTransformer.encode_query ` 4. :meth:`SentenceTransformer.encode_document ` 5. :meth:`SentenceTransformer.similarity ` + 6. :doc:`Speeding up Inference ` :: @@ -153,5 +154,3 @@ These methods accept all the same input types as :meth:`~sentence_transformers.s ../../../examples/sentence_transformer/applications/embedding-quantization/README custom_models mteb_evaluation - efficiency - diff --git a/docs/sparse_encoder/usage/usage.rst b/docs/sparse_encoder/usage/usage.rst index 2d0733c8f..e497d702b 100644 --- a/docs/sparse_encoder/usage/usage.rst +++ b/docs/sparse_encoder/usage/usage.rst @@ -16,6 +16,7 @@ Once you have `installed <../../installation.html>`_ Sentence Transformers, you 2. :meth:`SparseEncoder.encode ` 3. :meth:`SparseEncoder.similarity ` 4. :meth:`SparseEncoder.sparsity ` + 5. :doc:`Speeding up Inference ` :: @@ -85,5 +86,3 @@ You can inspect or set the available prompts via the ``prompts`` and ``default_p ../../../examples/sparse_encoder/applications/semantic_search/README ../../../examples/sparse_encoder/applications/retrieve_rerank/README ../../../examples/sparse_encoder/evaluation/README - efficiency - diff --git a/index.rst b/index.rst index dc0260264..0271cbf2b 100644 --- a/index.rst +++ b/index.rst @@ -270,23 +270,23 @@ Consider reading one of the following sections to answer the related questions: * Embedding Models: * How to **use** Sentence Transformer models? `Sentence Transformers > Usage `_ * What Sentence Transformer **models** can I use? `Sentence Transformers > Pretrained Models `_ - * How do I make Sentence Transformer models **faster**? `Sentence Transformers > Usage > Speeding up Inference `_ + * How do I make Sentence Transformer models **faster**? `Sentence Transformers > Speeding up Inference `_ * How do I **train/finetune** a Sentence Transformer model? `Sentence Transformers > Training Overview `_ * Reranker Models: * How to **use** Cross Encoder models? `Cross Encoder > Usage `_ * What Cross Encoder **models** can I use? `Cross Encoder > Pretrained Models `_ - * How do I make Cross Encoder models **faster**? `Cross Encoder > Usage > Speeding up Inference `_ + * How do I make Cross Encoder models **faster**? `Cross Encoder > Speeding up Inference `_ * How do I **train/finetune** a Cross Encoder model? `Cross Encoder > Training Overview `_ * Sparse Encoder Models: * How to **use** Sparse Encoder models? `Sparse Encoder > Usage `_ * What Sparse Encoder **models** can I use? `Sparse Encoder > Pretrained Models `_ - * How do I make Sparse Encoder models **faster**? `Sparse Encoder > Usage > Speeding up Inference `_ + * How do I make Sparse Encoder models **faster**? `Sparse Encoder > Speeding up Inference `_ * How do I **train/finetune** a Sparse Encoder model? `Sparse Encoder > Training Overview `_ * How do I **integrate** Sparse Encoder models with search engines? `Sparse Encoder > Vector Database Integration `_ * Multi-Vector Encoder Models: * How to **use** Multi-Vector Encoder models? `Multi-Vector Encoder > Usage `_ * What Multi-Vector Encoder **models** can I use? `Multi-Vector Encoder > Pretrained Models `_ - * How do I make Multi-Vector Encoder models **faster**? `Multi-Vector Encoder > Usage > Speeding up Inference `_ + * How do I make Multi-Vector Encoder models **faster**? `Multi-Vector Encoder > Speeding up Inference `_ * How do I **train/finetune** a Multi-Vector Encoder model? `Multi-Vector Encoder > Training Overview `_ Companion Blog Posts @@ -385,6 +385,7 @@ If you use the code for `data augmentation Usage > Speeding up Inference `_ - - `Cross Encoder > Usage > Speeding up Inference `_ + - `Sentence Transformer > Speeding up Inference `_ + - `Cross Encoder > Speeding up Inference `_ Args: model (SentenceTransformer | SparseEncoder | CrossEncoder | MultiVectorEncoder): The SentenceTransformer, diff --git a/sentence_transformers/backend/quantize.py b/sentence_transformers/backend/quantize.py index c4ed71d16..bc83e4138 100644 --- a/sentence_transformers/backend/quantize.py +++ b/sentence_transformers/backend/quantize.py @@ -39,8 +39,8 @@ def export_dynamic_quantized_onnx_model( See the following pages for more information & benchmarks: - - `Sentence Transformer > Usage > Speeding up Inference `_ - - `Cross Encoder > Usage > Speeding up Inference `_ + - `Sentence Transformer > Speeding up Inference `_ + - `Cross Encoder > Speeding up Inference `_ Args: model (SentenceTransformer | SparseEncoder | CrossEncoder | MultiVectorEncoder): The SentenceTransformer, @@ -128,8 +128,8 @@ def export_static_quantized_openvino_model( See the following pages for more information & benchmarks: - - `Sentence Transformer > Usage > Speeding up Inference `_ - - `Cross Encoder > Usage > Speeding up Inference `_ + - `Sentence Transformer > Speeding up Inference `_ + - `Cross Encoder > Speeding up Inference `_ Args: model (SentenceTransformer | SparseEncoder | CrossEncoder | MultiVectorEncoder): The SentenceTransformer,