Skip to content

support b300 - #875

Open
deepindeed2022 wants to merge 1 commit into
huggingface:mainfrom
deepindeed2022:dev
Open

support b300#875
deepindeed2022 wants to merge 1 commit into
huggingface:mainfrom
deepindeed2022:dev

Conversation

@deepindeed2022

Copy link
Copy Markdown

What does this PR do?

This PR adds support for running Text Embeddings Inference (TEI) on NVIDIA B300 GPUs.

The main changes focus on enabling the correct CUDA architecture targeting for the B300 platform by updating both the build environment and FlashAttention compilation configuration.

Changes included

Update Docker build configuration to support B300 CUDA compute capability.
Extend FlashAttention build settings to include B300-compatible compute capability targets.
Ensure TEI images can be built and executed correctly on B300 environments.
Keep existing GPU architecture support unchanged for backward compatibility.

Motivation

Current TEI build artifacts and dependencies are configured for existing supported GPU architectures and do not include the compute capability required by NVIDIA B300 GPUs.

Without these changes:

Docker images may not compile optimized kernels for B300.
FlashAttention kernels may fail to build or run suboptimally.
TEI deployment on B300 hardware is not fully supported.

This PR enables TEI users to build and deploy inference workloads on B300 while preserving compatibility with existing GPU targets.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant