support b300 - #875
Open
deepindeed2022 wants to merge 1 commit into
Open
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
This PR adds support for running Text Embeddings Inference (TEI) on NVIDIA B300 GPUs.
The main changes focus on enabling the correct CUDA architecture targeting for the B300 platform by updating both the build environment and FlashAttention compilation configuration.
Changes included
Update Docker build configuration to support B300 CUDA compute capability.
Extend FlashAttention build settings to include B300-compatible compute capability targets.
Ensure TEI images can be built and executed correctly on B300 environments.
Keep existing GPU architecture support unchanged for backward compatibility.
Motivation
Current TEI build artifacts and dependencies are configured for existing supported GPU architectures and do not include the compute capability required by NVIDIA B300 GPUs.
Without these changes:
Docker images may not compile optimized kernels for B300.
FlashAttention kernels may fail to build or run suboptimally.
TEI deployment on B300 hardware is not fully supported.
This PR enables TEI users to build and deploy inference workloads on B300 while preserving compatibility with existing GPU targets.