🎯
Focusing
- Hangzhou
Pinned Loading
-
xlite-dev/ffpa-attn
xlite-dev/ffpa-attn PublicFast and Memory-Efficient Exact Attention (BF16/FP16/FP8/FP4) for Large Headdim, 1.5x~15x speedup over PyTorch SDPA.
-
NVIDIA-NeMo/Automodel
NVIDIA-NeMo/Automodel Public🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
-
NVIDIA/cudnn-frontend
NVIDIA/cudnn-frontend PubliccuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
-
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.


