Skip to content
#

sparse-attention

Here are 83 public repositories matching this topic...

NVIDIA Sol-Attn for ComfyUI / Triton kernel on SM89 - SM121, with zero-copy MiniMax H3 nodes: memory-efficient attention, scheduled tau with graph preview, and feed-forward chunking. Measured 1.14–1.44× vs SageAttention and −37% MLP peak VRAM on H3

  • Updated Aug 13, 2026
  • Python

🦀 Decoder-only LLM built from scratch in pure Rust using Candle — no Python, no PyTorch. Gated DeltaNet + sparse attention, fine-grained MoE, native video/document understanding, long-horizon tool agents, quantization-aware training. Scales: Tiny (25M) to Large (1.3B).

  • Updated Aug 10, 2026
  • Rust

Improve this page

Add a description, image, and links to the sparse-attention topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the sparse-attention topic, visit your repo's landing page and select "manage topics."

Learn more