A walk through of the DeltaNet family of linear attention variants

Reverse-engineering Nvidia's CUDA-checkpoint for faster cold starts

Width vs. Depth: Speculating on the Margin

The gap between open weights LLMs and closed source LLMs