Python AxmlParser
Fast CUDA Kernels for ResNet Inference. Using Winograd algorithm to optimize the efficiency of co...
SGLang is a fast serving framework for large language models and vision language models.
PTX ISA 9.1 documentation converted to searchable markdown. Includes Claude Code skill for CUDA d...
AI-powered drum removal tool using Meta's Demucs. Drop in any song, get a drumless backing track ...
tabnotes
Benchmarking code for running quantized kernels from vLLM and other libraries
📚LeetCUDA: Modern CUDA Learn Notes with PyTorch for Beginners🐑, 200+ CUDA Kernels, Tensor Cores, ...
IQ of AI