207 WebGPU kernels put browser AI’s low-level work in the open
A new JavaScript library and versioned kernel collection targeted the operations underneath browser inference.
Local AI, faster kernels, and getting models into applications.
A new JavaScript library and versioned kernel collection targeted the operations underneath browser inference.
Late-interaction models preserve finer matching information at the cost of larger indexes.
Nunchaku integration brings another low-precision option into Diffusers, with hardware compatibility and image quality still central to adoption.
Attention profiling turns a vague performance problem into a sequence of operations that can be measured and compared.
The Transformers backend in vLLM narrows the gap between readable model code and optimized production inference.
A redesigned project added a repository type, revised tools, and broader backend coverage.
An experimental storage proposal explores sharing model resources across web origins.
A Chrome extension guide separates the model runtime from the page interface.
Modular Diffusers opens up the steps inside a generation pipeline.
Sparse activation makes parameter counts a less complete guide to runtime behavior.