pwilkin
PulseAugur coverage of pwilkin — every cluster mentioning pwilkin across labs, papers, and developer communities, ranked by signal.
-
llama-bench defaults corrected for flash attention and GPU layers
A recent build, b9437, for the llama-bench tool has corrected default settings related to flash attention and GPU layer counts. Previously, the tool hard-coded flash attention off, even on compatible hardware, and used …
-
DeepSeek V4 Flash model gains early support in llama.cpp
A pull request is in progress to add support for the DeepSeek V4 Flash model to the llama.cpp library. While currently in an early, slow, and unstable stage, the model is praised for its intelligence relative to its siz…
-
StepFun 3.5 MTP model integrated into llama.cpp
A new model called StepFun 3.5 MTP has been introduced via a pull request to the llama.cpp project. This model appears to be a successor to Gemma MTP, with its integration into llama.cpp being a key development.