llama-cpp-python
PulseAugur coverage of llama-cpp-python — every cluster mentioning llama-cpp-python across labs, papers, and developer communities, ranked by signal.
6 day(s) with sentiment data
-
Microsoft releases VibeVoice-ASR-BitNet for real-time CPU inference
Microsoft has released VibeVoice-ASR-BitNet, a highly compressed automatic speech recognition model designed for real-time inference on edge CPUs without requiring a GPU. This model achieves significant speedups over ex…
-
Qwopus3.6-27B-Fusion-GGUF model available for local AI inference
The KyleHessling1/Qwopus3.6-27B-Fusion-GGUF model is now available on Hugging Face, offering users a new option for local AI inference. The model is compatible with various popular libraries and applications, including …
-
Hugging Face hosts Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF with integration guides
A specific version of the Qwen model, named "DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF", has been made available on Hugging Face. The model is presented with detailed instructions …
-
Mac kernel panic caused by LLM model switching bug
A developer encountered a severe issue where switching between two large language models on their Mac caused a kernel panic, rebooting the entire system. The problem stemmed from the memory management of the llama.cpp P…
-
Unsloth releases two new optimized LLMs for local deployment
Unsloth has released two new models, unsloth/Laguna-S-2.1-GGUF and unsloth/Ornith-1.0-35B-GGUF, optimized for efficient local deployment. The models are compatible with various popular inference libraries and applicatio…
-
Qwen3.6-27B-Fable-Fusion Model Now Usable with Llama.cpp, vLLM, Ollama
The DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF model is now available for use with various popular libraries and applications. Instructions are provided for integrating it with tools…
-
unsloth/inkling-GGUF model now available across multiple AI platforms
The unsloth/inkling-GGUF model is now available for use with various libraries and inference providers, including llama-cpp-python, llama.cpp, vLLM, and Ollama. Instructions are provided for integrating the model into G…
-
AngelSlim/Hy3-GGUF model now supports multiple AI tools and libraries
The AngelSlim/Hy3-GGUF model is now available for use with various popular AI tools and libraries. Instructions are provided for integrating it with llama-cpp-python, llama.cpp, vLLM, Ollama, Unsloth Studio, and Pi. The…
-
New Qwen Model Version Released on Hugging Face with Broad Integration Support
A new version of the Qwen model, specifically LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V6-GGUF, has been released on Hugging Face. The model is designed for use with various libraries and inference provider…
-
New uncensored Qwen3.6-35B model released on Hugging Face
A new, uncensored version of the Qwen3.6-35B model, named LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V5-GGUF, has been released on Hugging Face. This model is designed to be compatible with various inference …
-
New Qwen3.6-35B model released on Hugging Face with broad integration support
A new model, LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V3-GGUF, has been released on Hugging Face, offering users instructions for integration with various libraries and applications. The model is compatible…
-
Empero AI releases Qwythos-9B-v2, fixing looping behavior and preserving reasoning
Empero AI has released Qwythos-9B-v2, an updated version of its Qwythos model. This new iteration addresses a looping behavior issue present in the previous version, reducing it from 6.7% to 0% without compromising reas…
-
Unsloth releases DeepSeek-V4-Flash model in GGUF format for local use
Unsloth has released the DeepSeek-V4-Flash model in GGUF format, making it available for local deployment. The model can be integrated with various tools and libraries, including llama-cpp-python, llama.cpp, LM Studio, …
-
Prism ML releases Bonsai 27B models in 1-bit and ternary formats
Prism ML has released two versions of its Bonsai 27B model: one utilizing 1-bit binary weights and another with ternary weights. The 1-bit Bonsai model, optimized for mobile devices like the iPhone 17 Pro Max, boasts a …
-
MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF model released with integration guides
A new model named MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF has been released on Hugging Face. The model is available in GGUF format and instructions are provided for its use with various libraries and applications, …
-
Poolside releases Laguna S 2.1 with 1M context and agentic capabilities · 4 sources tracked
Poolside has released Laguna S 2.1, a 118B parameter Mixture-of-Experts model with 8B activated parameters per token. This model is designed for agentic coding and long-horizon tasks, featuring a 1M token context window…
-
InternScience/Agents-A1-Q4_K_M-GGUF model integrated with multiple AI tools
The InternScience/Agents-A1-Q4_K_M-GGUF model is now available for use with various popular AI tools and libraries. Instructions are provided for integrating it with llama-cpp-python, llama.cpp, LM Studio, Jan, Ollama, …
-
bottlecapai releases multimodal Qwen3.6-27B model on Hugging Face
The bottlecapai/ThinkingCap-Qwen3.6-27B model, based on Qwen3.6-27B, is now available on Hugging Face. It offers multimodal capabilities, allowing users to process both text and images. The model can be integrated with …
-
MaralGPT/MaralGPT-Mythos-9B-2606-GGUF model now available for integration
The MaralGPT/MaralGPT-Mythos-9B-2606-GGUF model is now available for use with various popular libraries and inference providers. Instructions are provided for integrating the model with tools such as Transformers, llama…
-
Ornith-1.0-9B-MTP-GGUF model now supports multiple AI tools and libraries
The protoLabsAI/Ornith-1.0-9B-MTP-GGUF model is now available for use with various popular AI tools and libraries. Instructions are provided for integrating it with llama-cpp-python, llama.cpp, vLLM, Ollama, and Unsloth…