inclusionAI
PulseAugur coverage of inclusionAI — every cluster mentioning inclusionAI across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
Edge0-AI streams 35B MoE models off SSD to fit in 3GB RAM
Edge0-AI has released Edge0, an open-source streaming inference engine designed to run large Mixture-of-Experts (MoE) models on consumer hardware. The engine achieves this by memory-mapping the entire model checkpoint o…
-
InclusionAI releases Ling 3.0 Flash VL with open weights and large context · 4 sources tracked
InclusionAI has released Ling 3.0 Flash VL, a vision-language model with open weights. This model supports a context window of up to 262.1k tokens and offers competitive pricing at $0.06/$0.18 per 1 million tokens. The …
-
Small LLM comparison: Spark-X2.5-4B, Ling 3.0 Tiny, Nanbeige4.2-3B
A discussion on the r/LocalLLaMA subreddit compares three small language models: XHToken's Spark-X2.5-4B, inclusionAI's Ling 3.0 Tiny, and Nanbeige's Nanbeige4.2-3B. Users are seeking to identify which of these models, …
-
Ant Group releases finance-focused LLM with 256K context, multimodal variant noted
Ant Group has released Ling-3.0-flash-Fin, a new large language model enhanced for financial tasks. This 124B parameter model, with 5.1B activated parameters and a 256K context window, excels at end-to-end financial res…
-
AI API Digest: Qwen prices surge, Z.ai adds 1M context, AI21 & Mancer models removed
The AI API Digest for August 20, 2026, highlights significant pricing changes and model updates across various providers. Qwen's Qwen3.6 27B model saw a substantial price increase for both prompt and completion tokens, …
-
Ling 3.0 Tiny model enhances primary LLMs with efficient auxiliary tasks
The Ling 3.0 Tiny model is being utilized as an auxiliary component for larger language models like Hermes and Qwen 3.8-27B. Users have found that Ling 3.0 Tiny can efficiently handle tasks such as context compression a…
-
inclusionAI releases lightweight Ling-3.0-tiny MoE model for local deployment
inclusionAI has released Ling-3.0-tiny, a new hybrid reasoning Mixture-of-Experts (MoE) model with 7.9 billion total parameters and 1.3 billion activated parameters per token. This model is designed for efficient local …
-
Ant Group's Ling 3.0 Flash model released for local use
Ling 3.0 Flash, a 124B parameter Mixture of Experts model from Ant Group's inclusionAI, has been released with MIT license and is available on Hugging Face. This model is designed for local execution, requiring signific…
-
User demotes routine AI tasks from frontier models to save costs
A Reddit user analyzed their OpenAI API spending and discovered that routine tasks like classification, document summarization, and routing constituted the majority of their bill. They found that these common tasks did …
-
AI Labs' Open-Source Model Releases Tied to Funding Structures
Several AI labs are releasing open-source models, but their funding structures dictate their ability to continue this practice. DeepSeek, backed by a wealthy founder and a significant personal investment, maintains cont…
-
inclusionAI releases Ling-3.0-flash model with official FP8 weights
inclusionAI has released its Ling-3.0-flash model, available in both BF16 and an official FP8 version, on Hugging Face. The model boasts 127.5 billion total parameters with 5.1 billion active parameters, featuring a fin…
-
Ant Group's inclusionAI releases Ling-3.0-flash with 256K context · 2 sources tracked
inclusionAI, an Ant Group initiative, has released Ling-3.0-flash, a 124 billion parameter Mixture-of-Experts model. This model boasts a 256,000 token context window and is designed for efficient agentic tasks and long-…
-
LLM Pricing Fluctuates: NVIDIA, Qwen, and Z.ai See Changes; New Models Added · 10 sources tracked
The Token Ledger has released daily updates on LLM pricing changes throughout early August 2026. Several models saw price adjustments, including NVIDIA Nemotron 3 Super and Ultra, Qwen variants, and Z.ai's GLM 5.2, with…
-
New AI guardrails challenge reasoning necessity and boost multimodal safety
Two new research papers explore the effectiveness and adaptability of AI safety guardrails. One paper, LeanGuard, questions the necessity of complex reasoning in moderation, demonstrating that a lightweight, label-only …
-
AI Model Pricing Shifts: NVIDIA, MoonshotAI, DeepSeek Cut Costs; Z.ai Adds Long-Context Model
Several AI model providers have announced pricing adjustments and new model releases. NVIDIA's Nemotron 3 Ultra has seen a completion price drop, benefiting long-form generation workloads. MoonshotAI's Kimi K2.7 Code an…
-
LLM pricing shifts: Kimi K2.7 up, Claude 3.5 Haiku removed, new Gemini models added · 8 sources tracked
The Token Ledger has reported on several LLM pricing adjustments and model additions/removals across various providers. Notably, MoonshotAI's Kimi K2.7 Code saw a price increase for completions, while its Kimi Latest an…
-
InclusionAI releases Ring 2.5 1T model for free
InclusionAI has made its Ring 2.5 1T model freely available on the kilocode platform. This release is accompanied by a tutorial on YouTube, aimed at users interested in AI coding.
-
inclusionAI releases Vista 9B/4B GUI-grounding models
inclusionAI has released Vista 9B and Vista 4B, new vision-language models designed for GUI grounding. These models are trained using a view-consistent GRPO approach and self-verified cross-view anchoring, building upon…
-
AI API pricing sees major cuts for inclusionAI's Ring-2.6-1T
inclusionAI has significantly reduced its pricing for the Ring-2.6-1T model, cutting both prompt and completion prices by 75%. This change offers substantial cost savings for teams utilizing this model for high-volume i…
-
Model catalog sees new additions, price changes, and removals
New models have been added to the model catalog, including StepFun's Step 3.7 Flash, which offers large-context generation at a moderate cost. Anthropic has released two new Claude Opus 4.8 variants: a "Fast" version fo…