Beechcraft King Air 350
PulseAugur coverage of Beechcraft King Air 350 — every cluster mentioning Beechcraft King Air 350 across labs, papers, and developer communities, ranked by signal.
8 day(s) with sentiment data
Nvidia B300 GPU smuggling to China continues despite charges
Despite Taiwanese authorities charging nine individuals, including Nvidia employees, for smuggling B300 GPUs to China, the evidence suggests that 74 servers were successfully delivered. This indicates that the smuggling operations may continue or have been difficult to fully disrupt, highlighting the persistent demand and challenges in controlling the flow of advanced AI hardware.
Nvidia's Qwen3.8-2.4T-A95B MoE model will see rapid adoption for long-context tasks
Nvidia's release of the Qwen3.8-2.4T-A95B MoE model, featuring 1 million token context length and optimization for their GB200/B300 hardware, suggests a strategic push towards enabling advanced, long-context AI applications. We hypothesize that this model will be quickly adopted by researchers and developers focused on complex tasks requiring extensive memory, such as detailed document analysis or extended conversational AI.
New LLM efficiency metrics like tok/s/MW will drive hardware optimization
The introduction of novel LLM efficiency metrics, such as tokens per second per megawatt (tok/s/MW), signifies a growing industry focus on energy consumption alongside performance. This trend is likely to drive further hardware innovation and optimization specifically targeting power efficiency, potentially influencing future chip designs and deployment strategies for large-scale AI models.
-
AMD MI355x hardware outperforms B300 on token efficiency for AgentX · 2 sources tracked
A new submission for AMD's MI355x hardware has demonstrated superior performance in terms of total tokens per Total Cost of Ownership (TCO) compared to the B300, particularly at lower interactivity ranges within the Age…
-
AMD's MI355X shows rapid performance gains on AI workloads · 2 sources tracked
SemiAnalysis has noted significant performance improvements for AMD's MI355X accelerator, particularly on agentic workloads. These enhancements, achieved in just a few weeks, have narrowed the performance gap between th…
-
NVIDIA releases quantized DeepSeek and Qwen LLMs for Blackwell hardware
NVIDIA has released quantized versions of two large language models, DeepSeek-V4-Pro-0813 and Qwen3.8-2.4T-A95B, utilizing their NVFP4 quantization method. The DeepSeek model, with 1.65 trillion parameters, employs Hybr…
-
Nvidia Groq 3 LPX ships, OpenAI probed, Uber fined, and Nvidia servers smuggled
Nvidia has announced its Groq 3 LPX inference accelerator is now in full production, designed for high-throughput AI tasks and integrated into the Vera Rubin platform. Separately, Taiwanese authorities have indicted nin…
-
Taiwan charges nine for smuggling AI servers to China · 3 sources tracked
Taiwanese prosecutors have charged nine individuals, including employees from Nvidia and Super Micro, for allegedly smuggling high-end AI servers to mainland China. The servers, identified as B300 GPUs, are subject to U…
-
Kimi K3 LLM hosted with 8 B300 GPUs achieves 92 tokens/sec
A user detailed their experience hosting the Kimi K3 large language model, which has 2.8 trillion parameters, using eight B300 GPUs. The setup achieved a throughput of 92 tokens per second with a time-to-first-token of …
-
LLM efficiency measured by tokens per Big Mac calorie
SemiAnalysis has introduced a novel metric for evaluating large language model (LLM) inference efficiency: tokens per second per megawatt (tok/s/MW). This metric allows for a comparison of energy consumption against out…
-
Top GPU Neoclouds 2026: CoreWeave, Nebius, Lambda, Crusoe, Groq Ranked
A 2026 ranking of GPU neocloud providers highlights CoreWeave as the sole Platinum-rated option, commanding premium pricing for high-end NVIDIA GPUs. Nebius leads in publishing on-demand pricing for the latest B300 chip…
-
AI infrastructure sees soaring B300 prices, new chip architectures, and executive departures
The AI infrastructure market is experiencing significant shifts, with soaring prices for key components like the NVIDIA B300, which has tripled in value over six months. This scarcity is leading to concerns about inform…
-
AMD MI355X GPUs offer better performance per dollar for Kimi K3 model
A blog post from Wafer.ai details how they achieved better performance per dollar by running the Kimi K3 model on AMD's MI355X GPUs. Despite Kimi K3's massive 2.8T parameter size requiring significant VRAM, the MI355X, …
-
Moonshot AI releases 2.8T Kimi K3 weights, largest ever, but impractical to run
Moonshot AI has released the full 2.8 trillion parameter weights for its Kimi K3 model, making it the largest open-weight model to date. Despite the massive size and open release, running Kimi K3 is practically impossib…
-
Kimi K3 2.8T model requires advanced hardware beyond NVIDIA DGX B200
The Kimi K3 2.8T model is exceptionally large, requiring specialized hardware beyond a single NVIDIA DGX B200 system. To accommodate its size, configurations involving GB300 NVL72, B300, or MI355X systems are necessary …
-
New FP4 attention kernels boost B300 performance by 1.69x
A new set of FP4 attention kernels has been developed for the B300, offering a significant speedup of up to 1.69x compared to FA4. This advancement aims to improve the performance of local large language models.
-
vLLM Kimi Architecture Matches xAI's Cursor Composer 2.5; NVIDIA Leads AMD
SemiAnalysis reports that vLLM Kimi shares a model architecture with xAI's Cursor Composer 2.5. The analysis also indicates that NVIDIA's hardware is outperforming AMD's, with the B300 showing superior speed compared to…
-
NVIDIA, SGLang, RadixArk achieve 3.7x faster AI interactivity
SemiAnalysis highlighted the performance of NVIDIA, SGLang, and RadixArk, noting their systems achieved up to 3.7 times faster interactivity compared to the B300 benchmark. This advancement suggests significant progress…
-
AI token costs to drop by 2027 amid hardware/software gains · 4 sources tracked
SemiAnalysis reports that the cost of AI tokens is projected to decrease significantly by 2027, driven by advancements in hardware and software optimization. These improvements, such as increased throughput and efficien…
-
AI chip demand surges, driving GPU prices and sparking funding rounds
The AI chip industry is experiencing significant shifts, with major internet companies directly procuring thousands of NVIDIA B300 GPUs, bypassing traditional channels. This surge in demand is driving up prices for high…
-
FP8 with reconstruction schemes matches FP64 accuracy in HPC
A new research paper challenges the long-held belief that double-precision (FP64) hardware is essential for high-performance computing (HPC). The authors propose that using FP8 tensor cores, combined with specific recon…
-
ZFLOW AI Boosts B300 Inference with DeepSeek V4-Pro Tuning
ZFLOW AI has enhanced the inference capabilities of NVIDIA's B300 hardware by employing simulation tuning. This optimization resulted in a peak throughput of 826 tokens per second and reduced tail latency when utilizing…
-
AI GPU demand soars, B300 prices climb; older chip trade-in fails
Demand for high-performance AI GPUs like the B300 is surging, with a major East China tech company reportedly planning to purchase over 10,000 units, driving prices up significantly. Concurrently, a previous "trade-in" …