NVIDIA Qwen3.8-Flash-Next Achieves Over 16,000 Tokens/Second Throughput on GB300 NVL72
NVIDIA has announced that Alibaba's latest preview model, Qwen3.8-Flash-Next, is now supported on the NVIDIA GB300 NVL72 platform. The model has a total parameter scale of 176 billion, with approximately 6 billion parameters activated per token. It natively supports a context of 262,000 tokens and can be extended to 1 million tokens via YaRN, primarily targeting long-context agent applications such as intelligent programming, document processing, and tool invocation. NVIDIA stated that Qwen3.8-Flash-Next employs a mixed architecture of Gated DeltaNet (GDN) and Qwen Sparse Attention (QSA) to reduce computational and KV cache overhead in long-context scenarios. Testing shows that on the GB300 NVL72, the model achieves a single GPU throughput of over 16,000 tokens per second, with single-user throughput exceeding 200 tokens per second; it also supports inference frameworks such as SGLang, vLLM, and TensorRT-LLM.
-- Price
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

China's Core Open Source Large Model ARR Estimated at $6-8 Billion

Qualcomm Launches IMSDK 2.0 to Drive Edge AI Application Development

Apple Launches M6 Chip Mac mini Starting at $899

Thomson Reuters Develops Legal AI Model 'Thomson' with $40 Million Investment

The source on Naura and 3D DRAM memory is unavailable

Ubuntu 26.10 Releases Snapshot 3 for Testing Ahead of October Launch

Kalshi faces $500,000 daily fines as Michigan forces sports event contracts offline

30-Year Bond Yield Hits 4.079%, Raising Funding Costs for MetaPlanet's Bitcoin Purchases

Quantum Memory: The Device That Breaks Bitcoin and Replaces It

Japan’s 4% bond yield spike threatens the low-cost borrowing strategy behind corporate Bitcoin buying

Anthropic's Mea Culpa: A Complete Autopsy of Claude's Missteps

Arthur Hayes calls EUR/JPY prices crypto’s smoke alarm, but the Fed’s plumbing still shows no fire

Debate Over $300 Bitcoin Tax Exemption and Estimated Revenue Increase

Copy Trading: How Does It Work in 2026?

PL Deputy Proposes Gun Carrying Rights for Cryptocurrency Investors and Industry Executives

The Executive Who Anticipates a New Era for Cryptocurrencies: "We Are Just Getting Started"

Why GENIUS could leave digital dollars vulnerable to sudden blockchain network ‘bank runs’

Robinhood Chain Down for 14 Minutes: The Blockchain That Was Supposed to Tokenize Wall Street First Blocked Itself

Netflix Hits British Wallets with Up to 33.4% Price Increase on Plans

Cracking 1.33 Trillion Daily Tokens: B.AI Powers the “AI Grid” with Full-Stack Infrastructure to Fuel the Agentic Era

Hyperliquid vs Drift Protocol Whitepaper Comparison (2026): Technology, Tokenomics, and Trading Infrastructure

Cybercrime, Child Gambling, and Underground Banking

PostGREShell: flaw in PostgreSQL turned backup accounts into backdoors

Fomo Earns $1.2 Million Daily, Why Are Two Major Exchanges Nervous?

Stocks, Bonds, Funds: Seoul Prepares for Their Arrival on the Blockchain

A7A5: The number of transactions with the ruble stablecoin increased by 4.4 times

US Employment Surprises Threefold, Renewing Tightening Concerns... Dollar and Interest Rates Rise Together

Shen Yu: Knowledge and Action in the Age of AI

US 10-Year Treasury Yield at 4.79%, Long-Term Bond Absorption Pressure Increases



