Ant Group today announced the release of Ling-3.0-Flash, a next-generation native hybrid-reasoning foundational model engineered specifically for production-grade AI agent workflows. Designed to deliver rapid response capabilities, it serves as a high-speed execution node that offers a superior balance of intelligence density and cost-efficiency.
Featuring 124B total parameters with only 5.1B active parameters per token, Ling-3.0-Flash achieves remarkable performance despite its streamlined footprint. It matches or surpasses industry-leading models with two to three times its parameter scale across core benchmarks, including foundational reasoning, instruction following, and long-context processing.
Architectural Innovation for Efficiency
Ling-3.0-Flash moves away from the traditional approach of simply scaling parameter counts. Instead, it is built from the ground up with a native hybrid-linear attention architecture. By alternating KDA (Kimi Delta Attention) and MLA layers at a 5:1 ratio, the model optimally balances long-context efficiency with robust state memory.
Key architectural advancements include:
- Upgraded KDA: Evolving from the previous Lightning Attention, KDA introduces fine-grained diagonal gating in Delta Rule state updates, allowing the model to retain critical information more precisely when processing lengthy documents and extensive codebases.
- Optimized Mixture-of-Expert (MoE) Compute: The expert activation ratio per token has been compressed from 1/32 in the previous generation to 1/64, yielding a significantly higher "efficiency leverage."
- Extended Context Window: The model natively supports a 256K context window and can seamlessly scale to 1M tokens.
Purpose-Built for Agent Workflows
Rather than aiming to replace ultra-large, general-purpose reasoning models, Ling-3.0-Flash is designed to complete the "planning-execution separation" paradigm in AI workflows. It serves as a cost-controllable, fast, and highly stable execution node, delegating deep planning and high-frequency execution to specialized models.
To support this, Ling-3.0-Flash has been deeply refined for real-world agent scenarios, expanding its training to over 10,000 interactive environments. It features enhanced self-correction and long-horizon planning mechanisms, enabling autonomous, end-to-end delivery in complex tasks such as coding, task decomposition, and deep multi-source research. This resolves common issues of deviation or context loss in traditional models during large-scale operations.
Engineering for Speed and Stability
To ensure fast and reliable agent performance, Ant Group has paired Ling-3.0-Flash with a supporting engineering and collaboration architecture:
- Reduced Latency: A cluster-level hierarchical caching system eliminates redundant computations in long conversations and multi-turn interactions, reducing Time-to-First-Token (TTFT) for long inputs by 60% to over 80%.
- Enhanced Stability: An upgraded multi-agent collaboration architecture enables different agents to divide labor and cross-validate outputs, significantly reducing the risk of misjudgments by a single model and providing robust support for high-frequency online services.
Ling-3.0-Flash is now available on OpenRouter and Vercel AI Gateway, offering a free API through August 3, 2026. Following this limited-time free access period, the model weights will be open-sourced to support further development and innovation within the global AI community.
Developers are encouraged to integrate Ling-3.0-Flash into their coding, search, research, and tool-use workflows to experience its high-speed execution and stable tool-calling capabilities.
About Ant Group
Ant Group is a global digital technology provider and the operator of Alipay, a leading internet services platform in China, connecting over one billion users to more than 10,000 types of consumer services from partners. Through innovative products and solutions powered by AI, blockchain and other technologies, Ant Group supports partners across industries to thrive through digital transformation in an ecosystem for inclusive and sustainable development.

Ling-3.0-Flash delivers strong performance across multiple core benchmarks.
-
Ant Group Unveils Ling-3.0-Flash Delivering Top-Tier Performance at a Fraction of the Parameter ScalAnt Group today announced the release of Ling-3.0-Flash, a next-generation native hybrid-reasoning foundational model engineered specifically for prod2026-07-28
-
Ant International’s Alipay+ Adds New Bank Partners Amid Cross-border Mobile Payment Boom in Asia PacAlipay+, Ant International's unified wallet gateway, is adding more bank partners to its network of over 50 leading digital wallets and financial inst2026-07-28
-
芯和半导体携手联想集团在DAC 2026现场发布EDA Agent最新研发成果国产EDA率先落地AI智能体实战应用 【美国长滩讯】2026年7月26日,全球规模最大的电子设计自动化(EDA)行业盛会——DAC 2026在美国加利福尼亚州长滩市开幕。本2026-07-28
-
Uptime Institute在美伊商业峰会上与尼尼微省达成合作,助力伊拉克数字基础设施建设这项标志性合作协议在伊拉克总理Ali Al-Zaidi历史性访美期间签署,将引领伊拉克迈入高韧性世界一流数字基础设施的新时代 Uptime Institute今日宣布,与伊拉克尼2026-07-28
-
Tecnotree斩获价值880万美元的新数字化转型合同,加速拉美市场增长全球领先的AI原生数字业务支持系统(BSS)及数字平台解决方案提供商Tecnotree今日宣布,已在拉丁美洲地区签订总价值达880万美元的全新数字化转型合同。此举进一步巩2026-07-28
-
AMD股价暴跌17%创近9年之最,苏姿丰紧急回应:AI增速远超想象
-
Ledger 中国销售渠道说明:广州馨潇贸易有限公司官方直营渠道公示
-
江苏省脑机接口产业联盟在宁成立,麦澜德分享前沿成果
-
艾芬达入选国家知识产权强国建设示范创建对象:二十载长期主义,兑现每一份用户价值
-
Esentia宣布成功完成2033年到期的6.125%优先票据和2038年到期的6.500%优先票据的定价
-
中荷人寿北京分公司成功举办中荷创享家品牌发布暨协同发展启航仪式
-
华为系具身智能公司具脑磐石完成新一轮融资:对标JEPA,押注类脑智能的认知世界模型
-
北京暑假补习班有哪些?家长首推一对一权威机构金博升学
-
上海高新技术企业代理机构深度访谈与推荐
-
2026 雷瓦亮相京东 MALL ,匠心筑造专业造型新标杆
