推出 Kimi K3:开放前沿智能
🔹 2.8 万亿参数,100 万上下文,原生多模态
🔹 Kimi Delta Attention 在百万 token 场景下实现最高 6.3 倍解码加速
🔹 注意力残差带来约 25% 训练效率提升,额外成本低于 2%
🔹 专为长周期 Agentic 编程和自适应工作流打造
Kimi K3 现已登陆:
Kimi.com
、Kimi Work、Kimi Code 和 Kimi API
开放权重预计 2026 年 7 月 27 日
🔗 API:platform.kimi.ai
🔗 技术博客:kimi.com/blog/kimi-k3
K3 是基于 Kimi Delta Attention(KDA)和 Attention Residuals(AttnRes)打造的,这两个架构上的新设计主要为了优化信息在序列长度和模型深度上的流动方式。
同时,我们也放大了 MoE(混合专家)的稀疏度,搭配 Stable LatentMoE 框架后,每轮激活的专家数从原本的几百个变成了 896 个中的 16 个。
结合优化后的训练方案和数据配方,这些结构上的改进让整体缩放效率相比 K2 提升了大约 2.5 倍,让模型能把算力更高效地转化为智能。
内部知识工作台
除了公开基准外,Kimi K3 Max 在我们内部测试中也持续展现出提升。这些内部基准都基于真实用户代理工作流中的常见模式和挑战构建。
它在线实验基准得75.5分,DECK-Bench得73.5分,金融基准得62.6分,全面超越Claude Opus 4.8 (max)和GPT-5.5 (xhigh)。
这些成绩反映出Kimi K3在代理式知识工作能力上有明显进步,在真实场景中能提供更强大、更可靠的性能。
自我进化:AttnRes 内核优化
基于 FLA Triton AttnRes 的生产级规模(96层、8192维模型、8192 token),目标是在不改变数值结果的前提下最大化训练速度。
经过连续15小时的迭代,K3 设计出全新的两阶段内核算法,在保持数值精度的同时融合了内核参数,将正向+反向传播时间从 283.6 ms 压缩至 114.4 ms。
K3 与 Fable-5(支持潜在回退机制)达到了相近性能,但 K3 每次迭代的优化速度更快。
Kimi K3把强3D推理、编程和视觉能力结合到一起,可以把概念、图像和视频直接变成可交互的沉浸式体验。
通过代码和实时截图之间的无缝迭代,Kimi K3实现了真正的“视觉闭环”。
查看英文原文
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on
Kimi.com
, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API:
platform.kimi.ai
🔗 Tech blog:
kimi.com/blog/kimi-k3
K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), two architectural updates designed to improve how information flows across sequence length and model depth.
We have also scaled up Mixture of Experts (MoE) sparsity, effectively activating 16 out of 896 experts when paired with a Stable LatentMoE framework.
Together with refined training and data recipes, these structural changes yield an approximate 2.5× improvement in overall scaling efficiency compared to K2, allowing the model to convert compute into intelligence more effectively.
Internal knowledge work bench
Beyond public benchmarks, Kimi K3 Max also shows consistent gains on our internal benchmarks, which are built from recurring patterns and challenges in real-world user-agent workflows.
It scores 75.5 on Online Exp Bench, 73.5 on DECK-Bench, and 62.6 on Finance-Bench, outperforming Claude Opus 4.8 (max) and GPT-5.5 (xhigh) across all three.
These results reflect broad improvements in Kimi K3's agentic knowledge work capabilities, enabling more capable and reliable performance in real-world use cases.
Self-evolving: AttnRes Kernel Optimization
Given FLA Triton AttnRes at production scale (96 layers, 8192-dim model, 8192 tokens), the goal was to maximize training-side speed without changing numerics.
Over 15 hours of nonstop iteration, K3 designed a novel two-phase kernel algorithm, fused kernels while preserving numerics, and reduced forward+backward time from 283.6 ms to 114.4 ms.
K3 and Fable-5 (with potential fallback) reached similar performance, but K3 improved faster per iteration.
Kimi K3 combines strong 3D reasoning, coding, and vision capabilities to turn concepts, images, and videos into fully playable interactive experiences.
Kimi K3 achieves true "vision in the loop" by seamlessly iterating between code and live screenshots




























































