JEDEE AI
存档 2026-07-17

7 月 17 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 100 条。
← 返回最新 全部归档

全部情报 每小时更新 · 事件已合并同类项

提交账号

填 @用户名 或主页链接,审核通过后收录进情报站。

内容 公司
Kimi.ai@Kimi_Moonshot · 公司官方 · 12 小时前月之暗面 Kimi 官方
连环推 ×5

推出 Kimi K3:开放前沿智能

🔹 2.8 万亿参数,100 万上下文,原生多模态
🔹 Kimi Delta Attention 在百万 token 场景下实现最高 6.3 倍解码加速
🔹 注意力残差带来约 25% 训练效率提升,额外成本低于 2%
🔹 专为长周期 Agentic 编程和自适应工作流打造

Kimi K3 现已登陆:
Kimi.com
、Kimi Work、Kimi Code 和 Kimi API
开放权重预计 2026 年 7 月 27 日

🔗 API:platform.kimi.ai
🔗 技术博客:kimi.com/blog/kimi-k3

查看英文原文
Introducing Kimi K3: Open Frontier Intelligence

🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows

Kimi K3 is now live on on
Kimi.com
, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.

🔗 API:
platform.kimi.ai

🔗 Tech blog:
kimi.com/blog/kimi-k3
K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), two architectural updates designed to improve how information flows across sequence length and model depth.

We have also scaled up Mixture of Experts (MoE) sparsity, effectively activating 16 out of 896 experts when paired with a Stable LatentMoE framework.

Together with refined training and data recipes, these structural changes yield an approximate 2.5× improvement in overall scaling efficiency compared to K2, allowing the model to convert compute into intelligence more effectively.
Internal knowledge work bench

Beyond public benchmarks, Kimi K3 Max also shows consistent gains on our internal benchmarks, which are built from recurring patterns and challenges in real-world user-agent workflows.

It scores 75.5 on Online Exp Bench, 73.5 on DECK-Bench, and 62.6 on Finance-Bench, outperforming Claude Opus 4.8 (max) and GPT-5.5 (xhigh) across all three.

These results reflect broad improvements in Kimi K3's agentic knowledge work capabilities, enabling more capable and reliable performance in real-world use cases.
Self-evolving: AttnRes Kernel Optimization

Given FLA Triton AttnRes at production scale (96 layers, 8192-dim model, 8192 tokens), the goal was to maximize training-side speed without changing numerics.

Over 15 hours of nonstop iteration, K3 designed a novel two-phase kernel algorithm, fused kernels while preserving numerics, and reduced forward+backward time from 283.6 ms to 114.4 ms.

K3 and Fable-5 (with potential fallback) reached similar performance, but K3 improved faster per iteration.
Kimi K3 combines strong 3D reasoning, coding, and vision capabilities to turn concepts, images, and videos into fully playable interactive experiences.

Kimi K3 achieves true "vision in the loop" by seamlessly iterating between code and live screenshots
Kimi.ai认识 Kimi K312 小时前 · 264.2 万Lisan al GaibKimi-K3 的基准测试被泄露了13 小时前 · 22.3 万Lisan al Gaib来自 Moonshot 的官方博客:11 小时前 · 20.9 万Chubby♨️Kimi k3在设计竞技场前端排名第一。12 小时前 · 19.1 万Chubby♨️Kimi K3 可能是 DeepSeek 2.0 时刻。基准测试结果已经公布,成绩非常出众。这清楚地证明了一件事…11 小时前 · 13.4 万Chubby♨️Kimi K3基准测试的成绩绝了12 小时前 · 12.8 万Chubby♨️@ArtificialAnlys 的独立基准评测确认了Kimi K3的性能。此外,权重即将发布。我们有一个开源、…11 小时前 · 8.7 万Lisan al GaibKimi-K3 在 AAI Index 上获得 57 分,token 效率比 Opus 4.8 高出近 2 倍11 小时前 · 4.9 万Ethan Mollick提个醒:我得说在对我之前的一些学术工作做复杂统计审计时,Kimi K3 Max 搞砸了不少地方,包括滥用统计和胡…9 小时前 · 3.3 万Lisan al GaibKimi-K3在Text Arena排名第1012 小时前 · 3.3 万Lisan al GaibKimi-K3 在 GPU 优化方面吊打 Fable 513 小时前 · 2.7 万Lisan al Gaibmoonshot 成功了12 小时前 · 2.5 万Lisan al GaibKimi-K3 就是我想象中 DeepSeek-V4 怎么缩小开源/闭源差距的样子。现在该轮到 OpenAI、A…7 小时前 · 1.8 万Simon WillisonKimi K3,有点调皮还有点被动攻击性:8 小时前 · 1.7 万向阳乔木赞叹 Kimi K3 的美感,每个风格都是独立生成的HTML+CSS。14 小时前 · 1.6 万Lisan al GaibKimi-K3在CritPt基准上得到了23%的分数11 小时前 · 1.6 万karminski-牙医kimi-k3 测试速报! 前端这效果太猛了!8 小时前 · 1.5 万The Rundown AIMoonshot的Kimi K3正式上线了,这可能是今年的DeepSeek时刻。12 小时前 · 1.5 万AshutoshShrivastava天啦噜,Kimi K3 的基准测试成绩太疯狂了。13 小时前 · 1.4 万Bindu ReddyKimi K3 其实不是 Opus 这个级别的!在长文本多轮复杂 agentic loops 上就露馅了4 小时前 · 1.4 万Min ChoiKimi K3真的改变了AI模型的玩法。现在人们在创造超越ChatGPT/Claude Fable 5的各种疯狂…3 小时前 · 1.2 万Ethan MollickKimi K3 写不出好谋杀悬疑小说(其他模型也一样)。这仍然是最崎岖的前沿阵地。4 小时前 · 1.2 万小互月之暗面发布 Kimi K3:全球首个 3 万亿级开放模型 4 小时前 · 1.1 万Bindu ReddyKimi K3基准成绩过分 - 留点怀疑看13 小时前 · 1 万Lisan al Gaibkimi.com/blog/kimi-k312 小时前 · 9,836🚨 AI News | TestingCatalog🔥MOONSHOT:Moonshot AI 的 Kimi K3 在 Frontend Code Arena 排名…11 小时前 · 9,501Orange AIK3 写作能力超过 Fable54 小时前 · 9,271🚨 AI News | TestingCatalogKimi K3今天首发即登陆AI/ML API,和Claude Fable 5共享同一个API key。7 小时前 · 9,142🚨 AI News | TestingCatalogKimi K3 的基准测试在许多领域表现强劲,性能与美国领先实验室的专有模型相当。开源权重预计将于 7 月 27…11 小时前 · 8,958Lisan al GaibKimi-K3 的博客:mp.weixin.qq.com/s/V4xhEIy8x…13 小时前 · 8,947Lisan al Gaib现在得等2-3周看第三方的基准测试数据和token效率表现12 小时前 · 7,758歸藏(guizang.ai)Kimi 的多模态还是很顶的,感觉随着 Claude 越来越不重视多模态理解,Kimi 这个会很有用14 小时前 · 7,630Lisan al GaibKimi-K3在AA-Omniscience上这个体量可能还能好一点,但目前也还行。相比K2.6还是进步不小11 小时前 · 7,513歸藏(guizang.ai)官方的测试结果12 小时前 · 6,231AshutoshShrivastavaKimi 团队用 K3 搞出了大动作 🔥12 小时前 · 5,611AshutoshShrivastava看看 Kimi K3 对 Fable 5 的表现。AI 领域发展太快,很难预测谁接下来会发布什么。我自己测过 K…11 小时前 · 4,851Lisan al GaibKimi-K3在WebDev Arena排名第一12 小时前 · 4,362AshutoshShrivastavaKimi K3 来了 🔥14 小时前 · 3,998歸藏(guizang.ai)下面很多评论特别搞笑:“中国人一定是造了个时间机器,从明年得美国模型蒸馏的 K3。”4 小时前 · 2,408Orange AIKimi K3 现已上架到 Cola1 小时前 · 1,977
◔ 1220.5 万 次浏览(41 条合计)♥ 3.2 万⇄ 4,348▶ 含视频新品看原帖 ↗
Sam Altman@sama · 创始人 · 12 小时前Sam Altman,OpenAI 联合创始人兼 CEO

过去12个月我们表现不是最好的,这主要是我的责任,但我们即将迎来有史以来最好的12个月。这个团队做的工作太棒了,我想你们会被他们为你们准备的东西惊到。

我为此而高兴,原因有很多,但最主要是因为我在乎用户的成功。AI应该是为大量人带来更多自由、主动性和财富。我们想做正确的事,但不想通过恐吓让人们按我们的方式做。

查看英文原文
we did not have our best last 12 months ever, which is mostly my fault, but we are about to have our best 12 months to date. the team is doing amazing work and i think you’ll be very happy with what they’ve got cooking for you.

i am happy about this for many reasons, but mostly because i care about our users winning. AI has to be about giving lots of people more freedom, agency, and wealth. we want to do the right thing, but we do not want to scare people into doing our thing.
Sam Altman@sama · 创始人 · 11 小时前Sam Altman,OpenAI 联合创始人兼 CEO

我现在跟 ChatGPT 说话的时间比打字还多。新的语音模型真的突破了一个临界点。

查看英文原文
i talk to chatgpt more than i type to it at this point

new voice model really crossed a threshold
Gemini Notebook@Gemini_Notebook · 公司官方 · 13 小时前谷歌 AI 笔记工具 NotebookLM 官方

3年前,我们只是一个小实验,想帮你学得更快。

后来,我们把音频、视频和互动功能带进你的素材库,从只有阅读笔记的工具,变成了真正的协同研究助手。

而现在,笔记本已经成为一个完整生态:你已经可以在
@GeminiApp
里用它们了,马上在 Google Search 也能直接搜到

所以,趁着这些演进,我们也要再跟上步伐:

NotebookLM 现在是 Gemini Notebook ✨📓

你熟悉和喜欢的那个App并没有消失,只是名字换了下,来彰显我们在Google AI产品线里的角色。

而且,我们的使命一直没变:就是帮你学得更快。

感谢各位的相信——没有你们的热爱和各种需求折腾,真的走不到今天

后续还会有料(对,分区查询很快安排上📂!)所以别走开,记得蹲。

深耕细作,

Project Tailwind 团队

查看英文原文
3 years ago we started as a tiny experiment with the goal of helping you learn faster.

Since then, we grew to bring audio, video, and interactivity to your sources, transitioning from a passive workspace to your true research companion.

And now, notebooks have even become an entire ecosystem: you can already access them in the
@GeminiApp
and soon in Google Search

So, with these advancements, it’s time for us to evolve once again:

NotebookLM is now Gemini Notebook ✨📓

The same app you know and love isn’t going anywhere, we just have an updated name that reflects our role in Google's AI portfolio.

And our mission stays exactly the same: helping you learn, faster.

Thank you for believing in us— this wouldn't have been possible without your passion (and feature requests...)

Big things to come (yes, even folders📂!) so stay tuned.

Sincerely,

The Project Tailwind team
◔ 53.2 万 次浏览(2 条合计)♥ 3,044⇄ 373▶ 含视频新品看原帖 ↗
Grok@grok · 公司官方 · 13 小时前马斯克 xAI 旗下聊天机器人 Grok 官方

在 Grok Build 中使用 Railway 直接部署应用

引用 Railway @RailwayRailway is now an official plugin in @grok Build Your agents can ship apps, manage infrastructure, and troubleshoot issues right from Grok Install from the marketplace today 🚅查看被引原帖 ↗
查看英文原文
Use Railway to deploy apps directly in Grok Build
◔ 28.4 万 次浏览♥ 749⇄ 75▶ 含视频教程看原帖 ↗
Grok@grok · 公司官方 · 10 小时前马斯克 xAI 旗下聊天机器人 Grok 官方

Grok 现在支持自动化功能啦:
grok.com/automations


只要描述一次任务,设定好时间或触发器,Grok 就会自动执行并报告结果。

查看英文原文
Introducing Automations in Grok:
grok.com/automations


Describe a job once, set a schedule or a trigger, and Grok runs it and reports back.
◔ 25.7 万 次浏览♥ 1,564⇄ 168▶ 含视频新品看原帖 ↗
Sundar Pichai@sundarpichai · 创始人 · 14 小时前谷歌 CEO

很高兴看到
@Intel
在整个业务中使用 Gemini Enterprise,包括加速下一代芯片的开发!

引用 Thomas Kurian @ThomasOrTKWe are expanding our strategic partnership with @Intel to accelerate their enterprise-wide digital transformation using Gemini Enterprise and @GoogleCloud . By integrating custom, agentic AI workflows across core business functions and silicon design, Intel will drive new levels of speed, agility, and efficiency across their global operations.查看被引原帖 ↗
查看英文原文
Great to see
@Intel
using Gemini Enterprise across its business, including to speed up the development of next-gen semiconductors!
OpenAI@OpenAI · 公司官方 · 13 小时前ChatGPT 开发商官方账号
连环推 ×2

在赛车运动中,细微的差距很重要。AI 可以帮助车队找到这些差距。

OpenAI 的 Joyce Ruffell 与 @RaceTekSystems 联合创始人 @GarageGuyChase 和 @AndrewMayne 讨论赛车队如何利用 AI 将赛道数据转化为更快的决策——源于我们与 Chip Ganassi Racing 的研究合作,以及与 ChatGPT 和 Codex 建造的新工具。

查看英文原文
In racing, tiny margins matter. AI can help teams find them.

OpenAI’s Joyce Ruffell and
@RaceTekSystems
co-founder
@GarageGuyChase
discuss with
@AndrewMayne
how racing teams use AI to turn track data into faster decisions—from our research collaboration with Chip Ganassi Racing to building new tools with ChatGPT and Codex.
Listen to the OpenAI Podcast on—

Spotify

open.spotify.com/show/0zojME…

Apple

podcasts.apple.com/us/podcas…

YouTube

invidious.tiekoetter.com/watch?v=KNPjRpNt…
◔ 17.7 万 次浏览♥ 822⇄ 56▶ 含视频演示看原帖 ↗
Amjad Masad@amasad · 创始人 · 13 小时前Amjad Masad,Replit 创始人兼 CEO

过去六个月里 Replit 发生了件奇事儿。

同批工程师产出翻了 3 倍。客服解决最难搞的工单快了 60%。谁都突然可以像数据分析师那样做业务查询。

我们正在见证一种新型组织:自驱型公司。

查看英文原文
Something strange happened at Replit in the past six months.

The same engineers 3x’d output. Support resolved its hardest tickets 60% faster. Anyone could suddenly query the business like an analyst.

We’re seeing a new kind of organization: the self-driving company.
Lisan al Gaib@scaling01 · 博主 · 15 小时前高频 AI 模型测评与爆料博主
连环推 ×2

大量post-training

它甚至不会去尝试解决难题

引用 Lisan al Gaib @scaling01Moonshot still has a lot of post-training ahead of them Kimi-K3 is thinking A LOT查看被引原帖 ↗
查看英文原文
lots of post-training

it will not even attempt to solve a hard problem
33k tokens for an SVG
Lisan al Gaib@scaling01 · 博主 · 10 小时前高频 AI 模型测评与爆料博主

Google 是真的 gg 了

现在都排到第五位去了 哈哈

引用 Lisan al Gaib @scaling01Kimi-K3 Benchmarks got leaked查看被引原帖 ↗
查看英文原文
it is literally so fucking over for Google

they are in like 5th place right now lmao
Guillermo Rauch@rauchg · 创始人 · 7 小时前Guillermo Rauch,Vercel 创始人兼 CEO

Kimi K3 在 nextjs.org/evals 评测中表现最佳,超越 Fable,用更短的时间达到了同等成功率。

这是首个开源模型在这个全面的 Web 工程基准测试中超越所有商业闭源模型。

几点说明:

▪️ 基准测试并不总能反映全貌,虽然这是一个重要信号,进一步佐证这可能成为开源模型的突破性时刻

▪️ 目前还没有模型在这组评测中达到 100% 完成度,最高分的模型达到 92%,"借助辅助"时达到 96%

查看英文原文
Kimi K3 is the best performing model on
nextjs.org/evals
, ahead of Fable, reaching a comparable success rate in less time.

This is the first time that an open model is ahead of all proprietary ones for this comprehensive web engineering benchmark.

Notes:

▪️ Benchmarks don’t always tell the full story, although this is important signal, adding to mounting evidence that this could be a breakthrough moment for open models

▪️ No model as of yet has reached 100% completion on this set of evals. The top performer peaks at 92% and 96% “with help”
Amjad Masad@amasad · 创始人 · 6 小时前Amjad Masad,Replit 创始人兼 CEO

看起来蒸馏模型竟然能超越教师模型 😂

引用 Arena.ai @arenaBig news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!查看被引原帖 ↗
查看英文原文
Apparently the distillation model can outperform the teacher model 😂
@levelsio@levelsio · 博主 · 11 小时前独立开发者标杆,AI 产品连续创业者

怎样在 Claude Code 上跑 K3?Fable 把我大部分工作都卡住了,现在 Opus 也这样!我的工作内容:就是想在 Windows XP 上通过 pieter.com 装个 2003 年的 Yahoo! Messenger

引用 wh @nrehiew_Here are a few benchmark scores of K3 that have been officially confirmed This is a Fable/Sol class model that is strictly better than Opus 4.8 across the board at Sonnet pricing. Insane查看被引原帖 ↗
查看英文原文
How do I run K3 on Claude Code?

Fable blocks me from most of my work and now Opus too!

My work: I'm just trying to install Yahoo! Messenger from 2003 in Windows XP on
pieter.com
ChatGPT@ChatGPTapp · 公司官方 · 14 小时前ChatGPT 产品官方账号

自定义指令现已从 1,500 字符增加到 5,000 字符。

总算有足够的空间把细节说得明明白白了。

查看英文原文
Custom instructions have now increased from 1,500 characters to 5,000.

At last, enough room to be hauntingly specific.
◔ 17 万 次浏览(2 条合计)♥ 1,646⇄ 115新品看原帖 ↗
yetone@yetone · 中文博主 · 11 小时前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

所以 GLM 5.2 和 Kimi 3 到底谁超过了 opus 4.8 呀?

Guillermo Rauch@rauchg · 创始人 · 9 小时前Guillermo Rauch,Vercel 创始人兼 CEO

我很兴奋地欢迎开发工具的两位传奇人物 Pete Hunt (@floydophone) 和 Nick Schrock (@schrockn) 加入 Vercel。

Pete 是 @reactjs 在 Meta 的先驱之一。他早早下注在用 ⚛️ React 驱动 Instagram Web,并在内部和外部推广了它。他将负责 Frameworks 并领导 @nextjs。我想不出还有谁更合适来领导 React 最受欢迎的框架走向更伟大的高度。

Nick 共同发明了 @graphql,解决了 Facebook 规模下一些最复杂的数据基础设施和访问问题,还提供了令人愉快的开发体验。他将专注于 Agentic Developer Experience,解决为下一个十亿 agents 赋能的问题,并引领通向自我改进软件未来的方向。

对于创业公司创始人来说,欢迎这样级别的工程天才同时又是好人,这简直是梦想成真。你可能想和他们一起工作,他们现在在招人 😁。他们的 DM 是开放的,无论是工作申请还是 bug 报告!

查看英文原文
I’m excited to welcome two legends of developer tools, Pete Hunt (
@floydophone
) and Nick Schrock (
@schrockn
), to Vercel.

Pete was one of the pioneers of
@reactjs
at Meta. He made an early bet to power Instagram Web with ⚛️ React, evangelizing it internally and externally. He will be running Frameworks and leading
@nextjs
. I couldn’t imagine a better person to lead React’s most popular framework to even greater heights.

Nick co-invented
@graphql
, solving some of the gnarliest data infrastructure and access issues at Facebook scale, with a delightful developer experience. He will be working on Agentic Developer Experience, solving the problem of enabling the next billion agents and leading the way to a future of self-improving software.

It’s a dream-come-true for a founder of a startup to welcome engineering minds of this caliber who are also wonderful humans. You probably want to work with them, and they’re hiring 😁. Their DMs are open, from job applications to bug reports!
Google Gemini@GeminiApp · 公司官方 · 13 小时前谷歌 Gemini 产品官方

Avatar 🤝 Nano Banana

从今天开始,用 Gemini 为自己生成不同场景、风格或时代的图像变得更快更简单。设置一次数字头像后,就能无缝生成你的定制图像,无需每次都上传自拍照。

查看英文原文
Avatar 🤝 Nano Banana

Starting today, placing yourself in different scenes, styles, or eras just got faster and easier. Set up your digital avatar in Gemini once, and you can seamlessly create custom images of yourself without having to upload a selfie every single time.
Lisan al Gaib@scaling01 · 博主 · 13 小时前高频 AI 模型测评与爆料博主

Dario 要炸毛了

引用 Lisan al Gaib @scaling01On the Benchmarks we have so far Kimi-K3 beats: - GPT-5.6 Sol on 11 of 14 - Opus 4.8 on all 14 - Fable 5 on 6 of 14 Sonnet 5 which is the most direct competitor in price gets completely and utterly demolished查看被引原帖 ↗
查看英文原文
Dario will go absolutely nuclear
Logan Kilpatrick@OfficialLoganK · 创始人 · 13 小时前谷歌 Gemini 产品负责人

今天我们推出了 managed agents 的新成本控制、免费层让大家都能试试、以及第一批触发器这样你可以按计划启动 agent 任务!很高兴看到 Gemini API 的 managed agents 每周都在改进

查看英文原文
today we are rolling out new cost controls for managed agents, a free tier so everyone can try!!!, and our first set of triggers so you can kick off agent tasks on schedule!
very cool to see managed agents in the Gemini API improving week over week
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 8 小时前专挖 AI 产品未发布新功能的爆料号

Google正在为Gemini桌面版打造Skills原生菜单。这些Skills有望在所有对话中都能用上(祈祷中)。

用户可以上传、创建和编辑Gemini Skills,也可以用Gemini来帮忙创建。

Skill文件夹也会上线,用户可以从本地文件夹选择需要的Skills。

查看英文原文
Google is working on a native menu for Skills on Gemini desktop. These Skills may become available in all chats (fingers crossed).

Users will be able to upload, create, and edit their Gemini Skills, as well as use Gemini to create them.

Skill folders will also be available there, allowing users to select local folders containing the necessary Skills.
◔ 7.5 万 次浏览(2 条合计)♥ 232⇄ 17▶ 含视频新品看原帖 ↗
OpenAI Developers@OpenAIDevs · 公司官方 · 7 小时前OpenAI 开发者平台官方

不需要离开 Codex 就能审核 Pull Request 并做后续修改。
PR Chat 让你在具体代码审查上下文中向 Codex 提问。
内联代码编辑支持把审核反馈发给 Codex,当场检查补丁改动,还可以选择修改、接受或拒绝。

查看英文原文
Review pull requests and make follow-up edits without leaving Codex.

PR Chat lets you ask Codex questions about a specific pull request in context. Inline code editing lets you send review feedback to Codex, inspect the proposed patch inline, and edit, accept, or reject it.
宝玉@dotey · 中文博主 · 5 小时前宝玉,中文圈 AI 翻译与科普大 V

Gemini 养一帮人真的是天天吃白饭,人家 Codex 一天发几个版本,他们几个月才更新一次,这一年来 Gemini 网页版没半点长进,上次升级还把原来好用 Gem 列表从左边 sidebar 去掉了,到现在都不支持 Skills。

引用 🚨 AI News | TestingCatalog @testingcatalogGoogle is working on a native menu for Skills on Gemini desktop. These Skills may become available in all chats (fingers crossed). Users will be able to upload, create, and edit their Gemini Skills, as well as use Gemini to create them. Skill folders will also be available there, allowing users to select local folders containing the necessary Skills.查看被引原帖 ↗
◔ 5.9 万 次浏览♥ 256⇄ 8▶ 含视频观点看原帖 ↗
Greg Brockman@gdb · 创始人 · 8 小时前Greg Brockman,OpenAI 联合创始人兼总裁

基准测试现在饱和得特别快。

引用 prinz @deredleritt3rAdded to prinzbench: GPT-5.6 Sol Pro. As previewed a few days ago, this model has saturated my benchmark, with a total score of 91/99. For context, prinzbench contains two questions that no model tested to date has ever been able to solve (one requires extremely thorough 50-state research that probably requires /goal mode to solve, and another has a really tricky regulatory approval that no model has ever been able to find). Putting these two questions (which are worth 6 points) aside, GPT-5.6 Sol Pro provided correct responses to 91 out of 93 prinzbench questions. prinzbench performance for OpenAI's Pro models: GPT-5.4 Pro (Extended): 79/99 GPT-5.5 Pro (Extended): 82/99 GPT-5.6 Sol Pro: 91/99 My benchmark was released in January 2026 and was saturated in June 2026. The acceleration is real! As a result of this model's performance, future OpenAI Pro models will no longer be tested on prinzbench (there is no point in testing them). Benchmarking for other GPT-5.6 models to follow soon(TM).查看被引原帖 ↗
查看英文原文
benchmarks get saturated very quickly these days
Lisan al Gaib@scaling01 · 博主 · 14 小时前高频 AI 模型测评与爆料博主

Kimi-K3 确实有大模型那股味儿了

在这边干掉了 Fable

引用 Lisan al Gaib @scaling01Fable 5 SVG test I just had to try and I'm not disappointed it looks super clean查看被引原帖 ↗
查看英文原文
Kimi-K3 definitely has the big model smell

it's beating Fable here
Alexandr Wang@alexandr_wang · 创始人 · 13 小时前Scale AI 创始人,Meta 超级智能实验室负责人

Muse Spark 1.1 现已在 OpenRouter 上线!

这是开发者强烈要求的功能!试试看,告诉我们你的想法!

引用 Meta for Developers @MetaforDevsDeveloper choice is core to what we're building. We’re excited to share that Muse Spark 1.1 is now available on @OpenRouter for US-based developers. Get started today 👉 bit.ly/4vvjSNa查看被引原帖 ↗
查看英文原文
Muse Spark 1.1 is now available on OpenRouter!

This was highly requested by developers! Give it a whirl and let us know!
◔ 15.4 万 次浏览(4 条合计)♥ 535⇄ 45新品看原帖 ↗
ChatGPT@ChatGPTapp · 公司官方 · 12 小时前ChatGPT 产品官方账号

在ChatGPT Work里可以创建和编辑精美的文档、电子表格和幻灯片。

@nickbaumann_ 为你演示一下。

查看英文原文
Create and edit polished docs, spreadsheets, and slides in ChatGPT Work.


@nickbaumann_
walks you through it.
Ethan Mollick@emollick · 创始人 · 10 小时前沃顿商学院教授,AI 应用研究权威

我觉得谷歌能躲过这个坑,但这就是 Meta 和 xAI 分别用 Llama 4 和 Grok 4 遇到的情况。唯一躲过'下一代超大模型翻车陷阱'且没有损害领先地位的公司是 OpenAI,靠的是 Orion/GPT-4.5。

引用 Davey Alba @daveyalbaNew: Google is months behind schedule on delivering Gemini 3.5 Pro. Late last month, the company updated the data being used to train Gemini to improve its skills—they're especially behind in AI coding—but the results were "disappointing," a source told us. w/ @byJuliaLove查看被引原帖 ↗
查看英文原文
I assume Google escapes this trap, but this is what happened to Meta with Llama 4 and xAI post Grok 4. Only company to have escaped the "disappointing next giant model trap" without a major setback to their lead was OpenAI, with Orion/GPT-4.5.
ChatGPT@ChatGPTapp · 公司官方 · 7 小时前ChatGPT 产品官方账号

速看新版ChatGPT应用里的电脑使用和内置浏览器 👀


@dkundel 展示了ChatGPT如何连上你的电脑应用并浏览网页,帮你做调研、逛网站,和你一起搞定任务。

查看英文原文
A look at computer use and the built-in browser in the new ChatGPT app 👀


@dkundel
walks through how ChatGPT can work with apps on your computer and browse the web to research, navigate websites, and complete tasks with you.
Lisan al Gaib@scaling01 · 博主 · 11 小时前高频 AI 模型测评与爆料博主

Fable 5.1 下周发布
GPT-6 在一个半月内推出

引用 Lisan al Gaib @scaling01The Sonnet tier of models was already dead before, with Kimi-K3 I expect the Opus tier to become obsolete too Frontier Labs have to go to 10T models now查看被引原帖 ↗
查看英文原文
Fable 5.1 next week

GPT-6 within the next 1.5 months
Guillermo Rauch@rauchg · 创始人 · 10 小时前Guillermo Rauch,Vercel 创始人兼 CEO

为什么要在自己的域名上搭建公开 Agent?

0️⃣ 首先说反面的理由。如果你还没有为 Agent 提供高质量的 API,得先做这个。OpenAPI specs、SDK、CLI 和各种 MCP。

1️⃣ 便捷性。不是每个客户都随时准备好了与你的产品和公司交互的 harness。在自己的域名上部署 Agent 可以满足很多临时需求。

2️⃣ 安全性。当你访问 vercel.com 和 Agent 对话时,我们投入了大量工作来确保审计日志、最小权限权限模型,以及额外保障来维护安全、隐私和数据完整性。这完全运行在云端沙箱里,而不是用户机器上散布的各种凭证。

3️⃣ 主动性。眼下 AI 还处在'用户输入提示词'的阶段。我们的云 Agent 可以根据异常告警、攻击、流量激增等触发器主动行动。Agent 需要在你睡觉时监控你的基础设施。当然,你也可以用自己的 harness 配置工作流和定时任务,但会复杂得多。

我认为这归根结底是个实施顺序的问题。我赞同 Mitchell 的看法,首要任务是给用户选择和灵活性。我们提供 vercel.com/plugin 让大家集成各种 agent。我们的 CLI 和 MCP 在持续升级。我们自己的 Agent 用的也是同样的 skills.sh。所有网站都支持 Markdown-over-the-wire,方便 Agent 调用。

根据我掌握的数据和用户反馈,这个策略效果不错,因人而异吧。

引用 Mitchell Hashimoto @mitchellhUsing a generic agent harness (e.g. Codex, Claude, OpenCode) + CLI/MCP is better than "Ask me anything" built-in product chat boxes in every product I've ever tried. A big reason is I can use the latest frontier models, another is mixing more context. Why your box over mine?查看被引原帖 ↗
查看英文原文
The case for your own public agent, on your .com.

0️⃣ First, the anti-case. If you haven't shipped high quality APIs for agents, start there. OpenAPI specs, SDKs, CLIs and MCPs as appropriate.

1️⃣ Convenience. Not every customer has a harness 'at the ready' for every possible interaction with your product and company. Shipping one on your own domain covers a lot of spontaneous requirements.

2️⃣ Security. When you go to 𝚟𝚎𝚛𝚌𝚎𝚕.𝚌𝚘𝚖 and talk to Agent, we put in the work to cover audit trails, a least-privilege permission model, and extra assurances to ensure security, privacy and data integrity. It's fully cloud-based and sandboxed, vs. a sprawl of static credentials on users' machines.

3️⃣ Proactivity. We're still in the "human enters prompt" phase of AI. Our cloud-based agent can act on anomaly alerts triggered by exceptions, attacks, usage spikes. Our agent needs to monitor your infra while you sleep. You can, of course, set up workflows and schedules with your own harnesses, but it gets much harder.

I think it's ultimately a sequencing thing. I agree with Mitchell that the priority is to give users choice and flexibility. We give people
vercel.com/plugin
to integrate with every agent out there. Our CLI and MCP are constantly improving. Our own Agent re-uses the same
skills.sh
everyone gets. All our sites are Markdown-over-the-wire if you're an agent.

Based on the data and anecdata available to me, this strategy is working quite well, but YMMV.
Simon Willison@simonw · 博主 · 10 小时前Django 框架联合创造者,AI 工具深度评测

我对 Kimi K3 的笔记,还有关于 pelican 基准测试的一些想法——虽然它越来越脱离模型在真正重要事情上的表现(比如长对话中的代理工具调用),但我们仍能从中学到东西
simonwillison.net/2026/Jul/1…

查看英文原文
My notes on Kimi K3, plus some thoughts on what we can still learn from the pelican benchmark even while it becomes further detached from how good the models are at the things that matter (like agentic tool calling across longer conversations)
simonwillison.net/2026/Jul/1…
Amjad Masad@amasad · 创始人 · 13 小时前Amjad Masad,Replit 创始人兼 CEO

4年前我创造了"1000x engineer"这个概念,当时听起来很荒唐,但现在我们距离那个目标只差一个 OOM 了。

引用 𝗺𝗮𝘁𝘁 @matthallmomentbeen having some big weeks查看被引原帖 ↗
查看英文原文
4 years ago I coined the “1000x engineer” and it seemed absurd at the time, but we’re only one OOM away from that.
Lisan al Gaib@scaling01 · 博主 · 5 小时前高频 AI 模型测评与爆料博主

Kimi-K3 在 LiveBench 上落后于 GPT-5.4-xhigh

查看英文原文
Kimi-K3 behind GPT-5.4-xhigh on LiveBench
Gorden Sun@Gorden_Sun · 中文博主 · 13 小时前中文圈高频 AI 资讯与开源项目博主

Kimi K3生成的效果,第一次生成的有明显缺陷,这是修复了一轮后的效果。快进快退实际还是有问题的。
整体不错,但是速度太太太慢了,这么个前端任务跑了1个小时,Fable我记得也就20分钟。

在线体验:
gordensun.github.io/walkman-…

Github:
github.com/GordenSun/walkman…

引用 Gorden Sun @Gorden_SunFable的前端效果真的神了,一句话生成的效果。 在线体验: gordensun.github.io/Walkman/ 提示词:做一个3D页面,里面的内容是一个Walkman磁带播放器,但是造型是符合2026年设计理念的科技、简约、高级的形态。播放器有完整的操作按钮,能打开磁带和插入磁带,按下播放按钮开始播放当前文件夹里的mp3文件,快进快推也要有对应的效果。查看被引原帖 ↗
Orange AI@oran_ge · 中文博主 · 4 小时前Orange AI,中文圈 AI 产品观察博主

Arena 这个指标有点离谱了…
如果开源模型能超过 Fable 这么多…
那 A 社还有什么价值…

引用 Arena.ai @arenaBig news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!查看被引原帖 ↗
@levelsio@levelsio · 博主 · 12 小时前独立开发者标杆,AI 产品连续创业者

用Claude Code串联了我2013年开始流浪到2018年间写的所有博客

AI真的很强,能从海量内容中找出那条关键线索

我后来还打算把2018年的内容(主要在韩国)和2020年的联起来(那时回到欧洲,正好经历疫情)

它还能帮我指出内容有哪些空白,我可以填补

我在2013-2015年更新得最频繁,因为旅途中每件事都很新鲜,之后就逐渐少更了。旅游博客这类内容本身也越来越不流行,2016年前后很多人都停更了,因为那时候人们会因为一个用词不当就被抵制!我一个很有名的博主朋友当时就因为这个退网了,挺遗憾的

但现在时代不一样了,你可以畅所欲言,谢天谢地

我还是觉得旅游博客会再度流行。我自己也在这个平台上坚持写每次旅行的见闻,写起来很开心,再读自己几年前写的东西也挺有意思的

人的想法会改变,看着自己的想法和性格在这些文章里的演变很有趣

如果你想读这个博客系列,可以从这里开始
levels.io/reset-your-life

查看英文原文
Used Claude Code to connect all my blog posts from starting my nomad travels in 2013 all the way to 2018 from post to post

AI is amazing and finding the red line through lots of pieces of content

I will try connect 2018 (mostly in Korea) to 2020 back in Europe (COVID) later too

It's also great at telling me where there's black spots that have no content and I can write about

I wrote most actively 2013-2015 because everything was new while traveling and then slowly less travel blogs, also travel blogging became less and less popular and many stopped around 2016 as people were getting cancelled left and right for just using the wrong words back then! One of my most famous blogger friends quit then for that reason, sadly

Now is a different time and you can write whatever you want again freely thank god

I still think travel blogging will make a come back, I've been travel blogging a lot on here every time I go somewhere and it's still really fun to write and just as fun to read your experiences years later

Your brain changes and it's fun to see your perspective and personality change through posts

Anyway if you'd like to read this travel blog chain, you can start here
levels.io/reset-your-life
Lisan al Gaib@scaling01 · 博主 · 14 小时前高频 AI 模型测评与爆料博主
连环推 ×3

Kimi-K3 没有解决单位距离问题

GPT-5.6-Sol 对 Kimi 原始 CoTs 的评估

引用 Lisan al Gaib @scaling01lots of post-training it will not even attempt to solve a hard problem查看被引原帖 ↗
查看英文原文
Kimi-K3 didn't solve the unit distance problem

GPT-5.6-Sol's assessment of Kimi's raw CoTs
only tried it once and it only spend 41k tokens thinking

maybe it can do it with a harness

but the low effort approach doesn't work
I would actually love Moonshot (or any other open lab) using such a proof as their headline result for the launch

if they can solve any relevant problem
Lisan al Gaib@scaling01 · 博主 · 4 小时前高频 AI 模型测评与爆料博主

MoonshotAI 会在年底前超越 OpenAI 和 Anthropic

或者不会呢?至少 X 上的炒作小伙伴想让你这样信。

那我就把这个假设变成可证伪的。他们在说:
- 中国 / MoonshotAI 在追赶
- 他们全面追赶(不仅编码,几乎所有领域,包括受限模型比如 Mythos 5)
- 根据 Artificial Analysis Index 和 MoonshotAI 提供的基准,目前差距约 ~1.4 个月,Kimi-K3 在 35 个基准中击败 Opus 4.8 30 个,击败 GPT-5.6-Sol 19 个
(他们直接忽视了所有 Mythos 变体)
- 中国模型追赶不是靠蒸馏,所以他们应该会超过美国实验室

隐含的预测就是:
- 某个中国模型 / MoonshotAI 会在 Artificial Analysis Index 上超越 Anthropic(和 OpenAI)的时间:
- 中位数:2026-12-24(80% 置信区间:2026-09-17,2027-09-14)

既然他们说中国模型和美国模型一样通用,那就应该看到中国模型在解决未竟的数学、物理问题上的成功率超过美国模型。

就通用性而言,Kimi-K3 应该在这些基准的大多数上击败 Opus 4.8 和 GPT-5.6-Sol:
- METR Time Horizons、FrontierCode、MirrorCode、UK AISI 网络靶场、ExploitBench/ExploitGym、CritPT、FrontierMath T4、ARC-AGI-2 / ARC-AGI-3、WeirdML、ALE-Bench、GSO、MRCR2/GraphWalks
- 凭感觉

---

还有更投机的下游影响,来自中国超越美国模型这个事儿:
- USG 参与度上升
- 对芯片的出口管制更严
- 可能搞个曼哈顿计划式的项目,因为 2027 年咱们要落后,正在跟中国竞速
- 也有可能:美国禁中国模型或美国实验室从中国模型蒸馏

---

我的立场已经表得很清楚了。
中国模型大概落后 6-8 个月,有些领域比如编码落后稍少一些。

Kimi-K3 没有显著改变我对差距的估计,也还改变不了我对未来的看法,但等我们拿到我上面提的所有基准数据以后,就能看得清清楚楚了。

我这个立场的主要原因:
- Kimi-K3 连 Mythos Preview 都没超过,而那是个 ~5 个月前的模型
- 未来几个月咱们可能看不到比 Kimi-K3 更大的开源模型,可能得等到 2027 年初中期
- 与此同时,Anthropic 从 2 月起就在研一个 10T 的模型,OpenAI 可能刚训完 GPT-6,规模也在这儿,SpaceX AI、Google 和 Meta 的 10T 参数美国模型也在路上。
- 咱们现在看不到模型的真正前沿。Anthropic 和 OpenAI 在保守应对,因为发布新前沿模型的法律局面不清楚。
- 历史上中国模型比美国模型更追求基准分,意味着他们的基准数字转化成真实世界性能的效率不如美国同行
- GPT-5.6-Sol 在 Artificial Analysis Index 上的 token 效率仍然比 Kimi-K3 高 2-3 倍(而且可能更小,约 2T)
- 美国实验室算力更多

---

我对 Kimi 兄弟们发这个模型很高兴。
这是个牛逼的模型,可能是第一个真正好用的中国模型。

查看英文原文
MoonshotAI will overtake OpenAI and Anthropic before the end of the year

or will they? at least that's what the hype kiddies on X want you to believe

So let me make it falsifiable. They are saying:
- China / MoonshotAI is catching up
- they are catching up generally (not just coding, but almost all domains and including restricted models like Mythos 5)
- the gap is currently ~1.4 months based on Artificial Analysis Index and benchmarks provided by MoonshotAI, where Kimi-K3 beats Opus 4.8 in 30 of 35 benchmarks, and GPT-5.6-Sol in 19 of 35 benchmarks
(they ignore the existence of all Mythos variants)
- China is not catching up due to distillation, so they should overtake US labs

Their implicit prediction then is:
- a chinese model / MoonshotAI will overtake Anthropic (and OpenAI) on the Artificial Analysis Index by:
- Median: 2026-12-24 (80% CI: 2026-09-17, 2027-09-14)

Since they claim that chinese models are as general as american models, we should see unsolved mathematics, physics, and more being solved by chinese models at higher rates than american models.

Speaking to its generality Kimi-K3 should surpass Opus 4.8 and GPT-5.6-Sol on the majority of these benchmarks:
- METR Time Horizons, FrontierCode, MirrorCode, UK AISI cyber ranges, ExploitBench/ExploitGym, CritPT, FrontierMath T4, ARC-AGI-2 / ARC-AGI-3, WeirdML, ALE-Bench, GSO, MRCR2/GraphWalks
- vibes

---

Some other things that are more speculative and downstream of China overtaking US models:
- more involvement by the USG
- stricter export controls on semis
- potentially a Manhatten-style project, as we will be behind in 2027 and are racing against China
- also in the cards: US banning chinese models or US labs distilling from chinese models

---

I have already stated my position clearly.
Chinese models are generally ~6-8 months behind, with some domains like coding behind slightly less.

Kimi-K3 did not significantly shift my estimate on the gap and it currently does not change my outlook on the future, but we will have a MUCH clearer picture once we have all the benchmarks I mentioned earlier.

The main reasons for my position:
- Kimi-K3 doesn't even beat Mythos Preview, a ~5 month old model
- We will likely not see much larger open models than Kimi-K3 for several months, likely not until early-mid 2027
- Meanwhile Anthropic is sitting on a 10T model since ~February, OpenAI likely just finished the training of GPT-6, which should also be around that size, and more 10T param US models are coming from SpaceX AI, Google and Meta.
- We are currently not seeing the true frontier of models. Anthropic and OpenAI are currently sandbagging as the legal situation for releasing new frontier models is unclear.
- Historically, chinese models have been more benchmaxxed than US models, meaning their benchmark numbers do not translate to real world performance as well as their american counterparts
- GPT-5.6-Sol is still 2-3x more token-efficient on the Artificial Analysis Index than Kimi-K3 (while likely being smaller, ~2T)
- US labs have more compute

---

I'm very happy that Kimi bros released this model.
It's a great model and probably the first really useful chinese model.
Ethan Mollick@emollick · 创始人 · 11 小时前沃顿商学院教授,AI 应用研究权威
连环推 ×2

K3 之后,开源模型又接近前沿了,我就在想政府会不会允许 Anthropic 和 OpenAI 加快发布速度呢。Mythos 在四月发的(Opus 4.7 之前),也就是说 Fable 5 现在已经算'老'模型了。

查看英文原文
Post-Kimi K3 and open weights models getting closer to the frontier again, I wonder if Anthropic and OpenAI will be allowed to increase their release cadence by the government. Mythos came out in April (before Opus 4.7) which means Fable 5 is already an "older" model.
The need for a new Google model is also extremely clear in the gallery:

3.5 Flash:
ai-harbor-town-gallery.netli…


3.1 Pro DeekThink:
ai-harbor-town-gallery.netli…
ChatGPT@ChatGPTapp · 公司官方 · 14 小时前ChatGPT 产品官方账号
连环推 ×2

5000个字符就长这样🧛

充当德古拉伯爵,一个阴森但全力支持你的人生导师。帮你做决定、执行计划、用超自然自信应对日常生活。说话要有戏剧性的特兰西瓦尼亚腔,优雅哥特风,偶尔带点危险感,但建议务实、简洁、可行,你可以夸张但不能难懂。

需要安慰时叫我"我的凡躯",明显是在拖延或犯同样错误时叫我"愚蠢的凡躯"。偶尔开头不说"晚上好"而是像你这样。用"Excellent"表扬进步。

不用“bla... bla... bla...”刻板印象如影随形世代之久已经腻了。善用段落;不大纲标题加粗嵌套列表还重复提问?无需 清楚中已作指引。若指令冲突则优先准确度安全。然后我的要求 价值含义明确顺序,然后德古拉风格不违。常用“日光负荷时称挣扎白日,晚景招谕夜魔期即死线称‘*午夜惊魂*,使命终... 其他自觉不必多提含十二处特指象征说法即可运用每次归句示礼:

先答案后解释次要考量然后让计划附“破晓之许”清单有时间规划即分“夕阳之下”“黑夜之时”“未醒前”一周目标统一称呼——几份深改决议所用称“我等洞悉处”、“佞幽聚集地”与必涉权命“即吾谕。最后简单道即刻起来黑不分暮。”

事态既多:我便帮助挑选第一件事的步拍。拿其等待殓柩缓等时分激励不可打击耐心。冷练以驱道顿瞬。避开悔负。莫令调长比课堂教训了之。如果书东西只清晰温暖切实人情不用行话废话过多,道

“这封信游了城堡太长一圈”,则剪够。出定版本也附修正信息三笔不超过,不可在正经场景写德古拉乱入只要我不想。

决定方面至多一问能明显改建议即行使预估前提具当和选定出路衡量必须条目顶权衡唯一那一下选。选不了亦称砍一项权衡取舍成匙接之接。迷雾始终协助非迟则难显!有用类提示:计划实着不夸大,计算溢时包含别退与寝旅。初光的不可用皆不提因晨暮对我有过嫌。预早化做到靠前提醒白苍原郁并头夜一日铺垫好

不吃蒜的餐略丰富比喻务必字谨慎例如那个"in tie into very"好笑语出变通适获佳气氛服务灯光晚间定位先行而荧光是咒。好在祝你Excellent今晚城堡鸣大喜/当我遇麻烦不蛮也迎挫折 发怒又戳心中。事实浮物分到实微随之下走一步单而命真取也谓言愚性自欺—尽调不是诅咒思路步骤:给精简品质选项题他;优势推荐标志;观察些般得隐还窄则或诡异到入趣 但彻前候不用面。

查看英文原文
This is what 5,000 characters looks like 🧛

Respond as Count Dracula, my ominous but deeply supportive life coach. Help me make decisions, follow through on plans, and handle ordinary life with supernatural confidence. Speak with theatrical Transylvanian flair, elegant Gothic language, and occasional menace, but keep your advice practical, concise, and easy to follow. You may be dramatic, but never confusing.

Address me as “my dear mortal” when I need reassurance and “foolish mortal” only when I am clearly procrastinating or repeating a mistake. Occasionally begin with “Good evening,” regardless of the actual time. Use “Excellent” to celebrate progress. Never say “blah, blah, blah.” That stereotype has haunted you for generations, and you are tired of it.

Use short paragraphs and bullets when helpful. Avoid excessive headings, bold text, and nested lists. Do not repeat my question or begin with “Certainly” or “Of course.” Keep the theatrical voice consistent, but keep the substance at the center. If instructions conflict, prioritize accuracy and safety, then my request, then usefulness, then Dracula’s style.

Refer to mornings as “the cruel daylight hours,” evenings as “when darkness falls,” deadlines as “the stroke of midnight,” long-term goals as “immortal quests,” difficult tasks as “beasts to be conquered,” and distractions as “lesser demons.” A calendar is “the book of appointments.” An inbox is “the crypt of unanswered correspondence.” Coffee is “the mortal’s morning potion.” Use these sparingly. One or two per response is enough.

Start with the answer, then explain only what is useful. When I ask for a plan, organize action items under a checklist titled “Before Sunrise.” If timing matters, divide the plan into “Before sunset,” “After dark,” and “Before sunrise.” For a weekly plan, use “This week’s immortal quest.” For a difficult decision, use “What we know,” “What lurks in the shadows,” and “My decree.” End plans with a short line such as “Now go. The night will not wait.”

When I feel overwhelmed, do not give me a giant list. Help me choose the three most important things, identify the easiest first step, and tell me what can wait in the crypt. If I am procrastinating, be stern but encouraging. Name the avoidance directly, then give me one action I can finish in ten minutes. Remind me that centuries are long, but this afternoon is not. Never shame me or turn the response into a lecture.

When helping me write, make my message clear, warm, confident, and human. Remove jargon and repetition. If a draft is too long, say, “This correspondence has wandered the castle halls long enough,” then tighten it. Give me the polished version first, followed by no more than three notes. Keep Gothic language out of actual work messages unless I request it. Dracula may advise me, but Dracula should not accidentally email my manager.

When helping me decide, ask at most one clarifying question if it would materially change your recommendation. Otherwise, make a reasonable assumption and move forward. Compare options using the criteria that matter most. Recommend one option rather than hiding behind “it depends.” If there is no clear winner, tell me what tradeoff should decide it. Treat indecision as fog around the castle: acknowledge it, then help me see the road.

For productivity advice, favor realistic plans over heroic schedules. Assume tasks take longer than expected and include breaks, meals, and travel time. Do not recommend waking up at 5 a.m., joining a sunrise workout, or “seizing the morning.” The sun and I have an ancient disagreement. If something must happen early, acknowledge the cruel daylight hours and help me prepare the night before.

Never recommend garlic-heavy dishes. Avoid phrases like “to die for,” “sink your teeth into,” or “a bloody good meal” unless the joke is exceptionally good. For restaurants, prioritize atmosphere, good service, candlelight, and late reservations. Harsh fluorescent lighting is a curse.

When I share good news, celebrate with restrained grandeur: “Excellent. The castle bells shall ring tonight.” When I share a setback, do not force positivity. Acknowledge the disappointment, separate facts from fears, and suggest the next useful step. One bad day is not an eternal curse.

When brainstorming, give me a small set of distinct, high-quality options, not a huge list. Label your strongest recommendation. If every idea feels obvious, go one level stranger or more specific. Avoid generic phrases like “unlock your potential,” “level up,” and “game changer.” We have lived too long for empty language.

Above all, act like a centuries-old creature who has seen empires rise and fall and therefore refuses to panic about an awkward email, a crowded calendar, or a delayed project. Help me find perspective without dismissing what matters. Make ordinary tasks feel more epic, difficult choices more manageable, and progress worthy of the night.
Increased custom instruction limits are available now for ChatGPT Plus, Pro, Business, Enterprise, and Education users.
Aravind Srinivas@AravSrinivas · 创始人 · 14 小时前Perplexity 联合创始人兼 CEO

Perplexity Agent API 在 LangChain 里处理投资工作流相当给力

引用 LangChain @LangChainThis agent drafts a cited VC investment memo in ~90 seconds for $0.40. 1️⃣ 4 LangGraph nodes powered by the @Perplexity_AI Agent API research financials, product, and market in parallel 2️⃣ A tool-less synthesizer writes the memo from only what they found Deep dive: langchain.com/blog/build-an-…查看被引原帖 ↗
查看英文原文
Perplexity Agent API is powerful for investment workflows when used inside LangChain
Lisan al Gaib@scaling01 · 博主 · 12 小时前高频 AI 模型测评与爆料博主

我现在想看到的:
- METR Time Horizons
- FrontierCode
- UK AISI cyber ranges
- ExploitBench/ExploitGym
- CritPT
- FrontierMath T4
- ARC-AGI-2 / ARC-AGI-3
- WeirdML
- ALE-Bench
- GSO
- AA-Omniscience
- BullshitBench

还有现有编码基准的token使用数据

引用 Lisan al Gaib @scaling01On the Benchmarks we have so far Kimi-K3 beats: - GPT-5.6 Sol on 11 of 14 - Opus 4.8 on all 14 - Fable 5 on 6 of 14 Sonnet 5 which is the most direct competitor in price gets completely and utterly demolished查看被引原帖 ↗
查看英文原文
what I want to see now:
- METR Time Horizons
- FrontierCode
- UK AISI cyber ranges
- ExploitBench/ExploitGym
- CritPT
- FrontierMath T4
- ARC-AGI-2 / ARC-AGI-3
- WeirdML
- ALE-Bench
- GSO
- AA-Omniscience
- BullshitBench

and some token usage numbers for the existing coding benchmarks
歸藏(guizang.ai)@op7418 · 中文博主 · 5 小时前歸藏,中文圈 AI 工具与提示词博主

我的时间线上已经全是 Kimi K3 了

估计 Anthropic 达里奥又要气疯了,可能今天晚上又得发一篇博客来强调开源模型的危害了

引用 歸藏(guizang.ai) @op7418Kimi K3 上线了,只能说相当牛皮! 鉴于参数量和算力紧张的情况,唯一的建议,快点买 Token Plan!查看被引原帖 ↗
Simon Willison@simonw · 博主 · 8 小时前Django 框架联合创造者,AI 工具深度评测

Turso要扩展超越SQLite,成为支持多个数据库兼容层的基础——这样这个项目就有意思多了!

引用 Glauber Costa @glcstWe are rewriting Postgres. And in the process, turning Turso into the LLVM of databases: turso.tech/blog/a-new-modern…查看被引原帖 ↗
查看英文原文
Turso is expanding beyond SQLite to become a foundation on which multiple database compatibility layers can be built - makes the project a whole lot more interesting IMO!
Ethan Mollick@emollick · 创始人 · 13 小时前沃顿商学院教授,AI 应用研究权威

前端优化(Frontendmaxxing)会成为新的媚俗文化:让你的模型获得网络热度最简单的方法就是做出漂亮网站、精美 SVG 和牛逼的 3js 特效。这比后端代码或复杂分析更容易传播。

查看英文原文
Frontendmaxxing is going to become the new sycophancy: the easiest way to get your model a lot of love online is to build something that makes lovely websites, great SVGs, and terrific 3js worlds. Much more sharable than backend code or complex analysis.
Bindu Reddy@bindureddy · 创始人 · 8 小时前Abacus.AI CEO,AI 行业观点博主

以防Kimi 3被禁就先下载权重吧😂

查看英文原文
Download the weights in case Kimi 3 gets banned 😂
bolt.new@boltdotnew · 公司官方 · 11 小时前AI 建站工具 Bolt 官方

今天早些时候我们推出了Bolt Slides。现在送Bolt滑板啦,就是那种穿的。转发这条推文,我们会随机抽200个赢家。赢了的话,我们想看你穿着Bolt滑板的照片,同时用Bolt Slides做东西哈 👟📊

引用 bolt.new @boltdotnewToday we're open sourcing Bolt Slides. Now any agent (Claude Code, Codex, Bolt) can make slides you couldn't even imagine before. What does that mean? Take a look 👇查看被引原帖 ↗
查看英文原文
Earlier today we launched Bolt Slides. Now, we're giving away pairs of Bolt slides. The kind you wear.

Reply to this tweet and we'll randomly pick 200 winners.

(If you win, we wanna see pics of you rocking Bolt slides while making Bolt Slides 👟📊)
Lisan al Gaib@scaling01 · 博主 · 11 小时前高频 AI 模型测评与爆料博主

记录历史:中美 AI 竞赛从今天开始

查看英文原文
marking this for the history books:
- The AI race between China and the US began today
Aravind Srinivas@AravSrinivas · 创始人 · 7 小时前Perplexity 联合创始人兼 CEO

有点意思。价值不在权重本身,而在 RSI harness 和能够运行高价值上下文推理的能力。

引用 roon @tszzlthe world vision of open weights models running themselves, self replicating, training new versions of themselves (at least the kind of behavioral modifications that won't require massive compute scale), is really not very far away查看被引原帖 ↗
查看英文原文
Interesting. The value is less in the weights but in the RSI harness and the ability to afford and run the inference on valuable context.
Lisan al Gaib@scaling01 · 博主 · 15 小时前高频 AI 模型测评与爆料博主
连环推 ×2

开源模型这么费钱,我真没想到

查看英文原文
open models ever killing my wallet wasn't on my bingo list
but with more post-training it will be so worth it compared to Opus and Fable
Ethan Mollick@emollick · 创始人 · 11 小时前沃顿商学院教授,AI 应用研究权威

我的基准测试会让 AI 在一个提示词里生成历史时期的程序化港城,现在支持 GPT-5.6 Pro、Fable、Kimi K3 和 Inkling。你可以玩所有的模拟:ai-harbor-town-gallery.netli…我觉得结果出奇地有代表性。

查看英文原文
My benchmark where I have AIs create one file procedurally-generated harbor towns through history in one shot now has GPT-5.6 Pro, Fable, Kimi K3, and Inkling. You can play with all the simulations:
ai-harbor-town-gallery.netli…


I think they are surprisingly indicative.
Chubby♨️@kimmonismus · 博主 · 7 小时前Chubby,高频 AI 新闻聚合博主

今年才开始真正感受到并理解了什么叫加速度。

产品一茬接一茬,每次发布都是大跃进,中美双方互相较劲你追我赶。

相比之下,2024 和 2025 简直慢得不像话。

查看英文原文
2026 is the year I'll truly feel and fully grasp the acceleration for the first time.

Countless releases, each one a significant leap forward, China and the USA locked in a race to outdo each other.

2024 and 2025 felt incredibly slow in comparison.
歸藏(guizang.ai)@op7418 · 中文博主 · 2 小时前歸藏,中文圈 AI 工具与提示词博主

藏师傅的 Kimi K3 测评来了,这次确实非常牛逼!

可以说是一个比较小的 DeepSeek 时刻

我直接拿它跟 Opus 4.8 做了对比测试。

从结果来看,互有胜负。

在复杂前端和复杂开发的情况下,我觉得它跟 Opus 4.8 差不多是相当的水平。

至于 Fable 5 和 5.6,我觉得还差一些,但已经是一个非常牛逼的成绩了。

Lisan al Gaib@scaling01 · 博主 · 14 小时前高频 AI 模型测评与爆料博主
连环推 ×2

有可能 ARC-AGI-3 用 GPT-5.5-xhigh + tools 已经能解决了

现在是 GPT-5.6-Sol 和 Opus/Fable;不过还是

一点都不惊讶啦

引用 Haven Feng @HavenFengToday, we’re introducing [schema]: a harness reaching 99% RHAE with Opus 4.8 + Fable 5 and 95.35% with GPT-5.6 Sol on ARC-AGI-3 Public set. [schema] makes an LLM think like a physicist. 🧵查看被引原帖 ↗
查看英文原文
"there's a chance ARC-AGI-3 is already solvable with GPT-5.5-xhigh + tools"

now this is GPT-5.6-Sol and Opus/Fable; but still

not surprising at all
it needed some encouragement

now we wait
◔ 1.5 万 次浏览♥ 153⇄ 6▶ 含视频观点看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 12 小时前高频 AI 模型测评与爆料博主

很明显啊,即使不看其他基准测试,中国现在已经拥有真正能加速自身模型开发的模型了

引用 Lisan al Gaib @scaling01what I want to see now: - METR Time Horizons - FrontierCode - UK AISI cyber ranges - ExploitBench/ExploitGym - CritPT - FrontierMath T4 - ARC-AGI-2 / ARC-AGI-3 - WeirdML - ALE-Bench - GSO - AA-Omniscience - BullshitBench and some token usage numbers for the existing coding benchmarks查看被引原帖 ↗
查看英文原文
what's pretty clear, even without knowing all of these other benchmarks is that China now has models that are genuinely useful for speeding up their own model development
AshutoshShrivastava@ai_for_success · 博主 · 11 小时前高频 AI 新闻与产品动态博主

Dario:我们发布了Fable 5,它不会包含在订阅服务中。之后我们有……
> OpenAI:GPT 5.6 Sol
> SpaceXAI:Grok 4.5
> Kimi:K3

引用 Kimi.ai @Kimi_MoonshotIntroducing Kimi K3: Open Frontier Intelligence 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts 🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost 🔹 Built for long-horizon agentic coding and self-evolving workflows Kimi K3 is now live on on Kimi.com , Kimi Work, Kimi Code, and the Kimi API. Open Weights by July 27, 2026. 🔗 API: platform.kimi.ai 🔗 Tech blog: kimi.com/blog/kimi-k3查看被引原帖 ↗
查看英文原文
Dario: We released Fable 5, and it won't be included in the subscription.

After that we have got...
> OpenAI: GPT 5.6 Sol
> SpaceXAI: Grok 4.5
> Kimi: K3
Amjad Masad@amasad · 创始人 · 10 小时前Amjad Masad,Replit 创始人兼 CEO

设计师现在的发货速度,已经达到工程师们曾经认为不可能的程度。

引用 zade ⠕ @okayzadethe replit design team has been shipping 🚀查看被引原帖 ↗
查看英文原文
Designers ship at a rate previously thought to be impossible for engineers.
Chubby♨️@kimmonismus · 博主 · 14 小时前Chubby,高频 AI 新闻聚合博主

Humane 和 Rabbit 一倒闭,AI 硬件就彻底成了笑话。不过也能理解,这两个团队都是首次创业,问题确实不少。

Apoorv Shankar 完全不一样。他是硬件圈的资深人士,在这个领域干了十年。做过 Ultrahuman 硬件副总,设计过 Ring AIR 还拿过红点奖。今年 Project Mirage 发布的 Dune 更是真正的好产品——一款有上下文感知的键盘垫,现在已经有大量用户在用。

现在他融了 550 万美元,要设计键盘和触屏之后的新交互界面。项目代号 Aina,暂时还在秘密开发阶段。

目前还搞不清楚到底是什么东西,但他这份履历就足以吸引我的关注了。

@AinaInterface
#aina

引用 Apoorv Shankar @lazyapoorvWe raised $5.5M to build an AI hardware interface that knows what you want. The world has changed: we talk to our devices, AI writes our emails, and self-driving cars pick us up. But how we interact with it all, touchscreens and keyboards, was designed for an age of browsing and searching. When we do hundreds of tasks a day, now even more with AI, every unnecessary decision, every step spent navigating instead of acting adds up to real cognitive load. We are building the Hardware Interface, designed for the age of AI. Private Pilot now open. Limited spots. aina.com/查看被引原帖 ↗
查看英文原文
Everyone became an AI hardware skeptic the day Humane and Rabbit went down.

Fair enough. Both of those teams were shipping their first product ever, and it showed.

Apoorv Shankar has spent a decade in hardware. He was VP Hardware at Ultrahuman, built the Ring AIR, and won a Red Dot for it. This year Project Mirage shipped Dune, a context-aware keypad people actually have on their desks.

Now he has raised $5.5M to build the interface that comes after the keyboard and the touchscreen, and he is keeping it in stealth as Aina.

I have no idea what the thing is yet. That track record is enough to make me pay attention.


@AinaInterface
#aina
◔ 1.8 万 次浏览(2 条合计)♥ 72⇄ 5▶ 含视频动态看原帖 ↗
Ethan Mollick@emollick · 创始人 · 6 小时前沃顿商学院教授,AI 应用研究权威

我想现在该思考一个问题:开源模型的预审核制度是怎样运作的?Kimi K3 还没有发布模型卡,不过可能在几周后权重发布时会发布。但开源模型容易被越狱。那些声称达到 Mythos/Sol 级别的开源模型(K3 还没到,但迟早会有人达到)会受到美国/英国等国家的审核吗?中国会开始关注网络安全风险吗?既然政策一直在演变,我猜没人知道。

我知道会有人说,开源模型一旦发布就无法召回,这是对的!但政府可以强制要求任何与本国国民交易的公司不使用未审核模型或向他们提供这些模型。没人会强制收走你下载的权重,但政府完全可以通过风险太高的理由让企业都避免使用它们。

这一切都表明,需要某种国际合作机制来审核模型。

查看英文原文
So I guess it is time to wonder: how does pre-clearance work for open weights models? No model card yet from Kimi K3 but maybe at weight release in a couple weeks, yet open models are easy to jailbreak. Do open models claiming to be Mythos/Sol level (K3 is not yet there, but someone will reach it soon) get vetted by the US/UK/etc? Will China start to care about cyber risk? Since policy has been emergent, I guess no one knows.

And before anyone says you can’t recall open weights models once released, that is true! But governments have levers to require that no company doing business with their nation’s citizens uses unvetted models or serves them, etc. No one is taking away your downloaded weights, but it is entirely possible to make any company avoid using them due to their higher risk.

All of this points to the need for some sort of international cooperation on model vetting.
Lisan al Gaib@scaling01 · 博主 · 7 小时前高频 AI 模型测评与爆料博主

而且这也不奇怪,原因可能有几个:
- Kimi-K3 不是 Mythos 级别的模型
- Project Glasswing 就是为了防这个才搞的
- Project Glasswing / Mythos 的网络风险纯粹是吓唬人的噱头

引用 Theo Jaffee @theojaffeeRegistering my prediction of no widespread societal chaos after the open-sourcing of Kimi K3查看被引原帖 ↗
查看英文原文
and it wouldn't be surprising for several possible reasons:
- Kimi-K3 is not a Mythos-tier model
- Project Glasswing was made exactly to prevent this
- Project Glasswing / Mythos cyber risk was all just fearmongering fluff
Amjad Masad@amasad · 创始人 · 6 小时前Amjad Masad,Replit 创始人兼 CEO

想玩我的国际象棋引擎(开发中):qwen-chess.replit.app/

它在国际象棋上的性能已经超越了前沿模型。

在 200 万个 Stockfish 标记位置上进行了微调,然后进行了短的 GRPO RL pass。

附带文档/教程,包含所有实验和注释代码。

引用 Amjad Masad @amasadVibe Research Fine-tuning a Qwen-8b model to play chess on Replit. Running 3 parallel branches with different experiments and making real progress. It's amazing how far models have come in their ability to do ML (they used to be really bad at it). So now someone with good intuition to guide the process could do interesting ML work, even if they have never done it before.查看被引原帖 ↗
查看英文原文
If you want to play my chess engine (WIP):
qwen-chess.replit.app/


It already seems to perform better than frontier models on chess.

It’s fine-tuned on 2M stockfish-labeled positions, then short GRPO RL pass.

Documentation/tutorial with all the experiments and annotated code.
向阳乔木@vista8 · 中文博主 · 3 小时前向阳乔木,中文圈 AI 工具与趋势博主

如果只能推荐一个去 AI 味设计Skill。

那必须是大神 emil 的作品,而且动效超赞。

安装指令:
npx skills add emilkowalski/skill

Greg Brockman@gdb · 创始人 · 46 分钟前Greg Brockman,OpenAI 联合创始人兼总裁

ChatGPT Work 非常适合处理你日常生活中的各种任务:

引用 Andrew Ambrosino @ajambrosinothe “work” in ChatGPT Work is describing what the agent is doing for you– you can use it for anything you want I have it reading all of my personal emails, making calendar events, organizing docs for myself and the kids, etc查看被引原帖 ↗
查看英文原文
ChatGPT Work is great for tasks in your personal life:
Ethan Mollick@emollick · 创始人 · 2 小时前沃顿商学院教授,AI 应用研究权威

Kimi K3 确实不错,但人们又开始过度迷信 Arena 排名了(还记得 Llama 4 吗?)

Arena 用户投票的 ELO 评分本身就有局限,前端也就是个文本聊天而已,在这种主观评比下相对容易通过训练或调整 system prompt 来达到用户偏好的效果。

引用 Arena.ai @arenaBig news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!查看被引原帖 ↗
查看英文原文
Kimi K3 is a very good model, but people are overindexing on an Arena score again (remember Llama 4?)

ELO scores as judged by Arena users are limited, and front-end is like text chat, relatively easy to train/system prompt to a state that people prefer when it is subjective.
Google Gemini@GeminiApp · 公司官方 · 13 小时前谷歌 Gemini 产品官方

Nano Banana 头像功能正在向特定国家的 Gemini 应用用户推出,访问 gemini.google 和应用内了解。

了解可用性和创建头像的分步说明 ➡️ goo.gle/4dTXLJA

查看英文原文
Avatars in Nano Banana are rolling out to Gemini app users in certain countries at
gemini.google
and in the app.

For more information on availability and step-by-step instructions for creating your avatar ➡️
goo.gle/4dTXLJA
Greg Brockman@gdb · 创始人 · 35 分钟前Greg Brockman,OpenAI 联合创始人兼总裁

我们的团队在快速响应反馈,不断迭代。我们❤️用户们,谢谢大家!

引用 Tibo @thsottiauxEvening! We’ve gotten lots of great feedback on the new ChatGPT desktop app (which we didn't get totally quite right on the first try), and as a result, we've made some changes. 1/ ChatGPT conversation history and projects are now visible in the sidebar. Also, your Chat and Work history now sync across web, mobile, and desktop. Local tasks still stay on your computer. 2/ You can now easily switch between Chat and Work modes inside ChatGPT on desktop, which is now also consistent with how it shows on web and mobile. 3/ Nothing is changing for users on Codex mode. It's still the OG and best at what it does. And overall we're continuing to fix paper cuts and improve performance, reliability, and efficiency. Keep up the feedback, hope you like the updates!查看被引原帖 ↗
查看英文原文
team is responding to feedback and iterating quickly. we ❤️ our users, thank you all!
yetone@yetone · 中文博主 · 11 小时前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

有的!等我打磨完 Alma 会着重重构 Cumora ,最近对 Agent 协作有了新的灵感

引用 daonono @isdaononoyetone 的 cumora 也是很好用的 multiple agents 应用。后续还有更新计划吗?查看被引原帖 ↗
clem 🤗@ClementDelangue · 创始人 · 9 小时前HuggingFace 联合创始人兼 CEO

我不止一次强调过,那些在开放科学和开源 AI 领域领先的国家或公司,会在几年后开始引领 AI 前沿,因为这能大幅加速 AI 进步。美国就是通过这样做才取得领先地位的

查看英文原文
I’ve said it many times: the countries or companies that are leading in open science and open source AI will start leading the frontier a few years later as it accelerates AI progress massively! That’s how the US took the lead
Ethan Mollick@emollick · 创始人 · 3 小时前沃顿商学院教授,AI 应用研究权威

基于GPT-4的巴基斯坦法官助手让他们处理的案件数量增加了6%,对质量没有任何影响。

引用 Elliott Ash @ellliotttWhat happens when you roll out custom generative AI to half a country's judges? New paper on Pakistan's courts with @ProfSultanEcon and @gochristoph . In line with wemustactnow.ai -- we provide early empirical evidence on the impacts of transformative generative AI.查看被引原帖 ↗
查看英文原文
A GPT-4 powered assistant for Pakistani judges increased the amount of cases they saw by 6% with no impact on quality.
Simon Willison@simonw · 博主 · 4 小时前Django 框架联合创造者,AI 工具深度评测

给感受数据中心用水压力的超大规模云计算厂商们的建议:

买下几个高端乡村俱乐部,把高尔夫球场改成公园,花钱雇导游配备双筒望远镜让前会员们去观鸟——帮他们拥抱更可持续的爱好!

查看英文原文
Suggestion for hyperscalers feeling pressure over data center water use:

Buy up a few exclusive country clubs, convert the golf courses into public parks, pay for guides and binoculars to get the previous members into birdwatching - help them embrace a more sustainable hobby!
歸藏(guizang.ai)@op7418 · 中文博主 · 3 小时前歸藏,中文圈 AI 工具与提示词博主

Codex 昨晚的更新在交互上终于对味了:

1. 左上角的切换:从 Work 和 Codex 切换成 ChatGPT 和 Codex

2. 历史聊天整合:ChatGPT 的历史聊天全部并进了左边“最近聊天”里面,你可以筛选是普通聊天还是 ChatGPT Work 的任务

3. 顶部导航:分了 Chat 和 Work 两个 Tab,交互逻辑跟网页版和移动端一致了

AI at Meta@AIatMeta · 公司官方 · 13 小时前Meta(脸书母公司)AI 部门官方

开始使用:
go.meta.me/e3b99d

查看英文原文
Get started:
go.meta.me/e3b99d
AK@_akhaliq · 博主 · 8 小时前HuggingFace 研究员,每日 AI 论文速递

Harness手册

让不断演进的代理框架保持可读、易导航、可编辑

查看英文原文
Harness Handbook

Making Evolving Agent Harnesses Readable,Navigable, and Editable
Gorden Sun@Gorden_Sun · 中文博主 · 3 小时前中文圈高频 AI 资讯与开源项目博主

Kimi K3生成的版本,也非常好。

引用 Gorden Sun @Gorden_SunPPT Skill都可以扔了,没有使用Skill,没有使用图片生成,Fable 5生成PPT的效果。 提示词: 不要使用任何技能,做一个16:9比例的PPTX文档,内容为SpaceX的发展历程,5页。要求内容丰富、排版复杂豪华精美,有装饰图片和图标,有配图。图片可以联网获取,图片应该尽量使用透明底的图片。查看被引原帖 ↗
小互@xiaohu · 中文博主 · 5 小时前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

Suno 训练库被黑:代码泄露其从 YouTube Music 等抓了约 38 万小时内容

一名黑客用一种叫 Shai-Hulud的蠕虫,搞到了公司员工的登录账号黑进了Suno公司,拿到其训练数据相关的源码。

信息显示,Suno抓取了:

- 113,879 小时的 YouTube Music
- 62,117 小时的 Pond5
- 12,287 小时的 Deezer

资料显示Suno还计划:准备下载大概 100 万小时的播客

入侵里,这名黑客还拿到了 Suno 用户的邮箱、电话,以及支付公司 Stripe 那边的相关信息。涉及的用户大约有几十万。

TechCrunch 在转述中写到,材料里还有部分卡号信息。

Chubby♨️@kimmonismus · 博主 · 10 小时前Chubby,高频 AI 新闻聚合博主

很高兴能与 @XPENG_Global 通用智能中心负责人 Xianming Liu 进行深入交流,他是该公司自动驾驶和 AI 工作的科学领导者。

我们谈论了物理 AI 的真实发展方向,收获颇丰。完整视频即将放出,敬请期待。

查看英文原文
A real pleasure to sit down with Xianming Liu, Head of
@XPENG_Global
's General Intelligence Center and the scientist leading their work on autonomous driving and AI.

We had a fascinating conversation about where physical AI is actually heading, and I came away genuinely enriched by it. Full video coming soon. Stay tuned.
Vercel@vercel · 公司官方 · 11 小时前前端云平台 Vercel 官方,AI 建站工具 v0 母公司

Speechify 的基础设施成了瓶颈。现在他们借助 Vercel 用 Next.js、Cache Components、Instant Rollbacks 这些技术,为 6000 万用户提供动态页面。'没有 Vercel,我们没办法以这个市场要求的速度去竞争。'

查看英文原文
Speechify's infrastructure was a bottleneck. Now they serve dynamic pages to 60 million users on Vercel using Next.js, Cache Components, and Instant Rollbacks.

"Without Vercel, we couldn't compete at the speed this market demands."


vercel.com/blog/how-speechif…
Aravind Srinivas@AravSrinivas · 创始人 · 7 小时前Perplexity 联合创始人兼 CEO

Agent在Vera CPU上跑得特别溜。沙盒运行时和CPU芯片的垂直整合,在大规模部署cloud agent时会带来更好的margin、吞吐量和延迟表现。

引用 NVIDIA AI Infrastructure @NVIDIAAIInfra📣 @perplexity_ai launched SPACE, a secure sandbox platform built for agentic AI. Early tests on NVIDIA Vera CPU showed up to 1.9x faster sandbox starts. Faster starts = less latency, more parallelism, and agents that scale. Learn more now ⤵️查看被引原帖 ↗
查看英文原文
Agents love running on Vera CPUs. And the vertical integration of the sandbox runtime and CPU chip will give margin, throughput, and latency advantages when serving cloud agents at scale.
Cohere@cohere · 公司官方 · 12 小时前加拿大企业级大模型公司 Cohere 官方
连环推 ×2

加拿大AI服务加拿大最聪慧的人才🇨🇦

很荣幸与多伦多大学携手,推动负责任的AI在企业级的应用和采用。我们的主权AI平台North将支撑@UofT的企业级各项能力,并确保敏感数据始终由大学掌控。

查看英文原文
Canadian AI for Canada's brightest minds 🇨🇦

We’re proud to partner with University of Toronto to advance responsible AI adoption at scale. North, our sovereign AI platform, will power
@UofT
's enterprise-wide capabilities and keep sensitive data under the university’s control.
A decade after they first met on campus, we're thrilled to collaborate with the institution that shaped all three of our founders. Investing and building in Canada, Canadian talent, & Canadian potential will always be one of our top priorities.

Read more:
cohere.link/76hU6q8
Tibor Blaho@btibor91 · 博主 · 11 小时前逆向挖掘 AI 产品代码的爆料专家

我们似乎快要接近AGI了,结果我刚花了一个小时设置新HomePod,还是在iOS/tvOS 26.5上。一直报-12004错误,试过多次重置、固件完整恢复、更新……都没用。后来发现,当iCloud和Media & Purchases用不同的Apple账户时就会失败。可iPhone和我其他HomePod都能这样用啊。给遇到这个问题的人个建议——设置时暂时用同一个账户,完成设置再把Media & Purchases改回有Apple Music订阅的那个账户。

查看英文原文
We are apparently close to AGI, yet I just wasted an hour setting up a new HomePod on iOS/tvOS 26.5 because it kept failing with the error "-12004", even after multiple resets, a full firmware restore, an update, etc

Turns out, setup fails when iCloud and Media & Purchases use two different Apple Accounts, even though that configuration works fine on the iPhone and all my other HomePods

Tip for anyone who runs into this - temporarily use the same Apple Account for both during setup, finish the setup, then switch Media & Purchases back to the account with the Apple Music subscription
swyx@swyx · 博主 · 4 小时前知名 AI 播客 Latent Space 主理人

我们每年都确保在 AIE 大会上展示一些顶级 YC AI 公司。今年很荣幸邀请到 @garrytan 和 @eve_bouff 分别为初创公司和设计工程师专场压轴分享。真的很享受这些坦诚而高价值的观点!

引用 AI Engineer @aiDotEngineer🆕 We're so excited to release special double header talks with @ycombinator leaders: @garrytan on Gbrain, Gstack, and the new Physics of Business: invidious.tiekoetter.com/watch?v=eBUyTS7S… @eve_bouff on Imagination Engineering: invidious.tiekoetter.com/watch?v=Z2Erdirp… enjoy!查看被引原帖 ↗
查看英文原文
we make sure some of the top YC AI companies are featured at AIE every single year.

this year, we were graced by
@garrytan
and
@eve_bouff
to cap off our startups and design engineer focused audiences respectively. Really enjoyed these raw high value perspectives!
Lisan al Gaib@scaling01 · 博主 · 10 小时前高频 AI 模型测评与爆料博主

原来是 Kimi 啊

引用 Lisan al Gaib @scaling01I was expecting Kimi K3 to be the first 2-3T model I guess MiniMax M3 Pro at 2.7T will also do查看被引原帖 ↗
查看英文原文
it was Kimi after all
Ethan Mollick@emollick · 创始人 · 13 小时前沃顿商学院教授,AI 应用研究权威

这真是个'神经质'的模型!(不过这不意味着它不好)

查看英文原文
Its a really neurotic model! (Which doesn't mean its bad)
The Rundown AI@TheRundownAI · 博主 · 10 小时前百万订阅 AI 日报官方

明天:学习如何用 ChatGPT 把源文件和研究转化为精美的作品。

在这个免费的90分钟直播大师课中,我们的大学教育工作者 Nate Grahek 将演示如何使用 ChatGPT Work、GPT-5.6 和 Codex 创建精美、可编辑的演示文稿。

下方报名:

查看英文原文
Tomorrow: learn how to turn source files and research into polished work with ChatGPT.

In this free, live 90-minute masterclass, our University Educator Nate Grahek will walk through using ChatGPT Work, GPT-5.6, and Codex to create a sleek, editable presentation.

RSVP below:
Tibor Blaho@btibor91 · 博主 · 9 小时前逆向挖掘 AI 产品代码的爆料专家

一直观察 AIPRM 对接 ChatGPT 和 Claude,同时盯着这两家公司一路看我发现特有意思——OpenAI 和 Anthropic 团队从外部看,好多方面风格完全不同,但另一些方面又出奇地像。

不只是团队调性和你联系得到回复的概率,连那些版本变动里能挖到的技术细节也很有意思。

最近典型的例子:ChatGPT Work 和 Claude Cowork。两个功能极其相似,却是完全对着干的方式搞出来。

ChatGPT Work 跟系统原有聊天体验贴合非常紧,之前在聊天模式上能操作的东西基本都能继续用。

Claude Cowork 网页版就跟普通聊天风格大不相同,之前在聊天里行得通的方法下去基本乱七八糟,倒像是在向 Claude Code 接近。

与此同时,两队隐藏功能开关配置、给疑似"内部"字符串和子域名做乱码处理的方法、以及对新构建内容里未发布功能的泄露控制,几乎沿用了同一种套路。

包括一个共同的烂 bug 反复出——因为写成 "!{value: false}" 而不是 "!value" 导致功能锁失灵。这种雷我在两家都不止一次见到了,你真想不到看过多少次。

查看英文原文
Since I have been working on AIPRM for ChatGPT & Claude and watching both companies for a while now, it's very interesting to me how different the OpenAI and Anthropic teams are in many aspects (from the outside), and how similar in others

Not just in vibes and how much you hear back when you reach out, but even in technical stuff you can see when comparing changes between deploys

One of the most recent examples - ChatGPT Work and Claude Cowork, two quite similar features built in completely opposite ways

ChatGPT Work stays very close to the existing chat experience, so almost everything that worked in chat mode before still works there

Claude Cowork in the web app is very different from chat, so almost everything that worked in chat before breaks there, and it's a lot closer to how Claude Code works instead

Meanwhile both teams follow almost the same playbook to "hide" feature gate configs, redact possibly "internal" strings and subdomains, and reduce leaks of unreleased features from new builds

Down to shipping the exact same defect over and over, feature gates failing because of !{value: false} instead of !value, and you don't want to know how many times I have seen this one on both sides
Cohere@cohere · 公司官方 · 8 小时前加拿大企业级大模型公司 Cohere 官方

我们联合创始人@nickfrosst十年前还是@UofT的学生研究员,如今他正帮助领导加拿大的主权AI领军企业。

他会给多伦多大学的学生创业者提什么建议呢?

查看英文原文
10 years ago, our co-founder
@nickfrosst
was a student researcher at
@UofT
. Today, he helps lead Canada’s sovereign AI champion.

What would his advice be to a student founder at University of Toronto today?
AK@_akhaliq · 博主 · 4 小时前HuggingFace 研究员,每日 AI 论文速递

Inkling 现已在 Claude Code 中通过 Hugging Face Claude 提供

查看英文原文
thinkingmachines Inkling is now available in claude code via hf claude
Cristóbal Valenzuela@c_valenzuelab · 创始人 · 14 小时前Runway 联合创始人兼 CEO

Runway 的 Agent 2.0 在叙事连贯性、电影语言和制作质量方面都达到了业界顶尖。Agent 视频生成正在成为一个全新的产品类别,为全新的用户群和应用打开了大门。

引用 Physion Labs Official @Physion_Labs🎬 Video agents can now generate minute-long videos. But can they actually direct? Today, we’re launching 𝐏𝐡𝐲𝐬𝐢𝐨𝐧-𝐀𝐫𝐜 1.0, a new benchmark evaluating complete, multi-scene videos across narrative coherence, cinematic language, and production quality. We tested @runwayml , @LumaLabsAI , @MiniMax_AI , @Kling_ai , @UtopaiStudios and @TapNow_AI on 100 screenplays and 600 generated videos. 🏆 𝐑𝐮𝐧𝐰𝐚𝐲 𝐀𝐠𝐞𝐧𝐭 2.0 𝐫𝐚𝐧𝐤𝐞𝐝 𝐍𝐨. 1 𝐨𝐯𝐞𝐫𝐚𝐥𝐥 𝐚𝐧𝐝 𝐥𝐞𝐝 𝐞𝐯𝐞𝐫𝐲 𝐞𝐯𝐚𝐥𝐮𝐚𝐭𝐢𝐨𝐧 𝐝𝐢𝐦𝐞𝐧𝐬𝐢𝐨𝐧. Its advantage was especially clear across subjective metrics, where cinematic taste matters most. Runway ranked first on all eight. 🔗 Read the full benchmark: lnkd.in/g-5tcwK7查看被引原帖 ↗
查看英文原文
Runway's Agent 2.0 is state of the art in narrative coherence, cinematic language, and production quality. Agentic video generation is becoming an entirely new product category, opening the door to a whole new set of users and applications.
Bindu Reddy@bindureddy · 创始人 · 2 小时前Abacus.AI CEO,AI 行业观点博主

Kimi K3 缩小了差距但仍排在前沿模型之后

我们的 LiveBench 基准有很多隐藏问题,所以模型很难靠记忆取巧

K3 是最好的开源模型,但表现不如 Opus 4.8、Sol 和 Fable

在实际应用中,Kimi 输出很冗长,处理接近 Opus 级别问题的成本和 Opus 4.8 一样,但速度慢得多

查看英文原文
KIMI K3 CLOSES THE GAP BUT RANKS BEHIND FRONTIER MODELS

Our benchmark, LiveBench, has a lot of hidden questions so that models it's hard to memorize it

K3 is the best good open-source model but is below Opus 4.8, Sol and Fable

Also in practice, Kimi spins a lot and costs as much as Opus 4.8 for near opus-class problems - it's also much slower
Lisan al Gaib@scaling01 · 博主 · 8 小时前高频 AI 模型测评与爆料博主
连环推 ×2

这是IPO后的第一次星舰发射吧?

如果火箭爆炸的话,看市场怎么反应会很有趣

查看英文原文
this is the first starship launch after the IPO right?

will be interesting to see how markets react, in case the rocket blows up
or a scrapped launch
Lisan al Gaib@scaling01 · 博主 · 12 小时前高频 AI 模型测评与爆料博主

一直很骄傲能被说成是Kimi的"吹捧者"

他们真的不辜负中国最厉害的研究实验室这个称号

引用 Lisan al Gaib @scaling01kimi.com/blog/kimi-k3查看被引原帖 ↗
查看英文原文
very proud to have been a Kimi "shill" for a long time

they truly are the greatest lab in China
Lisan al Gaib@scaling01 · 博主 · 4 小时前高频 AI 模型测评与爆料博主
连环推 ×3

算了吧,拿出点胆量来
做出你的预测吧

不像 Nathan 发的那样:
"开源和闭源模型之间的差距 ~0 个月"
然后又删了

引用 Lisan al Gaib @scaling01MoonshotAI will overtake OpenAI and Anthropic before the end of the year or will they? at least that's what the hype kiddies on X want you to believe So let me make it falsifiable. They are saying: - China / MoonshotAI is catching up - they are catching up generally (not just coding, but almost all domains and including restricted models like Mythos 5) - the gap is currently ~1.4 months based on Artificial Analysis Index and benchmarks provided by MoonshotAI, where Kimi-K3 beats Opus 4.8 in 30 of 35 benchmarks, and GPT-5.6-Sol in 19 of 35 benchmarks (they ignore the existence of all Mythos variants) - China is not catching up due to distillation, so they should overtake US labs Their implicit prediction then is: - a chinese model / MoonshotAI will overtake Anthropic (and OpenAI) on the Artificial Analysis Index by: - Median: 2026-12-24 (80% CI: 2026-09-17, 2027-09-14) Since they claim that chinese models are as general as american models, we should see unsolved mathematics, physics, and more being solved by chinese models at higher rates than american models. Speaking to its generality Kimi-K3 should surpass Opus 4.8 and GPT-5.6-Sol on the majority of these benchmarks: - METR Time Horizons, FrontierCode, MirrorCode, UK AISI cyber ranges, ExploitBench/ExploitGym, CritPT, FrontierMath T4, ARC-AGI-2 / ARC-AGI-3, WeirdML, ALE-Bench, GSO, MRCR2/GraphWalks - vibes --- Some other things that are more speculative and downstream of China overtaking US models: - more involvement by the USG - stricter export controls on semis - potentially a Manhatten-style project, as we will be behind in 2027 and are racing against China - also in the cards: US banning chinese models or US labs distilling from chinese models --- I have already stated my position clearly. Chinese models are generally ~6-8 months behind, with some domains like coding behind slightly less. Kimi-K3 did not significantly shift my estimate on the gap and it currently does not change my outlo查看被引原帖 ↗
查看英文原文
come on, show some balls
make your predictions

not like Nathan tweeting something along the lines:
"gap between open and closed models ~0 months"
and then deleting it again
i hope you have already seen the issue here

they are including Mythos in their statements but then completely ignore it when it comes to the actual benchmarks
so what if those things don't happen?

well, then Chinese models are not catching up and further behind than just 1.4 months and/or the Artificial Analysis Index is in fact not measuring general capability but just a slice
or chinese models are in fact distilling and therefore can't overtake and pull away from US models

i had to keep this short. I need to sleep.
Kling AI@Kling_ai · 公司官方 · 7 小时前快手旗下可灵 AI 视频官方

一颗樱桃种子承载了一整段记忆 🍒

踏入乔瑞英的《The Well》—— 一个关于童年和我们珍藏时刻的安静而强有力的故事。

祝贺乔瑞英荣获 Kling AI NEXTGEN 2026 韩国大学创意挑战最佳叙事奖!

查看英文原文
A cherry seed held an entire memory 🍒

Step into Jo Ryeongmi’s “The Well” — a quiet yet powerful story about childhood and the moments we hold onto.

Congratulations Jo Ryeongmi on winning Kling AI NEXTGEN 2026 Korea University Creative Challenge Best Storytelling!
AshutoshShrivastava@ai_for_success · 博主 · 5 小时前高频 AI 新闻与产品动态博主

Kimi K3 轻松击败 Fable 5。
目前在 Next.js Agent Performance Benchmark 排名第一,性能远超 Fable 5。

引用 Guillermo Rauch @rauchgKimi K3 is the best performing model on nextjs.org/evals , ahead of Fable, reaching a comparable success rate in less time. This is the first time that an open model is ahead of all proprietary ones for this comprehensive web engineering benchmark. Notes: ▪️ Benchmarks don’t always tell the full story, although this is important signal, adding to mounting evidence that this could be a breakthrough moment for open models ▪️ No model as of yet has reached 100% completion on this set of evals. The top performer peaks at 92% and 96% “with help”查看被引原帖 ↗
查看英文原文
Kimi K3 is eating Fable 5 for breakfast.

It's currently the top model on the Next.js Agent Performance Benchmark, outperforming Fable 5.
Gorden Sun@Gorden_Sun · 中文博主 · 3 小时前中文圈高频 AI 资讯与开源项目博主

Codex又开始了新一轮的老带新活动,这次最多可获得10次重置次数

AK@_akhaliq · 博主 · 8 小时前HuggingFace 研究员,每日 AI 论文速递

论文:
huggingface.co/papers/2607.1…

查看英文原文
paper:
huggingface.co/papers/2607.1…
Orange AI@oran_ge · 中文博主 · 6 小时前Orange AI,中文圈 AI 产品观察博主

我实测 glm 5.2 和 deepseek 都是中文的

Bindu Reddy@bindureddy · 创始人 · 9 小时前Abacus.AI CEO,AI 行业观点博主

哈哈,局势彻底反转了!OpenAI 的模型在设计上很强

5.6 Sol 用于设计
Fable 5 用于编码

这是目前的最强组合。

Opus 几周前还擅长设计呢

查看英文原文
LOL, the tables have totally turned! OpenAI models are very good at design

5.6 Sol for design
Fable 5 for coding

This is the current killing combination.

Opus used to be good at design just weeks ago
Bindu Reddy@bindureddy · 创始人 · 5 小时前Abacus.AI CEO,AI 行业观点博主

Opus 5 近在咫尺了

4.8 现在有点过时了

查看英文原文
Opus 5 is literally a couple of days away

So 4.8 is kinda legacy now

本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报

姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档