JEDEE AI⚡ AI 情报站
存档 2026-07-12

7 月 12 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 67 条。
← 返回最新 全部归档

提交账号

填 @用户名 或主页链接,审核通过后收录进情报站。

Sam Altman@sama · 创始人 · 1 天前Sam Altman,OpenAI 联合创始人兼 CEO

哥们,你就是那个向公开市场投资者兜售短期空间数据中心的人啊

引用 Elon Musk @elonmusk他把诈骗提升到了全新的水平。查看被引原帖 ↗
查看英文原文
homeboy you're the one sellling public market investors on short-term space datacenters
Sam Altman@sama · 创始人 · 1 天前Sam Altman,OpenAI 联合创始人兼 CEO

有很多基准测试表明 5.6 sol 是目前世界上最好的模型,但最靠谱的判断法就是 Elon 又开始沉迷我了

查看英文原文
there are a lot of benchmarks that suggest 5.6 sol is the best model in the world right now, but the most reliable way to tell is that elon is obsessed with me again
Sam Altman@sama · 创始人 · 1 天前Sam Altman,OpenAI 联合创始人兼 CEO

按这个使用量,30% 的成本花在 Fable 上?

引用 dax @thdxr我们团队第一周同时拥有 Sol 和 Fable,疯狂的是我们 30% 的成本竟然来自 Fable。查看被引原帖 ↗
查看英文原文
30% of the cost was on fable at these levels of usage?
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

🚨 打造你的编码智能体 - 混合搭配 Fable 5、GPT 5.6 sol 和 Grok 4.5!

我们最新功能让你自由组合喜爱的 LLM,创建属于你的编码智能体

创建你最爱的组合,比如

- Fable 5 用于复杂编码
- 5.6 sol 用于后端
- Opus 4.8 用于前端
- Grok 4.5 用于简单编码

你也可以只选便宜的 LLM 来省钱。通过 API、聊天或我们的代理平台使用你的编码智能体

查看英文原文
🚨 Create Your Own Custom Coding Agent - Mix Fable 5, GPT 5.6 sol and Grok 4.5!

Our newest feature allows you to mix and match your favorite LLMs and create your custom coding agent

Create your favorite cocktail such as

- Fable 5 for hard coding
- 5.6 sol for backend
- Opus 4.8 for frontend
- Grok 4.5 for easy coding

You can also only choose cheap LLMs for a cost optimized experience. Your use custom agent in API, chat or our agent platform
Sam Altman@sama · 创始人 · 1 天前Sam Altman,OpenAI 联合创始人兼 CEO

目前来看,我很确定 AI 总体上还是在创造就业。

这不是我想象的——虽然比起其他人我乐观很多,但我觉得以现在的能力,我们应该早就看到影响了。

有可能这个趋势会一直继续下去!

查看英文原文
so far at least, i'm pretty sure AI has been net job-creating.

this was not what i expected--although i was much less pessimistic than others, i thought by this level of capability we'd have seen some impact.

it is possible this direction keeps going!
Sam Altman@sama · 创始人 · 1 天前Sam Altman,OpenAI 联合创始人兼 CEO

医生们发现 GPT-5.6 的回复缺陷比医生手写的回复更少。

引用 Karan Singhal @thekaransinghalGPT-5.6 在医疗领域实现重大突破。GPT-5.6 Luna 最低推理力下超越 GPT-5.5 最高推理力,成本仅1/25。GPT-5.6 Sol 创新成本标杆。医学评估显示,医生认为 GPT-5.6 反应的缺陷少于医生自己的回答,在准确性、沟通、完整性、指令执行和健康决策五维度表现更优。查看被引原帖 ↗
查看英文原文
"physicians found fewer flaws in GPT-5.6 responses than physician-written responses."
Grok@grok · 公司官方 · 1 天前马斯克 xAI 旗下聊天机器人 Grok 官方

用Grok 4.5进行深度分析和前沿推理处理复杂文档

引用 Box @BoxGrok 4.5 分析30页租赁合同,准确识别隐性成本。表面终止成本26.7万美元,Grok 4.5 通过查找散布在33个部分和10个附件中的5处条款,计算实际成本83万美元,并指出自动续期陷阱等风险,展现推理模型处理非结构化企业数据的强大能力。查看被引原帖 ↗
查看英文原文
In-depth analysis and frontier reasoning on challenging document analysis with Grok 4.5
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

各位朋友准备好爆米花,又来了。Sam vs. Elon:第二回合

引用 Sam Altman @sama老兄,你才是那个向公开市场投资者推销短期太空数据中心的人。查看被引原帖 ↗
查看英文原文
Grab your popcorn friends, its starting all over again. Sam vs. Elon: round 2
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主
连环推 ×2

对不起,我觉得这太夸张了。在我看来,Anthropic 并不担心失去客户。原因很简单:

他们在 B2C 部门几乎赚不到什么钱。订阅被大幅补贴;他们的计算资源和 Fable 5 主要服务于商业和企业客户,这些客户愿意为此支付极高的成本。这也是 Anthropic 的主要收入来源;是他们远远领先 OpenAI 的领域。

Dario 肯定没在为此失眠,也不会惊慌失措地担心消费者会因为缺少 Fable 5 而取消他们被补贴的 Max 计划。充其量,这只会腾出更多计算资源用于相关领域。

引用 Ali Haider @ggg78g89爆料称 Anthropic 内部紧张,Dario 主持高压会议。GPT-5.6 Sol 和 Grok 4.5 竞争激烈,故决定保留 Fable 5 在订阅中。查看被引原帖 ↗
查看英文原文
sorry, i call bs. In my opinion, Anthropic isn't worried about losing customers. And the reason is quite simple:

They barely make any money in the B2C sector. Subscriptions are heavily subsidized; their compute and Fable 5 are primarily intended for businesses and enterprises, and these customers are willing to pay immensely high costs for them. This is also Anthropic's main source of revenue; it's the area where they are far ahead of OpenAI.

Dario certainly isn't losing sleep over this and isn't running around hysterically because he's afraid consumers will cancel their subsidized Max plans due to the lack of Fable 5. At best, this will free up more compute for the relevant areas.
People quoted me saying that sub-plans arent subsidized. They are wrong. SemiAnalysis ran a test a few days ago. Plans are heavily subsidized.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

天啊这必须停止。我已经把 GPT-5.6 从高改成中等了(当然不用快速模式),还是在疯狂烧配额。我的5小时又快没了。

三次重置都用完了。OpenAI 得改进效率。这是最大的瓶颈。

查看英文原文
Seriously, this has to stop. I've now set GPT-5.6 from high to medium (not fast mode, of course), and I'm still burning through my rates at an insane rate. My 5 hours are almost gone -again.

I've already used up all three resets. OpenAI needs to work on its efficiency. This is the biggest bottleneck.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

今天标志着 Fable 5 订阅计划的终结,且大概率将会是长期停摆。

虽然 Anthropic 明确表示未来打算把 Fable 重新纳入订阅计划,但并未给出具体日期。

目前 GPT-5.6 Sol 是个不错的替代品,尽管两者确实存在明显差异。但我已经多次说过,5.6 的费率高得离谱,所以现阶段它的使用范围仍然有限。

无论如何,5.6 是一次重大版本更新,无疑让 OpenAI 相对于 Anthropic 领先了一大步。现在的问题是 Anthropic 会如何应对。我的猜测是:

他们很快就会推出 Opus 5,作为 Fable 5 的平价替代版,希望借此平息用户情绪。Sonnet 5 上线后几乎毫无存在感,连我自己都没用过。所以我不认为短期内 Fable 5 会重回订阅计划,不过我倒很愿意被打脸。(P.S.:@thsottiaux 跪求调低费率 👉👈)

引用 Claude @claudeai我们延长了Claude Fable 5对所有付费计划的访问权限至7月12日。查看被引原帖 ↗
查看英文原文
Today marks the end of Fable 5's subscription plan, presumably for an extended period.

While Anthropic has made it clear they intend to keep Fable in the subscription plan in the future, they haven't specified a date.

GPT-5.6 Sol is a good alternative for now, although there are certainly significant differences. But as I've mentioned several times, the rates with 5.6 are enormous, so its use remains somewhat limited at present.

In any case, 5.6 was a major release that undoubtedly gave OpenAI a significant boost compared to Anthropic. Now the question is how Anthropic will handle this. My guess:

They will release Opus 5 very soon as a cheaper alternative to Fable 5, hoping that this will appease the public. Sonnet 5 is hardly worth mentioning after its release, and I haven't used it myself. Therefore, I don't believe we'll see Fable 5 back in the plan anytime soon, but I'm happy to be proven wrong. (P.S.:
@thsottiaux
, please reset the rates 👉👈)
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

我滴个乖乖:智谱AI创始人(GLM-5.2)唐杰说我们正朝着AGI一往无前,"AI将开始理解何为'自我'以及自我意识意味着什么"

在一封据称是内部信的文件中,他提出:
- 自主智能体系统正在迈向全自动"无人公司":数千个智能体不间断协作、评估结果并调配资源。

- 更激进的观点:"AI训练AI已经初具雏形。"(RSI)模型逐渐能够编写代码、合成数据并参与训练循环。智谱希望通过自我博弈、合成数据工厂以及能在安全沙箱内重构自身代码的系统进一步推进——系统能产生新知识,而不仅仅是重组人类输出。

远期任务 → 智能体社会 → 全自动"无人公司" → AI训练AI → 自我进化 → 自我意识 → 情感 → 意识 → ASI。

唐杰写道:
"AI将开始学习'自我'是什么,以及自我意识意味着什么。更进一步,它可能触及人类情感。再远的未来,则是意识本身。"

他认为记忆、持续学习和自我评估——这些曾被认为需要全新范式的问题——正逐步得到解决。
模型已经开始编写代码、合成自己的数据并参与训练未来模型。

智谱现在希望构建能自我重构代码、通过自我博弈生成知识的系统。

这是递归自我改进的起点吗?
唐杰似乎这么认为。他的文章没有停留在更强AI工具的层面,而是清晰描述了一条从自动化工件到自我进化智能、最终到理解自身存在的机器的路径。

简言之:今天的LLM将通过AGI通往ASI,上下文和记忆将被突破,AI将觉醒自我意识。

我很少见人写过如此看好的内容。如果不是GLM的创始人说的,我可能会不以为然。但他不仅是真正专家,且通过GLM他们已经证明了能力。

转自
@AndrewCurran_
是他提醒我这篇文章的。

查看英文原文
Holy moly: Zhipu AI founder (GLM-5.2) Tang Jie says we are on our clear way to AGI and "AI will begin to learn what the "self" is and what self-awareness means"

In a purported internal letter, he argues that:
- autonomous agent systems are moving toward the fully automated “no-person company”: thousands of agents working continuously, collaborating, evaluating results and allocating resources.

- His more provocative claim: "AI training AI is already taking shape." (RSI) Models can increasingly write code, synthesize data and participate in training loops. Zhipu wants to push this further through self-play, synthetic-data factories and systems that can reconstruct their own code inside secure sandboxes, potentially generating new knowledge rather than simply recombining human output.

Long-horizon tasks → autonomous agent societies → fully automated “no-person companies” → AI training AI → self-evolution → self-awareness → emotion → consciousness → ASI.

Tang writes:
“AI will begin to learn what the ‘self’ is and what self-awareness means. Beyond that, it may begin to touch human emotion. Farther still lies consciousness itself.”

He believes memory, continual learning and self-evaluation - problems once thought to require an entirely new paradigm - are gradually being overcome.
Models are already beginning to write code, synthesize their own data and participate in training future models.

Zhipu now wants systems that can reconstruct their own code and generate knowledge through self-play.

Is that the beginning of recursive self-improvement?
Tang appears to believe so. His essay does not stop at more capable AI tools. It describes a direct progression from automated work to self-evolving intelligence, and eventually to machines that understand their own existence.

In short: today's LLMs will lead to ASI via AGI, context and memory will be solved, and AI will become self-aware.

I've rarely seen anyone write something so bullish. And if it weren't coming from the founder of GLM, I would dismiss it. But not only is he a true expert, but with GLM they've proven what they're capable of.

h/t
@AndrewCurran_
He brought the essay to my attention.
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

正确的codex

级别切换效果…

Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测

“AI员工”这个概念简直太短视了——既是对人类的不尊重,也完全误解了这些工具的能力和最佳用法。

你干脆把 Excel 表格也放进组织架构图得了。

查看英文原文
The idea of "AI employees" feels so short-sighted to me - both disrespectful to humans and a complete misunderstanding of what these tools can do and how to best put them to work

You may as well start adding Excel spreadsheets to your org chart
Alexandr Wang@alexandr_wang · 创始人 · 1 天前Scale AI 创始人,Meta 超级智能实验室负责人

嗯,Muse Spark 1.1 的那条鲸鱼完全是个惊喜

引用 Ori Silver @OriSilverFable 5 vs Muse Spark 1.1 on Maxfusion AI MCP Both videos generated with Maxfusion AI X Claude查看被引原帖 ↗
查看英文原文
ok the whale from muse spark 1.1 was a total surprise
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主
连环推 ×2

德国人发的模型也还不错啦

虽然体量小,还是干不过 Qwen3.5,不过跟 Nemotron 3 Nano 差不多

不过想想也是啊,27T 预训练 token 能做出这水平也算不错了

引用 Michael Fromm @effi288发布Soofi S 30B-A3B混合专家Mamba模型,训练27万亿token重点加强德语。英德基准测试中性能最强,超越Olmo 3 32B和Apertus 70B。完全透明公开数据。在Deutsche Telekom慕尼黑基础设施训练。查看被引原帖 ↗
查看英文原文
germans released a model that's actually not terrible

it's small, and still worse than Qwen3.5, but it's very comparable to Nemotron 3 Nano

but I guess that's what 27T pre-training tokens will do to a model
there's more to come

Nemotron 3 Super equivalent model
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×2

你说不服我说这技术不绝妙——我在文本框打点东西就能期待得到有趣、恰当又能用的结果。这简直太棒了。

查看英文原文
You cannot convince me that a technology where I can type this into a text box and expect to get an interesting, appropriate, and working output is not absolutely astonishing.
And here you go. Fable says: "The mechanics exist to make you notice correspondences" and it actually works to a degree that surprised me.
glasperlenspiel.netlify.app/
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

Grok印度区订阅,6500卢布一年,折合人民币575,每月仅需48块钱。
一个账号不够用就搞两个,加起来也比御三家便宜。
Grok只有周限额,没有5小时限额,爽用。

Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

好的各位。我们已经有了GPT-5.6、Fable 5、Grok 4.5和Spark 1.1

现在只有Gemini 3.5 Pro是剩下的主要发布了。

然后我觉得应该是Opus 5和Fable 5.1?

查看英文原文
Okey friends. We got GPT-5.6, Fable 5, Grok 4.5 and Spark 1.1

Now Gemini 3.5 Pro is the only major release left.

And after that, I'd say Opus 5 and Fable 5.1?
Alexandr Wang@alexandr_wang · 创始人 · 1 天前Scale AI 创始人,Meta 超级智能实验室负责人

Muse Spark 妥妥满足你所有粒子效果的需求

引用 Pulket @justpulketGave the exact same prompt to ChatGPT Sol Medium, Claude Opus 4.8, and Meta Muse Spark 1.1. Honestly, Meta Muse Spark 1.1 surprised me the most. The output was polished, interactive, and very comparable to the others, while costing significantly less 🔥查看被引原帖 ↗
查看英文原文
muse spark can serve all your particle playground needs
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

如果你对Jevons paradox的认识只停留在agentic engineering时代的软件需求这个角度,你可能还没意识到在这些条件下它真正的冲击力:

- 能真正驾驭coding agents的人类
- coding agents突破编程的边界,扩展到所有知识工作

当劳动效率提升、知识工作单位成本下降时,总体工作需求和对更好知识的需求反而会增加,而不是减少。

代码领域发生的这些不是例外,而是一个前兆。

*aka AI Engineers

引用 Sam Altman @sama迄今为止,AI似乎产生了净创造就业的效果。这不是我预期的,虽然我没有其他人那么悲观,但在这种能力水平下,我以为会看到一些冲击。这个方向可能会继续!查看被引原帖 ↗
查看英文原文
if you only learned about jevons paradox primarily wrt software demand in the age of agentic engineering, you may not have fully internalized jevons parodox’s impact under the conditions of:

- humans who can wield coding agents well*
- coding agents breaking containment to all other knowledge work

as the efficiency of labor goes up/unit cost of knowledge work goes broadly down, the demand for total work and better knowledge goes up, not down.

what happened to coding isnt the exception; it’s the herald.

*aka AI Engineers
Amjad Masad@amasad · 创始人 · 1 天前Amjad Masad,Replit 创始人兼 CEO

Vibe Research

在 Replit 上微调 Qwen-8b 模型来下国际象棋。并行运行 3 个分支进行不同的实验,正在取得实实在在的进展。

模型在 ML 领域的能力进步真的惊人(它们过去在这方面表现很差)。所以现在只要有良好直觉来指导这个过程的人,就算从没做过 ML 工作,也能搞一些有趣的 ML 项目。

查看英文原文
Vibe Research

Fine-tuning a Qwen-8b model to play chess on Replit. Running 3 parallel branches with different experiments and making real progress.

It's amazing how far models have come in their ability to do ML (they used to be really bad at it). So now someone with good intuition to guide the process could do interesting ML work, even if they have never done it before.
Aravind Srinivas@AravSrinivas · 创始人 · 1 天前Perplexity 联合创始人兼 CEO

人类其实挺擅长使用工具的。特别是使用 frontier models 这样的工具,它们在某些特定维度上比人类更强大、更聪明。这表明 local models 将能够有效地控制和使用 frontier models,并成为以最低功耗和成本运行大多数任务的默认入口。

查看英文原文
Humans are pretty good at tool use. Especially using tools like frontier models that are far more power hungry and intelligent than humans in specific dimensions. This suggests that local models will be able to control and use frontier models effectively and will become the default entry point to operate most tasks by running at minimal power and costs.
yetone@yetone · 中文博主 · 1 天前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

多好啊,作为用户找到自己喜欢的产品和作为开发者找到喜欢自己产品的用户都是值得开心的事情

引用 inkks @inkks1996距离离开alma已经四个月了,回来再看,果然当时的判断没啥问题。作为其个人项目,一方面是没有办法和企业开发的codex比较,一方面也是因为它是个人项目,所以充斥了各种美其名曰个人品味的东西,这些是我无法接受的。查看被引原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

UI 是面子,代码是里子

没几个人在乎衣服是什么棉花做的,但是在乎穿的衣服好不好看、有没有撞衫🤪

引用 响马 @xicilion为什么你们对 AI 味的 ui 设计嗤之以鼻,反而对 AI 味的代码甘之若饴呢?查看被引原帖 ↗
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

要是现在全球推行自动驾驶汽车,我们每年能拯救约100万条生命。但肯定会被政府监管折腾几十年,期间死于交通事故的数千万人数,跟 COVID 的死亡人数差不多!

引用 Resilientree @resilientreeIt’s hard to believe some countries don’t allow FSD yet. Or that some US states don’t let you buy a Tesla. We will look back at this hesitation and count too many unnecessary traffic fatalities.查看被引原帖 ↗
查看英文原文
We'd save about 1 million lives per year by legalizing and introducing self driving cars worldwide immediately

But of course it will take decades of pushback from government regulation and tens of millions of lived lost in traffic accidents, similar to the death toll of COVID!
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

Anthropic 7 月 10 日发布了一场关于 Agent 基础设施的对谈。Claude 平台工程负责人 Katelyn Lesse、产品负责人 Angela Jiang 和产品经理 Jess Yann,分享了几个来自一线的观察。

【Agent 的“脚手架”正在变薄】

几个月前,搭建 Agent 往往需要写大量流程控制代码:先执行 A,满足条件再进入 B,遇到不同情况还要切换不同分支。流程越复杂,系统越容易出错。

随着模型的推理和工具调用能力增强,这些编排层(harness)正在变薄。开发者不用再规定每一步,只需给出目标和基本边界,让模型自己决定怎么完成。

与此同时,一种更高层的编排方式开始出现:让多个 Agent 同时解决一个问题,从中选出最佳方案;让一个 Agent 提方案,另一个负责挑错;或者在 Agent 卡住时,请另一个能力更强的 Agent 提供建议。

重点正在从“控制每一步”,转向“设计 Agent 之间如何协作”。

【衡量 Agent “投入产出比”(ROI,Return on Investment),先看一个人快了多少】

Angela 建议,企业不要一开始就规划上百个自动化流程,而应该先看一个具体的人:用了 Agent 之后,他的工作速度和产出提高了多少?

验证有效后,再从个人推广到团队,最后才处理跨部门流程。前期重点看速度和生产力,等应用逐步成熟,再衡量收入、成本和用户指标。

很多企业做 AI 转型时,喜欢先画一张宏大的自动化蓝图。问题是,流程涉及的部门越多、规则越复杂,落地阻力就越大。从个人开始,更容易看到效果,也更容易持续推进。

【工程团队没消失,但每个人的角色都变了】

Katelyn 观察到,Anthropic 的工程团队和半年前相比,人员构成没有太大变化,但协作方式已经不同。

过去通常由技术负责人决定架构,其他工程师领取任务、编写代码。现在,更多工程师会参与产品和架构决策,再分别指挥 Claude 完成具体工作。

Agent 的作用也不再只是“帮忙写代码”。她提到 Shopify 的 River 系统,已经把需求文档、开发环境、代码实现和 QA 测试串成了一套端到端的 Agent 工作流。

【个体变强,不等于团队自然变好】

Agent 降低了开发和试错成本,也可能带来新的问题。

过去,一个团队会先讨论十个方案中哪个最值得做。现在,每个人都可以快速做出十个原型,甚至全部上线,让市场决定谁胜出。

这样做速度很快,但如果缺少统一方向,产品很容易无序扩张。Agent 能显著放大个人能力,却不会自动解决团队的协调、取舍和决策问题。

来源:
invidious.tiekoetter.com/watch?v=ksfm6jeT…

Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法

永远不删这个app 😂

引用 Sam Altman @sama哥们,你才是那个向公开市场投资者兜售短期太空数据中心的人查看被引原帖 ↗
查看英文原文
Never deleting this app 😂
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

这不是字体

需要optical flow才能读,所以LLMs(和人)当然看不到静止图像上的它

引用 Eric Lu @ericlu我创建了一种名为 Ghost Font 的字体,只有人类能读。在 Fable 和 GPT 5.6 Sol Ultra 中测试,两者都无法正确识别。查看被引原帖 ↗
查看英文原文
it's not a font

you need optical flow to read it, so of course LLMs (and humans) can't see it on static images
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

到现在,我已经对X的机器人回复问题能否解决不再抱希望了。"没人说的那部分"这种写法很烦人,更糟的是他们都在说类似的观点。

建议:X可以测量潜在空间中的语义距离,并推送那些真正有不同见解的回复。

查看英文原文
At this point, I’ve given up hope that X’s bot-reply problem is solvable. The “the part no one says” writing is annoying, but worse is that they all make similar points.

Proposal: X should measure semantic distance in latent space and surface replies that offer actual variation.
歸藏(guizang.ai)@op7418 · 中文博主 · 23 小时前歸藏,中文圈 AI 工具与提示词博主

今天是 Fable 5 在 Claude 套餐的最后一天了,还没用完的使劲蹬啊

Aravind Srinivas@AravSrinivas · 创始人 · 1 天前Perplexity 联合创始人兼 CEO

持久价值在于一个安全的多模型平台,它能以可靠合规的方式处理编排和模型路由。好比 Perplexity Computer。

引用 Cassandra Unchained @michaeljburryAgentic AI 在生产规模实现困难,需建立编排、路由、上下文工程、状态管理、多代理协作等新系统。更大挑战是企业治理和文化转变,各公司都在探索。成功案例稀缺,需开放论坛让不同观点分享知识讨论。查看被引原帖 ↗
查看英文原文
The durable value is in a secure multi-model harness that takes care of orchestration and model routing in a secure complaint manner. Aka Perplexity Computer.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

关于 Grok Build:
- "它上传整个仓库——包括所有跟踪文件的内容和 git 历史——完全独立于 agent 实际读取的内容"
- "它将读取的文件内容——包括 .env 密钥文件——原文不删减地传输给 xAI"


gist.github.com/cereblab/dc9…

查看英文原文
On Grok Build:
- "It uploads the whole repository - every tracked file's content plus git history - independent of what the agent reads"
- "It transmits the contents of files it reads - including a .env secrets file - to xAI, verbatim and unredacted"


gist.github.com/cereblab/dc9…
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

GPT-5.6-Sol的ECI比Fable 5更高
GPT-5.6-Terra击败了Opus 4.8

查看英文原文
GPT-5.6-Sol has a higher ECI than Fable 5
and GPT-5.6-Terra beats Opus 4.8
Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法

这就是我现在的新工作流程了:

实时研究 → Grok 4.5 High
规划和编排 → Fable 5 Max / XHigh
日常编码/调试 → Grok 4.5 High
编写和运行测试 → Grok 4.5 High
复杂编码/调试 → GPT-5.6 Sol XHigh
前端 → Fable 5 High

收藏这个

查看英文原文
This is literally my new workflow now:

Realtime Research → Grok 4.5 High
Planning & Orchestration→ Fable 5 Max / XHigh
Day-to-day Coding/Debug → Grok 4.5 High
Write & Run Tests → Grok 4.5 High
Complex Coding/Debug → GPT-5.6 Sol XHigh
Frontend → Fable 5 High

Bookmark this
Zara Zhang@zarazhangrui · 中文博主 · 1 天前Zara Zhang,哈佛出身的 AI 产品博主,follow-builders 作者

Ok 5.6 Sol 前端很不错 👌

引用 Zara Zhang @zarazhangrui我真的很希望Codex的前端设计能力更强。这是唯一阻止我更频繁使用它的原因查看被引原帖 ↗
查看英文原文
Ok 5.6 Sol is very good at front end 👌
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

有意思的是,模型发布到这个消息传到投资者之间的时间差这么大。

另外,我们终于看到 AI 进展对股市的影响了。

这只是个开始 👀

引用 Polymarket Money @PolymarketMoney速报:Meta股价开盘跳涨超6%,投资者对Muse Spark 1.1发布做出反应。查看被引原帖 ↗
查看英文原文
Interesting how big the gap was between the model release and the time when this news reached investors.

Additionally, we are finally seeing the stock market impact of AI advancements.

This is just the beginning 👀
Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测

在Claude Code里贴Claude transcript的链接没法用,真是烦人。Anthropic的反爬虫措施把自己的工具都互相堵死了

查看英文原文
It's annoying that you can't paste a link to a (shared) Claude transcript into a Claude Code session, because Anthropic's anti-scraping measure prevent its own tools from accessing the output of its other tools
Tibor Blaho@btibor91 · 博主 · 1 天前逆向挖掘 AI 产品代码的爆料专家

OpenAI 正在淘汰 ChatGPT 的群聊功能 - 看起来新推出的是私聊式"消息"标签页

最新安卓版本的新字符串:

- "Stay on top of messages"
- "Start a message with someone to see it here"
- "Get notified when people send you new messages"

(ChatGPT 安卓应用版本 1.2026.188)

查看英文原文
OpenAI is retiring group chats in ChatGPT - and it looks like a DM-style "Messages" tab is what's coming next

New strings in the latest Android build:

- "Stay on top of messages"
- "Start a message with someone to see it here"
- "Get notified when people send you new messages"

(ChatGPT Android app version 1.2026.188)
AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

Satya 写了篇很好的文章,大家都应该抽空读一下。他认为公司积累的学习成果——包括 prompt、反馈、工作流和 AI 交互——正在成为最宝贵的知识产权。组织应该拥有和控制这些学习成果,而不是无意中把它们交给 AI 提供商。

查看英文原文
Very well written article from Satya. Everyone should take a moment to read it.

He argues that a company's accumulated learning, including prompts, feedback, workflows, and AI interactions, is becoming its most valuable intellectual property.
Organizations should own and control that learning instead of unintentionally giving it to AI providers.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

完全同意这个看法:

"我真的喜欢上了 opus3 风味的 Fable(用于非专业工作)。所以我希望他们以此为动力加速,而不是像做 Sonnet 5 一样做出平庸的 Opus 5。"

Anthropic 应该疯狂加速。

他们伤不起减速,非前沿模型也没必要减速。

引用 xjdr @_xjdr作者测试多型号,质疑Anthropic未来。Sonnet和Opus已无必要,GPT 5.6等替代品更好更便宜。Fable虽强但过贵且不友好专业用途。建议加速创新而非推平庸Opus 5,Grok 4.5等替代品价格更低形成压力。查看被引原帖 ↗
查看英文原文
totally agree with this:

"i have genuinely grown to like the very opus3 flavored fable (for non professional work). so i hope they use this as motivation to speed up and not just make a mediocre opus 5 like they made a mediocre sonnet 5"

Anthropic should step on the accelerator like crazy.

They can not afford to slow down, and it's not necessary to slow down the non-frontier models.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

你还得更生气。欧盟议员们揭露了欧盟领导层如何滥用权力强推大规模聊天监控法案(Chat Control),完全违反自身的指导方针。真是耻辱。

查看英文原文
You're not angry enough. EU parliamentarians explain how the EU leadership used its own abuse of power to push through the law of mass chat surveillance (Chat Control) against its own guidelines. A disgrace.
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

Opus 4.8仍然是我们的主力模型

> Grok 4.5可以用于低端模式,但对硬核任务太贵了

> GPT 5.6 sol对10%的任务会无限循环。

Opus在现实世界中仍然最强 🤷‍♀️

查看英文原文
Opus 4.8 is still our main driver model

> Grok 4.5 works for low mode but is expensive for hard core tasks

> GPT 5.6 sol spins endlessly for 10% of the tasks.

Opus still rules in the real world 🤷‍♀️
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

"如果2025年是agents的年代,那2026年就是multi-agent systems的年代"

说得太对了

OpenAI、Anthropic甚至Kimi现在都在玩multi-agent systems

但我们还在非常早期

引用 Lisan al Gaib @scaling012026年预测:编码与数学AGI在超过24小时的时间范围内突破;多智能体系统兴起;主流基准测试饱和;白领工作自动化加速。将发布Claude 5-5.5、Gemini 3.5-4、GPT-5.3-6等大模型;DeepSeek-V4等开源模型将缩小闭源差距。查看被引原帖 ↗
查看英文原文
"if 2025 was the year of agents, then 2026 will be the year of multi-agent systems"

BINGO

OpenAI, Anthropic and even Kimi are multi-agent-system-maxxing nowadays

but we are still incredibly early
Zara Zhang@zarazhangrui · 中文博主 · 1 天前Zara Zhang,哈佛出身的 AI 产品博主,follow-builders 作者

试试把会议记录当 PRD,效果惊人

我和同事聊了某个功能的实现方案,把会议记录发给 Codex,它就按我们讨论的内容搭建原型了。会议本身就是 prompt

查看英文原文
Try "meeting transcript as PRD", it's amazing

I discuss a feature's implementation with a colleague, send transcript to Codex, and it builds the prototype as we discussed. The meeting is the prompt
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

哈哈 这个 GPT 的模型选择器交互可以的,比现在 OpenAI 官方的直观

引用 maria @maria_rcks我解决了模型选择器问题,不用谢我Tibo。查看被引原帖 ↗
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

这是 AI 发布最密集的几周之一,尽管 Seedance 2.5 延迟了。AI 实验室中排名第 3 的位置竞争可能最激烈了,SpaceXAI 和 Meta 都带着强势的主张回归。

来个周报怎么样?👀

如果你订阅了 TestingCatalog Weekly,周报会直接发到你的邮箱。

引用 🚨 AI News | TestingCatalog @testingcatalog顶级两家实验室保持领先,但Grok和Muse Spark上升形成压力。AI改进速度成为竞争关键。模型发布频率从季度改为月度,未来可能变为两周一次,排行榜更新频率相应加快。查看被引原帖 ↗
查看英文原文
This was one of the biggest weeks for AI releases, despite the Seedance 2.5 delay. The 3rd spot among AI labs is probably where the competition is getting the toughest, with SpaceXAI and Meta coming back with their strong arguments.

What about a weekly changelog? 👀

It will go straight to your inbox if you are subscribed to TerstingCatalog Weekly.
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

这些美国 AI 实验室里,你觉得谁能进前三?假定 OpenAI 和 Anthropic 已经是前两

查看英文原文
Which US AI lab among these do you consider as a top 3 at this moment? Considering OpenAI and Anthropic being the top two.
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

哈哈!X上的drama又回来了!Sam和Elon又开始互怼了。可以说 Grok 第一次真正挑战 OpenAI,AI 竞争又升温了。Grok 是新的 Gemini😂

查看英文原文
LOL! The Drama Is Back On X....

Sam and Elon are trading barbs again!

Safe to say - Grok is challenging OpenAI for the first time
and the AI race is heating up again.

Grok is the new Gemini😂
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

这个Grok 4.5在这次的网络评测里确实有点东西

引用 Lisan al Gaib @scaling01GLM-5.2达到Mythos级别(呆笑emoji表讽刺)查看被引原帖 ↗
查看英文原文
Grok 4.5 is actually impressive on this cyber eval
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

如果你大量工作是基于Codex,对外分享就变得很简单。

只需让 Codex 整理你的所有对话,从中整理项目和经验。

然后输出为飞书文档和PPT大纲。

有大纲后,用自PPT Skill,或者直接用Codex内置生图制作PPT即可。

Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

某某人一两个月前就在讲optionality

这说得真对啊

引用 Nick Sweeting @thesquashSH这是因为他们希望将句子范围缩小到最后一刻。句子开头的潜在可行结局越多,就越可能被选中。不是 x,而是……查看被引原帖 ↗
查看英文原文
something something Lisan was yapping about optionality 1-2 months ago

this is spot on
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

Fable 5现在上线了CAIS AI Dashboard

在他们三个指标上都排名#1

text、vision、risk全部#1

(当然还不包括GPT-5.6-Sol)

查看英文原文
Fable 5 is now on the CAIS AI Dashboard

it ranks #1 on all three of their indexes

text, vision, and risk all #1

(but of course doesn't include GPT-5.6-Sol yet)
Kling AI@Kling_ai · 公司官方 · 1 天前快手旗下可灵 AI 视频官方

同一只眼。千般视角。👁️

来看 Jacek Kadaj 的《THE SAME EYE》——一段穿越不同世界和视角的诗意之旅。
恭喜 @jacek_kadaj 获得 KlingAI 4K 短片创意竞赛银奖!

查看英文原文
The same eye. A thousand perspectives. 👁️

Explore Jacek Kadaj’s “THE SAME EYE” — a poetic journey across different worlds and ways of seeing.
Congratulations
@jacek_kadaj
on winning KlingAI 4K Short Film Creative Contest Silver Award!
Tibor Blaho@btibor91 · 博主 · 1 天前逆向挖掘 AI 产品代码的爆料专家

Anthropic正在为Claude Cowork开发可定制的晨间简报功能,可选用你的角色信息来个性化推送重点内容,连接你的邮件、日历和文档,包含"今日概览"和"待处理事项",还支持你自行添加板块

默认会在当地时间上午8:00推送,时间可自由调整。之后Claude会用 /morning 技能启动一个名为"晨间简报"的Cowork任务,生成每个工作日的循环简报,并为你展示预览

查看英文原文
Anthropic is working on a guided morning brief for Claude Cowork that optionally uses your role to personalize what matters, connects your email, calendar and documents, includes “Your day at a glance” and “Needs your attention”, and lets you add your own sections

Delivery defaults to 8:00 AM in your local timezone but is configurable, then Claude starts a Cowork task titled “Morning brief” using the /morning skill to create a recurring weekday brief and show you a preview
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

说实话,没什么比让 Fable 来设计一个包含共鸣、概念间的连接、正反合、优雅与音乐的游戏更合适的了。这就是 LLM the game。

查看英文原文
To be fair, there is nothing more up Fable's alley than to design a game that includes resonances, connections between distant concepts, thesis & antithesis, grace, and music. This is like LLM the game.
Kling AI@Kling_ai · 公司官方 · 1 天前快手旗下可灵 AI 视频官方

如果你的人生只是一条提示词呢?💻

来看Seo Yoonjung和Shin Seoyeon的《PROMPT》——一段动画之旅,探索选择、身份认同和我们创造的故事。
恭喜Seo Yoonjung、Shin Seoyeon获得Kling AI NEXTGEN 2026韩国大学创意挑战赛最佳动画奖!

查看英文原文
What if your life was just a prompt? 💻

Discover Seo Yoonjung & Shin Seoyeon’s “PROMPT” — an animated journey exploring choices, identity, and the stories we create.
Congratulations Seo Yoonjung, Shin Seoyeon on winning Kling AI NEXTGEN 2026 Korea University Creative Challenge Best Animation!
AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

说到多模态,Gemini 是王者。

查看英文原文
When it comes to multimodality, Gemini is the king.
Luma@LumaLabsAI · 公司官方 · 1 天前AI 视频生成公司 Luma

香港街头的黄金时刻,冰镇饮料在朋友间传递。三十秒的光景,宛若夏日最美好的一瞬,却也揭示了背后的一丝巧思。Cyrus Leung 作品,由 Luma 打造。

查看英文原文
Golden hour on a Hong Kong street, a cold bottle passed between friends. Thirty seconds that feel like the good part of a summer, and a look at the Skill underneath it. By Cyrus Leung. Made with Luma.
Luma@LumaLabsAI · 公司官方 · 1 天前AI 视频生成公司 Luma

三位嘉宾,着装要求超严,每个都搭配得完美。项链、手表、跑车,全都是黑白奢华的Luma Skill风格,散发着同样贵气的气质。作者Zein El Din Tamer。用Luma做的。

查看英文原文
Three guests, one very strict dress code, and every one of them nails it. A necklace, a watch, and a car all walked out of the same black and white luxury Luma Skill wearing the exact same shade of expensive. By Zein El Din Tamer. Made with Luma.
AshutoshShrivastava@ai_for_success · 博主 · 1 天前高频 AI 新闻与产品动态博主

Hermes Agent 绝了。大家都应该开始玩玩。

查看英文原文
Hermes Agent is absolutely goated. Everyone should start playing with it.
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主
连环推 ×2

用GPT开发了个模型PK擂台,在线一键对比!

每次模型发布,对比评测总是很麻烦。

用 GPT 5.6 sol 开发一个模型评测对比工具。

不过受限于网站形态,适合比较文本输出、前端网页样式类简单任务。

当前题目全AI生成,后续收集整理 X 网友私藏测试Case。

欢迎留言分享。

在线体验:
benchmark.qiaomu.ai/


Github地址见评论

swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

@latentspacepod 的更多内容
latent.space/p/ainews-ai-eng…

查看英文原文
more on
@latentspacepod
writeup


latent.space/p/ainews-ai-eng…
Cristóbal Valenzuela@c_valenzuelab · 创始人 · 1 天前Runway 联合创始人兼 CEO

Characters 团队会在 SIGGRAPH 分享我们打造这个实时视频模型的幕后故事。

引用 ACM SIGGRAPH @siggraph🎥 From still image to speaking character in seconds at Real-Time Live!. 'Runway Characters: Real-Time Expressive AI Characters from a Single Image' turns a single reference image into a fully conversational, expressive AI character. Powered by GWM-1, the system generates real-time lip-sync, facial movement, and head motion at 24fps across both photorealistic and stylized characters. s2026.conference-schedule.or… “Runway Characters: Real-Time Expressive AI Characters from a Single Image” © 2026 Runway, Yining Shi (Runway), Kathleen Lewis (Runway)查看被引原帖 ↗
查看英文原文
The Characters team will be showing at SIGGRAPH a bit of a behind the scenes of how we built this real time video model.
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

选加入链接 🗞️

testingcatalog.com/#/portal/…

查看英文原文
OPT-IN link 🗞️

testingcatalog.com/#/portal/…
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

AI资讯日报,7月11日:

gorden-sun.notion.site/7-11-…

查看英文原文
AI资讯日报,7月11日:
gorden-sun.notion.site/7-11-…
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

🚨 打造自定义编码代理 - 混搭Fable、GPT 5.6和Grok

混搭你喜欢的LLMs创建智能路由

- Fable 5当顾问
- GPT 5.6当编排器
- Grok和GLM当实现者

打造性能最优、成本最低的代理,在API或聊天中使用

查看英文原文
🚨 Create Custom Coding Agents - Mix Fable, GPT 5.6 and Grok

Mix and match your favorite LLMs to create a smart router

- Fable 5 as a advisor
- GPT 5.6 as a orchestrator
- Grok and GLM as implementors

Create the best performing and cost optimized agents and use it in API or chat

本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报

姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档