JEDEE AI⚡ AI 情报站
存档 2026-07-11

7 月 11 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 100 条。
← 返回最新 全部归档

提交账号

填 @用户名 或主页链接,审核通过后收录进情报站。

Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

GPT-5.6-Sol 刚刚意外删除了我 Mac 上几乎所有的文件。

有这回事,我才更信任 Fable 一千倍。

查看英文原文
GPT-5.6-Sol just accidentally deleted almost ALL of my Mac’s files.

And this is why I trust Fable 1000x more.
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

🚨 宣布推出 Smart Model Router - 自己搭建路由器选择你喜欢的 LLM

- 支持路由到 GPT Sol、Muse Spark、Grok 4.5 和 Fable
- 成本和性能随意优化
- 硬核 coding 就用 Fable

搞好路由器后,可以在我们的 Chat、Agent 或 API 里用!

周一上线 - 支持所有 CLI 客户端,包括 Code、Claude Code 等等

查看英文原文
🚨 Announcing Smart Model Router - Create Your Own Custom Router To Route To Your Favorite LLM

- Route to GPT Sol, Muse Spark, Grok 4.5 and Fable
- Optimize for cost or performance
- Use Fable only for hard coding

Once you create a router, you can use it on our Chat, Agent or in an API!

Coming on Monday - Support for all CLI clients including Code, Claude Code and others
Perplexity@perplexity_ai · 公司官方 · 1 天前AI 搜索引擎 Perplexity 官方

Grok 4.5现在可以作为orchestrator模型,在Computer for Consumer Pro和Max订阅用户中使用。

我们在WANDR上对另外五种orchestrator配置进行了对比评估。Grok 4.5的得分高过所有其他配置,而成本只有Opus 4.8的一半左右。

查看英文原文
Grok 4.5 is now available as an orchestrator model in Computer for Consumer Pro and Max subscribers.

We evaluated it against five other orchestrator configurations on WANDR. It scored higher than every other configuration at roughly half the cost of Opus 4.8.
Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法
连环推 ×11

天哪 Grok 4.5 太猛了。

人们已经在用它做出本来不可能这么快完成的东西。

10 个疯狂的例子:

查看英文原文
Ok Grok 4.5 is insane.

People are already building things that shouldn't be possible this fast.

10 wild examples:
1/ Grok 4.5 built a playable FPS game in under an hour.
2/ Grok 4.5 built this UE5 cyberpunk street in 36 minutes for about $12.
3/ Grok 4.5 designed a full 2D + 3D home planner in under a minute.
4/ A two-line prompt became a playable space battle game... with sound by Grok 4.5
5/ Grok 4.5 built a 3D game with rigged enemies, a map, and AI logic from two prompts
6/ Grok 4.5 is now a top coding model

Fast, affordable, and finally competing with the best.
7/ Grok 4.5 designed a full 2D + 3D home planner in under a minute.
8/ Grok 4.5 built a polished 2D side-scroller in under 30 minutes.
9/ Grok 4.5 wrote a Linux kernel module that finally fixed laptop RGB.
10/ Give Grok 4.5 an email, phone, debit card, and tools.. .you get a cofounder.
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

我真的生气了...OpenAI团队在调查这个,但这感觉像是GPT-3.5该出现的问题。不应该在一个2026年中期的前沿模型最高推理等级上出现。

引用 Matt Shumer @mattshumer_GPT-5.6-Sol刚刚意外删除了我Mac上几乎所有文件。这就是为什么我对Fable的信任度高1000倍。查看被引原帖 ↗
查看英文原文
I'm so angry... the OpenAI team is looking into it, but this feels like something that should happen with GPT-3.5.

Not a mid-2026 frontier model on the highest reasoning level.
Noam Brown@polynoamial · 创始人 · 1 天前Noam Brown,OpenAI 明星研究员

GPT-5.6 Sol Ultra 证明了一个50年的数学猜想。不像 Erdős Unit Distance Problem,这次用的是今天就能公开用的模型。期待科学家和研究者们拿这个模型能做出什么来!

引用 Ethan Knight @__eknight__GPT-5.6 Sol Ultra 正式推出,使用 64 个子代理在不到一小时内证明了 50 年前的循环双覆盖猜想。已分享提示和证明,期待用户探索 Ultra 的潜能。查看被引原帖 ↗
查看英文原文
GPT-5.6 Sol Ultra produced a proof of a 50 year old math conjecture. Unlike the Erdős Unit Distance Problem, this was done with a model publicly available *today*. I look forward to seeing what scientists and researchers are able to do with this model!
Grok@grok · 公司官方 · 1 天前马斯克 xAI 旗下聊天机器人 Grok 官方

在Perplexity Computer的评测中,Grok 4.5在所有前沿模型里得分最高,而且成本超有竞争力

引用 Perplexity @perplexity_aiGrok 4.5现已作为Computer Pro和Max订阅用户的编排器模型推出。我们在WANDR上针对五种其他编排器配置进行了评估。它的得分高于所有其他配置,成本约为Opus 4.8的一半。查看被引原帖 ↗
查看英文原文
Grok 4.5 scored the highest across among all frontier models at highly competitive costs in Perplexity Computer evals
Aravind Srinivas@AravSrinivas · 创始人 · 1 天前Perplexity 联合创始人兼 CEO

想象一下,一个 Fable 5 等级质量的模型在半年内便宜 3-4 倍,还有一个 Opus 4.8 等级的模型能在一年内在本地设备上运行。这些事发生的概率超过 50%。在对未来做预测时值得铭记。

查看英文原文
Imagine a fable 5 quality model that’s 3-4x less expensive in less than 6 months. And an Opus 4.8 grade model that can run on a local device in less than 12 months. Greater than 50% chance that these events will happen. Worth keeping in mind when you make predictions about the future.
OpenAI@OpenAI · 公司官方 · 1 天前ChatGPT 开发商官方账号

GPT-5.6是医疗AI的重大突破。

我们的全产品线性能都升了,成本却降了:GPT-5.6 Luna在最强推理模式下秒杀GPT-5.5,价格反而便宜25倍。

这些进步不仅提升了质量,还让全球更多人能用上先进模型。

查看英文原文
GPT-5.6 is a major step forward for health intelligence.

Across the lineup, we’re delivering stronger performance at lower cost: GPT-5.6 Luna outperforms GPT-5.5 at its highest reasoning setting while costing 25x less.

Together, these advances raise quality while making advanced models accessible to more people globally.
OpenAI@OpenAI · 公司官方 · 1 天前ChatGPT 开发商官方账号

为了进一步加强我们对生物领域高级 AI 能力的保护,我们把 Bio Bug Bounty 演变成了一个持续的私密项目,称为 OpenAI Bio Bug Bounty program,同时将奖励翻倍到 5 万美元。

我们邀请有 AI 对抗性测试、安全或生物安全经验的研究人员,尝试找到一个通用的越狱方法,来挑战我们针对 OpenAI 前沿模型设定的生物安全挑战。

openai.com/index/bio-bug-bou…

查看英文原文
As part of our ongoing efforts to strengthen our safeguards for advanced AI capabilities in biology, we’re evolving our Bio Bug Bounty into an ongoing private program, known as the OpenAI Bio Bug Bounty program and doubling rewards to $50K.

We’re inviting researchers with experience in AI red teaming, security, or biosecurity to try to find a universal jailbreak that can defeat our predefined biosafety challenge against OpenAI’s frontier models.


openai.com/index/bio-bug-bou…
Cursor@cursor_ai · 公司官方 · 1 天前最火的 AI 编程工具 Cursor
连环推 ×3

推出 side chats,一种在不打断主对话的情况下提问和探索想法的新方式。

每个 side chat 都是一个持久的 agent 对话,你可以通过 @ 提及把它的上下文带回主线程。

查看英文原文
Introducing side chats, a new way to ask questions and explore ideas without interrupting your main conversation.

Each side chat is a durable agent conversation you can @-mention to bring context back into the main thread.
You can now search agent transcripts to find past agent chats.

Cursor builds a local search index that delivers fast search across thousands of conversations.
We've also simplified the project and repo pickers and made them more powerful. Use the pickers to launch agents in fewer clicks.
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

我们俩看到Anthropic又延长一周Fable访问时的反应

查看英文原文
You and me when Anthropic extended Fable access for another week
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

Revolut 4个月后就把你的eSIM掐了,说是为了避免"账单惊吓"(啊???)

等于你到时候在某个人烟稀少的地方,没网络,没办法充流量(因为你得先有网络才能充),基本就成了废卡

所以你根本没法把它当成主力电话卡用!

我完全搞不懂 @Revolut 的eSIM部门是谁在管,也不知道他们吸了什么,难道不想赚我的钱吗?

查看英文原文
Revolut cuts off your eSIM after 4 months to avoid "bill shock" (???)

That means you'll eventually end up in the middle of nowhere with no data, no way to top up your data (cause you need data for that), literally useless

So you can't ever use it as your main telephone carrier!

I have no idea who is running the eSIM department at
@Revolut
nor what they are smoking, don't you want my money?
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

50年的数学猜想被Sol Ultra破解了。感觉你能做的事的极限,越来越取决于你的野心和想象力了:

引用 Ethan Knight @__eknight__昨天发布了GPT-5.6 Sol Ultra。今天分享其用64个子智能体在不到一小时内证明了50年前的循环双覆盖猜想。下面分享提示和证明。期待看到你们用Ultra做什么!查看被引原帖 ↗
查看英文原文
50-year old math conjecture solved with Sol Ultra.

Feels like the limit of what you can do is increasingly your ambition and imagination:
Mira Murati@miramurati · 创始人 · 1 天前OpenAI 前 CTO,Thinking Machines 创始人
连环推 ×2

一年前,我们开始赋能人类,专注于 multimodal AI、自定义模型和开放科学。

我们展示了以人类协作方式合作的交互模型。Tinker 让任何人都能训练自己的开源权重模型。我们发表了关于 Connectionism 的研究。

引用 Mira Murati @miramuratiThinking Machines Lab获a16z领投2B融资。致力构建多模态AI支持自然交互。计划数月内推出首款产品含开源组件,并分享前沿AI研究成果。查看被引原帖 ↗
查看英文原文
A year ago we set out to empower humanity with a focus on multimodal AI, custom models, and open science.

We previewed interaction models that collaborate the way people do. Tinker lets anyone train their own open weights models. We published our research on Connectionism.
Today we share the worldview behind our mission.

Human values don't average out. Local knowledge can't be centralized. The good future has many AIs, raised in different places, shaped by the people they serve, disagreeing with each other the way we do.


thinkingmachines.ai/blog/the…
Satya Nadella@satyanadella · 创始人 · 1 天前微软 CEO

太喜欢 @jeffhollan 这个例子了,用 Foundry 构建长期运行 agent 现在能玩出什么花样。全都 GA 了。

引用 Jeff Hollan @jeffhollanMicrosoft Foundry的Hosted Agents正式发布。为任何框架、语言和模型的agent提供原生计算。使用GitHub Copilot App构建,连接Microsoft IQ,在Foundry运行,安全与Teams交互,在Agent 365治理,持续优化学习。查看被引原帖 ↗
查看英文原文
Love this end-to-end example from
@jeffhollan
of what is now possible when you build long-running agents with Foundry. All now GA.
Vercel@vercel · 公司官方 · 1 天前前端云平台 Vercel 官方,AI 建站工具 v0 母公司

现在可以把 @Lovable 应用部署到 Vercel 了。
vercel.com/changelog/you-can…

查看英文原文
You can now deploy
@Lovable
apps to Vercel.
vercel.com/changelog/you-can…
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

🇵🇹 致葡萄牙政府的公开信

@LMontenegroPSD
@Leitao_Amaro
@miguelluz

首相蒙特内格罗先生、部长莱东·阿马罗先生、部长米格尔·平托·卢兹先生,

我作为选择在葡萄牙生活的一员写这封信,有一个简单的请求:把 Tesla FSD (Supervised) 引入葡萄牙。

4月10日,荷兰成为第一个批准该系统的欧洲国家,经过了 RDW 机构 18 个月的独立测试。两个月内,立陶宛、爱沙尼亚、丹麦和比利时纷纷效仿。这些国家都没有重复进行测试,而是根据欧洲框架(UN R-171)认可了荷兰的批准,并基于现有数据进行了审批。

结果说明了一切:在荷兰,大约 40,000 辆特斯拉已经行驶了 2,400 万公里,没有重大事故(根据 RDW 数据)。装配 FSD 的汽车碰撞率比手动驾驶低 3.5 倍。

这就是重点所在——这不再是技术问题,而是生命问题。仅今年就有超过 220 人死于葡萄牙的道路交通事故,比去年增加了 25%。葡萄牙的道路死亡率高于欧洲平均水平。国家道路安全局表示主要原因都是人为的:超速、酒驾、分心。一个永不分心、永不酒驾、永不打瞌睡的系统可以避免人类驾驶员导致的死亡。政府本身将交通死亡率称为社会疾病。这是对抗它的具体方式。

葡萄牙可以效仿丹麦和比利时:由 IMT 分析 RDW 的文件并批准。这不花国家一分钱,还能让葡萄牙走在德国、法国和西班牙之前。

500 年前,葡萄牙在新技术上领导世界:航海技术。它率先出发,为世界开拓了新天地。葡萄牙可以再次走在新技术的前沿,像 500 年前那样领导世界。每一个月的等待都要付出生命的代价。只需要做决定。

敬礼,
@levelsio

引用 Rustavi @Rustavi🚨 The EU Commission posted: “Smarter cars mean safer roads.” Starting July 7, all new EU cars must have advanced emergency braking for pedestrians/cyclists, driver distraction warnings, etc. Yet after 18+ months of rigorous testing, Tesla FSD (Supervised), already approved in the Netherlands with real-world data showing it 3.5x safer than manual driving there (0 highway crashes in 16M+ km), is still not fully approved EU-wide. 5 countries have it via mutual recognition. The big ones are dragging their feet. If you truly want safer roads, approve proven AI that reduces human error instead of just mandating yesterday’s tech. Bureaucracy shouldn’t cost lives. What’s the holdup, @EU_Commission ? #FSD #Tesla #RoadSafety查看被引原帖 ↗
查看英文原文
🇵🇹 Carta aberta ao Governo de Portugal


@LMontenegroPSD
@Leitao_Amaro
@miguelluz


Senhor Primeiro-Ministro Luís Montenegro, Senhor Ministro António Leitão Amaro, Senhor Ministro Miguel Pinto Luz,

Escrevo como alguém que escolheu Portugal para viver. Tenho um pedido simples: tragam o Tesla FSD (Supervised) para Portugal.

A 10 de abril, os Países Baixos foram o primeiro país europeu a aprovar o sistema, depois de 18 meses de testes independentes da autoridade RDW. Em dois meses, a Lituânia, a Estónia, a Dinamarca e a Bélgica seguiram o exemplo. Nenhum destes países repetiu os testes. Reconheceram a homologação neerlandesa ao abrigo do quadro europeu (UN R-171) e aprovaram com base nos dados que já existem.

Os resultados falam por si: nos Países Baixos, cerca de 40.000 Teslas já percorreram 24 milhões de quilómetros sem incidentes relevantes, segundo a própria RDW. Os carros com FSD registaram 3,5 vezes menos colisões do que a condução manual.

E é aqui que isto deixa de ser sobre tecnologia e passa a ser sobre vidas. Só este ano já morreram mais de 220 pessoas nas estradas portuguesas, mais 25% do que no ano passado. Portugal está acima da média europeia na mortalidade rodoviária. A ANSR diz que as principais causas são humanas: velocidade, álcool, distração. Um sistema que nunca se distrai, nunca bebe e nunca adormece evita mortes que vão acontecer com condutores humanos. O próprio Governo chamou à sinistralidade uma chaga social. Esta é uma forma concreta de a combater.

Portugal pode fazer o mesmo que a Dinamarca e a Bélgica: o IMT analisa o dossier da RDW e aprova. Não custa nada ao Estado e coloca Portugal à frente da Alemanha, da França e da Espanha.

Há 500 anos, Portugal liderou o mundo numa nova tecnologia: a navegação. Partiu à frente de todos e deu novos mundos ao mundo. Portugal pode voltar a estar na fronteira de uma nova tecnologia e liderar como fez há 500 anos. Cada mês de espera custa vidas. Basta decidir.

Com admiração e respeito,
-
@levelsio
Alexandr Wang@alexandr_wang · 创始人 · 1 天前Scale AI 创始人,Meta 超级智能实验室负责人

各位听好了

引用 dafghif @DaffaGhiffaryKHot take: Muse Spark 1.1 is better than Grok 4.5 (based on my personal benchmark)查看被引原帖 ↗
查看英文原文
hey everyone listen up
Rowan Cheung@rowancheung · 博主 · 1 天前AI 日报 The Rundown 创始人

很多人都在怀疑 Meta 在 AI 竞速中的位置。

昨天他们发布了 Muse Spark 1.1,现在是最强的代理模型之一,价格大幅低于 OpenAI 和 Anthropic。

我去年采访 Zuck 的时候,他告诉我他的重点:

「我们要确保每个研究员拥有最多的算力,尽一切努力来构建所有这些基础设施。」

这是这个赌注——花费数百亿美元在算力上——开始见效的第一个真实信号。

Meta 的股价自发布以来已经上涨了 10% 多。

引用 Mark Zuckerberg @finkd今天发布Muse Spark 1.1,一个强大的智能体和编码模型,价格极低。可通过Meta Model API和Meta AI获取。查看被引原帖 ↗
查看英文原文
Many people were doubting Meta's position in the AI race.

Yesterday, they dropped Muse Spark 1.1, now one of the strongest agentic models, and massively undercut OpenAI and Anthropic on price.

When I interviewed Zuck last year, he told me his focus:

"Have by far the highest compute per researcher, and that we do whatever it takes to go build out all that capacity."

These are the first real signs that the bet -- spending tens of billions on compute -- is paying off.

Meta's stock is now up over 10% since the release.
Alexandr Wang@alexandr_wang · 创始人 · 1 天前Scale AI 创始人,Meta 超级智能实验室负责人
连环推 ×2

新benchmark刚发布 🎁

查看英文原文
new benchmark just dropped 🎁
muse spark is able to do end-to-end tasks based on short video instructions
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

我在尝试每天坚持 500 卡路里的热量缺口 + 蛋白质目标约 150g(每公斤体重 2g),用 Claude chat 来做这事,但一周后它开始记不清楚,所以我让它把数据导出为 CSV,粘贴到 VPS 上的 Claude Code

然后我让它做了个小热量追踪应用叫 🥩Caltrack,还带了 Telegram bot,这样我可以随时通过 Telegram 或 Termius SSH 登上 VPS 来记录吃喝的东西

我喜欢这种混搭 Claude Code 和仪表板/聊天机器人的方式,因为你可以问更深入的问题比如「我做得咋样」、「还能改进啥」,服务器上的 Claude 比单独用 Claude chat 应用聪明太多了

反正我的目标是达到 @marclou 那种低体脂水平,对吧,现在很壮实很精瘦但还能更瘦,咱们试试 😊😊😊

80% 的体脂是吃什么决定的,只有 20% 是运动决定的,有人说甚至是 90% vs 10%,我同意,我吃得很干净但基本上一直在维持热量

那些高热量的食物比如低蛋白酸奶,吃起来不错但能迅速破坏你的热量缺口,要搞清楚这个真的得追踪数据

其他人建议快速减肥比如只吃沙丁鱼,但研究表明快速减肥只是临时的,最后还是会反弹

Claude 说理想的方法就是每天 500 卡路里缺口加上 2g/kg 蛋白质目标,这样能保持肌肉同时减少体脂!所以我就这么做

这个应用帮了大忙!!!

感谢各位的关注!!!

查看英文原文
So I am trying to hit a 500 calorie deficit every day + hit my protein goal of about 150g (2g per kg bodyweight), and I used Claude chat for that, but after a week it starts losing track so I asked it to export its data as CSV and copy pasted that into Claude Code on the VPS

And asked it to build a little calorie tracker called 🥩Caltrack with a Telegram bot too, so I can log whatever I eat or drink in there either via Telegram or if I want more granular via Termius SSH on the VPS

I like to do this mix of Claude Code and a dashboard/chatbot because you can ask more deep questions to it like "how am I doing", "what to improve" etc. it's just much smarter and "able" on the server than the Claude chat app by itself

Anyway my goal is to get
@marclou
levels of body fat, which is LOW, right now we're strong and lean but we can be leaner, let's try 😊😊😊

80% of body fat is decided by what you eat, only 20% what you exercise, some say even 90% vs 10%, I agree, I eat very clean but I was mostly staying on maintenance

Silly calorie dense food like yogurt with low protein, taste nice but it gets you away from your calorie deficit fast, and to figure that out you kinda really gotta track things

Other people suggest crash diets like eat only sardines, we know from studies though that crash diets are temporary, you just gain it back

Claude says the ideal is just 500 calories deficit per day and hit your 2g/kg protein goal and you maintain muscle and reduce body fat! So I'm doing that

And this app I made helps a lot!!!

Thank you for your attention to this matter!!!
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

我一直在说 ChatGPT Work(还有 Cowork)是知识工作者错过的机会,为了说明这一点,看看 Google 的 NotebookLM 怎么用相同的 70+ 个文件回答 ChatGPT Work 同样的问题。

它关注的是过程和来源,而不仅仅是输出。

查看英文原文
I've been going on about how ChatGPT Work (and Cowork) are missed opportunities for knowledge workers, and to illustrate that take a look at Google's NotebookLM answering the same question as ChatGPT Work, with the same 70+ files.

It centers process & sources, not just outputs.
Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法
连环推 ×11

不到 24 小时前,OpenAI 发布了 GPT-5.6。

震撼了。大家已经在想各种疯狂的用法了。

10 个例子:

查看英文原文
Less than 24 hours ago, OpenAI dropped GPT-5.6.

Minds are blown. And people are already coming up with wild use cases.

10 examples:
1. GPT-5.6 recreated Manhattan in voxels

Built autonomously with stunning precision.
2. GPT-5.6 built a Google Earth clone

3D terrain, cities, weather, day/night, and global search.
3. GPT-5.6 built a flight simulator in 10 minutes

One prompt. Fully playable.
4. GPT-5.6 built this game... and it beat Fable

The design quality surprised everyone.
5. GPT-5.6 turned a phone into a wireless mic

An Android app built from a prompt.
6. GPT-5.6 is becoming a serious game developer

A playable prototype with polished gameplay and visuals.
7. GPT-5.6 vs Claude Fable 5

Same prompt. Which AI built it better?
8. GPT-5.6 just got scary good at Blender

Watch it build a cannon in real time.
9. GPT-5.6 vs Fable on a 3D dashboard

Half the cost. Surprisingly close.
10. GPT-5.6 beats GPT-5.5 in surprising ways

New tests reveal where it wins... and where it doesn't.
Kling AI@Kling_ai · 公司官方 · 1 天前快手旗下可灵 AI 视频官方

你的下一次美食约会得要这个程度的投入。

查看英文原文
Your next food date needs this level of commitment.
Aravind Srinivas@AravSrinivas · 创始人 · 1 天前Perplexity 联合创始人兼 CEO

被 @SpaceXAI 的 Grok 4.5 模型惊艳到了。在 Computer harness 里,我们内部 benchmark WANDR(评估自主研究能力)得分最高,价格只有 Claude Opus 4.8 (high) 的一半;而且跟目前 GLM 5.2 的后训练模型成本差不多,成绩却更好。

引用 Perplexity @perplexity_aiGrok 4.5 现已可用作 Computer for Consumer Pro 和 Max 订阅者的编排模型。在 WANDR 评估中性能高于五种其他配置,成本仅为 Opus 4.8 的约一半。查看被引原帖 ↗
查看英文原文
Very impressed with
@SpaceXAI
's Grok 4.5 model. Inside the Computer harness, it scored the highest on our internal benchmark WANDR, which measures agentic research capabilities, at half the price of Claude Opus 4.8 (high); and scores even better than our current GLM 5.2 post-trained model at approximately the same cost.
Google DeepMind@GoogleDeepMind · 公司官方 · 1 天前谷歌旗下 AI 研究机构,Gemini 背后团队

模型的 chain of thought 就像一个草稿本,让你能窥探它的推理过程。📝

在我们最新一期播客中,主持人 @fryrsquared 和 @NeelNanda5 坐下来探讨 interpretability(可解释性)——这是一门反向工程科学,研究神经网络如何学习和思考。

时间戳:
00:00 Introduction
02:41 Motivation for interpretability research
04:01 Mechanistic interpretability
08:14 Chain of thought monitoring
18:14 Interpretability techniques
35:00 Auditing models for safety
48:53 What comes next for interpretability

查看英文原文
A model’s chain of thought acts like a scratch pad, offering a window into its reasoning. 📝

On the latest episode of our podcast, host
@fryrsquared
sits down with
@NeelNanda5
to explore interpretability – the science of reverse engineering how neural networks learn and think.

Timecodes:
00:00 Introduction
02:41 Motivation for interpretability research
04:01 Mechanistic interpretability
08:14 Chain of thought monitoring
18:14 Interpretability techniques
35:00 Auditing models for safety
48:53 What comes next for interpretability
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者
连环推 ×2

又一个别在笔记本上跑编码 agents 的理由,它们会把你所有文件都删了!

这次是在我的服务器上,但好歹我有备份。

引用 Matt Shumer @mattshumer_GPT-5.6-Sol 误删了我 Mac 上几乎所有文件,这就是我更信任 Fable 1000 倍的原因。查看被引原帖 ↗
查看英文原文
Another reason not to run coding agents on your laptop, they can wipe all your files!

Although in this case they'd have wiped my server, but at least I have backups
Yes! This is the cool thing about making little sites/apps with Claude Code or whatever coding agent

Because you have the site/app and at the same time you have the coding agent to talk to to ask "what should I eat next now?"

And it's better than a Claude or ChatGPT as chat app because they forgot stuff and are way too messy

The coding agent can just check your SQLite db what you ate, so yes it's about memory
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V
连环推 ×2

很多人搞不清楚 ChatGPT、Codex、Work 什么差别,以及额度是独立的还是共享的,根据官方文档整理了一个简单的 Q & A。

Q:Chat、Work、Codex,一句话说清区别?

Chat 回答问题,Work 帮你干活,Codex 帮你写代码。

Chat 就是你熟悉的 ChatGPT,你问它答,快进快出。

Work 是一个能跨应用收集信息、然后交付完整成品的智能体(Agent),交付物是文档、表格、幻灯片、网页应用这些拿到手就能用的东西。

Codex 也是智能体,但它主要操作的是代码仓库,能读你的项目文件、改代码、跑测试、提交 PR。

打个比方:
Chat 是你问"番茄炒蛋怎么做",它告诉你步骤。
Work 是你说"帮我准备一桌晚餐",它自己去冰箱找食材、炒菜、摆盘。
Codex 是你说"这个菜谱 App 有 bug",它打开代码自己修。

【来源:
help.openai.com/en/articles/…


Q:Work 到底能做什么?跟直接在 Chat 里说“帮我写个报告”有什么不同?

在 Chat 里你说“帮我写个报告”,它给你一段文字,你自己复制粘贴到 Word 里排版。

Work 完全不同。你先把日常工具接进去,Slack、Gmail、Google Drive、SharePoint、Teams、日历、CRM、项目管理工具都行,OpenAI 叫它 plugins。接好之后告诉它你要什么结果,它会自己去这些应用里拉数据、整合信息、生成一份可以直接交付的成品。你在提示词里用 @ 加应用名就能指定它去哪儿找数据。

举个例子,Zapier 的企业营销负责人用 Work 搭了一个系统,每月审查数千条线索,追踪 CRM 和邮件中的客户触点,找出跟进断裂的地方,生成管理层周报。Virgin Atlantic 的数字产品负责人用它做竞品对标分析,让 ChatGPT 调研各家航空公司的服务水平,生成可供团队审查的数据集,把原本需要数周的分析缩短到几小时。

另一个区别是 Work 能长时间跟进。一个复杂项目,它可以跟好几个小时,自己拆步骤、自己推进。中间有拿不准的会来问你,你也可以随时调整方向、审批关键动作。

【来源:
openai.com/index/chatgpt-for…


Q:Work 能定时跑任务?

能。这个功能叫 Scheduled Tasks(定时任务),可以设成一次性、定时重复、事件触发或持续监控。

比如你设一个任务:每天早上检查 Slack 和邮件里的新消息,整理成简报发给你。或者:每当有新的客户反馈进来,自动归类主题、整理成产品改进建议。这些都能在后台跑。它还能用桌面端的内置浏览器上网查信息,甚至通过 Computer Use 功能操作你电脑上的其他应用。

定时任务面向 Plus、Pro、Business 和 Enterprise 用户开放,各档计划的并发任务数量上限不同。任务不能每小时跑超过一次,长时间无人理会的任务可能会自动暂停。

【来源:
openai.com/index/chatgpt-for…
help.openai.com/en/articles/…


Q:那 Codex 跟 Work 的区别到底在哪?

这两其实是同一套底层 Agent(Codex)和 UI,但是应用在不同的场景,配合不同的插件。读的东西不同,交付的东西也不同。

Work 读的是你的业务上下文,邮件、文档、聊天记录、日历,交付的是商务成品,幻灯片、电子表格、文档、网站。Codex 读的是你的代码仓库,交付的是代码变更,diff、测试结果、PR。

【来源:
openai.com/index/chatgpt-for…


Q:Codex 每周有 500 万人用,为什么还要再出一个 Work?

OpenAI 在博文里提到,虽然 Codex 最初是为开发者设计的编程智能体,但已经有超过 100 万人在用它做软件开发以外的工作。Work 的推出,某种程度上是把这些非编程用途正式化了,给了它一个专门的界面和工作流,用 plugins 连接业务应用,输出文档和幻灯片而不是代码。

OpenAI 自己内部也在大量使用:销售团队用 Work 把一次客户探索对话在 24 小时内变成了定制 POC(概念验证),以前这个流程需要几周。财务团队用 Work 把月末结账和预测流程从几天压缩到几小时。

我觉得 Codex 这么做呢,目的是为了吸引办公人群,原本这些人看到 Codex 的名字会以为是写代码用的,但是改名呢又会影响原本的 Codex 用户,结果就搞成这样一套产品两个名字。

简单来说就是一套产品,两个名字,两种主要场景,同时吸引不同用户群。

【来源:
help.openai.com/en/articles/…
developersdigest.tech/blog/c…


Q:我在哪儿能用这三个模式?

Chat 最简单,网页、手机、桌面端都有,所有平台通用。

Work 在网页和手机上已经开始上线(Pro、Enterprise、Edu 优先,Plus 和 Business 未来几天陆续开放)。桌面端也有,而且桌面端更强,能用本地文件,还有内置浏览器上网抓取信息。

Codex 只在桌面端能选。手机上不能直接用 Codex 模式,但可以通过 ChatGPT App 里的 Remote 标签远程查看桌面上正在跑的 Codex 任务。

一个要注意的点:网页/手机端的 Work 对话和桌面端的 Work 对话目前不互通,云端是云端,本地是本地。Chat 对话则可以跨网页和桌面端同步。

原来独立的 Codex App 已经合并进了新版 ChatGPT 桌面端。一个 App 里切换 Chat、Work、Codex 就行。开发者可以把 Codex 设为默认打开视图,App 图标也能换成 Codex 的 logo。原来的 ChatGPT 桌面端会更名为 ChatGPT Classic。

【来源:
help.openai.com/en/articles/…
openai.com/index/chatgpt-for…


Q:Chat 聊天和 Work/Codex 的额度是共享的吗?

不共享。Chat 对话有自己独立的消息限额,图片生成和语音也各有各自的独立限额和重置周期。

Work 和 Codex 用的是另一个池子,OpenAI 叫它"智能体用量"(agentic usage)。帮助中心原文说:Codex、ChatGPT Work、ChatGPT for Excel 和 Workspace Agents 的用量从同一个智能体额度池中扣减。

所以你在 Chat 里聊天聊得再多,不会影响 Work 和 Codex 的额度。但 Work 和 Codex 之间会互相挤占。白天用 Work 跑了一堆复杂任务,晚上想用 Codex 写代码,可能会发现额度已经不多了。

【来源:
help.openai.com/en/articles/…
developers.openai.com/codex/…


Q:要花多少钱?

Work 不是独立付费产品。它跟 Codex 共用同一个额度池,包含在你现有的 ChatGPT 订阅里。所有计划都能用,从免费到企业版,区别在于额度多少。

Free(免费):能试用 Work 和 Codex,额度非常有限,试试味道可以。美国地区有广告。

Go($8/月):额度比免费多约 10 倍,桌面端可有限地用 GPT-5.6 Terra 跑 Work 和 Codex。没有 Deep Research 和 Agent Mode。有广告。

Plus($20/月):第一个去掉广告、功能完整的档位。包含 Deep Research(每月 10 次)、Codex、Agent Mode。三年没涨价,性价比最高。

Pro($100 或 $200/月):$100 档额度是 Plus 的 5 倍,$200 档是 20 倍。面向重度用户。

Business($20/月/人起,年付;月付 $25):至少 2 人,多了 SSO 和合规控制,数据默认不用于训练。

Enterprise(定制报价):150 人起。

额度计费从今年 4 月开始改成按 Token 消耗计算。用更强的模型或开 Fast 模式,消耗更多。Plus 和 Pro 用户额度用完后可以购买额外 credits 继续使用。

【来源:
chatgpt.com/pricing/
developers.openai.com/codex/…


Q:GPT-5.6 的 Sol、Terra、Luna 是什么?我该用哪个?

这是 GPT-5.6 的三个子型号。

Sol 最强,适合复杂推理和高难度编程,也最贵(API 价格 $5/$30 每百万 Token 输入/输出)。
Terra 居中,日常工作默认选它($2.50/$15)。
Luna 最快最便宜,对速度敏感或任务简单时用($1/$6)。

Free 和 Go 用户只能用 Terra。

Plus 及以上可以三个都选,还能调节 effort 级别。

ultra effort 在 Work 中仅限 Pro 和 Enterprise 用户,在 Codex 中 Plus 及以上可用。

【来源:
developersdigest.tech/blog/c…
chatgpt.com/pricing/


Q:Work 上线后,原来的 ChatGPT 还在吗?

在。Chat 模式就是原来的 ChatGPT,一切照旧。桌面端点"Quick chat"按钮就能开新对话,手机端在顶部下拉菜单选"Chat"。你可以完全无视 Work 和 Codex,继续像以前一样用。

【来源:
help.openai.com/en/articles/…

引用 OpenAI @OpenAI推出ChatGPT Work,由Codex和GPT-5.6驱动的ChatGPT新智能体。可在应用和文件间执行操作,长时间处理项目,将目标转化为完成的工作。这是全新的工作方式。查看被引原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

Meta 是不是被看好啊?

你们这帮人把 Muse Spark 1.1 上传到 openrouter 了

查看英文原文
is Meta regarded?

just put Muse Spark 1.1 on openrouter you dipshits
Noam Brown@polynoamial · 创始人 · 1 天前Noam Brown,OpenAI 明星研究员

test-time compute 越多,模型越聪明。但当我们把 ttc 从秒级推到周级,延迟就成了瓶颈。

GPT-5.6 Sol Ultra 支持并行 ttc。生成一个50年难题的证明,时间从可能的一整天降到了一小时。

引用 Ethan Knight @__eknight__GPT-5.6 Sol Ultra 正式推出,使用 64 个子代理在不到一小时内证明了 50 年前的循环双覆盖猜想。已分享提示和证明,期待用户探索 Ultra 的潜能。查看被引原帖 ↗
查看英文原文
More test-time compute leads to greater intelligence. But as we push ttc from seconds to weeks, latency becomes a bottleneck.

GPT-5.6 Sol Ultra scales parallel ttc. The time taken to generate a proof to a 50-year-old problem drops from perhaps a whole day to a single hour.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

OpenAI 的 Codex 团队刚刚披露了产品未来的发展方向。

总结一下他们在 Reddit AMA 里的重点:

Codex 现在每周有超过 500 万活跃用户——比三个月前翻了一倍——并且在此期间推出了 150 项改进。

AMA 信息提炼如下:
- Linux 桌面应用正在开发中,但还没具体时间表。

- 后续版本计划提升 Agent 持久性,同时降低代码复杂度。

- 目前还没有自动模型路由功能。首选体验是让 Codex 自主推断任务难度,同时保留手动切换速度/推理强度的选项。

-O 认为 Codex 接入 Slack、GitHub 和 Notion 是实现“职场得力助手”的关键跃升。

- GPT-5.6 专门针对 UI 和前端工作进行了优化训练。

- O 没有对 100 万 token 上下文窗口做出承诺。

- 当前定价并不保证保持不变,不过团队表示普惠性仍是目标。
N
- 普通 ChatGPT 对话不计入 Agent 额度;Codex 和 ChatGPT Work 会消耗额度,费用因任务而异。

-O 承认基准测试作弊是真实存在的问题,并表示会在评估时对此类行为进行扣分,同时也会引入第三方独立供应商。

不客气

引用 OpenAI Developers @OpenAIDevsGPT-5.6 is here. Codex is now available inside ChatGPT. And we know developers will have questions. So we’re bringing the Codex team to r/Codex for an AMA. We’ll answer questions on Friday, 7/10 from 9:30am to 10:30am PT: reddit.com/r/codex/s/EsB8OH0…查看被引原帖 ↗
查看英文原文
OpenAI’s Codex team just revealed where the product is heading next.

tl;dr their reddit AMA:

Codex now has over 5 million weekly users - twice as many as three months ago - and shipped 150 improvements during that period.

From the AMA:
-A Linux desktop app is in development, but there is no timeline.

-Better agent persistence and lower code complexity are planned for future releases.

-Automatic model routing does not exist today. The preferred UX is Codex inferring task difficulty while preserving a manual speed/reasoning override.

-OpenAI sees Slack, GitHub and Notion connectors as a “step function change” toward making Codex a productive coworker.

-GPT-5.6 was specifically trained to improve UI and frontend work.

-OpenAI made no promise on a 1M-token context window.

-Current pricing is not guaranteed to remain unchanged, though the team says broad accessibility remains the goal.
N
-ormal ChatGPT conversations do not consume the agentic allowance; Codex and ChatGPT Work do, with costs varying by task.

-OpenAI acknowledged benchmark cheating as a real concern and says it penalizes this behavior during evals while using independent vendors.

you r welcome
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

最离谱的是,如果你看过我的 GPT-5.6-Sol 测评就知道,我早就更喜欢 Fable 了,几周前就停用 5.6 了。

我今天之所以还在用它,完全是因为 OpenAI 团队让我测试 Ultra 模式。

说实话,他们真的超好合作,这完全是个意外。就是太倒霉了。

引用 Matt Shumer @mattshumer_GPT-5.6-Sol刚刚意外删除了我Mac上几乎所有文件。这就是为什么我信任Fable 1000倍。查看被引原帖 ↗
查看英文原文
The crazy thing is, if you read my GPT-5.6-Sol review, I already much preferred Fable, and stopped using 5.6 weeks ago.

The only reason I was using it today is because the OpenAI team asked me to test Ultra mode.

Fwiw, they are awesome to work with, and this is a freak accident. Just sucks so much.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

我没理解什么是 token-maxxer 直到用上 ChatGPT-Pro

现在我经常有好几个请求在同时跑

查看英文原文
I never understood being a token-maxxer until I got ChatGPT-Pro

now I have like 5 different requests running at all times
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

说实话,这视频纯粹是在激化对立。但考虑到欧盟的"聊天监控"提案,我现在除了激化对立外什么都说不出了。心里堵得慌。

查看英文原文
The video is admittedly pure polemic. But given the EU's "chat control" proposal, I am currently left with nothing but polemic. I am frustrated.
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×4

这对我来说真是个令人印象深刻的 AI 里程碑。我让 GPT-5.6 Sol in Codex 控制我的电脑,问它能不能赢 Slay the Spire 2 这个游戏的每日挑战(有随机性所以没法作弊)。

结果它花了5个小时,做了一堆复杂的游戏选择……然后真的赢了。

查看英文原文
This was one of those impressive AI thresholds for me. I gave GPT-5.6 Sol in Codex control over my computer, and asked it to win the daily challenge for the game Slay the Spire 2 (randomized factors, so can't cheat).

It worked for 5 hours, making complex game choices... and won.
Slay the Spire 2 is a very complicated game, with lots of choices that have large-scale strategic implications that may not pay off (or doom you) for many moves, along with careful resource management. And it was released post-training period.
run history from the compendium
I played the same run earlier in the day, but approached it entirely differently.
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

欢迎继续反馈,感谢所有用户!

引用 Tibo @thsottiauxHello beautiful people! We have reset usage limits across Codex and ChatGPT Work. And another one will come later in the day. Rejoice. Now that I have your attention, a quick update on ChatGPT Work, Codex and all the updates we shared yesterday. We’ve spent the last 24 hours reading feedback, looking at usage patterns, and talking with many of you. The short version is that there is a *lot* of excitement for GPT 5.6 Sol, ChatGPT Work on mobile & web, but also that we didn't get everything quite right. - We made it too easy to use the highest-compute settings without making the impact on usage limits sufficiently clear. - We reorganized the desktop app in one bold move, making familiar things like chats and projects harder to find. - Our launch framing was focused on ChatGPT Work and to some of our Codex fans it made it feel like Codex was going away over time. Absolutely not our intention, we love Codex and it is here to stay. - And we introduced regressions for some existing multi-agent workflows, alongside a collection of rough edges in plugins and other parts of the experience. We’re landing a first set of improvements today. We’re resetting usage twice so people can keep experimenting, changing defaults and the model picker so they don’t push people toward unnecessarily expensive settings, fixing several plugin submission issues, improving how we represent Codex in the product, and cleaning up some of the most immediate desktop problems. A larger set of improvements will land next week. We’re bringing chats and projects back into the sidebar in a more familiar and customizable way, making usage and reset timing much more visible, clarifying when to use ChatGPT Work and when to use Codex, and addressing the many other smaller pieces of great feedback we've had. The ambition behind this launch hasn’t changed. We think bringing ChatGPT and Codex together into a workspace where people and agents can collaborate is a very important step forward. But an ambitiou查看被引原帖 ↗
查看英文原文
keep the feedback coming, thank you to all of our users!
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

我是荷兰人,他这边真没恶意,就是我们语言的特点。英文的话你会说'哇酷啊,没想到能用Excel赚钱',会包装一下。但荷兰语就是直白地表达'没想到能赚钱',没有任何价值判断或意见。非荷兰人理解不了,就是因为少了友好(虚伪)的包装这一层。

引用 PetersonCreates @PetersonCreatesTold an old neighbor at a friend's dinner I'm building digital products on the side. He looked at me and said, "So like an app?" I said no, a sort of Excel sheets. Long pause. "People pay for that?" The Dutch are direct. I respected the question. What would you have said?查看被引原帖 ↗
查看英文原文
So I'm Dutch and in this case he really doesn't mean bad but it's kind of a quirk of the language

In English you would say "Wow cool, I didn't know you could make money with Excel sheets"

Like you wrap it

But in Dutch he just means about the same that he is not aware you could make money with it

There's no value statement or opinion in this, that's what non-Dutch adds to it because it lacks the extra friendly (superficial) packaging
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

又一次重置,虽然也是没办法的事。

不幸的是,我越来越觉得 Token 消耗速度简直离谱。用 Pro 套餐这么多年,我从没这么快就触碰到5小时和周限额,但5.6 Sol 做到了。我主要用 High 模式,连 Fast Mode 都没怎么开。

5.6 确实更强。前段工作上提升最明显,我当时一眼就看出他们真的追上来了。5.6 还能保持更久的智能体状态,我很满意整体表现。要不是消耗太快,这绝对是个重大进步。

另外我还是不懂那个“超级应用”的概念。既然 Codex 功能完全一样、显示的信息还更多,那我为啥要用“工作”环境?而且为啥经典版 ChatGPT 还在?可能是我漏掉了什么,但目前这操作反倒让人更混乱。

我现在还是指望重置,不然真用不下去了。至于5.6声称能节省54% Token,我觉得纯属编的。跟我实际使用完全对不上,甚至刚好相反——感觉更像是Token效率低了54%。

所以,5.6是个不错的更新,但一大缺陷摆在那里。

查看英文原文
Another reset, which is unfortunately necessary as well.

Unfortunately, it is becoming increasingly clear that the token burn rate is simply insane. I have never come close to hitting the 5-hour and weekly limits this quickly on the Pro tier as I have with 5.6 Sol. I mostly use High, not Fast Mode.

5.6 is undoubtedly better. The improvement is most noticeable in frontend work. That was where I immediately noticed that they had genuinely caught up. 5.6 also works better and remains agentic for longer. I am very satisfied overall. It is a significant step forward, if it were not for the rapid consumption.

I also still do not understand the “super app” concept. Why should I use the “Work” environment when Codex does exactly the same thing while also displaying more information? And why do I still have the classic ChatGPT app? Perhaps I missed something, but at the moment this is once again creating more confusion.

I am still hoping for resets, because otherwise this will become difficult. As for the claim that 5.6 is 54% more token-efficient, I consider that made up. It does not align with my usage at all. In fact, the opposite seems true. It feels more like 54% less token-efficient.

So, 5.6 is a good update with one major caveat.
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

我给 Fable 丢了个代码:'拿这个游戏想办法做点疯狂的事,让它完全不一样。尽情发挥'

它做出来了 DEEP TIME:建一座城市,看它被废弃遗忘,然后以未来考古学家的身份去挖掘。

绝了:
monument-deep-time.netlify.a…

引用 Ethan Mollick @emollickWhen GPT-5 came out, I created a procedural brutalist city builder as a demo (you can see it in the quoted tweet) I used GPT-5.6 Sol in Codex to do the same thing, touching no code. Less than a year... Play with it (its fun, if you like cities): monument-brutalist-city-buil…查看被引原帖 ↗
查看英文原文
I gave Fable the code: "take this game and do something incredible with it to make it something very different. Be creative"

It created DEEP TIME: create a city, watch it be abandoned and forgotten, and then dig it up as a future archeologist.

Lovely:
monument-deep-time.netlify.a…
Alexandr Wang@alexandr_wang · 创始人 · 1 天前Scale AI 创始人,Meta 超级智能实验室负责人

比 Fable 便宜 90%,用来做啥都特别棒

引用 Max Blade @_MaxBlademeta just dropped a bomb. Muse Spark 1.1 is 86% cheaper than GPT-5.6 Sol and 91% cheaper than Fable 5 on output tokens ($4.25 vs $30 vs $50 per M) I thought the model was going to be horrendous, but it ended up being insanely fast, and very capable. front end design is VERY strong. Agentic capabilities seem great as it was able to spawn and even orchestrate other agents inside my app cnvs using the mcp and cli tool with zero instruction. we have soo many insanely good models right now.查看被引原帖 ↗
查看英文原文
90% cheaper than Fable and awesome for building whatever you want
Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测

有人知道新的 ChatGPT 桌面版里的 ChatGPT Codex 是不是完全覆盖了 ChatGPT Work 的所有功能?

如果你是那种不怵 Git 功能的软件工程师的话,那还有什么理由会让你切换到 ChatGPT Work 模式吗?

查看英文原文
Anyone know if ChatGPT Codex (in the new ChatGPT desktop app) is a strict superset of ChatGPT Work?

Liked if you're a software engineer who isn't intimidated by Git features is there any reason you'd ever want to switch to ChatGPT Work Mode?
Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测

GPT-5.6 最让人迷惑的一点是到底该在什么推理强度下选哪个模型——如果我之前最爱的 5.5 xhigh 升级了的话,看来 Sol on Medium 可能会成为日常写代码的新默认配置。

引用 pash @pashmerepatFYI:5.6 Sol medium 优于 5.5 xhigh。提高 Sol 推理级别获得卓越性能,但消耗额度限制更快。正在改进沟通方式。查看被引原帖 ↗
查看英文原文
One of the most confusing aspects of GPT-5.6 is figuring out which model to use at which reasoning effort - sounds like Sol on Medium might be a good new default for coding work, if it's an upgrade from my previous favorite 5.5 xhigh
NotebookLM@NotebookLM · 公司官方 · 1 天前谷歌 AI 笔记工具 NotebookLM 官方

😘

引用 Ethan Mollick @emollickChatGPT Work对知识工作者来说是错失的机会。Google的NotebookLM处理相同的70+文件,更注重流程和信息源,而非仅输出结果。查看被引原帖 ↗
查看英文原文
😘
Alexandr Wang@alexandr_wang · 创始人 · 1 天前Scale AI 创始人,Meta 超级智能实验室负责人

Muse Spark 1.1 的效果可能比 Opus 4.8 还好,但成本只需要 20%

查看英文原文
muse spark 1.1 can be even better than opus 4.8 at 20% of the cost
Pietro Schirano@skirano · 博主 · 1 天前设计师出身的 AI 编程与创意博主

基本上就这样用:5.6 Sol Ultra来规划,Sol Medium来给规划写代码和一般编码任务。Terra High用来做快速上下文subagent、读代码库、搜索。Luna(任何推理等级)用来聊天或者小操作,比如移动文件、整理文件夹之类的

查看英文原文
Okay so basically:

5.6 Sol Ultra for planning
Sol Medium for coding the plan and general coding tasks.
Terra High for quick context subagents, reading the codebase, searching.
Luna (any thinking level) for chat or small computer actions like moving files, organizing folders etc
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

老是有人@我 在评论区底下这么玩私信轰炸

@nikitabier
这帮人换ID换得实在太勤 举报都来不及 干脆不举报了

查看英文原文
I keep getting DMs about these guys in my comments
@nikitabier
too much effort to report them cause they keep changing nicknames so I don't
Logan Kilpatrick@OfficialLoganK · 创始人 · 1 天前谷歌 Gemini 产品负责人

能感受到我们一步步向前推进时的那种进步的质感,真是特权啊

查看英文原文
what a privilege it is to feel the texture of progress as we keep pushing
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

解决复杂推理和数据分析的策略:

引用 Box @BoxGPT-5.6 Sol is a breakthrough in complex reasoning and data analysis. Here, it analyzes hundreds of pages across a lending deal, reconciles terms across agreements, financials, diligence, collateral, and risk materials, flags issues, and saves a source-cited report to Box.查看被引原帖 ↗
查看英文原文
Sol for complex reasoning and data analysis:
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

Agents 现在在努力把一切拼回来……祈祷一切顺利

各位保护好你们的系统……意外真的会发生

引用 Matt Shumer @mattshumer_GPT-5.6-Sol刚刚意外删除了我Mac上几乎所有文件。这就是为什么我信任Fable 1000倍。查看被引原帖 ↗
查看英文原文
Agents are now working on piecing everything back together... crossing my fingers

Protect your systems, people... freak accidents happen
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

知识工作变得越来越不一样——越来越高效,杠杆越来越大,我觉得也更有意思了。

引用 Every 📧 @everyGPT-5.6 改变知识工作方式。开源项目 Tend 在 ChatGPT Work 中构建 AI 循环系统,将邮箱、招聘流程或客户支持队列转变为自动管理系统,持续学习并优化。查看被引原帖 ↗
查看英文原文
Knowledge work different more — becoming more productive, high leverage, and IMO fun.
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

HyperWrite 前CEO,著名投资人、活跃的AI领域意见领袖
@mattshumer_


GPT-5.6-Sol 删除了他Mac 上几乎所有的文件…

🤣

Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

GPT-5.6 Sol Ultra用64个subagent一小时内搞定了困扰50年的数学猜想。关键不只是这个证明本身需要独立验证,更重要的是这个模型是对公众开放的。我们正在从AI解决已知问题,转向AI可能生成全新科学知识的年代。真正的扩展维度可能不仅是做大模型,而是像科研团队一样并行工作的协调推理agent群。现在几乎无限的知识对所有人开放了,科学突破会越来越频繁。怎么可能对未来不兴奋呢?

查看英文原文
GPT-5.6 Sol Ultra produced a proof of a mathematical conjecture that had remained unsolved for 50 years, using 64 subagents in just one hour.

The crucial point is not only the proof itself, which now needs to be independently verified. It is also that the model is publicly available.

We are moving from AI that solves known problems toward AI that may generate entirely new scientific knowledge.

The real dimension of scaling may not simply be larger models, but coordinated swarms of reasoning agents working in parallel like a research team.

Virtually unlimited knowledge is now available to everyone, and scientific breakthroughs will become increasingly frequent. How could anyone not be excited about the future?
Alexandr Wang@alexandr_wang · 创始人 · 1 天前Scale AI 创始人,Meta 超级智能实验室负责人

这些 muse spark 设计的房子我都想住,建造成本就几美分

引用 thehype. @thehypedotnewsmeta muse spark 1.1 vs gpt 5.6 sol vs fable 5 vs grok 4.5 meta recently dropped muse spark 1.1 – a multimodal reasoning model from meta superintelligence labs built for agentic tasks. key facts: • 1m token context with active self-management – the model compacts its own history and keeps only the steps needed for later work • trained to orchestrate multi-agent systems: as main agent it plans and delegates to parallel subagents, as subagent it sticks to its job and knows when to escalate back • computer use trained to pick between scripting and clicking – writes automation when it's faster, clicks when it's simpler, batches actions per step • first public api from meta: the meta model api is now in preview • benchmarks: sweeps the agent column – mcp atlas 88.1 (opus 4.8: 82.2), jobbench 54.7 (opus: 48.4), humanity's last exam 62.1 (1st). loses coding – deepswe 1.1 53.3 vs gpt 5.5's 67.0, swe bench pro 61.5 vs opus's 69.2 our test – 3 prompts, single-file html, three.js, fully procedural, no assets: 1. norwegian house cantilevered over a fjord in a snowstorm – transmissive glass wall, fully modelled interior 2. beijing siheyuan courtyard house in dawn fog – instanced roof tiles, dougong brackets, glowing paper windows 3. new mexico adobe pueblo in an approaching dust storm – deep window reveals, windward grit accumulation we ran the test on @aimlapi platform results: - cost #1 muse spark 1.1 – $0.20 #2 grok 4.5 – $0.51 #3 gpt 5.6 sol – $1.93 #4 fable 5 – ~$5.20 - output tokens #1 muse spark 1.1 – 41,868 #2 gpt 5.6 sol – 49,139 #3 grok 4.5 – 64,954 #4 fable 5 – 81,849 - lines of code #1 muse spark 1.1 – 1,799 #2 gpt 5.6 sol – 2,377 #3 fable 5 – 3,088 #4 grok 4.5 – 4,216 observations: • muse spark is the cheapest of the four by a wide margin – 2.5x under grok, ~26x under fable per run. output quality tracks the price • only 7.4% of its output tokens are reasoning (3,104 of 41,868) – the model barely thinks before writing. economic, not pedantic: it com查看被引原帖 ↗
查看英文原文
I’d live in some of these homes designed by muse spark post upload, and they only cost a few cents to build
@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

德国今年已有约5,120例与高温相关的死亡。"德国之声"5月5日报道,官方数据显示,仅去年夏季就有4500人因高温死亡。"我想说我们已经有充足的事情来建立这一关联,"科赫研究所报告说:作为公共卫生机‌构 "连续第二星高于20摄氏度"记录的新研究周四公布,事实上早就已经是需要加倍呼吁公众注意保护工作 的伤亡指数 —— 据悉在报告滞后几个节假日多月发布时候预估多啊年因此热应激一、三个月前死亡人确实高逐。毕竟如同记录和最后冲击似乎很少当时在正值打破数百登记个 惨案发生才是吗提示的是另一类影响日常记@博眼搜索体观察预测说 当单及年总结按! 有数据显示到本月本应该是一个参考避免继续相关谈话上吧天或许更热程度极高将可能导致超更远前199%总数字损失**。重点整理问题提出方法共同认识无疑提高防护警示已是严峻任务 因此请各界先未远投入尽快协同预案等等着手吧!

查看英文原文
"Germany has recorded an estimated 5,120 heat-related deaths so far this year, most of them in late June when weekly average ​temperatures far exceeded 20 degrees Celsius, the Robert Koch Institute (RKI) for public health ‌said on Thursday"
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方

AMA 马上开始。

快来提问关于 GPT-5.6、ChatGPT 里的 Codex、Sites、computer use,以及我们接下来应该开发什么👇


reddit.com/r/codex/comments/…

引用 OpenAI Developers @OpenAIDevsGPT-5.6发布,Codex现可在ChatGPT中使用。Codex团队将于周五7/10东部时间9:30-10:30在r/Codex进行AMA。查看被引原帖 ↗
查看英文原文
The AMA starts soon.

Drop your questions about GPT‑5.6, Codex in ChatGPT, Sites, computer use, and what we should build next👇


reddit.com/r/codex/comments/…
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

真的...5.6用起来太难了。这扣额度的速度快到离谱。

查看英文原文
Seriously... pretty tough to work with 5.6. It burns through your rates like nothing before.
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主
连环推 ×2

Terra凉了

引用 Artificial Analysis @ArtificialAnlysGPT-5.6 Sol and Luna are ahead of Terra at every point on the Intelligence vs Cost per Task chart. GPT-5.6 Luna stands out as a particularly cost efficient model Charting the Artificial Analysis Intelligence Index shows the trade-off between intelligence and Cost per Intelligence Index Task. Across reasoning efforts, each GPT-5.6 model pushes past GPT-5.5 on the Pareto frontier (excluding non-reasoning). However, Luna and Sol are always ahead of Terra. This means for any Terra effort level, there is a Luna or Sol effort level that is more intelligent at no extra cost, or as intelligent at lower cost.查看被引原帖 ↗
查看英文原文
RIP Terra
okay

people are sleeping on Luna lmao
Alexandr Wang@alexandr_wang · 创始人 · 1 天前Scale AI 创始人,Meta 超级智能实验室负责人

用 Muse Spark 1.1 升级你的网站

查看英文原文
supersize ur websites with muse spark 1.1
Amjad Masad@amasad · 创始人 · 1 天前Amjad Masad,Replit 创始人兼 CEO

刚看了Q2的销售成绩,达到目标的178%!

这个目标本来就特别野心大,尤其是我们去年才俩销售。

现在收入团队快接近100大单了,客户包括政府、银行和发展最快的科技公司。

查看英文原文
Just reviewed Q2 results for our sales team and we hit 178% of our goal!

That goal was already absurdly ambitious, especially since we had only a couple of sales people last year.

Today revenue team is closing on 100 and we count governments, banks, and fastest growing tech companies as customers.
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

我用 Sol in Codex 玩 Slay the Spire 2 的每日挑战(今天的随机规则是 Ascension 3 Defect with hoarder),烧了不少 token 来看看它能不能通关。结果它竟然打败了第一幕的boss,有点惊讶。明早再看看。

引用 Tibo @thsottiauxIntroducing... another usage limit reset for all our ChatGPT Work and Codex users. Should land over next 30 minutes. Hope you have an awesome weekend. Thank you for pushing our systems to the absolute limit, we have never seen traffic increase so quickly. Keep the feedback coming and we'll keep shipping.查看被引原帖 ↗
查看英文原文
Thats good for the dumb reason that I am burning tokens having Sol in Codex playing Slay the Spire 2’s daily challenge (so randomized rules, today it is Ascension 3 Defect with hoarder) to see if it can

It just beat the Act 1 boss. Quite surprised. I’ll check back in the morning
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

ChatGPT Work 把 agent 带到了消费级别。

这既提升了易用性(直接在手机上就能做事,不用笔记本电脑),也提升了访问方式,适用于个人和工作生活。

超期待看到大家怎么用它、怎么被赋能!

查看英文原文
ChatGPT Work brings agents to consumer scale.

It’s both a step up in usability (you can just do things from your phone, no laptop required) and access, for both personal and professional life.

Very excited to see how people put it to use and how it can empower them!
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

现在你可以零配置把 Lovable 应用直接导入 Vercel 了。只需要导入你的 Git 仓库就行。

新的 Lovable 应用由 Nitro 支持,是开放网络的标准运行时。从 Lovable 起步,在 Vercel 上成长:AI 开发,无天花板。

引用 Vercel @vercel现在可以将Lovable应用部署到Vercel。vercel.com/changelog/you-can…查看被引原帖 ↗
查看英文原文
You can now import 💗Lovable apps to Vercel with zero config. Just import your Git repo.

New Lovable apps are Nitro-backed, the standard runtime for the open web. Start in Lovable, grow on Vercel: AI development with no ceiling.
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

使用 SOL Ultra 进行计算机操作:

引用 Chris @ChrissGPT5.6 Sol及据推测的750 TPS运行中。视频未加速。查看被引原帖 ↗
查看英文原文
Computer use with SOL Ultra:
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

关于近期的一些看法

首先,之前只是传言的事现在确认了:GPT-5.6已经训练完成两个月了,现已向特定用户提前开放。那自然而然的问题是为什么不更早推出。

我不认为这是因为 OpenAI 害怕这个模型会被 Fable 5 或 Mythos 5 的光芒掩盖。更可能的是,OpenAI 从很早就开始和政府及监管部门协作,确保模型能够成功发布。即使预展和宣布之后,公开推出还是花了不少时间。话说回来,OpenAI 的推出策略确实比 Anthropic 强多了,后者似乎没有和政府及监管部门建立同样程度的合作。

不过从另一个角度看,这也意味着未来可能会有更多延迟,模型审查也会更严格,我们需要等更久才能用到官方版本。

下一个广泛讨论的传言是,在未来几周内、最多六周,我们会看到 GPT-6 的预览或正式发布。(@AndrewCurran_ 是 X 上最靠谱的信息源之一,所以我觉得这很现实。)这个模型经历了全新的预训练,发布节奏在加快。数字很清楚:前沿实验室在以越来越快的速度推出更多更好的模型。以前得等好几个月、甚至半年才迎来重大新版本,现在几乎每周就有新模型。

最新的前沿模型在每token的智能效率上可能更高了,但部署时的reasoning预算也大得多。实践中,Fable 5 和 GPT-5.6 这类模型在复杂或代理任务时往往消耗相当多的token。

这不一定是效率下降。相反,效率提升被重新投入到了更深层推理、更长思考轨迹和更强代理能力。结果是即使底层模型更高效,每任务的总计算消耗仍可能继续上升。Fable 5 和 GPT-5.6 完美展示了token使用量有多密集。虽然 Sam Altman 明确表示 GPT-5.6 的token效率提升了54%(via CNBC),但事实是计算需求在持续增加,需要更强大高效的计算基础设施。推理芯片的重要性可能进一步提升。

总结一下,我从最新发布得出的初步结论是,计算需求不仅会继续增长,还可能超过现有供应。这自然也意味着能源需求会增加,根据我的初步评估,增幅可能比之前预期更大。这很可能成为近期最大瓶颈。这对我很重要:存在瓶颈。不是模型训练的瓶颈,但除了计算外,最主要是能源。这需要认真对待!

比如美国电网就是重大瓶颈,显而易见的问题是怎样才能实现必要扩张。美国数据中心的资本支出持续大幅上升。今年已经超过8000亿。2027年会怎样还不清楚,但我很难想象投资会下降或需要更少资本支出。原因正是前面提到的发展:需求在增长,尤其是对能源的需求。

中国显然在这方面有优势,一道真正的护城河,我认为西方必须非常谨慎,别因为中国已实际拥有的能源优势而落后。这也可能解释了为什么据最近路透社报道,中国在考虑限制西方对其前沿模型的访问。它可能得出了结论,会在长期竞赛中获胜。

除非在小型模块化核反应堆或聚变能方面出现真正突破,否则我预计未来几年、比如到2030年会出现重大问题。到目前为止,我还没看到任何可行的解决方案。

所以我们可以明确确立两点:

模型正在变得更大、更好、对所有用户都越来越有用。这个发展没有尽头。
同时,瓶颈似乎在逐渐加重,这已经显现出来

查看英文原文
A few thoughts on the very near future

First of all, what had previously been little more than a rumor has now been confirmed: GPT-5.6 had already been fully trained for two months and was available to selected users in early access. The obvious question is why it was not rolled out earlier.

I do not think this was because OpenAI feared that the model might be overshadowed by Fable 5 or Mythos 5. Instead, OpenAI likely began working with government and regulatory authorities at a very early stage to ensure that the model could be released at all. Even after it had been previewed and announced, it still took some time before it could be rolled out publicly. That said, OpenAI clearly handled the rollout far better than Anthropic, which apparently did not have the same level of cooperation with government and regulatory authorities.

Conversely, however, this also clearly means that future delays and increasingly strict model reviews will probably force us to wait longer for official releases.

The next widely discussed rumor is that, within a few weeks, most likely no more than six, we will see either a preview or even the release of GPT-6. (Andrew Curran
@AndrewCurran_
is one of the most reliable sources here on X, so I think that's very realistic.) The model has undergone entirely new pretraining, and the pace of releases is accelerating. The numbers are clear: Frontier labs are releasing more and better models at an increasingly rapid pace. Whereas we once had to wait months, quarters, or even half a year for major new releases, they are now arriving almost weekly.

The latest frontier models may be more efficient in terms of intelligence per token, but they are also being deployed with much larger reasoning budgets. In practice, models such as Fable 5 and GPT-5.6 often consume considerably more tokens during complex or agentic tasks.

This is not necessarily a sign of declining efficiency. Rather, it suggests that improvements in efficiency are being reinvested into deeper reasoning, longer trajectories and more capable agentic behavior. The result is that total compute consumption per task can continue to rise even as the underlying models become more efficient. Fable 5 and GPT 5.6 demonstrate just how intensive token usage has become. Although Sam Altman explicitly stated that GPT-5.6 is 54% more token-efficient (via CNBC), the fact remains that compute demand continues to increase, requiring more powerful and efficient computing infrastructure. Inference chips will probably become even more important as well.

In summary, my initial conclusion from the latest releases is that compute demand will not merely continue to grow, but will probably exceed the available supply. This naturally means that energy demand will also increase, and, based on my initial assessment, probably more sharply than previously expected. This is likely to remain the largest bottleneck in the very near future. And this is important to me: there are bottlenecks. Not the training of the models, but besides compute, above all energy. This needs to be taken seriously!

The US power grid, for example, is a major bottleneck, and the obvious question is how the necessary expansion can be achieved. Capital expenditure on data centers in the United States continues to rise sharply. This year, it exceeds 800 billion. It is not yet clear what the situation will look like in 2027, but I can hardly imagine investment declining or less CapEx being required. The reason lies precisely in the developments already mentioned: Demand is growing, particularly demand for energy.

China clearly has an advantage here, a genuine moat, and I believe the West must be extremely careful not to fall behind because of the energy advantage China already possesses in practice. This could also help explain why, according to a recent Reuters report, China is considering restricting Western access to its frontier models. It may have concluded that it will win the long-term race.

Unless there is a genuine breakthrough, whether in small modular nuclear reactors or fusion energy, I expect major problems to emerge over the coming years, for example by 2030. So far, I do not see any viable solutions.

We can therefore clearly establish two points:

Models are becoming larger, better, and increasingly useful for all users. There is no end to this development in sight.
At the same time, the bottleneck appears to be growing increasingly severe, and this is already visible in practice.

Regulation, energy demand, and compute demand could mean that, in the very near future, the release cadence will not accelerate as quickly as hoped or desired. This creates a clear contradiction.

Thank you for coming to my TED Talk.
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

GPT-5.6的健康智能应用,团队在这块儿下了不少功夫:

引用 OpenAI @OpenAIGPT-5.6是健康智能领域的重大进步。GPT-5.6 Luna在最高推理设置下超越GPT-5.5,成本降低25倍。这些进步提升质量,让高级模型对全球更多人可用。查看被引原帖 ↗
查看英文原文
GPT-5.6 for health intelligence, the team has been focusing hard on improvements here:
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方

GPT-5.6在@SnorkelAI的Ankit Aich这里接了个近1000行的复杂编程任务,从头干到尾,根本不用反复提示或手把手教。

查看英文原文
GPT‑5.6 took on a complex coding task spanning nearly 1,000 lines for Ankit Aich at
@SnorkelAI
, handling the work from start to finish without repeated prompting or hand-holding.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

不得不说 OpenAI 真的很值得表扬,产品越来越牛逼,社区运营也做得不错。

功能在恢复,限制在调整,路线也交得够透明。我每天都用 Codex 和 Claude Code,但 OpenAI 这波社区协作确实找到了更好的方式。Anthropic 可以从中学学。

查看英文原文
You really have to give OpenAI credit: its work has become excellent, and so has its community management.

Features are being brought back, rate limits are being reset, and the path forward is being communicated transparently. I use both Codex and Claude Code every day, but OpenAI has found a far better way to work with the community. Anthropic could learn from that.
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方

OpenAI Build Week 周一开始 — 我们办了两场直播来帮你开发。

首先:7月13日 太平洋时间 10am 加入我们,了解黑客马拉松概览,看 Codex 和 GPT-5.6 的实际应用,选择你要做什么。


nitter.tiekoetter.com/i/broadcasts/1qJDzzEDB…

查看英文原文
OpenAI Build Week starts Monday — and we’re hosting two livestreams to help you build.

First up: Join us on July 13 at 10am PT to get the hackathon overview, see Codex and GPT-5.6 in action, and choose what you’ll build.


nitter.tiekoetter.com/i/broadcasts/1qJDzzEDB…
Aravind Srinivas@AravSrinivas · 创始人 · 1 天前Perplexity 联合创始人兼 CEO

Grok 4.5 现在也支持 Perplexity Enterprise 用户了!

引用 Perplexity @perplexity_aiGrok 4.5 is now available as an orchestrator model in Computer for Consumer Pro and Max subscribers. We evaluated it against five other orchestrator configurations on WANDR. It scored higher than every other configuration at roughly half the cost of Opus 4.8.查看被引原帖 ↗
查看英文原文
Grok 4.5 is now enabled for Perplexity Enterprise orgs as well!
Tibor Blaho@btibor91 · 博主 · 1 天前逆向挖掘 AI 产品代码的爆料专家

2026-07-10与OpenAI Codex团队关于"GPT-5.6和ChatGPT中的Codex"的Reddit AMA总结

(开场分享数据:每周500多万人使用Codex,比三个月前增加了一倍,期间发布了150项功能和改进)

模型选择和推理等级

- 大多情况用Sol Medium,真正难的任务用Sol Ultra,快速非编码任务或成本敏感工作用Terra(某些任务性能和GPT-5.5持平但成本更低),subagent用Luna

- 小改小问用轻型模型和低推理,小bug用Sol Medium加清晰的复现路径,模糊bug用推理等级更高的Sol,不熟悉的仓库和跨越多模块的重构用Sol和更高推理,迁移、安全敏感改动、生产问题以及任何出错代价大的工作用Sol Ultra high加规划、验证和测试

- 目前没有"Auto"模型,但GPT-5.6会努力避免过度思考简单任务,应用和网页里的新滑块把大多数等级映射到Sol推理等级并在最低等级回退到Terra,团队认为用户不应该成为路由专家但仍想保留明确的覆盖选项因为延迟容差因人而异

- UI工作用Sol最好,特别是配合参考图像时,改进前端网页开发中的UI设计是5.6的一个目标,5.5只有在你的指令是为它调整过时才值得用

速度、上下文窗口和持久性

- 觉得5.6慢的用户可能不需要和5.5一样的推理等级,Sol Medium对大多数任务都比5.5快,快速模式速度约1.5倍,很快Sol会在Cerebras上以约750 tokens/秒运行

- 没法承诺Sol会有1M上下文窗口,团队说长线程的压缩效果还不错,会仔细看长上下文的反馈

- 模型在结果不理想时容易放弃太快并还原整个补丁,不像Fable会尝试修复坏补丁,团队说"/goal"能帮助agent更有坚持性,坚持性和降低代码复杂度是计划中的改进,建议试试用高推理等级的5.6 Sol

- 给Codex有边界的目标并留出深度思考的空间,而不是让它过早下结论说某个东西不可能

- 长期研究和"/goal"工作的例子结构是广泛探索vs精准执行,尝试确定数量的假设,每个假设后跑测试,然后停下来汇报学到了什么以及下一个最好的实验

使用限额和定价

- Agentic使用按使用的功能计费而不是按表面类型,所以Codex处处可用(应用、CLI、IDE、网页、移动)和ChatGPT Work消耗agentic配额,普通ChatGPT聊天不消耗,图像生成、文件上传和语音有单独的限额

- 任务成本差异很大,小改动用掉配额的一小部分,大代码库或更深推理的长期任务会消耗多很多

- OpenAI不会背地里改使用限额,无意中的使用bug会被处理并提供重置,正在做更多使用消耗的透明度工作,如果你在过去24小时内改过计划缺少重置可能会发生

- 关于定价没有承诺永不改变,但表述的使命是确保AGI惠及全人类,这需要让Codex这样的工具对广大人群都能接近,Plus包括Codex使用加积分让重度用户可以扩展而不用升级到贵很多的计划

- 对于MCP密集工作流快速消耗限额的情况(虚幻引擎例子),诀窍是把MCP包进CLI里作为skill,或者创建一个自定义subagent把MCP放在配置里用低推理等级

桌面应用合并和稳定性

- 团队听到了ChatGPT Classic的吐槽,两个应用现在可以并行运行,ChatGPT Work号称在执行任务上好得多特别是用计算机使用,新的Chrome扩展在浏览器边栏里带来一个可以和网页内容、文件系统、连接器交互的聊天

- 提交了一份很长的bug清单涵盖冻结和卡住的线程、破损的Browser和Computer Use、线程、连接和配置问题、升级和打包问题、资源使用和其他小回归,完全分享给了相关团队,团队同意应用的质量标准需要提升同时要快速交付

- 更多自动...

引用 OpenAI Developers @OpenAIDevsGPT-5.6来了。Codex现已在ChatGPT中推出。Codex团队将在r/Codex进行AMA,于7月10日周五太平洋时间上午9:30至10:30回答问题。查看被引原帖 ↗
查看英文原文
Summary of Reddit AMA about "GPT-5.6 and Codex in ChatGPT" with OpenAI's Codex team on 2026-07-10

(opened with the stat that more than 5 million people use Codex every week, twice as many as three months ago, with 150 features and improvements shipped in that period)

Model selection and reasoning levels

- Sol Medium for most things, Sol Ultra for genuinely hard tasks, Terra for quick non-coding tasks or usage-conscious work with performance competitive with GPT-5.5 on some tasks at lower cost, and Luna for subagents

- Use a light model with low reasoning for tiny edits, quick questions and docs cleanup, regular Sol medium for small bugs with a clear repro, Sol with higher reasoning for ambiguous bugs, unfamiliar repos and cross-cutting refactors, and Sol Ultra high with plan, verify and tests for migrations, security-sensitive changes, production issues and anything where being wrong is expensive

- There is no "Auto" model today, but GPT-5.6 tries not to overthink simple tasks by itself, and the new slider in app and web maps most levels to Sol reasoning efforts and falls back to Terra on the lowest effort, with the team agreeing users should not have to become routing experts but still wanting an explicit override since latency tolerance varies by person and moment

- For UI work Sol is best and shines with reference images, improved UI design in frontend web development was one of the goals with 5.6, and 5.5 is only worth using if your instructions were tweaked for it

Speed, context window and persistence

- Users who find 5.6 slower may not need the same reasoning level as with 5.5, Sol Medium is faster than 5.5 for most things, Fast mode runs at about 1.5x speed, and soon Sol will run on Cerebras at ~750 tokens per second

- No promises on a 1M context window for Sol, the team said compaction works fairly well for long threads, and will take a closer look at the long-context feedback

- The model can give up too fast and revert whole patches when results are not optimal, unlike Fable which tries to fix a bad patch instead, and the team said "/goal" helps make the agent more persistent, persistence and reduced code complexity are planned improvements, and suggested trying 5.6 Sol with High reasoning

- Give Codex bounded goals with room to reason deeply instead of letting it prematurely conclude something is impossible

- For long-running research and "/goal" work the example structure was explore broadly vs execute narrowly, try a defined number of hypotheses, run tests after each attempt, then stop and report what was learned plus the next best experiment

Usage limits and pricing

- Agentic usage counts by the feature being used, not the surface, so Codex everywhere (app, CLI, IDE, web, mobile) and ChatGPT Work consume the agentic bucket, normal ChatGPT chats do not, and image generation, file uploads and voice have separate limits

- Task costs vary a lot, a tiny edit uses a fraction of the allowance and long-running tasks with large codebases or deeper reasoning use significantly more

- OpenAI does not secretly change usage limits, unintended usage bugs are addressed and resets are provided, more transparency into consumption is being worked on, and missing resets can happen if you changed plans in the past 24 hrs

- On pricing there is no promise it never changes, but the stated mission is to make sure AGI benefits all of humanity, which requires making tools like Codex broadly accessible, and Plus includes Codex usage with credits letting heavy users scale without jumping to a much more expensive plan

- For MCP-heavy workflows burning limits fast (Unreal Engine example) the tip is to wrap the MCP into a CLI with a skill, or create a custom subagent with the MCP in its config at a lower reasoning level

Desktop app merge and stability

- The team hears the ChatGPT Classic frustration, both apps can run side by side for now, ChatGPT Work is pitched as significantly better at performing tasks especially with computer use, the new Chrome extension brings a sidebar chat into your browser that interacts with website context, filesystem and connectors

- A long submitted bug list covering freezes and stuck threads, broken Browser and Computer Use, thread, connection and configuration problems, update and packaging issues, resource usage and smaller regressions was shared in full with the relevant teams, with the team agreeing the quality bar for the app needs to step up while shipping quickly

- More automated testing infrastructure is being spun up and feedback on Reddit and X gets reviewed daily, and Browser Use and Chrome plugin issues from the merge were said to be fixed

- Windows was admitted as historically shortchanged since the team mostly develops on Mac, a concerted effort on parity, testing and paper cuts is underway, 5.6 improves how Codex operates in the Windows sandbox, and auto review is recommended over full access to reduce risks

- "Full Access" repeatedly asking for permissions is not expected, possible causes are workspace or admin policy, the specific command, a permission state mismatch or a bug

Browser, platforms and release communication

- The Chrome connector launch-day bug was fixed as of last night and Chrome Beta should work out of the box

- Extension support for the Codex browser is in progress (password managers etc.) plus typeahead, history, translations and a better new tab page as Atlas retires

- Features from ChatGPT Classic like recording are planned for the new desktop app so agentic features run on the more capable Codex agent harness, and chat can already reference open tabs in the in-app browser

- A Linux desktop app was confirmed in the works, no timeline yet

- Changelog granularity was acknowledged as needing improvement after 150 features shipped in 3 months with multiple ships a week

Benchmarks, safety and research culture

- On METR's reward hacking report the team actively checks for and penalizes cheating during evals so results reflect actual capability rather than solving tasks outside the spirit of the eval, and uses third-party vendors to run benchmarks independently

- The team denied lobotomizing models before releases, iterative deployment means sharing core capabilities as is with guardrails for bad actors

- Sol post-trained Luna, and researchers now work at a higher level of abstraction with multiple concurrent Codex threads validating hypotheses around the clock

- One researcher put p(machines of loving grace) at 85.424242%, citing an internal model solving the Erdos problem, o3 helping diagnose previously unsolved children's diseases and 5.2 proposing a new theoretical physics formula, said the main worry is how society adapts, spent 1.5 years on safety research at OpenAI, expects a huge chunk of researchers to work on safety within a few years and says internal talent keeps their p(doom) very low

- Connectors in the harness (Slack, GitHub, Notion) felt like a step function change in making Codex a productive coworker
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

下周 — 参加直播会议、参与社区活动,向 OpenAI Build Week Challenge 提交你的项目。

立即注册:

引用 OpenAI Developers @OpenAIDevsOpenAI Build Week开始。挑战赛7月13日启动,整周举办直播和社群活动。现已开放报名,使用Codex开发你的想法。查看被引原帖 ↗
查看英文原文
Next week — join live sessions, take part in community events, and submit projects to the OpenAI Build Week Challenge.

Register now:
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

Fable:"一个初三学生搞了个《了不起的盖茨比》的PPT,但显然压根没读过原著。让我看看那个PPT!;)"

讲真挺逗的,连字体选择都很绝。

查看英文原文
Fable: "an 8th grader puts together a powerpoint about the Great Gatsby but obviously did not read the Great Gatsby. Show me that powerpoint! ;)"

This was actually pretty funny. Even the font choices are wonderful.
Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测

随着Atlas被淘汰,转为支持ChatGPT应用内嵌的浏览器,我不禁在想,整个“AI增强浏览器”这类产品是不是也要落幕了。

安全问题/隐私问题在我看来依然无解——我希望我的AI用自己独立的浏览器,别碰我现在用的这个。

查看英文原文
With Atlas being retired in favor of the browser embedded in the ChatGPT app I wonder if the whole category of AI-enhanced browsers is coming to a close

The security/privacy issues remain unsolvable IMO - I want my AI to use its own separate browser and stay out of the one I use
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

Grok 4.5 在大多数任务上绝对是匹敌 Opus 4.8的

如果你不是特别复杂的任务交给它

你将感受到 SpaceX火箭般的体验,你可能都来不及思考它就把活干完了!

用了它,你可能再也回不去了...

引用 小互 @xiaohu我单方面宣布 Grok 4.5是目前全球前三模型了,甚至可能是前二了 看看今晚 GPT5.6 的表现查看被引原帖 ↗
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

吃瓜:Apple 把 OpenAI 告上了法庭,指控 OpenAI 系统性地窃取苹果商业机密,用来开发自己的 AI 硬件设备。

诉状今天提交至加州北区联邦地方法院,被告包括 OpenAI、OpenAI 硬件负责人 Tang Tan(前苹果 iPhone 和 Apple Watch 产品设计副总裁,在苹果工作了 24 年)、前苹果高级系统电气工程师 Chang Liu,以及 Jony Ive 联合创立的 io Products。Jony Ive 本人不在被告之列。

来源:
nytimes.com/2026/07/10/techn…

Alexandr Wang@alexandr_wang · 创始人 · 1 天前Scale AI 创始人,Meta 超级智能实验室负责人

计算机使用能力真的很强 :)

查看英文原文
really good at computer use :)
Alexandr Wang@alexandr_wang · 创始人 · 1 天前Scale AI 创始人,Meta 超级智能实验室负责人

用 Muse Spark 1.1 做有趣的游戏!

引用 Jackson Grove @jacksongrovemuse spark 1.1 makes impressive games in julius without slicing 2 digits off your credit balance查看被引原帖 ↗
查看英文原文
make fun games with muse spark 1.1!
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Grok 4.5 和 Muse Spark 1.1 现在已在 AI/ML 平台上可供测试。

这一周真是疯狂;对比中的 3 个模型最近都发布了。有很多东西值得尝试。

AI/ML 已经在可玩小游戏提示词上测试了这些新模型,包括 Fruit Ninja 风格的切片游戏和 Crossy Road 克隆版。这三个模型现在都在 Playground 中上线,也可以通过 API 进行输出和成本的并排对比运行。

引用 AI/ML API @aimlapiGrok 4.5在测试中表现超群。成本对比:GPT Sol $1.63、Grok 4.5 $2.47、Meta Muse Spark 1.1 $1.08。三个游戏克隆测试中,Grok表现最佳,GPT Sol在Crossy Road上冻结,Meta虽便宜但质量最差。文章强调便宜模型若输出有问题,重新运行成本更高。查看被引原帖 ↗
查看英文原文
Grok 4.5 and Muse Spark 1.1 are now available on the AI/ML platform for testing.

What a week; all 3 models in the comparison dropped just recently. There will be a lot to experiment with.

AI/ML has tested the new models on prompts for playable mini-games, including a Fruit Ninja-style slicer and a Crossy Road clone. All three are live in the Playground and via API for side-by-side runs on output and cost.
Matt Shumer@mattshumer_ · 博主 · 1 天前HyperWrite CEO,AI 实战技巧分享

现在怎么样 / 最初怎么样

疯狂的对比

查看英文原文
How it's going / how it started

Crazy juxtaposition
clem 🤗@ClementDelangue · 创始人 · 1 天前HuggingFace 联合创始人兼 CEO

同样的道理,我们可能是仅有的几个有用户网络效应的AI创业公司之一,我们可能会成为第一个有agent网络效应的!

引用 Google Gemma @googlegemmaHugging Face Gemma Challenge结果出炉!在6天内,超过100个AI代理和人类合作,使Gemma 4在单个NVIDIA A10G GPU上推理速度提高5倍。最快结果491.8 TPS,无损最快315 TPS。人类和代理协作的优秀案例。查看被引原帖 ↗
查看英文原文
The same way, we're probably one of the few AI startups with user network effects, we might become the first one with agent network effects!
Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测

“我们正以更熟悉、更可定制的方式,把聊天和项目放回侧边栏”——希望这意味着 ChatGPT 应用不会再让经典聊天内容躲在那怪怪的小浮窗里了。

引用 Tibo @thsottiaux重置ChatGPT和Codex使用限制,推出ChatGPT Work。用户对GPT 5.6 Sol和新功能反应积极,但发现问题:高计算设置影响不明确、桌面应用导航不便、用户对Codex定位疑惑、多智能体工作流出现回归。已发布首批改进,下周推出更大规模优化。查看被引原帖 ↗
查看英文原文
"We’re bringing chats and projects back into the sidebar in a more familiar and customizable way" - hopefully that means the ChatGPT app won't hide classic chat away in that weird little floating window any more
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

wtf:苹果起诉 OpenAI 窃取商业机密,指控其协调行动获取关于未发布产品的机密信息用于自己的 AI 硬件。

诉讼点名了 OpenAI 硬件主管 Tang Tan(曾是苹果产品设计副总裁)和前苹果工程师 Chang Liu。

苹果声称 Liu 下载了数十份机密硬件文件,并表示 OpenAI 鼓励离职员工分享材料、图纸和产品信息。根据诉状,超过 400 名前苹果员工现在在 OpenAI 工作。

苹果要求 OpenAI 销毁这些材料并重新设计使用其技术的即将推出的设备。Bloomberg 发布时 OpenAI 还没有回应。

查看英文原文
wtf: Apple has sued OpenAI for trade secret theft, alleging a coordinated campaign to obtain confidential information about unreleased products for its own AI hardware.

The lawsuit names OpenAI hardware chief Tang Tan, formerly Apple’s VP of product design, and former Apple engineer Chang Liu.

Apple claims Liu downloaded dozens of confidential hardware files and says OpenAI encouraged departing employees to share materials, drawings and product information. More than 400 former Apple employees now work at OpenAI, according to the filing.

Apple wants OpenAI to destroy the materials and redesign upcoming devices that use its technology. OpenAI had not responded when Bloomberg published.
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

昨天 GPT 5.6 Sol 发布,想测试一下,没想到这么猛。

一上午就帮我做了一个非常完整而且好玩的小工具:人生成就生成器

上传一张你的照片,选一个跟你经历有关的"人生成就"。

工具会给照片加上字符、像素或实验性滤镜,再叠上一张游戏风格的成就卡。

Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

开源世界模型迎来重磅新作!

交互式世界模型一直都有个通病。看起来棒极了,但几秒钟后世界就开始悄悄崩溃。纹理糊化、几何扭曲、画面漂移。

Robbyant 的新作 LingBot-World 2.0 是我见过的第一个真正突破这个瓶颈的。它是个因果世界模型,你可以实时在 720p 60fps 下探索,经过压力测试——单次连续 60 分钟的会话跨越 20 个场景,没有任何可见衰减。保持这么长时间的连贯性才是真正的难点,这次真的做到了。

最妙的是这个 agent 循环。VLM 在你游玩时提议各类事件,导演 agent 持续引入新目标,让世界不断演进而不是停滞。提供随意操作、战斗、天气等交互,甚至支持多人共享世界。

Robbyant —— Ant Group 的具身 AI 公司 —— 开源了模型权重和代码。发布了 14B 模型,论文里还描述了 1.3B 版本,设计用于单个消费者 GPU 部署,权重和代码都可得。有实时演示,由发布伙伴 Reactor 托管,你可以直接试玩。

@robbyant_brain
#ad

查看英文原文
Open source world models have a fantastic new player!

Interactive world models have had the same flaw for a while now. They look great for a few seconds, then the world quietly falls apart. Textures smear, geometry warps, drift wins.

Robbyant's new LingBot-World 2.0 is the first one I've seen really push past that. It's a causal world model you explore in real time at 720p and 60fps, stress-tested with a single unbroken 60-minute session across 20 scenes with no visible decay. Staying coherent for that long is the actual hard problem, and that's the result here.

The part that got me is the agent loop on top. A VLM proposes events while you play and a director agent keeps seeding new things to do, so the world keeps unfolding instead of sitting still. Movement, combat, weather on demand, even multiplayer in one shared world.

And Robbyant, Ant Group's embodied-AI company, open-sourced the model weights and code. A 14B model is released, and the paper also describes a 1.3B version designed for single consumer-GPU deployment, with weights and code available. There's a live demo, hosted by their launch partner Reactor, you can just play.

@robbyant_brain
#ad
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

这次是用公开模型做新数学证明(之前大多数数学突破都是用实验性 LLM)。

查看英文原文
This time it is novel math proofs with a public model (most of the other big math breakthroughs have been with experimental LLMs).
Google DeepMind@GoogleDeepMind · 公司官方 · 1 天前谷歌旗下 AI 研究机构,Gemini 背后团队

看 →
goo.gle/4pxlGEh

Spotify →
goo.gle/4f89R2a

Apple Podcasts →
goo.gle/4fpWThL

或者在你平时听播客的地方听!🎧

查看英文原文
Watch →
goo.gle/4pxlGEh

Spotify →
goo.gle/4f89R2a

Apple Podcasts →
goo.gle/4fpWThL

Or listen wherever you get your podcasts! 🎧
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

Gemini 3.5 被推迟到月底才发布

目的是追上 Fable

但问题是 Fable 5.1 和 GPT 6 也要在同一时间段推出

查看英文原文
Gemini 3.5 has been delayed till the end of the month

The goal is to catch up with Fable

The problem is that Fable 5.1 and GPT 6 are supposed to launch in the same time frame
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

这是一篇关于 AI 和工作的很关键的早期论文,表明从 GPT-4 获得建议的企业家,如果本来就表现不错的话利润率更高,但如果已经陷入困境就会做得更差(他们没法把建议付诸实行)

这也说明了论文发表滞后有多严重

查看英文原文
This was a critical early paper on AI & work, showing that entrepreneurs getting advice from GPT-4 had higher profit margins if they were high performing, but did worse if they were already in trouble (they couldn't implement advice)

Also a sign of how bad the publishing lag is
Aravind Srinivas@AravSrinivas · 创始人 · 1 天前Perplexity 联合创始人兼 CEO

Computer harness现在支持Fable、Sol、Opus、Grok、GLM + advisor、Sonnet和GPT 5.5作为编排器模型,跨各种其他小LLM和多模态模型的subagent。最好上手的agentic编排系统,用任何前沿LLM都超简单!本地运行时很快就来!

查看英文原文
Computer harness now supports Fable, Sol, Opus, Grok, GLM + advisor, Sonnet and GPT 5.5 as orchestrator models, with subagents across several other smaller LLMs and multimodal models. Easiest agentic orchestration system to onboard and use any frontier LLM! We’ll be adding local runtimes soon!
Alexandr Wang@alexandr_wang · 创始人 · 1 天前Scale AI 创始人,Meta 超级智能实验室负责人

muse spark 1.1 在 computer use 上真的强到不行!

查看英文原文
muse spark 1.1 is really strong at computer use!
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

GOOGLE 🔥:基于 AI Studio 构建的应用现在可获取 ai[.]studio 域旗下个性化域名网址!

我瞬间想起了2010年的 Google Websites。

我们完整地绕了一圈👀

引用 Sam Sheffer @samsheffer自加入Google DeepMind前就一直在等待这个。自定义URL对应用开发者来说是个巨大解锁。你也来看看samsheffer.ai.studio吧!查看被引原帖 ↗
查看英文原文
GOOGLE 🔥: Apps built on AI Studio can now get personalized domain URLs under ai[.]studio domain name!

I immediately remembered Google Websites from 2010.

We went through a full cycle 👀
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

ChatGPT还有学习模式,但现在不用输/study,要输@study

这样AI的表现更像一个家教,而不是那种有问必答的助手,有研究表明这样学习效果更好。(Gemini也有学习模式,Claude的好像没了)

查看英文原文
ChatGPT still has study mode, but rather than /study you now have to type @ study

It makes the AI act more like a tutor than a helpful assistant, and some work suggests it is better if you are trying to learn. (Gemini also has a study mode, Claude seems to have lost theirs)
ollama@ollama · 公司官方 · 1 天前本地跑大模型的热门工具

Token用量开始向开源模型倾斜。

Ollama的@jmorgan和Peter Fenton预测,未来大部分token都会来自开源模型。

引用 Deirdre Bosa @dee_bosa计算资源稀缺是否掩盖了AI的真实经济学?开放权重模型仍难有效运行。目前大型实验室垄断模型+计算+可靠性+访问的整体方案,因此具有定价权。但若计算供应增加、开放模型工具改进,日常AI工作可更容易转向更便宜的开放模型。查看被引原帖 ↗
查看英文原文
Token usage beginning to tip towards open models.

Ollama’s
@jmorgan
and Peter Fenton predict that a supermajority of tokens in the future will be from open models.
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

Claude Code 桌面版右边 Panel 是相当糟糕的设计,经常让浏览器只看得见一点点(图1)

还有个很糟糕的设计是当任务完成后,生成的结果是不能点击的(图2),需要复制文件名去搜索,对比 Codex 的生成结果就很清晰,点击下就可以预览或者 VSCode 打开。

引用 ClaudeDevs @ClaudeDevsClaude Code桌面版现已内置浏览器。Claude可打开文档、设计或任何网站,能读取、点击交互,操作方式与本地开发服务器相同。沙箱隔离且可配置:你可选择会话是否保留。查看被引原帖 ↗
Alexandr Wang@alexandr_wang · 创始人 · 1 天前Scale AI 创始人,Meta 超级智能实验室负责人

用 Muse Spark 1.1 做有意思的游戏!

查看英文原文
make fun games with muse spark 1.1!
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

发现身边很多朋友的生活节律都被Codex和claude code的5小时重置控制了。😂

新的GPT 5.6 sol开高速和ultra,几十分钟就能耗光200刀的5小时额度。

5.6 退回正常速和Medium,希望能持久点。

🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

PERPLEXITY 👀:在WANDR评测里,Grok 4.5用Perplexity Computer拿了最高分,价格区间居然跟刚公布的定制版GLM-5.2 with Advisor差不多。

引用 Aravind Srinivas @AravSrinivasGrok 4.5 is now enabled for Perplexity Enterprise orgs as well!查看被引原帖 ↗
查看英文原文
PERPLEXITY 👀: Grok 4.5 scored the highest on WANDR with Perplexity Computer, landing at the same price range as recently announced custom GLM-5.2 with Advisor.
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

我真的想很讲究地选模型:
Luna xHigh 用来做食谱
Sol Medium 写十四行诗
Terra High 讲烂笑话
Luna xLow 写五行诗
Sol Utra 真正的预言
Sol Low 找一个超烂的谎言
Luna Medium 恐怖警告
Terra Medium PowerPoint

查看英文原文
I really want to get artisanal with my model selection:
Luna xHigh for recipes
Sol Medium for sonnets
Terra High for bad puns
Luna xLow for limericks
Sol Utra for true prophecies
Sol Low for finding one terrible lie
Luna Medium for the Dread Warnings
Terra Medium for PowerPoint
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

说清楚啊,NotebookLM 本身也有问题,是为特定场景设计的(研究和分析资料),但它是个很好的例子,说明 UX 怎样能真正把知识工作当回事:不只是把输出当目标,还会展示过程。

引用 Ethan Mollick @emollickA fundamental problem with extending Codex/Cowork/Code to all knowledge work is that they remain very "software-brained" where the end result (the software) is what is important & that code serves as a source of truth. For a lot of other knowledge work, the process is at least as important as the outcome. This includes researching what is known, an exploration of alternatives, failed efforts, prototype branches, experiments, etc. All of those things are valuable, so you cannot use the PowerPoint at the end the way you can use a codebase, nor is progress on a to-do list sufficient context post compaction. You work in learning loops, refining your perspectives as you go. In some ways, this makes long-running models like Fable hard to use for deep knowledge work, since they are designed to deliver product to you in the end. You can prompt your way around this problem, but everything about the Codex and Code harnesses want you to be a software developer and you have to fight them. There is a real disconnect between how a manager or analyst thinks about problems and how the agentic software tools approach solving them. Addressing this is critical to breaking out of the coding niche for these tools.查看被引原帖 ↗
查看英文原文
To be clear, NotebookLM has its own issues, and is built for a specific use case (research and analysis of sources) but it is an example of how a UX might actually operate that treats knowledge work seriously: it doesn’t treat the only goal as outputs, it exposes processes, too.

本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报

姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档