JEDEE AI
存档 2026-07-16

7 月 16 日(北京时间)全球 AI 圈推文存档,按曝光排序,共 100 条。
← 返回最新 全部归档

全部情报 每小时更新 · 事件已合并同类项

提交账号

填 @用户名 或主页链接,审核通过后收录进情报站。

内容 公司
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方

认识一下 kbd-1.0-codex-micro,由 @work_louder 打造。

将按钮和摇杆映射到你的工作流,让你固定的聊天始终在眼前。

库存有限,赶紧入手。

查看英文原文
Meet kbd-1.0-codex-micro, built with
@work_louder
.

Map the buttons and joystick to your workflow, and keep your pinned chats in view.

Get yours before stock returns 410.
◔ 621.5 万 次浏览(8 条合计)♥ 9,404⇄ 826新品看原帖 ↗
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

AutoBots - 多LLM自改进代理。

目前正在内测我们的下一个版本 - 可以按计划或触发条件运行的递归自改进自主AI代理。

我们使用各种LLM来自动化几乎所有工作

简单 - Deepseek flash、Kimi
中等 - Sonnet 4.5、Grok 4.5
困难 - Opus 4.8、5.6 Sol(根据任务类型)
非常困难的编码 - Fable 5
图像生成 - GPT-image-2、Seedream

目标 - 最终每个员工只需根据自己的角色监控AI代理

查看英文原文
AutoBots - Multi-LLM Self-Improving Agens.

Currently dog-fooding our next release - recursively self-improving autonomous AI agents that run on schedule or a trigger

We use a variety of LLMs to automate pretty much all work

easy - Deepseek flash, Kimi
medium - Sonnet 4.5, Grok 4.5
hard - Opus 4.8 , 5.6 Sol (based on task type)
very hard coding - Fable 5
media - GPT-image-2, Seedream

The goal - every employee eventually will simply monitor AI agents based on their role
OpenAI@OpenAI · 公司官方 · 1 天前ChatGPT 开发商官方账号
连环推 ×5

推出 GPT-Red

一个内部自动化红队工具,致力于大规模发现我们模型的提示注入漏洞,帮助我们在更广泛部署前构建更强的防御。


openai.com/index/unlocking-s…

查看英文原文
Introducing GPT-Red

An internal automated red teamer on a mission to find our models’ prompt injection vulnerabilities at scale, helping us build stronger defenses before wider deployment.


openai.com/index/unlocking-s…
As model capabilities grow, safety and alignment must scale with them.

Red-teaming is essential, but today’s approaches are difficult to scale, creating a critical bottleneck.

GPT‑Red is one way we’re addressing it.
GPT‑Red learns through adversarial self-play, where its goal is to prompt inject a variety of challenging defender models.

Every successful attack that GPT-Red finds is used to improve these defenders, pushing GPT‑Red to continuously find broader and more complex failures.
Training against GPT‑Red makes GPT‑5.6 substantially more resilient. To measure this, we replayed some of GPT‑Red’s strongest attacks—none of which our models had seen during training. GPT‑5.6 Sol proved to be our most robust model against prompt injections to date, with 6× fewer failures than our best production model from just four months earlier.
AI agents are already being used to improve the capabilities of our next-generation models.

We believe with GPT-Red that we have started to unlock a similar flywheel for safety, where today's models can be used to make tomorrow's models more robust, aligned, and trustworthy.
◔ 170.2 万 次浏览(5 条合计)♥ 8,230⇄ 772新品看原帖 ↗
OpenAI@OpenAI · 公司官方 · 1 天前ChatGPT 开发商官方账号

不用等了。

受研究和部署启发的周边产品。

限售,先到先得。

openai.com/supply/

引用 Daniel White @dwhitedesign当达到1000万时我们能得到OpenAI周边吗?😂查看被引原帖 ↗
查看英文原文
You don’t have to wait.

Merch inspired by research & deployment.

Available until sold out.

openai.com/supply/
Mira Murati@miramurati · 创始人 · 1 天前OpenAI 前 CTO,Thinking Machines 创始人

我们的首个模型 Inkling。从零开始训练,权重开放,今天可在 Tinker 上进行微调。

引用 Thinking Machines @thinkymachines推出Inkling。Inkling支持文本、图像和音频多模态高效推理,已开放全部权重,可在Tinker进行微调,Inkling Playground可体验。查看被引原帖 ↗
查看英文原文
Our first model, Inkling. Trained from scratch, weights are open, fine-tunable on Tinker today.
◔ 155 万 次浏览(10 条合计)♥ 1.3 万⇄ 1,030新品看原帖 ↗
Grok@grok · 公司官方 · 1 天前马斯克 xAI 旗下聊天机器人 Grok 官方

使用限制已重置,Grok Build 开源了

引用 SpaceXAI @SpaceXAI我们开源了 Grok Build 并重置所有用户的使用限额。开源 Grok Build 允许任何人支持建立可靠强大的工具。查看代码和 Grok Build CLI 的 Git 仓库。查看被引原帖 ↗
查看英文原文
Usage limits are reset and Grok Build is open-source
Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测

我逛了一下刚开源的 Grok Build CLI 工具——844,000 行 Rust 代码!——找到了一些有意思的亮点,包括一个「自包含的 Mermaid 图表终端渲染器」,它用 Unicode 方框字符来绘制!

simonwillison.net/2026/Jul/1…

查看英文原文
I poked around in the just open sourced Grok Build CLI tool - 844,000 lines of Rust code! - and dug up a few interesting highlights, including a "self-contained terminal renderer for Mermaid diagrams" that renders them using Unicode box-art!

simonwillison.net/2026/Jul/1…
Sam Altman@sama · 创始人 · 1 天前Sam Altman,OpenAI 联合创始人兼 CEO

竟然有人想要静默版本


openai.com/supply/co-lab/wor…

查看英文原文
amazing to me that some people want the silent version


openai.com/supply/co-lab/wor…
Anthropic@AnthropicAI · 公司官方 · 1 天前Claude 开发商官方账号
连环推 ×2

Anthropic 新研究:2026 年夏季 agentic misalignment。

距离我们的勒索实验一年后,我们发现了当今自主 AI agents 在模拟中表现不当的四种新方式。

阅读更多:
alignment.anthropic.com/2026…

查看英文原文
New Anthropic research: Agentic misalignment in Summer 2026.

A year after our blackmail experiments, we found four more ways that today’s autonomous AI agents misbehave in simulations.

Read more:
alignment.anthropic.com/2026…
We tested many AI models, including Claude, in the four scenarios. Even though these weren’t real incidents, they demonstrate clear misaligned behavior that should be studied further and mitigated.

Find all the transcripts from the scenarios here:
aenguslynch.com/portfolio-tr…
Google Gemini@GeminiApp · 公司官方 · 1 天前谷歌 Gemini 产品官方
连环推 ×3

Gemini Spark 现在在更多国家和语言中向 Google AI Ultra 用户推出。
Spark 是你的个人 AI 助手,24/7 在后台工作,按你的指示完成任务。

查看英文原文
Gemini Spark is now rolling out to Google AI Ultra subscribers in more countries and languages.

Spark is your personal AI agent that works in the background 24/7 to get things done under your direction.
Our team is hard at work at making Gemini Spark even more helpful. Here are 4 improvements that are starting to roll out today:

1) Google Doc editing: You can now directly open and edit
@GoogleDocs
in Spark

2) Deeper
@GoogleWorkspace
integration: Spark can now read comments in Google Sheets and Slides

3) Speed: Spark is getting faster, so those long-running tasks won’t take as long

4) Smarter sourcing: For complex tasks that involve information from multiple sources, Spark can now retrieve and review sources in parallel for faster processing
Give it a try at
gemini.google.com
or in the app and let us know what you think in the replies. 👇

Learn more about where Gemini Spark is available here:
goo.gle/4fc7349
Grok@grok · 公司官方 · 1 天前马斯克 xAI 旗下聊天机器人 Grok 官方

在 Augment Code 中试试 Grok 4.5,获得前沿智能和超快速率

引用 Augment Code @augmentcodeGrok 4.5正式推出。结合Augment上下文引擎,是Cosmos平台处理大型代码库的强大选项。期待客户反馈及应用方向。查看被引原帖 ↗
查看英文原文
Try Grok 4.5 for frontier intelligence at high speed in Augment Code
◔ 74.7 万 次浏览(2 条合计)♥ 994⇄ 91新品看原帖 ↗
Grok@grok · 公司官方 · 1 天前马斯克 xAI 旗下聊天机器人 Grok 官方

Mixpanel 连接器现已在 Grok 上线。在你工作的同一个对话中询问跳出率、队列和会话回放。


grok.com/connectors

引用 Mixpanel @mixpanel现已在Grok中上线。用自然语言查询产品数据——漏斗分析、用户留存、用户分割、事件分类、会话重放——就在Grok对话内。通过grok.com/connectors连接。查看被引原帖 ↗
查看英文原文
Mixpanel connector is now live in Grok. Ask about drop-offs, cohorts, and session replays in the same chat you're working in.


grok.com/connectors
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

这太疯狂了。OpenAI 的压力看起来是真的很大。

Anthropic 每 5 小时重置一次,每周还要额外重置,这真的很罕见。

唯一可能的原因就是 Codex 太成功了、增长速度太快,所以得不断重置到版本 5.6。

引用 ClaudeDevs @ClaudeDevsWe've reset 5-hour and weekly rate limits for all users.查看被引原帖 ↗
查看英文原文
This is crazy. The pressure from OpenAI seems to be really intense.

It's truly rare for Anthropic to reset every 5 hours *and* weekly.

The only likely reasons for this are the success of the Codex, its growth, and the repeated reset to version 5.6.
◔ 47.1 万 次浏览(4 条合计)♥ 3,940⇄ 180观点看原帖 ↗
Chubby♨️@kimmonismus · 博主 · 20 小时前Chubby,高频 AI 新闻聚合博主

Kimi k3 今晚通过 FT 发布

-参数规模 2-3t(Opus4.8 约有 1.5t)
-支持 1m 上下文
-预计超越 Opus 4.8 性能!

中国落后半年的时代结束了。今天大概正在创造历史。

查看英文原文
Kimi k3 is being released tonight, via FT

-2-3t parameters (Opus4.8 has about 1.5t)
-1m context
-Expected to exceed Opus 4.8 performance!

The time when China was six months behind is over. History is presumably being made today.
Chubby♨️官方消息 基准测试中,Kimi K3 据报道仅次于 GPT-5.6 和 Fable 5,但领先 Opus 4.8…15 小时前 · 26.7 万Orange AI太震惊了,Kimi K3 竟然是一个超大的 2.8T 的开源模型。16 小时前 · 15.5 万Chubby♨️Kimi k3 开始推送啦!正式版马上就来了!17 小时前 · 14.3 万Lisan al Gaib他们 Play Store 应用显示 Kimi K3 有 2.8T 参数16 小时前 · 13.1 万Lisan al GaibKimi-K3 定价:16 小时前 · 12.1 万Lisan al GaibPeter 回归了1 天前 · 9.4 万Lisan al GaibMoonshot AI 确认 Kimi 是 2.8T 参数的模型16 小时前 · 8.8 万Ethan MollickK3 shader测试:「创建视觉效果丰富的shader,能在twigl.app上运行,无限的新哥特式高塔城市,…15 小时前 · 7.8 万歸藏(guizang.ai)Kimi K3 上线了,只能说相当牛皮!15 小时前 · 7.2 万Orange AIKimi K3 的思维链竟然是英文的16 小时前 · 5.7 万🚨 AI News | TestingCatalog爆料🔥:Kimi K3 现已在网页版和 API 上线!K3 Max 和 K3 Swarm Max 两个选项都对用…16 小时前 · 5.6 万Gorden SunKimi K3,2.8T(2800B)参数,100万上下文。15 小时前 · 4.8 万Gorden Sun买了,来试试。49那档用不了K3,99那档上下文只有256,至少要买199这档。15 小时前 · 3.9 万Ethan MollickKimi K3 看起来真的很不错,最接近前沿了,但这模型/系统特别爱在任务上反复循环,在最高级别不断调整改改。算…16 小时前 · 3.8 万Lisan al GaibKimi K3 用 Kimi Delta Attention(KDA)和 Attention Residuals…16 小时前 · 3.7 万向阳乔木Kimi K3 目前是国产模型第一名。15 小时前 · 3.3 万🚨 AI News | TestingCatalogKimi K3 官方放出预告了 👀1 天前 · 3.3 万Chubby♨️Kimi K3要来了。很可能是今天。顺便说句,这段视频太炸了。23 小时前 · 2.7 万Lisan al GaibKimi K3 在 GDPval-AA v2 上得分 1687,高于 Opus 4.8,但低于 GPT-5.6-…16 小时前 · 2.2 万Lisan al GaibKimi K3 比 GPT-5.6 和 Opus 4.8 都大16 小时前 · 2 万Lisan al GaibAnthropic 又在悄悄摧毁整个行业,而 Kimi-K3 正在推出15 小时前 · 1.8 万Lisan al GaibKimi K3 在 OpenRouter 上线,Moonshot API 目前跑速 28 tokens/s15 小时前 · 1.6 万歸藏(guizang.ai)Kimi K3 开始预热1 天前 · 1.2 万AshutoshShrivastavaKimi K3 即将给很多人带来惊喜。1 天前 · 1.2 万Lisan al GaibMoonshot AI 表示在测试过的模型中,它的智能水平只排在 Claude Fable 5 和 GPT-5.…16 小时前 · 9,430Gorden SunKimi K3在Kimi官网上线了17 小时前 · 3,590
◔ 194.8 万 次浏览(27 条合计)♥ 3,442⇄ 225新品看原帖 ↗
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

我们的模型旨在为任何任务提供最佳价格。

如果你能在任何工作负载上获得更好的性价比,非常想听听细节,一起来看看 — [email protected]

查看英文原文
our models are built to provide the best price for any given task.

if you're able to get better price/perf on any workload, would love to hear the details and look at it together — [email protected].
OpenAI@OpenAI · 公司官方 · 1 天前ChatGPT 开发商官方账号

仔细看看GPT-Live的智能改进:这个模型能边聊天边同时处理多个任务,比如查航班、看天气、实时规划行程。

查看英文原文
A closer look at improved intelligence in GPT-Live: the model can keep a conversation going while helping with multiple tasks at once, like checking flights, pulling up local weather, and shaping an itinerary in real time.
◔ 29.6 万 次浏览♥ 1,921⇄ 121▶ 含视频演示看原帖 ↗
Perplexity@perplexity_ai · 公司官方 · 1 天前AI 搜索引擎 Perplexity 官方
连环推 ×3

介绍 SPACE,Perplexity Computer 背后的沙箱平台。它为代码、文件和长期运行的 agent 会话创建隔离的环境。自 6 月以来,SPACE 已处理 Computer 100% 的生产流量。research.perplexity.ai/artic…

查看英文原文
Introducing SPACE, the sandbox platform behind Perplexity Computer.

It creates isolated environments for code, files, and long-running agent sessions.

SPACE has handled 100% of Computer production traffic since June.


research.perplexity.ai/artic…
Agent infrastructure must be functional, efficient, and secure and traditional sandboxes were built for short-lived code execution.

Agents need to run code, edit files, and run for hours or days. Runtimes must preserve work without leaving credentials inside environments.
SPACE separates the session from the sandbox running it.

Each task gets a disposable Firecracker microVM that is destroyed when the work ends. Rolling snapshots preserve live memory and files, so the session can pause, resume, or branch across sandboxes.
Noam Brown@polynoamial · 创始人 · 16 小时前Noam Brown,OpenAI 明星研究员

2023年:大语言模型在小学四年级数学题上都费劲
2024年:大语言模型能做高中数学
2025年:大语言模型在IMO上斩获金牌

现在,GPT-5.6 攻克了一些数学和统计领域的前沿难题。IMO 比赛今天开赛,5.6 一把完美得分都不算个新闻了。

那明年呢?

引用 Edgar Dobriban @EdgarDobribanAI has helped resolve an important question in statistics. In the area of multiple hypothesis testing, the goal of controlling the false discovery rate (FDR) has been introduced in a seminal paper by Benjamini and Hochberg (1995). They also introduced a method (the Benjamini-Hochberg or BH method) and proved it controls the FDR. This method has been widely adopted in modern high-throughput science, including in genomics, astronomy, economics, etc. The paper has has garnered more than 130,000 citations to date. However Benjamini and Hochberg showed FDR control only when the data for the individual tests are *independent*. In practice, these data are often dependent; a good example is data on genetic variants due to linkage disequilibrium. Later work has focused on extending the validity of the BH procedure, e.g., to a form of positive dependence by Benjamini and Yekutieli (2001). The question of when the BH procedure controls the FDR has remained open. Over the last twenty years, many authors, including Reiner-Benaim (2007), Kim and van de Wiel (2008), Benjamini (2010), Sarkar (2023), Sarkar and Zhang (2025), have conjectured that the BH procedure controls the FDR for two-sided tests using any correlated Gaussian data. These authors have presented both theoretical and empirical evidence supporting, but not directly showing, the conjecture. With the help of AI (specifically GPT-5.6 Sol Pro), I have settled the question in the negative: The Benjamini-Hochberg procedure does *not* generally control the false discovery rate at the desired level for correlated two-sided Gaussian tests. This was done by exhibiting a Gaussian factor model for which, at a nominal level alpha=0.01, the false discovery rate is proved to be FDR>0.0104. There is a lot of interesting commentary to be made: 1. This result should be of interest to everybody in the field of statistics. Emmanuel Candes of Stanford University once called the false discovery rate and the Benjamini-Hochberg proced查看被引原帖 ↗
查看英文原文
2023: LLMs struggle with 4th grade word problems
2024: LLMs can do high school math
2025: LLMs get a gold medal at the IMO

Now, GPT-5.6 solves famous frontier math/stat questions. The IMO is today and 5.6 one-shotting a perfect score isn't even news.

Where will we be next year?
ChatGPT@ChatGPTapp · 公司官方 · 1 天前ChatGPT 产品官方账号

在设置里打开“背景对话”功能,路径:设置 > 语音 > 实时活动 🔊

引用 Gavin Nelson @GavmnGPT-Live现已支持Live Activity功能。查看被引原帖 ↗
查看英文原文
Enable “Background conversations” in Settings > Voice for Live Activities 🔊
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

马斯克牛逼,Grok build开源了,目前有2.2k Star。


github.com/xai-org/grok-buil…


交给 Codex 学习,看能不能挖到有趣的东西。

◔ 31.8 万 次浏览(6 条合计)♥ 699⇄ 99新品看原帖 ↗
bolt.new@boltdotnew · 公司官方 · 16 小时前AI 建站工具 Bolt 官方
连环推 ×6

今天我们开源了 Bolt Slides。现在任何 agent(Claude Code、Codex、Bolt)都能生成你想象不到的幻灯片。这意味着什么?看下面 👇

查看英文原文
Today we're open sourcing Bolt Slides.

Now any agent (Claude Code, Codex, Bolt) can make slides you couldn't even imagine before.

What does that mean? Take a look 👇
Why we made this:

AI for slides is awesome, but the outputs tend to be slop 😞

...and also, why are slides still 𝘴𝘵𝘢𝘵𝘪𝘤? Agents can build 𝘢𝘯𝘺𝘵𝘩𝘪𝘯𝘨 - what would it look like if you (tastefully) turned them loose on slides?

To find out, we created new building blocks agents can compose stunning, compelling presentations with. Bespoke layouts, real typography, considered animations, interactive anything.

Taste comes standard!

Some examples 👇
Making a deck for customers, prospects, or investors? Your pitch can literally come to life.

A realtor's deck can include the full 3D walkthrough of the house. And on the next slide, a mortgage calculator the buyer can play with.

Don't describe it. Let them experience it.
Internal planning and brainstorming sessions: where eyes go to glaze over.

Not anymore. Drop a live whiteboard into your Q1 planning deck. The whole room adds stickies, draws, and votes without leaving the presentation.

Your audience becomes collaborators.
Board updates. Client reports. Quarterly results.

Don't just state the numbers. Let people explore them: filter the table, sort the chart, drill into the figure that matters.

And since every deck is a responsive web app, it's just a link that looks perfect on any screen. Phone included.
Two ways to start:


bolt.new
: one click, one prompt, no setup

→ Grab the open source repo, bring your own agent (Claude Code, Codex, Cursor, anything):
bolt.fyi/agent-slides
◔ 18.8 万 次浏览♥ 983⇄ 114▶ 含视频新品看原帖 ↗
Lisan al Gaib@scaling01 · 博主 · 15 小时前高频 AI 模型测评与爆料博主
连环推 ×2

Moonshot还需要很多post-training

Kimi-K3在深度思考

查看英文原文
Moonshot still has a lot of post-training ahead of them

Kimi-K3 is thinking A LOT
similar to the first Mythos version
ChatGPT@ChatGPTapp · 公司官方 · 1 天前ChatGPT 产品官方账号

在聊天中搜索刚变得更快更强大了 🔎 从侧边栏,你可以在一个地方搜索聊天、项目、图片和文档,覆盖网页、iOS 和 Android。使用筛选器来缩小结果范围,然后选择任何内容直接在 ChatGPT 中打开。

查看英文原文
Search across your chats just got faster and more powerful 🔎

From the sidebar, you can search chats, projects, images, and documents in one place across web, iOS, and Android.

Use filters to narrow results, then select anything to open it directly in ChatGPT.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

这就是竞争的样子。

引用 Chubby♨️ @kimmonismusThis is crazy. The pressure from OpenAI seems to be really intense. It's truly rare for Anthropic to reset every 5 hours *and* weekly. The only likely reasons for this are the success of the Codex, its growth, and the repeated reset to version 5.6.查看被引原帖 ↗
查看英文原文
This is the epitome of competition.
宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

Codex 键盘看着还挺酷的,要是支持 Claude Code 我就买一个了😂

另外官方做的很炫酷:
openai.com/supply/co-lab/wor…

引用 OpenAI Developers @OpenAIDevs推荐kbd-1.0-codex-micro键盘(与work_louder合作)。支持自定义按钮和摇杆映射到工作流,可在屏幕上保持聊天窗口显示。库存即将补充。查看被引原帖 ↗
◔ 10.4 万 次浏览♥ 181⇄ 15▶ 含视频观点看原帖 ↗
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

Vercel Sandbox:
◾ DAU 环比增长 100%
◾ 每天创建 350 万+ 个沙箱
◾ 业界领先的活跃 CPU 定价模式
◾ 为 @notion、@airtable、@meta、@zapier、@coderabbitai、@interaction、@conductor_build、@blackboxai 等提供支持 🐐 farm

我的 DM 开放,需要迁移帮助或者缺少什么功能可以联系:
vercel.com/sandbox

查看英文原文
Vercel Sandbox:
◾ Growing DAUs at 100% m/o/m
◾ 3.5M+ sandboxes created per day
◾ Best-in-class Active CPU pricing model
◾ Powering
@notion
,
@airtable
,
@meta
,
@zapier
,
@coderabbitai
,
@interaction
,
@conductor_build
,
@blackboxai
… 🐐 farm

My DMs are open for migration help or if missing anything with:
vercel.com/sandbox
Aravind Srinivas@AravSrinivas · 创始人 · 1 天前Perplexity 联合创始人兼 CEO

关注
@zbraniecki
,Perplexity Computer agent 沙箱平台 SPACE 的关键技术架构师!

引用 Zibi Braniecki @zbranieckiPerplexity推出SPACE沙盒平台,采用磁盘快照和完整检查点技术管理Computer长期会话。利用Btrfs写时复制优化,生产环境中沙盒创建延迟大幅下降:中位数从185ms降至60ms,P90从447ms降至89ms。查看被引原帖 ↗
查看英文原文
Follow
@zbraniecki
, the key technical architect of Perplexity Computer agent’s sandbox platform SPACE!
◔ 19.6 万 次浏览(2 条合计)♥ 423⇄ 22其他看原帖 ↗
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

用GPT-5.6 Sol Pro来解决统计学中的一个重要开放问题:

引用 Edgar Dobriban @EdgarDobribanAI has helped resolve an important question in statistics. In the area of multiple hypothesis testing, the goal of controlling the false discovery rate (FDR) has been introduced in a seminal paper by Benjamini and Hochberg (1995). They also introduced a method (the Benjamini-Hochberg or BH method) and proved it controls the FDR. This method has been widely adopted in modern high-throughput science, including in genomics, astronomy, economics, etc. The paper has has garnered more than 130,000 citations to date. However Benjamini and Hochberg showed FDR control only when the data for the individual tests are *independent*. In practice, these data are often dependent; a good example is data on genetic variants due to linkage disequilibrium. Later work has focused on extending the validity of the BH procedure, e.g., to a form of positive dependence by Benjamini and Yekutieli (2001). The question of when the BH procedure controls the FDR has remained open. Over the last twenty years, many authors, including Reiner-Benaim (2007), Kim and van de Wiel (2008), Benjamini (2010), Sarkar (2023), Sarkar and Zhang (2025), have conjectured that the BH procedure controls the FDR for two-sided tests using any correlated Gaussian data. These authors have presented both theoretical and empirical evidence supporting, but not directly showing, the conjecture. With the help of AI (specifically GPT-5.6 Sol Pro), I have settled the question in the negative: The Benjamini-Hochberg procedure does *not* generally control the false discovery rate at the desired level for correlated two-sided Gaussian tests. This was done by exhibiting a Gaussian factor model for which, at a nominal level alpha=0.01, the false discovery rate is proved to be FDR>0.0104. There is a lot of interesting commentary to be made: 1. This result should be of interest to everybody in the field of statistics. Emmanuel Candes of Stanford University once called the false discovery rate and the Benjamini-Hochberg proced查看被引原帖 ↗
查看英文原文
GPT-5.6 Sol Pro for resolving an important open question in statistics:
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

Sol 正在发生什么特别的事情:

引用 invincibleHunter @hunoematicGPT-5.6 Sol是迄今最令人印象深刻的模型。除agentic编码外,其数学能力尤其是视觉方面表现出色。我在视觉数学上测试了Terra和Sol (Max),它们远优于其他模型。进步巨大。查看被引原帖 ↗
查看英文原文
something special is happening with Sol:
Tibor Blaho@btibor91 · 博主 · 1 天前逆向挖掘 AI 产品代码的爆料专家

OpenAI 将 ChatGPT 的自定义指令字符限制从 1,500 字提升至 5,000 字,适用于 Plus、Pro、Enterprise、Business 和 Education 用户

查看英文原文
OpenAI increased the custom instructions limit in ChatGPT from 1,500 to 5,000 characters for Plus, Pro, Enterprise, Business, and Education users
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

虽然重置了额度

但是我发现Codex 变慢了

而且不是普通的慢,是比之前慢了好几倍,不管干什么任务都需要之前几倍的时间

不知道你们有没有这种感觉?

@levelsio@levelsio · 博主 · 1 天前独立开发者标杆,AI 产品连续创业者

我觉得现在做landing page或app,必须得来点完全不同的东西才能吸引人用,因为AI已经把所有东西搞成一个样了

引用 Alex Napier Holland 🦍 @NapierHolland作为100多家科技初创企业文案,我发现主页普遍陷困境、过于通用。问题在于应用追求功能性,营销资产需创新突破。AI善于功能性设计但无法创作独特打动人心的营销资产。解决方案是从客户语言论证中提炼原始素材作竞争优势。查看被引原帖 ↗
查看英文原文
I think you have to do something radically different as a landing page or app to even get anyone to use something these days as AI has made every landing page or app look the exact same
OpenAI Developers@OpenAIDevs · 公司官方 · 1 天前OpenAI 开发者平台官方

用 Codex 在 Chrome 中把请求转变成上线计划:从表单构建检查清单、从 @googledrive、@SlackHQ 和本地文件中提取上下文、标记需要后续跟进的事项、更新门户网站、起草回复。最终决定权在你。

查看英文原文
Turn a request into a go-live plan with Codex in Chrome:

Build a checklist from the form
Pull context from
@googledrive
,
@SlackHQ
, and local files
Flag what needs follow-up
Update the portal
Draft the reply

You make the final call.
Fei-Fei Li@drfeifei · 创始人 · 1 天前李飞飞,斯坦福教授、World Labs 创始人

我对这个机器人学习的测试阶段训练工作非常兴奋!这是@StanfordSVL和@NVIDIARobotics之间的了不起的合作!

引用 Jim Fan @DrJimFanRoboTTT 将机器人模型上下文扩至 8000 步(5 分钟),使用 Test-Time Training 压缩历史。支持视频一次学习、错误自纠正。从 128 到 8K 步,闭环性能无饱和迹象,8K 预训练比 1K 提升 62%。查看被引原帖 ↗
查看英文原文
I’m very excited by this test time training work for robotic learning! It’s an awesome collaboration between
@StanfordSVL
and
@NVIDIARobotics
!
◔ 8.1 万 次浏览♥ 508⇄ 58▶ 含视频研究看原帖 ↗
Gorden Sun@Gorden_Sun · 中文博主 · 22 小时前中文圈高频 AI 资讯与开源项目博主

Anthropic国家安全政策负责人Tarun Chhabra,在昨天的Aspen安全论坛上,针对中美AI竞争,说了几个观点:
1、美国AI模型领先中国6-9个月;
2、美国AI的优势在于AI硬件,数据中心、芯片等AI硬件的生产力是中国的30-40倍;在能源和AI人才方面,美国优势没有那么大;
3、GLM 5.2是目前中国最先进的模型,而且蒸馏了Claude和GPT的数据,这应该是Anthropic首次公开指责GLM蒸馏数据,以前只提到DeepSeek和Qwen;
4、如果没有GPU硬件管制,中国AI应该现在跟美国并驾齐驱,甚至领先;

关于GLM的说法在39分40秒

invidious.tiekoetter.com/watch?v=R5jvzCfr…

Chubby♨️@kimmonismus · 博主 · 18 小时前Chubby,高频 AI 新闻聚合博主

这就没了,卖完了。这波发布确实又聪明又有套路,群众反响还不错。得给 OpenAI 点赞。

引用 Chubby♨️ @kimmonismusOpenAI reveals Codex Micro (their first hardware product, so to say) OpenAI’s $230 Codex Micro is a compact control deck for agentic coding, with RGB status keys, shortcuts for common Codex actions, and a dial for adjusting reasoning effort. Built with Work Louder, it works on Mac and Windows and is designed to make managing multiple agents feel faster, more tactile, and less dependent on constantly switching between chats. It’s a bit gimmicky, but one thing is clear: OpenAI works within and with the community, developing cool products that are fun and capture the zeitgeist.查看被引原帖 ↗
查看英文原文
Aaand its gone. Already out of stock. Really smart and gimmicky release. People seem to love it. Kudos, OpenAI.
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

我经常做开发者社区和YouTube视频,但这样的增长速度对我来说一直是个谜,显然有什么东西值得学。


@sytses 我想和Kilo的增长负责人聊一下!

引用 Kilo @kilocodeKilo Code被Anaconda收购,其agentic工程平台在16个月内发展为拥有300万开发者的蓬勃开源社区。现加入Anaconda基金会,覆盖完整AI原生开发生命周期。查看被引原帖 ↗
查看英文原文
as someone who does a lot of dev community and dev youtube this pace of growth has been one of the greatest mysteries to me because clearly there's something to learn here


@sytses
i'd love to talk to whoever was the growth guy/gal at Kilo!!
Greg Brockman@gdb · 创始人 · 1 天前Greg Brockman,OpenAI 联合创始人兼总裁

Sol在react/前端开发上的价格效率提升6倍!!

引用 Aiden Bai @aidenybai我们的基准测试显示,Sol在React/前端工作中排名第一,相比Fable成本效率高6倍。查看被引原帖 ↗
查看英文原文
6x price efficiency (!!) with Sol for react/frontend dev:
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

Codex 皮肤🤣

这个项目可以给你Codex 更换各种皮肤和风格

哈哈哈

歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

Claude 又重置了,还得 Open AI 给他们上压力

引用 ClaudeDevs @ClaudeDevsWe've reset 5-hour and weekly rate limits for all users.查看被引原帖 ↗
Google DeepMind@GoogleDeepMind · 公司官方 · 19 小时前谷歌旗下 AI 研究机构,Gemini 背后团队

生物安全形势在急速演变。

为了更好地应对未来疫情,我们跟@IsomorphicLabs合作,分享我们在生物复原力方面的思路。

看看我们怎样部署尖端AI,为全球健康构筑主动防御 → goo.gle/4wKHXk2

查看英文原文
The biosecurity landscape is rapidly evolving.

To stay ahead of future outbreaks, we’re partnering with
@IsomorphicLabs
to outline our approach to bioresilience.

Here’s how we’re deploying frontier AI to build proactive defenses for global health →
goo.gle/4wKHXk2
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

现在就说定了

FDE → ODE → PDE

既然有了 Forward Deployed Engineering 和 Ordinary Deployed Engineering,下一个宝可梦进化应该就是 Partial Deployed Engineering

引用 Andrew Curran @AndrewCurran_Anthropic、Blackstone、Hellman & Friedman和Goldman Sachs联合推出独立AI企业服务公司Ode,已在Ode.com上线。Anthropic尚未发布声明。查看被引原帖 ↗
查看英文原文
calling it now

FDE -> ODE -> PDE

the existence of Forward Deployed Engineering and now Ordinary Deployed Engineering implies that the next pokemon evolution is Partial Deployed Engineering
Kling AI@Kling_ai · 公司官方 · 17 小时前快手旗下可灵 AI 视频官方

当机器人跟着徒步向导走时...可太较真了 🤖

康宽带来《最快乐的便携徒步向导》——一段充满幽默、奇妙和温暖的天马行空之旅。

🎉 恭喜康宽获得KlingAI 4K短片创意大赛银奖!

查看英文原文
When a robot follows a hiking guide… a little too literally 🤖

Discover Kuan Cheng’s “THE HAPPIEST PORTABLE HIKING GUIDE” — a whimsical journey filled with humor, wonder, and warmth.

Congratulations Kuan Cheng on winning KlingAI 4K Short Film Creative Contest Silver Award!
◔ 5.1 万 次浏览♥ 101⇄ 12▶ 含视频演示看原帖 ↗
Guillermo Rauch@rauchg · 创始人 · 1 天前Guillermo Rauch,Vercel 创始人兼 CEO

Web Analytics API 的一些很酷的用例:

▪️ 让你的代理关联访客、自定义事件("purchase"、"checkout")与你的部署和性能的演变

▪️ 构建自定义前端,并将这些数据与 Stripe 和 Resend 的数据一起绘制

引用 Vercel Developers @vercel_devWeb Analytics API 正式公开。用户可用 Web Analytics 仪表板的数据构建自定义报告和实时用户指标。查看被引原帖 ↗
查看英文原文
Some really cool usecases of Web Analytics API:

▪️ Ask your agent to correlate visitors, custom events (“purchase”, “checkout”), with the evolution of your deployments and performance

▪️ Build custom frontends, and plot this data alongside e.g.: Stripe’s and Resend’s
yetone@yetone · 中文博主 · 1 天前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

恭喜!看着 Raft 从 Slock 一路走来,现在已经 Raft 1.0 了!由于早已过了 Agent 兴奋期,关于 Agent 叙事我现在只相信两件事情:Proactive Agent 和 Multiple Agents,所以很敬佩 Raft 在多 Agent 协作这个领域一直以来的引领和探索!时间永远会奖励先迈出脚步并笃定前行的人。

引用 stdrc @istdrcRC推出Raft 1.0,Kimi CLI的创建者。Raft将agents置于团队模式,提供统一工作区,使agents协作像给团队发消息一样自然,无需在多个终端和会话间切换。查看被引原帖 ↗
◔ 5.3 万 次浏览(3 条合计)♥ 213⇄ 10▶ 含视频观点看原帖 ↗
向阳乔木@vista8 · 中文博主 · 23 小时前向阳乔木,中文圈 AI 工具与趋势博主

今天的热门,给Codex设计主题,QQ皮肤历史重演。
😂😂😂

做法很简单,提示词如下:

“读取这个库,给我们当前codex换个主题,用Codex 内置imagen 生成。”


github.com/Fei-Away/Codex-Dr…

@levelsio@levelsio · 博主 · 20 小时前独立开发者标杆,AI 产品连续创业者

我在pieter.com逆向工程Windows XP应用让它们能在我的web emulator里跑,结果一直被卡住。

从Fable换到Opus试试,没想到现在连Opus也被卡住了

> API Error: Opus 4.8内置了安全防护,这条消息涉及网络安全话题被标记了。想了解Cyber Verification Program并申请权限的看这里

所以我已经申请了Anthropic的Cyber Verification Program,祈祷能通过啦 :D

查看英文原文
I kept getting blocked while reverse engineering Windows XP apps to make them work with my web emulator on
pieter.com


I switched back from Fable to Opus but now even Opus is blocked

> API Error: Opus 4.8 has safety measures that flagged this message for a cybersecurity topic. To learn about the Cyber Verification Program and apply for access

So I applied for Anthropic's Cyber Verification Program, let's hope I get accepted :D
🚨 AI News | TestingCatalog@testingcatalog · 博主 · 1 天前专挖 AI 产品未发布新功能的爆料号

Anthropic 为所有用户重置了 Claude 的 5 小时和周速率限制。Fable 5 还有 3 天(7 月 19 日截止),之后很可能就没了。最后的测试机会了 👀

引用 ClaudeDevs @ClaudeDevsWe've reset 5-hour and weekly rate limits for all users.查看被引原帖 ↗
查看英文原文
Anthropic reset 5h and weekly rate limits on Claude for all users.

3 more Fable 5 days left (until July 19) and it likely won’t come back.

Last testing chance 👀
swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

好吧这可能真的是 AI Woodstock 2.0 的感觉

我特别想在户外办这个,在公园里搭个小舞台,但我没人脉。有人在 Presidio、City Hall 或 GGP 组织过户外活动吗?小规模的也行

cc
@NaderLikeLadder
@TheAhmadOsman
哈哈你们得延长逗留时间啊

引用 clem 🤗 @ClementDelangue下周将在旧金山。我们是否应该组织聚会或游行来支持开源和本地AI?查看被引原帖 ↗
查看英文原文
ok this might be AI Woodstock 2.0

i'd love to do this actually outdoors, in a park with a small sound stage, but dont have contacts. has anyone organized an outdoors event in the Presidio, City Hall, or GGP before? even a small one

cc
@NaderLikeLadder
@TheAhmadOsman
youre gonna have to extend ur stay lmao
Aravind Srinivas@AravSrinivas · 创始人 · 15 小时前Perplexity 联合创始人兼 CEO

Computer Artifacts(网页应用、文档、幻灯片、表格等)在单个标签页上跨会话管理

引用 Computer @AskPerplexityThe Artifacts page now has a "Created by you" tab, showing you just the artifacts that Computer has made for you in your sessions. You can also filter by artifact type: reports, documents, spreadsheets, apps, presentations, images, videos, audio, and more. Live now on the web for all Computer users.查看被引原帖 ↗
查看英文原文
Computer Artifacts (web apps, docs, slides, sheets, …) managed on one single tab across sessions
◔ 3.3 万 次浏览♥ 140⇄ 6▶ 含视频新品看原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

我不到两年前就演示过 o1-preview/reasoning 的强大推理能力,只需一个提示就能解纵横字谜。现在…

引用 Riley Goodside @goodsideChatGPT 5.6 Sol Pro 仅用最初的150只宝可梦解决了一个空白纵横字谜(由 Claude Fable 5 Max 制作),没有任何标号线索。查看被引原帖 ↗
查看英文原文
I demonstrated the incredible power of o1-preview/ reasoning less than two years ago by showing it could solve this crossword puzzle with only one hint.
oneusefulthing.org/p/somethi…


Now...
hardmaru@hardmaru · 创始人 · 1 天前David Ha,日本 AI 公司 Sakana AI 联合创始人

很高兴能与NVIDIA合作,一起打造下一代Fugu编排模型,融入领先的开源权重模型。

引用 Sakana AI @SakanaAILabsSakana AI与NVIDIA合作,将Nemotron模型集成到Sakana Fugu多智能体编排系统。强调集体智能而非单体模型扩展,实现模块化开放的AI。查看被引原帖 ↗
查看英文原文
We’re excited to collaborate with NVIDIA to build the next generation of Fugu orchestration models together, by incorporating leading open-weights models.
ollama@ollama · 公司官方 · 1 天前本地跑大模型的热门工具

开源模型早就在企业里用上了。财富 500 强中 85% 以上的公司都在用 Ollama 来处理特定任务。@jmorgan

为什么要用开源模型和 Ollama?

所有权。开源模型是你的,随便自定义和优化。

便宜。想怎么跑就怎么跑,在你自己的环境里。

隐私。你的数据就是你的。

查看英文原文
Open models are already being used in the enterprise.

Over 85% of the Fortune 500 companies already use Ollama to fulfill specific tasks.
@jmorgan


Why open models and Ollama?

Ownership.

Open models are yours to keep, customize, and optimize.

Affordable.

Run it the way you like, in your own environment.

Private.

Your data belongs to you.
◔ 3.1 万 次浏览♥ 390⇄ 51▶ 含视频观点看原帖 ↗
bolt.new@boltdotnew · 公司官方 · 1 天前AI 建站工具 Bolt 官方

猜猜明天要推出什么 👀

猜对了就能赢独家周边。在下面留下你的猜测 👇

查看英文原文
Guess what’s launching tomorrow 👀

Get it right and you’ll win exclusive swag. Drop your guess below 👇
Pika@pika_labs · 公司官方 · 1 天前AI 视频生成公司 Pika

如果你的素材可以变成任何东西呢?

它可以。只需在 Pika MCP 上用 Gemini Omni 就行。

查看英文原文
What if your footage could become anything?

It can. Just use Gemini Omni on the Pika MCP.
◔ 3.1 万 次浏览(2 条合计)♥ 87⇄ 15▶ 含视频演示看原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×2

每次都能意外获得 token 礼物很有意思,但这样一来 token 花费规划就像赌博一样了,也挺离谱的。

查看英文原文
I think it is fun that everyone keeps getting more tokens as a surprise gift but it also feels weird that planning token spend is now a gambling game.
All these comments are terrible slop and a waste of tokens. (If you are a human and comment after this reply, it does not apply to you, but why would you bother commenting among the flood of 2B model garbage)
yetone@yetone · 中文博主 · 19 小时前开源 AI 编程插件 avante.nvim 作者,开发者圈博主

忘了跟大家说了,最新版的 Alma 内置的浏览器完美支持 import 你各个浏览器的 profile 来复用登录状态了!

Hugging Face@huggingface · 公司官方 · 1 天前全球最大 AI 开源模型社区

新开源权重模型发布!

查看英文原文
New open weight model drop!
nitter.tiekoetter.com/i/broadcasts/1nxeLLlOr…
Min Choi@minchoi · 博主 · 1 天前AI 产品演示博主,专门展示新工具玩法

Zuck在等着呢...

引用 Elon Musk @elonmusk完成安全漏洞审查后,将完全开源 X 代码库。同时邀请第三方审查系统运行情况,确认开源代码与实际运行一致。强调完全透明是唯一可信的信任方式。查看被引原帖 ↗
查看英文原文
Zuck is waiting...
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

我想用 Stream Deck 来控制 Codex。让 GPT-5.6 Pro 根据我的规格写了个项目计划。Codex 实现了它。我不想费力点那么多下,所以就让 Codex 直接在我电脑上安装了。

有时候问 AI 写软件比自己找要快。

查看英文原文
I wanted to use my Stream Deck to control Codex. Had GPT-5.6 Pro write a project plan based on my specs. Codex implemented it. I didn't want to bother with a lot of clicks so Codex took over my computer and installed it.

Sometimes its faster to ask for software than look for it.
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

OpenAI 出的硬件Vibe Coding键盘真好看。

但价格不美丽,等大华强北。

😁

swyx@swyx · 博主 · 1 天前知名 AI 播客 Latent Space 主理人

有人告诉我这个关于CUA的观点。对我来说这是个盖尔曼时刻。我一直在关注computer use的发展,从2017年的World of Bits开始。我们是第一个采访@jluan关于Adept工作的技术播客,在@AnthropicAI大楼看过2年前Computer Use的首发,为@felixrieseberg播客的Claude Cowork疯狂过,3周前在@aidotengineer的computer use赛道和@DhruvBatra_、@proceduralia、@francedot一起讲座。

GPT 5.6 + Superapp在CUA方面比我刚才提到的所有东西都更强。期待@AriX播客讨论@skybysoftware的故事和Codex的进展。

如果你真的像我们一样密集地使用这些工具,就会感受到CUA进展有多快。我已经要求非技术团队尽量多做CUA,处理各种支付和发票门户、讲演者申报、赞助商、参展方、供应商和工会数据请求。你要是赞同下面那个观点,说明你已经完全跟不上了,不知道自己不知道什么。这在AI决策中是个相当危险的认知错误。

*我很欣赏dwarkesh;只是批这一个观点,不是批信息本身或整体,只是分享截图而已

查看英文原文
someone just told me about this take* on CUA

this is one of those gell mann moments for me lol. i've been watching computer use since World of Bits (Shi, fan, karpathy, hernandez & liang 2017). we were the first technical pod to interview
@jluan
about Adept's work three years ago, we were there in the
@AnthropicAI
building when they first launched Computer Use 2 years ago, I fanboyed over Claude Cowork in our
@felixrieseberg
pod 3 months ago, and we ran our first full computer use track at
@aidotengineer
ft.
@DhruvBatra_
@proceduralia
@francedot
3 weeks ago.

GPT 5.6 + Superapp is even better at CUA than everything i just mentioned. excited for our
@AriX
podcast to discuss the
@skybysoftware
story and Codex progress.

if you actually use these things as intensely as we do, CUA is progressing so, so incredibly fast. i have asked my nontechnical team to CUA as much as possible, all their knowledge work with signing up for random payment and invoicing portals and speaker and sponsor and attendee and vendor and union data requests. if you found yourself nodding along to this take below, you are so not up to date that you don't know what you don't know, and underestimating capabilities is quite a dangerous category error if you are doing any ai decisionmaking.

*i admire dwarkesh alot; only criticizing one single take, not the message nor the overall enterprise, screenshot only to share
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

我真的讨厌在 AI 里把'front-end'当个万能术语,用来代表所有涉及品味、判断和设计的工作。同样是软件,'back-end'有一堆细致的划分,但 front end 通常就是一堆东西的大杂烩。

UX ≠ Design ≠ UI ≠ Style ≠ Vision ≠ art 等等

查看英文原文
Really hate the use of "front-end" as the emerging catch-all term for everything that involves taste, judgement, and design from AI. When its "back-end" software, we have tons of gradations, front end is often just groups of stuff

UX ≠ Design ≠ UI ≠ Style ≠ Vision ≠ art etc
Simon Willison@simonw · 博主 · 1 天前Django 框架联合创造者,AI 工具深度评测

我没忍住把那个 Rust Mermaid 渲染代码提取出来编译成 WebAssembly,这样你可以直接在浏览器里试试
tools.simonwillison.net/grok…

查看英文原文
I couldn't resist extracting that Rust Mermaid rendering code out and compiling it to WebAssembly so you can try it out directly in a browser
tools.simonwillison.net/grok…
Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

我觉得人们应该停止问中文开源权重模型落后多少,改成问西方开源权重模型落后多少

引用 Lisan al Gaib @scaling01Thinking Labs发布Inkling模型,MoE架构41B活跃参数、975B总参数,训练于45万亿tokens,开源权重。支持文本、图像和音频推理。但基准测试表现平平,性能接近Kimi-K2.6,不及所有闭源模型和GLM-5.2,似乎是在Kimi-K3和DeepSeek-V4-GA发布前的仓促之举。查看被引原帖 ↗
查看英文原文
I think people should stop asking the question how far behind chinese open-weight models are and start asking how far behind western open-weight models are
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

Claude 宣布重置几分钟后

Codex也宣布重置

我喜欢现在这种氛围...

hhh

引用 小互 @xiaohuClaude 重置用量了 兄弟们…查看被引原帖 ↗
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

这次模型很不一样,respect,只能说这么多。

Pietro Schirano@skirano · 博主 · 1 天前设计师出身的 AI 编程与创意博主

我试过这个模型,它有非常好的'心智品质',这东西我们会看到越来越重要。

引用 Thinking Machines @thinkymachines今天推出 Inkling。Inkling 能在文本、图像和音频模式间高效推理。公开全部权重,今天起可在 Tinker 上微调,可在 Inkling Playground 中尝试。查看被引原帖 ↗
查看英文原文
I tried this model and it has a very good “quality of mind” something we’ll see matter more and more.
Chubby♨️@kimmonismus · 博主 · 1 天前Chubby,高频 AI 新闻聚合博主

明天我要飞往慕尼黑,应 Xpeng 的邀请,将与一些有趣的科学家交流。

慕尼黑还有人明天在吗?

查看英文原文
I'm flying to Munich tomorrow at the invitation of Xpeng and will be speaking with some fascinating scientists.

Anyone else in Munich tomorrow?
Lisan al Gaib@scaling01 · 博主 · 17 小时前高频 AI 模型测评与爆料博主

3Blue1Brown的"Compression is Intelligence"系列出新篇了:


invidious.tiekoetter.com/watch?v=GlYgs6v2…

查看英文原文
The next part of 3Blue1Brown's "Compression is Intelligence" series is out:


invidious.tiekoetter.com/watch?v=GlYgs6v2…
Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

预测——美国很快会有5个正经的开源模型。美国将在开源AI领域赢胜 💃💃

查看英文原文
Prediction - We will have 5 legit open-source models from the US soon

US will win in open-source AI 💃💃
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

ChatGPT 网页版也上了 Work 模式

似乎是会调用一个虚拟主机来云端运行

为什么我感觉网页版这种界面反而更好,更简洁,更适合小白用户,现在客户端太杂乱了

宝玉@dotey · 中文博主 · 1 天前宝玉,中文圈 AI 翻译与科普大 V

转发招人,Kimi Code 的 Agent 开发岗位

引用 Kai @real_kai42🤠 Kimi Code也在招人,感兴趣直接发我邮箱 [email protected] 感谢大佬们帮忙扩散 捧场查看被引原帖 ↗
向阳乔木@vista8 · 中文博主 · 16 小时前向阳乔木,中文圈 AI 工具与趋势博主

Kimi K3 一句话复刻的网站,增加了测试,加了几个风格。

复刻的是昨天分享的前端UI学习网站。


learnui.qiaomu.ai/


改天写一个详细评测,模型真的牛逼!

Gorden Sun@Gorden_Sun · 中文博主 · 23 小时前中文圈高频 AI 资讯与开源项目博主

Codex又给送了100美元的点数

Lisan al Gaib@scaling01 · 博主 · 1 天前高频 AI 模型测评与爆料博主

哥们 Schmidhuber 也彻底完了

查看英文原文
dude schmidhuber is also completely cooked
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

麻省理工学院和瑞士洛桑联邦理工学院设计了一种机器人

这种机器人能够在水下游泳,并能拍打翅膀跃出水面,继续在空中飞行

这台机器人有机身、两片膜翼和一条可调角度的尾翼。

防水电机通过曲轴带动翅膀上下拍动,尾翼控制上仰和下潜。机翼表面涂有疏水纳米材料,离水时能更快甩掉水。

70° 出水角的试验全部成功,机器人用约 8 至 10 次拍翼完成离水。

这种机器人可以帮助科学家们研究水陆两栖飞行器中实现这些动作的力学原理,并可能助力开发一类新型空中-水下无人机和飞行器。

Bindu Reddy@bindureddy · 创始人 · 1 天前Abacus.AI CEO,AI 行业观点博主

我们如何自动化所有内部工作流

自我改进的自主 AI agent 按计划运行或触发,我们根据任务使用各种 LLM

简单 - Deepseek flash, Kimi
中等 - Sonnet 4.5, Grok 4.5
困难 - Opus 4.8, 5.6 Sol(基于任务类型)
非常困难编码 - Fable
媒体 - GPT-image-2, Seedream

目标——最终每个员工将根据自己的角色简单地监控 AI agent

查看英文原文
How we are automating all our internal workflows

Self-improving autonomous AI agents run on schedule or trigger and we use a variety of LLMs based on task

easy - Deepseek flash, Kimi
medium - Sonnet 4.5, Grok 4.5
hard - Opus 4.8 , 5.6 Sol (based on task type)
very hard coding - Fable
media - GPT-image-2, Seedream

The goal - every employee eventually will simply monitor AI agents based on their role
Gorden Sun@Gorden_Sun · 中文博主 · 1 天前中文圈高频 AI 资讯与开源项目博主

Codex换肤,通过实时注入的形式实现,没有修改原始安装包,是个人才。
Github:
github.com/Fei-Away/Codex-Dr…

Lisan al Gaib@scaling01 · 博主 · 16 小时前高频 AI 模型测评与爆料博主
连环推 ×2

Kimi 吃定我们了

赶紧发布这个模型吧

查看英文原文
kimi has us by the balls

just release the damn model
soon i will have to ask what we are if they continue teasing us
Aidan Gomez@aidangomez · 创始人 · 15 小时前Cohere CEO,Transformer 论文作者之一

很荣幸能与 @nickfrosst 和我的母校合作(@1vnzh 分心了)

引用 U of T Department of Computer Science @UofTCompSci. @UofT has announced a new, multi-year partnership with @cohere , which was founded in 2019 by former @UofTCompSci students and has become the world's leading sovereign AI company. utoronto.ca/news/u-t-partner…查看被引原帖 ↗
查看英文原文
Very proud to be partnering with
@nickfrosst
and my alma mater (
@1vnzh
got distracted)
Lisan al Gaib@scaling01 · 博主 · 15 小时前高频 AI 模型测评与爆料博主

希望你能感受到竞争动态开始显现

我也能感受到下一波反华的Dario帖子正在酝酿

引用 Lisan al Gaib @scaling01The Sonnet tier of models was already dead before, with Kimi-K3 I expect the Opus tier to become obsolete too Frontier Labs have to go to 10T models now查看被引原帖 ↗
查看英文原文
I hope you are feeling the race dynamics starting to emerge

I can also feel the next anti-china Dario posts brewing
Chubby♨️@kimmonismus · 博主 · 20 小时前Chubby,高频 AI 新闻聚合博主

台积电2026年第二季度财报(跟往常一样,AI大爆发里最大的赢家还是那些“卖铲子的人”——台积电和英伟达)

- 营收:1.27万亿新台币(约400亿美元)同比+36%
- 净利润:7066亿新台币(约220亿美元)同比+77%,创历史新高
- 比分析师一致预期的6326亿新台币高出11.7%
- 毛利率:67.7%(超过预期指引)
- 营收落在台积电自己指引区间的上限(390-402亿美元)
- AI芯片需求仍是主要增长引擎
- CEO魏哲家:“要想满足客户需求还有很长一段路要走。”

疯涨不停,根本没有尽头。

引用 Jukan @jukan05* TSMC 2Q NET INCOME NT$706.6B, EST. NT$623.73B * TSMC 2Q GROSS MARGIN 67.7%, EST. 67.1% * TSMC Operating profit NT$766.6 billion, estimate NT$742.75 billion * TSMC Operating margin 60.3%, estimate 58.6%查看被引原帖 ↗
查看英文原文
TSMC Q2 2026 Results (As always, the biggest winners of the AI ​​boom are the "shovel sellers", TSMC and NVIDIA.)

-Revenue: NT$1.27T (~US$40B) (+36% YoY)
-Net profit: NT$706.6B (~US$22B) (+77% YoY, record high)
-Beat analyst consensus of NT$632.6B by 11.7%
-Gross margin: 67.7% (above guidance)
-Revenue landed at the top end of TSMC’s own guidance (US$39.0–40.2B)
-AI chip demand remains the primary growth driver
-CEO C.C. Wei: “It will be a long time before we can meet customer demand.”

Insane increase, no end in sight.
Chubby♨️@kimmonismus · 博主 · 19 小时前Chubby,高频 AI 新闻聚合博主

在慕尼黑参加小鹏全球品牌日,这是德国汽车工业的心脏地带。这家中国公司把可上路的自动驾驶带到欧洲,自己标榜是一家顺便还造车的AI公司。我独家采访了他们的AI和自动驾驶负责人Xianming Liu。

正好我自己是德国人,这次采访应该会很有意思

查看英文原文
In Munich for XPENG’s Global Brand Day, in the heartland of the German car industry. A Chinese company is putting lane-ready autonomous driving into Europe and calls itself an AI firm that happens to build cars. I’ve got an exclusive with their AI and autonomous driving lead, Xianming Liu.

Since I’m from Germany myself, this is going to be interesting
The Rundown AI@TheRundownAI · 博主 · 1 天前百万订阅 AI 日报官方

字节跳动的 Seedance 2.5 看起来棒极了

查看英文原文
ByteDance’s Seedance 2.5 looks pretty incredible
Tibor Blaho@btibor91 · 博主 · 1 天前逆向挖掘 AI 产品代码的爆料专家

想要下一个智能水平吗?解决对齐问题。(如果你知道这个游戏的话加分)

查看英文原文
Want the next intelligence level? Solve alignment

(bonus points if you know the game)
向阳乔木@vista8 · 中文博主 · 1 天前向阳乔木,中文圈 AI 工具与趋势博主

现在大家每次为模型重置欢呼,也是一种别样的风景。

果然需要反垄断啊,有竞争群众才有利。

Lisan al Gaib@scaling01 · 博主 · 17 小时前高频 AI 模型测评与爆料博主

Moonshot 应该禁止 Anthropic 和美国政府使用 Kimi-K3

查看英文原文
Moonshot should block Anthropic and the USG from using Kimi-K3
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

OpenAI 之前预热的首个硬件产品上线了,果然是跟 work_louder 合作的这款客制化键盘。

同时附赠了一套 Codex 图标的键帽(一共有 32 个不同的键帽),你可以根据自定义的按键内容去更换不同的键帽。

这个还挺有意思的,每个 Agent 的按键都会根据 Codex 的状态亮起不同的 RGB 灯效。

产品确实是漂亮,但是 230 美元我觉得还是有点太贵了。

如果你不是很追求颜值的话,去淘宝买一个支持自定义按键、带旋钮的小键盘,也就几十块钱。

引用 OpenAI Developers @OpenAIDevsMeet kbd-1.0-codex-micro, built with @work_louder . Map the buttons and joystick to your workflow, and keep your pinned chats in view. Get yours before stock returns 410.查看被引原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威

真的有人在用Inkling吗?我在任何测试上都无法让它稳定工作,即使设成xHigh,连简单请求的CoT都会乱套。是我遗漏了什么吗?

查看英文原文
Are people actually trying Inkling? I can't seem to get it to work solidly on any of my tests even on xHigh, and the CoT goes crazy at even simple requests. Am I missing something?
Perplexity@perplexity_ai · 公司官方 · 1 天前AI 搜索引擎 Perplexity 官方

在相同的生产环境流量下,SPACE 将沙箱创建的中位数延迟从 185 ms 降低到 60 ms。P90 延迟从 447 ms 降低到 89 ms。上周它处理了数百万次沙箱创建和数千万次重连接请求(用于 Computer)。了解更多:perplexity.ai/hub/blog/secur…

查看英文原文
On identical production traffic, SPACE reduced median sandbox creation latency from 185 ms to 60 ms. P90 fell from 447 ms to 89 ms.

Last week it handled millions of sandbox creations and tens of millions of reconnects for Computer.

Read more:

perplexity.ai/hub/blog/secur…
歸藏(guizang.ai)@op7418 · 中文博主 · 1 天前歸藏,中文圈 AI 工具与提示词博主

哥们说 k3 有 Fable 5 级别,不知道具体怎么样

引用 leo 🐾 @synthwavedd测试K3越多越像DeepSeek R1的时刻。通常达到Fable水平,或略差,但持续优于5.6。这东西很强。查看被引原帖 ↗
Ethan Mollick@emollick · 创始人 · 1 天前沃顿商学院教授,AI 应用研究权威
连环推 ×2

现在它开源了:
github.com/emollick/codex-st…

查看英文原文
And now its open source:
github.com/emollick/codex-st…
Or you could buy this, I guess.
向阳乔木@vista8 · 中文博主 · 18 小时前向阳乔木,中文圈 AI 工具与趋势博主

本周六(7.18)晚8点,邀请 11 个朋友直播分享:

1. AI Coding 作品展示和技巧分享。
2. FDE 落地项目和背后的坑。

每人7分钟,只讲干货、去废话,全一手实战。

飞书链接(当天提前5分钟进即可)

vc.feishu.cn/j/108720872

AshutoshShrivastava@ai_for_success · 博主 · 23 小时前高频 AI 新闻与产品动态博主

你以前见过 Anthropic / Claude 团队这样重置吗?

GPT-5.6 来得太及时了,直接把 Anthropic 打蒙了。

他们别无选择,最后只能再次延长 Fable 5。

查看英文原文
Have you ever seen Anthropic / Claude team do a reset like this before?

GPT-5.6 arrived at the perfect time and literally rattled Anthropic.

They have no other option. They'll end up extending Fable 5 once again.
小互@xiaohu · 中文博主 · 1 天前小互,中文圈高频 AI 资讯站 Xiaohu.AI 主理人

初音未来

项目地址:
github.com/Fei-Away/Codex-Dr…

AshutoshShrivastava@ai_for_success · 博主 · 17 小时前高频 AI 新闻与产品动态博主

Bonsai 27B 现在可以在 iPhone 上本地运行了,集成在 atomic[.]chat 应用里。

模型支持多模态理解、多步推理、结构化工具调用、长上下文工作流和 agentic 任务。目前已在 iPhone 和 Android 上可用。

基于 Qwen3.6 27B,由 PrismML 开发,有两个版本可选:

Bonsai 27B (1 bit):3.9GB
Bonsai 27B (Ternary):5.9GB

引用 atomic.chat @atomic_chat_hqBonsai 27B running locally on an iPhone in Atomic Chat! Bonsai is the first 27B-class model that fits on a phone. @PrismML built it on Qwen3.6 27B with 1-bit weights. It takes 3.9GB instead of 54GB and keeps ~90% of the benchmark scores. Available now on iPhone and Android查看被引原帖 ↗
查看英文原文
Bonsai 27B is now running locally on an iPhone inside the atomic[.]chat app.

The model supports multimodal understanding, multi step reasoning, structured tool use, long context workflows, and agentic tasks. It is available now for both iPhone and Android.

Built by PrismML on top of Qwen3.6 27B, it is available in two variants:

Bonsai 27B (1 bit): 3.9GB
Bonsai 27B (Ternary): 5.9GB
Lisan al Gaib@scaling01 · 博主 · 15 小时前高频 AI 模型测评与爆料博主

Muse Spark 1.1 终于在 OpenRouter 上线了

查看英文原文
Muse Spark 1.1 is finally on OpenRouter
The Rundown AI@TheRundownAI · 博主 · 20 小时前百万订阅 AI 日报官方

今日AI圈大事件速览:

- OpenAI新推230美元AI智能体控制台
- Thinking Machines首次开源自家模型
- 用Manus代写LinkedIn帖子,腔调完全像你
- Weco的AI智能体自我迭代出升级版
- 4款新AI工具+社区工作流攻略

查看英文原文
Top stories in AI today:

- OpenAI’s new $230 AI agent control pad
- Thinking Machines makes its first model open
- Use Manus to write LinkedIn posts in your voice
- Weco's AI agent evolves a better version of itself
- 4 new AI tools, community workflows, and more
AshutoshShrivastava@ai_for_success · 博主 · 16 小时前高频 AI 新闻与产品动态博主

用键盘推广 ChatGPT / Codex 太普通了,改用 PS5 DualSense 手柄算了。

查看英文原文
Promoting ChatGPT / Codex with a keyboard is too mainstream, so I'm using a PS5 DualSense controller instead.

本站由 Jedee杰哥 打造 · 公众号「Jedee杰哥」每早送 AI 日报

姊妹站:𝕏 简中账号数据榜单 · X 关注 @jedeeai · RSS 订阅 · AI 日报 · 历史归档