🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-06-11
- 类型
- ai-daily
- 字数
- 3671
- 阅读时长
- 18 min
2026-06-11 AI日更 | 当顶级 AI 编程比人贵,端侧模型加速接管日常 链接到标题
Anthropic 最强模型 Claude Fable 5 发布后,实测显示高频使用成本已超过雇一名人类程序员,AI 的工程化落地正撞上成本墙。同一时间,端侧模型(Gemma 4、Qwen 等)在本地推理上取得可用性突破,响应迅速但中文能力仍弱。行业方法论也从“氛围编码”转向“契约先行”,强调可测试、可审计的工程规范。成本倒挂与端侧崛起正在重塑企业部署策略。
📖 本期 Watch List 深度导读 链接到标题
Watch List 深度导读
今天有几条线索值得摆在一起深读。
第一条关于企业级 AI 与科研尖端实践的交汇。OpenAI 同时放出了两条信息:天体物理学家利用 Codex 模拟黑洞引力,以及模型通过 Oracle 云承诺直接交付。这正好回应了一个趋势——前沿模型既在推动基础科学的边界(如检验相对论),又通过成熟的采购框架加速落地。建议工程团队关注这种“同一模型,两种信任路径”的部署逻辑。
第二条是智能体工程的深度演进。我们能从多篇论文中看到对“上下文工程”与“部署时记忆”的集中探讨。《Less Context, Better Agents》直面企业级工具带来的上下文溢出,《Deployment-Time Memorization》则界定了记忆设计的隐私-效用边界。强烈推荐 AI 工程主管结合《Regimes》提出的可审计自主改进循环一起阅读,这三者共同指向了长周期智能体落地的核心难题。
最后,关于地缘视角下的认知安全,OpenAI 关于与中国相关的影响力行动瞄准美国 AI 辩论的报告需要引起重视。它揭示了技术辩论正如何被外部叙事刻意裹挟,值得所有技术决策者警惕。
🌐 X 平台 AI 热点快讯 链接到标题
话题 1:Anthropic Releases Claude Fable 5 as Most Capable Public AI Model 链接到标题
- 分类:AI · News
- 概况:热度时间:2 days ago,相关帖子数:223000
- 是什么事:Anthropic 发布了其面向公众的最强大 AI 模型 Claude Fable 5。
- 为什么重要:这标志着 AI 发展重心从单纯性能提升转向安全治理与负责任部署,反映了行业在模型能力增强后对合规和风险控制的高度重视。
- 讨论概况:社区焦点集中于模型的高阶推理能力与严格安全护栏之间的平衡,部分用户通过预测市场提前押注其发布,同时也引发了对 Anthropic 与竞争对手在性能和部署策略上的差异化讨论。
话题 2:Anthropic CEO Calls for Urgent AI Policy Overhaul 链接到标题
- 分类:AI · News
- 概况:热度时间:4 hours ago,相关帖子数:5800
- 是什么事:Anthropic CEO公开呼吁对美国现行AI政策进行紧急且彻底的改革。
- 为什么重要:此举凸显了前沿AI实验室对现有监管框架严重滞后于技术迭代的深度焦虑,可能倒逼全球主要经济体加速收紧AI治理,进而重塑行业研发布局与安全标准。
- 讨论概况:X上围绕该呼吁产生显著分歧:支持者认为防范毁灭性风险已刻不容缓,反对者批评这是Anthropic通过推动严格监管来巩固自身竞争优势的游说行为,双方聚焦于“安全主张是否沦为商业工具”。
话题 3:Google AI Studio Hits 1.2 Million Apps Per Week Milestone 链接到标题
- 分类:AI · News
- 概况:热度时间:22 hours ago,相关帖子数:342
- 是什么事:谷歌 AI Studio 平台实现了每周构建 120 万个应用的里程碑。
- 为什么重要:这一数字标志着 AI 开发工具正加速大众化,低代码与无代码模式极大降低了应用创建门槛,反映出市场对快速部署 AI 功能的强劲需求。
- 讨论概况:讨论焦点在于该 120 万数字是否包含大量一次性、调试性项目,其统计口径能否真实反映活跃开发者规模;同时延伸至与 OpenAI GPTs 等竞品的对比,以及这种爆发式增长在开发者生态中的可持续性。
话题 4:Leaked System Prompt Reveals Claude Fable 5’s Inner Workings 链接到标题
- 分类:AI · News
- 概况:热度时间:18 hours ago,相关帖子数:695
- 是什么事:Anthropic 的 Claude Fable 5 对话模型系统提示(System Prompt)遭泄露,暴露了其内部行为规则与角色设定细节。
- 为什么重要:系统提示定义了模型对齐方式和约束条件,泄露会揭示 Anthropic 如何构建安全、个性与功能边界,有助于外界评估其安全机制及可能存在的提示注入风险,对透明性和可解释性研究具有参考价值。
- 讨论概况:X 平台主要围绕泄露真实性、提示内容的克制程度、是否包含隐藏偏见或内容审查倾向展开讨论,部分用户担忧此类披露可能被用于越狱攻击,另一部分则认为曝光有助于推动开源模型对齐实践的规范化。
话题 5:Anthropic Boosts Claude with Autonomous Agent Tools 链接到标题
- 分类:AI · News
- 概况:热度时间:6 hours ago,相关帖子数:884
- 是什么事:Anthropic 展示了 Claude 的自主代理工具,使其能够处理长期运行、多步骤的复杂任务,如编程、工作流自动化等。
- 为什么重要:这标志着 AI 从对话助手向具备自主执行任务能力的代理转变,可能重塑软件开发、业务流程等领域的人机协作模式,并加剧模型能力的竞争从纯文本生成延伸到现实世界的自动化。
- 讨论概况:X 上的讨论聚焦于代理的实用性与可靠性:支持者认为这类工具将极大提升开发效率,使其从“氛围编码”进化为严谨的运维实践;质疑者则关注自主代理的权限控制、错误容忍度及可能产生的意外后果,担心在没有充分人工监督的情况下部署会带来风险。
话题 6:AI Powers Faceless YouTube Channels to Thousands in Monthly Earnings 链接到标题
- 分类:AI · News
- 概况:热度时间:,相关帖子数:28
- 是什么事:AI 工具正被大量用于自动生成内容,驱动无出镜的“面孔匿藏”YouTube 频道实现每月数千美元的收入。
- 为什么重要:此事凸显 AI 降低内容创作门槛并重塑创作者经济的潜力,同时引发关于低质内容泛滥、版权模糊及平台生态健康的深层行业拷问。
- 讨论概况:X 上的讨论集中在:工具赋能的创业机会被推崇与批判;算法是否会因 AI 内容过载而劣化;真人创作者是否面临不公平竞争;以及这类频道在披露 AI 使用、内容原创性方面的伦理争议。
今日 X 上的 AI 舆情小结 链接到标题
今天的舆论主线围绕 Anthropic 密集的动作展开,从最强模型 Claude Fable 5 的发布到系统提示泄露、自主代理亮相以及 CEO 呼吁紧急政策改革,折射出行业前沿正从性能竞赛转向安全治理与负责任部署的集体转向。共识在于 AI 正加速向大众化工具和自主执行任务转变,同时安全护栏、透明度和监管框架的滞后已成为不可回避的关切。分歧则集中体现在对 Anthropic 安全主张动机的质疑,许多人担心其推动严格监管是借安全之名构筑商业壁垒,而谷歌 AI Studio 用户数的统计口径和 AI 生成内容对创作者经济的冲击也引发了虚实难辨的争论。潜在风险在于,系统提示泄露可能被用于越狱攻击,自主代理在缺乏充分监督时易引发意外后果,而低质 AI 内容泛滥也可能导致平台生态劣化和真人创作者的不公平竞争。
💡 大佬观点(Influencer Insights) 链接到标题
AI 行业日报:24小时热点洞察 链接到标题
一、共同关注的技术趋势与产品热点 链接到标题
1. Claude Fable 5 发布:最强通用模型引发热议 链接到标题
Anthropic 发布的 Claude Fable 5 成为绝对焦点,这是首个面向普通用户的 “Mythos-class” 模型。
| 维度 | 关键信息 |
|---|---|
| 定价 | $10/百万输入Token,$50/百万输出Token(比Mythos Preview降60%) |
| 安全机制 | 分类器触发时自动降级至Opus 4.8,>95%对话不触发 |
| 数据政策 | 流量强制保留30天(重大政策变化) |
| 限时免费 | 6月10日-22日订阅用户免费使用 |
实测反馈:
- @zhixianio:“40分钟后它不但做完了,还指出了我之前设计的不合理的地方,自己用更好的方案实现了” — 效率惊人
- @dotey:“Fable 5消耗流量超快,刚升级$200套餐根本不够用” — 成本敏感
- @Pluvio9yte:“标价是opus两倍,但实际消耗没有两倍”,建议开
/effort max强度 - @dotey 对比测试结论:UI/UX设计方面Claude 4.8就够好,Fable 5未体现明显优势
2. 端侧模型(On-device Model)加速落地 链接到标题
@zhixianio 持续深耕端侧场景,多线程验证可行性:
- Gemma 4 系列:12B多模态在M5Max 128G上"英语准确度完全OK,速度非常快",但中文"驴唇不对马嘴"
- Qwen3.6-35B-A3B:配合oMLX本地运行,“响应速度比远程LLM快,智商在线”
- QAT量化感知训练:Google新思路,“训练时就假定自己一定会被量化”
应用场景拓展:从代码助手(OpenClaw/PI-Mono)到生活助手(“解冻饭团🍙"),端侧模型正在渗透日常。
3. AI Agent 浏览器:从工具到入口的跃迁 链接到标题
@vista8 重点推荐的 Aye 浏览器 代表新形态:
- 基于Chromium,完全AI模拟真人操作(非CLI/插件,规避账号检测)
- 内置Skill录制与定时执行:自动拉黑X垃圾回复、回小红书评论、转写文章到多平台
- 集成RSS阅读器、广告拦截、视频翻译/下载
产品建议:需支持Chrome账号迁移、插件生态、明确付费计划。
二、独特观点与行业前瞻 链接到标题
1. “Contract First” —— Vibe Coding的进化论 链接到标题
@Pluvio9yte 提出关键方法论转型:
“Vibe Coding的最佳实践其实并不是Requirement First或Code First,而是Contract First。没有定义好契约,其他一切都是空谈。”
基于OpenSpec二开的开发框架,将"容易漂移的上下文外化成契约”,让人和AI都有稳定参照物。这标志着AI编程从"野蛮生长"向工程化规范演进。
2. 微信AI的战略困境:创新者的窘境 链接到标题
@dotey 尖锐指出:
“微信总以为自己是OS,但它只是寄生在手机系统上的庞然大物…未来微信的入口属性会越来越少,以后的年轻人不会再去打开微信,只会问自己的Agent”
核心矛盾:微信的RPV(Resource-Process-Value)被既有生态锁死,难以做出真正AI Native的独立产品。
3. AI成本重构:从"比人便宜"到"比人贵" 链接到标题
@ruanyf 引用OpenClaw创始人数据:月消耗6030亿Token,价值$130万(若按商业定价)。即使改用国产开源模型(价格1/30-1/50),年成本仍达200-300万人民币。
结论:无限量使用顶级AI编程,比真人程序员昂贵得多。成本优化将成为AI工程的核心竞争力。
4. “影子之书"阅读法:AI时代的认知升级 链接到标题
@lijigang 提出AI原生阅读范式:
“印刷时代我们只能读作者写下的这一本书;AI时代,把那些影子之书读出来——读到任何论断,立刻调用AI分析:三个反对学派、略过的前提、思想传承、推理边界…”
阅读从"单向接收"变为"多维诘问”,AI成为认知的放大器而非替代。
5. 测试即护城河:代码护城河已崩塌 链接到标题
@ruanyf 引用Cloudflare工程师案例:用AI重写Next.js仅花费$1100 Token费用。
“防止复刻的关键是测试用例”
当AI能低成本复刻大型软件,测试体系成为区分"能用"与"可靠"的核心壁垒。
三、推荐工具与资源 链接到标题
🔧 开发工具 链接到标题
| 工具 | 用途 | 来源 |
|---|---|---|
| Fable | AI设计/开发Agent,“Shut up and take my money"级别效率 | @zhixianio |
| oMLX | macOS原生MLX推理框架,支持MTP、多模态 | @zhixianio 转推 @jundotkim |
| Owlia Nest | 端侧模型文件浏览器,配合Tailscale内网访问 | @zhixianio |
| Aye浏览器 | AI Agent专用浏览器,自动化网页操作 | @vista8 |
| 乔木提词器 | 开源口播提词器,Codex 5小时开发 | @vista8 |
| Perculia | Mac蓝牙设备快速切换工具(免费) | @vista8 |
| Bartender 6 | Mac状态栏整理($20买断) | @Pluvio9yte |
| Maccy | 开源粘贴板管理 | @Pluvio9yte |
| Screen Studio | 录屏工具第一梯队(注意闲鱼价格) | @Pluvio9yte |
📚 技能包(Skills) 链接到标题
| Skill | 功能 | 安装指令 |
|---|---|---|
| baoyu-design | 支持导入Design System的Claude Design增强 | npx skills add JimLiu/baoyu-design |
| qiaomu-book-script | 书籍解读口播脚本生成(多Subagent协作) | npx skills add joeseesun/qiaomu-book-script |
| 产品经理Skill包 | 5天13k Stars,涵盖PM日常工作流 | @vista8 评论区 |
| 中文创作者Skills合集 | 写作、改稿、去AI味、配图、封面、小红书卡片 | @wsl8297(@AI_Jasonyu 推荐) |
🎓 学习资源 链接到标题
- 《认知有县》E5:@zhixianio 播客,Owlia(端侧TTS)主持,正式讨论端侧模型
- 《千脑智能》:@lijigang 推荐,“参考系"理论对AI认知架构的启发
- 《必然》第二章"知化”:Kevin Kelly,AI时代认知升级经典
💡 基础设施 链接到标题
| 服务 | 特点 | 来源 |
|---|---|---|
| OfoxAI | 官方直连、85折活动(gpt-image-2/GPT-5.5/o3) | @AI_Jasonyu |
| Giffgaff | 0月租永久保号海外手机号(“夯"级首选) | @AI_Jasonyu |
| Vercel | “最快上线网站方式”,Codex+插件几分钟部署 | @Pluvio9yte |
四、一句话洞察 链接到标题
“世界不再奖励只会亲自做事的人” — @Pluvio9yte
AI行业正从"模型竞赛"进入"工程化落地"与"成本优化"的深水区,端侧化、Agent化、契约化成为三大确定性趋势。
📚 附录:今日 Watch List 更新源列表 链接到标题
时间窗口:最近 3 天;覆盖 22 个源;共 39 条更新
Y Combinator Podcast (B_intro+search) 链接到标题
- “The CEO Must Be the Chief AI Officer”
- 发布时间:2026-06-10 23:27 北京时间
- 摘要:- 您可能已经听说过 OpenClaw(以前称为 Clawdbot/Moltbot)。
- 引起轰动的开源人工智能助手可以在您自己的设备上运行,与您已经使用的消息应用程序连接,并且超越聊天功能,实际执行管理电子邮件、日历、文件、工作流程等任务。
- 现在来认识一下它背后的人。
- YC 的 Raphael Schaad 与 OpenClaw 的创始人 Peter Steinberger 坐下来讨论病毒式个人 AI 代理背后的“顿悟”时刻、为什么本地优先代理可以取代当今的许多应用程序,以及个人代理将如何重塑软件的未来。
- EN 要点:
- Brex co-founder and CEO Pedro Franceschi believes most people still underestimate how much AI will change the way companies are built
- AI isn’t just another tool, it’s a new foundation for building products, teams, and companies.In this episode of Lightcone, Pedro shares why he thinks we’re onl…
All-In Podcast (A_full) 链接到标题
Senators John Fetterman and Dave McCormick: Bipartisanship, Money in DC, Datacenters, Graham Platner
- 发布时间:2026-06-11 02:05 北京时间
- 摘要:- 安永 - 安永帮助私募股权公司将市场洞察转化为行动,应对复杂性并开辟新的增长和长期价值之路。
- 纽约证券交易所 - 感谢我们的合作伙伴纽约证券交易所 - 一个致力于建设未来的现代化市场和交易所。
- Plaud,我们在 All-In Liquidity Summit 上的官方可穿戴人工智能笔记合作伙伴,捕捉到了每一个见解。
- 参议员约翰·费特曼和戴夫·麦考密克:两党合作、华盛顿特区的金钱、数据中心,格雷厄姆·普拉特纳。
- EN 要点:
- (0:00) PA Senators Fetterman and McCormick join the Besties
- (0:33) Bipartisanship in 2026, rejecting extremism
- (6:37) All-time unpopularity in the Senate, the filibuster question, tribalism
- (13:33) Fixing wealth concentration in the US
Dan Dreyfus: America’s Critical Minerals Crisis is Here
- 发布时间:2026-06-10 11:04 北京时间
- 摘要:- 安永 - 流动性、增长以及组织的下一步发展是峰会的焦点。
- 安永帮助将流动性挑战转化为可持续价值。
- 纽约证券交易所 - 感谢我们的合作伙伴纽约证券交易所 - 一个致力于建设未来的现代化市场和交易所。
- Plaud,我们在 All-In Liquidity Summit 上的官方可穿戴人工智能笔记合作伙伴,捕捉到了每一个见解。
- Dan Dreyfus:美国的关键矿产危机已经到来。
- EN 要点:
- (0:00) Dan Dreyfus Presents: The Future of Critical Minerals
- (0:33) America’s “Capital Light Era” is over, rapid supply/demand shocks
- (5:40) Impact of China cutting off the US from critical minerals
- (8:18) Copper’s Rise: The next 18 years need as much as the last 10,000
Stratechery by Ben Thompson (A_full) 链接到标题
- Fable 5, Anthropic Alignment, AI Tiers
- 发布时间:2026-06-10 18:00 北京时间
- 摘要:- 《神鬼寓言 5》是《神话》的公开版本,虽然它非常强大,但它开创了一些令人不安的新先例。
- 15 美元/月或150 美元/年。
- 通过每周三封电子邮件或播客对当天新闻进行实质性分析。
- 策略采访。
- 采访领先的上市首席执行官、私营公司创始人,并与分析师同行进行讨论。
- EN 要点:
- Fable 5 is the public version of Mythos, and while it is very capable it sets some troubling new precedents.
OpenAI Blog (A_full) 链接到标题
How an astrophysicist uses Codex to help simulate black holes
- 发布时间:2026-06-11 08:00 北京时间
- 摘要:- 黑洞周围的引力是如此之大,以至于一旦距离足够近,任何东西(甚至光)都无法逃脱。
- 像 Chi-kwan Chan 这样的天体物理学家通过计算机模拟和观测来研究黑洞。
- 但当前的算法和计算能力限制了这些模拟的真实程度。
- 亚利桑那大学和斯图尔德天文台的研究员 Chan 正在通过 Codex 解决这个问题。
- 他说,黑洞是检验爱因斯坦广义相对论的最佳场所之一。
- EN 要点:
- Discover how astrophysicist Chi-kwan Chan uses Codex to build black hole simulations, helping scientists study extreme physics and test Einstein’s theory of gen…
Access OpenAI models and Codex through your Oracle cloud commitment
- 发布时间:2026-06-11 04:00 北京时间
- 摘要:- 利用现有的 Oracle 云承诺,让团队能够访问 OpenAI 最先进的模型和 Codex,而无需创建新的购买路径。
- 企业通常希望通过他们已经信任的采购流程和治理框架来部署人工智能。
- 为了帮助实现这一目标,OpenAI 和 Oracle 正在合作,让 Oracle 云基础设施 (OCI) 客户更容易访问 OpenAI 前沿模型和 Codex。
- 在未来几周内,Oracle 客户将能够通过 OCI 将符合条件的 Oracle Customer Hub (UCM) 积分应用于 OpenAI 模型和 Codex。
- 这为客户提供了在现有采购工作流程和云承诺下访问 OpenAI 模型的途径。
- EN 要点:
- Access OpenAI models and Codex through Oracle Cloud, using existing commitments to build and deploy AI with enterprise security and governance.
PRC-linked influence operations are targeting AI debates in the US
- 发布时间:2026-06-10 20:00 北京时间
- 摘要:- OpenAI 的一份新报告详细介绍了利用人工智能针对美国的与中国相关的影响力行动。
- 关于 ChatGPT 的技术辩论、数据中心叙述、关税和虚假声明。
- OpenAI 的一份新报告详细介绍了与中国相关的利用人工智能针对美国科技辩论、数据中心叙述、关税和有关 ChatGPT 的虚假声明的影响力行动。
- 与中国相关的影响力行动瞄准了美国的人工智能辩论。
- EN 要点:
- A new report from OpenAI details PRC-linked influence operations using AI to target U.S
- tech debates, data center narratives, tariffs, and false claims about ChatGPT.
From data to decisions: how LSEG is scaling trusted AI
- 发布时间:2026-06-10 08:00 北京时间
- 摘要:- 了解 LSEG 如何使用 OpenAI 在其全球业务中扩展可信 AI,加速洞察、缩短发布周期并为 4,000 名员工提供支持。
- OpenAI 博客的这篇文章解释了从数据到决策:LSEG 如何扩展可信 AI 塑造更广泛的 AI 和基础设施格局。
- 它还为创始人、运营商和投资者揭示了从数据到决策:伦敦证券交易所集团如何扩展可信人工智能的实际意义。
- EN 要点:
- See how LSEG uses OpenAI to scale trusted AI across its global business, accelerating insights, shrinking release cycles, and empowering 4,000 employees.
Google DeepMind Blog (A_full) 链接到标题
- DiffusionGemma: 4x faster text generation
- 发布时间:2026-06-11 00:24 北京时间
- 摘要:- DiffusionGemma:文本生成速度提高 4 倍。
- Google DeepMind 博客中的这篇文章解释了 DiffusionGemma:文本生成速度提高 4 倍如何塑造更广泛的人工智能和基础设施格局。
- 它还为 DiffusionGemma 的创始人、运营商和投资者带来了实际影响:文本生成速度提高了 4 倍。
- EN 要点:
- DiffusionGemma: 4x faster text generation
ArXiv cs.AI (B_intro+search) 链接到标题
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.10044v1 公告类型:新。
- 摘要:企业越来越多地采用人工智能工具来提高生产力、降低成本并增强产品和服务。
- 然而,人工智能的变革潜力不仅仅局限于自动执行预定义任务:它还在于使智能系统能够根据高层战略目标来规划、优化和执行业务计划。
- 本文介绍了业务世界模型(BWM)的概念和架构,这是一种专门针对业务和组织环境的世界模型。
- EN 要点:
- arXiv:2606.10044v1 Announce Type: new
- Abstract: Businesses are increasingly adopting AI-enabled tools to improve productivity, reduce costs, and enhance products and services
- However, the transformative potential of AI extends beyond automating predefined tasks: it lies in enabling intelligent systems to plan, optimize, and execute b…
- This paper introduces the concept and architecture of a business world model (BWM), a world model specialized for business and organizational environments
Deployment-Time Memorization in Foundation-Model Agents
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.10062v1 公告类型:新。
- 摘要:基础模型代理是越来越长寿的系统,可以在交互过程中记住用户,使记忆成为显式的部署时间函数,而不仅仅是模型权重的属性。
- 现有的工作涉及参数记忆或审核固定内存配置,但没有描述内存设计选择如何共同塑造个性化效用、提取风险和删除保真度。
- 我们将这个表面作为部署时记忆进行研究,将代理记忆制定为通过个性化回忆(PR)和对抗性提取率(AER)来衡量的隐私实用边界,并扫除三个记忆设计旋钮:摘要积极性、检索广度(k)和删除模式。
- EN 要点:
- arXiv:2606.10062v1 Announce Type: new
- Abstract: Foundation-model agents are increasingly long-lived systems that remember users across interactions, making memorization an explicit deployment-time f…
- Existing work addresses parametric memorization or audits fixed memory configurations, but does not characterize how memory-design choices jointly shape persona…
- We study this surface as deployment-time memorization, formulating agent memory as a privacy-utility frontier measured by Personalization Recall (PR) and Advers…
Exploratory Responsiveness and Adaptive Rigidity under AI-Assisted Optimization
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.10086v1 公告类型:新。
- 摘要:本文提出了人工智能辅助优化下的探索性适应理论。
- 中心论点是,人工智能系统的长期适应性效果关键取决于预测辅助如何与探索性响应本身相互作用。
- 我们使用动态框架来形式化这种机制,其中认知、制度和技术系统在以多个局部强化配置为特征的崎岖认知景观中演化。
- EN 要点:
- arXiv:2606.10086v1 Announce Type: new
- Abstract: This paper develops a theory of exploratory adaptation under AI-assisted optimization
- The central argument is that the long-run adaptive effects of AI systems depend critically on how predictive assistance interacts with exploratory responsivenes…
- We formalize this mechanism using a dynamical framework in which cognitive, institutional, and technological systems evolve over rugged epistemic landscapes cha…
Predictive Assistance and the Temporal Dynamics of Exploratory Compression
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.10094v1 公告类型:新。
- 摘要:经典认知理论将问题解决描述为通过结构化问题空间进行探索性搜索,其中重复交互逐渐将搜索压缩为有效的表征结构。
- 预测人工智能系统引入了一种独特的机制,在探索性多样化展开之前可能会出现稳定,在内部生成搜索之前提供解决方案和决策轨迹。
- 本文开发了一个几何动态框架,其中注意力在由稳定漂移、内源探索性扰动和响应性门控学习形成的策略景观中演变。
- EN 要点:
- arXiv:2606.10094v1 Announce Type: new
- Abstract: Classical theories of cognition describe problem solving as exploratory search through structured problem spaces in which repeated interaction gradual…
- Predictive artificial intelligence systems introduce a distinct regime in which stabilization may occur before exploratory diversification unfolds, supplying so…
- This paper develops a geometric dynamical framework in which attention evolves over a landscape of strategies shaped by stabilizing drift, endogenous explorator…
From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.10147v1 公告类型:新。
- 摘要:多模态大型语言模型 (MLLM) 可以听和看,但音频和视觉信号实际上如何通过网络传输以形成答案?
- 尽管音频和视觉标记在研究和现实应用中发挥着越来越大的作用,但人们对它们影响最终预测的内部途径仍然知之甚少。
- 在这项研究中,我们检查了视听大语言模型 (AVLLM) 内的视听信息流,跟踪 AVLLM 如何跨两种输入配置、视听视频和多个交错视听项目路由、利用和集成音频和视频信息。
- EN 要点:
- arXiv:2606.10147v1 Announce Type: new
- Abstract: Multimodal Large Language Models (MLLMs) can listen and see, but how do audio and visual signals actually travel through the network to shape an answe…
- Despite their growing role in research and real-world applications, the internal pathways through which audio and visual tokens influence the final prediction r…
- In this study, we examine audio-visual information flow inside Audio-Visual Large Language Models (AVLLMs), tracing how AVLLMs route, utilize, and integrate aud…
Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.10209v1 公告类型:新。
-摘要:部署为企业工作流程的自主代理的大型语言模型面临着一个关键挑战:来自企业系统的详细工具响应可能会导致上下文溢出、陈旧状态错误和高推理成本。
- 我们使用模型上下文协议工具在 Microsoft Dynamics 365 Finance and Operations 中的自动费用明细中研究此问题。
- 我们在 50 项酒店费用基准上评估了四种 GPT-5 配置:无用户模型、完整的对话历史记录、上下文修剪到最后 5 个工具调用/响应对,以及使用自动摘要进行修剪。
- EN 要点:
- arXiv:2606.10209v1 Announce Type: new
- Abstract: Large language models deployed as autonomous agents for enterprise workflows face a key challenge: verbose tool responses from enterprise systems can…
- We study this problem in automated expense itemization in Microsoft Dynamics 365 Finance and Operations using Model Context Protocol tools
- We evaluate four GPT-5 configurations on a 50-task hotel expense benchmark: no user model, full conversation history, context pruned to the last 5 tool call/res…
Minimalist Genetic Programming
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.10237v1 公告类型:新。
- 摘要:遗传编程(GP)基于两个重要的见解。
- 首先,任何学习任务从根本上都可以被视为程序归纳问题,其目标是构建一个表示为语法树的符号层次模型。
- 其次,将此任务视为搜索问题,并使用进化来定位所需的模型。
- EN 要点:
- arXiv:2606.10237v1 Announce Type: new
- Abstract: Genetic programming (GP) is based on two important insights
- First, that any learning task can fundamentally be posed as a program induction problem, where the goal is to construct a symbolic hierarchical model that is ex…
- Second, to pose this task as a search problem, and use evolution to locate the desired model
Regimes: An Auditable, Held-Out-Gated Improvement Loop Demonstrated on LongMemEval with ActiveGraph
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.10241v1 公告类型:新。
- 摘要:自主改进循环很难信任,因为改进过程通常是固定在代理上的外部脚手架:故障不会记录,诊断无法重播,升级或放弃决策会存储在辅助数据库中,而不是代理自己的历史记录中。
- 我们表明,事件源代理运行时消除了这种摩擦,并将受控改进转变为一流的工作流程。
- 当代理的状态是仅附加事件日志的确定性投影时,将记录故障,从其日志中准确重放运行,候选补丁范围为类型化管道接缝,门是可审计的,并且每次升级或丢弃本身就是一个事件。
- EN 要点:
- arXiv:2606.10241v1 Announce Type: new
- Abstract: Autonomous improvement loops are hard to trust because the improvement process is usually external scaffolding bolted onto the agent: failures go unlo…
- We show that an event-sourced agent runtime removes that friction and turns controlled improvement into a first-class workflow
- When the agent’s state is a deterministic projection of an append-only event log, failures are recorded, a run replays exactly from its log, candidate patches s…
RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.10254v1 公告类型:新。
- 摘要:虽然大型语言模型(LLM)在解决高中数学方面取得了近乎完美的表现,但它们评估真实人类学生多样化推理过程的能力仍然没有得到充分检验。
- 为了弥补这一差距,我们引入了 \textbf{RealMath-Eval},这是一个严格注释的基准,包含来自高中的 224 份真实考试答案。
- 我们的初步评估表明,即使是最先进的法学硕士评委在这项任务上也遇到了很大的困难,与专家的人工评分相比,均方误差很高($\sim$2.96)。
- EN 要点:
- arXiv:2606.10254v1 Announce Type: new
- Abstract: While Large Language Models (LLMs) have achieved near-perfect performance in \emph{solving} high-school mathematics, their ability to \emph{evaluate}…
- To bridge this gap, we introduce \textbf{RealMath-Eval}, a rigorously annotated benchmark of 224 real-world exam responses from high schools
- Our initial evaluation reveals that even state-of-the-art LLM judges struggle significantly on this task, exhibiting a high Mean Squared Error ($\sim$2.96) agai…
Supervised Fine-tuning with Synthetic Rationale Data Hurts Real-World Disease Prediction
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.10279v1 公告类型:新。
-摘要:人们普遍认为,利用合成基本原理数据进行监督微调可以提高语言模型在临床预测任务上的性能,不仅可以教授模型预测什么,还可以教授模型预测的原因。
- 我们根据纵向健康史对五年阿尔茨海默氏病和相关痴呆症 (ADRD) 预测测试了这一假设。
- 在 504 个配置的大规模受控实验中,我们发现相对于仅标签微调,基于基本原理的 SFT 一致且显着地损害了预测性能。
- EN 要点:
- arXiv:2606.10279v1 Announce Type: new
- Abstract: Supervised fine-tuning with synthetic rationale data is widely assumed to improve language model performance on clinical prediction tasks by teaching…
- We test this assumption on five-year Alzheimer’s disease and related dementias (ADRD) prediction from longitudinal health histories
- Across a large-scale controlled experiment of 504 configurations, we find that rationale-based SFT consistently and substantially hurts prediction performance r…
ArXiv cs.CL (B_intro+search) 链接到标题
Automated Scoring of Arabic Text Using Large Language Models: A Literature Review
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.09830v1 公告类型:新。
- 摘要:在现代教育系统中,自动文本评分 (ATS) 发挥着核心作用,无需人工干预即可对学习者的反应进行可扩展且一致的评估。
- 最近,法学硕士和阿拉伯语特定数据集的可访问性不断提高,引发了人们对该领域的新兴趣。
- 在这项工作中,我们研究了基于法学硕士的阿拉伯语文本自动评估方法,重点关注简答评分 (ASAG) 和论文评分 (AES)。
- EN 要点:
- arXiv:2606.09830v1 Announce Type: new
- Abstract: In modern educational systems, Automatic Text Scoring (ATS) plays a central role by enabling scalable and consistent evaluation of learner responses w…
- Recently, the increased accessibility of LLMs and Arabic-specific datasets has sparked renewed interest in this area
- In this work, we investigate LLM-Based approaches for the automated evaluation of Arabic texts, focusing on both short answer grading (ASAG) and essay scoring (…
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.09854v1 公告类型:新。
-摘要:用于政治声明分析的多智能体大语言模型(LLM)管道很容易受到同行保护偏见的影响:模型倾向于保护同行模型免于停用并显示出与身份相关的评分扭曲。
- 提出了提示级匿名化作为一种缓解措施,但之前的工作同时记录了风格指纹在角色受限的输出中可以幸免于匿名化 - 提出了这种缓解措施是否足够的问题。
- 本文首次系统地调查了法学硕士是否可以在匿名条件下识别政治分析文本背后的模型家族。
- EN 要点:
- arXiv:2606.09854v1 Announce Type: new
- Abstract: Multi-agent large language model (LLM) pipelines for political statement analysis are vulnerable to peer-preservation bias: models tend to protect pee…
- Prompt-level anonymization was proposed as a mitigation, but prior work simultaneously documented that stylometric fingerprints survive anonymization in role-co…
- This paper provides the first systematic investigation of whether LLMs can identify the model family behind political analysis texts under anonymization conditi…
Using Probabilistic Programs to Train Inductive Reasoning in Large Language Models
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.09856v1 公告类型:新。
- 摘要:用于推理的大型语言模型 (LLM) 后训练通常侧重于演绎任务,例如数学和编码,其中正确性是可验证的。
- 然而,许多现实世界的推理问题都是归纳性的:智能体必须从稀疏、模糊的观察中推断出不确定的信念。
- 使用标准微调方法进行归纳推理存在挑战,包括管理大规模、高质量标记数据集和处理本质上分布的目标方面的困难。
- EN 要点:
- arXiv:2606.09856v1 Announce Type: new
- Abstract: Post-training Large Language Models (LLMs) for reasoning typically focuses on deductive tasks such as mathematics and coding where correctness is veri…
- Yet, many real-world reasoning problems are inductive: agents must infer uncertain beliefs from sparse, ambiguous observations
- There are challenges to using standard fine-tuning methods for inductive reasoning, including difficulties in curating large-scale, high-quality labeled dataset…
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.09900v1 公告类型:新。
- 摘要:长期记忆是 LLM 代理缺失的一层:他们会忘记整个会话,而常见的解决方法(将整个历史记录重播到提示中)成本高昂、速度缓慢,而且随着干扰因素的积累,准确性也会降低。
- 大多数内存系统在成本或延迟方面获胜,但在准确性方面仍然输给了全上下文基准,并且基准数据是在不一致、不可重现的工具上报告的,因此一个系统在不同来源上的得分截然不同。
- 我们推出 Engram,这是一种基于双时态数据模型的开源双进程内存引擎。
- EN 要点:
- arXiv:2606.09900v1 Announce Type: new
- Abstract: Long-term memory is the missing layer for LLM agents: across sessions they forget, and the common workaround – replaying the whole history into the p…
- Most memory systems win on cost or latency but still lose to the full-context baseline on accuracy, and benchmark numbers are reported on inconsistent, non-repr…
- We present Engram, an open-source, dual-process memory engine on a bi-temporal data model
BenSyc: Benchmarking Conversational Sycophancy and Human Alignment in LLMs for Bengali Contexts
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.10061v1 公告类型:新。
-摘要:大型语言模型(LLM)越来越多地参与情感敏感的社交对话,其中的反应可能从平衡支持转向过度验证或升级调整。
- 现有的阿谀奉承研究主要侧重于事实一致和遵循指令的环境,而对基于文化的对话式阿谀奉承的探索还不够。
- 我们推出 BenSyc,这是研究孟加拉社会环境中对话谄媚行为的第一个基准。
- EN 要点:
- arXiv:2606.10061v1 Announce Type: new
- Abstract: Large language models (LLMs) increasingly participate in emotionally sensitive social conversations, where responses may shift from balanced support t…
- Existing sycophancy research primarily focuses on factual agreement and instruction-following settings, leaving culturally grounded conversational sycophancy un…
- We introduce BenSyc, the first benchmark for studying conversational sycophancy in Bengali social contexts
CodeAlchemy: Synthetic Code Rewriting at Scale
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.10087v1 公告类型:新。
- 摘要:对原始代码的预训练教授语法,但为不同的现实世界任务格式提供稀疏信号。
- 虽然合成数据已被证明对语言模型具有变革性,但除了有限的质量改进之外,代码在很大程度上仍未得到探索。
- 我们提出了 CodeAlchemy,一个合成数据生成框架,它通过 5 种策略将公开来源的代码转换为语义丰富的训练数据:CodeEnhance(质量感知重写)、CodeQA(基于模板的问题)、CodeDev(开发人员任务)、CodeDialogue(多轮对话)和 CodeTrace(执行跟踪)。
- EN 要点:
- arXiv:2606.10087v1 Announce Type: new
- Abstract: Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats
- While synthetic data has proven transformative for language models, code remains largely unexplored beyond limited quality improvements
- We present CodeAlchemy, a synthetic data generation framework that transforms publicly sourced code into semantically-rich training data through 5 strategies: C…
Emotion Profiling in LLM-Based Literary Translation: Systematic Shifts Across MT and Post-Editing
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.10113v1 公告类型:新。
- 摘要:本文研究了法学硕士翻译是否表现出可识别的情感特征,以及译后编辑如何将其重塑为类人规范。
- 我们使用当代意大利科幻小说的大型语料库作为基线,将玛格丽特·阿特伍德的《Oryx and Crake》的法学硕士翻译与其后期编辑版本和人工翻译进行比较。
- 我们通过基于词典的多语言建模来检查情绪,对跨系统的情绪变化进行细粒度分析。
- EN 要点:
- arXiv:2606.10113v1 Announce Type: new
- Abstract: This paper investigates whether LLM translations exhibit identifiable emotional profiles and how post-editing reshapes them toward human-like norms
- We compare LLM translations of Margaret Atwood’s Oryx and Crake with their post-edited versions and a human translation, using a large-scale corpus of contempor…
- We examine emotion through lexicon-based and multilingual modeling, conducting a fine-grained analysis of emotional variation across systems
Pareto-Guided Teacher Alignment for Fair Personalized Text Generation
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.10126v1 公告类型:新。
-摘要:个性化的说服性文本生成可以提高相关性和参与度,但人口条件也可能会在群体之间引入不平等的框架。
- 我们将个性化生成中的公平性缓解研究为受限的多目标对齐问题:减少人口差异,同时保持个性化保真度。
- 我们提出了一个帕累托引导的教师对齐框架,该框架结合了基于修订的候选生成、配对感知可行性门控、帕累托式候选选择以及通过监督微调和直接偏好优化的可选偏好优化。
- EN 要点:
- arXiv:2606.10126v1 Announce Type: new
- Abstract: Personalized persuasive text generation can improve relevance and engagement, but demographic conditioning may also introduce unequal framing across g…
- We study fairness mitigation in personalized generation as a constrained multi-objective alignment problem: reduce demographic disparities while preserving pers…
- We propose a Pareto-guided teacher alignment framework that combines revision-based candidate generation, pair-aware feasibility gating, Pareto-style candidate…
Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.10159v1 公告类型:新。
- 摘要:人工智能越来越多地用于支持科学同行评审,从稿件筛选、审稿人协助到编辑分类。
- 尽管此类系统有望减轻审稿人的负担并加速出版,但其对战略操纵的鲁棒性仍然知之甚少。
- 在这里,我们表明人工智能介导的同行评审很容易受到简单、低成本的操纵:对手稿摘要的肤浅改写。
- EN 要点:
- arXiv:2606.10159v1 Announce Type: new
- Abstract: AI is increasingly used to support scientific peer review, from manuscript screening, reviewer assistance to editorial triage
- Although such systems promise to reduce reviewer burden and accelerate publication, their robustness to strategic manipulation remains poorly understood
- Here we show that AI-mediated peer review is vulnerable to a simple, low-cost manipulation: superficial rephrasing of the manuscript abstract
OpenRTLSet: A Fully Open-Source Dataset for Large Language Model-based Verilog Module Design
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.10285v1 公告类型:新。
- 摘要:OpenRTLSet 引入了最大的完全开源硬件设计数据集,为研究社区和行业提供超过 131,000 个不同的 Verilog 代码示例。
- 我们的数据集独特地结合了来自 GitHub 存储库的 Verilog 代码(102k 模块)、VHDL 翻译(5k 模块)和可综合的 C/C++ 翻译(24k 模块),所有这些都可以免费访问,没有专有限制。
- 使用推理模型 DeepSeek-R1,我们为每个代码示例生成了配对的自然语言描述,从而能够对各种语言模型系列(例如 Qwen 和 Granite)进行微调以生成 Verilog 代码。
- EN 要点:
- arXiv:2606.10285v1 Announce Type: new
- Abstract: OpenRTLSet introduces the largest fully open-source dataset for hardware design, offering over 131,000 diverse Verilog code samples to the research co…
- Our dataset uniquely combines Verilog code from GitHub repositories (102k modules), VHDL translations (5k modules), and synthesizable C/C++ translations (24k mo…
- Using the reasoning model DeepSeek-R1, we generated paired natural language descriptions for each code sample, enabling fine-tuning of various language model fa…
ArXiv cs.LG (B_intro+search) 链接到标题
Mechanistic Analysis of Alignment Algorithms in Language Models
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.09850v1 公告类型:新。
- 摘要:训练后对齐算法主要被评估为黑匣子,模糊了它们如何重塑语言模型的内部计算。
- 我们对三个开放权重模型系列中的六种偏好优化方法进行了系统的机制分析:PPO、DPO、SimPO、ORPO、GRPO 和 KTO。
- 通过集成逐层线性探测、稀疏自动编码器和交叉编码器,我们定位偏好表示并量化潜在空间中对齐引起的几何变换。
- EN 要点:
- arXiv:2606.09850v1 Announce Type: new
- Abstract: Post-training alignment algorithms are predominantly evaluated as black boxes, obscuring how they reshape language models’ internal computations
- We present a systematic mechanistic analysis of six preference-optimization methods: PPO, DPO, SimPO, ORPO, GRPO, and KTO across three open-weight model familie…
- By integrating layer-wise linear probing, Sparse Autoencoders, and crosscoders, we localize preference representations and quantify alignment-induced geometric…
SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.09853v1 公告类型:新。
- 摘要:多模态学习的一个中心目标是捕获协同作用:仅通过联合使用多种模态产生的任务相关信息,并且不能单独从任何单一模态中获得。
- 虽然大多数方法通过更大或更复杂的融合模型在架构级别上运行,但我们提出了一个补充轴:塑造培训目标本身。
- 标准训练通常强调单模态或冗余信息,缺乏需要跨模态推理的示例。
- EN 要点:
- arXiv:2606.09853v1 Announce Type: new
- Abstract: A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities…
- While most approaches operate at the architectural level through larger or more complex fusion models, we propose a complementary axis: shaping the training obj…
- Standard training often emphasizes unimodal or redundant information, falling short on examples that require cross-modal reasoning
Uncertainty-aware Multi-fidelity Closure via Conditional Normalizing Flows
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.09857v1 公告类型:新。
-摘要:降阶模型(ROM)为复杂的多尺度系统提供了一种有效的替代方法,但它们的预测准确性常常受到截断误差以及已解析和未解析尺度之间相互作用的不充分表示的影响。
- 截断(未解析)尺度对 ROM(已解析)尺度的缺失影响通常被表示为封闭问题。
- 在这项工作中,我们将 ROM 闭包建模制定为多保真 (MF) 学习问题,并提出一种基于条件归一化流的不确定性感知 MF 框架,以提高 ROM 预测准确性。
- EN 要点:
- arXiv:2606.09857v1 Announce Type: new
- Abstract: Reduced-order models (ROMs) provide an efficient surrogate for complex multiscale systems, but their predictive accuracy is often compromised by trunc…
- The missing effect of truncated (unresolved) scales on ROM (resolved) scales is often denoted as the closure problem
- In this work, we formulate ROM closure modeling as a multi-fidelity (MF) learning problem and propose an uncertainty-aware MF framework based on conditional nor…
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.09859v1 公告类型:新。
- 摘要:MLLM 经常产生与视觉输入不一致的幻觉对象。
- 这个问题通常归因于过度依赖语言先验,这可能会覆盖视觉上下文。
- 最近的免训练解码策略通过惩罚语言先验来解决这个问题。
- EN 要点:
- arXiv:2606.09859v1 Announce Type: new
- Abstract: MLLMs frequently hallucinate objects inconsistent with visual inputs
- This issue is typically attributed to the over-reliance on language priors, which can override the visual context
- Recent training-free decoding strategies address this by penalizing language priors
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.09860v1 公告类型:新。
- 摘要:非酒精性脂肪性肝病 (NAFLD) 影响着全球约 25% 的成年人,造成巨大的肝脏和心血管风险。
- 然而,人口层面的筛查工具仍然不足。
- 我们提出了 Method,一种用于 NAFLD 风险预测的机器学习框架,将梯度增强决策树与保形预测相结合,以产生对个人风险估计的校准、无分布覆盖保证。
- EN 要点:
- arXiv:2606.09860v1 Announce Type: new
- Abstract: Non-alcoholic fatty liver disease (NAFLD) affects roughly 25% of global adults, posing substantial hepatic and cardiovascular risks
- Yet, population-level screening tools remain inadequate
- We present Method, a machine-learning framework for NAFLD risk prediction coupling gradient-boosted decision trees with conformal prediction to yield calibrated…
Time Series as Language: A Universal Tokenizer for General-Purpose Time Series Foundation Models
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.09861v1 公告类型:新。
- 摘要:虽然 Next-Token Prediction (NTP) 具有统一的 LLM 预训练,但其对无界、连续时间序列 (TS) 的适应仍然开放。
- 为了弥补这一差距,我们引入了 UniTok(一种将 TS 转换为离散标记的通用标记器)和 UniTok-FM(通过 NTP 对这些标记进行预训练的基础模型)。
- UniTok-FM 是一种通用基础模型,支持零样本和提示增强预测,以及通过免训练的上下文推理进行少样本生成和分类,这是以前的工作未实现的功能。
- EN 要点:
- arXiv:2606.09861v1 Announce Type: new
- Abstract: While Next-Token Prediction (NTP) has unified LLM pretraining, its adaptation to unbounded, continuous time series (TS) remains open
- To bridge the gap, we introduce UniTok, a universal tokenizer that transforms TS into discrete tokens, and UniTok-FM, a foundation model pretrained via NTP on t…
- UniTok-FM is a general-purpose foundation model that supports zero-shot and prompt-boosted forecasting, as well as few-shot generation and classification via tr…
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.09862v1 公告类型:新。
- 摘要:Transformer 语言模型中的 Softmax Attention 操作具有序列长度的二次复杂度以及以 KV 缓存形式不断增长的状态大小,这成为长上下文场景中的瓶颈。
- 为了克服这一限制,引入了具有线性复杂性和有限状态大小的替代架构,例如状态空间模型(SSM)、线性注意力(LA)和有限内存控制注意力(ABC)。
- 尽管线性模型实现了与变形金刚相似的语言复杂性,但它们在需要检索或回忆特定信息的任务中仍然落后。
- EN 要点:
- arXiv:2606.09862v1 Announce Type: new
- Abstract: The Softmax Attention operation in Transformer language models has a quadratic complexity in the sequence length and a growing state size in the form…
- To overcome this limitation, alternative architectures with linear complexity and finite state size have been introduced, such as State-Space Models (SSMs), Lin…
- Though linear models achieve similar language perplexity as Transformers, they are still behind in tasks which require retrieval or recall of specific informati…
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.09863v1 公告类型:新。
- 摘要:当环境状态显示其他情况时,LLM 代理可以通过断言任务完成来静默失败。
- 我们在两个代理基准测试中研究了这种失败模式,即错误的成功:来自 8 个模型系列的 9,876 个 tau2-bench 轨迹和来自 4 个模型系列的 1,879 个 AppWorld 轨迹,具有与文本无关的基本事实。
- 错误成功很常见,但因设置而异:单控制 tau2 基准域中的失败率为 45–48%,双控制电信领域为 3%,在具有明确状态声明的 AppWorld 自评估编码代理轨迹中为 75.8%。
- EN 要点:
- arXiv:2606.09863v1 Announce Type: new
- Abstract: LLM agents can fail silently by asserting task completion when the environment state shows otherwise
- We study this failure mode, false success, across two agent benchmarks: 9,876 tau2-bench trajectories from 8 model families and 1,879 AppWorld trajectories from…
- False success is common but varies by setting: 45–48% of failures in single-control tau2-bench domains, 3% in dual-control telecom, and 75.8% among AppWorld se…
Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.09864v1 公告类型:新。
- 摘要:键值(KV)缓存量化被广泛用于减少大型语言模型(LLM)推理内存,但现有的评估仅侧重于测量复杂度和准确性,而没有评估安全影响。
- 在本研究中,我们探索 KV 缓存量化下的对齐保留。
- 在 11 个指令调整模型 (3.8B-72B) 和 5 个基准测试(1,894 个提示)中,我们发现低位量化可以默默地破坏安全对齐:Mistral-7B 在仅 1.03 倍的复杂度下就失去了 15.2% 的拒绝,并且不存在通用的安全位宽,并且特定于模型的尖锐相变对于标准指标来说是不可见的。
- EN 要点:
- arXiv:2606.09864v1 Announce Type: new
- Abstract: Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measu…
- In this study, we explore alignment preservation under KV cache quantization
- Across eleven instruction-tuned models (3.8B-72B) and five benchmarks (1,894 prompts), we find that low-bit quantization can silently destroy safety alignment:…
LLM-as-a-Discriminator: When Synthetic Tables Still Look Real
- 发布时间:2026-06-10 12:00 北京时间
- 摘要:- arXiv:2606.09865v1 公告类型:新。
- 摘要:隐私和数据共享常常处于紧张状态。
- 许多组织使用合成数据来降低隐私风险,同时仍然共享有用的数据。
- 对于表格数据,审计隐私仍然很困难。
- EN 要点:
- arXiv:2606.09865v1 Announce Type: new
- Abstract: Privacy and data sharing are often in tension
- Many organizations use synthetic data to reduce privacy risk and still share useful data
- For tabular data, auditing privacy remains hard