🤖 AI 速览

OpenAI 将 Codex 重新定位为 ChatGPT 的新形态,焦点从聊天转向可交付工作流;企业侧开始用每美元产出而非 token 价格衡量 AI 投资。与此同时,定制芯片与单卡模型继续压低部署门槛,算力与成本仍是下一阶段竞争核心。
📋 文章元数据
发布时间
2026-07-15
类型
ai-daily
字数
3415
阅读时长
17 min

2026-07-15 AI日更 | OpenAI 把 Codex 推成新 ChatGPT,AI 竞争开始看每美元产出 链接到标题

OpenAI 将 Codex 重新定位为 ChatGPT 的新形态,焦点从聊天转向可交付工作流;企业侧开始用每美元产出而非 token 价格衡量 AI 投资。与此同时,定制芯片与单卡模型继续压低部署门槛,算力与成本仍是下一阶段竞争核心。

📖 本期 Watch List 深度导读 链接到标题

今天最值得追的是“AI 从聊天入口走向工作流入口”。OpenAI Super App 与 Codex/ChatGPT 的讨论,和“agentic era 投资管理”形成呼应:真正要看的不再是 token 单价,而是每美元产出的可验证任务、节省时间与可扩展流程。

工程团队建议重点读编码智能体与推理效率几篇:Coding Agent 到底需要多少上下文、KV-Cache 压缩、本地 MoE 推理、Index 1.9B 小模型,都指向同一件事——把智能体做成可控、低成本、可部署的系统。

第三条线是可信 AI。Ground Truth 并非客观真理、AuditWeave 的证据层、蒸馏检测、临床时间序列基准,都在提醒我们:模型能力之外,数据来源、审计链路与评测边界正在成为下一阶段落地的硬门槛。

🌐 X 平台 AI 热点快讯 链接到标题

话题 1:OpenAI’s Codex and ChatGPT Work Hit 8 Million Users with Free Resets 链接到标题

  • 分类:AI · News
  • 概况:热度时间:1 day ago,相关帖子数:13000
  • 是什么事:OpenAI 的 Codex 与 ChatGPT Work 用户规模据称达到 800 万,并通过“免费重置”等方式降低使用门槛。
  • 为什么重要:这显示面向编程与办公场景的 AI 工具正在加速普及,免费或低成本策略可能进一步推动开发者和企业用户迁移到 AI 辅助工作流。
  • 讨论概况:X 上讨论集中在免费额度是否会改变 AI 工具付费模式、OpenAI 是否借此扩大生态优势,以及所谓“100+ 高级模型免费使用”的信息是否可靠、是否存在限制或营销夸大。

话题 2:OpenAI’s GPT-5.6 Sol Tops Benchmarks Amid Efficiency Fixes and Bug Reports 链接到标题

  • 分类:AI · News
  • 概况:热度时间:2 days ago,相关帖子数:30000
  • 是什么事:OpenAI 的 GPT-5.6 Sol 被曝在多项基准测试中取得领先,同时伴随效率优化进展和部分用户反馈的错误报告。
  • 为什么重要:这显示前沿大模型竞争继续围绕性能、推理效率和稳定性展开,若结果属实,可能影响企业和开发者对模型选型、成本控制及可靠性的判断。
  • 讨论概况:X 上的讨论主要集中在基准成绩是否能代表真实能力、效率修复能否降低使用成本,以及当前 bug 是否会影响实际部署;部分用户看好其领先表现,另一些人则质疑测试透明度和稳定性。

话题 3:Samsung to Produce Custom AI Chips for Anthropic, Report Says 链接到标题

  • 分类:AI · News
  • 概况:热度时间:14 hours ago,相关帖子数:753
  • 是什么事:据报道,三星将为 Anthropic 生产定制 AI 芯片,双方合作指向大模型公司自研算力供应链。
  • 为什么重要:这显示头部 AI 公司正加速降低对英伟达通用 GPU 的依赖,推动定制 ASIC、先进制程、HBM 和封装产能成为 AI 基础设施竞争核心。
  • 讨论概况:X 上讨论集中在三星能否借此缩小与台积电、SK 海力士在 AI 芯片供应链中的差距,以及 Anthropic、亚马逊等是否正在构建“英伟达替代方案”;分歧在于定制芯片短期能否真正取代 GPU,还是只会成为特定推理和云端场景的补充。

话题 4:Tencent Releases Quantized Hy3 AI Model for Single GPUs 链接到标题

  • 分类:AI · News
  • 概况:热度时间:5 hours ago,相关帖子数:186
  • 是什么事:腾讯发布了量化版 Hy3 AI 模型,主打可在单张 GPU 上运行。
  • 为什么重要:这表明大模型部署正在进一步向低成本、本地化和边缘计算场景靠拢,有助于降低企业和开发者使用 AI 模型的硬件门槛。
  • 讨论概况:X 上的讨论主要集中在模型量化后性能损失有多大、与同类开源模型相比是否具备竞争力,以及单 GPU 部署对中小开发者和本地 AI 应用的实际价值。

今日 X 上的 AI 舆情小结 链接到标题

今天舆论主线是:AI 正从“拼参数和榜单”转向“拼可用性、成本和落地门槛”,无论是 OpenAI 通过免费重置扩大编程与办公用户,还是腾讯推动单卡可跑的量化模型,都被视为在加速 AI 普及。整体共识是,低成本使用、推理效率提升和更低硬件门槛会继续推动开发者与企业迁移到 AI 工作流,同时也会促使大模型厂商争夺生态和算力供应链。分歧主要集中在两点:一是基准测试和“领先”宣传到底能否代表真实场景能力,二是定制芯片与量化模型究竟是会真正改变产业格局,还是仍只适合作为局部补充。潜在风险则在于,免费与低价策略可能伴随功能限制或营销夸大,前沿模型若稳定性不足会影响实际部署,而算力和芯片供应链的重构也可能带来新的依赖与竞争壁垒。

💡 大佬观点(Influencer Insights) 链接到标题

好的,作为资深 AI 行业分析师,我已仔细梳理了各位 AI Influencers 在过去 24 小时内的推文。以下是基于数据的洞察日报。


AI 行业每日洞察 (基于 X Influencers 推文分析) 链接到标题

1. 今日核心技术与产品热点 链接到标题

今日大佬们的注意力主要集中在前端模型能力进化、AI Agent 产品形态的边界拓展,以及对新兴社区现象的观察

  • 腾讯混元 Hy3 模型引发高度关注:

    • @ruanyf 和 @Pluvio9yte 都重点介绍了腾讯发布的旗舰模型 Hy3。@ruanyf 指出,其参数仅为 295B,远小于业界的 GLM 5.2 (744B),主打“速度快、成本低”,API 定价极具竞争力(输入 $0.15/百万token),性能达到甚至超越了 GLM 5.1 的水平,适合作为日常主力模型。
    • @Pluvio9yte 提供了更深度的应用案例,强调 Hy3 在 “工程实现、美学判断与内容策展” 方面的综合能力出色。一个典型案例是仅用两句话提示,就在 8 分钟内生成了一个完整、美观、可直接上线的《英雄联盟》角色展示网站,这在以往需要数天手工开发。
  • AI Agent 产品形态的持续演进与深度融合:

    • Codex/Work 的增长与进化:@dotey 转推指出,Codex 和 Work 的合并用户已突破 800万,并在继续重置使用配额,其增长势头迅猛。同时,他详细解释了 Chat、Work、Codex 三者定位的差异与融合:Chat 是对话,Work 是跨应用交付成品的智能体,Codex 是专注代码仓库的智能体,三者共享部分底层的 Agent 框架但用途不同。
    • Claude 的策略调整:Anthropic 宣布将高端模型 Claude Fable 5 的访问权限再度延长至 7 月 19 日。@dotey 对此评论为“儿戏一般的决策”,这从侧面反映出在 OpenAI 强大产品攻势下,Anthropic 为留住用户所展现出的竞争性姿态和策略的摇摆。
    • Apple vs OpenAI:@dotey 报道了苹果正式起诉 OpenAI 及其前员工窃取商业机密的新闻,指控其用于开发 AI 硬件。这起诉讼为科技巨头的顶级人才争夺和 AI 硬件入口之争增添了新的紧张维度。
  • 新兴社区现象的崛起:小红书 REDSkill

    • @ruanyf 敏锐地观察到一个独特的跨界现象:生活方式平台小红书开始构建 REDSkill 社区,允许用户上传和分享 AI Skill 文件。他评论称,这相当于将社媒平台与 “Skill Hub” 结合,全球尚无先例,可能是小红书为平台 AI 转型和增加技术内容铺垫,对开发者而言则是一个触及海量用户的巨大新分发渠道。

2. 值得注意的独特观点与行业前瞻 链接到标题

  • 关于“省 Token”的迷思 (by @dotey): 宝玉对流行的“电报体 Skill”(如 Caveman 项目,号称节省 65% Token)进行了深度剖析。他引用 JetBrains 的测试结果指出,在真实编程任务中,输出 Token 仅节省 8.5%。他认为这种优化是针对“聊天场景”的,而 Agent 真正的成本在工具调用和系统提示。在 API 价格持续下行的趋势下,与其纠结“说话像穴居人”,不如优化上下文管理和减少路径返工。这为“省钱型”Skill 的开发泼了一盆冷水,点明了优化的主次矛盾。

  • 个人角色与能力的重塑 (by @dotey, @vista8):

    • @dotey 分享了 Anthropic 的内部观察:Agent 的“脚手架”正在变薄,重点从“控制每一步”转向“设计 Agent 间的协作”。同时他强调,Agent 能放大个人能力,但不会自动解决团队协调问题,产品可能因个人快速试错而无序扩张。
    • @vista8 引用 IBM 1979 年的幻灯片指出,“计算机永远不能被追究责任,因此绝不能做出管理决策”。AI 同理,未来真正稀缺的是在信息不足时做决策并为后果负责的能力。他观察到的“老板想沉淀员工经验为 Skill,而员工抵制”的现象,也深刻揭示了AI时代下企业内部知识和价值归属的新矛盾。
  • “端侧模型”与“模型卡带”的想象力 (by @zhixianio):

    • @zhixianio 通过对 Gemma 4 12B Coder 的实际评测,给出了一个反直觉的结论:尽管社区热捧,但 12B 参数模型的“天花板”明显,难以胜任需要长篇、一次性生成的复杂程序。这提醒我们,小模型的能力边界依然清晰,微调提升效率但抬不高上限
    • 他赞同 @geekbb 的 “Model-Pak” 畅想,认为未来的端侧模型可以像插拔卡带一样分发和使用,这与 Google 正在大力推动的端侧 AI 趋势(如 Gemma QAT 模型)相呼应,预示着一个新的离网 AI 应用形态。

3. 推荐的实用工具与资源 链接到标题

  • 本地代码生成模型评测参考:@zhixianio 提供了 Gemma 4 12B Coder vs. Qwen 3.6-35B-A3B MoE 的详细对比测评,结论是 35B 的 MoE 模型在当前仍是本地代码生成的“甜点”级别,可以作为开发者选择本地模型的重要参考。
  • AI 剪辑工具 Skill
    • @Pluvio9yte 和 @dotey 都关注了 ChatCut。它是一个可与 Codex/Claude Code 集成的 AI 剪辑 Skill,能根据转录自动剪辑视频、删除口癖等。
    • @dotey 发布了自己开发的 BaoCut 字幕转录翻译剪辑 Skill(仅 Mac),并特别强调了其解决 Agent 生成后“二次编辑”难题的思路——通过 CLI 与 GUI 配合,为 Agent 提供友好的操作界面。
  • AI 视频复刻与综合创作:@Pluvio9yte 分享了其“视频复刻 Skill”在 Sol 模型加持下的强大能力,并推荐用 腾讯 Hy3 进行需要兼顾工程、审美和内容策划的复杂创作。
  • 开源项目与教学资源
    • @vista8 推荐了用 AI 开发的 模型 PK 擂台,可一键对比多个模型的文本和前端输出。
    • @vista8 推荐了 fireworks-tech-graph (8.5k Star),一个由社区自发推广的专业技术图绘制开源项目。
    • @vista8 同时提醒,AI 剪辑工具的效果可能一般,自己搭建 Listenhub CLI + Remotion 的工作流依然是追求更高质量的可靠选择。

📚 附录:今日 Watch List 更新源列表 链接到标题

时间窗口:最近 3 天;覆盖 22 个源;共 34 条更新

Stratechery by Ben Thompson (A_full) 链接到标题

  • The OpenAI Super App, ChatGPT = Codex, Whither Chat
    • 发布时间:2026-07-14 18:00 北京时间
    • 摘要:- OpenAI 将 Codex 重塑为新的 ChatGPT;该公司是否正在放弃他们开创的聊天类别?
      • 15 美元/月150 美元/年。
      • 通过每周三封电子邮件或播客对当天新闻进行实质性分析。
      • 策略采访
      • 采访领先的上市首席执行官、私营公司创始人,并与分析师同行进行讨论。
    • EN 要点:
      • OpenAI has refashioned Codex as the new ChatGPT; is the company abandoning the chat category they pioneered

OpenAI Blog (A_full) 链接到标题

  • How to manage AI investments in the agentic era

    • 发布时间:2026-07-14 18:00 北京时间
    • 摘要:- OpenAI 的目标是随着时间的推移,让人工智能变得更容易获取、更有能力、更经济实惠。
      • 从 GPT-4 到 GPT-5.4,每百万代币的价格下降了 97%。
      • 但仅凭代币价格并不能表明人工智能是否正在创造价值。
      • 领导者应该关注每一美元的有用工作:完成的任务、节省的时间、改进的决策以及准备扩展的工作流程。
      • 随着团队从聊天转向运行时间较长的工作流程,管理员需要更清晰地了解需求、支出和风险。
    • EN 要点:
      • Learn how enterprises can manage AI investments in the agentic era by measuring useful work per dollar, improving efficiency, and scaling high-value workflows.
  • How data science teams use ChatGPT Work

    • 发布时间:2026-07-14 08:00 北京时间
    • 摘要:- 了解数据科学团队如何使用 ChatGPT Work 将问题、仪表板和原始数据转化为可供审查的分析资产。
      • 借助 ChatGPT Work,数据科学团队可以更快地将分散的输入转化为可用的分析资产。
      • 从仪表板、指标定义、导出、实验注释和业务上下文开始,ChatGPT Work 帮助组装可交付成果的初稿(包括图表、说明、源链接和审核问题),以便团队可以验证工作并自信地共享。
      • 观看点播网络研讨会。 链接到标题

      • 注意:本网络研讨会是在这些工作流程位于以前的 Codex 应用程序中时录制的。
    • EN 要点:
      • See how data science teams can use ChatGPT Work to build root-cause briefs, impact readouts, KPI memos, scoped analyses, and dashboard specs from real work inpu…
  • How sales teams use ChatGPT Work

    • 发布时间:2026-07-14 08:00 北京时间
    • 摘要:- 了解销售团队如何使用 ChatGPT Work 创建渠道简报、会议准备包、预测审查、客户计划和真实的停滞交易诊断……。
      • OpenAI 博客中的这篇文章解释了销售团队如何使用 ChatGPT Work 塑造更广泛的人工智能和基础设施格局。
      • 遵循销售团队如何使用 ChatGPT Work,它还为创始人、运营商和投资者提供了实际意义。
    • EN 要点:
      • See how sales teams can use ChatGPT Work to create pipeline briefs, meeting prep packets, forecast reviews, account plans, and stalled-deal diagnoses from real…

ArXiv cs.AI (B_intro+search) 链接到标题

  • From ML Predictions to Informed Diagnostic Assistance Using the Toulmin Model of Argumentation

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09664v1 公告类型:新。
      • 摘要:为了提供结构化且可解释的评估,我们将基于图像的诊断分解为遵循图尔明论证模型的组件。
      • 该模型由主张、理由、保证、限定词、反驳和支持组成。
      • 考虑用于视网膜诊断的机器学习 (ML) 模型生成的声明。
    • EN 要点:
      • arXiv:2607.09664v1 Announce Type: new
      • Abstract: To provide a structured and interpretable assessment, we decompose the image-based diagnosis into components following the Toulmin model of argumentat…
      • This model consists of a claim, grounds, warrant, qualifier, rebuttal, and backing
      • Consider a claim generated by a machine learning (ML) model for retinal diagnosis
  • Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09665v1 公告类型:新。
      • 摘要:提示包装器通常仅在格式上有所不同,但它们可以改变模型分数足以推翻排行榜结论。
      • 我们在令牌控制协议下研究这种方差,并引入两个补充指标:格式敏感度指数(FSI),由包装器选择引起的准确度范围,以及可解析性敏感度指数(PSI),答案可解析性的相应范围。
      • 在跨越 7 个 QA 任务、5 个包装器系列和 4 个从 7B 到 72B 参数的指令模型的 140,000 个 OpenRouter 代中,我们发现模型之间的平均 FSI 变化超过 30 倍,这在很大程度上是由合规性失败造成的。
    • EN 要点:
      • arXiv:2607.09665v1 Announce Type: new
      • Abstract: Prompt wrappers often differ only in formatting, yet they can change model scores enough to flip leaderboard conclusions
      • We study this variance under a token-controlled protocol and introduce two complementary metrics: the Format Sensitivity Index (FSI), the accuracy range induced…
      • Across 140,000 OpenRouter generations spanning 7 QA tasks, 5 wrapper families, and 4 instruct models from 7B to 72B parameters, we find that mean FSI varies by…
  • Faithful, Not Corrective: Message-Format Effects in Multi-Hop Agent Relays Are Tier-Dependent

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09678v1 公告类型:新。
      • 摘要:当LLM代理相互传递信息时,消息格式重要吗?
      • 两篇文献存在分歧:格式优化工作报告称,结构化消息可以在不损害准确性的情况下降低成本,而格式限制工作发现,强加结构会降低生成速度——并且两者都没有衡量当消息穿越多个跃点时会发生什么,在这种情况下,复制保真度而不是一次性生成占主导地位。
      • 我们引入了一个受控中继测试平台:十二个以编程方式生成的原子事实的摘要在六跳上以五种格式(免费 NL、精确指导的 NL、JSON、三元组、键值)逐跳重新编码,由固定的强分级器针对编程地面事实进行评分,跨两个中继能力层、认知负载条件和配对叉错误注入。
    • EN 要点:
      • arXiv:2607.09678v1 Announce Type: new
      • Abstract: When LLM agents hand off information to one another, does the message format matter
      • Two literatures disagree: format-optimization work reports that structured messages cut cost without hurting accuracy, while format-restriction work finds that…
      • We introduce a controlled relay testbed: briefs of twelve programmatically generated atomic facts are re-encoded hop-by-hop in five formats (free NL, precision-…
  • Boltzmann MapReduce: A Partition-Function Reduce for Forkable Sandboxes

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09689v1 公告类型:新。
      • 摘要:对于局部渐近正态性(LAN)下的领先顺序,工人在大小为$n$的块上发出的置信密度是吉布斯-玻尔兹曼测度$\exp{-\beta E(\theta)}$,其反温度是样本大小$\beta=n$。
      • 三个结果在高斯/线性情况下是精确的,否则是一阶的:不相交的块带有独立的玻尔兹曼因子,因此从字面上看,MapReduce \emph{reduce} 是一个分区函数 $Z=\int\prod_k h_k,d\theta$ ,其模式是精度加权(逆方差)池化;频率一致性是零温度限制 $T=1/n\to0$。
      • arXiv:2607.09689v1 公告类型:新摘要:对于局部渐近正态性(LAN)下的领先顺序,工作人员在大小为 $n$ 的块上发出的置信密度是吉布斯 - 玻尔兹曼测度…三个结果在高斯/线性情况下是精确的,否则是一阶的:不相交的块带有独立的玻尔兹曼因子,因此 MapReduce \emph{….
    • EN 要点:
      • arXiv:2607.09689v1 Announce Type: new
      • Abstract: To leading order under local asymptotic normality (LAN), the confidence density a worker emits over a chunk of size $n$ is a Gibbs–Boltzmann measure…
      • Three consequences are exact in the Gaussian/linear case and first-order otherwise: disjoint chunks carry independent Boltzmann factors, so the MapReduce \emph{…
  • Interpreting Latent CoT Reasoning as Dynamical Systems

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09698v1 公告类型:新。
      • 摘要:最近的潜在推理方法,例如 CODI 和 COCONUT,面临着一个基本的可解释性问题:它们在每一步的隐藏空间中维护多个叠加的候选轨迹,这与显式 CoT 不同,显式 CoT 遵循单个透明推理轨迹。
      • 现有的机械方法显示了压缩、快捷方式和叠加,但没有解释推理如何跨潜在步骤演变。
      • 为了解决这一差距,我们将潜在标记序列建模为表示空间中的轨迹,并应用动态系统分析来表征推理的演化。
    • EN 要点:
      • arXiv:2607.09698v1 Announce Type: new
      • Abstract: Recent latent reasoning methods, such as CODI and COCONUT, face a fundamental interpretability problem: they maintain multiple superimposed candidate…
      • Existing mechanistic methods show compression, shortcuts, and superposition without explaining how reasoning evolves across latent steps
      • To address this gap, we model latent token sequences as trajectories in representation space and apply dynamical systems analysis to characterize the evolution…
  • YUKTI: From Natural-Language Situations to Robust, Verifiable Decisions An Uncertainty-Typed Proposition IR, Assumption-Robust Pareto Frontiers, and a Regret Certificate

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09706v1 公告类型:新。
      • 摘要:语言模型将措辞情况转化为数字计划,并且主要管道(NL4Opt、OptiMUS、ORLM、OR-LLM-Agent)致力于单个目标和点值系数,然后求解一次。
      • 对于分配实际预算、工作量或临床注意力的决策,这种信心就是失败模式:每个客观化的数字都是一个假设,只有猜测完全正确的最佳计划是脆弱的——模拟计算。
      • YUKTI 改变了自动配制的目标。
    • EN 要点:
      • arXiv:2607.09706v1 Announce Type: new
      • Abstract: Language models turn a worded situation into a numeric plan, and the dominant pipelines (NL4Opt, OptiMUS, ORLM, OR-LLM-Agent) commit to a single objec…
      • For decisions that allocate real budget, effort, or clinical attention, that confidence is the failure mode: every objectified number is an assumption, and a pl…
      • YUKTI changes the target of autoformulation
  • GES-TSP: Graph Edge Sparsification for TSP

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09708v1 公告类型:新。
      • 摘要:解决旅行商问题(TSP)的大规模实例在计算上是昂贵的。
      • 研究人员经常采用图稀疏方法来提高计算效率。
      • 传统的稀疏方法通常依赖于固定的启发式方法,无法充分利用特定于实例的结构信息。
    • EN 要点:
      • arXiv:2607.09708v1 Announce Type: new
      • Abstract: Solving large-scale instances of the Traveling Salesman Problem (TSP) exactly is computationally expensive
      • Researchers often employ graph sparsification methods to improve computational efficiency
      • Traditional sparsification methods typically rely on fixed heuristics and fail to fully exploit instance-specific structural information
  • The Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game Generation

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09709v1 公告类型:新。
      • 摘要:根据学习过的判断对代码生成器进行后训练可以优化代理功能,从而在不改进工件的情况下提高分数。
      • 我们研究相反的信号:确定性的、无判断的、不可游戏的过滤器——生成的项目是否在无头引擎下干净地启动(严格启动)。
      • 在这个门下,拒绝采样自蒸馏复合了家族外泛化。
    • EN 要点:
      • arXiv:2607.09709v1 Announce Type: new
      • Abstract: Post-training a code generator against a learned judge can optimize proxy features that raise the score without improving the artifact
      • We study the opposite signal: a deterministic, judge-free, ungameable filter – whether a generated project launches cleanly under a headless engine (strict-lau…
      • Under this gate, rejection-sampling self-distillation compounds out-of-family generalization
  • Closed-Loop Control with Rule-Aligned Small Language Models and Multi-Agent Self-Correction

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09713v1 公告类型:新。
      • 摘要:实现自主工业运营的关键一步是能够根据自然语言需求规范创建和重新配置控制策略,而无需或很少进行手动重新设计。
      • 在这种情况下,当与植物感知验证器(例如数字孪生)配合使用时,人工智能代理生成的策略可以是一条可靠的路径,该验证器可以在执行前检查生成的候选操作。
      • 然而,实际部署受到推理延迟和计算占用的限制:基于云的大型模型对于边缘闭环使用通常太慢、不透明或数据敏感。
    • EN 要点:
      • arXiv:2607.09713v1 Announce Type: new
      • Abstract: A key step toward autonomous industrial operation is the ability to create and reconfigure control policies from natural-language requirement specific…
      • In this setting, policy generation by AI agents can be a credible path when paired with a plant-aware validator (e.g., a digital twin) that can check generated…
      • However, practical deployment is constrained by inference latency and compute footprint: large cloud-based models are often too slow, opaque, or data-sensitive…
  • Feedback-Coupled Memory Systems in Continuous Time

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09714v1 公告类型:新。
      • 摘要:反馈耦合内存系统 (FCMS) 架构通过四个抽象运算符形式化闭环协调,其中两个 - 代理更新运算符 $f_i$ 和环境更新运算符 $\Psi$ - 在原始框架中未公理地定义。
      • 为了解决这个问题,$f_i$ 由基于机制的智能(MBI)定义,其中代理通过去中心化的价格机制和经济原理进行本地更新,而 $\Psi$ 由耦合记忆图过程(CMGP)定义,这是一个非马尔可夫框架,其中环境被视为物理基质,可以在没有外部强迫的情况下连贯地记录和响应轨迹历史。
      • 由此产生的连续时间 FCMS 实例化实现了由可计算阈值 $4\beta^2 < 2\eta\mu\gamma^2$ 控制的 Lyapunov 全局耗散性。
    • EN 要点:
      • arXiv:2607.09714v1 Announce Type: new
      • Abstract: The Feedback-Coupled Memory Systems (FCMS) architecture formalizes closed-loop coordination through four abstract operators, two of which - the agent…
      • To address this, $f_i$ is defined by Mechanism-Based Intelligence (MBI), where agents update locally through a decentralized price mechanism and economic princi…
      • The resulting continuous-time FCMS instantiation achieves Lyapunov global dissipativity governed by the computable threshold $4\beta^2 < 2\eta\mu\gamma^2$

ArXiv cs.CL (B_intro+search) 链接到标题

  • CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09880v1 公告类型:新。
      • 摘要:临床时间序列对于患者监测、风险评估和临床决策支持至关重要。
      • 然而,它们通常是稀疏的、不规则采样且异步的,使得模型难以识别临床问答(QA)所需的时间证据。
      • 现有基准主要关注静态数据上的定期采样时间序列 QA 或医学 QA,因此很少评估模型是否能够忠实地在不规则时间观察中得出答案。
    • EN 要点:
      • arXiv:2607.09880v1 Announce Type: new
      • Abstract: Clinical time series are central to patient monitoring, risk assessment, and clinical decision support
      • However, they are often sparse, irregularly sampled, and asynchronous, making it difficult for models to identify the temporal evidence required for clinical Qu…
      • Existing benchmarks primarily focus on regularly sampled time-series QA or medical QA over static data, and therefore rarely assess whether models can faithfull…
  • Index SLM Technical Report

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09885v1 公告类型:新。
      • 摘要:我们推出 Index-1.9B,这是 Bilibili 开发的一系列开放小语言模型。
      • 该系列包括四个模型: Index-1.9B-Base,一个基础模型,具有 19 亿个非嵌入参数,在 2.8 万亿个主要是中文和英文的标记上进行预训练; Index-1.9B-Pure,一种使用相同配方训练的控制变体,但所有类似指令的数据都从语料库中严格过滤; Index-1.9B-Chat,从基础模型出发,进行监督微调和直接偏好优化; Index-1.9B-Character,它通过检索增强生成来增强聊天模型,以实现少量角色扮演定制。
      • 预训练采用了 Warmup-Stable-Decay 学习率计划,其中在衰减阶段大幅提高了策划数据的浓度,以及在大学习率下稳定训练的 Norm-Head 输出层。
    • EN 要点:
      • arXiv:2607.09885v1 Announce Type: new
      • Abstract: We present Index-1.9B, a series of open small language models developed at Bilibili
      • The series comprises four models: Index-1.9B-Base, a foundation model with 1.9 billion non-embedding parameters pre-trained on 2.8 trillion predominantly Chines…
      • Pre-training employs a Warmup-Stable-Decay learning-rate schedule in which the concentration of curated data is raised substantially during the decay phase, tog…
  • RouteRec: Strict Evaluation of Recommender-Agent Selection and Aggregation

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09908v1 公告类型:新。
      • 摘要:推荐系统越来越多地面临异构代理(协作过滤器、顺序模型、基于内容的检索器和基于 LLM 的重新排序器)之间的选择,但没有一个代理是一致最好的。
      • 我们使用 RouteRec 将这种选择作为成本约束下的任务感知代理排名进行研究,RouteRec 是一个框架,该框架将四个传统推荐代理和一个 LLM 重新排序代理的请求级硬选择与项目级学习聚合进行比较。
      • 在 MovieLens-1M 上,全质量预言机具有很大的余量 (HR@10 = 0.584),确认存在有用的跨代理信号。
    • EN 要点:
      • arXiv:2607.09908v1 Announce Type: new
      • Abstract: Recommender systems increasingly face a choice among heterogeneous agents – collaborative filters, sequential models, content-based retrievers, and L…
      • We study this choice as task-aware agent ranking under cost constraints using RouteRec, a framework that compares request-level hard selection with item-level l…
      • On MovieLens-1M, the full quality oracle has substantial headroom (HR@10 = 0.584), confirming that useful cross-agent signal exists
  • Global Merger-Arbitrage Forecasting with Language Models

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09921v1 公告类型:新。
      • 摘要:我们提出了一种用于合并套利的语言模型预测系统,这是一种专门的高风险金融环境,其任务是预测已宣布的并购交易的结果。
      • 与之前的法学硕士判断预测工作不同,之前的工作侧重于广泛的混合主题基准和新闻片段等简短上下文,我们研究的是需要对数百页技术文档进行长上下文推理的设置。
      • 我们的系统将专家引导的上下文工程与对源自历史交易的事后引导的推理轨迹进行微调相结合。
    • EN 要点:
      • arXiv:2607.09921v1 Announce Type: new
      • Abstract: We present a language-model forecasting system for merger arbitrage, a specialized high-stakes financial setting in which the task is to predict the o…
      • Unlike prior work on judgmental forecasting with LLMs, which has focused on broad mixed-topic benchmarks and short context such as news snippets, we study a set…
      • Our system combines expert-guided context engineering with finetuning on hindsight-guided reasoning traces derived from historical deals
  • Faithful by Design: Evaluating and Improving LLM-Generated Clinical Trial Summaries for Multi-Stakeholder Audiences

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09932v1 公告类型:新。 -摘要:大型语言模型越来越多地用于为医疗保健提供者、患者和付款人总结临床试验结果,但它们的幻觉倾向在这种高风险背景下构成了重大风险。
      • 这项研究引入了一个基准评估框架,用于衡量法学硕士生成的临床试验摘要在三个利益相关者受众中的可信度。
      • 该框架由来自 ClinicalTrials.gov 数据库聚合分析的 200 项分层试验组成,使用特定于受众的提示模板和六维忠实度注释模式进行评估。
    • EN 要点:
      • arXiv:2607.09932v1 Announce Type: new
      • Abstract: Large language models are increasingly used to summarize clinical trial results for healthcare providers, patients, and payers, but their tendency to…
      • This study introduces a benchmark evaluation framework for measuring the faithfulness of LLM-generated clinical trial summaries across three stakeholder audienc…
      • The framework consists of 200 stratified trials drawn from the Aggregate Analysis of ClinicalTrials.gov database, evaluated using audience-specific prompt templ…
  • Workload-Driven Optimization for On-Device Real-Time Subtitle Translation

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09957v1 公告类型:新。
      • 摘要:本报告研究了台湾在短输入、短输出、批量大小推理、低延迟和隐私约束下的设备上英译繁体字幕翻译。
      • 这些条件限制了为长上下文或高吞吐量语言模型服务设计的优化的价值。
      • 从 LMT-60-0.6B 开始,初步分析表明,在 GGUF 量化降低 Transformer 块的相对成本之后,词汇投影成为更重要的解码时间成本。
    • EN 要点:
      • arXiv:2607.09957v1 Announce Type: new
      • Abstract: This report studies on-device English-to-Traditional-Chinese subtitle translation for Taiwan under short inputs, short outputs, batch-size-one inferen…
      • These conditions limit the value of optimizations designed for long-context or high-throughput language-model serving
      • Starting from LMT-60-0.6B, preliminary profiling suggests that vocabulary projection becomes a more important decode-time cost after GGUF quantization reduces t…
  • Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09999v1 公告类型:新。
      • 摘要:我们表明,即使任务准确性得以保留,训练后量化也可以默默地改变大型语言模型的推理量。
      • 使用由两名独立人工注释者验证的六类故障分类法(Cohen 的 $\kappa$ = 0.906),我们对来自五个指令调整的 LLM(3B–14B 参数)的 30,000 个思想链输出进行了分类,涉及三个量化精度(FP32、FP16、NF4)和四个推理基准。
      • 我们发现,虽然精度在各个精度上都很稳健(最大下降 3.1 pp),但空心收敛(通过不完整或无法验证的推理得出的正确答案)在 NF4 下显示出显着的大小相关变化,测试的两个最小模型急剧下降,但 12B 参数及以上的模型保持不变。
    • EN 要点:
      • arXiv:2607.09999v1 Announce Type: new
      • Abstract: We show that post-training quantization can silently alter how large language models reason even when task accuracy is preserved
      • Using a six-category failure taxonomy validated by two independent human annotators (Cohen’s $\kappa$ = 0.906), we classify 30,000 chain-of-thought outputs from…
      • We find that while accuracy is robust across precisions (maximum 3.1 pp drop), Hollow Convergence (correct answers reached through incomplete or unverifiable re…
  • Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.10020v1 公告类型:新。
      • 摘要:我们推出 FindMyText,这是一个开源 Python 包,旨在有效评估给定文本是否部分或全部出现在文本语料库中。
      • 该工具建立在现有的文档指纹识别技术的基础上,但通过一种新颖的机制对其进行了扩展,以显式捕获匹配指纹的序列。
      • 通过识别此类链,该工具可以更可靠地检测给定文本的近乎逐字副本,而不仅仅是文本相似性。
    • EN 要点:
      • arXiv:2607.10020v1 Announce Type: new
      • Abstract: We present FindMyText, an open-source Python package designed to efficiently assess whether a given text appears, in part or in full, within a text co…
      • The tool builds on prior techniques for document fingerprinting, but extends them with a novel mechanism to explicitly capture sequences of matching fingerprint…
      • By identifying such chains, the tool can more reliably detect near-verbatim copies of a given text rather than mere textual similarities
  • Efficiently Adapting Spoken Language Models for the Singaporean Context

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.10092v1 公告类型:新。
      • 摘要:口语模型 (SLM) 统一了语音感知和推理,但将其适应敏感领域的研究尚未充分,特别是当原始训练数据无法访问且用例需要多语言、口语查询交互时。
      • 我们将开源 SLM 应用于新加坡主队环境,涵盖新加坡四种官方语言的五个语音任务,结合 LoRA 微调、防止灾难性遗忘的替代文本 QA 数据集,以及使 CoBa 重新加权方案适应语音的多任务目标。
      • 我们还构建了 HTD-multilingual-QA,这是一个包含 504,853 个文本和语音形式的多语言 QA 样本数据集。
    • EN 要点:
      • arXiv:2607.10092v1 Announce Type: new
      • Abstract: Spoken language models (SLMs) unify speech perception and reasoning, but adapting them to sensitive domains is underexplored, especially when the orig…
      • We adapt an open-source SLM to the Singaporean Home Team context across five speech tasks in Singapore’s four official languages, combining LoRA fine-tuning, a…
      • We also build HTD-multilingual-QA, a 504,853 sample multilingual QA dataset in text and spoken form
  • Cost of Reasoning in non-English Languages: A Case Study on Japanese

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.10114v1 公告类型:新。
      • 摘要:推理语言模型(RLM)在用英语进行推理时能达到最强的性能,而英语是面向推理的训练数据最丰富的语言。
      • 然而,推理轨迹是模型可解释性和安全性的线索,并且在实践中对于模型用户和模型开发人员都很有用。
      • 因此,希望能够开发一种以用户选择的语言进行推理的模型,同时仍然保持强大的推理性能。
    • EN 要点:
      • arXiv:2607.10114v1 Announce Type: new
      • Abstract: Reasoning Language Models (RLMs) achieve their strongest performance when they reason in English, the language for which reasoning-oriented training d…
      • However, reasoning trace is a clue for model interpretability and safety, and useful in practice for both the model users and for model developers
      • Thus, it is desirable to be able to develop a model that reasons in a language of the user’s choice, while still maintaining strong reasoning performance

ArXiv cs.LG (B_intro+search) 链接到标题

  • Knowledge Graphs Meet Graph Neural Networks: A Comprehensive Survey

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09666v1 公告类型:新。
      • 摘要:图神经网络(GNN)由于其对图结构数据进行建模的内在能力,已成为知识图(KG)中的强大范例。
      • 然而,仍然缺乏对整个知识图谱技术管道中基于 GNN 的方法的系统回顾。
      • 为了解决这一差距,我们首先提出了一种基于 GNN 的知识图技术的新型两级分类框架:知识图谱技术管道和基于 GNN 的视角。
    • EN 要点:
      • arXiv:2607.09666v1 Announce Type: new
      • Abstract: Graph Neural Networks (GNNs) have emerged as a powerful paradigm in Knowledge Graphs (KGs) due to their intrinsic ability to model graph-structured da…
      • However, there remains a lack of a systematic review about GNN-based methodologies across the entire knowledge graph technologies pipeline
      • To address this gap, we first propose a novel two-level taxonomy framework for GNN-based knowledge graph technologies: the KG technologies pipeline and GNN-base…
  • Position: Every Ground Truth is a Human Construction, not an Objective Truth

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09668v1 公告类型:新。
      • 摘要:地面实况数据集在机器学习模型的训练和评估中作为参考值发挥着基础作用。
      • 本立场文件认为,基本事实并不是自然给出的中立客观测量,而是由人类和技术的安排构建的。
      • 我们认为,阐明和讨论这些通常不可见或未报告的选择,并承认参考数据集是偶然的,而不是通用的,机器学习社区将从中受益。
    • EN 要点:
      • arXiv:2607.09668v1 Announce Type: new
      • Abstract: Ground truth datasets play a fundamental role as reference values in the training and evaluation of machine learning models
      • This position paper argues that ground truths are not neutral objective measurements that are naturally given, but instead that they are constructed by arrangem…
      • We argue that the ML community will benefit from articulating and discussing these often invisible or unreported choices and acknowledging that reference data s…
  • AuditWeave: A Tamper-Evident, Auditor-Navigable Evidence Layer for AI-Assisted and Data-Transformation Workflows

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09682v1 公告类型:新。
      • 摘要:人工智能系统越来越多地用于协助审计、金融和医疗保健等监管领域的后续决策。
      • 这就产生了一项经常性的义务:组织必须能够在事后重建哪些证据表明了给定的结论,并表明该推理的记录没有被改变。
      • 现有工具解决相关但不同的问题 - 模型可观察性、漂移监控、治理报告 - 并且是为操作系统的机器学习工程师构建的,而不是为必须将一个特定结论追溯到其支持证据的审阅者构建的。
    • EN 要点:
      • arXiv:2607.09682v1 Announce Type: new
      • Abstract: AI systems are increasingly used to assist consequential decisions in regulated domains such as auditing, finance, and healthcare
      • This creates a recurring obligation: an organization must be able to reconstruct, after the fact, which evidence informed a given conclusion, and to show that t…
      • Existing tools address related but distinct problems - model observability, drift monitoring, governance reporting - and are built for the machine-learning engi…
  • Ablation, Statistical Inference, and Validation for KV-Cache Compression

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09683v1 公告类型:新。
      • 摘要:本研究系统地比较了 Turbo-Quant 和 SpectralQuant KV 缓存压缩,评估非支配方案,包括使用 Beta Lloyd-Max 和 QJL 进行 WHT 旋转,通过统计验证方法将系统编解码器差异与实现差异分开。
      • 主要发现表明,虽然基于特征基的方法由于协方差不稳定而无法处理重尾数据,但它们在结构化体系中表现出色,有效语义维度 ($d_{eff}$) 适应校准预算而不是真实数据排名。 -(这是摘要中的摘要,谢谢)。
    • EN 要点:
      • arXiv:2607.09683v1 Announce Type: new
      • Abstract: This study systematically compares Turbo-Quant and SpectralQuant KV-cache compression, evaluating non-dominated schemes, including WHT rotation with B…
      • Key findings reveal that while eigenbasis-based methods fail on heavy-tailed data due to covariance instability, they excel in structured regimes, with the effe…
      • (this is an abstract of the abstract thank you )
  • SciML in the Wild: A Diagnostic Study of When Structural Priors Help and When They Hurt

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09684v1 公告类型:新。
      • 摘要:当结构先验反映可靠的动态控制时,神经常微分方程 (NODE)、物理信息神经网络 (PINN) 和通用微分方程 (UDE) 等科学机器学习 (SciML) 方法最为有效。
      • 我们问当这个假设被违反时会发生什么。
      • 使用宏观经济预测作为压力测试领域,我们使用稀疏年度数据、多个时间分割和五个随机种子来评估 23 个国家/地区的五个模型系列:ARIMA、LSTM、NODE、PINN 和 UDE。
    • EN 要点:
      • arXiv:2607.09684v1 Announce Type: new
      • Abstract: Scientific Machine Learning (SciML) methods such as Neural Ordinary Differential Equations (NODEs), Physics-Informed Neural Networks (PINNs), and Univ…
      • We ask what happens when this assumption is violated
      • Using macroeconomic forecasting as a stress-test domain, we evaluate five model families, ARIMA, LSTM, NODE, PINN, and UDE, across 23 countries using sparse ann…
  • MawForge: Memory-Bounded Expert Materialization for Local Mixture-of-Experts Inference

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09686v1 公告类型:新。
      • 摘要:稀疏专家混合 (MoE) 语言模型将总参数计数与每个令牌的主动计算分开,但本地推理系统通常仍然需要完整模型、键值缓存、运行时缓冲区和操作系统空间以适应快速内存。
      • MawForge 测试了不同的系统假设:通过将完整模型存储在磁盘上、保持通用张量驻留以及按需将路由专家张量具体化到有界执行缓存中,本地 MoE 服务可以在受限统一内存机器上变得实用。
      • 主要发现是 MawForge 作为本地 MoE 推理的有界执行机制和测量基础是有效的,但不是缓存最大化策略。
    • EN 要点:
      • arXiv:2607.09686v1 Announce Type: new
      • Abstract: Sparse Mixture-of-Experts (MoE) language models separate total parameter count from per-token active computation, but local inference systems often st…
      • MawForge tests a different systems hypothesis: local MoE serving can be made practical on constrained unified-memory machines by storing the full model on disk,…
      • The central finding is that MawForge is effective as a bounded execution mechanism and measurement substrate for local MoE inference, but not as a cache-maximiz…
  • Prioritizing Search Space Regions in the Low Autocorrelation Binary Sequences Problem

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09688v1 公告类型:新。
      • 摘要:低自相关二进制序列问题 (LABS) 是一项硬组合优化挑战,在通信、信号处理和卫星导航中具有重要应用。
      • 本文提出了一种混合搜索框架,它将汤普森采样与并行自回避行走相结合,以跨 LABS 搜索空间的限制类别自适应地分配计算量。
      • 通过将分区建模为多臂老虎机设置中的臂,所提出的方法动态地将搜索资源转移到凭经验产生更高优点因子的分区,同时保持对采样较少区域的探索。
    • EN 要点:
      • arXiv:2607.09688v1 Announce Type: new
      • Abstract: Low autocorrelation binary sequences problem (LABS) is a hard combinatorial optimization challenge with important applications in communications, sign…
      • This paper proposes a hybrid search framework that combines Thompson sampling with parallel self-avoiding walks to adaptively allocate computational effort acro…
      • By modeling partitions as arms in a multi-armed bandit setting, the proposed method dynamically shifts search resources toward partitions that empirically produ…
  • What Context Does a Coding Agent Actually Need to Act?

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09691v1 公告类型:新。
      • 摘要:现代编码代理可以在其上下文窗口中保存整个存储库。
      • 它的大部分阅读都被浪费了 - 有趣的问题不是代理可以使用多少上下文,而是它实际上\emph{需要}。
      • 我们现在研究这个最重要的问题:代理何时必须\emph{编辑}代码。
    • EN 要点:
      • arXiv:2607.09691v1 Announce Type: new
      • Abstract: A modern coding agent can hold an entire repository in its context window
      • Most of its reading is wasted – and the interesting question is not how much context an agent can use, but what it actually \emph{needs}
      • We study that question at the moment it matters most: when the agent must \emph{edit} code
  • Reference-Based Distillation Detection in LLMs

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09692v1 公告类型:新。
      • 摘要:模型蒸馏——对更强大的第三方模型的输出进行训练——被广泛用于提高性能,但引起了对不公平优势和政策违规的担忧。
      • 这就引发了一个基本问题:我们能否检测一个模型是否是从另一个模型中提取出来的?
      • 我们表明,虽然从孤立的学生中识别教师模型非常具有挑战性,但在基于参考的环境中它变得容易处理:给定一个模型和来自同一谱系的早期检查点,我们可以识别用于训练后期检查点的教师模型。
    • EN 要点:
      • arXiv:2607.09692v1 Announce Type: new
      • Abstract: Model distillation – training on outputs from stronger third-party models – is widely used to boost performance, but raises concerns about unfair ad…
      • This motivates a fundamental question: can we detect whether a model was distilled from another
      • We show that, while identifying a teacher model from a student in isolation is highly challenging, it becomes tractable in a reference-based setting: given a mo…
  • Depth-Entropy Guided Sampling for Training-Free LLM Reasoning

    • 发布时间:2026-07-14 12:00 北京时间
    • 摘要:- arXiv:2607.09693v1 公告类型:新。
      • 摘要:强化学习(RL)已成为提高大型语言模型推理能力的主导范式,但它需要昂贵的训练、精选数据和奖励信号。
      • 最近的工作表明,在测试时从锐化的基本模型分布中进行采样可以恢复大部分 RL 增益,但现有方法仅依赖于输出层可能性,而忽略了变压器的内部前向传递动态。
      • 我们引入深度熵引导采样(DEGS),这是一种免训练的测试时间方法,利用逐层熵崩溃作为内在质量信号。
    • EN 要点:
      • arXiv:2607.09693v1 Announce Type: new
      • Abstract: Reinforcement learning (RL) has become the dominant paradigm for improving the reasoning capabilities of large language models, but it requires expens…
      • Recent work shows that sampling from sharpened base-model distributions at test time recovers much of the RL gain, yet existing methods rely solely on output-la…
      • We introduce Depth-Entropy Guided Sampling (DEGS), a training-free, test-time method that exploits layer-wise entropy collapse as an intrinsic quality signal