🤖 AI 速览

今天的主线是 AI Agent 从演示走向真实工作流:手机代理评测、端到端开发插件与托管智能体都在强化可执行能力。与此同时,行业开始更重视可靠推理、测试与契约,关注模型答案如何被验证。端侧模型和垂直小模型也继续升温,显示效率、成本与本地部署正成为新竞争点。
📋 文章元数据
发布时间
2026-06-17
类型
ai-daily
字数
3394
阅读时长
16 min

2026-06-17 AI日更 | AI Agent 开始进入真实工作流:从手机操作到物理实验 链接到标题

今天的主线是 AI Agent 从演示走向真实工作流:手机代理评测、端到端开发插件与托管智能体都在强化可执行能力。与此同时,行业开始更重视可靠推理、测试与契约,关注模型答案如何被验证。端侧模型和垂直小模型也继续升温,显示效率、成本与本地部署正成为新竞争点。

📖 本期 Watch List 深度导读 链接到标题

今天最值得关注的是“AI Agent 从演示走向真实工作流”。PhoneHarness 重新定义手机代理评测,强调 GUI、CLI 与工具混合使用;Web Agents 研究则提醒我们,记忆与技能模块并非免费午餐,必须放进 token 预算里重新衡量。

第二条主线是“可靠推理与可解释性”。CoRA 关注置信度是否真的被推理过程支撑,Lean 4 自动形式化与 LLM 解释定义的论文,也都在追问同一个问题:模型答案如何被验证、被信任。

此外,DeepMind 与英国政府的 AI 规划原型值得产品和政策团队阅读,它展示了 AI 正在进入住房审批这类高摩擦公共基础设施场景。Context Compression、多语言 tokenizer 与 Nemotron 3 Ultra,则可作为模型效率方向的技术补充。

🌐 X 平台 AI 热点快讯 链接到标题

话题 1:SpaceX Acquires Cursor Maker Anysphere in $60 Billion Stock Deal 链接到标题

  • 分类:AI · News
  • 概况:热度时间:13 hours ago,相关帖子数:78000
  • 是什么事:据 X 平台热传消息,SpaceX 以 600 亿美元股票交易收购 AI 编程工具 Cursor 的开发商 Anysphere。
  • 为什么重要:如果属实,这将显示航天与硬科技公司正加速整合 AI 编程能力,凸显代码生成工具在工程研发、自动化开发和企业生产力中的战略价值。
  • 讨论概况:X 上讨论集中在交易真实性、600 亿美元估值是否过高、SpaceX 收购 AI 开发工具的战略意图,以及 Cursor 是否会被深度用于航天软件和内部工程流程。

话题 2:Matt Shumer Seeks Mobile Control for Claude Code AI Agent 链接到标题

  • 分类:AI · Other
  • 概况:热度时间:,相关帖子数:33
  • 摘要:Matt Shumer Seeks Mobile Control for Claude Code AI Agent:

话题 3:OpenAI Fixes Codex Outage and Promises Rate Limit Resets 链接到标题

  • 分类:AI · News
  • 概况:热度时间:11 hours ago,相关帖子数:2400
  • 是什么事:OpenAI 修复了 Codex 服务中断,并表示将为受影响用户重置相关速率限制。
  • 为什么重要:Codex 是开发者使用 AI 编程工具的重要入口,此次故障凸显了 AI 基础设施稳定性、可用性和配额管理对生产力工具落地的关键影响。
  • 讨论概况:X 上的讨论主要集中在服务中断对开发工作流的影响、OpenAI 补偿措施是否足够,以及高频使用者对速率限制透明度和可靠性的担忧。

话题 4:OpenAI Launches Developers Plugin for Codex Coding App 链接到标题

  • 分类:AI · News
  • 概况:热度时间:23 hours ago,相关帖子数:294
  • 是什么事:OpenAI 为 Codex 编程应用推出开发者插件,支持在 Codex 内构建、预览和测试 iOS 应用,并结合 SwiftUI Preview 与热重载优化开发流程。
  • 为什么重要:这显示 AI 编程工具正从代码生成走向端到端开发环境,能够在同一工作流中完成编写、运行、调试和反馈,提升智能体参与真实软件开发的能力。
  • 讨论概况:X 上的讨论集中在 Codex 与 Claude Code、Cursor 等工具的竞争,以及插件生态是否会成为 AI 编程产品的关键壁垒;支持者认为它让开发流程更紧凑,质疑者则关注复杂项目中的稳定性、可控性和实际效率。

话题 5:AI Builders Embrace Agentic Loops for Self-Reliant Tasks 链接到标题

  • 分类:AI · News
  • 概况:热度时间:6 hours ago,相关帖子数:105
  • 是什么事:AI 开发者正在更多采用“代理式循环”来构建可自主规划、执行、检查并迭代任务的 AI 系统。
  • 为什么重要:这标志着 AI 应用从单次问答走向更具自主性的工作流,有望提升复杂任务处理、软件开发、数据分析和自动化运营的效率。
  • 讨论概况:X 上的讨论集中在代理式循环是否真的可靠:支持者认为它能减少人工干预、推动 AI 助手实用化;质疑者则担心错误累积、成本失控、安全边界和评估标准仍不成熟。

话题 6:Anthropic’s Claude Managed Agents Speed Up AI Production Deployment 链接到标题

  • 分类:AI · News
  • 概况:热度时间:2 hours ago,相关帖子数:217
  • 是什么事:Anthropic 围绕 Claude 推出面向企业和开发者的托管智能体与工具链,加速 AI 应用从原型走向生产部署。
  • 为什么重要:这表明大模型竞争正从单纯模型能力转向企业级集成、自动化部署和开发者生态,可能影响 AI 在电商、工程和办公场景中的落地速度。
  • 讨论概况:X 上讨论焦点集中在 Anthropic 是否正在追赶甚至超越 OpenAI、Claude 托管智能体对企业 AI 部署的实际价值,以及 Shopify 等生态合作是否会带来新的开发工作流;分歧则在于这些工具是生产力跃迁,还是仍被市场宣传和“vibe coding”热潮放大。

今日 X 上的 AI 舆情小结 链接到标题

今天的舆论主线集中在 AI 编程工具从“辅助写代码”快速升级为“端到端开发与代理式执行平台”:无论是 Cursor 传闻被 SpaceX 高价收购、Codex 推出 iOS 开发插件,还是 Anthropic 推托管智能体,都指向开发者生态和工程工作流正在成为大模型竞争的新战场。共识是,代码生成、预览、测试、部署和自主迭代能力正在被视为企业生产力与硬科技研发的战略基础设施,工具链整合能力的重要性不亚于模型本身。分歧主要在估值与实际效用上:支持者认为这些产品会显著压缩开发周期、推动 AI 助手进入生产环境,质疑者则认为交易传闻、插件生态和“代理式循环”可能被资本叙事与 vibe coding 热潮过度放大。潜在风险则集中在三方面:基础设施稳定性和速率限制会直接影响开发工作流,智能体自主执行可能带来错误累积、成本失控和安全边界问题,而平台生态一旦高度绑定,也可能削弱开发者对工具链的可控性与迁移自由。

💡 大佬观点(Influencer Insights) 链接到标题

AI 行业动态日报分析 (2026年6月中旬) 链接到标题

1. 今日热点:本地端侧模型与编程 Agent 的物理世界扩张 链接到标题

今日大佬们讨论的焦点发生了显著的“重力转移”:从云端超级模型转向本地端侧部署,以及从纯数字编程转向物理世界操控。

  • 端侧模型进入“甜点”时刻:

    • @zhixianio 进行了高强度测试,认为 Qwen3.6-35B-A3B 在 Mac 本地的运行效果已稳坐“甜点宝座”,速度和智商超越远程 LLM,原生多模态体验甚至优于云端大模型。
    • 针对 Google 新发布的 Gemma 4 12B Coder,@zhixianio 指出其虽有优化,但在“长篇、有状态、一次成型”的复杂生成任务(如完整游戏编写)中,12B 体量仍是瓶颈,与 35B MoE 模型差距明显。他还测试了 Gemma 4 的音频能力,指出中文识别不佳但日英 OK,且量化感知训练(QAT)是提升端侧效率的新思路。
    • @AI_Jasonyu 从另一个角度印证了“端侧智能”:百度 PP-OCRv6 以极小参数(1.5MB)在浏览器端实现了超越 GPT-5.5 等大模型的 OCR 准确率,证明了“小模型吃透垂直场景”的巨大优势。
  • AI Agent 进入物理世界(AutoResearch):

    • @dotey 重点解读了 NVIDIA GEAR 实验室的 ENPIRE 项目。这是首次将 AI 编程 Agent 的全自主科研循环(设计、实验、失败分析、代码迭代)落地到真实物理环境。Agent 可自主操控机器人完成高精度任务,并发现了“物理规模化法则”:多机器人并行能加速研究。这标志着 AI 能力从数字代码生成向物理世界生产力的跨越。

2. 独特观点与行业前瞻 链接到标题

  • AI 编程的真正护城河:“契约”与“测试”,而非代码逻辑

    • @Pluvio9yte 提出,Vibe Coding 的奥秘在于 “Contract First” (契约优先)。他通过实战总结出,只有外化定义好 API 与数据模型的契约,人机协作才能避免上下文漂移,这是比单纯需求或代码更关键的框架。
    • @ruanyf 则引用 Cloudflare 工程师复刻 Next.js 的案例,尖锐指出:代码本身已无护城河,“测试是新的护城河”。因为 AI 能轻易复刻大型项目,但能否通过高质量测试并保证稳定运行才是核心壁垒。
  • 对 Claude Fable 5 的理性审视与“反共识”

    • @Pluvio9yte 提供了与炒作不同的“反共识”深度体验:Fable 5 速度极慢,Token 消耗没有想象中离谱(约 Opus 的 1.5 倍),能力边界更广但未到“惊艳”程度,更像 Opus 4.6++ 与 GPT-5.5++ 的结合体,提倡理性消费。
    • @dotey 则指出了 Claude Code 新 Dynamic Workflows 的 Token 消耗黑洞,一个简单任务就消耗 130 万 Token。
  • AI 时代的生产关系与经济学

    • @ruanyf 提出了尖锐的社会学问题:AI 大幅提升效率后,员工能否放假?如果既不加薪也不放假,AI 对员工的意义何在?同时他通过计算指出,若无限制使用顶级模型进行 AI 编程,其算力成本已远超人类程序员薪资,暗示未来企业需权衡 AI 投入产出比。
    • @vista8 分享了 Factory AI CEO 的前瞻观点:未来最值钱的是能端到端搞定业务结果的工程师,而非单纯写代码的人;且三年内员工的 Token 中位数支出将与薪资持平。
  • AI 产品设计与流量之道

    • @Pluvio9yte 提出消除 UI “AI 味”的方法:使用知名品牌的 DESIGN.md 设计文件作为 AI 的生成约束,提升质感。
    • @gefei55 分享了 SEO 实战经验,强调与谷歌算法共成长,短期 AIGC 低质量内容终将遭反噬。并分享域名投资的暴利故事及 SaaS 海外高定价策略的可行性。
    • @AI_Jasonyu 发现 AI 视频赛道的付费墙已从卷功能转向卷“积分的解释方式”,OpusClip 这种“垂直场景 + 已有长剪短”的逻辑最适合独立开发者。

3. 推荐的工具与资源 链接到标题

工具/资源推荐人核心亮点与用途
baoyu-design Skill@dotey本地设计转代码。支持 Figma 文件导入,本地生成设计系统与 PPT,甚至可导出为可编辑的 PPTX 文件。
info-digest Skill@doteyAI 资讯整理助手。宝玉公开的日常写作 Skill,包含读者视角、事实核查、精炼格式等可借鉴策略。
getdesign.md@Pluvio9yteUI 设计去 AI 味神器。收集了 Linear, Vercel, Apple 等真实品牌的 DESIGN.md 设计系统文件,可喂给 AI 生成高质量 UI。
Papr@vista8轻量开源 RSS 客户端。支持接入自有 API Key 进行 AI 总结与问答。
Figma Chrome 插件@vista8网页仿站降维打击。一键将任意网页元素变为可编辑图层并导入 Figma。
PP-OCRv6@AI_Jasonyu极致端侧 OCR。1.5MB 模型可跑在浏览器,速度与准确率反超 GPT-5.5 等大模型,完全开源。
App Store 评价分析工具@vista8开源用户反馈挖掘器。可抓取任意 App 评论并用 LLM 进行痛点、机会点分析。
GPT Image 提示词@dotey (from @Ciri_ai)照片转涂鸦插画。将照片转化为“装饰性民间平面涂鸦风”的特定提示词。
Maccy / Mos@Pluvio9yteMac 效率两件套。Maccy 为开源剪贴板工具,Mos 用于解决外接鼠标滚动方向反人类问题。

📚 附录:今日 Watch List 更新源列表 链接到标题

时间窗口:最近 3 天;覆盖 22 个源;共 34 条更新

Stratechery by Ben Thompson (A_full) 链接到标题

  • Fox Buys Roku, The Problem With Fox’s Smart Strategy, Streaming That Works
    • 发布时间:2026-06-16 18:00 北京时间
    • 摘要:- 市场讨厌福克斯收购 Roku,但该公司正在用从版权所有者那里获取的收益来换取作为承租人的杠杆。
      • 15 美元/月150 美元/年。
      • 通过每周三封电子邮件或播客对当天新闻进行实质性分析。
      • 策略采访
      • 采访领先的上市首席执行官、私营公司创始人,并与分析师同行进行讨论。
    • EN 要点:
      • The market hates Fox’s acquisition of Roku, but the company is trading extraction from rights holders for leverage as a renter.

OpenAI Blog (A_full) 链接到标题

  • Predicting model behavior before release by simulating deployment
    • 发布时间:2026-06-16 08:00 北京时间
    • 摘要:- 在发布新模型之前,实验室不仅需要了解它可以做什么,还需要了解它在实际使用中的表现,包括它可能在哪里引入新风险。
      • 随着能力的增强,这一点变得更加重要。
      • 作为部署前安全审查的一部分,我们利用有针对性的评估、红队和其他检查来了解模型行为。
      • 我们现在开始使用一种方法在模型部署发生之前对其进行模拟,这增加了一个补充信号:在候选模型到达用户之前对候选模型的行为进行类似部署的预览。
      • 部署模拟是一种在未来部署发生之前对其进行模拟的方法。
    • EN 要点:
      • OpenAI introduces Deployment Simulation, a method to predict AI model behavior before deployment using real conversation data to improve safety and evaluation a…

Google DeepMind Blog (A_full) 链接到标题

  • Unlocking UK house-building with AI-accelerated planning
    • 发布时间:2026-06-17 05:29 北京时间
    • 摘要:- 英国政府与 Google DeepMind 合作构建了一个新的人工智能原型,旨在更快地做出住房决策。
      • 这篇来自 Google DeepMind 博客的文章解释了如何通过 AI 加速规划解锁英国房屋建筑,塑造更广泛的 AI 和基础设施景观。
      • 在通过人工智能加速规划解锁英国房屋建筑之后,它还为创始人、运营商和投资者带来了实际影响。
    • EN 要点:
      • UK government partners with Google DeepMind to build a new AI-powered prototype aimed at faster housing decisions.

Two Minute Papers (B_intro+search) 链接到标题

  • They Looked Inside Claude’s AI’s Mind. It Got Weird
    • 发布时间:2026-06-16 23:53 北京时间
    • 摘要:- ❤️ 在这里查看 Lambda 并注册他们的 GPU Cloud:。
      • 📝 该论文可在此处获取:.
      • Adam Bridges、Benji Rabhan、B Shang、Cameron Navor、Charles Ian Norman Venn、Christian Ahlin、Eric T、Fred R、Gordon Child、Juan Benet、Michael Tedder、Owen Skarpness、Richard Sundvall、Ryan Stankye、Shawn Becker、Steef、Taras Bobrovytsky、Tazaur Sagenclaw、Tybie Fitzhugh、Ueli Gallizzi。
      • 他们深入了解了克劳德人工智能的思维。
    • EN 要点:
      • ❤️ Check out Lambda here and sign up for their GPU Cloud:
      • 📝 The paper is available here:
      • 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
      • Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Ska…

ArXiv cs.AI (B_intro+search) 链接到标题

  • A Definition of Good Explanations and the Challenges Explaining LLM Outputs

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14838v1 公告类型:新。
      • 摘要:如何定义一个好的解释是一个长期存在的哲学争论,最近人们对人工智能输出的背景重新产生了兴趣。
      • 可解释性对于在许多情况下采用人工智能至关重要,但为了对人工智能系统产生良好的解释,我们必须首先了解什么是好的解释。
      • 在本文中,我们提出了一个受反事实解释概念启发的定义,但我们认为,还必须考虑对话者对解释中可能提供的每个事实的先前信念。
    • EN 要点:
      • arXiv:2606.14838v1 Announce Type: new
      • Abstract: How to define a good explanation is a long-standing philosophical debate which has found recent renewed interest in the context of AI outputs
      • Explainability is crucial for AI adoption in many contexts, but in order to produce good explanations of AI systems, we must first have an understanding of what…
      • In this paper we propose a definition inspired by the notion of counterfactual explanations, however we argue that one must also take into account the interlocu…
  • Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14885v1 公告类型:新。
      • 摘要:大型语料库上的代理搜索依赖于检索器介导的接口(例如 BM25 或 ColBERT)来实现可扩展的候选发现。
      • 虽然可以有效地对相关文档进行排名,但这些界面仅以排名结果或有界文档视图的形式公开证据,从而限制了代理重新组织材料和验证跨文档约束的能力。
      • 直接语料库交互 (DCI) 通过公开 shell 可执行语料库操作以进行灵活的搜索、过滤、比较和验证来解决此限制。
    • EN 要点:
      • arXiv:2606.14885v1 Announce Type: new
      • Abstract: Agentic search over large corpora relies on retriever-mediated interfaces (e.g., BM25 or ColBERT) for scalable candidate discovery
      • While effective at ranking relevant documents, these interfaces expose evidence only as ranked results or bounded document views, limiting agents’ ability to re…
      • Direct Corpus Interaction (DCI) addresses this limitation by exposing shell-executable corpus operations for flexible search, filtering, comparison, and verific…
  • Relational Structural Causal Models

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14892v1 公告类型:新。
      • 摘要:人工智能必须拥有一个因果环境模型,支持对干预和反事实的推理,同时也必须具有组合环境模型,支持对未见过的对象组合的概括。
      • 在这项工作中,我们正式研究何时以及如何学习这样的模型。
      • 我们开发关系结构因果模型,将结构因果模型(Pearl 2009)扩展到对象及其关系变化的环境。
    • EN 要点:
      • arXiv:2606.14892v1 Announce Type: new
      • Abstract: An artificial intelligence must have a model of its environment that is causal, supporting reasoning about interventions and counterfactuals, and also…
      • In this work, we formally study when and how such a model can be learned
      • We develop relational structural causal models, extending structural causal models (Pearl 2009) to settings where objects and their relations vary
  • Trust Between AI Agents: Measuring Formation, Breakage, and Recovery, with Implications for Governing Multi-Agent Systems

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14923v1 公告类型:新。
      • 摘要:随着语言模型智能体越来越多地在团队中工作,每个智能体必须决定对队友的信任程度。
      • 然而我们缺乏衡量人工智能代理之间信任的标准方法。
      • 我们提出了一种基于昂贵验证的行为措施。
    • EN 要点:
      • arXiv:2606.14923v1 Announce Type: new
      • Abstract: As language-model agents increasingly work in teams, each agent must decide how much to trust its teammates
      • Yet we lack a standard way to measure trust between AI agents
      • We propose a behavioral measure based on costly verification
  • PrologMCP: A Standardized Prolog Tool Interface for LLM Agents

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14935v1 公告类型:新。 -摘要:前沿推理调整的语言模型在深度演绎任务上仍然失败,并且通过扩展内部推理来提高性能的成本也很差。
      • 符号委托提供了一种补充途径:语言模型翻译问题,而求解器执行推理。
      • 然而,当前用于逻辑编程的自动形式化管道通常是与特定任务或代理相关的定制集成。
    • EN 要点:
      • arXiv:2606.14935v1 Announce Type: new
      • Abstract: Frontier reasoning-tuned language models still fail on deductive tasks at depth, and the cost of improved performance through extended internal reason…
      • Symbolic delegation offers a complementary route: a language model translates the problem, while a solver performs the inference
      • However, current autoformalization pipelines for logic programming are typically bespoke integrations tied to particular tasks or agents
  • Semantics-Enhanced Retrieval-Augmented Time Series Forecasting

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14941v1 公告类型:新。
      • 摘要:时间序列预测模型通常受益于历史模式。
      • 受检索增强生成 (RAG) 的启发,最近的研究探索检索相关的历史时间序列片段以增强预测。
      • 然而,仅依靠时间序列相似性往往不足以进行非平稳性下的检索。
    • EN 要点:
      • arXiv:2606.14941v1 Announce Type: new
      • Abstract: Time series forecasting models often benefit from historical patterns
      • Inspired by Retrieval-Augmented Generation (RAG), recent research explored retrieving relevant historical time series segments to enhance forecasting
      • However, relying solely on time series similarity is often insufficient for retrieval under non-stationarity
  • AI Engram: In Search of Memory Traces in Artificial Intelligence

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14997v1 公告类型:新。
      • 摘要:记忆形成是智力的基础,但深层神经网络是否保留类似于生物记忆单元的可识别记忆痕迹仍然是一个悬而未决的问题。
      • 这项工作引入了一个几何框架,通过将特异性、重新激活、充分性和必要性的神经科学标准形式化为受约束的逆问题来识别此类“人工智能印迹”。
      • 我们推导了一个封闭形式的估计器,它将个体记忆痕迹与全局纠缠参数隔离开来,并表明这种生物衍生的解决方案对应于参数流形上的自然梯度更新。
    • EN 要点:
      • arXiv:2606.14997v1 Announce Type: new
      • Abstract: Memory formation is fundamental to intelligence, yet whether deep neural networks preserve identifiable memory traces analogous to biological memory u…
      • This work introduces a geometric framework to identify such “AI engrams” by formalizing the neuroscientific criteria of specificity, reactivation, sufficiency,…
      • We derive a closed-form estimator that isolates individual memory traces from globally entangled parameters, and show that this biologically-derived solution co…
  • Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.15029v1 公告类型:新。
      • 摘要:LLM 法官用于减少评估开放式文本生成时对昂贵人力的需求。
      • 然而,这些评委的可靠性很大程度上取决于他们与人类评分者的一致性——这一特性本身就依赖于昂贵的人类注释。
      • 在这项工作中,我们开发了一种方法(指标匹配),用于根据有限的注释估计法学硕士法官基于相关性的可靠性指标。
    • EN 要点:
      • arXiv:2606.15029v1 Announce Type: new
      • Abstract: LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation
      • However, the reliability of these judges depends critically on their alignment with human raters – a property that itself depends on costly human annotations
      • In this work, we develop a method (Metric Match) for estimating correlation-based reliability metrics of LLM judges from limited annotations
  • OSGuard: A Benchmark for Safety in Computer-Use Agents

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.15034v1 公告类型:新。
      • 摘要:越来越多地通过计算机使用代理是否完成现实的桌面和 Web 任务来对其进行评估。
      • 然而,仅任务成功可能会错过代理通过不安全的捷径达到名义目标的失败。
      • 我们引入了 OSGuard,这是一个双粒度基准套件,用于在良性、未更改的用户指令下评估计算机使用代理的安全性。
    • EN 要点:
      • arXiv:2606.15034v1 Announce Type: new
      • Abstract: Computer-use agents are increasingly evaluated by whether they complete realistic desktop and web tasks
      • However, task success alone can miss failures in which an agent reaches the nominal goal through an unsafe shortcut
      • We introduce OSGuard, a dual-granularity benchmark suite for evaluating safety in computer-use agents under benign, unchanged user instructions
  • Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.15038v1 公告类型:新。
      • 摘要:由于模态不平衡和分布变化,根据多模态临床数据准确预测事件发生时间 (TTE) 仍然具有挑战性。
      • 我们引入了一个基础模型驱动框架,用于 CT 成像和纵向 EHR 数据之间的跨模式表示对齐,旨在跨任务和机构进行泛化。
      • CT 和 EHR 模式使用特定领域的基础模型独立编码,并通过四种原则融合策略在共享潜在空间中对齐:后期融合、对比对齐、交叉注意和共同注意。
    • EN 要点:
      • arXiv:2606.15038v1 Announce Type: new
      • Abstract: Accurate time-to-event (TTE) prediction from multimodal clinical data remains challenging due to modality imbalance and distribution shift
      • We introduce a foundation model-driven framework for cross-modal representation alignment between CT imaging and longitudinal EHR data, designed to generalize a…
      • CT and EHR modalities are encoded independently using domain-specific foundation models and aligned in a shared latent space through four principled fusion stra…

ArXiv cs.CL (B_intro+search) 链接到标题

  • PhoneHarness: Harnessing Phone-Use Agents through Mixed GUI, CLI, and Tool Actions

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14832v1 公告类型:新。
      • 摘要:人们越来越期望电话代理能够完成真实的移动工作流程,而不仅仅是预测下一个屏幕操作。
      • 然而,当前的许多移动代理文献仍然主要将代理评估为 GUI 控制器,这些控制器观察屏幕、发出点击和滑动操作,并根据目标应用程序状态进行评分。
      • 真正的手机使用任务更广泛:它们需要决定何时使用应用程序 GUI、设备端命令或结构化工具,同时留下证据表明预期的副作用确实发生了。
    • EN 要点:
      • arXiv:2606.14832v1 Announce Type: new
      • Abstract: Phone agents are increasingly expected to complete real mobile workflows rather than merely predict the next screen action
      • However, much of the current mobile-agent literature still evaluates agents primarily as GUI controllers that observe a screen, emit taps and swipes, and are sc…
      • Real phone-use tasks are broader: they require deciding when to use app GUIs, device-side commands, or structured tools, while leaving evidence that the intende…
  • Evaluating the Robustness of Proof Autoformalization in Lean 4

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14867v1 公告类型:新。
      • 摘要:证明自动形式化旨在将以自然语言编写的数学非正式证明转换为正式语言(例如 Lean~4)的形式证明。
      • 一些作品开发了基于 LLM 的证明自动形式化模型。
      • 然而,现有的评估通常侧重于从整理的数据集中翻译格式良好的非正式证明。
    • EN 要点:
      • arXiv:2606.14867v1 Announce Type: new
      • Abstract: Proof autoformalization aims to translate a mathematical informal proof written in natural language into a formal proof in a formal language such as L…
      • Several works have developed LLM-based models for proof autoformalization
      • However, existing evaluations have typically focused on translating well-formed informal proofs from curated datasets
  • Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14875v1 公告类型:新。
      • 摘要:我们研究使用小语言模型进行多跳问答的上下文压缩。
      • 我们提出 Telegraph English,一种可读的符号格式,可将检索到的段落重写为结构化实体关系语句,以较低的令牌成本保留推理证据。
      • 在 MuSiQue、TwoWiki 和 HotpotQA 的对照实验中,Telegraph English 在每个数据集上都优于三个匹配预算压缩基线(字符级删除、截断和随机子采样),增益提高了 13 到 20 个 F1 个百分点。
    • EN 要点:
      • arXiv:2606.14875v1 Announce Type: new
      • Abstract: We study context compression for multi-hop question answering with small language models
      • We propose Telegraph English, a readable symbolic format that rewrites retrieved passages into structured entity-relation statements, preserving reasoning evide…
      • In controlled experiments on MuSiQue, TwoWiki, and HotpotQA, Telegraph English outperforms three matched-budget compression baselines (character-level deletion,…
  • Simplifying the Modeling of Arbitrary Conditionals in Natural Language

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14943v1 公告类型:新。
      • 摘要:因果变换器通过联合分布的自回归分解对序列进行建模,从而实现高效的从左到右解码和条件似然计算。
      • 然而,它们无法轻松地从任意条件中采样或评估——例如,以过去和未来标记为条件的文本块。
      • 最近的工作旨在通过新颖的架构来解决这个问题,但它们通常会导致此类条件的次优建模和退化的生成。
    • EN 要点:
      • arXiv:2606.14943v1 Announce Type: new
      • Abstract: Causal Transformers model sequences through an autoregressive factorization of the joint distribution, which enables efficient left-to-right decoding…
      • However, they cannot tractably sample from or evaluate arbitrary conditionals – e.g., a block of text conditioned on past and future tokens
      • Recent work aims to solve this problem through novel architectures, but they often lead to sub-optimal modeling of such conditionals and degraded generations
  • CoRA: Confidence-Rationale Alignment for Reliable Chain-of-Thought Reasoning

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14961v1 公告类型:新。
      • 摘要:思想链 (CoT) 推理可以提高 LLM 的成绩,但当伴随的 CoT 基本原理看似合理但不完整或缺乏支持时,高答案置信度可能会产生误导。
      • 我们研究信心-基本原理对齐:模型对其承诺答案的信心是否由其生成的基本原理证明是合理的。
      • 我们引入了基于 GRPO 的强化学习框架,该框架联合奖励答案正确性、承诺答案概率和基于评分标准的理由支持,其中评分标准评估基础、连贯性、任务匹配以及与所选答案的联系,而不向法官透露黄金答案。
    • EN 要点:
      • arXiv:2606.14961v1 Announce Type: new
      • Abstract: Chain-of-thought (CoT) reasoning can improve LLM performance, but high answer confidence may be misleading when the accompanying CoT rationale is plau…
      • We study confidence–rationale alignment: whether a model’s confidence in its committed answer is justified by its generated rationale
      • We introduce a GRPO-based reinforcement learning framework that jointly rewards answer correctness, committed-answer probability, and rubric-based rationale sup…
  • Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.15007v1 公告类型:新。
      • 摘要:我们介绍 Nemotron 3 Ultra,这是一个总量为 5500 亿、活动参数为 550 亿的 Mixture-of-Experts 混合 Mamba-Attention 语言模型。
      • 我们在 20 万亿个文本标记上对 Nemotron 3 Ultra 进行预训练,然后将上下文长度扩展到 1M 个标记,并使用监督微调 (SFT)、强化学习 (RL) 和多教师按策略蒸馏 (MOPD) 进行后训练。
      • Nemotron 3 Ultra 是我们迄今为止最强大的模型,采用了多项关键技术 - LatentMoE、多令牌预测 (MTP)、NVFP4 预训练、多环境 RLVR、MOPD 和推理预算控制。
    • EN 要点:
      • arXiv:2606.15007v1 Announce Type: new
      • Abstract: We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model
      • We pre-trained Nemotron 3 Ultra on 20 trillion text tokens, then extended the context length to 1M tokens, and post-trained using Supervised Fine Tuning (SFT),…
      • Nemotron 3 Ultra is our most capable model yet, employing multiple key technologies - LatentMoE, Multi Token Prediction (MTP), NVFP4 pre-training, multi-environ…
  • Are Online Skill and Memory Modules Always Worth Their Tokens? A Budget-Constrained Study of Web Agents

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.15017v1 公告类型:新。
      • 摘要:在线网络代理通常会通过内存、工作流程或技能模块来增强基础参与者。
      • 这些模块可以提高性能,但它们也会消耗测试时间令牌,这种成本很少与参与者的推理成本一起报告。
      • 我们研究在线增强,这种开销是在每项任务上支付的,并在固定的总推理预算下重新评估其好处。
    • EN 要点:
      • arXiv:2606.15017v1 Announce Type: new
      • Abstract: Online web agents often augment a base actor with memory, workflow, or skill modules
      • These modules can improve performance, but they also consume test-time tokens, a cost rarely reported alongside the actor’s inference cost
      • We study online augmentation, where this overhead is paid on every task, and re-evaluate its benefits under a fixed total inference budget
  • Deep Temporal Modeling and Ensemble Fusion for Multimodal Emotion Recognition from Physiological Signals

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.15026v1 公告类型:新。
      • 摘要:生理压力和情绪识别对于健康监测和情感计算非常重要。
      • 在这项工作中,我们对 WESAD 数据集上的长短期记忆 (LSTM)、时间卷积网络 (TCN) 和 Transformer 等深度学习模型进行了全面评估,以使用手腕和胸部传感器信号进行多模态情感识别。
      • 我们进行消融研究,通过仅手腕和仅胸部输入训练模型来评估每种模式的个体贡献。
    • EN 要点:
      • arXiv:2606.15026v1 Announce Type: new
      • Abstract: Physiological stress and emotion recognition are important for health monitoring and affective computing
      • In this work, we present a comprehensive evaluation of deep learning models such as Long Short-Term Memory (LSTM), Temporal Convolutional Networks (TCN), and Tr…
      • We perform ablation studies to assess the individual contributions of each modality by training models on wrist-only and chest-only inputs
  • ReportQA: QA-Based Radiology Report Evaluation

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.15037v1 公告类型:新。
      • 摘要:放射学报告评估对于推进自动化报告生成至关重要。
      • 自然语言生成指标的临床相关性有限。
      • 临床疗效 (CE) 指标评估重要的医学发现,但主要关注存在并仅涵盖有限的一组实体。
    • EN 要点:
      • arXiv:2606.15037v1 Announce Type: new
      • Abstract: Radiology report evaluation is essential for advancing automated report generation
      • Natural language generation metrics have limited clinical relevance
      • Clinical efficacy (CE) metrics evaluate important medical findings, but focus mainly on presence and cover only a limited set of entities
  • Equity with Efficiency: An Empirical Study of Tokenizers for Multilingual Large Language Models

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.15044v1 公告类型:新。
      • 摘要:多语言大语言模型 (LLM) 依赖于子词标记化来连接离散文本和连续神经表示。
      • 最先进的多语言法学硕士通常使用字节级字节对编码 (BPE) 分词器,该分词器在结构上有利于高资源语言和拉丁文字。
      • 对于使用代表性不足的语言的人来说,尤其是东南亚地区的语言使用者,这种偏见会增加推理成本并扩大跨语言能力差距。
    • EN 要点:
      • arXiv:2606.15044v1 Announce Type: new
      • Abstract: Multilingual large language models (LLMs) depend on subword tokenization to bridge discrete text and continuous neural representation
      • State-of-the-art multilingual LLMs often use Byte-level Byte-Pair Encoding (BPE) tokenizers that structurally favor high-resource languages and Latin scripts
      • For speakers of underrepresented languages, particularly those across Southeast Asia, this bias inflates inference costs and widens cross-lingual capability gap…

ArXiv cs.LG (B_intro+search) 链接到标题

  • QPILOTS: Efficient Test-Time Q-Steering for Flow Policies

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14801v1 公告类型:新。 -摘要:流匹配和扩散策略是表达动作生成器,但使用时差强化学习(RL)对其进行优化仍然很困难。
      • 有效的策略提取需要利用批评家的动作梯度,但通过多步去噪过程直接反向传播该信号可能在数值上不稳定。
      • 现有方法通过丢弃梯度信息、将策略提炼为更简单的单步参与者或随着批评者的改进而反复微调去噪策略来解决此问题。
    • EN 要点:
      • arXiv:2606.14801v1 Announce Type: new
      • Abstract: Flow-matching and diffusion policies are expressive action generators, but optimizing them with temporal-difference reinforcement learning (RL) remain…
      • Effective policy extraction requires exploiting the critic’s action gradient, yet directly backpropagating this signal through a multi-step denoising process ca…
      • Existing methods work around this either by discarding gradient information, distilling the policy into a simpler one-step actor, or repeatedly fine-tuning the…
  • GRAPE: Guided Parameter-Space Evolution for Compact Adversarial Robustness

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14865v1 公告类型:新。
      • 摘要:对抗训练(AT)提高了神经网络的鲁棒性,但大多数方法从一开始就训练固定的参数空间。
      • 本文询问参数可优化的顺序是否会影响最终的鲁棒解决方案,即使最终的架构或计算预算受到控制。
      • 我们提出 GRAPE(引导参数空间演化),一种紧凑对抗鲁棒性的训练框架。
    • EN 要点:
      • arXiv:2606.14865v1 Announce Type: new
      • Abstract: Adversarial Training (AT) improves neural network robustness, but most methods train a fixed parameter space from the start
      • This paper asks whether the order in which parameters become optimizable can affect the final robust solution, even when the final architecture or computation b…
      • We propose GRAPE, Guided Parameter-Space Evolution, a training framework for compact adversarial robustness
  • {\alpha}-Fair Insurance Pricing: A Fairness Continuum

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14898v1 公告类型:新。
      • 摘要:保险定价的公平性仍然是一个长期存在且备受争议的难题。
      • 一方面,保险公司在盈利能力考虑的驱动下,根据个人风险设定不同的保费,以实现精算公平。
      • 另一方面,保险通过分散人群的风险、激励群体之间的交叉补贴以促进团结公平来发挥重要的社会功能。
    • EN 要点:
      • arXiv:2606.14898v1 Announce Type: new
      • Abstract: Fairness in insurance pricing remains a long-standing and deeply debated puzzle
      • On one hand, insurers, driven by profitability considerations, set premiums that differentiate across individual risks to achieve actuarial fairness
      • On the other hand, insurance serves a critical societal function by pooling risks across a population, motivating cross-subsidization among groups to promote so…
  • GRASP: Gradient-Aligned Sequential Parameter Transfer for Memory-Efficient Multi-Source Learning

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14900v1 公告类型:新。
      • 摘要:多源迁移学习面临一个基本的可扩展性瓶颈:现有方法要么需要在参数融合期间将所有 K 个源模型同时加载到内存中,需要 O(K) 内存,要么在推理时部署所有模型,这使得生产部署不可行。
      • 我们提出GRASP(梯度对齐顺序参数传输),它通过三个关键创新实现卓越的知识集成,同时保持 O(1) 内存消耗:(1) 顺序处理,一次将一个源合并到不断发展的目标模型中;(2) 参数梯度对齐,有选择地仅传输优化方向与目标域一致的参数,避免负迁移;(3) 迭代微调,在集成下一个源之前适应传输的知识。
      • 跨越 10 至 108 年时间分布变化和四种架构(1.3M 至 25.6M 参数)的三个连续学习基准(Yearbook、CLEAR-10、CLEAR-100)的广泛实验表明,GRASP 在所有数据集和架构上实现了 93.5% 的平均准确度,而集成方法的准确度为 71.7%,同时仅需要恒定内存,而标准多源的 K 模型则需要恒定内存融合。
    • EN 要点:
      • arXiv:2606.14900v1 Announce Type: new
      • Abstract: Multi-source transfer learning faces a fundamental scalability bottleneck: existing approaches require either loading all K source models into memory…
      • We propose GRASP (Gradient-Aligned Sequential Parameter Transfer), which achieves superior knowledge integration while maintaining O(1) memory consumption throu…
      • Extensive experiments across three continual learning benchmarks (Yearbook, CLEAR-10, CLEAR-100) spanning 10 to 108-year temporal distribution shifts and four a…
  • Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14929v1 公告类型:新。
      • 摘要:现代推荐系统越来越依赖于将不同的查询动态路由到多个嵌入模型。
      • 尽管具有实际意义,但在对抗性查询、强盗反馈和模型的有限可观察性等现实条件下,这个问题仍然知之甚少。
      • 我们将嵌入模型路由形式化为具有低等级专家的对抗性上下文线性强盗,其中上下文是查询,动作是项目,专家是在低等级潜在表示空间上工作的嵌入模型。
    • EN 要点:
      • arXiv:2606.14929v1 Announce Type: new
      • Abstract: Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models
      • Despite its practical significance, this problem remains poorly understood under realistic conditions like adversarial queries, bandit feedback, and limited obs…
      • We formalize embedding model routing as an adversarial contextual linear bandit with low-rank experts, where contexts are queries, actions are items, and expert…
  • Separable Neural Architectures as Physical World Models: from Mathematical Theory to Applications

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14934v1 公告类型:新。
      • 摘要:这项工作介绍了可分离神经架构(SNA),这是一种将神经逼近与张量分解相结合的函数表示类。
      • SNA 将局部坐标函数(原子)与由稀疏、低阶交互对象控制的全局交互解耦。
      • 该架构具有紧凑且平滑的归纳偏置,非常适合求解偏微分方程 (PDE)。
    • EN 要点:
      • arXiv:2606.14934v1 Announce Type: new
      • Abstract: This work introduces the Separable Neural Architecture (SNA), a function representational class combining neural approximation with tensor decompositi…
      • The SNA decouples localized coordinate functions (atoms) from global interactions governed by a sparse, low-rank interaction object
      • This architecture possesses a compact and smooth inductive bias well-suited for solving partial differential equations (PDEs)
  • Remember, Don’t Re-read: Stateful ReAct Agents for Token-Efficient Autonomous Experimentation

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14945v1 公告类型:新。
      • 摘要:自动研究模式通过大型语言模型 (LLM) 迭代修改代码以优化目标指标来实现自主实验。
      • 然而,其无状态设计在每次迭代时从头开始重建实验上下文,每次迭代产生 $O(n)$ 代币成本,总共 $O(n^{2})$。
      • 这项工作使用 LangGraph 将模式重新表述为有状态的 ReAct 代理,其中类型化的持久状态通过工具调用接口在迭代中携带实验历史记录。
    • EN 要点:
      • arXiv:2606.14945v1 Announce Type: new
      • Abstract: The autoresearch pattern enables autonomous experimentation by having a large language model (LLM) iteratively modify code to optimize a target metric
      • Its stateless design, however, reconstructs experimental context from scratch at every iteration, incurring $O(n)$ token cost per iteration and $O(n^{2})$ total
      • This work reformulates the pattern as a stateful ReAct agent using LangGraph, where typed persistent state carries experimental history across iterations via a…
  • A Comparative Study of Graph Neural Network Layer Selection for Interaction Modelling in Driving Trajectory Prediction

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14956v1 公告类型:新。
      • 摘要:自动驾驶系统依靠精确的轨迹预测来规划安全高效的运动。
      • 图神经网络(GNN)已成为对道路代理之间的时空交互进行建模的一种有前途的方法。
      • 然而,设计用于轨迹预测的 GNN 架构仍然是非标准化的,对于哪些图层有效捕获空间交互和时间动态几乎没有指导。
    • EN 要点:
      • arXiv:2606.14956v1 Announce Type: new
      • Abstract: Autonomous driving systems rely on precise trajectory prediction to plan safe and efficient movement
      • Graph Neural Networks (GNNs) have become a promising approach for modelling spatiotemporal interactions among road agents
      • However, designing GNN architectures for trajectory prediction remains non-standardized, with little guidance on which graph layers effectively capture spatial…
  • Leveraging Physiological Signals to Predict Exam Outcomes with Machine Learning

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14960v1 公告类型:新。
      • 摘要:本研究调查了机器学习模型的应用,以利用考试期间收集的生理数据来预测考试结果。
      • 分析生理压力指标,包括皮肤电活动、心率和皮肤温度,以揭示它们与学业成绩的关系。
      • 采用了各种机器学习方法,从逻辑回归、随机森林和支持向量机等标准模型到更先进的架构,包括变压器、长短期记忆 (LSTM) 和门控循环单元 (GRU) 模型。
    • EN 要点:
      • arXiv:2606.14960v1 Announce Type: new
      • Abstract: This study investigates the application of machine learning models to predict exam outcomes using physiological data collected during examination sess…
      • Physiological stress indicators, including electrodermal activity, heart rate, and skin temperature, were analyzed to uncover their association with academic pe…
      • A variety of machine learning approaches were employed, ranging from standard models like logistic regression, random forest, and support vector machines to mor…
  • Benchmarking Instance-Dependent Label Noise with Controlled Corruptions

    • 发布时间:2026-06-16 12:00 北京时间
    • 摘要:- arXiv:2606.14965v1 公告类型:新。 -摘要:合成实例相关标签噪声(IDN)基准广泛用于评估噪声标签学习方法,但现有方法通常通过不完善的注释器或分类器评估器生成噪声,从而隐含了模糊性的来源。
      • 我们引入了 CILN,这是一个基准生成框架,可通过受控的输入损坏创建 IDN。
      • 多样化的选民池标记损坏的实例,生成基准数据集,其中模糊性的来源和严重性都是明确且可控的。
    • EN 要点:
      • arXiv:2606.14965v1 Announce Type: new
      • Abstract: Synthetic instance-dependent label noise (IDN) benchmarks are widely used to evaluate noisy-label learning methods, yet existing approaches typically…
      • We introduce CILN, a benchmark generation framework that creates IDN through controlled input corruptions
      • A diverse voter pool labels corrupted instances, producing benchmark datasets in which both the source and severity of ambiguity are explicit and controllable