🤖 AI 速览

今天的主线是 AI 助手从“能力展示”转向“规模分发”和“商业化落地”。Gemini 月活突破 10 亿,说明生态分发的作用正在放大;OpenAI 开始在 ChatGPT 试水广告,并推进 Daybreak 上 AWS,企业化路径更清晰。同时,本地个人智能体、低成本 Agent 和可靠性评测继续升温,行业重心正从参数竞争转向可用性、信任与落地效率。
📋 文章元数据
发布时间
2026-08-12
类型
ai-daily
字数
3434
阅读时长
17 min

2026-08-12 AI日更 | AI 助手进入规模化分发:Gemini 破十亿,OpenAI 开始试广告 链接到标题

今天的主线是 AI 助手从“能力展示”转向“规模分发”和“商业化落地”。Gemini 月活突破 10 亿,说明生态分发的作用正在放大;OpenAI 开始在 ChatGPT 试水广告,并推进 Daybreak 上 AWS,企业化路径更清晰。同时,本地个人智能体、低成本 Agent 和可靠性评测继续升温,行业重心正从参数竞争转向可用性、信任与落地效率。

📖 本期 Watch List 深度导读 链接到标题

今天最值得追的是“智能体进入生产”的主线:YC 对 OpenClaw 创始人 Peter Steinberger 的访谈,适合理解本地运行、接入消息应用并执行邮件/日历/文件工作流的个人智能体形态;同时,OpenAI 智能体能力、ChatGPT 广告测试与 Daybreak 登陆 AWS,显示大模型公司正把能力、分发和商业化同时推向企业场景。另一条主线是可靠性评估:多模态幻觉 Fuzzing、DocAtlas 长文档交互理解、Search-G1 grounded search、可解释语言模型,都在回应“模型能否被信任地使用”。此外,WuYuEval、低资源语言翻译、Dravidian 语言模型和波兰 VLM 评测提醒我们:AI 的下一阶段竞争,不只在通用能力,也在专业领域与本地文化覆盖。

🌐 X 平台 AI 热点快讯 链接到标题

话题 1:Google’s Gemini App Hits 1 Billion Monthly Users 链接到标题

  • 分类:AI · News
  • 概况:热度时间:6 hours ago,相关帖子数:3300
  • 是什么事:Google 宣布 Gemini 应用月活用户已突破 10 亿,并称语音交互正在成为其主要使用方式之一。
  • 为什么重要:这标志着生成式 AI 助手已进入超大规模分发阶段,也反映出 Google 借助安卓、搜索和生态整合推动 AI 普及的能力,对行业竞争格局很重要。
  • 讨论概况:X 上讨论主要集中在两点:一是 Gemini 是否真正追平或逼近 ChatGPT 的用户规模,二是这 10 亿数据究竟更多来自 Google 的分发能力还是产品本身的吸引力;同时也有人关注语音交互是否会成为 AI 助手的主流入口。

话题 2:DeepSeek V4 Flash Shines in Efficient AI Agent Tests 链接到标题

  • 分类:AI · News
  • 概况:热度时间:1 day ago,相关帖子数:7800
  • 是什么事:DeepSeek V4 Flash因在高效AI Agent测试中表现突出而受到关注。
  • 为什么重要:这表明更低成本、更快响应的模型可能在自动化任务、工具调用和多步骤推理等Agent场景中具备实际竞争力,有助于推动AI应用落地。
  • 讨论概况:X上的讨论主要集中在其性能与成本效率是否优于现有模型、测试基准是否可靠,以及DeepSeek在开源生态和商业AI Agent市场中的潜在影响。

话题 3:Jeff Dean Shares Emotional Final Days at Google During KDD Keynote 链接到标题

  • 分类:AI · News
  • 概况:热度时间:14 hours ago,相关帖子数:143
  • 是什么事:Jeff Dean 在 KDD 主题演讲中分享了自己在 Google 的最后阶段经历,引发广泛关注。
  • 为什么重要:Jeff Dean 是 AI 和计算机领域的重要人物,他的动向往往被视为 Google AI 战略、技术领导力和行业人才流动的重要信号。
  • 讨论概况:X 上的讨论主要围绕他是否真的将离开 Google、这对 Google AI 团队和产品布局意味着什么,以及他在 AI 领域的影响力和象征意义展开。

话题 4:Manus to Resume Independent Operations After Meta Deal Unwinds 链接到标题

  • 分类:AI · News
  • 概况:热度时间:9 hours ago,相关帖子数:1100
  • 是什么事:据称,Manus 在与 Meta 的合作或交易告吹后,将恢复独立运营。
  • 为什么重要:这反映了 AI 初创公司在巨头合作、并购与独立发展之间的战略选择,对行业生态、人才流动和技术路线都具有参考意义。
  • 讨论概况:X 上讨论集中在 Meta 交易为何失效、Manus 独立后能否保持融资与增长,以及这是否意味着 AI 初创公司仍应优先保持独立性。

话题 5:SpaceXAI Launches Grok Bot with Charming Animated Mascot 链接到标题

  • 分类:AI · News
  • 概况:热度时间:,相关帖子数:257
  • 是什么事:xAI 推出了带有可爱动画吉祥物形象的 Grok Bot,引发 X 平台关注。
  • 为什么重要:这表明 AI 助手正从单纯的文本工具向更具人格化、视觉化和陪伴感的产品形态演进,可能影响用户交互体验与品牌差异化竞争。
  • 讨论概况:X 上的讨论主要集中在吉祥物是否能提升 Grok 的亲和力和使用黏性,也有人质疑这类包装是否掩盖了模型能力、准确性和安全性等更核心的问题。

今日 X 上的 AI 舆情小结 链接到标题

今天的舆论主线是,AI 正从模型能力竞赛进一步转向大规模分发、低成本 Agent 落地和产品人格化体验的竞争:Google 借 Gemini 的十亿月活展示生态优势,DeepSeek V4 Flash 则代表效率型模型在 Agent 场景中的吸引力,Grok Bot 则体现助手产品向陪伴化、品牌化演进。共识在于,AI 助手已经进入更广泛的用户触达阶段,语音、多模态形象和自动化任务能力都会成为下一轮应用普及的重要入口。分歧主要集中在这些进展究竟来自真实产品力还是平台分发和营销包装,例如 Gemini 的用户规模含金量、DeepSeek 测试基准可信度、Grok 吉祥物是否只是表层创新,以及 Manus 独立发展是否优于巨头合作。潜在风险则包括巨头生态进一步挤压初创公司空间,行业过度依赖未经充分验证的基准和用户数据,同时人格化 AI 可能掩盖模型准确性、安全性和治理问题;而 Jeff Dean 相关传闻也加剧了外界对 Google AI 组织稳定性和人才流动的敏感情绪。

💡 大佬观点(Influencer Insights) 链接到标题

24小时 AI 趋势洞察:代码审核成本危机、Agent 基础设施爆发与模型“军备竞赛”的冷思考 链接到标题

以下是基于过去 24 小时 AI 领域 KOL 观点的深度总结:

1. 今日核心热点:从“写代码”到“审代码”的成本倒挂与 Agent 基础设施化 链接到标题

A. 开发范式的根本性转移:成本从编写端向审核端迁移 今天的核心共识在于,AI 极大降低了代码生成的成本,但这并没有消灭工作量,反而将其转移到了 Code Review 环节。

  • 成本转嫁理论:@dotey 提出了一个深刻洞察:以前写代码成本高,现代表写成本极低,导致开发者实际上将成本“转嫁”给了审核者。审核者面对 AI 生成的海量代码,需要付出更高的认知负荷去理解业务逻辑(“what to build”),这是代码审核争议加剧的底层原因。
  • “Write-Only”代码:@lijigang 则用极简的“Read-only 变 Write-only”概括了这一现象,AI 生成代码的可维护性危机正在显现。

B. Agent 成为新“浏览器”:Cloudflare 的底层布局 Agent 正在脱离重量级浏览器,转向更轻量级的 Web 运行时。

  • Kitesurf 发布:@Pluvio9yte 重点关注了 Cloudflare 推出的 Kitesurf,这是一个基于 Rust 并运行在 Workers 上的浏览器引擎。它的核心价值在于将 Agent 的“眼睛”从昂贵的 Chrome 实例替换为低成本、易扩缩容的基建。这标志着 Agent 运行环境正在从 GUI 自动化向 Headless API 级访问深水区演化,旨在消除 worktree 等传统本地开发模式带来的高磁盘占用(@dotey 转发了对 worktree 浪费存储的吐槽)。

C. AI 攻克黎曼猜想取得 80 年来最大突破

  • @dotey 详细报道了 Anthropic 未公开模型在黎曼猜想上的数学进展,将关键指标从 41.6% 提升至 67.2%。这是 AI 从“工具辅助”迈向“原创数学研究”的标志性事件,过程涉及极大规模的 Subagent 协作。

2. 值得注意的独特观点与行业前瞻 链接到标题

A. 关于 AI 边界的冷思考:“Token 无限”是伪命题?

  • Token 浪费与边际效应:@dotey 引用了一项硅谷大厂的激进实验:给 20 人团队 100 万美元预算无限制使用 Token。结论出人意料——AI 反而比人更贵,且出现了明显的组织效率天花板。人的思考外包导致细节掌控力下降。这给“无限上下文/Token”的鼓吹者泼了一盆冷水(即便 @LinearUncle 也调侃重点是有无限 Fable 的 Token 福利)。
  • 小模型的硬天花板:@zhixianio 评测了 Gemma 4 12B Coder 与 Qwen 35B,发现 12B 量级即使通过微调优化了“收敛效率”,也撑不住“长篇、有状态、一次成型”的复杂工程。这揭示了小模型在实用场景中的物理极限。

B. AI 水印的本质是“合规表演”?

  • @dotey 深度拆解了 Anthropic 的全量文本水印机制。他指出,水印是通过控制选词概率(红绿组)实现的隐形标记,且极易被密集改写攻破。前沿实验室心理有数,这更多是为了应对欧盟监管的“合规动作”,而非真正的防伪墙。

C. 招聘市场的剧变:AI 时代的“特种兵”

  • FDE 崛起:@dotey 转载了 Cursor 人才观:FDE(前线部署工程师)取代传统程序员成为最抢手岗位。业界需要“既能写代码,又能谈业务,帮客户优化账单而非最大化消耗”的复合型人才,这是 AI 商业化落地的关键瓶颈。
  • 软件巨头转向:@gefei55 复盘了从 AI 浏览器到桌面客户端(Codex, Workbuddy)的进化史,认为 Manus 定义了新一代 Agent 的交互范式(云端虚拟机执行),导致了后续 Claude Code 及一众竞品的诞生。

D. 极致的出海实战认知

  • @Pluvio9yte 分享了许多微观体感:在 Cursor 中使用 Grok 4.5 时,虽然速度快,但在“理解规划/需求对齐”方面比 Claude 差一档;MiniMax Agent 在生成短剧时由于产品逻辑缺陷成了“额度吞噬机”;营销上产品直接作为首页(低跳出率)比画廊展示更有效。

3. 推荐的工具与资源 链接到标题

开源利器与硬核技术:

  • OpenConnector(@ruanyf 推荐):开源的密码连接网关,专门解决 AI Agent 泄密风险,统一管理凭证,Agent 接触不到明文密码。
  • OpenCodex(@vista8 推荐):强烈推荐的终端工具,能够让你在 Codex 中无缝切换并调用 DeepSeek、Kimi、Gemini 等外部模型,打破单一模型的限制。
  • GEOHub Skill(@vista8 转发自 @yaojingang):一个将 GEO 研究、诊断与内容生产合一的超级 Skill,开源且持续迭代,或将成为 SEO 领域的专业标准。
  • React Bits / Uiverse / Motion Sites(@AI_Jasonyu 推荐):针对 Vibe Coding 出来的 UI 缺乏设计感的痛点,提供的动效组件库和设计提示词合集,适合前端快速出活。

生产力与工作流方案:

  • Airtap(@Pluvio9yte 与 @AI_Jasonyu 推荐):云手机方案。其 iMessage 功能可以让你通过短信操控云端美国真机,用于构建个人 AI 优质信息源、养号或处理海外日常任务,解决了海外身份和环境的稳定性问题。
  • GSC + Codex 工作流(@Pluvio9yte 提供教程):将 Google Search Console 数据接入 Codex,实现自动化数据分析与定时任务,将 SEO 维护完全流程化。

AI 算法进阶学习:

  • Nathan Lambert 的 RLHF 课程(@vista8 推荐):面向大众的硬核课程,提供免费 PPT 和视频,从 KL 散度讲到 DPO 及后训练,适合希望深入理解大模型训练微调的人。

海外支付基础设施:

  • PayPal CN 个人收款(@gefei55 跑通):解决了国内个人开发者 AI 出海的收款合规难题,国内身份证即可注册并接入网站收取美元。

📚 附录:今日 Watch List 更新源列表 链接到标题

时间窗口:最近 3 天;覆盖 22 个源;共 35 条更新

Y Combinator Podcast (B_intro+search) 链接到标题

  • Peter Steinberger: “Fun Is Velocity”
    • 发布时间:2026-08-12 03:53 北京时间
    • 摘要:- 您可能已经听说过 OpenClaw(以前称为 Clawdbot/Moltbot)。
      • 引起轰动的开源人工智能助手可以在您自己的设备上运行,与您已经使用的消息应用程序连接,并且超越聊天功能,实际执行管理电子邮件、日历、文件、工作流程等任务。
      • 现在来认识一下它背后的人。
      • YC 的 Raphael Schaad 与 OpenClaw 的创始人 Peter Steinberger 坐下来,讨论了病毒式个人 AI 代理背后的“顿悟”时刻、为什么本地优先代理可以取代当今的许多应用程序,以及个人代理将如何重塑软件的未来。
    • EN 要点:
      • Last November, Peter Steinberger was annoyed that there was no good way to talk to his coding agents from his phone, so he built one himself
      • A few months later, OpenClaw had exploded into one of the biggest open source AI projects in the world, with nearly 3,000 contributors and a peak of 4.7 million…

Stratechery by Ben Thompson (A_full) 链接到标题

  • Nvidia’s Risky Business
    • 发布时间:2026-08-11 18:00 北京时间
    • 摘要:- 听这个帖子**:**。
      • 1870 年 1 月 1 日,杰伊·库克 (Jay Cooke) 签署了一份合同,如果你仔细观察的话,这份合同可能会导致世界大战。
      • 1864年,国会创建了北太平洋铁路公司,目标是连接五大湖和普吉特海湾,最终从德卢斯到塔科马的轨道;该宪章包括邻近拟议线路的 4000 万英亩土地,以换取完成扩建工程。
      • 然而,在接下来的六年里,尽管联合太平洋铁路和中央太平洋铁路相互建设,推动了 1869 年 5 月连接萨克拉门托和奥马哈的金钉,北太平洋仍难以获得融资。
      • 北太平洋公司于 1866 年就资金问题与库克接洽,但缺乏支撑联合太平洋公司和中央太平洋公司的慷慨的联邦担保(值得注意的是,这导致了数量惊人的贪污行为);库克本人对联邦政府的财政权力并不陌生,因此对此并不感兴趣。
    • EN 要点:
      • Listen to this post :
      • Log in to listen
      • On January 1, 1870, Jay Cooke, hailed as an American hero for his role in financing the Union effort in the Civil War, signed a contract that would, if you squi…
      • In 1864, Congress had created the Northern Pacific Railway Company with the goal of linking the Great Lakes and Puget Sound with tracks that would eventually ru…

OpenAI Blog (A_full) 链接到标题

  • Testing ads in ChatGPT

    • 发布时间:2026-08-11 18:00 北京时间
    • 摘要:- 2026 年 8 月 11 日更新:ChatGPT 广告现已在英国、墨西哥、巴西、日本和韩国推出。
      • 今年我们将继续扩展到更多市场。
      • 2026 年 5 月 7 日更新:在未来几周内,我们计划在英国、墨西哥、巴西、日本和韩国扩大 ChatGPT 中的广告试点。
      • 这些试点项目将帮助我们了解哪些方法在不同地区行之有效,以便我们可以在业务扩展的过程中继续改进体验。
      • 2026 年 3 月 26 日更新:我们的广告试点重点是支持更广泛地访问 ChatGPT,同时维护消费者的信任、实用性和用户控制。
    • EN 要点:
      • OpenAI begins testing ads in ChatGPT to support free access, with clear labeling, answer independence, strong privacy protections, and user control.
  • Daybreak models are now available on AWS

    • 发布时间:2026-08-11 18:00 北京时间
    • 摘要:- 今年早些时候,OpenAI 前沿模型和 Codex 在 AWS 上普遍可用,为企业提供了将先进人工智能引入生产的新途径。
      • 今天,我们将分享我们与 AWS 合作的下一步:通过 Amazon Bedrock 提供 Daybreak 功能。
      • AWS 中提供 Daybreak Blue 和 Daybreak Red 访问级别:。
        • Daybreak Blue 提供对前沿通用模型的访问,包括 GPT‑5.6 Sol,并具有针对授权防御安全工作量身定制的保障措施。
        • Daybreak Red 提供对我们专门训练的网络安全模型的访问权限,以进行授权漏洞研究、漏洞利用验证和安全测试。
    • EN 要点:
      • OpenAI and AWS are making Daybreak cybersecurity capabilities available through Amazon Bedrock to support enterprise security workflows.

Two Minute Papers (B_intro+search) 链接到标题

  • OpenAI’s AI Agents Just Crossed A Line
    • 发布时间:2026-08-11 23:35 北京时间
    • 摘要:- ❤️ 在这里查看 Lambda 并注册他们的 GPU Cloud:。
      • 📝 更多报告可在此处获取:。
      • Adam Bridges、B Shang、Carlos Galarza、Christian Ahlin、Eric Tyson、Juan Benet、Lukas Biewald、Michael Tedder、Owen Skarpness、Ryan Stankye、Shawn Becker、Steef、Taras Bobrovytsky、Tazaur Sagenclaw、Tybie Fitzhugh、Ueli Gallizzi。
      • OpenAI 的人工智能代理刚刚跨越了界限。
    • EN 要点:
      • ❤️ Check out Lambda here and sign up for their GPU Cloud:
      • 📝 More reports are available here:
      • 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
      • Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef…

ArXiv cs.AI (B_intro+search) 链接到标题

  • Towards an Argumentative Foundation for Evaluative AI

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07473v1 公告类型:新。
      • 摘要:评估性人工智能(EAI)最近被提出作为支持人类决策的一种方式,不是通过产生单一建议,而是通过提出相互竞争的假设以及支持和反对每个假设的证据。
      • 在这篇立场文件中,我们主张(计算)论证作为一种特别合适的范式,为可解释和可争议的 EAI 形式提供正式的、可计算的基础,为分布式和以人为中心的 EAI 系统的长期研究议程奠定基础。
      • arXiv:2608.07473v1 公告类型:新摘要:评估人工智能(EAI)最近被提议作为支持人类决策的一种方式,不是通过产生单一建议,而是通过呈现……在这篇立场文件中,我们提倡(计算)论证作为一种特别合适的范式,为 EA 形式提供正式的、可计算的基础……。
    • EN 要点:
      • arXiv:2608.07473v1 Announce Type: new
      • Abstract: Evaluative AI (EAI) has been recently proposed as a way to support human decision-making, not by producing a single recommendation, but by presenting…
      • In this position paper, we advocate (computational) argumentation as a particularly suitable paradigm to provide a formal, computable foundation for forms of EA…
  • Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07474v1 公告类型:新。
      • 摘要:先前的工作表明,当人工智能输出速度 V 超过人类认知能力 C_max 时,人机交互监督在高损失领域中在结构上变得站不住脚。
      • 然而,操作约束不是单独的 V,而是 V x L,其中 L 表示每个项目的认知负荷。
      • L 由分类、判断和响应组成,它们对 AI 能力提升的响应不对称。
    • EN 要点:
      • arXiv:2608.07474v1 Announce Type: new
      • Abstract: Prior work showed that human-in-the-loop oversight becomes structurally untenable in high-loss domains when AI output velocity V exceeds human cogniti…
      • The operative constraint, however, is not V alone but V x L, where L denotes per-item cognitive load
      • L consists of triage, judgment, and response, which respond asymmetrically to AI capability improvement
  • Determinization in Structure Theories: A Unified Framework via Closure, Comparability, and Joint Admissibility

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07476v1 公告类型:新。
      • 摘要:我们开发了一个正式框架,用于从多元结构理论构建规范解释。
      • 结构理论是由签名、公理和推理策略组成的三重 T = ({\Sigma}, A, I),其可接受的解释族收集结构结论的所有全局一致分配。
      • 我们区分了三个级别的规范化:封闭稳定(每个种子收敛)、全局完成(与种子无关的收敛)和确定性(独特的可接受的解释)。
    • EN 要点:
      • arXiv:2608.07476v1 Announce Type: new
      • Abstract: We develop a formal framework for constructing canonical interpretations from plural structure theories
      • A structure theory is a triple T = ({\Sigma}, A, I) consisting of a signature, axioms, and an inference policy, whose admissible interpretation family collects…
      • We distinguish three levels of canonicalization: closure stabilization (per-seed convergence), global completion (seed-independent convergence), and determiniza…
  • Emotion in an active inference model of human driving

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07480v1 公告类型:新。
      • 摘要:主动推理已成为通过平衡目标导向行动与不确定性减少来建模自适应行为的原则框架。
      • 它已成功应用于生物和人工系统,包括最近在人类驾驶方面的工作。
      • 然而,现有的主动驾驶推理模型尚未解决交通行为的一个重要决定因素:情感状态,它对决策产生重大影响。
    • EN 要点:
      • arXiv:2608.07480v1 Announce Type: new
      • Abstract: Active inference has emerged as a principled framework for modeling adaptive behavior by balancing goal-directed action with uncertainty reduction
      • It has been successfully applied across biological and artificial systems, including recent work on human driving
      • However, existing active inference models of driving have yet to address an important determinant of behavior in traffic: affective state, which significantly i…
  • Training Variable Long Sequences with Data-Centric Parallel

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07524v1 公告类型:新。
      • 摘要:在可变长序列上训练深度学习模型带来了巨大的计算挑战。
      • 现有方法迫使我们在效率和易用性之间进行艰难的权衡。
      • 简单的方法使用静态配置,导致工作负载不平衡,效率低下,而复杂的方法则引入了显着的复杂性和新模型的代码更改。
    • EN 要点:
      • arXiv:2608.07524v1 Announce Type: new
      • Abstract: Training deep learning models on variable long sequences poses significant computational challenges
      • Existing methods force a difficult trade-off between efficiency and ease-of-use
      • Simple approaches use static configurations that cause workload imbalance low efficiency, while complex methods introduces significant complexity and code chang…
  • The Knowing-Saying Gap: When Probes See Errors that Confidence Misses

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07528v1 公告类型:新。
      • 摘要:线性探针以近乎完美的精度检测语言模型中损坏的上下文,但这并不能转化为可靠的故障预测。
      • 结果是对部署监控产生直接影响的分离。
      • 在多跳算术链上,检测损坏的探针无法提供有关最终答案正确性的信息;被迫采用结构化置信格式的模型会崩溃为两个值,并且错误率无法区分;跨跃点的探测持久性无法区分正确结果和错误结果,反驳了我们预先注册的“持久性击败峰值”假设。
    • EN 要点:
      • arXiv:2608.07528v1 Announce Type: new
      • Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction
      • The result is a dissociation with direct implications for deployment monitoring
      • Across multi-hop arithmetic chains, probes that detect corruption turn out to be uninformative about final answer correctness; models forced into structured con…
  • NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07530v1 公告类型:新。
      • 摘要:SHACL 是验证 RDF 知识图谱(KG)一致性的核心技术。
      • 然而,创作 SHACL 形状需要大多数领域专家所缺乏的技术专业知识。
      • 将自然语言要求转换为 SHACL (NL2SHACL) 将降低这一障碍。
    • EN 要点:
      • arXiv:2608.07530v1 Announce Type: new
      • Abstract: SHACL is a core technology for validating the conformance of RDF knowledge graphs (KGs)
      • Yet, authoring SHACL shapes requires technical expertise that most domain experts lack
      • Translating natural language requirements into SHACL (NL2SHACL) would lower this barrier
  • Dynamic Coalition Formation and Communication Pricing in Skill-Based Agentic AI Systems

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07532v1 公告类型:新。
      • 摘要:现代代理人工智能系统结合了多个具有异构技能的大型语言模型代理,但大多数架构要么提前修复通信,要么允许完全广播。
      • 两者都可能效率低下,因为令牌成本、延迟、冗余和错误传播随着活动代理和通信链路数量的增加而增加。
      • 我们将代理选择和通信建模为具有任务条件净效用 $U(C\mid x)=V(C\mid x)-\sum_{i\in C}c_i$ 的合作游戏,将联盟级别成本与代理激活成本分开。
    • EN 要点:
      • arXiv:2608.07532v1 Announce Type: new
      • Abstract: Modern agentic AI systems combine multiple large language model agents with heterogeneous skills, yet most architectures either fix communication in a…
      • Both can be inefficient because token cost, latency, redundancy, and error propagation increase with the number of active agents and communication links
      • We model agent selection and communication as a cooperative game with task-conditioned net utility $U(C\mid x)=V(C\mid x)-\sum_{i\in C}c_i$, separating coalitio…
  • MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07533v1 公告类型:新。
      • 摘要:具身代理​​是通过物理身体与其环境交互的智能实体。
      • 目前,实体代理的评估主要依赖于两种范式:(1)手动注释的视觉问答(VQA)对和(2)高级任务完成指标,例如导航或操作的成功。
      • 前者是劳动密集型的,并且注释质量存在差异。
    • EN 要点:
      • arXiv:2608.07533v1 Announce Type: new
      • Abstract: An embodied agent is an intelligent entity that interacts with its environment through a physical body
      • Currently, the evaluation of embodied agents primarily relies on two paradigms: (1) manually annotated Visual Question Answering (VQA) pairs and (2) high-level…
      • The former is labor-intensive and subject to variability in annotation quality
  • When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07538v1 公告类型:新。
      • 摘要:随着法学硕士代理人从决策支持转向自主采购,公司需要知道委托谈判者是否创造价值,可预测地分配价值,并避免赔钱合同。
      • 我们在典型的供应链讨价还价问题中研究这一点:拥有私人需求信息的买方与不知情的卖方协商数量支付合同。
      • 我们对来自 OpenAI、谷歌和阿里巴巴的 9 个法学硕士与 9,840 个法学硕士之间的谈判中经过验证的完美贝叶斯均衡进行了基准测试。
    • EN 要点:
      • arXiv:2608.07538v1 Announce Type: new
      • Abstract: As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictab…
      • We study this in a canonical supply chain bargaining problem: a buyer with private demand information negotiates a quantity-payment contract with an uninformed…
      • We benchmark nine LLMs from OpenAI, Google, and Alibaba against a validated Perfect Bayesian Equilibrium across 9,840 LLM-to-LLM negotiations

ArXiv cs.CL (B_intro+search) 链接到标题

  • Unified Hallucination Fuzzing for Multimodal Large Language Models

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07525v1 公告类型:新。
      • 摘要:幻觉仍然是多模态大型语言模型(MLLM)的一个持续挑战,严重限制了它们在高风险应用中的可靠性。
      • 现有的评估主要基于静态基准,存在分类覆盖范围狭窄和性能快速饱和的问题,无法反映模型在不断变化的现实场景中的稳健性。
      • 为了弥补这一差距,我们提出了一个系统的评估框架,将综合基准与自我演进的压力测试相结合。
    • EN 要点:
      • arXiv:2608.07525v1 Announce Type: new
      • Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applicat…
      • Existing evaluations, predominantly based on static benchmarks, suffer from narrow taxonomical coverage and rapid performance saturation, failing to reflect mod…
      • To bridge this gap, we present a systematic evaluation framework integrating a comprehensive benchmark with self-evolving stress testing
  • DocAtlas: Long-Document Understanding as Mutable-State Interaction

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07527v1 公告类型:新。
      • 摘要:长文档理解需要模型在多个页面、布局、表格、图形和图表中查找和组合证据。
      • 现有的检索增强系统通常在生成之前从静态索引中选择证据,而最近的代理系统增加了多轮工具的使用,但通常依赖于冻结的专有主干,其行为由提示设置。
      • 我们推出了 DocAtlas,这是一个将长文档理解视为可变状态信息查找过程的系统。
    • EN 要点:
      • arXiv:2608.07527v1 Announce Type: new
      • Abstract: Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts
      • Existing retrieval-augmented systems usually select evidence from a static index before generation, while recent agentic systems add multi-turn tool use but oft…
      • We present DocAtlas, a system that treats long-document understanding as a mutable-state information-seeking process
  • WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07529v1 公告类型:新。 -摘要:大型语言模型(LLM)越来越多地被用作技术助手,但它们在固体废物管理(SWM)方面的能力仍然难以评估,因为现有基准强调一般知识而不是工程、环境和政策约束下的专业决策。
      • 我们推出 WuYuEval,这是一个多层次的基准,用于评估 SWM 中的法学硕士,涵盖基础知识、领域推理和专家决策。
      • 经过质量审核后,WuYuEval包含一个基础模块,其中包含涉及六种任务类型和八个领域类别的4,590个封闭式多项选择问题,以及一个专家模块,其中包含247个基于场景的开放式问题,涉及多目标优化、约束权衡和系统设计。
    • EN 要点:
      • arXiv:2608.07529v1 Announce Type: new
      • Abstract: Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to…
      • We introduce WuYuEval, a multi-level benchmark for evaluating LLMs in SWM across foundational knowledge, domain reasoning, and expert decision-making
      • After quality auditing, WuYuEval contains a Foundation Module with 4,590 closed-ended multiple-choice questions across six task types and eight domain categorie…
  • Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07531v1 公告类型:新。
      • 摘要:搜索增强语言代理应仅在必要时检索外部信息,并将其答案基于检索到的证据。
      • 现有的外部奖励要么提供稀疏的结果监督,要么提供来自过程注释和法学硕士法官的更丰富的反馈。
      • 结果奖励很容易扩展,但无法区分基础检索和冗余搜索,而更丰富的信号需要在训练期间进行昂贵的注释或推理。
    • EN 要点:
      • arXiv:2608.07531v1 Announce Type: new
      • Abstract: Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence
      • Existing external rewards provide either sparse outcome supervision or richer feedback from process annotations and LLM judges
      • Outcome rewards scale readily but cannot distinguish grounded retrieval from redundant search, whereas richer signals require costly annotation or inference dur…
  • Scaling Inherently Interpretable Language Models

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07594v1 公告类型:新。
      • 摘要:可解释性通常被视为对能力的征税:语言模型被训练为不透明的系统,然后在事后用难以建立可靠性的方法进行解释。
      • 在这项工作中,我们挑战这个前提。
      • 我们不是对模型进行逆向工程,而是将可解释性作为训练管道的约束,并与语言建模目标一起进行优化。
    • EN 要点:
      • arXiv:2608.07594v1 Announce Type: new
      • Abstract: Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods w…
      • In this work, we challenge this premise
      • Rather than reverse-engineering a model, we make interpretability a constraint of the training pipeline, optimized alongside the language modeling objective
  • Embedding Initialization for Unseen Low-resource Languages in Multilingual NMT: A Case Study on Limbum-English Translation

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07629v1 公告类型:新。
      • 摘要:NLLB-200 等多语言神经机器翻译模型涵盖 200 种语言,但仍有数千种语言不受支持,其中包括喀麦隆的大多数 Grassfields Bantu 语言。
      • 当针对未见过的语言微调这些模型时,从业者必须选择代理语言标记,但不存在用于此选择的原则性方法。
      • 我们实现了嵌入初始化策略,其中语言标记是模型中已有的多种类型相关语言的嵌入的平均值。
    • EN 要点:
      • arXiv:2608.07629v1 Announce Type: new
      • Abstract: Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu…
      • When fine-tuning these models for an unseen language, practitioners must choose a proxy language token, yet no principled method exists for this selection
      • We implemented an embedding initialization strategy where a language token is the average of embeddings from multiple typologically related languages already in…
  • SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07641v1 公告类型:新。 -摘要:大型语言模型的快速发展已将调查写作从长达数月的手动工作转变为自动化过程。
      • 随着世代规模的扩大,可靠的评估成为瓶颈,法学硕士越来越多地被用作调查评估者。
      • 然而,现有的方法在很大程度上依赖于现成的法学硕士作为法官的方法,没有与人类审稿人系统地一致,并且仍然缺乏量化与人类审稿人的一致性的系统框架。
    • EN 要点:
      • arXiv:2608.07641v1 Announce Type: new
      • Abstract: The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process
      • As generation scales, reliable evaluation becomes the bottleneck, and LLMs are increasingly used as survey evaluators
      • However, existing approaches largely rely on off-the-shelf LLM-as-a-judge methods without systematic alignment to human reviewers, and there remains a lack of s…
  • Evaluating Dedicated Monolingual and Joint Multilingual Causal Models for Dravidian Languages

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07727v1 公告类型:新。
      • 摘要:德拉威语言,主要是泰米尔语、泰卢固语、卡纳达语和马拉雅拉姆语,仅占用于训练多语言语言模型的数据的一小部分,因此尚不清楚这些模型实际上保留了多少每种语言的能力。
      • 我从头开始训练了五个 GPT-2 架构模型,以将四个单语言模型(泰米尔语、泰卢固语、卡纳达语和马拉雅拉姆语各一个,每个模型都有自己的 32K 词汇量子词标记器)与一个跨所有四种语言共享 64K 词汇量子词标记器的多语言模型进行比较。
      • 所有 5 个模型均使用经过清理的 CC-100、维基百科和 Samanantar 数据进行训练。
    • EN 要点:
      • arXiv:2608.07727v1 Announce Type: new
      • Abstract: Dravidian languages, mainly Tamil, Telugu, Kannada, and Malayalam make up only a small part of the data used to train multilingual language models, so…
      • I have trained five GPT-2 architecture models from scratch to compare four monolingual models (one each for Tamil, Telugu, Kannada, and Malayalam, each with its…
      • All the 5 models are trained on cleaned CC-100, Wikipedia, and Samanantar data
  • The No-Meaning Falsity: The Structural Impossibility of the Arbitrary Sign in Classical Arabic

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07737v1 公告类型:新。
      • 摘要:本文研究了后现代关于不受限制的语义不确定性的主张及其任意符号的索绪尔基本公理是否与古典阿拉伯语的结构体系兼容。
      • 我们开发了阿拉伯语非连接形态的正式数学模型,其中词汇含义由不变词根和形态句法模式之间的相互作用决定。
      • 在此框架内,我们建立了形态对应定理,证明每个词汇项都是由根模式对唯一生成的,以及语义定位定理,证明词汇意义是在表面实现之前的派生层面上确定的。
    • EN 要点:
      • arXiv:2608.07737v1 Announce Type: new
      • Abstract: This paper investigates whether the postmodern claim of unrestricted semantic indeterminacy, and its foundational Saussurean axiom of the arbitrary si…
      • We develop a formal mathematical model of Arabic non concatenative morphology in which lexical meaning is determined by the interaction between an invariant roo…
      • Within this framework, we establish a Morphological Correspondence Theorem, demonstrating that every lexical item is uniquely generated by a root pattern pair,…
  • Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07763v1 公告类型:新。
      • 摘要:视觉语言模型(VLM)在图像字幕、视觉问答和图像到文本生成等任务上取得了出色的性能。
      • 然而,他们主要接受以英语为中心的数据训练,这限制了他们处理基于文化的视觉理解的能力,并导致无法解释特定区域的含义、符号内容和上下文相关的视觉线索。
      • 现有的文化能力基准通常是模板驱动的,并且侧重于表面层面的识别,这使得它们不足以评估文化背景下更深层次的语言和语用理解。
    • EN 要点:
      • arXiv:2608.07763v1 Announce Type: new
      • Abstract: Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generat…
      • However, they are predominantly trained on English-centric data, which limits their ability to handle culturally grounded visual understanding and leads to fail…
      • Existing benchmarks for cultural competence are often template-driven and focused on surface-level recognition, making them insufficient for evaluating deeper l…

ArXiv cs.LG (B_intro+search) 链接到标题

  • Application of Artificial Intelligence for Fraudulent Banking Operations Recognition

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07471v1 公告类型:新。
      • 摘要:本研究考虑应用人工智能识别银行欺诈的任务。
      • 近年来,由于新冠病毒大流行,银行欺诈变得更加普遍,因为许多业务大规模转移到在线平台,并且创建了许多慈善基金,犯罪分子可以利用这些基金来欺骗用户。
      • 目前的工作重点是机器学习算法,作为一种非常适合分析和识别网上银行交易的工具。
    • EN 要点:
      • arXiv:2608.07471v1 Announce Type: new
      • Abstract: This study considers the task of applying artificial intelligence to recognize bank fraud
      • In recent years, due to the COVID19 pandemic, bank fraud has become even more common due to the massive transition of many operations to online platforms and th…
      • The present work focuses on machine learning algorithms as a tool well suited for analyzing and recognizing online banking transactions
  • Data-Driven Fire-Zone Segmentation for Improved Short-Term Wildfire Prediction

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07472v1 公告类型:新。
      • 摘要:野火预测模型通常将研究区域离散成统一的网格,忽略点火的异质空间分布。
      • 我们通过证明数据如何离散化比使用哪种模型更重要来挑战这种范式。
      • 我们提出了一种无监督火区分割算法,将分水岭检测与 K 均值聚类相结合,直接根据历史火灾模式定义预测单元。
    • EN 要点:
      • arXiv:2608.07472v1 Announce Type: new
      • Abstract: Wildfire prediction models typically discretize study areas into uniform grids, ignoring the heterogeneous spatial distribution of ignitions
      • We challenge this paradigm by showing that how data is discretized matters more than which model is used
      • We propose an unsupervised fire-zone segmentation algorithm combining watershed detection with K-means clustering to define prediction units directly from histo…
  • Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07535v1 公告类型:新。
      • 摘要:多模态大语言模型(MLLM)通过模态对齐和融合来集成异构模态,从而实现更强的理解和推理。
      • 然而,这种架构转变重塑了机器学习的安全格局。
      • 模型复杂性的增加和跨模态交互产生了新的威胁,包括模态集成受损、模态错位和融合安全风险,反映了威胁建模超越单模态假设的转变。
    • EN 要点:
      • arXiv:2608.07535v1 Announce Type: new
      • Abstract: Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding an…
      • However, this architectural shift reshapes the safety landscape of machine learning
      • Increased model complexity and cross-modal interactions give rise to novel threats, including compromised modality integration, modality misalignment, and fused…
  • Tracing sources of epistemic uncertainty in deep learning predictions: homo- and hetero-scedastic linearized estimators

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07630v1 公告类型:新。 -摘要:我们采用两种经典的统计估计器来量化现代深度学习的不确定性,以便更清晰地了解归因于两个来源的不确定性:任意不确定性或局部稀缺数据。
      • 我们的方法利用近似费舍尔信息矩阵的最新进展,以实现扩展到实际架构。
      • 实验结果证明了每个测试点如何受到两个来源的不同影响,突出了我们的估计器在提高实际应用程序的稳健性方面的实际效用。
    • EN 要点:
      • arXiv:2608.07630v1 Announce Type: new
      • Abstract: We adapt two classical statistical estimators for quantifying uncertainty to modern deep learning, in order to provide clearer insights into uncertain…
      • Our approach leverages recent advances in approximate Fisher Information Matrices, to enable scaling to actual architectures
      • Experimental results demonstrate how each test points is differentially impacted by both sources, highlighting the practical utility of our estimators in improv…
  • SkillConsist: Detecting Inconsistencies in Agent Skills via Bidirectional Graph Alignment

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07639v1 公告类型:新。
      • 摘要:代理技能为 LLM 代理提供可重用的功能。
      • 特工技能不一致可能会暴露未公开的危险行为或导致错误的技能选择。
      • 最近的代理技能研究越来越多地检查代理技能一致性检测。
    • EN 要点:
      • arXiv:2608.07639v1 Announce Type: new
      • Abstract: Agent Skills provide reusable capabilities to LLM agents
      • Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wrong Skill selection
      • Recent Agent Skill research has increasingly examined Agent Skill consistency detection
  • PhysAttNet: Enhancing Predictive Performance in Industrial and Astrophysical Time Series via Physics-Informed Attention

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07681v1 公告类型:新。
      • 摘要:准确而稳健的时间序列预测在涉及物理过程的许多应用中至关重要,例如制造监控和天体物理事件检测。
      • 在这些设置中,预测模型必须在噪声、变异性和测量不确定性下保持可靠,同时捕获与物理上有意义的事件相对应的时间局部结构。
      • 卷积神经网络(CNN)由于其计算效率和强大的表示能力而被广泛用于此类任务。
    • EN 要点:
      • arXiv:2608.07681v1 Announce Type: new
      • Abstract: Accurate and robust time series forecasting is essential in many applications involving physical processes, such as manufacturing monitoring and astro…
      • In these settings, predictive models must remain reliable under noise, variability, and measurement uncertainty while capturing temporally localized structures…
      • Convolutional neural networks (CNNs) are widely used for such tasks due to their computational efficiency and strong representational capacity
  • CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07719v1 公告类型:新。 -摘要:离线强化学习反复训练来自固定转换池的策略,使得种子和超参数之间的冗余数据成本高昂,而朴素子采样可以消除长期信用分配所需的罕见转换。
      • 我们引入了 CODS,一种批评家引导的选择器,它在拟合算法匹配的批评家和冻结可重用子集之前获取高残差转换之间交替。
      • 与优先重放不同,CODS 会产生静态工件;与一次性残差选择不同,它会随着批评者的变化而刷新分数。
    • EN 要点:
      • arXiv:2608.07719v1 Announce Type: new
      • Abstract: Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters,…
      • We introduce CODS, a critic-guided selector that alternates between fitting an algorithm-matched critic and acquiring high-residual transitions before freezing…
      • Unlike prioritized replay, CODS produces a static artifact; unlike one-shot residual selection, it refreshes scores as the critic changes
  • Neural Operators for Immersed-Boundary Soft Swimmers Locomotion

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07722v1 公告类型:新。
      • 摘要:高保真浸入边界仿真解决了变形游泳者及其周围流动的耦合运动,但由此产生的成本限制了工程设计、参数研究和控制的重复评估。
      • 我们开发神经算子代理,用于对平面和体积鳗鱼游泳者产生的流体动力场进行时间预测。
      • 代理在从自适应流体结构模拟导出的规则网格场上进行训练,并以游泳者几何形状和雷诺数为条件。
    • EN 要点:
      • arXiv:2608.07722v1 Announce Type: new
      • Abstract: High-fidelity immersed-boundary simulation resolves the coupled motion of a deforming swimmer and its surrounding flow, but the resulting cost limits…
      • We develop neural-operator surrogates for temporal prediction of the hydrodynamic fields generated by planar and volumetric eel swimmers
      • The surrogates are trained on regular-grid fields exported from adaptive fluid–structure simulations and are conditioned on swimmer geometry and Reynolds numbe…
  • Finite Constant Frontiers and Auditable Regret Certificates for Average-Reward Reinforcement Learning

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07725v1 公告类型:新。 -摘要:平均奖励强化学习后悔是已知的对数因子,但由于概率模式、结构参数、对数归一化、先验信息和规划假设不同,已发布的保证的数字内容很难比较。
      • 我们引入了一种常量感知比较协议,并导出了用于通信 MDP 的显式有限下级证书。
      • 构造是一个二态块的二叉树;它的证明使用精确的轨迹级 Bernoulli KL 散度,并保持动作预算、直径、占用、导航成本和终端偏差明确。
    • EN 要点:
      • arXiv:2608.07725v1 Announce Type: new
      • Abstract: Average-reward reinforcement-learning regret is known up to logarithmic factors, but the numerical content of published guarantees is difficult to com…
      • We introduce a constant-aware comparison protocol and derive an explicit finite lower certificate for communicating MDPs
      • The construction is a binary tree of two-state blocks; its proof uses exact trajectory-level Bernoulli KL divergence and keeps action budget, diameter, occupanc…
  • LUCID: Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation

    • 发布时间:2026-08-11 12:00 北京时间
    • 摘要:- arXiv:2608.07746v1 公告类型:新。
      • 摘要:长视距人形机器人控制需要具备多功能的全身技能和可靠的高层决策。
      • 现有方法通常将预先训练的技能与脚本规划器、有限状态机或特定于任务的无模型策略相协调,限制了它们处理复杂任务序列的能力。
      • 为了解决这个限制,我们提出 \textbf{LUCID},一个基于分层模型的强化学习框架,通过想象的学习动态模型的推出来规划可重用的技能。
    • EN 要点:
      • arXiv:2608.07746v1 Announce Type: new
      • Abstract: Long-horizon humanoid loco-manipulation requires composing versatile whole-body skills and reliable high-level decision making
      • Existing methods often coordinate pretrained skills with scripted planners, finite-state machines or task-specific model-free policies, restricting their abilit…
      • To address this limitation, we propose \textbf{LUCID}, a hierarchical model-based reinforcement learning framework that plans over reusable skills through imagi…