🤖 AI 速览

今天的焦点从模型参数转向两端:一边是 Stripe 收购 OpenRouter,AI 路由、计费和入口继续上移;另一边是 OpenAI 把国家安全监督与青少年 AI 教育纳入产品外延。工程侧则更关注提效、评测和成本控制。
📋 文章元数据
发布时间
2026-08-19
类型
ai-daily
字数
3626
阅读时长
18 min

2026-08-19 AI日更 | AI 基础设施与治理同时升温:Stripe 收购 OpenRouter,OpenAI 推监督与教育合作 链接到标题

今天的焦点从模型参数转向两端:一边是 Stripe 收购 OpenRouter,AI 路由、计费和入口继续上移;另一边是 OpenAI 把国家安全监督与青少年 AI 教育纳入产品外延。工程侧则更关注提效、评测和成本控制。

📖 本期 Watch List 深度导读 链接到标题

今天最值得深读的有三条线。首先是 OpenAI 连续更新安全与治理议题:从国家安全中的民主监督、网络关键能力下的模型开发节奏,到面向青少年的 ChatGPT 与 CodeAI 教育合作,主线很清晰——AI 能力扩张后,制度、教育和产品保护必须同步升级。

第二条是工程落地。Asana 用 Codex 两周清掉五年技术债,是今天最值得工程团队细看的案例;同时关于 LSP 是否能为 Coding Agent 节省 token 的论文,也把“智能体写代码”的成本问题落到了可测量层面。

第三条是效率与可靠性研究:token 膨胀路由、长上下文 BCMT、提示压缩、潜在推理可解释性、多语言 GRPO,都在回答同一个问题——下一阶段的模型竞争,不只是更强,而是更省、更可控、更可验证。

🌐 X 平台 AI 热点快讯 链接到标题

话题 1:Stripe Acquires OpenRouter for Over $7 Billion in Major AI Deal 链接到标题

  • 分类:AI · News
  • 概况:热度时间:2 days ago,相关帖子数:16000
  • 是什么事:据报道,Stripe 以超过 70 亿美元收购 AI 模型路由与网关平台 OpenRouter,成为今年引发广泛关注的一笔 AI 基础设施交易。
  • 为什么重要:这笔交易反映出 AI 生态竞争正从模型本身转向调用、路由、计费和基础设施层,Stripe 也在加深对 AI 支付与开发者入口的布局。
  • 讨论概况:X 上讨论主要集中在三点:70 亿美元估值是否过高、OpenRouter 作为连接 400 多个模型的中间层价值有多大,以及这是否意味着 Stripe 正在从支付公司转向更完整的 AI 基础设施平台。

话题 2:Tech Teams Build Custom Harnesses to Scale Claude AI Agents 链接到标题

  • 分类:AI · News
  • 概况:热度时间:22 hours ago,相关帖子数:409
  • 是什么事:技术团队正在为 Claude 等 AI Agent 构建定制化运行框架和测试/调度工具,以支持更大规模的自动化任务执行。
  • 为什么重要:这表明 AI Agent 的落地重点正从模型能力本身转向工程化基础设施,包括数据管线、评测、监控和成本控制,可能影响企业级 AI 应用的扩展速度。
  • 讨论概况:X 上的讨论主要集中在定制 harness 是否会成为 AI Agent 规模化的关键,以及大型“AI 数据工厂”模式能否带来可靠产出;分歧在于有人看好其提升自动化效率,也有人担心复杂度、成本和可控性问题。

话题 3:Anthropic CEO Predicts AI Will Cure Most Diseases in 5-10 Years 链接到标题

  • 分类:AI · Other
  • 概况:热度时间:2 days ago,相关帖子数:51000
  • 是什么事:Anthropic CEO 预测,AI 可能在未来 5 到 10 年内帮助治愈大多数疾病。
  • 为什么重要:这一判断反映了 AI 在生物医学、药物研发和临床研究中的高预期,也关系到 AI 能否真正加速科研转化并影响医疗产业格局。
  • 讨论概况:X 上讨论主要分为乐观与质疑两派:支持者强调 AI 在蛋白质结构、药物筛选和医学数据分析上的进展,质疑者则认为疾病机制复杂、临床试验周期长,监管、安全和责任问题难以被 AI 快速解决。

今日 X 上的 AI 舆情小结 链接到标题

今天舆论主线明显指向一个共识:AI 竞争正在从“谁的模型更强”转向“谁掌握更关键的基础设施、调度能力和场景入口”,无论是 Stripe 收购 OpenRouter,还是围绕 Claude Agent 的定制框架,都体现出行业对调用、路由、评测和成本控制的重视。分歧主要集中在估值与落地速度上,有人认为 OpenRouter 这类中间层和 Agent harness 会成为规模化的关键,也有人担心这类平台价值被高估、工程复杂度过高,最终难以形成稳定回报。Anthropic CEO 对“5 到 10 年内治愈大多数疾病”的乐观判断,则把这种分歧推向更高维度:支持者相信 AI 将显著加速科研转化,质疑者则认为生物医学的临床验证、监管与责任链条远比模型能力更难突破。潜在风险在于,资本和预期可能跑在落地前面,若基础设施投入、Agent 自动化和医疗叙事都被过度乐观定价,后续容易出现估值回调、技术兑现不足与安全合规争议。

💡 大佬观点(Influencer Insights) 链接到标题

今日 AI 大佬关注动态简报 链接到标题

一、共同关注的技术趋势与产品热点 链接到标题

1. Agent 型操作系统与编码平台成为主战场

  • DeepSeek Harness (DSH) 生态爆发:@vista8 与 @dotey 均大量讨论 DSH 的插件生态、GUI 客户端、二次元头像开发者群体,称“DSH 刚开源没几天,插件生态已经这么繁荣,未来可期”。@Pluvio9yte 甚至转发了民间打包的开箱即用桌面客户端。
  • 豆包全面 Agent 化:@dotey 称豆包“越来越像 Codex”,其 Windows 版通过虚拟桌面实现 GUI 操作,手机可远程遥控电脑 Agent,@vista8 也表示其体验“跟 Codex 很像”,“有点超预期”。
  • Cursor 代码托管平台 Origin 上线:@dotey 详解 Origin 以 AI Agent 为第一用户,支持每秒 22.6 次提交、内置 AI 自动解决合并冲突,形成“编辑器-云智能体-代码审查-代码托管”闭环。
  • Omarchy(Agent 优先 Linux 系统)走红:@vista8 跟风安装,并引述 DHH 称这是其最满意作品。

2. 端侧模型与本地化运行持续升温

  • @zhixianio 在 Mac Studio 上测试 DeepSeek V4 Flash 4bit 量化版,同时对比 Gemma 4 12B CoderQwen 3.6-35B-A3B MoE,结论是 12B 规模无法支撑“长篇、有状态、一次成型”的复杂编程任务,35B 仍为甜点。
  • @zhixianio 体验 MiniCPM-o 4.5 的音视频全双工交互,表示“很难想象这是一个 9B 模型能达到的效果”。
  • @zixianio 公布正在开发的“舔狗模块”(主动式 memory 系统)及消费级设备多端侧模型 workload scheduler。
  • @ruanyf 对比本地 AI 硬件:RTX 5090 与 AMD Strix Halo 板载方案,指出“很多时候板载芯片组才是更好的本地 AI 解决方案”。

3. AI 视频生成走向实用化与精细化

  • @Pluvio9yte 提出视频生成不能纯靠提示词抽卡,需用已有片段迭代投喂“视频生视频”模型,分段反复修改;并展示 MiniMax H3 与 Seedance 2.5 同提示词对比,后者微表情控制更佳。
  • @vist8 分享用 AI 视频工具快速生成恐龙科普视频。
  • @Pluvio9yte 发现采样步数从 4 提到 8 反而导致表情对称化,出现反直觉结果。

4. 多 Agent 协作与新型工作流平台

  • @dotey 介绍 Cumora(yetone 开源):将 AI Agent 变成聊天群正式成员,有人设、会主动发言,支持云端及本地 BYOA 模式,并内置协调机制防冲突。
  • @dotey 转发 Claude Code 跨 session 自动发消息功能引发吐槽,并指出最新版默认开启且难以关闭。

5. 大模型记忆、上下文积累与数据飞轮

  • @zhixianio “舔狗模块”主动式记忆、@vista8 付费使用 Obsidian 同步并强调“AI 时代需要认真对待上下文积累”、@dotey 提及 ZCode 等产品可能形成“产品使用→数据→模型升级”飞轮,均指向记忆和上下文正成为产品护城河。
  • @ruanyf 撰文解释大模型 缓存命中输入价格 是未命中的 1/50,呼吁充分利用缓存降费。

二、值得注意的独特观点与行业前瞻 链接到标题

  • “小模型+工具”路线的有力反驳 (@dotey 转 @_jasonwei):Jason Wei 认为仅靠 1B 认知核心加工具检索无法替代大模型内化知识,速度、理解深度、可靠性均不足,追求最高质量始终需要更大模型,符合“苦涩的教训”。
  • 代码即真相,Bash 足矣 (@dotey 转 Pi 作者):Pi 的两位作者认为代码不需要记忆系统/RAG,Bash 可任意组合,大部分场景无需 MCP,skill+脚本即可。
  • 美团反思“全员养虾” (@dotey 引述):美团高管称 2-3 月全员使用进口 AI 工具导致账单日消耗千万级且产生谬误干扰经营。
  • AI 编程可能比真人更昂贵 (@ruanyf):OpenClaw 创始人月消耗 Token 估值 130 万美元,无限量使用旗舰模型的企业成本惊人。
  • 小红书成为 Skill 发布平台 (@ruanyf):小红书上新 RedSkill 功能,笔记可附 Skill 文件一键复制安装,想做“Skill 的 GitHub”。
  • Anthropic:AI 开源是伪命题 (@ruanyf):创始人称只公开权重看不到内部运作,无法参与开发,不宜称开源。
  • AI 时代的免费成本与道德责任 (@ruanyf):SQLite 作者拒绝外部 PR,比喻为“收留免费小狗”,维护责任长达二十五年。
  • AI 提升效率后是否该放假 (@ruanyf):文章提出既然 AI 几小时干完一周活,放假一两天合乎逻辑。
  • Claude 限额的纠结心态 (@dotey):Anthropic 数次延期提额又反复,被评价“不爽利”,宁可一直维持 50% 加额。

三、推荐的工具与资源 链接到标题

  • Deep Seak Harness 插件生态 (@vista8):推荐 modlens(识图)、dsh-at-file(@引用文件)、dsh-paste-input(粘贴文件)、dsh-cc-tui(终端风格界面)。聚合站及 GUI 客户端均已出现。
  • Cumora (@yetone/@dotey):开源多 Agent 协作聊天室,网址 cumora.com,GitHub 开源,可本地部署或云端使用。
  • Omarchy OS (@vista8):Agent 优先的 Linux 发行版,由 DHH 打造,镜像可从官方下载。
  • 豆包 PC 端工作任务 (@dotey/@vista8):支持 GUI 操作、手机远程连电脑,国内用户可直接体验高级 Agent 功能。
  • OpenConnector (@ruanyf):开源密码连接网关,防止 AI Agent 泄漏凭据,支持 Cloudflare Workers 部署。
  • fireworks-tech-gaph skill (@vista8):用于生成技术类图片的 skill,已获万星,支持 12 种风格和 svg/png/gif。
  • 牛来.skill (@Pluvio9yte):用于小红书/抖音起号搞抽象的开源 Skill,素材取自龙餐馆。
  • Kimi K3 作为无魔法的强力大模型 (@vista8):透过它配置网络与 Agent 完成任务。
  • Codex 开启 1M 上下文方法 (@dotey):修改 config.toml 中 model_context_window 与 model_auto_compact_token_limit 即可。
  • Mail Agent 与 Codex 定时邮件总结 (@Pluvio9yte):用 GPT-5.4/5.6-luna 每晚总结邮件,高效处理多业务邮箱。
  • Starryblu 新加坡银行卡 (@AI_Jasonyu):部分博主正在免限免开卡费窗口期推广。
  • Giffgaff 转网 Lebara 保号教程 (@AI_Jasonyu):针对封号与漫游停服的自救实操。
  • CapWords (@nishuang):趣味实拍动效的 AI 学外语 App,小红书用户和背词死忠粉均适用。

📚 附录:今日 Watch List 更新源列表 链接到标题

时间窗口:最近 3 天;覆盖 22 个源;共 37 条更新

All-In Podcast (A_full) 链接到标题

  • Flock CEO Garrett Langley on Controversy, “Surveillance State” Claims, and Privacy vs Safety
    • 发布时间:2026-08-18 08:47 北京时间
    • 摘要:- AppLovin Ads 是 AppLovin 的 AI 广告平台,覆盖移动游戏领域超过 10 亿的每日活跃用户。
      • 全屏视频广告,观看时间中位数为 35 秒。
      • 广告商每天花费数十万美元获利。
      • 全球 3,000 多家企业使用销售税、增值税和商品及服务税。
      • 他们负责处理注册、备案和税率,因此您可以领先于风险。
    • EN 要点:
      • (0:00) The most controversial company in privacy right now, Flock CEO joins the show
      • (7:23) License plate data retention: 7 days solves 90% of crimes
      • (13:00) Camera vandalism, felony charges, and privacy concerns
      • (18:15) Dirty cops exposed: Flock’s audit tool got 9 Georgia officers fired

Stratechery by Ben Thompson (A_full) 链接到标题

  • Nvidia Backs OpenAI Data Center, Anthropic News, Google Buys Spirit Airlines Data
    • 发布时间:2026-08-18 18:00 北京时间
    • 摘要:- Nvidia 又达成一项协议,这次是与前沿实验室合作; Anthropic 的收入继续令人惊叹;也许数据最终就是石油。
      • 15 美元/月150 美元/年。
      • 通过每周三封电子邮件或播客对当天新闻进行实质性分析。
      • 策略采访
      • 采访领先的上市首席执行官、私营公司创始人,并与分析师同行进行讨论。
    • EN 要点:
      • Nvidia makes another deal, this time with a frontier lab; Anthropic’s revenue continues to amaze; and maybe data finally is oil.

OpenAI Blog (A_full) 链接到标题

  • Strengthening democratic oversight in national security

    • 发布时间:2026-08-19 03:00 北京时间
    • 摘要:- 人工智能正在改变民主政府保护人民的方式。
      • 它可以帮助阻止网络攻击,保护关键基础设施,更早地检测威胁,并让公务员在危机中更清楚地了解情况。
      • 使用得当,这些工具可以加强国家安全。
      • 民主监督有助于确保公共权力的使用对其所服务的人民​​负责。
      • 随着人工智能使国家安全工作更快、更有能力,负责监督这项工作的机构也需要跟上步伐。
    • EN 要点:
      • OpenAI launches an initiative to strengthen democratic oversight of AI in national security, supporting government institutions with tools, training, and expert…
  • Partnering with CodeAI to prepare the first AI generation

    • 发布时间:2026-08-18 19:00 北京时间
    • 摘要:- 今天的学生将是第一代在人工智能的陪伴下成长的人。
      • 对于家长和教育工作者来说,问题不仅仅是年轻人是否会使用人工智能,而是他们是否会学会批判性地评估其输出、了解其局限性并负责任地使用它。
      • 目前,使用和理解之间存在差距。
      • 年轻人需要了解人工智能的工作原理,批判性地思考其输出,并通过使用优先考虑其安全和发展的工具来培养塑造未来的技能。
      • 因此,OpenAI 和 CodeAI 宣布建立标志性合作伙伴关系,为学生和教育工作者提供工具和资源,帮助他们学习如何使用人工智能并从中受益。
    • EN 要点:
      • OpenAI and CodeAI are partnering to help students build AI literacy, think critically about AI, and develop the skills to use and shape it responsibly.
  • Pacing model development in an era of cyber-critical capabilities

    • 发布时间:2026-08-18 19:00 北京时间
    • 摘要:- OpenAI 正在加强对前沿人工智能模型的监控、协调和安全。
      • 了解新的保障措施如何指导模型开发的步伐。
      • OpenAI 正在加强前沿 AI 模型的监控、调整和安全性 了解新的保障措施如何指导模型开发的步伐。
      • 网络关键能力时代的模型开发节奏。
    • EN 要点:
      • OpenAI is strengthening monitoring, alignment, and security for frontier AI models
      • See how new safeguards are guiding the pace of model development.
  • Introducing ChatGPT for Teens: Built for learning, backed by protections

    • 发布时间:2026-08-18 19:00 北京时间
    • 摘要:- ChatGPT for Teens 帮助青少年自信地学习、批判性思考和使用人工智能,提供更强大的内置保护、健康使用功能以及家长的额外控制。
      • ChatGPT for Teens 帮助青少年自信地学习、批判性思考和使用 AI,具有更强大的内置保护、健康使用功能和附加内容……。
      • 推出面向青少年的 ChatGPT:专为学习而设计,并有保护措施支持。
    • EN 要点:
      • ChatGPT for Teens helps teens learn, think critically, and use AI with confidence, with stronger built-in protections, healthy-use features, and additional cont…
  • Asana cleared 5 years of engineering work in 2 weeks with Codex

    • 发布时间:2026-08-18 15:00 北京时间
    • 摘要:- Asana 使用 OpenAI Codex 在两周内替换了过时的测试系统,以约 12000 美元的价格完成了预计需要五年的工作。
      • OpenAI 博客中的这篇文章解释了 Asana 如何在 2 周内完成了 5 年的工程工作,并利用 Codex 塑造了更广泛的人工智能和基础设施格局。
      • Asana 在 2 周内通过 Codex 完成了 5 年的工程工作,这也为创始人、运营商和投资者带来了实际影响。
    • EN 要点:
      • Asana used OpenAI Codex to replace an outdated testing system in two weeks, completing work expected to take five years for about $12K.

ArXiv cs.AI (B_intro+search) 链接到标题

  • FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14550v1 公告类型:新。
      • 摘要:由于模型规模庞大、能源需求高和环境成本高,人工智能效率最近受到学术界和工业界的关注。
      • 虽然报告浮点运算 (FLOP) 是评估计算成本的传统方法,但 FLOP 和执行时间之间的关系并不简单,因为具有相同数量 FLOP 的层可能不具有相同的执行时间,因为某些操作比其他操作更容易并行化。
      • 本文着手复制一项研究中的原始实验,该研究提出了 $\alpha-FLOPs$ 估计公式,以验证结果是否仍然适用于更新、更强大的硬件。
    • EN 要点:
      • arXiv:2608.14550v1 Announce Type: new
      • Abstract: AI efficiency has recently taken the spotlight in both academy and industry due to massive model scales, high energy demands, and environmental costs
      • While reporting Floating Point Operations (FLOPs) is a traditional approach for assessing computational costs, the relationship between FLOPs and execution time…
      • This paper sets out to replicate the original experiments from a study that proposed the $\alpha-FLOPs$ estimation formula to verify whether the results remain…
  • Large Language Models Show Metacognitive Sensitivity in Medical Reasoning

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14552v1 公告类型:新。
      • 摘要:大语言模型 (LLM) 在医学中得到越来越多的评估和使用,但临床实用性取决于答案的准确性以及置信度是否跟踪证据质量和不确定性。
      • 我们开发了一个受控的、受心理物理学启发的临床基准来测试医学法学硕士的诊断选择和信心行为。
      • 该基准重点关注可能的阿尔茨海默型神经认知障碍 (AT-NCD) 与抑郁相关认知障碍 (DRCI)。
    • EN 要点:
      • arXiv:2608.14552v1 Announce Type: new
      • Abstract: Large language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whether confidenc…
      • We developed a controlled, psychophysics-inspired clinical benchmark to test diagnostic choice and confidence behavior in a medical LLM
      • The benchmark focused on probable Alzheimer-type neurocognitive disorder (AT-NCD) versus depression-related cognitive impairment (DRCI)
  • The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14558v1 公告类型:新。
      • 摘要:当前的多模态模型在识别静态视觉和听觉内容方面表现出了卓越的能力。
      • 然而,他们的抽象感知推理能力,即从动态生成过程中推断出看不见的信息的能力,仍然是一个关键且尚未充分探索的前沿领域。
      • 在本文中,我们介绍了不成文基准,这是一项旨在探讨这种抽象感知和认知能力的新挑战。
    • EN 要点:
      • arXiv:2608.14558v1 Announce Type: new
      • Abstract: Current multimodal models have demonstrated remarkable proficiency in recognizing static visual and auditory content
      • However, their capacity for abstract perceptual reasoning, inferring unseen information from dynamic, generative processes, remains a critical and underexplored…
      • In this paper, we introduce The Unwritten Benchmark, a new challenge designed to probe this abstract perceptual and cognitive ability
  • When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14559v1 公告类型:新。
      • 摘要:多智能体强化学习中的有效通信要求智能体不仅要决定\textit{什么}进行通信,还要决定何时进行通信?
      • 现有方法要么在每个时间步进行通信,要么通过强化策略梯度学习二元门\cite{singh2019},这是一种产生不稳定且无法解释的门控行为的高方差信号。
      • 我提出了一个有原则的替代方案:仅当代理学习到的信念分布之间的 KL 散度超过固定阈值时才进行通信。
    • EN 要点:
      • arXiv:2608.14559v1 Announce Type: new
      • Abstract: Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when
      • Existing approaches either communicate at every timestep or learn a binary gate through REINFORCE policy gradients \cite{singh2019}, a high-variance signal that…
      • I propose a principled alternative: agents communicate only when the KL divergence between their learned belief distributions exceeds a fixed threshold
  • Global AI Regulations for FAIR and Ethics in High-Risk Use Cases: A Comparative Review

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14562v1 公告类型:新。
      • 摘要:人工智能治理正在从自愿道德转向可执行的、基于风险的监管,但跨司法管辖区的分歧给高风险人工智能运营商带来了合规不确定性。
      • 我们提出了欧盟、美国和中国的比较矩阵,其中映射了 (i) 风险分类触发因素,(ii) 具有约束力的义务,(iii) 执行和问责机制,以及 (iv) FAIR 原则在实践中的实施程度。
      • 我们在三个高影响力领域对矩阵进行了压力测试:脑电图 (EEG) 引导的康复机器人、未来央行数字货币 (CBDC) 生态系统中人工智能支持的债务催收,以及新兴人工智能工厂基础设施中人工智能驱动的稀缺图形处理单元 (GPU) 资源分配。
    • EN 要点:
      • arXiv:2608.14562v1 Announce Type: new
      • Abstract: AI governance is shifting from voluntary ethics to enforceable, risk-based regulation, yet cross-jurisdictional divergence creates compliance uncertai…
      • We present a comparative matrix for the EU, US, and China that maps (i) risk classification triggers, (ii) binding obligations, (iii) enforcement and accountabi…
      • We stress-test the matrix on three high-impact domains: Electroencephalography (EEG)-guided rehabilitation robotics, AI-enabled debt collection in prospective C…
  • Position: AI Lock-In Is in Progress, and We Must Be Prepared

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14565v1 公告类型:新。
      • 摘要:人工智能安全研究主要集中在两个领域:技术一致性(确保人工智能系统产生与人类一致的输出)和生成式人工智能社会影响的监管(包括失业风险和劳动力市场混乱)。
      • 然而,一个同样重要的维度仍未得到充分探索:依赖人工智能系统本身所固有的风险。
      • 在这篇立场文件中,我们认为人工智能安全研究应该解决人工智能锁定问题,即过度依赖人工智能系统导致人类技能下降、削弱人类独立运作的能力,并在人工智能系统不可用或受到损害时产生系统漏洞。
    • EN 要点:
      • arXiv:2608.14565v1 Announce Type: new
      • Abstract: AI safety research has mainly focused on two areas: technical alignment (ensuring AI systems produce human-aligned outputs) and the regulation of gene…
      • However, an equally important dimension remains underexplored: the risk inherent in dependence on AI systems themselves
      • In this position paper, we argue that AI safety research should address AI Lock-In, the phenomenon whereby excessive reliance on AI systems leads to human deski…
  • Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14566v1 公告类型:新。
      • 摘要:最近评估大型语言模型(LLM)道德能力的工作主要集中在我们所说的道德价值问题,即模型输出是否符合人类道德价值观。
      • 相比之下,道德规范问题,即模型是否能够识别并正确应用情境相关的道德规范,仍然没有得到充分探索。
      • 我们认为这种不平衡源于该领域对描述性伦理框架的依赖,例如道德基础理论和科尔伯格的道德发展阶段,它们强调价值表征而不是规范应用。
    • EN 要点:
      • arXiv:2608.14566v1 Announce Type: new
      • Abstract: Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily on what we call the moral value problem, i.e., wh…
      • In contrast, the moral norm problem, i.e., whether models can identify and correctly apply context-sensitive moral norms, remains underexplored
      • We posit that this imbalance stems from the field’s reliance on descriptive ethics frameworks, such as Moral Foundations Theory and Kohlberg’s stages of moral d…
  • From Doyle to AGM: A Survey and an Implementation Roadmap for Belief Change

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14567v1 公告类型:新。
      • 摘要:本文提出了有针对性的叙述性回顾,为计算信念改变的实施奠定了历史和理论基础。
      • 以 Doyle 和 London 的 1980 年基础分类法为基础,我们追溯了信念修正从计算起源到 AGM 框架到当代方法的理论转变的演变。
      • 我们的分析展示了年度股东大会前的计算实用主义与年度股东大会理论结构的关系,揭示了这一演变的连续性和转变。
    • EN 要点:
      • arXiv:2608.14567v1 Announce Type: new
      • Abstract: This paper presents a targeted narrative review establishing the historical and theoretical foundations for computational belief change implementation
      • Seeded by Doyle and London’s foundational 1980 taxonomy, we trace the evolution of belief revision from computational origins through the theoretical transforma…
      • Our analysis demonstrates how pre-AGM computational pragmatism relates to AGM theoretical constructs, revealing both continuities and transformations across thi…
  • Position: AI Governance Needs ISO-like Interoperability Protocols, Not Just Laws

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14568v1 公告类型:新。
      • 摘要:随着人工智能 (AI) 系统深入融入关键的全球基础设施,建立健全的治理框架的紧迫性日益增强。
      • 然而,当前的方法,以特定司法管辖区的法律、政策和自愿框架(例如欧盟人工智能法案、中国的算法治理和美国的 NIST 人工智能风险管理框架)为主导,造成了碎片化的监管环境。
      • 在这份立场文件中,我们认为 \textbf{\textit{人工智能治理必须不仅仅建立在法律之上,而必须建立在类似 ISO 的互操作性协议之上,以实现标准化、机器可读的跨境风险沟通}}。
    • EN 要点:
      • arXiv:2608.14568v1 Announce Type: new
      • Abstract: As Artificial Intelligence (AI) systems become deeply integrated into critical global infrastructure, the urgency for robust governance frameworks has…
      • However, current approaches, led by jurisdiction-specific laws, policies, and voluntary frameworks such as the EU AI Act, China’s algorithm governance, and the…
      • In this position paper, we argue that \textbf{\textit{AI governance must be built not on laws alone, but on ISO-like interoperability protocols that enable stan…
  • Position: Certified Correctness in Neural Constraint Reasoning Requires Symbolic Integration

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14569v1 公告类型:新。
      • 摘要:用于约束满足问题的神经求解器已经实现了显着的分布内精度,但它们存在一个基本限制:即使模型报告高置信度,在分布变化下也会发生持续的约束违规。
      • 本立场文件认为,当存在硬约束且验证成本相对较低时,神经约束推理必须优先考虑符号集成而不是纯学习。
      • 我们证明我们对数独作为代表性 NP 完全测试平台的关注是合理的,因为它在简单验证和困难求解之间表现出尖锐的不对称性:检查候选解决方案仅需要多项式时间 $O(n^{2})$,而找到解决方案可能需要指数搜索。
    • EN 要点:
      • arXiv:2608.14569v1 Announce Type: new
      • Abstract: Neural solvers for constraint satisfaction problems have achieved remarkable in-distribution accuracy, yet they suffer from a fundamental limitation p…
      • This position paper argues that when hard constraints exist and the cost of verification is relatively low, neural constraint reasoning must prioritize symbolic…
      • We justify our focus on Sudoku as a representative NP-complete testbed because it exhibits a sharp asymmetry between easy verification and hard solving: checkin…

ArXiv cs.CL (B_intro+search) 链接到标题

  • Does a Language Server Save Tokens for Coding Agents? A Measurement Methodology and Preliminary Study

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.13568v1 公告类型:新。
      • 摘要:编码代理将大部分上下文预算用于检索。
      • 词汇检索 (grep) 是通用的、即时的、零设置的,但噪音很大:它无法区分定义、调用和注释。
      • 通过语言服务器协议 (LSP) 进行的语义检索是精确且类型化的,但需要一个正在运行的索引服务器并支付每个符号的往返费用。
    • EN 要点:
      • arXiv:2608.13568v1 Announce Type: new
      • Abstract: Coding agents spend most of their context budget on retrieval
      • Lexical retrieval (grep) is universal, instant, and zero-setup, but noisy: it cannot tell a definition from a call from a comment
      • Semantic retrieval via the Language Server Protocol (LSP) is precise and typed, but needs a running, indexed server and pays a per-symbol round-trip
  • Think in Latent, Explain in Language: Self-Explainable Latent Reasoning

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.13570v1 公告类型:新。
      • 摘要:潜在推理已成为基于文本的思想链 (CoT) 的强大替代方案,通过将冗长的推理压缩为紧凑的嵌入,显着提高了计算效率。
      • 然而,将推理压缩到潜在空间会使思维变得不透明,阻碍其可解释性。
      • 当前的方法存在明显的权衡:它们要么充当无法解释的“黑匣子”(例如,Coconut),其中潜在推理不是人类可读的,要么依赖单独的事后解码器来实现可解释性(例如,Heima),引入架构开销并将解释与实际推理过程脱钩。
    • EN 要点:
      • arXiv:2608.13570v1 Announce Type: new
      • Abstract: Latent reasoning has emerged as a powerful alternative to text-based Chain-of-Thought (CoT), offering significant gains in computational efficiency by…
      • However, compressing reasoning into the latent space renders the thinking opaque, hindering its interpretability
      • Current methods present a stark trade-off: they either function as unexplainable ‘‘black boxes’’ (e.g., Coconut), where the latent reasoning is not human-readab…
  • Not All Tokens Are Equal: Inflation-Aware Routing for Agentic LLM Systems

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.13571v1 公告类型:新。
      • 摘要:当语言模型第一次尝试无法回答查询时,代理系统会重试,每次都会消耗额外的令牌。
      • 这种重试开销在模型的每个代币价格暗示的价格与完整工作流程的实际成本之间造成了差距。
      • 我们将这个差距称为\emph{代币膨胀},并将其定义为真实工作流程成本与单次调用成本的比率。
    • EN 要点:
      • arXiv:2608.13571v1 Announce Type: new
      • Abstract: When a language model fails to answer a query on the first attempt, an agentic system retries, consuming additional tokens each time
      • This retry overhead creates a gap between what a model’s per-token price implies and what a full workflow actually costs
      • We call this gap \emph{token inflation} and define it as the ratio of true workflow cost to single-call cost
  • BCMT: Blockwise Causal Memory Transformer

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.13578v1 公告类型:新。
      • 摘要:Transformer 架构依赖于密集的自注意力来建模远程依赖关系,但这种机制在序列长度方面表现出二次复杂度。
      • 我们引入了 BCMT(Blockwise Causal Memory Transformer),这是一种用于长上下文语言建模的架构,可将本地令牌交互与全局上下文传播解耦。
      • 密集因果自注意力在局部块内独立应用,而每个块都会通过指数因果记忆聚合生成自适应摘要。
    • EN 要点:
      • arXiv:2608.13578v1 Announce Type: new
      • Abstract: Transformer architectures rely on dense self-attention to model long-range dependencies, but this mechanism exhibits quadratic complexity with respect…
      • We introduce BCMT (Blockwise Causal Memory Transformer), an architecture for long-context language modeling that decouples local token interactions from global…
      • Dense causal self-attention is applied independently within local blocks, while each block produces an adaptive summary aggregated through an exponential causal…
  • Jais 2: A Family of Arabic-Centric Open Large Language Models

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.13580v1 公告类型:新。
      • 摘要:Jais 2 是由 MBZUAI、Cerebras 和 Inception 联合开发的一系列以阿拉伯语为中心的大型语言模型,旨在推进以阿拉伯语为中心的语言建模,在本报告评估的阿拉伯语和文化基准中具有强大的性能。
      • 据我们所知,该系列包括最大的以阿拉伯语为中心的开放法学硕士,在 70B 参数上从头开始训练,以及评估的开放模型中具有竞争力的 8B 参数变体。
      • 以阿拉伯语为中心的定制词汇可以实现高效的训练和推理。
    • EN 要点:
      • arXiv:2608.13580v1 Announce Type: new
      • Abstract: Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric la…
      • The family includes, to our knowledge, the largest open Arabic-centric LLM trained from scratch at 70B parameters, and a competitive 8B-parameter variant among…
      • A custom Arabic-centric vocabulary enables efficient training and inference
  • IterCOMP: Reasoning-aware Adaptive Prompt Compression for Multi-hop Question Answering

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.13588v1 公告类型:新。
      • 摘要:多跳问答需要跨多个证据片段进行复杂的推理,这通常会淹没具有冗长且嘈杂的上下文的检索增强生成系统,从而损害效率和准确性。
      • 虽然现有的提示压缩方法试图解决这个问题,但它们通常是为单轮查询而设计的,无法捕获相互依赖的推理步骤。
      • 我们提出 IterCOMP,一个统一的、免训练的即时压缩框架,它将多跳推理纳入迭代压缩循环中。
    • EN 要点:
      • arXiv:2608.13588v1 Announce Type: new
      • Abstract: Multi-hop question answering requires complex reasoning across multiple evidence segments, which often overwhelms retrieval-augmented generation syste…
      • While existing prompt compression methods attempt to address this issue, they are typically designed for single-turn queries and fail to capture interdependent…
      • We propose IterCOMP, a unified, training-free prompt compression framework that incorporates multi-hop reasoning within an iterative compression loop
  • Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.13624v1 公告类型:新。
      • 摘要:大型音频语言模型 (LALM) 在语音识别和音频问答等音频理解任务中的使用越来越多,引发了对跨人口群体公平性的担忧。
      • 由于混杂因素,包括口语内容的语义变化和特定于说话者的特征,口语输入设置中的公平性评估具有挑战性。
      • 忽略这些因素可能会导致有关模型偏差的误导性结论。
    • EN 要点:
      • arXiv:2608.13624v1 Announce Type: new
      • Abstract: Large Audio Language Models (LALMs) have seen increasing use for audio understanding tasks such as speech recognition and audio question answering, ra…
      • Fairness evaluation in spoken-input settings is challenging due to confounding factors, including semantic variation in spoken content and speaker-specific char…
      • Ignoring these factors can result in misleading conclusions about model bias
  • GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.13698v1 公告类型:新。 -摘要:具有可验证奖励的强化学习(RLVR)通常通过组相对策略优化(GRPO)进行优化,已成为提高预训练语言模型推理能力的核心方法,但目前的研究仍然主要以英语为中心。
      • 我们对多语言和非英语 GRPO 进行了大规模的实证研究,涉及广泛的基础模型、训练语言和不同的推理语言奖励。
      • 我们发现母语推理训练通常与英语推理训练差距很小。
    • EN 要点:
      • arXiv:2608.13698v1 Announce Type: new
      • Abstract: Reinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for…
      • We conduct a large-scale empirical study of multilingual and non-English GRPO across a wide range of base models, training languages, and different reasoning la…
      • We find that training to reason in the native language often leaves only a small gap to training for English reasoning
  • CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.13706v1 公告类型:新。 -摘要:在检索增强和多代理管道中,现有的针对幻觉的防御措施仍然是片面的:尽管模态存在分歧,证据仍是可信的,辩论验证的是总体报告而不是个人主张,而且这种验证仅在起草后进行,导致代理间错误在最终文本之前未被发现。
      • 为了弥补这一差距,我们提出了 CLAIR-Fin,这是一个九代理框架,它将每个问题分解为在类型化的金融索赔账本中维护的原子索赔。
      • 每项索赔均通过不对称证据权威解决,该机构根据索赔类型确定证据信任,而不是将所有模式视为同等可靠;监管链验证,在起草和对抗性审查之间的交接处检查接地情况,而不仅仅是在管道出口处;适应性反驳循环,通过对抗性辩论来引导有争议的主张,辩论的深度与辩论发现的内容成比例;以及与连续幻觉风险指数相结合的最终必然审计,该指数将通过审查的主张与从未提出异议的主张区分开来。
    • EN 要点:
      • arXiv:2608.13706v1 Announce Type: new
      • Abstract: Existing defenses against hallucination in retrieval-augmented and multi-agent pipelines remain partial: evidence is trusted despite modality disagree…
      • To close this gap, we present CLAIR-Fin, a nine-agent framework that decomposes each question into atomic claims maintained in a typed Financial Claim Ledger
      • Each claim is resolved through Asymmetric Evidence Authority, which conditions evidence trust on claim type rather than treating all modalities as equally relia…
  • TeachMateGPT: A Multi-Agent Knowledge-Grounded Framework for Pedagogical Assessment Generation from Science Curriculum Materials

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.13708v1 公告类型:新。
      • 摘要:自动生成基于教科书的评估项目可以减少科学教师的工作量,但现有的检索增强生成(RAG)系统依赖于平面检索,仅支持单问题生成,缺乏针对弱证据的保障措施,并且不适合资源匮乏、考试结构的课程。
      • 我们通过 TeachMateGPT 解决了这些限制,这是一个多智能体系统,为基于课程的科学评估创作做出了四项进步。
      • (i) COPE,一个分层知识库,用多分辨率索引取代令牌窗口分块,该索引沿着教学大纲结构对文档进行分段,并通过可遍历的基于图形的谱系以三个粒度将它们链接起来,将证据与每个主题的教学水平相匹配。
    • EN 要点:
      • arXiv:2608.13708v1 Announce Type: new
      • Abstract: Automatically generating textbook-grounded assessment items can reduce science teachers’ workload, but existing retrieval-augmented generation (RAG) s…
      • We address these limitations with TeachMateGPT, a multi-agent system contributing four advances to curriculum-grounded science-assessment authoring
      • (i) COPE, a hierarchical knowledge base replacing token-window chunking with a multi-resolution index that segments documents along syllabus structure and links…

ArXiv cs.LG (B_intro+search) 链接到标题

  • Learning Discrete Riemannian Metrics for Physical Fields with Cochain-Frame Equivarianc

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14556v1 公告类型:新。
      • 摘要:网格上的物理场需要拓扑和几何的分离:守恒定律是拓扑的并且应该是精确的,而几何、材料响应和各向异性耦合必须从数据中学习。
      • 现有的神经代理经常在不受约束的消息传递中混合这些角色。
      • 我们引入了黎曼霍奇消息传递(RHMP),它将这种分离变成了一种架构原则。
    • EN 要点:
      • arXiv:2608.14556v1 Announce Type: new
      • Abstract: Physical fields on meshes require a separation between topology and geometry: conservation laws are topological and should be exact, while geometry, m…
      • Existing neural surrogates often mix these roles inside unconstrained message passing
      • We introduce Riemannian Hodge Message Passing (RHMP), which turns this separation into an architectural principle
  • Forward Pass Domain Adaptation (Without Cross-Layer Backpropagation)

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14563v1 公告类型:新。
      • 摘要:仅前向传递 MLP 训练 (FPO) 无需向后传递模型主体即可适应大型语言模型,在峰值训练内存减少约 40% 的情况下实现标准微调吞吐量的 2.7–3.2 倍,同时将域外基准保留在基线的种子噪声内,这是全网络微调无法可靠再现的属性。
      • FPO 依赖于单一的经验观察:在 Transformer 的后期层,输出层预测误差在我们调查的六个公共模型中以余弦相似度 0.47–0.59 近似真实梯度。
      • 我们引入了一个两分钟的诊断,可以量化任何模型每层的近似值,从而确定后期层适应的可行性。
    • EN 要点:
      • arXiv:2608.14563v1 Announce Type: new
      • Abstract: Forward-Pass-Only MLP training (FPO) adapts large language models without a backward pass through the model body, achieving 2.7–3.2x the throughput o…
      • FPO rests on a single empirical observation: at late layers of a transformer, the output-layer prediction error approximates the true gradient with cosine simil…
      • We introduce a two-minute diagnostic that quantifies this approximation per layer for any model, identifying where late-layer adaptation is viable
  • Coarse-to-Fine Multi-Resolution Diffusion Models for Trajectory Generation in Urban Systems

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14570v1 公告类型:新。
      • 摘要:了解人员流动对于交通管理、流行病控制和城市规划等广泛的城市应用至关重要。
      • 然而,由于隐私问题,大规模公共轨迹数据的可用性仍然有限,这给下游出行分析带来了挑战。
      • 现有的合成轨迹生成方法主要侧重于匹配全局分布相似性,而常常忽视不同空间和时间分辨率下的移动模式,这对于实际应用至关重要。
    • EN 要点:
      • arXiv:2608.14570v1 Announce Type: new
      • Abstract: Understanding human mobility is critical for a wide range of urban applications, including traffic management, epidemic control, and urban planning
      • However, due to privacy concerns, the availability of large-scale public trajectory data remains limited, posing challenges for downstream mobility analysis
      • Existing methods for synthetic trajectory generation primarily focus on matching global distribution similarity, while often overlooking mobility patterns acros…
  • Geometry Is Not Robustness: A Trajectory-Level Study of PGD Evaluation

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14594v1 公告类型:新。
      • 摘要:投影梯度下降(PGD)广泛用于评估对抗鲁棒性,通常通过最终对抗精度来评估,但它不会捕获整个攻击过程中的模型行为。
      • 最近的工作提出了轨迹级诊断,例如损失演化、梯度对齐和失败步骤,以更深入地了解对抗性优化动态。
      • 然而,这些诊断是否可靠地表明鲁棒性强度仍不清楚。
    • EN 要点:
      • arXiv:2608.14594v1 Announce Type: new
      • Abstract: Projected Gradient Descent (PGD) is widely used to evaluate adversarial robustness, typically via final adversarial accuracy, which does not capture m…
      • Recent work proposes trajectory-level diagnostics, such as loss evolution, gradient alignment, and steps-to-failure, for deeper insight into adversarial optimis…
      • However, whether these diagnostics reliably indicate robustness strength remains unclear
  • DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14614v1 公告类型:新。
      • 摘要:随着人工智能数据中心淘汰功能性 GPU,大量仍具有功能的加速器进入二级市场。
      • 本文研究了这些退役的 GPU 是否可以找到一个富有成效的来世,以形成一个可以为现代 LLM 推理服务的 DumpsterCluster,以及在什么条件下这种重新利用在经济上可行且环境上可持续。
      • 我们仅使用二手组件从头开始实际构建了一个 128-GPU DumpsterCluster,并运行了一年。
    • EN 要点:
      • arXiv:2608.14614v1 Announce Type: new
      • Abstract: As AI datacenters retire functional GPUs, vast quantities of still capable accelerators enter secondary markets
      • This paper investigates whether these retired GPUs can find a productive afterlife to form a DumpsterCluster that can serve modern LLM inference, and under what…
      • We physically built a 128-GPU DumpsterCluster from scratch using only second-hand components and ran it for one year
  • Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14617v1 公告类型:新。
      • 摘要:法律人工智能领域的一项反复出现的提议是通过将不确定性工具(具有信念传播的证据图、顺序贝叶斯赔率更新、Dempster-Shafer 组合和保形预测)融合到一个管道中来改进案件结果预测。
      • 我们在来自 LexGLUE 和 FairLex 的 1,000 个真实的欧洲人权法院案例中对此进行了测试,预测法院是否从案件的事实段落中发现了违反《公约》的情况。
      • 我们比较了两个前沿法学硕士(Claude Opus 4.8 和 GPT-5.5)的三个系列作为事实证据估计器:(A)原始法学硕士,(B)通过融合管道路由的法学硕士,以及(C)通过同一管道的术语频率基线。
    • EN 要点:
      • arXiv:2608.14617v1 Announce Type: new
      • Abstract: A recurring proposal in legal AI is to improve case-outcome prediction by fusing uncertainty tools (evidence graphs with belief propagation, sequentia…
      • We test this on 1,000 real European Court of Human Rights cases from LexGLUE and FairLex, predicting whether the Court found a Convention violation from the cas…
      • We compare three families across two frontier LLMs (Claude Opus 4.8 and GPT-5.5) as per-fact evidence estimators: (A) the raw LLM, (B) the LLM routed through th…
  • PIKFNO: An Interpretable Neural Operator Based on Physics Informed Kernel Function

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14619v1 公告类型:新。
      • 摘要:这项工作提出了一种新的可解释的神经算子框架,称为物理通知核函数神经算子(PIKFNO),它明确地将从控制方程导出的物理通知核函数合并到神经算子架构中。
      • 与 DeepONet 等依赖深度网络隐式学习基函数的传统神经算子不同,PIKFNO 通过物理通知的核函数来约束主干网络,从而使其算子结构与无网格配置方法中使用的核扩展保持一致。
      • 引入了两种构建策略:一种直接从数据中学习核函数,其中学习的核可以被视为非奇异基本解,而另一种通过解析基本解的变换来构建它们。
    • EN 要点:
      • arXiv:2608.14619v1 Announce Type: new
      • Abstract: This work proposes a new interpretable neural operator framework, termed the Physics Informed Kernel Function Neural Operator (PIKFNO), which explicit…
      • Unlike traditional neural operators such as DeepONet, which rely on deep networks to implicitly learn basis functions, PIKFNO constrains the trunk network throu…
      • Two construction strategies are introduced: one learns kernel functions directly from data, where the learned kernel can be regarded as a nonsingular fundamenta…
  • Explaining Reinforcement Learning Decisions in Self-adaptive Systems

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14620v1 公告类型:新。
      • 摘要:强化学习(RL)已广泛应用于自治和自*系统中,但强化学习策略,尤其是依赖神经网络的深度强化学习策略,缺乏透明度且难以理解。
      • 这可能会导致用户信任度下降,并使系统验证更具挑战性。
      • 为了应对这一挑战,本文介绍了使用强化学习替代现实 (EARL) 的解释,这是一个在 RL 设置中生成反事实解释的 Python 库。
    • EN 要点:
      • arXiv:2608.14620v1 Announce Type: new
      • Abstract: Reinforcement Learning (RL) has been extensively used in autonomous and self-* systems, but RL policies, especially deep RL ones relying on neural net…
      • This can lead to diminished user trust, and makes for a more challenging verification of systems
      • To address this challenge, this paper introduces Explanations using Alternative Realities for Reinforcement Learning (EARL), a Python library to produce counter…
  • Metaplasticity as adaptive gradient preconditioning for incremental learning

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14634v1 公告类型:新。
      • 摘要:生物智能通过补充学习系统(CLS)理论自然地防止灾难性遗忘,这是一种由突触化塑性在局部水平驱动的宏观巩固过程:个体突触的连续的、依赖于历史的神经调节。
      • 虽然人工神经网络在非平稳环境中努力解决稳定性-可塑性困境,但现有的解决方案通常需要任务标签或产生大量内存开销,与生物现实背道而驰。
      • 将这种局部神经调节重新定义为优化驱动的过程,我们引入了$\textbf{SynGAP}$:$\textbf{Syn}$aptic $\textbf{G}$eometric $\textbf{A}$daptive $\textbf{P}$reconditioning。
    • EN 要点:
      • arXiv:2608.14634v1 Announce Type: new
      • Abstract: Biological intelligence naturally prevents catastrophic forgetting through Complementary Learning Systems (CLS) theory, a macroscopic consolidation pr…
      • While artificial neural networks struggle with the stability-plasticity dilemma in non-stationary environments, existing solutions often require task labels or…
      • Re-framing this localized neuromodulation as an optimization-driven process, we introduce $\textbf{SynGAP}$: $\textbf{Syn}$aptic $\textbf{G}$eometric $\textbf{A…
  • Fractional Optimizers Meet Fractal Activation Functions: An Empirical Study of Multi-Scale Optimization in Neural Network

    • 发布时间:2026-08-18 12:00 北京时间
    • 摘要:- arXiv:2608.14636v1 公告类型:新。
      • 摘要:分数优化方法和分形激活函数是改进神经网络训练的两个独立方向。
      • 分数优化器通过分数导数和记忆效应扩展一阶优化,而分形激活则引入基于自相似 Weierstrass 和 Blancmange 型函数的多尺度非线性表示。
      • 在这里,我们在统一的实验框架内研究它们的相互作用。
    • EN 要点:
      • arXiv:2608.14636v1 Announce Type: new
      • Abstract: Fractional optimization methods and fractal activation functions are two independent directions for improving neural network training
      • Fractional optimizers extend first-order optimization through fractional derivatives and memory effects, whereas fractal activations introduce multi-scale nonli…
      • Here, we investigate their interaction within a unified experimental framework