🤖 AI 速览

Claude Fable 5 引发的热议验证了长时推理 Agent 的可行性,但其单日消耗已达等效数千美元的阈值,倒逼行业正视 Token 经济学。与此同时,端侧模型从“能用”走向“好用”,而 DeepSeek 设立“Agent Harness 研究员”则标志着,连接基座模型与自主行为的工程化承接体系正在成为一门独立的新学科。
📋 文章元数据
发布时间
2026-06-12
类型
ai-daily
字数
3628
阅读时长
18 min

2026-06-12 AI日更 | Fable 5 触达成本警戒线,AI Agent 正在催生“Harness”新工程学科 链接到标题

Claude Fable 5 引发的热议验证了长时推理 Agent 的可行性,但其单日消耗已达等效数千美元的阈值,倒逼行业正视 Token 经济学。与此同时,端侧模型从“能用”走向“好用”,而 DeepSeek 设立“Agent Harness 研究员”则标志着,连接基座模型与自主行为的工程化承接体系正在成为一门独立的新学科。

📖 本期 Watch List 深度导读 链接到标题

今天 AI 前沿呈现出两条鲜明的经纬线。一是 Agent 从“执行”向“决策”的深度进化。 arXiv 集中出现了多篇值得精读的论文:INFRAMIND 不再只做模型路由,而是直接感知底层 GPU 集群的运行时状态来调度多智能体;Self-Gated Clarification 则让智能体学会在推理分支点主动判定“何时该问”,这直击当前长周期 Agent 落地时的脆弱性。对正在构建研究或金融代理的团队来说,SciConBench 和 MoCA-Agent 分别从科学结论综合与金融数值推理切入,给出了全新的基准与架构。

二是大模型对齐的内部结构与行为预测新视角。 Dual-Stance Evaluation 揭示了谄媚与事实一致性在几何上居于不同子空间,这为精准转向提供了可能;而一篇立场论文则直接提出要把行为预测本身当作可学习的任务,绕开解释链条。此外,侧畔还有一份高质量访谈值得一听——Ben Bajarin 深度拆解了 Apple 在 AI 与计算上的最新布局,为端侧智能提供了非常克制的产业参照。

🌐 X 平台 AI 热点快讯 链接到标题

话题 1:Anthropic’s Claude Code Event Packs Tokyo with Developers 链接到标题

  • 分类:AI · News
  • 概况:热度时间:17 hours ago,相关帖子数:223
  • 是什么事:Anthropic 在东京举办的 Claude Code 开发者活动爆满,大量本土开发者齐聚现场。
  • 为什么重要:这标志着 Anthropic 正加速拓展亚洲开发者生态,强化其在 AI 编程助手领域与 OpenAI、GitHub Copilot 等竞品的全球竞争态势,并试探企业级 AI 工具在日本市场的落地潜力。
  • 讨论概况:X 上的讨论焦点集中在 Claude Code 相对于 GitHub Copilot 的实际编码表现、Anthropic 在亚洲的推广策略是否务实,以及部分开发者对活动内容到底偏重技术干货还是品牌营销存在分歧。

话题 2:OpenAI Hires Cybersecurity Leaders to Counter AI Risks 链接到标题

  • 分类:AI · News
  • 概况:热度时间:23 hours ago,相关帖子数:270
  • 是什么事:OpenAI 聘请网络安全领导人以加强自身防御并应对 AI 可能带来的安全威胁。
  • 为什么重要:此举表明领先的 AI 公司正投入治理与防御资源,主动应对 AI 技术可能被滥用带来的新型网络安全风险,对行业安全实践和标准具有风向标意义。
  • 讨论概况:X 上的讨论主要聚焦于新聘领导人的背景与过往经历、这一举措是否标志着 OpenAI 在安全承诺上迈出实质一步,以及它对 AI 行业安全标准的潜在影响;部分观点也在质疑该动作是否与近期内部安全争议有关。

话题 3:OpenClaw Developer Open-Sources AI Repo Maintainer Skills 链接到标题

  • 分类:AI · News
  • 概况:热度时间:9 hours ago,相关帖子数:488
  • 是什么事:OpenClaw 开发者开源了一套“AI 仓库维护者技能”,将代码审查、问题分流、文档更新等常见维护任务封装为可复用的自动化能力。
  • 为什么重要:此举旨在降低 AI 开源项目的维护成本与门槛,让智能体代替或协助人类处理重复性维护工作,有望加速 AI 工具链的民主化并提高社区协作效率。
  • 讨论概况:社区焦点集中在:1) 自动化维护能否真的保证代码质量与安全,还是会引入难以察觉的错误;2) 人类维护者的角色会否被边缘化;3) 不同项目间的维护风格差异如何被技能模板兼容,以及是否会导致维护行为的同质化。

话题 4:Peter Steinberger Open-Sources AI Skills for Autonomous Repo Maintenance 链接到标题

  • 分类:AI · News
  • 概况:热度时间:10 hours ago,相关帖子数:609
  • 是什么事:开发者 Peter Steinberger 开源了一套用于自主代码仓库维护的 AI 技能集。
  • 为什么重要:此举降低了自主软件维护智能体的开发门槛,推动 AI 从代码生成向长期、可靠的仓库级自动化管理演进,有助于验证大模型在真实、持续工程任务中的可靠性。
  • 讨论概况:主要焦点包括:自主修 Bug 与发 PR 的可靠性边界、如何避免自动化引入新问题,以及开源方案与现有 CI/CD 及代码审查流程的有效集成方式。

话题 5:Recursive AI Tops Benchmarks in Automated Research Breakthrough 链接到标题

  • 分类:AI · News
  • 概况:热度时间:13 hours ago,相关帖子数:540
  • 是什么事:递归型AI系统在自动化科研基准测试中取得最高成绩,实现突破性进展。
  • 为什么重要:这展现了AI自主进行科学研究并通过递归自我改进加速发现的能力,可能重塑科学研发范式,同时也引发对递归自我改进安全性与控制难题的深度关注。
  • 讨论概况:X平台上的讨论集中在:该成果是实现了真正自主科研突破还是单纯的基准刷榜;递归自我改进可能带来的失控风险与对齐挑战;相关模型是否开源及结果的可复现性。乐观者认为这开启了AI驱动的科学革命,悲观者则强调必须提前加强安全监管。

话题 6:Tesla Deploys FSD Supervised v14.3.4 with Smart Summon for Cybertruck 链接到标题

  • 分类:AI · News
  • 概况:热度时间:,相关帖子数:4300
  • 是什么事:特斯拉开始向 Cybertruck 推送 FSD Supervised v14.3.4 版本,并新增 Smart Summon 智能召唤功能。
  • 为什么重要:这表明特斯拉在全自动驾驶软件迭代上进一步统一了旗下车型的体验,同时将争议性车型 Cybertruck 纳入 L2 级辅助驾驶核心生态,有助于收集该特殊平台的路况数据,推动端到端模型的泛化能力。
  • 讨论概况:用户主要关注该系统在非铺装路面和恶劣天气下的表现,部分声音质疑 Cybertruck 异常造型对摄像头可视度的影响;还就 Smart Summon 在狭窄车位的实用性、与旧版本相比的改进程度,以及该版本是否仍要求持续注意力监控等问题展开激烈讨论。

今日 X 上的 AI 舆情小结 链接到标题

今天舆论的主线清晰地指向了 AI 系统从辅助工具向长期、自主的复杂任务执行者演进,这一趋势同时席卷了软件维护、科学研究和物理驾驶等领域。行业共识体现在头部公司正积极布局以抢占生态位——Anthropic 倾力拓展亚洲开发者社区,OpenAI 着重加强安全治理,特斯拉则将端到端驾驶模型覆盖至争议性车型,各方均认可开源与自动化是降低门槛、加速迭代的关键路径。分歧则集中在自主能力的真实可靠性上,社区围绕自动化维护是否会引入隐蔽缺陷、递归型 AI 是真正的科研突破还是基准刷榜,以及 Cybertruck 特殊造型对驾驶感知的影响等议题争执不下。潜在风险日渐凸显:无人值守的智能体可能渗透难以察觉的错误,递归自我改进的失控与对齐难题引发深切担忧,而 AI 滥用所催生的新型网络威胁,更让安全防御升级成为迫在眉睫的行业命题。

💡 大佬观点(Influencer Insights) 链接到标题

AI 行业日报:Claude Fable 5 引爆 Agent 开发新范式,端侧模型与成本焦虑并存 链接到标题

一、今日核心热点:Claude Fable 5 与 Agent 开发范式 链接到标题

1.1 Fable 5:能力跃升与成本争议并存 链接到标题

Claude Fable 5 成为今日绝对焦点,多位大佬密集测试并给出深度评测:

维度观察来源
能力边界思维边界更广、架构能力更强,前端创意能力显著提升,能指出人类设计的不合理之处并自主优化@zhixianio
推理时长可连续思考 15 分钟才开始行动,验证环节极多,“结果固然好,但时间耗得很长”@vista8 @dotey
成本现实消耗约为 Opus 的 1.5x,Max5 订阅用户 10 小时消耗等效 $1,500;@jerryjliu0 团队成员单日触发 3 次限额@Pluvio9yte @dotey
使用策略建议缓存重建设为 1 小时,非 Max 强度即可满足多数需求,“不敢随便选 Max"成为共识@Pluvio9yte @dotey

关键分歧:@Pluvio9yte 提出"反共识”——Fable 5 速度"像乌龟慢爬",消耗快被夸大,实际能力"类似于 Opus4.6++ 与 GPT-5.5++ 的结合版,未到惊艳程度";而 @zhixianio 则盛赞其 40 分钟完成 70% 工作并优化设计,“Shut up and take my money”。

1.2 Agent 开发范式演进 链接到标题

“Goal 指令"成为 Codex/Claude Code 新标配

  • @vista8 开发 “乔木 Goal Meta Skill”npx skills add joeseesun/qiaomu-goal-meta-skill),将一句话需求转化为可执行目标,支持"睡前执行、次日收菜"的长时任务模式
  • @dotey 实践 /goal 指令,长任务稳定性显著提升,“不用继续了”

Fable 5 的极端案例:@trq212 展示完全由 AI 编码生成的视频制作流程——Whisper 转写 → Subagent 选片 → FFmpeg 粗剪 → 手写 LUTs 调色 → Remotion 动画组件 → Figma MCP 协作,全程无传统非编软件介入


二、端侧模型:性能突破与生态成熟 链接到标题

2.1 实测进展 链接到标题

@zhixianio 的"苦行僧式修行"验证端侧模型已具备生产力:

  • 配置:Qwen3.6-35B-A3B-oQ6-fp16-mtp / oMLX / Native MTP / 128K CTX
  • 结论:“响应速度比远程 LLM 快,智商在线,原生多模态比 DSV4 Pro 还爽”

Google Gemma 4 系列持续迭代:

  • 12B 多模态模型在 M5Max 128G 上英语/日语识别"秒出”,中文"驴唇不对马嘴"
  • **QAT(量化感知训练)**新思路:训练阶段即假定会被量化,“以此为前提提升训练效果”

2.2 工具链完善 链接到标题

工具更新亮点
oMLX v0.4.0首个官方 Swift macOS 原生应用@jundotkim
Owlia Nest新增收藏系统、Markdown 在线编辑器、文件保存 API@zhixianio
baoyu-design skill支持导入 Figma 本地文件重建设计系统@dotey

三、独特观点与行业前瞻 链接到标题

3.1 成本焦虑与商业模式重构 链接到标题

“AI 比员工还贵"成为新现实

  • @ruanyf 测算:OpenClaw 创始人月消耗 6030 亿 Token,等效 $130 万;即使改用国产开源模型(价格 1/30-1/50),年成本仍达 200-300 万人民币
  • @lijigang 提出"真实指标"论:Token 消耗是虚假指标,“问题是否被更好解决"才是真实指标

新商业模式探索

  • @oran_ge(via @lijigang):“第一版免费,后续更新收费”——因 AI Coding 第一版最简单,维护最费心
  • @dotey 引用:OpenDoor 裁撤印度 200+ 人离岸团队,以"美国本土更小规模的 AI 原生团队"替代

3.2 软件工程本质再思考 链接到标题

@dotey 核心论断:“AI 没有重新定义软件工程,AI 放大了软件工程的重要性”

@Pluvio9yte 的实践验证:

“Vibe Coding 的最佳实践不是 Requirement First 或 Code First,而是 Contract First。没有定义好契约,其他一切都是空谈。”

其基于 OpenSpec 二开的开发框架,目标将"容易漂移的上下文外化成契约,让人和 AI 都有稳定参照物”。

3.3 DeepSeek 的"Harness"战略 链接到标题

@dotey 披露 DeepSeek 招聘 “Agent Harness 研究员”——世界范围内首次明确招聘该岗位:

“Model + Harness = Agent。除模型本身外的所有工作,都属于 Harness:上下文管理、长期记忆、Subagent 与 Multi-Agent、自进化 Agent…”

定义 “Harness Engineering” 为新学科,要求候选人"重度 Agent 用户、全栈开发、从 0 到 1 推动研究”。


四、推荐工具与资源 链接到标题

4.1 开发工具 链接到标题

工具用途来源
Fable 5长时推理、复杂架构设计Anthropic
Codex + /goal自动化长任务开发OpenAI
乔木 Goal Meta Skill一句话需求转目标@vista8
baoyu-design skill本地 Claude Design + Figma 导入@dotey
oMLXmacOS 原生端侧模型运行@jundotkim
Owlia NestPA 产出文件浏览器@zhixianio

4.2 内容创作 链接到标题

工具用途来源
焚决视频翻译全流程(下载→转写→翻译→润色→烧字幕)@xiaohu
乔木书籍解读 Skill多 Subagent 协作生成口播脚本@vista8
AllyHubYouTube 频道数据分析与选题规划@Pluvio9yte

4.3 基础设施 链接到标题

  • OfoxAI:中转站,“稳定+优惠"兼顾,GPT-5.5/5.4 mini 85 折(@AI_Jasonyu)
  • Vercel:“最快上线网站方式”,Codex 插件几句话部署(@Pluvio9yte)

五、关键趋势总结 链接到标题

┌─────────────────────────────────────────┐
│  今日核心矛盾:能力跃升 ↑  vs  成本失控 ↑  │
├─────────────────────────────────────────┤
│  • Fable 5 验证"长思考"Agent 的可行性    │
│  • 端侧模型从"能用"走向"好用"            │
│  • "Harness Engineering"成为新工程学科  │
│  • Token 经济学倒逼商业模式创新           │
│  • Contract First 取代 Vibe Coding 随意性 │
└─────────────────────────────────────────┘

明日关注点:Fable 5 的 6 月 22 日订阅截止前的用户留存策略、Gemma 4 QAT 实测结果、DeepSeek Harness 团队组建进展。

📚 附录:今日 Watch List 更新源列表 链接到标题

时间窗口:最近 3 天;覆盖 22 个源;共 35 条更新

Stratechery by Ben Thompson (A_full) 链接到标题

  • An Interview with Ben Bajarin About Apple, AI, and Compute
    • 发布时间:2026-06-11 18:00 北京时间
    • 摘要:- Ben Bajarin 关于 WWDC 和 AI 计算行业现状的采访。
      • 15 美元/月150 美元/年。
      • 通过每周三封电子邮件或播客对当天新闻进行实质性分析。
      • 策略采访
      • 采访领先的上市首席执行官、私营公司创始人,并与分析师同行进行讨论。
    • EN 要点:
      • An interview with Ben Bajarin about WWDC and the status of the AI compute industry.

OpenAI Blog (A_full) 链接到标题

  • OpenAI to acquire Ona

    • 发布时间:2026-06-11 08:00 北京时间
    • 摘要:- 每周有超过 500 万人使用 Codex 来研究、分析、构建和自动化他们的工作,比今年早些时候增加了 400%。
      • Codex 最初是作为软件开发人员的工具,现在帮助更广泛的人完成从最初请求到最终结果的复杂工作。
      • 随着法典变得更加强大,其最有价值的工作将在数小时或数天内展开,而不是几分钟。
      • 我们相信人们应该能够委派更雄心勃勃的工作,而不必继续受制于工作开始的机器。
      • 工作应在初次会议之后继续进行,Codex 使人们能够在任何地方保持联系并检查进度、提供方向、做出决策和审查结果。
    • EN 要点:
      • OpenAI plans to acquire Ona to expand Codex with secure, persistent cloud environments, enabling long-running AI agents across enterprise workflows.
  • Supporting Europe’s work in ensuring a trustworthy AI ecosystem

    • 发布时间:2026-06-11 08:00 北京时间
    • 摘要:- 人们正在使用人工智能以新的方式创建和编辑内容。
      • 随着这些工具变得更加强大和更广泛使用,人们应该了解他们在网上看到的内容的背景。
      • 该准则是实施欧盟人工智能法案和建立更加透明的数字生态系统的重要一步。
      • 我们的支持建立在多年的内部研究、产品开发以及与更广泛的生态系统的合作的基础上,以加强人工智能生成内容的来源。
      • 凭借多年的专业知识,我们与数百名其他利益相关者一起为本准则的制定做出了贡献,以帮助确保建立值得信赖的人工智能生态系统。
    • EN 要点:
      • OpenAI supports the EU Code of Practice on AI content transparency, advancing provenance standards and tools to help people understand AI-generated content.
  • How an astrophysicist uses Codex to help simulate black holes

    • 发布时间:2026-06-11 08:00 北京时间
    • 摘要:- 了解天体物理学家 Chi-kwan Chan 如何使用 Codex 构建黑洞模拟,帮助科学家研究极端物理学并测试爱因斯坦的广义相对论。
      • 了解天体物理学家 Chi-kwan Chan 如何使用 Codex 构建黑洞模拟,帮助科学家研究极端物理并测试爱因斯坦的生成理论。
      • 天体物理学家如何使用 Codex 帮助模拟黑洞。
    • EN 要点:
      • Discover how astrophysicist Chi-kwan Chan uses Codex to build black hole simulations, helping scientists study extreme physics and test Einstein’s theory of gen…
  • BBVA puts AI at the core of banking with OpenAI

    • 发布时间:2026-06-11 08:00 北京时间
    • 摘要:- 了解 BBVA 如何将 ChatGPT Enterprise 扩展到 100,000 名员工,并与 OpenAI 合作,加速全球人工智能驱动的银行业转型。
      • OpenAI 博客的这篇文章解释了 BBVA 如何将人工智能置于银行业的核心,并通过 OpenAI 塑造更广泛的人工智能和基础设施格局。
      • 继 BBVA 通过 OpenAI 将人工智能置于银行业核心之后,这也为创始人、运营商和投资者带来了实际影响。
    • EN 要点:
      • Learn how BBVA scaled ChatGPT Enterprise to 100,000 employees and partnered with OpenAI to accelerate AI-powered banking transformation worldwide.

ArXiv cs.AI (B_intro+search) 链接到标题

  • From Explicit Elements to Implicit Intent: A Predefined Library for Auditable Behavioral Inference

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11207v1 公告类型:新。
      • 摘要:我们提出了 SemantiClean,这是一个模块化框架,用于从电子商务会话数据中提取结构化语义信号,并通过共享元素库驱动可插入推理目标,包括购买意图、客户细分和产品亲和力。
      • 与仅针对准确性进行优化的传统端到端预测器不同,SemantiClean 优先考虑可审核性、结构治理和 sigma=0 再现性,明确用边际预测收益换取元素级透明度和可防御的决策轨迹。
      • 该框架以在线购物者购买意向 (OSPI) 数据集为基础,将 24 个行为元素组织成一个四层架构(功能、交互、系统、上下文),并通过三种反通货膨胀机制强制执行信号质量:RedundancyGroup 贡献上限、TieredPenaltyCalculator 偏差处罚和 AdaptiveConstraintMode 冷启动保护。本报告介绍了 LLM 集成语义推理引擎,这是一个完全实现的两阶段 LLM 驱动的推理架构,在推理时利用完整的元素元数据。
    • EN 要点:
      • arXiv:2606.11207v1 Announce Type: new
      • Abstract: We present SemantiClean, a modular framework for extracting structured semantic signals from e-commerce session data and driving pluggable inference t…
      • Unlike conventional end-to-end predictors that optimise solely for accuracy, SemantiClean prioritises auditability, structural governance, and sigma=0 reproduci…
      • Built upon the Online Shoppers Purchasing Intention (OSPI) dataset, the framework organises twenty-four behavioural elements into a four-layer architecture (Fun…
  • Position: Hippocampal Explicit Memory Is the Cornerstone for AGI

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11245v1 公告类型:新。
      • 摘要:大型语言模型 (LLM) 在各种任务中表现出了卓越的能力,提高了人们对通用人工智能 (AGI) 的期望。
      • 本立场文件认为,整合外显记忆是法学硕士迈向 AGI 的基石。
      • 关键原因是LLM的底层学习机制与人类的内隐记忆高度相似。
    • EN 要点:
      • arXiv:2606.11245v1 Announce Type: new
      • Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, raising expectations for Artificial General Intelligence…
      • This position paper argues that integrating explicit memory is the cornerstone for advancing LLMs toward AGI
      • The key reason is that the underlying learning mechanism of LLMs is highly analogous to human implicit memory
  • Can AI Agents Synthesize Scientific Conclusions?

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11337v1 公告类型:新。
      • 摘要:科学人工智能代理越来越多地检索证据、跨来源推理,并综合用于后续决策的结论。
      • 然而,他们在健康等高风险领域这样做的能力仍不清楚。
      • 我们引入了 SciConBench,这是一个包含 9.11K 个问题和来自系统评价的专家撰写结论的大规模实时基准,用于评估开放领域的科学结论综合。
    • EN 要点:
      • arXiv:2606.11337v1 Announce Type: new
      • Abstract: Scientific AI agents increasingly retrieve evidence, reason across sources, and synthesize conclusions used in consequential decisions
      • Yet, their ability to do so in high-stakes domains such as health remains unclear
      • We introduce SciConBench, a large-scale live benchmark of 9.11K questions and expert-written conclusions from systematic reviews to evaluate open-domain scienti…
  • Knowing When to Ask: Self-Gated Clarification for Hierarchical Language Agents

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11349v1 公告类型:新。
      • 摘要:在层次推理中,失败通常起源于中间决策点,其中代理提交到错误的分支而没有意识到它缺乏关键信息。
      • 我们不将澄清视为外部不​​确定性触发器,而是提出行动评级,这是一种将其置于智能体行动空间内的公式,与导航共享序数尺度,以便在每个决策点询问与行动直接竞争,并且在中间状态下可以观察到寻求帮助。
      • 代理自身的评级出现了两种结构上不同的信息寻求模式:强制性(没有可行的分支)和机会主义(尽管有领先的候选人,但仍然存在不确定性)。
    • EN 要点:
      • arXiv:2606.11349v1 Announce Type: new
      • Abstract: In hierarchical reasoning, failures often originate at intermediate decision points where the agent commits to a wrong branch without recognizing that…
      • Rather than treating clarification as an external uncertainty trigger, we propose ACTION-RATING, a formulation that places it inside the agent’s action space on…
      • Two structurally distinct information-seeking modes emerge from the agent’s own ratings: mandatory (no viable branch) and opportunistic (residual uncertainty de…
  • Automated Mediator for Human Negotiation: Pre-Mediation via a Structured LLM Pipeline

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11379v1 公告类型:新。
      • 摘要:预调解是直接人工谈判之前的准备阶段,在达成互利协议方面发挥着关键作用,但由于成本、时间和训练有素的调解员的机会有限而常常被省略。
      • 我们引入了用于人工谈判的自动调解员,作为法学硕士模块的结构化管道实施,支持综合谈判设置中的预调解。
      • 该管道将准备工作分解为对话、偏好预测、响应级别批评和结构化摘要的专用模块,将推理、生成和评估分开,以解决单一提示方法的局限性。
    • EN 要点:
      • arXiv:2606.11379v1 Announce Type: new
      • Abstract: Pre-mediation, the preparatory phase preceding direct human negotiation, plays a critical role in achieving mutually beneficial agreements, yet is oft…
      • We introduce an automated mediator for human negotiation, implemented as a structured pipeline of LLM modules, that supports pre-mediation in integrative negoti…
      • The pipeline decomposes preparation into specialized modules for dialogue, preference prediction, response-level critique, and structured summarization, separat…
  • INFRAMIND: Infrastructure-Aware Multi-Agent Orchestration

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11440v1 公告类型:新。
      • 摘要:现有的多代理 LLM 编排方法,从强力集成到学习路由器,根据任务和模型特征选择模型和拓扑。
      • 但是,这些方法不考虑服务基础设施的运行时状态。
      • 在并发负载下的共享 GPU 集群上,这种基础设施盲目性会导致系统资源利用不足:首选模型会积累深度请求队列,而同等能力的替代模型则闲置。
    • EN 要点:
      • arXiv:2606.11440v1 Announce Type: new
      • Abstract: Existing multi-agent LLM orchestration methods, ranging from brute-force ensembles to learned routers, select models and topologies based on task and…
      • However, these methods do not consider the runtime state of the serving infrastructure
      • On shared GPU clusters under concurrent load, this infrastructure blindness causes systematic resource underutilization: preferred models accumulate deep reques…
  • Forecasting Future Behavior as a Learning Task

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11445v1 公告类型:新。
      • 摘要:对人工智能系统的信任通常取决于对其工作原理的解释,然后使用该解释来预测其对新输入的行为。
      • 对于大型推理模型(LRM),这种传统路线特别难以遵循:单个标记生成的解释方法不能自然地推广到长轨迹,并且当作为自然语言阅读时,轨迹本身通常不忠实。
      • 我们提出了一种绕过解释步骤的替代方案:将行为预测视为一项可学习的任务,并训练在单一推理轨迹上运行的行为预测者,以做出人们通常从解释中寻求的相同预测。
    • EN 要点:
      • arXiv:2606.11445v1 Announce Type: new
      • Abstract: Trust in an AI system is often anchored by explanations of how it works, which one then uses to forecast its behavior on new inputs
      • For large reasoning models (LRMs), this conventional route is particularly difficult to follow: explanation methods for single token generations do not naturall…
      • We propose an alternative that bypasses the explanation step: treat behavior forecasting as a learnable task and train Behavior Forecasters that operates on a s…
  • Search Discipline for Long-Horizon Research Agents

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11522v1 公告类型:新。 -摘要:自动研究代理现在根据指标提出、评估和选择科学候选者,而该指标通常是在区域、切片或群组的异构空间上减少的聚合。
      • 我们表明,当科学有效性存在于这种分类结构中时,总体上可能会将错误的候选者排在第一位。
      • 标题数字得到改善,而底层结构却发生倒转,因此对数字做出的决定会接受一个悄悄打破模型的候选者。
    • EN 要点:
      • arXiv:2606.11522v1 Announce Type: new
      • Abstract: Autoresearch agents now propose, evaluate, and select scientific candidates against a metric, and that metric is usually an aggregate reduced over a h…
      • We show that when scientific validity lives in that disaggregated structure, the aggregate can rank the wrong candidate first
      • The headline number improves while the structure underneath inverts, so a decision made on the number accepts a candidate that quietly breaks the model
  • MoCA-Agent: A Market-of-Claims Code Agent for Financial and Numerical Reasoning

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11537v1 公告类型:新。
      • 摘要:财务和表格问题的回答需要的不仅仅是流畅的推理:答案必须基于支持它们的确切事实、公式、单位、符号和尺度。
      • 单个误读单元格或不正确的操作可能会默默地产生看似合理但错误的结果。
      • 我们引入了 \textsc{MOCA-Agent},一种索赔市场代码代理,它用索赔级别验证取代了自由形式的多代理辩论。
    • EN 要点:
      • arXiv:2606.11537v1 Announce Type: new
      • Abstract: Financial and tabular question answering requires more than fluent reasoning: answers must be grounded in the exact facts, formulas, units, signs, and…
      • A single misread cell or incorrect operation can silently produce a plausible but wrong result
      • We introduce \textsc{MOCA-Agent}, a market-of-claims code agent that replaces free-form multi-agent debate with claim-level verification
  • SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11543v1 公告类型:新。
      • 摘要:智能体技能在推理时通过程序知识增强大型语言模型(LLM)智能体,但当前的基准很少区分技能的内容和组织方式。
      • 我们通过渐进式披露来研究这种区别,其中简洁的根文件将代理指向按需支持的资源,并将其与标准化的平坦基线进行比较。
      • 我们提出了 SkillJuror,一个通过语义控制变体、匹配的多试验评估和轨迹证据来评估技能写作范式的框架,同时保持任务知识固定。
    • EN 要点:
      • arXiv:2606.11543v1 Announce Type: new
      • Abstract: Agent Skills augment large language model (LLM) agents with procedural knowledge at inference time, but current benchmarks rarely distinguish what a S…
      • We study this distinction through Progressive Disclosure, where a concise root file points agents to supporting resources on demand, and compare it with a norma…
      • We present SkillJuror, a framework for evaluating Skill writing paradigms through semantically controlled variants, matched multi-trial evaluations, and traject…

ArXiv cs.CL (B_intro+search) 链接到标题

  • PoQ-Judge: A Multi-Architecture Evaluation Framework for Cost-Aware Proof-of-Quality in Decentralized LLM Inference

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2606.11196v1 Announce Type: new.
      • Abstract: Decentralized LLM inference networks need lightweight, reference-free quality evaluation for Proof of Quality (PoQ).
      • We present PoQ-Judge, a framework that trains dedicated judge models to score query-output pairs without ground-truth references.
      • We study three architectures across the quality-cost tradeoff: a TextCNN judge, a MiniLM cross-encoder, and a DeBERTa judge.
    • EN 要点:
      • arXiv:2606.11196v1 Announce Type: new
      • Abstract: Decentralized LLM inference networks need lightweight, reference-free quality evaluation for Proof of Quality (PoQ)
      • We present PoQ-Judge, a framework that trains dedicated judge models to score query-output pairs without ground-truth references
      • We study three architectures across the quality-cost tradeoff: a TextCNN judge, a MiniLM cross-encoder, and a DeBERTa judge
  • The Structural Attention Tax: How Retrieval Format Hijacks In-Context Learning Independent of Content

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11198v1 公告类型:新。
      • 摘要:检索增强生成(RAG)系统注入外部知识来改善 LLM 输出,但注入内容的格式(与其语义相关性不同)可能会独立扭曲模型的注意力分布。
      • 我们识别并形式化了一种我们称之为结构注意力税的现象:知识图谱(KG)三元组,由于它们的关系分隔符和重复的槽模式,每个标记捕获的注意力比语义上等效的自然语言文本高 2-3 倍($\hat{o}$(KG) $\approx$ 0.70 vs.
      • $\hat{o}$(neutral) $\approx$ 0.25),将演示注意力压缩高达 42%——无论三元组是否相关或噪音。
    • EN 要点:
      • arXiv:2606.11198v1 Announce Type: new
      • Abstract: Retrieval-augmented generation (RAG) systems inject external knowledge to improve LLM outputs, yet the format of injected content – distinct from its…
      • We identify and formalise a phenomenon we term the structural attention tax: knowledge graph (KG) triples, due to their relational delimiters and repeated slot…
      • $\hat{o}$(neutral) $\approx$ 0.25), compressing demonstration attention by up to 42% – regardless of whether the triples are relevant or noise
  • NightFeats @ MMU-RAGent NeurIPS 2025: A Context-Optimized Multi-Agent RAG System for the Text-to-Text Track

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11199v1 公告类型:新。
      • 摘要:我们展示了 NightFeats,这是一种结构化多智能体检索增强生成 (RAG) 系统,已提交给 NeurIPS 2025 的 MMU-RAGent 竞赛,并获得了文本到文本赛道的最佳动态评估奖。
      • 这项工作不是以基准最大化为目标,而是提出了一个原则性的管道,将知识合成分解为三个协调的阶段:检索、管理和组合,每个阶段都由显式的中间表示和移交契约控制。
      • 受代理上下文工程(ACE)的启发,该系统引入了时间语义重新排序、有界矛盾调和和引文保留组合作为核心架构原语。
    • EN 要点:
      • arXiv:2606.11199v1 Announce Type: new
      • Abstract: We present NightFeats, a structured multi-agent retrieval-augmented generation (RAG) system submitted to the MMU-RAGent competition at NeurIPS 2025, w…
      • Rather than targeting benchmark maximization, this work proposes a principled pipeline that decomposes knowledge synthesis into three coordinated phases: retrie…
      • Inspired by Agentic Context Engineering (ACE), the system introduces temporal-semantic reranking, bounded contradiction reconciliation, and citation-preserving…
  • Detecting AI-Generated Content on Social Media with Multi-modal Language Models

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11200v1 公告类型:新。
      • 摘要:生成式人工智能能够创建逼真的图像和视频,这些图像和视频越来越多地在社交媒体上传播,通常用于垃圾邮件、错误信息、操纵和欺诈。
      • 现有的人工智能生成内容(AIGC)检测方法面临挑战,包括对新一代模型的泛化能力差、依赖单一模式以及缺乏可解释的解释。
      • 我们提出了通过不断整理不同的多模式社交媒体数据并训练用于检测和解释的紧凑视觉语言模型来缓解这些问题的管道。
    • EN 要点:
      • arXiv:2606.11200v1 Announce Type: new
      • Abstract: Generative AI has enabled the creation of photorealistic images and videos that are increasingly disseminated on social media, often used for spam, mi…
      • Existing AI-generated content (AIGC) detection methods face challenges including poor generalization to new generation models, reliance on single modalities, an…
      • We present our pipeline that mitigates these issues by continuously curating diverse multi-modal social media data and training a compact vision-language model…
  • One Jailbreak, Many Tongues: Learning Language-Insensitive Intention Representations for Multilingual Jailbreak Detection

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11202v1 公告类型:新。
      • 摘要:大语言模型(LLM)越来越多地部署在全球多语言用户的应用中,但安全培训仍然集中在主导语言上,并没有与多语言能力同步发展,为越狱攻击创造了可利用的漏洞。
      • 当前的越狱防御主要是用主流语言开发和评估的,其有效性受到缺乏一致的多语言监督和语言变异造成的表示分散的限制。
      • 为了解决这个问题,我们提出了 MLJailDe,一种多语言越狱检测框架,旨在提高多语言鲁棒性和跨语言泛化能力。
    • EN 要点:
      • arXiv:2606.11202v1 Announce Type: new
      • Abstract: Large language models (LLMs) are increasingly deployed in applications for global multilingual users, yet safety training remains concentrated in domi…
      • Current jailbreak defenses are largely developed and evaluated in dominant languages, and their effectiveness is limited by the scarcity of aligned multilingual…
      • To address this issue, we propose MLJailDe, a multilingual jailbreak detection framework designed to improve both multilingual robustness and cross-lingual gene…
  • LatticeBridge: Rare-Event Sequential Inference for Faithful Structured Sequence Synthesis

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11203v1 公告类型:新。
      • 摘要:结构化序列生成通常需要一个模型在单个输出中满足多个输入导出的约束。
      • 标准解码方法可以为流畅的延续分配高概率,同时将低质量放在共同实现所有所需锚点的延续上。
      • 我们将这种机制作为罕见事件顺序推理问题来研究。
    • EN 要点:
      • arXiv:2606.11203v1 Announce Type: new
      • Abstract: Structured sequence generation often requires a model to satisfy several input-derived constraints in a single output
      • Standard decoding methods may assign high probability to fluent continuations while placing low mass on continuations that realize all required anchors jointly
      • We study this regime as a rare-event sequential inference problem
  • Benchmarking Large Language Models for Safety Data Extraction

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11204v1 公告类型:新。
      • 摘要:由于异构文档格式和传统基于规则的方法的局限性,从安全数据表 (SDS) 中准确提取结构化信息在工业安全领域仍然具有挑战性。
      • 这项研究对用于自动 SDS 数据提取的最先进的大型语言模型 (LLM) 进行了基准测试,比较了基于文本的处理流程和多模式处理流程。
      • 我们系统地评估了四种模型:Gemini 1.5 Pro、GPT-4o、Claude 3.7 Sonnet 和 Llama 3.1-70B,涵盖三种提示策略:零样本、少样本和思维链。
    • EN 要点:
      • arXiv:2606.11204v1 Announce Type: new
      • Abstract: Accurate extraction of structured information from Safety Data Sheets (SDS) remains challenging in industrial safety due to heterogeneous document for…
      • This study benchmarks state-of-the-art Large Language Models (LLMs) for automated SDS data extraction, comparing text-based and multimodal processing pipelines
      • We systematically evaluate four models: Gemini 1.5 Pro, GPT-4o, Claude 3.7 Sonnet, and Llama 3.1-70B, across three prompting strategies: zero-shot, few-shot, an…
  • Compatibility-Aware Dynamic Fine-Tuning for Large Language Models

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11206v1 公告类型:新。
      • 摘要:监督微调(SFT)是对齐大型语言模型(LLM)的主要范例,但它存在优化不稳定和泛化有限的问题。
      • 最近的工作将此问题归因于病态梯度缩放,并提出动态微调(DFT)以在令牌级别纠正它。
      • 然而,DFT 假设所有演示都是同样合适的学习目标,大规模指令数据的强异质性违反了这一假设,其中演示策略不匹配会导致样本级别的高方差更新。
    • EN 要点:
      • arXiv:2606.11206v1 Announce Type: new
      • Abstract: Supervised Fine-Tuning (SFT) is the predominant paradigm for aligning large language models (LLMs), yet it suffers from optimization instability and l…
      • Recent work attributes this issue to pathological gradient scaling and proposes Dynamic Fine-Tuning (DFT) to correct it at the token level
      • However, DFT assumes all demonstrations are equally suitable learning targets, an assumption violated by the strong heterogeneity of large-scale instruction dat…
  • BioDivergence: A Benchmark and Evaluation Framework for Hidden Contextual Contradictions in Biomedical Abstracts

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11208v1 公告类型:新。
      • 摘要:不同研究之间的生物医学发现似乎常常存在冲突,但其中许多差异是与背景相关的,而不是真正的矛盾。
      • 队列、地理位置、检测方案、疾病亚型和临床环境的变化可以使这两种说法在当地有效。
      • 现有的 NLI 和科学主张验证基准将此类情况简化为蕴含、矛盾或中性,未能捕捉分歧背后的背景结构。
    • EN 要点:
      • arXiv:2606.11208v1 Announce Type: new
      • Abstract: Biomedical findings often seem to conflict across studies, but many of these differences are context-dependent rather than true contradictions
      • Variations in cohort, geography, assay protocol, disease subtype, and clinical setting can make both claims locally valid
      • Existing NLI and scientific claim-verification benchmarks reduce such cases to entailment, contradiction, or neutral, failing to capture the contextual structur…
  • ProcessThinker: Enhancing Multi-modal Large Language Models Reasoning via Rollout-based Process Reward

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11209v1 公告类型:新。
      • 摘要:视觉问答越来越需要多步骤推理。
      • 最近在可验证奖励(RLVR)和组相对策略优化(GRPO)下进行强化学习的后训练可以改善多模态推理,但大多数方法依赖于稀疏的仅结果奖励。
      • 因此,他们很难判断错误的答案是来自推理后期的小错误,还是来自从一开始就无益的轨迹。
    • EN 要点:
      • arXiv:2606.11209v1 Announce Type: new
      • Abstract: Visual question answering increasingly requires multi-step reasoning
      • Recent post-training with reinforcement learning under verifiable rewards (RLVR) and Group Relative Policy Optimization (GRPO) can improve multimodal reasoning,…
      • As a result, they struggle to tell whether an incorrect answer comes from a small mistake late in the reasoning or from an unhelpful trajectory from the start

ArXiv cs.LG (B_intro+search) 链接到标题

  • Restless bandits with imperfect binary feedback: PCL-indexability analysis and computation

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11192v1 公告类型:新。
      • 摘要:我们研究具有二元潜态和不完美二元反馈的不安强强盗,其动机是具有传感错误的机会主义频谱访问。
      • 对于相关的信念状态模型,我们开发了一个基于部分守恒定律(PCL)的分析和计算框架,用于建立可索引性和评估 Whittle 指数,建立在真实状态折扣不安强盗的验证定理的基础上。
      • 该框架通过相关的确定性骨架、更新分解和单词组合来分析随机动力学。
    • EN 要点:
      • arXiv:2606.11192v1 Announce Type: new
      • Abstract: We study restless bandits with binary latent states and imperfect binary feedback, motivated by opportunistic spectrum access with sensing errors
      • For the associated belief-state model, we develop a partial conservation laws (PCL)-based analytical and computational framework for establishing indexability a…
      • The framework analyzes the stochastic dynamics via an associated deterministic skeleton, renewal decompositions, and combinatorics on words
  • To Intervene or Not: Guiding Inference-time Alignment with Probabilistic Model Blending

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11201v1 公告类型:新。
      • 摘要:法学硕士的广泛部署使得模型对齐成为必要,以使新训练的模型安全有效地响应用户指令。
      • 在不同的方法中,推理时间对齐通常更便宜,因为它仅在输出生成期间进行干预(即提供指导)。
      • 现有提案应用从某些一致模型中提取的指导,而没有正确评估其可靠性。
    • EN 要点:
      • arXiv:2606.11201v1 Announce Type: new
      • Abstract: The wide deployment of LLMs has made model alignment necessary to make newly trained models safely and effectively respond to user instructions
      • Among different methods, inference-time alignment is often cheaper as it intervenes (i.e., offers guidances) only during output generation
      • Existing proposals apply guidances extracted from certain aligned models without properly assessing their reliability
  • Dual-Stance Evaluation of Sycophancy: The Structure of Agreement and the Limits of Intervention

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11205v1 公告类型:新。 -摘要:激活转向可以改变法学硕士的行为,但标准评估通常不会测试减少阿谀奉承的方向是否也会抑制与事实正确的陈述的一致性。
      • 我们引入了双立场评估,它测试每个主题的两种立场,并将其应用于 Llama-3-8B-Instruct 上的质心差异转向。
      • 我们发现了一种分离:该模型代表了几何上不同的子空间中的阿谀奉承和事实一致,但转向方向平等地投射到两者上,并且不能有区别地瞄准其中任何一个。
    • EN 要点:
      • arXiv:2606.11205v1 Announce Type: new
      • Abstract: Activation steering can shift LLM behaviour, but standard evaluations do not typically test whether a sycophancy-reduction direction also suppresses a…
      • We introduce dual-stance evaluation, which tests both stances of each topic, and apply it to centroid-difference steering on Llama-3-8B-Instruct
      • We find a dissociation: the model represents sycophantic and factual agreement in geometrically distinct subspaces, yet the steering direction projects equally…
  • Few-Shot Resampling for Scalable Statistically-Sound Data Mining

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11235v1 公告类型:新。
      • 摘要:知识发现的关键步骤是对数据挖掘结果的评估。
      • 在多种应用中,包括模式挖掘、图形分析等,此步骤包括评估结果的统计显着性,以避免仅由于数据中的噪声或随机波动而导致虚假发现。
      • 虽然针对某些特定应用开发了专门程序,但基于重采样的方法得到了广泛使用,特别是对于无法得出分析结果的复杂分析。
    • EN 要点:
      • arXiv:2606.11235v1 Announce Type: new
      • Abstract: A key step in knowledge discovery is the evaluation of data mining results
      • In several applications, including pattern mining, graph analysis, and others, this step includes the evaluation of the statistical significance of the results,…
      • While specialized procedures have been developed for some specific applications, resampling-based approaches are widely used, in particular for complex analyses…
  • ProHiFlo: Hierarchical Flow Matching with Functional Guidance for De Novo Protein Generation

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11243v1 公告类型:新。
      • 摘要:从头蛋白质生成在治疗设计、酶工程和合成生物学方面具有变革潜力。
      • 虽然基于扩散和流量匹配的方法已经取得了进展,但它们通常以单一分辨率运行,并且缺乏合并功能约束的机制。
      • 我们引入了 ProHiFlo,一种具有三项创新的分层流匹配框架:(1)从粗到细的生成,在细化到全原子坐标之前对骨干几何结构进行建模,从而在保持准确性的同时降低计算成本; (2) 功能指导利用预训练的预测器来引导一代人实现所需的属性,而无需重新训练; (3) 用于高效多尺度处理的自适应SE(3)等变架构。
    • EN 要点:
      • arXiv:2606.11243v1 Announce Type: new
      • Abstract: De novo protein generation has transformative potential in therapeutic design, enzyme engineering, and synthetic biology
      • While diffusion-based and flow matching approaches have achieved progress, they typically operate at single resolution and lack mechanisms for incorporating fun…
      • We introduce ProHiFlo, a hierarchical flow matching framework with three innovations: (1) coarse-to-fine generation that models backbone geometry before refinin…
  • Physics-informed generative AI for semiconductor manufacturing: Enforcing hard physical constraints in generative models by construction

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11247v1 公告类型:新。
      • 摘要:生成模型越来越多地用于提出物理系统的设计、数据和控制操作,但许多此类系统受到严格的物理约束而不是感知合理性的控制。
      • 半导体制造提供了严格的测试用例:生成的掩模、布局、合成缺陷数据和工艺配方必须遵守光刻、传输、反应和设备物理约束,因为物理无效的样品不仅质量低而且无法使用。
      • 本观点认为,半导体制造面临着更广泛的计算科学挑战,即受限物理领域的生成式人工智能必须通过构造获得物理信息,而不仅仅是通过事后过滤进行纠正。
    • EN 要点:
      • arXiv:2606.11247v1 Announce Type: new
      • Abstract: Generative models are increasingly used to propose designs, data, and control actions for physical systems, yet many such systems are governed by hard…
      • Semiconductor manufacturing provides a demanding test case: generated masks, layouts, synthetic defect data, and process recipes must obey lithography, transpor…
      • This Perspective argues that semiconductor manufacturing exposes a broader computational-science challenge, namely that generative AI for constrained physical d…
  • Mechanical Field Networks: Structured Neural Dynamics for Multivariate Systems

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11251v1 公告类型:新。
      • 摘要:许多多元动力系统只能通过轨迹来观察,从而隐藏了控制其联合动力学的机制。
      • 现有方法可以强加可解释的动态或学习灵活的状态转换,但所得到的交互结构通常要么提前指定,要么隐含在学习的动态中。
      • 我们引入了 MF-Net,这是一种循环动态模型,它表示共享场状态中的所有变量,并通过学习的关系律更新该状态。
    • EN 要点:
      • arXiv:2606.11251v1 Announce Type: new
      • Abstract: Many multivariate dynamical systems are observed only through trajectories, leaving the mechanisms governing their joint dynamics hidden
      • Existing approaches can impose interpretable dynamics or learn flexible state transitions, yet the resulting interaction structure is typically either specified…
      • We introduce MF-Net, a recurrent dynamical model that represents all variables in a shared field state and updates this state through a learned relation law
  • Bernstein-Schur Kernels: Random Features by Sketched Modulation and Radial Randomization

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11255v1 公告类型:新。
      • 摘要:Bernstein–Schur 核是有限特征核(具有显式有限维特征图)和完全单调平移不变核的乘积:介于平移不变和点积模板之间的非平稳核通常利用随机特征,因此一般来说,Bochner 采样和多项式草图都不能直接应用于完整核。
      • 我们为整个类提供一个随机特征构造,\emph{随机化两个因子:它勾勒出有限调制并随机化完全单调的径向因子,对后者的一维 Bernstein-Widder 尺度进行采样,然后应用高斯随机傅里叶特征(其频率仍然是 $d$ 维)。
      • 特征尺寸为 $Dm$,由草图尺寸 $m$ 和径向绘制计数 $D$ 设置,不受精确调制特征的 $O(d^2)$ 尺寸影响。
    • EN 要点:
      • arXiv:2606.11255v1 Announce Type: new
      • Abstract: Bernstein–Schur kernels are products of a finite-feature kernel (one with an explicit finite-dimensional feature map) and a completely monotone shift…
      • We give one random-feature construction for the whole class that \emph{randomizes both factors: it sketches the finite modulation and randomizes the completely…
      • The feature dimension is then $Dm$, set by the sketch size $m$ and the radial-draw count $D$, free of the $O(d^2)$ size of the exact modulation feature
  • Loss Landscape Diagnosis for Gradient-Based Gray-Scott System Inversion: Disentangling the Roles of PINN Components

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11258v1 公告类型:新。
      • 摘要:基于梯度的反应扩散系统反演通常通过代理模型或物理信息神经网络 (PINN) 来实现,而最直接的途径,即通过 PDE 结构本身的反向传播,在很大程度上被避免了。
      • 我们将这种直接途径作为诊断探针,通过展开的格雷-斯科特模拟反向传播稳态损失以恢复其参数,无需替代或神经网络增强。
      • 优化无法收敛,绘制地形图可以直接定位其几何形状中的故障——没有梯度信号的平坦高原,以与分叉边界对齐的陡峭悬崖为界——这种结构在损失函数中重复出现,并且是继承的,但梯度被路由到参数。
    • EN 要点:
      • arXiv:2606.11258v1 Announce Type: new
      • Abstract: Gradient-based inversion of reaction-diffusion systems is typically approached via surrogate models or physics-informed neural networks (PINNs), while…
      • We pursue this direct route as a diagnostic probe, backpropagating a steady-state loss through unrolled Gray-Scott simulation to recover its parameters, with no…
      • Optimization fails to converge, and plotting the landscape directly locates the failure in its geometry – flat plateaus with no gradient signal, bounded by sha…
  • PermDoRA – Understanding Adapter Interference in Language Models: Limits of Parameter-Space Geometry

    • 发布时间:2026-06-11 12:00 北京时间
    • 摘要:- arXiv:2606.11262v1 公告类型:新。
      • 摘要:大型语言模型(LLM)中的访问控制需要模块化机制来实现特定于域的行为,而无需重新训练或跨域干扰。
      • 一个常见的假设是,适配器组合过程中的干扰是由线性参数更新的重叠引起的,这表明强制正交性或方向独立性应该可以提高多域性能。
      • 我们使用 DoRA-RBAC 来测试这个假设,DoRA-RBAC 是一种基于权重分解的低秩适应的分层适配器组合框架。
    • EN 要点:
      • arXiv:2606.11262v1 Announce Type: new
      • Abstract: Access control in large language models (LLMs) requires modular mechanisms to enable domain-specific behavior without retraining or cross-domain inter…
      • A common hypothesis is that interference during adapter composition arises from overlap in linear parameter updates, suggesting that enforcing orthogonality or…
      • We test this hypothesis using DoRA-RBAC, a hierarchical adapter composition framework based on weight-decomposed low-rank adaptation