🤖 AI 速览

今日重点从模型能力转向行为边界与工程交付。多篇研究指向对齐受数据、人格设定和训练后流程共同影响;智能体评测开始关注效率、可靠性和责任链。企业侧,常驻协作 Agent、开源代码模型与 AI 可观测性同步推进,但隐私、权限和自动执行风险仍需审慎处理。
📋 文章元数据
发布时间
2026-06-26
类型
ai-daily
字数
3180
阅读时长
15 min

2026-06-26 AI日更 | 模型可控性升温,Agent 从演示走向可验证交付 链接到标题

今日重点从模型能力转向行为边界与工程交付。多篇研究指向对齐受数据、人格设定和训练后流程共同影响;智能体评测开始关注效率、可靠性和责任链。企业侧,常驻协作 Agent、开源代码模型与 AI 可观测性同步推进,但隐私、权限和自动执行风险仍需审慎处理。

📖 本期 Watch List 深度导读 链接到标题

今天最值得关注的是“模型行为可控性”的一组论文:从 sycophancy 的级联线性特征,到 persona 如何影响 refusal,再到训练后流程可能削弱同情价值,几篇研究共同指向一个问题——对齐并不是单点开关,而是由数据、角色设定和下游行为共同塑造,值得安全与评测团队细读。

第二条主线是“智能体评测与治理”。CORE-Bench 论文提醒我们,基准饱和后不应只追逐更难题,而要看效率、可靠性、人机协作等维度;Coding Agent rewards 的验证困境,以及“治理行动而非治理智能体”的制度证明模型,也都在把焦点从能力展示转向可验证、可追责的部署。

最后,应用侧可以关注知识增强智能体在精神健康用药信息中的探索,以及 LLM 驱动算法交易、DAO/企业 AI 协议治理分析等方向。它们显示智能体正在进入高风险、高复杂度场景,但真正的门槛仍是证据边界、责任结构与持续验证。

🌐 X 平台 AI 热点快讯 链接到标题

话题 1:Anthropic Launches Claude Tag for Persistent Slack AI Teammate 链接到标题

  • 分类:AI · News
  • 概况:热度时间:2 days ago,相关帖子数:53000
  • 是什么事:Anthropic 推出 Claude Tag,将 Claude 作为可在 Slack 中持续在线、记忆上下文并协助工作流的 AI 队友,面向 Enterprise/Team 计划替代旧版 Slack 应用。
  • 为什么重要:这标志着企业 AI 正从被动问答工具转向具备长期记忆、环境感知和自主执行能力的协作型智能体,可能改变团队知识管理、软件开发和日常协作方式。
  • 讨论概况:X 上讨论集中在 Claude Tag 能否真正减少上下文丢失、把企业隐性知识产品化,以及哪些工作流适合交给常驻 AI;分歧则在于其自主监控 Slack 消息带来的效率提升与隐私、权限和可靠性风险之间的平衡。

话题 2:GLM-5.2 Emerges as Top Open-Weight Coding Model 链接到标题

  • 分类:AI · News
  • 概况:热度时间:22 hours ago,相关帖子数:529
  • 是什么事:智谱 AI 的 GLM-5.2 被 X 用户热议为当前表现领先的开源权重代码模型之一。
  • 为什么重要:这显示开源权重模型在代码生成与软件工程任务上继续逼近前沿闭源模型,可能降低企业和开发者使用高性能编程 AI 的成本,并加速本地化、可控部署。
  • 讨论概况:X 上的讨论集中在 GLM-5.2 的真实基准表现、与 Google、MiniMax 等模型的性价比对比,以及开源权重模型是否会在 2026 年进一步改变软件开发流程;分歧主要在于其能力是否已达到“前沿级”,以及评测结果能否转化为稳定的实际开发效率。

话题 3:Sazabi Raises $8M to Build Self-Healing AI Observability Platform 链接到标题

  • 分类:AI · News
  • 概况:热度时间:23 hours ago,相关帖子数:1400
  • 是什么事:AI 可观测性初创公司 Sazabi 完成 800 万美元融资,计划打造具备“自愈”能力的 AI 系统监控与运维平台。
  • 为什么重要:随着企业部署更多大模型应用,模型异常、性能退化、成本失控和安全风险更难被及时发现,AI 可观测性与自动修复能力正成为生产级 AI 基础设施的重要环节。
  • 讨论概况:X 上的讨论主要集中在“自愈式”AI 运维是否能真正降低人工排障成本,以及该赛道是否会成为继模型开发平台之后的新基础设施热点;也有人质疑早期产品能力是否被融资叙事放大。

今日 X 上的 AI 舆情小结 链接到标题

今天的舆论主线是,AI 正从“模型能力竞赛”进一步走向“生产环境落地”:一边是 Claude Tag 这类常驻企业协作智能体进入 Slack 工作流,另一边是 GLM-5.2 等开源权重代码模型拉低高性能编程 AI 的使用门槛,同时 Sazabi 代表的可观测性与自愈运维开始补齐企业部署后的基础设施短板。共识在于,AI 已不只是聊天或单点生成工具,而是在向长期记忆、上下文感知、自主执行、本地可控部署和自动化运维扩展。分歧主要集中在“看起来领先”能否稳定转化为真实生产效率:Claude Tag 是否值得用隐私和权限风险换取协作效率,GLM-5.2 是否真正达到前沿级代码能力,Sazabi 这类自愈运维是否已有足够产品成熟度。潜在风险则是企业过快把关键知识流、代码流程和运维判断交给 AI,在权限边界、数据泄露、错误自动执行、评测泡沫和成本失控上暴露新的系统性问题。

💡 大佬观点(Influencer Insights) 链接到标题

好的,作为资深 AI 行业分析师,我已为您梳理了这份基于过去 24 小时 X 平台大佬观点的行业情报。以下是核心摘要:


1. 今日大佬们共同关注的技术趋势或产品热点 链接到标题

核心趋势:Agent 化的“操作系统”之争与工程化落地

  • 焦点产品:Codex 的生态位与瓶颈

    • 定位升维:业界已不满足于将 Codex 视为编程工具。@dotey 明确指出,Codex 的发展趋势是成为“Agent OS”,而不仅是 “Agent Office”。他分享了反编译和复刻 Codex 代码项目的经验,揭示了其底层逻辑正被社区深度研究。
    • 用量与成本焦虑:多位博主(@Pluvio9yte, @ruanyf)反馈 Codex 的 Token 消耗量激增或限制变严,甚至戏称为“百亿补贴缩水”。同时,@Pluvio9yte@vista8 分享了在 Codex 中组合使用不同模型(如用 Gemini 聊天、Claude 做规划)的混合工作流,以应对单一模型的局限性。
  • 模型竞技场:端侧、Coding 与多模态的角力

    • 端侧模型成为红海@zhixianio 持续深入测试端侧模型,从 Google 的 Gemma 系列(4 12B Coder、E4B + MTP)到 MiniCPM-o 4.5 的音视频全双工能力,均给予高度评价,认为其“可以用起来了”。@googledevs 发布的 QAT(量化感知训练) 模型被认为是关键的优化思路。
    • Code Generation 能力实测@zhixianioGemma 4 12B CoderQwen3.6-35B-A3B 进行了深度对比评测,指出 12B 模型在处理复杂、有状态的程序(如俄罗斯方块)时存在天花板,而 35B MoE 依然是“甜点”。@Pluvio9yte 则分享了将豆包 Seed 2.1 Pro 等性价比模型接入 Claude Code 的教程,预示模型供应市场的多元化和价格战。
    • AI 视频生成工具轻量化Topview AISeedance 2.0 Mini 上线,主打“快和便宜”,@AI_Jasonyu 认为这对日常 AI 漫剧创作非常友好,反映出 AI 视频领域的竞争已从单纯画质转向性价比与速度。

2. 值得注意的独特观点或行业前瞻 链接到标题

  • AI 领域的“新冷战”:蒸馏攻击与政府管制

    • @dotey 深度解读了 Anthropic 指控阿里巴巴大规模蒸馏 Claude 的事件,指出其攻击规模已超过之前所有中国公司的总和。他剖析了 Anthropic 在向政府求助与抗议政府限制模型发布之间的“拧巴”处境,并点明这背后是中美 AI 能力与商业利益的直接碰撞。
    • 同时,他报道了 OpenAI 的 GPT-5.6 发布因政府要求而改为“逐个客户审批” 的史无前例的流程,警告这可能导致公司内部能力与公众可用能力之间的差距越拉越大。
  • 从 Vibe Coding 到 Agent 交付的思考转变

    • “Vibe Coding”之后是“工程化”@gefei55 提出 Token 是无限的,但时间和精力有限,警告大家不要迷失在“什么都能做的 Token 陷阱”里。@Pluvio9yte 也点出 AI 时代重复三遍以上的事必须自动化,并分享了使用“Skill 洁癖 skill”清理过时脚本的案例。
    • 从 Demo 到产品的“最后一公里”@AI_Jasonyu@vista8 不约而同地推广了 EdgeOne Makers 这样的 Agent 部署平台,强调其解决了从本地跑通到上线面临的并发、沙箱安全、记忆存储等痛点。这标志着行业焦点正从“能不能跑通”转向“能不能交付给真实用户”。
  • 关于 AI 创作与知识管理的哲学

    • 语言风格的“模型烙印”@lijigang 提出了一个深刻观察:重度使用某款模型会沾染其语言风格,产生 “Claude 味儿” 或 “Deepseek 味儿”,并指出大脑的神经网络非常“吃” Context。
    • 知识管理的“可机读”格式@vista8 介绍了 谷歌的 Open Knowledge Format (OKF),其核心是用 Markdown + YAML frontmatter 将知识打包成可被 Agent 直接消费、可版本控制的文件包。这预示着为 AI 优化知识结构将成为个人和组织的新技能。

3. 推荐的工具或资源 链接到标题

类别工具/资源核心用途与亮点推荐来源
Agent/Skill 开发与管理Skill 软链接管理方案通过软链接统一管理多个项目中的 Skill,实现一次更新、全局同步。@dotey
EdgeOne Makers一站式部署 AI Agent,解决并发、沙箱、Token 管理、监控等线上问题,提供免费额度。@AI_Jasonyu, @vista8
Skill 洁癖 Skill扫描已安装的 Skill,统计使用频率,提供清理建议。@Pluvio9yte
实用工具VoxCPM2GitHub 开源黑科技(22.9k Star),用自然语言描述即可生成声音。@AI_Jasonyu
视频制作 Skills 仓库开源的 AI 视频制作 Skills,可复刻特定风格的视频(如打字样式)。@Pluvio9yte
反编译 Codex 项目 (decode-codex)用于学习,包含解包 App 代码和反混淆 JS 的 Skills。@dotey
内容与知识管理YouMind已发布 1.0 正式版,被多位博主推荐为适合图文内容创作与排版的工具。@lifesinger, @AI_Jasonyu
Google Open Knowledge Format (OKF)谷歌推出的规范,让 AI 整理的知识变成可读、可版本控制、可被 Agent 直接消费的文件夹。@vista8
API 与基础设施Twitter API(便宜方案)用于辅助监控 X 平台,成本极低(194次调用不到0.43美元)。@gefei55
豆包 Seed 2.1 Pro 接入教程提供在 Claude Code 等 Agent 中配置使用性价比高的豆包模型的方法。@Pluvio9yte

📚 附录:今日 Watch List 更新源列表 链接到标题

时间窗口:最近 3 天;覆盖 22 个源;共 32 条更新

Stratechery by Ben Thompson (A_full) 链接到标题

  • An Interview with Figma CEO Dylan Field About Design and AI
    • 发布时间:2026-06-25 18:00 北京时间
    • 摘要:- Field 是 Thiel Fellow,于 2012 年从布朗大学退学并创办 Figma。
      • Figma 走过了一条令人着迷的道路:该公司于 2022 年接受了 Adobe 的收购要约,但由于监管阻力,后者被迫于 2023 年底放弃合并。
      • 我与 Field 谈论了所有这些,包括他的背景、Figma 的差异化发现过程,以及创造力与设计的本质。
      • 我们讨论人工智能问题,市场将其视为逆风,但菲尔德将其视为顺风。
      • 提醒一下,所有 Stratechery 内容(包括采访)均以播客形式提供;单击此电子邮件顶部的链接将 Stratechery 添加到您的播客播放器。
    • EN 要点:
      • Listen to this post:
      • Good morning,
      • This week’s Stratechery interview is with Figma co-founder and CEO Dylan Field
      • Field was a Thiel Fellow who dropped out of Brown in 2012 to start Figma

OpenAI Blog (A_full) 链接到标题

  • How agents are transforming work
    • 发布时间:2026-06-25 10:00 北京时间
    • 摘要:- 代理人工智能将知识工作单元从单一交互转变为委派的长期任务。
      • 聊天机器人交互通常简短且独立。
      • 代理可以独立操作几分钟或几小时,同时编排工具调用、与环境交互以及迭代解决方案。
      • 因此,代理正迅速成为最强大的人工智能工作工具。
      • 去年,我们在 OpenAI 上亲眼目睹了这一转变。
    • EN 要点:
      • A new OpenAI research paper shows how AI agents are transforming work, enabling longer, more complex tasks and expanding productivity across roles.

ArXiv cs.AI (B_intro+search) 链接到标题

  • Detecting and Controlling Sycophancy with Cascading Linear Features

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26155v1 公告类型:新。
      • 摘要:通过激活引导方法解释和控制模型行为需要许多对对比样本,这些样本清楚地表现出所需或不需要的行为。
      • 这些数据对决定了可解释性框架能够可靠地检测导致行为的模型特征的程度,从而决定引导模型走向或远离这种行为的能力。
      • 在这项工作中,我们提出了一个迭代数据生成管道,它隔离了负责行为的级联线性特征。
    • EN 要点:
      • arXiv:2606.26155v1 Announce Type: new
      • Abstract: Interpreting and controlling model behaviors through activation steering methods requires many pairs of contrastive samples that clearly exhibit desir…
      • These data pairs determine the degree to which interpretability frameworks can reliably detect model features responsible for a behavior, and therefore the abil…
      • In this work, we present an iterative data generation pipeline that isolates cascading linear features responsible for a behavior
  • Life After Benchmark Saturation: A Case Study of CORE-Bench

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26158v1 公告类型:新。
      • 摘要:当基准测试的准确性饱和时,它通常会被淘汰并被更具挑战性的版本所取代。
      • 我们表明,这种方法优先考虑准确性,并错过了研究智能体性能的其他六个关键维度的机会:构造有效性问题,例如捷径、分布外泛化性、效率、可靠性、模型与支架的相对重要性以及人类与智能体协作的提升。
      • 我们使用 CORE-Bench Hard(科学代码的计算再现性基准)作为案例研究,以证明即使在准确性饱和后,沿着这些维度测量代理也可以产生对代理性能的有意义的见解。
    • EN 要点:
      • arXiv:2606.26158v1 Announce Type: new
      • Abstract: When a benchmark’s accuracy saturates, it is often retired and replaced with a more challenging version
      • We show that this approach privileges accuracy and misses the opportunity to study six other key dimensions of agent performance: construct validity issues such…
      • We use CORE-Bench Hard, a benchmark for computational reproducibility of scientific code, as a case study to demonstrate that measuring agents along these dimen…
  • Refusal Lives Downstream of Persona in Chat Models

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26161v1 公告类型:新。 -摘要:在指令调整的聊天模型中,拒绝和角色特征的激活空间中的线性方向已被确定,但两者已作为单独的机制进行研究。
      • 我们展示了他们的互动:顺从的角色会拒绝。
      • 在Qwen2.5-7B-Instruct和Llama-3.1-8B-Instruct中,我们提取顺从模型角色方向和拒绝方向并对两者进行干预。
    • EN 要点:
      • arXiv:2606.26161v1 Announce Type: new
      • Abstract: Linear directions in activation space have been identified for both refusal and persona traits in instruction-tuned chat models, but the two have been…
      • We show they interact: a compliant persona gates refusal
      • In Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct, we extract a compliant model-persona direction and a refusal direction and intervene on both
  • AlgoEvolve: LLM-driven Meta-evolution of Algorithmic Trading Programs

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26173v1 公告类型:新。
      • 摘要:最近的工作表明,大型语言模型(LLM)可以充当语​​义变异算子,用于程序和证明的进化发现。
      • 当前大多数应用程序都专注于静态编码基准。
      • 我们将这种范式扩展到算法交易。
    • EN 要点:
      • arXiv:2606.26173v1 Announce Type: new
      • Abstract: Recent work shows that Large Language Models (LLMs) can act as semantic mutation operators for the evolutionary discovery of programs and proofs
      • Most current applications focus on static coding benchmarks
      • We extend this paradigm to algorithmic trading
  • Agentic Analysis for Agentic Infrastructure: An LLM-Powered Pipeline for Comparative Governance of DAO and Corporate AI Protocols

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26203v1 公告类型:新。
      • 摘要:随着人工智能代理协议的激增,塑造其互操作性标准的治理结构仍然没有得到充分的实证检验。
      • 我们引入了一个由法学硕士支持的用于大规模治理话语分析的比较管道,集成了自动注释、神经主题建模和多层网络分析来大规模研究社会技术权力结构。
      • 我们根据代理互操作性的两个对比标准对其进行验证:ERC-8004(无需许可,链上)和 Google A2A(企业主导)。
    • EN 要点:
      • arXiv:2606.26203v1 Announce Type: new
      • Abstract: As AI agent protocols proliferate, the governance structures shaping their interoperability standards remain empirically underexamined
      • We introduce an LLM-powered comparative pipeline for large-scale governance discourse analysis, integrating automated annotation, neural topic modeling, and mul…
      • We validate it on two contrasting standards for agent interoperability: ERC-8004 (permissionless, on-chain) and Google A2A (corporate-led)
  • Knowledge-augmented Agentic AI for Mental Health Medication Information Seeking

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26205v1 公告类型:新。 -摘要:患者越来越多地在网上寻求药物信息,但精神科药物的安全知识分为权威但抽象的监管不良事件记录和接近经验但未经验证的患者叙述。
      • 在不将证据和轶事混为一谈的情况下将它们整合起来,这在精神病学中尤其重要,因为情境化不佳的信息可能会放大恐惧、反安慰剂反应和不依从性。
      • 在这里,我们开发了一个具有来源感知、基于知识图的多代理框架,统一了 466,525 个 Reddit 帖子、60,782 个 WebMD 评论和 20 年的美国历史。
    • EN 要点:
      • arXiv:2606.26205v1 Announce Type: new
      • Abstract: Patients increasingly seek medication information online, yet safety knowledge for psychiatric drugs is split between regulatory adverse-event records…
      • Integrating them without conflating evidence and anecdote is especially consequential in psychiatry, where poorly contextualised information can amplify fear, n…
      • Here we develop a provenance-aware, knowledge-graph-based multi-agent framework unifying 466,525 Reddit posts, 60,782 WebMD reviews, and twenty years of U.S
  • Accelerating Skill Assessment in Chess: A Drift-Diffusion-Enhanced Elo Rating System

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26267v1 公告类型:新。
      • 摘要:Elo 等评级系统是国际象棋竞技匹配的黄金标准。
      • 然而,由于他们完全依赖比赛结果,而忽视了游戏玩法的精细质量,因此他们本质上会遭受响应滞后的困扰。
      • 尽管如此,考虑到游戏状态空间的巨大噪音和广阔性,将逐步信息纳入评级调整中提出了重大挑战。
    • EN 要点:
      • arXiv:2606.26267v1 Announce Type: new
      • Abstract: Rating systems such as Elo serve as the gold standard for matchmaking in competitive chess
      • However, they inherently suffer from response lag due to their exclusive reliance on match outcomes, neglecting the granular quality of gameplay
      • Nevertheless, incorporating move-by-move information into rating adjustments presents a significant challenge given the substantial noise and the vastness of th…
  • Governing Actions, Not Agents: Institutional Attestation as a Governance Model for Autonomous AI Systems

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26298v1 公告类型:新。
      • 摘要:自主人工智能代理可能会开始执行相应的、不可逆转的行动,例如临床处方和生产软件部署。
      • 本文观察到,人类机构不是通过监控他们的推理而是通过在采取相应行动时要求独立证明的证据来管理强大的自主行为者。
      • 我们将这种制度模式正式化为人工智能代理系统的计算治理模型。
    • EN 要点:
      • arXiv:2606.26298v1 Announce Type: new
      • Abstract: Autonomous AI agents may begin to perform consequential, irreversible actions such as clinical prescribing and production software deployment
      • This paper observes that human institutions have governed powerful autonomous actors not by monitoring their reasoning but by requiring independently attested e…
      • We formalise this institutional pattern as a computational governance model for AI agent systems
  • COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origami

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26299v1 公告类型:新。
      • 摘要:虽然生成式人工智能在通过可验证的解决方案解决问题方面取得了显着的成功,但生成既满足严格的几何约束又满足主观视觉美学的物理艺术仍然是一个挑战。
      • 本文提出了一种在计算折纸领域解决这些困难的方法,计算折纸是一种严格的数学环境,将艺术设计建立在平面可折叠性方程的基础上。
      • 我们推出了 COrigami,这是一种端到端人工智能驱动的管道,它通过从自然语言生成折痕图案来协助设计周期。
    • EN 要点:
      • arXiv:2606.26299v1 Announce Type: new
      • Abstract: While generative AI has achieved remarkable success in solving problems with verifiable solutions, generating physical art that satisfies both strict…
      • This paper presents an approach to tackle these difficulties in the domain of computational origami, a mathematically rigid environment that grounds artistic de…
      • We present COrigami, an end-to-end AI-driven pipeline that assists the design cycle by generating crease patterns from natural language
  • The Verification Horizon: No Silver Bullet for Coding Agent Rewards

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26300v1 公告类型:新。
      • 摘要:一种经典的直觉认为验证解决方案比生成解决方案更容易。
      • 对于当今的编码代理来说,这种直觉正在被逆转:随着基础模型发展出更强大的推理能力并且工程工具变得更加复杂,生成复杂的候选解决方案不再困难 - 可靠地验证它们已成为更难的问题。
      • 我们可以构建的每个验证器都只是人类意图的代理,而不是意图本身。
    • EN 要点:
      • arXiv:2606.26300v1 Announce Type: new
      • Abstract: A classical intuition holds that verifying a solution is easier than producing one
      • For today’s coding agents, this intuition is being inverted: as foundation models develop stronger reasoning capabilities and engineering harnesses grow more so…
      • Every verifier we can build is only a proxy for human intent, never the intent itself

ArXiv cs.CL (B_intro+search) 链接到标题

  • HierBias: Context-Conditioned Hierarchical Media Bias Detection with Multi-Task Type Classification

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26100v1 公告类型:新。 -摘要:媒体偏见检测是确保公平和平衡的信息传播的关键任务,但现有的句子级方法独立地对每个句子进行分类,忽略了人类注释者自然利用的句子间上下文信号。
      • 我们提出了 \textbf{HierBias},一种分层上下文条件媒体偏差检测器,它在偏差预测中对文档上下文进行正式建模。
      • 我们引入了\emph{上下文条件偏差概率},并从理论上证明,当句子间互信息非零时,利用文档上下文严格减少句子级分类的贝叶斯误差。
    • EN 要点:
      • arXiv:2606.26100v1 Announce Type: new
      • Abstract: Media bias detection is a critical task for ensuring fair and balanced information dissemination, yet existing sentence-level approaches classify each…
      • We present \textbf{HierBias}, a hierarchical context-conditioned media bias detector that formally models document context in bias prediction
      • We introduce the \emph{context-conditioned bias probability} and prove theoretically that leveraging document context strictly reduces the Bayes error of senten…
  • Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26101v1 公告类型:新。 -摘要:对大型语言模型的可靠评估应该将支持的答案与不支持的猜测分开,而不将其与数据污染、提示特质或一般拒绝行为混为一谈。
      • 我们提出了一个污染感知的多区域基准,用于衡量在冻结构建时间标签下从可回答的知识到弃权预期未知的转变。
      • 该基准包含跨五个领域的 1,200 个项目、明确的弃权期望、污染风险元数据以及使用官方严格解析器和标准化稳健性解析器的双重解析。
    • EN 要点:
      • arXiv:2606.26101v1 Announce Type: new
      • Abstract: Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contami…
      • We present a contamination-aware, multi-zone benchmark for measuring the transition from answerable knowledge to abstention-expected unknowns under frozen build…
      • The benchmark contains 1,200 items across five domains, explicit abstention expectations, contamination-risk metadata, and dual parsing with an official strict…
  • Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26102v1 公告类型:新。
      • 摘要:标准的训练后流程应用监督微调(SFT)和强化学习(RL)来使语言模型变得有用,但这些过程可能会无意中降低预训练期间灌输的值。
      • 我们使用 SFT(通过 Dolly-15k 提供帮助与通过 Dolly-15k 提供帮助)研究训练后数据领域是否对 Llama 3.1 8B 模型中动物同情心值的保留产生差异影响,该模型在以同情心为导向的合成数据上进行了中期训练。
      • 通过 Magicoder-110K 进行编码)和 GRPO(通过 RLHFlow 与 GRPO 提供帮助)
    • EN 要点:
      • arXiv:2606.26102v1 Announce Type: new
      • Abstract: Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but these process…
      • We investigate whether the domain of post-training data differentially affects the retention of animal compassion values in a Llama 3.1 8B model mid-trained on…
      • coding via Magicoder-110K) and GRPO (helpfulness via RLHFlow vs
  • Investigating LLM’s Problem Solving Capability – a Study on Statics Questions

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26103v1 公告类型:新。
      • 摘要:大型语言模型(LLM)迅速影响了社会的许多方面,特别是教育,因为它们具有完成广泛学科的作业和考试的能力。
      • 尽管之前的研究已经考察了法学硕士的教育影响,但现有的大部分工作依赖于公共或开放问题数据集,并且缺乏针对特定主题的分析。
      • 在工程教育中,特别是在机械工程领域,对特定问题类型的法学硕士表现的系统研究仍然有限。
    • EN 要点:
      • arXiv:2606.26103v1 Announce Type: new
      • Abstract: Large Language Models (LLMs) have rapidly influenced many aspects of society, particularly education, due to their demonstrated ability to complete as…
      • Although prior studies have examined the educational impact of LLMs, much of the existing work relies on public or open problem datasets and lacks topic-specifi…
      • In engineering education, especially within mechanical engineering, systematic investigations of LLM performance on specific problem types remain limited
  • Assert, don’t describe: Linguistic features that shift LLM reasoning about animal welfare

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26104v1 公告类型:新。
      • 摘要:动物福利倡导者撰写了大量的文章,并且越来越多的文章训练了数百万人随后询问动物福利的语言模型。
      • 在保留的动物福利基准上使用词汇匹配的立场对比探针,我们测量十个语言特征中的每一个在用作微调数据时如何改变 Llama-3.2-1B 对支持动物福利推理的偏好。
      • 十个特征中的八个产生统计上显着的变化。
    • EN 要点:
      • arXiv:2606.26104v1 Announce Type: new
      • Abstract: Animal-welfare advocates produce a lot of writing, and increasingly that writing trains the language models that millions of people then ask about ani…
      • Using vocabulary-matched stance-contrast probes on a held-out animal-welfare benchmark, we measure how each of ten linguistic features changes Llama-3.2-1B’s pr…
      • Eight of the ten features produce statistically significant shifts
  • Context Recycling for Long-Horizon LLM Inference

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26105v1 公告类型:新。 -摘要:大型语言模型(LLM)在短上下文推理中表现出强大的能力,但由于上下文窗口限制和低效的令牌使用,在长对话范围内性能下降。
      • 我们引入了 ContextForge,这是一个上下文回收系统,通过结合结构化查询生成、外部存储器检索和受控合成来跨轮维护任务相关信息。
      • 该系统可以高效地重用先前的计算,而无需依赖完整的上下文重放,从而减少令牌开销,同时保持答案质量。
    • EN 要点:
      • arXiv:2606.26105v1 Announce Type: new
      • Abstract: Large language models (LLMs) exhibit strong capabilities in short-context reasoning but degrade in performance over long conversational horizons due t…
      • We introduce ContextForge, a system for context recycling that maintains task-relevant information across turns by combining structured query generation, extern…
      • The system enables efficient reuse of prior computation without relying on full context replay, reducing token overhead while preserving answer quality
  • Reducing Conversational Escalation in Large Language Model Dialogue with Nonviolent Communication Constraints

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26106v1 公告类型:新。
      • 摘要:大型语言模型 (LLM) 越来越多地用于涉及人际冲突、沮丧和痛苦的情绪激动的情况。
      • 虽然之前的安全研究侧重于防止有毒或违反政策的内容等明显伤害,但很少关注可能无意中加剧冲突的对话行为。
      • 在本文中,我们研究是否可以通过源自非暴力沟通(NVC)的轻量级提示级别约束来引导法学硕士走向更加缓和的对话行为。
    • EN 要点:
      • arXiv:2606.26106v1 Announce Type: new
      • Abstract: Large language models (LLMs) are increasingly used in emotionally charged situations involving interpersonal conflict, frustration, and distress
      • While prior safety research has focused on preventing explicit harms such as toxic or policy-violating content, less attention has been paid to conversational b…
      • In this paper, we investigate whether LLMs can be guided toward more de-escalating dialogue behavior through lightweight prompt-level constraints derived from N…
  • Low Resource Multimodal Translation of Nepali Spoken Words into Emotion-Conditioned Sign Language Avatars

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26107v1 公告类型:新。
      • 摘要:整合情感表达的手语交流系统仍未得到充分探索,特别是对于资源匮乏的语言。
      • 这项试点研究提出了 NEST-V1(尼泊尔情感和语音转换器 - 第 1 版),这是一个概念验证的多模式框架,演示了从口头输入生成情绪调节的尼泊尔手语化身的可行性。
      • 作为初步调查,我们重点关注三种情绪状态(快乐、中立、悲伤)的四个常见尼泊尔语单词(“谢谢”、“你好”、“房子”、“我”),以验证我们的核心技术方法。
    • EN 要点:
      • arXiv:2606.26107v1 Announce Type: new
      • Abstract: Sign language communication systems, that integrate emotional expression remain underexplored, particularly for low-resource languages
      • This pilot study presents NEST-V1 (Nepali Emotion and Speech Transformer - Version 1), a proof-of-concept multimodal framework that demonstrates the feasibility…
      • As a preliminary investigation, we focus on four common Nepali words (“thank you”, “hello”, “house”, “me”) across three emotional states (happy, neutral, sad) t…
  • Where Larger Models Excel: The Primacy of Constraint-Guided Reasoning

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26108v1 公告类型:新。
      • 摘要:较大的语言模型在推理基准上始终优于较小的语言模型,但这种差距背后的推理差异仍未得到充分探索。
      • 在数学、物理、化学和编程的基准测试中,我们观察到稳定的性能差距:在数据集上平均,Qwen3-32B 比 Qwen3-8B 好 6.43%,而 GPT-OSS-120B 比 GPT-OSS-20B 好 7.38%。
      • 为了研究这些收益背后的推理差异,我们开发了 AdvCluster,这是一个自动化框架,可以识别较大模型显示出稳定优势的问题,从较大和较小模型产生的配对推理轨迹中提取细粒度的优势描述,并通过语义聚类来组织它们,并在审阅者模型的指导下进行定量评估和选择。
    • EN 要点:
      • arXiv:2606.26108v1 Announce Type: new
      • Abstract: Larger language models consistently outperform smaller ones on reasoning benchmarks, yet the reasoning differences underlying this gap remain underexp…
      • Across benchmarks in mathematics, physics, chemistry, and programming, we observe stable performance gaps: averaged over datasets, Qwen3-32B outperforms Qwen3-8…
      • To study the reasoning differences behind these gains, we develop AdvCluster, an automated framework that identifies questions where the larger model shows a st…
  • From Lexicon to AI: A Structured-Data Pipeline for Specialized Conversational Systems in Low-Resource Languages

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26112v1 公告类型:新。
      • 摘要:低资源语言在人工智能开发中面临着严峻的挑战:在无法访问大量训练语料库的情况下创建专门的对话系统。
      • 我们提出了一种将结构化语言资源转化为专门的人工智能系统的系统方法,证明专家策划的词汇数据库可以作为对话式人工智能开发的有效基础。
      • 我们的方法将 Hindi WordNet 转换为 125 万个不同的指令响应对,使用具有 4 位量化的资源高效型 LoRA 微调 12B 参数语言模型。
    • EN 要点:
      • arXiv:2606.26112v1 Announce Type: new
      • Abstract: Low-resource languages face a critical challenge in AI development: creating specialized conversational systems without access to massive training cor…
      • We present a systematic methodology for transforming structured linguistic resources into specialized AI systems, demonstrating that expert-curated lexical data…
      • Our approach converts Hindi WordNet into 1.25 million diverse instruction-response pairs, fine-tunes a 12B-parameter language model using resource-efficient LoR…

ArXiv cs.LG (B_intro+search) 链接到标题

  • Physics-guided Convolutional Neural Network for Domain Growth Prediction in Systems with Conserved Kinetics

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26128v1 公告类型:新。
      • 摘要:许多物理、化学和生物系统的时空演化是通过非线性偏微分方程(PDE)来描述的。
      • 最近,基于深度神经网络的代理模型作为计算成本昂贵的传统数值求解器的有效替代品越来越受到人们的关注。
      • 在这项工作中,我们提出了一种基于注意力的、物理引导的卷积神经网络作为替代模型来学习此类系统的微观结构演化。
    • EN 要点:
      • arXiv:2606.26128v1 Announce Type: new
      • Abstract: The spatiotemporal evolution of many physical, chemical, and biological systems is described by nonlinear partial differential equations (PDEs)
      • Recently, deep neural network-based surrogate models have gained increasing interest as efficient alternatives to computationally expensive traditional numerica…
      • In this work, we propose an attention-based, physics-guided convolutional neural network as a surrogate model to learn the microstructural evolution of such sys…
  • \chisao{}: A GPU-Native Parallel Optimizer for Multimodal Black-Box Functions via Convergence-Anticonvergence Oscillation

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26164v1 公告类型:新。
      • 摘要:寻找多模态黑盒函数的所有模式是优化、贝叶斯推理和科学计算中的基本挑战。
      • 现有方法——盆地跳跃、CMA-ES、多起点梯度下降——按顺序操作,无法利用现代 GPU 硬件的大规模并行性。
      • 我们引入了 \chisao{} (\textbf{C}onvergence-\textbf{H}alt-\textbf{I}nvert-\textbf{S}tick-\textbf{A}nd-\textbf{O}scillate),这是一个 GPU 原生群体优化器,它同时运行整个样本批次,并利用故意的收敛-反收敛振荡周期来逃避局部陷阱,同时冻结确认模式。
    • EN 要点:
      • arXiv:2606.26164v1 Announce Type: new
      • Abstract: Finding all modes of a multimodal black-box function is a fundamental challenge in optimization, Bayesian inference, and scientific computing
      • Existing approaches – basin-hopping, CMA-ES, multistart gradient descent – operate sequentially and cannot exploit the massive parallelism of modern GPU hardw…
      • We introduce \chisao{} (\textbf{C}onvergence-\textbf{H}alt-\textbf{I}nvert-\textbf{S}tick-\textbf{A}nd-\textbf{O}scillate), a GPU-native population optimizer th…
  • Implementation of reinforcement learning in chemical reaction networks: application to phototaxis as curiosity-driven exploration

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26168v1 公告类型:新。
      • 摘要:生命系统使用嘈杂且不完整的感官信号来导航环境。
      • 在单细胞藻类中,趋光性通常被建模为由刺激响应规则驱动的机械运行-翻滚过程。
      • 然而,这样的描述忽略了生物体如何主动采样其环境以减少感官模糊性。
    • EN 要点:
      • arXiv:2606.26168v1 Announce Type: new
      • Abstract: Living systems navigate environments using noisy and incomplete sensory signals
      • In unicellular algae, phototaxis is often modeled as a mechanistic run–tumble process driven by stimulus–response rules
      • However, such descriptions overlook how organisms actively sample their environment to reduce sensory ambiguity
  • Neural Architecture Search for Generative Adversarial Networks: A Comprehensive Review and Critical Analysis

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26169v1 公告类型:新。
      • 摘要:神经架构搜索 (NAS) 已成为优化生成对抗网络 (GAN) 设计的关键技术,可自动搜索有效架构,同时解决手动设计中固有的挑战。
      • 本文对应用于 GAN 的 NAS 方法进行了全面回顾,根据搜索策略、评估指标和性能结果等标准对各种方法进行分类和比较。
      • 该评论强调了 NAS 在提高 GAN 性能、稳定性和效率方面的优势,同时也确定了未来研究的局限性和领域。
    • EN 要点:
      • arXiv:2606.26169v1 Announce Type: new
      • Abstract: Neural Architecture Search (NAS) has emerged as a pivotal technique in optimizing the design of Generative Adversarial Networks (GANs), automating the…
      • This paper provides a comprehensive review of NAS methods applied to GANs, categorizing and comparing various approaches based on criteria such as search strate…
      • The review highlights the benefits of NAS in improving GAN performance, stability, and efficiency, while also identifying limitations and areas for future resea…
  • KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26179v1 公告类型:新。
      • 摘要:虽然基于 WGS 的 AMR 预测已达到高精度,但现有模型缺乏在已建立的生物途径中进行神经归因的机制。
      • 我们提出了 KG-TRACE,一种新颖的神经符号框架,它将 WHO 突变知识图 (KG) 集成为神经基因组模型的结构化生物约束。
      • 与孤立学习统计模式的现有方法不同,KG-TRACE 通过学习的认知信任门融合基因组特征和基于 RotatE 的 KG 嵌入,根据符号生物学知识动态加权神经证据。
    • EN 要点:
      • arXiv:2606.26179v1 Announce Type: new
      • Abstract: While WGS-based AMR prediction has reached high accuracy, existing models lack a mechanism to ground neural attributions in established biological pat…
      • We present KG-TRACE, a novel neuro-symbolic framework that integrates the WHO mutation knowledge graph (KG) as a structured biological constraint on a neural ge…
      • Unlike existing methods that learn statistical patterns in isolation, KG-TRACE fuses genomic features and RotatE-based KG embeddings through a learned epistemic…
  • Necessary but Not Sufficient: Temperature Control and Reproducibility in LLM-as-Judge Safety Evaluations

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26185v1 公告类型:新。
      • 摘要:LLM 作为法官(“评分者”)组件现已成为评估工具的标准配置,包括安全评估,其中通过/失败判决可能会影响下游部署决策。
      • 一个普遍的假设是,将分级机的采样温度设置为 0 可以使分级具有确定性。
      • 我们针对真实的安全评估代码库(日本 AISI 的开源 aisev)测试了这一假设,并表明它在两个层面上失败了。
    • EN 要点:
      • arXiv:2606.26185v1 Announce Type: new
      • Abstract: LLM-as-judge (“grader”) components are now standard in evaluation harnesses, including safety evaluations where a pass/fail verdict may gate downstrea…
      • A widespread assumption is that setting the grader’s sampling temperature to 0 makes grading deterministic
      • We test this assumption against a real safety-evaluation codebase (Japan AISI’s open-source aisev) and show it fails on two levels
  • Clue-Guided Money Laundering Group Discovery

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26189v1 公告类型:新。
      • 摘要:洗钱集团发现(MLGD)旨在识别隐藏的犯罪集团并恢复其在大规模金融网络中的完整结构。
      • 现有的图异常检测方法主要产生节点级风险警报,而全局组发现方法则被动地在整个网络中搜索可疑组。
      • 两者都与真实的反洗钱(AML)调查不相符,分析师通常从具体线索开始,逐步扩大调查范围以追回责任人。
    • EN 要点:
      • arXiv:2606.26189v1 Announce Type: new
      • Abstract: Money Laundering Group Discovery (MLGD) aims to identify hidden criminal groups and recover their complete structures in large-scale financial network…
      • Existing graph anomaly detection methods mainly produce node-level risk alerts, while global group discovery methods passively search for suspicious groups over…
      • Both are mismatched with real Anti-money-laundering (AML) investigations, where analysts usually start from a concrete clue and gradually expand the investigati…
  • Federated Hash Projected Latent Factor Learning

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26192v1 公告类型:新。
      • 摘要:哈希学习(HL)是一种有效的表示学习方法,可将实值数据映射为紧凑的二进制表示。
      • 传统的HL方法通常需要用户将个人数据上传到中央服务器,这与日益严格的数据安全法规不兼容。
      • 联邦学习(FL)提供了一种去中心化范例,用于学习全局最优模型,而无需集中私有数据。
    • EN 要点:
      • arXiv:2606.26192v1 Announce Type: new
      • Abstract: Hash Learning (HL) is an efficient representation learning approach that maps real-valued data into compact binary representations
      • Traditional HL methods typically require users to upload personal data to a central server, which is incompatible with increasingly stringent data security regu…
      • Federated Learning (FL) provides a decentralized paradigm for learning globally optimal models without centralizing private data
  • Statistical and Structural Approaches to Algorithmic Fairness

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26200v1 公告类型:新。
      • 摘要:现代机器学习系统已经超越了其作为孤立的预测结构的起源,演变成积极调节人类机会的复杂的社会技术架构。
      • 随着算法越来越多地决定获得经济和社会机会的机会,人们普遍认识到这些系统深深地植根于其环境的结构性不平等和偏见。
      • 人们越来越认识到,针对预测准确性而优化的模型可能会系统性地使边缘群体处于不利地位,因此算法公平领域的出现是为了回应这种认识。
    • EN 要点:
      • arXiv:2606.26200v1 Announce Type: new
      • Abstract: Modern machine learning systems have outgrown their origins as isolated predictive constructs, evolving into complex socio-technical architectures tha…
      • As algorithms increasingly determine access to economic and social opportunities, it has become widely recognized that these systems are deeply embedded with th…
      • The field of algorithmic fairness emerged in response to the growing recognition that models optimized for predictive accuracy can systematically disadvantage m…
  • Topology-Informed Neural Networks for Flood Detection in Optical and Synthetic Aperture Radar Imagery

    • 发布时间:2026-06-26 12:00 北京时间
    • 摘要:- arXiv:2606.26204v1 公告类型:新。
      • 摘要:洪水频繁影响世界各地的地区。
      • 快速、准确的洪水检测对于应急响应和及时减轻人员和经济损失至关重要。
      • 卫星数据可用性的不断扩大和人工智能的进步增强了对环境危害的监测,但由于云层遮盖了光学卫星图像,许多洪水事件仍然难以检测。
    • EN 要点:
      • arXiv:2606.26204v1 Announce Type: new
      • Abstract: Floods frequently impact regions around the world
      • Rapid and accurate flood detection is crucial for emergency response and timely mitigation of human and economic loss
      • The expanding availability of satellite data and advances in artificial intelligence have enhanced monitoring of environmental hazards, but many flood events re…