🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-06-20
- 类型
- ai-daily
- 字数
- 3409
- 阅读时长
- 17 min
2026-06-20 AI日更 | 智能体进入治理时刻,端侧模型开始重塑工作流 链接到标题
今日主线从模型能力转向可控落地:智能体需要权限、审计与运行时治理;端侧模型在多模态、编程和个人助理场景加速成熟;AI 编码工具则从辅助生成走向可复用自动化流程。
📖 本期 Watch List 深度导读 链接到标题
今天最值得深读的主线,是“智能体从演示走向治理”。OpenClaw 创始人访谈很适合产品和工程团队听:真正有价值的 Agent,正在由懂场景的领域专家定义;而关于运行时治理、DeFi 风险监督和澄清式不确定性的几篇论文,则提醒我们,能调用工具的系统必须被纳入权限、义务和审计框架。
第二条主线是模型范式与可靠性。扩散语言模型、ITNet 等工作继续挑战 Transformer/自回归的默认假设;多智能体审议中的“隐藏锚点”、临床表格数据里的认知盲区、随机路径聚合揭示偏见,则共同指向一个问题:模型不仅要更强,还要知道自己为何错、何时该停。
最后,科技权力结构也值得关注。从 SpaceX IPO、Cursor 收购传闻到 Anthropic 寓言争议,AI 产业正在同时重塑资本、平台与公共叙事。
🌐 X 平台 AI 热点快讯 链接到标题
话题 1:Transformer Pioneer Noam Shazeer Leaves Google for OpenAI 链接到标题
- 分类:AI · News
- 概况:热度时间:1 day ago,相关帖子数:15000
- 是什么事:Transformer 共同作者、MoE 先驱 Noam Shazeer 据称离开 Google 加入 OpenAI,引发外界对顶级 AI 人才流动的关注。
- 为什么重要:Shazeer 参与奠定了现代大模型架构基础,其去向被视为衡量前沿模型研发能力、组织吸引力和下一代 AI 技术路线的重要信号。
- 讨论概况:X 上讨论集中在 Google 是否持续流失关键 AI 人才、OpenAI 是否进一步巩固研究优势,以及 Transformer 之后的下一代架构或训练范式将由谁主导;也有人质疑相关消息细节和时间线是否准确。
话题 2:Loop Engineering Turns AI Agents into Self-Sustaining Coders 链接到标题
- 分类:AI · News
- 概况:热度时间:10 hours ago,相关帖子数:336
- 是什么事:Loop Engineering 提出一种让 AI 代理通过循环式开发流程自主编写、测试和改进代码的方法,被称为“自我维持的编码者”。
- 为什么重要:这反映了 AI 编程代理从辅助生成代码走向持续执行工程任务的趋势,可能影响软件开发效率、自动化程度以及人类工程师的角色分工。
- 讨论概况:X 上的讨论主要集中在这类代理是否真的具备长期自主开发能力、代码质量和安全性如何保障,以及它会加速开发者生产力提升还是带来过度自动化和岗位替代风险。
话题 3:Anthropic Fixes Claude Code Usage Bug for Premium Users 链接到标题
- 分类:AI · News
- 概况:热度时间:20 hours ago,相关帖子数:3700
- 是什么事:Anthropic 修复了影响 Claude Code 高级订阅用户的用量统计或额度扣减问题。
- 为什么重要:Claude Code 面向开发者的高频工作流,额度和计费准确性直接影响用户信任、企业采用和 AI 编程工具的可持续商业模式。
- 讨论概况:X 上讨论主要集中在 Anthropic 是否及时补偿受影响用户、Claude Code 的用量规则是否足够透明,以及在 OpenAI 等竞争对手加速推出编程代理能力时,Anthropic 能否维持开发者信心。
话题 4:Z.ai’s GLM-5.2 Tops Open AI Model Charts with Strong Benchmarks 链接到标题
- 分类:AI · News
- 概况:热度时间:1 day ago,相关帖子数:19000
- 是什么事:智谱旗下 Z.ai 发布的 GLM-5.2 在多项开放模型基准测试中取得领先成绩,引发 AI 社区关注。
- 为什么重要:这显示中国开源大模型在推理、代码和通用能力评测上继续逼近或超越国际主流模型,可能加速开放模型生态的竞争与应用落地。
- 讨论概况:X 上讨论主要集中在基准成绩是否真实反映实际能力、GLM-5.2 与 DeepSeek、Qwen、Llama 等模型的差距,以及开放权重、可商用性和部署成本对开发者的吸引力。
今日 X 上的 AI 舆情小结 链接到标题
今天的舆论主线围绕“前沿 AI 能力正在向更强组织、更自主工具和更开放生态集中”展开:顶级研究人才流动被视为判断 OpenAI、Google 等机构竞争力的风向标,而编码代理和开源模型进展则显示 AI 正从模型能力竞赛走向实际工程生产力竞争。共识是,AI 编程、推理和开放模型生态都在快速成熟,开发者工作流会被深度重塑,人才、算力、产品体验和商业信任将共同决定平台优势。分歧主要在于相关消息和基准是否可靠、自主编码代理是否真正具备长期工程能力,以及中国开源模型的领先成绩能否转化为真实场景中的稳定优势。潜在风险则集中在过度依赖未经充分验证的代理系统、计费和额度不透明损害用户信任,以及人才和技术资源进一步向少数头部机构集中,放大行业竞争和治理压力。
💡 大佬观点(Influencer Insights) 链接到标题
AI 行业动态日报分析 链接到标题
日期:2026年6月18-19日(综合近期)
1. 今日焦点:端侧模型的崛起与工程化实践 链接到标题
过去24小时,AI大佬们的讨论高度指向一个核心趋势:高性能端侧模型(On-device Models)正在从“能用”走向“好用”,并开始重塑开发者和高级用户的工作流。
1.1 端侧模型能力已验证,性能与质量双双突破 链接到标题
- @zhixianio 对 MiniCPM-o 4.5 的音视频全双工效果表示“很满意”,惊叹“很难想象这是一个 9B 模型能达到的效果”。尽管长时间运行仍有稳定性问题,但其质量已初步可用。
- @zhixianio 将 Qwen3.6-35B-A3B MoE(oMLX) 称为“甜点🍮宝座”,其在个人助理(PA)和编程(Coding)场景下,响应速度优于远程LLM,且原生多模态使用体验“比 DSV4 Pro 还要爽”。
- Google Gemma 4 家族成为端侧焦点:
- Gemma 4 E4B + MTP:@zhixianio 实测其在日文邮件解析分类任务中性能出色。
- Gemma 4 12B Coder:@zhixianio 进行了严苛的代码生成对比测试。结论是,在与 Qwen3.6-35B-A3B MoE 的对决中,Gemma 12B Coder 在复杂、有状态的程序生成(如俄罗斯方块)上存在明显“天花板”,12B参数量难以支撑长篇复杂逻辑。
- Gemma 4 QAT 量化感知模型:@zhixianio 特别强调了Google的量化感知训练(QAT)思路,认为这是关键的端侧优化方向,“Android 快能用上自带的模型了”。
1.2 AI编码工具进入“双雄”争霸与自动化新阶段 链接到标题
- Claude Code vs. OpenAI Codex:@ruanyf 的提问“你是用 Codex 还是 Claude Code?”引发讨论。@vista8 表示,Codex 产品更优,但特定场景仍需 Claude Code,并用自研MCP实现两者协同,甚至实现“双倍Codex额度”的玩法。@gefei55 同样分享了一个通过MCP让ChatGPT网页版获得Codex本地代码操控能力的开源项目,本质也是为了实现“双倍额度”和调用最强模型。
- Claude Code 推出 Artifact 可视化协作功能:@dotey 详细解读了此功能,称其解决了AI编程成果“只有操作者自己看得到”的协作难题,让调试、PR走查、架构说明等工作成果能以实时网页形式分享给团队。
- OpenAI Codex “Record & Replay”功能:@dotey 和 @AI_Jasonyu 高度评价此功能,认为其本质是“超级版本的RPA + 按键精灵 + Computer Use的结合体”。用户只需演示一遍操作流程,Codex就能自动生成可复用的Skill,极大降低了自动化门槛。
2. 值得注意的独特观点与行业前瞻 链接到标题
2.1 从“相关性”到“因果性”的技术范式转移 链接到标题
- @Pluvio9yte 深度分析了黄碧薇教授创办的 Aether AI,并指出其“因果大模型(Causal World Models)”指出了一个关键的下一阶段方向。他认为,当前大模型在生成“往底部有洞的杯子里倒水”这类画面时会失败,因为它们学习的是“倒水”和“满杯”的数据相关性,而非“水会从洞中漏出”的物理因果机制。Aether AI 旨在让AI理解环境变量和干预结果,这对具身智能、新材料研发等要求逻辑严谨性的领域至关重要。
2.2 Vibe Coding 的进化:从“需求优先”到“契约优先” 链接到标题
- @Pluvio9yte 分享了从安全从业者转型全栈开发的深度思考。他提出,Vibe Coding的最佳实践并非 Requirement First 或 Code First,而是 Contract First(契约优先)。他将此经验结合开源项目OpenSpec,形成了一套“将容易漂移的上下文外化成契约,让人和AI都有稳定参照物”的开发框架,目标是让开发经验少的人也能规范地进行大型项目开发。
2.3 AI时代的个人发展哲学 链接到标题
- @ruanyf 转发《今天可以放假吗》一文,提出AI大幅提高白领工作效率后,员工的收益何在。他预测,AI将提高全社会平均薪资或福利,这是长期趋势。
- @lijigang 从“菩萨畏因,凡夫畏果”的哲学出发,指出个体的 三观(人生观、世界观、价值观)是处理人生事件的底层“函数f”,其设计优于对单一结果(f(x))的祈求。同时,他提醒Token消耗量是“虚假指标”,解决问题的效果才是“真实指标”。
- @gefei55 分享了AI辅助的“抢时间差”商业情报玩法:利用X API筛选近期高互动带链推文,在Google Trends出现信号前就发现新词、新产品,实现了“当别人在等Trends曲线时,你已经提前上站了”,并将此方法开源。
3. 推荐的工具与资源 链接到标题
3.1 开发与设计工具 链接到标题
- baoyu-design Skill(by @dotey):功能强大的本地设计Skill,支持生成动画视频并导出MP4,还能一键生成带配图的PPTX,并可二次编辑。其动画引擎采用声明式设计
f(t),支持精确到帧的导出。GitHub地址已发布。 - Meta Skill 2.0(by @yaojingang,@vista8推荐):基于Anthropic泄露的Claude code源码整合而成的“元Skill”,用于创建高质量Skill,被@vista8称为“用过的最好的Meta Skill”。
- Figma Chrome插件(by @vista8推荐):可将任意网页元素转为Figma可编辑图层,对设计师仿站和精准截图极为实用。
- Fable(by @zhixianio推荐):AI驱动的开发助手,能在40分钟内完成70%的Demo工作并优化原方案,被评价为“Shut up and take my money”。
3.2 效率与自动化 链接到标题
- YouMind 1.0(by @lifesinger,@AI_Jasonyu、@gefei55推荐):历时两年正式发布的内容创作工具,亮点是高质量的文章输出和极佳的平台(X、公众号)排版兼容性,配合配图功能一站式解决长文创作痛点。
- 本地视频翻译工具(by @xiaohu,@Pluvio9yte推荐):开源的视频翻译一条龙工具,可实现下载、转写、翻译、润色、烧字幕全流程自动化。
- EvoMap(@vista8推荐):平台活动,GitHub开源项目的Star可兑换大模型API Token,鼓励开发者封装工作流、Prompt并上传以获取更多奖励。
- Giffgaff英国手机卡指南(by @AI_Jasonyu):解决海外AI服务(Codex, Claude Code)手机号验证难题的详细教程。
3.3 前沿模型与框架 链接到标题
- oMLX v0.4.0(@jundotkim发布,@zhixianio转发):首个支持原生Swift macOS App的正式版本,是流畅运行MLX模型的关键工具。
- PP-OCRv6(by @AI_Jasonyu):百度发布的超轻量OCR模型,仅1.5MB即可在浏览器运行,单图最快97ms,逐字识别准确率反超GPT-5.5等大模型,非常适合端侧集成。
- OpenClaw(@zhixianio提及):一款集成本地模型的个人助理框架,配合端侧模型(如Qwen)可实现快速、智能的原生多模态交互。
📚 附录:今日 Watch List 更新源列表 链接到标题
时间窗口:最近 3 天;覆盖 22 个源;共 34 条更新
Y Combinator Podcast (B_intro+search) 链接到标题
- Why Domain Experts Are Winning In The Age Of AI
- 发布时间:2026-06-19 23:00 北京时间
- 摘要:- 您可能已经听说过 OpenClaw(以前称为 Clawdbot/Moltbot)。
- 引起轰动的开源人工智能助手可以在您自己的设备上运行,与您已经使用的消息应用程序连接,并且超越聊天功能,实际执行管理电子邮件、日历、文件、工作流程等任务。
- 现在来认识一下它背后的人。
- YC 的 Raphael Schaad 与 OpenClaw 的创始人 Peter Steinberger 坐下来讨论病毒式个人 AI 代理背后的“顿悟”时刻、为什么本地优先代理可以取代当今的许多应用程序,以及个人代理将如何重塑软件的未来。
- EN 要点:
- Bryant Chou co-founded Webflow, which today powers around 1% of all websites on the internet
- Now he’s back in the current YC batch with Ploy, an AI-powered website and marketing platform that doesn’t just build your site — it connects to your analytics,…
- In this episode of the Lightcone he explains how he built Ploy to be “anti slop,” how building today compares to his first startup, and why founders with domain…
- Now26:01 — First Three Months: Webflow 2013 vs
All-In Podcast (A_full) 链接到标题
- World’s First Trillionaire, Anthropic Fable Banned, The New Oligarchs, Iran Peace Deal
- 发布时间:2026-06-20 06:07 北京时间
- 摘要:- (0:00) 闺蜜介绍。
- (2:41) 新寡头、美国即将上任的政治局以及习得性无助。
- (14:18) SpaceX 破纪录的 IPO、60B 美元的 Cursor 收购以及万亿富翁的反应。
- (33:34) Anthropic 的《寓言》禁令的幕后花絮。
- EN 要点:
- (0:00) Bestie intros
- (2:41) The New Oligarchs, America’s incoming politburo, and learned helplessness
- (14:18) SpaceX’s record breaking IPO, $60B Cursor acquisition, and trillionaire reactions
- (33:34) Behind the scenes of Anthropic’s Fable ban
Stratechery by Ben Thompson (A_full) 链接到标题
- 2026.25: The Stuff of Myth(os)
- 发布时间:2026-06-20 01:00 北京时间
- 摘要:-(罗纳德·科尔特斯/盖蒂图片社拍摄)。
- 欢迎回到本周的Stratechery!
- 提醒一下,每周、每周五,我们都会发送 Stratechery 捆绑包中的内容概述;突出显示的链接对所有人免费。
- 此外,您可以完全控制我们发送给您的内容。
- 就此而言,这是本周我们最喜欢的一些。
- EN 要点:
- (Photo by Ronald Cortes/Getty Images)
- Welcome back to This Week in Stratechery
- As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone
- Additionally, you have complete control over what we send to you
Two Minute Papers (B_intro+search) 链接到标题
- Scientists Found A Better Language For AI Agents
- 发布时间:2026-06-19 22:06 北京时间
- 摘要:- ❤️ 查看权重和偏差并在此处注册免费演示:。
- 📝 该论文可在此处获取:.
- Adam Bridges、Benji Rabhan、B Shang、Cameron Navor、Charles Ian Norman Venn、Christian Ahlin、Eric T、Fred R、Gordon Child、Juan Benet、Michael Tedder、Owen Skarpness、Richard Sundvall、Ryan Stankye、Shawn Becker、Steef、Taras Bobrovytsky、Tazaur Sagenclaw、Tybie Fitzhugh、Ueli Gallizzi。
- 科学家为人工智能代理找到了更好的语言。
- EN 要点:
- ❤️ Check out Weights & Biases and sign up for a free demo here:
- 📝 The paper is available here:
- Brain reading video:
- 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
ArXiv cs.AI (B_intro+search) 链接到标题
Deontic Policies for Runtime Governance of Agentic AI Systems
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19464v1 公告类型:新。
- 摘要:由大型语言模型 (LLM) 驱动的自主代理人工智能系统引入了新一类安全、隐私和合规性挑战:可以调用工具、操作数据、安装软件以及跨组织边界与对等代理进行协调的代理不仅必须受到身份验证和访问控制的约束,还必须受到企业治理的完整结构的约束。
- 这包括指定代理允许和禁止做什么、在采取某些行动后他们有义务做什么(例如通知 CISO)、在什么条件下可以放弃长期义务,以及当政策发生冲突时哪些规则优先。
- 这个治理问题超出了当前政策引擎所能提供的范围。
- EN 要点:
- arXiv:2606.19464v1 Announce Type: new
- Abstract: Autonomous agentic AI systems driven by Large Language Models (LLMs) introduce a new class of security, privacy, and compliance challenges: an agent t…
- This includes specifying what agents are permitted and prohibited from doing, what they areobliged to do after certain actions (e.g., notify the CISO), under wh…
- This governance problem exceeds what current policy engines provide
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19469v1 公告类型:新。
- 摘要:本科计算机科学受大约每十年修订一次的国际课程指南管辖,但课程缺乏可靠、可重复的方法来衡量它们覆盖当前指南的完全程度以及指南重组时覆盖范围如何变化。
- 我们通过人机交互管道来解决这个问题,该管道测量程序对外部知识体系的覆盖范围,并根据 2013 年 (CS2013) 和 2023 年 (CS2023) 计算机科学课程纵向应用于经认可的计算机科学学士学位。
- 管道将程序和每个指南表示为结构化语料库,通过语义检索生成候选课程到知识单元的匹配,并在明确的覆盖范围定义下通过人类判断来确认它们。
- EN 要点:
- arXiv:2606.19469v1 Announce Type: new
- Abstract: Undergraduate computer science is governed by international curricular guidelines revised about once a decade, yet programs lack a reliable, reproduci…
- We address this with a human-in-the-loop pipeline that measures a program’s coverage of an external body of knowledge, applied longitudinally to one accredited…
- The pipeline represents the program and each guideline as structured corpora, generates candidate course-to-knowledge-unit matches by semantic retrieval, and co…
Diffusion Language Models: An Experimental Analysis
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19475v1 公告类型:新。
- 摘要:大型语言模型 (LLM) 通过自回归生成彻底改变了语言建模,在广泛的任务中实现了强大的性能。
- 最近,扩散语言模型(DLM)作为一种替代范式出现,它通过迭代去噪而不是下一个标记预测来生成文本,从而允许并行细化整个序列。
- 虽然已经提出了许多基于扩散的架构,但评估协议、数据集、推理预算和生成超参数的差异使得很难比较它们的功能并理解它们提供的权衡。
- EN 要点:
- arXiv:2606.19475v1 Announce Type: new
- Abstract: Large Language Models (LLMs) have revolutionized language modeling through autoregressive generation, enabling strong performance across a wide range…
- Recently, Diffusion Language Models (DLMs) have emerged as an alternative paradigm that generates text through iterative denoising rather than next-token predic…
- While numerous diffusion-based architectures have been proposed, differences in evaluation protocols, datasets, inference budgets, and generation hyperparameter…
Hidden Anchors in Multi-Agent LLM Deliberation
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19494v1 公告类型:新。
-摘要:多智能体 LLM 审议(智能体在多轮中交换和修改答案)越来越多地用于提高推理和准确性,但其工作方式和原因却很少被建模。
- 这种深思熟虑反映了人类如何做出决定。
- 作为社会性动物,我们既受到群体的拉动,即德格鲁特和弗里德金-约翰森等经典舆论动态模型捕捉到的羊群效应,也受到我们自己的内在信念的拉动,而它们却没有。
- EN 要点:
- arXiv:2606.19494v1 Announce Type: new
- Abstract: Multi-agent LLM deliberation, where agents exchange and revise answers over several rounds, is increasingly used to improve reasoning and accuracy, ye…
- Such deliberation mirrors how humans reach decisions
- As social animals we are pulled both by the group, the herd effect that classical opinion-dynamics models such as DeGroot and Friedkin–Johnsen capture, and by…
DeXposure-Claw: An Agentic System for DeFi Risk Supervision
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19501v1 公告类型:新。
- 摘要:去中心化金融使监管者面临快速变化、网络化的信用风险。
- 通用法学硕士代理人不太适合这种环境:他们过度阅读了薄弱的证据并建议采取高风险的干预措施,而现有的评估没有提供与监管机构一致的方法来衡量由此产生的误报。
- 我们引入了 DeXposure-Claw,一种基于预测的代理监督系统,通过结构化证据引导 LLM 决策:(1)DeXposure-FM,一种图形时间序列基础模型,预测未来的暴露网络; (2) 确定性监控器和压力场景然后将这些预测转化为类型警报、归因信号和场景证据; (3) 在 DeXposure-Claw 发出带有理由的可审计监管罚单之前,数据健康状况和信任门会限制升级。
- EN 要点:
- arXiv:2606.19501v1 Announce Type: new
- Abstract: Decentralized finance exposes supervisors to fast-moving, networked credit risks
- General-purpose LLM agents fit this setting poorly: they over-read weak evidence and recommend high-stakes interventions, while existing evaluations offer no re…
- We introduce DeXposure-Claw, a forecast-grounded agentic supervision system that routes LLM decisions through structured evidence: (1) DeXposure-FM, a graph tim…
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19509v1 公告类型:新。
-摘要:大型语言模型(LLM)越来越多地应用于结构化临床数据,但它们是否能够认识到自己在此类任务上的知识的局限性仍有待探索。
- 我们通过跨模型归因分歧的视角研究这个问题,目的是减少结构化任务的认知不确定性,通过归因分歧分析在预测任务上比较 Qwen 2.5 7B 和 XGBoost。
- 首先,LLM 语言化置信度在认知上是空洞的,无论准确度是 49% 还是 75.3%,它都会输出接近常数 (0.856-0.937),跟踪提示格式而不是预测质量。
- EN 要点:
- arXiv:2606.19509v1 Announce Type: new
- Abstract: Large language models (LLMs) are increasingly applied to structured clinical data, yet whether they can recognize the limits of their own knowledge on…
- We study this question through the lens of cross-model attribution divergence with the goal of reducing epistemic uncertainty for structured tasks, comparing Qw…
- We report four findings
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19522v1 公告类型:新。
- 摘要:视网膜为神经退行性疾病提供了一个非侵入性的窗口,捕捉与未来认知能力下降风险相关的微妙结构模式。
- REVEAL 等视觉语言对齐框架表明,将视网膜眼底图像与结构化临床风险叙述配对可以改善阿尔茨海默病 (AD) 的早期预测。
- 这些方法中的一个关键设计选择是使用表型分组,其中具有相似风险状况的个体在对比学习期间被视为多阳性对。
- EN 要点:
- arXiv:2606.19522v1 Announce Type: new
- Abstract: The retina offers a noninvasive window into neurodegenerative disease, capturing subtle structural patterns associated with a risk of future cognitive…
- Vision-language alignment frameworks such as REVEAL have shown that pairing retinal fundus images with structured clinical risk narratives improves early predic…
- A key design choice in these approaches is the use of phenotypic grouping, where individuals with similar risk profiles are treated as multi-positive pairs duri…
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19527v1 公告类型:新。
- 摘要:大型语言模型 (LLM) 能否辨别自己的输出何时与人类道德不一致?
- 我们赋予法学硕士一个良心步骤,审查其自身的推理和输出,并且我们使用直接偏好优化(DPO)通过对齐组件扩展训练损失,以引导模型远离非道德输出。
- 结果是一种在线技术,可以在广泛的应用中调整模型:训练、微调、对抗性提示和零样本学习。
- EN 要点:
- arXiv:2606.19527v1 Announce Type: new
- Abstract: Can Large Language Models (LLMs) discern when their own outputs are misaligned with human ethics
- And can they self-correct
- We endow an LLM with a conscience step that reviews its own reasoning and outputs, and we extend the training loss with an alignment component using Direct Pref…
ITNet: A Learnable Integral Transform That Subsumes Convolution, Attention, and Recurrence
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19538v1 公告类型:新。
-摘要:卷积网络、循环网络和变压器各自编码不同的归纳偏差——局部性、顺序记忆和内容相关的成对交互——并且自诞生以来在数学上一直保持着不同。
- 我们表明,这种碎片反映的不是信号处理方式的根本多样性,而是单个基础数学对象的不完整视图:可学习的积分变换。
- 我们引入了积分变换网络(ITNet),这是一个围绕可学习内核构建的统一架构,该内核共同依赖于位置和特征。
- EN 要点:
- arXiv:2606.19538v1 Announce Type: new
- Abstract: Convolutional networks, recurrent networks, and transformers each encode different inductive biases – locality, sequential memory, and content-depend…
- We show that this fragmentation reflects not a fundamental diversity in how signals should be processed, but rather incomplete views of a single underlying math…
- We introduce the Integral Transform Network (ITNet), a unified architecture built around a learnable kernel that depends jointly on positions and features
Uncertainty Decomposition for Clarification Seeking in LLM Agents
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19559v1 公告类型:新。
- 摘要:最近的立场文件认为,经典的任意/认知不确定性框架对于交互式大语言模型(LLM)代理来说是不够的,并呼吁缺乏规范感知、分解和可交流的不确定性表示,这些表示可以解锁新的代理能力,例如主动寻求澄清和共享心智模型构建。
- 实际部署限制——黑盒API、交互式延迟预算和缺乏标记轨迹——排除了基于对数概率、多重采样和基于训练的方法,使基于提示的估计成为在部署时呈现此类信号的最可行的系列。
- 我们通过一个简单的基于提示的分解来回答这个调用,该分解将动作置信度与请求不确定性 (u) 分开,使代理能够在任务规范不明确时要求澄清。
- EN 要点:
- arXiv:2606.19559v1 Announce Type: new
- Abstract: Recent position papers argue that the classical aleatoric/epistemic uncertainty framework is insufficient for interactive large language model (LLM) a…
- Practical deployment constraints – black-box APIs, interactive latency budgets, and the absence of labeled trajectories – rule out logprob-based, multi-sampli…
- We answer this call with a simple prompt-based decomposition that separates action confidence from request uncertainty (u), enabling the agent to ask for clarif…
ArXiv cs.CL (B_intro+search) 链接到标题
Exposing the Unsaid: Visualizing Hidden LLM Bias through Stochastic Path Aggregation
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19344v1 公告类型:新。
- 摘要:大型语言模型 (LLM) 表现出表征和句法偏差,由于文本生成的随机性,这些偏差很难评估。
- 标准审核方法依赖于单一输出检查或静态自动化指标。
- 这些方法掩盖了潜在的概率分布,并且无法捕获隐藏在较低概率生成分支中的偏差。
- EN 要点:
- arXiv:2606.19344v1 Announce Type: new
- Abstract: Large Language Models (LLMs) exhibit representational and syntactic biases that are difficult to evaluate due to the stochastic nature of text generat…
- Standard auditing methods rely on a single output inspection or static automated metrics
- These approaches obscure the underlying probability distributions and fail to capture biases hidden in lower-probability generation branches
Ensembles of Large Language Models for Identifying EQ-5D Studies in PubMed Based on Their Abstracts
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19345v1 公告类型:新。
-摘要:科学出版物的快速增加导致系统文献综述(SLR)中的手动研究筛选越来越消耗资源、效率低下且不一致。
- 对清楚报告与健康相关的生活质量结果(例如 EQ-5D 数据)的研究进行分类,需要高水平的临床解释,这给人类审查人员带来了挑战。
- 这项研究调查了 Google 的 Gemini 和 Gemma 大语言模型 (LLM) 在仅基于已发表的摘要的 PubMed 生物医学数据库中自动进行 EQ-5D 检测的用途。
- EN 要点:
- arXiv:2606.19345v1 Announce Type: new
- Abstract: The rapid increase in scientific publications leads to the fact that manual study screening in systematic literature reviews (SLRs) is increasingly re…
- Classifying studies that clearly report health-related quality-of-life results, such as EQ-5D data, requires a high level of clinical interpretation and poses c…
- This study investigates the use of Google’s Gemini and Gemma large language models (LLMs) in automating EQ-5D detection in the PubMed biomedical database based…
Disentangling Linguistic Relatedness from Task Alignment in Cross-Lingual Transfer
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19346v1 公告类型:新。
- 摘要:我们通过微调阿拉伯语的七个大型语言模型(4B–671B 参数)并评估闪米特语言和非闪米特对照的零样本阅读理解来研究跨语言迁移。
- 在密集和专家混合架构中,我们没有发现闪米特特定迁移的证据:具有弱基线的模型在所有语言中都有显着改善,而强基线模型仅显示出边际收益,无论语系如何。
- 思想链消融强化了这一发现——从微调中受益最多的相同模型同样从推理时间推理中受益,这表明这两种机制都解决任务格式对齐问题,而不是跨语言知识转移。
- EN 要点:
- arXiv:2606.19346v1 Announce Type: new
- Abstract: We study cross-lingual transfer by fine-tuning seven large language models (4B–671B parameters) on Arabic and evaluating zero-shot reading comprehens…
- Across dense and Mixture-of-Experts architectures, we find no evidence of Semitic-specific transfer: models with weak baselines improve dramatically across all…
- A chain-of-thought ablation reinforces this finding – the same models that benefit most from fine-tuning benefit equally from inference-time reasoning, suggest…
How LLMs Fail and Generalize in RTL Coding for Hardware Design?
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19347v1 公告类型:新。
- 摘要:将顺序编程先验转换为硬件设计的并行时序逻辑仍然是大型语言模型(LLM)的关键瓶颈。
- 为了研究这一点,我们受认知理论的启发,引入了一种基于问题可解决性的新错误分类法。
- 我们的分类法将失败分为句法、语义、可解函数和不可解函数类型。
- EN 要点:
- arXiv:2606.19347v1 Announce Type: new
- Abstract: Translating sequential programming priors into the parallel temporal logic of hardware design remains a crucial bottleneck for large language models(L…
- To investigate this, we introduce a new error taxonomy grounded in problem solvability, inspired by cognitive theory
- Our taxonomy categorizes failures into syntactic, semantic, solvable functional, and unsolvable functional types
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19348v1 公告类型:新。
- 摘要:我们推出了 DeepSeek-V4 系列的预览版,包括两个强大的 Mixture-of-Experts (MoE) 语言模型 - 具有 1.6T 参数(49B 激活)的 DeepSeek-V4-Pro 和具有 284B 参数(13B 激活)的 DeepSeek-V4-Flash - 均支持 100 万个令牌的上下文长度。
- DeepSeek-V4系列在架构和优化方面进行了多项关键升级:(1)混合注意力架构,结合压缩稀疏注意力(CSA)和重压缩注意力(HCA),提高长上下文效率; (2) 流形约束超连接(mHC),增强传统的残差连接; (3) 和 Muon 优化器可实现更快的收敛和更高的训练稳定性。
- 我们在超过 32T 的多样化和高质量代币上对这两个模型进行预训练,然后通过全面的后训练管道来解锁并进一步增强其功能。
- EN 要点:
- arXiv:2606.19348v1 Announce Type: new
- Abstract: We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models – DeepSeek-V4-Pro with 1.6T paramet…
- DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attent…
- We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline that unlocks and further enhances…
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19349v1 公告类型:新。
- 摘要:虽然上下文学习 (ICL) 在自回归 (AR) 法学硕士中得到了广泛研究,但其在扩散大语言模型 (dLLM) 中的机制仍然很大程度上未被探索。
- 与受单向因果屏蔽限制的 AR 模型不同,dLLM 本质上利用双向注意力,为查询放置提供广泛的空间灵活性。
- 不幸的是,当前的实践通常继承 AR 风格的尾随查询模板,常常忽视结构范式的转变。
- EN 要点:
- arXiv:2606.19349v1 Announce Type: new
- Abstract: While In-Context Learning (ICL) is extensively studied in Autoregressive (AR) LLMs, its mechanism within Diffusion Large Language Models (dLLMs) remai…
- Unlike AR models restricted by unidirectional causal masking, dLLMs intrinsically utilize bidirectional attention, offering extensive spatial flexibility for qu…
- Unfortunately, current practices conventionally inherit AR-style trailing-query templates, often overlooking the structural paradigm shift
Pruning via Causal Attribution Preserves Reasoning Performance in Large Language Models
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19350v1 公告类型:新。
- 摘要:大型语言模型(LLM)擅长多步推理,但会产生大量推理成本。
- 我们引入了因果归因修剪(CAP),这是一种无需训练的方法,通过测量关键注意力头对推理任务的因果影响来识别关键注意力头,并使用这些头级分数来指导细粒度的权重修剪。
- 对于每个注意力头,CAP 估计在前向传递一小部分推理问题时该头被屏蔽时的预期性能下降。
- EN 要点:
- arXiv:2606.19350v1 Announce Type: new
- Abstract: Large language models (LLMs) excel at multi-step reasoning but incur substantial inference cost
- We introduce Causal Attribution Pruning (CAP), a training-free method that identifies critical attention heads by measuring their causal impact on reasoning tas…
- For each attention head, CAP estimates the expected performance degradation when the head is masked during forward passes on a small calibration set of reasonin…
Detecting Hallucinations for Large Language Model-based Knowledge Graph Reasoning
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19351v1 公告类型:新。
- 摘要:知识图(KG)推理从现有事实中推断出新知识,广泛应用于问答、推荐和决策支持。
- 随着大语言模型(LLM)的快速发展,基于LLM的知识图谱推理框架通过利用检索到的知识图谱信息变得越来越流行。
- 然而,法学硕士的幻觉仍然是一个关键问题。
- EN 要点:
- arXiv:2606.19351v1 Announce Type: new
- Abstract: Knowledge graph (KG) reasoning infers new knowledge from existing facts and is widely applied in question answering, recommendation, and decision supp…
- With the rapid development of large language models (LLMs), LLM-based KG reasoning frameworks have become increasingly popular by leveraging retrieved KG inform…
- However, hallucinations in LLMs remain a critical issue
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19352v1 公告类型:新。
- 摘要:手语是聋人和听力障碍 (DHH) 社区使用的表达性视觉语言。
- 尽管手语识别、翻译和制作方面取得了实质性进展,但进步仍然受到分散的数据集、不一致的注释和有限的语言覆盖范围的限制。
- 现有的基准通常无法反映现实世界的通信需求,并且对这些限制的系统分析仍然有限。
- EN 要点:
- arXiv:2606.19352v1 Announce Type: new
- Abstract: Sign languages are expressive visual languages used by Deaf and Hard-of-Hearing (DHH) communities
- Despite substantial progress in sign-language recognition, translation, and production, advances remain constrained by fragmented datasets, inconsistent annotat…
- Existing benchmarks often fail to reflect real-world communication needs, and systematic analyses of these limitations remain limited
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19353v1 公告类型:新。
-摘要:情境学习(ICL)允许法学硕士通过一些演示来适应新任务,但其可靠性仍然是一个问题:预测对提示设计和模型理解上下文的能力都高度敏感,模糊了失败是由数据属性还是模型限制引起的。
- 不确定性分解(将任意的认知来源分开)在这种情况下尤其重要,但为标准生成任务设计的现有方法无法捕获 ICL 的独特动态。
- 为了解决这个问题,我们引入了自函数向量的概念,该概念建立在贝叶斯观点和 ICL 的机械解释性之上。
- EN 要点:
- arXiv:2606.19353v1 Announce Type: new
- Abstract: In-Context Learning (ICL) allows LLMs to adapt to new tasks from a few demonstrations, but its reliability remains a concern: predictions are highly s…
- Uncertainty decomposition-separating aleatoric from epistemic sources-is particularly crucial in this setting, yet existing methods, designed for standard gener…
- To address this, we introduce a concept of self-function vectors, built upon Bayesian views and the mechanistic interpretability of ICL
ArXiv cs.LG (B_intro+search) 链接到标题
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19361v1 公告类型:新。
- 摘要:识别条件描述了目标查询或感兴趣参数的可计算性,作为可用信息类型和数量的函数。
- 在因果识别中,此信息通常以因果图的形式表示,并且针对图中的某些变量子集观察或收集数据。
- 目标查询可能仅针对单个效果,也可能针对给定模型中的一类效果。
- EN 要点:
- arXiv:2606.19361v1 Announce Type: new
- Abstract: Identification conditions describe the computability of a target query or parameter of interest as a function of the type and amount of information av…
- In causal identification, this information is often expressed in the form of a causal graph, and data are observed or collected for some subset of variables in…
- Target queries may be for a single effect alone or for a class of effects in a given model
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19363v1 公告类型:新。
-摘要:时间序列基础模型(TSFM)在物理科学中的部署受到了一个关键权衡的阻碍:虽然这些模型编码了丰富的、普遍的时间动态,但当零样本应用于特定科学领域时,它们会遭受严重的分布错位,而且它们的计算成本阻碍了在边缘计算传感器网络中的部署。
- 我们解决一个根本性的挑战:如何从错位的基础模型(FM)中提取潜在的结构知识来训练轻量级的专业预报员?
- 我们提出门控不确定性感知路由蒸馏(Guard),这是一种新颖的框架,它将多教师蒸馏重新构建为具有两种自适应机制的实例决策过程:(1)上下文路由器,根据本地输入统计数据动态选择最相关的教师,利用不同基础模型之间的互补性; (2) 不确定性门控温度机制,充当“断路器”,当教师信心偏离领域现实时,自动减弱蒸馏强度。
- EN 要点:
- arXiv:2606.19363v1 Announce Type: new
- Abstract: The deployment of Time-Series Foundation Models (TSFMs) in physical sciences is hindered by a critical trade-off: while these models encode rich, univ…
- We address a fundamental challenge: How can we extract latent structural knowledge from misaligned foundation models (FM) to train lightweight, specialized fore…
- We propose Gated Uncertainty-Aware Routing for Distillation (Guard), a novel framework that reframes multiteacher distillation as an instance-wise decision proc…
Closing the Social-Semantic Gap: SPSD for Edge-Based Prompt Compression in Cloud LLM Inference
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19364v1 公告类型:新。
- 摘要:大型语言模型 (LLM) 推理的预填充阶段对云规模能源成本的影响越来越大。
- 许多消费者支持和对话提示包含社交脚手架:礼貌标记、道歉序言、重复和建立融洽关系的语言,这些语言对人类交流很重要,但对机器推理来说边缘信息很少。
- 我们将这种差异称为社会语义差距。
- EN 要点:
- arXiv:2606.19364v1 Announce Type: new
- Abstract: The prefill stage of Large Language Model (LLM) inference is a growing contributor to cloud-scale energy cost
- Many consumer-support and conversational prompts contain social scaffolding: politeness markers, apologetic preamble, repetition, and rapport-building language…
- We call this discrepancy the Social-Semantic Gap
Performance Analysis and Optimization of 3D Generative Diffusion Models across GPU Architectures
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19365v1 公告类型:新。
- 摘要:扩散模型对于高保真 3D MRI 合成至关重要,但其部署仍然受到每个样本数百次 U-Net 评估和高度异构内核行为所产生的大量 GPU 资源需求的限制。
- 本文对三代 NVIDIA 架构中最先进的医疗扩散模型 Med-DDPM 进行了全面的性能分析,以研究内核级运行时故障、指令混合特征、内存系统利用率、扭曲级活动和分析器优先级分数估计。
- 我们表明训练绝大多数由 cuDNN 卷积和隐式 GEMM 内核主导,内存访问模式、张量布局转换和有限的张量核心利用率导致效率低下。
- EN 要点:
- arXiv:2606.19365v1 Announce Type: new
- Abstract: Diffusion models have become essential for high-fidelity 3D MRI synthesis, yet their deployment remains constrained by substantial GPU resource demand…
- This paper performs a comprehensive performance analysis of the state-of-the-art medical diffusion model, Med-DDPM, across three generations of NVIDIA architect…
- We show that training is overwhelmingly dominated by cuDNN convolution and implicit-GEMM kernels, with inefficiencies arising from memory-access patterns, tenso…
Information Lattice Learning as Probabilistic Graphical Model Structure Learning
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19366v1 公告类型:新。
- 摘要:信息格学习(ILL)通过交替地将信号投影到对抽象层次结构进行编码的分区格上并将选定的规则提升回信号域来学习信号的可解释规则。
- 当信号是概率质量函数时,我们展示了 ILL 学习到的概率规则,承认自然概率图形模型 (PGM) 解释并详细开发这种解释。
- ILL 中的分区会产生确定性商变量,而规则就是该商变量的边际法则。
- EN 要点:
- arXiv:2606.19366v1 Announce Type: new
- Abstract: Information lattice learning (ILL) learns interpretable rules of a signal by alternately projecting the signal onto a partition lattice that encodes a…
- When the signal is a probability mass function, we show the probabilistic rules learned by ILL admit a natural probabilistic graphical model (PGM) interpretatio…
- A partition in ILL induces a deterministic quotient variable, and a rule is the marginal law of that quotient variable
Weibull Weight-Scale Parameter Evolution under AdamW Training Dynamics
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19367v1 公告类型:新。
- 摘要:基于用于诊断变压器权重分布的二参数 Weibull 框架,我们研究了为什么 Weibull 权重尺度参数 $\lambda$ 在 AdamW 训练期间增长、超调,然后松弛。
- 我们从 AdamW 更新中推导出平方权重范数的前序三力分解:测量权重和自适应更新方向之间相关性的对齐力、来自自适应步幅的注入力以及来自解耦权重衰减的衰减力。
- 在具有地面实况优化器矩的自训练 Pythia-70M 模型上,对齐在上升阶段占主导地位,在四个随机种子中贡献了 88-94% 的绝对力预算,并且对超重去除保持稳健。
- EN 要点:
- arXiv:2606.19367v1 Announce Type: new
- Abstract: Building on a two-parameter Weibull framework for diagnosing transformer weight distributions, we study why the Weibull weight-scale parameter $\lambd…
- We derive a leading-order three-force decomposition of the squared weight norm from the AdamW update: an alignment force measuring the correlation between weigh…
- On self-trained Pythia-70M models with ground-truth optimizer moments, alignment dominates the rise phase, contributing 88-94% of the absolute force budget acro…
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19369v1 公告类型:新。
- 摘要:分布估计算法 (EDA) 是一类强大的黑盒优化进化方法,尤其是在对目标结构知之甚少的情况下。
- 经典进化算法依赖于手工设计的变异和交叉算子,很难针对未知的问题结构进行设计,并且存在偏差来源,而 EDA 完全回避算子设计:它们将概率分布拟合到最佳个体,并从中采样下一代。
- EDA 在连续参数空间上得到了很好的建立,但它们以前没有被推广到稀疏参数空间,在稀疏参数空间中,好的解决方案的大多数系数恰好为零。
- EN 要点:
- arXiv:2606.19369v1 Announce Type: new
- Abstract: Estimation-of-distribution algorithms (EDAs) are a powerful class of evolutionary methods for black-box optimization, especially when little is known…
- Whereas classical evolutionary algorithms rely on hand-designed mutation and crossover operators, hard to devise for unknown problem structures, and a source of…
- EDAs are well established on continuous parameter spaces, but they have not previously been generalized to sparse ones, in which most coefficients of a good sol…
Human-like autonomy emerges from self-play and a pinch of human data
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19370v1 公告类型:新。
- 摘要:自我对战强化学习最近已成为一种无需任何人类数据即可训练驾驶策略的方法。
- 它使用廉价的大规模模拟来替代昂贵的大规模人类驾驶演示。
- 这种方法的一个关键限制是,通过纯粹的自我游戏训练的政策可以学习有效但与人类不相容的外来驾驶惯例。
- EN 要点:
- arXiv:2606.19370v1 Announce Type: new
- Abstract: Self-play reinforcement learning has recently emerged as a way to train driving policies without any human data
- It uses cheap, large-scale simulations to substitute expensive, large-scale human driving demonstrations
- A key limitation of this approach is that policies trained through pure self-play can learn effective but alien driving conventions incompatible with people
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19371v1 公告类型:新。
- 摘要:阿尔茨海默病(AD)是一种致命性疾病,会破坏老年人的记忆和认知能力。
- 大多数 AD 治疗在早期阶段是有效的,导致对 AD 早期诊断的需求不断增加。
- AD 诊断越来越依赖于多模态数据,例如临床评估、结构磁共振成像 (MRI) 和正电子发射断层扫描 (PET) 成像。
- EN 要点:
- arXiv:2606.19371v1 Announce Type: new
- Abstract: Alzheimer’s disease (AD) is a fatal disorder that destroys memory and cognitive skills in the elderly population
- Most treatments for AD are effective in the early stage, leading to an increasing demand for early AD diagnosis
- AD diagnosis increasingly relies on multimodal data such as clinical assessments, structural Magnetic Resonance Imaging (MRI), and Positron Emission Tomography…
cAPM: Continual AI-Assisted Pace-Mapping with Active Learning
- 发布时间:2026-06-19 12:00 北京时间
- 摘要:- arXiv:2606.19373v1 公告类型:新。
- 摘要:室性心动过速是一种危及生命的心律失常,也是心源性猝死的主要原因。
- 起搏映射是一种临床程序,用于在 VT 导管消融过程中识别干预目标。
- 它要求临床医生对心室的不同部位进行起搏,并快速解释产生的心电图,以确定下一步在哪里起搏或是否已识别出目标部位。
- EN 要点:
- arXiv:2606.19373v1 Announce Type: new
- Abstract: Ventricular tachycardia is a life-threatening rhythm disorder and a major cause of sudden cardiac death
- Pace-mapping is a clinical procedure for identifying the intervention target during catheter ablation of VT
- It requires clinicians to pace different sites in the ventricles and rapidly interpret the resulting electrocardiograms to determine where to pace next or wheth…