🤖 AI 速览

vLLM 等开源推理引擎被重新定义为 AI 栈的关键底座;Jeff Dean 离开谷歌,也提示顶级人才与研究资源正在向独立组织流动。与此同时,模型进入医疗、青少年心理与法律等高风险场景后,评估、治理和可信部署的重要性明显上升。
📋 文章元数据
发布时间
2026-08-07
类型
ai-daily
字数
3635
阅读时长
18 min

2026-08-07 AI日更 | Jeff Dean 出走之后,AI 竞争转向推理基础设施 链接到标题

vLLM 等开源推理引擎被重新定义为 AI 栈的关键底座;Jeff Dean 离开谷歌,也提示顶级人才与研究资源正在向独立组织流动。与此同时,模型进入医疗、青少年心理与法律等高风险场景后,评估、治理和可信部署的重要性明显上升。

📖 本期 Watch List 深度导读 链接到标题

今天最值得优先看的,是“智能体走向个人设备与基础设施化”这条线:YC 对 OpenClaw 创始人的访谈,和 vLLM 作为开源推理引擎的深度讨论,可以连在一起读——前者关注个人 AI 助手如何真正接管邮件、日历、文件等工作流,后者则解释这些应用背后的推理层为何正成为 AI 栈的关键基础设施。

第二条主线是“模型能力进入高风险场景后的评估与治理”。OpenAI 更新 ChatGPT,并与 APA 合作关注青少年心理健康;多篇论文则从医学安全、隐私过滤、法律基准、LLM-as-judge 可复现性等角度提醒:偏好、更高分或最终答案正确,并不等同于安全可靠。

此外,Google DeepMind 的 WeatherNext 气旋预测,以及楔形文字、肿瘤多模态等新基准,值得作为“AI 深入专业领域”的观察样本。今天的信号很清晰:下一阶段竞争不只在模型本身,而在可信部署、评估体系和真实工作流整合。

🌐 X 平台 AI 热点快讯 链接到标题

话题 1:Google’s Jeff Dean Leaves After 27 Years to Co-Found AI Research Venture 链接到标题

  • 分类:AI · News
  • 概况:热度时间:1 day ago,相关帖子数:43000
  • 是什么事:谷歌资深计算机科学家 Jeff Dean 在任职 27 年后离开公司,转而共同创办一家新的 AI 研究机构。
  • 为什么重要:Jeff Dean 长期参与谷歌搜索、分布式系统和深度学习基础设施建设,其离职被视为顶尖 AI 人才与研究资源从大型科技公司向新型研究组织流动的标志。
  • 讨论概况:X 上讨论集中在这是否意味着谷歌 AI 人才流失加剧、新机构可能采取何种研究路线,以及独立 AI 研究组织能否在算力、资金和模型能力上与科技巨头竞争。

话题 2:Meta AI Models Win Gold in Five Major STEM Olympiads 链接到标题

  • 分类:AI · News
  • 概况:热度时间:7 hours ago,相关帖子数:670
  • 是什么事:Meta 宣布其 AI 模型在五项重要 STEM 奥赛中获得金牌成绩,展示了在数学、推理和综合解题上的强能力。
  • 为什么重要:这表明大模型正从通用对话能力进一步迈向高难度学术推理任务,对衡量 AI 的真实推理水平、教育应用和未来科研辅助能力都很重要。
  • 讨论概况:X 上的讨论主要集中在两点:一是这是否说明模型推理能力有实质提升,二是这种“金牌”成绩是否足够可信、是否可能受题目泄露或评测方式影响;同时也有人将其视为 Meta 在 AI 竞赛中追赶领先者的信号。

话题 3:Garry Tan Pushes Personal AGI for Solo Founders at Startup School 2026 链接到标题

  • 分类:AI · News
  • 概况:热度时间:,相关帖子数:70
  • 是什么事:Garry Tan 在 Startup School 2026 上倡导“个人 AGI”,鼓励独立创始人借助 AI 以更小团队甚至单人模式创业。
  • 为什么重要:这反映出 AI 正从工具走向创业能力放大器,可能重塑初创公司的组织形态、融资逻辑和产品开发效率。
  • 讨论概况:X 上的讨论主要集中在个人 AGI 是否真能让单人创业成为主流,以及这会不会降低创业门槛、加速创新;也有人质疑其可行性、夸大程度和对传统团队协作模式的冲击。

话题 4:Meta Launches Muse Code Beta for Complex Coding Tasks 链接到标题

  • 分类:AI · News
  • 概况:热度时间:1 day ago,相关帖子数:15000
  • 是什么事:Meta 发布了 Muse Code Beta,这是一款基于 Muse Spark 1.2 的终端式 AI 编程代理,面向复杂、多文件、长周期的软件工程任务。
  • 为什么重要:这表明 AI 正从代码补全走向可执行的工程协作,能在大型代码库中进行规划、修改和验证,影响软件开发流程、效率与人机分工。
  • 讨论概况:X 上讨论主要集中在其多智能体架构、是否真能胜过现有编码模型、长任务稳定性和实际生产可用性,也有人关注定价、安装方式以及对开发者工作的替代还是辅助作用。

今日 X 上的 AI 舆情小结 链接到标题

今天 X 上的舆论主线,是 AI 正从“会聊天”进一步走向“能做事、能创业、能研究”的阶段:一边是 Jeff Dean 出走引发对顶尖人才和研究资源从大厂外流的关注,另一边是 Meta 的奥赛金牌、Muse Code 和“个人 AGI”都在强化“模型正在接近真实工作能力”的叙事。相对一致的共识是,AI 的能力边界确实在扩展,尤其在推理、软件工程和创业效率上,已经开始改变组织形态和生产方式。分歧主要集中在这些成绩到底有多“真”——比如金牌是否可信、编码代理是否足够稳定、个人创业是否真会成为主流,以及独立 AI 研究机构能否在算力和资金上与巨头正面竞争。潜在风险则在于,外界对能力跃升的预期可能过快,若评测和演示被过度包装,会放大泡沫;同时,人才流动、自动化开发和单人创业热潮也可能让行业更集中于少数资源方,增加技术失控和工作替代的不确定性。

💡 大佬观点(Influencer Insights) 链接到标题

好的,基于过去24小时内多位AI大佬在X平台的推文汇总,以下是作为资深AI行业分析师的我,为你带来的深度洞察与总结。


1. 今日大佬们共同关注的技术趋势或产品热点 链接到标题

今日的核心关注点高度集中在模型能力进化、AI Agent 的落地形态以及端侧智能的突破上。

  • DeepSeek V4 正式版引发“性能-成本”风暴: 多个博主热议 DeepSeek V4 正式版。

    • @zhixianio 亲自在 Mac Studio 上测试了 DeepSeek V4 Flash 的 4bit 量化版,展示了其本地运行的可行性。
    • @vista8 援引数据和融资消息,指出 DeepSeek V4 Pro 性能已“跻身全球第一梯队”,编程能力仅比 Claude 旗舰模型低 0.3%,但其API定价仅为海外竞品的 1/10 到 1/100,性价比优势形成了“断层领先”。
    • 同时,@Pluvio9yte 观察到 DeepSeek 官方尚未推出自研 CLI 编程工具,并推荐了在 DeepSeek 官方文档中被提及、专门为其前缀缓存机制深度优化的开源 Agent 集成工具 Reasonix (已获32k Star),这反映出围绕 DeepSeek 的开发者生态正在快速成熟。
  • AI Agent 进入“桌面上脑”与“框架融合”新阶段

    • Agent 框架百花齐放@vista8 测评了名为 bb 的 Agent 框架,其亮点在于能自动识别并调用本机已安装的 Codex CLI、Claude Code CLI 等多种编程工具,实现了“框架即平台”的产品思路。@dotey 则关注到 Hermes Desktop 支持内置浏览器,让 Agent 能更直接地感知和操作 Web 环境。
    • 从终端走向桌面@dotey 观察到一个行业共识,即最强大的 Agent 都需要一台“自己的电脑”而不仅是容器,这推动了 Claude Code 等工具从命令行走向桌面客户端,Workbuddy 等产品也借此拿到大量市场。
    • 字节 SeedRealtime 重新定义多模态交互@vista8 详细体验并解读了字节的 SeedRealtime 模型,强调其实现了原生音视频全双工交互,模型不仅能听、能看,还能在连续音视频流中主动感知、决策和交互,比如在逛博物馆时主动提醒用户感兴趣的展品。这被**@vista8** 认为是超越了 GPT-4o 仅限于音频全双工的突破,可能影响具身机器人的发展速度。
  • 端侧模型部署:极限压榨硬件,叙事被不断重写: 端侧运行大模型的能力正以意想不到的速度突破。

    • @Pluvio9yte 分享了一个名为 Swiftlet 的项目,它将 80B 参数的 Qwen 模型塞进 4.3GB 内存的 Mac 上运行,甚至在 iPhone 上号称能跑 35B 模型。其核心技术是“按需流式加载”MoE 的专家权重,这为消费级硬件运行超大模型提供了全新路径。
    • @zhixianio 则通过实际测试得出结论:在本地代码生成任务上,一个部署在 M5 Max 上的 Qwen3.6-35B-A3B MoE 模型,在生成复杂程序(如俄罗斯方块)的能力上,显著优于最新的 Gemma 4 12B Coder。他指出,模型参数量带来的“天花板”比微调技术更重要,12B 模型撑不住“长篇、有状态、一次成型”的复杂程序。

2. 值得注意的独特观点或行业前瞻 链接到标题

  • Vibe Coding 的未来分工:专家与“全栈平民”共存(来自 @dotey) @dotey 提出了一个前瞻性观点:未来的编程世界,“专业人士不仅要救火,给 Vibe Coding 出的问题擦屁股,还需要搭好基础设施,方便大家高效安全地去 Vibe”。这意味着会保留少数专业前端、后端岗位,他们负责底层架构和安全,而大量中间层、由运营或产品人员兼任的“全栈”岗位将利用 AI 完成具体实现。这与当前“人人都是开发者”的叙事不同,更加强调专业人士在 AI 时代作为“基础设施构建者”和“守门人”的关键作用。

  • AI 写作的“AI 味”悖论(来自 @dotey) @dotey 对“消除 AI 味的 Skill”这一热门需求进行了深刻反思。他认为,人天然讨厌 AI 写的文章,但又想用 AI 写出没有 AI 味的文章,这本身是矛盾的。他坦言自己已放弃寻找完美的“去 AI 味 Skill”,因为“AI 不能代替人写作”。他将自己的方法论定位为“大量使用 AI 辅助写作”:利用 AI 帮忙收集资料、调整结构、在“卡住”时提供灵感,最终由人本身完成核心表达。这是一种从“替代”转向“协作者”的更务实态度。

  • 大模型公司的“营销新战场”:比拼谁能黑掉别人的系统(来自 @vista8) @vista8 调侃性地发现,多家大模型公司开始以“自家模型能黑入别人系统”作为营销亮点,这或许预示着 AI 安全问题将从纯粹的技术挑战,演变为一场带有公关色彩的公开竞赛。

  • 算力市场的“拼装厂”模式兴起(来自 @Pluvio9yte) Anthropic 与成立仅数月、硬件几乎全是租来的云初创公司 Volta 签下约 100 亿美元算力协议,这一事件被 @Pluvio9yte 敏锐地捕捉。他认为,这标志着算力不再只是自建或找大云商排队,一种“谁先把电和 GPU 拼起来,谁就能拿到前沿大单”的“拼装厂”模式开始出现,这将重塑 AI 基础设施的供应格局。

3. 推荐的工具或资源 链接到标题

  • AI 编程与开发工具

    • Reasonix: 专为 DeepSeek 模型优化的开源 Agent 编程框架,尤其适合利用其前缀缓存特性降低长会话成本,Star 数超 32k。 (@Pluvio9yte)
    • bb Agent 框架: 一个创新的 Agent IDE,能自动识别并集成本机安装的各种 CLI 编程工具(如 Codex, Claude Code, PI Agent 等),实现一站式管理。 (@vista8)
    • CodeBuddy NPC (腾讯云): 一款将 AI 模型封装为游戏 NPC 概念的产品,可在代码托管平台用自然语言指派“NPC”完成任务,为编程教育或自动化带来新玩法。 (@ruanyf)
    • QLMarkdown: 一个 Mac 小工具,安装后即可用“空格键”快速预览 Markdown 文件,提升开发效率。 (@vista8)
  • AI 原生应用与服务

    • @makeplayai: 一个免费的 AI 游戏构建平台,输入一句话指令即可生成包含美术、音效和动效的完整小游戏,支持分支开发以对比玩法。 (@Pluvio9yte)
    • 小红书 REDSkill 社区: 小红书推出的 AI Skill 托管与分享功能,允许用户将 Skill 文件上传至笔记,他人可一键安装。被**@ruanyf** 认为是全球首个将社媒平台与 Skill Hub 结合的尝试,为开发者提供了接触海量用户的新渠道。
  • AI 辅助写作与效率

    • mattpocock/skills: 包含 /hand off 等实用功能,能生成交接文档,让 AI 编程的上下文在新旧会话间无缝衔接。 (@Pluvio9yte)
    • 活人感写作 .skill (By @Khazix0918): 旨在帮助用户写出没有 AI 味的文字的开源 Skill。尽管存在争议,但仍被推荐。 (@Pluvio9yte)
    • qiaomu-campus-resume (By @vista8): 一个基于访谈模式生成和优化 PDF 简历的 Skill,参考了清华、MIT 等名校就业中心的最佳实践,适合大学生求职使用。 (@vista8)
  • AI 安全与基础设施

    • OpenConnector: 一个开源的密码连接网关,专为防止 AI Agent 泄漏密码到上下文而设计,能统一管理所有外部应用的连接授权,Agent 只能拿到执行结果而无法接触凭证。 (@ruanyf)

📚 附录:今日 Watch List 更新源列表 链接到标题

时间窗口:最近 3 天;覆盖 22 个源;共 36 条更新

a16z Podcast (A_full) 链接到标题

  • Inside vLLM: The Engine Powering Open-Source AI
    • 发布时间:2026-08-06 18:00 北京时间
    • 摘要:- Elena Burger 和 Matt Bornstein 以及 Inferact 联合创始人兼首席执行官 Simon Mo 加入,Inferact 是为当今许多最先进的人工智能应用程序提供支持的开源推理引擎。
      • 他们共同探讨开源人工智能如何从研究项目演变为关键基础设施、推理为何成为人工智能堆栈最重要的层之一,以及如何为世界各地的开发人员带来前沿智能。
      • 对话内容涵盖 vLLM 的起源、开放权重模型的兴起、公司为何越来越希望控制其 AI 基础设施,以及开源推理如何支持下一代 AI 应用程序。
      • 他们还讨论了模型许可、开放权重人工智能的经济学、Kimi K3、蒸馏、人工智能基础设施,以及为什么西蒙认为开放模型和封闭模型之间的差距正在迅速消失。
      • 在这里查看 a16z 使用人工智能所做的一切,包括文章、项目和更多播客。
    • EN 要点:
      • Elena Burger and Matt Bornstein are joined by Simon Mo, co-founder and CEO of Inferact, the open-source inference engine powering many of today’s most advanced…
      • Together, they explore how open-source AI evolved from a research project into critical infrastructure, why inference has become one of the most important layer…
      • The conversation covers vLLM’s origins, the rise of open-weight models, why companies increasingly want control over their AI infrastructure, and how open-sourc…
      • They also discuss model licensing, the economics of open-weight AI, Kimi K3, distillation, AI infrastructure, and why Simon believes the gap between open and cl…

Y Combinator Podcast (B_intro+search) 链接到标题

  • Garry Tan: Own Your Intelligence
    • 发布时间:2026-08-07 03:28 北京时间
    • 摘要:- 您可能已经听说过 OpenClaw(以前称为 Clawdbot/Moltbot)。
      • 引起轰动的开源人工智能助手可以在您自己的设备上运行,与您已经使用的消息应用程序连接,并且超越聊天功能,实际执行管理电子邮件、日历、文件、工作流程等任务。
      • 现在来认识一下它背后的人。
      • YC 的 Raphael Schaad 与 OpenClaw 的创始人 Peter Steinberger 坐下来,讨论了病毒式个人 AI 代理背后的“顿悟”时刻、为什么本地优先代理可以取代当今的许多应用程序,以及个人代理将如何重塑软件的未来。
    • EN 要点:
      • The next generation of startups will be built by smaller teams than ever before
      • At Startup School 2026, YC President & CEO Garry Tan explains why we’re entering the era of personal AGI: AI agents that run on your own infrastructure, compoun…
      • He shares the tools and workflows he uses every day, why every founder should own their intelligence instead of renting it, and what it means to build under you…

OpenAI Blog (A_full) 链接到标题

  • Improving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free users

    • 发布时间:2026-08-06 18:00 北京时间
    • 摘要:- 我们的使命是确保通用人工智能造福全人类。
      • 我们正在推出 ChatGPT 更新,以改善日常对话,同时扩大免费用户的访问范围。
      • 对于 Plus 和 Pro 用户,我们正在更新聊天中的 GPT‑5.6 Sol,以便更可靠地提供事实并提供更有针对性的答案。
      • 新的滑块可让您选择 ChatGPT 对每个响应的思考程度。
      • 对于免费用户,我们将默认模型更新为 GPT‑5.6 Luna,并通过无限制的文本聊天扩展访问权限。
    • EN 要点:
      • ChatGPT introduces improved GPT-5.6 Sol with better accuracy and consistency, plus expanded access for free users and unlimited everyday chats with GPT-5.6 Luna…
  • Working with the American Psychological Association on youth mental health and AI

    • 发布时间:2026-08-06 14:00 北京时间
    • 摘要:- 年轻人已经使用人工智能来学习、创造、提问和寻求建议。
      • 随着这种用途的增长,家庭、学校、临床医生和社区需要更清晰的证据、更好的资源和更强有力的保障措施。
      • 这就是为什么我们与美国心理学会 (APA) 合作,将心理科学带入我们对年轻人负责任的人工智能开发和使用的思考。
      • APA 是美国领先的心理学科学和专业组织,其基于证据的工作将有助于澄清什么是已知的、什么是不确定的,以及随着技术的发展,负责任的人工智能应该是什么样子。
      • “技术是青少年生活的一部分,我们的责任是通过安全、适合年龄且为家庭设计的体验来满足这一现实。
    • EN 要点:
      • OpenAI and the American Psychological Association advance evidence-based guidance, resources, and safeguards for responsible AI use and youth mental health.
  • From asking to doing: How the world is putting ChatGPT to work

    • 发布时间:2026-08-06 08:00 北京时间
    • 摘要:- 新的 OpenAI Signals 数据显示人们如何在全球范围内使用 ChatGPT,并提供有关采用情况、使用趋势和不断变化的行为的国家级见解。
      • OpenAI 博客的这篇文章解释了从要求到行动:世界如何让 ChatGPT 发挥作用,塑造更广泛的人工智能和基础设施格局。
      • 它还为创始人、运营商和投资者揭示了从要求到做:世界如何让 ChatGPT 发挥作用的实际意义。
    • EN 要点:
      • New OpenAI Signals data shows how people use ChatGPT worldwide, with country-level insights on adoption, usage trends, and evolving behavior.

Google DeepMind Blog (A_full) 链接到标题

  • WeatherNext: AI model achieves breakthrough in forecasting cyclones
    • 发布时间:2026-08-06 23:06 北京时间
    • 摘要:- WeatherNext:人工智能模型在预测气旋方面取得突破。
      • 这篇来自 Google DeepMind 博客的文章解释了 WeatherNext:AI 模型如何在预测气旋方面取得突破,塑造更广泛的 AI 和基础设施景观。
      • 它还为 WeatherNext 的创始人、运营商和投资者带来了实际意义:人工智能模型在预测气旋方面取得了突破。
    • EN 要点:
      • WeatherNext: AI model achieves breakthrough in forecasting cyclones

ArXiv cs.AI (B_intro+search) 链接到标题

  • A Long-Run Persistence Theory for AI Systems under the Redundancy-Adjusted Artificial Age Score (AAS)

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04012v1 公告类型:新。 -摘要:人们越来越期望人工智能系统能够在交互、适应和更新的重复循环中运行,而不是通过孤立的一次性输出。
      • 这就提出了一个基本的理论问题:人工智能系统能否无限期地持续存在,而不会导致无限制的结构老化?
      • 本文基于冗余调整人工年龄评分(AAS),为人工智能系统开发了一个长期持久性框架。
    • EN 要点:
      • arXiv:2608.04012v1 Announce Type: new
      • Abstract: Artificial intelligence systems are increasingly expected to operate over repeated cycles of interaction, adaptation, and update rather than through i…
      • This raises a fundamental theoretical question: can an AI system persist indefinitely without incurring unbounded structural aging
      • This paper develops a long-run persistence framework for AI systems based on the redundancy-adjusted Artificial Age Score (AAS)
  • The LLM Proposes, the Executive Disposes: A Self-Verifying Agent Instrument that Dissociates Commitment Drift from Binding Drift in Long-Horizon Agents

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04066v1 公告类型:新。
      • 摘要:当长期代理的状态和自我报告完全无法信任时,如何验证其真实性?
      • 我们提出了一种代理工具,使验证是结构性的而不是事后的。
      • 一个确定性的执行者拥有所有的信念;语言模型只能提交键入的建议,并且只有在行动之前预先注册的预测与代码观察相匹配时,声明才会被承认。
    • EN 要点:
      • arXiv:2608.04066v1 Announce Type: new
      • Abstract: How do you verify a long-horizon agent when its own state and self-reports are exactly what you cannot trust
      • We present an agent instrument built so that verification is structural rather than post-hoc
      • A deterministic Executive owns all belief; a language model may only file typed proposals, and a claim is admitted only when a prediction pre-registered before…
  • Monte Carlo Tree Search for Table-to-Multimodal Report Generation

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04071v1 公告类型:新。
      • 摘要:从结构化表格数据自动生成包含文本分析和可视化图表的专业多模式报告是数据智能中的一项关键挑战。
      • 现有方法存在固定的线性管道和孤立的子任务处理,这阻碍了事实准确性、视觉质量和叙述连贯性的联合优化。
      • 为了解决这些问题,本文提出了 MCTS-Report,这是一种蒙特卡罗树搜索 (MCTS) 驱动的框架,它将多模式表到报告的生成制定为结构化搜索空间上的渐进式构建过程。
    • EN 要点:
      • arXiv:2608.04071v1 Announce Type: new
      • Abstract: Automatically generating professional multimodal reports comprising both textual analysis and visual charts from structured tabular data is a critical…
      • Existing methods suffer from fixed linear pipelines and isolated subtask processing, which hinder joint optimization of factual accuracy, visual quality, and na…
      • To address these issues, this paper proposes MCTS-Report, a Monte Carlo Tree Search (MCTS)-driven framework that formulates multimodal table-to-report generatio…
  • FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04077v1 公告类型:新。
      • 摘要:评估金融人工智能代理需要符合实际专业工作的标准。
      • 现有的评估标准方法通常从任务提示或模型输出中得出标准,而忽略了仅在从业者可交付成果中可见的默认标准。
      • 我们引入了 FinProBench(专业财务任务的基准)和基于角色的评分标准构建(RGRC),这是一个可重复使用的管道,可从具有相同角色的从业者生成的可交付成果中派生评分标准。
    • EN 要点:
      • arXiv:2608.04077v1 Announce Type: new
      • Abstract: Evaluating financial AI agents requires criteria aligned with real professional work
      • Existing rubric methods typically derive criteria from task prompts or model outputs, overlooking tacit standards visible only in practitioner deliverables
      • We introduce FinProBench, a benchmark for professional financial tasks, and Role-Grounded Rubric Construction (RGRC), a reusable pipeline that derives rubrics f…
  • FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04095v1 公告类型:新。 -摘要:大型语言模型(LLM)代理越来越多地用作财务咨询等高风险领域的个性化助手,但仍不清楚它们是否可以长期维护和更新个性化用户模型。
      • 现有的个性化记忆基准主要测试事实保留或依赖于弱约束的模型生成的轨迹,而事件驱动的偏好适应尚未得到充分探索。
      • 我们推出 FinPerMA,这是一个基于事件的基准,可根据冻结的纵向投资者轨迹评估个性化记忆。
    • EN 要点:
      • arXiv:2608.04095v1 Announce Type: new
      • Abstract: Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains u…
      • Existing personalized-memory benchmarks primarily test factual retention or rely on weakly constrained model-generated trajectories, leaving event-driven prefer…
      • We introduce FinPerMA, an event-grounded benchmark that evaluates personalized memory against frozen longitudinal investor trajectories
  • BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04156v1 公告类型:新。
      • 摘要:脑电图 (EEG) 分析不仅限于为录音分配预定义标签;它需要连接自然语言指令、信号处理、定量证据和科学解释的工作流程。
      • 我们将这种能力称为\emph{全面的脑电图理解}。
      • 然而,现有的评估主要针对孤立的解码任务或特定于系统的演示,使得大型语言模型(LLM)的能力没有得到充分量化。
    • EN 要点:
      • arXiv:2608.04156v1 Announce Type: new
      • Abstract: Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural-language inst…
      • We term this capability \emph{comprehensive EEG understanding}
      • Existing evaluations, however, primarily target isolated decoding tasks or system-specific demonstrations, leaving the competence of large language models (LLMs…
  • Adversarially Robust Abductive Fusion of Pre-trained Transformer-based Perception Models

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04190v1 公告类型:新。 -摘要:在新环境中部署预先训练的感知模型会降低分布变化下的准确性,并且单独组装它们并不能恢复它:诸如多数投票交易之类的组合器会为了精度而进行召回,并且很容易出现协调故障。
      • 先前的元认知方法学习标记模型错误的逻辑规则,但依赖于手工编写的领域知识线索(对象大小先验、分割掩模),这些线索不会转移到真正新颖的场景。
      • 我们表明,通过利用向量空间几何,可以在没有任何领域知识的情况下学习该元认知层:根据每个模型自己的训练嵌入构建的每个模型标签向量池(LVP),从相对于训练确定的原型的检测几何中产生错误检测规则,与领域知识规则达到同等水平,测试集上每个 F1 的成本在 0.002 美元以内。
    • EN 要点:
      • arXiv:2608.04190v1 Announce Type: new
      • Abstract: Deploying pre-trained perception models in novel environments degrades their accuracy under distributional shift, and assembling them alone does not r…
      • Prior metacognitive methods learn logical rules that flag a model’s errors, but rely on hand-authored domain-knowledge cues (object-size priors, segmentation ma…
      • We show that this metacognitive layer can be learned without any domain knowledge by exploiting vector-space geometry: per-model Label Vector Pools (LVP), built…
  • MatrAIx: Simulating the World with 8.3 Billion Persona Agents

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04205v1 公告类型:新。
      • 摘要:人类对人工智能系统和数字产品的评估成本高昂、速度缓慢且难以扩展。
      • 离线评估更具可扩展性,但通常会抽象出人类的多样性和交互行为。
      • 因此,我们推出了MatrAIx,这是一种人口规模的模拟用户评估基础设施,用于使用异构用户测试人工智能系统和数字产品。
    • EN 要点:
      • arXiv:2608.04205v1 Announce Type: new
      • Abstract: Human evaluation of AI systems and digital products is costly, slow, and difficult to scale
      • Offline evaluations are more scalable but often abstract away human diversity and interactive behavior
      • We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users
  • Interoceptive Attention as Dynamic Homeostatic Prioritization in a Foraging Agent

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04232v1 公告类型:新。
      • 摘要:生物系统必须在有限的感知带宽下调节竞争需求,其中锐化一种估计会消耗锐化其他估计的能力。
      • 因此,任何固定预算系统都必须决定在哪里分配其感知精度。
      • 我们在觅食剂中研究这一点,它必须满足多种身体需求才能生存,并以主动推理为模型。
    • EN 要点:
      • arXiv:2608.04232v1 Announce Type: new
      • Abstract: Biological systems must regulate competing needs under limited perceptual bandwidth, where sharpening one estimate costs the capacity to sharpen the o…
      • Any fixed-budget system therefore has to decide where to allocate its perceptual precision
      • We study this in a foraging agent that must keep several bodily needs satisfied to survive, modelled with active inference
  • The RAIL Principles for Neurosymbolic AI: Reasoning, Assurances, Interfacing and Learning

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04285v1 公告类型:新。
      • 摘要:集成机器学习和符号推理的神经符号人工智能系统正在迅速引起人们的关注。
      • 它们用符号推理算法补充了神经网络和语言模型的数据密集型统计方法,以在高风险领域或代表许多现实世界应用的低数据体系中发挥作用。
      • 我们认为,机器学习和形式推理的神经符号组合并不是人工智能中的小众方法,而是包括许多已经成功的技术,这些技术对于开发可靠、高效且最终值得信赖的系统至关重要。
    • EN 要点:
      • arXiv:2608.04285v1 Announce Type: new
      • Abstract: Neurosymbolic AI systems that integrate machine learning and symbolic reasoning are rapidly gaining attention
      • They complement the data-intensive statistical approaches of neural networks and language models with symbolic reasoning algorithms to function in high-stakes d…
      • We argue that the neurosymbolic combination of machine learning and formal reasoning is not a niche approach within AI, but rather includes many already success…

ArXiv cs.CL (B_intro+search) 链接到标题

  • TabletCraft: Bridging a 4,000-Year Cultural Gap with Bidirectional Akkadian NMT and Cuneiform Rendering

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.02609v1 公告类型:新。
      • 摘要:世界各地的博物馆中保存着 50 万块楔形文字泥板,但现代用户既无法使用世界上最古老的书写系统进行读写,留下了 4000 年的文化障碍,而现有的 NLP 工具仅部分解决了这一障碍。
      • 先前的工作实现了从阿卡德语到英语的单向、面向学者的翻译,但没有提供相反方向的路径:非专业用户无法用楔形文字撰写新内容,因此仍然是古代文化的被动消费者,而不是积极的参与者。
      • 我们推出 TabletCraft,这是第一个能够与美索不达米亚书写进行双向交互的开源系统。
    • EN 要点:
      • arXiv:2608.02609v1 Announce Type: new
      • Abstract: Half a million cuneiform clay tablets survive in museums worldwide, yet modern users can neither read nor write in the world’s oldest writing system,…
      • Prior work enables one-way, scholar-oriented translation from Akkadian to English, but offers no path in the reverse direction: non-specialist users cannot comp…
      • We present TabletCraft, the first open-source system that enables bidirectional interaction with Mesopotamian writing
  • BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.02612v1 公告类型:新。
      • 摘要:制定优化问题会极大地影响最终解决方案的质量,而良好的制定通常需要大量的专业知识。
      • 因此,最近的研究研究了如何从自然语言描述中自动导出优化问题,但现有基准侧重于目标和约束可以明确编写为数学表达式的设置。
      • 许多实际重要的问题自然被视为黑盒优化(BBO)问题,其中只能观察到客观值,而无法获得函数形式。
    • EN 要点:
      • arXiv:2608.02612v1 Announce Type: new
      • Abstract: Formulating an optimization problem strongly affects the quality of the final solution, yet good formulations usually require substantial expertise
      • Recent studies have therefore examined how to automatically derive optimization problems from natural-language descriptions, but existing benchmarks focus on se…
      • Many practically important problems are naturally treated as black-box optimization (BBO) problems, in which only objective values are observable, and the funct…
  • MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.02613v1 公告类型:新。
      • 摘要:边缘部署的个人记忆助手必须使用开放权重模型处理设备上的私人人际对话。
      • 然而,现有的记忆基准测试常常未能充分测试活动密集的交互、以自我为中心的视角和连贯的多会话世界的组合。
      • MemArena 通过其 MASim 代理模拟器构建的单一世界对话基准填补了这些空白,在 15 天内针对 50 个代理(1030 万个对话文本令牌、24100 个纯文本自我观察令牌/代理/天)。
    • EN 要点:
      • arXiv:2608.02613v1 Announce Type: new
      • Abstract: Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models
      • Yet, existing memory benchmarks often under-test the combination of activity-dense interaction, ego-centric perspective, and coherent multi-session worlds
      • MemArena fills these gaps with a single-world conversational benchmark built with its MASim agent simulator, for 50 agents over 15 days (10.3M dialog-text token…
  • OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.02615v1 公告类型:新。
      • 摘要:癌症诊断和表征需要整合来自放射学、病理学、基因组学和临床元数据的补充证据。
      • 然而,大多数医学大语言模型 (LLM) 和视觉语言模型 (VLM) 基准侧重于孤立的模式或狭窄的图像文本任务,使得跨多个证据流的患者级肿瘤学评估在很大程度上未经测试。
      • 我们推出 OncoTriad-QA,这是一种用于泛癌症问答的患者级放射学-病理学-基因组学基准。
    • EN 要点:
      • arXiv:2608.02615v1 Announce Type: new
      • Abstract: Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata
      • However, most medical large language model (LLM) and vision-language model (VLM) benchmarks focus on isolated modalities or narrow image-text tasks, leaving pat…
      • We introduce OncoTriad-QA, a patient-level radiology-pathology-genomics benchmark for pan-cancer question answering
  • Evaluating OpenAI’s Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.02616v1 公告类型:新。
      • 摘要:我们首次对 OpenAI 的隐私过滤器 (OPF)(一种 1.5B 参数双向 PII 检测器)进行了独立、系统的评估,涵盖 22 种语言和 5 个领域的 42 个综合基准。
      • 零样本,OPF 在 AI4Privacy 上实现 F1=0.855,在 SPY Medical 上实现 0.464,在 PII 注释基准上优于 Presidio(0.431, 0.273)和 XLM-RoBERTa(0.269, 0.111);在多语言 NER 上,XLM-RoBERTa 在所有 13 种印度语和非拉丁语上领先 OPF。
      • GPT-4o 在医疗、法律和金融 PII 方面领先(SPY:平均 0.643,Gretel:0.527),而 OPF 在结构化合成 PII(平均 0.71)和客户支持(平均 0.60)方面领先。
    • EN 要点:
      • arXiv:2608.02616v1 Announce Type: new
      • Abstract: We present the first independent, systematic evaluation of OpenAI’s Privacy Filter (OPF), a 1.5B-parameter bidirectional PII detector, across 42 synth…
      • Zero-shot, OPF achieves F1=0.855 on AI4Privacy and 0.464 on SPY medical, outperforming Presidio (0.431, 0.273) and XLM-RoBERTa (0.269, 0.111) on PII-annotated b…
      • GPT-4o leads on medical, legal, and financial PII (SPY: 0.643 avg, Gretel: 0.527), while OPF leads on structured synthetic PII (0.71 avg) and customer support (…
  • Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.02617v1 公告类型:新。
      • 摘要:我们使用 MOOVE(大规模开放在线验证和评估)的专家反馈来评估临床医生配对偏好是否在大语言模型(LLM)评估中提供可靠的临床安全信号,MOOVE 是一个临床医生主导的平台,收集盲法配对偏好以及多标准评分。
      • 临床医生按照离散的 $[-2, +2]$ 等级进行评分,其中负值表示临床上不安全或具有误导性的内容。
      • 使用来自 13 个法学硕士的 26{,}804 个成对判断,这些判断由来自 28 个以上国家/地区的超过 736 名临床医生提供,我们发现临床医生的偏好并不能很好地代表安全关键绩效。
    • EN 要点:
      • arXiv:2608.02617v1 Announce Type: new
      • Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert…
      • Clinicians assign scores on a discrete $[-2, +2]$ scale, where negative values indicate clinically unsafe or misleading content
      • Using 26{,}804 pairwise judgments across outputs from 13 LLMs, contributed by more than 736 clinicians across 28+ countries, we find that clinician preference i…
  • JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.02620v1 公告类型:新。
      • 摘要:LLM 作为法官评估已成为对语言模型进行排名的主导范例,但生态系统仍然支离破碎:大多数基准测试都有自己的代码库,对特定的封闭模型法官进行硬编码,并支持单一评估协议。
      • 这种碎片化使得研究设计选择(基准、判断模型、提示、推理后端)如何影响我们得出的有关模型质量的结论变得困难。
      • 我们推出 JudgeArena,这是一个开源框架,它将主要的 LLM 法官基准(AlpacaEval、Arena-Hard、MT-Bench 和 m-Arena-Hard)统一在一个界面下,具有可交换的法官和全面的元数据记录,以提高报告和可重复性的透明度。
    • EN 要点:
      • arXiv:2608.02620v1 Announce Type: new
      • Abstract: LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their…
      • This fragmentation makes it difficult to study how design choices–the benchmark, the judge model, the prompt, the inference backend–affect the conclusions we…
      • We introduce JudgeArena, an open-source framework that unifies major LLM-judge benchmarks (AlpacaEval, Arena-Hard, MT-Bench, and m-Arena-Hard) under a single in…
  • Knowing the Form, Not the Function: Automatically Auditing Answer–Authority Decoupling in Legal Benchmarks

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.02621v1 公告类型:新。
      • 摘要:即使模型也规定了法律权威,法律基准通常也会得出最终答案。
      • 我们测试答案的正确性是否可以作为权威基础的代理。
      • 在不要求法定引用的普通推理提示下,四位法学硕士自发地为 238 个台湾律师考试项目制作了权威标记。
    • EN 要点:
      • arXiv:2608.02621v1 Announce Type: new
      • Abstract: Legal benchmarks typically score final answers even when models also state legal authority
      • We test whether answer correctness can serve as a proxy for authority grounding
      • Under ordinary reasoning prompts that did not request statutory citations, four LLMs spontaneously produced authority markers across 238 Taiwan bar-examination…
  • Speculative Correction: Draft-then-Refine Decoding for Diffusion Language Models

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.02625v1 公告类型:新。 -摘要:扩散语言模型(DLM)可以双向修改标记,但标准解码程序通常通过逐块生成文本来使它们适应从左到右的生成。
      • 我们研究一种简单的即插即用推理模式:首先生成完整的草稿,然后使用双向扩散完善完整的响应。
      • 使用LLaDA2.1-Flash和LLaDA2.1-Mini,我们评估了两种配置。
    • EN 要点:
      • arXiv:2608.02625v1 Announce Type: new
      • Abstract: Diffusion language models (DLMs) can revise tokens bidirectionally, but standard decoding procedures often adapt them to left-to-right generation by p…
      • We study a simple plug-and-play inference pattern: first generate a complete draft, then refine the full response using bidirectional diffusion
      • Using LLaDA2.1-Flash and LLaDA2.1-Mini, we evaluate two configurations
  • Stuck on “A”: Diagnosing and Repairing Interface Injury in Attention-to-KDA Linearization of a 0.6B Language Model

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.02689v1 公告类型:新。
      • 摘要:我们在单个消费级 GPU 预算上将 Qwen3-0.6B-Base 的 28 个全注意力层中的 21 个转换为 KDA(Kimi Delta Attention)线性注意力层,并提出一个简单的问题:转换到底会破坏什么?
      • 手术后,隐藏状态对齐和端到端 KL 蒸馏使学生在困惑中接近老师,但多项选择的准确性保持接近随机机会(25-29% vs.
      • 老师的 C-Eval 得分为 50.6%)。
    • EN 要点:
      • arXiv:2608.02689v1 Announce Type: new
      • Abstract: We convert 21 of 28 full-attention layers of Qwen3-0.6B-Base into KDA (Kimi Delta Attention) linear-attention layers on a single consumer-grade GPU bu…
      • After surgery, hidden-state alignment and end-to-end KL distillation drive the student close to its teacher in perplexity, yet multiple-choice accuracy stays ne…
      • the teacher’s 50.6% on C-Eval)

ArXiv cs.LG (B_intro+search) 链接到标题

  • C$^2$MOE: Consistency and Complementarity-guided Mixture of Experts for Incomplete Multimodal Emotion Learning

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04013v1 公告类型:新。
      • 摘要:对话中多模态情绪识别(MERC)的最新进展凸显了其对完整多模态输入的依赖。
      • 然而,现实世界的数据经常由于传输错误或用户行为而丢失模态,从而严重降低模型性能。
      • 现有方法通过跨模态一致性学习增强鲁棒性,但很大程度上忽略了模态互补性,导致重建有偏差。
    • EN 要点:
      • arXiv:2608.04013v1 Announce Type: new
      • Abstract: Recent advances in Multimodal Emotion Recognition in Conversations (MERC) highlight its reliance on complete multimodal inputs
      • However, real-world data often suffer from missing modalities due to transmission errors or user behavior, severely degrading model performance
      • Existing methods enhance robustness via cross-modal consistency learning but largely ignore modality complementarity, leading to biased reconstructions
  • On Hamming-Lipschitz Type Stability of the Subdominant (Minmax) Ultrametric: Theory and Simple Proofs

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04014v1 公告类型:新。 -摘要:次主导(最小最大)超度量是相异矩阵的规范树结构摘要,与单链接聚类引发的超度量等效。
      • 虽然其经典稳定性理论通常用 $\ell_\infty$ 或 Gromov–Hausdorff 项来表述,但这样的界限不太适合仅改变几个成对距离的稀疏扰动。
      • 我们为该算子开发了 $\ell_0$ 型稳定性理论。
    • EN 要点:
      • arXiv:2608.04014v1 Announce Type: new
      • Abstract: The subdominant (minmax) ultrametric is a canonical tree-structured summary of a dissimilarity matrix, arising equivalently as the ultrametric induced…
      • While its classical stability theory is usually formulated in $\ell_\infty$ or Gromov–Hausdorff terms, such bounds are poorly suited to sparse perturbations th…
      • We develop an $\ell_0$-type stability theory for this operator
  • A Trust-region Framework for Moment Estimation

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04026v1 公告类型:新。
      • 摘要:在本文中,我们开发了一个信任域框架,用于理解随机梯度优化中自适应矩估计机制(例如 \textsc{Adam})的行为。
      • 具体来说,在这个框架中,每个单独权重的更新步骤的大小被限制在由 $p\in[2,4]$ 阶矩约束控制的信任区域内。
      • 由此产生的推导得出一系列基于二阶矩估计和标准化 $p$ 矩估计的学习率机制。
    • EN 要点:
      • arXiv:2608.04026v1 Announce Type: new
      • Abstract: In this paper, we develop a trust-region framework for understanding the behavior of adaptive moment estimation mechanisms, such as \textsc{Adam}, in…
      • Specifically, in this framework, the magnitude of the update step for each individual weight is constrained within a trust-region governed by a moment constrain…
      • The resulting derivation then leads to a family of learning-rate mechanisms based on second-moment estimation and a normalized $p$-th moment estimation
  • Learning to Resolve Neutron Resonances with Fully Convolutional Neural Networks

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04027v1 公告类型:新。
      • 摘要:这项工作研究了使用强大的机器学习框架增强传统 R 矩阵代码以自动检测透射谱中的中子共振的可行性。
      • 中子传输数据通常复杂且嘈杂,因此很难使用传统的峰值识别方法进行分析。
      • 物理学家目前用来拟合这些数据的最先进的 R 矩阵代码通常取决于先前的评估,并且需要大量的手动工作。
    • EN 要点:
      • arXiv:2608.04027v1 Announce Type: new
      • Abstract: This work investigates the feasibility of augmenting traditional R-Matrix codes with a robust machine learning framework for automatically detecting n…
      • Neutron transmission data are often complex and noisy, making them difficult to analyze using traditional peak-identification methods
      • The state-of-the-art R-Matrix codes currently used by physicists to fit these data often depend on prior evaluations and require substantial manual effort
  • Lindblad-Inspired Multi-Timescale Reservoir Computing with Separable Rotation and Dissipation

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04028v1 公告类型:新。
      • 摘要:回声状态网络通过修复循环动态并仅训练线性读数来实现高效的时间学习。
      • 然而,传统的储存器通常在单个随机循环矩阵内容纳信号混合、记忆保留和稳定性。
      • 现有的结构化设计改进了拓扑、范数保存、泄漏或深度,但通常不提供可逆混合和不可逆遗忘的单独模态控制以及直接的全局稳定性保证。
    • EN 要点:
      • arXiv:2608.04028v1 Announce Type: new
      • Abstract: Echo-state networks enable efficient temporal learning by fixing the recurrent dynamics and training only a linear readout
      • However, conventional reservoirs typically accommodate signal mixing, memory retention, and stability within a single random recurrent matrix
      • Existing structured designs improve topology, norm preservation, leakage, or depth, but generally do not provide separate modal control of reversible mixing and…
  • An Explainable LLM Agent Layer for Open-World Anomaly Detection in Oil Wells

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04041v1 公告类型:新。
      • 摘要:最近表明,用于油井异常检测的开放世界学习 (OWL) 管道在公共 3W 数据集上结合了基于自动编码器的检测、多类分类和基于 Mahalanobis 的新颖性检测。
      • 这些管道回答 \textit{发生了什么},但它们不解释 \textit{为什么模型相信它} 或 \textit{操作员下一步应该做什么},并且它们不会在他们发现的新奇簇上放置人类可读的名称。
      • 本文评估了位于 OWL 管道下游的大型语言模型 (LLM) 代理层,该代理层被设计为已发布的上游方法的 \textbf{companion},而不是替代品。
    • EN 要点:
      • arXiv:2608.04041v1 Announce Type: new
      • Abstract: Open-World Learning (OWL) pipelines for oil well anomaly detection have recently been shown to combine autoencoder-based detection, multiclass classif…
      • These pipelines answer \textit{what happened}, but they do not explain \textit{why the model believes it} or \textit{what the operator should do next}, and they…
      • This paper evaluates a Large Language Model (LLM) agent layer placed downstream of the OWL pipeline, designed as a \textbf{companion} to the published upstream…
  • Tactus: Open-Vocabulary Object Recognition from Low-Cost Pressure Arrays

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04043v1 公告类型:新。
      • 摘要:电阻压力阵列是最便宜且运输最广泛的触觉传感器,但触觉表示学习主要集中在对变形凝胶进行成像的光学传感器上。
      • 我们提出了 Tactus,一个仅根据压力数据回答文本查询的开放模型:在 STAG 基准(27 个对象,保留的记录)上,它在四次运行中达到 0.771 +/- 0.062 top-1(top-3 0.935),匹配并且最好超过数据集的有监督闭集 CNN(没有经过训练的分类器头)为 0.76。
      • 秘诀是小数据:187 个训练记录、在 144k 个未标记的相同传感器帧上进行掩蔽自动编码器预训练,以及传感器自己的校准仿射,它恢复的精度比每个架构更改的总和更高。
    • EN 要点:
      • arXiv:2608.04043v1 Announce Type: new
      • Abstract: Resistive pressure arrays are the cheapest and most widely shipped tactile sensors, yet tactile representation learning has concentrated on optical se…
      • We present Tactus, an open model that answers text queries from pressure data alone: on the STAG benchmark (27 objects, held-out recordings), it reaches 0.771 +…
      • The recipe is small-data: 187 training recordings, masked-autoencoder pretraining on 144k unlabeled same-sensor frames, and the sensor’s own calibration affine,…
  • Robust and Personalized Federated Learning for Aircraft-Engine Prognostics under Benign and Adversarial Client Heterogeneity

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04045v1 公告类型:新。
      • 摘要:联合学习 (FL) 使飞机机队运营商能够通过发动机传感器遥测技术联合训练剩余使用寿命 (RUL) 模型,而无需共享原始数据。
      • 这项研究研究了两个互补的挑战:良性异质性(诚实的操作员观察不同的操作条件和故障模式)和对抗性异质性(其中受损的操作员提交有毒更新)。
      • 我们使用多任务一维卷积神经网络和商业模块化航空推进系统仿真 (C-MAPSS) 基准的结构非 IID 分区进行受控的、以安全为导向的评估。
    • EN 要点:
      • arXiv:2608.04045v1 Announce Type: new
      • Abstract: Federated learning (FL) enables aircraft fleet operators to jointly train remaining-useful-life (RUL) models from engine sensor telemetry without shar…
      • This study examines two complementary challenges: benign heterogeneity, where honest operators observe different operating conditions and fault modes, and adver…
      • We conduct a controlled, safety-oriented evaluation using a multi-task one-dimensional convolutional neural network and a structurally non-IID partition of the…
  • Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04048v1 公告类型:新。
      • 摘要:在不同的部署约束下为大型语言模型 (LLM) 提供服务需要在准确性、内存占用和吞吐量之间进行灵活的权衡。
      • 然而,传统的量化方法通常需要针对每个目标位宽的单独检查点。
      • 我们引入了循环残差量化(RRQ),这是一种训练后量化(PTQ)框架,它将权重表示为低位量化基础以及一系列量化残差校正,从而从单个检查点实现多个有效精度。
    • EN 要点:
      • arXiv:2608.04048v1 Announce Type: new
      • Abstract: Serving large language models (LLMs) under diverse deployment constraints requires flexible trade-offs between accuracy, memory footprint, and through…
      • However, conventional quantization methods typically require a separate checkpoint for each target bit-width
      • We introduce Recurrent Residual Quantization (RRQ), a post-training quantization (PTQ) framework that represents weights as a low-bit quantized base together wi…
  • CAMP: A Cycle-Aware Multi-Scale Patch Mixer for Time Series Forecasting

    • 发布时间:2026-08-06 12:00 北京时间
    • 摘要:- arXiv:2608.04051v1 公告类型:新。
      • 摘要:现实世界的时间序列通常受重复模式的控制,但其主导周期可能因数据集、预测设置和单个输入窗口而异。
      • 现有的周期感知预测器通常依赖于在数据集级别选择的单个周期,当周期性行为随时间变化或多个周期共存时,这可能会受到限制。
      • 此外,基于补丁的模型通常统一处理所有补丁位置,尽管远离预测边界的补丁可能需要更广泛的上下文细化,而最近的补丁包含应该更直接保留的信息。
    • EN 要点:
      • arXiv:2608.04051v1 Announce Type: new
      • Abstract: Real-world time series are often governed by recurring patterns, but their dominant periods may vary across datasets, forecasting settings, and indivi…
      • Existing cycle-aware forecasters commonly rely on a single period selected at the dataset level, which can be restrictive when periodic behavior changes over ti…
      • Moreover, patch-based models typically process all patch positions uni- formly, although patches farther from the forecast boundary may require broader contextu…