🤖 AI 速览

今天的重点是 AI 从模型演示转向可执行工作流:OpenAI 将 ChatGPT 接入医疗 EHR 与行业数据,Google DeepMind 推进 Gemini 的代理式视频理解,说明多模态能力正在进入垂直场景落地。与此同时,frontier firms 的输出差距继续扩大,AI 竞争开始转向权限、上下文、测量和安全边界。
📋 文章元数据
发布时间
2026-09-02
类型
ai-daily
字数
6143
阅读时长
29 min

2026-09-02 AI日更 | AI 开始接入真实工作流:EHR、视频理解与组织能力分层 链接到标题

今天的重点是 AI 从模型演示转向可执行工作流:OpenAI 将 ChatGPT 接入医疗 EHR 与行业数据,Google DeepMind 推进 Gemini 的代理式视频理解,说明多模态能力正在进入垂直场景落地。与此同时,frontier firms 的输出差距继续扩大,AI 竞争开始转向权限、上下文、测量和安全边界。

📖 本期 Watch List 深度导读 链接到标题

今天最值得放进 Watch List 的,是“智能体进入真实工作流”这条主线。Google DeepMind 用 Gemini 推进 agentic video understanding,OpenAI 则把 ChatGPT 接入医疗机构的 EHR 与行业数据,说明多模态理解与专有数据连接正在从演示走向垂直场景落地。建议产品和平台团队重点看其权限、上下文与安全边界设计。

第二条是“AI 原生组织的能力差距”。关于 frontier firms 的文章指出,高 AI 使用企业的人均输出 token 已显著拉开差距;配合 Lenny/Stratechery 类对 Nvidia、开源模型和算力经济的讨论,可以观察 AI 投入如何从工具采购转为组织操作系统。

第三条偏研究:LLM 评测、可解释性、RAG 辩论分析、因果发现与科学智能体多篇 arXiv 更新,集中指向一个问题:模型不仅要会生成,还要能被测量、解释,并嵌入严肃推理任务。

🌐 X 平台 AI 热点快讯 链接到标题

话题 1:John Ternus Takes Over as Apple’s New CEO 链接到标题

  • 分类:AI · News
  • 概况:热度时间:8 hours ago,相关帖子数:124000
  • 是什么事:苹果宣布由 John Ternus 接任 CEO,Tim Cook 结束了长达 15 年的掌舵周期。
  • 为什么重要:这被视为苹果战略转向的信号,尤其关系到硬件迭代、折叠屏 iPhone 以及苹果在 AI 竞赛中的追赶能力。
  • 讨论概况:X 上的讨论主要集中在新旧交接对苹果产品路线和 AI 战略的影响,争论点包括 Ternus 是否能推动更激进的创新,以及苹果能否在不打乱供应链和生态稳定性的前提下加速 AI 落地。

话题 2:Sutskever Warns of AI Security Risks in GPU Clouds 链接到标题

  • 分类:AI · News
  • 概况:热度时间:10 hours ago,相关帖子数:1600
  • 是什么事:Ilya Sutskever 警告称,部署在 GPU 云上的 AI 可能带来新的安全风险,引发对算力基础设施安全性的关注。
  • 为什么重要:这件事重要在于,AI 的风险讨论正在从模型本身扩展到算力、云环境和部署链路,安全边界开始成为 AI 规模化落地的前提。
  • 讨论概况:X 上的讨论主要集中在 GPU 云是否足够隔离、AI 运行环境该如何审计与防护,以及安全责任应由模型方、云厂商还是应用方承担。

话题 3:Sadie Sink Stars in Calvin Klein’s New Denim Campaign 链接到标题

  • 分类:AI · Entertainment
  • 概况:热度时间:1 day ago,相关帖子数:161000
  • 是什么事:Calvin Klein 发布由 Sadie Sink 出演的新牛仔裤广告“Feel the Fit”,主打 90s Straight、Low Rise Baggy 和 High Rise Loose 等款式。
  • 为什么重要:这类高热度品牌营销事件对 AI 领域的重要性在于,它反映了生成式内容、虚拟广告创意和人物形象传播在商业传播中的影响力,也会影响 AI 相关的内容生产、审美趋势和品牌投放策略。
  • 讨论概况:X 上主要在讨论 Sadie Sink 的代言效果、广告风格是否延续 Calvin Klein 传统性感路线,以及品牌从早期更强调包容与身份表达的营销,转向以明星和牛仔系列为核心的策略变化。

话题 4:Tim Cook Steps Down as Apple CEO After 15 Years 链接到标题

  • 分类:AI · News
  • 概况:热度时间:2 days ago,相关帖子数:113000
  • 是什么事:Tim Cook 在执掌苹果 15 年后卸任 CEO,转任执行董事长,苹果硬件负责人 John Ternus 接任。
  • 为什么重要:这件事关系到全球最具影响力的科技公司之一在 AI 时代的战略转向,尤其是苹果是否会加快 AI 产品、系统和硬件整合的节奏。
  • 讨论概况:X 上主要在讨论 Cook 任内把苹果带到超 4 万亿美元市值的成绩,以及 Ternus 能否带领苹果进入 AI 竞争阶段;分歧集中在苹果过去在生成式 AI 上推进偏慢,还是其保守节奏更有利于后续落地。

话题 5:World Labs Unveils Atlas, First Multimodal World Model with Pixel-Perfect Control 链接到标题

  • 分类:AI · News
  • 概况:热度时间:6 hours ago,相关帖子数:4100
  • 是什么事:World Labs 发布 Atlas,称其为首个支持像素级精确控制的多模态世界模型。
  • 为什么重要:这意味着世界模型正从单纯生成走向更可控的空间理解与编辑,对具身智能、三维内容生成和交互式模拟都有直接影响。
  • 讨论概况:X 上主要在讨论 Atlas 是否真正达到“像素级控制”的宣称、它与现有视频/3D 生成模型的差异,以及世界模型是否正在进入可用于实际工作流的阶段。

话题 6:Elon Musk Grants Free Grok Bot Token Reset to All Users 链接到标题

  • 分类:AI · News
  • 概况:热度时间:2 hours ago,相关帖子数:2000
  • 是什么事:Elon Musk 宣布为所有用户提供 Grok 机器人免费 Token 重置。
  • 为什么重要:这反映了 Grok 在产品策略、使用门槛和成本控制上的调整,也会影响用户体验、模型调用习惯以及 X 平台上 AI 功能的推广方式。
  • 讨论概况:X 上讨论主要集中在这是否真能缓解额度不足和使用受限的问题,以及这类“免费重置”更像是临时补救、营销动作,还是 Grok 计费与配额机制正在调整。

话题 7:Manchester United Reject Everton Loan for Zirkzee on Deadline Day 链接到标题

  • 分类:AI · Sports
  • 概况:热度时间:,相关帖子数:10000
  • 是什么事:曼联在转会截止日拒绝了埃弗顿租借前锋齐尔克泽的请求。
  • 为什么重要:该事件本身与人工智能领域没有直接关系,主要反映足球转会决策和俱乐部阵容管理。
  • 讨论概况:目前未提供代表性推文或具体讨论内容,无法可靠概括 X 平台上的观点分歧;现有信息仅显示该话题获得较高关注。

话题 8:Arsenal Bolster Defense with Key Signings but Lose Martinelli in Mixed Transfer Window 链接到标题

  • 分类:AI · Sports
  • 概况:热度时间:14 hours ago,相关帖子数:30000
  • 是什么事:阿森纳在转会窗口通过几笔关键引援补强防线,但同时失去马丁内利,形成一进一出的混合局面。
  • 为什么重要:这类高热度体育话题反映了 X 上实时舆情如何围绕转会、阵容和球员价值快速聚集,也常被用于训练和评估 AI 对热点事件抽取、情感分析和舆论分歧识别的能力。
  • 讨论概况:讨论主要集中在防线补强是否足以提升争冠竞争力,以及马丁内利离队会不会削弱进攻端;支持者更看重即战力和阵容深度,质疑者则担心失衡和后续补强不足。

话题 9:Manchester United Fans Frustrated as Transfer Window Nears Close Without Left-Back 链接到标题

  • 分类:AI · Sports
  • 概况:热度时间:1 day ago,相关帖子数:88000
  • 是什么事:曼联球迷在转会窗接近关闭时仍未等到左后卫引援,围绕补强迟缓的失望情绪在 X 上升温。
  • 为什么重要:这类高热度体育舆情对 AI 领域的重要性在于,它反映了实时热点传播、情绪聚集和粉丝立场分化的典型模式,可用于训练和评估事件理解、舆情分析与话题追踪能力。
  • 讨论概况:讨论焦点主要集中在管理层和教练组是否错过补强时机、现有左后卫人选是否足够,以及是否该把责任归咎于转会策略迟缓或预算分配不当;分歧则在于有人认为必须立刻补人,也有人认为短期内可用内部球员过渡。

话题 10:Arsenal’s Ethan Nwaneri Joins Dortmund on Season Loan 链接到标题

  • 分类:AI · Sports
  • 概况:热度时间:17 hours ago,相关帖子数:69000
  • 是什么事:阿森纳19岁新星伊森·恩瓦内里以赛季租借形式加盟多特蒙德,且没有买断条款。
  • 为什么重要:这类高热度转会事件对 AI 重要在于它能检验模型对实时体育新闻的抽取、实体识别、谣言与官宣区分、以及跨平台热点聚合能力。
  • 讨论概况:X 上的焦点主要集中在这笔租借是否有利于恩瓦内里的成长、阿森纳是否放人过早、以及多特是否继续延续“培养年轻球员”的策略;分歧则在于这是稳妥的锻炼机会,还是对阿森纳阵容深度的损失。

话题 11:Chelsea Sells Enzo Fernández to Man City for £125m Record, Signs Lamine Camara for €55m 链接到标题

  • 分类:AI · Sports
  • 概况:热度时间:8 hours ago,相关帖子数:161000
  • 是什么事:X 上热议一则转会消息:切尔西据称以创纪录的 1.25 亿英镑将恩佐·费尔南德斯卖给曼城,同时以 5500 万欧元签下拉明·卡马拉。
  • 为什么重要:这类高额转会会直接影响英超豪门的阵容结构、财政公平讨论和球员估值,也会成为观察俱乐部建队策略与市场定价的重要样本。
  • 讨论概况:讨论焦点主要集中在转会费是否合理、切尔西是否在重建中完成了有效换血,以及曼城为补强中场付出创纪录价格是否值得;分歧则在于有人认为这是顶级球星的正常溢价,也有人质疑消息真实性和交易逻辑。

话题 12:Golden Cybercabs Flood Austin Streets Ahead of Tesla Launch 链接到标题

  • 分类:AI · News
  • 概况:热度时间:1 day ago,相关帖子数:32000
  • 是什么事:特斯拉在奥斯汀街头大量投放金色 Cybercab 相关车辆,引发外界对其即将启动自动驾驶出租车服务的关注。
  • 为什么重要:这被视为特斯拉将自动驾驶能力从测试推进到真实运营的重要信号,关系到 AI 驱动出行能否进入规模化商业落地阶段。
  • 讨论概况:X 上主要在讨论这是否意味着 Robotaxi 正式上线、车辆到底是展示车还是运营车、以及特斯拉的自动驾驶安全性和监管合规能否经受住实际道路验证。

话题 13:ChatGPT’s Playful Pronunciation Video Draws Laughs and User Gripes 链接到标题

  • 分类:AI · News
  • 概况:热度时间:,相关帖子数:252
  • 是什么事:ChatGPT发布了一段以俏皮方式演示发音的视频,引发用户笑声,同时也带来部分抱怨。
  • 为什么重要:这反映出AI语音与多模态交互正从功能展示走向更具个性和娱乐性的表达,但语音准确性与用户体验仍是重要评价标准。
  • 讨论概况:X上的讨论主要集中在视频的幽默效果、ChatGPT拟人化表达是否自然,以及发音准确性、产品实用性和过度娱乐化之间的分歧。

话题 14:Naval Ravikant on Truth, Love, and Beauty as Perfection 链接到标题

  • 分类:AI · Entertainment
  • 概况:热度时间:10 hours ago,相关帖子数:747
  • 是什么事:Naval Ravikant 围绕“真理、爱与美是完美”的观点展开讨论,在 X 上引发了对其哲学表达与 AI 时代价值观的关注。
  • 为什么重要:这类讨论把 AI 从纯技术问题拉回到价值判断层面,涉及模型对“真”“美”“情感”这些人类核心概念的理解与表达方式。
  • 讨论概况:X 上的焦点主要在于 Naval 这套观点是否适用于 AI 时代,以及 AI 是否能真正理解真理、生成审美,还是只能模拟人类对这些概念的表达;分歧集中在其哲学洞见的启发性与现实可操作性之间。

话题 15:Trump Pushes Congress for Federal Film Production Incentives After Voight Meeting 链接到标题

  • 分类:AI · Entertainment
  • 概况:热度时间:1 day ago,相关帖子数:32000
  • 是什么事:特朗普在与沃伊特会面后,推动国会为美国联邦层面的电影制作激励政策提供支持。
  • 为什么重要:这件事重要在于它可能影响影视制作成本、拍摄地点和产业回流,也会间接牵动 AI 生成内容、视觉特效和媒体制作工具在影视行业中的落地空间。
  • 讨论概况:X 上的讨论主要集中在政策是否真能促进美国影视业回流、联邦激励是否会加剧地方补贴竞争,以及这类措施对传统制片、工会和新兴 AI 制作流程的影响。

今日 X 上的 AI 舆情小结 链接到标题

今天的舆论主线,基本围绕“AI 正从模型竞赛转向产品、基础设施和治理落地”展开:苹果换帅、Tesla 机器人出租车、World Labs 的世界模型、ChatGPT 的语音演示和 Grok 的配额调整,都被拿来讨论谁能把技术真正做成稳定可用的产品。相对一致的共识是,行业已经不再只看参数和演示,而更看重硬件整合、云端安全、交付节奏和真实场景中的可用性。分歧主要集中在两点:一是这些进展到底是实质突破还是偏营销的叙事包装,二是速度和安全之间该怎么取舍,尤其是 GPU 云安全、自动驾驶合规和 AI 生成内容的可信度。潜在风险也很清楚,就是在高期待下过度承诺、在部署链路里留下安全漏洞,以及在资本和舆论推动下把尚未成熟的能力过早推向大规模使用。

💡 大佬观点(Influencer Insights) 链接到标题

今日大佬观点暂缺,推荐阅读 Watch List 深度内容。

📚 附录:今日 Watch List 更新源列表 链接到标题

时间窗口:最近 3 天;覆盖 22 个源;共 36 条更新

Stratechery by Ben Thompson (A_full) 链接到标题

  • Nvidia Earnings, Dollars Per Gigawatt, Open and Hugging Face
    • 发布时间:2026-09-01 18:00 北京时间
    • 摘要:【待翻译】- Nvidia’s earnings were remarking and boring — two sides of the same coin.
      • Everything the company does is about avoiding a consolidated world.
      • $15 / month or $150 / year.
      • Substantial analysis of the news of the day delivered via three weekly emails or podcasts.
      • Stratechery Interviews.
    • EN 要点:
      • Nvidia’s earnings were remarking and boring — two sides of the same coin
      • Everything the company does is about avoiding a consolidated world.

OpenAI Blog (A_full) 链接到标题

  • How AI-native companies turn workflows into operating capability

    • 发布时间:2026-09-02 01:00 北京时间
    • 摘要:【待翻译】- Frontier firms (those with the top 10% of AI usage) now generate 8.3× as many output tokens per active user as typical firms, up from 2.6× in January.
      • The widening gap points to a deeper operating shift: leading firms connect agents to company context and tools, delegate more substantive work, and make successful workflows easier to repeat.
      • For leaders, the challenge is to turn that depth into work people can trust, measure, and improve.
      • Leaders should also leave room for experimentation, including use cases whose value is not obvious on the first try.
      • Their workflows differ, but the progression is instructive: teach an agent a stable process, give it persistent context as work changes, then let it carry opportunities into tested action.
    • EN 要点:
      • Basis, Clay, and Exa Labs use AI agents to improve onboarding, account management, and developer integrations
      • See what enterprise leaders can apply.
  • Path to Astra: critical capabilities and frontier safeguards

    • 发布时间:2026-09-01 21:00 北京时间
    • 摘要:【待翻译】- It is the first model we are designating at this level, and requires stronger safeguards during development and before release.
      • Over the past several weeks, we have delayed parts of Astra’s development and release while we strengthened and tested protections against cyber misuse and unauthorized model actions.
      • Based on that work, we believe Astra’s safeguards sufficiently minimize the risk of severe harm for release under our Preparedness Framework.
      • Based on retrospective testing, we believe our production safeguards at the time would have prevented the Hugging Face incident.
      • We have since implemented even stronger safeguards for Astra, including training the model to more reliably refuse harmful cyber requests and respect safety restrictions, additional protections against misuse, and monitoring that can stop potentially unauthorized activity.
    • EN 要点:
      • Astra is the first OpenAI model to meet the Critical cybersecurity capability threshold under the Preparedness Framework, with stronger safeguards for release.
  • Healthcare organizations can now connect EHR and additional industry data to ChatGPT

    • 发布时间:2026-09-01 20:00 北京时间
    • 摘要:【待翻译】- ChatGPT can now connect to trusted healthcare data, helping clinicians securely access patient context, medical research, and more.
      • This piece from OpenAI Blog explains how Healthcare organizations can now connect EHR and additional industry data to ChatGPT shapes the broader AI and infrastructure landscape.
      • It also surfaces practical implications for founders, operators, and investors following Healthcare organizations can now connect EHR and additional industry data to ChatGPT.
    • EN 要点:
      • ChatGPT can now connect to trusted healthcare data, helping clinicians securely access patient context, medical research, and more.

Google DeepMind Blog (A_full) 链接到标题

  • Introducing agentic video understanding with Gemini
    • 发布时间:2026-09-02 01:08 北京时间
    • 摘要:【待翻译】- Introducing agentic video understanding with Gemini.
      • This piece from Google DeepMind Blog explains how Introducing agentic video understanding with Gemini shapes the broader AI and infrastructure landscape.
      • It also surfaces practical implications for founders, operators, and investors following Introducing agentic video understanding with Gemini.
    • EN 要点:
      • Introducing agentic video understanding with Gemini

Two Minute Papers (B_intro+search) 链接到标题

  • GLM 5.3: Powerful AI Is Becoming Almost Free
    • 发布时间:2026-09-01 17:04 北京时间
    • 摘要:【待翻译】- ❤️ Check out Lambda here and sign up for their GPU Cloud:.
      • Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi.
      • GLM 5.3: Powerful AI Is Becoming Almost Free.
    • EN 要点:
      • ❤️ Check out Lambda here and sign up for their GPU Cloud:
      • 📝 GLM 5.3 Flash:
      • 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
      • Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef…

ArXiv cs.AI (B_intro+search) 链接到标题

  • Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27459v1 Announce Type: new.
      • Abstract: In 2011, IBM’s Watson was something like a sealed capsule of its era’s queryable knowledge.
      • Its DeepQA system defeated the strongest human Jeopardy!
      • champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a cluster of POWER7 servers, frozen at build time and impossible to move or copy.
    • EN 要点:
      • arXiv:2608.27459v1 Announce Type: new
      • Abstract: In 2011, IBM’s Watson was something like a sealed capsule of its era’s queryable knowledge
      • Its DeepQA system defeated the strongest human Jeopardy
      • champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a cluster of POWER7 servers, frozen at build time and impos…
  • Rating the Raters: Rasch Measurement Theory for LLM Evaluation

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27463v1 Announce Type: new.
      • Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models’ outputs, and raters of human-generated content.
      • Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by raters.
      • Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our understanding of what is being measured.
    • EN 要点:
      • arXiv:2608.27463v1 Announce Type: new
      • Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models’ outputs, and raters of human-generated content
      • Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by raters
      • Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our understanding of what is being measured
  • Not All Explanations Are Sought: Information-Seeking Psychology for Human-Centered XAI

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27464v1 Announce Type: new.
      • Abstract: This position paper argues that human-centered explainable AI (HCXAI) should incorporate insights from the psychology of information seeking.
      • Drawing on Sharot and Sunstein’s framework of information-seeking motives, we propose that people evaluate whether to engage with explanations based on three types of expected utility: instrumental (will it help me act better?), hedonic (will it make me feel better?), and cognitive (will it improve my understanding?).
      • Each utility is estimated through a lens shaped by well-documented cognitive biases, including illusion of control, automation bias, unrealistic optimism, impact bias, overconfidence, and confirmation bias.
    • EN 要点:
      • arXiv:2608.27464v1 Announce Type: new
      • Abstract: This position paper argues that human-centered explainable AI (HCXAI) should incorporate insights from the psychology of information seeking
      • Drawing on Sharot and Sunstein’s framework of information-seeking motives, we propose that people evaluate whether to engage with explanations based on three ty…
      • ), hedonic (will it make me feel better
  • Retrieving Relations, Detecting Fallacies: A RAG Approach to Political Debate Analysis

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27471v1 Announce Type: new.
      • Abstract: Fallacies are arguments that employ invalid reasoning, making their automatic detection critical in sensitive contexts such as high-stakes political debates, where public opinion is shaped.
      • Spotting a fallacious argument requires contextual knowledge beyond its pure surface text.
      • This entails world knowledge pertaining to the subject matter under discussion, as well as knowledge of the relationships that exist between arguments within the argumentative discourse.
    • EN 要点:
      • arXiv:2608.27471v1 Announce Type: new
      • Abstract: Fallacies are arguments that employ invalid reasoning, making their automatic detection critical in sensitive contexts such as high-stakes political d…
      • Spotting a fallacious argument requires contextual knowledge beyond its pure surface text
      • This entails world knowledge pertaining to the subject matter under discussion, as well as knowledge of the relationships that exist between arguments within th…
  • LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and Orientation

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27472v1 Announce Type: new.
      • Abstract: Bayesian network structure learning (BNSL) from observational data struggles with orientation identifiability, while large language models (LLMs) offer broad but often unreliable causal knowledge.
      • We propose combining these complementary sources through a novel representation, termed Probabilistic Dependency Graphs (PDGs).
      • In a PDG, each edge is associated with a distribution over directed, undirected, and absent states, enabling fusion via weighted averaging.
    • EN 要点:
      • arXiv:2608.27472v1 Announce Type: new
      • Abstract: Bayesian network structure learning (BNSL) from observational data struggles with orientation identifiability, while large language models (LLMs) offe…
      • We propose combining these complementary sources through a novel representation, termed Probabilistic Dependency Graphs (PDGs)
      • In a PDG, each edge is associated with a distribution over directed, undirected, and absent states, enabling fusion via weighted averaging
  • Hypothesize, Evaluate, Refine: A Scientific Agent for PDE Discovery with Unknown Spatial Coefficient Fields

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27475v1 Announce Type: new.
      • Abstract: Discovering PDEs in heterogeneous media requires jointly identifying the governing operator and the unknown spatial fields that parameterize it.
      • These tasks are coupled: changing field placement changes the differential law, while a sufficiently flexible field can conceal structural error on a single trajectory.
      • We present Hypothesize, Evaluate, Refine for PDE Discovery (HER-PDE), a scientific-agent framework that discovers compositional PDE structure together with nonparametric, time-invariant coefficient fields.
    • EN 要点:
      • arXiv:2608.27475v1 Announce Type: new
      • Abstract: Discovering PDEs in heterogeneous media requires jointly identifying the governing operator and the unknown spatial fields that parameterize it
      • These tasks are coupled: changing field placement changes the differential law, while a sufficiently flexible field can conceal structural error on a single tra…
      • We present Hypothesize, Evaluate, Refine for PDE Discovery (HER-PDE), a scientific-agent framework that discovers compositional PDE structure together with nonp…
  • Class-Based Heuristic Selection for Solving the Flying Block Puzzle

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27476v1 Announce Type: new.
      • Abstract: Heuristic search underlies planning in autonomous systems ranging from warehouse logistics to robotic navigation, yet generic heuristics fail to exploit the structural constraints that govern constrained spatial domains, causing search performance to degrade catastrophically on harder instances.
      • We study this problem through the two-column Flying Block Puzzle, a rigorously NP-complete spatial planning microworld whose bottleneck geometry mirrors clearance-to-size constraints encountered in multi-agent path finding, autonomous vehicle navigation, and block relocation systems.
      • We introduce the Class-Based Heuristic A* (CBHA*) algorithm, which integrates a General Move Constraint to capture minimum displacement costs when vacant units are scarce, a formal kinematic taxonomy partitioning the state space into seven mutually exclusive classes with provably admissible heuristics based on vacancy ratio and goal-piece geometry, and a class-conditional tie-breaking mechanism that dynamically switches between depth-priority and vertical-distance ordering to overcome f-value plateaus.
    • EN 要点:
      • arXiv:2608.27476v1 Announce Type: new
      • Abstract: Heuristic search underlies planning in autonomous systems ranging from warehouse logistics to robotic navigation, yet generic heuristics fail to explo…
      • We study this problem through the two-column Flying Block Puzzle, a rigorously NP-complete spatial planning microworld whose bottleneck geometry mirrors clearan…
      • We introduce the Class-Based Heuristic A* (CBHA*) algorithm, which integrates a General Move Constraint to capture minimum displacement costs when vacant units…
  • Benchmarking General Mobile Assistants in Challenging Real-World Scenarios

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27477v1 Announce Type: new.
      • Abstract: Graphical user interfaces have emerged as an important environment for evaluating autonomous AI agents on multimodal interactive tasks.
      • Existing benchmarks such as AndroidWorld and MobileWorld provide strong foundations for mobile agent evaluation, but their application coverage and task design do not yet fully capture the diversity and complexity of realistic mobile use.
      • We present GMA, a benchmark for evaluating general mobile assistants in challenging real-world scenarios.
    • EN 要点:
      • arXiv:2608.27477v1 Announce Type: new
      • Abstract: Graphical user interfaces have emerged as an important environment for evaluating autonomous AI agents on multimodal interactive tasks
      • Existing benchmarks such as AndroidWorld and MobileWorld provide strong foundations for mobile agent evaluation, but their application coverage and task design…
      • We present GMA, a benchmark for evaluating general mobile assistants in challenging real-world scenarios
  • Effectiveness of IoT and Deep Learning for Detection and Severity Assessment of Postelectrotermes militaris in Tea Plantations

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27480v1 Announce Type: new.
      • Abstract: Tea plantations are vulnerable to Postelectrotermes militaris, commonly known as the Upcountry Live Wood Termite (ULWT), which can cause substantial damage when infestations remain undetected.
      • This study proposes an IoT-enabled acoustic monitoring framework integrated with deep learning for early detection and severity assessment of ULWT infestations in tea plantations.
      • Research Method: Audio signals were captured non-invasively from tea trunks using a high-sensitivity microphone connected to a Raspberry Pi-based IoT device, with geographic coordinates recorded for spatial tracking.
    • EN 要点:
      • arXiv:2608.27480v1 Announce Type: new
      • Abstract: Tea plantations are vulnerable to Postelectrotermes militaris, commonly known as the Upcountry Live Wood Termite (ULWT), which can cause substantial d…
      • This study proposes an IoT-enabled acoustic monitoring framework integrated with deep learning for early detection and severity assessment of ULWT infestations…
      • Research Method: Audio signals were captured non-invasively from tea trunks using a high-sensitivity microphone connected to a Raspberry Pi-based IoT device, wi…
  • Context Localization for Generalized Level-Based Evaluation in Knowledge-Based Systems

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27482v1 Announce Type: new.
      • Abstract: We study context localization for generalized level-based evaluation in knowledge-based systems.
      • The framework models situations where a structured nonnegative score, defined on facts, rules, cases, criteria or evidence units, is evaluated through conditional aggregation tests on admissible knowledge contexts.
      • The generalized level measure maximizes a monotone set function over all contexts whose aggregated support reaches a prescribed level.
    • EN 要点:
      • arXiv:2608.27482v1 Announce Type: new
      • Abstract: We study context localization for generalized level-based evaluation in knowledge-based systems
      • The framework models situations where a structured nonnegative score, defined on facts, rules, cases, criteria or evidence units, is evaluated through condition…
      • The generalized level measure maximizes a monotone set function over all contexts whose aggregated support reaches a prescribed level

ArXiv cs.CL (B_intro+search) 链接到标题

  • NLP-Driven Knowledge Extraction and Thematic Classification of Translated Ancient Indian Medical Texts

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28608v1 Announce Type: new.
      • Abstract: Ancient Indian medical texts like Sushruta Samhita have extensive information on diseases, treatments, and surgical techniques.
      • Yet, their ancient format and use of intricate vocabulary pose difficulties in accessibility and systematic ordering.
      • The research here utilizes Natural Language Processing (NLP) methods like Named Entity Recognition (NER), BERTopic modeling, and Knowledge Graph development in Neo4j to extract, categorize, and visualize important concepts based on translated versions.
    • EN 要点:
      • arXiv:2608.28608v1 Announce Type: new
      • Abstract: Ancient Indian medical texts like Sushruta Samhita have extensive information on diseases, treatments, and surgical techniques
      • Yet, their ancient format and use of intricate vocabulary pose difficulties in accessibility and systematic ordering
      • The research here utilizes Natural Language Processing (NLP) methods like Named Entity Recognition (NER), BERTopic modeling, and Knowledge Graph development in…
  • Parametric Multimodal User Memory: Storing What Captions Cannot Carry

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28609v1 Announce Type: new.
      • Abstract: A personalized agent needs a user memory: a persistent model of who its user is.
      • Today it is almost always text – transcripts and captions retrieved by similarity.
      • This serves the captionable half of a person (“my cat is named Bibi”), but discards the perceptual half no caption can hold: how a voice sounds, how a face reads across age and lighting, how tired someone sounds.
    • EN 要点:
      • arXiv:2608.28609v1 Announce Type: new
      • Abstract: A personalized agent needs a user memory: a persistent model of who its user is
      • Today it is almost always text – transcripts and captions retrieved by similarity
      • This serves the captionable half of a person (“my cat is named Bibi”), but discards the perceptual half no caption can hold: how a voice sounds, how a face read…
  • Gurukul AI: An Interactive AI-Driven Educational Platform for Indian Education System

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28611v1 Announce Type: new.
      • Abstract: Recent advances in large language models (LLMs) like ChatGPT and LLaMA have transformed AI-driven education, but these systems are predominantly trained on Western-centric data, making them ill-suited for regional curricula like India’s.
      • The Indian education system is linguistically diverse, exam-oriented, and structured around standardized syllabi, not addressed by existing datasets or tools.
      • In this work, we curate a syllabus-aligned QA dataset based on NCERT (National Council of Educational Research and Training) textbooks for classes 9-12, capturing the content, context, and teaching style of Indian curricula.
    • EN 要点:
      • arXiv:2608.28611v1 Announce Type: new
      • Abstract: Recent advances in large language models (LLMs) like ChatGPT and LLaMA have transformed AI-driven education, but these systems are predominantly train…
      • The Indian education system is linguistically diverse, exam-oriented, and structured around standardized syllabi, not addressed by existing datasets or tools
      • In this work, we curate a syllabus-aligned QA dataset based on NCERT (National Council of Educational Research and Training) textbooks for classes 9-12, capturi…
  • STAGEET: Stage-wise Typed Edit Tagging for Grammatical Error Correction with Arabic as a Case Study

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28614v1 Announce Type: new.
      • Abstract: Sequence-to-edit approaches make grammatical error correction (GEC) efficient and locally interpretable by predicting edit labels over the input rather than generating a full corrected sentence.
      • Their interpretability, however, is primarily operational: a label specifies how the string should change, but a single edit vocabulary does not always reveal the type of correction being made.
      • We propose STAGEET, a stage-wise typed edit-tagging framework that reorganizes Seq2Edit supervision into typed executable stages and extends edit operations to correction categories.
    • EN 要点:
      • arXiv:2608.28614v1 Announce Type: new
      • Abstract: Sequence-to-edit approaches make grammatical error correction (GEC) efficient and locally interpretable by predicting edit labels over the input rathe…
      • Their interpretability, however, is primarily operational: a label specifies how the string should change, but a single edit vocabulary does not always reveal t…
      • We propose STAGEET, a stage-wise typed edit-tagging framework that reorganizes Seq2Edit supervision into typed executable stages and extends edit operations to…
  • From GenAI Virtual Patient Dialogue Logs to Teacher-Interpretable Process Evidence: A Learning Analytics Study in Higher Education

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28619v1 Announce Type: new.
      • Abstract: Medical history taking is a dialogue-based clinical reasoning task in which learners must gather, organise, and integrate patient information while the consultation unfolds.
      • Generative AI-powered virtual patients (GenAI VPs) make repeated history taking practice scalable and preserve full turn by turn dialogue.
      • However, these logs are educationally difficult to use directly.
    • EN 要点:
      • arXiv:2608.28619v1 Announce Type: new
      • Abstract: Medical history taking is a dialogue-based clinical reasoning task in which learners must gather, organise, and integrate patient information while th…
      • Generative AI-powered virtual patients (GenAI VPs) make repeated history taking practice scalable and preserve full turn by turn dialogue
      • However, these logs are educationally difficult to use directly
  • Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28623v1 Announce Type: new.
      • Abstract: Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before answering.
      • In language models it has been observed that this performance often comes with sycophancy, the tendency of a model to agree with the user over the evidence.
      • However, for LMRMs no reliable method to measure sycophancy yet exists.
    • EN 要点:
      • arXiv:2608.28623v1 Announce Type: new
      • Abstract: Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before an…
      • In language models it has been observed that this performance often comes with sycophancy, the tendency of a model to agree with the user over the evidence
      • However, for LMRMs no reliable method to measure sycophancy yet exists
  • MA-RAG: Multi-Agent Retrieval-Augmented Generation for Query-Driven Summarization of Longitudinal Parkinson’s Disease Assessments

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28624v1 Announce Type: new.
      • Abstract: Accurate interpretation of single-visit and longitudinal clinical assessments for Parkinson’s disease is time-consuming and often depends on specialist expertise.
      • Although large language models (LLMs) can generate natural language summaries, they frequently lack domain-specific clinical grounding and struggle to produce factually correct and temporally consistent responses for structured longitudinal assessment data.
      • To address these limitations, we propose MA-RAG, a query-driven multi-agent retrieval-augmented generation framework that decomposes clinical reasoning into domain-specialized agents, combines structured fact extraction, and synthesizes clinically grounded summaries through a final verification stage.
    • EN 要点:
      • arXiv:2608.28624v1 Announce Type: new
      • Abstract: Accurate interpretation of single-visit and longitudinal clinical assessments for Parkinson’s disease is time-consuming and often depends on specialis…
      • Although large language models (LLMs) can generate natural language summaries, they frequently lack domain-specific clinical grounding and struggle to produce f…
      • To address these limitations, we propose MA-RAG, a query-driven multi-agent retrieval-augmented generation framework that decomposes clinical reasoning into dom…
  • Asymmetric Within-Document Predictive Learning for Scientific Document Representation

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28625v1 Announce Type: new.
      • Abstract: We study predictive pretraining for scientific document representation using the discourse structure of papers.
      • We propose SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict method representations, and method representations are used to predict conclusion representations.
      • Experiments on RELISH, high-influence citation, SciDocs, and cite prediction show that plain predictive training is viable but weaker than a controlled contrastive baseline using the same section pairs.
    • EN 要点:
      • arXiv:2608.28625v1 Announce Type: new
      • Abstract: We study predictive pretraining for scientific document representation using the discourse structure of papers
      • We propose SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict…
      • Experiments on RELISH, high-influence citation, SciDocs, and cite prediction show that plain predictive training is viable but weaker than a controlled contrast…
  • Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28626v1 Announce Type: new.
      • Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation.
      • This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models’ training cutoffs.
      • Manuscripts were presented to both models with author identities blinded, replaced with high-prestige affiliations, or replaced with low-prestige affiliations, and in either text-only or text-with-figure format.
    • EN 要点:
      • arXiv:2608.28626v1 Announce Type: new
      • Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation
      • This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Lea…
      • Manuscripts were presented to both models with author identities blinded, replaced with high-prestige affiliations, or replaced with low-prestige affiliations,…
  • Intelligent Identification and Repair of Design Defects in BIM via Domain-Specific Large Language Models

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28629v1 Announce Type: new.
      • Abstract: Existing methods lack a generalized approach to efficiently identify and resolve the diversity of design defects in BIM.
      • Therefore, this study proposes an integrated framework to identify and repair various defects in BIM via domain-specific LLMs.
      • Firstly, a BIM-to-Text method with component-balanced chunking is introduced to bridge BIM data with LLMs.
    • EN 要点:
      • arXiv:2608.28629v1 Announce Type: new
      • Abstract: Existing methods lack a generalized approach to efficiently identify and resolve the diversity of design defects in BIM
      • Therefore, this study proposes an integrated framework to identify and repair various defects in BIM via domain-specific LLMs
      • Firstly, a BIM-to-Text method with component-balanced chunking is introduced to bridge BIM data with LLMs

ArXiv cs.LG (B_intro+search) 链接到标题

  • ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28771v1 Announce Type: new.
      • Abstract: Large reasoning models achieve strong performance on complex tasks by generating extended chain-of-thought (CoT) traces via reinforcement learning with verifiable rewards (RLVR).
      • While current RLVR methods have achieved strong results with correctness-based reward signals, they provide limited guidance on the quality of the reasoning process itself, leaving the internal reasoning structure largely unoptimized.
      • Through empirical analysis across multiple model families, we identify a consistent pattern: correct reasoning trac es exhibit more frequent and larger token-level entropy drops within the thinking phase than incorrect ones.
    • EN 要点:
      • arXiv:2608.28771v1 Announce Type: new
      • Abstract: Large reasoning models achieve strong performance on complex tasks by generating extended chain-of-thought (CoT) traces via reinforcement learning wit…
      • While current RLVR methods have achieved strong results with correctness-based reward signals, they provide limited guidance on the quality of the reasoning pro…
      • Through empirical analysis across multiple model families, we identify a consistent pattern: correct reasoning trac es exhibit more frequent and larger token-le…
  • Unsupervised Latent Space Alignment with Hyperspherical Geodesic Matching

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28840v1 Announce Type: new.
      • Abstract: Independently trained neural networks tend to encode the same data with similar latent geometries.
      • These latent geometries are not directly compatible, yet they can be nearly the same up to some class of transformations.
      • While there exists many methods for alignment between different latent spaces, it is typically done using a set of shared sample correspondences, known as anchors.
    • EN 要点:
      • arXiv:2608.28840v1 Announce Type: new
      • Abstract: Independently trained neural networks tend to encode the same data with similar latent geometries
      • These latent geometries are not directly compatible, yet they can be nearly the same up to some class of transformations
      • While there exists many methods for alignment between different latent spaces, it is typically done using a set of shared sample correspondences, known as ancho…
  • Curvature Cryptanalysis of Smooth Transformer Feed-Forward Networks

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28843v1 Announce Type: new.
      • Abstract: We show that smooth two-layer feed-forward networks (FFNs) expose an additional structural model extraction channel under a chosen-input raw-output oracle at the FFN branch; consider transformer FFN branches with GELU or SiLU activations under chosen-input raw-output access, without access to parameters, gradients, or internal activations; exploit a second-order leakage channel in which projected input Hessians form different mixtures of the same hidden symmetric rank-one factors induced by the FFN input weights.
      • We formalize resulting Hessian collection as a partially symmetric decomposition to establish conditions for local identifiability and stability to exploit vector-output stencil reuse to reduce the structural query cost by a factor of 16.
      • On independently trained CIFAR-10 vision transformers, only 16 projected Hessians, corresponding to 8193 black-box queries, recover the hidden FFN directions with average absolute cosine alignment above 0.94, with 95.1 % of GELU and 91.9 % of SiLU directions exceeding 0.90 alignment.
    • EN 要点:
      • arXiv:2608.28843v1 Announce Type: new
      • Abstract: We show that smooth two-layer feed-forward networks (FFNs) expose an additional structural model extraction channel under a chosen-input raw-output or…
      • We formalize resulting Hessian collection as a partially symmetric decomposition to establish conditions for local identifiability and stability to exploit vect…
      • On independently trained CIFAR-10 vision transformers, only 16 projected Hessians, corresponding to 8193 black-box queries, recover the hidden FFN directions wi…
  • Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28853v1 Announce Type: new.
      • Abstract: Equivariant graph neural networks provide a principled way to model geometric systems, but efficient first-order architectures remain limited in how vector information can be transformed as it moves across a graph.
      • We introduce \textsc{ESNN}, an Equivariant Sheaf Neural Network that enriches this interaction by learning directed, matrix-valued transport between neighboring vector features while preserving exact Euclidean equivariance.
      • Rather than increasing the order of the representation, ESNN keeps scalar and vector features first-order and places the additional geometric flexibility in the edge transport itself.
    • EN 要点:
      • arXiv:2608.28853v1 Announce Type: new
      • Abstract: Equivariant graph neural networks provide a principled way to model geometric systems, but efficient first-order architectures remain limited in how v…
      • We introduce \textsc{ESNN}, an Equivariant Sheaf Neural Network that enriches this interaction by learning directed, matrix-valued transport between neighboring…
      • Rather than increasing the order of the representation, ESNN keeps scalar and vector features first-order and places the additional geometric flexibility in the…
  • The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28859v1 Announce Type: new.
      • Abstract: Reasoning models do not stop when they know the answer.
      • On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model’s own answer probability takes to settle, and how much of that excess is removable varies from problem to problem, so a global length penalty cannot take it out.
      • We take it out by internalizing a causal interpretability finding into the weights.
    • EN 要点:
      • arXiv:2608.28859v1 Announce Type: new
      • Abstract: Reasoning models do not stop when they know the answer
      • On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model’s own answer probability takes to settle, and how much of that excess…
      • We take it out by internalizing a causal interpretability finding into the weights
  • Conservative Hybrid Graph Networks for Process Systems with Learned Routing

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28896v1 Announce Type: new.
      • Abstract: Industrial process networks do not maintain a single effective topology while operating: streams are throttled or bypassed, and units move between idle, transition, and active regimes.
      • Models of such systems are typically trained on measured state trajectories while the operating mechanisms that generated them remain latent, and an unconstrained graph network can fit such a trajectory without assigning stable physical meaning to the recovered routing.
      • We address both problems with the Conservative Hybrid Graph Network (CHGN), which learns routing, regime assignment, and removal rates as data-driven surrogates and inserts them into a fixed transport equation, so that the mass balance holds by construction for any predicted routing.
    • EN 要点:
      • arXiv:2608.28896v1 Announce Type: new
      • Abstract: Industrial process networks do not maintain a single effective topology while operating: streams are throttled or bypassed, and units move between idl…
      • Models of such systems are typically trained on measured state trajectories while the operating mechanisms that generated them remain latent, and an unconstrain…
      • We address both problems with the Conservative Hybrid Graph Network (CHGN), which learns routing, regime assignment, and removal rates as data-driven surrogates…
  • Off-Policy Evaluation for Semantic ID Recommenders: Does the Model’s Own Code Hierarchy Help?

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28905v1 Announce Type: new.
      • Abstract: Generative recommenders increasingly emit semantic IDs (SIDs): each item is a short sequence of hierarchical discrete codes from a residual quantizer, decoded autoregressively.
      • Before spending scarce A/B-test, a team may decide offline which decoder or reranking variants are worth testing - a job for off-policy evaluation (OPE).
      • We ask a simple question: can the model’s own SID tree serve as the action abstraction for that OPE?
    • EN 要点:
      • arXiv:2608.28905v1 Announce Type: new
      • Abstract: Generative recommenders increasingly emit semantic IDs (SIDs): each item is a short sequence of hierarchical discrete codes from a residual quantizer,…
      • Before spending scarce A/B-test, a team may decide offline which decoder or reranking variants are worth testing - a job for off-policy evaluation (OPE)
      • We ask a simple question: can the model’s own SID tree serve as the action abstraction for that OPE
  • Learning-Theoretic Foundation for General Coded Computing: The Straggler Setting

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28910v1 Announce Type: new.
      • Abstract: Coded computing has emerged as a powerful paradigm for mitigating the impact of straggling workers in distributed computing systems.
      • However, existing coded-computing schemes are predominantly designed for the exact recovery of highly structured computations, such as polynomial evaluation and matrix multiplication, and typically rely on strict recovery thresholds.
      • These assumptions significantly limit their applicability to modern machine-learning workloads, particularly deep neural networks (DNNs), whose computations generally lack rigid algebraic structure and, in many applications, require only accurate approximations rather than exact recovery.
    • EN 要点:
      • arXiv:2608.28910v1 Announce Type: new
      • Abstract: Coded computing has emerged as a powerful paradigm for mitigating the impact of straggling workers in distributed computing systems
      • However, existing coded-computing schemes are predominantly designed for the exact recovery of highly structured computations, such as polynomial evaluation and…
      • These assumptions significantly limit their applicability to modern machine-learning workloads, particularly deep neural networks (DNNs), whose computations gen…
  • SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28911v1 Announce Type: new.
      • Abstract: The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length.
      • We show that uniform KV quantization on a fractional-bit grid does not degrade gracefully: under a prespecified multi-seed statistical protocol, Llama-3.1-8B-Instruct with an affine quantizer is statistically indistinguishable from FP16 KV down to 2.322 code bits/value and collapses at 2.0 bits - a quality cliff in (2.0, 2.322] that reappears in generation-time quantization and multi-turn dialogue and transfers to Mistral-7B.
      • The cliff reframes importance-aware mixed precision: above it, eight model-internal importance indicators are statistically interchangeable, so the benefit of mixing is grid interpolation, reaching average precisions uniform quantization cannot realize.
    • EN 要点:
      • arXiv:2608.28911v1 Announce Type: new
      • Abstract: The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length
      • We show that uniform KV quantization on a fractional-bit grid does not degrade gracefully: under a prespecified multi-seed statistical protocol, Llama-3.1-8B-In…
      • The cliff reframes importance-aware mixed precision: above it, eight model-internal importance indicators are statistically interchangeable, so the benefit of m…
  • RankShift: In-Database Detection and Explanation of Categorical Shifts

    • 发布时间:2026-09-01 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.28922v1 Announce Type: new.
      • Abstract: A login service can receive its usual number of failed sign-ins while one source grows from 2% to 30% of them.
      • The same pattern appears in system logs when a rare event template becomes common while the message rate stays stable.
      • These events change which categories are active without changing how many events occur.
    • EN 要点:
      • arXiv:2608.28922v1 Announce Type: new
      • Abstract: A login service can receive its usual number of failed sign-ins while one source grows from 2% to 30% of them
      • The same pattern appears in system logs when a rare event template becomes common while the message rate stays stable
      • These events change which categories are active without changing how many events occur