{
  "title": "2026-08-26 AI日更 | OpenAI 把 AI 竞争推向全栈：自研推理芯片、算力系统与可信智能体同步加速",
  "url": "https://miaok.ong/ai-daily/ai-daily-2026-08-26/",
  "date": "2026-08-26T07:00:00+08:00",
  "lastmod": "2026-08-26T07:00:00+08:00",
  "type": "ai-daily",
  "kind": "page",
  "language": "zh",
  "description": "今天的主线从模型能力转向系统能力。OpenAI 继续强化从数据中心、芯片到产品的全栈策略，Jalapeño 指向更低成本推理；同时，智能体治理、引用归因与运行时证据协议升温，可靠性评测也开始深入长尾语言和文化场景。",
  "keywords": null,
  "tags": [],
  "categories": [],
  "author": "孔淼",
  "image": "https://miaok.ong/images/avatar.jpg",
  "content": "\u003ch1 id=\"2026-08-26-ai日更--openai-把-ai-竞争推向全栈自研推理芯片算力系统与可信智能体同步加速\"\u003e\n  2026-08-26 AI日更 | OpenAI 把 AI 竞争推向全栈：自研推理芯片、算力系统与可信智能体同步加速\n  \u003ca class=\"heading-link\" href=\"#2026-08-26-ai%e6%97%a5%e6%9b%b4--openai-%e6%8a%8a-ai-%e7%ab%9e%e4%ba%89%e6%8e%a8%e5%90%91%e5%85%a8%e6%a0%88%e8%87%aa%e7%a0%94%e6%8e%a8%e7%90%86%e8%8a%af%e7%89%87%e7%ae%97%e5%8a%9b%e7%b3%bb%e7%bb%9f%e4%b8%8e%e5%8f%af%e4%bf%a1%e6%99%ba%e8%83%bd%e4%bd%93%e5%90%8c%e6%ad%a5%e5%8a%a0%e9%80%9f\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h1\u003e\n\u003cblockquote\u003e\n\u003cp\u003e今天的主线从模型能力转向系统能力。OpenAI 继续强化从数据中心、芯片到产品的全栈策略，Jalapeño 指向更低成本推理；同时，智能体治理、引用归因与运行时证据协议升温，可靠性评测也开始深入长尾语言和文化场景。\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch2 id=\"-本期-watch-list-深度导读\"\u003e\n  📖 本期 Watch List 深度导读\n  \u003ca class=\"heading-link\" href=\"#-%e6%9c%ac%e6%9c%9f-watch-list-%e6%b7%b1%e5%ba%a6%e5%af%bc%e8%af%bb\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003e今天最值得跟进的主线有三条。第一条是 AI 基础设施全面走向“全栈化”：OpenAI 继续把算力、芯片与推理系统打通，Jalapeño 的早期结果也指向更高效率的推理架构，值得工程团队重点看。第二条是智能体进入治理与安全的硬问题：从 sycophancy、引用归因，到 agentic security 和运行时证据协议，说明“能做事”之后，可信、可审计开始成为核心门槛。第三条是模型可靠性正向长尾语言和文化场景下沉，今天几篇关于 Khmer、Nigerian Pidgin、Cyrillic 及视觉语言偏差的研究，提醒我们评测体系仍远未成熟。顺带一提，Netflix 探索把自己做成流媒体聚合入口，也值得关注，它反映的其实是平台型分发逻辑正在重排。\u003c/p\u003e\n\u003ch2 id=\"-x-平台-ai-热点快讯\"\u003e\n  🌐 X 平台 AI 热点快讯\n  \u003ca class=\"heading-link\" href=\"#-x-%e5%b9%b3%e5%8f%b0-ai-%e7%83%ad%e7%82%b9%e5%bf%ab%e8%ae%af\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003ch3 id=\"话题-1spacexai-and-cursor-boost-grok-model-usage-limits-again\"\u003e\n  话题 1:SpaceXAI and Cursor Boost Grok Model Usage Limits Again\n  \u003ca class=\"heading-link\" href=\"#%e8%af%9d%e9%a2%98-1spacexai-and-cursor-boost-grok-model-usage-limits-again\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e分类:AI · News\u003c/li\u003e\n\u003cli\u003e概况:热度时间:7 hours ago,相关帖子数:5400\u003c/li\u003e\n\u003cli\u003e是什么事:xAI 和 Cursor 再次上调 Grok 模型的使用额度，引发开发者和 AI 用户对可用性与成本的关注。\u003c/li\u003e\n\u003cli\u003e为什么重要:这反映出 AI 产品正在通过更高调用额度争夺开发者入口和工作流占用时长，也说明推理模型的算力供给、成本控制和商业化竞争正在加剧。\u003c/li\u003e\n\u003cli\u003e讨论概况:X 上的讨论主要围绕额度提升是否意味着 Grok 基础设施和模型服务能力真的增强，Cursor 用户是否会获得更稳定的编码体验，以及这是否会压缩 Claude、GPT 等模型在开发者工具中的份额；同时也有人质疑这种提升的成本可持续性和实际质量提升幅度。\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"话题-2shopify-ceo-pushes-for-claude-code-to-adopt-agentsmd-standard\"\u003e\n  话题 2:Shopify CEO Pushes for Claude Code to Adopt AGENTS.md Standard\n  \u003ca class=\"heading-link\" href=\"#%e8%af%9d%e9%a2%98-2shopify-ceo-pushes-for-claude-code-to-adopt-agentsmd-standard\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e分类:AI · News\u003c/li\u003e\n\u003cli\u003e概况:热度时间:9 hours ago,相关帖子数:2900\u003c/li\u003e\n\u003cli\u003e是什么事:Shopify CEO 公开呼吁 Claude Code 采用 AGENTS.md 标准，用统一文件描述代码库中 AI 编程代理的规则、上下文和协作约定。\u003c/li\u003e\n\u003cli\u003e为什么重要:这关系到 AI 编程工具的标准化和互操作性：如果不同代理能读取同一套项目约定，就能降低配置成本、减少误操作，并提升多工具、多代理在真实代码库中的协作稳定性。\u003c/li\u003e\n\u003cli\u003e讨论概况:X 上讨论焦点集中在 Claude Code 是否会带动 AGENTS.md 成为事实标准，以及统一规范能否改善编码代理对项目上下文的理解；分歧在于它究竟是必要的行业基础设施，还是又一个增加维护负担、可能造成配置碎片化的项目文件。\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"话题-3openai-engineer-surprises-with-changed-appearance-in-interview\"\u003e\n  话题 3:OpenAI Engineer Surprises with Changed Appearance in Interview\n  \u003ca class=\"heading-link\" href=\"#%e8%af%9d%e9%a2%98-3openai-engineer-surprises-with-changed-appearance-in-interview\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e分类:AI · News\u003c/li\u003e\n\u003cli\u003e概况:热度时间:2 days ago,相关帖子数:4600\u003c/li\u003e\n\u003cli\u003e是什么事:有 X 用户热议一位 OpenAI 工程师在采访中的外貌变化，引发对其身份、状态和背景的讨论。\u003c/li\u003e\n\u003cli\u003e为什么重要:这类话题之所以重要，在于它反映了公众对 OpenAI 及其核心员工的高度关注，容易影响公司形象、人才叙事以及外界对 AI 行业文化与工作压力的判断。\u003c/li\u003e\n\u003cli\u003e讨论概况:讨论焦点主要集中在外貌变化是否属实、原因是什么，以及这类关注是正常的人物观察还是对个人隐私的过度解读；也有人借题发挥，联想到 AI 公司高强度工作环境与舆论放大效应。\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"话题-4openai-unveils-jalapeño-chip-outpacing-nvidia-in-speed-and-efficiency\"\u003e\n  话题 4:OpenAI Unveils Jalapeño Chip Outpacing Nvidia in Speed and Efficiency\n  \u003ca class=\"heading-link\" href=\"#%e8%af%9d%e9%a2%98-4openai-unveils-jalape%c3%b1o-chip-outpacing-nvidia-in-speed-and-efficiency\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e分类:AI · News\u003c/li\u003e\n\u003cli\u003e概况:热度时间:9 hours ago,相关帖子数:13000\u003c/li\u003e\n\u003cli\u003e是什么事:OpenAI 据称公布了一款名为 Jalapeño 的芯片，并宣称其在速度和能效上超过了 Nvidia 的同类产品。\u003c/li\u003e\n\u003cli\u003e为什么重要:如果这一说法成立，说明头部 AI 公司正在通过自研芯片降低对通用 GPU 的依赖，这会影响算力成本、供应链和未来模型部署方式。\u003c/li\u003e\n\u003cli\u003e讨论概况:X 上的讨论主要集中在三个点：性能对比是否有公开基准支撑、这是否意味着 OpenAI 正在推进更深层的硬件自研，以及它对 Nvidia 在 AI 算力市场中的地位会产生多大冲击。\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch4 id=\"今日-x-上的-ai-舆情小结\"\u003e\n  今日 X 上的 AI 舆情小结\n  \u003ca class=\"heading-link\" href=\"#%e4%bb%8a%e6%97%a5-x-%e4%b8%8a%e7%9a%84-ai-%e8%88%86%e6%83%85%e5%b0%8f%e7%bb%93\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h4\u003e\n\u003cp\u003e今天 X 上的主线是：AI 竞争正在从“谁的模型更强”转向“谁能更便宜、更稳定、更深地嵌入开发者工作流”，无论是 Grok 提高额度、Claude Code 被呼吁统一 \u003ccode\u003eAGENTS.md\u003c/code\u003e，还是 OpenAI 传出自研芯片，都在指向算力、接口标准和入口控制这三件事。共识大致是，开发者工具的核心不只是模型能力本身，还包括可用性、成本和项目上下文的协作效率；分歧则集中在这些动作到底是实质能力提升，还是营销式放量、标准化负担或未经验证的性能叙事。对 \u003ccode\u003eAGENTS.md\u003c/code\u003e 的看法也明显分裂：支持者把它看成减少误操作、提升多代理协作的基础设施，反对者担心它会变成额外维护成本并制造新的碎片化。潜在风险主要有三类：一是高额度和自研芯片背后的算力成本是否可持续，二是未经充分基准支撑的性能宣称可能误导市场判断，三是围绕个体工程师外貌和状态的放大讨论，容易把行业压力、隐私边界和舆论猎奇混在一起。\u003c/p\u003e\n\u003ch2 id=\"-大佬观点influencer-insights\"\u003e\n  💡 大佬观点(Influencer Insights)\n  \u003ca class=\"heading-link\" href=\"#-%e5%a4%a7%e4%bd%ac%e8%a7%82%e7%82%b9influencer-insights\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cblockquote\u003e\n\u003cp\u003e今日大佬观点暂缺,推荐阅读 Watch List 深度内容。\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch2 id=\"-附录今日-watch-list-更新源列表\"\u003e\n  📚 附录:今日 Watch List 更新源列表\n  \u003ca class=\"heading-link\" href=\"#-%e9%99%84%e5%bd%95%e4%bb%8a%e6%97%a5-watch-list-%e6%9b%b4%e6%96%b0%e6%ba%90%e5%88%97%e8%a1%a8\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cblockquote\u003e\n\u003cp\u003e时间窗口:最近 3 天;覆盖 22 个源;共 35 条更新\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch3 id=\"stratechery-by-ben-thompson-a_full\"\u003e\n  Stratechery by Ben Thompson (A_full)\n  \u003ca class=\"heading-link\" href=\"#stratechery-by-ben-thompson-a_full\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003e\u003ca href=\"https://stratechery.com/2026/netflix-to-sell-streaming-services-streamers-as-aggregators-revisiting-roku/\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eNetflix to Sell Streaming Services?, Streamers as Aggregators, Revisiting Roku\u003c/a\u003e\u003c/strong\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 18:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- Netflix is considering selling other streaming services, and I think it’s a good idea; it’s also a let-down for Netflix’s original goals and potential pivots.\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003e$15\u003c/strong\u003e / month \u003cem\u003eor\u003c/em\u003e \u003cstrong\u003e$150\u003c/strong\u003e / year.\u003c/li\u003e\n\u003cli\u003eSubstantial analysis of the news of the day delivered via three weekly emails or podcasts.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eStratechery Interviews\u003c/strong\u003e.\u003c/li\u003e\n\u003cli\u003eInterviews with leading public CEOs, private company founders, and discussions with fellow analysts.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003eNetflix is considering selling other streaming services, and I think it\u0026rsquo;s a good idea; it\u0026rsquo;s also a let-down for Netflix\u0026rsquo;s original goals and potential pivots.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"openai-blog-a_full\"\u003e\n  OpenAI Blog (A_full)\n  \u003ca class=\"heading-link\" href=\"#openai-blog-a_full\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://openai.com/index/the-full-stack-behind-abundant-intelligence\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eThe full stack behind abundant intelligence\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 15:05 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- Progress in AI compounds fastest when the entire system improves together.\n\u003cul\u003e\n\u003cli\u003eThat is how I think about OpenAI’s compute strategy: one integrated system spanning data centers and chips, frontier models, our developer platform, consumer and enterprise products, and AI-native devices, with each layer strengthening the next.\u003c/li\u003e\n\u003cli\u003eBetter software makes hardware more productive.\u003c/li\u003e\n\u003cli\u003eHardware designed for our workloads improves speed and efficiency.\u003c/li\u003e\n\u003cli\u003eMore capable models unlock better products, which generate more demand, usage, and learning.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003eOpenAI CFO Sarah Friar explains how advances across chips, compute, models, and products compound to deliver more useful intelligence at greater scale and lower…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://openai.com/index/jalapeno-first-results\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eJalapeño’s first results show industry-leading speed and efficiency in AI inference\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 15:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- Since announcing Jalapeño, OpenAI’s first custom inference chip, we have been testing the chip and the system built around it.\n\u003cul\u003e\n\u003cli\u003eThe results show a significant performance advance: Jalapeño can serve more AI work per unit of power while also returning responses more quickly.\u003c/li\u003e\n\u003cli\u003eJalapeño delivers both higher throughput and lower latency with one architecture, where existing hardware systems often have to make a tradeoff between the two.\u003c/li\u003e\n\u003cli\u003eFor customers, that can mean faster responses, more responsive agents, and more reliable access as demand grows.\u003c/li\u003e\n\u003cli\u003eOur mission is to ensure that artificial general intelligence benefits all of humanity.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003eJalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern mod…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://openai.com/index/disrupting-malicious-uses-of-ai-influence-campaign-russia\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDisrupting a new covert influence campaign from Russia\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 08:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- OpenAI banned Russia-origin accounts using AI to promote a fake Israel-based think tank and a “sovereignty” index praising Russia and criticizing the West.\n\u003cul\u003e\n\u003cli\u003eThis piece from OpenAI Blog explains how Disrupting a new covert influence campaign from Russia shapes the broader AI and infrastructure landscape.\u003c/li\u003e\n\u003cli\u003eIt also surfaces practical implications for founders, operators, and investors following Disrupting a new covert influence campaign from Russia.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003eOpenAI banned Russia-origin accounts using AI to promote a fake Israel-based think tank and a “sovereignty” index praising Russia and criticizing the West.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://openai.com/index/introducing-admin-plugin\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eIntroducing the Admin plugin for ChatGPT Work and Codex\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 08:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- Use the Admin plugin for ChatGPT Work and Codex to analyze workspace usage, manage members and permissions, adjust limits, and act on admin requests.\n\u003cul\u003e\n\u003cli\u003eThis piece from OpenAI Blog explains how Introducing the Admin plugin for ChatGPT Work and Codex shapes the broader AI and infrastructure landscape.\u003c/li\u003e\n\u003cli\u003eIt also surfaces practical implications for founders, operators, and investors following Introducing the Admin plugin for ChatGPT Work and Codex.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003eUse the Admin plugin for ChatGPT Work and Codex to analyze workspace usage, manage members and permissions, adjust limits, and act on admin requests.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-csai-b_introsearch\"\u003e\n  ArXiv cs.AI (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-csai-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21362\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eKVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21362v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Transformer-based large language models (LLMs) incur high prefill latency because key-value (KV) tensors must be recomputed for each request.\u003c/li\u003e\n\u003cli\u003eExisting prefix-caching systems reduce this cost but require prompts to share a leading contiguous prefix, limiting effectiveness when shared content appears at arbitrary positions.\u003c/li\u003e\n\u003cli\u003eWe present KVBoost, a chunk-level KV cache reuse system for HuggingFace-compatible decoder models that enables reuse regardless of content position.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21362v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Transformer-based large language models (LLMs) incur high prefill latency because key-value (KV) tensors must be recomputed for each request\u003c/li\u003e\n\u003cli\u003eExisting prefix-caching systems reduce this cost but require prompts to share a leading contiguous prefix, limiting effectiveness when shared content appears at…\u003c/li\u003e\n\u003cli\u003eWe present KVBoost, a chunk-level KV cache reuse system for HuggingFace-compatible decoder models that enables reuse regardless of content position\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21363\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eAIREP: A Protocol for Per-Decision Evidence in AI Runtime Governance\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21363v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: A protocol is presented for recording the governance decisions of automated AI runtimes.\u003c/li\u003e\n\u003cli\u003eWhen a runtime releases, blocks, defers, redacts, or escalates an individual output, AIREP records that decision as a single signed object that any party can check offline, independent of the runtime that produced it.\u003c/li\u003e\n\u003cli\u003eA record carries the decision as one of a closed set of verbs under a stated policy basis, references its input, output, and evidence by hash rather than by value, and declares both what its evidence covers and what it does not.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21363v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: A protocol is presented for recording the governance decisions of automated AI runtimes\u003c/li\u003e\n\u003cli\u003eWhen a runtime releases, blocks, defers, redacts, or escalates an individual output, AIREP records that decision as a single signed object that any party can ch…\u003c/li\u003e\n\u003cli\u003eA record carries the decision as one of a closed set of verbs under a stated policy basis, references its input, output, and evidence by hash rather than by val…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21366\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eReviewing Model Collapse and Countermeasures\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21366v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Driven by massive amounts of web-scale data, generative AI (GenAI) has achieved remarkable progress, enabling various applications in diverse sectors.\u003c/li\u003e\n\u003cli\u003eThe advances of GenAI have actuated practitioners to use AI-synthesized data for training next-generation AI models.\u003c/li\u003e\n\u003cli\u003eUndeniably, using synthetic data has alleviated the increasing stringent demand for data supply.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21366v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Driven by massive amounts of web-scale data, generative AI (GenAI) has achieved remarkable progress, enabling various applications in diverse sectors\u003c/li\u003e\n\u003cli\u003eThe advances of GenAI have actuated practitioners to use AI-synthesized data for training next-generation AI models\u003c/li\u003e\n\u003cli\u003eUndeniably, using synthetic data has alleviated the increasing stringent demand for data supply\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21372\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eAI Learning and Conceptual Transfer in the Game of Hidden Rules\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21372v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning agents trained to infer hidden rules from trial-and-error feedback, representation design, rule difficulty analysis, transfer learning, generalization, and pseudo-bot-assisted human learning analysis.\u003c/li\u003e\n\u003cli\u003eThe report focuses on the Transformer-based A2C framework, Feature-Centric and Object-Centric representations, experimental findings, and classification of human learning data.\u003c/li\u003e\n\u003cli\u003earXiv:2608.21372v1 Announce Type: new Abstract: This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning agents trained to infer hidden rules… The report focuses on the Transformer-based A2C framework, Feature-Centric and Object-Centric representations, experimental findings, and classification of huma….\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21372v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning agents trained to infer hidden rules…\u003c/li\u003e\n\u003cli\u003eThe report focuses on the Transformer-based A2C framework, Feature-Centric and Object-Centric representations, experimental findings, and classification of huma…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21374\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eLitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21374v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics.\u003c/li\u003e\n\u003cli\u003eWe introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-writing experience compare anonymized drafts, are matched to topics within their expertise, and provide dimension-wise outcomes over five literature-review-specific criteria.\u003c/li\u003e\n\u003cli\u003eFrom this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs such as Sonar Deep Research substantially outperform base language models by over 60%.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21374v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspe…\u003c/li\u003e\n\u003cli\u003eWe introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-…\u003c/li\u003e\n\u003cli\u003eFrom this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21375\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eSchemaRouter: Field-Aware Tool Routing for Efficient Heterogeneous Agentic RAG\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21375v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Heterogeneous agentic retrieval-augmented generation (RAG) systems increasingly orchestrate external APIs, internal databases, vector stores, and graph stores.\u003c/li\u003e\n\u003cli\u003eExposing all tool descriptions to an LLM agent, or selecting tools only by vector similarity, causes two costly failures: over-fetching, which increases payload size, token use, and latency, and under-fetching, which omits fields needed to answer the query.\u003c/li\u003e\n\u003cli\u003eWe present SchemaRouter, a lightweight routing layer that represents tools, endpoints, parameters, response fields, domain concepts, units, provenance, and license policies as a schema graph.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21375v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Heterogeneous agentic retrieval-augmented generation (RAG) systems increasingly orchestrate external APIs, internal databases, vector stores, and grap…\u003c/li\u003e\n\u003cli\u003eExposing all tool descriptions to an LLM agent, or selecting tools only by vector similarity, causes two costly failures: over-fetching, which increases payload…\u003c/li\u003e\n\u003cli\u003eWe present SchemaRouter, a lightweight routing layer that represents tools, endpoints, parameters, response fields, domain concepts, units, provenance, and lice…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21379\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eRIACT: A Responsible AI System for Personalized Study Habit Tracking and Early Burnout Signal Detection in University Students\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21379v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Student burnout is highly prevalent in higher education, with reported rates ranging from 12% to over 70% and consistently exceeding those of the working population - yet it is typically identified only retrospectively, after academic decline has already occurred.\u003c/li\u003e\n\u003cli\u003eA contributing factor is that students have little structured visibility into their own study behaviour, and existing productivity tools record activity without interpreting it.\u003c/li\u003e\n\u003cli\u003eThis paper presents RIACT (Record, Insight, Analyze, Coach, Track), a web-based application that combines structured study session logging with a hybrid AI architecture to surface personalized insights and early burnout signals.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21379v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Student burnout is highly prevalent in higher education, with reported rates ranging from 12% to over 70% and consistently exceeding those of the work…\u003c/li\u003e\n\u003cli\u003eA contributing factor is that students have little structured visibility into their own study behaviour, and existing productivity tools record activity without…\u003c/li\u003e\n\u003cli\u003eThis paper presents RIACT (Record, Insight, Analyze, Coach, Track), a web-based application that combines structured study session logging with a hybrid AI arch…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21382\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eThere Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21382v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language model\u0026rsquo;s answer is read from generated text or from per-option likelihoods.\u003c/li\u003e\n\u003cli\u003eWork on this harness sensitivity reports it as aggregate score variance, leaving unexamined which items the variance falls on and whether they are the items that separate one model from the next.\u003c/li\u003e\n\u003cli\u003eWe treat the evaluation harness of large language models (LLMs) as an independent variable and resolve its effect to single items.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21382v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and wh…\u003c/li\u003e\n\u003cli\u003eWork on this harness sensitivity reports it as aggregate score variance, leaving unexamined which items the variance falls on and whether they are the items tha…\u003c/li\u003e\n\u003cli\u003eWe treat the evaluation harness of large language models (LLMs) as an independent variable and resolve its effect to single items\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21393\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eSpyre-Accelerated Retrieval-Augmented Generation on IBM LinuxONE: A Cloud-Native Architecture for Secure, High-Throughput Enterprise AI Inference\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21393v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Running large language models inside enterprise environments has always bumped up against a practical wall: the data lives in one place, the AI horsepower sits somewhere else, and moving sensitive records between the two creates real headaches around latency, security, and regulatory exposure.\u003c/li\u003e\n\u003cli\u003eIBM\u0026rsquo;s Spyre accelerator PCIe inference card built for LinuxONE and the broader IBM Z family changes that equation.\u003c/li\u003e\n\u003cli\u003eIn this paper we lay out a six-subsystem RAG architecture that runs entirely on IBM LinuxONE, using Spyre for generative inference, the Telum II on-chip accelerator for lightweight classification tasks, and Red Hat OpenShift for container orchestration.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21393v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Running large language models inside enterprise environments has always bumped up against a practical wall: the data lives in one place, the AI horsep…\u003c/li\u003e\n\u003cli\u003eIBM\u0026rsquo;s Spyre accelerator PCIe inference card built for LinuxONE and the broader IBM Z family changes that equation\u003c/li\u003e\n\u003cli\u003eIn this paper we lay out a six-subsystem RAG architecture that runs entirely on IBM LinuxONE, using Spyre for generative inference, the Telum II on-chip acceler…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21408\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eHate Speech Classification In Roman Urdu: A Comparative Study On Parameter Efficient Fine-Tuning And Prompt Engineering\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21408v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Due to the widespread accessibility of the internet and social media, toxic and hateful con-tent has grown exponentially, causing significant distress and negative societal impacts.\u003c/li\u003e\n\u003cli\u003eRo-man Urdu, a low-resource language used in Pakistan and among Urdu-speaking communities worldwide, presents additional challenges because of its informal grammar, inconsistent sen-tence structures, and multiple variations in word spellings.\u003c/li\u003e\n\u003cli\u003eThis research aims to identify the most effective techniques for hate speech classification in such low-resource settings with limited data.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21408v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Due to the widespread accessibility of the internet and social media, toxic and hateful con-tent has grown exponentially, causing significant distress…\u003c/li\u003e\n\u003cli\u003eRo-man Urdu, a low-resource language used in Pakistan and among Urdu-speaking communities worldwide, presents additional challenges because of its informal gram…\u003c/li\u003e\n\u003cli\u003eThis research aims to identify the most effective techniques for hate speech classification in such low-resource settings with limited data\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-cscl-b_introsearch\"\u003e\n  ArXiv cs.CL (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-cscl-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21364\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDistinguishing Revision and Delayed Elaboration in Incremental Narrative Interpretation\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21364v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Both human and AI systems that process narrative or long-form content operate incrementally: input is received over time, and internal representations must be updated accordingly.\u003c/li\u003e\n\u003cli\u003eIncremental interpretation, therefore, depends not only on what is represented but also on how the representational state evolves under new evidence.\u003c/li\u003e\n\u003cli\u003eWe distinguish two structurally different update operators that arise in narrative interpretation: revision-driven update and delayed elaboration.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21364v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Both human and AI systems that process narrative or long-form content operate incrementally: input is received over time, and internal representations…\u003c/li\u003e\n\u003cli\u003eIncremental interpretation, therefore, depends not only on what is represented but also on how the representational state evolves under new evidence\u003c/li\u003e\n\u003cli\u003eWe distinguish two structurally different update operators that arise in narrative interpretation: revision-driven update and delayed elaboration\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21365\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eKSE-Web: An Analysis of Hybrid Retrieval and LLM-Assisted Query Expansion for Low-Resource Khmer Semantic Search\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21365v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: As a low-resource language, Khmer presents several retrieval challenges, including limited annotated data, ambiguous word boundaries, weak support in multilingual embedding models, and frequent mixed Khmer-English usage.\u003c/li\u003e\n\u003cli\u003eThis paper presents KSE-Web, an analysis of hybrid retrieval and LLM-assisted query expansion for Khmer semantic search.\u003c/li\u003e\n\u003cli\u003eWe construct the dataset from approximately 17K candidate Khmer titles and retain 3K cleaned full-text Khmer documents after filtering, normalization, deduplication, and document-length control.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21365v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: As a low-resource language, Khmer presents several retrieval challenges, including limited annotated data, ambiguous word boundaries, weak support in…\u003c/li\u003e\n\u003cli\u003eThis paper presents KSE-Web, an analysis of hybrid retrieval and LLM-assisted query expansion for Khmer semantic search\u003c/li\u003e\n\u003cli\u003eWe construct the dataset from approximately 17K candidate Khmer titles and retain 3K cleaned full-text Khmer documents after filtering, normalization, deduplica…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21369\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eWazobia Eval: A Benchmark for Nigerian Pidgin Emotion Understanding, Sarcasm Detection, and Cultural Reasoning\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21369v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Nigerian Pidgin is one of Africa\u0026rsquo;s most widely spoken languages, yet remains severely underrepresented in language model evaluation.\u003c/li\u003e\n\u003cli\u003eExisting benchmarks primarily focus on translation, transcription, or generic sentiment analysis, leaving critical aspects of culturally grounded language understanding unmeasured.\u003c/li\u003e\n\u003cli\u003eWe introduce Wazobia Eval, a benchmark for evaluating Nigerian Pidgin emotion understanding, sarcasm detection, and cultural reasoning.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21369v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Nigerian Pidgin is one of Africa\u0026rsquo;s most widely spoken languages, yet remains severely underrepresented in language model evaluation\u003c/li\u003e\n\u003cli\u003eExisting benchmarks primarily focus on translation, transcription, or generic sentiment analysis, leaving critical aspects of culturally grounded language under…\u003c/li\u003e\n\u003cli\u003eWe introduce Wazobia Eval, a benchmark for evaluating Nigerian Pidgin emotion understanding, sarcasm detection, and cultural reasoning\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21376\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eOn the Role of Citations in Preference Data\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21376v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Many NLP tasks require systems to provide attribution in their outputs\u0026ndash;i.e.\u003c/li\u003e\n\u003cli\u003ecitations to grounding sources.\u003c/li\u003e\n\u003cli\u003eAttribution serves as a bulwark against model hallucination and as a means for users to verify the credibility of model outputs.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21376v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Many NLP tasks require systems to provide attribution in their outputs\u0026ndash;i.e\u003c/li\u003e\n\u003cli\u003ecitations to grounding sources\u003c/li\u003e\n\u003cli\u003eAttribution serves as a bulwark against model hallucination and as a means for users to verify the credibility of model outputs\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21377\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eAgentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21377v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Sycophancy in large language models, the tendency to prioritize user agreement over truthful responses, has been documented extensively but studied primarily in single-turn settings.\u003c/li\u003e\n\u003cli\u003eThis paper investigates a critical question: does subjecting LLMs to greater interaction scaffolding make sycophancy better or worse?\u003c/li\u003e\n\u003cli\u003eAcross 4,800 veracity judgments (200 statements $\\times$ 6 models $\\times$ 4 conditions), we find that the interaction scaffolding characteristic of agentic systems (feedback loops, reconsideration checkpoints, and iterative refinement) systematically amplifies sycophantic behavior.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21377v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Sycophancy in large language models, the tendency to prioritize user agreement over truthful responses, has been documented extensively but studied pr…\u003c/li\u003e\n\u003cli\u003eThis paper investigates a critical question: does subjecting LLMs to greater interaction scaffolding make sycophancy better or worse\u003c/li\u003e\n\u003cli\u003eAcross 4,800 veracity judgments (200 statements $\\times$ 6 models $\\times$ 4 conditions), we find that the interaction scaffolding characteristic of agentic sys…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21384\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eBeyond Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systems\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21384v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Modern multilingual tokenizers often fragment Ukrainian and other underrepresented Cyrillic-script languages more heavily than English, creating disparities in cost and context capacity.\u003c/li\u003e\n\u003cli\u003eWe quantify this overhead across nine production tokenizers and five languages with standardized Cyrillic and Latin representations, covering 8.37 million word forms.\u003c/li\u003e\n\u003cli\u003eOn a corpus benchmark, Ukrainian shows 68-121% token overhead on modern tokenizers and 220% on the older cl100k, measured through full-text fertility on the BrUK and Brown corpora.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21384v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Modern multilingual tokenizers often fragment Ukrainian and other underrepresented Cyrillic-script languages more heavily than English, creating dispa…\u003c/li\u003e\n\u003cli\u003eWe quantify this overhead across nine production tokenizers and five languages with standardized Cyrillic and Latin representations, covering 8.37 million word…\u003c/li\u003e\n\u003cli\u003eOn a corpus benchmark, Ukrainian shows 68-121% token overhead on modern tokenizers and 220% on the older cl100k, measured through full-text fertility on the BrU…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21385\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eA Social Media Analysis of Discourse on the Israel\u0026ndash;Palestine Conflict on Telegram\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21385v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Social media has become a central arena in which armed conflicts are contested, yet the pro-Israel and pro-Palestine communities on Telegram, whose broadcast architecture yields an unusually direct record of deliberate political communication, have not been systematically compared at scale.\u003c/li\u003e\n\u003cli\u003eThis study presents a multi-method computational analysis of 87,617 messages from sixteen Telegram channels, eight pro-Israel and eight pro-Palestine, spanning May 2021 to June 2026 and covering multiple conflict escalations.\u003c/li\u003e\n\u003cli\u003eIt combines sentiment analysis, three stance detection methods drawn from distinct paradigms (keyword matching, zero-shot DeBERTa via natural language inference, and a fine-tuned BERTweet model), and a framing analysis, all evaluated against 736 manually annotated messages.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21385v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Social media has become a central arena in which armed conflicts are contested, yet the pro-Israel and pro-Palestine communities on Telegram, whose br…\u003c/li\u003e\n\u003cli\u003eThis study presents a multi-method computational analysis of 87,617 messages from sixteen Telegram channels, eight pro-Israel and eight pro-Palestine, spanning…\u003c/li\u003e\n\u003cli\u003eIt combines sentiment analysis, three stance detection methods drawn from distinct paradigms (keyword matching, zero-shot DeBERTa via natural language inference…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21415\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eMitigating Bias in Large Vision-Language Models via Counterfactual Ensemble Decoding\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21415v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Large Vision-Language Models (LVLMs) have achieved remarkable performance across a wide range of tasks; however, they often inherit social biases from their training data, resulting in biased behavior when processing portraits from different social groups.\u003c/li\u003e\n\u003cli\u003eExisting debiasing approaches typically compare token probabilities between the original and biased generations during decoding, but they are fundamentally limited by their reliance on a single, stereotyped viewpoint and fail to account for the diversity of social perspectives.\u003c/li\u003e\n\u003cli\u003eInspired by the social science principle that diversity fosters fairness, we propose Counterfactual Ensemble Decoding (CED), a novel framework that constructs multi-group counterfactual perspectives within the visual representation space and integrates them during decoding to promote equitable model behavior.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21415v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Large Vision-Language Models (LVLMs) have achieved remarkable performance across a wide range of tasks; however, they often inherit social biases from…\u003c/li\u003e\n\u003cli\u003eExisting debiasing approaches typically compare token probabilities between the original and biased generations during decoding, but they are fundamentally limi…\u003c/li\u003e\n\u003cli\u003eInspired by the social science principle that diversity fosters fairness, we propose Counterfactual Ensemble Decoding (CED), a novel framework that constructs m…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21423\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eAgentic Security: A Systematization of Tools, Failure Modes, and Design Laws for LLM-Driven Penetration Testing\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21423v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools.\u003c/li\u003e\n\u003cli\u003eAs these systems move from demonstrations to deployed products, practitioners repeatedly encounter the same operational failures.\u003c/li\u003e\n\u003cli\u003eWe systematize these failures through a hands-on evaluation of ten widely used static, dynamic, cloud, orchestration, and AI red-teaming tools for unattended pipelines.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21423v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools\u003c/li\u003e\n\u003cli\u003eAs these systems move from demonstrations to deployed products, practitioners repeatedly encounter the same operational failures\u003c/li\u003e\n\u003cli\u003eWe systematize these failures through a hands-on evaluation of ten widely used static, dynamic, cloud, orchestration, and AI red-teaming tools for unattended pi…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21462\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eCyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21462v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alphabet with large speaker populations, while disadvantaging other language varieties.\u003c/li\u003e\n\u003cli\u003eNevertheless, they can also be a versatile tool for preserving precisely such endangered languages.\u003c/li\u003e\n\u003cli\u003eBut do they also possess the necessary creativity and capacity for abstraction to decode phonetically encoded language the same way humans do?\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21462v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alph…\u003c/li\u003e\n\u003cli\u003eNevertheless, they can also be a versatile tool for preserving precisely such endangered languages\u003c/li\u003e\n\u003cli\u003eBut do they also possess the necessary creativity and capacity for abstraction to decode phonetically encoded language the same way humans do\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-cslg-b_introsearch\"\u003e\n  ArXiv cs.LG (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-cslg-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21386\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eModel of Models: When Does Emitting a Specialist Beat Attending, Adapting, or Tuning?\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21386v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Given a task described by a few examples, how should a model be specialized to it?\u003c/li\u003e\n\u003cli\u003eFour mechanisms are available \u0026ndash; zero-shot, in-context attention, test-time gradient adaptation, and emitting specialist weights from a hypernetwork \u0026ndash; yet the operating regime of the last is rarely mapped.\u003c/li\u003e\n\u003cli\u003eWe run the identical four-way comparison across six tasks spanning regression, generation, language modeling, reinforcement learning, and clinical and genomic classification, holding the specialist, the context, and (where we can) the training budget fixed.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21386v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Given a task described by a few examples, how should a model be specialized to it\u003c/li\u003e\n\u003cli\u003eFour mechanisms are available \u0026ndash; zero-shot, in-context attention, test-time gradient adaptation, and emitting specialist weights from a hypernetwork \u0026ndash; yet the…\u003c/li\u003e\n\u003cli\u003eWe run the identical four-way comparison across six tasks spanning regression, generation, language modeling, reinforcement learning, and clinical and genomic c…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21398\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eRuntime Action Interference for AI Control of AlphaStar in StarCraft II\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21398v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: A trained reinforcement learning policy does not determine the complete behavior that users encounter: deployment code still schedules, admits, suppresses, or replaces its proposed actions.\u003c/li\u003e\n\u003cli\u003eWe contribute \\emph{runtime action interference} (RAI), an AI control mechanism that preserves policy parameters while regulating action pacing and filtering configured action patterns after inference.\u003c/li\u003e\n\u003cli\u003eRAI releases a proposed action only when its cooldown condition is satisfied and its content detector does not flag the action; otherwise, it dispatches a no-op.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21398v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: A trained reinforcement learning policy does not determine the complete behavior that users encounter: deployment code still schedules, admits, suppre…\u003c/li\u003e\n\u003cli\u003eWe contribute \\emph{runtime action interference} (RAI), an AI control mechanism that preserves policy parameters while regulating action pacing and filtering co…\u003c/li\u003e\n\u003cli\u003eRAI releases a proposed action only when its cooldown condition is satisfied and its content detector does not flag the action; otherwise, it dispatches a no-op\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21399\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eFederated Ensemble Forecasting Under Supply-Chain Market Volatility\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21399v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Supply chain forecasting systems increasingly operate under market shocks, non-identically distributed regional demand, and limited willingness to centralize commercial data.\u003c/li\u003e\n\u003cli\u003eThis work proposes Federated Ensemble Forecasting with Negative-Correlation Learning (FEF NCL), a distributed method that trains specialized forecasting experts across client nodes while discouraging redundant model errors.\u003c/li\u003e\n\u003cli\u003eThe framework combines temporal feature encoders, client level drift scoring, reliability-weighted aggregation, and an explain ability layer that exposes the market and supplier variables most responsible for each forecast.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21399v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Supply chain forecasting systems increasingly operate under market shocks, non-identically distributed regional demand, and limited willingness to cen…\u003c/li\u003e\n\u003cli\u003eThis work proposes Federated Ensemble Forecasting with Negative-Correlation Learning (FEF NCL), a distributed method that trains specialized forecasting experts…\u003c/li\u003e\n\u003cli\u003eThe framework combines temporal feature encoders, client level drift scoring, reliability-weighted aggregation, and an explain ability layer that exposes the ma…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21473\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eClass-Conditioned Gaussian Mixture Modeling for Imbalanced Time Series Quantification\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21473v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Quantification, estimating class prevalences in bags of unlabeled instances is vital in domains where aggregate statistics are more important than individual instance labels, such as biosignal monitoring, fall detection, and activity recognition.\u003c/li\u003e\n\u003cli\u003eWe investigate this issue in the challenging setting of imbalanced time series data and develop CC-GMNet-TS, a class-conditioned Gaussian mixture quantifier that combines a Transformer-based feature extractor with per-class latent mixtures.\u003c/li\u003e\n\u003cli\u003eUnlike previous mixture-based quantifiers, which use a single Gaussian mixture shared by all classes, CC-GMNet-TS assigns each class its own compact mixture in a bounded latent space and scores segment embeddings against these class-specific components to create bag-level representations that emphasize rare but informative patterns.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21473v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Quantification, estimating class prevalences in bags of unlabeled instances is vital in domains where aggregate statistics are more important than ind…\u003c/li\u003e\n\u003cli\u003eWe investigate this issue in the challenging setting of imbalanced time series data and develop CC-GMNet-TS, a class-conditioned Gaussian mixture quantifier tha…\u003c/li\u003e\n\u003cli\u003eUnlike previous mixture-based quantifiers, which use a single Gaussian mixture shared by all classes, CC-GMNet-TS assigns each class its own compact mixture in…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21485\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eCongruence Decomposition with Neural Block Solvers for Large-Scale PCI Assignment\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21485v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Physical Cell Identity (PCI) assignment is essential for interference management in dense 5G networks.\u003c/li\u003e\n\u003cli\u003eAs cellular networks scale, PCI reuse becomes unavoidable, which may cause collisions, confusions, and multiple forms of modular interference.\u003c/li\u003e\n\u003cli\u003eJointly mitigating these effects gives rise to a large-scale, multi-objective combinatorial optimization problem that is difficult to solve efficiently at practical network scales.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21485v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Physical Cell Identity (PCI) assignment is essential for interference management in dense 5G networks\u003c/li\u003e\n\u003cli\u003eAs cellular networks scale, PCI reuse becomes unavoidable, which may cause collisions, confusions, and multiple forms of modular interference\u003c/li\u003e\n\u003cli\u003eJointly mitigating these effects gives rise to a large-scale, multi-objective combinatorial optimization problem that is difficult to solve efficiently at pract…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21488\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eKAN-Robust-Bench: A Benchmark for Evaluating the Robustness of Kolmogorov-Arnold Networks\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21488v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: While machine learning models have demonstrated strong performance in many domains, these models have shown profound vulnerabilities when they are exposed to adversarial threats.\u003c/li\u003e\n\u003cli\u003eWhile adversarial attacks fall into various categories, the most prominent category in research studies is evasion.\u003c/li\u003e\n\u003cli\u003eIn evasion attacks, the adversary generates perturbed versions of samples, which might not be observable by human eyes.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21488v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: While machine learning models have demonstrated strong performance in many domains, these models have shown profound vulnerabilities when they are exp…\u003c/li\u003e\n\u003cli\u003eWhile adversarial attacks fall into various categories, the most prominent category in research studies is evasion\u003c/li\u003e\n\u003cli\u003eIn evasion attacks, the adversary generates perturbed versions of samples, which might not be observable by human eyes\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21496\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eThe geometry of AI validation: Exact certification limits for iid best-of-N search\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21496v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: AI systems increasingly generate alternatives, inspect evidence, and deploy a selected output.\u003c/li\u003e\n\u003cli\u003eValidation is therefore target-relative: evidence certifies deployment only in directions resolved by the interventions that produced it.\u003c/li\u003e\n\u003cli\u003eWe represent validation and deployment rules as kernels over a reliability surface.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21496v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: AI systems increasingly generate alternatives, inspect evidence, and deploy a selected output\u003c/li\u003e\n\u003cli\u003eValidation is therefore target-relative: evidence certifies deployment only in directions resolved by the interventions that produced it\u003c/li\u003e\n\u003cli\u003eWe represent validation and deployment rules as kernels over a reliability surface\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21499\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eSelection of Heart Sound Segments for Synchronous Classification of Multi-channel Heart Sounds\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21499v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Cardiac auscultation remains the most cost-effective screening procedure for cardiovascular diseases, and requires listening at the four main auscultation spots.\u003c/li\u003e\n\u003cli\u003eDespite this, automatic heart sound analysis algorithms mostly classify patients using a single heart sound (single-channel), or, when using more than one (multi-channel), analyze each channel individually.\u003c/li\u003e\n\u003cli\u003eTo our knowledge, no prior work classifies patients through the synchronous analysis of multi-channel heart sounds, following the procedure used by physicians.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21499v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Cardiac auscultation remains the most cost-effective screening procedure for cardiovascular diseases, and requires listening at the four main ausculta…\u003c/li\u003e\n\u003cli\u003eDespite this, automatic heart sound analysis algorithms mostly classify patients using a single heart sound (single-channel), or, when using more than one (mult…\u003c/li\u003e\n\u003cli\u003eTo our knowledge, no prior work classifies patients through the synchronous analysis of multi-channel heart sounds, following the procedure used by physicians\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21504\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eChemDIRT: A Diversified Instruction, Representation, and Task Benchmark for Robust Chemistry-LLM Evaluation\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21504v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: The rapid advancement of large language models (LLMs) has led to increasing interest in their application to scientific domains such as chemistry.\u003c/li\u003e\n\u003cli\u003eHowever, existing chemistry benchmarks often provide only a narrow view of model capability, focusing on limited task sets while overlooking robustness to variations in problem formulation and chemical representation.\u003c/li\u003e\n\u003cli\u003eAs a result, reported performance may overestimate a model\u0026rsquo;s true ability to reason consistently across realistic settings.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21504v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: The rapid advancement of large language models (LLMs) has led to increasing interest in their application to scientific domains such as chemistry\u003c/li\u003e\n\u003cli\u003eHowever, existing chemistry benchmarks often provide only a narrow view of model capability, focusing on limited task sets while overlooking robustness to varia…\u003c/li\u003e\n\u003cli\u003eAs a result, reported performance may overestimate a model\u0026rsquo;s true ability to reason consistently across realistic settings\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.21530\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eMultimodal Injury Risk and Performance Prediction in Tennis Using Weighted Ensemble Learning\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-25 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.21530v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Machine learning has had a positive impact on the sports industry, with one of its most promising applications being the prediction of athlete performance and injury risk.\u003c/li\u003e\n\u003cli\u003eRecent advances have employed state-of-the-art models to improve prediction accuracy, yet progress remains limited by data availability and the reliance on subjective observations or expert assessments.\u003c/li\u003e\n\u003cli\u003eTo address these limitations, researchers in sports such as soccer, basketball, and wrestling have begun integrating heterogeneous data sources, such as wearable device readings, with traditional subjective assessments.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.21530v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Machine learning has had a positive impact on the sports industry, with one of its most promising applications being the prediction of athlete perform…\u003c/li\u003e\n\u003cli\u003eRecent advances have employed state-of-the-art models to improve prediction accuracy, yet progress remains limited by data availability and the reliance on subj…\u003c/li\u003e\n\u003cli\u003eTo address these limitations, researchers in sports such as soccer, basketball, and wrestling have begun integrating heterogeneous data sources, such as wearabl…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n",
  "wordCount": 5674,
  "readingTime": 27,
  "tableOfContents": "\u003cnav id=\"TableOfContents\"\u003e\n  \u003cul\u003e\n    \u003cli\u003e\u003ca href=\"#-本期-watch-list-深度导读\"\u003e📖 本期 Watch List 深度导读\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-x-平台-ai-热点快讯\"\u003e🌐 X 平台 AI 热点快讯\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#话题-1spacexai-and-cursor-boost-grok-model-usage-limits-again\"\u003e话题 1:SpaceXAI and Cursor Boost Grok Model Usage Limits Again\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#话题-2shopify-ceo-pushes-for-claude-code-to-adopt-agentsmd-standard\"\u003e话题 2:Shopify CEO Pushes for Claude Code to Adopt AGENTS.md Standard\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#话题-3openai-engineer-surprises-with-changed-appearance-in-interview\"\u003e话题 3:OpenAI Engineer Surprises with Changed Appearance in Interview\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#话题-4openai-unveils-jalapeño-chip-outpacing-nvidia-in-speed-and-efficiency\"\u003e话题 4:OpenAI Unveils Jalapeño Chip Outpacing Nvidia in Speed and Efficiency\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-大佬观点influencer-insights\"\u003e💡 大佬观点(Influencer Insights)\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-附录今日-watch-list-更新源列表\"\u003e📚 附录:今日 Watch List 更新源列表\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#stratechery-by-ben-thompson-a_full\"\u003eStratechery by Ben Thompson (A_full)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#openai-blog-a_full\"\u003eOpenAI Blog (A_full)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-csai-b_introsearch\"\u003eArXiv cs.AI (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-cscl-b_introsearch\"\u003eArXiv cs.CL (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-cslg-b_introsearch\"\u003eArXiv cs.LG (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n  \u003c/ul\u003e\n\u003c/nav\u003e",
  "isDraft": false
}
