{
  "title": "2026-08-28 AI日更 | AI 从能用走向可验证：双盲评测落地，语音与视频能力继续前推",
  "url": "https://miaok.ong/ai-daily/ai-daily-2026-08-28/",
  "date": "2026-08-28T07:00:00+08:00",
  "lastmod": "2026-08-28T07:00:00+08:00",
  "type": "ai-daily",
  "kind": "page",
  "language": "zh",
  "description": "今天的重点不在单点模型刷新，而在能力进入更严格的验证与交付阶段。Google DeepMind 推进双盲评测和更可控的生成能力，同时语音转写、视频生成等产品继续向低延迟、长上下文和生产工作流靠拢。另一条线索是，评测可靠性被重新审视，许多指标开始暴露出测试本身的偏差与失真。",
  "keywords": null,
  "tags": [],
  "categories": [],
  "author": "孔淼",
  "image": "https://miaok.ong/images/avatar.jpg",
  "content": "\u003ch1 id=\"2026-08-28-ai日更--ai-从能用走向可验证双盲评测落地语音与视频能力继续前推\"\u003e\n  2026-08-28 AI日更 | AI 从能用走向可验证：双盲评测落地，语音与视频能力继续前推\n  \u003ca class=\"heading-link\" href=\"#2026-08-28-ai%e6%97%a5%e6%9b%b4--ai-%e4%bb%8e%e8%83%bd%e7%94%a8%e8%b5%b0%e5%90%91%e5%8f%af%e9%aa%8c%e8%af%81%e5%8f%8c%e7%9b%b2%e8%af%84%e6%b5%8b%e8%90%bd%e5%9c%b0%e8%af%ad%e9%9f%b3%e4%b8%8e%e8%a7%86%e9%a2%91%e8%83%bd%e5%8a%9b%e7%bb%a7%e7%bb%ad%e5%89%8d%e6%8e%a8\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h1\u003e\n\u003cblockquote\u003e\n\u003cp\u003e今天的重点不在单点模型刷新，而在能力进入更严格的验证与交付阶段。Google DeepMind 推进双盲评测和更可控的生成能力，同时语音转写、视频生成等产品继续向低延迟、长上下文和生产工作流靠拢。另一条线索是，评测可靠性被重新审视，许多指标开始暴露出测试本身的偏差与失真。\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch2 id=\"-本期-watch-list-深度导读\"\u003e\n  📖 本期 Watch List 深度导读\n  \u003ca class=\"heading-link\" href=\"#-%e6%9c%ac%e6%9c%9f-watch-list-%e6%b7%b1%e5%ba%a6%e5%af%bc%e8%af%bb\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003e今天最值得跟进的主线有三条。第一条是“模型可控性与后训练”：Google DeepMind 的 Gemini Omni 1.1 Flash、双盲 AI 评测，以及关于无监督后训练、激活 steering 和 fine-tuning 影响的几篇论文，合在一起指向同一个问题——模型越来越强，但真正可控、可验证的边界仍在重画。第二条是“评测可靠性”：从 dialect bias、语义回复一致性，到 imperfective paradox 的 benchmark 失真，今天多篇工作都在提醒，很多看似稳固的指标，先坏掉的可能是测试本身。第三条是“结构化任务落地”：ESQ-Bench、DataKernelBench 以及记忆/RAG 评估的新框架，说明企业级 NL2SQL、数据库优化和检索系统，正在进入更严苛的真实场景检验。\u003c/p\u003e\n\u003ch2 id=\"-x-平台-ai-热点快讯\"\u003e\n  🌐 X 平台 AI 热点快讯\n  \u003ca class=\"heading-link\" href=\"#-x-%e5%b9%b3%e5%8f%b0-ai-%e7%83%ad%e7%82%b9%e5%bf%ab%e8%ae%af\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003ch3 id=\"话题-1cursor-launches-scratch-to-deploy-web-app-builder\"\u003e\n  话题 1:Cursor Launches Scratch-to-Deploy Web App Builder\n  \u003ca class=\"heading-link\" href=\"#%e8%af%9d%e9%a2%98-1cursor-launches-scratch-to-deploy-web-app-builder\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e分类:AI · News\u003c/li\u003e\n\u003cli\u003e概况:热度时间:,相关帖子数:203\u003c/li\u003e\n\u003cli\u003e是什么事:Cursor 发布了一款可从零开始构建并直接部署的网页应用生成器，把 AI 编程从原型创建推进到上线交付。\u003c/li\u003e\n\u003cli\u003e为什么重要:这表明 AI 编程工具正在从“辅助写代码”转向“端到端交付应用”，对开发流程、产品迭代速度和 AI 原生应用的商业化都具有直接影响。\u003c/li\u003e\n\u003cli\u003e讨论概况:X 上的讨论主要集中在两点：一是这种 Scratch-to-Deploy 能否真正降低非工程用户做出可用应用的门槛；二是它与现有低代码、IDE 和 AI 编程助手相比，究竟是新范式还是能力整合。\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"话题-2tech-giants-urge-urgent-ai-cyber-defense-action\"\u003e\n  话题 2:Tech Giants Urge Urgent AI Cyber Defense Action\n  \u003ca class=\"heading-link\" href=\"#%e8%af%9d%e9%a2%98-2tech-giants-urge-urgent-ai-cyber-defense-action\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e分类:AI · News\u003c/li\u003e\n\u003cli\u003e概况:热度时间:6 hours ago,相关帖子数:8200\u003c/li\u003e\n\u003cli\u003e是什么事:多家科技巨头呼吁各方尽快采取更紧急的 AI 网络防御措施，以应对生成式 AI 被用于攻击和防护失衡的风险。\u003c/li\u003e\n\u003cli\u003e为什么重要:这件事重要在于，AI 正在同时放大攻击与防御能力，若防护体系跟不上，模型、数据和基础设施都可能成为更大规模网络威胁的目标。\u003c/li\u003e\n\u003cli\u003e讨论概况:X 上的讨论主要集中在两点：一是企业和政府是否已低估 AI 带来的安全风险，二是应优先依靠行业自律、技术标准，还是更强监管来推动 AI 网络防御落地。\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"话题-3google-launches-gemini-35-transcribe-for-precise-speech-to-text\"\u003e\n  话题 3:Google Launches Gemini 3.5 Transcribe for Precise Speech-to-Text\n  \u003ca class=\"heading-link\" href=\"#%e8%af%9d%e9%a2%98-3google-launches-gemini-35-transcribe-for-precise-speech-to-text\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e分类:AI · News\u003c/li\u003e\n\u003cli\u003e概况:热度时间:1 day ago,相关帖子数:5500\u003c/li\u003e\n\u003cli\u003e是什么事:Google 发布 Gemini 3.5 Transcribe，用于高精度语音转文字，支持 85 余种语言自动识别、去除口头语、区分多说话人，并提供离线与低延迟直播两种模式。\u003c/li\u003e\n\u003cli\u003e为什么重要:这意味着语音识别正从通用能力走向可直接嵌入生产工作流的基础设施，对会议纪要、客服、媒体转写和多语言应用都有直接影响，也体现了 AI 产品化能力在向实时性、准确率和开发集成能力竞争。\u003c/li\u003e\n\u003cli\u003e讨论概况:X 上讨论集中在几个点：它的多语言和低延迟表现是否足以对标现有转写方案，85 语言与自定义词表对行业场景的价值有多大，以及谷歌是否借此把语音能力进一步嵌入 Gemini 生态和开发者工具中。\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"话题-4google-launches-gemini-omni-11-flash-for-advanced-video-creation\"\u003e\n  话题 4:Google Launches Gemini Omni 1.1 Flash for Advanced Video Creation\n  \u003ca class=\"heading-link\" href=\"#%e8%af%9d%e9%a2%98-4google-launches-gemini-omni-11-flash-for-advanced-video-creation\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e分类:AI · News\u003c/li\u003e\n\u003cli\u003e概况:热度时间:7 hours ago,相关帖子数:3000\u003c/li\u003e\n\u003cli\u003e是什么事:Google 发布了 Gemini Omni 1.1 Flash，用于更高级的视频创作，新增场景延展等能力，并可分析最多 10 秒素材以保持叙事连贯。\u003c/li\u003e\n\u003cli\u003e为什么重要:这表明生成式视频模型正在从单段生成走向更强的上下文理解与连续性控制，直接关系到 AI 视频创作的可用性、成本和生产效率。\u003c/li\u003e\n\u003cli\u003e讨论概况:X 上的讨论主要集中在它是否真的提升了视频叙事一致性、与现有视频生成模型相比的实际效果，以及这类能力会如何影响内容创作流程和版权、真实性等问题。\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"话题-5salesforce-beats-earnings-expectations-with-claude-ai-integration\"\u003e\n  话题 5:Salesforce Beats Earnings Expectations with Claude AI Integration\n  \u003ca class=\"heading-link\" href=\"#%e8%af%9d%e9%a2%98-5salesforce-beats-earnings-expectations-with-claude-ai-integration\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e分类:AI · News\u003c/li\u003e\n\u003cli\u003e概况:热度时间:1 day ago,相关帖子数:8100\u003c/li\u003e\n\u003cli\u003e是什么事:Salesforce 公布的财报超出市场预期，并将 Claude AI 集成作为业绩亮点之一。\u003c/li\u003e\n\u003cli\u003e为什么重要:这说明生成式 AI 正在从概念验证走向企业软件的实际收入和产品差异化，对 AI 商业化路径和企业级应用落地具有信号意义。\u003c/li\u003e\n\u003cli\u003e讨论概况:X 上的讨论主要集中在 AI 集成是否真正推动了 Salesforce 的增长、Claude 在企业场景中的竞争力，以及这类合作对营收、利润率和对 Anthropic 依赖度的影响。\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"话题-6nvidia-acquires-hugging-face-for-129-billion-in-major-ai-deal\"\u003e\n  话题 6:Nvidia Acquires Hugging Face for $12.9 Billion in Major AI Deal\n  \u003ca class=\"heading-link\" href=\"#%e8%af%9d%e9%a2%98-6nvidia-acquires-hugging-face-for-129-billion-in-major-ai-deal\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e分类:AI · News\u003c/li\u003e\n\u003cli\u003e概况:热度时间:19 hours ago,相关帖子数:22000\u003c/li\u003e\n\u003cli\u003e是什么事:X 平台热议一则消息称，Nvidia 以 129 亿美元收购 Hugging Face，成为一笔引发广泛关注的 AI 交易。\u003c/li\u003e\n\u003cli\u003e为什么重要:Hugging Face 是开源模型与工具生态的重要入口，若被 Nvidia 收购，可能重塑 AI 基础设施、模型分发和开发者生态的权力结构。\u003c/li\u003e\n\u003cli\u003e讨论概况:讨论主要集中在这笔交易对开源中立性的影响、Nvidia 是否会进一步强化其在 AI 全栈中的控制力，以及这是否会改变模型社区和企业用户的选择。\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"话题-7neil-movva-breaks-down-ai-inference-economics-on-invest-like-the-best\"\u003e\n  话题 7:Neil Movva Breaks Down AI Inference Economics on Invest Like the Best\n  \u003ca class=\"heading-link\" href=\"#%e8%af%9d%e9%a2%98-7neil-movva-breaks-down-ai-inference-economics-on-invest-like-the-best\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e分类:AI · News\u003c/li\u003e\n\u003cli\u003e概况:热度时间:2 days ago,相关帖子数:4100\u003c/li\u003e\n\u003cli\u003e是什么事:Neil Movva 在《Invest Like the Best》中讨论并拆解了 AI 推理的经济模型，重点关注算力成本、定价和商业化路径。\u003c/li\u003e\n\u003cli\u003e为什么重要:推理成本正在直接决定 AI 产品的毛利、规模化速度和竞争格局，因此这类分析会影响模型厂商、云服务商和应用层公司的商业决策。\u003c/li\u003e\n\u003cli\u003e讨论概况:X 上的讨论主要集中在推理成本是否会快速下降、谁能在算力和基础设施上获得优势，以及 AI 公司能否在高算力消耗下建立可持续的收入模型。\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"话题-8tesla-adds-79-model-ys-to-texas-robotaxi-fleet-in-one-day\"\u003e\n  话题 8:Tesla Adds 79 Model Ys to Texas Robotaxi Fleet in One Day\n  \u003ca class=\"heading-link\" href=\"#%e8%af%9d%e9%a2%98-8tesla-adds-79-model-ys-to-texas-robotaxi-fleet-in-one-day\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e分类:AI · News\u003c/li\u003e\n\u003cli\u003e概况:热度时间:,相关帖子数:823\u003c/li\u003e\n\u003cli\u003e是什么事:特斯拉被指在一天内向德州 Robotaxi 车队新增了 79 辆 Model Y，用于自动驾驶出行服务的扩张。\u003c/li\u003e\n\u003cli\u003e为什么重要:这说明特斯拉正在加速把量产车转化为可运营的自动驾驶车队，对 AI 在真实世界交通场景中的落地、规模化部署和商业化验证都很关键。\u003c/li\u003e\n\u003cli\u003e讨论概况:X 上的讨论主要集中在这是否意味着 Robotaxi 进展已进入实质扩张阶段，以及这些车辆的自动驾驶能力、监管合规性、运营范围和数据回流是否足以支撑特斯拉的叙事；分歧则在于这更像是真正的商业化突破，还是一次有限的车队补充与市场宣传。\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch4 id=\"今日-x-上的-ai-舆情小结\"\u003e\n  今日 X 上的 AI 舆情小结\n  \u003ca class=\"heading-link\" href=\"#%e4%bb%8a%e6%97%a5-x-%e4%b8%8a%e7%9a%84-ai-%e8%88%86%e6%83%85%e5%b0%8f%e7%bb%93\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h4\u003e\n\u003cp\u003e今天的舆论主线很一致：AI 正从“会演示”转向“能交付”，无论是 Cursor 的从零到部署、谷歌的语音转写和视频生成，还是 Salesforce 的企业收入验证，讨论焦点都落在可用性、集成度和商业化落地上。比较明确的共识是，真正决定胜负的已经不是单点模型能力，而是能否进入工作流、压低使用门槛，并把推理成本和产品收入跑通。分歧主要在两类问题上：一类是这些发布到底是范式变化，还是既有 IDE、低代码、云和模型能力的重新打包；另一类是特斯拉 Robotaxi、Nvidia 收购 Hugging Face 这类消息究竟代表实质扩张，还是资本与叙事先行。潜在风险也被反复提到：AI 网络攻防失衡会让模型、数据和基础设施暴露在更高频、更大规模的威胁下；视频与转写能力增强则会进一步放大版权、真实性和内容滥用问题；而基础设施和分发入口一旦继续向少数巨头集中，开源中立性、开发者选择和市场竞争都会承压。整体看，今天的讨论不是在争论 AI 会不会落地，而是在争论谁能用可持续的成本、合规边界和生态控制力把它真正做成生意。\u003c/p\u003e\n\u003ch2 id=\"-大佬观点influencer-insights\"\u003e\n  💡 大佬观点(Influencer Insights)\n  \u003ca class=\"heading-link\" href=\"#-%e5%a4%a7%e4%bd%ac%e8%a7%82%e7%82%b9influencer-insights\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cblockquote\u003e\n\u003cp\u003e今日大佬观点暂缺,推荐阅读 Watch List 深度内容。\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch2 id=\"-附录今日-watch-list-更新源列表\"\u003e\n  📚 附录:今日 Watch List 更新源列表\n  \u003ca class=\"heading-link\" href=\"#-%e9%99%84%e5%bd%95%e4%bb%8a%e6%97%a5-watch-list-%e6%9b%b4%e6%96%b0%e6%ba%90%e5%88%97%e8%a1%a8\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cblockquote\u003e\n\u003cp\u003e时间窗口:最近 3 天;覆盖 22 个源;共 34 条更新\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch3 id=\"openai-blog-a_full\"\u003e\n  OpenAI Blog (A_full)\n  \u003ca class=\"heading-link\" href=\"#openai-blog-a_full\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://openai.com/index/what-students-gain-from-chatgpt-critical-thinking-training\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eBetter answers, broader thinking: What students gain from ChatGPT and critical-thinking training\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 17:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- What happens when students use ChatGPT on a real-world assignment?\n\u003cul\u003e\n\u003cli\u003eDo quality improvements come at the expense of originality?\u003c/li\u003e\n\u003cli\u003eA new experiment from researchers at Bocconi University, in collaboration with OpenAI Economic Research, found distinct and complementary effects from ChatGPT access and critical-thinking training.\u003c/li\u003e\n\u003cli\u003eAccess to ChatGPT improved the quality and coherence of students’ work, while an exercise in causal reasoning—a form of critical thinking—led students to generate more unique ideas.\u003c/li\u003e\n\u003cli\u003eStudents who received both ChatGPT access and the training showed both effects.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003eA randomized study of more than 1,000 students examines ChatGPT, critical thinking, originality, and student performance on a real-world university assignment.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://openai.com/index/expanding-our-presence-in-brazil\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eExpanding OpenAI’s presence in Brazil\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 11:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- We’re excited to expand our work in Brazil with the launch of our commercial operations.\n\u003cul\u003e\n\u003cli\u003eBased in São Paulo, our local team will work with Brazilian businesses, developers, researchers, and public institutions to help translate the country’s rapid adoption of AI into economic growth and meaningful progress.\u003c/li\u003e\n\u003cli\u003eBrazil is one of ChatGPT’s three largest markets by weekly active users.\u003c/li\u003e\n\u003cli\u003eThe number of users in the country has nearly doubled over the past year, and people in Brazil now send approximately 215 million messages to ChatGPT each day.\u003c/li\u003e\n\u003cli\u003e\n\u003cblockquote\u003e\n\u003cp\u003e“The most exciting part isn’t just the scale of adoption.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003eOpenAI is expanding its presence in Brazil, deepening engagement with developers, businesses, and communities to support AI adoption across the country.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"google-deepmind-blog-a_full\"\u003e\n  Google DeepMind Blog (A_full)\n  \u003ca class=\"heading-link\" href=\"#google-deepmind-blog-a_full\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://deepmind.google/blog/gemini-omni-1-1-flash-lets-you-build-with-more-control/\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eGemini Omni 1.1 Flash lets you build with more control\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-28 00:11 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- Gemini Omni 1.1 Flash lets you build with more control.\n\u003cul\u003e\n\u003cli\u003eThis piece from Google DeepMind Blog explains how Gemini Omni 1.1 Flash lets you build with more control shapes the broader AI and infrastructure landscape.\u003c/li\u003e\n\u003cli\u003eIt also surfaces practical implications for founders, operators, and investors following Gemini Omni 1.1 Flash lets you build with more control.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003eGemini Omni 1.1 Flash lets you build with more control\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://deepmind.google/blog/piloting-the-worlds-first-double-blind-ai-evaluations/\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003ePiloting the world\u0026rsquo;s first double-blind AI evaluations\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 20:59 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- Piloting the world\u0026rsquo;s first double-blind AI evaluations.\n\u003cul\u003e\n\u003cli\u003eThis piece from Google DeepMind Blog explains how Piloting the world\u0026rsquo;s first double-blind AI evaluations shapes the broader AI and infrastructure landscape.\u003c/li\u003e\n\u003cli\u003eIt also surfaces practical implications for founders, operators, and investors following Piloting the world\u0026rsquo;s first double-blind AI evaluations.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003ePiloting the world\u0026rsquo;s first double-blind AI evaluations\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-csai-b_introsearch\"\u003e\n  ArXiv cs.AI (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-csai-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23568\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eRENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.23568v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Memory and RAG evaluations often treat the answering model\u0026rsquo;s input as an implementation detail, even though systems may render the same history as a memory entry, summary, typed record, or raw excerpt.\u003c/li\u003e\n\u003cli\u003eWe introduce RENDER, a benchmark control that fixes the conversation while varying the reader-facing artifact.\u003c/li\u003e\n\u003cli\u003eRENDER combines a five-level packet ladder, localizing when answer-bearing content enters the input, with deterministic templates approximating ChatGPT-style entries, LangChain summaries, MemGPT-style typed records, and raw conversation.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23568v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Memory and RAG evaluations often treat the answering model\u0026rsquo;s input as an implementation detail, even though systems may render the same history as a m…\u003c/li\u003e\n\u003cli\u003eWe introduce RENDER, a benchmark control that fixes the conversation while varying the reader-facing artifact\u003c/li\u003e\n\u003cli\u003eRENDER combines a five-level packet ladder, localizing when answer-bearing content enters the input, with deterministic templates approximating ChatGPT-style en…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23569\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.23569v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on established benchmarks such as Spider and BIRD.\u003c/li\u003e\n\u003cli\u003eHowever, these benchmarks rely on simplified academic schemas and open-source SQL dialects that do not reflect the complexity of enterprise database environments.\u003c/li\u003e\n\u003cli\u003eWe introduce ESQ-Bench, an Oracle-first NL2SQL benchmark with systematic complexity tiers and silent-divergence evaluation across three enterprise schema complexity tiers.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23569v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on established benchmarks such as Spider and B…\u003c/li\u003e\n\u003cli\u003eHowever, these benchmarks rely on simplified academic schemas and open-source SQL dialects that do not reflect the complexity of enterprise database environment…\u003c/li\u003e\n\u003cli\u003eWe introduce ESQ-Bench, an Oracle-first NL2SQL benchmark with systematic complexity tiers and silent-divergence evaluation across three enterprise schema comple…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23622\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eLLM Agents Perform Controlled Experiments Using Simulation Models\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.23622v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require more than plausible text and code generation.\u003c/li\u003e\n\u003cli\u003eThey require understanding how a system responds to intervention, which in practice depends on controlled experimentation.\u003c/li\u003e\n\u003cli\u003eIn this work, we propose a multi-agent framework that enables LLM agents to conduct controlled experiments with scientific simulation models for pharmaceutical process design.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23622v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require mo…\u003c/li\u003e\n\u003cli\u003eThey require understanding how a system responds to intervention, which in practice depends on controlled experimentation\u003c/li\u003e\n\u003cli\u003eIn this work, we propose a multi-agent framework that enables LLM agents to conduct controlled experiments with scientific simulation models for pharmaceutical…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23626\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eA survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.23626v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Foundation models for astronomy are trained on survey pixels together with the catalogue products derived from those pixels.\u003c/li\u003e\n\u003cli\u003eThose catalogues are incomplete at a measurable rate, and a model trained on both inherits that incompleteness as a systematic.\u003c/li\u003e\n\u003cli\u003eWe audit AION-1, a 39-modality transformer trained on more than 200 million objects, using causal interventions on its inputs.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23626v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Foundation models for astronomy are trained on survey pixels together with the catalogue products derived from those pixels\u003c/li\u003e\n\u003cli\u003eThose catalogues are incomplete at a measurable rate, and a model trained on both inherits that incompleteness as a systematic\u003c/li\u003e\n\u003cli\u003eWe audit AION-1, a 39-modality transformer trained on more than 200 million objects, using causal interventions on its inputs\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23631\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eTRACE: Transition-Aware Residual Control for Multi-Objective Materials Discovery\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.23631v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Multi-objective materials discovery with LLM agents is often limited not only by how many candidates can be proposed, but by how effectively each costly property evaluation informs the next search step.\u003c/li\u003e\n\u003cli\u003eExisting agents mainly store evaluated candidates and their scores, so they know which materials succeeded but not which executable edits caused useful property changes.\u003c/li\u003e\n\u003cli\u003eThis makes local refinement difficult when objectives compete and an edit that improves one property may damage another.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23631v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Multi-objective materials discovery with LLM agents is often limited not only by how many candidates can be proposed, but by how effectively each cost…\u003c/li\u003e\n\u003cli\u003eExisting agents mainly store evaluated candidates and their scores, so they know which materials succeeded but not which executable edits caused useful property…\u003c/li\u003e\n\u003cli\u003eThis makes local refinement difficult when objectives compete and an edit that improves one property may damage another\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23632\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eFunction-Level Execution Feedback for Code Preference Optimization\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.23632v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Process supervision has improved mathematical reasoning, where intermediate steps are naturally expressed as chains of thought.\u003c/li\u003e\n\u003cli\u003eIn code generation, however, process supervision remains underexplored because there is no standard notion of a step.\u003c/li\u003e\n\u003cli\u003eSupervision can target lines, reasoning traces, or program states, making it unclear what to label and optimize.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23632v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Process supervision has improved mathematical reasoning, where intermediate steps are naturally expressed as chains of thought\u003c/li\u003e\n\u003cli\u003eIn code generation, however, process supervision remains underexplored because there is no standard notion of a step\u003c/li\u003e\n\u003cli\u003eSupervision can target lines, reasoning traces, or program states, making it unclear what to label and optimize\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23640\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eAuditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.23640v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: When a large language model (LLM) is asked to write a person\u0026rsquo;s life, how much of what it writes actually happened?\u003c/li\u003e\n\u003cli\u003eWe present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are aware of, based on an unsystematic literature search.\u003c/li\u003e\n\u003cli\u003eThe subject and the author of this paper are the same person: a 366-day \u0026ldquo;page-a-day\u0026rdquo; book of first-person anecdotal entries was drafted with a conversational LLM whose documented inputs were a template, two exemplar days, and each day\u0026rsquo;s quote - not her corpus - and every day was subsequently audited at the anecdote-scene level against an independent verification corpus using a four-level rubric fixed before analysis.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23640v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: When a large language model (LLM) is asked to write a person\u0026rsquo;s life, how much of what it writes actually happened\u003c/li\u003e\n\u003cli\u003eWe present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are…\u003c/li\u003e\n\u003cli\u003eThe subject and the author of this paper are the same person: a 366-day \u0026ldquo;page-a-day\u0026rdquo; book of first-person anecdotal entries was drafted with a conversational LL…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23641\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eHow much of a measured AI preference is the model, and how much is the instrument?\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.23641v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences.\u003c/li\u003e\n\u003cli\u003e(2025), Tagliabue and Dung (2025) and Trhlik et al.\u003c/li\u003e\n\u003cli\u003e(2026) have built four instruments for that purpose, and their findings disagree.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23641v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences\u003c/li\u003e\n\u003cli\u003eKeeling et al\u003c/li\u003e\n\u003cli\u003e(2024), Mazeika et al\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23642\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eAI Agents Push Humans Out of the Loop\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.23642v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: AI agents pose significant risks as they are granted increasing autonomy.\u003c/li\u003e\n\u003cli\u003eA commonly proposed solution is human oversight and keeping a \u0026lsquo;\u0026lsquo;human in the loop\u0026rsquo;\u0026rsquo;, but this is not a simple solution: Not only do current approaches to AI agent design impede effective human oversight, but the cognitive capacities required for it are also themselves degraded by extended use of AI systems.\u003c/li\u003e\n\u003cli\u003eThis position paper argues that current approaches to the development and deployment of AI agent systems do not support effective human oversight \u0026ndash; they contribute to its degradation.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23642v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: AI agents pose significant risks as they are granted increasing autonomy\u003c/li\u003e\n\u003cli\u003eA commonly proposed solution is human oversight and keeping a \u0026lsquo;\u0026lsquo;human in the loop\u0026rsquo;\u0026rsquo;, but this is not a simple solution: Not only do current approaches to AI age…\u003c/li\u003e\n\u003cli\u003eThis position paper argues that current approaches to the development and deployment of AI agent systems do not support effective human oversight \u0026ndash; they contri…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23643\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eFLARE: A Systematic, Uncertainty-Aware Framework for Evidence-Based Adoption of Artificial Intelligence in Healthcare\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.23643v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Artificial intelligence is increasingly being introduced into healthcare workflows, yet most evaluations emphasize model accuracy rather than whether adoption is economically worthwhile in real clinical settings.\u003c/li\u003e\n\u003cli\u003eThis study proposes FLARE, a systematic and uncertainty-aware framework for evaluating the financial and operational implications of adopting AI in healthcare.\u003c/li\u003e\n\u003cli\u003eFLARE combines fuzzy logic, time-driven activity-based costing, and return on investment analysis to estimate the cost of clinical service delivery, the cost of AI development and operation, and the economic consequences of workflow integration under uncertainty.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23643v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Artificial intelligence is increasingly being introduced into healthcare workflows, yet most evaluations emphasize model accuracy rather than whether…\u003c/li\u003e\n\u003cli\u003eThis study proposes FLARE, a systematic and uncertainty-aware framework for evaluating the financial and operational implications of adopting AI in healthcare\u003c/li\u003e\n\u003cli\u003eFLARE combines fuzzy logic, time-driven activity-based costing, and return on investment analysis to estimate the cost of clinical service delivery, the cost of…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-cscl-b_introsearch\"\u003e\n  ArXiv cs.CL (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-cscl-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24901\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDetection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.24901v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: A decodable \u0026ldquo;empathy\u0026rdquo; direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change.\u003c/li\u003e\n\u003cli\u003eWe test this for two EPITOME-derived facets \u0026ndash; Recognition (cognitive) and Resonance (affective) \u0026ndash; in three instruction-tuned LLMs, scoring every intervention with two LLM judges and a discriminative EPITOME classifier, each gated by an emotional-vs-neutral positive control.\u003c/li\u003e\n\u003cli\u003eThe control passes for the affective facet across all automated instruments, but cognitive range is inconsistent across them.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24901v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: A decodable \u0026ldquo;empathy\u0026rdquo; direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change\u003c/li\u003e\n\u003cli\u003eWe test this for two EPITOME-derived facets \u0026ndash; Recognition (cognitive) and Resonance (affective) \u0026ndash; in three instruction-tuned LLMs, scoring every intervention…\u003c/li\u003e\n\u003cli\u003eThe control passes for the affective facet across all automated instruments, but cognitive range is inconsistent across them\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24920\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eSemantic Variability of Replies Across LLMs: Implications for Designing Conversation-Based Assessment\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.24920v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes.\u003c/li\u003e\n\u003cli\u003eUsing messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and without preceding chat history.\u003c/li\u003e\n\u003cli\u003eResults show that model choice and conversational context both affect response similarity and alignment with human replies.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24920v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes\u003c/li\u003e\n\u003cli\u003eUsing messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and withou…\u003c/li\u003e\n\u003cli\u003eResults show that model choice and conversational context both affect response similarity and alignment with human replies\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24952\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eThe Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.24952v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language modeling pipeline remains unclear.\u003c/li\u003e\n\u003cli\u003eOur study traces this \u0026ldquo;dialect tax\u0026rdquo; across the natural language processing pipeline.\u003c/li\u003e\n\u003cli\u003eUsing parallel English dialect corpora that hold meaning fixed while varying surface form, we first confirm that LMs recognize matched Standard American English (SAE) and dialectal texts as semantically equivalent.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24952v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language mod…\u003c/li\u003e\n\u003cli\u003eOur study traces this \u0026ldquo;dialect tax\u0026rdquo; across the natural language processing pipeline\u003c/li\u003e\n\u003cli\u003eUsing parallel English dialect corpora that hold meaning fixed while varying surface form, we first confirm that LMs recognize matched Standard American English…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24982\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eUnsupervised Post-Training of Foundation Models: A Survey\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.24982v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers.\u003c/li\u003e\n\u003cli\u003eWe study Unsupervised Post-Training (UPT): update-bearing adaptation on unlabeled inputs whose learning signal is derived from same-lineage model artifacts rather than an external oracle.\u003c/li\u003e\n\u003cli\u003eWe catalog 80 strict UPT methods and organize them by the object that supplies the update signal: a prediction statistic, a sample relation, a self-generated target, or an internal evaluator.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24982v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers\u003c/li\u003e\n\u003cli\u003eWe study Unsupervised Post-Training (UPT): update-bearing adaptation on unlabeled inputs whose learning signal is derived from same-lineage model artifacts rath…\u003c/li\u003e\n\u003cli\u003eWe catalog 80 strict UPT methods and organize them by the object that supplies the update signal: a prediction statistic, a sample relation, a self-generated ta…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24988\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDoes Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.24988v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Activation steering can be embedded directly into a language model\u0026rsquo;s weights, shaping behaviour without inference-time intervention and offering a way to encode alignment prior to release.\u003c/li\u003e\n\u003cli\u003eHowever, models are routinely fine-tuned after deployment, and it is unknown whether embedded interventions survive this.\u003c/li\u003e\n\u003cli\u003eWe study the stability of embedded steering for refusal suppression and brevity induction across five instruction-tuned models (3B-14B) under non-adversarial SFT and RLHF.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24988v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Activation steering can be embedded directly into a language model\u0026rsquo;s weights, shaping behaviour without inference-time intervention and offering a way…\u003c/li\u003e\n\u003cli\u003eHowever, models are routinely fine-tuned after deployment, and it is unknown whether embedded interventions survive this\u003c/li\u003e\n\u003cli\u003eWe study the stability of embedded steering for refusal suppression and brevity induction across five instruction-tuned models (3B-14B) under non-adversarial SF…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.25005\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eThe Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.25005v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: The imperfective paradox provides a useful test of compositional semantic analysis.\u003c/li\u003e\n\u003cli\u003eRecent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing this behavior to a Teleological Bias.\u003c/li\u003e\n\u003cli\u003eIt further argues that prompting interventions cause a Calibration Crisis.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.25005v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: The imperfective paradox provides a useful test of compositional semantic analysis\u003c/li\u003e\n\u003cli\u003eRecent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing this behavior…\u003c/li\u003e\n\u003cli\u003eIt further argues that prompting interventions cause a Calibration Crisis\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.25022\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eA Primer on Computational Semantics for Artificial Intelligence Systems\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.25022v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: As people adopt transformer-based language models (e.g., ChatGPT and Gemini) for an increasing number of use-cases, it is important to know how such models learn and represent the meaning of the language, and to be more informed about what language is.\u003c/li\u003e\n\u003cli\u003eThis document is an attempt to help the reader understand how linguistic meaning (i.e., semantics) is approached from different fields of scientific and philosophical examination.\u003c/li\u003e\n\u003cli\u003eI also explain three primary semantic theories: formal semantics, grounded semantics, and distributional semantics then compare how transformer-based language models differ from how humans learn language.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.25022v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: As people adopt transformer-based language models (e.g., ChatGPT and Gemini) for an increasing number of use-cases, it is important to know how such m…\u003c/li\u003e\n\u003cli\u003eThis document is an attempt to help the reader understand how linguistic meaning (i.e., semantics) is approached from different fields of scientific and philoso…\u003c/li\u003e\n\u003cli\u003eI also explain three primary semantic theories: formal semantics, grounded semantics, and distributional semantics then compare how transformer-based language m…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.25028\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eBehind the [MASK]: Disentangling Representation and Faithfulness in DAPF-Based Dementia Detection\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.25028v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Spoken-language analysis via prompt-based domain-adaptive models is a promising direction for low-resource, non-invasive dementia screening, but such models remain internally opaque.\u003c/li\u003e\n\u003cli\u003eWe study the interpretability of the Domain-Adapted models via Prompt-based Fine-tuning (DAPF) framework, which casts dementia detection as diagnosis-related masked-token prediction.\u003c/li\u003e\n\u003cli\u003eWe interpret DAPF and strong baselines using a variety of probing and analysis techniques, finding that DAPF achieved the best overall performance (accuracy=0.83 and macro-F1=0.83) with diagnosis most recoverable from its [MASK] representation.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.25028v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Spoken-language analysis via prompt-based domain-adaptive models is a promising direction for low-resource, non-invasive dementia screening, but such…\u003c/li\u003e\n\u003cli\u003eWe study the interpretability of the Domain-Adapted models via Prompt-based Fine-tuning (DAPF) framework, which casts dementia detection as diagnosis-related ma…\u003c/li\u003e\n\u003cli\u003eWe interpret DAPF and strong baselines using a variety of probing and analysis techniques, finding that DAPF achieved the best overall performance (accuracy=0.8…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.25038\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003ePadamitra: Grounded Glossary Generation for Classical Sanskrit\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.25038v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation-grounded meanings from a sloka-translation pair, formalizing the traditional patha commentary practice as an evaluable NLP objective.\u003c/li\u003e\n\u003cli\u003eWe construct a benchmark of 31,316 sloka-translation-glossary triples from the Valmiki Ramayana and Srimad Bhagavatam, paired with two metrics: Jaccard for phrase recovery and Meaning Faithfulness for semantic consistency.\u003c/li\u003e\n\u003cli\u003eAcross zero-shot, few-shot, and instruction fine-tuned variants of Gemma-3n-E4B, Gemma-3-12B, Phi-4, and Qwen3.5-9B, instruction fine-tuning substantially outperforms prompting, while explicit segmentation yields gains.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.25038v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translat…\u003c/li\u003e\n\u003cli\u003eWe construct a benchmark of 31,316 sloka-translation-glossary triples from the Valmiki Ramayana and Srimad Bhagavatam, paired with two metrics: Jaccard for phra…\u003c/li\u003e\n\u003cli\u003eAcross zero-shot, few-shot, and instruction fine-tuned variants of Gemma-3n-E4B, Gemma-3-12B, Phi-4, and Qwen3.5-9B, instruction fine-tuning substantially outpe…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.25061\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDataKernelBench: Can LLMs Optimize Database Queries on GPUs?\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.25061v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels.\u003c/li\u003e\n\u003cli\u003eExisting LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operators untested.\u003c/li\u003e\n\u003cli\u003eWe introduce DataKernelBench, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimize either the core tensor-bounded snippet or the full query in CUDA or Triton through execution-guided repair.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.25061v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels\u003c/li\u003e\n\u003cli\u003eExisting LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operators untested\u003c/li\u003e\n\u003cli\u003eWe introduce DataKernelBench, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimize either the core tensor-bounded sni…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-cslg-b_introsearch\"\u003e\n  ArXiv cs.LG (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-cslg-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"链接到标题\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003e链接到标题\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24904\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDynamic Influence-Weighted Distillation for Single-IMU Activity Recognition\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.24904v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Inertial sensors at multiple body locations can improve activity recognition, but requiring every sensor at inference increases the deployment burden.\u003c/li\u003e\n\u003cli\u003eWe study whether four synchronized IMUs available during training can improve a student that uses only the right-arm IMU during fitting and inference.\u003c/li\u003e\n\u003cli\u003eA frozen four-IMU teacher provides logit and feature targets.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24904v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Inertial sensors at multiple body locations can improve activity recognition, but requiring every sensor at inference increases the deployment burden\u003c/li\u003e\n\u003cli\u003eWe study whether four synchronized IMUs available during training can improve a student that uses only the right-arm IMU during fitting and inference\u003c/li\u003e\n\u003cli\u003eA frozen four-IMU teacher provides logit and feature targets\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24936\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eGreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.24936v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: We present GreenLeaf Law Embed Tiny, a 0.6B parameter embedding model for legal domain retrieval.\u003c/li\u003e\n\u003cli\u003eGreenLeaf-Tiny achieves 75.11% on the Massive Legal Embedding Benchmark (MLEB) and 64.38% on MTEB(Law, v1),demonstrating competitive performance among models under 1B parameters.\u003c/li\u003e\n\u003cli\u003eOur approach combines a two-stage training pipeline that first distills knowledge from a larger teacher model into a compact student architecture, then applies domain-specific fine-tuning with hard negative mining; a carefully curated dataset of 3.4 million query-passage pairs, including 150,000 human-curated samples across diverse legal jurisdictions; and an efficient inference architecture supporting multiple quantization levels (BF16, INT8, binary) enabling deployment in resource-constrained environments.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24936v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: We present GreenLeaf Law Embed Tiny, a 0.6B parameter embedding model for legal domain retrieval\u003c/li\u003e\n\u003cli\u003eGreenLeaf-Tiny achieves 75.11% on the Massive Legal Embedding Benchmark (MLEB) and 64.38% on MTEB(Law, v1),demonstrating competitive performance among models un…\u003c/li\u003e\n\u003cli\u003eOur approach combines a two-stage training pipeline that first distills knowledge from a larger teacher model into a compact student architecture, then applies…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24937\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eMulti-Modal Anomaly Detection: A Survey\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.24937v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Multi-Modal Anomaly Detection (MMAD) detects rare abnormal events from heterogeneous data sources and is increasingly used in safety- and reliability-critical applications such as industrial inspection and cybersecurity.\u003c/li\u003e\n\u003cli\u003eYet the literature is fragmented across domains and modality combinations, and existing surveys usually group methods by architecture rather than by how abnormality is defined and separated in multi-modal settings.\u003c/li\u003e\n\u003cli\u003eWe survey MMAD from an assumption-driven perspective.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24937v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Multi-Modal Anomaly Detection (MMAD) detects rare abnormal events from heterogeneous data sources and is increasingly used in safety- and reliability-…\u003c/li\u003e\n\u003cli\u003eYet the literature is fragmented across domains and modality combinations, and existing surveys usually group methods by architecture rather than by how abnorma…\u003c/li\u003e\n\u003cli\u003eWe survey MMAD from an assumption-driven perspective\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24938\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.24938v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation.\u003c/li\u003e\n\u003cli\u003eYet low-latency MoE serving is increasingly challenging, because it spans two inference phases with fundamentally different bottlenecks: prefill is dominated by token-wise expert computation, whereas decode is constrained by memory traffic from the batch-wise activated expert set.\u003c/li\u003e\n\u003cli\u003eHowever, existing training-free acceleration methods optimize only a single resource proxy, either the experts each token executes or the experts a batch activates, and either discard the excluded experts\u0026rsquo; contribution or leave it only implicitly approximated.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24938v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation\u003c/li\u003e\n\u003cli\u003eYet low-latency MoE serving is increasingly challenging, because it spans two inference phases with fundamentally different bottlenecks: prefill is dominated by…\u003c/li\u003e\n\u003cli\u003eHowever, existing training-free acceleration methods optimize only a single resource proxy, either the experts each token executes or the experts a batch activa…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24940\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eWhen Does Frequency Decomposition Benefit Physics-Informed Neural Networks? A Preliminary Ablation Study\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.24940v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Partial differential equations (PDEs) often have high-frequency and multi-scale features that neural networks struggle to approximate.\u003c/li\u003e\n\u003cli\u003ePhysics-Informed Neural Networks (PINNs) build the governing equations directly into training, but suffer from spectral bias: they learn low-frequency components faster than high-frequency ones.\u003c/li\u003e\n\u003cli\u003eTechniques such as Fourier feature embeddings and sinusoidal activations address this, but most studies assume they help across the board without checking which spectral regimes actually benefit.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24940v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Partial differential equations (PDEs) often have high-frequency and multi-scale features that neural networks struggle to approximate\u003c/li\u003e\n\u003cli\u003ePhysics-Informed Neural Networks (PINNs) build the governing equations directly into training, but suffer from spectral bias: they learn low-frequency component…\u003c/li\u003e\n\u003cli\u003eTechniques such as Fourier feature embeddings and sinusoidal activations address this, but most studies assume they help across the board without checking which…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24945\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eFAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.24945v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resource requirements of LLMs hinder the deployment on resource-constrained devices.\u003c/li\u003e\n\u003cli\u003eAlthough model quantization stands out as an effective approach, conventional quantization approaches typically incur severe performance degradation due to uniform bit-width or simple heuristic sensitivity evaluation.\u003c/li\u003e\n\u003cli\u003eIn this paper, we propose a novel Fisher information-based Adaptive Mixed Precision Weight Quantization approach, i.e., FAMPWQ, which performs layer-adaptive weight quantization for effective LLM inference on commodity GPUs.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24945v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resource requirements of…\u003c/li\u003e\n\u003cli\u003eAlthough model quantization stands out as an effective approach, conventional quantization approaches typically incur severe performance degradation due to unif…\u003c/li\u003e\n\u003cli\u003eIn this paper, we propose a novel Fisher information-based Adaptive Mixed Precision Weight Quantization approach, i.e., FAMPWQ, which performs layer-adaptive we…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24946\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eMacroAgent: Regularity-Aware Macro Legalization with LLM-Agent-Designed Contour Algorithms\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.24946v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Macros constitute a large part of the core area in modern very large-scale integration (VLSI) designs.\u003c/li\u003e\n\u003cli\u003eMoreover, macro positions have a significant impact on the final quality of result (QoR), and macro legalization is typically the final step in determining the macro positions.\u003c/li\u003e\n\u003cli\u003eHowever, existing approaches related to macro legalization either lack robustness or incur substantial computational costs or neglect the regularity between macros.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24946v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Macros constitute a large part of the core area in modern very large-scale integration (VLSI) designs\u003c/li\u003e\n\u003cli\u003eMoreover, macro positions have a significant impact on the final quality of result (QoR), and macro legalization is typically the final step in determining the…\u003c/li\u003e\n\u003cli\u003eHowever, existing approaches related to macro legalization either lack robustness or incur substantial computational costs or neglect the regularity between mac…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24947\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eCAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.24947v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: End-to-end training of multimodal neural networks often exhibits unstable neural dynamics characterized by three coupled failure modes that degrade learning: (i) modality imbalance, where one branch dominates gradient-based optimization; (ii) unstable gating, where noisy confidence cues induce erratic modality selection; and (iii) fusion interference, where modality-specific gradients conflict at the shared fusion layer.\u003c/li\u003e\n\u003cli\u003eWe propose CAT-GS (Calibrated, Adaptive, Thresholded Gating with Fusion Surgery), a neural dynamics-based optimization controller for intelligent computing applications.\u003c/li\u003e\n\u003cli\u003eCAT-GS operates during backpropagation without modifying model architectures, fusion modules, or task losses.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24947v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: End-to-end training of multimodal neural networks often exhibits unstable neural dynamics characterized by three coupled failure modes that degrade le…\u003c/li\u003e\n\u003cli\u003eWe propose CAT-GS (Calibrated, Adaptive, Thresholded Gating with Fusion Surgery), a neural dynamics-based optimization controller for intelligent computing appl…\u003c/li\u003e\n\u003cli\u003eCAT-GS operates during backpropagation without modifying model architectures, fusion modules, or task losses\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24949\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDemystifying Reinforcement Learning Post-Training of Language Models\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.24949v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Reinforcement learning (RL) post-training has emerged as a powerful framework for enhancing the capabilities of large language models (LLMs), enabling impressive reasoning, math, and coding capabilities.\u003c/li\u003e\n\u003cli\u003eYet for many researchers and practitioners, the principles behind classical RL remain a \u0026ldquo;black box\u0026rdquo;.\u003c/li\u003e\n\u003cli\u003eIn this work, we deconstruct the RL post-training algorithm, investigating each step to clarify what is actually happening beneath the surface.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24949v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Reinforcement learning (RL) post-training has emerged as a powerful framework for enhancing the capabilities of large language models (LLMs), enabling…\u003c/li\u003e\n\u003cli\u003eYet for many researchers and practitioners, the principles behind classical RL remain a \u0026ldquo;black box\u0026rdquo;\u003c/li\u003e\n\u003cli\u003eIn this work, we deconstruct the RL post-training algorithm, investigating each step to clarify what is actually happening beneath the surface\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24954\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eAFDBench: A Reasoning-First AI Scientist for NationalWeather Service Forecast Discussions\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e发布时间:2026-08-27 12:00 北京时间\u003c/li\u003e\n\u003cli\u003e摘要:【待翻译】- arXiv:2608.24954v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Large language models (LLMs) hallucinate numerical values when generating high-stakes meteorological text, posing risks for weather communication.\u003c/li\u003e\n\u003cli\u003eWe present AFDBench, an AI meteorologist that generates professional Area Forecast Discussions (AFDs) by reasoning through structured AI weather forecast data from Google\u0026rsquo;s WeatherNext 2.\u003c/li\u003e\n\u003cli\u003eWe introduce AFDBench, the first benchmark for evaluating generative meteorological reasoning, comprising 7,732 expert written discussions from 13 National Weather Service (NWS) offices paired with real AI weather forecast inputs, and three complementary metrics: Met-Align (numerical accuracy), Style-Align (professional dialect adherence), and Input-Grounding (fidelity to source weather data).\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24954v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Large language models (LLMs) hallucinate numerical values when generating high-stakes meteorological text, posing risks for weather communication\u003c/li\u003e\n\u003cli\u003eWe present AFDBench, an AI meteorologist that generates professional Area Forecast Discussions (AFDs) by reasoning through structured AI weather forecast data f…\u003c/li\u003e\n\u003cli\u003eWe introduce AFDBench, the first benchmark for evaluating generative meteorological reasoning, comprising 7,732 expert written discussions from 13 National Weat…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n",
  "wordCount": 5435,
  "readingTime": 26,
  "tableOfContents": "\u003cnav id=\"TableOfContents\"\u003e\n  \u003cul\u003e\n    \u003cli\u003e\u003ca href=\"#-本期-watch-list-深度导读\"\u003e📖 本期 Watch List 深度导读\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-x-平台-ai-热点快讯\"\u003e🌐 X 平台 AI 热点快讯\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#话题-1cursor-launches-scratch-to-deploy-web-app-builder\"\u003e话题 1:Cursor Launches Scratch-to-Deploy Web App Builder\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#话题-2tech-giants-urge-urgent-ai-cyber-defense-action\"\u003e话题 2:Tech Giants Urge Urgent AI Cyber Defense Action\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#话题-3google-launches-gemini-35-transcribe-for-precise-speech-to-text\"\u003e话题 3:Google Launches Gemini 3.5 Transcribe for Precise Speech-to-Text\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#话题-4google-launches-gemini-omni-11-flash-for-advanced-video-creation\"\u003e话题 4:Google Launches Gemini Omni 1.1 Flash for Advanced Video Creation\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#话题-5salesforce-beats-earnings-expectations-with-claude-ai-integration\"\u003e话题 5:Salesforce Beats Earnings Expectations with Claude AI Integration\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#话题-6nvidia-acquires-hugging-face-for-129-billion-in-major-ai-deal\"\u003e话题 6:Nvidia Acquires Hugging Face for $12.9 Billion in Major AI Deal\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#话题-7neil-movva-breaks-down-ai-inference-economics-on-invest-like-the-best\"\u003e话题 7:Neil Movva Breaks Down AI Inference Economics on Invest Like the Best\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#话题-8tesla-adds-79-model-ys-to-texas-robotaxi-fleet-in-one-day\"\u003e话题 8:Tesla Adds 79 Model Ys to Texas Robotaxi Fleet in One Day\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-大佬观点influencer-insights\"\u003e💡 大佬观点(Influencer Insights)\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-附录今日-watch-list-更新源列表\"\u003e📚 附录:今日 Watch List 更新源列表\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#openai-blog-a_full\"\u003eOpenAI Blog (A_full)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#google-deepmind-blog-a_full\"\u003eGoogle DeepMind Blog (A_full)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-csai-b_introsearch\"\u003eArXiv cs.AI (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-cscl-b_introsearch\"\u003eArXiv cs.CL (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-cslg-b_introsearch\"\u003eArXiv cs.LG (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n  \u003c/ul\u003e\n\u003c/nav\u003e",
  "isDraft": false
}
