{
  "title": "2026-08-28 AI Daily | AI Evolves from Usable to Verifiable: Double-Blind Evaluations Implemented, Speech and Video Capabilities Continue to Advance",
  "url": "https://miaok.ong/en/ai-daily/ai-daily-2026-08-28/",
  "date": "2026-08-28T07:00:00+08:00",
  "lastmod": "2026-08-28T07:00:00+08:00",
  "type": "ai-daily",
  "kind": "page",
  "language": "en",
  "description": "Today\u0026rsquo;s focus is not on single-point model refreshes, but on capabilities entering a more rigorous validation and delivery phase. Google DeepMind is advancing double-blind evaluations and more controllable generation capabilities, while products like speech transcription and video generation continue to move towards low-latency, long-context, and production workflows. Another thread is that evaluation reliability is being re-examined, with many metrics starting to expose biases and distortions in the tests themselves.",
  "keywords": null,
  "tags": [],
  "categories": [],
  "author": "Mark (Miao) Kong",
  "image": "https://miaok.ong/images/avatar.jpg",
  "content": "\u003ch1 id=\"2026-08-28-ai-daily--from-usable-to-verifiable-double-blind-evaluation-arrives-as-voice-and-video-capabilities-advance\"\u003e\n  2026-08-28 AI Daily | From Usable to Verifiable: Double-Blind Evaluation Arrives as Voice and Video Capabilities Advance\n  \u003ca class=\"heading-link\" href=\"#2026-08-28-ai-daily--from-usable-to-verifiable-double-blind-evaluation-arrives-as-voice-and-video-capabilities-advance\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h1\u003e\n\u003cblockquote\u003e\n\u003cp\u003eToday\u0026rsquo;s focus isn\u0026rsquo;t on isolated model updates, but on capabilities entering a more rigorous phase of validation and delivery. Google DeepMind is advancing double-blind evaluations and more controllable generation capabilities, while products for speech-to-text and video generation continue to move towards low latency, long context, and production workflows. Another key thread is the re-evaluation of benchmark reliability, as many metrics are beginning to reveal the biases and distortions of the tests themselves.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch2 id=\"-in-depth-guide-to-this-issues-watch-list\"\u003e\n  📖 In-depth Guide to This Issue\u0026rsquo;s Watch List\n  \u003ca class=\"heading-link\" href=\"#-in-depth-guide-to-this-issues-watch-list\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003eThere are three main themes to follow today. The first is \u0026ldquo;Model Controllability and Post-Training\u0026rdquo;: Google DeepMind\u0026rsquo;s Gemini Omni 1.1 Flash, double-blind AI evaluations, and several papers on unsupervised post-training, activation steering, and the effects of fine-tuning all point to the same issue—models are growing more powerful, but the boundaries of what is truly controllable and verifiable are still being redrawn. The second is \u0026ldquo;Benchmark Reliability\u0026rdquo;: From dialect bias and semantic reply consistency to the \u0026ldquo;imperfective paradox\u0026rdquo; distorting benchmarks, multiple studies today warn that for many seemingly stable metrics, the tests themselves may be the first to break. The third is \u0026ldquo;Structured Task Implementation\u0026rdquo;: ESQ-Bench, DataKernelBench, and new frameworks for evaluating memory/RAG show that enterprise-level NL2SQL, database optimization, and retrieval systems are now undergoing more rigorous real-world scenario testing.\u003c/p\u003e\n\u003ch2 id=\"-ai-hot-topics-on-x\"\u003e\n  🌐 AI Hot Topics on X\n  \u003ca class=\"heading-link\" href=\"#-ai-hot-topics-on-x\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003ch3 id=\"topic-1-cursor-launches-scratch-to-deploy-web-app-builder\"\u003e\n  Topic 1: Cursor Launches Scratch-to-Deploy Web App Builder\n  \u003ca class=\"heading-link\" href=\"#topic-1-cursor-launches-scratch-to-deploy-web-app-builder\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eSummary: Trending Time: , Related Posts: 203\u003c/li\u003e\n\u003cli\u003eWhat it is: Cursor has launched a web app builder that can build from scratch and deploy directly, advancing AI programming from prototype creation to live delivery.\u003c/li\u003e\n\u003cli\u003eWhy it matters: This indicates that AI programming tools are shifting from \u0026lsquo;code assistance\u0026rsquo; to \u0026rsquo;end-to-end application delivery,\u0026rsquo; which directly impacts development workflows, product iteration speed, and the commercialization of AI-native applications.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: The discussion on X focuses on two main points: first, whether this Scratch-to-Deploy approach can genuinely lower the barrier for non-engineering users to create functional applications; and second, whether it represents a new paradigm or merely an integration of capabilities compared to existing low-code platforms, IDEs, and AI coding assistants.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-2-tech-giants-urge-urgent-ai-cyber-defense-action\"\u003e\n  Topic 2: Tech Giants Urge Urgent AI Cyber Defense Action\n  \u003ca class=\"heading-link\" href=\"#topic-2-tech-giants-urge-urgent-ai-cyber-defense-action\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eSummary: Trending Time: 6 hours ago, Related Posts: 8200\u003c/li\u003e\n\u003cli\u003eWhat it is: Several major tech companies are calling for more urgent action on AI cyber defense measures to address the risk of generative AI being used for attacks and the resulting imbalance in security.\u003c/li\u003e\n\u003cli\u003eWhy it matters: This is important because AI is simultaneously amplifying both offensive and defensive capabilities. If protection systems don\u0026rsquo;t keep pace, models, data, and infrastructure could become targets for larger-scale cyber threats.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: The discussion on X centers on two points: whether businesses and governments have underestimated the security risks posed by AI, and whether the focus should be on industry self-regulation, technical standards, or stricter regulations to drive the implementation of AI cyber defenses.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-3-google-launches-gemini-35-transcribe-for-precise-speech-to-text\"\u003e\n  Topic 3: Google Launches Gemini 3.5 Transcribe for Precise Speech-to-Text\n  \u003ca class=\"heading-link\" href=\"#topic-3-google-launches-gemini-35-transcribe-for-precise-speech-to-text\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eSummary: Trending Time: 1 day ago, Related Posts: 5500\u003c/li\u003e\n\u003cli\u003eWhat it is: Google has released Gemini 3.5 Transcribe for high-precision speech-to-text. It supports automatic recognition of over 85 languages, removal of filler words, multi-speaker differentiation, and offers both offline and low-latency live streaming modes.\u003c/li\u003e\n\u003cli\u003eWhy it matters: This signifies that speech recognition is evolving from a general capability into an infrastructure-level service that can be directly embedded into production workflows. It has a direct impact on meeting minutes, customer service, media transcription, and multilingual applications, and reflects the competition in AI productization shifting towards real-time performance, accuracy, and ease of developer integration.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: The discussion on X focuses on several points: whether its multilingual and low-latency performance is sufficient to compete with existing transcription solutions, the value of its 85-language support and custom vocabularies for industry-specific scenarios, and whether Google is using this to further embed voice capabilities into the Gemini ecosystem and developer tools.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-4-google-launches-gemini-omni-11-flash-for-advanced-video-creation\"\u003e\n  Topic 4: Google Launches Gemini Omni 1.1 Flash for Advanced Video Creation\n  \u003ca class=\"heading-link\" href=\"#topic-4-google-launches-gemini-omni-11-flash-for-advanced-video-creation\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eSummary: Trending Time: 7 hours ago, Related Posts: 3000\u003c/li\u003e\n\u003cli\u003eWhat it is: Google has released Gemini Omni 1.1 Flash for more advanced video creation, adding new capabilities like scene extension and the ability to analyze up to 10 seconds of footage to maintain narrative coherence.\u003c/li\u003e\n\u003cli\u003eWhy it matters: This shows that generative video models are moving from single-clip generation towards stronger contextual understanding and continuity control, which directly impacts the usability, cost, and production efficiency of AI video creation.\u003c/li\u003e\n\u003cli\u003eDiscussion Summary: Discussions on X are focused on whether it truly improves video narrative consistency, its actual performance compared to existing video generation models, and how such capabilities will impact content creation workflows, copyright, and authenticity issues.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-5-salesforce-beats-earnings-expectations-with-claude-ai-integration\"\u003e\n  Topic 5: Salesforce Beats Earnings Expectations with Claude AI Integration\n  \u003ca class=\"heading-link\" href=\"#topic-5-salesforce-beats-earnings-expectations-with-claude-ai-integration\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending Time: 1 day ago, Related Posts: 8100\u003c/li\u003e\n\u003cli\u003eWhat Happened: Salesforce announced earnings that exceeded market expectations, highlighting its Claude AI integration as a key performance driver.\u003c/li\u003e\n\u003cli\u003eWhy It Matters: This indicates that generative AI is moving from proof-of-concept to driving actual revenue and product differentiation in enterprise software, signaling progress in AI\u0026rsquo;s commercialization path and enterprise-level adoption.\u003c/li\u003e\n\u003cli\u003eDiscussion Summary: Discussions on X are centered on whether the AI integration is genuinely driving Salesforce\u0026rsquo;s growth, Claude\u0026rsquo;s competitiveness in the enterprise sector, and the impact of such collaborations on revenue, profit margins, and dependence on Anthropic.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-6-nvidia-acquires-hugging-face-for-129-billion-in-major-ai-deal\"\u003e\n  Topic 6: Nvidia Acquires Hugging Face for $12.9 Billion in Major AI Deal\n  \u003ca class=\"heading-link\" href=\"#topic-6-nvidia-acquires-hugging-face-for-129-billion-in-major-ai-deal\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending Time: 19 hours ago, Related Posts: 22000\u003c/li\u003e\n\u003cli\u003eWhat Happened: A trending report on X claims that Nvidia has acquired Hugging Face for $12.9 billion, a deal that has attracted widespread attention in the AI community.\u003c/li\u003e\n\u003cli\u003eWhy It Matters: Hugging Face is a critical gateway to the open-source model and tool ecosystem. If acquired by Nvidia, it could reshape the power structure of AI infrastructure, model distribution, and the developer ecosystem.\u003c/li\u003e\n\u003cli\u003eDiscussion Summary: The discussion is focused on the impact of this deal on open-source neutrality, whether Nvidia will further strengthen its control over the full AI stack, and if this will change the choices of the model community and enterprise users.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-7-neil-movva-breaks-down-ai-inference-economics-on-invest-like-the-best\"\u003e\n  Topic 7: Neil Movva Breaks Down AI Inference Economics on Invest Like the Best\n  \u003ca class=\"heading-link\" href=\"#topic-7-neil-movva-breaks-down-ai-inference-economics-on-invest-like-the-best\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending Time: 2 days ago, Related Posts: 4100\u003c/li\u003e\n\u003cli\u003eWhat Happened: On the \u0026ldquo;Invest Like the Best\u0026rdquo; podcast, Neil Movva discussed and broke down the economic model of AI inference, focusing on compute costs, pricing, and commercialization paths.\u003c/li\u003e\n\u003cli\u003eWhy It Matters: Inference costs directly determine the gross margins, scaling speed, and competitive landscape of AI products. Therefore, this type of analysis influences the business decisions of model vendors, cloud service providers, and application-layer companies.\u003c/li\u003e\n\u003cli\u003eDiscussion Summary: Discussions on X are centered on whether inference costs will decline rapidly, who will gain an advantage in compute and infrastructure, and whether AI companies can build sustainable revenue models amid high compute consumption.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-8-tesla-adds-79-model-ys-to-texas-robotaxi-fleet-in-one-day\"\u003e\n  Topic 8: Tesla Adds 79 Model Ys to Texas Robotaxi Fleet in One Day\n  \u003ca class=\"heading-link\" href=\"#topic-8-tesla-adds-79-model-ys-to-texas-robotaxi-fleet-in-one-day\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending Time:, Related Posts: 823\u003c/li\u003e\n\u003cli\u003eWhat Happened: Tesla reportedly added 79 Model Y vehicles to its Texas Robotaxi fleet in a single day for the expansion of its autonomous ride-hailing service.\u003c/li\u003e\n\u003cli\u003eWhy It Matters: This suggests Tesla is accelerating the conversion of its mass-produced cars into an operational autonomous fleet, which is crucial for the real-world application, large-scale deployment, and commercial validation of AI in transportation.\u003c/li\u003e\n\u003cli\u003eDiscussion Summary: Discussions on X are focused on whether this marks a substantial expansion phase for the Robotaxi project, and whether the vehicles\u0026rsquo; autonomous capabilities, regulatory compliance, operational range, and data feedback are sufficient to support Tesla\u0026rsquo;s narrative. The main point of contention is whether this represents a genuine commercial breakthrough or merely a limited fleet expansion and marketing effort.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch4 id=\"todays-ai-public-opinion-summary-on-x\"\u003e\n  Today\u0026rsquo;s AI Public Opinion Summary on X\n  \u003ca class=\"heading-link\" href=\"#todays-ai-public-opinion-summary-on-x\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h4\u003e\n\u003cp\u003eToday\u0026rsquo;s main narrative is consistent: AI is shifting from \u0026ldquo;demo-ready\u0026rdquo; to \u0026ldquo;delivery-ready.\u0026rdquo; Whether it\u0026rsquo;s Cursor\u0026rsquo;s zero-to-deployment, Google\u0026rsquo;s speech-to-text and video generation, or Salesforce\u0026rsquo;s enterprise revenue validation, the focus is on usability, integration, and commercialization. The clear consensus is that the deciding factor for success is no longer standalone model capabilities, but rather the ability to integrate into workflows, lower the barrier to entry, and create a viable business model based on inference costs and product revenue. Disagreements primarily fall into two categories: first, whether these announcements represent a paradigm shift or simply a repackaging of existing IDE, low-code, cloud, and model capabilities; and second, whether news like the Tesla Robotaxi fleet and the Nvidia-Hugging Face acquisition signifies substantial expansion or is driven more by capital and narrative. Potential risks are also repeatedly mentioned: the imbalance in AI-driven cyber offense and defense will expose models, data, and infrastructure to more frequent and larger-scale threats; enhanced video and transcription capabilities will amplify issues of copyright, authenticity, and content misuse; and if infrastructure and distribution channels continue to consolidate among a few giants, open-source neutrality, developer choice, and market competition will be under pressure. Overall, today\u0026rsquo;s debate is not about \u003cem\u003eif\u003c/em\u003e AI will be implemented, but about \u003cem\u003ewho\u003c/em\u003e can turn it into a real business with sustainable costs, within regulatory boundaries, and with control over the ecosystem.\u003c/p\u003e\n\u003ch2 id=\"-influencer-insights\"\u003e\n  💡 Influencer Insights\n  \u003ca class=\"heading-link\" href=\"#-influencer-insights\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cblockquote\u003e\n\u003cp\u003eInfluencer insights are unavailable today. We recommend reading the in-depth content from the Watch List.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch2 id=\"-appendix-todays-watch-list-update-source-list\"\u003e\n  📚 Appendix: Today\u0026rsquo;s Watch List Update Source List\n  \u003ca class=\"heading-link\" href=\"#-appendix-todays-watch-list-update-source-list\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cblockquote\u003e\n\u003cp\u003eTime window: last 3 days; covers 22 sources; 34 updates in total.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch3 id=\"openai-blog-a_full\"\u003e\n  OpenAI Blog (A_full)\n  \u003ca class=\"heading-link\" href=\"#openai-blog-a_full\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://openai.com/index/what-students-gain-from-chatgpt-critical-thinking-training\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eBetter answers, broader thinking: What students gain from ChatGPT and critical-thinking training\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-08-27 17:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary: [Translation Pending] - What happens when students use ChatGPT on a real-world assignment?\n\u003cul\u003e\n\u003cli\u003eDo quality improvements come at the expense of originality?\u003c/li\u003e\n\u003cli\u003eA new experiment from researchers at Bocconi University, in collaboration with OpenAI Economic Research, found distinct and complementary effects from ChatGPT access and critical-thinking training.\u003c/li\u003e\n\u003cli\u003eAccess to ChatGPT improved the quality and coherence of students’ work, while an exercise in causal reasoning—a form of critical thinking—led students to generate more unique ideas.\u003c/li\u003e\n\u003cli\u003eStudents who received both ChatGPT access and the training showed both effects.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key points:\n\u003cul\u003e\n\u003cli\u003eA randomized study of more than 1,000 students examines ChatGPT, critical thinking, originality, and student performance on a real-world university assignment.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://openai.com/index/expanding-our-presence-in-brazil\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eExpanding OpenAI’s presence in Brazil\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-08-27 11:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary: [Translation Pending] - We’re excited to expand our work in Brazil with the launch of our commercial operations.\n\u003cul\u003e\n\u003cli\u003eBased in São Paulo, our local team will work with Brazilian businesses, developers, researchers, and public institutions to help translate the country’s rapid adoption of AI into economic growth and meaningful progress.\u003c/li\u003e\n\u003cli\u003eBrazil is one of ChatGPT’s three largest markets by weekly active users.\u003c/li\u003e\n\u003cli\u003eThe number of users in the country has nearly doubled over the past year, and people in Brazil now send approximately 215 million messages to ChatGPT each day.\u003c/li\u003e\n\u003cli\u003e\n\u003cblockquote\u003e\n\u003cp\u003e“The most exciting part isn’t just the scale of adoption.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key points:\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eOpenAI is expanding its presence in Brazil, deepening engagement with developers, businesses, and communities to support AI adoption across the country.\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"google-deepmind-blog-a_full\"\u003e\n  Google DeepMind Blog (A_full)\n  \u003ca class=\"heading-link\" href=\"#google-deepmind-blog-a_full\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://deepmind.google/blog/gemini-omni-1-1-flash-lets-you-build-with-more-control/\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eGemini Omni 1.1 Flash lets you build with more control\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished at: 2026-08-28 00:11 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary: [To be translated] - Gemini Omni 1.1 Flash lets you build with more control.\n\u003cul\u003e\n\u003cli\u003eThis piece from Google DeepMind Blog explains how Gemini Omni 1.1 Flash lets you build with more control shapes the broader AI and infrastructure landscape.\u003c/li\u003e\n\u003cli\u003eIt also surfaces practical implications for founders, operators, and investors following Gemini Omni 1.1 Flash lets you build with more control.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003eGemini Omni 1.1 Flash lets you build with more control\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://deepmind.google/blog/piloting-the-worlds-first-double-blind-ai-evaluations/\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003ePiloting the world\u0026rsquo;s first double-blind AI evaluations\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished at: 2026-08-27 20:59 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary: [To be translated] - Piloting the world\u0026rsquo;s first double-blind AI evaluations.\n\u003cul\u003e\n\u003cli\u003eThis piece from Google DeepMind Blog explains how Piloting the world\u0026rsquo;s first double-blind AI evaluations shapes the broader AI and infrastructure landscape.\u003c/li\u003e\n\u003cli\u003eIt also surfaces practical implications for founders, operators, and investors following Piloting the world\u0026rsquo;s first double-blind AI evaluations.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003ePiloting the world\u0026rsquo;s first double-blind AI evaluations\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-csai-b_introsearch\"\u003e\n  ArXiv cs.AI (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-csai-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23568\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eRENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished at: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary: [To be translated] - arXiv:2608.23568v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Memory and RAG evaluations often treat the answering model\u0026rsquo;s input as an implementation detail, even though systems may render the same history as a memory entry, summary, typed record, or raw excerpt.\u003c/li\u003e\n\u003cli\u003eWe introduce RENDER, a benchmark control that fixes the conversation while varying the reader-facing artifact.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eRENDER combines a five-level packet ladder, localizing when answer-bearing content enters the input, with deterministic templates approximating ChatGPT-style entries, LangChain summaries, MemGPT-style typed records, and raw conversation.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23568v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Memory and RAG evaluations often treat the answering model\u0026rsquo;s input as an implementation detail, even though systems may render the same history as a m…\u003c/li\u003e\n\u003cli\u003eWe introduce RENDER, a benchmark control that fixes the conversation while varying the reader-facing artifact\u003c/li\u003e\n\u003cli\u003eRENDER combines a five-level packet ladder, localizing when answer-bearing content enters the input, with deterministic templates approximating ChatGPT-style en…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23569\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [To be translated] - arXiv:2608.23569v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on established benchmarks such as Spider and BIRD.\u003c/li\u003e\n\u003cli\u003eHowever, these benchmarks rely on simplified academic schemas and open-source SQL dialects that do not reflect the complexity of enterprise database environments.\u003c/li\u003e\n\u003cli\u003eWe introduce ESQ-Bench, an Oracle-first NL2SQL benchmark with systematic complexity tiers and silent-divergence evaluation across three enterprise schema complexity tiers.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23569v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on established benchmarks such as Spider and B…\u003c/li\u003e\n\u003cli\u003eHowever, these benchmarks rely on simplified academic schemas and open-source SQL dialects that do not reflect the complexity of enterprise database environment…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe introduce ESQ-Bench, an Oracle-first NL2SQL benchmark with systematic complexity tiers and silent-divergence evaluation across three enterprise schema comple…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23622\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eLLM Agents Perform Controlled Experiments Using Simulation Models\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [To be translated] - arXiv:2608.23622v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require more than plausible text and code generation.\u003c/li\u003e\n\u003cli\u003eThey require understanding how a system responds to intervention, which in practice depends on controlled experimentation.\u003c/li\u003e\n\u003cli\u003eIn this work, we propose a multi-agent framework that enables LLM agents to conduct controlled experiments with scientific simulation models for pharmaceutical process design.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23622v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require mo…\u003c/li\u003e\n\u003cli\u003eThey require understanding how a system responds to intervention, which in practice depends on controlled experimentation\u003c/li\u003e\n\u003cli\u003eIn this work, we propose a multi-agent framework that enables LLM agents to conduct controlled experiments with scientific simulation models for pharmaceutical…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23626\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eA survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [To be translated] - arXiv:2608.23626v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Foundation models for astronomy are trained on survey pixels together with the catalogue products derived from those pixels.\u003c/li\u003e\n\u003cli\u003eThose catalogues are incomplete at a measurable rate, and a model trained on both inherits that incompleteness as a systematic.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe audit AION-1, a 39-modality transformer trained on more than 200 million objects, using causal interventions on its inputs.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eEN Key points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23626v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Foundation models for astronomy are trained on survey pixels together with the catalogue products derived from those pixels\u003c/li\u003e\n\u003cli\u003eThose catalogues are incomplete at a measurable rate, and a model trained on both inherits that incompleteness as a systematic\u003c/li\u003e\n\u003cli\u003eWe audit AION-1, a 39-modality transformer trained on more than 200 million objects, using causal interventions on its inputs\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23631\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eTRACE: Transition-Aware Residual Control for Multi-Objective Materials Discovery\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [To be translated] - arXiv:2608.23631v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Multi-objective materials discovery with LLM agents is often limited not only by how many candidates can be proposed, but by how effectively each costly property evaluation informs the next search step.\u003c/li\u003e\n\u003cli\u003eExisting agents mainly store evaluated candidates and their scores, so they know which materials succeeded but not which executable edits caused useful property changes.\u003c/li\u003e\n\u003cli\u003eThis makes local refinement difficult when objectives compete and an edit that improves one property may damage another.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23631v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Multi-objective materials discovery with LLM agents is often limited not only by how many candidates can be proposed, but by how effectively each cost…\u003c/li\u003e\n\u003cli\u003eExisting agents mainly store evaluated candidates and their scores, so they know which materials succeeded but not which executable edits caused useful property…\u003c/li\u003e\n\u003cli\u003eThis makes local refinement difficult when objectives compete and an edit that improves one property may damage another\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23632\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eFunction-Level Execution Feedback for Code Preference Optimization\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eSummary: arXiv:2608.23632v1 Announce Type: new.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eAbstract: Process supervision has improved mathematical reasoning, where intermediate steps are naturally expressed as chains of thought.\u003c/li\u003e\n\u003cli\u003eIn code generation, however, process supervision remains underexplored because there is no standard notion of a step.\u003c/li\u003e\n\u003cli\u003eSupervision can target lines, reasoning traces, or program states, making it unclear what to label and optimize.\u003c/li\u003e\n\u003cli\u003eKey Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23632v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Process supervision has improved mathematical reasoning, where intermediate steps are naturally expressed as chains of thought\u003c/li\u003e\n\u003cli\u003eIn code generation, however, process supervision remains underexplored because there is no standard notion of a step\u003c/li\u003e\n\u003cli\u003eSupervision can target lines, reasoning traces, or program states, making it unclear what to label and optimize\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23640\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eAuditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary: arXiv:2608.23640v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: When a large language model (LLM) is asked to write a person\u0026rsquo;s life, how much of what it writes actually happened?\u003c/li\u003e\n\u003cli\u003eWe present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are aware of, based on an unsystematic literature search.\u003c/li\u003e\n\u003cli\u003eThe subject and the author of this paper are the same person: a 366-day \u0026ldquo;page-a-day\u0026rdquo; book of first-person anecdotal entries was drafted with a conversational LLM whose documented inputs were a template, two exemplar days, and each day\u0026rsquo;s quote - not her corpus - and every day was subsequently audited at the anecdote-scene level against an independent verification corpus using a four-level rubric fixed before analysis.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eKey Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23640v1 Announce Type: new\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: When a large language model (LLM) is asked to write a person\u0026rsquo;s life, how much of what it writes actually happened\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eThe subject and the author of this paper are the same person: a 366-day \u0026ldquo;page-a-day\u0026rdquo; book of first-person anecdotal entries was drafted with a conversational LL…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23641\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eHow much of a measured AI preference is the model, and how much is the instrument?\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublishing Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [To be translated] - arXiv:2608.23641v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences.\u003c/li\u003e\n\u003cli\u003e(2025), Tagliabue and Dung (2025) and Trhlik et al.\u003c/li\u003e\n\u003cli\u003e(2026) have built four instruments for that purpose, and their findings disagree.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23641v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences\u003c/li\u003e\n\u003cli\u003eKeeling et al\u003c/li\u003e\n\u003cli\u003e(2024), Mazeika et al\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23642\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eAI Agents Push Humans Out of the Loop\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublishing Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [To be translated] - arXiv:2608.23642v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: AI agents pose significant risks as they are granted increasing autonomy.\u003c/li\u003e\n\u003cli\u003eA commonly proposed solution is human oversight and keeping a \u0026lsquo;\u0026lsquo;human in the loop\u0026rsquo;\u0026rsquo;, but this is not a simple solution: Not only do current approaches to AI agent design impede effective human oversight, but the cognitive capacities required for it are also themselves degraded by extended use of AI systems.\u003c/li\u003e\n\u003cli\u003eThis position paper argues that current approaches to the development and deployment of AI agent systems do not support effective human oversight \u0026ndash; they contribute to its degradation.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23642v1 Announce Type: new\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: AI agents pose significant risks as they are granted increasing autonomy\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eA commonly proposed solution is human oversight and keeping a \u0026lsquo;\u0026lsquo;human in the loop\u0026rsquo;\u0026rsquo;, but this is not a simple solution: Not only do current approaches to AI age…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eThis position paper argues that current approaches to the development and deployment of AI agent systems do not support effective human oversight \u0026ndash; they contri…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.23643\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eFLARE: A Systematic, Uncertainty-Aware Framework for Evidence-Based Adoption of Artificial Intelligence in Healthcare\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [Translation Pending] - arXiv:2608.23643v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Artificial intelligence is increasingly being introduced into healthcare workflows, yet most evaluations emphasize model accuracy rather than whether adoption is economically worthwhile in real clinical settings.\u003c/li\u003e\n\u003cli\u003eThis study proposes FLARE, a systematic and uncertainty-aware framework for evaluating the financial and operational implications of adopting AI in healthcare.\u003c/li\u003e\n\u003cli\u003eFLARE combines fuzzy logic, time-driven activity-based costing, and return on investment analysis to estimate the cost of clinical service delivery, the cost of AI development and operation, and the economic consequences of workflow integration under uncertainty.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.23643v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Artificial intelligence is increasingly being introduced into healthcare workflows, yet most evaluations emphasize model accuracy rather than whether…\u003c/li\u003e\n\u003cli\u003eThis study proposes FLARE, a systematic and uncertainty-aware framework for evaluating the financial and operational implications of adopting AI in healthcare\u003c/li\u003e\n\u003cli\u003eFLARE combines fuzzy logic, time-driven activity-based costing, and return on investment analysis to estimate the cost of clinical service delivery, the cost of…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-cscl-b_introsearch\"\u003e\n  ArXiv cs.CL (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-cscl-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24901\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDetection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: arXiv:2608.24901v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: A decodable \u0026ldquo;empathy\u0026rdquo; direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change.\u003c/li\u003e\n\u003cli\u003eWe test this for two EPITOME-derived facets \u0026ndash; Recognition (cognitive) and Resonance (affective) \u0026ndash; in three instruction-tuned LLMs, scoring every intervention with two LLM judges and a discriminative EPITOME classifier, each gated by an emotional-vs-neutral positive control.\u003c/li\u003e\n\u003cli\u003eThe control passes for the affective facet across all automated instruments, but cognitive range is inconsistent across them.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eKey English Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24901v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: A decodable \u0026ldquo;empathy\u0026rdquo; direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change\u003c/li\u003e\n\u003cli\u003eWe test this for two EPITOME-derived facets \u0026ndash; Recognition (cognitive) and Resonance (affective) \u0026ndash; in three instruction-tuned LLMs, scoring every intervention…\u003c/li\u003e\n\u003cli\u003eThe control passes for the affective facet across all automated instruments, but cognitive range is inconsistent across them\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24920\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eSemantic Variability of Replies Across LLMs: Implications for Designing Conversation-Based Assessment\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: arXiv:2608.24920v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes.\u003c/li\u003e\n\u003cli\u003eUsing messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and without preceding chat history.\u003c/li\u003e\n\u003cli\u003eResults show that model choice and conversational context both affect response similarity and alignment with human replies.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eKey English Points:\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003earXiv:2608.24920v1 Announce Type: new\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eUsing messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and withou…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eResults show that model choice and conversational context both affect response similarity and alignment with human replies\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24952\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eThe Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [To be translated] - arXiv:2608.24952v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language modeling pipeline remains unclear.\u003c/li\u003e\n\u003cli\u003eOur study traces this \u0026ldquo;dialect tax\u0026rdquo; across the natural language processing pipeline.\u003c/li\u003e\n\u003cli\u003eUsing parallel English dialect corpora that hold meaning fixed while varying surface form, we first confirm that LMs recognize matched Standard American English (SAE) and dialectal texts as semantically equivalent.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24952v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language mod…\u003c/li\u003e\n\u003cli\u003eOur study traces this \u0026ldquo;dialect tax\u0026rdquo; across the natural language processing pipeline\u003c/li\u003e\n\u003cli\u003eUsing parallel English dialect corpora that hold meaning fixed while varying surface form, we first confirm that LMs recognize matched Standard American English…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24982\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eUnsupervised Post-Training of Foundation Models: A Survey\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [To be translated] - arXiv:2608.24982v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe study Unsupervised Post-Training (UPT): update-bearing adaptation on unlabeled inputs whose learning signal is derived from same-lineage model artifacts rather than an external oracle.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eWe catalog 80 strict UPT methods and organize them by the object that supplies the update signal: a prediction statistic, a sample relation, a self-generated target, or an internal evaluator.\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24982v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers\u003c/li\u003e\n\u003cli\u003eWe study Unsupervised Post-Training (UPT): update-bearing adaptation on unlabeled inputs whose learning signal is derived from same-lineage model artifacts rath…\u003c/li\u003e\n\u003cli\u003eWe catalog 80 strict UPT methods and organize them by the object that supplies the update signal: a prediction statistic, a sample relation, a self-generated ta…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24988\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDoes Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [Translation Pending] - arXiv:2608.24988v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Activation steering can be embedded directly into a language model\u0026rsquo;s weights, shaping behaviour without inference-time intervention and offering a way to encode alignment prior to release.\u003c/li\u003e\n\u003cli\u003eHowever, models are routinely fine-tuned after deployment, and it is unknown whether embedded interventions survive this.\u003c/li\u003e\n\u003cli\u003eWe study the stability of embedded steering for refusal suppression and brevity induction across five instruction-tuned models (3B-14B) under non-adversarial SFT and RLHF.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24988v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Activation steering can be embedded directly into a language model\u0026rsquo;s weights, shaping behaviour without inference-time intervention and offering a way…\u003c/li\u003e\n\u003cli\u003eHowever, models are routinely fine-tuned after deployment, and it is unknown whether embedded interventions survive this\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe study the stability of embedded steering for refusal suppression and brevity induction across five instruction-tuned models (3B-14B) under non-adversarial SF…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.25005\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eThe Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [Translation Pending] - arXiv:2608.25005v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: The imperfective paradox provides a useful test of compositional semantic analysis.\u003c/li\u003e\n\u003cli\u003eRecent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing this behavior to a Teleological Bias.\u003c/li\u003e\n\u003cli\u003eIt further argues that prompting interventions cause a Calibration Crisis.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.25005v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: The imperfective paradox provides a useful test of compositional semantic analysis\u003c/li\u003e\n\u003cli\u003eRecent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing this behavior…\u003c/li\u003e\n\u003cli\u003eIt further argues that prompting interventions cause a Calibration Crisis\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.25022\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eA Primer on Computational Semantics for Artificial Intelligence Systems\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [Translation Pending] - arXiv:2608.25022v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: As people adopt transformer-based language models (e.g., ChatGPT and Gemini) for an increasing number of use-cases, it is important to know how such models learn and represent the meaning of the language, and to be more informed about what language is.\u003c/li\u003e\n\u003cli\u003eThis document is an attempt to help the reader understand how linguistic meaning (i.e., semantics) is approached from different fields of scientific and philosophical examination.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eI also explain three primary semantic theories: formal semantics, grounded semantics, and distributional semantics then compare how transformer-based language models differ from how humans learn language.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.25022v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: As people adopt transformer-based language models (e.g., ChatGPT and Gemini) for an increasing number of use-cases, it is important to know how such m…\u003c/li\u003e\n\u003cli\u003eThis document is an attempt to help the reader understand how linguistic meaning (i.e., semantics) is approached from different fields of scientific and philoso…\u003c/li\u003e\n\u003cli\u003eI also explain three primary semantic theories: formal semantics, grounded semantics, and distributional semantics then compare how transformer-based language m…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.25028\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eBehind the [MASK]: Disentangling Representation and Faithfulness in DAPF-Based Dementia Detection\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [Translation Pending] - arXiv:2608.25028v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Spoken-language analysis via prompt-based domain-adaptive models is a promising direction for low-resource, non-invasive dementia screening, but such models remain internally opaque.\u003c/li\u003e\n\u003cli\u003eWe study the interpretability of the Domain-Adapted models via Prompt-based Fine-tuning (DAPF) framework, which casts dementia detection as diagnosis-related masked-token prediction.\u003c/li\u003e\n\u003cli\u003eWe interpret DAPF and strong baselines using a variety of probing and analysis techniques, finding that DAPF achieved the best overall performance (accuracy=0.83 and macro-F1=0.83) with diagnosis most recoverable from its [MASK] representation.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.25028v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Spoken-language analysis via prompt-based domain-adaptive models is a promising direction for low-resource, non-invasive dementia screening, but such…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe study the interpretability of the Domain-Adapted models via Prompt-based Fine-tuning (DAPF) framework, which casts dementia detection as diagnosis-related ma…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe interpret DAPF and strong baselines using a variety of probing and analysis techniques, finding that DAPF achieved the best overall performance (accuracy=0.8…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.25038\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003ePadamitra: Grounded Glossary Generation for Classical Sanskrit\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary: [Awaiting Translation] - arXiv:2608.25038v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation-grounded meanings from a sloka-translation pair, formalizing the traditional patha commentary practice as an evaluable NLP objective.\u003c/li\u003e\n\u003cli\u003eWe construct a benchmark of 31,316 sloka-translation-glossary triples from the Valmiki Ramayana and Srimad Bhagavatam, paired with two metrics: Jaccard for phrase recovery and Meaning Faithfulness for semantic consistency.\u003c/li\u003e\n\u003cli\u003eAcross zero-shot, few-shot, and instruction fine-tuned variants of Gemma-3n-E4B, Gemma-3-12B, Phi-4, and Qwen3.5-9B, instruction fine-tuning substantially outperforms prompting, while explicit segmentation yields gains.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.25038v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translat…\u003c/li\u003e\n\u003cli\u003eWe construct a benchmark of 31,316 sloka-translation-glossary triples from the Valmiki Ramayana and Srimad Bhagavatam, paired with two metrics: Jaccard for phra…\u003c/li\u003e\n\u003cli\u003eAcross zero-shot, few-shot, and instruction fine-tuned variants of Gemma-3n-E4B, Gemma-3-12B, Phi-4, and Qwen3.5-9B, instruction fine-tuning substantially outpe…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.25061\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDataKernelBench: Can LLMs Optimize Database Queries on GPUs?\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary: [Awaiting Translation] - arXiv:2608.25061v1 Announce Type: new.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eExisting LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operators untested.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe introduce DataKernelBench, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimize either the core tensor-bounded snippet or the full query in CUDA or Triton through execution-guided repair.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eEN Key Points:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003earXiv:2608.25061v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels\u003c/li\u003e\n\u003cli\u003eExisting LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operators untested\u003c/li\u003e\n\u003cli\u003eWe introduce DataKernelBench, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimize either the core tensor-bounded sni…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-cslg-b_introsearch\"\u003e\n  ArXiv cs.LG (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-cslg-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24904\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDynamic Influence-Weighted Distillation for Single-IMU Activity Recognition\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [Translation Pending] - arXiv:2608.24904v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Inertial sensors at multiple body locations can improve activity recognition, but requiring every sensor at inference increases the deployment burden.\u003c/li\u003e\n\u003cli\u003eWe study whether four synchronized IMUs available during training can improve a student that uses only the right-arm IMU during fitting and inference.\u003c/li\u003e\n\u003cli\u003eA frozen four-IMU teacher provides logit and feature targets.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24904v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Inertial sensors at multiple body locations can improve activity recognition, but requiring every sensor at inference increases the deployment burden\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe study whether four synchronized IMUs available during training can improve a student that uses only the right-arm IMU during fitting and inference\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eA frozen four-IMU teacher provides logit and feature targets\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24936\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eGreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eRelease Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [Pending Translation] - arXiv:2608.24936v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: We present GreenLeaf Law Embed Tiny, a 0.6B parameter embedding model for legal domain retrieval.\u003c/li\u003e\n\u003cli\u003eGreenLeaf-Tiny achieves 75.11% on the Massive Legal Embedding Benchmark (MLEB) and 64.38% on MTEB(Law, v1),demonstrating competitive performance among models under 1B parameters.\u003c/li\u003e\n\u003cli\u003eOur approach combines a two-stage training pipeline that first distills knowledge from a larger teacher model into a compact student architecture, then applies domain-specific fine-tuning with hard negative mining; a carefully curated dataset of 3.4 million query-passage pairs, including 150,000 human-curated samples across diverse legal jurisdictions; and an efficient inference architecture supporting multiple quantization levels (BF16, INT8, binary) enabling deployment in resource-constrained environments.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24936v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: We present GreenLeaf Law Embed Tiny, a 0.6B parameter embedding model for legal domain retrieval\u003c/li\u003e\n\u003cli\u003eGreenLeaf-Tiny achieves 75.11% on the Massive Legal Embedding Benchmark (MLEB) and 64.38% on MTEB(Law, v1),demonstrating competitive performance among models un…\u003c/li\u003e\n\u003cli\u003eOur approach combines a two-stage training pipeline that first distills knowledge from a larger teacher model into a compact student architecture, then applies…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24937\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eMulti-Modal Anomaly Detection: A Survey\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eRelease Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [Pending Translation] - arXiv:2608.24937v1 Announce Type: new.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: Multi-Modal Anomaly Detection (MMAD) detects rare abnormal events from heterogeneous data sources and is increasingly used in safety- and reliability-critical applications such as industrial inspection and cybersecurity.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eYet the literature is fragmented across domains and modality combinations, and existing surveys usually group methods by architecture rather than by how abnormality is defined and separated in multi-modal settings.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe survey MMAD from an assumption-driven perspective.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24937v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Multi-Modal Anomaly Detection (MMAD) detects rare abnormal events from heterogeneous data sources and is increasingly used in safety- and reliability-\u0026hellip;\u003c/li\u003e\n\u003cli\u003eYet the literature is fragmented across domains and modality combinations, and existing surveys usually group methods by architecture rather than by how abnorma\u0026hellip;\u003c/li\u003e\n\u003cli\u003eWe survey MMAD from an assumption-driven perspective\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24938\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [Translation Pending] - arXiv:2608.24938v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation.\u003c/li\u003e\n\u003cli\u003eYet low-latency MoE serving is increasingly challenging, because it spans two inference phases with fundamentally different bottlenecks: prefill is dominated by token-wise expert computation, whereas decode is constrained by memory traffic from the batch-wise activated expert set.\u003c/li\u003e\n\u003cli\u003eHowever, existing training-free acceleration methods optimize only a single resource proxy, either the experts each token executes or the experts a batch activates, and either discard the excluded experts\u0026rsquo; contribution or leave it only implicitly approximated.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24938v1 Announce Type: new\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eYet low-latency MoE serving is increasingly challenging, because it spans two inference phases with fundamentally different bottlenecks: prefill is dominated by…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eHowever, existing training-free acceleration methods optimize only a single resource proxy, either the experts each token executes or the experts a batch activa…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24940\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eWhen Does Frequency Decomposition Benefit Physics-Informed Neural Networks? A Preliminary Ablation Study\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: arXiv:2608.24940v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Partial differential equations (PDEs) often have high-frequency and multi-scale features that neural networks struggle to approximate.\u003c/li\u003e\n\u003cli\u003ePhysics-Informed Neural Networks (PINNs) build the governing equations directly into training, but suffer from spectral bias: they learn low-frequency components faster than high-frequency ones.\u003c/li\u003e\n\u003cli\u003eTechniques such as Fourier feature embeddings and sinusoidal activations address this, but most studies assume they help across the board without checking which spectral regimes actually benefit.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eKey Points (EN):\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24940v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Partial differential equations (PDEs) often have high-frequency and multi-scale features that neural networks struggle to approximate\u003c/li\u003e\n\u003cli\u003ePhysics-Informed Neural Networks (PINNs) build the governing equations directly into training, but suffer from spectral bias: they learn low-frequency component…\u003c/li\u003e\n\u003cli\u003eTechniques such as Fourier feature embeddings and sinusoidal activations address this, but most studies assume they help across the board without checking which…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24945\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eFAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: [To be translated] - arXiv:2608.24945v1 Announce Type: new.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eAbstract: Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resource requirements of LLMs hinder the deployment on resource-constrained devices.\u003c/li\u003e\n\u003cli\u003eAlthough model quantization stands out as an effective approach, conventional quantization approaches typically incur severe performance degradation due to uniform bit-width or simple heuristic sensitivity evaluation.\u003c/li\u003e\n\u003cli\u003eIn this paper, we propose a novel Fisher information-based Adaptive Mixed Precision Weight Quantization approach, i.e., FAMPWQ, which performs layer-adaptive weight quantization for effective LLM inference on commodity GPUs.\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24945v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resource requirements of…\u003c/li\u003e\n\u003cli\u003eAlthough model quantization stands out as an effective approach, conventional quantization approaches typically incur severe performance degradation due to unif…\u003c/li\u003e\n\u003cli\u003eIn this paper, we propose a novel Fisher information-based Adaptive Mixed Precision Weight Quantization approach, i.e., FAMPWQ, which performs layer-adaptive we…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24946\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eMacroAgent: Regularity-Aware Macro Legalization with LLM-Agent-Designed Contour Algorithms\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [To be translated] - arXiv:2608.24946v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Macros constitute a large part of the core area in modern very large-scale integration (VLSI) designs.\u003c/li\u003e\n\u003cli\u003eMoreover, macro positions have a significant impact on the final quality of result (QoR), and macro legalization is typically the final step in determining the macro positions.\u003c/li\u003e\n\u003cli\u003eHowever, existing approaches related to macro legalization either lack robustness or incur substantial computational costs or neglect the regularity between macros.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003earXiv:2608.24946v1 Announce Type: new\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eAbstract: Macros constitute a large part of the core area in modern very large-scale integration (VLSI) designs\u003c/li\u003e\n\u003cli\u003eMoreover, macro positions have a significant impact on the final quality of result (QoR), and macro legalization is typically the final step in determining the…\u003c/li\u003e\n\u003cli\u003eHowever, existing approaches related to macro legalization either lack robustness or incur substantial computational costs or neglect the regularity between mac…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24947\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eCAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: [To be translated] - arXiv:2608.24947v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: End-to-end training of multimodal neural networks often exhibits unstable neural dynamics characterized by three coupled failure modes that degrade learning: (i) modality imbalance, where one branch dominates gradient-based optimization; (ii) unstable gating, where noisy confidence cues induce erratic modality selection; and (iii) fusion interference, where modality-specific gradients conflict at the shared fusion layer.\u003c/li\u003e\n\u003cli\u003eWe propose CAT-GS (Calibrated, Adaptive, Thresholded Gating with Fusion Surgery), a neural dynamics-based optimization controller for intelligent computing applications.\u003c/li\u003e\n\u003cli\u003eCAT-GS operates during backpropagation without modifying model architectures, fusion modules, or task losses.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24947v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: End-to-end training of multimodal neural networks often exhibits unstable neural dynamics characterized by three coupled failure modes that degrade le…\u003c/li\u003e\n\u003cli\u003eWe propose CAT-GS (Calibrated, Adaptive, Thresholded Gating with Fusion Surgery), a neural dynamics-based optimization controller for intelligent computing appl…\u003c/li\u003e\n\u003cli\u003eCAT-GS operates during backpropagation without modifying model architectures, fusion modules, or task losses\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24949\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDemystifying Reinforcement Learning Post-Training of Language Models\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary: [Translation pending] - arXiv:2608.24949v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Reinforcement learning (RL) post-training has emerged as a powerful framework for enhancing the capabilities of large language models (LLMs), enabling impressive reasoning, math, and coding capabilities.\u003c/li\u003e\n\u003cli\u003eYet for many researchers and practitioners, the principles behind classical RL remain a \u0026ldquo;black box\u0026rdquo;.\u003c/li\u003e\n\u003cli\u003eIn this work, we deconstruct the RL post-training algorithm, investigating each step to clarify what is actually happening beneath the surface.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24949v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Reinforcement learning (RL) post-training has emerged as a powerful framework for enhancing the capabilities of large language models (LLMs), enabling…\u003c/li\u003e\n\u003cli\u003eYet for many researchers and practitioners, the principles behind classical RL remain a \u0026ldquo;black box\u0026rdquo;\u003c/li\u003e\n\u003cli\u003eIn this work, we deconstruct the RL post-training algorithm, investigating each step to clarify what is actually happening beneath the surface\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2608.24954\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eAFDBench: A Reasoning-First AI Scientist for NationalWeather Service Forecast Discussions\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-08-27 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary: [Translation pending] - arXiv:2608.24954v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Large language models (LLMs) hallucinate numerical values when generating high-stakes meteorological text, posing risks for weather communication.\u003c/li\u003e\n\u003cli\u003eWe present AFDBench, an AI meteorologist that generates professional Area Forecast Discussions (AFDs) by reasoning through structured AI weather forecast data from Google\u0026rsquo;s WeatherNext 2.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe introduce AFDBench, the first benchmark for evaluating generative meteorological reasoning, comprising 7,732 expert written discussions from 13 National Weather Service (NWS) offices paired with real AI weather forecast inputs, and three complementary metrics: Met-Align (numerical accuracy), Style-Align (professional dialect adherence), and Input-Grounding (fidelity to source weather data).\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2608.24954v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Large language models (LLMs) hallucinate numerical values when generating high-stakes meteorological text, posing risks for weather communication\u003c/li\u003e\n\u003cli\u003eWe present AFDBench, an AI meteorologist that generates professional Area Forecast Discussions (AFDs) by reasoning through structured AI weather forecast data f…\u003c/li\u003e\n\u003cli\u003eWe introduce AFDBench, the first benchmark for evaluating generative meteorological reasoning, comprising 7,732 expert written discussions from 13 National Weat…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n",
  "wordCount": 6937,
  "readingTime": 33,
  "tableOfContents": "\u003cnav id=\"TableOfContents\"\u003e\n  \u003cul\u003e\n    \u003cli\u003e\u003ca href=\"#-in-depth-guide-to-this-issues-watch-list\"\u003e📖 In-depth Guide to This Issue\u0026rsquo;s Watch List\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-ai-hot-topics-on-x\"\u003e🌐 AI Hot Topics on X\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#topic-1-cursor-launches-scratch-to-deploy-web-app-builder\"\u003eTopic 1: Cursor Launches Scratch-to-Deploy Web App Builder\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-2-tech-giants-urge-urgent-ai-cyber-defense-action\"\u003eTopic 2: Tech Giants Urge Urgent AI Cyber Defense Action\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-3-google-launches-gemini-35-transcribe-for-precise-speech-to-text\"\u003eTopic 3: Google Launches Gemini 3.5 Transcribe for Precise Speech-to-Text\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-4-google-launches-gemini-omni-11-flash-for-advanced-video-creation\"\u003eTopic 4: Google Launches Gemini Omni 1.1 Flash for Advanced Video Creation\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-5-salesforce-beats-earnings-expectations-with-claude-ai-integration\"\u003eTopic 5: Salesforce Beats Earnings Expectations with Claude AI Integration\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-6-nvidia-acquires-hugging-face-for-129-billion-in-major-ai-deal\"\u003eTopic 6: Nvidia Acquires Hugging Face for $12.9 Billion in Major AI Deal\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-7-neil-movva-breaks-down-ai-inference-economics-on-invest-like-the-best\"\u003eTopic 7: Neil Movva Breaks Down AI Inference Economics on Invest Like the Best\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-8-tesla-adds-79-model-ys-to-texas-robotaxi-fleet-in-one-day\"\u003eTopic 8: Tesla Adds 79 Model Ys to Texas Robotaxi Fleet in One Day\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-influencer-insights\"\u003e💡 Influencer Insights\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-appendix-todays-watch-list-update-source-list\"\u003e📚 Appendix: Today\u0026rsquo;s Watch List Update Source List\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#openai-blog-a_full\"\u003eOpenAI Blog (A_full)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#google-deepmind-blog-a_full\"\u003eGoogle DeepMind Blog (A_full)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-csai-b_introsearch\"\u003eArXiv cs.AI (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-cscl-b_introsearch\"\u003eArXiv cs.CL (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-cslg-b_introsearch\"\u003eArXiv cs.LG (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n  \u003c/ul\u003e\n\u003c/nav\u003e",
  "isDraft": false
}
