{
  "title": "2026-07-15 AI Daily | OpenAI Pushes Codex as New ChatGPT, AI Competition Now Focuses on Output Per Dollar",
  "url": "https://miaok.ong/en/ai-daily/ai-daily-2026-07-15/",
  "date": "2026-07-15T07:00:00+08:00",
  "lastmod": "2026-07-15T07:00:00+08:00",
  "type": "ai-daily",
  "kind": "page",
  "language": "en",
  "description": "OpenAI is repositioning Codex as a new form of ChatGPT, shifting the focus from chat to deliverable workflows; enterprises are beginning to measure AI investments by output per dollar rather than token price. Meanwhile, custom chips and single-card models continue to lower deployment thresholds, with computing power and cost remaining at the core of the next phase of competition.",
  "keywords": null,
  "tags": [],
  "categories": [],
  "author": "Mark (Miao) Kong",
  "image": "https://miaok.ong/images/avatar.jpg",
  "content": "\u003ch1 id=\"2026-07-15-ai-daily--openai-repositions-codex-as-the-new-chatgpt-ai-competition-begins-to-focus-on-output-per-dollar\"\u003e\n  2026-07-15 AI Daily | OpenAI Repositions Codex as the New ChatGPT, AI Competition Begins to Focus on Output Per Dollar\n  \u003ca class=\"heading-link\" href=\"#2026-07-15-ai-daily--openai-repositions-codex-as-the-new-chatgpt-ai-competition-begins-to-focus-on-output-per-dollar\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h1\u003e\n\u003cblockquote\u003e\n\u003cp\u003eOpenAI is repositioning Codex as a new form of ChatGPT, shifting the focus from chat to deliverable workflows. Enterprises are beginning to measure AI investment by output per dollar rather than token price. Meanwhile, custom chips and single-card models continue to lower the deployment threshold, with computing power and cost remaining central to the next phase of competition.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch2 id=\"-deep-dive-this-issues-watch-list\"\u003e\n  📖 Deep Dive: This Issue\u0026rsquo;s Watch List\n  \u003ca class=\"heading-link\" href=\"#-deep-dive-this-issues-watch-list\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003eThe most noteworthy trend today is \u0026ldquo;AI\u0026rsquo;s shift from a chat interface to a workflow entry point.\u0026rdquo; Discussions around the OpenAI Super App and Codex/ChatGPT resonate with \u0026ldquo;investment management in the agentic era\u0026rdquo;: the real metric is no longer the price per token, but the verifiable tasks, time saved, and scalable processes generated per dollar.\u003c/p\u003e\n\u003cp\u003eEngineering teams are advised to focus on the articles covering coding agents and inference efficiency: topics like how much context a Coding Agent needs, KV-Cache compression, local MoE inference, and the Index 1.9B small model all point to one thing—making agents controllable, low-cost, and deployable systems.\u003c/p\u003e\n\u003cp\u003eThe third theme is trustworthy AI. Concepts like Ground Truth not being objective truth, AuditWeave\u0026rsquo;s evidence layer, distillation detection, and clinical time-series benchmarks all serve as reminders: beyond model capabilities, data provenance, audit trails, and evaluation boundaries are becoming hard requirements for the next phase of implementation.\u003c/p\u003e\n\u003ch2 id=\"-ai-hot-topics-on-x\"\u003e\n  🌐 AI Hot Topics on X\n  \u003ca class=\"heading-link\" href=\"#-ai-hot-topics-on-x\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003ch3 id=\"topic-1-openais-codex-and-chatgpt-work-hit-8-million-users-with-free-resets\"\u003e\n  Topic 1: OpenAI\u0026rsquo;s Codex and ChatGPT Work Hit 8 Million Users with Free Resets\n  \u003ca class=\"heading-link\" href=\"#topic-1-openais-codex-and-chatgpt-work-hit-8-million-users-with-free-resets\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending: 1 day ago, Related Posts: 13,000\u003c/li\u003e\n\u003cli\u003eWhat happened: OpenAI\u0026rsquo;s Codex and ChatGPT Work have reportedly reached 8 million users, lowering the barrier to entry with features like \u0026ldquo;free resets.\u0026rdquo;\u003c/li\u003e\n\u003cli\u003eWhy it matters: This indicates the rapid adoption of AI tools for programming and office scenarios. Free or low-cost strategies could further drive developers and enterprise users to migrate to AI-assisted workflows.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: Discussions on X focus on whether free tiers will change the payment models for AI tools, whether OpenAI is using this to expand its ecosystem advantage, and whether the claim of \u0026ldquo;free access to 100+ advanced models\u0026rdquo; is reliable or subject to limitations and marketing hype.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-2-openais-gpt-56-sol-tops-benchmarks-amid-efficiency-fixes-and-bug-reports\"\u003e\n  Topic 2: OpenAI\u0026rsquo;s GPT-5.6 Sol Tops Benchmarks Amid Efficiency Fixes and Bug Reports\n  \u003ca class=\"heading-link\" href=\"#topic-2-openais-gpt-56-sol-tops-benchmarks-amid-efficiency-fixes-and-bug-reports\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending: 2 days ago, Related Posts: 30,000\u003c/li\u003e\n\u003cli\u003eWhat happened: OpenAI\u0026rsquo;s GPT-5.6 Sol has reportedly taken the lead in several benchmarks, accompanied by progress in efficiency optimization and some user-reported bugs.\u003c/li\u003e\n\u003cli\u003eWhy it matters: This shows that competition among frontier large models continues to revolve around performance, inference efficiency, and stability. If the results are true, it could influence decisions by enterprises and developers regarding model selection, cost control, and reliability.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: Discussions on X are centered on whether benchmark scores represent real-world capabilities, whether efficiency fixes can reduce usage costs, and if current bugs will affect actual deployment. Some users are optimistic about its leading performance, while others question the transparency and stability of the tests.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-3-samsung-to-produce-custom-ai-chips-for-anthropic-report-says\"\u003e\n  Topic 3: Samsung to Produce Custom AI Chips for Anthropic, Report Says\n  \u003ca class=\"heading-link\" href=\"#topic-3-samsung-to-produce-custom-ai-chips-for-anthropic-report-says\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending: 14 hours ago, Related Posts: 753\u003c/li\u003e\n\u003cli\u003eWhat happened: Samsung will reportedly produce custom AI chips for Anthropic, a collaboration that points toward large model companies developing their own computing power supply chains.\u003c/li\u003e\n\u003cli\u003eWhy it matters: This indicates that leading AI companies are accelerating their efforts to reduce dependence on NVIDIA\u0026rsquo;s general-purpose GPUs, making custom ASICs, advanced manufacturing processes, HBM, and packaging capacity the new core of competition in AI infrastructure.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: Discussions on X focus on whether Samsung can close the gap with TSMC and SK Hynix in the AI chip supply chain, and whether companies like Anthropic and Amazon are building an \u0026ldquo;NVIDIA alternative.\u0026rdquo; The debate is whether custom chips can truly replace GPUs in the short term or will merely supplement specific inference and cloud scenarios.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-4-tencent-releases-quantized-hy3-ai-model-for-single-gpus\"\u003e\n  Topic 4: Tencent Releases Quantized Hy3 AI Model for Single GPUs\n  \u003ca class=\"heading-link\" href=\"#topic-4-tencent-releases-quantized-hy3-ai-model-for-single-gpus\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending: 5 hours ago, Related Posts: 186\u003c/li\u003e\n\u003cli\u003eWhat happened: Tencent has released a quantized version of its Hy3 AI model, designed to run on a single GPU.\u003c/li\u003e\n\u003cli\u003eWhy it matters: This shows that large model deployment is moving further towards low-cost, localized, and edge computing scenarios, which helps lower the hardware barrier for enterprises and developers to use AI models.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: Discussions on X primarily focus on the extent of performance loss after model quantization, its competitiveness against similar open-source models, and the practical value of single-GPU deployment for small to medium-sized developers and local AI applications.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch4 id=\"summary-of-ai-public-opinion-on-x-today\"\u003e\n  Summary of AI Public Opinion on X Today\n  \u003ca class=\"heading-link\" href=\"#summary-of-ai-public-opinion-on-x-today\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h4\u003e\n\u003cp\u003eThe main narrative today is that AI is shifting from a \u0026ldquo;competition of parameters and leaderboards\u0026rdquo; to a \u0026ldquo;competition of usability, cost, and implementation barriers.\u0026rdquo; Whether it\u0026rsquo;s OpenAI expanding its user base in programming and office work through free resets, or Tencent promoting quantized models that can run on a single card, these moves are seen as accelerating the popularization of AI. The general consensus is that low-cost usage, improved inference efficiency, and lower hardware barriers will continue to drive developers and enterprises to migrate to AI workflows. This will also prompt large model manufacturers to compete for ecosystem and computing power supply chains. The main points of disagreement are twofold: first, whether benchmarks and claims of \u0026ldquo;leading\u0026rdquo; performance can truly represent capabilities in real-world scenarios; second, whether custom chips and quantized models will genuinely change the industry landscape or will remain suitable only as supplementary solutions. The potential risks are that free and low-cost strategies may come with functional limitations or marketing exaggerations. If cutting-edge models lack stability, it will affect actual deployment. Furthermore, the restructuring of the computing power and chip supply chains could create new dependencies and competitive barriers.\u003c/p\u003e\n\u003ch2 id=\"-influencer-insights\"\u003e\n  💡 Influencer Insights\n  \u003ca class=\"heading-link\" href=\"#-influencer-insights\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003eOkay, as a senior AI industry analyst, I have carefully reviewed the tweets from various AI influencers over the past 24 hours. Here is a data-driven daily insight report.\u003c/p\u003e\n\u003chr\u003e\n\u003ch3 id=\"ai-industry-daily-insights-based-on-analysis-of-x-influencer-tweets\"\u003e\n  AI Industry Daily Insights (Based on Analysis of X Influencer Tweets)\n  \u003ca class=\"heading-link\" href=\"#ai-industry-daily-insights-based-on-analysis-of-x-influencer-tweets\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003ch4 id=\"1-todays-core-technology-and-product-hotspots\"\u003e\n  1. Today\u0026rsquo;s Core Technology and Product Hotspots\n  \u003ca class=\"heading-link\" href=\"#1-todays-core-technology-and-product-hotspots\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h4\u003e\n\u003cp\u003eToday, influencers\u0026rsquo; attention is primarily focused on \u003cstrong\u003ethe evolution of front-end model capabilities, the expansion of AI Agent product forms, and observations of emerging community phenomena\u003c/strong\u003e.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eTencent\u0026rsquo;s Hunyuan Hy3 Model Attracts High Attention:\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eBoth @ruanyf and @Pluvio9yte highlighted Tencent\u0026rsquo;s newly released flagship model, Hy3. @ruanyf pointed out that it has only \u003ccode\u003e295B\u003c/code\u003e parameters, far smaller than the industry\u0026rsquo;s \u003ccode\u003eGLM 5.2 (744B)\u003c/code\u003e. It focuses on \u0026ldquo;high speed and low cost,\u0026rdquo; with highly competitive API pricing (input \u003ccode\u003e$0.15/百万token\u003c/code\u003e). Its \u003cstrong\u003eperformance reaches or even surpasses that of GLM 5.1\u003c/strong\u003e, making it suitable as a primary model for daily use.\u003c/li\u003e\n\u003cli\u003e@Pluvio9yte provided a more in-depth use case, emphasizing Hy3\u0026rsquo;s excellent comprehensive abilities in \u003cstrong\u003e\u0026ldquo;engineering implementation, aesthetic judgment, and content curation.\u0026rdquo;\u003c/strong\u003e A typical example is generating a complete, beautiful, and ready-to-launch \u0026ldquo;League of Legends\u0026rdquo; character showcase website in just 8 minutes with a two-sentence prompt, a task that would have previously required days of manual development.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eThe Continuous Evolution and Deep Integration of AI Agent Product Forms:\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eGrowth and Evolution of Codex/Work\u003c/strong\u003e: @dotey retweeted that the combined user base of Codex and Work has exceeded \u003ccode\u003e800万\u003c/code\u003e and continues to reset usage quotas, showing rapid growth momentum. He also explained in detail the \u003cstrong\u003edifferentiation and integration of Chat, Work, and Codex\u003c/strong\u003e: Chat is for conversation, Work is an agent that delivers finished products across applications, and Codex is an agent focused on code repositories. The three share parts of the underlying Agent framework but have different purposes.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eClaude\u0026rsquo;s Strategy Adjustment\u003c/strong\u003e: Anthropic announced it would again extend access to its high-end model, \u003cstrong\u003eClaude Fable 5, until July 19th\u003c/strong\u003e. @dotey commented on this as a \u0026ldquo;child\u0026rsquo;s play decision,\u0026rdquo; which reflects Anthropic\u0026rsquo;s competitive posture and strategic wavering in its efforts to retain users under the pressure of OpenAI\u0026rsquo;s strong product offensive.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eApple vs. OpenAI\u003c/strong\u003e: @dotey reported the news that Apple has formally sued OpenAI and its former employees for stealing trade secrets, alleging they were used to develop AI hardware. This lawsuit adds a new dimension of tension to the top-tier talent war and the battle for the AI hardware entry point among tech giants.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eThe Rise of an Emerging Community Phenomenon: Xiaohongshu REDSkill\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e@ruanyf keenly observed a unique cross-disciplinary phenomenon: the lifestyle platform Xiaohongshu (Little Red Book) is starting to build a \u003cstrong\u003eREDSkill community\u003c/strong\u003e, allowing users to upload and share AI Skill files. He commented that this is equivalent to combining a social media platform with a \u0026ldquo;Skill Hub,\u0026rdquo; an approach \u003cstrong\u003eunprecedented globally\u003c/strong\u003e. It might be Xiaohongshu\u0026rsquo;s way of paving the path for the platform\u0026rsquo;s AI transformation and increasing its technical content. For developers, it represents a massive new distribution channel to reach a vast user base.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch4 id=\"2-noteworthy-unique-perspectives-and-industry-foresight\"\u003e\n  2. Noteworthy Unique Perspectives and Industry Foresight\n  \u003ca class=\"heading-link\" href=\"#2-noteworthy-unique-perspectives-and-industry-foresight\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h4\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eThe Myth of \u0026ldquo;Saving Tokens\u0026rdquo; (by @dotey)\u003c/strong\u003e:\nBaoyu conducted an in-depth analysis of the popular \u0026ldquo;telegraph-style Skills\u0026rdquo; (like the Caveman project, which claims to save 65% of tokens). He cited test results from JetBrains, pointing out that in actual programming tasks, \u003cstrong\u003eoutput tokens are only reduced by 8.5%\u003c/strong\u003e. He argues that this type of optimization is targeted at \u0026ldquo;chat scenarios,\u0026rdquo; whereas the real cost of an Agent lies in tool calls and system prompts. With the ongoing trend of falling API prices, \u003cstrong\u003eit\u0026rsquo;s better to optimize context management and reduce rework on paths than to obsess over \u0026ldquo;talking like a caveman.\u0026rdquo;\u003c/strong\u003e This pours cold water on the development of \u0026ldquo;money-saving\u0026rdquo; Skills, highlighting the primary and secondary contradictions in optimization.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eReshaping of Individual Roles and Capabilities (by @dotey, @vista8)\u003c/strong\u003e:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e@dotey shared an internal observation from Anthropic: The \u0026ldquo;scaffolding\u0026rdquo; for Agents is becoming thinner, with the focus shifting from \u0026ldquo;controlling every step\u0026rdquo; to \u0026ldquo;designing collaboration between Agents.\u0026rdquo; He also emphasized that \u003cstrong\u003eAgents can amplify individual capabilities but won\u0026rsquo;t automatically solve team coordination problems\u003c/strong\u003e, and products might expand chaotically due to rapid individual trial-and-error.\u003c/li\u003e\n\u003cli\u003e@vista8 referenced a 1979 IBM slide, stating, \u0026ldquo;A computer can never be held accountable, therefore a computer must never make a management decision.\u0026rdquo; The same logic applies to AI; \u003cstrong\u003ewhat will be truly scarce in the future is the ability to make decisions with incomplete information and take responsibility for the consequences\u003c/strong\u003e. His observation of the phenomenon where \u0026ldquo;bosses want to distill employee experience into a Skill, but employees resist\u0026rdquo; also profoundly reveals the new conflicts over knowledge and value attribution within companies in the AI era.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eThe Potential of \u0026ldquo;On-Device Models\u0026rdquo; and \u0026ldquo;Model Cartridges\u0026rdquo; (by @zhixianio)\u003c/strong\u003e:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e@zhixianio, through a hands-on review of \u003cstrong\u003eGemma 4 12B Coder\u003c/strong\u003e, reached a counter-intuitive conclusion: despite community hype, the \u0026ldquo;ceiling\u0026rdquo; for a 12B parameter model is apparent, making it unsuitable for complex programs that require long, single-pass generation. This reminds us that \u003cstrong\u003ethe capability boundaries of small models remain clear; fine-tuning improves efficiency but doesn\u0026rsquo;t raise the upper limit\u003c/strong\u003e.\u003c/li\u003e\n\u003cli\u003eHe agreed with @geekbb\u0026rsquo;s vision of a \u003cstrong\u003e\u0026ldquo;Model-Pak,\u0026rdquo;\u003c/strong\u003e suggesting that future on-device models could be distributed and used like plug-in cartridges. This aligns with the on-device AI trend Google is heavily promoting (e.g., Gemma QAT models) and hints at a new form of offline AI applications.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch4 id=\"3-recommended-tools-and-resources\"\u003e\n  3. Recommended Tools and Resources\n  \u003ca class=\"heading-link\" href=\"#3-recommended-tools-and-resources\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h4\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eReference for Local Code Generation Model Evaluation\u003c/strong\u003e: @zhixianio provided a detailed comparative review of \u003cstrong\u003eGemma 4 12B Coder vs. Qwen 3.6-35B-A3B MoE\u003c/strong\u003e. The conclusion is that the 35B MoE model is currently the \u0026ldquo;sweet spot\u0026rdquo; for local code generation and can serve as an important reference for developers choosing a local model.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eAI Video Editing Tools (Skills)\u003c/strong\u003e:\n\u003cul\u003e\n\u003cli\u003e@Pluvio9yte and @dotey both noted \u003cstrong\u003eChatCut\u003c/strong\u003e, an AI video editing Skill that integrates with Codex/Claude Code. It can automatically edit videos based on transcription, remove filler words, and more.\u003c/li\u003e\n\u003cli\u003e@dotey released his self-developed \u003cstrong\u003eBaoCut\u003c/strong\u003e, a subtitle transcription, translation, and editing Skill (Mac only). He specifically highlighted its approach to solving the \u0026ldquo;post-generation re-editing\u0026rdquo; problem for Agents by combining a CLI with a GUI to provide an Agent-friendly interface.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eAI Video Replication and Integrated Creation\u003c/strong\u003e: @Pluvio9yte shared the powerful capabilities of his \u0026ldquo;Video Replication Skill\u0026rdquo; enhanced by the Sol model and recommended using \u003cstrong\u003eTencent Hy3\u003c/strong\u003e for complex creative work that requires a balance of engineering, aesthetics, and content planning.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eOpen Source Projects and Learning Resources\u003c/strong\u003e:\n\u003cul\u003e\n\u003cli\u003e@vista8 recommended a \u003cstrong\u003eModel \u0026ldquo;PK\u0026rdquo; Arena\u003c/strong\u003e developed with AI, which allows for one-click comparison of text and front-end output from multiple models.\u003c/li\u003e\n\u003cli\u003e@vista8 recommended \u003cstrong\u003efireworks-tech-graph\u003c/strong\u003e (8.5k stars), an open-source project for creating professional technical diagrams that has gained popularity through community promotion.\u003c/li\u003e\n\u003cli\u003e@vista8 also noted that \u003cstrong\u003ethe effectiveness of AI editing tools can be average\u003c/strong\u003e, and building your own workflow with Listenhub CLI + Remotion remains a reliable option for pursuing higher quality.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch2 id=\"-appendix-todays-watch-list-source-updates\"\u003e\n  📚 Appendix: Today\u0026rsquo;s Watch List Source Updates\n  \u003ca class=\"heading-link\" href=\"#-appendix-todays-watch-list-source-updates\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cblockquote\u003e\n\u003cp\u003eTimeframe: Last 3 days; covers 22 sources; 34 updates in total\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch3 id=\"stratechery-by-ben-thompson-a_full\"\u003e\n  Stratechery by Ben Thompson (A_full)\n  \u003ca class=\"heading-link\" href=\"#stratechery-by-ben-thompson-a_full\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003e\u003ca href=\"https://stratechery.com/2026/the-openai-super-app-chatgpt-codex-whither-chat/\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eThe OpenAI Super App, ChatGPT = Codex, Whither Chat\u003c/a\u003e\u003c/strong\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-14 18:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - OpenAI has refashioned Codex as the new ChatGPT; is the company abandoning the chat category they pioneered?\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003e$15\u003c/strong\u003e/month \u003cem\u003eor\u003c/em\u003e \u003cstrong\u003e$150\u003c/strong\u003e/year.\u003c/li\u003e\n\u003cli\u003eSubstantive analysis of the day’s news via three weekly emails or a podcast.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eStrategy Interviews\u003c/strong\u003e.\u003c/li\u003e\n\u003cli\u003eInterviews with leading public company CEOs, private company founders, and discussions with fellow analysts.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003eOpenAI has refashioned Codex as the new ChatGPT; is the company abandoning the chat category they pioneered\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"openai-blog-a_full\"\u003e\n  OpenAI Blog (A_full)\n  \u003ca class=\"heading-link\" href=\"#openai-blog-a_full\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://openai.com/index/managing-ai-investments-in-agentic-era\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eHow to manage AI investments in the agentic era\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-14 18:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary:\n\u003cul\u003e\n\u003cli\u003eOpenAI aims to make AI more accessible, capable, and affordable over time.\u003c/li\u003e\n\u003cli\u003eFrom GPT-4 to GPT-5.4, the price per million tokens has decreased by 97%.\u003c/li\u003e\n\u003cli\u003eHowever, token prices alone do not indicate whether AI is creating value.\u003c/li\u003e\n\u003cli\u003eLeaders should focus on useful work per dollar: tasks completed, time saved, improved decisions, and workflows ready to scale.\u003c/li\u003e\n\u003cli\u003eAs teams shift from chat to longer-running workflows, administrators need clearer visibility into demand, spending, and risks.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003eLearn how enterprises can manage AI investments in the agentic era by measuring useful work per dollar, improving efficiency, and scaling high-value workflows.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://openai.com/academy/codex-for-work/how-data-science-teams-use-codex\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eHow data science teams use ChatGPT Work\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-14 08:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary:\n\u003cul\u003e\n\u003cli\u003eLearn how data science teams use ChatGPT Work to transform problems, dashboards, and raw data into analytical assets ready for review.\u003c/li\u003e\n\u003cli\u003eWith ChatGPT Work, data science teams can more quickly turn disparate inputs into usable analytical assets.\u003c/li\u003e\n\u003cli\u003eStarting with dashboards, metric definitions, exports, experiment annotations, and business context, ChatGPT Work helps assemble first drafts of deliverables (including charts, explanations, source links, and audit questions) so teams can validate work and share with confidence.\u003c/li\u003e\n\u003cli\u003e\n\u003ch2 id=\"watch-the-on-demand-webinar\"\u003e\n  Watch the on-demand webinar.\n  \u003ca class=\"heading-link\" href=\"#watch-the-on-demand-webinar\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003c/li\u003e\n\u003cli\u003eNote: This webinar was recorded when these workflows were located in the former Codex application.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003eSee how data science teams can use ChatGPT Work to build root-cause briefs, impact readouts, KPI memos, scoped analyses, and dashboard specs from real work inpu…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://openai.com/academy/codex-for-work/how-sales-teams-use-codex\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eHow sales teams use ChatGPT Work\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-14 08:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary:\n\u003cul\u003e\n\u003cli\u003eLearn how sales teams use ChatGPT Work to create pipeline briefs, meeting prep packets, forecast reviews, customer plans, and real stalled-deal diagnoses\u0026hellip;\u003c/li\u003e\n\u003cli\u003eThis article from the OpenAI blog explains how sales teams use ChatGPT Work to shape the broader AI and infrastructure landscape.\u003c/li\u003e\n\u003cli\u003eFollowing how sales teams use ChatGPT Work, it also provides practical implications for founders, operators, and investors.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003eSee how sales teams can use ChatGPT Work to create pipeline briefs, meeting prep packets, forecast reviews, account plans, and stalled-deal diagnoses from real…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-csai-b_introsearch\"\u003e\n  ArXiv cs.AI (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-csai-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09664\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eFrom ML Predictions to Informed Diagnostic Assistance Using the Toulmin Model of Argumentation\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09664v1 Announce Type: new.\u003c/li\u003e\n\u003cli\u003eAbstract: To provide structured and interpretable evaluations, we decompose image-based diagnostics into components following the Toulmin model of argumentation.\u003c/li\u003e\n\u003cli\u003eThe model consists of claim, grounds, warrant, qualifier, rebuttal, and backing.\u003c/li\u003e\n\u003cli\u003eConsider claims generated by machine learning (ML) models for retinal diagnosis.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09664v1 Announce Type: new\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: To provide a structured and interpretable assessment, we decompose the image-based diagnosis into components following the Toulmin model of argumentat…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eThis model consists of a claim, grounds, warrant, qualifier, rebuttal, and backing\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eConsider a claim generated by a machine learning (ML) model for retinal diagnosis\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09665\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eFormat Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09665v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Prompt wrappers often differ only in formatting, yet they can change model scores enough to flip leaderboard conclusions.\u003c/li\u003e\n\u003cli\u003eWe study this variance under a token-controlled protocol and introduce two complementary metrics: the Format Sensitivity Index (FSI), the accuracy range induced by wrapper choice, and the Parsability Sensitivity Index (PSI), the corresponding range for answer parsability.\u003c/li\u003e\n\u003cli\u003eAcross 140,000 OpenRouter generations spanning 7 QA tasks, 5 wrapper families, and 4 instruct models from 7B to 72B parameters, we find that mean FSI varies by over 30x between models, driven largely by compliance failures.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09665v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Prompt wrappers often differ only in formatting, yet they can change model scores enough to flip leaderboard conclusions\u003c/li\u003e\n\u003cli\u003eWe study this variance under a token-controlled protocol and introduce two complementary metrics: the Format Sensitivity Index (FSI), the accuracy range induced…\u003c/li\u003e\n\u003cli\u003eAcross 140,000 OpenRouter generations spanning 7 QA tasks, 5 wrapper families, and 4 instruct models from 7B to 72B parameters, we find that mean FSI varies by…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09678\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eFaithful, Not Corrective: Message-Format Effects in Multi-Hop Agent Relays Are Tier-Dependent\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09678v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: When LLM agents hand off information to one another, does the message format matter?\u003c/li\u003e\n\u003cli\u003eTwo literatures disagree: format optimization work reports that structured messages can reduce cost without harming accuracy, while format constraint work finds that imposing structure degrades generation\u0026ndash;and neither measures what happens when messages traverse multiple hops, where copy-fidelity rather than one-shot generation dominates.\u003c/li\u003e\n\u003cli\u003eWe introduce a controlled relay testbed: a digest of twelve programmatically-generated atomic facts is serially re-encoded over six hops in five formats (free NL, precisely-instructed NL, JSON, triples, key-value), scored by a fixed strong grader against the programmatic ground truth, across two relay capability tiers, cognitive load conditions, and paired-fork error injection.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09678v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: When LLM agents hand off information to one another, does the message format matter\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eTwo literatures disagree: format-optimization work reports that structured messages cut cost without hurting accuracy, while format-restriction work finds that…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe introduce a controlled relay testbed: briefs of twelve programmatically generated atomic facts are re-encoded hop-by-hop in five formats (free NL, precision-…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09689\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eBoltzmann MapReduce: A Partition-Function Reduce for Forkable Sandboxes\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09689v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: To leading order under local asymptotic normality (LAN), the confidence density a worker emits over a chunk of size $n$ is a Gibbs\u0026ndash;Boltzmann measure $\\exp{-\\beta E(\\theta)}$, where the inverse temperature is the sample size $\\beta=n$.\u003c/li\u003e\n\u003cli\u003eThree consequences are exact in the Gaussian/linear case and first-order otherwise: disjoint chunks carry independent Boltzmann factors, so the MapReduce \\emph{reduce} is literally a partition function $Z=\\int\\prod_k h_k,d\\theta$ whose mode is precision-weighted (inverse-variance) pooling; frequentist consistency is the zero-temperature limit $T=1/n\\to0$.\u003c/li\u003e\n\u003cli\u003earXiv:2607.09689v1 Announce Type: new Abstract: To leading order under local asymptotic normality (LAN), the confidence density a worker emits over a chunk of size $n$ is a Gibbs\u0026ndash;Boltzmann measure\u0026hellip; Three consequences are exact in the Gaussian/linear case and first-order otherwise: disjoint chunks carry independent Boltzmann factors, so the MapReduce \\emph{\u0026hellip;}.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09689v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: To leading order under local asymptotic normality (LAN), the confidence density a worker emits over a chunk of size $n$ is a Gibbs\u0026ndash;Boltzmann measure…\u003c/li\u003e\n\u003cli\u003eThree consequences are exact in the Gaussian/linear case and first-order otherwise: disjoint chunks carry independent Boltzmann factors, so the MapReduce \\emph{…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09698\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eInterpreting Latent CoT Reasoning as Dynamical Systems\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09698v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Recent latent reasoning methods, such as CODI and COCONUT, face a fundamental interpretability problem: they maintain multiple superimposed candidate trajectories in hidden space at each step, unlike explicit CoT which follows a single, transparent reasoning trajectory.\u003c/li\u003e\n\u003cli\u003eExisting mechanistic approaches have shown compression, shortcuts, and superposition, but do not explain how reasoning evolves across latent steps.\u003c/li\u003e\n\u003cli\u003eTo address this gap, we model the sequence of latent tokens as a trajectory in representation space and apply dynamical systems analysis to characterize the evolution of reasoning.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09698v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Recent latent reasoning methods, such as CODI and COCONUT, face a fundamental interpretability problem: they maintain multiple superimposed candidate…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eExisting mechanistic methods show compression, shortcuts, and superposition without explaining how reasoning evolves across latent steps\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eTo address this gap, we model latent token sequences as trajectories in representation space and apply dynamical systems analysis to characterize the evolution…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09706\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eYUKTI: From Natural-Language Situations to Robust, Verifiable Decisions An Uncertainty-Typed Proposition IR, Assumption-Robust Pareto Frontiers, and a Regret Certificate\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:- arXiv:2607.09706v1 Announce Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Language models convert worded situations into numerical plans, and dominant pipelines (NL4Opt, OptiMUS, ORLM, OR-LLM-Agent) commit to a single objective and point-value coefficients, then solve once.\u003c/li\u003e\n\u003cli\u003eFor decisions that allocate real budget, effort, or clinical attention, this confidence is a failure mode: every objectified number is an assumption, and the optimal plan that presumes perfect correctness is fragile—a simulated calculation.\u003c/li\u003e\n\u003cli\u003eYUKTI changes the target of autoformulation.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09706v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Language models turn a worded situation into a numeric plan, and the dominant pipelines (NL4Opt, OptiMUS, ORLM, OR-LLM-Agent) commit to a single objec…\u003c/li\u003e\n\u003cli\u003eFor decisions that allocate real budget, effort, or clinical attention, that confidence is the failure mode: every objectified number is an assumption, and a pl…\u003c/li\u003e\n\u003cli\u003eYUKTI changes the target of autoformulation\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09708\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eGES-TSP: Graph Edge Sparsification for TSP\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:- arXiv:2607.09708v1 Announce Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Solving large-scale instances of the Traveling Salesman Problem (TSP) is computationally expensive.\u003c/li\u003e\n\u003cli\u003eResearchers often employ graph sparsification methods to improve computational efficiency.\u003c/li\u003e\n\u003cli\u003eTraditional sparsification methods typically rely on fixed heuristics and fail to fully exploit instance-specific structural information.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09708v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Solving large-scale instances of the Traveling Salesman Problem (TSP) exactly is computationally expensive\u003c/li\u003e\n\u003cli\u003eResearchers often employ graph sparsification methods to improve computational efficiency\u003c/li\u003e\n\u003cli\u003eTraditional sparsification methods typically rely on fixed heuristics and fail to fully exploit instance-specific structural information\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09709\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eThe Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game Generation\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003ePublication Time: 2026-07-14 12:00 Beijing Time\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: - arXiv:2607.09709v1 Announcement Type: New.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eAbstract: Post-training a code generator against a learned judge can optimize proxy features that raise the score without improving the artifact.\u003c/li\u003e\n\u003cli\u003eWe study the opposite signal: a deterministic, judge-free, ungameable filter \u0026ndash; whether a generated project launches cleanly under a headless engine (strict-lau…).\u003c/li\u003e\n\u003cli\u003eUnder this gate, rejection-sampling self-distillation compounds out-of-family generalization.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09713\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eClosed-Loop Control with Rule-Aligned Small Language Models and Multi-Agent Self-Correction\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09713v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: A key step toward autonomous industrial operation is the ability to create and reconfigure control policies from natural-language requirement specific….\u003c/li\u003e\n\u003cli\u003eIn this setting, policy generation by AI agents can be a credible path when paired with a plant-aware validator (e.g., a digital twin) that can check generated….\u003c/li\u003e\n\u003cli\u003eHowever, practical deployment is constrained by inference latency and compute footprint: large cloud-based models are often too slow, opaque, or data-sensitive….\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09714\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eFeedback-Coupled Memory Systems in Continuous Time\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09714v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: The Feedback-Coupled Memory Systems (FCMS) architecture formalizes closed-loop coordination through four abstract operators, two of which—the agent update operator $f_i$ and the environment update operator $\\Psi$—are not axiomatically defined in the original framework.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eTo address this, $f_i$ is defined by Mechanism-Based Intelligence (MBI), where agents update locally through a decentralized price mechanism and economic principles, while $\\Psi$ is defined by the Coupled Memory Graph Process (CMGP), a non-Markovian framework in which the environment is treated as a physical substrate that can coherently record and respond to trajectory history without external forcing.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eThe resulting continuous-time FCMS instantiation achieves Lyapunov global dissipativity governed by the computable threshold $4\\beta^2 \u0026lt; 2\\eta\\mu\\gamma^2$.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09714v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: The Feedback-Coupled Memory Systems (FCMS) architecture formalizes closed-loop coordination through four abstract operators, two of which - the agent…\u003c/li\u003e\n\u003cli\u003eTo address this, $f_i$ is defined by Mechanism-Based Intelligence (MBI), where agents update locally through a decentralized price mechanism and economic princi…\u003c/li\u003e\n\u003cli\u003eThe resulting continuous-time FCMS instantiation achieves Lyapunov global dissipativity governed by the computable threshold $4\\beta^2 \u0026lt; 2\\eta\\mu\\gamma^2$\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-cscl-b_introsearch\"\u003e\n  ArXiv cs.CL (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-cscl-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09880\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eCLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eRelease Time: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09880v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Clinical time series are crucial for patient monitoring, risk assessment, and clinical decision support.\u003c/li\u003e\n\u003cli\u003eHowever, they are often sparse, irregularly sampled, and asynchronous, making it difficult for models to identify the temporal evidence required for clinical Question Answering (QA).\u003c/li\u003e\n\u003cli\u003eExisting benchmarks primarily focus on regularly sampled time-series QA or medical QA over static data, and therefore rarely assess whether models can faithfully derive answers from irregularly timed observations.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09880v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Clinical time series are central to patient monitoring, risk assessment, and clinical decision support\u003c/li\u003e\n\u003cli\u003eHowever, they are often sparse, irregularly sampled, and asynchronous, making it difficult for models to identify the temporal evidence required for clinical Qu…\u003c/li\u003e\n\u003cli\u003eExisting benchmarks primarily focus on regularly sampled time-series QA or medical QA over static data, and therefore rarely assess whether models can faithfull…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09885\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eIndex SLM Technical Report\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eRelease Time: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09885v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: We introduce Index-1.9B, a series of open small language models developed by Bilibili.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eThe series includes four models: Index-1.9B-Base, a foundation model with 1.9 billion non-embedding parameters pre-trained on 2.8 trillion predominantly Chinese and English tokens; Index-1.9B-Pure, a control variant trained with the same recipe but with all instruction-like data strictly filtered from the corpus; Index-1.9B-Chat, which starts from the base model and undergoes supervised fine-tuning and direct preference optimization; and Index-1.9B-Character, which enhances the chat model with retrieval-augmented generation for few-shot role-playing customization.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003ePre-training employed a Warmup-Stable-Decay learning rate schedule, where the concentration of curated data was substantially increased during the decay phase, along with a Norm-Head output layer to stabilize training at high learning rates.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eEN Highlights:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09885v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: We present Index-1.9B, a series of open small language models developed at Bilibili\u003c/li\u003e\n\u003cli\u003eThe series comprises four models: Index-1.9B-Base, a foundation model with 1.9 billion non-embedding parameters pre-trained on 2.8 trillion predominantly Chines…\u003c/li\u003e\n\u003cli\u003ePre-training employs a Warmup-Stable-Decay learning-rate schedule in which the concentration of curated data is raised substantially during the decay phase, tog…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09908\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eRouteRec: Strict Evaluation of Recommender-Agent Selection and Aggregation\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eRelease Time: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09908v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Recommender systems increasingly face a choice among heterogeneous agents (collaborative filters, sequential models, content-based retrievers, and LLM-based rerankers), with no single agent being consistently the best.\u003c/li\u003e\n\u003cli\u003eWe study this choice as task-aware agent ranking under cost constraints using RouteRec, a framework that compares request-level hard selection with item-level learned aggregation of four traditional recommender agents and one LLM reranking agent.\u003c/li\u003e\n\u003cli\u003eOn MovieLens-1M, the full-quality oracle has substantial headroom (HR @ai_daily_20260510.md = 0.584), confirming that useful cross-agent signals exist.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09908v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Recommender systems increasingly face a choice among heterogeneous agents \u0026ndash; collaborative filters, sequential models, content-based retrievers, and L…\u003c/li\u003e\n\u003cli\u003eWe study this choice as task-aware agent ranking under cost constraints using RouteRec, a framework that compares request-level hard selection with item-level l…\u003c/li\u003e\n\u003cli\u003eOn MovieLens-1M, the full quality oracle has substantial headroom (HR @ai_daily_20260510.md = 0.584), confirming that useful cross-agent signal exists\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09921\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eGlobal Merger-Arbitrage Forecasting with Language Models\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eRelease Time: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09921v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: We propose a language model forecasting system for merger arbitrage, a specialized, high-stakes financial setting where the task is to predict the outcomes of announced merger and acquisition deals.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eUnlike previous LLM judgmental forecasting work, which focused on broad mixed-topic benchmarks and short contexts such as news snippets, we investigate a setting that requires long-context reasoning over hundreds of pages of technical documentation.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eOur system combines expert-guided context engineering with finetuning on hindsight-guided reasoning traces derived from historical deals.\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09921v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: We present a language-model forecasting system for merger arbitrage, a specialized high-stakes financial setting in which the task is to predict the o…\u003c/li\u003e\n\u003cli\u003eUnlike prior work on judgmental forecasting with LLMs, which has focused on broad mixed-topic benchmarks and short context such as news snippets, we study a set…\u003c/li\u003e\n\u003cli\u003eOur system combines expert-guided context engineering with finetuning on hindsight-guided reasoning traces derived from historical deals\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09932\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eFaithful by Design: Evaluating and Improving LLM-Generated Clinical Trial Summaries for Multi-Stakeholder Audiences\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09932v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Large language models are increasingly used to summarize clinical trial results for healthcare providers, patients, and payers, but their tendency towards hallucination poses significant risks in this high-stakes context.\u003c/li\u003e\n\u003cli\u003eThis study introduces a benchmark evaluation framework for measuring the faithfulness of LLM-generated clinical trial summaries across three stakeholder audiences.\u003c/li\u003e\n\u003cli\u003eThe framework consists of 200 stratified trials drawn from the ClinicalTrials.gov database aggregated analysis, evaluated using audience-specific prompt templates and a six-dimensional faithfulness annotation scheme.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09932v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Large language models are increasingly used to summarize clinical trial results for healthcare providers, patients, and payers, but their tendency to…\u003c/li\u003e\n\u003cli\u003eThis study introduces a benchmark evaluation framework for measuring the faithfulness of LLM-generated clinical trial summaries across three stakeholder audienc…\u003c/li\u003e\n\u003cli\u003eThe framework consists of 200 stratified trials drawn from the Aggregate Analysis of ClinicalTrials.gov database, evaluated using audience-specific prompt templ…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09957\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eWorkload-Driven Optimization for On-Device Real-Time Subtitle Translation\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09957v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: This report studies on-device English-to-Traditional-Chinese subtitle translation in Taiwan under short-input, short-output, batch-size-one inference, low-latency, and privacy constraints.\u003c/li\u003e\n\u003cli\u003eThese conditions limit the value of optimizations designed for long-context or high-throughput language model serving.\u003c/li\u003e\n\u003cli\u003eStarting with LMT-60-0.6B, preliminary analysis suggests that vocabulary projection becomes a more significant decoding-time cost after GGUF quantization reduces the relative cost of Transformer blocks.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003earXiv:2607.09957v1 Announce Type: new\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eAbstract: This report studies on-device English-to-Traditional-Chinese subtitle translation for Taiwan under short inputs, short outputs, batch-size-one inferen…\u003c/li\u003e\n\u003cli\u003eThese conditions limit the value of optimizations designed for long-context or high-throughput language-model serving\u003c/li\u003e\n\u003cli\u003eStarting from LMT-60-0.6B, preliminary profiling suggests that vocabulary projection becomes a more important decode-time cost after GGUF quantization reduces t…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09999\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eSilent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09999v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: We show that post-training quantization can silently alter how large language models reason, even when task accuracy is preserved.\u003c/li\u003e\n\u003cli\u003eUsing a six-category failure taxonomy validated by two independent human annotators (Cohen\u0026rsquo;s κ = 0.906), we classify 30,000 chain-of-thought outputs from five instruction-tuned LLMs (3B\u0026ndash;14B parameters) across three quantization precisions (FP32, FP16, NF4) and four reasoning benchmarks.\u003c/li\u003e\n\u003cli\u003eWe find that while accuracy is robust across precisions (maximum 3.1 pp drop), Hollow Convergence (correct answers reached through incomplete or unverifiable reasoning) shows significant size-dependent changes under NF4, with the two smallest models tested dropping sharply, while models 12B parameters and larger remain unchanged.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09999v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: We show that post-training quantization can silently alter how large language models reason even when task accuracy is preserved\u003c/li\u003e\n\u003cli\u003eUsing a six-category failure taxonomy validated by two independent human annotators (Cohen\u0026rsquo;s $\\kappa$ = 0.906), we classify 30,000 chain-of-thought outputs from…\u003c/li\u003e\n\u003cli\u003eWe find that while accuracy is robust across precisions (maximum 3.1 pp drop), Hollow Convergence (correct answers reached through incomplete or unverifiable re…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.10020\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eRobust, Scalable Detection of Text Containment in Large Web-Crawled Corpora\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.10020v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: We introduce FindMyText, an open-source Python package designed to efficiently assess whether a given text appears partially or entirely within a text corpus.\u003c/li\u003e\n\u003cli\u003eThe tool builds upon existing document fingerprinting techniques but extends them with a novel mechanism to explicitly capture sequences of matching fingerprints.\u003c/li\u003e\n\u003cli\u003eBy identifying such chains, the tool can more reliably detect near-verbatim copies of a given text, not just text similarity.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.10020v1 Announce Type: new\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: We present FindMyText, an open-source Python package designed to efficiently assess whether a given text appears, in part or in full, within a text co…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eThe tool builds on prior techniques for document fingerprinting, but extends them with a novel mechanism to explicitly capture sequences of matching fingerprint…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eBy identifying such chains, the tool can more reliably detect near-verbatim copies of a given text rather than mere textual similarities\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.10092\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eEfficiently Adapting Spoken Language Models for the Singaporean Context\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.10092v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Spoken language models (SLMs) unify speech perception and reasoning, but their adaptation to sensitive domains is underexplored, especially when the original training data is inaccessible and use cases require multilingual, spoken-query interactions.\u003c/li\u003e\n\u003cli\u003eWe apply an open-source SLM to the Singaporean Home Team context, covering five speech tasks across Singapore\u0026rsquo;s four official languages, combining LoRA fine-tuning, an alternative text QA dataset to prevent catastrophic forgetting, and a multi-task objective that adapts the CoBa re-weighting scheme to speech.\u003c/li\u003e\n\u003cli\u003eWe also build HTD-multilingual-QA, a multilingual QA dataset containing 504,853 samples in both text and speech formats.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.10092v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Spoken language models (SLMs) unify speech perception and reasoning, but adapting them to sensitive domains is underexplored, especially when the orig…\u003c/li\u003e\n\u003cli\u003eWe adapt an open-source SLM to the Singaporean Home Team context across five speech tasks in Singapore\u0026rsquo;s four official languages, combining LoRA fine-tuning, a…\u003c/li\u003e\n\u003cli\u003eWe also build HTD-multilingual-QA, a 504,853 sample multilingual QA dataset in text and spoken form\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.10114\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eCost of Reasoning in non-English Languages: A Case Study on Japanese\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.10114v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Reasoning Language Models (RLMs) achieve their strongest performance when reasoning in English, the language for which reasoning-oriented training data is most abundant.\u003c/li\u003e\n\u003cli\u003eHowever, the reasoning trajectory serves as a clue for model interpretability and safety, and is useful in practice for both model users and developers.\u003c/li\u003e\n\u003cli\u003eTherefore, it is desirable to develop models that can reason in a user\u0026rsquo;s chosen language while still maintaining strong reasoning performance.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.10114v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Reasoning Language Models (RLMs) achieve their strongest performance when they reason in English, the language for which reasoning-oriented training d…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eHowever, reasoning trace is a clue for model interpretability and safety, and useful in practice for both the model users and for model developers\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eThus, it is desirable to be able to develop a model that reasons in a language of the user\u0026rsquo;s choice, while still maintaining strong reasoning performance\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-cslg-b_introsearch\"\u003e\n  ArXiv cs.LG (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-cslg-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09666\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eKnowledge Graphs Meet Graph Neural Networks: A Comprehensive Survey\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09666v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Graph Neural Networks (GNNs) have emerged as a powerful paradigm in Knowledge Graphs (KGs) due to their intrinsic ability to model graph-structured data.\u003c/li\u003e\n\u003cli\u003eHowever, there remains a lack of a systematic review of GNN-based methods across the entire knowledge graph technology pipeline.\u003c/li\u003e\n\u003cli\u003eTo address this gap, we first propose a novel two-level classification framework for GNN-based knowledge graph technologies: the KG technologies pipeline and the GNN-based perspective.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09666v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Graph Neural Networks (GNNs) have emerged as a powerful paradigm in Knowledge Graphs (KGs) due to their intrinsic ability to model graph-structured da…\u003c/li\u003e\n\u003cli\u003eHowever, there remains a lack of a systematic review about GNN-based methodologies across the entire knowledge graph technologies pipeline\u003c/li\u003e\n\u003cli\u003eTo address this gap, we first propose a novel two-level taxonomy framework for GNN-based knowledge graph technologies: the KG technologies pipeline and GNN-base…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09668\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003ePosition: Every Ground Truth is a Human Construction, not an Objective Truth\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09668v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Ground truth datasets play a fundamental role as reference values in the training and evaluation of machine learning models.\u003c/li\u003e\n\u003cli\u003eThis position paper argues that ground truths are not neutral objective measurements naturally given, but instead that they are constructed by human and technological arrangements.\u003c/li\u003e\n\u003cli\u003eWe believe that the machine learning community will benefit from elucidating and discussing these often invisible or unreported choices, and acknowledging that reference datasets are contingent rather than universal.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09668v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Ground truth datasets play a fundamental role as reference values in the training and evaluation of machine learning models\u003c/li\u003e\n\u003cli\u003eThis position paper argues that ground truths are not neutral objective measurements that are naturally given, but instead that they are constructed by arrangem…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe argue that the ML community will benefit from articulating and discussing these often invisible or unreported choices and acknowledging that reference data s…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09682\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eAuditWeave: A Tamper-Evident, Auditor-Navigable Evidence Layer for AI-Assisted and Data-Transformation Workflows\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09682v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: AI systems are increasingly used to assist consequential decisions in regulated domains such as auditing, finance, and healthcare.\u003c/li\u003e\n\u003cli\u003eThis creates a recurring obligation: an organization must be able to reconstruct, after the fact, which evidence informed a given conclusion, and to show that the record of that reasoning has not been altered.\u003c/li\u003e\n\u003cli\u003eExisting tools address related but distinct problems - model observability, drift monitoring, governance reporting - and are built for machine-learning engineers of operating systems, not for reviewers who must trace a specific conclusion back to its supporting evidence.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09682v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: AI systems are increasingly used to assist consequential decisions in regulated domains such as auditing, finance, and healthcare\u003c/li\u003e\n\u003cli\u003eThis creates a recurring obligation: an organization must be able to reconstruct, after the fact, which evidence informed a given conclusion, and to show that t…\u003c/li\u003e\n\u003cli\u003eExisting tools address related but distinct problems - model observability, drift monitoring, governance reporting - and are built for the machine-learning engi…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09683\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eAblation, Statistical Inference, and Validation for KV-Cache Compression\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09683v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: This study systematically compares Turbo-Quant and SpectralQuant KV-cache compression, evaluating non-dominated schemes, including WHT rotation with Beta Lloyd-Max and QJL, separating system codec differences from implementation differences through statistical validation methods.\u003c/li\u003e\n\u003cli\u003eKey findings reveal that while eigenbasis-based methods fail on heavy-tailed data due to covariance instability, they excel in structured regimes, where the effective semantic dimension ($d_{eff}$) adapts to the calibration budget rather than the true data rank.\u003c/li\u003e\n\u003cli\u003e(This is an abstract of the abstract, thank you).\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09683v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: This study systematically compares Turbo-Quant and SpectralQuant KV-cache compression, evaluating non-dominated schemes, including WHT rotation with B…\u003c/li\u003e\n\u003cli\u003eKey findings reveal that while eigenbasis-based methods fail on heavy-tailed data due to covariance instability, they excel in structured regimes, with the effe…\u003c/li\u003e\n\u003cli\u003e(this is an abstract of the abstract thank you )\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09684\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eSciML in the Wild: A Diagnostic Study of When Structural Priors Help and When They Hurt\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePosted: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09684v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Scientific Machine Learning (SciML) methods such as Neural Ordinary Differential Equations (NODEs), Physics-Informed Neural Networks (PINNs), and Universal Differential Equations (UDEs) are most effective when structural priors reflect reliable dynamic controls.\u003c/li\u003e\n\u003cli\u003eWe ask what happens when this assumption is violated.\u003c/li\u003e\n\u003cli\u003eUsing macroeconomic forecasting as a stress-test domain, we evaluate five model families (ARIMA, LSTM, NODE, PINN, and UDE) across 23 countries using sparse annual data, multiple time splits, and five random seeds.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09684v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Scientific Machine Learning (SciML) methods such as Neural Ordinary Differential Equations (NODEs), Physics-Informed Neural Networks (PINNs), and Univ…\u003c/li\u003e\n\u003cli\u003eWe ask what happens when this assumption is violated\u003c/li\u003e\n\u003cli\u003eUsing macroeconomic forecasting as a stress-test domain, we evaluate five model families, ARIMA, LSTM, NODE, PINN, and UDE, across 23 countries using sparse ann…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09686\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eMawForge: Memory-Bounded Expert Materialization for Local Mixture-of-Experts Inference\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePosted: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09686v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Sparse Mixture-of-Experts (MoE) language models separate the total parameter count from the active computation per token, but local inference systems often still require the full model, key-value cache, runtime buffers, and operating system space to fit into fast memory.\u003c/li\u003e\n\u003cli\u003eMawForge tests a different systems hypothesis: local MoE serving can be made practical on constrained unified-memory machines by storing the full model on disk, keeping common tensors resident, and materializing routed expert tensors on-demand into a bounded execution cache.\u003c/li\u003e\n\u003cli\u003eThe central finding is that MawForge is effective as a bounded execution mechanism and measurement substrate for local MoE inference, but not as a cache-maximization strategy.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09686v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Sparse Mixture-of-Experts (MoE) language models separate total parameter count from per-token active computation, but local inference systems often st…\u003c/li\u003e\n\u003cli\u003eMawForge tests a different systems hypothesis: local MoE serving can be made practical on constrained unified-memory machines by storing the full model on disk,…\u003c/li\u003e\n\u003cli\u003eThe central finding is that MawForge is effective as a bounded execution mechanism and measurement substrate for local MoE inference, but not as a cache-maximiz…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09688\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003ePrioritizing Search Space Regions in the Low Autocorrelation Binary Sequences Problem\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePosted: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: - arXiv:2607.09688v1 Announce Type: new.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eAbstract: The Low Autocorrelation Binary Sequences (LABS) problem is a hard combinatorial optimization challenge with important applications in communication, signal processing, and satellite navigation.\u003c/li\u003e\n\u003cli\u003eThis paper proposes a hybrid search framework that combines Thompson sampling with parallel self-avoiding walks to adaptively allocate computational effort across restricted classes of the LABS search space.\u003c/li\u003e\n\u003cli\u003eBy modeling partitions as arms in a multi-armed bandit setting, the proposed method dynamically shifts search resources toward partitions that empirically produce higher merit factors, while maintaining exploration of less-sampled regions.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eEN Highlights:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09688v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Low autocorrelation binary sequences problem (LABS) is a hard combinatorial optimization challenge with important applications in communications, sign…\u003c/li\u003e\n\u003cli\u003eThis paper proposes a hybrid search framework that combines Thompson sampling with parallel self-avoiding walks to adaptively allocate computational effort acro…\u003c/li\u003e\n\u003cli\u003eBy modeling partitions as arms in a multi-armed bandit setting, the proposed method dynamically shifts search resources toward partitions that empirically produ…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09691\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eWhat Context Does a Coding Agent Actually Need to Act?\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eRelease Time: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09691v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: A modern coding agent can hold an entire repository in its context window.\u003c/li\u003e\n\u003cli\u003eMost of its reading is wasted - and the interesting question is not how much context an agent can use, but what it actually \\emph{needs}.\u003c/li\u003e\n\u003cli\u003eWe study that question at the moment it matters most: when the agent must \\emph{edit} code.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09691v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: A modern coding agent can hold an entire repository in its context window\u003c/li\u003e\n\u003cli\u003eMost of its reading is wasted \u0026ndash; and the interesting question is not how much context an agent can use, but what it actually \\emph{needs}\u003c/li\u003e\n\u003cli\u003eWe study that question at the moment it matters most: when the agent must \\emph{edit} code\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09692\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eReference-Based Distillation Detection in LLMs\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eRelease Time: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09692v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Model distillation—training on the output of a more powerful third-party model—is widely used to improve performance but raises concerns about unfair advantages and policy violations.\u003c/li\u003e\n\u003cli\u003eThis raises a fundamental question: can we detect if one model has been distilled from another?\u003c/li\u003e\n\u003cli\u003eWe show that while identifying the teacher model from an isolated student is very challenging, it becomes tractable in a reference-based setting: given a model and an earlier checkpoint from the same lineage, we can identify the teacher model used to train the later checkpoint.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09692v1 Announce Type: new\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.09693\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDepth-Entropy Guided Sampling for Training-Free LLM Reasoning\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-14 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.09693v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Reinforcement learning (RL) has become the dominant paradigm for improving the reasoning capabilities of large language models, but it requires expensive training, curated data, and reward signals.\u003c/li\u003e\n\u003cli\u003eRecent work shows that sampling from sharpened base-model distributions at test time recovers much of the RL gain, yet existing methods rely solely on output-layer likelihoods and ignore the transformer\u0026rsquo;s internal forward-pass dynamics.\u003c/li\u003e\n\u003cli\u003eWe introduce Depth-Entropy Guided Sampling (DEGS), a training-free, test-time method that exploits layer-wise entropy collapse as an intrinsic quality signal.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.09693v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Reinforcement learning (RL) has become the dominant paradigm for improving the reasoning capabilities of large language models, but it requires expens…\u003c/li\u003e\n\u003cli\u003eRecent work shows that sampling from sharpened base-model distributions at test time recovers much of the RL gain, yet existing methods rely solely on output-la…\u003c/li\u003e\n\u003cli\u003eWe introduce Depth-Entropy Guided Sampling (DEGS), a training-free, test-time method that exploits layer-wise entropy collapse as an intrinsic quality signal\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n",
  "wordCount": 7432,
  "readingTime": 35,
  "tableOfContents": "\u003cnav id=\"TableOfContents\"\u003e\n  \u003cul\u003e\n    \u003cli\u003e\u003ca href=\"#-deep-dive-this-issues-watch-list\"\u003e📖 Deep Dive: This Issue\u0026rsquo;s Watch List\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-ai-hot-topics-on-x\"\u003e🌐 AI Hot Topics on X\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#topic-1-openais-codex-and-chatgpt-work-hit-8-million-users-with-free-resets\"\u003eTopic 1: OpenAI\u0026rsquo;s Codex and ChatGPT Work Hit 8 Million Users with Free Resets\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-2-openais-gpt-56-sol-tops-benchmarks-amid-efficiency-fixes-and-bug-reports\"\u003eTopic 2: OpenAI\u0026rsquo;s GPT-5.6 Sol Tops Benchmarks Amid Efficiency Fixes and Bug Reports\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-3-samsung-to-produce-custom-ai-chips-for-anthropic-report-says\"\u003eTopic 3: Samsung to Produce Custom AI Chips for Anthropic, Report Says\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-4-tencent-releases-quantized-hy3-ai-model-for-single-gpus\"\u003eTopic 4: Tencent Releases Quantized Hy3 AI Model for Single GPUs\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-influencer-insights\"\u003e💡 Influencer Insights\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#ai-industry-daily-insights-based-on-analysis-of-x-influencer-tweets\"\u003eAI Industry Daily Insights (Based on Analysis of X Influencer Tweets)\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-appendix-todays-watch-list-source-updates\"\u003e📚 Appendix: Today\u0026rsquo;s Watch List Source Updates\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#stratechery-by-ben-thompson-a_full\"\u003eStratechery by Ben Thompson (A_full)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#openai-blog-a_full\"\u003eOpenAI Blog (A_full)\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#watch-the-on-demand-webinar\"\u003eWatch the on-demand webinar.\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-csai-b_introsearch\"\u003eArXiv cs.AI (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-cscl-b_introsearch\"\u003eArXiv cs.CL (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-cslg-b_introsearch\"\u003eArXiv cs.LG (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n  \u003c/ul\u003e\n\u003c/nav\u003e",
  "isDraft": false
}
