{
  "title": "2026-07-25 AI Daily | Open-source weights enter policy debate, voice multi-agent assistants beginning to take shape",
  "url": "https://miaok.ong/en/ai-daily/ai-daily-2026-07-25/",
  "date": "2026-07-25T07:00:00+08:00",
  "lastmod": "2026-07-25T07:00:00+08:00",
  "type": "ai-daily",
  "kind": "page",
  "language": "en",
  "description": "Today\u0026rsquo;s main theme shifted from model capabilities to ecosystem control. Tech companies call for the protection of open-weight models, and the Anthropic settlement case also brought renewed attention to copyright, regulation, and competitive boundaries. On the product side, ChatGPT desktop voice, Codex real-time voice, and multi-agent scheduling demonstrate that AI assistants are evolving from text tools to commandable workflow entry points.",
  "keywords": null,
  "tags": [],
  "categories": [],
  "author": "Mark (Miao) Kong",
  "image": "https://miaok.ong/images/avatar.jpg",
  "content": "\u003ch1 id=\"2026-07-25-ai-daily--open-weight-models-enter-the-policy-arena-voice-based-multi-agent-assistants-start-to-take-shape\"\u003e\n  2026-07-25 AI Daily | Open-Weight Models Enter the Policy Arena, Voice-Based Multi-Agent Assistants Start to Take Shape\n  \u003ca class=\"heading-link\" href=\"#2026-07-25-ai-daily--open-weight-models-enter-the-policy-arena-voice-based-multi-agent-assistants-start-to-take-shape\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h1\u003e\n\u003cblockquote\u003e\n\u003cp\u003eToday\u0026rsquo;s main theme shifts from model capabilities to ecosystem control. Tech companies are calling for the protection of open-weight models, and the Anthropic settlement case brings renewed attention to the boundaries of copyright, regulation, and competition. On the product side, ChatGPT\u0026rsquo;s desktop voice features, real-time voice in Codex, and multi-agent orchestration indicate that AI assistants are evolving from text-based tools into commandable workflow entry points.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch2 id=\"-in-depth-guide-to-this-issues-watch-list\"\u003e\n  📖 In-depth Guide to This Issue\u0026rsquo;s Watch List\n  \u003ca class=\"heading-link\" href=\"#-in-depth-guide-to-this-issues-watch-list\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003eThe most important reads today follow three threads. First, podcasts and Stratechery are both probing the regulatory, copyright, and competitive boundaries behind the \u0026ldquo;Open AI War\u0026rdquo; and the Anthropic 1.5B settlement. This isn\u0026rsquo;t just legal news; it\u0026rsquo;s about the realignment of power in the industry. Second, several arXiv papers focus on deconstructing MoE, hallucinations, compression, and knowledge editing. Topics like whether routing approaches Huffman coding and how to perform fixes without compromising core capabilities are worth close attention from algorithm teams. Third, agent-based AI is moving from concept to vertical implementation. Human-machine collaboration is being explored in clinical review, materials literature, and test management, but evaluation, controllability, and trustworthiness remain key hurdles.\u003c/p\u003e\n\u003ch2 id=\"-ai-hot-topics-on-x\"\u003e\n  🌐 AI Hot Topics on X\n  \u003ca class=\"heading-link\" href=\"#-ai-hot-topics-on-x\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003ch3 id=\"topic-1-tech-giants-urge-us-to-protect-open-weight-ai-models\"\u003e\n  Topic 1: Tech Giants Urge U.S. to Protect Open-Weight AI Models\n  \u003ca class=\"heading-link\" href=\"#topic-1-tech-giants-urge-us-to-protect-open-weight-ai-models\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending time: 10 hours ago, Related posts: 111,000\u003c/li\u003e\n\u003cli\u003eWhat it is: Several major tech giants are urging the U.S. government to protect open-weight AI models and avoid policies or regulations that would restrict their release and use.\u003c/li\u003e\n\u003cli\u003eWhy it matters: This affects whether AI models can be widely reused, fine-tuned, and deployed, directly impacting the speed of technological innovation, the openness of the ecosystem, and the competitive landscape between the U.S. and China in AI.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: Discussions on X are mainly focused on whether open-weight or closed-source models are more beneficial for innovation and safety, whether the U.S. should support open models through policy, and whether the emergence of Chinese models like Kimi K3 signifies a new shift in the U.S.-China AI model race.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-2-uncle-bob-skips-ai-code-reviews-with-strict-testing-gauntlet\"\u003e\n  Topic 2: Uncle Bob Skips AI Code Reviews with Strict Testing Gauntlet\n  \u003ca class=\"heading-link\" href=\"#topic-2-uncle-bob-skips-ai-code-reviews-with-strict-testing-gauntlet\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · Other\u003c/li\u003e\n\u003cli\u003eOverview: Trending time: 1 day ago, Related posts: 11,000\u003c/li\u003e\n\u003cli\u003eWhat it is: Robert C. Martin (Uncle Bob) stated that he does not use AI for code reviews, instead relying on a rigorous testing pipeline to ensure code quality.\u003c/li\u003e\n\u003cli\u003eWhy it matters: This touches on the boundaries of AI in software engineering: whether AI can replace or assist in code reviews, and the central role of automated testing in ensuring code reliability.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: The discussion on X is divided into two main camps: one side agrees with the pragmatic \u0026ldquo;tests first, AI later\u0026rdquo; approach, believing that code quality should ultimately be determined by verifiable tests; the other side argues that AI can still be used to improve review efficiency, questioning this stance as overly conservative or dismissive of AI\u0026rsquo;s assistive value.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-3-anthropic-launches-claude-opus-5-with-frontier-performance-at-half-the-cost\"\u003e\n  Topic 3: Anthropic Launches Claude Opus 5 with Frontier Performance at Half the Cost\n  \u003ca class=\"heading-link\" href=\"#topic-3-anthropic-launches-claude-opus-5-with-frontier-performance-at-half-the-cost\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending time: 6 hours ago, Related posts: 39,000\u003c/li\u003e\n\u003cli\u003eWhat it is: Anthropic has released Claude Opus 5, claiming it achieves near-frontier performance at about half the cost. It also introduces low, medium, and high inference intensity options to control task expenses.\u003c/li\u003e\n\u003cli\u003eWhy it matters: This indicates that the competition among large models is shifting from a pure focus on peak performance to a balanced optimization of performance, cost, and controllability. This could influence budget decisions for enterprises deploying AI Agents and high-intensity inference tasks.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: Discussions on X center on whether Claude Opus 5 truly delivers \u0026ldquo;frontier performance,\u0026rdquo; whether its cost advantages can be realized in practical scenarios, and if the benchmarks are inflated for marketing. Others are focused on the practical value of the inference intensity switch for developers and enterprise users.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-4-xai-announces-grok-46-and-47-releases-weeks-apart\"\u003e\n  Topic 4: xAI Announces Grok 4.6 and 4.7 Releases Weeks Apart\n  \u003ca class=\"heading-link\" href=\"#topic-4-xai-announces-grok-46-and-47-releases-weeks-apart\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending time: 5 hours ago, Related posts: 5,300\u003c/li\u003e\n\u003cli\u003eWhat it is: xAI announced on X that Grok 4.6 and Grok 4.7 will be released sequentially, just weeks apart.\u003c/li\u003e\n\u003cli\u003eWhy it matters: This reflects that large model products are entering a phase of rapid iteration. It also suggests that xAI is accelerating its efforts to catch up and close the gap with major competitors in terms of capabilities, speed, and product cadence.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: The discussion on X is focused on two points: first, whether these two updates will bring significant improvements in reasoning, code, and multimodal capabilities; second, whether such a rapid release schedule is a sign of enhanced strength or merely a marketing tactic to capture attention.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-5-etched-raises-300m-at-103-billion-valuation-for-ai-inference-chips\"\u003e\n  Topic 5: Etched Raises $300M at $10.3 Billion Valuation for AI Inference Chips\n  \u003ca class=\"heading-link\" href=\"#topic-5-etched-raises-300m-at-103-billion-valuation-for-ai-inference-chips\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending time: 1 day ago, Related posts: 3,900\u003c/li\u003e\n\u003cli\u003eWhat it is: AI chip startup Etched has completed a $300 million financing round, reaching a valuation of $10.3 billion, with a focus on AI inference chips.\u003c/li\u003e\n\u003cli\u003eWhy it matters: The cost of AI inference and the supply of computing power are becoming critical bottlenecks for the commercialization of large models. Etched\u0026rsquo;s high-valuation financing shows that capital continues to bet on specialized inference chips to challenge general-purpose GPU solutions like NVIDIA\u0026rsquo;s.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: Discussions on X are focused on whether Etched\u0026rsquo;s valuation is too high, whether specialized ASICs can maintain an advantage amid rapidly changing model architectures, and whether it has a chance to shake NVIDIA\u0026rsquo;s dominant position in the inference market.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-6-ai-community-awaits-opus-5-as-openai-rolls-out-voice-on-desktop\"\u003e\n  Topic 6: AI Community Awaits Opus 5 as OpenAI Rolls Out Voice on Desktop\n  \u003ca class=\"heading-link\" href=\"#topic-6-ai-community-awaits-opus-5-as-openai-rolls-out-voice-on-desktop\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending: 2 days ago, Related posts: 18,000\u003c/li\u003e\n\u003cli\u003eWhat it is: As OpenAI launches its voice feature on desktop, the AI community is closely watching for the potential release of Anthropic\u0026rsquo;s Claude Opus 5.\u003c/li\u003e\n\u003cli\u003eWhy it matters: This reflects how major AI companies are accelerating competition in both model capability upgrades and multimodal interactive experiences. Voice and more powerful models could further push AI assistants into daily office and productivity scenarios.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: Discussions on X are centered on whether Opus 5 will significantly surpass existing models, whether OpenAI\u0026rsquo;s desktop voice experience is practical enough, and the competitive gap between the two companies in multimodality, inference capabilities, and speed of product implementation.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-7-class-of-2027-prospects-land-first-division-i-offers\"\u003e\n  Topic 7: Class of 2027 Prospects Land First Division I Offers\n  \u003ca class=\"heading-link\" href=\"#topic-7-class-of-2027-prospects-land-first-division-i-offers\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · Other\u003c/li\u003e\n\u003cli\u003eOverview: Trending:, Related posts: 437\u003c/li\u003e\n\u003cli\u003eWhat it is: On X, some are discussing Vicor Corporation ($VICR) as an under-the-radar \u0026ldquo;next-generation AI power architecture\u0026rdquo; stock, debating its opportunities in AI infrastructure.\u003c/li\u003e\n\u003cli\u003eWhy it matters: As the computing power of AI chips continues to increase, so do the demands for power supply and cooling. Power architecture has become a critical component in the expansion of AI servers and data centers.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: The current discussion focuses on whether Vicor is undervalued by the market, whether it can benefit from upgrades in AI power supply, and whether its valuation and performance realization pace are sufficient to support a bullish outlook.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-8-ai-clip-of-green-eyed-woman-at-blue-jays-game-divides-opinions-on-beauty\"\u003e\n  Topic 8: AI Clip of Green-Eyed Woman at Blue Jays Game Divides Opinions on Beauty\n  \u003ca class=\"heading-link\" href=\"#topic-8-ai-clip-of-green-eyed-woman-at-blue-jays-game-divides-opinions-on-beauty\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · Entertainment\u003c/li\u003e\n\u003cli\u003eOverview: Trending:, Related posts: 40\u003c/li\u003e\n\u003cli\u003eWhat it is: A short video, allegedly generated by AI, of a \u0026ldquo;green-eyed woman at a Blue Jays game\u0026rdquo; has spread on X, sparking discussions among users about her appearance and its authenticity.\u003c/li\u003e\n\u003cli\u003eWhy it matters: This event reflects the increasing realism of generative AI in creating entertainment content and human images. It also highlights the impact of synthetic media on aesthetics, authenticity recognition, and platform distribution.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: Discussions on X mainly focus on whether the video is real, whether AI-generated beautiful women reinforce a single standard of beauty, and why people are attracted to virtual figures. Some also believe it\u0026rsquo;s just harmless entertainment and shouldn\u0026rsquo;t be over-analyzed.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch4 id=\"ai-public-opinion-summary-on-x-today\"\u003e\n  AI Public Opinion Summary on X Today\n  \u003ca class=\"heading-link\" href=\"#ai-public-opinion-summary-on-x-today\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h4\u003e\n\u003cp\u003eThe main theme of today\u0026rsquo;s public opinion is that AI competition is shifting from single-point model capabilities to a full-chain battle encompassing \u0026ldquo;open ecosystems, cost efficiency, product experience, and infrastructure.\u0026rdquo; Policies on open-sourcing weights, the high-frequency iterations of Claude Opus 5 and Grok, OpenAI\u0026rsquo;s voice feature, and investments in inference chips and power architecture all indicate that AI is accelerating towards large-scale deployment. A clear consensus is emerging that inference cost, controllability, computing power, and deployment efficiency will be key in the next stage of competition. Companies and developers are no longer just looking at the highest benchmark scores but are also paying more attention to practical usability, cost, and the risk of ecosystem lock-in. Points of disagreement are centered on openness versus security, whether AI should be deeply involved in code review, whether model releases represent genuine performance breakthroughs or just marketing hype, and whether the valuations of specialized chips and AI infrastructure stocks are already overdrawn. Potential risks include innovation-security imbalances caused by excessive or insufficient regulation, investment bubbles and technological misjudgments due to overly rapid model and hardware iterations, and the further blurring of reality and fiction by generative imagery, which amplifies aesthetic homogenization and information credibility issues.\u003c/p\u003e\n\u003ch2 id=\"-influencer-insights\"\u003e\n  💡 Influencer Insights\n  \u003ca class=\"heading-link\" href=\"#-influencer-insights\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003eOkay, based on the tweet content from multiple AI influencers over the past 24 hours, here is an in-depth analysis report combined with recent hot topics.\u003c/p\u003e\n\u003ch3 id=\"1-key-technology-trends-or-product-hotspots-followed-by-influencers-today\"\u003e\n  1. Key Technology Trends or Product Hotspots Followed by Influencers Today\n  \u003ca class=\"heading-link\" href=\"#1-key-technology-trends-or-product-hotspots-followed-by-influencers-today\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003eA. The Comprehensive Outbreak of Multimodality, Multi-Agent Collaboration, and Voice Interaction\u003c/strong\u003e\nThis is the core theme today, signaling that the interaction dimension of AI assistants is evolving from singular text/code towards a more anthropomorphic and parallelized direction.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003e\u0026ldquo;Voice Control for Everything\u0026rdquo; on ChatGPT Desktop\u003c/strong\u003e: @dotey and @vista8 both highlighted the update to the ChatGPT desktop client (formerly the Codex App). It integrates a voice mode based on a \u003ccode\u003eGPT-Live\u003c/code\u003e full-duplex architecture. The core breakthrough is that it \u003cstrong\u003eallows you to use natural language to orchestrate multiple background Agents simultaneously\u003c/strong\u003e (such as Codex for writing code and ChatGPT Work for running tasks), and can see the screen context via \u003ccode\u003eAppshots\u003c/code\u003e (as mentioned by @dotey). This is akin to a commander issuing voice commands to a digital team.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eCodex\u0026rsquo;s Real-Time Voice Mode\u003c/strong\u003e: @Pluvio9yte cited a leak claiming that Codex is about to launch \u003ccode\u003eRealtime Voice Mode\u003c/code\u003e, where the main assistant handles conversation while \u003ccode\u003eworker agents\u003c/code\u003e processes tasks like Slack, Spotify, and web browsing in the background. This aligns with the strategy for the OpenAI Chat desktop client, moving towards a \u0026ldquo;voice + multi-Agent collaboration\u0026rdquo; super personal assistant.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003e\u0026ldquo;Dictation Programming\u0026rdquo; and Stream-of-Consciousness Input\u003c/strong\u003e: @Pluvio9yte relayed @karpathy\u0026rsquo;s unique workflow: when faced with complex ideas and not wanting to type, he switches to voice mode for a 10-minute \u0026ldquo;stream-of-consciousness\u0026rdquo; dump, letting the AI grasp the original intent. This indicates that voice is not just for commands but is also becoming a \u0026ldquo;high-bandwidth\u0026rdquo; input channel for unstructured thoughts, solving the information loss problem associated with typing.\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e\u003cstrong\u003eB. The \u0026ldquo;Warring States Period\u0026rdquo; of the Programming Agent Ecosystem and Paradigm Shifts\u003c/strong\u003e\nProgramming Agents remain the most competitive field, and today\u0026rsquo;s discussion has shifted from model evaluation to toolchains, cost control, and a new philosophy of code supervision.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eThe Model Battle: Claude Opus 5\u0026rsquo;s \u0026ldquo;Cost-Effectiveness\u0026rdquo; Positioning\u003c/strong\u003e: @dotey provided a detailed analysis of Anthropic\u0026rsquo;s release of Claude Opus 5. It is positioned to \u0026ldquo;provide near-Fable 5 frontier intelligence at half the price,\u0026rdquo; performing impressively across multiple benchmarks and particularly excelling at long-range tasks like autonomously building testing frameworks. This indicates that model vendors are starting to differentiate their product lines, balancing peak performance with cost.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eOpen-Source Tool Competition Heats Up\u003c/strong\u003e:\n\u003cul\u003e\n\u003cli\u003e@Pluvio9yte shared \u0026ldquo;This week\u0026rsquo;s top 10 trending AI open-source projects on GitHub,\u0026rdquo; the vast majority of which are Agent-related. Examples include \u003ccode\u003emattpocock/skills\u003c/code\u003e (composable Agent engineering patterns), \u003ccode\u003eorca\u003c/code\u003e (managing multiple coding Agents in parallel), and \u003ccode\u003ecode-review-graph\u003c/code\u003e (parsing codebases into knowledge graphs to reduce token consumption).\u003c/li\u003e\n\u003cli\u003e@AI_Jasonyu mentioned SpaceX\u0026rsquo;s open-source terminal AI coding Agent, \u003ccode\u003egrok-build\u003c/code\u003e, which is feature-complete, has 14k stars, and competes directly with tools like Claude Code.\u003c/li\u003e\n\u003cli\u003e@Pluvio9yte also discovered \u003ccode\u003eOpenCodex\u003c/code\u003e, which allows Codex applications to connect to other large models like Kimi, Grok, and GLM, reflecting the developer demand to avoid single-model lock-in.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eThe New \u0026ldquo;Don\u0026rsquo;t Read the Code\u0026rdquo; Programming Philosophy\u003c/strong\u003e: @dotey quoted Uncle Bob (@unclebobmartin), author of \u003cem\u003eClean Code\u003c/em\u003e, who said: \u0026ldquo;\u003cstrong\u003eI don\u0026rsquo;t read AI-written code because humans read code too slowly, which defeats the purpose of using AI.\u003c/strong\u003e\u0026rdquo; His new method involves setting up a series of gates for the Agent (tests, quality metrics, mutation testing) to manage code quality through metrics rather than manual inspection. This signals that the core competency in programming is shifting from \u0026ldquo;reading and writing code\u0026rdquo; to \u0026ldquo;defining constraints, writing tests, and interpreting metrics.\u0026rdquo;\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e\u003cstrong\u003eC. Practical Application of On-Device Models and Hardware Choices\u003c/strong\u003e\n@zhixianio and @ruanyf continue to focus on on-device models, with today\u0026rsquo;s topic delving into specific applications and hardware comparisons. @ruanyf suggested that for running AI locally, a mini PC with an on-board chipset like the AMD Strix Halo is often a better choice than a high-end dedicated graphics card, thanks to its 128GB unified memory advantage. @zhixianio, on the other hand, considers Google\u0026rsquo;s \u003cstrong\u003eGemma 4 Quantization-Aware Training (QAT)\u003c/strong\u003e to be an important optimization approach for on-device models that will accelerate their deployment on Android devices.\u003c/p\u003e\n\u003ch3 id=\"2-notable-unique-perspectives-or-industry-foresight\"\u003e\n  2. Notable Unique Perspectives or Industry Foresight\n  \u003ca class=\"heading-link\" href=\"#2-notable-unique-perspectives-or-industry-foresight\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003e\u0026ldquo;The Pharmaceutical Business with a Ten-Month Patent Period\u0026rdquo;\u003c/strong\u003e: @dotey forwarded a view from @xleaps that compares the large model industry to the pharmaceutical business but with extremely short patent periods, starkly revealing the harsh reality of rapid model iterations and brief windows for monetization.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eAI Hasn\u0026rsquo;t Brought Leisure, But Stronger Shackles\u003c/strong\u003e: In a podcast, @vista8 reflected that many people are chained to the credit reset times of their AI programming tools, fostering a \u0026ldquo;scarcity mindset\u0026rdquo; akin to farmers being domesticated by wheat. \u003cstrong\u003eAI was supposed to bring abundance, but instead, it has trapped people in a higher-frequency production rhythm\u003c/strong\u003e—a profound warning about the relationship between technology and humanity.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eProduct Design from \u0026lsquo;Gamification\u0026rsquo; to \u0026lsquo;Game Sense\u0026rsquo;\u003c/strong\u003e: @nishuang, through a language learning App \u003ccode\u003eCapWords\u003c/code\u003e, incisively differentiated between dopamine-driven \u0026lsquo;gamification\u0026rsquo; (e.g., Duolingo\u0026rsquo;s reward mechanism) and endorphin-driven \u0026lsquo;game sense\u0026rsquo; (e.g., the joy of collecting in Pokémon-style games), pointing the direction for AI product interaction design.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eEtymological Research on the Term \u0026lsquo;TikTok\u0026rsquo;\u003c/strong\u003e: @ruanyf discovered that the term \u003ccode\u003eTikTok\u003c/code\u003e is actually the name of a robot in \u0026lsquo;The Wizard of Oz\u0026rsquo; series of novels, and not coined by ByteDance.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eNew Professions in the AI Era\u003c/strong\u003e: @gefei55 observed that by rapidly learning cutting-edge knowledge with AI and combining it with deliberate practice, becoming an \u003cstrong\u003eoffline conference speaker\u003c/strong\u003e is emerging as a free, high-income side hustle.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eControversy over the Global Competitiveness of Chinese Open-Source Models\u003c/strong\u003e: @vista8 mentioned in an article that multiple US AI startups jointly called for not banning Chinese models, as a ban would only protect the high pricing of US frontier models, indirectly confirming the huge competitiveness of Chinese models in terms of cost-effectiveness.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"3-recommended-tools-or-resources\"\u003e\n  3. Recommended Tools or Resources\n  \u003ca class=\"heading-link\" href=\"#3-recommended-tools-or-resources\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003eOpen Source / Tools:\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eSkills and Workflows:\u003c/strong\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ccode\u003eAgent Skills\u003c/code\u003e Collection (@mattpocock): Compiles engineering experience such as TDD and debugging into composable skill packs, recommended by @Pluvio9yte.\u003c/li\u003e\n\u003cli\u003eXiangyang Qiaomu (@vista8)\u0026rsquo;s Skill Series: Includes \u003cstrong\u003evideo editing/download, frontend design, AI PRD generation, server deployment\u003c/strong\u003e, etc., which can be installed with one click via \u003ccode\u003enpx skills add\u003c/code\u003e, highly practical.\u003c/li\u003e\n\u003cli\u003e\u003ccode\u003eTopview MCP\u003c/code\u003e (@TopviewAIhq): A full-stack marketing MCP integrating Amazon, YouTube, and TikTok Shop data, enabling automation from data analysis to content generation. Recommended by @AI_Jasonyu.\u003c/li\u003e\n\u003cli\u003e\u003ccode\u003eclaude-tap\u003c/code\u003e (recommended by @seekjourney, retweeted by @dotey): A local observability platform for Claude Code and other Agents, making it convenient for developers to gain insight into the actual running process of Agents.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eDownload Tools:\u003c/strong\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ccode\u003evista8\u003c/code\u003e developed a Video Account download Skill, solving the problem of difficulty in downloading Video Account content.\u003c/li\u003e\n\u003cli\u003e\u003ccode\u003eFlclash\u003c/code\u003e (recommended by @AI_Jasonyu): An Android open-source VPN tool based on Clash, with a user-friendly interface and card-style layout.\u003c/li\u003e\n\u003cli\u003e\u003ccode\u003eOfficeCLI\u003c/code\u003e: A rising star project from @Pluvio9yte\u0026rsquo;s weekly report, allowing Agents to directly read and write Word, Excel, and PPT without installing Office.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e\u003cstrong\u003ePlatforms / Resources:\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eAIHOT\u003c/strong\u003e (@Khazix0918): An AI news aggregation platform with over 600,000 monthly active users, jointly recommended by @dotey and @vista8, who called it a \u0026lsquo;crystallization of taste and experience\u0026rsquo;.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eXiaohongshu REDSkill Community\u003c/strong\u003e: @ruanyf discovered that Xiaohongshu is building a social media-based Skill Hub that supports uploading and sharing Agent Skills, considered the Github of the Skill domain, providing a new channel for programmers to reach a massive number of C-end users.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eAPI Relay Services\u003c/strong\u003e: @ruanyf and @Pluvio9yte respectively mentioned @fennoAI and self-built relay stations, which have become a choice for many developers to solve service stability issues amidst increased risk of overseas model account bans.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eBolivian Exchange Rate Difference Loophole\u003c/strong\u003e: @Pluvio9yte discovered that the sharp drop in the Bolivian exchange rate could be exploited to subscribe to Codex 20x services at extremely low prices, but this carries high risks.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch2 id=\"-appendix-todays-watch-list-update-sources\"\u003e\n  📚 Appendix: Today\u0026rsquo;s Watch List Update Sources\n  \u003ca class=\"heading-link\" href=\"#-appendix-todays-watch-list-update-sources\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cblockquote\u003e\n\u003cp\u003eTime Window: Recent 3 days; Covering 22 sources; Total 32 updates\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch3 id=\"all-in-podcast-a_full\"\u003e\n  All-In Podcast (A_full)\n  \u003ca class=\"heading-link\" href=\"#all-in-podcast-a_full\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003e\u003ca href=\"https://allinchamathjason.libsyn.com/the-fight-over-open-source-ai-anthropics-15b-payout-nyc-socialists-evictions-violence\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eThe Fight Over Open Source AI, Anthropic\u0026rsquo;s $1.5B Payout, NYC Socialists: Evictions = Violence?\u003c/a\u003e\u003c/strong\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-25 04:46 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary:\n\u003cul\u003e\n\u003cli\u003e(0:00) Bestie intros.\u003c/li\u003e\n\u003cli\u003e(0:18) The Fight to Save Open Source AI: Kimi K3 panic, Anthropic/OpenAI regulatory capture.\u003c/li\u003e\n\u003cli\u003e(27:38) Anthropic/OpenAI historical growth rates, China\u0026rsquo;s protracted war.\u003c/li\u003e\n\u003cli\u003e(48:29) Anthropic\u0026rsquo;s $1.5B piracy settlement and massive IP theft hypocrisy.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003e(0:00) Bestie intros\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e(0:18) The fight to save open source AI: Kimi K3 panic, Anthropic/OpenAI regulatory capture\u003c/li\u003e\n\u003cli\u003e(27:38) Anthropic/OpenAI historic growth rates, China\u0026rsquo;s long game\u003c/li\u003e\n\u003cli\u003e(48:29) Anthropic\u0026rsquo;s $1.5B piracy settlement and the great IP theft hypocrisy\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"stratechery-by-ben-thompson-a_full\"\u003e\n  Stratechery by Ben Thompson (A_full)\n  \u003ca class=\"heading-link\" href=\"#stratechery-by-ben-thompson-a_full\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003e\u003ca href=\"https://stratechery.com/2026/the-copium-wars/\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003e2026.30: The Copium Wars\u003c/a\u003e\u003c/strong\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-25 01:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - (Photo by Ng Hanguan-Pool/Getty Images).\n\u003cul\u003e\n\u003cli\u003eWelcome back to This Week in Stratechery!\u003c/li\u003e\n\u003cli\u003eAs a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone.\u003c/li\u003e\n\u003cli\u003eAdditionally, you have complete control over what we send to you.\u003c/li\u003e\n\u003cli\u003eWith that said, here are some of our favorites from this week.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003e(Photo by Ng Han Guan-Pool/Getty Images)\u003c/li\u003e\n\u003cli\u003eWelcome back to This Week in Stratechery\u003c/li\u003e\n\u003cli\u003eAs a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone\u003c/li\u003e\n\u003cli\u003eAdditionally, you have complete control over what we send to you\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-csai-b_introsearch\"\u003e\n  ArXiv cs.AI (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-csai-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20452\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eAINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20452v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Modern software quality assurance demands intelligent, autonomous systems capable of adaptive decision-making across distributed cloud environments.\u003c/li\u003e\n\u003cli\u003eThis paper presents AINTMA (Agentic Intelligent Test Management Architecture), a multi-agent agentic AI system that transforms traditional test management into an autonomous quality intelligence ecosystem.\u003c/li\u003e\n\u003cli\u003eAINTMA deploys six specialized AI agents (Test Discovery, Risk Assessment, Reinforcement Learning Prioritization, Execution Orchestration, Generative Quality Intelligence, and Cloud Security Monitor), coordinated through a secure multi-agent communication framework on a cloud-native microservices infrastructure.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20452v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Modern software quality assurance demands intelligent, autonomous systems capable of adaptive decision-making across distributed cloud environments\u003c/li\u003e\n\u003cli\u003eThis paper presents AINTMA (Agentic Intelligent Test Management Architecture), a multi-agent agentic AI system that transforms traditional test management into…\u003c/li\u003e\n\u003cli\u003eAINTMA deploys six specialized AI agents (Test Discovery, Risk Assessment, Reinforcement Learning Prioritization, Execution Orchestration, Generative Quality In…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20462\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eMarking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20462v1 Announcement Type: New.\u003c/li\u003e\n\u003cli\u003eAbstract: Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need for reliable traceability of watermarked model-generated outputs.\u003c/li\u003e\n\u003cli\u003eHowever, most watermarks are evaluated on general-purpose benchmarks, leaving domains like medicine, where small token-level perturbations can result in significant semantic changes, largely unexplored.\u003c/li\u003e\n\u003cli\u003eIn this work, we present the first rigorous study of how LLM watermarks affect medical performance, benchmarking 5 watermarking schemes across 11 LLMs and 7 VLMs, encompassing various tasks in unimodal and multimodal clinical reasoning.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eKey Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20462v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need for reliable traceability of model-generated outp…\u003c/li\u003e\n\u003cli\u003eYet, most watermarks are evaluated on general-purpose benchmarks, leaving domains like medicine, where small token-level perturbations can result in significant…\u003c/li\u003e\n\u003cli\u003eIn this work, we present the first rigorous study of how LLM watermarks affect medical performance, benchmarking 5 watermarking schemes across 11 LLMs and 7 VLM…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20463\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eClickGuard: Detecting and Spoiling Clickbait News with Informativeness Measures and Large Language Models\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20463v1 Announcement Type: New.\u003c/li\u003e\n\u003cli\u003eAbstract: This paper proposes an AI-driven browser extension that identifies clickbait to help users avoid misleading internet articles.\u003c/li\u003e\n\u003cli\u003eThe application goes beyond traditional detection, employing a hybrid machine learning architecture that combines transformer-based embeddings with linguistically-driven features and a custom \u0026lsquo;bait\u0026rsquo; score.\u003c/li\u003e\n\u003cli\u003eAfter evaluating various natural language processing techniques (from classic vectorizers to large language model (LLM) embeddings), an XGBoost-based model was developed, achieving a 91% F1 score on an open composite dataset.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eKey Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20463v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: This paper presents an AI-driven browser extension that identifies clickbait to help users avoid misleading Internet articles\u003c/li\u003e\n\u003cli\u003eMoving beyond traditional detection, the application employs a hybrid machine learning architecture that combines transformer-based embeddings with linguistical…\u003c/li\u003e\n\u003cli\u003eAfter evaluating various natural language processing techniques \u0026ndash; from classic vectorizers to large language model (LLM) embeddings \u0026ndash; an XGBoost-based model w…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20464\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eStochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary: - arXiv:2607.20464v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: When a language model gives different answers on repeated runs, does that variation reveal what it does not know?\u003c/li\u003e\n\u003cli\u003eSelf-consistency turns the variation into a per-question uncertainty estimate via majority voting.\u003c/li\u003e\n\u003cli\u003eBut does the same variation reveal cross-question structure \u0026ndash; related questions flipping together, the way a diverse ensemble does?\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20464v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: When a language model gives different answers on repeated runs, does that variation reveal what it does not know\u003c/li\u003e\n\u003cli\u003eSelf-consistency turns the variation into a per-question uncertainty estimate via majority voting\u003c/li\u003e\n\u003cli\u003eBut does the same variation reveal cross-question structure \u0026ndash; related questions flipping together, the way a diverse ensemble does\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20466\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eJAXBench: Benchmarking Autonomous TPU Kernel Optimization\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary: - arXiv:2607.20466v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equivalent target exists for TPUs.\u003c/li\u003e\n\u003cli\u003eWe introduce JAXBench, a TPU-native benchmark suite for optimizing AI-generated kernels on Google Cloud TPUs.\u003c/li\u003e\n\u003cli\u003eJAXBench includes 50 JAX workloads that are both relevant and offer room for optimization.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20466v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equ…\u003c/li\u003e\n\u003cli\u003eWe present JAXBench, a TPU-native benchmark suite for AI-generated kernel optimization on Google Cloud TPUs\u003c/li\u003e\n\u003cli\u003eJAXBench comprises 50 JAX workloads that are both relevant and provide headroom for optimization\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20467\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary: - arXiv:2607.20467v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: While parallel decoding is crucial for the efficiency of diffusion large language models (dLLMs), current strategies are often hindered by overly conservative confidence thresholds.\u003c/li\u003e\n\u003cli\u003eThese thresholds, required by the Joint Probability Discrepancy Error (JPDE), lead to redundant denoising iterations and suboptimal inference speeds.\u003c/li\u003e\n\u003cli\u003eTo overcome this problem, we propose DC-Leap, a training-free framework that can reliably accelerate dLLMs under moderate confidence.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20467v1 Announce Type: new\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: While parallel decoding is central to the efficiency of Diffusion Large Language Models (dLLMs), current strategies are often hindered by overly conse…\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eThese thresholds, necessitated by the Joint Probability Dependence Error (JPDE), result in redundant denoising iterations and suboptimal inference speeds\u003c/li\u003e\n\u003cli\u003eTo overcome this, we propose DC-Leap, a training-free framework that enables reliable acceleration of dLLMs in the moderate-confidence regime\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20468\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eInferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20468v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: AI agents are increasingly used to automate research and development tasks, but existing benchmarks typically evaluate them on prescribed workflows or narrow action spaces.\u003c/li\u003e\n\u003cli\u003eEven nominally open-ended tasks can often be solved by retrieving well-known recipes and adjusting some hyperparameters, making it unclear whether strong results reflect true optimization or memorized solutions.\u003c/li\u003e\n\u003cli\u003eWe introduce InferenceBench, where an agent must deploy an OpenAI-compatible inference server and optimize the speed of LLM inference.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20468v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or…\u003c/li\u003e\n\u003cli\u003eEven nominally open-ended tasks can often be solved by retrieving a well-known recipe and tuning a few hyperparameters, making it unclear whether strong results…\u003c/li\u003e\n\u003cli\u003eWe introduce InferenceBench, where an agent must deploy an OpenAI-compatible inference server and optimize the speed of LLM inference\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20469\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20469v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Large Language Models (LLMs) use a set of parameters to handle many tasks, but under KV cache inference, it is unclear what task-general structure, if any, is used during decoding rather than during prefill.\u003c/li\u003e\n\u003cli\u003eWe propose DecodeShare, a protocol that identifies a consistently shared low-dimensional subspace among tasks in the hidden states at decode time, and then tests its causal role by ablating only this subspace during decoding.\u003c/li\u003e\n\u003cli\u003eIn our experiments, under the same intervention budget, interfering with the discovered shared subspace degrades decision performance more than interfering with prefill-derived subspaces or random subspaces.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20469v1 Announce Type: new\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: Large language models (LLMs) handle many tasks with one set of parameters, but under KV-cached inference it is unclear what task-general structure, if…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe propose DecodeShare, a protocol that identifies a low-dimensional subspace consistently shared across tasks in decode-time hidden states, and then tests its…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eIn our experiments, disturbing the discovered shared subspace degrades decision performance far more than disturbing either a prefill-derived or random subspace…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20470\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003ePlanE: Meta Planning of Data, Tuning, and Inference for Extractive-based LLMs\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20470v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Enhancing the task-specific capabilities of Large Language Models (LLMs) primarily requires substantial instruction-tuning datasets.\u003c/li\u003e\n\u003cli\u003eHowever, the sheer volume of such data imposes a considerable annotation cost, and there is a lack of optimization methods for tailoring LLMs for specific tasks.\u003c/li\u003e\n\u003cli\u003eTo address the above issues, we propose a framework for building extractive-based LLMs called \\textbf{PlanE}, which includes data decomposition, instruction tuning, and prompt inference.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20470v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Enhancing the task-specific capabilities of Large Language Models (LLMs) primarily requires substantial instruction-tuning datasets\u003c/li\u003e\n\u003cli\u003eHowever, the sheer volume of such data imposes a considerable annotation cost, and a lack of optimization methods for tailoring LLMs to specific tasks\u003c/li\u003e\n\u003cli\u003eTo address the above issues, we propose a \\textbf{Plan}ning framework for constructing \\textbf{E}xtractive-based LLMs called \\textbf{PlanE}, which includes data…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20471\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eBenchmarking the Personalization Capabilities of Large Language Models\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20471v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Personalization is the act of altering a message to induce action in a specific recipient while keeping the sender, channel, and time fixed. It has a long tradition in psychology and marketing as a two-party problem where the sender and receiver have independent goals.\u003c/li\u003e\n\u003cli\u003eLarge language models remove the finite inventory constraints of classic retrieval and ranking methods by generating a continuum of message variants conditioned on the inferred state of the recipient, raising the question of how effectively current models can perform personalization in the classic sense.\u003c/li\u003e\n\u003cli\u003eExisting LLM personalization benchmarks measure the adaptability of the sender, where the recipient is the same user being served by the model.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20471v1 Announce Type: new\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long trad…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eLarge language models remove the bounded-inventory constraint of classical retrieval-and-ranking approaches by generating a continuum of message variants condit…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eExisting LLM personalization benchmarks measure sender-side adaptation, in which the receiver is the same user the model is serving\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-cscl-b_introsearch\"\u003e\n  ArXiv cs.CL (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-cscl-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20425\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eWhat is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eRelease Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20425v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: What makes writing \u0026ldquo;good\u0026rdquo; remains a long-standing problem in literary studies and computational linguistics.\u003c/li\u003e\n\u003cli\u003eWe present a two-study investigation into how reasoning-enabled LLMs evaluate literary quality.\u003c/li\u003e\n\u003cli\u003eIn Study 1, we construct a benchmark of 30 real texts spanning six quality tiers (from canonical literature to anonymous forum posts) and extract the model\u0026rsquo;s implicit theories of quality from its reasoning traces.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20425v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: What makes writing \u0026ldquo;good\u0026rdquo; remains a persistent question in literary studies and computational linguistics\u003c/li\u003e\n\u003cli\u003eWe present a two-study investigation of how reasoning-enabled LLMs evaluate literary quality\u003c/li\u003e\n\u003cli\u003eIn Study 1, we construct a benchmark of 30 real texts spanning six quality tiers, from canonical literature to anonymous forum posts, and extract the model\u0026rsquo;s im…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20426\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eKnowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs\u0026rsquo;Hallucinations\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eRelease Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20426v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Existing LLM hallucination mitigation methods, including prompt engineering and model optimization, either find it difficult to alter the model\u0026rsquo;s internal knowledge or have poor cross-domain generalization.\u003c/li\u003e\n\u003cli\u003eContrastive decoding mitigates hallucinations by using layer-wise differences in LLMs.\u003c/li\u003e\n\u003cli\u003eHowever, previous research has only explored Transformer-based models (e.g., GPT), neglecting other effective frameworks like Mixture-of-Experts (MoE) models.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20426v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Existing LLM hallucination mitigation methods, including prompt engineering and model optimization, either hardly alter models\u0026rsquo;internal knowledge or h…\u003c/li\u003e\n\u003cli\u003eContrastive decoding mitigates hallucinations by using layer-wise differences in LLMs\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eHowever, prior studies only explore transformer-based models (e.g., GPT), ignoring other effective frameworks like mixture-of-experts (MoE) models\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20427\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eIs MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20427v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Mixture-of-Experts architectures have revolutionized scaling, yet the underlying logic of their routing remains a black box.\u003c/li\u003e\n\u003cli\u003eIn this paper, we uncover a fundamental governing principle: MoE routing is not merely selection, but a manifestation of Huffman Coding.\u003c/li\u003e\n\u003cli\u003eWe introduce the Frequency-Diversity Law, revealing that state-of-the-art models, such as Phi-3.5-MoE and Gemma-4-27B-A4B, spontaneously act as information-theoretic engines.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20427v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Mixture-of-Experts architectures have revolutionized scaling, yet the underlying logic of their routing remains a black box\u003c/li\u003e\n\u003cli\u003eIn this paper, we uncover a fundamental governing principle: MoE routing is not merely selection, but a manifestation of Huffman Coding\u003c/li\u003e\n\u003cli\u003eWe introduce the Frequency-Diversity Law, revealing that state-of-the-art models, such as Phi-3.5-MoE and Gemma-4-27B-A4B, spontaneously act as information-theo…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20428\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eHuman-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20428v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: This study evaluated a retrieval-augmented, multi-agent large language model (LLM)-driven, human-in-the-loop framework for detecting cutaneous immune-related adverse events (cirAE) from clinical records.\u003c/li\u003e\n\u003cli\u003eCompared with unassisted manual review, the LLM-assisted workflow improved accuracy (F1 = 0.88 vs 0.77), inter-rater agreement as measured by Cohen\u0026rsquo;s kappa (kappa = 0.82 vs 0.50), and reduced the average review time by approximately half.\u003c/li\u003e\n\u003cli\u003eThis framework pilots the application of LLMs to identify immune-related toxicities across organ systems and, more broadly, to achieve accurate, scalable, and transparent extraction of adverse event data.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20428v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: This study evaluated a retrieval-augmented, multi-agent large language model (LLM)-driven, human-in-the-loop framework for detecting cutaneous immune-…\u003c/li\u003e\n\u003cli\u003eCompared with unassisted manual review, the LLM-assisted workflow improved accuracy (F1 = 0.88 vs 0.77), inter-rater agreement measured by Cohen\u0026rsquo;s kappa (kappa…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eThis framework pilots how LLMs can be applied to identify immune-related toxicities across organ systems and, more broadly, enable accurate, scalable, and trans…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20429\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eMore Is Not More: What Matters for Diversity in LLM Opinions?\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20429v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group modeling, and public opinion forecasting.\u003c/li\u003e\n\u003cli\u003eHowever, LLM outputs exhibit systematic opinion homogenization.\u003c/li\u003e\n\u003cli\u003ePractitioners have explored various interventions to increase diversity, but the landscape remains fragmented: different methods are evaluated in isolation with incomparable metrics, and in practice they are often deployed and scaled concurrently, making it difficult to attribute gains to specific components.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20429v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group modeling, an…\u003c/li\u003e\n\u003cli\u003eHowever, LLM outputs exhibit systematic opinion homogenization\u003c/li\u003e\n\u003cli\u003ePractitioners have explored various interventions to increase diversity, but the landscape remains fragmented: different methods are evaluated in isolation with…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20430\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eLLM-INSTRUCT at UZH Shared Task 2026: Constraint-Aware Retrieval and Selective Debate for Paragraph-Level Argument Mining\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20430v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: We introduce LLM-INSTRUCT, the winning system for the UZH Shared Task at ArgMining 2026, which involves paragraph-level argument mining in UN and UNESCO resolutions.\u003c/li\u003e\n\u003cli\u003eThe task requires paragraph-type classification, prediction of a subset of 141 official tags, and directed relation prediction under a strict JSON schema setting using only open-weight models with up to 8B parameters.\u003c/li\u003e\n\u003cli\u003eWe define the task as constrained structured prediction.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20430v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: We present LLM-INSTRUCT, the winning system for the UZH Shared Task at ArgMining 2026 on paragraph-level argument mining in UN and UNESCO resolutions\u003c/li\u003e\n\u003cli\u003eThe task requires paragraph-type classification, prediction of a subset of 141 official tags, and directed relation prediction under a strict JSON schema settin…\u003c/li\u003e\n\u003cli\u003eWe frame the task as constrained structured prediction\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20431\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eSkill-Contracted Agents for Evidence-Aware Materials Literature Analysis\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20431v1 Announcement Type: new.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: Materials science literature analysis requires simultaneous attention to composition, processing, characterization, and property relationships, but traditional retrieval-augmented generation pipelines struggle to coordinate heterogeneous tasks in a single retrieve-then-generate architecture.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eHere, we introduce AlphaAgent, a skill-driven agent framework that decouples retrieval-based question answering from paper-level report generation through an explicit skill contract.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eA dedicated retrieval skill rewrites user requests into material-specific search intents, queries a curated index of over 300,000 papers from the Journal Citation Reports metallurgy and metallurgical engineering category, and reformulates queries when initial evidence is insufficient.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eEN Highlights:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20431v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Materials science literature analysis requires simultaneous attention to composition, processing, characterization, and property relationships, yet co…\u003c/li\u003e\n\u003cli\u003eHere we present AlphaAgent, a skill-driven agent framework that decouples retrieval-based question answering from paper-level report generation through explicit…\u003c/li\u003e\n\u003cli\u003eA dedicated retrieval skill rewrites user requests into material-specific search intents, queries a curated index of more than 300,000 papers from the Journal C…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20432\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003ePosition: Natural Language Should Not Fully Replace Formal Languages\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:- arXiv:2607.20432v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Recent advances in large language models and their widespread adoption have prompted claims that natural language could entirely replace formal languages, such as programming languages for software design.\u003c/li\u003e\n\u003cli\u003eIn this position paper, we argue that this perspective overlooks fundamental linguistic properties of natural language, specifically that it is optimized for irregularities in open-ended contexts.\u003c/li\u003e\n\u003cli\u003eWe introduce a formal framework centered on \u0026ldquo;task specificity,\u0026rdquo; defining it as the information-theoretic reduction of uncertainty in an output space (e.g., all possible images) according to the user\u0026rsquo;s specific requirements.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20432v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Recent advances in large language models and their widespread adoption have prompted claims that natural language could entirely replace formal langua…\u003c/li\u003e\n\u003cli\u003eIn this position paper, we argue that this perspective overlooks fundamental linguistic properties of natural language, specifically that it is optimized for un…\u003c/li\u003e\n\u003cli\u003eWe introduce a formal framework centered on \u003cem\u003etask specificity\u003c/em\u003e, defining it as the information-theoretic reduction of uncertainty in an output space \u0026ndash; such as…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20433\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eMoir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:- arXiv:2607.20433v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: While language models remain in a static training state, the world continues to evolve.\u003c/li\u003e\n\u003cli\u003eKnowledge editing has emerged as a key alternative to full retraining, but its deployment is bottlenecked by the erosion of core capabilities: mathematical and procedural reasoning collapse, while encyclopedic recall remains intact.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe trace this asymmetric degradation to a distributional mismatch.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20433v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: While language models remain frozen at their training state, the world evolves continuously\u003c/li\u003e\n\u003cli\u003eKnowledge editing has emerged as a key alternative to full retraining, but its deployment is bottlenecked by the erosion of core capabilities: mathematical and…\u003c/li\u003e\n\u003cli\u003eWe trace this asymmetric degradation to a distributional mismatch\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20434\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eBreak Through the Compression Bottleneck: From Theory to Practice\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:- arXiv:2607.20434v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: As the parameter size of language models continues to grow, effective model compression is required to reduce their computational and memory overhead.\u003c/li\u003e\n\u003cli\u003eExisting compression methods suffer from bottleneck issues: when the compression ratio is increased, performance degrades significantly.\u003c/li\u003e\n\u003cli\u003eLow-rank decomposition and quantization are two important compression methods that have been proven to significantly reduce the computational and memory requirements of large language models (LLMs) while maintaining model accuracy.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20434v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: As the parameter size of language models continues to grow, effective model compression is required to reduce their computational and memory overhead\u003c/li\u003e\n\u003cli\u003eExisting compression methods suffer from bottleneck issues: when the compression ratio is increased, performance degrades significantly\u003c/li\u003e\n\u003cli\u003eLow-rank decomposition and quantization are two prominent compression methods that have been proven to significantly reduce the computational and memory require…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-cslg-b_introsearch\"\u003e\n  ArXiv cs.LG (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-cslg-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20465\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDataPrep-Bench: Benchmarking LLMs as Training Data Preparators\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:- arXiv:2607.20465v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), but there is no unified benchmark to measure how LLMs, agents, and data-centric workflows actually prepare training data end-to-end.\u003c/li\u003e\n\u003cli\u003eWe believe LLM-driven data preparation includes two complementary capabilities: data construction, which transforms raw sources into supervised training data, and data quality evaluation, which predicts the training value of candidate datasets before downstream training; throughout, \u0026ldquo;quality\u0026rdquo; refers to downstream training utility, not superficial text attributes.\u003c/li\u003e\n\u003cli\u003eWe introduce DataPrep-Bench, the first unified benchmark to jointly evaluate these two capabilities under a shared downstream foundation protocol across six domains and multiple foundation models.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20465v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe view LLM-driven data preparation as comprising two complementary capabilities: data construction, which transforms raw sources into supervised training data,…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe introduce DataPrep-Bench, the first unified benchmark that jointly evaluates both capabilities under a shared downstream-grounded protocol over six domains a…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20492\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003ePhantomFill: When the Form Demands an Answer, Language Models Invent One\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20492v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Language models in production do not write prose.\u003c/li\u003e\n\u003cli\u003eThey fill forms: JSON fields, function arguments, extraction templates.\u003c/li\u003e\n\u003cli\u003eWe show that the form itself causes hallucination.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20492v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Language models in production do not write prose\u003c/li\u003e\n\u003cli\u003eThey fill forms: JSON fields, function arguments, extraction templates\u003c/li\u003e\n\u003cli\u003eWe show that the form itself causes hallucination\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20512\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eThe Active Ingredient in Muon\u0026rsquo;s Grokking\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20512v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: The Muon optimizer reaches the grokking threshold on modular arithmetic faster than AdamW.\u003c/li\u003e\n\u003cli\u003ePrior work attributes this to \u0026ldquo;spectral-norm constraints plus orthogonalized momentum\u0026rdquo; but does not isolate which mechanism matters.\u003c/li\u003e\n\u003cli\u003eTo better understand Moun\u0026rsquo;s behavior, we run multi-seed and multi-learning-rate sweeps to decompose and stress-test the effect.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20512v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: The Muon optimizer reaches the grokking threshold on modular arithmetic faster than AdamW\u003c/li\u003e\n\u003cli\u003ePrior work attributes this to \u0026ldquo;spectral-norm constraints plus orthogonalized momentum\u0026rdquo; but does not isolate which mechanism matters\u003c/li\u003e\n\u003cli\u003eTo better understand Moun\u0026rsquo;s behavior, we run multi-seed and multi-learning-rate sweeps to decompose and stress-test the effect\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20516\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eScaling Closed-Loop Feature Channel Configuration with LLMs\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20516v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Initial results from closed-loop large language model-based channel configuration search have shown that neural network width can be directly optimized via executable code generation and accuracy feedback.\u003c/li\u003e\n\u003cli\u003eHowever, these results were obtained from a relatively sparse set of valid evaluations, and it remains unknown whether the observed optimization behavior transfers to a more dense sampling regime and if additional architectural regularities emerge when more generated networks are evaluated.\u003c/li\u003e\n\u003cli\u003eTo test this, the same search setup is extended to 250 candidate networks per fine-tuning cycle.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20516v1 Announce Type: new\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: Promising initial results in closed-loop large-language-model-based channel-configuration search demonstrated that neural-network widths can be optimi…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eHowever, those results were obtained from a relatively sparse set of valid evaluations, leaving open whether the observed optimization behavior transfers to a d…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eTo test this, the same search setting is scaled to 250 candidate networks per fine-tuning cycle\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20517\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eMultimodal CoLRAG-TF: Triple-Filtered Retrieval for Complex PDFs\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20517v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Retrieval-augmented generation (RAG) over heterogeneous PDF collections remains challenging due to multimodal content, domain-specific terminology, and the need for multi-hop reasoning across disparate evidence.\u003c/li\u003e\n\u003cli\u003eWe propose Multimodal CoLRAG-TF, a four-axis fusion architecture that integrates dense text embeddings, BM25 keyword matching, knowledge graph triple filtering, and image-based similarity for robust retrieval on complex documents.\u003c/li\u003e\n\u003cli\u003eOur system builds a multimodal index of 2,403 chunks extracted from 43 Japanese disaster lesson PDFs, supported by a hybrid OCR pipeline and LLM-based caption generation.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20517v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Retrieval-augmented generation (RAG) over heterogeneous PDF collections remains challenging due to multimodal content, domain-specific terminology, an…\u003c/li\u003e\n\u003cli\u003eWe present Multimodal CoLRAG-TF, a four-axis fusion architecture that integrates dense text embeddings, BM25 keyword matching, knowledge-graph triple filtering,…\u003c/li\u003e\n\u003cli\u003eOur system constructs a multimodal index of 2,403 blocks extracted from 43 Japanese disaster lesson PDFs, supported by a hybrid OCR pipeline and LLM-based capti…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20519\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eAdaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20519v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Looped Transformers increase test-time computation by repeatedly applying a shared recurrent block.\u003c/li\u003e\n\u003cli\u003eThe learned halting objective in a Looped Transformer typically uses a single exit distribution to both serve as an inference-time stopping rule and a training-time weighting over per-depth losses.\u003c/li\u003e\n\u003cli\u003eThis entangles exit selection with trajectory shaping: the gate not only chooses which recurrent state to use, but also determines the supervision strength on each intermediate state.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20519v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Looped Transformers increase test-time computation by repeatedly applying a shared recurrent block\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eLearned halting objectives in looped Transformers typically use a single exit distribution both as the inference-time stopping rule and as the training-time wei…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eThis entangles exit selection with trajectory formation: the gate not only chooses which recurrent state to use, but also determines how strongly each intermedi…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20521\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eGenerative Bayesian Filtering for State Estimation\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20521v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: The state of a dynamic system evolves over time, switching among several latent modes that govern its observable behavior.\u003c/li\u003e\n\u003cli\u003eFiltering methods infer the latent state from observations.\u003c/li\u003e\n\u003cli\u003eClassical filtering approaches, including Kalman filters, typically rely on simple observation models, such as linear-Gaussian models, that are incapable of characterizing the increasingly nonlinear and heterogeneous patterns in high-dimensional sensor signals.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20521v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: The state of a dynamic system evolves over time, switching among several latent modes that govern its observable behavior\u003c/li\u003e\n\u003cli\u003eFiltering methods infer the latent state from observations\u003c/li\u003e\n\u003cli\u003eClassical filtering approaches, including Kalman filters, typically rely on simple observation models, such as linear-Gaussian models, that are incapable of cha…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20522\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDo Active SAE Feature Planes Carry More Holonomy? A Preregistered Reversal in Gemma\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20522v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: This paper tests whether holonomy concentrates on active sparse-autoencoder (SAE) feature planes in Gemma 2 2B, a concrete operationalization of the broader semantic concentration prediction.\u003c/li\u003e\n\u003cli\u003eHolonomy is measured at the final-token layer-12 to layer-13 residual-stream readout by carrying a local frame around small loops using the instrument\u0026rsquo;s restricted Jacobian transport rule, then standardizing the resulting rotation by the enclosed area.\u003c/li\u003e\n\u003cli\u003eThe design, materiality threshold, analysis, and verdict rules were preregistered and frozen before the analysed measurements were inspected.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20522v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: This paper tests whether holonomy concentrates on active sparse-autoencoder (SAE) feature planes in Gemma 2 2B, a concrete operationalization of the b…\u003c/li\u003e\n\u003cli\u003eHolonomy is measured at the final-token layer-12 to layer-13 residual-stream readout by carrying a local frame around small loops using the instrument\u0026rsquo;s restric…\u003c/li\u003e\n\u003cli\u003eThe design, materiality threshold, analysis, and verdict rules were preregistered and frozen before the analysed measurements were inspected\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20529\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eUncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20529v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Large Language Model (LLM) ensembles are increasingly used to improve reliability by combining predictions from multiple LLMs.\u003c/li\u003e\n\u003cli\u003eHowever, existing aggregation methods typically assume that all models are equally trustworthy, overlooking differences in uncertainty quality.\u003c/li\u003e\n\u003cli\u003eThis assumption is poorly suited to heterogeneous LLMs, whose reliability and capability vary significantly, making naive aggregation vulnerable to unreliable or adversarial experts.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20529v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Large Language Model (LLM) ensembles are increasingly used to improve reliability by combining predictions from multiple LLMs\u003c/li\u003e\n\u003cli\u003eHowever, existing aggregation methods typically assume that all models are equally trustworthy, overlooking differences in uncertainty quality\u003c/li\u003e\n\u003cli\u003eThis assumption is poorly suited to heterogeneous LLMs, whose reliability and capability vary significantly, making naive aggregation vulnerable to unreliable o…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2607.20530\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eCLOE: Christoffel Loss Autoencoder for Anomaly Detection\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-07-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2607.20530v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Semi-supervised anomaly detection plays a key role in diverse fields such as process monitoring, healthcare, and finance.\u003c/li\u003e\n\u003cli\u003eHowever, lightweight methods often struggle with high-dimensional data and typically require careful tuning of multiple hyperparameters.\u003c/li\u003e\n\u003cli\u003eAmong existing approaches, Christoffel Function\u0026ndash;based methods are attractive due to their simplicity, requiring at most a single hyperparameter.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2607.20530v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Semi-supervised anomaly detection plays a key role in diverse fields such as process monitoring, healthcare, and finance\u003c/li\u003e\n\u003cli\u003eHowever, lightweight methods often struggle with high-dimensional data and typically require careful tuning of multiple hyperparameters\u003c/li\u003e\n\u003cli\u003eAmong existing approaches, Christoffel Function\u0026ndash;based methods are attractive due to their simplicity, requiring at most a single hyperparameter\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n",
  "wordCount": 7749,
  "readingTime": 37,
  "tableOfContents": "\u003cnav id=\"TableOfContents\"\u003e\n  \u003cul\u003e\n    \u003cli\u003e\u003ca href=\"#-in-depth-guide-to-this-issues-watch-list\"\u003e📖 In-depth Guide to This Issue\u0026rsquo;s Watch List\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-ai-hot-topics-on-x\"\u003e🌐 AI Hot Topics on X\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#topic-1-tech-giants-urge-us-to-protect-open-weight-ai-models\"\u003eTopic 1: Tech Giants Urge U.S. to Protect Open-Weight AI Models\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-2-uncle-bob-skips-ai-code-reviews-with-strict-testing-gauntlet\"\u003eTopic 2: Uncle Bob Skips AI Code Reviews with Strict Testing Gauntlet\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-3-anthropic-launches-claude-opus-5-with-frontier-performance-at-half-the-cost\"\u003eTopic 3: Anthropic Launches Claude Opus 5 with Frontier Performance at Half the Cost\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-4-xai-announces-grok-46-and-47-releases-weeks-apart\"\u003eTopic 4: xAI Announces Grok 4.6 and 4.7 Releases Weeks Apart\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-5-etched-raises-300m-at-103-billion-valuation-for-ai-inference-chips\"\u003eTopic 5: Etched Raises $300M at $10.3 Billion Valuation for AI Inference Chips\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-6-ai-community-awaits-opus-5-as-openai-rolls-out-voice-on-desktop\"\u003eTopic 6: AI Community Awaits Opus 5 as OpenAI Rolls Out Voice on Desktop\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-7-class-of-2027-prospects-land-first-division-i-offers\"\u003eTopic 7: Class of 2027 Prospects Land First Division I Offers\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-8-ai-clip-of-green-eyed-woman-at-blue-jays-game-divides-opinions-on-beauty\"\u003eTopic 8: AI Clip of Green-Eyed Woman at Blue Jays Game Divides Opinions on Beauty\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-influencer-insights\"\u003e💡 Influencer Insights\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#1-key-technology-trends-or-product-hotspots-followed-by-influencers-today\"\u003e1. Key Technology Trends or Product Hotspots Followed by Influencers Today\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#2-notable-unique-perspectives-or-industry-foresight\"\u003e2. Notable Unique Perspectives or Industry Foresight\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#3-recommended-tools-or-resources\"\u003e3. Recommended Tools or Resources\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-appendix-todays-watch-list-update-sources\"\u003e📚 Appendix: Today\u0026rsquo;s Watch List Update Sources\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#all-in-podcast-a_full\"\u003eAll-In Podcast (A_full)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#stratechery-by-ben-thompson-a_full\"\u003eStratechery by Ben Thompson (A_full)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-csai-b_introsearch\"\u003eArXiv cs.AI (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-cscl-b_introsearch\"\u003eArXiv cs.CL (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-cslg-b_introsearch\"\u003eArXiv cs.LG (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n  \u003c/ul\u003e\n\u003c/nav\u003e",
  "isDraft": false
}
