{
  "title": "2026-06-25 AI Daily | Agents enter infrastructure competition: From browser operations to inference chips",
  "url": "https://miaok.ong/en/ai-daily/ai-daily-2026-06-25/",
  "date": "2026-06-25T07:00:00+08:00",
  "lastmod": "2026-06-25T07:00:00+08:00",
  "type": "ai-daily",
  "kind": "page",
  "language": "en",
  "description": "Today\u0026rsquo;s main theme is the transition of intelligent agents from concept to systematic implementation: Gemini is delegating computer use to lightweight models, while evaluation and security frameworks are simultaneously gaining traction. OpenAI\u0026rsquo;s partnership with Broadcom on inference chips shows that model competition is extending to the silicon level. On-device models, Coding Agents, and organizational collaboration tools are also rapidly maturing.",
  "keywords": null,
  "tags": [],
  "categories": [],
  "author": "Mark (Miao) Kong",
  "image": "https://miaok.ong/images/avatar.jpg",
  "content": "\u003ch1 id=\"2026-06-25-ai-daily--agents-enter-infrastructure-competition-from-browser-operations-to-inference-chips\"\u003e\n  2026-06-25 AI Daily | Agents Enter Infrastructure Competition: From Browser Operations to Inference Chips\n  \u003ca class=\"heading-link\" href=\"#2026-06-25-ai-daily--agents-enter-infrastructure-competition-from-browser-operations-to-inference-chips\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h1\u003e\n\u003cblockquote\u003e\n\u003cp\u003eToday\u0026rsquo;s main theme is the transition of agents from concept to systematic implementation: Gemini is delegating \u0026ldquo;computer use\u0026rdquo; capabilities to lightweight models, while evaluation and security frameworks are gaining traction. The OpenAI and Broadcom inference chip shows that model competition is reaching the silicon level. Edge models, Coding Agents, and organizational collaboration tools are also maturing rapidly.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch2 id=\"-in-depth-guide-to-this-issues-watch-list\"\u003e\n  📖 In-depth Guide to This Issue\u0026rsquo;s Watch List\n  \u003ca class=\"heading-link\" href=\"#-in-depth-guide-to-this-issues-watch-list\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003eThe most noteworthy theme today is \u0026ldquo;agents moving from concept to infrastructure.\u0026rdquo; Mirendil\u0026rsquo;s interview on self-accelerating AI is a must-read for those interested in the organizational structures of cutting-edge research. Google DeepMind\u0026rsquo;s introduction of \u0026ldquo;computer use\u0026rdquo; into Gemini 3.5 Flash indicates that browser/desktop operation capabilities are being delegated to more lightweight models. Correspondingly, RIFT-Bench, AgenticInterpBench, and several papers on agent boundaries, security, and interpretability remind us: the faster agent capabilities expand, the more systematic evaluation, red-teaming, and responsibility boundaries are needed.\u003c/p\u003e\n\u003cp\u003eThe second theme is \u0026ldquo;the continued acceleration of the full-stack AI.\u0026rdquo; OpenAI and Broadcom\u0026rsquo;s launch of a chip for LLM inference, with an emphasis on a nine-month design-to-production cycle, deserves close attention from infrastructure teams: model companies are extending the competition to the silicon level, data centers, and energy efficiency.\u003c/p\u003e\n\u003cp\u003eFinally, inference and reinforcement learning remain research hotspots. From SGPO\u0026rsquo;s policy distillation to the long-term alignment of beneficial models, multi-objective recommendations, and safe multi-agent RL, today\u0026rsquo;s papers all point to one issue: AI must not only complete tasks but also act in a stable, interpretable, and generalizable manner under complex constraints.\u003c/p\u003e\n\u003ch2 id=\"-x-platform-ai-hot-topics-quick-brief\"\u003e\n  🌐 X Platform AI Hot Topics Quick Brief\n  \u003ca class=\"heading-link\" href=\"#-x-platform-ai-hot-topics-quick-brief\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003ch3 id=\"topic-1-loop-engineering-emerges-as-new-way-to-automate-ai-coding-agents\"\u003e\n  Topic 1: Loop Engineering Emerges as New Way to Automate AI Coding Agents\n  \u003ca class=\"heading-link\" href=\"#topic-1-loop-engineering-emerges-as-new-way-to-automate-ai-coding-agents\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · Other\u003c/li\u003e\n\u003cli\u003eOverview: Trending 13 hours ago, Related posts: 3600\u003c/li\u003e\n\u003cli\u003eWhat it is: A new idea called \u0026ldquo;Loop Engineering\u0026rdquo; is trending on X, which involves automating the workflow of AI coding agents through a cyclic mechanism of feedback, evaluation, and correction.\u003c/li\u003e\n\u003cli\u003eWhy it\u0026rsquo;s important: It could significantly improve the stability, controllability, and task completion rate of AI coding agents, driving the shift from \u0026ldquo;single-shot generation\u0026rdquo; to \u0026ldquo;sustainably executing\u0026rdquo; engineered agents.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: Current discussions focus on whether this loop framework can genuinely improve performance on complex coding tasks, its advantages over traditional agent workflows, and whether it will become a standard paradigm for future AI programming automation.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-2-developer-migrates-36k-ai-agents-from-openclaw-to-hermes-for-better-reliability\"\u003e\n  Topic 2: Developer Migrates $36K AI Agents from OpenClaw to Hermes for Better Reliability\n  \u003ca class=\"heading-link\" href=\"#topic-2-developer-migrates-36k-ai-agents-from-openclaw-to-hermes-for-better-reliability\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending 11 hours ago, Related posts: 83\u003c/li\u003e\n\u003cli\u003eAbstract: Developer Migrates $36K AI Agents from OpenClaw to Hermes for Better Reliability:\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-3-anthropic-accuses-alibaba-linked-group-of-massive-claude-ai-distillation-attack\"\u003e\n  Topic 3: Anthropic Accuses Alibaba-Linked Group of Massive Claude AI Distillation Attack\n  \u003ca class=\"heading-link\" href=\"#topic-3-anthropic-accuses-alibaba-linked-group-of-massive-claude-ai-distillation-attack\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending 7 hours ago, Related posts: 4700\u003c/li\u003e\n\u003cli\u003eWhat it is: Anthropic has accused an organization linked to Alibaba of using Claude outputs on a large scale for model distillation to replicate or enhance their own AI system\u0026rsquo;s capabilities.\u003c/li\u003e\n\u003cli\u003eWhy it\u0026rsquo;s important: This incident highlights the risks related to intellectual property, model security, and API abuse for frontier large models. It may also intensify regulatory pressure on AI companies regarding training data, access control, and cross-border technology competition.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: Discussions on X are focused on whether Anthropic\u0026rsquo;s evidence is sufficient, whether model distillation should be considered a violation or a common industry practice, the potential for politicized interpretations in the context of US-China AI competition, and how platforms should balance open access with preventing capability replication.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-4-chinas-glm-52-tops-coding-benchmarks-at-fraction-of-us-ai-costs\"\u003e\n  Topic 4: China\u0026rsquo;s GLM-5.2 Tops Coding Benchmarks at Fraction of U.S. AI Costs\n  \u003ca class=\"heading-link\" href=\"#topic-4-chinas-glm-52-tops-coding-benchmarks-at-fraction-of-us-ai-costs\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending 9 hours ago, Related posts: 2500\u003c/li\u003e\n\u003cli\u003eWhat it is: China\u0026rsquo;s GLM-5.2 has been reported to lead in code benchmarks, with its training or usage costs being only a fraction of comparable U.S. AI models.\u003c/li\u003e\n\u003cli\u003eWhy it\u0026rsquo;s important: This suggests that in the AI field, especially for programming models, performance advantages do not necessarily depend solely on greater computing power and higher costs. It could change the industry\u0026rsquo;s perspective on model efficiency, R\u0026amp;D investment, and the competitive landscape.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: The discussion on X centers on whether its benchmark scores are representative of real-world programming capabilities, whether its cost advantage is replicable, and if this signals a shift in the US-China AI competition from a \u0026ldquo;compute power race\u0026rdquo; to a \u0026ldquo;battle of efficiency\u0026rdquo; and engineering capabilities.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-5-users-build-ai-powered-second-brains-with-claude-and-obsidian\"\u003e\n  Topic 5: Users Build AI-Powered Second Brains with Claude and Obsidian\n  \u003ca class=\"heading-link\" href=\"#topic-5-users-build-ai-powered-second-brains-with-claude-and-obsidian\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending 8 hours ago, 144 related posts\u003c/li\u003e\n\u003cli\u003eWhat happened: Discussions on X regarding tutorials, plugins, and workflows for building \u0026ldquo;AI Second Brains\u0026rdquo; with Claude and Obsidian are heating up. Users are sharing methods to organize, associate, and retrieve personal knowledge bases using AI, all without coding.\u003c/li\u003e\n\u003cli\u003eWhy it\u0026rsquo;s important: This reflects a shift in large language models from being mere Q\u0026amp;A tools to becoming personal knowledge management and productivity infrastructure, enabling non-technical users to build sustainable, accumulating intelligent workflows with a low barrier to entry.\u003c/li\u003e\n\u003cli\u003eDiscussion overview: Discussions focus on how to structure Obsidian notes for AI pattern recognition, the practicality of Claude plugins and GitHub projects, and whether such tools truly enhance thinking ability or merely create new efficiency illusions and information organization burdens.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch4 id=\"todays-ai-public-opinion-summary-on-x\"\u003e\n  Today\u0026rsquo;s AI Public Opinion Summary on X\n  \u003ca class=\"heading-link\" href=\"#todays-ai-public-opinion-summary-on-x\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h4\u003e\n\u003cp\u003eToday\u0026rsquo;s main thread of AI discussion on X has shifted from \u0026ldquo;model capabilities themselves\u0026rdquo; to \u0026ldquo;how to transform models into sustainable, controllable, and deployable productivity systems\u0026rdquo;: on one side, cyclical feedback frameworks like Loop Engineering emphasize moving coding agents from one-time generation to repeated evaluation and correction; on the other, the Claude + Obsidian \u0026ldquo;Second Brain\u0026rdquo; workflow pushes AI deeper into personal knowledge management. The consensus is that everyone acknowledges AI is evolving from a chat tool into engineered and daily infrastructure, with significant potential especially in programming, retrieval, and organization tasks. However, there is disagreement on whether these solutions truly enhance stability and thinking quality, or primarily create an illusion of efficiency. Meanwhile, debates surrounding Anthropic\u0026rsquo;s accusations of distillation and GLM-5.2\u0026rsquo;s low-cost, high-performance capabilities also draw focus back to the underlying logic of model competition: whether success comes from computing power, data, and closed access barriers, or from efficiency, engineering, and workflow innovation. Potential risks are concentrated in three areas: intellectual property and security issues arising from model capabilities being replicated or misused, misjudgments caused by a disconnect between benchmark scores and real-world scenarios, and the potential for over-reliance on AI tools to amplify information organization costs and cognitive outsourcing.\u003c/p\u003e\n\u003ch2 id=\"-influencer-insights\"\u003e\n  💡 Influencer Insights\n  \u003ca class=\"heading-link\" href=\"#-influencer-insights\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003eBased on a summary of activity from several AI influencers on the X platform over the past 24 hours, here is a core observation and in-depth analysis report.\u003c/p\u003e\n\u003chr\u003e\n\u003ch1 id=\"-ai-industry-daily-briefing-on-device-models-rise-coding-agent-ecosystem--regulatory-storm\"\u003e\n  📊 AI Industry Daily Briefing: On-device Models Rise, Coding Agent Ecosystem \u0026amp; Regulatory Storm\n  \u003ca class=\"heading-link\" href=\"#-ai-industry-daily-briefing-on-device-models-rise-coding-agent-ecosystem--regulatory-storm\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h1\u003e\n\u003ch2 id=\"1--todays-focus-technology-trends-and-product-hotspots\"\u003e\n  1. 🔥 Today\u0026rsquo;s Focus: Technology Trends and Product Hotspots\n  \u003ca class=\"heading-link\" href=\"#1--todays-focus-technology-trends-and-product-hotspots\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003ch3 id=\"on-device-models-enter-practical-stage\"\u003e\n  On-device Models Enter Practical Stage\n  \u003ca class=\"heading-link\" href=\"#on-device-models-enter-practical-stage\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cp\u003eLocal deployment is no longer just a toy for geeks but is becoming a mainstream choice for high-productivity scenarios.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eQwen Series firmly holds \u0026ldquo;sweet spot\u0026rdquo;\u003c/strong\u003e: After \u0026ldquo;ascetic\u0026rdquo; testing, @zhixianio stated that running \u003cstrong\u003eQwen3.6-35B-A3B (oMLX)\u003c/strong\u003e on an M5 Max offers an excellent experience, with faster response times than remote APIs and \u0026ldquo;smart\u0026rdquo; performance in multimodal personal assistant (OpenClaw) and programming (PI-Mono) scenarios. He even believes on-device performance has surpassed DSV4 Pro.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eAstonishing potential and ceiling of small models\u003c/strong\u003e: @zhixianio conducted brutal code generation tests on \u003cstrong\u003eGemma 4 12B\u003c/strong\u003e and its code-tuned version. The results show that while 12B models perform well in handling daily scripts, they frequently encounter logical failures when generating complex stateful programs like \u0026ldquo;Tetris,\u0026rdquo; constrained by their size. This proves that \u003cstrong\u003e12B is the current \u0026ldquo;entry-level programming model,\u0026rdquo; but cannot replace the complex engineering capabilities of 35B+ models\u003c/strong\u003e.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eGoogle\u0026rsquo;s on-device ambition\u003c/strong\u003e: @zhixianio noted Google\u0026rsquo;s release of the \u003cstrong\u003eQAT (Quantization-Aware Training) model\u003c/strong\u003e, considering it a critical step for Google to pave the way for Android on-device AI. By optimizing quantization during training, Google aims to make AI run seamlessly on mobile phones.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"coding-agent-ecosystem-from-usable-to-ultimate-cost-and-experience\"\u003e\n  Coding Agent Ecosystem: From \u0026ldquo;Usable\u0026rdquo; to \u0026ldquo;Ultimate Cost and Experience\u0026rdquo;\n  \u003ca class=\"heading-link\" href=\"#coding-agent-ecosystem-from-usable-to-ultimate-cost-and-experience\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cp\u003eThe discussion around AI programming has shifted from \u0026ldquo;can it write code\u0026rdquo; to \u003cstrong\u003ecost control, Skill workflow management, and cross-model orchestration\u003c/strong\u003e.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eCodex usage anxiety\u003c/strong\u003e: Multiple bloggers (@Pluvio9yte, @msjiaozhu) reported a sharp increase in \u003cstrong\u003eCodex\u0026rsquo;s Token consumption speed\u003c/strong\u003e, with the same tasks consuming several times more than before. @Pluvio9yte, a $200 subscriber, expressed strong uneasiness. This hints that OpenAI might be adjusting backend billing weights or model inference consumption.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eDomestic models\u0026rsquo; coding surprise attack\u003c/strong\u003e: @Pluvio9yte and @vista8 strongly recommended ByteDance\u0026rsquo;s \u003cstrong\u003eDoubao-Seed-2.1-Pro\u003c/strong\u003e model. Tests showed significant progress in its frontend and coding capabilities, and @Pluvio9yte provided a specific tutorial for integrating this model into Claude Code. \u003cstrong\u003eSwitching between multiple models and finding the most cost-effective API has become an essential skill for developers\u003c/strong\u003e.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eFable\u0026rsquo;s Disruptive Productivity\u003c/strong\u003e: @zhixianio shared his astounding experience with \u003cstrong\u003eFable\u003c/strong\u003e, an agent that completed 70% of a demo project in just 40 minutes. It even identified flaws in the human-led design and optimized it autonomously. This marks an evolution in AI programming, shifting from a mere \u0026ldquo;executor\u0026rdquo; to an \u0026ldquo;architecture reviewer.\u0026rdquo;\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"deeper-integration-of-ai-and-office-software\"\u003e\n  Deeper Integration of AI and Office Software\n  \u003ca class=\"heading-link\" href=\"#deeper-integration-of-ai-and-office-software\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eThe Launch of Claude Tag\u003c/strong\u003e: @dotey provided a detailed overview of \u003cstrong\u003eClaude Tag\u003c/strong\u003e, launched by Anthropic. It\u0026rsquo;s not just a simple chatbot, but a persistent \u0026ldquo;colleague\u0026rdquo; within Slack, featuring contextual memory, proactive notifications (Ambient mode), and cross-thread collaboration. @dotey noted that 65% of Anthropic\u0026rsquo;s internal code is already generated by it, signaling that agents are now entering the complex domain of organizational collaboration.\u003c/li\u003e\n\u003c/ul\u003e\n\u003chr\u003e\n\u003ch2 id=\"2--unique-perspectives-and-industry-foresight\"\u003e\n  2. 🧠 Unique Perspectives and Industry Foresight\n  \u003ca class=\"heading-link\" href=\"#2--unique-perspectives-and-industry-foresight\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eThe Future of \u0026ldquo;Causal Large Models\u0026rdquo;\u003c/strong\u003e (from @Pluvio9yte, quoting @huang_biwei):\nToday\u0026rsquo;s LLMs operate on \u0026ldquo;correlation prediction\u0026rdquo; and cannot grasp physical causality, like understanding that \u0026ldquo;pouring water into a cup with a hole will cause it to leak.\u0026rdquo; \u003cstrong\u003eAether AI\u0026rsquo;s Causal World Models\u003c/strong\u003e aim to shift AI from data fitting to understanding underlying mechanisms. In fields like embodied intelligence and scientific R\u0026amp;D, this is seen as an alternative path to breaking the bottlenecks of Scaling Laws.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eToken Heterogeneity and New Factors of Production\u003c/strong\u003e (@lijigang \u0026amp; @vista8):\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e@lijigang offered a profound insight: \u003cstrong\u003eElectricity is homogeneous, but Tokens are heterogeneous\u003c/strong\u003e (the value of Tokens from different models varies). This disparity can easily create \u0026ldquo;tax-like\u0026rdquo; supply bottlenecks rather than becoming a simple commoditized infrastructure.\u003c/li\u003e\n\u003cli\u003e@vista8 argued that \u003cstrong\u003ewhile Agents are becoming a nearly-free digital labor force, human \u0026ldquo;attention, trust, and branding\u0026rdquo; will not depreciate; instead, they become more valuable due to information overload\u003c/strong\u003e. Personal media will become the new era\u0026rsquo;s \u0026ldquo;trusted context entry point.\u0026rdquo;\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eBeware the \u0026ldquo;Token Trap\u0026rdquo; and Embrace an Automation Mindset\u003c/strong\u003e (@gefei55 \u0026amp; @Pluvio9yte):\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e@gefei55 cautioned that while Tokens may seem infinite, human energy is finite. Don\u0026rsquo;t let Vibe Coding turn into a life-draining \u0026ldquo;Token trap\u0026rdquo;; prioritize what\u0026rsquo;s important.\u003c/li\u003e\n\u003cli\u003e@Pluvio9yte summarized a core principle for the AI era: \u003cstrong\u003eIf you\u0026rsquo;ve done something manually three times, you must think about how to automate it\u003c/strong\u003e.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eOn Time Gaps and Information Arbitrage\u003c/strong\u003e (@gefei55): @gefei55 shared an SEO strategy that uses AI to monitor viral links on Twitter to discover new domain names. By mining social media signals (likes, links) before a topic trends on Google, one can achieve \u003cstrong\u003e\u0026ldquo;time-gap arbitrage.\u0026rdquo;\u003c/strong\u003e This demonstrates that business opportunities in the AI era lie in preemptive insight.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eReturning to the Essence of Thinking\u003c/strong\u003e (@lijigang):\nIn an age of tool proliferation, @lijigang repeatedly emphasized that \u0026ldquo;taste is a person\u0026rsquo;s loss function\u0026rdquo; and recommended \u0026ldquo;thematic reading\u0026rdquo; to gain disciplinary perspectives. Tools are merely means; \u003cstrong\u003ethe real defensible moat is a person\u0026rsquo;s ability to define problems and formulate decision functions\u003c/strong\u003e.\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003chr\u003e\n\u003ch2 id=\"3--recommended-tools-and-resources\"\u003e\n  3. 🛠️ Recommended Tools and Resources\n  \u003ca class=\"heading-link\" href=\"#3--recommended-tools-and-resources\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003e\u003cstrong\u003eProgramming and Development Agents\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eDoubao-Seed-2.1-Pro\u003c/strong\u003e: Recommended by @Pluvio9yte, this domestic model excels in coding and front-end capabilities and can be integrated with various Agents via the Volcano Engine API.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eFable\u003c/strong\u003e: A widely discussed AI programming tool capable of handling complex demo development and proposing solutions superior to human ones.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eTencent WeKnora\u003c/strong\u003e: Recommended by @Pluvio9yte, an open-source enterprise knowledge platform that integrates RAG, Agent, and automatic knowledge graph generation.\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e\u003cstrong\u003eSkill and Workflow Management\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eSymbolic Link-Based Skill Management\u003c/strong\u003e: @dotey shared a geek-style approach to Skill management, using symbolic links to unify configuration files across multiple projects and Agents, combined with Git for efficient iteration.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eInterview Analysis Skill Suite\u003c/strong\u003e: @dotey open-sourced two Skills, \u003ccode\u003einterview-analysis\u003c/code\u003e and \u003ccode\u003einterview-writing\u003c/code\u003e, designed to transform long podcasts into high-quality articles.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003ePort Management Tool\u003c/strong\u003e: @vista8 recommended a free macOS menu bar utility for intuitively checking local port usage—a great assistant for Vibe Coding.\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e\u003cstrong\u003eEfficiency and Data\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eObsidian-WeChat Sync Assistant\u003c/strong\u003e: Shared by @AI_Jasonyu, this tool allows one-click synchronization of WeChat articles and valuable content from paid communities to Obsidian for building a personal knowledge base.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eCapWords\u003c/strong\u003e: Recommended by @nishuang, a foreign language learning app that uses AI to recognize real-world objects for gamified learning, setting a benchmark for the integration of design and AI.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eGoogle Workspace CLI\u003c/strong\u003e: Although the author\u0026rsquo;s dismissal caused controversy (reported by @dotey), the tool remains powerful, allowing you to manage Gmail and Drive from the command line, and it comes with a built-in MCP service.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch2 id=\"-appendix-todays-watch-list-update-source-list\"\u003e\n  📚 Appendix: Today\u0026rsquo;s Watch List Update Source List\n  \u003ca class=\"heading-link\" href=\"#-appendix-todays-watch-list-update-source-list\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cblockquote\u003e\n\u003cp\u003eTime window: Last 3 days; covers 22 sources; 34 updates in total.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch3 id=\"a16z-podcast-a_full\"\u003e\n  a16z Podcast (A_full)\n  \u003ca class=\"heading-link\" href=\"#a16z-podcast-a_full\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003e\u003ca href=\"https://ai-a16z.simplecast.com/episodes/building-self-accelerating-ai-with-mirendil-rt3KPYBk\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eBuilding Self-Accelerating AI with Mirendil\u003c/a\u003e\u003c/strong\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-25 03:12 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - Matt Bornstein speaks with Mirendil co-founders Behnam Neyshabur and Harsh Mehta about their vision for building self-accelerating artificial intelligence.\n\u003cul\u003e\n\u003cli\u003eAfter leading research efforts at Google and Anthropic, the founders started Mirendil around a simple question: what happens when AI systems can meaningfully contribute to their own development?\u003c/li\u003e\n\u003cli\u003eThey believe that the most important application may be accelerating scientific and technological progress itself, rather than just focusing on AI as a productivity tool.\u003c/li\u003e\n\u003cli\u003eThe conversation explores AI research, scaling laws, automated engineering, scientific discovery, and the challenges of building systems that can improve over time.\u003c/li\u003e\n\u003cli\u003eThey discuss the future of AI-assisted research, why they believe scientific progress is still bottlenecked by intelligence, and how more powerful AI systems can help unlock advancements in medicine, engineering, and the natural sciences.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003eMatt Bornstein speaks with Mirendil cofounders Behnam Neyshabur and Harsh Mehta about their vision for building self-accelerating AI\u003c/li\u003e\n\u003cli\u003eAfter leading research efforts at Google and Anthropic, the founders started Mirendil around a simple question: what happens when AI systems can meaningfully co…\u003c/li\u003e\n\u003cli\u003eRather than focusing solely on AI as a tool for productivity, they argue that the most important application may be accelerating scientific and technological pr…\u003c/li\u003e\n\u003cli\u003eThe conversation explores AI research, scaling laws, automated engineering, scientific discovery, and the challenges of building systems that can improve over t…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"stratechery-by-ben-thompson-a_full\"\u003e\n  Stratechery by Ben Thompson (A_full)\n  \u003ca class=\"heading-link\" href=\"#stratechery-by-ben-thompson-a_full\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003e\u003ca href=\"https://stratechery.com/2026/my-vibe-coding-adventure-the-app-and-the-experience-ten-takeaways/\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eMy Vibe Coding Adventure, The App and the Experience, Ten Takeaways\u003c/a\u003e\u003c/strong\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 18:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - My experience and reflections on vibe coding an app that I plan on actually using regularly.\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003e$15\u003c/strong\u003e/month* or *\u003cstrong\u003e$150\u003c/strong\u003e/year.\u003c/li\u003e\n\u003cli\u003eSubstantive analysis of the day\u0026rsquo;s news via three weekly emails or a podcast.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eStrategy Interviews\u003c/strong\u003e.\u003c/li\u003e\n\u003cli\u003eInterviews with leading public company CEOs, private company founders, and discussions with fellow analysts.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003eMy experience and reflections on vibe coding an app that I plan on actually using regularly.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"openai-blog-a_full\"\u003e\n  OpenAI Blog (A_full)\n  \u003ca class=\"heading-link\" href=\"#openai-blog-a_full\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003e\u003ca href=\"https://openai.com/index/openai-broadcom-jalapeno-inference-chip\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eOpenAI and Broadcom unveil LLM-optimized inference chip\u003c/a\u003e\u003c/strong\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 14:00 Beijing Time\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eAbstract:\n\u003cul\u003e\n\u003cli\u003eEarly tests show that the first generation accelerator\u0026rsquo;s performance per watt will significantly outperform current state-of-the-art technology.\u003c/li\u003e\n\u003cli\u003eBuilt from the ground up for current and future LLMs across the industry.\u003c/li\u003e\n\u003cli\u003eThe development process from design to production took only nine months, with OpenAI\u0026rsquo;s models accelerating the process.\u003c/li\u003e\n\u003cli\u003eExpands OpenAI\u0026rsquo;s full-stack platform, from product to model, and now to chip.\u003c/li\u003e\n\u003cli\u003eWill be deployed in gigawatt-scale, multi-generational deployments with data center partners.\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003eOpenAI and Broadcom introduce Jalapeño, a custom AI chip built for LLM inference to improve performance, efficiency, and scale across AI systems.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"google-deepmind-blog-a_full\"\u003e\n  Google DeepMind Blog (A_full)\n  \u003ca class=\"heading-link\" href=\"#google-deepmind-blog-a_full\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003e\u003ca href=\"https://deepmind.google/blog/introducing-computer-use-in-gemini-3-5-flash/\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eIntroducing computer use in Gemini 3.5 Flash\u003c/a\u003e\u003c/strong\u003e\n\u003cul\u003e\n\u003cli\u003ePublished Time:2026-06-25 00:30 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:\n\u003cul\u003e\n\u003cli\u003eIntroducing computer use in Gemini 3.5 Flash.\u003c/li\u003e\n\u003cli\u003eThis piece from Google DeepMind Blog explains how Introducing computer use in Gemini 3.5 Flash shapes the broader AI and infrastructure landscape.\u003c/li\u003e\n\u003cli\u003eAfter introducing computer use in Gemini 3.5 Flash, it also brings practical implications for founders, operators, and investors.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003eIntroducing computer use in Gemini 3.5 Flash\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-csai-b_introsearch\"\u003e\n  ArXiv cs.AI (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-csai-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23927\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eRIFT-Bench: Dynamic Red-teaming For Agentic AI Systems\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished Time:2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23927v1 Announce Type: new.\u003c/li\u003e\n\u003cli\u003eAbstract: Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyond traditional LLM vulnerabilities.\u003c/li\u003e\n\u003cli\u003eExisting security evaluations are often tied to specific implementations or domains, limiting unified comparison across heterogeneous systems.\u003c/li\u003e\n\u003cli\u003eTo address this gap, we introduce RIFT-Bench, a graph representation-driven methodology for dynamic red-teaming that enables unified evaluations across diverse agent architectures.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23927v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyon…\u003c/li\u003e\n\u003cli\u003eExisting security evaluations are often tied to specific implementations or domains, limiting unified comparison across heterogeneous systems\u003c/li\u003e\n\u003cli\u003eTo address this gap, we introduce RIFT-Bench, a graph representation-driven methodology for dynamic red-teaming that enables unified evaluations across diverse…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23938\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eNeuro-Symbolic Drive: Rule-Grounded Faithful Reasoning for Driving VLAs\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished Time:2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23938v1 Announce Type: new.\u003c/li\u003e\n\u003cli\u003eAbstract: Driven VLA models that incorporate Chain-of-Thought (CoT) reasoning are appealing because they leverage pre-trained VLM representations and expose intermediate decisions in natural language, but current rationales often lack the step-by-step decision semantics required to maintain causal relationships between rationales and planned movements.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe introduce Neuro-Symbolic Drive, a neuro-symbolic driving framework that supervises a driving VLA with rule-grounded reasoning traces extracted directly from classic rule-based planners.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eOur key observation is that rule-based planners are symbolic AI systems that already function as executable reasoning engines: they reason about active safety constraints, search for candidate actions, and select the final trajectory.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eEN Highlights:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23938v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Driving VLA models incorporating Chain-of-Thought (CoT) reasoning are attractive because they leverage pretrained VLM representations and expose inter…\u003c/li\u003e\n\u003cli\u003eWe introduce Neuro-Symbolic Drive, a neuro-symbolic driving framework that supervises a driving VLA with rule-grounded reasoning traces extracted directly from…\u003c/li\u003e\n\u003cli\u003eOur key observation is that rule-based planners are symbolic AI systems that already function as executable reasoning engines: they reason about active safety c…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23991\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eCritique of Agent Model\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eRelease Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23991v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eWith the rise of Large Language Model (LLM) systems marketed as \u0026ldquo;coding agents,\u0026rdquo; \u0026ldquo;AI co-scientists,\u0026rdquo; and other \u0026ldquo;agentic\u0026rdquo; tools that promise to increase productivity, and with \u0026ldquo;existential\u0026rdquo; concerns about, for instance, AI having destructive power under a speculative \u0026ldquo;machine agency\u0026rdquo; directed against humanity and escaping human control, it becomes critical to clarify where automation ends and agency begins, both for building capable systems and for understanding whether and what to fear.\u003c/li\u003e\n\u003cli\u003eDrawing on the Cartesian foundation of independent thought and depictions of autonomous existence in science fiction, we survey the state of AI agents and analyze agent architectures across five dimensions: goals, identity, decision-making, self-regulation, and learning.\u003c/li\u003e\n\u003cli\u003eSpecifically, we argue that true agency requires these constructs to be \\emph{internalized within the system itself} rather than assembled through external scaffolding.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23991v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: What is an agent\u003c/li\u003e\n\u003cli\u003eWhat constitutes agency\u003c/li\u003e\n\u003cli\u003eWith the rise of Large Language Model (LLM) systems marketed as \u003ccode\u003ecoding agents'', \u003c/code\u003eAI co-scientists\u0026rsquo;\u0026rsquo;, and other ``agentic\u0026quot; tools that promise to drive up pro…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.24010\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eSafe and Generalizable Hierarchical Multi-Agent RL via Constraint Manifold Control\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eRelease Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.24010v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Multi-agent systems are widely used in safety-critical applications that require coordinated behavior under strict safety constraints.\u003c/li\u003e\n\u003cli\u003eExisting approaches face a fundamental trade-off: learning-based methods achieve strong empirical performance but lack theoretical safety guarantees, while control-theoretic methods enhance safety but often lead to overly conservative and inefficient behavior.\u003c/li\u003e\n\u003cli\u003eWe propose a hierarchical multi-agent reinforcement learning framework that enforces hard safety constraints under mild assumptions at a low level via a constraint manifold, while enabling effective coordination through high-level policy learning.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.24010v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Multi-agent systems are widely used in safety-critical applications that require coordinated behavior under strict safety constraints\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eExisting approaches face a fundamental trade-off: learning-based methods achieve strong empirical performance but lack theoretical safety guarantees, while cont…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe propose a hierarchical multi-agent reinforcement learning framework that enforces hard safety constraints under mild assumptions at low level via a constrain…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.24014\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eReinforcement Learning Towards Broadly and Persistently Beneficial Models\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.24014v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: As AI systems are deployed in increasingly diverse and high-risk environments, model alignment must generalize beyond the tasks and domains seen during training.\u003c/li\u003e\n\u003cli\u003eThis is especially important for Reinforcement Learning (RL), as RL can introduce unexpected misalignment through reward hacking, deception, or other unintended strategies.\u003c/li\u003e\n\u003cli\u003eWe investigate whether reinforcement learning of beneficial behaviors, instantiated in realistic domains, can produce broad and persistent alignment generalization beyond the training distribution.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.24014v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domains seen dur…\u003c/li\u003e\n\u003cli\u003eThis is especially important for reinforcement learning (RL), which can introduce unexpected misalignment through reward hacking, deception, or other unintended…\u003c/li\u003e\n\u003cli\u003eWe study whether RL on beneficial behavior, instantiated in realistic domains, can produce broad and persistent alignment generalization beyond the training dis…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.24026\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eCan Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.24026v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining the role of localized components remains labor-intensive and difficult to standardize.\u003c/li\u003e\n\u003cli\u003eIn this work, we investigate whether Language Model (LM) agents can help solve this explanation problem once a circuit has already been identified.\u003c/li\u003e\n\u003cli\u003eWe introduce AgenticInterpBench, a circuit explanation benchmark constructed from 84 semi-synthetic transformer circuits and 163 component-level annotations.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.24026v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining what localized components do remains l…\u003c/li\u003e\n\u003cli\u003eIn this work, we study whether language model (LM) agents can assist with this explanation problem once a circuit has already been identified\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe introduce AgenticInterpBench, a benchmark for circuit explanation built from 84 semi-synthetic transformer circuits with 163 component-level annotations\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.24042\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eBreaking the Filter Bubble: A Semantic Pareto-DQN Framework for Multi-Objective Recommendation\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.24042v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Recommender systems often induce filter bubbles and semantic homogenization by monolithically optimizing for immediate user engagement.\u003c/li\u003e\n\u003cli\u003eStandard single-objective models, including traditional Deep Q-Networks, are ill-equipped to navigate the trade-offs between platform retention and critical social values such as information diversity and provider fairness.\u003c/li\u003e\n\u003cli\u003eTo address these limitations, we introduce a multi-objective reinforcement learning framework that formalizes recommendation as a semantic multi-objective Markov Decision Process.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.24047\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eEnsemble Feature Selection and Harris Hawks Optimization for Explainable Mental Health Risk Prediction in Female Sex Workers\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.24047v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: One of the significant mental health issues affecting female sex workers (FSWs) is mental disorders, especially depression.\u003c/li\u003e\n\u003cli\u003eExposure to violence, stigma, and economic hardship further increases their psychological risk.\u003c/li\u003e\n\u003cli\u003eCurrent machine learning (ML) models are typically ineffective at capturing the high-dimensional and complex risk patterns that exist in this marginalized group.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.24064\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eBeyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003ePublished: 2026-06-24 12:00 Beijing Time\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eAbstract: - arXiv:2606.24064v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectively transferring what to answer instead of how to reason.\u003c/li\u003e\n\u003cli\u003eThis trajectory-level imitation encourages memorization of instance-specific steps rather than the acquisition of transferable problem-solving skills, thus limiting generalization to new problems.\u003c/li\u003e\n\u003cli\u003eWe propose Strategy-Guided Policy Optimization (SGPO), which replaces instance-level trajectory imitation with reusable strategy distillation.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.24064v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectively transfe…\u003c/li\u003e\n\u003cli\u003eThis trajectory-level imitation encourages memorization of instance-specific steps rather than acquisition of transferable problem-solving skills, limiting gene…\u003c/li\u003e\n\u003cli\u003eWe propose Strategy-Guided Policy Optimization (SGPO), which replaces instance-level trajectory imitation with reusable strategy distillation\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.24099\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eExploring Academic Influence of Algorithms by Co-occurrence Network Based on Full-text of Academic Papers\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.24099v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Algorithms have become central to scientific research in the era of Artificial Intelligence (AI).\u003c/li\u003e\n\u003cli\u003eAlthough algorithms mentioned in papers are often used to indicate popularity and influence, existing research typically evaluates individual algorithms in isolation, with limited focus on the collective influence formed by their interconnections.\u003c/li\u003e\n\u003cli\u003eThis study constructs a large-scale algorithm co-occurrence network in Natural Language Processing (NLP) based on the full text of academic papers and investigates algorithmic influence from a network perspective.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.24099v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Algorithms have become central to scientific research in the era of artificial intelligence (AI)\u003c/li\u003e\n\u003cli\u003eAlthough algorithm mentions in papers are often used to indicate popularity and influence, existing studies usually evaluate individual algorithms in isolation…\u003c/li\u003e\n\u003cli\u003eThis study constructs large-scale algorithm co-occurrence networks in natural language processing (NLP) based on the full text of academic papers and investigat…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-cscl-b_introsearch\"\u003e\n  ArXiv cs.CL (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-cscl-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23693\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eEXPO-SQL: Execution-based Clause-level Policy Optimization for Text-to-SQL\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23693v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Text-to-SQL enables users to query databases using natural language by generating executable SQL queries.\u003c/li\u003e\n\u003cli\u003eRecent methods increasingly adopt Reinforcement Learning (RL) based on large language models to leverage execution feedback for training.\u003c/li\u003e\n\u003cli\u003eHowever, existing RL methods assign a uniform query-level reward to all clauses in an SQL query, treating correct and incorrect clauses equally.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eEN 要点:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23693v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Text-to-SQL enables users to query databases using natural language by generating executable SQL queries\u003c/li\u003e\n\u003cli\u003eRecent methods have increasingly adopted Large Language Models based reinforcement learning (RL) to leverage execution feedback for training\u003c/li\u003e\n\u003cli\u003eHowever, existing RL methods assign uniform query-level rewards to all clauses in a SQL query, treating correct and incorrect clauses equally\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23694\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eModTGCN: Modularity-aware Graph Neural Networks for Text Classification\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23694v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Despite semantic document graphs exhibiting strong class-consistent clustering, graph-based text classification models typically rely on local neighborhood aggregation and overlook global community structures.\u003c/li\u003e\n\u003cli\u003eIgnoring this can blur class boundaries and lead to over-smoothing.\u003c/li\u003e\n\u003cli\u003eWe propose ModTGCN, a modularity-aware graph neural network for text classification that jointly optimizes cross-entropy and a modularity-based auxiliary objective to facilitate class-consistent document communities while preserving discriminative representations.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23694v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Graph-based text classification models typically rely on local neighborhood aggregation and overlook global community structure, despite semantic docu…\u003c/li\u003e\n\u003cli\u003eIgnoring this can blur class boundaries and lead to over-smoothing\u003c/li\u003e\n\u003cli\u003eWe propose ModTGCN, a modularity-aware graph neural network for text classification that jointly optimizes cross-entropy and a modularity-based auxiliary object…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23695\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eQuantifying Prior Dominance in RAG Systems\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23695v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Retrieval-Augmented Generation (RAG) grounds Large Language Models in external knowledge, yet current evaluations rely on discrete heuristics suffering from \u0026ldquo;cognitive blindness\u0026rdquo;—unable to distinguish genuine contextual information extraction from parametric memory recall.\u003c/li\u003e\n\u003cli\u003eTo address this, we introduce Normalized Context Utilization (NCU), a metric leveraging continuous token log-probabilities under zero-shot, oracle, and adversarial conditions to rigorously quantify contextual information gain.\u003c/li\u003e\n\u003cli\u003eEvaluations across architectures from 1.5B to 72B parameters and proprietary commercial APIs reveal that for strict factual extraction (without chain-of-thought reasoning), traditional scaling laws exhibit extreme diminishing returns: efficient Small Language Models (SLMs) match or outperform high-capacity architectures.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23695v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Retrieval-Augmented Generation (RAG) grounds Large Language Models in external knowledge, yet current evaluations rely on discrete heuristics that suf…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eTo address this, we introduce the Normalized Context Utilization (NCU) metric, leveraging continuous token log-probabilities across zero-shot, oracle, and adver…\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eEvaluating architectures ranging from 1.5B to 72B parameters alongside a proprietary commercial API reveals that for strict factual extraction (without Chain-of…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23700\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eSelf-Recognition Finetuning can Prevent and Reverse Emergent Misalignment\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23700v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Emergent misalignment (EM) is associated with the activation of misaligned persona vectors and malevolent personality traits, suggesting that EM operates by disrupting the model\u0026rsquo;s consistent persona rather than directly learning harmful content.\u003c/li\u003e\n\u003cli\u003eMotivated by this connection, we investigate self-generated text recognition (SGTR) finetuning as a character-targeted intervention distinct from existing in-training defenses.\u003c/li\u003e\n\u003cli\u003eWe conduct two-stage finetuning experiments on three models (GPT-4.1, Qwen2.5-32B-Instruct, Seed-OSS-36B-Instruct) and multiple EM datasets, comparing SGTR finetuning with a benign finetuning baseline (correct domain-specific data, common sense, and word count statistics) to find it an effective defense in both reversal and prevention settings.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23700v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates thro…\u003c/li\u003e\n\u003cli\u003eMotivated by this connection, we study self-generated text recognition (SGTR) finetuning as a character-targeted intervention that is distinct from existing in-…\u003c/li\u003e\n\u003cli\u003eWe conduct two-stage finetuning experiments across three models (GPT-4.1, Qwen2.5-32B-Instruct, Seed-OSS-36B-Instruct) and multiple EM datasets to compare SGTR…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23701\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eEvaluating LLM Usage for Efficient and Explainable Numerical and Classified Implicit Sentiment Analysis of Product Desirability\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23701v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Qualitative product feedback can reveal nuanced user experiences, but its implicit sentiment is difficult to measure.\u003c/li\u003e\n\u003cli\u003eThis paper proposes a scalable and explainable framework that uses Large Language Models (LLMs) to quantify product desirability in such data.\u003c/li\u003e\n\u003cli\u003eUsing two Product Desirability Toolkit (PDT) datasets from ZORQ and CARMA, which include 106 respondent term groupings along with gold-standard human annotations, evaluation is performed with zero-shot continuous numerical sentiment scoring and categorical sentiment classification, without relying on explicit review scores.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23701v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Qualitative product feedback can reveal nuanced user experiences, but its implicit sentiment is difficult to measure\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eThis paper presents a scalable and interpretable framework that uses large language models (LLMs) to quantify product desirability from such data\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eUsing two Product Desirability Toolkit (PDT) datasets from ZORQ and CARMA comprising 106 respondent term groupings with gold-standard human annotation, zero-sho…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23881\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eGround Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23881v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Knowledge-Based Visual Question Answering (KB-VQA) requires grounding visual queries to external knowledge beyond directly observable content in the image.\u003c/li\u003e\n\u003cli\u003eWhile recent multi-modal large language models (MLLMs) show strong perceptual abilities, they struggle with KB-VQA tasks that require grounding at both the fine-grained entity and evidence levels.\u003c/li\u003e\n\u003cli\u003eMost existing multi-modal retrieval-augmented generation (MM-RAG) methods tightly couple entity discrimination and section-level evidence ranking into a single re-ranking stage, resulting in high costs and limited generalization.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23881v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Knowledge-Based Visual Question Answering (KB-VQA) requires grounding visual queries to external knowledge beyond directly observable content in image…\u003c/li\u003e\n\u003cli\u003eWhile recent multi modal large language models (MLLMs) show strong perceptual abilities, they struggle on KB-VQA tasks requiring groundings from both fine-grain…\u003c/li\u003e\n\u003cli\u003eMost existing multi-modal retrieval augmented generation (MM-RAG) methods tightly couple entity discrimination and section-level evidence ranking into a single…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23884\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eOne Year Later\u0026hellip;The Harms Persist, But So Do We!\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23884v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, but safety safeguards remain inadequate and inconsistent across different clinical conditions.\u003c/li\u003e\n\u003cli\u003eThis study evaluates six proprietary LLMs across 16 DSM-5 conditions using four adversarial attack variants, introducing an eight-dimension harm taxonomy and a multidimensional evaluation framework.\u003c/li\u003e\n\u003cli\u003eThe results show that safeguards are only effective for suicide and self-harm, while failure rates reach up to 100% for conditions such as eating disorders, substance use disorders, and major depressive disorder.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23884v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, yet safety safeguards remain inadequate an…\u003c/li\u003e\n\u003cli\u003eThis study evaluates six proprietary LLMs across 16 DSM-5 conditions using four adversarial attack variants, introducing an eight-dimension harm taxonomy and a…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eResults show that safeguards hold reliably only for suicide and self-harm, while conditions such as eating disorders, substance use disorder, and major depressi…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23915\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDo LLM Attribution Metrics Transfer? Auditing Retrieval-Augmented Generation Evaluation Across Datasets and Constructs\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23915v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: In practice, automatic metrics for attribution in LLM retrieval-augmented generation are often treated as interchangeable.\u003c/li\u003e\n\u003cli\u003eWe audited eight automatic scorers—lexical, embedding, and BERTScore baselines alongside entailment/grounding-trained models (clean and FEVER NLI, the checker MiniCheck)—across three evaluation constructs (provenance/topicality, generated-answer attribution, and fact-checking entailment), asking if any scorer transfers: remaining within the 95% confidence interval of the best-audited scorer on each dataset of a multi-dataset construct.\u003c/li\u003e\n\u003cli\u003eIn the construct with the most multi-dataset human-labeled coverage—generated-answer attribution (AttributionBench\u0026rsquo;s four source datasets, n = 1,610, with independent HAGRID, n = 2,150)—none did: per-dataset metric rankings inverted (Kendall tau = -0.64, p = 0.031 on AttributedQA vs. p = 0.031 on AttributedQA).\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23915v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Practice often treats automatic metrics for attribution in LLM retrieval-augmented generation as interchangeable\u003c/li\u003e\n\u003cli\u003eWe audit eight automatic scorers \u0026ndash; lexical, embedding, and BERTScore baselines alongside entailment/grounding-trained models (clean and FEVER NLI, the checker…\u003c/li\u003e\n\u003cli\u003eIn the construct with the most multi-dataset human-labeled coverage \u0026ndash; generated-answer attribution (AttributionBench\u0026rsquo;s four source datasets, n = 1,610, with in…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23937\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eWhen Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23937v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Exact-match retrieval recall is often used as a proxy for whether a retriever provides useful policy context for a downstream decision model.\u003c/li\u003e\n\u003cli\u003eWe test this proxy for pre-action policy classification in tau-bench using Qwen2.5-3B/7B classifiers.\u003c/li\u003e\n\u003cli\u003eUnder golden policy conditions, a compact structured state improves macro F1 by 0.13-0.17 over the original trajectory after tuning.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23937v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream decision model\u003c/li\u003e\n\u003cli\u003eWe test this proxy for pre-action policy classification in tau-bench using Qwen2.5-3B/7B classifiers\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eUnder gold-policy conditioning, a compact structured state improves macro-F1 over raw trajectories by 0.13-0.17 after tuning\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23943\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eQuechuaTok: Morphological Boundary Accuracy as a Necessary Metric for Tokenizer Evaluation in Agglutinative Low-Resource Languages\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23943v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Tokenization is a foundational step in NLP pipelines, yet standard evaluation metrics such as fertility rate fail to capture the morphological correctness of agglutinative languages.\u003c/li\u003e\n\u003cli\u003eWe present QuechuaTok, a systematic benchmark comparing four tokenization strategies (BPE, Unigram LM, WordPiece, and a morphology-aware PRPE tokenizer) for Southern Quechua (quz), a low-resource agglutinative language spoken by 8-10 million people in South America.\u003c/li\u003e\n\u003cli\u003eUsing a 200k-sentence corpus and the SQUOIA finite-state morphological analyzer (Rios, 2016) as a silver standard, we evaluate three metrics: fertility rate, OOV rate, and morphological boundary accuracy (MorphAcc).\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23943v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Tokenization is a foundational step in NLP pipelines, yet standard evaluation metrics such as fertility rate fail to capture morphological correctness…\u003c/li\u003e\n\u003cli\u003eWe present QuechuaTok, a systematic benchmark comparing four tokenization strategies - BPE, Unigram LM, WordPiece, and a morphology-aware PRPE tokenizer - for S…\u003c/li\u003e\n\u003cli\u003eUsing a 200k-sentence corpus and the SQUOIA finite-state morphological analyzer (Rios, 2016) as silver standard, we evaluate three metrics: fertility rate, OOV…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-cslg-b_introsearch\"\u003e\n  ArXiv cs.LG (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-cslg-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23739\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eSystematic Exploration of 4-Expert Heterogeneous Mixture-of-Experts via Automated Pipeline Search\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23739v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: We present an automated large-scale search pipeline for heterogeneous 4-Expert Mixture-of-Experts (MoE4) architectures within the LEMUR neural network dataset ecosystem.\u003c/li\u003e\n\u003cli\u003eBuilding on a hand-crafted heterogeneous MoE reference model, we replaced manual design with a deterministic code assembly generator that systematically combines base architecture families extracted from the LEMUR database into MoE4 ensembles, each controlled by a convolutional gating network with temperature scaling, mix-up augmentation, and a cosine-annealed learning rate schedule.\u003c/li\u003e\n\u003cli\u003eIn a 28-day campaign on an NVIDIA RTX 4090, the pipeline generated 4,463 candidate models in 197 batches, of which 1,021 were successfully evaluated.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23739v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: We present an automated large-scale search pipeline for heterogeneous 4-Expert Mixture-of-Experts (MoE4) architectures within the LEMUR neural network…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eBuilding on a hand-crafted heterogeneous MoE reference model, we replace manual design with a deterministic code-assembly generator that systematically combines…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eOver a 28-day campaign on an NVIDIA RTX 4090, the pipeline generated 4,463 candidate models across 197 batches, of which 1,021 were evaluated successfully\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23740\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eWeight-Space Geometry of Offline Reasoning Training\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23740v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Offline reinforcement-learning losses (RFT, RIFT, DFT, offline GRPO, DPO) are widely used to distill reasoning from large teachers into smaller students, and are typically compared only on downstream accuracy.\u003c/li\u003e\n\u003cli\u003eWe ask whether they are mechanistically distinct or converge to a similar weight update.\u003c/li\u003e\n\u003cli\u003eUsing attention-only LoRA to train six methods (SFT, RFT, DFT, RIFT, offline GRPO, DPO) from the same math rollouts of a single base model (Qwen3-4B), we analyze the generated deltas via cosine similarity, principal subspace analysis, linear mode connection, and CKA.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23740v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Offline reinforcement-learning losses (RFT, RIFT, DFT, Offline GRPO, DPO) are widely used to distill reasoning from large teachers into smaller studen…\u003c/li\u003e\n\u003cli\u003eWe ask whether they are mechanistically distinct or converge to a similar weight update\u003c/li\u003e\n\u003cli\u003eTraining six methods (SFT, RFT, DFT, RIFT, Offline GRPO, DPO) on identical math rollouts from a single base model (Qwen3-4B) with attention-only LoRA, we analyz…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23741\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eA Survey on Federated Causal Discovery and Inference\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23741v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Causal inference, which includes the discovery of causal structures and the inference of causal effects, is the foundation of data-driven decision-making.\u003c/li\u003e\n\u003cli\u003eIn practice, data for reliable causal analysis is often distributed across various institutions and cannot be centralized due to privacy regulations or communication constraints.\u003c/li\u003e\n\u003cli\u003eFederated Learning (FL) addresses this issue by enabling collaborative analysis without sharing raw data, which has led to the rapid development in the fields of Federated Causal Discovery (FCD) and Inference (FCI).\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23741v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Causal reasoning, which encompasses the discovery of causal structures and the inference of causal effects, is fundamental to data-driven decision mak…\u003c/li\u003e\n\u003cli\u003eIn practice, data for reliable causal analysis are often distributed across institutions and cannot be centralized due to privacy regulations or communication c…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eFederated learning (FL) addresses this by enabling collaborative analysis without raw data sharing, giving rise to the rapidly growing field of federated causal…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23742\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eLow-power analogue neural networks with trainable nonlinear connections for continuous control\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23742v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Physical neural networks promise low-power machine learning by computing directly with analog device physics, but most architectures force nonlinear device responses to act as scalar weights.\u003c/li\u003e\n\u003cli\u003eInspired by Kolmogorov-Arnold networks, we place trainable nonlinear functions on the connections, making each physical connection a learnable computational element.\u003c/li\u003e\n\u003cli\u003eBy implementing these functions as analog band-pass filters on field-programmable analog arrays, we find that their benefits are task-dependent and stem from the smoothness of the physical basis: the network represents smooth, continuously valued targets—including robot kinematics, continuous control, and photovoltaic maximum power point tracking—with far fewer nodes and connections than a multilayer perceptron, but offers no parametric efficiency advantage on classification-like decision boundaries.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23742v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Physical neural networks promise low-power machine learning by computing directly with analogue device physics, but most architectures force nonlinear…\u003c/li\u003e\n\u003cli\u003eInspired by Kolmogorov-Arnold networks, we place trainable nonlinear functions on the connections, making each physical connection a learnable computational ele…\u003c/li\u003e\n\u003cli\u003eRealising these functions as analogue band-pass filters on field-programmable analogue arrays, we find that the benefit is task-dependent and follows from the s…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23757\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eSynergizing Physically Constrained MCMC and Chemical-Informed Gaussian Processes for Reaction Network Discovery\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23757v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Extracting interpretable governing equations from sparse, noisy chemical time-series data remains difficult because the discrete reaction topology and continuous kinetic parameters are tightly coupled.\u003c/li\u003e\n\u003cli\u003eWe present PC-MCMC-CIGP, a reproducible gray-box workflow that combines spike-and-slab topology sampling, hard conservation and thermodynamic screening, and a chemically-informed Gaussian process (CIGP) residual model for parameter calibration and experimental design.\u003c/li\u003e\n\u003cli\u003eThe methodological contribution is not a new MCMC or GP family in isolation; rather, it is the integration of these components into a physically-constrained workflow with explicit uncertainty-aware acquisition choices.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23757v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Extracting interpretable governing equations from sparse, noisy chemical time-series data remains difficult because discrete reaction topology and con…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe present PC-MCMC-CIGP, a reproducible gray-box workflow that combines spike-and-slab topology sampling, hard conservation and thermodynamic screening, and a C…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eThe methodological contribution is not a new MCMC or GP family in isolation; rather, it is the integration of these components into a physically constrained wor…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23758\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eExploring Dualistic Meta-Learning to Enhance Domain Generalization in Open Set Scenarios\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23758v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Domain generalization learns from multiple source domains to generalize to unseen target domains.\u003c/li\u003e\n\u003cli\u003eHowever, it often neglects the realistic case of label mismatch between source and target.\u003c/li\u003e\n\u003cli\u003eOpen set domain generalization is then proposed to recognize unseen classes in unseen domains.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23758v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Domain generalization learns from multiple source domains to generalize to unseen target domains\u003c/li\u003e\n\u003cli\u003eHowever, it often neglects the realistic case of label mismatch between source and target\u003c/li\u003e\n\u003cli\u003eOpen set domain generalization is then proposed to recognize unseen classes in unseen domains\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23767\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eOne Ruler: A Same-Hands Re-Evaluation of Bivariate Causal Direction on Tuebingen, with a Parameter-Free Compression Baseline\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublished Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23767v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Headline accuracies on the Tuebingen cause-effect pairs are routinely compared across papers even though each is measured under its authors\u0026rsquo; own proto…\u003c/li\u003e\n\u003cli\u003eWe argue this is the wrong comparison and run the right one: a same-hands re-evaluation in which every method is run by us on the identical 102 pairs, with one…\u003c/li\u003e\n\u003cli\u003eAs a clean reference point we introduce a deliberately minimal baseline: sorted-conditional compression, which feeds quantized, sorted, first-differenced data t…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23767v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Headline accuracies on the Tuebingen cause-effect pairs are routinely compared across papers even though each is measured under its authors\u0026rsquo; own proto…\u003c/li\u003e\n\u003cli\u003eWe argue this is the wrong comparison and run the right one: a same-hands re-evaluation in which every method is run by us on the identical 102 pairs, with one…\u003c/li\u003e\n\u003cli\u003eAs a clean reference point we introduce a deliberately minimal baseline: sorted-conditional compression, which feeds quantized, sorted, first-differenced data t…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23830\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDeciphering Fingerprints of 3D Molecular Surfaces for Accurate Epitope Prediction\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23830v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Molecular surfaces encode the geometric and physicochemical patterns that determine antibody-antigen recognition, which is crucial for epitope prediction.\u003c/li\u003e\n\u003cli\u003eHowever, existing methods rely on sequences or backbone structures and struggle to capture discontinuous, surface-driven epitopes.\u003c/li\u003e\n\u003cli\u003eThis study presents SurfBind, a surface-centric learning framework for epitope prediction that operates directly on molecular surface representations.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23830v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Molecular surfaces encode the geometric and physicochemical patterns that determine antibody-antigen recognition, central to epitope prediction\u003c/li\u003e\n\u003cli\u003eHowever, existing methods rely on sequences or backbone structures and struggle to capture discontinuous, surface-driven epitopes\u003c/li\u003e\n\u003cli\u003eThis study presents SurfBind, a surface-centric learning framework for epitope prediction that operates directly on molecular surface representations\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23833\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eReconstructing GRACE Terrestrial Water Storage with Spatio-Temporal Graph Neural Networks: An Application to South America\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23833v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Terrestrial Water Storage (TWS) integrates snow, soil moisture, surface water, and groundwater, and is a key indicator of how climate change and human activities are reshaping the global water cycle.\u003c/li\u003e\n\u003cli\u003eThe GRACE and GRACE-FO satellite missions provide the only direct, globally consistent observations of TWS changes, but their records only begin in 2002, which is too short for many climate-scale analyses.\u003c/li\u003e\n\u003cli\u003eWe propose a deep learning application that reconstructs monthly GRACE-like TWS anomalies (TWSA) back to 1940 by learning the relationship between daily ERA5 meteorological forcings (precipitation, evapotranspiration, runoff) and monthly GRACE observations.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23833v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Terrestrial water storage (TWS) integrates snow, soil moisture, surface water, and groundwater and is a key indicator of how climate variability and h…\u003c/li\u003e\n\u003cli\u003eThe GRACE and GRACE-FO satellite missions provide the only direct, globally consistent observations of TWS change, but their record only begins in 2002 which is…\u003c/li\u003e\n\u003cli\u003eWe present a deep learning application that reconstructs monthly GRACE-like TWS anomalies (TWSA) back to 1940 by learning the relationship between daily ERA5 me…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.23838\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eThe Degeneracy Distillery\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-24 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.23838v1 Announcement Type: New.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: When two or more parameters or labels produce similar data, they are degenerate, or hard to distinguish.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eDegeneracy makes both label prediction and inverse problems difficult, because machine learning algorithms and probabilistic samplers both rely on the distinguishability of data and its gradients with respect to parameters.\u003c/li\u003e\n\u003cli\u003eHowever, identifying degeneracies in physical models or real-world datasets can elucidate the choice of model or the underlying process that generates the data.\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.23838v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: When two or more parameters or labels produce similar data, they are degenerate, or hard to distinguish\u003c/li\u003e\n\u003cli\u003eDegeneracies render both label prediction and inverse problems difficult, since both machine machine learning algorithms and probabilistic samplers rely on the distingu…\u003c/li\u003e\n\u003cli\u003eHowever, identifying degeneracies in physical models or real-world datasets can be elucidating about the choice of model or the underlying process that produces…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n",
  "wordCount": 7580,
  "readingTime": 36,
  "tableOfContents": "\u003cnav id=\"TableOfContents\"\u003e\n  \u003cul\u003e\n    \u003cli\u003e\u003ca href=\"#-in-depth-guide-to-this-issues-watch-list\"\u003e📖 In-depth Guide to This Issue\u0026rsquo;s Watch List\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-x-platform-ai-hot-topics-quick-brief\"\u003e🌐 X Platform AI Hot Topics Quick Brief\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#topic-1-loop-engineering-emerges-as-new-way-to-automate-ai-coding-agents\"\u003eTopic 1: Loop Engineering Emerges as New Way to Automate AI Coding Agents\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-2-developer-migrates-36k-ai-agents-from-openclaw-to-hermes-for-better-reliability\"\u003eTopic 2: Developer Migrates $36K AI Agents from OpenClaw to Hermes for Better Reliability\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-3-anthropic-accuses-alibaba-linked-group-of-massive-claude-ai-distillation-attack\"\u003eTopic 3: Anthropic Accuses Alibaba-Linked Group of Massive Claude AI Distillation Attack\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-4-chinas-glm-52-tops-coding-benchmarks-at-fraction-of-us-ai-costs\"\u003eTopic 4: China\u0026rsquo;s GLM-5.2 Tops Coding Benchmarks at Fraction of U.S. AI Costs\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-5-users-build-ai-powered-second-brains-with-claude-and-obsidian\"\u003eTopic 5: Users Build AI-Powered Second Brains with Claude and Obsidian\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-influencer-insights\"\u003e💡 Influencer Insights\u003c/a\u003e\u003c/li\u003e\n  \u003c/ul\u003e\n\n  \u003cul\u003e\n    \u003cli\u003e\u003ca href=\"#1--todays-focus-technology-trends-and-product-hotspots\"\u003e1. 🔥 Today\u0026rsquo;s Focus: Technology Trends and Product Hotspots\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#on-device-models-enter-practical-stage\"\u003eOn-device Models Enter Practical Stage\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#coding-agent-ecosystem-from-usable-to-ultimate-cost-and-experience\"\u003eCoding Agent Ecosystem: From \u0026ldquo;Usable\u0026rdquo; to \u0026ldquo;Ultimate Cost and Experience\u0026rdquo;\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#deeper-integration-of-ai-and-office-software\"\u003eDeeper Integration of AI and Office Software\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#2--unique-perspectives-and-industry-foresight\"\u003e2. 🧠 Unique Perspectives and Industry Foresight\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#3--recommended-tools-and-resources\"\u003e3. 🛠️ Recommended Tools and Resources\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-appendix-todays-watch-list-update-source-list\"\u003e📚 Appendix: Today\u0026rsquo;s Watch List Update Source List\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#a16z-podcast-a_full\"\u003ea16z Podcast (A_full)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#stratechery-by-ben-thompson-a_full\"\u003eStratechery by Ben Thompson (A_full)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#openai-blog-a_full\"\u003eOpenAI Blog (A_full)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#google-deepmind-blog-a_full\"\u003eGoogle DeepMind Blog (A_full)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-csai-b_introsearch\"\u003eArXiv cs.AI (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-cscl-b_introsearch\"\u003eArXiv cs.CL (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-cslg-b_introsearch\"\u003eArXiv cs.LG (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n  \u003c/ul\u003e\n\u003c/nav\u003e",
  "isDraft": false
}
