{
  "title": "2026-06-16 AI Daily Update | Agents Enter Real Workflows, Enterprises Begin to Re-evaluate \"Token Capital\"",
  "url": "https://miaok.ong/en/ai-daily/ai-daily-2026-06-16/",
  "date": "2026-06-16T07:00:00+08:00",
  "lastmod": "2026-06-16T07:00:00+08:00",
  "type": "ai-daily",
  "kind": "page",
  "language": "en",
  "description": "Today\u0026rsquo;s main theme is the transition of AI Agents from demos to real-world production: the completion rate of office agents has significantly improved, and their implementation in scenarios such as software factories, BI, and multimodal orchestration is accelerating. Meanwhile, benchmarking, security, refusal control, and the accumulation of proprietary enterprise knowledge are becoming the new infrastructure, while on-device code models are beginning to undergo more stringent productivity tests.",
  "keywords": null,
  "tags": [],
  "categories": [],
  "author": "Mark (Miao) Kong",
  "image": "https://miaok.ong/images/avatar.jpg",
  "content": "\u003ch1 id=\"2026-06-16-ai-daily--agents-enter-real-workflows-companies-begin-to-re-evaluate-token-capital\"\u003e\n  2026-06-16 AI Daily | Agents Enter Real Workflows, Companies Begin to Re-evaluate \u0026ldquo;Token Capital\u0026rdquo;\n  \u003ca class=\"heading-link\" href=\"#2026-06-16-ai-daily--agents-enter-real-workflows-companies-begin-to-re-evaluate-token-capital\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h1\u003e\n\u003cblockquote\u003e\n\u003cp\u003eToday\u0026rsquo;s main theme is the transition of AI Agents from demos to real production: the completion rate of office agents has significantly increased, and their adoption in scenarios like software factories, BI, and multimodal orchestration is accelerating. Meanwhile, evaluation, security, refusal control, and the accumulation of proprietary enterprise knowledge are becoming new forms of infrastructure, while on-device code models are starting to undergo more rigorous productivity tests.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch2 id=\"-in-depth-guide-to-this-issues-watch-list\"\u003e\n  📖 In-depth Guide to This Issue\u0026rsquo;s Watch List\n  \u003ca class=\"heading-link\" href=\"#-in-depth-guide-to-this-issues-watch-list\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003eThe most noteworthy trend to follow closely today is \u0026ldquo;agents moving into real workflows\u0026rdquo;: A two-year follow-up on WorkBench shows that the task completion rate for top-tier office agents has increased from 43% to 89%, with a significant reduction in errors. Simultaneously, TwinBI, WebDecept, and Orchestra-o1 are pushing agents into BI analytics, multimodal orchestration, and e-commerce security testing, respectively, making this a key area for product and platform teams to track systematically.\u003c/p\u003e\n\u003cp\u003eThe second trend is that \u0026ldquo;evaluation and security are becoming more granular.\u0026rdquo; Updates like Anthropic\u0026rsquo;s safety narrative, research on the stability of LLM-as-a-Judge, interventions in refusal mechanisms, and model collapse due to sample selection bias collectively remind us that as model capabilities expand, reliable evaluation, refusal control, and data feedback loops become fundamental infrastructure problems.\u003c/p\u003e\n\u003cp\u003eFinally, in creative AI, the interview about Ideogram\u0026rsquo;s open-weight image model is worth noting. It shifts the focus from \u0026ldquo;generation quality\u0026rdquo; to text rendering, layout, controllable editing, and design workflows, signaling that image models are entering a more engineering-driven stage of product competition.\u003c/p\u003e\n\u003ch2 id=\"-ai-hot-topics-on-x\"\u003e\n  🌐 AI Hot Topics on X\n  \u003ca class=\"heading-link\" href=\"#-ai-hot-topics-on-x\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003ch3 id=\"topic-1-ai-coding-agents-speed-past-human-review-limits\"\u003e\n  Topic 1: AI Coding Agents Speed Past Human Review Limits\n  \u003ca class=\"heading-link\" href=\"#topic-1-ai-coding-agents-speed-past-human-review-limits\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending Time:, Related Posts: 75\u003c/li\u003e\n\u003cli\u003eSummary: AI Coding Agents Speed Past Human Review Limits: A.I News A.I Def: artificial intelligence (AI), the ability of a digital computer or computer-controlled robot to perform tasks commonly associated with intelligent beings. The term is frequently applied to the project of developing systems endowed w\u0026hellip;\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-2-factory-ai-launches-factory-20-as-fully-autonomous-software-factories\"\u003e\n  Topic 2: Factory AI Launches Factory 2.0 as Fully Autonomous Software Factories\n  \u003ca class=\"heading-link\" href=\"#topic-2-factory-ai-launches-factory-20-as-fully-autonomous-software-factories\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending Time: 5 hours ago, Related Posts: 1600\u003c/li\u003e\n\u003cli\u003eWhat happened: Factory AI announced Factory 2.0, claiming to upgrade the software development process into autonomously operating \u0026ldquo;software factories.\u0026rdquo;\u003c/li\u003e\n\u003cli\u003eWhy it matters: This reflects the evolution of AI Agents from coding assistants to end-to-end software delivery systems, which could potentially transform R\u0026amp;D efficiency, team structures, and the software startup model.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: Discussions on X are focused on whether its autonomous development capabilities are truly reliable, whether new metrics like DAA can effectively measure Agent output, and whether such platforms will shift Silicon Valley\u0026rsquo;s investment logic from human-intensive teams to highly automated software factories.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-3-vercel-extends-serverless-functions-to-30-minutes-for-ai-workloads\"\u003e\n  Topic 3: Vercel Extends Serverless Functions to 30 Minutes for AI Workloads\n  \u003ca class=\"heading-link\" href=\"#topic-3-vercel-extends-serverless-functions-to-30-minutes-for-ai-workloads\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · News\u003c/li\u003e\n\u003cli\u003eOverview: Trending Time: 5 hours ago, Related Posts: 184\u003c/li\u003e\n\u003cli\u003eWhat happened: Vercel has extended the maximum execution time for its Serverless Functions to 30 minutes to better support long-running AI inference, generation, and background tasks.\u003c/li\u003e\n\u003cli\u003eWhy it matters: This reduces the architectural complexity for developers deploying AI applications on Vercel, allowing tasks like long-text generation, agent workflows, and multi-step inference to be executed without needing to be prematurely split into queues or separate backend services.\u003c/li\u003e\n\u003cli\u003eDiscussion summary: Discussions on X are centered on whether this makes Vercel more suitable for AI prototyping and production deployment. Supporters believe it simplifies full-stack AI application development, while critics raise concerns about cost, cold starts, stability, and the cost-effectiveness compared to dedicated inference platforms or traditional backends.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-4-le-chaton-fat-mistral-ais-massive-meme-hoax\"\u003e\n  Topic 4: Le Chaton Fat: Mistral AI\u0026rsquo;s Massive Meme Hoax\n  \u003ca class=\"heading-link\" href=\"#topic-4-le-chaton-fat-mistral-ais-massive-meme-hoax\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · Entertainment\u003c/li\u003e\n\u003cli\u003eOverview: Trending Time: 22 hours ago, Related Posts: 10000\u003c/li\u003e\n\u003cli\u003eWhat happened: A topic called \u0026ldquo;Le Chaton Fat,\u0026rdquo; suspected of being an impersonation or parody of Mistral AI, went viral on X. The content is widely considered to be a meme-based prank or hoax related to a new AI model/product.\u003c/li\u003e\n\u003cli\u003eWhy it matters: It reflects how, under the intense spotlight on the AI industry, model releases, brand announcements, and technical rumors can easily be turned into memes and spread rapidly, highlighting the importance of information verification and platform dissemination mechanisms.\u003c/li\u003e\n\u003cli\u003eDiscussion Summary: The discussion focuses on whether this is harmless entertainment or a misleading rumor. Some view it as a satire on the over-hyping within the AI community, while others are concerned that such false information could obscure actual product launches, harm company reputations, and mislead investors and users.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"topic-5-ai-coded-wow-clone-draws-12000-players-in-days\"\u003e\n  Topic 5: AI-Coded WoW Clone Draws 12,000 Players in Days\n  \u003ca class=\"heading-link\" href=\"#topic-5-ai-coded-wow-clone-draws-12000-players-in-days\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003eCategory: AI · Entertainment\u003c/li\u003e\n\u003cli\u003eSummary: Trending since: 17 hours ago, Related posts: 501\u003c/li\u003e\n\u003cli\u003eWhat happened: A World of Warcraft-style clone game, coded with AI assistance, attracted approximately 12,000 players within days.\u003c/li\u003e\n\u003cli\u003eWhy it matters: This shows that generative AI is lowering the barrier to entry for game development and may accelerate the cycle from prototyping to live operations, impacting content production, independent development, and the division of labor in the gaming industry.\u003c/li\u003e\n\u003cli\u003eDiscussion Summary: The discussion on X centers on whether AI genuinely enhances development efficiency, the originality and copyright risks associated with such projects, whether the game\u0026rsquo;s quality can retain players in the long run, and whether AI tools will empower individual developers or further disrupt traditional jobs in the gaming sector.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch4 id=\"summary-of-ai-public-opinion-on-x-today\"\u003e\n  Summary of AI Public Opinion on X Today\n  \u003ca class=\"heading-link\" href=\"#summary-of-ai-public-opinion-on-x-today\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h4\u003e\n\u003cp\u003eThe main narrative today is that AI is evolving from an \u0026ldquo;assistive tool\u0026rdquo; into a production system capable of independently completing complex processes. Software development, application deployment, and game creation are all being reimagined as more automated, lower-barrier workflows. The consensus is that Agents, long-running AI workflows, and generative development tools are indeed reducing the cost from prototype to launch and may alter the scale of startup teams, the division of R\u0026amp;D labor, and the velocity of content creation. Disagreements primarily focus on whether these capabilities are reliable enough and hold real production value, and whether platform-based solutions are genuinely more cost-effective than traditional backends, specialized inference services, or human teams. Potential risks include an expectation bubble fueled by hype, copyright and originality controversies surrounding AI-generated content, the cost and stability of deploying long-running tasks, and false news or meme-driven narratives that obscure actual product launches and mislead users and investors.\u003c/p\u003e\n\u003ch2 id=\"-influencer-insights\"\u003e\n  💡 Influencer Insights\n  \u003ca class=\"heading-link\" href=\"#-influencer-insights\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003ch1 id=\"ai-industry-frontier-insights-daily-2026-06-15\"\u003e\n  AI Industry Frontier Insights Daily (2026-06-15)\n  \u003ca class=\"heading-link\" href=\"#ai-industry-frontier-insights-daily-2026-06-15\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h1\u003e\n\u003cp\u003eBased on posts from several AI influencers on the X platform over the past 24 hours, here is a quick overview and in-depth summary of today\u0026rsquo;s industry dynamics.\u003c/p\u003e\n\u003ch2 id=\"1-todays-tech-trends-and-product-hotspots\"\u003e\n  1. Today\u0026rsquo;s Tech Trends and Product Hotspots\n  \u003ca class=\"heading-link\" href=\"#1-todays-tech-trends-and-product-hotspots\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003ch3 id=\"-hot-topic-the-ultimate-showdown-and-capability-limits-of-on-device-models\"\u003e\n  🔥 Hot Topic: The \u0026ldquo;Ultimate Showdown\u0026rdquo; and Capability Limits of On-Device Models\n  \u003ca class=\"heading-link\" href=\"#-hot-topic-the-ultimate-showdown-and-capability-limits-of-on-device-models\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cp\u003eToday\u0026rsquo;s most heated technical discussion centers on the \u003cstrong\u003epractical, real-world capabilities of code models\u003c/strong\u003e, particularly in performance comparisons on local hardware.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eReal-World Test of Gemma 4 12B Coder\u003c/strong\u003e:\nGoogle\u0026rsquo;s new Gemma 4 12B Coder has attracted widespread attention, but its performance has been controversial. \u003cstrong\u003e@zhixianio\u003c/strong\u003e conducted a rigorous test on an M5 Max, comparing it against the \u0026ldquo;daily driver\u0026rdquo; Qwen3.6-35B-A3B MoE. The results were as follows:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eSimple tasks (Matplotlib plotting)\u003c/strong\u003e: The two were on par.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eComplex, long-running tasks (Three.js galaxy/Tetris)\u003c/strong\u003e: The 12B Gemma experienced critical issues like black screens and logic flaws, completely losing to the 35B Qwen. \u003cstrong\u003e@zhixianio\u003c/strong\u003e noted that a 12B model size inherently cannot support complex programs that are \u0026ldquo;long-form, stateful, and single-shot,\u0026rdquo; and while community fine-tuning can improve efficiency, it cannot raise the capability ceiling.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eHighlights\u003c/strong\u003e: He observed that the fine-tuned Gemma Coder version learned to \u0026ldquo;think briefly then act,\u0026rdquo; converging faster than the original. In contrast, the original Gemma 4, after enabling its thinking mode, showed an extreme case of using 12,000 tokens entirely for \u0026ldquo;thinking\u0026rdquo; without producing a single line of code.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eNew Frontiers in Multimodality and On-Device Applications\u003c/strong\u003e:\n\u003cstrong\u003e@zhixianio\u003c/strong\u003e also deeply tested the QAT (Quantization Aware Training) version and multimodal capabilities of Gemma 4. On the M5 Max, its English and Japanese recognition performance was excellent and fast, but \u003cstrong\u003eChinese recognition was \u0026ldquo;completely nonsensical.\u0026rdquo;\u003c/strong\u003e He spoke highly of Google\u0026rsquo;s on-device strategy, suggesting that enabling models to be \u0026ldquo;natively adapted for quantization\u0026rdquo; through QAT will drastically reduce memory footprint, heralding an era where Android devices can smoothly run their own high-performance models.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026ldquo;Ascetic-Style\u0026rdquo; Local-First Practice\u003c/strong\u003e:\n\u003cstrong\u003e@zhixianio\u003c/strong\u003e shared his \u0026ldquo;monastic\u0026rdquo; experience of completely switching to a local model (Qwen3.6-35B-A3B), and the results were stunning: in PA and Coding scenarios, its response speed is faster than remote LLMs, its \u0026ldquo;IQ\u0026rdquo; is reliable, and the native multimodal experience felt \u0026ldquo;even better than DSV4 Pro.\u0026rdquo; This signifies that \u003cstrong\u003eon-device intelligence on high-end consumer hardware has officially crossed the threshold into practical productivity\u003c/strong\u003e.\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"-tools-and-platforms-claude-design-sparks-a-design-paradigm-shift\"\u003e\n  🛠️ Tools and Platforms: Claude Design Sparks a Design Paradigm Shift\n  \u003ca class=\"heading-link\" href=\"#-tools-and-platforms-claude-design-sparks-a-design-paradigm-shift\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eIn-Depth Analysis of Claude Design\u003c/strong\u003e:\n\u003cstrong\u003e@dotey (Bao Yu)\u003c/strong\u003e published a long-form article deeply analyzing why GPT-5.5 (Codex) can\u0026rsquo;t yet create a product like Claude Design. The core argument is that Claude Design delivers not just UI mockups, but \u003cstrong\u003ehigh-fidelity interactive prototypes with a complete data architecture and state management logic\u003c/strong\u003e. This requires the model to thoroughly design the entire interactive system before taking action, something only Claude Opus 4.8 has achieved so far.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eBest Practices for Design-to-Code\u003c/strong\u003e:\n\u003cstrong\u003e@dotey\u003c/strong\u003e gave a live demonstration of how to integrate Claude Design into the development workflow: modify the UI in the design file, view the changes via \u003ccode\u003egit diff\u003c/code\u003e after downloading, and then have Claude Code automatically sync the Swift code. This flow completely transforms the traditional communication model in front-end development, turning designers into \u0026ldquo;design managers\u0026rdquo; who direct agents using natural language.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eFable 5 Review and Leak Controversy\u003c/strong\u003e:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003ePolarized Reviews\u003c/strong\u003e: \u003cstrong\u003e@Pluvio9yte (Xueta Wuyun)\u003c/strong\u003e offered a non-mainstream conclusion: it\u0026rsquo;s \u003cstrong\u003eextremely slow\u003c/strong\u003e, but its conceptual boundaries and architectural capabilities are exceptionally strong, like a mix of Claude Opus 4.6++ and GPT-5.5++. They also noted the consumption rate wasn\u0026rsquo;t alarming. In contrast, \u003cstrong\u003e@zhixianio\u003c/strong\u003e was amazed by Fable\u0026rsquo;s autonomy, claiming it completed 70% of a development task in 40 minutes and even identified flaws in the original plan.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eThe Legendary System Prompt Leak\u003c/strong\u003e: \u003cstrong\u003e@Pluvio9yte\u003c/strong\u003e shared a document purported to be the Fable 5 system prompt, noting its immense value for understanding AI Agent design.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch2 id=\"2-unique-perspectives-and-industry-foresight\"\u003e\n  2. Unique Perspectives and Industry Foresight\n  \u003ca class=\"heading-link\" href=\"#2-unique-perspectives-and-industry-foresight\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003ch3 id=\"-a-new-concept-token-capital\"\u003e\n  🧠 A New Concept: \u0026ldquo;Token Capital\u0026rdquo;\n  \u003ca class=\"heading-link\" href=\"#-a-new-concept-token-capital\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cp\u003eThis concept from Microsoft CEO Satya Nadella was highlighted by \u003cstrong\u003e@dotey\u003c/strong\u003e. The core argument is that in the future, companies will need not just human capital, but also \u003cstrong\u003eToken Capital\u003c/strong\u003e—the proprietary knowledge and experience that the company builds and embeds into its own AI systems.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eThe Key Litmus Test\u003c/strong\u003e: Can you replace the underlying foundation model at any time without losing the company\u0026rsquo;s accumulated experience? If not, you are merely \u0026ldquo;renting\u0026rdquo; intelligence.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eThe Learning Flywheel / Compounding Effect\u003c/strong\u003e: The biggest crisis for a company isn\u0026rsquo;t technological obsolescence, but outsourcing its \u0026ldquo;capacity to learn.\u0026rdquo; Nadella warned against allowing a few models to monopolize all industry value, which could lead to a hollowing-out of industries, much like early globalization and outsourcing did.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"-counter-intuitive-observations-and-deep-reflections\"\u003e\n  📉 Counter-intuitive Observations and Deep Reflections\n  \u003ca class=\"heading-link\" href=\"#-counter-intuitive-observations-and-deep-reflections\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eIs AI Programming More Expensive Than Humans?\u003c/strong\u003e: \u003cstrong\u003e@ruanyf\u003c/strong\u003e calculated the token consumption of an OpenAI employee (equivalent to $1.3 million per person per month) and pointed out that unrestricted use of top-tier models is prohibitively expensive for companies. Even with Chinese models that are 30-50 times cheaper, the annual cost would still be millions of RMB.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eThe Engineering Pitfalls of \u0026ldquo;Vibe Coding\u0026rdquo;\u003c/strong\u003e: \u003cstrong\u003e@Pluvio9yte\u003c/strong\u003e shared his personal journey from a \u0026ldquo;Vibe Coder\u0026rdquo; to a disciplined engineer. He proposed the \u003cstrong\u003e\u0026ldquo;Contract First\u0026rdquo;\u003c/strong\u003e principle: if you have an AI write code without first defining API contracts and data models, the entire project becomes a \u0026ldquo;drafty wall.\u0026rdquo;\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eTeam Culture in the Age of AI\u003c/strong\u003e: \u003cstrong\u003e@dotey\u003c/strong\u003e quoted seven lessons from the head of design at Lovable, the most insightful being \u003cstrong\u003e\u0026ldquo;Get senior people to be hands-on again.\u0026rdquo;\u003c/strong\u003e AI gives experienced managers who have long been detached from front-line work a chance to regain the immense leverage of an individual contributor.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eThe Evolution of Skills\u003c/strong\u003e: \u003cstrong\u003e@lijigang (Li Jigang)\u003c/strong\u003e reflected on the future AI ecosystem, suggesting that Skills will evolve in two directions: one is \u003cstrong\u003e\u0026ldquo;downward atomization,\u0026rdquo;\u003c/strong\u003e breaking down human abilities into specialized skill packs; the other is \u003cstrong\u003e\u0026ldquo;upward componentization,\u0026rdquo;\u003c/strong\u003e encapsulating end-to-end best practices for specific scenarios.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch2 id=\"3-recommended-tools-and-resources\"\u003e\n  3. Recommended Tools and Resources\n  \u003ca class=\"heading-link\" href=\"#3-recommended-tools-and-resources\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eDesign \u0026amp; Prototyping\u003c/strong\u003e:\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eClaude Design\u003c/strong\u003e: Highly recommended by \u003cstrong\u003e@dotey\u003c/strong\u003e and \u003cstrong\u003e@vista8\u003c/strong\u003e for generating high-fidelity interactive prototypes from a single prompt.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eDeveloper Tools \u0026amp; Skill Frameworks\u003c/strong\u003e:\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eSpec-Code (Development Framework)\u003c/strong\u003e: Launched by \u003cstrong\u003e@Pluvio9yte\u003c/strong\u003e, this is a full-stack framework built on OpenSpec that emphasizes a \u0026ldquo;Contract First\u0026rdquo; approach, making it especially suitable for less experienced developers. (\u003ca href=\"https://t.co/UTeCTrjerA\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eGitHub\u003c/a\u003e)\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eQiaomu Series Skills\u003c/strong\u003e: Open-sourced by \u003cstrong\u003e@vista8\u003c/strong\u003e, including:\n\u003cul\u003e\n\u003cli\u003e\u003ccode\u003eqiaomu-ai-prd\u003c/code\u003e: Specifically designed for AI Agents to generate structured development requirement documents. (\u003ca href=\"https://t.co/MM8oY3WV5f\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eGitHub\u003c/a\u003e)\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\u003ccode\u003eqiaomu-novel-generator\u003c/code\u003e: A Skill for automatically generating novel synopses, character settings, and full text. (\u003ca href=\"https://t.co/qNzQGKuRkU\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eGitHub\u003c/a\u003e)\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003e\u0026ldquo;Illustrated Skill\u0026rdquo; Companion Repo\u003c/strong\u003e: The companion repository for \u003cstrong\u003e@dotey\u003c/strong\u003e\u0026rsquo;s new book, \u0026ldquo;Illustrated Skill,\u0026rdquo; which includes practical Skills like \u003ccode\u003einfo-digest\u003c/code\u003e (AI news compilation and writing). (\u003ca href=\"https://t.co/Pa4Ah9d4nl\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eGitHub\u003c/a\u003e)\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eData Analysis \u0026amp; Efficiency\u003c/strong\u003e:\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eApp Store Review Analysis Tool\u003c/strong\u003e: An open-source tool by \u003cstrong\u003e@vista8\u003c/strong\u003e that automatically scrapes App Store reviews and uses DeepSeek for sentiment and product opportunity analysis. (\u003ca href=\"https://t.co/STK2hbpYfv\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eGitHub\u003c/a\u003e)\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eSitedata Plugin\u003c/strong\u003e: Shared by group member \u003cstrong\u003e@AI_Jasonyu\u003c/strong\u003e, this plugin can reverse-lookup other websites owned by the same advertiser via a Google AdSense ID.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003emacOS Productivity Tools\u003c/strong\u003e: \u003cstrong\u003e@Pluvio9yte\u003c/strong\u003e recommended tools like \u003cstrong\u003eMaccy\u003c/strong\u003e (open-source clipboard), \u003cstrong\u003eMos\u003c/strong\u003e (smooth mouse scrolling), and \u003cstrong\u003eScreen Studio\u003c/strong\u003e (screen recording).\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eOther Cool Applications\u003c/strong\u003e:\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eYouMind 1.0\u003c/strong\u003e: A creation tool recommended by both \u003cstrong\u003e@vista8\u003c/strong\u003e and \u003cstrong\u003e@gefei55\u003c/strong\u003e, suitable for making presentations and mind maps.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eWorld Cup Viewing Calendar\u003c/strong\u003e: Developed by \u003cstrong\u003e@vista8\u003c/strong\u003e in just 24 minutes using Codex + Goal Skill, this tool supports personalized match schedule calendar subscriptions.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch2 id=\"-appendix-todays-watch-list-update-sources\"\u003e\n  📚 Appendix: Today\u0026rsquo;s Watch List Update Sources\n  \u003ca class=\"heading-link\" href=\"#-appendix-todays-watch-list-update-sources\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cblockquote\u003e\n\u003cp\u003eTimeframe: Last 3 days; 22 sources covered; 32 updates in total.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch3 id=\"a16z-podcast-a_full\"\u003e\n  a16z Podcast (A_full)\n  \u003ca class=\"heading-link\" href=\"#a16z-podcast-a_full\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003e\u003ca href=\"https://ai-a16z.simplecast.com/episodes/ideograms-open-weights-image-model-and-the-future-of-ai-design-qXdoKLhL\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eIdeogram’s Open-Weights Image Model and the Future of AI Design\u003c/a\u003e\u003c/strong\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 23:40 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:\n\u003cul\u003e\n\u003cli\u003eYoko Li and Justine Moore speak with Ideogram founder and CEO Mohammad Norouzi about image generation models, design workflows, and the evolving relationship between artificial intelligence and creative work.\u003c/li\u003e\n\u003cli\u003eThe conversation covers Ideogram\u0026rsquo;s decision to release an open-weights model, the challenges of generating text and layouts within images, and why controllability has become an increasingly important area of research.\u003c/li\u003e\n\u003cli\u003eThey discuss prompting, customization, editing, and the tradeoffs between general-purpose models and systems optimized for specific creative tasks.\u003c/li\u003e\n\u003cli\u003eAlong the way, Norouzi shares his views on open-source AI, design tools, agentic workflows, and how image generation models may evolve as creators and enterprises seek greater control over their output.\u003c/li\u003e\n\u003cli\u003eListen to the a16z Podcast on Apple Podcasts.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003eYoko Li and Justine Moore speak with Ideogram founder and CEO Mohammad Norouzi about image generation models, design workflows, and the evolving relationship be…\u003c/li\u003e\n\u003cli\u003eThe conversation covers Ideogram\u0026rsquo;s decision to release an open-weight model, the challenges of generating text and layouts within images, and why controllabilit…\u003c/li\u003e\n\u003cli\u003eThey discuss prompting, customization, editing, and the tradeoffs between general-purpose models and systems optimized for specific creative tasks\u003c/li\u003e\n\u003cli\u003eAlong the way, Norouzi shares his views on open-source AI, design tools, agentic workflows, and how image generation models may evolve as creators and enterpris…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"stratechery-by-ben-thompson-a_full\"\u003e\n  Stratechery by Ben Thompson (A_full)\n  \u003ca class=\"heading-link\" href=\"#stratechery-by-ben-thompson-a_full\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003e\u003ca href=\"https://stratechery.com/2026/anthropics-safety-superpower/\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eAnthropic’s Safety Superpower\u003c/a\u003e\u003c/strong\u003e\n\u003cul\u003e\n\u003cli\u003ePosted: 2026-06-15 18:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - \u003cstrong\u003eListen to this\u003c/strong\u003e post**:**.\n\u003cul\u003e\n\u003cli\u003eI\u0026rsquo;m sympathetic to the cynics who consistently characterize Anthropic\u0026rsquo;s public statements (especially those surrounding their model releases) as fear-mongering for marketing purposes.\u003c/li\u003e\n\u003cli\u003eJust two months ago, Anthropic announced Mythos Preview, a model they deemed too dangerous for public release, particularly due to its advanced cybersecurity capabilities.\u003c/li\u003e\n\u003cli\u003eThen, two months later, the company publicly released Fable, a version of Mythos with various safety guardrails.\u003c/li\u003e\n\u003cli\u003eIn my limited experience, Fable is a very impressive model.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003eListen to this post :\u003c/li\u003e\n\u003cli\u003eLog in to listen\u003c/li\u003e\n\u003cli\u003eI’m sympathetic to the cynics who consistently characterize Anthropic’s public statements, particularly those surrounding their model releases, as scare-mongeri…\u003c/li\u003e\n\u003cli\u003eIt was only two months ago that Anthropic announced Mythos Preview, a model that they said was too dangerous to make publicly available, thanks in particular to…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-csai-b_introsearch\"\u003e\n  ArXiv cs.AI (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-csai-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13682\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eA Deep Reinforcement Learning (DRL)-Based Transformer Method for Solving the Open Shop Scheduling Problem\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePosted: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13682v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: The Open Shop Scheduling Problem (OSSP) appears in many industrial and service environments, but remains computationally challenging as the number of jobs and machines increases.\u003c/li\u003e\n\u003cli\u003eWhile exact methods quickly become intractable, classical dispatching rules and metaheuristics may require substantial tuning to maintain solution quality at scale.\u003c/li\u003e\n\u003cli\u003eThis study develops a Transformer-based scheduling policy for OSSP using an encoder-decoder architecture with multi-head attention.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13682v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: The open shop scheduling problem (OSSP) arises in many industrial and service settings but remains computationally challenging as the number of jobs a…\u003c/li\u003e\n\u003cli\u003eWhile exact methods quickly become intractable, classical dispatching rules and metaheuristics may require substantial tuning to maintain solution quality at la…\u003c/li\u003e\n\u003cli\u003eThis study develops a Transformer-based scheduling policy for OSSP using an encoder-decoder architecture with multi-head attention\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13683\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eUP-NRPA: User Portrait based Nested Rollout Policy Adaptation for Planning with Large Language Models in Goal-oriented Dialogue Systems\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePosted: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13683v1 Announcement Type: new.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: To address the challenge that current dialogue policy planning methods struggle to dynamically adapt to diverse user characteristics, this paper proposes an online framework, User Portrait-based Nested Rollout Policy Adaptation (UP-NRPA), which features large language models.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eIn contrast to conventional approaches that depend on model training and require offline reinforcement learning policy models for user groups, UP-NRPA enables the dynamic customization of dialogue policies through an adaptive mechanism.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eThis is achieved by leveraging real-time user feedback alongside personality, preferences, and objectives mapped from the current user portrait, thereby adapting to user characteristics without the need for offline reinforcement learning.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eEN Highlights:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13683v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: To address the challenge that current dialogue policy planning methods struggle to dynamically adapt to diverse user characteristics, this paper propo…\u003c/li\u003e\n\u003cli\u003eIn contrast to conventional approaches dependent on model training and require offline reinforcement learning policy models for user groups, UP-NRPA enables dyn…\u003c/li\u003e\n\u003cli\u003eThis is achieved by leveraging real-time user feedback alongside personality, preferences, and objectives mapped from the current user portrait, thereby adaptin…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13703\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eHistory of the Muddy Children Puzzle\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13703v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: The Muddy Children Puzzle is a puzzle about knowledge and ignorance that has been inspiring for the development of epistemic logic.\u003c/li\u003e\n\u003cli\u003eWe trace the origins of the Muddy Children Puzzle through logical and literary publications of the past two centuries.\u003c/li\u003e\n\u003cli\u003eThe puzzle has inspired many variations, for example involving numbers or colored hats.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13703v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: The Muddy Children Puzzle is a puzzle about knowledge and ignorance that has been inspiring for the development of epistemic logic\u003c/li\u003e\n\u003cli\u003eWho came up with it first\u003c/li\u003e\n\u003cli\u003eThis is unclear\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13707\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eOrchestra-o1: Omnimodal Agent Orchestration\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13707v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: The recent success of agent swarms has shifted the paradigm of large language model (LLM)-based agents from single-agent workflows to multi-agent systems, highlighting the importance of agent orchestration for task decomposition and collaboration.\u003c/li\u003e\n\u003cli\u003eHowever, existing orchestration frameworks are limited to a narrow set of modalities and struggle to generalize to more complex settings where heterogeneous modalities coexist and interact.\u003c/li\u003e\n\u003cli\u003eThis limitation becomes particularly pronounced in omnimodal scenarios, where tasks require a unified understanding and coordination of diverse inputs such as text, images, audio, and video.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13707v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: The recent success of agent swarms has shifted the paradigm of large language model (LLM)-based agents from single-agent workflows to multi-agent syst…\u003c/li\u003e\n\u003cli\u003eHowever, existing orchestration frameworks are limited to a narrow set of modalities and struggle to generalize to more complex settings where heterogeneous mod…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eThis limitation becomes particularly pronounced in omnimodal scenarios, where tasks require the unified understanding and coordination of diverse inputs such as…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13710\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eHybrid Open-Ended Tri-Evolution Makes Better Deep Researcher\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13710v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Deep research and agent evolution are practical tasks for AI agents in real-world applications of artificial general intelligence.\u003c/li\u003e\n\u003cli\u003eThe former can autonomously retrieve and integrate information in an open environment to solve open-ended research tasks, but it is limited by the static parameterization of the agent system\u0026rsquo;s deep research capabilities.\u003c/li\u003e\n\u003cli\u003eThe latter allows agents to autonomously interact with the environment to gain experience for developing model capabilities.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13710v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Deep research and agent evolution serve as de-facto tasks for AI agents in real-world applications toward artificial general intelligence\u003c/li\u003e\n\u003cli\u003eThe former enables autonomous retrieval and integration of information in open-ended environments to tackle open-ended research tasks, yet it is constrained by…\u003c/li\u003e\n\u003cli\u003eThe latter allows agents to autonomously interact with the environment to gain experiences that evolve model capabilities\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13715\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eWorkBench Revisited: Workplace Agents Two Years On\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13715v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: In March 2024, the best agent on WorkBench, GPT-4, completed 43% of tasks and took unintentional harmful actions on 26% of them, such as sending emails to the wrong person.\u003c/li\u003e\n\u003cli\u003eWe revisited the benchmark in June 2026 and found that the best agent to date, Claude Opus 4.8, completes 89% of tasks and takes unintended harmful actions on 2.5% of them.\u003c/li\u003e\n\u003cli\u003eAside from the significant progress in frontier agent performance, three things are noteworthy.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13715v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: The best agent on WorkBench in March 2024, GPT-4, completed 43% of tasks and took an unintended harmful action, such as emailing the wrong person, on…\u003c/li\u003e\n\u003cli\u003eWe re-visit the benchmark in June 2026 and find that the best agent to date, Claude Opus 4.8, completes 89% and takes an unintended harmful action on 2.5%\u003c/li\u003e\n\u003cli\u003eAside from this considerable progress in frontier agent performance, three things stand out\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13720\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eRefusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13720v1 Announcement Type: New.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e(2024) has shown that refusal in safety fine-tuned chat models is mediated by a single linear direction in the residual stream, recoverable by a difference in means (DiM) of harmful and harmless activations.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe compare DiM-based interventions (activation addition and directional ablation) with two interventions derived from Iterative Nullspace Projection (INLP) (nullspace projection and counterfactual flipping) on five open-weight chat models, asking whether INLP can match DiM in turning off refusal and whether its richer parameterization yields more tunable interventions.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eINLP counterfactual flipping is competitive with DiM directional ablation in refusal suppression, while nullspace projection is consistently weaker.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13720v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Arditi et al\u003c/li\u003e\n\u003cli\u003e(2024) has shown that refusal in safety fine-tuned chat models is mediated by a single linear direction in the residual stream, recoverable by a difference-in-m…\u003c/li\u003e\n\u003cli\u003eWe compare DiM-based interventions (activation addition and directional ablation) with two interventions derived from Iterative Nullspace Projection (INLP) \u0026ndash; n…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13722\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eYeasierAgent: Agentic Social Sandbox as a Canvas for Intent-Driven Creation of Platform-Agnostic Symbiotic Agent-Native Applications\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:- arXiv:2606.13722v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: This paper introduces YeasierAgent, an application-building paradigm based on symbiotic agents, narrative worlds, and scene-aware interaction.\u003c/li\u003e\n\u003cli\u003eIt challenges the conventional device-coupled software model by redefining applications as collaborative spaces among users, agents, and worlds.\u003c/li\u003e\n\u003cli\u003eWe present a system architecture that achieves two primary contributions: (1) enabling the rapid, cross-platform construction of agent-native applications by utilizing platform-agnostic interaction units (agents, scenes, conversations) rather than fixed graphical layouts; (2) unifying the emotional companionship and practical tool execution attributes of intelligent agents within a single experiential sandbox.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13722v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: This paper introduces YeasierAgent, an application-building paradigm based on symbiotic agents, narrative worlds, and scene-aware interaction\u003c/li\u003e\n\u003cli\u003eIt challenges the conventional device-coupled model of software by redefining applications as collaborative spaces among users, agents, and worlds\u003c/li\u003e\n\u003cli\u003eWe present a system architecture that achieves two primary contributions: (1) enabling the rapid, cross-platform construction of agent-native applications by ut…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13731\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eTwinBI: An Agentic Digital Twin for Efficient Augmented Interactions with Business Intelligence Dashboards\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:- arXiv:2606.13731v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Business Intelligence (BI) increasingly combines dashboard interaction with LLM-based assistance, but these two modalities are often out of sync during multi-step analytical processes.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWhen users switch between direct dashboard manipulation and natural language queries, it becomes difficult to maintain a consistent analytical state across filters, hierarchies, metrics, and chart contexts.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe introduce TwinBI, an agentic digital twin framework that combines an LLM-based agent system with an executable BI dashboard state.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13731v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Business intelligence (BI) increasingly combines dashboard interaction with LLM-based assistance, but these two modes often fall out of sync during mu…\u003c/li\u003e\n\u003cli\u003eAs users switch between direct dashboard manipulation and natural-language queries, it becomes difficult to preserve a consistent analytical state across filter…\u003c/li\u003e\n\u003cli\u003eWe present TwinBI, an agentic digital-twin framework that couples an LLM-based agent system with an executable BI dashboard state\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13732\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eWhen Sample Selection Bias Precipitates Model Collapse\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13732v1 Announce Type: new.\u003c/li\u003e\n\u003cli\u003eAbstract: The proliferation of recursive training on synthetic data can alleviate data scarcity but risks model collapse, where repeated training erodes the distribution tails and homogenizes outputs.\u003c/li\u003e\n\u003cli\u003eData selection is widely viewed as a remedy, but its reliability critically depends on the reference distribution used by the verifier.\u003c/li\u003e\n\u003cli\u003eWe show that in low-resource verification regimes, where each verifier observes only a small, fragmented, and biased slice of the target manifold, selection itself generates bias.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13732v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: The proliferation of recursive training on synthetic data can alleviate data scarcity but risks model collapse, where repeated training erodes distrib…\u003c/li\u003e\n\u003cli\u003eData selection is widely viewed as a remedy, yet its reliability depends critically on the reference distribution used by the verifier\u003c/li\u003e\n\u003cli\u003eWe show that in low-resource verification regimes, where each verifier observes only a small, fragmented, and biased slice of the target manifold, selection its…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-cscl-b_introsearch\"\u003e\n  ArXiv cs.CL (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-cscl-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13685\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eThe Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13685v1 Announce Type: new.\u003c/li\u003e\n\u003cli\u003eAbstract: LLM-as-a-Judge is now widely used for ranking model outputs, training reward models, and populating public leaderboards, yet its inter-run reliability remains underexplored.\u003c/li\u003e\n\u003cli\u003eWe investigate repeated identical evaluations across 29 tasks spanning 10 categories, using two OpenAI judge models (GPT-4o-mini and GPT-4.1-mini) with 50 pairwise and 50 pointwise trials per question, complemented by temperature and prompt sensitivity ablations.\u003c/li\u003e\n\u003cli\u003eAcross judges, the flip rate for pairwise preferences averaged 13.6%, with 28% of questions exhibiting flip rates over 20%, and one question reaching 56%.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN 要点:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13685v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: LLM-as-a-Judge is now widely used for ranking model outputs, training reward models, and populating public leaderboards, yet its inter-run reliability remains underexplored.\u003c/li\u003e\n\u003cli\u003eWe investigate repeated identical evaluations across 29 tasks spanning 10 categories, using two OpenAI judge models (GPT-4o-mini and GPT-4.1-mini) with 50 pairwise and 50 pointwise trials per question, complemented by temperature and prompt sensitivity ablations.\u003c/li\u003e\n\u003cli\u003eAcross judges, the flip rate for pairwise preferences averaged 13.6%, with 28% of questions exhibiting flip rates over 20%, and one question reaching 56%.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003earXiv:2606.13685v1 Announce Type: new\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eAbstract: LLM-as-a-Judge is now widely used to rank model outputs, train reward models, and populate public leaderboards, but its run-to-run reliability remains…\u003c/li\u003e\n\u003cli\u003eWe study repeated identical evaluations on 29 tasks spanning 10 categories using two OpenAI judge models (GPT-4o-mini and GPT-4.1-mini), with 50 pairwise trials…\u003c/li\u003e\n\u003cli\u003eAcross judges, pairwise preferences flip on average 13.6% of the time, with 28% of questions exceeding a 20% flip rate and one question reaching 56%\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13686\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eBenchmarking Web Agent Safety under E-commerce Deceptive Interfaces\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:- arXiv:2606.13686v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: As more and more autonomous web agents are deployed to perform real-world tasks, ensuring their safety has become a critical concern.\u003c/li\u003e\n\u003cli\u003eIn this work, we study web agent behavior under realistic deceptive interfaces in the e-commerce domain.\u003c/li\u003e\n\u003cli\u003eWe introduce WebDecept, a lightweight and configurable plugin framework that enables controlled injection of deceptive interface patterns into existing Web envi…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13686v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: As autonomous web agents are increasingly deployed to perform real-world tasks, ensuring their safety has become a critical concern\u003c/li\u003e\n\u003cli\u003eIn this work, we study web agent behavior under realistic deceptive interfaces in the e-commerce domain\u003c/li\u003e\n\u003cli\u003eWe introduce WebDecept, a lightweight and configurable plugin framework that enables controlled injection of deceptive interface patterns into existing web envi…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13751\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eWhich Models Perform Better in Inheritance Reasoning?\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract:- arXiv:2606.13751v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: This paper presents the participation of team PSL in the QIAS 2026 Shared Task on Arabic Islamic inheritance reasoning.\u003c/li\u003e\n\u003cli\u003eThis task evaluates the ability of large language models to solve inheritance cases requiring legal interpretation, multi-step reasoning, and precise numerical calculation.\u003c/li\u003e\n\u003cli\u003eWe compare \\textit{commercial} and \\textit{open-source} models under a unified prompting strategy to evaluate their effectiveness in structured legal reasoning and minimize task-specific adaptation where possible.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13751v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: This paper presents the participation of team PSL in the QIAS 2026 Shared Task on Arabic Islamic inheritance reasoning\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eThe task evaluates the ability of large language models to solve inheritance cases that require legal interpretation, multi-step reasoning, and precise numerical…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe compare \\textit{commercial} and \\textit{open-source} models under a unified prompting strategy to assess their effectiveness in structured legal reasoning wi…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13756\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eQIAS 2026: Overview of the Shared Task on Islamic Inheritance Reasoning\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13756v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: This paper provides a comprehensive overview of the QIAS 2026 shared task, organized as part of the OSACT7 Workshop and co-located with LREC 2026.\u003c/li\u003e\n\u003cli\u003eThe shared task is designed to evaluate the ability of large language models to perform complex reasoning in the religious and legal domain of Islamic inheritance.\u003c/li\u003e\n\u003cli\u003eUnlike traditional question-answering benchmarks, QIAS 2026 focuses on end-to-end reasoning from natural language cases, requiring systems to perform the full inheritance calculation process, from identifying qualified heirs to assigning the correct shares for each beneficiary.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13756v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: This paper presents a comprehensive overview of the QIAS 2026 shared task, organized as part of the OSACT7 Workshop and co-located with LREC 2026\u003c/li\u003e\n\u003cli\u003eThe shared task was designed to evaluate the ability of large language models to perform complex reasoning in the religious and legal domain of Islamic inherita…\u003c/li\u003e\n\u003cli\u003eUnlike conventional question-answering benchmarks, QIAS 2026 focuses on end-to-end reasoning from natural language cases, requiring systems to perform the full…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13808\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eThe Culture Funnel: You Can\u0026rsquo;t Align What isn\u0026rsquo;t in the Data\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13808v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Current cultural alignment methods focus on inference-time interventions, assuming that models already contain sufficient cultural knowledge.\u003c/li\u003e\n\u003cli\u003eWe argue that modern LLM pipelines are impacted by a cultural data funnel.\u003c/li\u003e\n\u003cli\u003eUsing a multi-dimensional tagging framework across pre-training, fine-tuning, alignment, and reasoning datasets, we find that explicit cultural signals decline sharply post-training, while geographically concentrated, task-specific data becomes dominant.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13808v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Current cultural alignment approaches focus on inference-time interventions, assuming models already contain sufficient cultural knowledge\u003c/li\u003e\n\u003cli\u003eWe argue modern LLM pipelines suffer from a cultural data funnel\u003c/li\u003e\n\u003cli\u003eUsing a multidimensional tagging framework across pretraining, fine-tuning, alignment, and reasoning datasets, we show explicit cultural signals decline sharply…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13835\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eWhen Plausible Is Not Realistic: Evaluating Human Mobility in LLM-Based Urban Simulation\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13835v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: LLM-based generative agents are increasingly used in urban simulators, but it remains unclear whether they reproduce empirically realistic human mobility patterns or merely generate plausible mobility narratives.\u003c/li\u003e\n\u003cli\u003eWe introduce a validation framework for evaluating the mobility of generative agents of LLM-based urban simulators against real-world mobility data.\u003c/li\u003e\n\u003cli\u003eFor this, we use mobility laws, temporal rhythms, network motifs, semantic activity transitions, and behavioral mobility profiles.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13835v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: LLM-based generative agents are increasingly used in urban simulators, yet it remains unclear whether they reproduce empirically realistic human mobil…\u003c/li\u003e\n\u003cli\u003eWe introduce a validation framework for evaluating the mobility of generative agents of LLM-based urban simulators against real-world mobility data\u003c/li\u003e\n\u003cli\u003eFor this, we use mobility laws, temporal rhythms, network motifs, semantic activity transitions, and behavioral mobility profiles\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13852\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eHybrid Classical-Quantum Variational Autoencoder for Neural Topic Modeling\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13852v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Neural topic models enable scalable semantic discovery, but their integration with quantum hardware remains largely unexplored.\u003c/li\u003e\n\u003cli\u003eWe present a proof-of-concept hybrid classical-quantum variational autoencoder (VAE) for topic modeling, embedding parameterized quantum circuits within the VAE inference network, while retaining a classical topic-word decoder.\u003c/li\u003e\n\u003cli\u003eTo address the resource constraints of quantum hardware, we propose a modified Gaussian Softmax posterior that decouples latent space dimensionality from the number of topics to be extracted, enabling the model to run on low-resource 10-qubit quantum devices.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13852v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Neural topic models enable scalable semantic discovery, but their integration with quantum hardware remains largely unexplored\u003c/li\u003e\n\u003cli\u003eWe present a proof-of-concept hybrid classical-quantum variational autoencoder (VAE) for topic modeling, embedding parameterized quantum circuits within the VAE…\u003c/li\u003e\n\u003cli\u003eTo address the resource constraints of quantum hardware, we propose a modified Gaussian Softmax posterior that decouples latent space dimensionality from the nu…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13904\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eSANA: What Matters for QA Agents over Massive Data Lakes?\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13904v1 Announcement Type: new.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: Exploratory Question Answering (EQA) over data lakes requires an LLM agent to discover relevant sources, analyze retrieved data, and adapt its actions based on intermediate results.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eEnd-to-end accuracy alone cannot distinguish failures in search, planning, data analysis, or the agent\u0026rsquo;s action policy: it determines what to do next and when to submit an answer.\u003c/li\u003e\n\u003cli\u003eWe present SANA (Search Agent Navigation Ablation framework), a diagnostic ablation framework that transforms EQA tasks into runtime profiles containing gold source sequences, sanitized sub-questions, and execution records.\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13904v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Exploratory question answering (EQA) over data lakes requires an LLM agent to discover relevant sources, analyze retrieved data, and adapt its actions…\u003c/li\u003e\n\u003cli\u003eEnd-to-end accuracy alone cannot distinguish failures in search, planning, data analysis, or the agent\u0026rsquo;s Action Policy: its decisions about what to do next and…\u003c/li\u003e\n\u003cli\u003eWe present SANA (Search Agent Navigation Ablation framework), a diagnostic ablation framework that transforms EQA tasks into runtime profiles containing gold so…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13931\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDLawBench: Evaluating LLMs Through Multi-Turn Legal Consultation\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13931v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Lawyer-client consultation is a critical starting point for legal services.\u003c/li\u003e\n\u003cli\u003eEffective legal assistance hinges on eliciting sufficient and truthful information from clients in order to devise strategies that best protect their interests.\u003c/li\u003e\n\u003cli\u003eThis task requires Large Language Models (LLMs) not only to perform robust legal reasoning, but also to strategically elicit material facts through multi-turn interactions, and effectively guide clients with diverse temperaments.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13931v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Lawyer-client consultation is a critical starting point for legal services\u003c/li\u003e\n\u003cli\u003eEffective legal assistance hinges on eliciting sufficient and truthful information from clients in order to devise strategies that best protect their interests\u003c/li\u003e\n\u003cli\u003eThis task requires Large Language Models (LLMs) not only to perform robust legal reasoning, but also to strategically elicit material facts through multi-turn i…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13940\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eCan Post-Training Turn LLMs into Good Medical Coders? An Empirical Study of Generative ICD Coding\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13940v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Automated International Classification of Diseases (ICD) coding is a core medical coding task for billing, epidemiology, and clinical decision support.\u003c/li\u003e\n\u003cli\u003eGenerative Large Language Models (LLMs) are often considered weaker medical coders, but this finding primarily stems from inference-time settings such as prompting, retrieval, re-ranking, or tool usage, thus the role of post-training for task-specific adaptation has not been fully explored.\u003c/li\u003e\n\u003cli\u003eWe propose a controlled empirical study on post-training for generative ICD coding, comparing discriminative baselines with LLM coders across prompting, supervised fine-tuning, and reinforcement learning under a common protocol and metric set.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Key Points:\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003earXiv:2606.13940v1 Announce Type: new\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAbstract: Automated International Classification of Diseases (ICD) coding is a core medical-coding task for billing, epidemiology, and clinical decision support\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eGenerative large language models (LLMs) are often reported as weak medical coders, but this finding mainly comes from inference-time settings such as prompting,…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe present a controlled empirical study of post-training for generative ICD coding, comparing discriminative baselines with LLM coders across prompting, supervi…\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"arxiv-cslg-b_introsearch\"\u003e\n  ArXiv cs.LG (B_intro+search)\n  \u003ca class=\"heading-link\" href=\"#arxiv-cslg-b_introsearch\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13705\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eCan Editing 1 Neuron Fix Repetition Loops in LLMs?\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13705v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eThe Gemma 4 instruction-tuned model has a reproducible failure: on longer factual enumeration prompts, such as listing every episode of a TV show, the 88 IAU constellations, or the 151 original Pokemon, they get stuck in repetition, either a strict verbatim loop or a list whose entries decay to a single answer.\u003c/li\u003e\n\u003cli\u003eThese loops occur with up to 95% incidence and survive prompt rewording, inference engine changes, and most sampling tweaks.\u003c/li\u003e\n\u003cli\u003eIn this paper, we explore whether this behavior is localized enough to be removed with weight editing.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13705v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Yes\u003c/li\u003e\n\u003cli\u003eCan it cure doom loops\u003c/li\u003e\n\u003cli\u003eProbably not\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13740\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eEfficient On-Device Diffusion LLM Inference with Mobile NPU\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13740v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Diffusion large language models (dLLMs) accelerate generation by denoising multiple tokens in parallel, making them attractive for latency-sensitive mobile inference.\u003c/li\u003e\n\u003cli\u003eHowever, repeated denoising introduces substantial computation on smartphones.\u003c/li\u003e\n\u003cli\u003eMobile neural processing units (NPUs) offer high-throughput dense matrix computation, but efficiently exploiting them remains challenging: token commitment shrinks the effective workload per block, token revision complicates KV cache reuse, and the limited NPU-visible address space causes expensive remapping and data transfer overhead.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13740v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Diffusion large language models (dLLMs) accelerate generation by denoising multiple tokens in parallel, making them attractive for latency-sensitive m…\u003c/li\u003e\n\u003cli\u003eHowever, repeated denoising introduces substantial computation on smartphones\u003c/li\u003e\n\u003cli\u003eMobile neural processing units (NPUs) offer high-throughput dense matrix computation, but efficiently exploiting them remains challenging: token commitment shri…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13741\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eHigh-Frequency Pricing at Scale for E-Commerce\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eSummary: - arXiv:2606.13741v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: This paper presents the design, development, and implementation of a specialized forecast-then-optimize algorithmic pricing tool for fashion e-commerce sales campaigns.\u003c/li\u003e\n\u003cli\u003eSales events present unique challenges for pricing, including volatile demand patterns, rapid pricing decisions, and the need to balance short-term revenue with long-term profitability.\u003c/li\u003e\n\u003cli\u003eWe describe our approach, which combines daily-resolution demand forecasting using gradient-boosted trees with a multi-objective optimization framework that maximizes long-term profit and net merchandise value for over 5 million items.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13741v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: This paper presents the design, development, and implementation of a specialized forecast-then-optimize algorithmic pricing tool for sales campaigns i…\u003c/li\u003e\n\u003cli\u003eSales events present unique challenges for pricing including volatile demand patterns, rapid pricing decisions, and the need to balance short-term revenue with…\u003c/li\u003e\n\u003cli\u003eWe describe our approach combining daily-resolution demand forecasting using gradient-boosted trees with a multi-objective optimization framework that maximizes…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13742\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eA fully GPU-based workflow for building physics emulators of hypersonic flows\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary: - arXiv:2606.13742v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: The ability to resolve complex physical phenomena with high fidelity and at low computational cost is central to addressing key challenges in modern engineering.\u003c/li\u003e\n\u003cli\u003eA prime example lies in hypersonic flows, where the precise prediction of the full flowfield topology, particularly with respect to shock wave location and intensity, is crucial.\u003c/li\u003e\n\u003cli\u003eYet supersonic and hypersonic flows continue to be a stumbling block for traditional reduced-order models and neural emulators that struggle to capture steep gradients of flow states with physical consistency in industrially relevant applications.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13742v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: The ability to resolve complex physical phenomena with high fidelity and at low computational cost is central to addressing key challenges in modern e…\u003c/li\u003e\n\u003cli\u003eA prime example lies in hypersonic flows, where the precise prediction of the full flowfield topology, in particular with respect to shock wave location and int…\u003c/li\u003e\n\u003cli\u003eYet supersonic and hypersonic flows continue to be a stumbling block for traditional reduced-order models and neural emulators that struggle to capture steep gr…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13748\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eFedSPC: Shared Parameter Correction for Personalized Federated Learning\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eSummary: - arXiv:2606.13748v1 Announcement Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Personalized Federated Learning (PFL) is one of the important methods in federated learning to address statistical heterogeneity while achieving client-specific adaptation.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eMany PFL methods divide models into shared and personalized parameters, which are jointly trained on each client.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eHowever, this creates an optimization problem: shared parameters are updated by clients optimizing different local objectives, which can lead to inconsistent shared updates and weaken the shared representation.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13748v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Personalized federated learning (PFL) is one of the important approaches in federated learning for addressing statistical heterogeneity while enabling…\u003c/li\u003e\n\u003cli\u003eMany PFL methods split the model into shared and personalized parameters, which are jointly trained on each client\u003c/li\u003e\n\u003cli\u003eHowever, this creates an optimization issue: shared parameters are updated by clients optimizing different local objectives, which can lead to inconsistent shar…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13753\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eThe Weight Norm Sets the Grokking Timescale: A Causal Delay Law\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eRelease Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13753v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Grokking is the delayed onset of generalization in neural networks, arising long after they fit the training data.\u003c/li\u003e\n\u003cli\u003eWhether the weight norm causes this delay is disputed: some studies report a critical norm at the transition, while others observe grokking with no fixed norm at all.\u003c/li\u003e\n\u003cli\u003eWe settle this by intervening on the norm during training rather than only observing it.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13753v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Grokking is the delayed onset of generalization in neural networks, arising long after they fit the training data\u003c/li\u003e\n\u003cli\u003eWhether the weight norm causes this delay is disputed: some studies report a critical norm at the transition, others observe grokking with no fixed norm at all\u003c/li\u003e\n\u003cli\u003eWe settle this by intervening on the norm during training rather than only observing it\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13754\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eD2H-AD: A Hybrid Model Utilizing Hyperdimensional Computing for Advanced Anomaly Detection\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eRelease Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13754v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Anomaly detection is a fundamental component of intelligent systems, applied in healthcare, cybersecurity, smart grids, and IoT environments.\u003c/li\u003e\n\u003cli\u003eAlthough traditional machine learning and deep learning methods have proven effective in identifying anomalies, they often rely on large labeled datasets, incur high computational costs, and face scalability challenges in edge and high-dimensional settings.\u003c/li\u003e\n\u003cli\u003eThis paper introduces D2H-AD, a novel anomaly detection framework based on Hyperdimensional Computing (HDC), a brain-inspired paradigm that represents information using high-dimensional distributed vectors.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13754v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Anomaly detection is a fundamental component of intelligent systems with applications in healthcare, cybersecurity, smart grids, and IoT environments\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAlthough conventional machine learning and deep learning methods have demonstrated effectiveness in identifying anomalies, they often rely on large labeled data…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eThis paper presents D2H-AD, a novel anomaly detection framework based on Hyperdimensional Computing (HDC), a brain-inspired paradigm that represents information…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13767\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eBeyond LoRA: Is Sparsity-Induced Adaptation Better?\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13767v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: Low-rank adaptation (LoRA) and its variants provide a memory- and compute-efficient alternative to the full fine-tuning of pre-trained models.\u003c/li\u003e\n\u003cli\u003eHowever, questions remain about the comparative generalizability of these approaches and how the structural restrictions on low-rank updates preserve effective adaptation performance.\u003c/li\u003e\n\u003cli\u003eWe present a historical framing, covering the past (full fine-tuning and original LoRA), the present (different variants of LoRA), and propose simpler, cheaper, and more parameter-efficient extensions by introducing sparsity into existing LoRA variants: cheap LoRA (cLA), which trains a single low-rank factor with another fixed one (deterministic, or random in stochastic variants), and a chained-cycle variant, ${c}^3$LA.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13767v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Low-rank adaptation (LoRA) and its variants provide a memory- and compute-efficient alternative to full fine-tuning of pre-trained models\u003c/li\u003e\n\u003cli\u003eHowever, questions remain about the comparative generalizability of these approaches and how the structural restrictions on low-rank updates preserve effective…\u003c/li\u003e\n\u003cli\u003eWe present a historical framing, covering the past (full fine-tuning and original LoRA), the present (different variants of LoRA), and propose simpler, cheaper,…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13795\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eDiffusion Policy Optimization without Drifting Apart\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13795v1 Announcement Type: New.\n\u003cul\u003e\n\u003cli\u003eAbstract: RL post-training has become increasingly pivotal for improving diffusion policies, but existing diffusion policy-gradient methods are often unstable and fail to achieve reliable policy improvement.\u003c/li\u003e\n\u003cli\u003eWe identify the cause as a dual-drift phenomenon: optimizing the variational surrogate can allow the ELBO to separate from the true log-likelihood, causing the resulting surrogate policy gradient to be inconsistent with the true policy gradient of the expected return.\u003c/li\u003e\n\u003cli\u003eWe propose \\textbf{DiPOD}, a diffusion policy optimization framework that maintains tight-bound behavior throughout the training process by intertwining self-distillation with policy improvement gradient updates.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13795v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: RL post-training has become increasingly pivotal for improving diffusion policies, but existing diffusion policy-gradient methods are often unstable a…\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe identify the cause as the double-drift phenomenon: optimizing a variational surrogate can let the ELBO separate from the true log-likelihood, which then make…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWe propose \\textbf{DiPOD}, a diffusion policy optimization framework that maintains tight-bound behavior throughout training by interleaving self-distillation w…\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e\u003ca href=\"https://arxiv.org/abs/2606.13801\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003eNeural Variability Enhances Artificial Network Robustness\u003c/a\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003ePublication Time: 2026-06-15 12:00 Beijing Time\u003c/li\u003e\n\u003cli\u003eAbstract: - arXiv:2606.13801v1 Announce Type: new.\n\u003cul\u003e\n\u003cli\u003eAbstract: Neural responses in the cortex exhibit significant trial-to-trial variability in response to repeated stimuli, while the responses of peripheral sensory neurons are more consistent, leading many to question whether this stochasticity might be meaningful.\u003c/li\u003e\n\u003cli\u003eExisting work suggests that noise and signal correlations can be optimized for discrimination in animals, while research in artificial neural networks (ANNs) indicates that noise provides similar benefits in machine learning tasks, although most ANN studies have overlooked the impact of correlations.\u003c/li\u003e\n\u003cli\u003eHere, we investigate whether correlated noise improves the robustness of artificial neural networks to adversarial attacks and naturalistic image modifications.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003eEN Highlights:\n\u003cul\u003e\n\u003cli\u003earXiv:2606.13801v1 Announce Type: new\u003c/li\u003e\n\u003cli\u003eAbstract: Neural responses in cortex exhibit substantial trial-to-trial variability in response to repeated stimuli, while peripheral sensory neurons respond fa…\u003c/li\u003e\n\u003cli\u003eExisting work has argued that noise and signal correlations may be optimized for discrimination in animals, whereas artificial neural network (ANN) studies have…\u003c/li\u003e\n\u003cli\u003eHere we investigate whether correlated noise improves the robustness of artificial neural networks to adversarial attacks and naturalistic image modifications\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n",
  "wordCount": 7707,
  "readingTime": 37,
  "tableOfContents": "\u003cnav id=\"TableOfContents\"\u003e\n  \u003cul\u003e\n    \u003cli\u003e\u003ca href=\"#-in-depth-guide-to-this-issues-watch-list\"\u003e📖 In-depth Guide to This Issue\u0026rsquo;s Watch List\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-ai-hot-topics-on-x\"\u003e🌐 AI Hot Topics on X\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#topic-1-ai-coding-agents-speed-past-human-review-limits\"\u003eTopic 1: AI Coding Agents Speed Past Human Review Limits\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-2-factory-ai-launches-factory-20-as-fully-autonomous-software-factories\"\u003eTopic 2: Factory AI Launches Factory 2.0 as Fully Autonomous Software Factories\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-3-vercel-extends-serverless-functions-to-30-minutes-for-ai-workloads\"\u003eTopic 3: Vercel Extends Serverless Functions to 30 Minutes for AI Workloads\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-4-le-chaton-fat-mistral-ais-massive-meme-hoax\"\u003eTopic 4: Le Chaton Fat: Mistral AI\u0026rsquo;s Massive Meme Hoax\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#topic-5-ai-coded-wow-clone-draws-12000-players-in-days\"\u003eTopic 5: AI-Coded WoW Clone Draws 12,000 Players in Days\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-influencer-insights\"\u003e💡 Influencer Insights\u003c/a\u003e\u003c/li\u003e\n  \u003c/ul\u003e\n\n  \u003cul\u003e\n    \u003cli\u003e\u003ca href=\"#1-todays-tech-trends-and-product-hotspots\"\u003e1. Today\u0026rsquo;s Tech Trends and Product Hotspots\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#-hot-topic-the-ultimate-showdown-and-capability-limits-of-on-device-models\"\u003e🔥 Hot Topic: The \u0026ldquo;Ultimate Showdown\u0026rdquo; and Capability Limits of On-Device Models\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#-tools-and-platforms-claude-design-sparks-a-design-paradigm-shift\"\u003e🛠️ Tools and Platforms: Claude Design Sparks a Design Paradigm Shift\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#2-unique-perspectives-and-industry-foresight\"\u003e2. Unique Perspectives and Industry Foresight\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#-a-new-concept-token-capital\"\u003e🧠 A New Concept: \u0026ldquo;Token Capital\u0026rdquo;\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#-counter-intuitive-observations-and-deep-reflections\"\u003e📉 Counter-intuitive Observations and Deep Reflections\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#3-recommended-tools-and-resources\"\u003e3. Recommended Tools and Resources\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#-appendix-todays-watch-list-update-sources\"\u003e📚 Appendix: Today\u0026rsquo;s Watch List Update Sources\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#a16z-podcast-a_full\"\u003ea16z Podcast (A_full)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#stratechery-by-ben-thompson-a_full\"\u003eStratechery by Ben Thompson (A_full)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-csai-b_introsearch\"\u003eArXiv cs.AI (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-cscl-b_introsearch\"\u003eArXiv cs.CL (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#arxiv-cslg-b_introsearch\"\u003eArXiv cs.LG (B_intro+search)\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n  \u003c/ul\u003e\n\u003c/nav\u003e",
  "isDraft": false
}
