System translated (Gemini)

🤖 AI 速览

Today’s main theme is agent engineering: local-first assistants, long-term memory, sandbox concurrency, and robot deployment collectively complete the execution layer. Model competition continues in coding capabilities and underlying architecture, but enterprises are more concerned with ROI, …
📋 文章元数据
发布时间
2026-07-18
类型
ai-daily
字数
8042
阅读时长
38 min

2026-07-18 AI Daily | Agents Enter the Infrastructure War: Local Assistants, Million-Container Sandboxes, and World Models Emerge Link to heading

Today’s main theme is the engineering of agents: local-first assistants, long-term memory, sandbox concurrency, and robotics deployment are collectively completing the execution layer. Model competition continues in coding capabilities and underlying architecture, but enterprises are more focused on ROI, data provenance, and reliability assessment.

📖 This Issue’s Watch List: In-Depth Guide Link to heading

The most important reading today is the “Agent Implementation” group: OpenClaw’s local assistant practices, Oracle’s long-term agent memory, the SPINE deployment framework for bimanual robots, and the survey on self-improving agents. Together, they point to a trend: agent competition is shifting from model capabilities to memory, tools, physical execution, and continuous adaptation.

The second main theme is “AI Value and Governance.” “A scorecard for the AI age” re-examines AI ROI from a CFO’s perspective, while OriginBlame advances training data provenance to the record and token level, making it a detailed read for teams focused on corporate compliance, right-to-be-forgotten requests, and cost accounting.

Finally, we recommend several papers on evaluation and safety: auditing premise dependencies in CoT, stability testing for multi-turn refutation in VLMs, and safety training based on human preferences and world models. These are all practical methodologies for determining if a model is “truly reliable.”

🌐 X Platform AI Hot Topics Link to heading

Topic 1: Fake Claim of Chinese AI Distilling Anthropic Model Spreads on X Link to heading

  • Category: AI · News
  • Overview: Trending time:, Related posts: 121
  • What it is: A claim circulated on X about a “Chinese AI company distilling an Anthropic model,” but the claim is reported to lack credible evidence or be misleading.
  • Why it matters: This reflects the growing sensitivity of disputes over data sources, model distillation, and intellectual property in the large model competition. It also highlights how misinformation in the AI field can impact corporate reputations and regulatory discussions.
  • Discussion summary: The discussion focuses on whether there is empirical evidence for the accusation, whether model similarity is sufficient to prove distillation, and whether the technological competition between Western and Chinese AI companies is being politicized. Some also call for verifiable evidence before spreading such allegations.

Topic 2: Moonshot AI’s Kimi K3 Tops Frontend Code Arena Leaderboard Link to heading

  • Category: AI · News
  • Overview: Trending time: 1 day ago, Related posts: 114,000
  • What it is: Moonshot AI’s Kimi K3 has reached the top of the Frontend Code Arena leaderboard, attracting attention from the developer community.
  • Why it matters: This demonstrates the improved competitiveness of Chinese large models in front-end code generation, UI implementation, and engineering tasks. It also reflects that coding ability is becoming a key metric for evaluating the practical value of AI models.
  • Discussion summary: Discussions on X are mainly focused on the real coding capabilities of Kimi K3, whether the leaderboard evaluation represents actual development experience, and the gap between it and models like Claude, GPT, and Gemini in terms of front-end generation quality, stability, and usability.

Topic 3: Boris Cherny Maps Path to 1,000-Agent AI Coding Teams Link to heading

  • Category: AI · News
  • Overview: Trending time: 4 hours ago, Related posts: 306
  • What it is: Boris Cherny has proposed a roadmap for expanding AI programming collaboration to “1,000-agent teams,” envisioning a large number of code agents completing software development tasks in parallel.
  • Why it matters: This reflects the trend of AI programming moving from single assistants to multi-agent collaboration, which could change the organization, development efficiency, and management processes of software engineering.
  • Discussion summary: Discussions on X focus on whether multi-agent programming is truly scalable, how to address code quality and coordination costs, and whether such systems will augment or replace human developers.

Topic 4: Moonshot AI’s Zhilin Yang Returns to China After CMU PhD, Sparks Talent Debate Link to heading

  • Category: AI · News
  • Overview: Trending time: 17 hours ago, Related posts: 6,300
  • What it is: Moonshot AI co-founder Zhilin Yang returned to China after completing his PhD at Carnegie Mellon University, sparking a discussion about the flow of top AI talent.
  • Why it matters: This move is seen as a signal of Chinese large model startups’ ability to attract top global research talent and enhance their foundational model R&D capabilities. It also reflects the importance of talent and the research ecosystem in the US-China AI competition.
  • Discussion summary: Discussions on X center on whether Chinese AI companies have the environment to retain top talent, whether the return of overseas PhDs will change the technological competition landscape, and whether talent mobility should be viewed as a normal career choice or part of geopolitical tech competition.

Topic 5: Modal Launches Sandbox Platform for 1 Million Concurrent AI Environments Link to heading

  • Category: AI · News
  • Overview:Trending since:22 hours ago,Related posts:274
  • What it is:Cloud computing platform Modal launched a Sandbox platform for AI workloads, claiming it can run 1 million isolated AI environments simultaneously.
  • Why it matters:This indicates that AI applications are moving from single model calls to scenarios involving massively concurrent, isolated execution of agents and code, placing higher demands on inference infrastructure, elastic scheduling, and secure sandbox capabilities.
  • Discussion highlights:Discussions on X focused on whether its concurrent scale is genuinely usable, its cost and latency performance, its differences from existing cloud services and container platforms, and whether this capability will accelerate the adoption of AI Agents, automated programming, and batch experiments.

Topic 6: Higgsfield AI Open-Sources Prompts for Seedance 2.0 Cinematic Videos Link to heading

  • Category:AI · News
  • Overview:Trending since:22 hours ago,Related posts:1500
  • What it is:Higgsfield AI open-sourced a set of prompt templates for generating cinematic videos with Seedance 2.0.
  • Why it matters:This helps lower the barrier to high-quality AI video generation, making it easier for creators to replicate camera language, motion control, and cinematic styles, while also promoting the sharing and standardization of prompt engineering in the video generation field.
  • Discussion highlights:Discussions on X primarily focused on whether these prompts can consistently produce professional-grade results, the performance comparison of Seedance 2.0 with models like Runway, Kling, and Veo, and whether open-sourcing prompts will enhance creative efficiency or exacerbate content homogenization.

Topic 7: Decart AI Launches Lucy 2.5 for Real-Time Video Edits Link to heading

  • Category:AI · News
  • Overview:Trending since:22 hours ago,Related posts:3000
  • What it is:Decart AI released Lucy 2.5, claiming it can perform real-time edits like special effects, outfit changes, scene swaps, and stylization on live video streams using prompts.
  • Why it matters:If real-time generative video editing becomes stable and usable, it will transform video from a fixed product into an interactive, dynamically rewritable content layer, potentially impacting AI application scenarios such as live streaming, e-commerce, advertising, social media, gaming, and streaming.
  • Discussion highlights:Discussions on X centered on whether Lucy 2.5 marks the productization of a “real-time world model” and its impact on creator workflows and commercial content customization. Concurrently, some questioned whether the demos were cherry-picked and if the actual latency and consistency are sufficient to support large-scale production use.

AI Public Opinion Summary on X Today Link to heading

Today’s main narrative revolves around the “rapid productization of AI capabilities” and “intensifying competitive narratives.” From Kimi K3 topping code leaderboards and multi-agent programming to million-scale sandboxes, real-time video editing, and cinematic prompts, the community generally agrees that AI is shifting from single-point model capabilities to engineering, scaling, and the reshaping of creative workflows. The consensus is that Chinese AI companies are gaining significant presence in code, video, and talent attraction, with infrastructure and agent-based execution environments becoming the next focal point of competition. Disagreements center on the authenticity and transferability of these advancements: whether leaderboards reflect real development experience, if demos are cherry-picked, if a million concurrent instances are usable, if multi-agent systems will be offset by coordination costs, and whether model similarities are sufficient to support allegations of “distillation.” Potential risks include unverified technical accusations becoming politicized and damaging corporate reputations, AI video and real-time editing leading to content homogenization, authenticity, and misuse issues, and large-scale agent/sandbox systems potentially exposing new engineering and governance challenges in cost, security isolation, and code quality control.

💡 Influencer Insights Link to heading

Based on tweet data from the last 24 hours, here is your AI Daily report.


AI Daily: Influencer Insights at a Glance (July 17, 2026) Link to heading

1. Core Trend: Model Competition Heats Up, Agent Engineering Becomes the Focus Link to heading

Today’s discussion centers on the fierce competition between two giants and the rapid infrastructure build-out of the Agent ecosystem.

  • Kimi K3 Released, Aiming at Claude Fable 5 The release of Kimi K3 is undoubtedly the biggest flashpoint in the community today. Many bloggers see it as a milestone for domestic models, directly comparing it to and even claiming it surpasses Claude Fable 5.

    • @Pluvio9yte pointedly stated: “Kimi K3 is out, rumored to surpass Opus…” and boldly predicted that Anthropic might once again extend the paid access for Fable 5 as a market defense tactic.
    • @vista8, after a quick test, praised Kimi K3’s “aesthetic sense” and front-end code generation capabilities, calling it “currently the number one domestic model,” and demonstrated its powerful “replicate a website with one sentence” feature.
  • User Benefits from Competition: The competition in model capabilities has directly led to better plans and strategies. Both @Pluvio9yte and @dotey observed that Claude and Codex frequently “reset quotas” around the time of each other’s new model releases. @vista8 summarized this phenomenon by saying, “It really shows the need for antitrust measures; competition benefits the public.”

  • The “Warring States Period” of Agent Frameworks and Developer Tools Agents are no longer just a showcase of model capabilities; a complete developer toolchain and ecosystem are now taking shape.

    • Grok Build Goes Open-Source: Musk’s xAI has open-sourced its terminal AI coding agent, grok-build. Both @AI_Jasonyu and @vista8 were quick to notice this development. @AI_Jasonyu commented that it is feature-complete, stating, “it has basically everything that Claude Code uses,” including MCP, Skills, and a sandbox, and is written in Rust. This injects strong open-source competitiveness into the agent race.
    • The Maturation and Expansion of the Codex Ecosystem: OpenAI’s Codex remains a central topic of discussion, but the focus has shifted to its ecosystem. The official Codex Micro physical keyboard was launched (noted by @dotey), while the community has seen projects for custom theme skins (@vista8) and open-source alternatives using hundred-yuan game controllers to replace the Codex Micro (retweeted by @dotey). This demonstrates developers’ enthusiasm and high level of engagement with the product.

2. Unique Perspectives and Industry Foresight Link to heading

  • The Frontier of AI Programming: Evolving from “Code Generation” to “World Models”

    • @Pluvio9yte published a long article profoundly explaining the next trend in AI video models: not generating longer videos, but building real-time interactive “World Models.” Using the recently open-sourced Alaya World as an example, he pointed out that AI is evolving from a content generation tool into an explorable, interactive environment, which will have a fundamental impact on fields like gaming and embodied intelligence. “AI Is Moving Beyond ‘Generating Videos’ — Toward ‘Generating Worlds’.”
    • Meanwhile, @Pluvio9yte also shared his “Frontend-First” AI development strategy: first, use AI to generate the front-end interface and the front-end/back-end contract. Once the UX and interactions are fully confirmed, complete the back end according to the established contract to avoid the high costs of repeated modifications. This is a pragmatic engineering approach.
  • The Technical Depth Behind Kimi K3: Rethinking the Foundations of Transformers

    • @dotey provided a detailed analysis of the speech by Yang Zhilin, CEO of Moonshot AI, pointing out that its core idea is to “rebuild three fundamental components of AI training that have been in use for nearly a decade”:
      1. Optimizer: Replace Adam with MuonClip to extract the value of each token.
      2. Attention Mechanism: Replace Full Attention with Kimi Linear (KDA) to improve long-context efficiency.
      3. Residual Connection: Propose Attention Residue as the next-generation solution for residual connections, allowing the model to actively select information from previous layers.
    • @dotey believes this indicates that, under the pressure of data walls and costs, redesigning the underlying architecture is a crucial direction for breakthroughs, second only to scaling up parameters.
  • Pragmatic Discussion on On-Device Models and Hardware

    • @ruanyf offered a counter-intuitive judgment: for running AI locally, “mini-PCs with on-board chipsets are often better than those with discrete graphics cards.” He pointed out that an AMD Strix Halo platform with 128GB of unified memory can run large models that a 5090 graphics card with its 32GB VRAM limit cannot, and at a lower cost. This provides a new perspective for developers choosing local devices.
    • @zhixianio shared his frustration when testing Gemma 4 12B Coder, noting that its 12B size ceiling makes it difficult to handle “long, stateful, single-pass” complex programs like Tetris. This highlights the importance of matching model size to task complexity.
    • @dotey strongly recommended using a Mac for developing Agent applications due to its superior ecosystem support, specifically emphasizing, “You need a lot of RAM; even 16GB is too small.”
  • Reflections on Human Competitiveness in the AI Era

    • In response to discussions about GPT 5.6’s poor design capabilities, @dotey pointed out: “AI can raise the floor for ordinary people, but it cannot raise the ceiling; AI can amplify and accelerate professional abilities.” Taste and design skills are still led by humans.
  • @gefei55 criticized the practice of using AI to build 37 apps in a month that no one cares about. He argues that while AI enhances productivity, entrepreneurs who only focus on coding without “researching before building” and “promoting their work” will fall into a “Token Trap,” wasting time and money. He urges developers to practice their marketing skills to become “super-individuals” capable of both production and sales.

3. Tools and Resources to Try Link to heading

  • Grok Build: An open-source terminal AI coding agent from Musk’s xAI. It’s written in Rust, is feature-rich, and supports MCP/Skills/sandbox. Find the GitHub link in tweets by @AI_Jasonyu or @vista8.
  • rn-wechat-extract (Skill): An open-source Claude Code/Codex skill developed by @Pluvio9yte. It solves the problem of agents being unable to directly read content from WeChat Official Account articles by simulating a WeChat user agent to fetch the content and convert it to Markdown. Installation command: npx -y skills add Pluviobyte/rnskill --skill rn-wechat-extract.
  • AnySearch: Recommended by @Pluvio9yte, this is a search infrastructure designed for AI agents. It supports vertical searches in fields like finance and academia and provides structured output, significantly boosting an agent’s research efficiency.
  • grillme (Skill): A project planning Skill highly recommended by @Pluvio9yte. Before coding, it “grills” you with questions to resolve all uncertain requirements, generating a meticulously detailed plan. “Spending an extra 20 minutes to perfect the plan is much more efficient than spending an extra 60 minutes changing modules after the code is finished.”
  • Qwen3 ASR: Recommended by @dotey. Paired with Qwen3-ForcedAligner, it achieves high-accuracy, word-level timestamp transcription. The 0.6B model can run locally, making it an excellent open-source alternative to OpenAI Whisper, especially for creating subtitles.
  • Emil’s UI Design Skill: Recommended by @vista8 as “the one skill to recommend if you want to avoid generic AI designs.” Created by designer Emil, it generates highly aesthetic UI with beautiful animations. Installation command: npx skills add emilkowalski/skill.

📚 Appendix: Today’s Watch List Source Updates Link to heading

Timeframe: Last 3 days; 22 sources covered; 33 updates in total.

Y Combinator Podcast (B_intro+search) Link to heading

  • World Models, Explained
    • Publication Time: 2026-07-17 22:43 Beijing Time
    • Summary: - You’ve probably heard of OpenClaw (formerly Clawdbot/Moltbot).
      • The open-source AI assistant that made a splash runs on your own device, connects with the messaging apps you already use, and goes beyond chat to actually perform tasks like managing email, calendars, files, workflows, and more.
      • Now meet the person behind it.
      • YC’s Raphael Schaad sits down with OpenClaw founder Peter Steinberger to talk about the “aha” moment behind the viral personal AI agent, why a local-first agent could replace many of today’s apps, and how personal agents will reshape the future of software.
    • EN Key Points:
      • Why do even our best AI models need tens of thousands of examples to learn skills that a human picks up in a handful of tries
      • Solving this problem is one of the great open challenges in modern AI
      • World models, which give AI an internal simulation of its environment, are one of the most promising paths forward.In this episode of Decoded, YC’s Ankit Gupta…
      • Full Transcript:

Stratechery by Ben Thompson (A_full) Link to heading

  • 2026.29: Mainframes and Main Characters
    • Publication Time: 2026-07-18 01:00 Beijing Time
  • Summary: - Jean Pigozzi via Andy Hertzfeld on left, GenAI via Muse Image on right.
    • Welcome back to This Week in Stratechery!
    • As a reminder, each week, every Friday, we send out this overview of content in the Stratechery bundle; highlighted links are free for everyone.
    • Additionally, you have complete control over what we send to you.
    • With that said, here are some of our favorites from this week.
    • EN Highlights:
      • Jean Pigozzi via Andy Hertzfeld on left, GenAI via Muse Image on right
      • Welcome back to This Week in Stratechery
      • As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone
      • Additionally, you have complete control over what we send to you

OpenAI Blog (A_full) Link to heading

  • A scorecard for the AI age
    • Published: 2026-07-17 18:00 Beijing Time
    • Summary: - The question I hear from CFOs everywhere is simple: How do we get more value from our AI spending?
      • For years, the market has measured software success by adoption: purchased seats, active users, and renewed licenses.
      • Understanding the value of AI requires a more robust metric: work completed.
      • The fundamental economic question for CFOs and other business leaders is whether the value of the work completed by AI is growing faster than its production costs.
      • Answering this question requires a deeper look than metrics like cost per token.
    • EN Highlights:
      • Sarah Friar, CFO of OpenAI, introduces a practical AI scorecard to measure ROI through useful work, cost per successful task, dependability, and return on compu…

ArXiv cs.AI (B_intro+search) Link to heading

  • OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets

    • Published: 2026-07-17 12:00 Beijing Time
    • Summary: - arXiv:2607.13037v1 Announcement Type: new.
      • Abstract: When data contributors request removal, model trainers face a practical gap: unlearning algorithms require a forget set, yet no tool can locate which training records belong to a given author.
      • Existing provenance systems operate at the file or dataset level, leading to catastrophic over-deletion.
      • We present ob, a record- and token-level data provenance system that propagates authorship through data processing pipelines and resolves revocation requests into precise forget sets via deterministic queries.
    • EN Highlights:
      • arXiv:2607.13037v1 Announce Type: new
      • Abstract: When a data contributor requests removal, model trainers face a practical gap: unlearning algorithms require a forget set, yet no tool can locate whic…
      • Existing provenance systems operate at file or dataset level, forcing catastrophic over-deletion
      • We present ob, a record- and token-level data provenance system that propagates author identity through data processing pipelines and resolves revocation reques…
  • SPINE: Bridging the Cyber-Physical Gap with Agentic AI

    • Publish Time: 2026-07-17 12:00 Beijing Time
    • Abstract:- arXiv:2607.13049v1 Announce Type: new.
      • Abstract: Foundation models provide robots with sophisticated brains for complex decision-making, but deploying this intelligence into physical platforms still requires tedious, expert-driven calibration.
      • This deployment gap, the robot’s spinal cord, remains a major bottleneck for scalable Embodied AI.
      • Therefore, we propose SPINE (Scalable Physical Integration with ageNtic Expertise): an agentic framework for systematically debugging and deploying bimanual robots with minimal robotic expertise.
    • EN 要点:
      • arXiv:2607.13049v1 Announce Type: new
      • Abstract: Foundation models have given robots a sophisticated brain for complex decision-making, yet deploying that intelligence into a physical platform still…
      • This deployment gap, the robot’s spinal cord, remains a primary bottleneck to scalable Embodied AI
      • Hence, we propose SPINE (Scalable Physical Integration with ageNtic Expertise): an agentic framework for systematically debugging and deploying bimanual robots…
  • Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution

    • Publish Time: 2026-07-17 12:00 Beijing Time
    • Abstract:- arXiv:2607.13069v1 Announce Type: new.
      • Abstract: Large language models generate Chain-of-Thought (CoT) reasoning that appears logically sound but may not truly depend on its stated premises.
      • We introduce Interventional Grounding Audits, a black-box, step-level test of premise dependence: we intervene on a single premise by replacing its target predicate with a new symbol, rerun the model, and check whether the normalized conclusion (canonical predicate form) of each reasoning step changes.
      • We evaluate on ProntoQA, a synthetic multi-hop deductive reasoning benchmark with gold proof trees, where step-level premise dependencies are known.
    • EN 要点:
      • arXiv:2607.13069v1 Announce Type: new
      • Abstract: Large language models produce chain-of-thought (CoT) reasoning that appears logically sound yet may not genuinely depend on its stated premises
      • We introduce interventional grounding audits, a black-box, step-level test of premise dependency: we intervene on a single premise by substituting its target pr…
      • We evaluate on ProntoQA, a synthetic multi-hop deductive reasoning benchmark with gold proof trees, where step-level premise dependencies are known
  • Probabilistic Extension of Neuro-Symbolic AGI Robots based on Belnap’s Typed Intensional FOL

    • Publish Time: 2026-07-17 12:00 Beijing Time
    • Abstract:- arXiv:2607.13073v1 Announce Type: new.
  • Abstract: Neuro-symbolic AI based on $IFOL_B$ is a method that combines neural learning and symbolic reasoning to overcome the limitations of purely neural systems (e.g., lack of interpretability and logical structure), and features a formal logic mechanism for self-reference.

  • In this paper, we expand the cognitive capabilities of $IFOL_B$ by performing probability calculations on currently unknown sentences, based on Nilsson’s probabilistic structure for $IFOL_B$.

  • We introduce global symmetry transformations that preserve the current knowledge database and logical derivations, as well as local symmetry transformations for real-time decision-making on specific (sub)problems involving only a very strict subset of $IFOL_B$ predicates.

  • EN Highlights:

    • arXiv:2607.13073v1 Announce Type: new
    • Abstract: Neuro-symbolic AI based on $IFOL_B$ is a way to combine neural learning and symbolic reasoning to overcome limitations of purely neural systems (like…
    • In this paper we expand the cognitive power of $IFOL_B$ by using the probability computation for the currently unknown sentences, based on Nilsson’s probability…
    • We introduce the global symmetry transformation that preserves the current knowledge database and logical deduction, and the local one used for real-time decisi…
  • Self-Improvements in Modern Agentic Systems: A Survey

    • Publication Time: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.13104v1 Announcement Type: New.
      • Abstract: Self-improving autonomous agents are moving from research prototypes to deployed systems.
      • The primary goal is controllable evolution or adaptation from experience with minimal or even no human input.
      • This survey frames modern self-improving agents as adaptive systems that convert experience into cumulative capability gains.
    • EN Highlights:
      • arXiv:2607.13104v1 Announce Type: new
      • Abstract: Self-improving autonomous agents are moving from research prototypes to deployed systems
      • The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input
      • This survey frames modern self-improving agents as adaptive systems that convert experience into accumulated capability gains
  • Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools

    • Publication Time: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.13115v1 Announcement Type: New.
      • Abstract: Small Language Models (SLMs) have shown promise for zero-shot molecular property prediction from SMILES strings, but they often suffer from structural blindness because the sequential representation does not specify key graph topological cues.
      • We propose a modular, context-augmented prompting framework that can use agentic tools at inference time: a trained GNN expert model to confidently provide prediction hints, and GNN-extracted, instance-specific explanation subgraphs (e.g., subgraph SMILES with accompanying explanatory paragraphs).
      • We evaluate three popular SLMs on MUTAG and Tox21 under five prompting configurations, ranging from SMILES-only to using all available tools at hand.
    • EN Highlights:
      • arXiv:2607.13115v1 Announce Type: new
  • Abstract: Small language models (SLMs) have shown promise for zero-shot molecular property prediction from SMILES strings, yet they often suffer from structural…

  • We propose a modular Context-Augmented Prompting framework that enables agentic tool use at inference time: a trained GNN expert model provides a predictive hin…

  • We evaluate three commonly used SLMs on MUTAG and Tox21 under five prompting configurations ranging from SMILES-only to using all available tools at hand

  • Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents

    • Release Date: 2026-07-17 12:00 Beijing Time
    • Abstract:- arXiv:2607.13157v1 Announce Type: New.
      • Abstract: Agent memory is a systems problem for long-horizon agents.
      • Practical deployments require retaining task state across extended conversations, recovering user-specific facts and preferences across sessions, and accumulating procedural knowledge from prior results.
      • These requirements extend beyond document retrieval: the memory layer must determine which interactions become durable state, how that state is scoped, how to retrieve it under latency constraints, and how to modify or delete it over time.
    • EN Highlights:
      • arXiv:2607.13157v1 Announce Type: new
      • Abstract: Agent memory is a systems problem for long-horizon agents
      • Practical deployments require retention of task state across extended conversations, recovery of user-specific facts and preferences across sessions, and accumu…
      • These requirements extend beyond document retrieval: a memory layer must determine which interactions become durable state, how that state is scoped, how it is…
  • Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models

    • Release Date: 2026-07-17 12:00 Beijing Time
    • Abstract:- arXiv:2607.13172v1 Announce Type: New.
      • Abstract: We address the problem of safely training agent policies and deploying good and safe policies in settings where environment dynamics are unknown and no suitable reward function is available.
      • In the context of safety-critical environments, we argue that traditional reinforcement learning is impractical and turn to human input resources.
      • We introduce DROPJ, a human-centric approach to safe training and deployment.
    • EN Highlights:
      • arXiv:2607.13172v1 Announce Type: new
      • Abstract: We address the problem of safely training an agent policy and deploying a good and safe policy, in settings where the environment dynamics are unknown…
  • In the context of safety-critical environments, we consider traditional reinforcement learning impractical and resort to the resource of human input

  • We introduce DROPJ, a human-centered method for both safe training and deployment

  • CayleyR: Solving the TopSpin puzzle via cycle intersection

    • Publication Time: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.13219v1 Announcement Type: New.
      • Abstract: We present cayleyR, an R package for solving permutation puzzles by detecting cycle intersections in Cayley graphs.
      • The core algorithm performs an iterative bidirectional search: from initial and target permutation states, random operation sequences generate cycles in the Cayley graph of the symmetric group Sn; their intersections yield a connecting path.
      • When no direct intersection is found, a distance-guided bridge selection narrows the gap, and the process repeats.
    • EN Key Points:
      • arXiv:2607.13219v1 Announce Type: new
      • Abstract: We present cayleyR, an R package for solving permutation puzzles by detecting cycle intersections in Cayley graphs
      • The core algorithm performs an iterative bidirectional search: from both the initial and target permutation states, random operation sequences generate cycles i…
      • When no direct intersection is found, a distance-guided bridge selection narrows the gap, and the process repeats
  • Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science

    • Publication Time: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.13220v2 Announcement Type: New.
      • Abstract: Most AI-for-science systems focus on scaling a single reasoning process by using better models, larger context windows, long-horizon agentic execution, or digital co-scientists collaborating with one main user.
      • However, challenging scientific problems are rarely solved by one reasoner alone.
      • They are solved by teams whose members carry different priors, experimental background, tacit knowledge, and domain-trained intuitions.
    • EN Key Points:
      • arXiv:2607.13220v2 Announce Type: new
      • Abstract: Most AI-for-science systems focus on scaling a single reasoning process by using better models, larger context windows, long-horizon agentic execution…
      • However, challenging scientific problems are rarely solved by one reasoner alone
      • They are solved by teams whose members carry different priors, experimental background, tacit knowledge, and domain-trained intuitions

ArXiv cs.CL (B_intro+search) Link to heading

  • Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs

    • Publication Time: 2026-07-17 12:00 Beijing Time
  • Abstract:

    • arXiv:2607.14099v1 Announce Type: new.
    • Abstract: Deploying Vision-Language Models (VLMs) in real-world environments requires not only strong visual reasoning capabilities but also stability under sustained conversational pressure.
    • We introduce Just Keep Prompting (JKP), a multi-turn evaluation framework that measures VLM epistemic stability when users repeatedly challenge, question, or refute the model’s answers.
    • JKP probes models for up to 10 follow-up turns using three strategies: Adversarial Negation (repeated rejection), Pure Socratic Interrogation (repeated calls to re-evaluate certainty), and Context-Aware Socratic Summarization (reflecting the model’s previous rationale before requesting reconsideration).
    • EN 要点:
      • arXiv:2607.14099v1 Announce Type: new
      • Abstract: Deploying Vision-Language Models (VLMs) in real-world settings requires not only strong visual reasoning but also stability under sustained conversati…
      • We introduce Just Keep Prompting (JKP), a multi-turn evaluation framework that measures VLM epistemic stability when users repeatedly challenge, question, or co…
      • JKP probes models for up to 10 follow-up turns using three strategies: Adversarial Negation (repeated rejection), Pure Socratic Interrogation (repeated calls to…
  • Quantum Compositional NLP for Arabic: Grammar, Morphology, and Word Sense in Circuit Topology

    • Publication Time: 2026-07-17 12:00 Beijing Time
    • Abstract:
      • arXiv:2607.14100v1 Announce Type: new.
      • Abstract: We present the first application of pregroup grammar-based quantum compositional natural language processing (QNLP) to Arabic; a morphologically rich, free word order language whose structural complexity provides a unique and demanding testbed for theories of meaning composition in quantum circuits.
      • Our system converts Arabic sentences into quantum circuits whose topology reflects grammatical structure: subjects, verbs, and objects become quantum gates, and type dependencies between them (pregroup grammar) determine how these gates connect.
      • We conduct three controlled experiments covering word order, morphological tense, and verb sense disambiguation, comparing quantum circuit methods with classical baselines (including AraVec (Arabic word embeddings) and AraBERT (pre-trained Arabic transformer)).
    • EN 要点:
      • arXiv:2607.14100v1 Announce Type: new
      • Abstract: We present the first application of pregroup grammar-based quantum compositional natural language processing (QNLP) to Arabic; a morphologically rich,…
      • Our system converts Arabic sentences into quantum circuits whose topology mirrors grammatical structure: subjects, verbs, and objects become quantum gates, and…
      • We conduct three controlled experiments spanning word order, morphological tense, and verb sense disambiguation, comparing quantum circuit methods against class…
  • LBA: Textual Hard-Label Adversarial Attack under Low Query Budgets

    • Publication Time: 2026-07-17 12:00 Beijing Time
  • Abstract: - arXiv:2607.14101v1 Announce Type: new.

  • Abstract: Generating high-quality adversarial texts with low query budgets remains a challenging problem in the hard-label scenario.

  • Most existing approaches rely on greedy algorithms, where one position in the text is selected for substitution, followed by the substitutions of other positions.

  • This local search approach may fail to discover high-quality adversarial examples and often leads to excessive query costs.

    • EN Key Points:
      • arXiv:2607.14101v1 Announce Type: new
      • Abstract: Generating high-quality adversarial texts with low query budgets remains a challenging problem in the hard-label scenario
      • Most existing approaches rely on greedy algorithms, where one position in the text is selected for substitution, followed by the substitutions of other position…
      • This local search approach may fail to discover high-quality adversarial examples and often leads to excessive query costs
  • UniSAGE: Unifying Static and Dynamic Attributes with Hyper-Structure

    • Published: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.14102v1 Announce Type: new.
      • Abstract: With the rapid growth of digital data, real-world applications increasingly involve hierarchical information that combines static attributes with dynamic records.
      • Modeling such heterogeneous data in a unified and generalizable manner remains challenging.
      • Existing approaches often rely on extensive manual design, are tightly coupled to specific data schemas, and typically process static and dynamic attributes in isolation, thus ignoring their implicit interactions.
    • EN Key Points:
      • arXiv:2607.14102v1 Announce Type: new
      • Abstract: With the rapid growth of digital data, real-world applications increasingly involve hierarchical information that combines static attributes with dyna…
      • Modeling such heterogeneous data in a unified and generalizable manner remains challenging
      • Existing approaches often rely on extensive manual design, are tightly coupled to specific data schemas, and typically process static and dynamic attributes in…
  • Latent Communication Between Language Model Agents: Channels, Alignment, and the Limits of Text

    • Published: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.14103v1 Announce Type: new.
      • Abstract: Multi-agent systems (MAS) are utilized in many contexts and many professions.
      • These MAS rely on inter-agent communication, typically achieved through plain-text message passing.
      • We hypothesize that when it is necessary to convey complex concepts, large language models may possess a world model that exceeds the expressive power of text.
    • EN Key Points:
      • arXiv:2607.14103v1 Announce Type: new
      • Abstract: Multi-agent systems (MAS) are utilized in many contexts and many professions
  • Those MAS rely on inter-agent communication, usually implemented by clear-text message passing

  • We hypothesize that Large Language Models may have a world model at their disposal that exceeds expressibility in text when complex concepts need to be communic…

  • UzWordnet and Generative AI for Learning Uzbek by Game Playing

    • Published: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.14104v1 Announcement Type: new.
      • Abstract: This paper proposes an educational system architecture that enables learners to practice the Uzbek language through games.
      • The architecture integrates UzWordnet and the largest currently available orthographic dictionary for Uzbek as core lexical resources, along with generative AI as a fundamental component for learning support.
      • We have designed four educational games to facilitate Uzbek language learning and propose a game-based method to improve UzWordnet as a direct byproduct of the game dynamics.
    • EN Highlights:
      • arXiv:2607.14104v1 Announce Type: new
      • Abstract: This paper presents an educational system architecture that enables learners to practice the Uzbek language through game-playing
      • The architecture integrates UzWordnet and the largest currently available orthographic dictionary for Uzbek as core lexical resources, together with generative…
      • We design four educational games to facilitate Uzbek language learning and propose a game-based methodology for improving UzWordnet as a direct by-product of ga…
  • Automatically Evolving Prompt Guidelines for Task-Specific Optimization

    • Published: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.14105v1 Announcement Type: new.
      • Abstract: To enable large language models to reliably answer user queries, users must clearly specify requirements, context, and constraints.
      • However, in practice, user queries are often underspecified, forcing the model to infer unstated assumptions that may not align with the actual user’s intent.
      • Existing prompt engineering guidelines aim to mitigate this issue, but they are often generic and task-agnostic, limiting their practical utility.
    • EN Highlights:
      • arXiv:2607.14105v1 Announce Type: new
      • Abstract: For Large Language Models to reliably answer user queries, users must clearly specify requirements, context, and constraints
      • In practice, however, user queries are often underspecified, forcing models to infer unstated assumptions that may misalign with the actual user intent
      • Existing prompt engineering guidelines aim to mitigate this issue, they are typically generic and task-agnostic, limiting their practical utility
  • Token Time Continuous Diffusion for Language Modeling

    • Publication Time: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.14106v1 Announce Type: new.
      • Abstract: In this paper, we introduce Token Time Continuous Diffusion (TTCD), a new diffusion language model which (a) operates in a continuous space, deterministically mapping Gaussian noise to a final token canvas without requiring further sampling, and crucially (b) incorporates the novel concept of per-token time, where some tokens proceed from noise to token at a faster rate than others.
      • Continuous space modeling helps TTCD avoid the parallel sampling of multiple tokens, which is a key source of inaccuracy at high speedups for models that iterate purely in discrete space.
      • The concept of per-token time helps TTCD to better model conditional generation, allows more certain tokens to proceed at a faster rate, and allows for differentiating inter-token influence during the refinement process.
    • EN Highlights:
      • arXiv:2607.14106v1 Announce Type: new
      • Abstract: In this paper we introduce token time continuous diffusion (TTCD), a new diffusion language model which (a) operates in continuous space, deterministi…
      • Continuous space modeling helps TTCD avoid the parallel sampling of multiple tokens, which is a key source of inaccuracy at high speedups for models that iterat…
      • The notion of per-token times helps TTCD to better model conditional generation, allows for more sure tokens to proceed at a faster rate, and allows for differe…
  • Polestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs

    • Publication Time: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.14107v1 Announce Type: new.
      • Abstract: The inference efficiency of diffusion large language models (dLLMs) is limited by two challenges: bidirectional attention hinders effective KV cache reuse, while using static confidence thresholds to increase decoding parallelism can harm generation quality.
      • We observe that both of these challenges stem from a common phenomenon: as tokens are decoded, their contextual integration through bidirectional attention causes token representations to drift (evolve) across decoding steps.
      • This insight motivates Polestar, a training-free inference framework that uses token representation drift as a unified signal to jointly address both challenges.
    • EN Highlights:
      • arXiv:2607.14107v1 Announce Type: new
      • Abstract: The inference efficiency of diffusion large language models (dLLMs) is constrained by two challenges: bidirectional attention precludes efficient KV-c…
      • We observe that both challenges arise from a shared phenomenon: as tokens are decoded, their contextual integration through bidirectional attention causes token…
      • This insight motivates Polestar, a training-free inference framework that uses token representation drift as a unified signal to jointly address both challenges Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility
    • Publish Date: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.14108v1 Announcement Type: new.
      • Abstract: This paper introduces tool efficiency, a new quantitative metric for evaluating the rate of useful tool calls in an LLM agent trajectory.
      • To ensure that tool efficiency is well-defined, we also introduce marginal tool utility, a new quantitative metric defined per tool call, indicating whether a tool is useful or can be safely removed from the tool suite without affecting accuracy, while improving tool efficiency; in this paper, we use LLM-as-a-Judge to determine the sign of the marginal tool utility for each tool call in a trajectory.
      • While much prior work has been done to develop techniques that improve LLM tool use and design evaluation methods measuring efficiency indirectly using accuracy as a proxy, our work focuses on directly measuring efficiency through the quantitative metrics proposed in this paper for post-hoc trajectory analysis.
    • EN Key Points:
      • arXiv:2607.14108v1 Announce Type: new
      • Abstract: This paper introduces tool efficiency, a new quantitative metric to evaluate the rate of useful tool calls in an LLM agent trajectory
      • To ensure that tool efficiency is well-defined, we also introduce marginal tool utility, a new quantitative metric defined per tool call indicating whether a to…
      • While much prior work has been done to develop techniques that improve tool use by LLMs and design evaluation methods measuring efficiency indirectly using accu…

ArXiv cs.LG (B_intro+search) Link to heading

  • Position: Explainability Research Must Prioritize Foundations over Ad-hoc Methods

    • Publish Date: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.14123v1 Announcement Type: new.
      • Abstract: Despite the proliferation of Explainable AI (XAI) techniques — from feature attributions to sparse autoencoders — explanations rarely influence real-world workflows.
      • In practice, they are often generated and discarded without guiding meaningful action.
      • This gap reflects foundational shortcomings: research has not yet established methodologies for integrating explanations into end-to-end, human-in-the-loop systems.
    • EN Key Points:
      • arXiv:2607.14123v1 Announce Type: new
      • Abstract: Despite the proliferation of Explainable AI (XAI) techniques – from feature attributions to sparse autoencoders – explanations rarely influence real…
      • In practice, they are often generated and discarded without guiding meaningful action
      • This gap reflects foundational shortcomings: research has not yet established methodologies for integrating explanations into end-to-end, human-in-the-loop syst…
  • CARPRT: Class-Aware Zero-Shot Prompt Reweighting for Black-Box Vision-Language Models

    • Publish Date: 2026-07-17 12:00 Beijing Time
  • Abstract:- arXiv:2607.14125v1 Announcement Type: new.

    • Abstract: Pre-trained vision-language models (VLMs) enable zero-shot image classification by computing the similarity score between an image and textual descriptions, typically formed by inserting class labels (e.g., “cat”) into prompts (e.g., “a photo of”).
    • Since the score for a given image-class pair is sensitive to the choice of prompt, existing research uses weighting vectors to integrate multiple prompts to aggregate scores from different prompts.
    • However, in current strategies, the weighting vector assigned to each prompt is shared across all classes, implicitly assuming that prompts are conditionally independent of classes, which often does not hold in practice, as prompts like “bird’s-eye view” may be suitable for “airport” but not for “apple”.
    • EN Key Points:
      • arXiv:2607.14125v1 Announce Type: new
      • Abstract: Pre-trained vision-language models (VLMs) enable zero-shot image classification by computing the similarity score between an image and textual descrip…
      • Since the score for a given image-class pair is sensitive to the choice of prompt, existing studies ensemble multiple prompts using a weighting vector to aggreg…
      • Yet, in current strategies, the weighting vector assigned to each prompt is shared across all classes, implicitly assuming that prompts are conditionally indepe…
  • Explainable Geospatial AI for Satellite Ground Station Siting Using LiDAR-Derived Terrain Intelligence

    • Publication Time: 2026-07-17 12:00 Beijing Time
    • Abstract:- arXiv:2607.14127v1 Announcement Type: new.
      • Abstract: Representative Clutter Height (RCH) is a critical parameter in radio propagation and interference analysis, as it captures the dominant height of local obstacles contributing to terminal clutter loss.
      • Current practices often rely on fixed clutter heights assigned to land use categories in ITU-R P.452-18 Recommendation, but this omits within-class variation and can lead to conservative exclusion zones and poor site ranking for LEO ground station siting and spectrum coordination.
      • We propose an interpretable, globally deployable machine learning framework for predicting RCH from open geospatial data.
    • EN Key Points:
      • arXiv:2607.14127v1 Announce Type: new
      • Abstract: Representative clutter height (RCH) is a key parameter in radio propagation and interference analysis because it captures the dominant height of local…
      • Current practice often relies on fixed clutter heights assigned to land use classes in Recommendation ITU-R P.452-18, but this misses within class variation and…
      • We present an interpretable, globally deployable machine learning framework for predicting RCH from open geospatial data
  • Certified Domain Consistency for Multi-Domain Retrieval: Label-Free Per-Domain Contamination Control with Conformal Risk Guarantees

    • Publication Time: 2026-07-17 12:00 Beijing Time
  • Abstract: - arXiv:2607.14157v1 Announce Type: new.

    • Abstract: Retrieval over corpora that mix multiple domains often returns relevant but wrong-domain evidence, which ranking metrics miss and for which conformal risk control provides only a small coverage that fails to include the worst-case domain.
    • This work introduces C3R, a drop-in control layer that, based on an inferred domain posterior and no query-time labels, certifies a per-domain contamination budget when feasible, otherwise abstaining rather than silently violating; in the most difficult domains, it guarantees a reduction, rather than a strict limit.
    • The core is a two-split scheme built on risk-controlling prediction sets, whose finite-sample transfer bound crosses from the inferred to the true domain with fully estimable slack, supporting heterogeneous budgets and deployment inversion.
    • EN Key Points:
      • arXiv:2607.14157v1 Announce Type: new
      • Abstract: Retrieval over corpora that mix several domains often returns relevant but wrong-domain evidence that ranking metrics miss and that conformal risk con…
      • This work introduces C3R, a drop-in control layer that, from an inferred domain posterior and no query-time label, certifies a per-domain contamination budget w…
      • The core is a two-split scheme built on risk-controlling prediction sets, whose finite-sample transfer bound crosses from the inferred to the true domain with f…
  • QFireNet: A Quantum-Enhanced U-Net for Wildfire Segmentation from Sentinel-2 Imagery

    • Release Time: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.14160v1 Announce Type: new.
      • Abstract: Wildfire detection from satellite imagery is a semantic image segmentation problem that has proven to be difficult due to challenges such as class imbalance, feature complexity, and atmospheric interference.
      • In this paper, we build on the foundational U-Net image segmentation model to develop a quantum-hybrid solution in hopes of more effectively modeling the high-dimensional spectral feature space of the Sen2Fire dataset.
      • We inject a variational quantum circuit in the bottleneck portion of U-Net, specifically the QuFeX and QB-Net ansatzes.
    • EN Key Points:
      • arXiv:2607.14160v1 Announce Type: new
      • Abstract: Wildfire detection from satellite imagery is a semantic image segmentation problem that has proven to be difficult due to challenges such as class imb…
      • In this paper, we build on the foundational U-Net image segmentation model to develop a quantum-hybrid solution in hopes of more effectively modeling the high-d…
      • We inject a variational quantum circuit in the bottleneck portion of U-Net, specifically the QuFeX and QB-Net ansatzes
  • Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning

    • Release Time: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.14171v1 Announce Type: new.
      • Abstract: Reinforcement learning has become the dominant paradigm for training large language model (LLM) agents that interact with executable sandboxes.
  • State-of-the-art algorithms, such as PPO, RLOO, and GRPO, inherit the rollout topology from RLHF: for each prompt, N independent trajectories are sampled from the initial state, and advantages are computed by subtracting a group baseline.

    • This design ignores a defining property of agent sandboxes.
    • EN Highlights:
      • arXiv:2607.14171v1 Announce Type: new
      • Abstract: Reinforcement learning has emerged as the dominant paradigm for training large language model (LLM) agents that interact with executable sandboxes
      • State-of-the-art algorithms such as PPO, RLOO, and GRPO inherit their rollout topology from RLHF: for each prompt, N independent trajectories are sampled from t…
      • This design ignores a defining property of agent sandboxes
  • How Much of a 10-K Matters? Aggregation-Dependent Value of Full-Text versus Risk-Factor Sentiment

    • Publication Time: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.14174v1 Announcement Type: New.
      • Abstract: Financial sentiment extraction has largely relied on news text and supervised extraction against return labels alone, while 10-K filings and volatility-targeted risk disclosures, arguably the most suitable sources, remain relatively unexplored.
      • We extend a supervised lexicon-learning approach to 10-K filings and their Item 1A risk-factor sections, training sentiment scores against both return and volatility labels at three levels of aggregation (industry, portfolio, and individual firm).
      • Across 1,383 filings from 94 Nasdaq-100 technology constituents (2006-2023), we evaluate the resulting twelve sentiment metrics on classification accuracy, correlation with realized market outcomes, and qualitative lexical content.
    • EN Highlights:
      • arXiv:2607.14174v1 Announce Type: new
      • Abstract: Financial sentiment extraction has largely relied on news text and supervised extraction against return labels alone, leaving 10-K filings – and vola…
      • We extend a supervised lexicon-learning approach to 10-K filings and their Item 1A risk-factor sections, training sentiment scores against both return and volat…
      • Across 1,383 filings from 94 Nasdaq-100 technology constituents (2006–2023), we evaluate the resulting twelve sentiment metrics on classification accuracy, cor…
  • Low-Latency Relay Selection in NR-V2X Vehicular Communications via Graph Isomorphism Networks with Edge Features

    • Publication Time: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.14176v1 Announcement Type: New.
      • Abstract: Reliable, low-latency uplink connectivity is a key requirement for C-V2X networks in dense urban environments, where rapid channel variations and blockages often degrade direct vehicle-to-infrastructure links.
      • Multi-hop relaying can restore coverage, but relay link activation under radio, capacity, and routing constraints leads to an NP-hard optimization problem, typically solved via mixed-integer linear programming (MILP), whose runtime scales poorly with graph size.
      • This paper introduces an edge-aware learning-optimization framework for real-time relay selection.
  • EN Key Points:

    • arXiv:2607.14176v1 Announce Type: new
    • Abstract: Reliable, low-latency uplink connectivity is a key requirement for C-V2X networks in dense urban environments, where fast channel variations and block…
    • Multi-hop relaying can restore coverage, but relay-link activation under radio, capacity, and routing constraints results in an NP-hard optimisation problem, ty…
    • This paper introduces an edge-aware Learning-to-Optimise framework for real-time relay selection
  • RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences

    • Publication Time: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.14180v1 Announce Type: new.
      • Abstract: World models are widely used in offline reinforcement learning (RL) to improve sample efficiency and generate experience beyond a fixed dataset.
      • However, they are vulnerable to model exploitation where data coverage is thin.
      • Prior work addresses this by either collecting more expert demonstrations, which is often expensive, unsafe, or unavailable, or by conservative algorithms that avoid uncertain regions, which limits generalization.
    • EN Key Points:
      • arXiv:2607.14180v1 Announce Type: new
      • Abstract: World models are widely used in offline reinforcement learning (RL) to improve sample efficiency and generate experience beyond a fixed dataset
      • However, they are vulnerable to model exploitation where data coverage is thin
      • Prior work addresses this either by collecting more expert demonstrations, which is often expensive, unsafe, or unavailable, or by conservative algorithms that…
  • Closed-Loop Knowledge Dynamics: An Operational Framework for Saturation and Escape

    • Publication Time: 2026-07-17 12:00 Beijing Time
    • Abstract: - arXiv:2607.14185v1 Announce Type: new.
      • Abstract: Feedback-driven loops support iterative improvement in large language models, reinforcement learning, and autonomous discovery, yet their gains often diminish under repeated internal feedback.
      • We investigate why closed-loop knowledge systems saturate and what external information can enable them to escape their current attractors.
      • We introduce a three-level operational framework where the knowledge state $x_t$ evolves via a transition kernel $K_{\theta}$ indexed by structural parameters $\theta$.
    • EN Key Points:
      • arXiv:2607.14185v1 Announce Type: new
      • Abstract: Feedback-driven loops support iterative improvement in large language models, reinforcement learning, and autonomous discovery, yet their gains often…
  • We study why closed-loop knowledge systems saturate and what external information can move them beyond their current attractors

    • We introduce a three-level operational framework in which knowledge states $x_t$ evolve through transition kernels $K_{\theta}$ indexed by a structural paramete…