System translated (Gemini)

🤖 AI 速览

The main theme today is the progression of AI from general conversation and tool use to high-judgment scenarios such as scientific research and life sciences. Claude Science, GeneBench-Pro, and multiple planning studies all point to a single trend: the industry is beginning to prioritize verifiable …
📋 文章元数据
发布时间
2026-07-01
类型
ai-daily
字数
8037
阅读时长
38 min

2026-07-01 AI Daily | AI Enters the Operational Layer of Scientific Research: From Model Capabilities to Real-World Judgment and Workbench Competition Link to heading

Today’s main theme is AI’s progression from general conversation and tool use to high-judgment scenarios like scientific research and life sciences. Claude Science, GeneBench-Pro, and multiple new studies on planning all point to a trend: the industry is starting to prioritize verifiable reasoning, specialized workflows, and tangible user value over simply model scores and engineering efficiency.

📖 In-depth Guide to This Issue’s Watch List Link to heading

The most noteworthy topic today is “how agents can achieve reliable planning.” Multiple new arXiv papers simultaneously address world models, symbolic feedback, self-correction, and hallucination propagation control. This indicates a shift in the industry from “making models that can call tools” to “making models that can simulate consequences and verify paths.” We recommend focusing on the papers about world model planning and grounded iterative planning.

On the product side, Google DeepMind launched Nano Banana 2 Lite and Gemini Omni Flash, while OpenAI used Signals data to review the global adoption and expansion of ChatGPT. These are useful for observing the trend of AI platformization from both the supply and demand sides.

Additionally, papers on multi-agent personality composition, explainable fact-checking, and multimodal emotion reasoning are worth exploring as extended reading. They collectively answer a key question: as models enter collaborative, judgmental, and social scenarios, factors beyond performance—such as explainability, robustness, and human alignment—are becoming more critical.

🌐 Quick Takes on AI Hot Topics from X Link to heading

Topic 1: Andrew Ng Shares Loop Engineering Framework for Faster AI Products Link to heading

  • Category: AI · News
  • Summary: Trending: 6 hours ago, Related posts: 1000
  • What it is: Andrew Ng shared a product development framework called “Loop Engineering,” which advocates for accelerating the deployment of AI products through rapid building, evaluation, and iteration.
  • Why it matters: This framework emphasizes a shift from focusing on model capabilities to a systematic iterative process, reflecting that competition in AI applications is increasingly reliant on data feedback, evaluation systems, and product engineering efficiency.
  • Discussion summary: The discussion on X is mainly focused on whether this framework is just a rebranding of agile development and MLOps, and whether it can help teams discover real user needs faster and reduce the failure rate of AI products moving from prototype to production.

Topic 2: Dario Amodei’s 2023 Open-Source AI Warning Resurfaces Amid Chinese Model Boom Link to heading

  • Category: AI · News
  • Summary: Trending: 2 days ago, Related posts: 35000
  • What it is: A warning made in 2023 by Anthropic CEO Dario Amodei about the potential security and geopolitical risks of open-source AI has regained attention due to the rapid rise of Chinese open-source models.
  • Why it matters: This highlights how the open-source AI path, while driving innovation and lowering barriers to entry, may also accelerate the proliferation of cutting-edge capabilities and affect the global AI competition and governance landscape.
  • Discussion summary: Discussions on X are centered on whether open-source AI is weakening the US’s lead, whether export and model release restrictions should be tightened, and whether the open ecosystem ultimately brings more benefits than risks to security and industrial innovation.

Topic 3: Linq Launches Scalable iMessage Apps API for AI Agents Link to heading

  • Category: AI · News
  • Summary: Trending: 3 hours ago, Related posts: 1100
  • What it is: Linq has launched a scalable iMessage Apps API for AI Agents, allowing developers to integrate intelligent agent capabilities into Apple’s iMessage application ecosystem.
  • Why it matters: This means AI Agents could enter more frequently used private communication scenarios, driving deeper integration between chat, task execution, automated services, and mobile messaging platforms.
  • Discussion summary: The discussion on X mainly focuses on whether this API can lower the barrier to entry for AI Agent deployment, the platform limitations within the closed iMessage ecosystem, and the risks related to privacy, security, and automated message abuse.

Topic 4: X Launches Hosted Servers for AI Agents to Access Real-Time Data Link to heading

  • Category: AI · News
  • Summary: Trending: 22 hours ago, Related posts: 9600
  • What it is: X has launched hosted MCP (Model Context Protocol) servers, enabling AI Agents like Grok, Cursor, and Claude Desktop to more easily access the X API for real-time data.
  • Why it matters: This lowers the barrier for AI Agents to connect to real-time information sources, which helps improve the timeliness and utility of agent systems in tasks like news, search, monitoring, and decision-making. It could also promote MCP as a key standard for agent tool integration.
  • Discussion Summary: Discussions on X are focused on whether developers will be able to build Agents with real-time awareness capabilities more quickly. Supporters call it a “game-changer,” while skeptics are concerned about API costs, data quality, platform dependency, and the risks of privacy and abuse.

Topic 5: Etched Emerges from Stealth with Transformer-Specific AI Chips Link to heading

  • Category: AI · News
  • Overview: Trending since: 8 hours ago, Related Posts: 4900
  • What it is: AI chip startup Etched has come out of stealth, releasing an AI inference chip specifically optimized for the Transformer architecture.
  • Why it matters: This indicates that the competition in AI hardware is shifting from general-purpose GPUs to specialized accelerators for mainstream model architectures, which could impact inference costs, energy efficiency, and Nvidia’s dominance.
  • Discussion Summary: Discussions on X are centered on whether specialized chips can truly reduce the inference costs of large models, whether the Transformer architecture will remain dominant long-term, and the performance, supply, and commercialization risks of Etched compared to GPUs and other ASIC solutions.

Summary of AI Public Opinion on X Today Link to heading

The main theme of today’s discourse is that the AI competition is shifting from a focus on individual model capabilities to a systemic competition composed of “engineered implementation, real-time data access, platform distribution, and specialized hardware.” There is a broad consensus that rapid iteration frameworks, Agent access to messaging and social data, and inference chip optimization will lower the deployment barrier for AI applications and enhance their practicality. However, the true value depends on evaluation systems, validation of user needs, data quality, and cost structures. The main points of divergence are centered on open vs. closed and general-purpose vs. specialized: Is open-source AI an innovation accelerator or an amplifier of security and geopolitical risks? Are platform interfaces like MCP and iMessage an ecosystem opportunity or a new platform dependency? Are Transformer-specific chips a judgment on a trend or a gamble on an architecture? Potential risks include the governance pressure from the proliferation of cutting-edge capabilities, privacy and abuse issues arising from Agents entering private communications and real-time information streams, and new lock-ins and commercial uncertainties for businesses regarding APIs, hardware roadmaps, and closed ecosystems.

💡 Influencer Insights Link to heading


AI Industry Daily Observations (06/30) Link to heading

Today’s core trend is clearly focused on the deep tooling and platformization of AI Agents, as well as practical skills surrounding the Codex ecosystem. Meanwhile, the model arms race has escalated from foundational models to vertical industry operating layers and the edge.

  • 🔥 Focus 1: Anthropic’s Double Release—Sonnet 5 and Claude Science Anthropic released two major products on the same day, dominating nearly half of today’s discussions.

    • Claude Sonnet 5: @dotey (Baoyu) provided a detailed analysis of its positioning—offering Agent capabilities close to top-tier models at a lower price (API price is only 40% of Opus 4.8). It replaces Sonnet 4.6 and aims to close the gap between standard and top-tier models in autonomous planning and multi-step tasks, with significant improvements, especially in Agent programming benchmarks.
    • Claude Science: @dotey pointed out that this is a strategic product from Anthropic, shifting AI from mere model capabilities to a “specific industry operating layer.” It is not a new model, but a workbench for researchers (especially in life sciences) that integrates over 60 scientific databases, computing resources, and collaborative Agents, aiming to replicate the success of Claude Code in the software engineering field. This marks the evolution of AI from a “conversational tool” to a “vertical domain operating system.”
  • 💻 Trend 2: Deep Dive and “Decryption” of the Codex Ecosystem @Pluvio9yte (Xueta Wuyun) and @dotey shared a wealth of in-depth content about Codex, showing that the tool has entered the stage of practical skills and ecosystem expansion.

    • Dissemination of Practical Skills: @Pluvio9yte shared and summarized multiple practical guides on Codex, including hardcore tips like credit management, memory functions, and finishing up in /goal mode, and even exposed a suspected bug for /goal unlimited credits.
    • Ecosystem Expansion: From the Windows desktop Dynamic Island tool shared by @AI_Jasonyu to the open-source project DevSpace introduced by @gefei55 (Gefei) (which allows the ChatGPT web version to manipulate local code like Codex), the Codex ecosystem is being rapidly enriched by the community.
    • Transparency Controversy: @dotey reported on a major allegation—a security researcher discovered through reverse engineering that Claude Code uses hidden Unicode characters to “watermark” the system prompts of users in China via proxies, sparking a heated discussion about trust and privacy in AI tools.
  • 🚀 Trend 3: The Continued Rise of Local/On-Device Models @zhixianio (Zhixian) continues to share insights from testing local models, noting that while the audio-video full-duplex performance of MiniCPM-o 4.5 is satisfactory, its stability needs improvement. This confirms that on-device intelligence is both promising and faces engineering challenges, a topic also focused on in @zhixianio’s podcast, “Cognitive County.”

2. Noteworthy Unique Perspectives or Industry Foresight Link to heading

  • 🙈 Disconnect Between AI Engineering Metrics and User Value: @dotey (Baoyu) keenly observed that a promotional video Anthropic made for Spotify backfired on X. The promotion highlighted engineering-side numbers like 4,500 daily deployments and 73% of PRs being AI-assisted, but users widely complained about a decline in product experience. This sharply points out a fundamental problem in current AI applications: a huge gap exists between the “productivity” metrics used to measure AI’s value (lines of code, number of deployments) and the “product quality” ultimately perceived by users. The industry needs new benchmarks for measuring value.

  • 🕵️ “Covert Channel”: A Questioning of Trust in AI Tools: @dotey (Baoyu) relayed in detail the allegations that Claude Code embeds hidden characters in its system prompt to mark users from Chinese proxies. Regardless of whether this is an anti-abuse measure or a privacy violation, it has raised public concern about AI tools that have deep system access (the ability to read code, run commands). Users have the right to know what the tool is doing behind the scenes, and this kind of undisclosed “marking” behavior causes a massive erosion of trust.

  • 🧬 Milestone in Non-Invasive Brain-Computer Interfaces: @dotey (Baoyu) reported on Meta’s Brain2Qwerty v2, which increased the word accuracy of non-invasive brain decoding from the industry average of 8% to 61%, with a peak of 78%. @dotey’s point is key: This proves that “it’s possible to achieve results close to invasive methods without surgery, and what remains are engineering problems, not fundamental principle problems.” This has profound implications for the vast number of brain injury patients who cannot undergo craniotomy.

  • 🧠 Dissenting Views on “AI Open Source”: @ruanyf (Ruan Yifeng) cited the viewpoint of Anthropic’s founder, who believes that AI models that only release weights without revealing their internal workings are merely “open-weight,” not “open-source” in the traditional sense. This viewpoint punctures the current hype bubble around “open-source” large models and prompts deep reflection on the true definition of openness.

  • 👴 The Symbiotic Relationship Between AI and Human Experts: @dotey (Baoyu) shared the case of Ford Motor Company rehiring 350 veteran engineers (“graybeards”) because its AI quality inspection system failed to meet expectations and needed these experienced masters to train the AI. This provides a valuable counter-narrative to the mainstream “AI will replace jobs” story: the successful implementation of AI may, in fact, depend even more on top-tier human experience and judgment.

  • Development & Tools

    • Orca (@LinearUncle / shared by @dotey): An open-source Coding IDE recommended as being comparable to or even surpassing Codex App, with cross-platform support.
    • Codex Utilities:
      • Windows Desktop Codex Dynamic Island (@mooyuking): Helps Windows users monitor status and quotas.
      • ChatGPT Batch Delete Plugin (@Pluvio9yte): A free, local plugin for clearing chat history on the ChatGPT web interface.
    • oMLX v0.4.0 (@jundotkim / shared by @zhixianio): A local application for running models on Apple Silicon, which has released its first native Swift macOS app.
    • Feishu’s Open-Source CLI Toolkit (recommended by @ruanyf): Allows AI Agents to call office functionalities, and has already surpassed 10,000 stars on GitHub.
  • Learning & Resources

    • “Claude Code From Scratch” (recommended by @dotey): An open-source e-book that replicates the core architecture of Claude Code in about 4,300 lines of code, serving as an excellent tutorial for understanding the principles of coding agents.
    • Codex Orange Paper (@bozhou_ai / shared by @AI_Jasonyu): Open-source, systematized learning materials for Codex.
    • GEO Content Engineering Resource Pack (@vista8): A set of systematic resources on Generative Engine Optimization (GEO), including an operation manual, skill packs, and demos.
    • Video Production Skills Repository (@Pluvio9yte): An open-source set of video production skills that can replicate effects similar to HyperFrames, suitable for users with no video editing experience.
  • Large Models & Applications

    • Apodex 4B (@Pluvio9yte): A local model positioned as a “personal deep research assistant.” A guide to pitfalls encountered during deployment has been published.
  • YouMind 1.0 (recommended by @lifesinger / @gefei55): A tool that helps creators efficiently produce graphic and text content and distribute it to X and public accounts with one click.

📚 Appendix: Today’s Watch List Update Source List Link to heading

Time window: Last 3 days; 22 sources covered; 37 updates in total.

Lex Fridman Podcast (A_full) Link to heading

  • #498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires
    • Published: 2026-07-01 05:33 Beijing Time
    • Summary: - Anthony Kaldellis is a historian of the Roman Empire and the author of “The New Roman Empire,” a comprehensive history of the Byzantine Empire (Eastern Roman Empire).
      • Please see the timestamps and transcript below to provide feedback, submit questions, contact Lex, etc.
      • Upwork: A platform for hiring freelancers.
      • Fin: An AI agent for customer service.
      • BetterHelp: Online therapy and counseling.
    • EN Key Points:
      • Anthony Kaldellis is a historian of the Roman Empire and author of “The New Roman Empire”, a comprehensive history of the Byzantine Empire (Eastern Roman Empire…
      • Thank you for listening ❤ Check out our sponsors:
      • See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc
      • CONTACT LEX:

OpenAI Blog (A_full) Link to heading

  • How ChatGPT adoption has expanded

    • Published: 2026-06-30 17:00 Beijing Time
    • Summary: - The adoption of ChatGPT is expanding and deepening globally.
      • New OpenAI Signals data shows that over time, people are using ChatGPT more frequently and for a wider range of tasks, while its user base is becoming more global and diverse.
      • OpenAI Signals uses aggregated data to measure how people interact with personal ChatGPT plans (including Free, Go, Plus, and Pro plans) over time.
      • This analysis provides a perspective on how personal AI usage is evolving as ChatGPT reaches a global scale.
      • The visualizations below highlight several core trends in global AI adoption: how ChatGPT usage has deepened since launch, the ways people incorporate it into their daily work, and the variations across different languages and regions.
    • EN Key Points:
      • New OpenAI Signals data shows how ChatGPT adoption is growing globally, with users increasing usage, exploring more capabilities, and driving growth across regi…
  • Introducing GeneBench-Pro

    • Published: 2026-06-30 08:00 Beijing Time
    • Summary: - Scientific data rarely comes with instructions.
      • Researchers must decide whether a pattern reflects biology or noise, whether the data can support the questions being asked, and how each result should inform their next steps.
      • AI agents are increasingly capable of performing complex analyses, but real scientific research relies not just on recalling facts or following predefined workflows, but also on making these kinds of higher-order judgments.
      • Today, we are introducing GeneBench-Pro—a challenging, research-grade benchmark for testing whether models can handle the kind of judgment-heavy analysis required for real-world computational biology.
      • To date, compelling evaluations of system-level judgment calls have been rare, making real-world computational research difficult.
    • EN Key Points:
  • Introducing GeneBench-Pro, a new benchmark testing AI performance in genomics, biology, and scientific research using complex, real-world datasets.

  • Core dump epidemiology: fixing an 18-year-old bug

    • Publication Time: 2026-06-30 08:00 Beijing Time
    • Abstract: - OpenAI engineers used large-scale core dump analysis to debug rare infrastructure crashes, discovering hardware failures and long-standing software bugs.
      • This article from the OpenAI blog explains how core dump epidemiology: fixing an 18-year-old bug shapes the broader AI and infrastructure landscape.
      • It also reveals the practical implications of core dump epidemiology: fixing an 18-year-old bug for founders, operators, and investors.
    • EN Highlights:
      • OpenAI engineers used large-scale core dump analysis to debug rare infrastructure crashes, uncovering both a hardware fault and a long-standing software bug.
  • Inside Genebench-Pro

    • Publication Time: 2026-06-30 08:00 Beijing Time
    • Abstract: - Inside Genebench-Pro.
      • This article from the OpenAI blog explains how Inside Genebench-Pro shapes the broader AI and infrastructure landscape.
      • It also provides practical implications for founders, operators, and investors following Inside Genebench-Pro.
    • EN Highlights:
      • Inside Genebench-Pro

Google DeepMind Blog (A_full) Link to heading

  • Start building with Nano Banana 2 Lite and Gemini Omni Flash
    • Publication Time: 2026-07-01 00:02 Beijing Time
    • Abstract: - Start building with Nano Banana 2 Lite and Gemini Omni Flash.
      • This article from the Google DeepMind blog explains how starting to build with Nano Banana 2 Lite and Gemini Omni Flash shapes the broader AI and infrastructure landscape.
      • It also has practical implications for founders, operators, and investors after starting to build with Nano Banana 2 Lite and Gemini Omni Flash.
    • EN Highlights:
      • Start building with Nano Banana 2 Lite and Gemini Omni Flash

Lex Fridman (B_intro+search) Link to heading

  • Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires | Lex Fridman Podcast #498
    • Publication Time: 2026-07-01 05:16 Beijing Time
    • Abstract: - Anthony Kaldellis is a historian of the Roman Empire and the author of “The New Roman Empire,” a comprehensive history of the Byzantine Empire (Eastern Roman Empire).
      • Please see the timestamps and transcript below, and provide feedback, submit questions, contact Lex, etc.
      • Feedback - Provide feedback to Lex:.
      • AMA - Submit questions, videos, or call in:.
    • EN Highlights:
      • Anthony Kaldellis is a historian of the Roman Empire and author of “The New Roman Empire”, a comprehensive history of the Byzantine Empire (Eastern Roman Empire…
      • Thank you for listening ❤ Check out our sponsors:
  • See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc
  • Transcript:

ArXiv cs.AI (B_intro+search) Link to heading

  • AI-Model Network: Concept, Current State and Future

    • Posted: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.27382v1 Announce Type: new.
      • Abstract: The primary function of computers lies in computation and processing, while the core value of the Internet is rooted in sharing and collaboration.
      • Computers create the Internet, and the Internet gives value to computers.
      • The rapid development of the Internet, cloud computing, and big data is pushing artificial intelligence into the era of large models (LMs).
    • EN Highlights:
      • arXiv:2606.27382v1 Announce Type: new
      • Abstract: While the primary function of computers lies in computation and processing, the core value of the Internet is rooted in sharing and collaboration
      • Computers create the Internet, and the Internet empowers the value of computers
      • The rapid development of the Internet, cloud computing, and big data is pushing artificial intelligence into the era of large models (LMs)
  • When Does Personality Composition Matter for Multi-Agent LLM Teams?

    • Posted: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.27443v1 Announce Type: new.
      • Abstract: Personality prompts determine how large language models communicate, but whether these behavioral shifts affect objective task outcomes remains to be explored.
      • Previous research has shown that agents prompted with low agreeableness produce adversarial language, while those prompted with high agreeableness become cooperative, but the relationship between communication style and task performance has not been systematically examined across multiple domains.
      • In this work, we investigate whether personality composition matters for multi-agent team performance by manipulating the personality traits of frontier LLMs in three task domains: structured coding, open-ended research collaboration, and competitive negotiation.
    • EN Highlights:
      • arXiv:2606.27443v1 Announce Type: new
      • Abstract: Personality prompting shapes how large language models communicate, yet whether these behavioral shifts affect objective task outcomes remains under-e…
      • Prior work shows that agents prompted with low agreeableness produce adversarial language, while those prompted with high agreeableness become cooperative, but…
      • In this work, we investigate whether personality composition matters for multi-agent team performance by manipulating personality traits across frontier LLMs on…
  • Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning

    • Posted: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.27483v1 Announce Type: new.
  • Abstract: Large language model (LLM) agents have demonstrated strong capability in sequential decision-making, yet they remain fundamentally reactive in long-term tasks.

  • Unlike humans who employ “what-if” reasoning to evaluate potential plans before commitment, standard agents lack an internal world model to simulate future outcomes.

  • Therefore, we propose to internalize future-aware planning by training a single autoregressive model to verbalize both a prospective state rollout and a plan-conditional success estimate (a textual simulation of Q-values).

  • EN Highlights:

    • arXiv:2606.27483v1 Announce Type: new
    • Abstract: Large language model (LLM) agents have demonstrated strong capability in sequential decision-making, yet they remains fundamentally reactive in long-h…
    • Unlike humans who employ “what-if” reasoning to evaluate potential plans before commitment, standard agents lack an internal world model to simulate future outc…
    • Therefore, we propose to internalize future-aware planning by training a single autoregressive model to verbalize both a prospective state rollout and a plan-co…
  • Odyssey: Constructing Verifiable Local Truth-Preserving Foundation Models

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.27593v1 Announce Type: new.
      • Abstract: We introduce a categorical framework called ODYSSEY for constructing verifiable, local truth-preserving foundation models as compositions of foundries: building block architectural components that specify local contexts, families of local representations, restriction graphs, gluing rules, obstruction policies, update obligations, and human-facing views.
      • A foundry is an organized knowledge base that contains an argumentation component.
      • Concrete foundries are built from generic foundries, such as those for evidence/argument, operational decisions, institutional/financial matters, market meaning, scientific challenges, research programs, assisted construction, and evaluation rigs.
    • EN Highlights:
      • arXiv:2606.27593v1 Announce Type: new
      • Abstract: We introduce a categorical framework called ODYSSEY for constructing verifiable, local truth-preserving foundation models as compositions of foundries…
      • A foundry is an organized sheaf of knowledge that carries within it an argumentation component
      • Concrete foundries are built from generic foundries such as evidence/argument, operational decision, institutional/financial, market meaning, scientific challen…
  • DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.27619v1 Announce Type: new.
      • Abstract: Dyslexic learners are increasingly using artificial intelligence (AI) tools to support tasks related to reading, writing, organization, and learning.
      • However, their lived experiences using these tools remain largely unexamined.
      • This paper presents DysLexLens, a low-resource LLM framework designed to analyze the experiences of dyslexic learners with AI through discussions on online forums.
    • EN Highlights:
  • arXiv:2606.27619v1 Announce Type: new

  • Abstract: Dyslexic learners increasingly use artificial intelligence (AI) tools to support reading, writing, organisation, and study-related tasks

  • However, their lived experiences with these tools remain largely underexamined

  • This paper proposes DysLexLens, a low-resource LLM framework, designed to analyse dyslexic learners experience with AI through online forum discussions

  • MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.27652v1 Announce Type: new.
      • Abstract: We find that explicit reasoning does not necessarily translate into better multimodal emotion recognition (MER) accuracy, even though it makes predictions more interpretable.
      • Specifically, for reasoning-based MLLMs, fast thinking by triggering direct answers often outperforms slow thinking after deliberative reasoning.
      • Our empirical analysis shows that fast thinking can improve recall with broader and more confident predictions, while slow thinking improves precision by conservatively filtering out incorrect categories.
    • EN Highlights:
      • arXiv:2606.27652v1 Announce Type: new
      • Abstract: We find that explicit reasoning does not necessarily translate into better multimodal emotion recognition (MER) accuracy, even though it makes predict…
      • Specifically, for reasoning-based MLLMs, fast thinking by triggering direct answers often outperforms slow thinking after deliberative reasoning
      • Our empirical analyses show that fast thinking improves recall with broader and more confident predictions, whereas slow thinking favors precision through conse…
  • ToE: A Hierarchical and Explainable Claim Verification Framework with Dynamic Multi-source Evidence Retrieval and Aggregation

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.27736v1 Announce Type: new.
      • Abstract: The rapid spread of fake news poses an increasing threat to the information ecosystem, especially as AI-generated misinformation under Generative Engine Optimization (GEO) poisoning causes retrieval systems to systematically present adversarially crafted content, thereby contaminating the reasoning of LLMs.
      • In this paper, we propose the Tree of Evidence (ToE), a hierarchical evidence reasoning framework for automated fact-checking that models each claim as a dynamically expanding argument tree.
      • ToE integrates a reinforcement learning-driven multi-source retrieval agent, an evidence evaluation agent, and an argument tree aggregation algorithm to iteratively decompose, retrieve, and verify claims through an explainable chain of evidence.
    • EN Highlights:
      • arXiv:2606.27736v1 Announce Type: new
      • Abstract: The rapid spread of fake news poses increasing threats to information ecosystems, especially as AI-generated misinformation under Generative Engine Op…
  • In this paper, we propose Tree of Evidence (ToE), a hierarchical evidence reasoning framework for automated fact-checking that models each claim as a dynamicall…

  • ToE integrates a reinforcement learning-driven multi-source retrieval agent, an evidence evaluation agent, and an argument tree aggregation algorithm to iterati…

  • Towards Reliable and Robust LLM Planning: Symbolic Feedback-Driven Iterative Self-Refinement Framework

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.27757v1 Announce Type: new.
      • Abstract: Large language models (LLMs) have attracted widespread attention from academia and industry, yet their deployment raises critical security concerns.
      • Planning, a core component of intelligent behavior, remains challenging for LLMs, which often produce infeasible or incorrect solutions in long-horizon decision-making tasks due to inherent complexities.
      • In this paper, we propose a symbolic feedback-driven iterative self-refinement framework to enhance the robustness and reliability of LLMs in long-horizon planning.
    • EN Key Points:
      • arXiv:2606.27757v1 Announce Type: new
      • Abstract: Large language models (LLMs) have attracted widespread attention from academia and industry, yet their deployment raises critical security concerns re…
      • Planning, a core component of intelligent behavior, remains challenging for LLMs, which often produce infeasible or incorrect solutions in long-horizon decision…
      • In this paper, we propose a symbolic feedback-driven iterative self-refinement framework to enhance the robustness and reliability of LLMs in long-horizon plann…
  • Understanding Rollout Error in Graph World Models

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.27780v1 Announce Type: new.
      • Abstract: World models are often used for planning by rolling learned dynamics forward.
      • Many planning environments, however, are not vectors or images; they are graphs of agents, tools, skills, routes, and dependencies.
      • In these settings, local prediction errors may remain local or propagate through the graph, and failure modes change again when edges are predicted rather than fixed.
    • EN Key Points:
      • arXiv:2606.27780v1 Announce Type: new
      • Abstract: World models are often used for planning by rolling learned dynamics forward
      • Many planning environments, however, are not vectors or images; they are graphs of agents, tools, skills, routes, and dependencies
  • In these settings, a local prediction error may stay local or spread through the graph, and the failure mode changes again when edges are predicted rather than…

  • Grounded Iterative Language Planning: How Parameterized World Models Reduce Hallucination Propagation in LLM Agents

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.27806v1 Announcement Type: new.
      • Abstract: World models for language agents come in two useful forms.
      • An agent-based world model calls an LLM API and reasons flexibly in language, but its errors manifest as hallucinated state changes that are difficult to score with ordinary regression loss.
      • A parameterized world model is a trained transition predictor; its errors are easier to measure with quantities such as NodeMSE, delta accuracy, and validity accuracy, but it is generally weaker as a standalone planner.
    • EN Highlights:
      • arXiv:2606.27806v1 Announce Type: new
      • Abstract: World models for language agents come in two useful forms
      • An agent-based world model calls an LLM API and reasons flexibly in language, but its errors appear as hallucinated state changes that are hard to score with or…
      • A parameterized world model is a trained transition predictor; its errors are easier to measure with quantities such as NodeMSE, delta accuracy, and validity ac…

ArXiv cs.CL (B_intro+search) Link to heading

  • Generating in the Limit with Infinitely Many Hallucinations

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.28354v1 Announcement Type: new.
      • Abstract: The classic paradigm of language identification in the limit models learning as a game between an adversary (who reveals strings from an unknown target language) and a learner (responsible for identifying that language).
      • The recently introduced framework of language generation in the limit shifted the objective to better reflect modern language modeling, requiring the learner to generate valid, unseen strings from the target language.
      • Related work highlighted a fundamental tension: broad coverage of the target often comes at the cost of validity.
    • EN Highlights:
      • arXiv:2606.28354v1 Announce Type: new
      • Abstract: The classic paradigm of language identification in the limit models learning as a game between an adversary, who reveals strings from an unknown targe…
      • The recently introduced framework of language generation in the limit shifted the objective to better reflect modern language modeling, requiring the learner to…
      • Related work highlighted a fundamental tension: a broad coverage of the target often comes at the cost of validity
  • Extracting Knowledge from an Arabic-English Machine-Readable Dictionary Using Information Extraction

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Summary: - arXiv:2606.28457v1 Announcement Type: New.
      • Abstract: Natural Language Processing (NLP) applications require a large and rich amount of linguistic knowledge.
      • Furthermore, electronic language resources such as dictionaries, encyclopedias, and corpora have become available.
      • Therefore, automatic methods have emerged to extract lexical information from these sources to overcome the knowledge acquisition bottleneck.
    • EN Key Points:
      • arXiv:2606.28457v1 Announce Type: new
      • Abstract: Natural language processing (NLP) applications need large and rich amount of linguistic knowledge
      • Furthermore, electronic language sources such as dictionaries, encyclopedia, and corpora became available
      • So, automatic methods are emerged to extract lexical information from those sources to overcome the knowledge acquisition bottleneck
  • Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Summary: - arXiv:2606.28524v1 Announcement Type: New.
      • Abstract: Recent work suggests that Large Language Models (LLMs) are sensitive to the belief states of subjects described in text, as measured by the False Belief Task (FBT), but concerns about structural validity persist.
      • We adopt a developmental perspective, tracing the patterns of mental state reasoning behavior—and the likely preconditions for this behavior—across multiple training stages of the Olmo2 and Pythia language model suites.
      • We find that above-chance FBT performance depends on both model size and sufficient training volume, emerges relatively late in pre-training, and is most improved by post-training interventions (SFT, DPO) in the cases most diagnostic of mentalizing (false belief, implicit).
    • EN Key Points:
      • arXiv:2606.28524v1 Announce Type: new
      • Abstract: Recent work suggests that Large Language Models (LLMs) are sensitive to the belief states of agents described by text, as measured by the false belief…
      • We adopt a developmental perspective, tracing the pattern of mental state reasoning behavior – and likely preconditions for this behavior – across mul…
      • We find that above-chance FBT performance depends both on model size and sufficient training volume, emerges relatively late in pretraining, and is most improve…
  • A French OSCE Dialogue Dataset and Controllable Virtual Patient System for Clinical Training

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Summary: - arXiv:2606.28526v1 Announcement Type: New.
      • Abstract: The clinical and communication skills of medical students are typically assessed through Objective Structured Clinical Examinations (OSCEs), which involve short, scenario-driven simulations of doctor-patient interactions.
  • However, training is often limited by the low availability of human standardized patients, motivating the development of realistic virtual patients (VPs).

  • To address this gap, we introduce a French OSCE dialogue dataset comprising 240 student-patient training interactions.

    • EN Highlights:
      • arXiv:2606.28526v1 Announce Type: new
      • Abstract: The clinical and communication skills of medical students are commonly assessed through Objective Structured Clinical Examinations (OSCEs), which cons…
      • However, training is often limited by the low availability of human standardized patients, motivating the development of realistic virtual patients (VPs)
      • To address this gap, we introduce a French OSCE dialogue dataset comprising 240 student-patient training interactions
  • Legal Domain Adaptation of Modern BERT Models

    • Release Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.28538v1 Announce Type: new.
      • Abstract: We investigate the domain adaptation of modern BERT models in the legal domain.
      • We use the masked language modeling objective to further pre-train ModernBERT on all US court opinions.
      • Although ModernBERT has been trained on roughly 500x more data than original BERT, we still find that this model benefits from further pre-training and domain adaptation in the legal domain: we report significant improvements on all datasets related to US court opinions compared to the plain ModernBERT.
    • EN Highlights:
      • arXiv:2606.28538v1 Announce Type: new
      • Abstract: We investigate domain adaptation of modern BERT models in the legal domain
      • We further pre-train ModernBERT on all US court opinions using the masked language modeling objective
      • Although ModernBERT has been trained on roughly 500x more data than original BERT, we still find that this model benefits from further pre-training and domain a…
  • Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution

    • Release Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.28548v1 Announce Type: new.
      • Abstract: Sparse autoencoders (SAEs) have become a useful tool for extracting interpretable features in language models.
      • However, standard SAE architectures operate on single token activations, which means the number of active features scales linearly with context length, and studying long model transcripts becomes difficult.
      • We introduce turn-averaged SAEs, which represent a single human or assistant turn with a fixed number of features by learning to reconstruct the average model activation over the entire turn.
    • EN Highlights:
      • arXiv:2606.28548v1 Announce Type: new
      • Abstract: Sparse autoencoders (SAEs) have become a useful tool for extracting interpretable features in language models
  • However, standard SAE architectures operate on individual token activations, meaning that the number of active features scales linearly with context length, and…

  • We introduce turn-averaged SAEs, which represent a single Human or Assistant turn with a fixed number of features by learning to reconstruct the average model a…

  • Depth-Staggered Fibonacci Spacing for Sparse Attention: Static Schedules Beat Learned Dilation and Extrapolate Where Dense Attention Fails

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.28560v1 Announcement Type: New.
      • Abstract: We study sparse self-attention, where each query attends to a dense local window plus a set of Fibonacci-spaced offsets, with a per-layer scalar alpha to compress or expand the spacing.
      • Across 21 language models trained under a matched recipe (60M parameters, 512 hidden, 16 layers, 426M tokens), we compare four ways of setting alpha across depth: fixed, learned per-layer, static linear staggering, a coprime (anti-grid) reassignment of that stagger, and a range-matched power-of-2 control.
      • First, static per-layer staggering improves perplexity over fixed and learned alphas, and the gain is base-agnostic: applying the same stagger to a power-of-2 base lifts it above fixed Fibonacci and on par with learned Fibonacci attention.
    • EN Highlights:
      • arXiv:2606.28560v1 Announce Type: new
      • Abstract: We study sparse self-attention in which each query attends to a dense local window plus a set of Fibonacci-spaced offsets, with a per-layer scalar alp…
      • Across 21 language models trained under one matched recipe (60M parameters, 512 hidden, 16 layers, 426M tokens), we compare four ways of setting alpha across de…
      • Three results stand out
  • SEAD: Competence-Aware On-Policy Distillation via Entropy-Guided Supervision

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.28562v1 Announcement Type: New.
      • Abstract: On-policy distillation (OPD) has a property absent in offline distillation and reinforcement learning: teacher supervision quality depends on student competence.
      • Incoherent rollouts yield noisy gradients; already-mastered tokens yield redundant ones.
      • This creates waste at three levels (token, training stage, and prompt), yet existing methods apply uniform supervision.
    • EN Highlights:
      • arXiv:2606.28562v1 Announce Type: new
      • Abstract: On-policy distillation (OPD) has a property absent in offline distillation and RL: teacher supervision quality depends on student competence
      • Incoherent rollouts yield noisy gradients; already-mastered tokens yield redundant ones
  • This creates waste at three scales (tokens, training phases, and prompts) yet existing methods supervise uniformly

  • Correct codes for the wrong reasons? validating LLMs as measurement instruments for theoretical constructs

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.28574v1 Announcement Type: New.
      • Abstract: When a large language model (LLM) codes a construct in text as a human annotator would, that agreement makes the LLM a reliable coder.
      • However, reliability leaves construct validity untouched.
      • The instrument may be theory-naive, reaching the code through a correlate that meets none of the demands the construct’s theory makes, and no current method can account for this apart from genuine measurement.
    • EN Highlights:
      • arXiv:2606.28574v1 Announce Type: new
      • Abstract: When a large language model (LLM) codes a construct in text as a human annotator would, that agreement makes the LLM a reliable coder
      • Yet reliability leaves construct validity untouched
      • The instrument may be theory-naive, reaching the code through a correlate that meets none of the demands the construct’s theory makes, and no current method tel…
  • Phonological Perception of Sign Language Models

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.28667v1 Announcement Type: New.
      • Abstract: Sign languages are compositional systems where meaning arises by combining sublexical phonological parameters, such as handshape, location, and movement.
      • While deep learning models for Sign Language Recognition (SLR) have achieved increased performance on translation benchmarks, it remains unclear whether these models distinguish abstract phonological features or merely rely on low-level statistical correlations.
      • This work evaluates the phonological perception of SLR models trained on American Sign Language (ASL) by probing phonological sensitivity using minimal pairs and assessing representational alignment with human behavioral data.
    • EN Highlights:
      • arXiv:2606.28667v1 Announce Type: new
      • Abstract: Sign languages are compositional systems where meaning arises by combining sublexical phonological parameters, such as handshape, location, and moveme…
      • While deep learning models for Sign Language Recognition (SLR) have achieved increased performance on translation benchmarks, it remains unclear whether these m…
      • This work evaluates the phonological perception of SLR models trained on American Sign Language (ASL) by probing phonological sensitivity using minimal pairs an…

ArXiv cs.LG (B_intro+search) Link to heading

  • Can AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal Models

  • Publication Time: 2026-06-30 12:00 Beijing Time

    • Abstract: - arXiv:2606.28406v1 Announcement Type: New.
      • Abstract: Text-to-image and multimodal generative models are increasingly used to generate scientific figures, such as mechanism diagrams, experimental design schematics, conceptual frameworks, and graphical abstracts.
      • However, existing image generation benchmarks (e.g., GenEval, T2I-CompBench, DPG-Bench) evaluate natural images and measure compositionality, object counting, or photorealism.
      • None of them measure the usability factors of generated scientific figures: correct and clear text labels, faithful depiction of entities and their relationships, coherent diagram structure, and adherence to disciplinary drawing conventions.
    • EN Highlights:
      • arXiv:2606.28406v1 Announce Type: new
      • Abstract: Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design sch…
      • Yet existing image-generation benchmarks (e.g., GenEval, T2I-CompBench, DPG-Bench) evaluate natural images and measure compositionality, object counting, or pho…
      • None of them measure what makes a generated scientific figure usable: correct and legible text labels, faithful depiction of entities and their relations, coher…
  • On the Necessity of a Liquid Substrate for Mesh Intelligence

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.28413v1 Announcement Type: New.
      • Abstract: A mesh of sovereign agents has no center: no shared clock, no shared model, and no coordinator to gather data or retrain.
      • Its competence depends on each agent folding the predictions emitted by its peers into a single online internal state, from observations that arrive at irregular, unscheduled times, on a substrate whose weights cannot be retrained.
      • Any one of these constraints can be handled individually; folding optimally under all three at once is not.
    • EN Highlights:
      • arXiv:2606.28413v1 Announce Type: new
      • Abstract: A mesh of sovereign agents has no center: no shared clock, no shared model, and no coordinator to gather data or retrain
      • Its competence rests on each agent folding the projections its peers emit into a single internal state, online, from observations that arrive at irregular, unsc…
      • Any one of these constraints is tractable on its own; folding optimally under all three at once is not
  • Position: RL Researchers Need to Distinguish Between Solving Simulators and Using Simulators as a Proxy

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.28433v1 Announcement Type: New.
      • Abstract: One of the goals of Reinforcement Learning (RL) research is to understand general sequential decision-making, using benchmark simulators as a proxy for learning in deployment settings.
      • However, when running experiments, the goal of achieving high performance in a simulator can turn into a focus on solving the simulator problem.
      • To get high scores, researchers may adopt solutions specifically designed to solve the simulator problem, rather than for learning when the agent is deployed outside the simulator.
  • EN Highlights:

    • arXiv:2606.28433v1 Announce Type: new
    • Abstract: One goal in reinforcement learning (RL) research is to understand general-purpose sequential decision-making, using benchmark simulators as a proxy fo…
    • When running experiments, however, the goal of achieving high performance in the simulator can mutate into focusing exclusively on solving the simulator
    • To achieve high scores, researchers may adopt solutions exclusively meant for solving simulators, rather than learning while the agent is deployed outside a sim…
  • Learning to Distributedly Estimate under Partially Known Dynamics: A Covariance-Agnostic Neural Kalman Consensus Filter

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.28441v1 Announce Type: new.
      • Abstract: Online latent state estimation constitutes a fundamental challenge within the artificial intelligence field, serving as a foundational tool for various applications, including sequential decision-making, anomaly and change point detection.
      • In this paper, a novel online distributed sensing framework is proposed, where agents collaborate and exchange information to perform latent state estimation.
      • The proposed estimator combines available partial domain knowledge with the representation capabilities of deep neural networks.
    • EN Highlights:
      • arXiv:2606.28441v1 Announce Type: new
      • Abstract: Online latent state estimation constitutes a fundamental challenge within the artificial intelligence field, serving as a foundational tool for divers…
      • In this paper, a novel online distributed sensing framework, where agents collaborate and exchange information to perform latent state estimation, is presented
      • The proposed estimator combines available partial domain knowledge with the representation capabilities of deep neural networks
  • S-GAI: Spectral Geometry-Aware Initialization for Sigmoidal MLPs – From Dataset Geometry to Network Weights

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.28444v1 Announce Type: new.
      • Abstract: The classic universal approximation theorems establish the expressive power of sigmoidal multilayer perceptrons, but they do not prescribe how initial weights should encode the geometry of the data distribution.
      • We propose S-GAI, a spectral geometry-aware initialization framework for single hidden layer sigmoidal MLPs.
      • Starting from the constructive idea that sigmoid units can act as smooth half-space gates, we move from hand-specified planar geometries to class-level spectral geometries estimated from image data.
    • EN Highlights:
      • arXiv:2606.28444v1 Announce Type: new
  • Abstract: Classical universal approximation theorems establish the expressive power of sigmoidal multilayer perceptrons, but they do not prescribe how initial w…

  • We propose S-GAI, a spectral geometry-aware initialization framework for one-hidden-layer sigmoidal MLPs

  • Starting from the constructive idea that sigmoid units can act as smooth half-space gates, we move from hand-specified planar geometry to class-wise spectral ge…

  • scKDGM: KAN-guided Dynamic Graph Masked Learning for Single-Cell RNA-seq Clustering

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.28459v1 Announcement Type: new.
      • Abstract: Single-cell RNA sequencing (scRNA-seq) clustering is crucial for identifying cell types, but high dimensionality, sparsity, dropout, and technical noise hinder robust expression representation and cell graph construction.
      • Existing masked autoencoders primarily use expression recovery for feature reconstruction, while graph clustering methods often rely on a fixed KNN graph and do not feed the recovered expression back into graph optimization.
      • We propose scKDGM, a KAN-guided dynamic graph masked learning framework for scRNA-seq clustering.
    • EN Key Points:
      • arXiv:2606.28459v1 Announce Type: new
      • Abstract: Single-cell RNA sequencing (scRNA-seq) clustering is essential for identifying cell types, but high dimensionality, sparsity, dropout, and technical n…
      • Existing masked autoencoders mainly use expression recovery for feature reconstruction, while graph clustering methods usually depend on fixed KNN graphs and do…
      • We propose scKDGM, a KAN-guided dynamic graph masked learning framework for scRNA-seq clustering
  • Counterfactual Residual Data Augmentation for Regression

    • Publication Time: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.28460v1 Announcement Type: new.
      • Abstract: Data-driven modeling in real-world regression tasks often suffers from limited training samples, high collection costs, and noisy observations.
      • Inspired by the impact of data augmentation in vision and language, we propose a novel Counterfactual Residual Data Augmentation (CRDA) technique for tabular regression.
      • Our main insight is that once a regressor has modeled the systematic components of the data, the remaining noise can be treated as an invariant residual that remains stable under small perturbations of carefully selected features.
    • EN Key Points:
      • arXiv:2606.28460v1 Announce Type: new
      • Abstract: Data-driven modeling in real-world regression tasks often suffers from limited training samples, high collection costs, and noisy observations
  • Inspired by the impact of data augmentation in vision and language, we propose a novel Counterfactual Residual Data Augmentation (CRDA) technique for tabular re…

  • Our key insight is that once a regressor has modeled the systematic component of the data, the remaining noise can be viewed as an invariant residual that remai…

  • Singular Learning and Occam’s Razor in Deep Monomial Networks

    • Posted: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.28464v1 Announcement Type: new.
      • Abstract: In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model’s architecture.
      • These critical points occur where the Jacobian of the model’s parametrization is rank-deficient, and are the most pronounced singularities studied in Singular Learning Theory.
      • We investigate such points in deep fully-connected networks with monomial activations via tools from polynomial algebra such as Mason’s Theorem.
    • EN Highlights:
      • arXiv:2606.28464v1 Announce Type: new
      • Abstract: In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model’s architecture
      • These critical points occur where the Jacobian of the model’s parametrization is rank-deficient, and are the most pronounced singularities studied in Singular L…
      • We investigate such points in deep fully-connected networks with monomial activations via tools from polynomial algebra such as Mason’s Theorem
  • An Agentic AI Pipeline for Appliance-Level Energy Anomaly Detection and LLM-Driven Recommendations

    • Posted: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.28467v1 Announcement Type: new.
      • Abstract: Appliance-level energy monitoring in office buildings produces noisy alerts that non-expert facility managers struggle to use.
      • This paper proposes an end-to-end agentic pipeline that combines deep time-series forecasting, variational anomaly detection, and LLM-based reasoning to generate prioritized, actionable maintenance recommendations.
      • The system uses a hybrid Singular Spectrum Analysis (SSA) and Long Short-Term Memory (LSTM) forecasting model to track seven types of office equipment and applies a per-appliance LSTM Variational Autoencoder (VAE), focusing on flagging anomalous daily consumption events.
    • EN Highlights:
      • arXiv:2606.28467v1 Announce Type: new
      • Abstract: Appliance-level energy monitoring in office buildings produces noisy alerts that non-expert facility managers struggle to use
      • This paper proposes an end-to-end agentic pipeline that combines deep time-series forecasting, variational anomaly detection, and LLM-based reasoning to generat…
  • The system tracks seven office appliances using a hybrid Singular Spectrum Analysis (SSA) and Long Short-Term Memory (LSTM) forecasting model, and applies a per…

  • Modelling Emotional Memory in Children with Tensor Networks

    • Published: 2026-06-30 12:00 Beijing Time
    • Abstract: - arXiv:2606.28470v1 Announcement Type: New.
      • We demonstrate how emotional valence influences the order-dependent structure of children’s recognition memory: correct recall of a sequence of emotionally-valenced toys depends not only on the valence of a given toy itself, but also on the valence of the toys presented before and after it.
      • Whilst standard psychological models confirm that order-dependence differs across an event (a set of toys shown in sequence), accuracy is low and the model fails to reflect how memory for an emotional object affects other objects within the set.
      • A classical tensor network model factoring in valence is able to achieve a 77.98% accuracy in modelling the results of the study.
    • EN Highlights:
      • arXiv:2606.28470v1 Announce Type: new
      • Abstract: We demonstrate how emotional valence influences the order-dependent structure of children’s recognition memory: correct recall of a sequence of emotio…
      • Whilst standard psychological models confirm that order-dependence differs across an event (a set of toys shown in sequence), accuracy is low and the model does…
      • A classical tensor network model factoring in valence is able to achieve a 77.98% accuracy in modelling the results of the study