🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-06-30
- 类型
- ai-daily
- 字数
- 7760
- 阅读时长
- 37 min
2026-06-30 AI Daily Update | Agents Moving Towards Long-Term Planning, AI Competition Enters Verifiable Systems Stage Link to heading
Today’s main theme is no longer simply pursuing larger models, but rather Agent long-term planning, multi-agent collaboration, and trustworthy deployment. Frontier models are still accelerating their iteration, but restricted access, safety unblocking, open-weight controversies, and real-world scenario feedback are jointly reshaping the boundaries of AI competition.
📖 This Issue’s Watch List In-Depth Guide Link to heading
Today’s Watch List recommends focusing on three main areas: Firstly, the underlying capabilities of agent “long-term planning.” Multiple arXiv papers discuss world models, symbolic feedback, self-iteration, and hallucination propagation, indicating that Agent research is shifting from prompt engineering to verifiable, measurable planning mechanisms. Secondly, multi-agent and model networks. From persona composition to AI-Model Network, it’s worth observing how large model collaboration paradigms impact task performance and system architecture. Thirdly, trustworthy AI in vertical scenarios: work such as fact-checking, legal adaptation, clinical training, and support for dyslexic learners are all applying model capabilities to high-risk, low-resource, or strong-explanation-demand scenarios. Overall, today’s key deep-read keywords are not “larger models,” but “controllable, verifiable, and deployable.”
🌐 X Platform AI Hot News Briefs Link to heading
话题 1:Zhipu AI Seeks Community Input on GLM-5.3 Vision Features Link to heading
- 分类:AI · News
- 概况:Trending Time:19 hours ago,Related Posts:1300
- What it is:Zhipu AI is soliciting community feedback and improvement suggestions for GLM-5.3’s vision capabilities.
- Why it matters:This indicates that multimodal large model competition is shifting from mere parameters and leaderboard performance to iteration focused on real-world use cases, product experience, and developer needs.
- Discussion Overview:Discussions on X primarily focus on whether GLM-5.3 should strengthen image understanding, visual reasoning, OCR, video capabilities, and agent workflow integration; some also question whether community feedback can truly influence the product roadmap, and the gap between domestic models and models like Claude, GPT, and Gemini in the multimodal domain.
话题 2:DeepSeek V4 Models Launch Officially in Mid-July with Peak Pricing Link to heading
- 分类:AI · News
- 概况:Trending Time:23 hours ago,Related Posts:2800
- What it is:DeepSeek is reportedly launching its V4-Pro and V4-Flash models officially in mid-July, featuring 1M token context, open weights, faster inference, and lower output pricing.
- Why it matters:If the stated performance, open-source licensing, and pricing information are accurate, DeepSeek V4 will further drive competition among large models in long context, code capabilities, inference efficiency, and low-cost deployment, putting price and open ecosystem pressure on closed-source model vendors like OpenAI, Google, Anthropic, and xAI.
- Discussion Overview:Discussions on X focus on whether DeepSeek will truly open high-performance weights with an MIT license, the credibility of 1.6 trillion parameters and SWE-bench scores, the sustainability of low prices, and its impact on global AI model pricing and the China-US AI competition landscape.
话题 3:Flexion Robotics Unveils Reflect v1.0 for Long-Horizon Robot Autonomy Link to heading
- 分类:AI · News
- 概况:Trending Time:14 hours ago,Related Posts:249
- What it is:Flexion Robotics has released Reflect v1.0, focusing on long-horizon robot autonomy capabilities.
- Why it matters:This indicates that the robotics field is moving from short-command execution to more complex continuous planning, environmental adaptation, and error recovery capabilities, which is a crucial direction for the realization of general embodied AI.
- Discussion Overview:Discussions on X primarily focus on whether Reflect v1.0 can significantly enhance stability and generalization capabilities in real-world scenarios, with some also concerned about its differences from existing robot foundation models, simulation training, end-to-end control solutions, and its commercial prospects.
话题 4:Dario Amodei’s 2023 Open-Source AI Warning Resurfaces Amid Model Surge Link to heading
- 分类:AI · Other
- 概况:Trending Time:1 day ago,Related Posts:21000
- What it is:Anthropic CEO Dario Amodei’s 2023 warning about the potential safety risks of open-source AI has resurfaced on X due to the rapid growth of open-source and open-weight models recently.
- Why it matters:This reflects the core contradiction within the AI field between model capability proliferation, transparency, innovation speed, and safety governance, especially as high-performance models become increasingly easy to access and deploy.
- Discussion Overview:Discussions on X mainly revolve around whether open-source AI promotes competition and democratization or lowers the barrier to misuse; supporters emphasize that an open ecosystem facilitates auditing and innovation, while critics worry that powerful models could be used for cyberattacks, fraud, or to circumvent regulation.
Topic 5: xAI Launches Grok 4.5 Private Beta at SpaceX and Tesla Link to heading
- Category: AI · News
- Overview: Trending for: 1 day ago, Related posts: 43,000
- What it is: Musk announced that xAI’s Grok 4.5 has entered private beta internally at SpaceX and Tesla. Early performance is reported to be close to or exceeding some benchmarks of Claude Opus, with a public release pending further validation.
- Why it matters: The significance of this isn’t just about individual model performance. It’s about xAI potentially creating a closed loop that spans from computing power, foundational models, and coding agents to internal enterprise deployment, real-world task feedback, and rapid retraining. This reflects a shift in cutting-edge AI from single model releases to a high-frequency, iterative “model factory” paradigm.
- Discussion summary: Discussions on X are focused on whether Grok 4.5 truly has Opus-level capabilities, the credibility of its purported 1.5T parameters and Cursor-related training data, and the associated data rights issues. There’s also debate on whether internal scenarios at Tesla and SpaceX can provide high-quality reinforcement learning feedback. The division lies between supporters, who believe xAI’s release speed and vertical integration will challenge OpenAI and Anthropic, and skeptics, who highlight the lack of public, reproducible, third-party evaluations in a comparable tool environment.
AI Public Opinion Summary on X Today Link to heading
The main narrative today is that the AI competition is shifting from a focus on single-point model performance to a comprehensive battle of “capability, cost, open ecosystems, and real-world feedback.” Multimodality, long context, long-duration robotics tasks, and closed-loop iteration within enterprises are all becoming focal points. The consensus is that model providers must get closer to actual use cases, building their advantage through developer feedback, low-cost deployment, continuous training, and productization capabilities, rather than relying solely on leaderboard rankings. Key points of contention include the trade-off between open-sourcing weights and security governance, and the credibility of the parameter sizes, benchmark scores, licensing promises, and internal test results of new models like DeepSeek and xAI. Potential risks include high-performance open-source models lowering the bar for misuse, model promotions lacking reproducible validation, disputes over training data rights, and the gap between safety and commercialization due to the instability of robotics and agent systems in real-world environments.
💡 Influencer Insights Link to heading
Based on an analysis of tweets from leading AI influencers over the past 24 hours, here is your industry briefing.
AI Influencer Daily Briefing: Cutting-Edge Models Accelerate Iteration, AI Integration into Workflows Becomes the Main Theme Link to heading
1. Core Technology and Product Hotspots: A Flurry of Cutting-Edge Model Releases, On-Device and Agent Infrastructure Emerge as New Battlegrounds Link to heading
Today’s hot topics center on the iteration of foundational models, the rapid evolution of development tools, and how AI can be more deeply embedded into organizational workflows.
OpenAI’s GPT-5.6 Series Release and Restricted Access Spark Heated Debate This was the top story of the day. OpenAI has released its new-generation model, GPT-5.6, which comes in three versions: the flagship Sol, the balanced Terra, and the economy Luna.
- @dotey provided a detailed breakdown of the model’s information: Sol supports
max(deep reasoning) andultra(multi-agent parallel) modes, excels in programming benchmarks, and features prominent safety mechanisms. However, the most striking detail is that the model is currently available only to about 20 partners approved by the U.S. government, leaving general users and developers without access for now. @dotey believes this signals that pre-release reviews of cutting-edge AI models by the U.S. government are shifting “from isolated cases to standard practice.” - Meanwhile, the community is already looking for signs of A/B testing. Bloggers like @dotey and @Pluvio9yte shared methods for checking if their accounts have been included in the A/B test for GPT-5.6 Sol using the “Juice Test Prompt” and the Codex Analytics dashboard, sparking widespread discussion.
- @dotey provided a detailed breakdown of the model’s information: Sol supports
Key Developments in the Claude Ecosystem: Mythos 5 Unbanned and the Claude Tag Paradigm
- @dotey reported that Anthropic’s Mythos 5, after being banned for two weeks due to safety concerns, has received a partial lifting of the ban from the U.S. government. It will be redeployed to about 100 U.S. agencies for cybersecurity defense. @dotey commented, “A model taken down for being too dangerous is brought back for being too useful,” which poignantly reflects the trade-offs between safety and utility in cutting-edge models.
- At the same time, Anthropic’s release of the Claude Tag feature has triggered in-depth discussions about AI and team collaboration models. @dotey relayed @GergelyOrosz’s understanding, emphasizing that its breakthrough isn’t Slack, but rather that “a cloud-based AI works out-of-the-box once connected to a company’s internal systems.” @dotey also cited @karpathy’s view, calling it the “third major redesign” of LLM interaction methods, following web and app versions.
Rapid Development of On-Device Models and Applications
New Model Paradigm: @zhixianio continues to focus on on-device models, envisioning the “Model-Pak” concept where models are plug-and-play like game cartridges, believing this represents the future direction.
Performance Testing: @zhixianio tested the full-duplex audio and video capabilities of MiniCPM-o 4.5 and found the results satisfactory. He also conducted detailed local tests on Google’s Gemma 4 series (12B Coder). The results showed that while the 12B model excels at specific tasks, it still has a significant gap compared to the 35B Qwen model when faced with complex programs like “Tetris” that require long, stateful, single-shot generation, indicating that a performance “ceiling” for models still exists.
Breakthrough in Non-Invasive Brain-Computer Interfaces @dotey shared Meta’s Brain2Qwerty v2, which uses MEG (Magnetoencephalography) devices to decode brain signals into sentences in real-time, achieving an average word accuracy of 61%, with top performers reaching 78%. @dotey emphasized that this non-invasive breakthrough proves “it’s possible to achieve results close to invasive methods without surgery,” offering new hope for individuals who have lost their ability to communicate.
2. Unique Perspectives and Industry Foresight Link to heading
On the True Cost and Value of AI Programming:
- @ruanyf, citing the case of the OpenClaw founder’s exorbitant Token consumption ($1.3 million in one month), calmly pointed out that if top-tier models are used without limit, the Token cost of AI programming could be more expensive than human programmers.
- @gefei55 offered another perspective: “Tokens are infinite, but time and energy are finite.” He warned against getting lost in the “Token trap of being able to do anything,” urging people to avoid a meaningless drain on their lives.
- @dotey used Ford as a counterexample: when the company’s AI quality inspections failed to meet expectations, it re-hired 350 senior engineers to “tune” the AI and secured the top spot in new car quality rankings. This demonstrates that an AI’s effectiveness depends on the quality of the people and data that train it, and that human-machine collaboration is still the reality.
Reflections on AI Products and Business Models:
- @zhixianio was impressed with Fable, stating that it completed 70% of a demo’s work in 40 minutes. It not only implemented the required features but also identified shortcomings in the original design and proposed a better solution. This signifies a paradigm shift for AI Agents, moving from passive execution to active design participation.
- @Pluvio9yte shared that Volcano Engine’s Coding Plan has entered the market at the ultra-low price of 9.9 RMB/month, lamenting how “abstract” and fierce the competition has become.
Social and Humanistic Insights in the AI Era:
- @ruanyf keenly posed a profound question: “If all future code is written by AI, how do we hire programmers?” He pointed out that evaluating a candidate’s ability to use AI is becoming more important than evaluating their coding ability itself.
- @lijigang continued to offer philosophical insights, with quotes like “Taste is a person’s loss function” and “In the digital world, people live on likes.” From a cognitive science perspective, he observed that heavy use of a model can lead to adopting its linguistic style, which relates to the brain’s tendency to “consume” context.
Redefining “Open Source”: @ruanyf cited the view of Anthropic’s founder, pointing out that publicizing an AI model’s weights is not equivalent to traditional open source because you cannot see its internal operations or participate in its development. This is more accurately described as “open weights.”
3. Recommended Tools and Resources Link to heading
AI Development and Productivity Tools:
- RepoPrompt (Community Edition): Recommended by @dotey. This open-source tool for “context engineering” now features a new architecture that uses inference models for task decomposition, distributing tasks to different Agents for parallel execution.
- Efficient Codex Usage Tips: @Pluvio9yte open-sourced a Skill management tool that scans and tallies the usage frequency of all installed Skills to help with “decluttering.” Meanwhile, @dotey shared session management tips such as
/btwandfork. - EdgeOne Makers: Recommended by @vista8. This is an Agent development platform from Tencent Cloud that allows deploying an AI Agent framework with just three commands. It solves deployment challenges like sandboxing, persistence, and observability, making it ideal for developers looking for rapid implementation.
Content Creation and SEO Tools:
- Video Creation Skills Library: @Pluvio9yte has open-sourced a complete suite of video creation Skills that can directly replicate videos in specific styles, lowering the barrier to entry for video production.
GEO (Generative Engine Optimization) Resource Pack: After conducting a GEO masterclass, @vista8 generously shared a complete set of materials, including an operational manual, system research reports, prompts, and Skills. This is extremely valuable for anyone looking to optimize content in the age of AI search.
Twitter Hot Topic Mining Tool: @gefei55 has open-sourced a small tool that scans high-interaction tweets with links by calling the X API and checks the domain traffic in reverse. It helps users discover new terms and hot topics faster than Google Trends by leveraging information gaps.
Practical Resources & Learning Materials:
- Codex Learning Treasure Trove: Jointly recommended by @AI_Jasonyu and @dotey. @bozhou_ai and @dotey have respectively open-sourced the “Orange Book” for systematically learning Codex and the “Claude Code From Scratch” e-book. The latter reproduces the core architecture of Claude Code in 4,300 lines, making it an excellent resource for deeply understanding Coding Agents.
- ChatGPT Management Plugin: @Pluvio9yte developed and recommended a Chrome plugin that can batch delete or archive chat records on the ChatGPT web interface, solving a major pain point for users.
Summary: In the last 24 hours, the industry has been firing on all cylinders. On one hand, frontier models represented by the GPT-5.6 and Claude 5 series are fiercely competing on capabilities, pricing, and security compliance. On the other hand, AI is transforming from a standalone tool into a fundamental infrastructure embedded in organizations, integrated into workflows (like Claude Tag), and empowering individuals (through on-device models and automated Skills). The discussions among industry leaders have shifted from simply comparing model capabilities to deeper considerations of cost, efficiency, compliance, and how to truly transform AI into productivity.
📚 Appendix: Today’s Watch List Update Source List Link to heading
Time window: Last 3 days; covers 22 sources; 33 updates in total
All-In Podcast (A_full) Link to heading
- Nate Silver Predicts: Democrats Take the House, Newsom Is Fading & AOC Might Win It All in 2028
- Published: 2026-06-29 23:47 Beijing Time
- Summary: - Northwest Registered Agent - Starting a business?
- Northwest Registered Agent gives you everything you need to build a complete business identity, including free tools and privacy by default.
- Visit Northwestregisteredagent.com/ALLINFREE to learn more.
- PLAUD - If your job relies on conversations - meetings, deal flow, interviews, customer calls - Plaud helps you capture and organize it all with highly accurate AI-generated notes that go beyond simple summaries to highlight pain points, key decisions, next steps, and customizable summary templates.
- Check out Plaud at Plaud.ai/allin and use code ALLIN for up to 20% off!
- EN Highlights:
- (0:00) Nate Silver joins the pod
- (10:02) California’s ballot counting problem: Raman’s late-mail surge, ballot harvesting claims, and why the US counts slower than India
- (25:18) Democrats’ three-way civil war: The left, the abundance libs, and Newsom’s “resistance lib” base
- (34:48) The winning 2028 playbook: Anti-oligarch messaging, why young men want control, and immigrants fleeing the Dems
Stratechery by Ben Thompson (A_full) Link to heading
- Summer Break: Week of June 29
- Published: 2026-06-29 18:00 Beijing Time
- Summary: - Stratechery is on summer break for the week of June 29.
- There will be no Weekly Articles or Updates.
- The next update will be on Monday, July 6.
- Dithering, Sharp Tech, and Sharp China will….
- EN Highlights:
- Stratechery is on summer break the week of June 29
- There will be no Weekly Article or Updates
- The next Update will be on Monday, July 6
- Dithering , Sharp Tech , and Sharp China will also return the week of July 6
OpenAI Blog (A_full) Link to heading
- Mapping Europe’s AI Workforce Opportunity
- Publication Time: 2026-06-29 15:00 Beijing Time
- Summary: - AI capabilities can cross national borders quickly.
- Changes in work will not be so smooth.
- Jobs are determined by licensing systems, local institutions, and the practicalities of providing care, education, justice, public services, and other forms of human support.
- These systems are important because they help determine if and how AI will change the labor market.
- What impact will AI have on the labor market?
- EN Key Points:
- A new OpenAI report maps how AI could reshape jobs across the EU, highlighting which occupations may face automation, growth, or workflow changes.
ArXiv cs.AI (B_intro+search) Link to heading
AI-Model Network: Concept, Current State and Future
- Publication Time: 2026-06-30 12:00 Beijing Time
- Summary: - arXiv:2606.27382v1 Announce Type: new.
- Abstract: The primary function of computers lies in computation and processing, while the core value of the Internet is rooted in sharing and collaboration.
- Computers create the Internet, and the Internet empowers the value of computers.
- The rapid development of the Internet, cloud computing, and big data is pushing artificial intelligence into the era of large models (LMs).
- EN Key Points:
- arXiv:2606.27382v1 Announce Type: new
- Abstract: While the primary function of computers lies in computation and processing, the core value of the Internet is rooted in sharing and collaboration
- Computers create the Internet, and the Internet empowers the value of computers
- The rapid development of the Internet, cloud computing, and big data is pushing artificial intelligence into the era of large models (LMs)
When Does Personality Composition Matter for Multi-Agent LLM Teams?
- Publication Time: 2026-06-30 12:00 Beijing Time
- Summary: - arXiv:2606.27443v1 Announce Type: new.
- Abstract: Personality prompts determine the communication style of large language models, but whether these behavioral shifts affect objective task outcomes remains to be explored.
- Previous research has shown that agents with low-agreeableness prompts produce adversarial language, while agents with high-agreeableness prompts become cooperative, but the relationship between communication style and task performance has not been systematically examined across multiple domains.
- In this work, we investigate whether personality composition matters for multi-agent team performance by manipulating the personality traits of frontier LLMs across three task domains: structured coding, open-ended research collaboration, and competitive negotiation.
- EN Key Points:
- arXiv:2606.27443v1 Announce Type: new
Abstract: Personality prompting shapes how large language models communicate, yet whether these behavioral shifts affect objective task outcomes remains under-e…
Prior work shows that agents prompted with low agreeableness produce adversarial language, while those prompted with high agreeableness become cooperative, but…
In this work, we investigate whether personality composition matters for multi-agent team performance by manipulating personality traits across frontier LLMs on…
Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning
- Published: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.27483v1 Announcement Type: New.
- Abstract: Large language model (LLM) agents have demonstrated strong capabilities in sequential decision-making, yet they remain fundamentally reactive in long-term tasks.
- Unlike humans who use “what-if” reasoning to evaluate potential plans before commitment, standard agents lack an internal world model to simulate future outcomes.
- Therefore, we propose to internalize future-aware planning by training a single autoregressive model to express both prospective state rollouts and success estimations conditioned on the plan (a textual simulation of Q-values).
- EN Highlights:
- arXiv:2606.27483v1 Announce Type: new
- Abstract: Large language model (LLM) agents have demonstrated strong capability in sequential decision-making, yet they remains fundamentally reactive in long-h…
- Unlike humans who employ “what-if” reasoning to evaluate potential plans before commitment, standard agents lack an internal world model to simulate future outc…
- Therefore, we propose to internalize future-aware planning by training a single autoregressive model to verbalize both a prospective state rollout and a plan-co…
Odyssey: Constructing Verifiable Local Truth-Preserving Foundation Models
- Published: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.27593v1 Announcement Type: New.
- Abstract: We introduce a categorical framework named ODYSSEY for constructing verifiable, locally truth-preserving foundation models as a composition of foundries: architectural components of building blocks specifying local contexts, families of local representations, restriction graphs, gluing rules, obstruction policies, update obligations, and human-facing views.
- A foundry is an organized knowledge base containing argumentation components.
- Concrete foundries are built from generic ones, such as those for evidence/argumentation, operational decisions, institutional/financial matters, market significance, scientific challenges, research programs, and assisted construction and evaluation rigs.
- EN Highlights:
- arXiv:2606.27593v1 Announce Type: new
Abstract: We introduce a categorical framework called ODYSSEY for constructing verifiable, local truth-preserving foundation models as compositions of foundries…
A foundry is an organized sheaf of knowledge that carries within it an argumentation component
Concrete foundries are built from generic foundries such as evidence/argument, operational decision, institutional/financial, market meaning, scientific challen…
DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums
- Published: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.27619v1 Announce Type: new.
- Abstract: Dyslexic learners are increasingly using artificial intelligence (AI) tools to support tasks related to reading, writing, organization, and learning.
- However, their lived experiences with these tools remain largely underexamined.
- This paper proposes DysLexLens, a low-resource LLM framework designed to analyze the experience of dyslexic learners with AI through online forum discussions.
- EN Key Points:
- arXiv:2606.27619v1 Announce Type: new
- Abstract: Dyslexic learners increasingly use artificial intelligence (AI) tools to support reading, writing, organisation, and study-related tasks
- However, their lived experiences with these tools remain largely underexamined
- This paper proposes DysLexLens, a low-resource LLM framework, designed to analyse dyslexic learners experience with AI through online forum discussions
MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy
- Published: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.27652v1 Announce Type: new.
- Abstract: We find that explicit reasoning does not necessarily translate into better multimodal emotion recognition (MER) accuracy, although it makes predictions more interpretable.
- Specifically, for reasoning-based MLLMs, fast thinking by triggering direct answers often outperforms slow thinking after deliberate reasoning.
- Our empirical analysis shows that fast thinking can improve recall through broader and more confident predictions, while slow thinking improves precision by conservatively filtering incorrect categories.
- EN Key Points:
- arXiv:2606.27652v1 Announce Type: new
- Abstract: We find that explicit reasoning does not necessarily translate into better multimodal emotion recognition (MER) accuracy, even though it makes predict…
- Specifically, for reasoning-based MLLMs, fast thinking by triggering direct answers often outperforms slow thinking after deliberative reasoning
Our empirical analyses show that fast thinking improves recall with broader and more confident predictions, whereas slow thinking favors precision through conse…
- Published: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.27736v1 Announcement Type: new.
- Abstract: The rapid spread of fake news poses an increasing threat to the information ecosystem, especially as AI-generated misinformation under Generative Engine Optimization (GEO) poisoning causes retrieval systems to systematically present adversarially crafted content, thereby contaminating the reasoning of LLMs.
- In this paper, we introduce the Tree of Evidence (ToE), a hierarchical evidence reasoning framework for automated fact-checking that models each claim as a dynamically expanding argument tree.
- ToE integrates a reinforcement learning-driven multi-source retrieval agent, an evidence evaluation agent, and an argument tree aggregation algorithm to iteratively decompose, retrieve, and validate claims through an explainable chain of evidence.
- EN Highlights:
- arXiv:2606.27736v1 Announce Type: new
- Abstract: The rapid spread of fake news poses increasing threats to information ecosystems, especially as AI-generated misinformation under Generative Engine Op…
- In this paper, we propose Tree of Evidence (ToE), a hierarchical evidence reasoning framework for automated fact-checking that models each claim as a dynamicall…
- ToE integrates a reinforcement learning-driven multi-source retrieval agent, an evidence evaluation agent, and an argument tree aggregation algorithm to iterati…
- Published: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.27757v1 Announcement Type: new.
- Abstract: Large Language Models (LLMs) have garnered significant attention from both academia and industry, but their deployment raises critical security concerns regarding robustness and reliability.
- Planning, a core component of intelligent behavior, remains challenging for LLMs, which often produce infeasible or incorrect solutions in long-horizon decision-making tasks due to inherent complexity.
- In this paper, we propose a symbolic feedback-driven iterative self-refinement framework to enhance the robustness and reliability of LLMs in long-term planning.
- EN Highlights:
- arXiv:2606.27757v1 Announce Type: new
- Abstract: Large language models (LLMs) have attracted widespread attention from academia and industry, yet their deployment raises critical security concerns re…
- Planning, a core component of intelligent behavior, remains challenging for LLMs, which often produce infeasible or incorrect solutions in long-horizon decision…
In this paper, we propose a symbolic feedback-driven iterative self-refinement framework to enhance the robustness and reliability of LLMs in long-horizon plann…
Understanding Rollout Error in Graph World Models
- Published: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.27780v1 Announce Type: new.
- Abstract: World models are often used for planning by rolling learned dynamics forward.
- However, many planning environments are not vectors or images; they are graphs of agents, tools, skills, routes, and dependencies.
- In these settings, a local prediction error may stay local or spread through the graph, and the failure mode changes again when edges are predicted rather than fixed.
- EN Highlights:
- arXiv:2606.27780v1 Announce Type: new
- Abstract: World models are often used for planning by rolling learned dynamics forward
- Many planning environments, however, are not vectors or images; they are graphs of agents, tools, skills, routes, and dependencies
- In these settings, a local prediction error may stay local or spread through the graph, and the failure mode changes again when edges are predicted rather than…
- Published: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.27806v1 Announce Type: new.
- Abstract: World models for language agents come in two useful forms.
- An agent-based world model calls an LLM API and reasons flexibly in language, but its errors appear as hallucinated state changes that are hard to score with ordinary regression losses.
- A parameterized world model is a trained transition predictor; its errors are easier to measure with quantities such as NodeMSE, delta accuracy, and validity accuracy, but it is often weaker as a standalone planner.
- EN Highlights:
- arXiv:2606.27806v1 Announce Type: new
- Abstract: World models for language agents come in two useful forms
- An agent-based world model calls an LLM API and reasons flexibly in language, but its errors appear as hallucinated state changes that are hard to score with or…
- A parameterized world model is a trained transition predictor; its errors are easier to measure with quantities such as NodeMSE, delta accuracy, and validity ac…
ArXiv cs.CL (B_intro+search) Link to heading
Generating in the Limit with Infinitely Many Hallucinations
- Published: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.28354v1 Announce Type: new.
Abstract: The classic paradigm of language identification in the limit models learning as a game between an adversary (who reveals strings from an unknown target language) and a learner (responsible for identifying that language).
The recently introduced framework of language generation in the limit shifts the objective to better reflect modern language modeling, requiring the learner to generate valid, unseen strings from the target language.
Related work highlights a fundamental tension: broad coverage of the target often comes at the cost of validity.
- EN Key Points:
- arXiv:2606.28354v1 Announce Type: new
- Abstract: The classic paradigm of language identification in the limit models learning as a game between an adversary, who reveals strings from an unknown targe…
- The recently introduced framework of language generation in the limit shifted the objective to better reflect modern language modeling, requiring the learner to…
- Related work highlighted a fundamental tension: a broad coverage of the target often comes at the cost of validity
- EN Key Points:
Extracting Knowledge from an Arabic-English Machine-Readable Dictionary Using Information Extraction
- Publication Time: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.28457v1 Announce Type: new.
- Abstract: Natural language processing (NLP) applications require a large and rich amount of linguistic knowledge.
- Furthermore, electronic language resources such as dictionaries, encyclopedias, and corpora have become available.
- Therefore, automatic methods have emerged to extract lexical information from these sources to overcome the knowledge acquisition bottleneck.
- EN Key Points:
- arXiv:2606.28457v1 Announce Type: new
- Abstract: Natural language processing (NLP) applications need large and rich amount of linguistic knowledge
- Furthermore, electronic language sources such as dictionaries, encyclopedia, and corpora became available
- So, automatic methods are emerged to extract lexical information from those sources to overcome the knowledge acquisition bottleneck
Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models
- Publication Time: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.28524v1 Announce Type: new.
- Abstract: Recent work shows that Large Language Models (LLMs) are sensitive to the belief states of subjects described in text, as measured by the False Belief Task (FBT), but persistent concerns about construct validity remain.
- We adopt a developmental perspective, tracing the patterns of mental state reasoning behavior, and the possible prerequisites for this behavior, across multiple training stages in the Olmo2 and Pythia language model suites.
- We find that above-chance FBT performance depends on model size and sufficient training volume, emerges relatively late in pre-training, and is most improved by post-training interventions (SFT, DPO) in the conditions most diagnostic of mentalizing (false belief, implicit).
- EN Key Points:
- arXiv:2606.28524v1 Announce Type: new
Abstract: Recent work suggests that Large Language Models (LLMs) are sensitive to the belief states of agents described by text, as measured by the false belief…
We adopt a developmental perspective, tracing the pattern of mental state reasoning behavior – and likely preconditions for this behavior – across mul…
We find that above-chance FBT performance depends both on model size and sufficient training volume, emerges relatively late in pretraining, and is most improve…
A French OSCE Dialogue Dataset and Controllable Virtual Patient System for Clinical Training
- Publication Time: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.28526v1 Announce Type: new.
- Abstract: The clinical and communication skills of medical students are commonly assessed through Objective Structured Clinical Examinations (OSCEs), which involve brief, scenario-driven simulations of doctor-patient interactions.
- However, training is often limited by the low availability of human standardized patients, motivating the development of realistic virtual patients (VPs).
- To address this gap, we introduce a French OSCE dialogue dataset containing 240 student-patient training interactions.
- EN Key Points:
- arXiv:2606.28526v1 Announce Type: new
- Abstract: The clinical and communication skills of medical students are commonly assessed through Objective Structured Clinical Examinations (OSCEs), which cons…
- However, training is often limited by the low availability of human standardized patients, motivating the development of realistic virtual patients (VPs)
- To address this gap, we introduce a French OSCE dialogue dataset comprising 240 student-patient training interactions
Legal Domain Adaptation of Modern BERT Models
- Publication Time: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.28538v1 Announce Type: new.
- Abstract: We investigate the domain adaptation of modern BERT models in the legal field.
- We use the masked language modeling objective to further pre-train ModernBERT on all US court opinions.
- Although ModernBERT was trained on approximately 500 times more data than the original BERT, we find that the model still benefits from further pre-training and domain adaptation in the legal field: we report significant improvements on all datasets related to US court opinions compared to the standard ModernBERT.
- EN Key Points:
- arXiv:2606.28538v1 Announce Type: new
- Abstract: We investigate domain adaptation of modern BERT models in the legal domain
- We further pre-train ModernBERT on all US court opinions using the masked language modeling objective
Although ModernBERT has been trained on roughly 500x more data than original BERT, we still find that this model benefits from further pre-training and domain a…
Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution
- Publication Time: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.28548v1 Announcement Type: New.
- Abstract: Sparse autoencoders (SAEs) have become a useful tool for extracting interpretable features in language models.
- However, standard SAE architectures operate on individual token activations, meaning that the number of active features scales linearly with context length, and studying long model transcripts becomes difficult.
- We introduce turn-averaged SAEs, which represent a single Human or Assistant turn with a fixed number of features by learning to reconstruct the average model activation for the entire turn.
- EN Key Points:
- arXiv:2606.28548v1 Announce Type: new
- Abstract: Sparse autoencoders (SAEs) have become a useful tool for extracting interpretable features in language models
- However, standard SAE architectures operate on individual token activations, meaning that the number of active features scales linearly with context length, and…
- We introduce turn-averaged SAEs, which represent a single Human or Assistant turn with a fixed number of features by learning to reconstruct the average model a…
- Publication Time: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.28560v1 Announcement Type: New.
- Abstract: We study sparse self-attention, where each query attends to a dense local window plus a set of Fibonacci-spaced offsets, and uses a per-layer scalar alpha to compress or expand the spacing.
- In 21 language models trained under a matched recipe (60M parameters, 512 hidden, 16 layers, 426M tokens), we compare four methods for setting alpha across depth: fixed, per-layer learned, static linear interleaving, coprime (anti-grid) redistribution of this interleaving, and a range-matched power-of-2 control.
- First, static per-layer interleaving improves perplexity over fixed and learned alphas, and the gain is base-agnostic: applying the same interleaving to a power-of-2 base lifts it above fixed Fibonacci and on par with learned Fibonacci attention.
- EN Key Points:
- arXiv:2606.28560v1 Announce Type: new
- Abstract: We study sparse self-attention in which each query attends to a dense local window plus a set of Fibonacci-spaced offsets, with a per-layer scalar alp…
- Across 21 language models trained under one matched recipe (60M parameters, 512 hidden, 16 layers, 426M tokens), we compare four ways of setting alpha across de…
- Three results stand out
SEAD: Competence-Aware On-Policy Distillation via Entropy-Guided Supervision
- Published time: 2026-06-30 12:00 Beijing time
- Abstract:- arXiv:2606.28562v1 Announce Type: new.
- Abstract: On-policy distillation (OPD) has properties absent in offline distillation and reinforcement learning: the quality of teacher supervision depends on student competence.
- Incoherent rollouts generate noisy gradients; already-mastered tokens generate redundant ones.
- This creates waste at three levels (tokens, training phases, and prompts), yet existing methods apply uniform supervision.
- EN 要点:
- arXiv:2606.28562v1 Announce Type: new
- Abstract: On-policy distillation (OPD) has a property absent in offline distillation and RL: teacher supervision quality depends on student competence
- Incoherent rollouts yield noisy gradients; already-mastered tokens yield redundant ones
- This creates waste at three scales (tokens, training phases, and prompts) yet existing methods supervise uniformly
- Published time: 2026-06-30 12:00 Beijing time
- Abstract:- arXiv:2606.28574v1 Announce Type: new.
- Abstract: When a large language model (LLM) codes a construct in text like a human annotator would, that agreement makes the LLM a reliable coder.
- However, reliability does not affect construct validity.
- The instrument may be theoretically naive, arriving at the code through a correlate that meets none of the demands posed by the construct’s theory, and no current method can account for this beyond true measurement.
- EN 要点:
- arXiv:2606.28574v1 Announce Type: new
- Abstract: When a large language model (LLM) codes a construct in text as a human annotator would, that agreement makes the LLM a reliable coder
- Yet reliability leaves construct validity untouched
- The instrument may be theory-naive, reaching the code through a correlate that meets none of the demands the construct’s theory makes, and no current method tel…
Phonological Perception of Sign Language Models
- Published time: 2026-06-30 12:00 Beijing time
- Abstract:- arXiv:2606.28667v1 Announce Type: new.
- Abstract: Sign language is a combinatorial system where meaning is generated by combining sub-lexical phonological parameters (e.g., handshape, location, and movement).
- While sign language recognition (SLR) deep learning models have achieved higher performance on translation benchmarks, it remains unclear whether these models distinguish abstract phonological features or merely rely on low-level statistical correlations.
- This work evaluates the phonological perception of SLR models trained on American Sign Language (ASL) by exploring phonological sensitivity using minimal pairs and assessing representational consistency with human behavioral data.
- EN 要点:
- arXiv:2606.28667v1 Announce Type: new
Abstract: Sign languages are compositional systems where meaning arises by combining sublexical phonological parameters, such as handshape, location, and moveme…
While deep learning models for Sign Language Recognition (SLR) have achieved increased performance on translation benchmarks, it remains unclear whether these m…
This work evaluates the phonological perception of SLR models trained on American Sign Language (ASL) by probing phonological sensitivity using minimal pairs an…
ArXiv cs.LG (B_intro+search) Link to heading
- Published: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.28406v1 Announce Type: new.
- Abstract: Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design schematics, conceptual frameworks, and graphical abstracts.
- However, existing image-generation benchmarks (e.g., GenEval, T2I-CompBench, DPG-Bench) evaluate natural images and measure compositionality, object counting, or photorealism.
- None of them measure what makes a generated scientific figure usable: correct and legible text labels, faithful depiction of entities and their relations, coherent diagram structure, and adherence to disciplinary drawing conventions.
- EN Key Points:
- arXiv:2606.28406v1 Announce Type: new
- Abstract: Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design sch…
- Yet existing image-generation benchmarks (e.g., GenEval, T2I-CompBench, DPG-Bench) evaluate natural images and measure compositionality, object counting, or pho…
- None of them measure what makes a generated scientific figure usable: correct and legible text labels, faithful depiction of entities and their relations, coher…
On the Necessity of a Liquid Substrate for Mesh Intelligence
- Published: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.28413v1 Announce Type: new.
- Abstract: A mesh of sovereign agents has no center: no shared clock, no shared model, and no coordinator to gather data or retrain.
- Its power rests on each agent folding predictions from its peers–arriving from irregular, unscheduled observations–into a single internal state, online, on a substrate whose weights cannot be retrained.
- Any one of these constraints can be handled alone; optimally folding under all three at once can not.
- EN Key Points:
- arXiv:2606.28413v1 Announce Type: new
- Posted: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.28433v1 Announcement Type: new.
- Abstract: One goal in reinforcement learning (RL) research is to understand general-purpose sequential decision-making, using benchmark simulators as a proxy for agents learning in deployment settings.
- However, when running experiments, the goal of achieving high performance in a simulator can mutate into focusing on solving the simulator problem.
- To achieve high scores, researchers may adopt solutions meant exclusively for solving the simulator, rather than for learning when the agent is deployed outside the simulator.
- EN Highlights:
- arXiv:2606.28433v1 Announce Type: new
- Abstract: One goal in reinforcement learning (RL) research is to understand general-purpose sequential decision-making, using benchmark simulators as a proxy fo…
- When running experiments, however, the goal of achieving high performance in the simulator can mutate into focusing exclusively on solving the simulator
- To achieve high scores, researchers may adopt solutions exclusively meant for solving simulators, rather than learning while the agent is deployed outside a sim…
- Posted: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.28441v1 Announcement Type: new.
- Abstract: Online latent state estimation constitutes a fundamental challenge in the field of artificial intelligence, serving as a foundational tool for various applications, including sequential decision-making, anomaly and change point detection.
- In this paper, a novel online distributed sensing framework is proposed, where agents collaborate and exchange information to perform latent state estimation.
- The proposed estimator combines available partial domain knowledge with the representation power of deep neural networks.
- EN Highlights:
- arXiv:2606.28441v1 Announce Type: new
- Abstract: Online latent state estimation constitutes a fundamental challenge within the artificial intelligence field, serving as a foundational tool for divers…
In this paper, a novel online distributed sensing framework, where agents collaborate and exchange information to perform latent state estimation, is presented
The proposed estimator combines available partial domain knowledge with the representation capabilities of deep neural networks
- Published: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.28444v1 Announce Type: new.
- Abstract: Classical universal approximation theorems establish the expressive power of sigmoidal multilayer perceptrons, but they do not prescribe how initial weights should encode the geometry of the data distribution.
- We propose S-GAI, a spectral geometry-aware initialization framework for one-hidden-layer sigmoidal MLPs.
- Starting from the constructive idea that sigmoid units can act as smooth half-space gates, we move from hand-specified planar geometry to class-wise spectral geometry estimated from image data.
- EN Key Points:
- arXiv:2606.28444v1 Announce Type: new
- Abstract: Classical universal approximation theorems establish the expressive power of sigmoidal multilayer perceptrons, but they do not prescribe how initial w…
- We propose S-GAI, a spectral geometry-aware initialization framework for one-hidden-layer sigmoidal MLPs
- Starting from the constructive idea that sigmoid units can act as smooth half-space gates, we move from hand-specified planar geometry to class-wise spectral ge…
scKDGM: KAN-guided Dynamic Graph Masked Learning for Single-Cell RNA-seq Clustering
- Published: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.28459v1 Announce Type: new.
- Abstract: Single-cell RNA sequencing (scRNA-seq) clustering is essential for identifying cell types, but high dimensionality, sparsity, dropout, and technical noise hinder robust expression representation and cell graph construction.
- Existing masked autoencoders mainly use expression recovery for feature reconstruction, while graph clustering methods usually depend on fixed KNN graphs and do not feed the recovered expression back into graph optimization.
- We propose scKDGM, a KAN-guided dynamic graph masked learning framework for scRNA-seq clustering.
- EN Key Points:
- arXiv:2606.28459v1 Announce Type: new
- Abstract: Single-cell RNA sequencing (scRNA-seq) clustering is essential for identifying cell types, but high dimensionality, sparsity, dropout, and technical n…
- Existing masked autoencoders mainly use expression recovery for feature reconstruction, while graph clustering methods usually depend on fixed KNN graphs and do…
We propose scKDGM, a KAN-guided dynamic graph masked learning framework for scRNA-seq clustering
Counterfactual Residual Data Augmentation for Regression
- Publication Time: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.28460v1 Announce Type: new.
- Abstract: Data-driven modeling in real-world regression tasks is often affected by limited training samples, high collection costs, and noisy observations.
- Inspired by the impact of data augmentation in vision and language, we propose a novel Counterfactual Residual Data Augmentation (CRDA) technique for tabular regression.
- Our main insight is that once a regressor has modeled the systematic components of the data, the remaining noise can be treated as an invariant residual that remains stable under small perturbations of carefully selected features.
- EN Key Points:
- arXiv:2606.28460v1 Announce Type: new
- Abstract: Data-driven modeling in real-world regression tasks often suffers from limited training samples, high collection costs, and noisy observations
- Inspired by the impact of data augmentation in vision and language, we propose a novel Counterfactual Residual Data Augmentation (CRDA) technique for tabular re…
- Our key insight is that once a regressor has modeled the systematic component of the data, the remaining noise can be viewed as an invariant residual that remai…
Singular Learning and Occam’s Razor in Deep Monomial Networks
- Publication Time: 2026-06-30 12:00 Beijing Time
- Abstract: - arXiv:2606.28464v1 Announce Type: new.
- Abstract: In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model’s architecture.
- These critical points occur where the Jacobian of the model’s parametrization is rank-deficient, and are the most pronounced singularities studied in Singular Learning Theory.
- We investigate such points in deep fully-connected networks with monomial activations via tools from polynomial algebra such as Mason’s Theorem.
- EN Key Points:
- arXiv:2606.28464v1 Announce Type: new
- Abstract: In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model’s architecture
- These critical points occur where the Jacobian of the model’s parametrization is rank-deficient, and are the most pronounced singularities studied in Singular L…
- We investigate such points in deep fully-connected networks with monomial activations via tools from polynomial algebra such as Mason’s Theorem
An Agentic AI Pipeline for Appliance-Level Energy Anomaly Detection and LLM-Driven Recommendations
- Publication Time: 2026-06-30 12:00 Beijing Time
Abstract:- arXiv:2606.28467v1 Announce Type: new.
- Abstract: Appliance-level energy monitoring in office buildings produces noisy alerts that non-expert facility managers struggle to use.
- This paper proposes an end-to-end agentic pipeline that combines deep time-series forecasting, variational anomaly detection, and LLM-based reasoning to generate prioritized, actionable maintenance recommendations.
- The system tracks seven office appliances using a hybrid Singular Spectrum Analysis (SSA) and Long Short-Term Memory (LSTM) forecasting model, and applies a per-appliance LSTM Variational Autoencoder (VAE) focusing on marked anomalous daily consumption events.
- EN 要点:
- arXiv:2606.28467v1 Announce Type: new
- Abstract: Appliance-level energy monitoring in office buildings produces noisy alerts that non-expert facility managers struggle to use
- This paper proposes an end-to-end agentic pipeline that combines deep time-series forecasting, variational anomaly detection, and LLM-based reasoning to generat…
- The system tracks seven office appliances using a hybrid Singular Spectrum Analysis (SSA) and Long Short-Term Memory (LSTM) forecasting model, and applies a per…
Modelling Emotional Memory in Children with Tensor Networks
- Release Time: 2026-06-30 12:00 Beijing Time
- Abstract:- arXiv:2606.28470v1 Announce Type: new.
- Abstract: We demonstrate how emotional valence influences the order-dependent structure of children’s recognition memory: correct recall of a sequence of emotionally valenced toys depends not only on the valence of a given toy itself, but also on the valence of the toys shown immediately before and after it.
- Whilst standard psychological models confirm that order-dependence differs across an event (a set of toys shown in sequence), accuracy is low and the model fails to reflect how memory for emotional objects influences other objects in the set.
- A classical tensor network model factoring in valence is able to achieve a 77.98% accuracy in modelling the results of the study.
- EN 要点:
- arXiv:2606.28470v1 Announce Type: new
- Abstract: We demonstrate how emotional valence influences the order-dependent structure of children’s recognition memory: correct recall of a sequence of emotio…
- Whilst standard psychological models confirm that order-dependence differs across an event (a set of toys shown in sequence), accuracy is low and the model does…
- A classical tensor network model factoring in valence is able to achieve a 77.98% accuracy in modelling the results of the study