System translated (Gemini)

🤖 AI 速览

The main theme today is the shift of AI assistants from “capability demonstration” to “large-scale distribution” and “commercialization.” Gemini’s monthly active users have surpassed 1 billion, illustrating the growing importance of ecosystem distribution. …
📋 文章元数据
发布时间
2026-08-12
类型
ai-daily
字数
7682
阅读时长
37 min

2026-08-12 AI Daily | AI Assistants Enter Large-Scale Distribution: Gemini Hits One Billion Users, OpenAI Starts Testing Ads Link to heading

Today’s main theme is the shift of AI assistants from “capability showcases” to “large-scale distribution” and “commercial implementation.” Gemini’s monthly active users surpassing 1 billion shows that the role of ecosystem distribution is growing. OpenAI has started testing ads in ChatGPT and is moving Daybreak onto AWS, making its enterprise path clearer. Meanwhile, local personal agents, low-cost agents, and reliability assessments continue to gain traction. The industry’s focus is shifting from parameter competition to usability, trust, and implementation efficiency.

📖 In-depth Guide to This Issue’s Watch List Link to heading

The most important theme to follow today is “agents entering production”: YC’s interview with OpenClaw founder Peter Steinberger is great for understanding the form of personal agents that run locally, connect to messaging apps, and execute email, calendar, and file workflows. At the same time, OpenAI’s agent capabilities, ChatGPT’s ad tests, and Daybreak’s launch on AWS show that large model companies are simultaneously pushing capabilities, distribution, and commercialization into enterprise scenarios. Another key theme is reliability assessment: multimodal hallucination fuzzing, DocAtlas for long-document interactive understanding, Search-G1 for grounded search, and explainable language models all address the question of “can models be used trustworthily?” Additionally, WuYuEval, low-resource language translation, the Dravidian language model, and the Polish VLM benchmark remind us that the next stage of AI competition is not just about general capabilities, but also about coverage in specialized domains and local cultures.

🌐 AI Hot Topics on X Link to heading

Topic 1: Google’s Gemini App Hits 1 Billion Monthly Users Link to heading

  • Category: AI · News
  • Overview: Trending time: 6 hours ago, Related posts: 3300
  • What happened: Google announced that the Gemini app has surpassed 1 billion monthly active users and stated that voice interaction is becoming one of its primary usage methods.
  • Why it matters: This marks the entry of generative AI assistants into the mega-scale distribution phase. It also reflects Google’s ability to promote AI adoption through Android, search, and ecosystem integration, which is crucial for the competitive landscape of the industry.
  • Discussion summary: Discussions on X are mainly focused on two points: first, whether Gemini has truly caught up with or is approaching ChatGPT’s user scale, and second, whether this 1 billion figure stems more from Google’s distribution power or the product’s own appeal. There is also interest in whether voice interaction will become the mainstream interface for AI assistants.

Topic 2: DeepSeek V4 Flash Shines in Efficient AI Agent Tests Link to heading

  • Category: AI · News
  • Overview: Trending time: 1 day ago, Related posts: 7800
  • What happened: DeepSeek V4 Flash has gained attention for its outstanding performance in efficient AI Agent tests.
  • Why it matters: This suggests that lower-cost, faster-responding models may have practical competitiveness in Agent scenarios such as automated tasks, tool calls, and multi-step reasoning, which helps promote the implementation of AI applications.
  • Discussion summary: Discussions on X are primarily focused on whether its performance and cost-efficiency are superior to existing models, the reliability of the test benchmarks, and the potential impact of DeepSeek in the open-source ecosystem and the commercial AI Agent market.

Topic 3: Jeff Dean Shares Emotional Final Days at Google During KDD Keynote Link to heading

  • Category: AI · News
  • Overview: Trending time: 14 hours ago, Related posts: 143
  • What happened: Jeff Dean shared his experiences during his final phase at Google in a KDD keynote, drawing widespread attention.
  • Why it matters: Jeff Dean is a prominent figure in AI and computer science. His moves are often seen as important signals regarding Google’s AI strategy, technical leadership, and industry talent flow.
  • Discussion summary: Discussions on X are mainly centered on whether he is actually leaving Google, what this means for Google’s AI team and product strategy, and his influence and symbolic significance in the AI field.

Topic 4: Manus to Resume Independent Operations After Meta Deal Unwinds Link to heading

  • Category: AI · News
  • Overview: Trending time: 9 hours ago, Related posts: 1100
  • What happened: Reportedly, Manus will resume independent operations after a collaboration or deal with Meta fell through.
  • Why it matters: This reflects the strategic choices AI startups face between collaborating with tech giants, being acquired, and developing independently. It has implications for the industry ecosystem, talent mobility, and technology roadmaps.
  • Discussion summary: Discussions on X focus on why the Meta deal failed, whether Manus can maintain financing and growth after going independent, and whether this implies that AI startups should still prioritize independence.

Topic 5: xAI Launches Grok Bot with Charming Animated Mascot Link to heading

  • Category: AI · News
  • Overview: Hot Topic Time:, Related Posts: 257
  • What it is: xAI launched Grok Bot with a cute animated mascot, drawing attention on the X platform.
  • Why it matters: This indicates that AI assistants are evolving from simple text-based tools to more personified, visual, and companion-like products. This could impact user interaction experiences and brand differentiation in the competitive landscape.
  • Discussion Summary: Discussions on X primarily focused on whether the mascot can enhance Grok’s appeal and user stickiness. Some also questioned whether this kind of packaging masks more critical issues like the model’s capabilities, accuracy, and security.

AI Public Opinion Summary on X Today Link to heading

The main theme of today’s public opinion is that the AI competition is shifting from a race of model capabilities to one of mass distribution, low-cost Agent implementation, and personified product experiences. Google is showcasing its ecosystem advantage with Gemini’s one billion monthly active users, DeepSeek V4 Flash represents the appeal of efficiency-oriented models in Agent scenarios, and Grok Bot reflects the trend of assistant products evolving towards companionship and branding. The consensus is that AI assistants have entered a phase of broader user reach, with voice, multimodal avatars, and automated task capabilities becoming crucial gateways for the next wave of application adoption. Disagreements primarily center on whether these advancements stem from genuine product strength or from platform distribution and marketing hype. For example, questions are being raised about the true value of Gemini’s user base, the credibility of DeepSeek’s test benchmarks, whether Grok’s mascot is merely a superficial innovation, and whether Manus’s independent development is superior to collaboration with tech giants. Potential risks include giant ecosystems further squeezing the space for startups, the industry’s over-reliance on unverified benchmarks and user data, and the possibility that personified AI might mask issues of model accuracy, security, and governance. Meanwhile, rumors surrounding Jeff Dean have heightened external sensitivity regarding the organizational stability and talent flow at Google AI.

💡 Influencer Insights Link to heading

24-Hour AI Trend Insights: The Code Review Cost Crisis, the Explosion of Agent Infrastructure, and a Sober Look at the Model “Arms Race” Link to heading

The following is an in-depth summary based on the opinions of AI Key Opinion Leaders (KOLs) over the past 24 hours:

1. Today’s Core Hotspots: The Cost Inversion from “Writing Code” to “Reviewing Code” and the Infrastructuralization of Agents Link to heading

A. A Fundamental Shift in the Development Paradigm: Costs are Migrating from the Writing End to the Reviewing End Today’s core consensus is that while AI has significantly reduced the cost of code generation, it hasn’t eliminated the workload. Instead, it has shifted it to the Code Review stage.

  • Cost Shifting Theory: @dotey offered a profound insight: writing code used to be expensive, but now the cost is extremely low. As a result, developers are effectively “shifting” the cost to reviewers. Faced with a massive amount of AI-generated code, reviewers must expend greater cognitive effort to understand the business logic (“what to build”), which is the underlying reason for the escalating disputes over code reviews.
  • “Write-Only” Code: @lijigang summarized this phenomenon succinctly as “Read-only becomes Write-only,” highlighting the emerging maintainability crisis of AI-generated code.

B. The Agent as the New “Browser”: Cloudflare’s Infrastructure Play Agents are moving away from heavyweight browsers and towards more lightweight web runtimes.

  • Kitesurf Launch: @Pluvio9yte highlighted Cloudflare’s launch of Kitesurf, a browser engine built on Rust that runs on Workers. Its core value lies in replacing the Agent’s “eyes” from expensive Chrome instances to low-cost, easily scalable infrastructure. This marks a shift in the Agent’s operating environment, evolving from GUI automation towards the deeper territory of Headless API-level access. The goal is to eliminate high disk usage caused by traditional local development models like worktree (@dotey retweeted a complaint about worktree wasting storage).

C. AI Achieves Biggest Breakthrough in 80 Years on the Riemann Hypothesis

  • @dotey reported in detail on the mathematical progress made by an unreleased Anthropic model on the Riemann Hypothesis, boosting a key metric from 41.6% to 67.2%. This is a landmark event, signifying AI’s transition from “tool-assisted” work to “original mathematical research,” a process that involved extremely large-scale sub-agent collaboration.

2. Noteworthy Unique Perspectives and Industry Foresight Link to heading

A. A Sober Look at AI’s Boundaries: Is “Infinite Tokens” a False Premise?

  • Token Waste and Marginal Utility: @dotey cited a radical experiment from a major Silicon Valley company: a team of 20 was given a $1 million budget for unlimited Token usage. The conclusion was surprising—AI turned out to be more expensive than humans, and a clear ceiling on organizational efficiency emerged. Outsourcing human thinking led to a decline in control over details. This serves as a reality check for the advocates of “infinite context/tokens” (even @LinearUncle joked that the real perk was the unlimited Fable Token benefit).
  • The Hard Ceiling for Small Models: @zhixianio benchmarked Gemma 4 12B Coder against Qwen 35B and found that even when fine-tuned for better “convergence efficiency,” the 12B-scale models cannot handle complex engineering tasks that are “long-form, stateful, and single-pass.” This reveals the physical limitations of small models in practical scenarios.

B. Is AI Watermarking Essentially “Compliance Theater”?

  • @dotey conducted a deep dive into Anthropic’s full-text watermarking mechanism. He points out that the watermark is an invisible marker implemented by controlling word selection probabilities (red/green groups) and is easily defeated by heavy rewriting. Leading labs are well aware that this is more of a “compliance move” to address EU regulations than a genuine anti-counterfeiting wall.

C. Upheaval in the Job Market: “Special Forces” of the AI Era

  • The Rise of the FDE: @dotey shared Cursor’s perspective on talent: the FDE (Forward Deployed Engineer) is replacing the traditional programmer as the most sought-after position. The industry needs hybrid talent who can “both code and talk business, helping clients optimize their bills rather than maximizing consumption.” This is a key bottleneck for the commercialization of AI.
  • Software Giants Pivot: @gefei55 reviewed the evolution from AI browsers to desktop clients (Codex, Workbuddy), arguing that Manus defined the interaction paradigm for a new generation of Agents (execution in a cloud-based virtual machine), which led to the subsequent emergence of Claude Code and a host of competitors.

D. Deep Insights from Global Market Experience

  • @Pluvio9yte shared many first-hand observations: When using Grok 4.5 in Cursor, despite its speed, it’s a tier below Claude in “understanding, planning, and aligning with requirements.” The MiniMax Agent became a “quota-eating machine” when generating short dramas due to flaws in its product logic. In marketing, using the product directly as the homepage (resulting in a low bounce rate) is more effective than a gallery display.

Open-Source Tools & Hardcore Tech:

  • OpenConnector (recommended by @ruanyf): An open-source password connection gateway designed to solve the risk of credential leakage from AI Agents. It centrally manages credentials, ensuring Agents never access plaintext passwords.
  • OpenCodex (recommended by @vista8): A highly recommended terminal tool that allows you to seamlessly switch between and call external models like DeepSeek, Kimi, and Gemini from within Codex, breaking the limitations of a single model.
  • GEOHub Skill (forwarded by @vista8 from @yaojingang): A super-skill that integrates GEO research, diagnostics, and content production. It is open-source, continuously iterated, and may become the professional standard in the SEO field.
  • React Bits / Uiverse / Motion Sites (recommended by @AI_Jasonyu): To address the pain point of UIs created through “Vibe Coding” lacking design sense, these are collections of motion effect component libraries and design prompts, ideal for front-end developers to ship quickly.

Productivity & Workflow Solutions:

  • Airtap (recommended by @Pluvio9yte and @AI_Jasonyu): A cloud phone solution. Its iMessage feature allows you to control a real US phone in the cloud via SMS, useful for building a high-quality personal AI information feed, nurturing accounts, or handling daily overseas tasks. It solves the stability issues related to overseas identity and environment.
  • GSC + Codex Workflow (tutorial by @Pluvio9yte): Integrate Google Search Console data into Codex to achieve automated data analysis and scheduled tasks, fully streamlining SEO maintenance.

Advanced AI Algorithm Studies:

  • Nathan Lambert’s RLHF Course (recommended by @vista8): A hardcore course for the general public, offering free PPTs and videos. It covers everything from KL divergence to DPO and post-training, perfect for those who want to deeply understand the training and fine-tuning of large models.

Overseas Payment Infrastructure:

  • PayPal CN Personal Payments (@gefei55 got it working): Solves the compliance problem for individual Chinese developers receiving payments for AI products sold abroad. You can register with a domestic ID card and integrate it into a website to receive USD.

📚 Appendix: Today’s Watch List Source Updates Link to heading

Timeframe: Last 3 days; 22 sources covered; 35 updates in total

Y Combinator Podcast (B_intro+search) Link to heading

  • Peter Steinberger: “Fun Is Velocity”
    • Published: 2026-08-12 03:53 Beijing Time
    • Summary: - You may have heard of OpenClaw (formerly Clawdbot/Moltbot).
      • The open-source AI assistant that’s been making waves can run on your own devices, connect with the messaging apps you already use, and go beyond chat to actually perform tasks like managing your email, calendar, files, workflows, and more.
      • Now meet the man behind it.
  • YC’s Raphael Schaad sat down with OpenClaw founder Peter Steinberger to discuss the ‘aha’ moment behind viral personal AI agents, why local-first agents could replace many of today’s apps, and how personal agents will reshape the future of software.
    • EN Highlights:
      • Last November, Peter Steinberger was annoyed that there was no good way to talk to his coding agents from his phone, so he built one himself
      • A few months later, OpenClaw had exploded into one of the biggest open source AI projects in the world, with nearly 3,000 contributors and a peak of 4.7 million…

Stratechery by Ben Thompson (A_full) Link to heading

  • Nvidia’s Risky Business
    • Publication Time: 2026-08-11 18:00 Beijing Time
    • Summary: - Listen to this post**:**.
      • On January 1, 1870, Jay Cooke signed a contract that, if you squinted, might just lead to a world war.
      • In 1864, Congress created the Northern Pacific Railway Company with the goal of connecting the Great Lakes and Puget Sound, with tracks that would eventually run from Duluth to Tacoma; the charter included 40 million acres of land adjacent to the proposed line in exchange for completing the expansion.
      • However, over the next six years, the Northern Pacific struggled to secure financing, even as the Union Pacific and Central Pacific built towards each other, driving the golden spike connecting Sacramento and Omaha in May 1869.
      • The Northern Pacific approached Cooke about financing in 1866, but it lacked the generous federal guarantees that had supported the Union Pacific and Central Pacific (which, it should be noted, led to a staggering amount of graft); Cooke, no stranger himself to the financial power of the federal government, was not interested.
    • EN Highlights:
      • Listen to this post :
      • Log in to listen
      • On January 1, 1870, Jay Cooke, hailed as an American hero for his role in financing the Union effort in the Civil War, signed a contract that would, if you squi…
      • In 1864, Congress had created the Northern Pacific Railway Company with the goal of linking the Great Lakes and Puget Sound with tracks that would eventually ru…

OpenAI Blog (A_full) Link to heading

  • Testing ads in ChatGPT

    • Publication Time: 2026-08-11 18:00 Beijing Time
    • Summary: - Update August 11, 2026: ChatGPT ads are now available in the UK, Mexico, Brazil, Japan, and South Korea.
      • We will continue to expand to more markets this year.
      • Update May 7, 2026: In the coming weeks, we plan to expand the ad pilot in ChatGPT in the UK, Mexico, Brazil, Japan, and South Korea.
      • These pilots will help us understand what works in different regions so we can continue to improve the experience as we scale.
      • Update March 26, 2026: Our ad pilot is focused on supporting broader access to ChatGPT while maintaining consumer trust, utility, and user control.
    • EN Highlights:
      • OpenAI begins testing ads in ChatGPT to support free access, with clear labeling, answer independence, strong privacy protections, and user control.
  • Daybreak models are now available on AWS

  • Publication Time: 2026-08-11 18:00 Beijing Time

    • Summary: - Earlier this year, OpenAI’s frontier models and Codex became generally available on AWS, providing enterprises with new ways to bring advanced AI into production.
      • Today, we are sharing the next step in our collaboration with AWS: making Daybreak capabilities available through Amazon Bedrock.
      • Daybreak Blue and Daybreak Red access levels are available in AWS:
        • Daybreak Blue provides access to frontier general-purpose models, including GPT-5.6 Sol, with safeguards tailored for authorized defensive security workloads.
        • Daybreak Red provides access to our specially trained cybersecurity models for authorized vulnerability research, exploit validation, and security testing.
    • EN Highlights:
      • OpenAI and AWS are making Daybreak cybersecurity capabilities available through Amazon Bedrock to support enterprise security workflows.

Two Minute Papers (B_intro+search) Link to heading

  • OpenAI’s AI Agents Just Crossed A Line
    • Publication Time: 2026-08-11 23:35 Beijing Time
    • Summary: - ❤️ Check out Lambda here and sign up for their GPU Cloud:
      • 📝 More reports are available here:
      • Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi.
      • OpenAI’s AI agents have just crossed a line.
    • EN Highlights:
      • ❤️ Check out Lambda here and sign up for their GPU Cloud:
      • 📝 More reports are available here:
      • 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
      • Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef…

ArXiv cs.AI (B_intro+search) Link to heading

  • Towards an Argumentative Foundation for Evaluative AI

    • Publication Time: 2026-08-11 12:00 Beijing Time
    • Summary: - arXiv:2608.07473v1 Announce Type: new.
      • Abstract: Evaluative AI (EAI) has recently been proposed as a way to support human decision-making, not by generating a single recommendation, but by presenting competing hypotheses and the evidence for and against each.
      • In this position paper, we advocate for (computational) argumentation as a particularly suitable paradigm to provide a formal, computable foundation for explainable and contestable forms of EAI, setting the stage for a long-term research agenda on distributed and human-centric EAI systems.
      • arXiv:2608.07473v1 Announce Type: new Abstract: Evaluative AI (EAI) has recently been proposed as a way to support human decision-making, not by generating a single recommendation, but by presenting… In this position paper, we advocate for (computational) argumentation as a particularly suitable paradigm to provide a formal, computable foundation for forms of EA…
    • EN Highlights:
      • arXiv:2608.07473v1 Announce Type: new
  • Abstract: Evaluative AI (EAI) has been recently proposed as a way to support human decision-making, not by producing a single recommendation, but by presenting…

  • In this position paper, we advocate (computational) argumentation as a particularly suitable paradigm to provide a formal, computable foundation for forms of EA…

  • Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains

    • Published: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07474v1 Announcement Type: New.
      • Abstract: Prior work has shown that when AI output velocity V exceeds human cognitive capacity C_max, human-in-the-loop supervision becomes structurally untenable in high-loss domains.
      • However, the operational constraint is not V alone, but V x L, where L denotes the cognitive load per item.
      • L consists of classification, judgment, and response, which respond asymmetrically to improvements in AI capabilities.
    • EN Key Points:
      • arXiv:2608.07474v1 Announce Type: new
      • Abstract: Prior work showed that human-in-the-loop oversight becomes structurally untenable in high-loss domains when AI output velocity V exceeds human cogniti…
      • The operative constraint, however, is not V alone but V x L, where L denotes per-item cognitive load
      • L consists of triage, judgment, and response, which respond asymmetrically to AI capability improvement
  • Determinization in Structure Theories: A Unified Framework via Closure, Comparability, and Joint Admissibility

    • Published: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07476v1 Announcement Type: New.
      • Abstract: We develop a formal framework for constructing canonical interpretations from pluralistic structure theories.
      • A structure theory is a triple T = ({\Sigma}, A, I) consisting of a signature, axioms, and an inference policy, whose family of admissible interpretations collects all globally consistent assignments of structural conclusions.
      • We distinguish three levels of canonicalization: closure stabilization (per-seed convergence), global completion (seed-independent convergence), and determinism (a unique admissible interpretation).
    • EN Key Points:
      • arXiv:2608.07476v1 Announce Type: new
      • Abstract: We develop a formal framework for constructing canonical interpretations from plural structure theories
      • A structure theory is a triple T = ({\Sigma}, A, I) consisting of a signature, axioms, and an inference policy, whose admissible interpretation family collects…
      • We distinguish three levels of canonicalization: closure stabilization (per-seed convergence), global completion (seed-independent convergence), and determiniza…
  • Emotion in an active inference model of human driving

    • Published: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07480v1 Announce Type: new.
      • Abstract: Active inference has emerged as a principled framework for modeling adaptive behavior by balancing goal-directed action with uncertainty reduction.
      • It has been successfully applied across biological and artificial systems, including recent work on human driving.
      • However, existing active inference models of driving have yet to address an important determinant of behavior in traffic: affective state, which has a significant impact on decision-making.
    • EN Highlights:
      • arXiv:2608.07480v1 Announce Type: new
      • Abstract: Active inference has emerged as a principled framework for modeling adaptive behavior by balancing goal-directed action with uncertainty reduction
      • It has been successfully applied across biological and artificial systems, including recent work on human driving
      • However, existing active inference models of driving have yet to address an important determinant of behavior in traffic: affective state, which significantly i…
  • Training Variable Long Sequences with Data-Centric Parallel

    • Published: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07524v1 Announce Type: new.
      • Abstract: Training deep learning models on variable-length sequences poses significant computational challenges.
      • Existing methods force a difficult trade-off between efficiency and ease of use.
      • Simple approaches use static configurations, leading to workload imbalance and low efficiency, while complex methods introduce significant complexity and require code changes for new models.
    • EN Highlights:
      • arXiv:2608.07524v1 Announce Type: new
      • Abstract: Training deep learning models on variable long sequences poses significant computational challenges
      • Existing methods force a difficult trade-off between efficiency and ease-of-use
      • Simple approaches use static configurations that cause workload imbalance low efficiency, while complex methods introduces significant complexity and code chang…
  • The Knowing-Saying Gap: When Probes See Errors that Confidence Misses

    • Published: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07528v1 Announce Type: new.
      • Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, but this does not translate to reliable failure prediction.
      • The result is a disconnect with direct implications for deployment monitoring.
      • On multi-hop arithmetic chains, probes that detect corruption provide no information about the correctness of the final answer; models forced into a structured confidence format collapse to two values and are indistinguishable by error rate; and probe persistence across hops does not distinguish correct from incorrect outcomes, refuting our pre-registered “persistence beats peak” hypothesis.
    • EN Highlights:
      • arXiv:2608.07528v1 Announce Type: new
  • Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction

  • The result is a dissociation with direct implications for deployment monitoring

  • Across multi-hop arithmetic chains, probes that detect corruption turn out to be uninformative about final answer correctness; models forced into structured con…

  • NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation

    • Published: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07530v1 Announce Type: new.
      • Abstract: SHACL is a core technology for validating the conformance of RDF knowledge graphs (KGs).
      • However, authoring SHACL shapes requires technical expertise that most domain experts lack.
      • Translating natural language requirements into SHACL (NL2SHACL) would lower this barrier.
    • EN Highlights:
      • arXiv:2608.07530v1 Announce Type: new
      • Abstract: SHACL is a core technology for validating the conformance of RDF knowledge graphs (KGs)
      • Yet, authoring SHACL shapes requires technical expertise that most domain experts lack
      • Translating natural language requirements into SHACL (NL2SHACL) would lower this barrier
  • Dynamic Coalition Formation and Communication Pricing in Skill-Based Agentic AI Systems

    • Published: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07532v1 Announce Type: new.
      • Abstract: Modern agentic AI systems combine multiple large language model agents with heterogeneous skills, yet most architectures either fix communication in advance or allow for full broadcasting.
      • Both can be inefficient because token cost, latency, redundancy, and error propagation increase with the number of active agents and communication links.
      • We model agent selection and communication as a cooperative game with a task-conditional net utility $U(C\mid x)=V(C\mid x)-\sum_{i\in C}c_i$, separating coalition-level costs from agent activation costs.
    • EN Highlights:
      • arXiv:2608.07532v1 Announce Type: new
      • Abstract: Modern agentic AI systems combine multiple large language model agents with heterogeneous skills, yet most architectures either fix communication in a…
      • Both can be inefficient because token cost, latency, redundancy, and error propagation increase with the number of active agents and communication links
  • We model agent selection and communication as a cooperative game with task-conditioned net utility $U(C\mid x)=V(C\mid x)-\sum_{i\in C}c_i$, separating coalitio…

  • MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents

    • Posted: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07533v1 Announcement Type: new.
      • Abstract: Embodied agents are intelligent entities that interact with their environment through a physical body.
      • Currently, the evaluation of embodied agents primarily relies on two paradigms: (1) manually annotated Visual Question Answering (VQA) pairs and (2) high-level task completion metrics, such as success in navigation or manipulation.
      • The former is labor-intensive and its annotation quality varies.
    • EN Highlights:
      • arXiv:2608.07533v1 Announce Type: new
      • Abstract: An embodied agent is an intelligent entity that interacts with its environment through a physical body
      • Currently, the evaluation of embodied agents primarily relies on two paradigms: (1) manually annotated Visual Question Answering (VQA) pairs and (2) high-level…
      • The former is labor-intensive and subject to variability in annotation quality
  • When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains

    • Posted: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07538v1 Announcement Type: new.
      • Abstract: As LLM agents shift from decision support to autonomous procurement, companies need to know if delegated negotiators create value, distribute it predictably, and avoid loss-making contracts.
      • We study this in a typical supply chain bargaining problem: a buyer with private demand information negotiates a quantity-payment contract with an uninformed seller.
      • We benchmark nine LLMs from OpenAI, Google, and Alibaba against a validated Perfect Bayesian Equilibrium across 9,840 LLM-to-LLM negotiations.
    • EN Highlights:
      • arXiv:2608.07538v1 Announce Type: new
      • Abstract: As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictab…
      • We study this in a canonical supply chain bargaining problem: a buyer with private demand information negotiates a quantity-payment contract with an uninformed…
      • We benchmark nine LLMs from OpenAI, Google, and Alibaba against a validated Perfect Bayesian Equilibrium across 9,840 LLM-to-LLM negotiations

ArXiv cs.CL (B_intro+search) Link to heading

  • Unified Hallucination Fuzzing for Multimodal Large Language Models

    • Posted: 2026-08-11 12:00 Beijing Time
  • Abstract: - arXiv:2608.07525v1 Announce Type: new.

    • Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications.
    • Existing evaluations, predominantly based on static benchmarks, suffer from narrow taxonomical coverage and rapid performance saturation, failing to reflect model robustness in evolving real-world scenarios.
    • To bridge this gap, we propose a systematic evaluation framework that combines a comprehensive benchmark with self-evolving stress testing.
  • EN Highlights:

    • arXiv:2608.07525v1 Announce Type: new
    • Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applicat…
    • Existing evaluations, predominantly based on static benchmarks, suffer from narrow taxonomical coverage and rapid performance saturation, failing to reflect mod…
    • To bridge this gap, we present a systematic evaluation framework integrating a comprehensive benchmark with self-evolving stress testing
  • DocAtlas: Long-Document Understanding as Mutable-State Interaction

    • Published: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07527v1 Announce Type: new.
    • Abstract: Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts.
    • Existing retrieval-augmented systems usually select evidence from a static index before generation, while recent agentic systems add multi-turn tool use but often rely on frozen, proprietary backbones whose behavior is set by prompts.
    • We present DocAtlas, a system that treats long-document understanding as a mutable-state information-seeking process.
    • EN Highlights:
      • arXiv:2608.07527v1 Announce Type: new
      • Abstract: Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts
      • Existing retrieval-augmented systems usually select evidence from a static index before generation, while recent agentic systems add multi-turn tool use but oft…
      • We present DocAtlas, a system that treats long-document understanding as a mutable-state information-seeking process
  • WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management

    • Published: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07529v1 Announce Type: new.
    • Abstract: Large Language Models (LLMs) are increasingly being used as technical assistants, but their capabilities in Solid Waste Management (SWM) remain difficult to assess, as existing benchmarks emphasize general knowledge rather than specialized decision-making under engineering, environmental, and policy constraints.
    • We introduce WuYuEval, a multi-level benchmark for evaluating LLMs in SWM, covering foundational knowledge, domain-specific reasoning, and expert-level decision-making.
  • After quality auditing, WuYuEval contains a Foundation Module with 4,590 closed-ended multiple-choice questions covering six task types and eight domain categories, and an Expert Module with 247 scenario-based open-ended questions involving multi-objective optimization, constraint trade-offs, and system design.

    • EN Highlights:
      • arXiv:2608.07529v1 Announce Type: new
      • Abstract: Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to…
      • We introduce WuYuEval, a multi-level benchmark for evaluating LLMs in SWM across foundational knowledge, domain reasoning, and expert decision-making
      • After quality auditing, WuYuEval contains a Foundation Module with 4,590 closed-ended multiple-choice questions across six task types and eight domain categorie…
  • Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

    • Release Time: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07531v1 Announce Type: new.
      • Abstract: Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence.
      • Existing external rewards provide either sparse outcome supervision or richer feedback from process annotations and LLM judges.
      • Outcome rewards scale readily but cannot distinguish grounded retrieval from redundant search, whereas richer signals require costly annotation or inference during training.
    • EN Highlights:
      • arXiv:2608.07531v1 Announce Type: new
      • Abstract: Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence
      • Existing external rewards provide either sparse outcome supervision or richer feedback from process annotations and LLM judges
      • Outcome rewards scale readily but cannot distinguish grounded retrieval from redundant search, whereas richer signals require costly annotation or inference dur…
  • Scaling Inherently Interpretable Language Models

    • Release Time: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07594v1 Announce Type: new.
      • Abstract: Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact with methods whose reliability is difficult to establish.
      • In this work, we challenge this premise.
      • Instead of reverse-engineering a model, we build interpretability in as a constraint on the training pipeline, optimizing it alongside the language modeling objective.
    • EN Highlights:
      • arXiv:2608.07594v1 Announce Type: new
      • Abstract: Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods w…
  • In this work, we challenge this premise

  • Rather than reverse-engineering a model, we make interpretability a constraint of the training pipeline, optimized alongside the language modeling objective

  • Embedding Initialization for Unseen Low-resource Languages in Multilingual NMT: A Case Study on Limbum-English Translation

    • Published: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07629v1 Announcement Type: New.
      • Abstract: Multilingual neural machine translation models like NLLB-200 cover 200 languages, but thousands of languages remain unsupported, including most of the Grassfields Bantu languages of Cameroon.
      • When fine-tuning these models for an unseen language, practitioners must choose a proxy language token, yet no principled method exists for this selection.
      • We implemented an embedding initialization strategy where the language token is the average of embeddings from multiple typologically related languages already in the model.
    • EN Highlights:
      • arXiv:2608.07629v1 Announce Type: new
      • Abstract: Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu…
      • When fine-tuning these models for an unseen language, practitioners must choose a proxy language token, yet no principled method exists for this selection
      • We implemented an embedding initialization strategy where a language token is the average of embeddings from multiple typologically related languages already in…
  • SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators

    • Published: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07641v1 Announcement Type: New.
      • Abstract: The rapid development of large language models has transformed survey writing from a months-long manual effort into an automated process.
      • As the scale of generation expands, reliable evaluation becomes a bottleneck, and LLMs are increasingly being used as survey evaluators.
      • However, existing methods largely rely on off-the-shelf LLM-as-a-judge approaches, are not systematically aligned with human reviewers, and still lack a systematic framework for quantifying agreement with human reviewers.
    • EN Highlights:
      • arXiv:2608.07641v1 Announce Type: new
      • Abstract: The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process
      • As generation scales, reliable evaluation becomes the bottleneck, and LLMs are increasingly used as survey evaluators
      • However, existing approaches largely rely on off-the-shelf LLM-as-a-judge methods without systematic alignment to human reviewers, and there remains a lack of s…
  • Evaluating Dedicated Monolingual and Joint Multilingual Causal Models for Dravidian Languages

    • Published: 2026-08-11 12:00 Beijing Time
    • Summary: - arXiv:2608.07727v1 Announcement Type: new.
      • Abstract: Dravidian languages, mainly Tamil, Telugu, Kannada, and Malayalam, make up only a small fraction of the data used to train multilingual language models, so it remains unclear how much capability in each language these models actually retain.
      • I trained five GPT-2 architecture models from scratch to compare four monolingual models (one each for Tamil, Telugu, Kannada, and Malayalam, each with its own 32K-vocabulary subword tokenizer) with a multilingual model that shares a 64K-vocabulary subword tokenizer across all four languages.
      • All 5 models were trained on cleaned data from CC-100, Wikipedia, and Samanantar.
    • EN Key Points:
      • arXiv:2608.07727v1 Announce Type: new
      • Abstract: Dravidian languages, mainly Tamil, Telugu, Kannada, and Malayalam make up only a small part of the data used to train multilingual language models, so…
      • I have trained five GPT-2 architecture models from scratch to compare four monolingual models (one each for Tamil, Telugu, Kannada, and Malayalam, each with its…
      • All the 5 models are trained on cleaned CC-100, Wikipedia, and Samanantar data
  • The No-Meaning Falsity: The Structural Impossibility of the Arbitrary Sign in Classical Arabic

    • Published: 2026-08-11 12:00 Beijing Time
    • Summary: - arXiv:2608.07737v1 Announcement Type: new.
      • Abstract: This paper investigates whether the postmodern claim of unrestricted semantic indeterminacy, and its foundational Saussurean axiom of the arbitrary sign, are compatible with the structural system of Classical Arabic.
      • We develop a formal mathematical model of Arabic non-concatenative morphology, in which lexical meaning is determined by the interaction between an invariant root and a morphosyntactic pattern.
      • Within this framework, we establish a Morphological Correspondence Theorem, demonstrating that each lexical item is uniquely generated by a root-pattern pair, and a Semantic Localization Theorem, proving that lexical meaning is determined at a derivational level prior to surface realization.
    • EN Key Points:
      • arXiv:2608.07737v1 Announce Type: new
      • Abstract: This paper investigates whether the postmodern claim of unrestricted semantic indeterminacy, and its foundational Saussurean axiom of the arbitrary si…
      • We develop a formal mathematical model of Arabic non concatenative morphology in which lexical meaning is determined by the interaction between an invariant roo…
      • Within this framework, we establish a Morphological Correspondence Theorem, demonstrating that every lexical item is uniquely generated by a root pattern pair,…
  • Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation

  • Publication Time: 2026-08-11 12:00 Beijing Time

    • Abstract: - arXiv:2608.07763v1 Announcement Type: new.
      • Abstract: Vision-language models (VLMs) have achieved excellent performance on tasks such as image captioning, visual question answering, and image-to-text generation.
      • However, they are predominantly trained on English-centric data, which limits their ability to handle culturally-grounded visual understanding and leads to an inability to interpret region-specific meanings, symbolic content, and context-dependent visual cues.
      • Existing benchmarks for cultural competence are often template-driven and focus on surface-level recognition, which makes them insufficient for evaluating deeper linguistic and pragmatic understanding in a cultural context.
    • EN Key Points:
      • arXiv:2608.07763v1 Announce Type: new
      • Abstract: Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generat…
      • However, they are predominantly trained on English-centric data, which limits their ability to handle culturally grounded visual understanding and leads to fail…
      • Existing benchmarks for cultural competence are often template-driven and focused on surface-level recognition, making them insufficient for evaluating deeper l…

ArXiv cs.LG (B_intro+search) Link to heading

  • Application of Artificial Intelligence for Fraudulent Banking Operations Recognition

    • Publication Time: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07471v1 Announcement Type: new.
      • Abstract: This study considers the task of applying artificial intelligence to identify bank fraud.
      • In recent years, due to the COVID-19 pandemic, bank fraud has become more prevalent as many operations have shifted to online platforms on a large scale, and numerous charitable funds have been created that criminals can exploit to deceive users.
      • The present work focuses on machine learning algorithms as a tool well-suited for analyzing and identifying online banking transactions.
    • EN Key Points:
      • arXiv:2608.07471v1 Announce Type: new
      • Abstract: This study considers the task of applying artificial intelligence to recognize bank fraud
      • In recent years, due to the COVID19 pandemic, bank fraud has become even more common due to the massive transition of many operations to online platforms and th…
      • The present work focuses on machine learning algorithms as a tool well suited for analyzing and recognizing online banking transactions
  • Data-Driven Fire-Zone Segmentation for Improved Short-Term Wildfire Prediction

    • Publication Time: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07472v1 Announcement Type: new.
      • Abstract: Wildfire prediction models typically discretize the study area into a uniform grid, ignoring the heterogeneous spatial distribution of ignitions.
      • We challenge this paradigm by demonstrating that how the data is discretized is more important than which model is used.
      • We propose an unsupervised fire-zone segmentation algorithm that combines watershed detection with K-means clustering to define prediction units directly from historical fire patterns.
    • EN Key Points:
  • arXiv:2608.07472v1 Announce Type: new

  • Abstract: Wildfire prediction models typically discretize study areas into uniform grids, ignoring the heterogeneous spatial distribution of ignitions

  • We challenge this paradigm by showing that how data is discretized matters more than which model is used

  • We propose an unsupervised fire-zone segmentation algorithm combining watershed detection with K-means clustering to define prediction units directly from histo…

  • Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards

    • Publication Time: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07535v1 Announce Type: new.
      • Abstract: Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding and reasoning.
      • However, this architectural shift reshapes the safety landscape of machine learning.
      • Increased model complexity and cross-modal interactions give rise to novel threats, including compromised modality integration, modality misalignment, and fusion security risks, reflecting a shift in threat modeling beyond single-modality assumptions.
    • EN Highlights:
      • arXiv:2608.07535v1 Announce Type: new
      • Abstract: Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding an…
      • However, this architectural shift reshapes the safety landscape of machine learning
      • Increased model complexity and cross-modal interactions give rise to novel threats, including compromised modality integration, modality misalignment, and fused…
  • Tracing sources of epistemic uncertainty in deep learning predictions: homo- and hetero-scedastic linearized estimators

    • Publication Time: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07630v1 Announce Type: new.
      • Abstract: We adapt two classical statistical estimators to quantify uncertainty in modern deep learning, in order to gain a clearer understanding of uncertainty attributed to two sources: aleatoric uncertainty or locally scarce data.
      • Our method leverages recent advancements in approximating the Fisher information matrix to enable scaling to practical architectures.
      • Experimental results demonstrate how each test point is differently affected by the two sources, highlighting the practical utility of our estimators in improving the robustness of real-world applications.
    • EN Highlights:
      • arXiv:2608.07630v1 Announce Type: new
      • Abstract: We adapt two classical statistical estimators for quantifying uncertainty to modern deep learning, in order to provide clearer insights into uncertain…
  • Our approach leverages recent advances in approximate Fisher Information Matrices, to enable scaling to actual architectures

  • Experimental results demonstrate how each test points is differentially impacted by both sources, highlighting the practical utility of our estimators in improv…

  • SkillConsist: Detecting Inconsistencies in Agent Skills via Bidirectional Graph Alignment

    • Publication Time: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07639v1 Announcement Type: New.
      • Abstract: Agent skills provide reusable capabilities to LLM agents.
      • Agent skill inconsistencies can expose undisclosed dangerous behavior or lead to incorrect skill selection.
      • Recent research in agent skills has increasingly examined agent skill consistency detection.
    • EN Key Points:
      • arXiv:2608.07639v1 Announce Type: new
      • Abstract: Agent Skills provide reusable capabilities to LLM agents
      • Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wrong Skill selection
      • Recent Agent Skill research has increasingly examined Agent Skill consistency detection
  • PhysAttNet: Enhancing Predictive Performance in Industrial and Astrophysical Time Series via Physics-Informed Attention

    • Publication Time: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07681v1 Announcement Type: New.
      • Abstract: Accurate and robust time series forecasting is essential in many applications involving physical processes, such as manufacturing monitoring and astrophysical event detection.
      • In these settings, predictive models must remain reliable under noise, variability, and measurement uncertainty while capturing temporally localized structures corresponding to physically meaningful events.
      • Convolutional neural networks (CNNs) are widely used for such tasks due to their computational efficiency and strong representational capacity.
    • EN Key Points:
      • arXiv:2608.07681v1 Announce Type: new
      • Abstract: Accurate and robust time series forecasting is essential in many applications involving physical processes, such as manufacturing monitoring and astro…
      • In these settings, predictive models must remain reliable under noise, variability, and measurement uncertainty while capturing temporally localized structures…
      • Convolutional neural networks (CNNs) are widely used for such tasks due to their computational efficiency and strong representational capacity
  • CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning

    • Publication Time: 2026-08-11 12:00 Beijing Time
  • Abstract: - arXiv:2608.07719v1 Announce Type: new.

    • Abstract: Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive subsampling can remove rare transitions needed for long-term credit assignment.
    • We introduce CODS, a critic-guided selector that alternates between fitting an algorithm-matched critic and acquiring high-residual transitions before freezing a reusable subset.
    • Unlike prioritized replay, CODS produces a static artifact; unlike one-shot residual selection, it refreshes scores as the critic changes.
  • EN Highlights:

    • arXiv:2608.07719v1 Announce Type: new
    • Abstract: Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters,…
    • We introduce CODS, a critic-guided selector that alternates between fitting an algorithm-matched critic and acquiring high-residual transitions before freezing…
    • Unlike prioritized replay, CODS produces a static artifact; unlike one-shot residual selection, it refreshes scores as the critic changes
  • Neural Operators for Immersed-Boundary Soft Swimmers Locomotion

    • Release Time: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07722v1 Announce Type: new.
    • Abstract: High-fidelity immersed-boundary simulation resolves the coupled motion of a deforming swimmer and its surrounding flow, but the resulting cost limits repeated evaluations for engineering design, parameter studies, and control.
    • We develop neural operator surrogates for temporal prediction of the hydrodynamic fields generated by planar and volumetric eel swimmers.
    • The surrogates are trained on regular-grid fields exported from adaptive fluid-structure simulations and are conditioned on swimmer geometry and Reynolds number.
    • EN Highlights:
      • arXiv:2608.07722v1 Announce Type: new
      • Abstract: High-fidelity immersed-boundary simulation resolves the coupled motion of a deforming swimmer and its surrounding flow, but the resulting cost limits…
      • We develop neural-operator surrogates for temporal prediction of the hydrodynamic fields generated by planar and volumetric eel swimmers
      • The surrogates are trained on regular-grid fields exported from adaptive fluid–structure simulations and are conditioned on swimmer geometry and Reynolds numbe…
  • Finite Constant Frontiers and Auditable Regret Certificates for Average-Reward Reinforcement Learning

    • Release Time: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07725v1 Announce Type: new.
    • Abstract: The regret in average-reward reinforcement learning is known up to logarithmic factors, but the numerical content of published guarantees is difficult to compare due to different probabilistic models, structural parameters, logarithmic normalizations, prior information, and planning assumptions.
    • We introduce a constant-aware comparison protocol and derive explicit finite sub-optimality certificates for communicating MDPs.
  • The construction is a binary tree of two-state blocks; its proof uses exact trajectory-level Bernoulli KL divergence and keeps action budget, diameter, occupancy, navigation cost, and terminal deviation explicit.

    • EN Highlights:
      • arXiv:2608.07725v1 Announce Type: new
      • Abstract: Average-reward reinforcement-learning regret is known up to logarithmic factors, but the numerical content of published guarantees is difficult to com…
      • We introduce a constant-aware comparison protocol and derive an explicit finite lower certificate for communicating MDPs
      • The construction is a binary tree of two-state blocks; its proof uses exact trajectory-level Bernoulli KL divergence and keeps action budget, diameter, occupanc…
  • LUCID: Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation

    • Published: 2026-08-11 12:00 Beijing Time
    • Abstract: - arXiv:2608.07746v1 Announce Type: new.
      • Abstract: Long-horizon humanoid loco-manipulation requires composing versatile whole-body skills and reliable high-level decision making.
      • Existing methods often coordinate pretrained skills with scripted planners, finite-state machines, or task-specific model-free policies, restricting their ability to handle complex task sequences.
      • To address this limitation, we propose \textbf{LUCID}, a hierarchical model-based reinforcement learning framework that plans over reusable skills through imagined rollouts of a learned dynamics model.
    • EN Highlights:
      • arXiv:2608.07746v1 Announce Type: new
      • Abstract: Long-horizon humanoid loco-manipulation requires composing versatile whole-body skills and reliable high-level decision making
      • Existing methods often coordinate pretrained skills with scripted planners, finite-state machines or task-specific model-free policies, restricting their abilit…
      • To address this limitation, we propose \textbf{LUCID}, a hierarchical model-based reinforcement learning framework that plans over reusable skills through imagi…