🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-09-04
- 类型
- ai-daily
- 字数
- 7544
- 阅读时长
- 36 min
2026-09-04 AI Daily | OpenAI Bets $1B on AI Security Defense, Astra Enters Legal & Gaming Workflows Link to heading
While OpenAI is launching its Daybreak program for critical infrastructure to bolster cyber defense and industry adoption, it is also deploying GPT-6 Astra into legal research, financial review, and game prototyping scenarios. The focus today isn’t on model hype, but on AI’s accelerating integration into professional workflows, which more clearly exposes engineering bottlenecks in evaluation, memory, and trusted execution.
📖 This Issue’s Watch List: A Deep Dive Link to heading
The most important trend to follow today is that “AI is moving from general-purpose models to implementable systems.” Google DeepMind’s WeatherNext 3 represents the continued push of foundational models toward high-precision industry forecasting, and it’s worth watching how it turns weather—a problem with strong spatiotemporal dependencies—into a deployable capability. Meanwhile, OpenAI is applying GPT-6 Astra to legal research and game prototyping, showing that agents are entering professional workflows and productivity scenarios. Another set of papers is even more crucial for engineering teams to read: research on evaluation awareness, persistent memory, the data layer for static LLM APIs, and the credibility of evidence from multi-agent systems collectively exposes the most pressing engineering bottlenecks in current agent systems.
🌐 AI Hot Topics on X Link to heading
Topic 1: OpenAI Launches GPT-6 Astra as Most Capable Model Yet Link to heading
- Category: AI · News
- Details: Trending 5 hours ago, 79,000 related posts
- What it is: OpenAI’s pre-release of a new model security update named Astra is trending on X, amid rumors that it may launch soon with capabilities significantly surpassing GPT-5.6.
- Why it matters: If true, this means OpenAI could once again widen the gap in model capability, inference efficiency, and cybersecurity offense-defense capabilities. It would also influence industry predictions about the timing, naming, and release cadence of the next generation of frontier models.
- Discussion Summary: The discussion is focused on whether Astra is in its final release stage, whether its official public name will be GPT-6 or another version, and the credibility of the leaked information versus official announcements. Another point of contention is that the current public evidence points more toward a breakthrough in cybersecurity scenarios rather than a comprehensive leap in general capabilities.
Topic 2: Lululemon Shares Plunge 15% After Weak Earnings and Cut Outlook Link to heading
- Category: AI · News
- Details: Trending 3 hours ago, 3,800 related posts
- What it is: Lululemon’s stock dropped about 15% during trading after the company reported earnings below expectations and lowered its outlook.
- Why it matters: Changes in the performance and guidance of consumer brands like this are often seen as a bellwether for macroeconomic demand, retail data, and market risk appetite, which can also affect AI-related investment valuations for consumer tech and retail automation narratives.
- Discussion Summary: The discussion on X centers on whether the weak performance is due to slowing demand, increased competition, or inventory and discount pressures. Some are also debating whether this is a short-term fluctuation or if the company’s growth logic has significantly cooled.
Topic 3: JD Vance Blames Iran Attacks for High Gas Prices in White House Briefing Link to heading
- Category: AI · Other
- Details: Trending 1 day ago, 37,000 related posts
- What it is: At a White House briefing, JD Vance attributed high gas prices to attacks on Iran, sparking widespread attention and discussion on X.
- Why it matters: Such statements can influence market judgments on energy prices, geopolitical risks, and policy narratives. These variables, in turn, can indirectly affect the AI industry’s computing costs, supply chain expectations, and the broader investment environment.
- Discussion Summary: The debate on X is mainly about whether this attribution is valid, if it’s a political explanation for rising gas prices, and what the actual link is between the situation in Iran and the energy market.
Topic 4: John Ternus Takes Over as Apple’s New CEO from Tim Cook Link to heading
- Category: AI · News
- Details: Trending 2 days ago, 267,000 related posts
- What it is: Reports have emerged that John Ternus is succeeding Tim Cook as the new CEO of Apple.
- Why it matters: Apple is one of the world’s most important technology companies. A change in its leadership will directly impact its AI product roadmap, chip strategy, device ecosystem, and the competitive landscape of the industry.
- Discussion Summary: The conversation on X focuses on whether this signals a more aggressive push by Apple into on-device AI, a potential shift from its current conservative product cadence, and whether the new CEO can sustain Apple’s growth and ecosystem dominance in the AI era.
Topic 5: JD Vance Details Trucking Fraud Crackdown Shutting Down 2,000 CDL Mills Link to heading
- Category: AI · Other
- Details: Trending 1 day ago, 46,000 related posts
- What it is: JD Vance publicly discussed cracking down on Commercial Driver’s License (CDL) fraud and “CDL mills,” stating that enforcement has already shut down approximately 2,000 non-compliant institutions.
- Why it matters: Such regulatory actions impact the data reliability, compliance costs, and talent supply in the logistics and transportation industries. They also indirectly affect the security and identity verification systems that AI applications like autonomous freight and fleet management rely on for implementation.
- Discussion summary: Discussions on X are mainly focused on two points: first, support for strongly combating fraud and improving road safety; second, questioning whether the scale of the shutdowns is excessive and whether this will further exacerbate the truck driver shortage and industry labor pressures.
Topic 6: NBA Hits Clippers with Historic Penalties in Kawhi Leonard Cap Probe Link to heading
- Category: AI · Sports
- Summary: Trending since: 1 day ago, Related posts: 265,000
- What it is: The NBA has imposed record penalties on the Clippers for allegedly circumventing the salary cap during the recruitment and signing of Kawhi Leonard.
- Why it matters: While not an AI event itself, it highlights the emphasis sports leagues place on transparency and governance in complex data, contract reviews, and rule enforcement. This could also affect the data environment relied upon by AI applications for sports analytics and business decision-making.
- Discussion summary: Discussions on X are mainly focused on whether the penalty is proportional to the violation, how much responsibility the Clippers’ management should bear, the consistency of the league’s investigation and enforcement, and whether this ruling will set a precedent for future salary cap circumvention cases.
Topic 7: Liverpool Names 25-Man Champions League Squad, Omits Chiesa and Endo Link to heading
- Category: AI · Other
- Summary: Trending since: 3 hours ago, Related posts: 4,600
- What it is: Liverpool announced its 25-man Champions League squad, with Chiesa and Endo not included.
- Why it matters: This is a football roster news item with no direct connection to AI technology, industry, or research. It primarily shows that trending topic classifications can include cross-domain noise.
- Discussion summary: The discussion focus on X is mainly on the reasons for the two players’ omission, the team’s Champions League squad trade-offs, and their subsequent playing opportunities. Due to a lack of representative tweets, the specific public opinion trend cannot be confirmed at this time.
Topic 8: Sophie Cunningham Enjoys Fresh Ground Chuck from Family Friend’s Ranch Link to heading
- Category: AI · Sports
- Summary: Trending since: 23 hours ago, Related posts: 1,800
- What it is: Sophie Cunningham shared or talked about enjoying fresh ground beef from a family friend’s ranch.
- Why it matters: Based on the given information, this topic has no direct connection to AI technology or industry. Its inclusion in the AI category may reflect a tagging bias or contextual recognition issue on the platform’s trending list.
- Discussion summary: The material does not provide representative tweets, making it impossible to reliably determine the specific focus of discussion on X. Currently, the confirmed content mainly revolves around the player’s life and food sources.
Topic 9: 1968 NYC Subway Photos Spark Debate on Change and Order Link to heading
- Category: AI · Sports
- Summary: Trending since: 2 hours ago, Related posts: 405
- Abstract: 1968 NYC Subway Photos Spark Debate on Change and Order:
Topic 10: Gabriel Martinelli Joins Al-Hilal in Arsenal’s Record €70m Sale Link to heading
- Category: AI · Sports
- Summary: Trending since: 2 days ago, Related posts: 192,000
- Abstract: Gabriel Martinelli Joins Al-Hilal in Arsenal’s Record €70m Sale: 🚨🔵⚪️ OFFICIAL: Gabriel Martinelli joins Al Hilal from Arsenal in a €70m package deal. 🇧🇷 It becomes Arsenal’s record sale, with a sell-on clause also included in the agreement.
Topic 11: Lisa Reveals Hidden Romances and K-pop Struggles in New Documentary Link to heading
- Category: AI · Entertainment
- Summary: Trending since: 20 hours ago, Related posts: 61,000
- Abstract: Lisa Reveals Hidden Romances and K-pop Struggles in New Documentary:
Topic 12: Elon Musk Documentary Sparks Family Defense and Fierce Backlash Link to heading
- Category: AI · Entertainment
- Overview: Trending for: 1 day ago, Related posts: 49,000
- What happened: After a documentary about Elon Musk drew attention, his family came forward to defend him, which also led to significant criticism and backlash on X.
- Why it matters: Musk is also the central figure at the AI company xAI. Such public opinion incidents can influence external judgments of his personal image, business decisions, and AI landscape, while also amplifying discussions around the responsibility and influence of AI leaders.
- Discussion summary: The main debate on X revolves around whether the documentary presented Musk fairly, whether his family’s defense is convincing, and whether his personal controversies will affect public trust in xAI, Tesla, and related AI topics.
Topic 13: Maye Musk Rejects Claim of Elon’s Misery from Anonymous Source Link to heading
- Category: AI · Entertainment
- Overview: Trending for: 6 hours ago, Related posts: 7,800
- What happened: Maye Musk publicly denied claims from an anonymous source that “Elon Musk is miserable/unhappy,” responding to rumors about his personal well-being.
- Why it matters: Musk is the central figure of xAI, Tesla, and X. His personal image, emotional state, and public narrative directly influence external attention on his AI business, management style, and strategic judgment.
- Discussion summary: Discussions on X are mainly focused on the reliability of the source, whether Maye Musk’s denial is sufficient to refute the rumors, and whether Musk’s personal life is being overly scrutinized and affecting the evaluation of his AI endeavors.
Topic 14: Elon Musk Warns Austin Heat Hinders Recruiting Link to heading
- Category: AI · News
- Overview: Trending for: 3 hours ago, Related posts: 1,700
- Summary: Elon Musk Warns Austin Heat Hinders Recruiting:
Topic 15: Healthcare Worker Leaves Voicemail for Joy as Hope Link to heading
- Category: AI · Entertainment
- Overview: Trending for: 3 hours ago, Related posts: 7,000
- What happened: A healthcare worker left a voicemail for someone named Joy, expressing hope, and the message gained attention on X.
- Why it matters: This topic highlights the viral potential of combining AI with emotional entertainment content and raises concerns about synthetic voices, real identities, and content authenticity.
- Discussion summary: The discussion focuses on whether the voicemail was AI-generated or voice-synthesized, the authenticity of the story, and whether such emotionally resonant content is spreading hope or exploiting emotions for engagement.
💡 Influencer Insights Link to heading
No influencer insights today. In-depth content from the Watch List is recommended.
📚 Appendix: Today’s Watch List Update Source List Link to heading
Time window: Last 3 days; covers 22 sources; 36 updates in total
OpenAI Blog (A_full) Link to heading
Daybreak for Frontline Defenders: $1B to protect essential services
- Published: 2026-09-03 21:15 Beijing Time
- Summary: [Translation Needed] - Today OpenAI is introducing Daybreak for Frontline Defenders, a new global initiative to help frontline defenders use frontier AI cyber capabilities to protect essential services in the United States and around the world.
- A $1 billion global commitment to expand subsidized access to Daybreak cyber models and products, training, technical support, and partnerships in the United States and internationally.
- Daybreak for America, bringing together all of OpenAI’s U.S.
work to protect the systems Americans rely on every day—from water and electricity to local government and banking—including a new pilot with the Multi-State Information Sharing and Analysis Center (MS-ISAC).
- More than 35 enterprise products and partner-operated services through the Daybreak Defense Network, bringing Daybreak cyber models into the tools, services, and workflows enterprise defenders already use.
- EN Key Points:
- OpenAI introduces Daybreak for Frontline Defenders
- A $1 billion commitment expands access to frontier cyber AI, training, and support for essential services.
Legora reviewed 41 documents in minutes with GPT-6 Astra
- Published: 2026-09-03 20:00 Beijing Time
- Summary: Legora is an agentic operating system for legal and professional work, used by more than 100,000 professionals across more than 1,800 in-house legal departments and law firms in over 50 markets.
- Its legal engineers work directly with customers to understand how they operate and adapt Legora to their end-to-end workflows, from contract and agreement review to legal research.
- One of the more tedious workflows is financial-statement tie-out: checking every figure in draft accounts against trial balances, a consolidation schedule, and the previous year’s accounts until each item agrees.
- As Legora Legal Engineer Percevale Perks says, the work “can take an entire evening, sometimes days.”.
Processing complex financial context at scale. Link to heading
- EN Key Points:
- Legora used GPT-6 Astra to review 41 documents in minutes, find all four planted errors, and improve performance by nearly 40% in this financial-review workflow…
Playco cut manual fixes 50% prototyping games with GPT-6 Astra
- Published: 2026-09-03 20:00 Beijing Time
- Summary: Using GPT-6 Astra, Playco built three themed game prototypes from one grey box foundation and reported 50% fewer manual fixes than with the previous model.
This piece from OpenAI Blog explains how Playco cut manual fixes 50% prototyping games with GPT-6 Astra shapes the broader AI and infrastructure landscape.
- It also surfaces practical implications for founders, operators, and investors following Playco cut manual fixes 50% prototyping games with GPT-6 Astra.
- Key takeaways:
- Using GPT-6 Astra, Playco built three themed game prototypes from one grey box foundation and reported 50% fewer manual fixes than with the previous model.
- Published: 2026-09-03 08:00 Beijing Time
- Summary: GPT-6 Astra is our most capable broadly deployed model and our first to reach the Critical level of cybersecurity capability under our Preparedness Framework.
- This piece from OpenAI Blog explains how Safety overview: GPT-6 Astra shapes the broader AI and infrastructure landscape.
- It also surfaces practical implications for founders, operators, and investors following Safety overview: GPT-6 Astra.
- Key takeaways:
- GPT-6 Astra is our most capable broadly deployed model and our first to reach the Critical level of cybersecurity capability under our Preparedness Framework.
Google DeepMind Blog (A_full) Link to heading
- Introducing WeatherNext 3, our most advanced and accurate global weather AI model
- Published: 2026-09-03 23:02 Beijing Time
- Summary: Introducing WeatherNext 3, our most advanced and accurate global weather AI model.
- This piece from Google DeepMind Blog explains how Introducing WeatherNext 3, our most advanced and accurate global weather AI model shapes the broader AI and infrastructure landscape.
- It also surfaces practical implications for founders, operators, and investors following Introducing WeatherNext 3, our most advanced and accurate global weather AI model.
- Key takeaways:
- Introducing WeatherNext 3, our most advanced and accurate global weather AI model
Two Minute Papers (B_intro+search) Link to heading
- Claude Fable AI Is Much Stranger Than The Headlines Suggest
- Publish Time: 2026-09-03 16:23 Beijing Time
- Summary: ❤️ Check out Lambda here and sign up for their GPU Cloud:.
- 📝 The Claude Fable 5.1 paper is available here:.
- Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi.
- Claude Fable AI Is Much Stranger Than The Headlines Suggest.
- EN Key Points:
- ❤️ Check out Lambda here and sign up for their GPU Cloud:
- 📝 The Claude Fable 5.1 paper is available here:
- 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
- Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef…
ArXiv cs.AI (B_intro+search) Link to heading
EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models
- Publish Time: 2026-09-03 12:00 Beijing Time
- Summary: arXiv:2609.01611v1 Announce Type: new.
- Abstract: Frontier large language models can often recognize when they are being evaluated, a capability known as evaluation awareness.
- If models behave differently in evaluations than in deployment, this undermines the validity of evaluation results, which are a crucial component of current AI safety frameworks.
- We introduce EvalDetectBench, an open pipeline and benchmark for measuring evaluation awareness that works with any Inspect-compatible evaluation, allowing practitioners to test against current and future benchmarks.
- EN Key Points:
- arXiv:2609.01611v1 Announce Type: new
- Abstract: Frontier large language models can often recognize when they are being evaluated, a capability known as evaluation awareness
If models behave differently in evaluations than in deployment, this undermines the validity of evaluation results, which are a crucial component of current AI…
We introduce EvalDetectBench, an open pipeline and benchmark for measuring evaluation awareness that works with any Inspect-compatible evaluation, allowing prac…
Meta-ethics and AI: exploring the novel meta-ethical questions in the era of AI
- Publication Time: 2026-09-03 12:00 Beijing Time
- Abstract: [TO BE TRANSLATED] - arXiv:2609.01685v1 Announce Type: new.
- Abstract: With the development of artificial intelligence (AI), the landscape of meta-ethics, which has largely centred on human ethics, faces pressures that may significantly reconfigure it.
- In particular, if future AI systems were to exhibit sufficiently integrated capacities for moral reasoning, moral intentionality, and moral reflection, novel meta-ethical questions would arise concerning what I call “AI’s own ethics”, as distinct from ethical principles merely imposed on AI by human designers.
- This paper offers a conditional and methodological framework for identifying the questions that would emerge if such AI systems were to arise.
- EN Key Points:
- arXiv:2609.01685v1 Announce Type: new
- Abstract: With the development of artificial intelligence (AI), the landscape of meta-ethics, which has largely centred on human ethics, faces pressures that ma…
- In particular, if future AI systems were to exhibit sufficiently integrated capacities for moral reasoning, moral intentionality, and moral reflection, novel me…
- This paper offers a conditional and methodological framework for identifying the questions that would emerge if such AI systems were to arise
When Can a Machine Trust a Statute? A Survival Certificate for Machine-Extracted Legal Logic
- Publication Time: 2026-09-03 12:00 Beijing Time
- Abstract: [TO BE TRANSLATED] - arXiv:2609.01741v1 Announce Type: new.
Abstract: Statutes are increasingly parsed by machines before people read them, and the parsers disagree: on Missouri’s statutes, two independently written extractors diverge on numeric-threshold presence at a false-negative rate of 0.43.
- We ask what formal logic survives such noise.
- We build a passive survival certificate for the Duquenne-Guigues implication basis of machine-extracted statutory contexts: per-attribute inter-extractor disagreement is measured, replayed against the basis in 1,000 Monte Carlo trials, and an implication is certified only when a one-sided Wilson 95% lower bound on survival reaches 0.95; every certified implication carries premise spans and a minimal counterexample.
- EN Highlights:
- arXiv:2609.01741v1 Announce Type: new
- Abstract: Statutes are increasingly parsed by machines before people read them, and the parsers disagree: on Missouri’s statutes, two independently written extr…
- We ask what formal logic survives such noise
- We build a passive survival certificate for the Duquenne-Guigues implication basis of machine-extracted statutory contexts: per-attribute inter-extractor disagr…
- Published: 2026-09-03 12:00 Beijing Time
- Abstract: - arXiv:2609.01814v1 Announce Type: new.
- Abstract: Information sharing can improve a pooled estimate while eliminating independent rescue actions.
- This paper separates those effects in exact finite discovery models.
- A centralized action-budget profile shows that equal one-person accuracy can coexist with different portfolio values.
- EN Highlights:
- arXiv:2609.01814v1 Announce Type: new
- Abstract: Information sharing can improve a pooled estimate while eliminating independent rescue actions
- This paper separates those effects in exact finite discovery models
A centralized action-budget profile shows that equal one-person accuracy can coexist with different portfolio values
Induction and Inquiry via Probabilistic Reasoning over Language and Code
- Publish Date: 2026-09-03 12:00 Beijing Time
- Abstract: arXiv:2609.01815v1 Announce Type: new.
- Abstract: How humans grow and maintain abstract knowledge from the sparse, streaming noisy data of experience is a longstanding challenge in cognitive science.
- Any computational account must satisfy at least three desiderata: It must be (1) data-efficient and compute-efficient, (2) capture gradations of uncertainty to support intelligent inquiry and information gathering, and (3) be flexible enough to mentally represent the endless range of concepts people can learn and think about.
- Here we introduce a computational model that captures these three properties, by encoding symbolic knowledge as mental programs that combine natural language with source code, and sequentially inferring mental programs using LLM-guided Bayesian learning algorithms.
- EN Key Points:
- arXiv:2609.01815v1 Announce Type: new
- Abstract: How humans grow and maintain abstract knowledge from the sparse, streaming noisy data of experience is a longstanding challenge in cognitive science
- Any computational account must satisfy at least three desiderata: It must be (1) data-efficient and compute-efficient, (2) capture gradations of uncertainty to…
- Here we introduce a computational model that captures these three properties, by encoding symbolic knowledge as mental programs that combine natural language wi…
Architecting Conversational Data Systems for Stateless LLM APIs: The Hydration Proxy Pattern
- Publish Date: 2026-09-03 12:00 Beijing Time
- Abstract: arXiv:2609.01834v1 Announce Type: new.
- Abstract: As enterprise platforms transition to conversational reasoning interfaces, the stateless nature of LLM APIs creates an architectural gap.
While statelessness enables horizontal scalability for AI providers, it forces client applications to manage the entire burden of conversational state and semantic memory.
- The work identifies the Hydration Proxy Pattern, an architecture that decouples session persistence from the reasoning engine.
- EN Key Points:
- arXiv:2609.01834v1 Announce Type: new
- Abstract: As enterprise platforms transition to conversational reasoning interfaces, the stateless nature of LLM APIs creates an architectural gap
- While statelessness enables horizontal scalability for AI providers, it forces client applications to manage the entire burden of conversational state and seman…
- The work identifies the Hydration Proxy Pattern, an architecture that decouples session persistence from the reasoning engine
- Publication Time: 2026-09-03 12:00 Beijing Time
- Abstract: [To be translated] - arXiv:2609.01849v1 Announce Type: new.
- Abstract: This article presents SSAKG 2.0, an open-source software package for constructing and operating Structural Sequential Associative Knowledge Graphs (SSAKGs).
- An SSAKG represents objects as graph vertices and ordered sequences as structural patterns of graph connections.
- The resulting sparse graph is used as an associative memory in which complete sequences can be reconstructed from a partial, unordered context.
- EN Key Points:
- arXiv:2609.01849v1 Announce Type: new
- Abstract: This article presents SSAKG 2.0, an open-source software package for constructing and operating Structural Sequential Associative Knowledge Graphs (SS…
- An SSAKG represents objects as graph vertices and ordered sequences as structural patterns of graph connections
- The resulting sparse graph is used as an associative memory in which complete sequences can be reconstructed from a partial, unordered context
The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents
- Published: 2026-09-03 12:00 Beijing Time
- Summary: [TO BE TRANSLATED] - arXiv:2609.01852v1 Announce Type: new.
- Abstract: Persistent memory supports personalized agents, but a stale stored fact can override current authoritative evidence without warning.
- We study when this harm begins as model capability changes.
- We evaluate a frozen, closed-set, action-scored benchmark with 2 suites that represent 2 different meanings of “no memory” (a Benefit suite, unsolvable without the stored fact, and a Safety suite, in which an authoritative tool always holds the correct value), on a same-family model-size series (Qwen3 0.6/1.7/4/8B).
- EN Highlights:
- arXiv:2609.01852v1 Announce Type: new
- Abstract: Persistent memory supports personalized agents, but a stale stored fact can override current authoritative evidence without warning
- We study when this harm begins as model capability changes
- We evaluate a frozen, closed-set, action-scored benchmark with 2 suites that represent 2 different meanings of “no memory” (a Benefit suite, unsolvable without…
Belief-Calibrated Optimization: An Explicit World Model for Agentic Optimization
- Published: 2026-09-03 12:00 Beijing Time
- Summary: [TO BE TRANSLATED] - arXiv:2609.01861v1 Announce Type: new.
- Abstract: The performance of an LLM agent depends on the scaffold around a frozen model.
- A common way to improve that scaffold is to use a coding agent as an optimizer: it reads current scores and traces and iteratively edits the source, producing a new candidate each round.
- Each edit is chosen according to a belief about how the environment will respond: what went wrong, and which change should help.
- EN Highlights:
- arXiv:2609.01861v1 Announce Type: new
- Abstract: The performance of an LLM agent depends on the scaffold around a frozen model
A common way to improve that scaffold is to use a coding agent as an optimizer: it reads current scores and traces and iteratively edits the source, producing a…
Each edit is chosen according to a belief about how the environment will respond: what went wrong, and which change should help
Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence
- Publish Date: 2026-09-03 12:00 Beijing Time
- Summary: [TO BE TRANSLATED] - arXiv:2609.01873v1 Announce Type: new.
- Abstract: Multi-agent AI systems improve inference by spawning agents and synthesizing reports.
- But another agent is not another observation: apparently independent reports may descend from the same evidence, and genuinely independent evidence can produce nearly identical reports.
- We formalize this as an epistemic Sybil problem.
- EN Highlights:
- arXiv:2609.01873v1 Announce Type: new
- Abstract: Multi-agent AI systems improve inference by spawning agents and synthesizing reports
- But another agent is not another observation: apparently independent reports may descend from the same evidence, and genuinely independent evidence can produce…
- We formalize this as an epistemic Sybil problem
ArXiv cs.CL (B_intro+search) Link to heading
PRO-Step: Step-level Process Reward Optimization for Retrieval-Augmented Generation
- Publish Date: 2026-09-03 12:00 Beijing Time
- Summary: [TO BE TRANSLATED] - arXiv:2609.01658v1 Announce Type: new.
- Abstract: Retrieval-Augmented Generation enhances Large Language Models by grounding responses in external knowledge, but multi-hop reasoning remains vulnerable to error propagation, where early retrieval failures confound subsequent steps.
- Standard outcome-based optimization only rewards the final answer, leaving intermediate retrieval and reasoning errors undetected.
While existing process-based methods introduce step-level signals, they still score each step against the final answer, rewarding spurious successes where flawed retrieval coincidentally produces the correct answer.
- EN Key Points:
- arXiv:2609.01658v1 Announce Type: new
- Abstract: Retrieval-Augmented Generation enhances Large Language Models by grounding responses in external knowledge, but multi-hop reasoning remains vulnerable…
- Standard outcome-based optimization only rewards the final answer, leaving intermediate retrieval and reasoning errors undetected
- While existing process-based methods introduce step-level signals, they still score each step against the final answer, rewarding spurious successes where flawe…
- EN Key Points:
Learning Evidence Sufficiency Boundaries for Selective Answering in Grounded Multi-Hop QA
- Publication Time: 2026-09-03 12:00 Beijing Time
- Abstract: [To be translated] - arXiv:2609.01687v1 Announce Type: new.
- Abstract: Grounded question answering systems should answer only when the supplied evidence supports the answer.
- In multi-hop QA, this requirement is difficult because partial evidence can make an unsupported answer appear plausible.
- We study selective answering through evidence sufficiency boundaries: for the same question, a model should abstain under unsupported or partially supported context, answer when the context first becomes sufficient, and keep the answer stable when redundant evidence is added.
- EN Key Points:
- arXiv:2609.01687v1 Announce Type: new
- Abstract: Grounded question answering systems should answer only when the supplied evidence supports the answer
- In multi-hop QA, this requirement is difficult because partial evidence can make an unsupported answer appear plausible
- We study selective answering through evidence sufficiency boundaries: for the same question, a model should abstain under unsupported or partially supported con…
- Published: 2026-09-03 12:00 Beijing Time
- Summary: [Translation pending] - arXiv:2609.01737v1 Announce Type: new.
- Abstract: Mobile payment applications in Nepal are graphically mediated and largely inaccessible to visually impaired users.
- This paper presents SpeakPay, a voice-first digital wallet, and documents the central technical contribution: a controlled study of domain adaptation for low-resource financial speech recognition.
- We introduce NepFinSpeech-403, a 403-utterance dataset of Nepali financial voice commands (send, load, and balance operations spanning 237 unique numerals), and fine-tune Whisper large-v2 with LoRA.
- EN Highlights:
- arXiv:2609.01737v1 Announce Type: new
- Abstract: Mobile payment applications in Nepal are graphically mediated and largely inaccessible to visually impaired users
- This paper presents SpeakPay, a voice-first digital wallet, and documents the central technical contribution: a controlled study of domain adaptation for low-re…
- We introduce NepFinSpeech-403, a 403-utterance dataset of Nepali financial voice commands (send, load, and balance operations spanning 237 unique numerals), and…
MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models
- Published: 2026-09-03 12:00 Beijing Time
- Summary: [Translation pending] - arXiv:2609.01772v1 Announce Type: new.
- Abstract: Meme understanding goes beyond recognizing visual content or literal text; it requires implicit cultural knowledge and pragmatic inference that most vision-language models still lack.
- We introduce MemeCULT-1K, a multilingual benchmark of 1,000 South Asian memes in Bengali, English, and Hindi, where each meme is paired with a cultural context note and three human-written explanations, along with a supplementary set of 54 Bengali regional dialect memes.
We evaluate thirteen popular Vision Language Models (VLMs) under two settings: meme-only and context-aware.
- Key Points:
- arXiv:2609.01772v1 Announce Type: new
- Abstract: Meme understanding goes beyond recognizing visual content or literal text; it requires implicit cultural knowledge and pragmatic inference that most v…
- We introduce MemeCULT-1K, a multilingual benchmark of 1,000 South Asian memes in Bengali, English, and Hindi, where each meme is paired with a cultural context…
- We evaluate thirteen popular Vision Language Models (VLMs) under two settings: meme-only and context-aware
- Key Points:
VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages
- Publish Time: 2026-09-03 12:00 Beijing Time
- Abstract: [To be translated] - arXiv:2609.01788v1 Announce Type: new.
- Abstract: Real-world communication often requires pragmatic reasoning: interpreting meanings implied through context and cultural convention rather than stated literally.
- Existing pragmatic evaluation remains largely limited to English and high-resource languages, leaving Indic languages unexplored despite their linguistic and cultural diversity.
- We introduce VakyArth, the first pragmatic benchmark for Indic languages, designed as a diagnostic evaluation covering Hindi, Punjabi, Tamil, and Malayalam.
- Key Points:
- arXiv:2609.01788v1 Announce Type: new
- Abstract: Real-world communication often requires pragmatic reasoning: interpreting meanings implied through context and cultural convention rather than stated…
- Existing pragmatic evaluation remains largely limited to English and high-resource languages, leaving Indic languages unexplored despite their linguistic and cu…
- We introduce VakyArth, the first pragmatic benchmark for Indic languages, designed as a diagnostic evaluation covering Hindi, Punjabi, Tamil, and Malayalam
- Published: 2026-09-03 12:00 Beijing Time
- Abstract: [Translation pending] - arXiv:2609.01794v1 Announce Type: new.
- Abstract: How do learners avoid overgeneralizations such as Tom laughed me without explicit negative evidence?
- Constructionists have posited two proposals that describe indirect negative evidence against overgeneralizations: preemption (which privileges exposure to near-synonymous construction—e.g., she made him laugh) vs.
- entrenchment (all exposures to a verb’s grammatical usages, including cases like He laughed).
- EN Key Points:
- arXiv:2609.01794v1 Announce Type: new
- Abstract: How do learners avoid overgeneralizations such as Tom laughed me without explicit negative evidence
- Constructionists have posited two proposals that describe indirect negative evidence against overgeneralizations: preemption (which privileges exposure to near-…
- entrenchment (all exposures to a verb’s grammatical usages, including cases like He laughed)
How Do Prompt Variations Affect Energy Consumption in On-Device LLMs?
- Published: 2026-09-03 12:00 Beijing Time
- Abstract: [Translation pending] - arXiv:2609.01798v1 Announce Type: new.
- Abstract: Large language models (LLMs) are increasingly deployed on mobile devices, making energy efficiency a key deployment constraint, yet the energy impact of prompt design remains underexplored.
- This paper aims to understand how two prompt properties, cognitive load and phrasing pattern, shape the energy behavior of on-device LLM inference.
- We conduct a broad empirical study covering prompt properties, datasets, models, and devices, with phase-level profiling that separates prefill and decode energy.
- EN Key Points:
- arXiv:2609.01798v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly deployed on mobile devices, making energy efficiency a key deployment constraint, yet the energy impact…
- This paper aims to understand how two prompt properties, cognitive load and phrasing pattern, shape the energy behavior of on-device LLM inference
- We conduct a broad empirical study covering prompt properties, datasets, models, and devices, with phase-level profiling that separates prefill and decode energ…
TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding
- Publication Date: 2026-09-03 12:00 Beijing Time
- Abstract: arXiv:2609.01810v1 Announce Type: new.
- Abstract: Farsi, spoken by more than 120 million people, lacks a comprehensive benchmark for dialogue generation and understanding.
- We introduce TALKFA, a unified benchmark comprising three complementary datasets: (1) WIKI-FADIAL, 4.2K Wikipedia-grounded dialogues for knowledge-grounded generation; (2) DAILYDIALOG-FA, 6.6K dialogues annotated for dialogue acts and emotions; and (3) PLAYDIAL-FA, 2.1K theatrical dialogues with sentiment labels.
- While LLMs assist data construction, every dialogue undergoes multi-stage review and revision by native Farsi speakers, and only the final human-approved dialogues are released.
- Key English Points:
- arXiv:2609.01810v1 Announce Type: new
- Abstract: Farsi, spoken by more than 120 million people, lacks a comprehensive benchmark for dialogue generation and understanding
- We introduce TALKFA, a unified benchmark comprising three complementary datasets: (1) WIKI-FADIAL, 4.2K Wikipedia-grounded dialogues for knowledge-grounded gene…
- While LLMs assist data construction, every dialogue undergoes multi-stage review and revision by native Farsi speakers, and only the final human-approved dialog…
AVERT: Audio-Verified Adjudication for Spoken Dialogue State Tracking
- Publication Date: 2026-09-03 12:00 Beijing Time
Summary: arXiv:2609.01828v1 Announce Type: new.
- Abstract: Spoken dialogue state tracking recovers slot-value pairs from speech, where ASR errors concentrate in entity values and persist across turns, making it both a generation and an editing problem.
- A strong per-turn text editor corrects much of this but, operating on the transcript alone, leaves three recoverable errors: a value predicted inconsistently across turns, an omitted slot, and a value the audio does not support.
- We present AVERT, which scores each candidate value by combining cross-turn agreement with a trained audio-conditioned verifier and resolves the three error types with three operators, vote, add, and swap, each restricted to the slots where its error is common.
- EN Key Points:
- arXiv:2609.01828v1 Announce Type: new
- Abstract: Spoken dialogue state tracking recovers slot-value pairs from speech, where ASR errors concentrate in entity values and persist across turns, making i…
- A strong per-turn text editor corrects much of this but, operating on the transcript alone, leaves three recoverable errors: a value predicted inconsistently ac…
- We present AVERT, which scores each candidate value by combining cross-turn agreement with a trained audio-conditioned verifier and resolves the three error typ…
Interpretable Symptom Vectors for Depression in a Large Language Model
- Published: 2026-09-03 12:00 Beijing Time
- Summary: arXiv:2609.01832v1 Announce Type: new.
- Abstract: Patients with depression present with diverse symptom profiles, yet clinical practice routinely reduces this variation to a single severity score.
- Large language models (LLMs) can potentially capture various symptoms and their severity from patient speech.
- However, how depressive symptoms are represented inside LLMs remains poorly understood, limiting clinical trust.
- EN Key Points:
- arXiv:2609.01832v1 Announce Type: new
Abstract: Patients with depression present with diverse symptom profiles, yet clinical practice routinely reduces this variation to a single severity score
Large language models (LLMs) can potentially capture various symptoms and their severity from patient speech
However, how depressive symptoms are represented inside LLMs remains poorly understood, limiting clinical trust
ArXiv cs.LG (B_intro+search) Link to heading
WMLLM: Self-Evolving Optimization Agents via Predict-Then-Act World Modeling
- Publication Time: 2026-09-03 12:00 Beijing Time
- Abstract: 【Translation Needed】- arXiv:2609.01608v1 Announce Type: new.
- Abstract: Black-box optimization problems remain challenging because of large, weakly structured, and high-dimensional search spaces.
- Existing methods often suffer from poor sample efficiency because they rely on direct candidate generation or trial-and-error refinement.
- A natural way to improve search efficiency is to use world modeling, which can help identify promising optimization directions before costly evaluation.
- EN Key Points:
- arXiv:2609.01608v1 Announce Type: new
- Abstract: Black-box optimization problems remain challenging because of large, weakly structured, and high-dimensional search spaces
- Existing methods often suffer from poor sample efficiency because they rely on direct candidate generation or trial-and-error refinement
- A natural way to improve search efficiency is to use world modeling, which can help identify promising optimization directions before costly evaluation
- Publication Time: 2026-09-03 12:00 Beijing Time
- Abstract: 【Translation Needed】- arXiv:2609.01609v1 Announce Type: new.
Abstract: While diffusion models effectively capture multimodal behavioral priors for autonomous driving, offline reinforcement learning (RL) policies remain susceptible to distribution shift, heavy-tailed risk signals, out-of-distribution (OOD) action generation, and high-dimensional state redundancy.
- To address these challenges, we propose DiDrive, a distribution-guided offline diffusion framework featuring two synergistic components: the Risk-Aware Hierarchical Diffusion (RHDif) architecture and the 3DICE policy optimization paradigm.
- In the state space, RHDif utilizes a low-level risk-gated encoder and a high-level contextual modulator to filter environmental redundancy and focus on safety-critical threats.
- EN Highlights:
- arXiv:2609.01609v1 Announce Type: new
- Abstract: While diffusion models effectively capture multimodal behavioral priors for autonomous driving, offline reinforcement learning (RL) policies remain su…
- To address these challenges, we propose DiDrive, a distribution-guided offline diffusion framework featuring two synergistic components: the Risk-Aware Hierarch…
- In the state space, RHDif utilizes a low-level risk-gated encoder and a high-level contextual modulator to filter environmental redundancy and focus on safety-c…
Prompt-Space Meta-Learning Does Not Transfer Across Users: A Frozen-LLM Negative Result
- Publication Time: 2026-09-03 12:00 Beijing Time
- Abstract: arXiv:2609.01615v1 Announce Type: new.
- Abstract: Personalizing a frozen large language model (LLM) to individual users is often framed as a meta-learning problem in prompt space: each user is a task, and one seeks a shared natural-language adaptation policy that, given a handful of the user’s labeled interactions, configures the frozen model for that user.
The framing is attractive because it is backbone-agnostic and reuses the machinery of prompt optimization, yet the field rarely tests whether the optimized meta-objective encodes transferable cross-user adaptation rather than generic instruction quality.
We study this question with Muse (Meta-learned User-adaptation via Shared Evolution), which evolves a single shared adaptation prompt over a meta-train user population by reflective prompt evolution, freezes it, and applies it zero-shot to held-out users; matched controls isolate learning from confounds of phrasing and selection.
- EN Key Points:
- arXiv:2609.01615v1 Announce Type: new
- Abstract: Personalizing a frozen large language model (LLM) to individual users is often framed as a meta-learning problem in prompt space: each user is a task,…
- The framing is attractive because it is backbone-agnostic and reuses the machinery of prompt optimization, yet the field rarely tests whether the optimized meta…
- We study this question with Muse (Meta-learned User-adaptation via Shared Evolution), which evolves a single shared adaptation prompt over a meta-train user pop…
- EN Key Points:
- Published: 2026-09-03 12:00 Beijing Time
- Abstract: [Translation pending] - arXiv:2609.01647v1 Announce Type: new.
- Abstract: The creation of telescope bibliographies is a crucial part of assessing the scientific impact of observatories and ensuring reproducibility in astronomy.
- This task involves identifying, categorizing, and linking scientific publications that reference or use specific telescopes.
- However, this process remains largely manual and resource intensive.
- EN Key Points:
- arXiv:2609.01647v1 Announce Type: new
- Abstract: The creation of telescope bibliographies is a crucial part of assessing the scientific impact of observatories and ensuring reproducibility in astrono…
This task involves identifying, categorizing, and linking scientific publications that reference or use specific telescopes
However, this process remains largely manual and resource intensive
CliffRank: A Dual-Branch Framework for Activity-Cliff Ranking Prediction
- Publication Time: 2026-09-03 12:00 Beijing Time
- Abstract: [Pending Translation] - arXiv:2609.01673v1 Announce Type: new.
- Abstract: Activity-cliff ranking remains difficult because local structural changes can cause large activity differences, while high-quality data that resolve the underlying mechanisms remain limited.
- To use available activity labels more effectively, we combine absolute-activity regression with ranking-consistency learning.
- CliffRank trains two parallel predictors with mean squared error, a thresholded listwise loss, and Pairwise Preference Consistency (PPC), which aligns relative ordering in the preference-probability space.
- EN Key Points:
- arXiv:2609.01673v1 Announce Type: new
- Abstract: Activity-cliff ranking remains difficult because local structural changes can cause large activity differences, while high-quality data that resolve t…
- To use available activity labels more effectively, we combine absolute-activity regression with ranking-consistency learning
- CliffRank trains two parallel predictors with mean squared error, a thresholded listwise loss, and Pairwise Preference Consistency (PPC), which aligns relative…
Sim2Signal: Sim-to-Real Benchmarks for Traffic Signal Control
- Publication Time: 2026-09-03 12:00 Beijing Time
- Abstract: [Pending Translation] - arXiv:2609.01676v1 Announce Type: new.
- Abstract: Reinforcement learning achieves strong traffic signal control performance in simulation, yet policies trained in simulators often fail once deployed in the real world, a failure known as the Sim-to-Real gap.
When RL is applied to traffic signal control, this gap arises from several sources: sensing, action execution, traffic dynamics, and the control objective.
Their relative impact and the reliability of existing Sim-to-Real mitigation methods remain insufficiently understood, and the field lacks a standard benchmark for systematically measuring the gap and evaluating mitigation methods.
EN Highlights:
- arXiv:2609.01676v1 Announce Type: new
- Abstract: Reinforcement learning achieves strong traffic signal control performance in simulation, yet policies trained in simulators often fail once deployed i…
- When RL is applied to traffic signal control, this gap arises from several sources: sensing, action execution, traffic dynamics, and the control objective
- Their relative impact and the reliability of existing Sim-to-Real mitigation methods remain insufficiently understood, and the field lacks a standard benchmark…
- Published: 2026-09-03 12:00 Beijing Time
- Abstract: [Translation Pending] - arXiv:2609.01679v1 Announce Type: new.
- Abstract: The ability of AI systems to improve their behavior during deployment is becoming increasingly important.
- As inference moves beyond the static execution of a fixed trained model, a growing body of work studies how models can refine their behavior on the fly by exploiting test-time information and additional computation.
- These developments have largely evolved along two directions: methods that modify the model’s state using test-time signals, and methods that improve predictions through extra inference-time resources such as more sampling and tool use.
- EN Highlights:
- arXiv:2609.01679v1 Announce Type: new
- Abstract: The ability of AI systems to improve their behavior during deployment is becoming increasingly important
As inference moves beyond the static execution of a fixed trained model, a growing body of work studies how models can refine their behavior on the fly by explo…
These developments have largely evolved along two directions: methods that modify the model’s state using test-time signals, and methods that improve prediction…
Reinforcement Learning and Rule-Based Peer-to-Peer Pricing in Residential PV-BES Communities
- Publication Time: 2026-09-03 12:00 Beijing Time
- Abstract: [Translation Pending] - arXiv:2609.01680v1 Announce Type: new.
- Abstract: This paper compares rule-based and learning-based pricing mechanisms for peer-to-peer (P2P) electricity trading in residential photovoltaic communities.
- The rule-based benchmarks comprise bill-sharing as an ex post allocation mechanism, the mid-market rate, and supply-demand-ratio pricing.
- The reinforcement-learning (RL) formulation is implemented through a Deep Q-Network and evaluated under multiplier-based and learnable SDR-shaped pricing, with a fixed-parameter SDR variant as a non-learning control.
- EN Key Points:
- arXiv:2609.01680v1 Announce Type: new
- Abstract: This paper compares rule-based and learning-based pricing mechanisms for peer-to-peer (P2P) electricity trading in residential photovoltaic communitie…
- The rule-based benchmarks comprise bill-sharing as an ex post allocation mechanism, the mid-market rate, and supply-demand-ratio pricing
- The reinforcement-learning (RL) formulation is implemented through a Deep Q-Network and evaluated under multiplier-based and learnable SDR-shaped pricing, with…
Median-of-Means as an Extremal Convex Estimator and a Nonconvex Route to the Trimmed Oracle
- Publication Time: 2026-09-03 12:00 Beijing Time
- Abstract: [Translation Pending] - arXiv:2609.01689v1 Announce Type: new.
Abstract: We revisit median-of-means estimation from a deterministic optimization viewpoint and develop a family of block-Lp estimators for robust learning with heavy-tailed and adversarially corrupted data.
In a block contamination model with at least a fraction 1 minus epsilon of good blocks, we first show that every convex block M-estimator has worst-case robustness constant at least 1 divided by 1 minus 2 epsilon.
This matches the classical median-of-means bound and proves that the trimmed-block oracle constant 1 divided by 1 minus epsilon cannot be attained within the convex class.
- Key Points (EN):
- arXiv:2609.01689v1 Announce Type: new
- Abstract: We revisit median-of-means estimation from a deterministic optimization viewpoint and develop a family of block-Lp estimators for robust learning with…
- In a block contamination model with at least a fraction 1 minus epsilon of good blocks, we first show that every convex block M-estimator has worst-case robustn…
- This matches the classical median-of-means bound and proves that the trimmed-block oracle constant 1 divided by 1 minus epsilon cannot be attained within the co…
- Key Points (EN):
Tri-Band Channel Measurement-Enabled Multi-Layer Digital Twin for Terahertz Wireless Data Centers
- Published: 2026-09-03 12:00 Beijing Time
- Abstract: arXiv:2609.01699v1 Announce Type: new.
- Abstract: The rapid growth of AI computing has driven increasing demands for flexible and high-capacity data-center interconnections.
- Owing to its ultra-wide bandwidth and high spatial reuse capability, terahertz (THz) communication has emerged as a promising solution for future wireless data centers, while digital twins (DTs) enable efficient wireless planning and real-time optimization.
In this work, a measurement-driven multi-layer DT framework is proposed for THz wireless data centers, where the physical, channel, evaluation, and manipulation layers are progressively constructed from bottom to top.
- EN Key points:
- arXiv:2609.01699v1 Announce Type: new
- Abstract: The rapid growth of AI computing has driven increasing demands for flexible and high-capacity data-center interconnections
- Owing to its ultra-wide bandwidth and high spatial reuse capability, terahertz (THz) communication has emerged as a promising solution for future wireless data…
- In this work, a measurement-driven multi-layer DT framework is proposed for THz wireless data centers, where the physical, channel, evaluation, and manipulation…
- EN Key points: