System translated (Gemini)

🤖 AI 速览

OpenAI offers zero data retention for eligible frontier model customers and previews private, secure processing for multi-turn tasks; Replit is lowering the barrier to software creation with GPT-5.6 Luna. The signal today is clear: AI is evolving from merely “answering questions” to …
📋 文章元数据
发布时间
2026-08-20
类型
ai-daily
字数
6448
阅读时长
31 min

2026-08-20 AI Daily | OpenAI Strengthens Data Boundaries, Replit Brings Cheaper, Advanced Intelligence into the Development Workflow Link to heading

OpenAI is offering zero data retention for eligible frontier model customers and previewing private, secure processing for multi-turn tasks, while Replit is lowering the barrier to software creation with GPT-5.6 Luna. The signal today is clear: AI is evolving from “being able to answer” to becoming “controllable, auditable, and production-ready,” as research in clinical trial programming and runtime governance simultaneously raises the standards for implementation.

📖 This Issue’s Watch List: A Deep Dive Link to heading

There are three key threads worth reading today. First, model commercialization and platform boundaries: When you consider OpenAI’s zero data retention promise and Replit embedding cheaper, advanced intelligence into the software creation workflow with GPT-5.6 Luna, alongside DeepSeek’s ongoing challenge to the “closed-source, high-price” narrative, it’s clear that model capabilities are rapidly being commoditized. Second, security and controllability: Safe RAG, MLLM decision-making under uncertainty, and risk identification in multi-turn tasks all remind us that a truly production-ready system is not just a model that can answer questions, but one that knows when to refuse an answer and when to isolate evidence. Third, healthcare and compliant implementation: From clinical trial programming to PHI de-identification and ICD prediction, LLMs are entering high-risk workflows, but the demands for evaluation and auditing are rising in parallel.

🌐 Top AI News on X Link to heading

Topic 1:Study Reveals Trade-Offs in AI Vibe Coding Tools Link to heading

  • Category: AI · News
  • Summary: Trending for: 23 hours ago, Related posts: 268
  • What happened: A study on AI “vibe coding” tools has drawn attention, pointing out that while these tools can rapidly generate applications and lower the barrier to coding, they come with trade-offs in controllability, stability, and consistency.
  • Why it’s important: This is significant for the AI field as it directly concerns whether AI programming tools can transition from “demo-ready” to “production-ready for the long term,” impacting development efficiency, code quality, product security, and enterprise adoption decisions.
  • Discussion summary: Discussions on X are mainly divided: one side believes tools like Cursor, Claude, and Google AI Studio are significantly boosting development speed, even enabling non-engineers to deliver applications quickly; the other side worries this will amplify problems with code review, debugging, maintenance, and reliability, especially in complex projects and production environments.

Topic 2:Stripe Acquires OpenRouter for Over $8 Billion in AI Push Link to heading

  • Category: AI · News
  • Summary: Trending for: 14 hours ago, Related posts: 6100
  • What happened: It’s trending on X that Stripe has acquired OpenRouter for over $8 billion, a move seen by outsiders as a significant step to bolster its AI business.
  • Why it’s important: This is important because it could combine payment infrastructure with an AI model routing platform, affecting how developers access models, control costs, and the overall ecosystem distribution landscape.
  • Discussion summary: The discussion focuses on the authenticity of the news, whether the $8 billion valuation is reasonable, OpenRouter’s strategic value, and whether the acquisition will weaken its platform neutrality and impact model providers and developers.

Topic 3:TrueFoundry Open Sources TrueForge for Cost-Effective AI Agents Link to heading

  • Category: AI · News
  • Summary: Trending for: , Related posts: 428
  • What happened: TrueFoundry announced it is open-sourcing TrueForge, a tool designed to help developers build, deploy, and manage AI Agents more cost-effectively.
  • Why it’s important: As AI Agent applications grow, inference costs, infrastructure complexity, and observability are becoming bottlenecks for implementation. Open-source tools help lower the barrier to entry and promote enterprise-grade Agent engineering.
  • Discussion summary: Discussions on X are centered on whether TrueForge can significantly reduce Agent operational costs, its compatibility with existing frameworks and cloud platforms, and whether its open-source strategy can attract a developer ecosystem. Some also question whether its actual effectiveness still needs to be validated in real production environments.

Topic 4:Etched Raises $700 Million at $21 Billion Valuation from Jane Street Link to heading

  • Category: AI · News
  • Summary: Trending for: 1 day ago, Related posts: 7200
  • What happened: AI inference chip startup Etched has raised $700 million at a $21 billion valuation, in a round led by Jane Street.
  • Why it’s important: This reflects the market’s continued enthusiasm for specialized AI inference hardware and shows that alternative paths to Nvidia are gaining validation from both capital and actual customers, potentially impacting the AI infrastructure landscape.
  • Discussion summary: The main topics of discussion on X are whether Etched’s valuation is overheated, whether Jane Street’s testing and purchasing of complete systems signifies commercial validation, and whether specialized inference chips can truly challenge existing chip giants.

Topic 5: Tesla Owners Share Full Self-Driving Real-World Wins Amid Cybercab Buzz Link to heading

  • Category: AI · News
  • Overview: Trending Time:, Related Posts: 303
  • What it is: Tesla owners are actively sharing their real-world experiences with FSD V14.2 on X, while the mass production and expansion plans for Cybercab and Robotaxi are drawing attention.
  • Why it matters: This indicates that end-to-end autonomous driving systems are transitioning from demonstrations to larger-scale real-world testing and commercialization expectations. This is crucial for autonomous driving safety validation, in-car AI computing power, Robotaxi business models, and Tesla’s competitive position in embodied intelligence and mobility services.
  • Discussion Overview: Discussions focus on whether FSD V14.2 has significantly improved its ability to handle complex road conditions, whether owner tests can represent general safety, the credibility of the Cybercab timeline, and whether European regulatory restrictions will slow down global deployment. Supporters highlight real-world cases and the speed of software iteration, while critics are concerned about regulation, liability attribution, and the risks of implementing unsupervised autonomous driving.

Summary of AI Public Opinion on X Today Link to heading

The main narrative on X today can be summarized as: AI is accelerating from “can demo” to “can deliver.” Discussions are centered on commercialization paths such as programming tools, Agent infrastructure, inference chips, and autonomous driving, with an overall enthusiastic sentiment. The broad consensus is that vibe coding, open-source Agent tools, and FSD real-world tests are all proving that AI’s efficiency and productization speed are indeed increasing, and capital continues to bet on underlying infrastructure and new entry points. The point of contention is whether these technologies can reliably enter production environments. Supporters emphasize the lower barrier to entry, increased speed, and real user feedback, while critics worry about code controllability, platform neutrality, potentially overheated valuations, and whether the optimism around autonomous driving and specialized chips is only temporary. The potential risks mainly stem from a disconnect between technological maturity and the business narrative, which could become particularly evident in issues related to code quality, system stability, supply chain/platform lock-in, regulatory compliance, and liability.

💡 Influencer Insights Link to heading

No influencer insights for today. We recommend reading the in-depth content from the Watch List.

📚 Appendix: Today’s Watch List Source Updates Link to heading

Time Window: Last 3 days; Covering 22 sources; 34 updates in total

Stratechery by Ben Thompson (A_full) Link to heading

  • Apple Settles With E.U., U.S. App Store Fees, ATT Rules in Germany
    • Published: 2026-08-19 18:00 Beijing Time
    • Summary: - Apple’s App Store is finally facing the reality of lower fees, and the EU should be satisfied with its work; it’s ok it’s late.
      • $15/month* or *$150/year.
      • Substantive analysis of the day’s news via three weekly emails or a podcast.
      • Strategy Interviews.
      • Interviews with leading public company CEOs, private company founders, and discussions with fellow analysts.
    • EN Highlights:
      • Apple’s App Store is finally facing the reality of lower fees, and the EU should be satisfied with its work; it’s ok it’s late.

OpenAI Blog (A_full) Link to heading

  • Offering Zero Data Retention for frontier models
    • Published: 2026-08-20 03:00 Beijing Time
    • Summary: - Zero Data Retention provides a clear commitment to eligible API customers: OpenAI will not retain their prompts or model responses after processing a request.
      • OpenAI personnel cannot view customer content1, and enterprise customer data is not used to train our models unless a customer explicitly opts in.
      • As models take on longer and more complex tasks, some serious risks may only become apparent over multiple interactions.
      • Existing ZDR-compatible safety systems evaluate each interaction individually.
      • Today, we are previewing Private Safety Processing, which aims to identify patterns across related interactions without giving OpenAI personnel access to the underlying content.
    • EN Highlights:
      • OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compromising data privacy.
  • Replit expands access to software creation with GPT-5.6 Luna
    • Published time: 2026-08-19 15:00 Beijing Time
    • Summary:
      • As models become more capable, their economics are rapidly shifting.
      • Better price-performance makes advanced intelligence practical in more products, more workflows, and more moments.
      • For software creation, this shift is narrowing the distance between having an idea and building something viable.
      • Replit was an early user of GPT-3, building with OpenAI models as natural language software development began to take shape.
      • Now, GPT-5.6 Luna is powering Replit Free Mode, demonstrating how the GPT-5.6 series’ price-performance and recent OpenAI price drops can translate to broader access at scale.
    • EN highlights:
      • Replit introduces Free Mode, powered by GPT-5.6 Luna, so anyone can turn ideas into working software without worrying about token costs.

Two Minute Papers (B_intro+search) Link to heading

  • DeepSeek Just Made Closed AI Look Ridiculous
    • Published time: 2026-08-20 02:02 Beijing Time
    • Summary:
      • ❤️ Check out Lambda here and sign up for their GPU Cloud:.
      • Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi.
      • DeepSeek just made closed AI look ridiculous.
    • EN highlights:
      • ❤️ Check out Lambda here and sign up for their GPU Cloud:
      • DeepSeek V4 Pro 0813:
      • DSpark full episode:
      • 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:

ArXiv cs.AI (B_intro+search) Link to heading

  • GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents

    • Published time: 2026-08-19 12:00 Beijing Time
    • Summary:
      • arXiv:2608.16890v1 Announce Type: new.
      • Abstract: Clinical trial programming—translating study protocols into analysis-ready datasets according to CDISC standards—is a bottleneck for regulatory submissions, yet LLM-based code generation catastrophically fails at this task: none of 11 single attempts using 5 cutting-edge models produced a valid subject-level analysis dataset.
      • We introduce GxP-Agent, a multi-agent system that encodes regulatory process orchestration as a Directed Acyclic Graph (DAG), decomposing holistic dataset generation into 15 domain-specific nodes executed by worker agents with pharmaverse skill context, validation gates, and conditional retries.
      • On CDISC-Bench, a new execution-based benchmark constructed from FDA pilot submission document CDISSCPilot01 (254 subjects, 49 true ADSL variables), GxP-Agent with Claude Sonnet 4.6 achieved 100% structural match (49/49 variables, 254 correct records) across three independent runs, compared to 59.2% for the best retrieval-augmented baseline and 0% for all baseline single-agent and flat multi-agent approaches.
    • EN highlights:
      • arXiv:2608.16890v1 Announce Type: new
  • Abstract: Clinical trial programming – transforming study protocols into analysis-ready datasets under CDISC standards – is a bottleneck in regulatory submiss…

  • We introduce GxP-Agent, a multi-agent system that encodes regulatory process ordering as a directed acyclic graph (DAG), decomposing monolithic dataset generati…

  • On CDISC-Bench, a new execution-based benchmark built from the FDA pilot submission CDISCPilot01 (254 subjects, 49 ground-truth ADSL variables), GxP-Agent with…

  • Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution

    • Published: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.16891v1 Announce Type: new.
      • Abstract: Agentic AI systems request tool actions that can modify files, send messages, launch jobs, or change workflow state.
      • This shifts the safety problem from harmful text generation to harmful operational side effects.
      • Prompt-level governance can shape model behavior, but it does not create an execution boundary.
    • EN Highlights:
      • arXiv:2608.16891v1 Announce Type: new
      • Abstract: Agentic AI systems request tool actions that can modify files, send messages, launch jobs, or change workflow state
      • This shifts the safety problem from harmful text generation to harmful operational side effects
      • Prompt-level governance can shape model behavior, but it does not create an execution boundary
  • The Price of Thinking: Reasoning Effort as a Model-Specific API Contract

    • Published: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.16956v1 Announce Type: new.
      • Abstract: API buyers purchase a dated contract, not just a model name: the contract includes the requested and served model, reasoning-effort terms or their omission, output tracks, service offerings, prompts, and a price list.
      • We study the reasoning-effort term through a registered paired contrast of Sonnet 5 with explicit high effort against the same model with effort omitted, using 30 AIME 2026 problems and 5 calls per problem.
      • Each paid attempt is assigned a frozen terminal category and reasoning resamples the item while preserving its duplicate calls.
    • EN Highlights:
      • arXiv:2608.16956v1 Announce Type: new
      • Abstract: API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omiss…
      • We study the reasoning-effort term through a registered paired contrast of Sonnet 5 with explicit high effort against the same model with effort omitted, using…
  • Every paid attempt was assigned one frozen terminal category, and inference resampled items while retaining their repeated calls

  • FedPref: Federated Preference Learning for Structured Radiology Report Extraction

    • Publication Time: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.16971v1 Announce Type: new.
      • Abstract: Radiology reports describe findings and locations in free text, but downstream search and analysis require these relations in a fixed schema.
      • Learning this extraction requires labels that are unevenly distributed across institutions: smaller hospitals have less local evidence, and pooling data may be infeasible.
      • We introduce FedPref: a frozen public language model proposes alternative JSON extractions, local annotations rank them, and sites collaboratively train a compact Qwen3-8B adapter while only sharing model updates.
    • EN Highlights:
      • arXiv:2608.16971v1 Announce Type: new
      • Abstract: Radiology reports describe findings and locations in free text, but downstream search and analysis require these relations in a fixed schema
      • Learning this extraction requires labels that are unevenly distributed across institutions: smaller hospitals have less local evidence, and pooling data may be…
      • We introduce FedPref: frozen public language models propose alternative JSON extractions, local annotations rank them, and sites collaboratively train compact Q…
  • The Problem Is the Problem: Towards Scalable Mathematical Discovery

    • Publication Time: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.16977v1 Announce Type: new.
      • Abstract: AI systems are increasingly capable of contributing to mathematical research.
      • In research practice, frontier model inference resources are limited, and expert mathematical review is even more severely constrained.
      • Therefore, the proper allocation of these scarce resources is crucial for improving the efficiency of AI-assisted mathematical discovery.
    • EN Highlights:
      • arXiv:2608.16977v1 Announce Type: new
      • Abstract: AI systems are increasingly capable of contributing to mathematical research
      • In research practice, frontier-model reasoning is a limited resource, and expert mathematical review is even more sharply constrained
      • Allocating these scarce resources well is therefore central to making AI-assisted mathematical discovery efficient
  • SkillEffect: Checked Lowering for Memory-Bounded Agent Tools

    • Publication Time: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.17007v1 Announce Type: new.
      • Abstract: Agent skills can specify procedures and resource obligations for tool use, which language models then instantiate into concrete programs.
  • However, when models translate this guidance into code for existing tool interfaces, even a semantically correct program may load an entire input and exceed the memory available for a single tool call.

  • We present SkillEffect, a checked-lowering runtime for computations, featuring a recoverable source relation, an audited bounded implementation, and registered output post-conditions.

    • EN Highlights:
      • arXiv:2608.17007v1 Announce Type: new
      • Abstract: Agent Skills can specify procedural and resource obligations for tool use, and language models instantiate them as concrete programs
      • However, when models turn this guidance into code for existing tool interfaces, even a semantically correct program may load an entire input and exceed the memo…
      • We present SkillEffect, a checked-lowering runtime for computations with a recoverable source relation, an audited bounded implementation, and a registered outp…
  • Memory Is Communication: The Frontier Between Remembering and Signaling

    • Published: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.17053v1 Announce Type: new.
      • Abstract: A bounded agent may obtain information for a decision from its own past, from peers, or from both sources.
      • Retaining task-relevant history can reduce later communication, while a peer message can supply what memory lacks.
      • Under limits on both resources, how should an agent allocate its information budget?
    • EN Highlights:
      • arXiv:2608.17053v1 Announce Type: new
      • Abstract: A bounded agent may obtain information for a decision from its own past, from peers, or from both sources
      • Retaining task-relevant history can reduce later communication, while a peer message can supply what memory lacks
      • Under limits on both resources, how should an agent allocate its information budget
  • DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

    • Published: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.17067v1 Announce Type: new.
      • Abstract: With the advancements in text-to-image generation models, they have raised serious security concerns, particularly the generation of Not Safe For Work (NSFW) content such as violence and nudity, a problem further exacerbated by red-teaming adversarial attacks.
      • Existing defenses primarily operate under a white-box assumption, relying on text encoder optimization, weight editing, or inference-time interventions, and are fundamentally not scalable to proprietary models.
      • Black-box alternatives based on LLM prompt rewriting offer broader applicability but fail in a key mechanism we identify as the \textit{benign adversarial} problem: where prompts are linguistically safe yet still trigger harmful generations due to the data distributions learned by the model.
    • EN Highlights:
      • arXiv:2608.17067v1 Announce Type: new
  • Abstract: As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such…

  • Existing defenses predominantly operate under white-box assumptions, relying on text encoder optimization, weight editing, or inference-time intervention, and f…

  • Black-box alternatives based on LLM prompt rewriting offer broader applicability, yet fail in a critical regime we identify as the \textit{benign adversarial} p…

  • KernelArc: A Multi-Agent Framework for GPU Kernel Optimization

    • Publication Time: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.17071v1 Announcement Type: New.
      • Abstract: We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads.
      • Strategy-specialized agents run in parallel and coordinate through conclusions-only shared memory, a deterministic benchmark guard, and a read-only cross-agent state with smooth-trigger drafting.
      • We evaluate \kernelarc{} on NVIDIA H100 and B200 GPUs using category-representative SOL-ExecBench workloads.
    • EN Key Points:
      • arXiv:2608.17071v1 Announce Type: new
      • Abstract: We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads
      • Strategy-specialized agents run in parallel and coordinate through conclusions-only shared memory, a deterministic benchmark guard, and read-only cross-agent st…
      • We evaluate \kernelarc{} on NVIDIA H100 and B200 GPUs using category-representative SOL-ExecBench workloads
  • A decodability criterion predicts when hidden-state selection beats majority voting in large language models

    • Publication Time: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.17124v1 Announcement Type: New.
      • Abstract: Combining the answers a large language model (LLM) samples for a question into one decision is a test-time information fusion problem, usually solved by majority voting.
      • For difficult problems, voting is unreliable when sampled answers have correlated errors, so an incorrect answer may win, and drawing more samples can make the decision worse.
      • Selecting candidates by reading a correctness signal from the model’s hidden states is a promising alternative, but its accuracy varies by model and task, and there is no measure to indicate when it can be trusted.
    • EN Key Points:
      • arXiv:2608.17124v1 Announce Type: new
      • Abstract: Combining the answers a large language model (LLM) samples for a question into one decision is a test-time information fusion problem, usually solved…
  • Voting is unreliable on difficult questions, where the sampled answers share correlated errors, so the wrong answer can win and drawing more samples makes the d…

  • Selecting a candidate by reading a correctness signal from the model’s hidden states is a promising alternative, but its accuracy varies across models and tasks…

ArXiv cs.CL (B_intro+search) Link to heading

  • Margin-Regularized Structured Semantic Alignment for Brain-Language Correspondence

    • Publication Time: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.16975v1 Announcement Type: New.
      • Abstract: With the rapid advancement of large language models, brain-language decoding has achieved remarkable progress.
      • However, it remains unclear whether decoded content genuinely reflects neural representations or is largely reconstructed by the language model itself.
      • This ambiguity limits interpretability and hinders the investigation of intrinsic brain-language correspondence.
    • EN Highlights:
      • arXiv:2608.16975v1 Announce Type: new
      • Abstract: With the rapid advancement of large language models, brain-language decoding has achieved remarkable progress
      • However, it remains unclear whether decoded content genuinely reflects neural representations or is largely reconstructed by the language model itself
      • This ambiguity limits interpretability and hinders the investigation of intrinsic brain-language correspondence
  • Cross-Model Memory Transfer via Target-Side Reader Adaptation

    • Publication Time: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.17050v1 Announcement Type: New.
      • Abstract: Methods for improving knowledge use in large language models typically fall into two regimes.
      • Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backbone.
      • Parametric adaptation is efficient at inference time, but entangles knowledge with model weights and can be hard to update, audit, or transfer.
    • EN Highlights:
      • arXiv:2608.17050v1 Announce Type: new
      • Abstract: Methods for improving knowledge use in large language models typically fall into two regimes
      • Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backb…
      • Parametric adaptation is efficient at inference time, but entangles knowledge with model weights and can be hard to update, audit, or transfer
  • Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss

    • Publication Time: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.17051v1 Announcement Type: New.
      • Abstract: The secondary use of electronic health records requires de-identification, yet existing systems miss \emph{institution-specific} protected health information (PHI), such as hospital abbreviations, building names, and internal codes whose status is locally determined.
      • We ask whether large language models (LLMs) with in-context learning (ICL) can close this gap and control the precision-recall trade-off.
      • On 100 annotated pediatric oncology notes (5,322 PHI spans) from Texas Children’s Hospital, we benchmarked eight LLMs against two specialized systems (Stanford TiDE, OpenMed PII) and two pattern-based baselines.
    • EN Highlights:
      • arXiv:2608.17051v1 Announce Type: new
      • Abstract: Secondary use of electronic health records requires de-identification, yet existing systems miss \emph{institutionally situated} protected health info…
      • We ask whether large language models (LLMs) with in-context learning (ICL) can close this gap and control the precision–recall trade-off
      • On 100 annotated pediatric oncology notes (5,322 PHI spans) from Texas Children’s Hospital, we benchmarked eight LLMs against two purpose-built systems (Stanfor…
  • Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting

    • Publication Time: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.17075v1 Announcement Type: New.
      • Abstract: Next-encounter ICD forecasting predicts which standardized diagnosis codes will be documented at a future visit based on the previously available longitudinal record.
      • The task is prospective and multi-label: the target annotation does not yet exist, and multiple codes may be correct.
      • Structured EHR foundation models capture recurrence and temporal progression, while language foundation models generate flexible diagnostic hypotheses.
    • EN Highlights:
      • arXiv:2608.17075v1 Announce Type: new
      • Abstract: Next-encounter ICD forecasting predicts which standardized diagnosis codes will be documented at a future visit from the longitudinal record available…
      • The task is prospective and multi-label: the target note does not yet exist, and several codes may be correct
      • Structured EHR foundation models capture recurrence and temporal progression, whereas language foundation models generate flexible diagnostic hypotheses
  • Uncertainty-Aware Decision Making in Multimodal Large Language Models

    • Publication Time: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.17084v1 Announcement Type: New.
  • Abstract: Multimodal large language models (MLLMs) increasingly answer questions whose correctness depends on visual, textual, temporal, acoustic, document, chart, or concrete evidence.

  • Therefore, their failures are not just linguistic.

  • A fluent answer may conceal poor input quality, perceptual errors, weak grounding, conflicts between modalities, unstable reasoning, distribution shifts, or an inability to answer from the provided evidence.

  • EN Highlights:

    • arXiv:2608.17084v1 Announce Type: new
    • Abstract: Multimodal large language models (MLLMs) increasingly answer questions whose correctness depends on visual, textual, temporal, acoustic, document, cha…
    • Their failures are therefore not only linguistic
    • A fluent answer may conceal poor input quality, a perceptual error, weak grounding, conflict between modalities, unstable reasoning, distribution shift, or a qu…
  • There is No Theoretical Curse of Multilinguality For Embedding Space Structure

    • Release Time: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.17088v1 Announce Type: new.
      • Abstract: A central goal of multilingual NLP is to achieve high monolingual performance per language and cross-lingual alignment through multilingual models to realize large-scale language coverage.
      • The curse of multilinguality describes the phenomenon of performance degradation in multilingual models as we increase language coverage, posing a threat to the aforementioned goal.
      • This paper asks whether multilingual embedding spaces are inherently incapable of achieving perfect multilinguality without a prohibitive increase in required capacity.
    • EN Highlights:
      • arXiv:2608.17088v1 Announce Type: new
      • Abstract: A central goal of multilingual NLP is to achieve high monolingual performance per language and cross-lingual alignment for large-scale language covera…
      • The curse of multilinguality describes the phenomenon of degradation in multilingual model performance as we increase language coverage, posing a threat to the…
      • This paper asks whether multilingual embedding spaces are inherently incapable of achieving perfect multilinguality without a prohibitive increase in required c…
  • A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space: What the Units of Voynichese Are Not

    • Release Time: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.17096v1 Announce Type: new.
      • Abstract: The Voynich manuscript (Beinecke MS 408) is commonly analyzed based on three unstated assumptions: its glyphs are letters, the strings between spaces are words, and every space is a word-space.
      • We test all three against the Zandbergen-Landini transliteration using matched prose, cryptographic, and pseudotext controls, as well as quire-level resampling.
      • None of them hold, and the failures have a common shape: order in Voynichese is located at the edges of the signs and at the graded boundaries between them, not in the continuity of the signs themselves.
    • EN Highlights:
      • arXiv:2608.17096v1 Announce Type: new
  • Abstract: The Voynich manuscript (Beinecke MS 408) is usually analysed on three unstated assumptions: that its glyphs are letters, that the strings between blan…

  • We test all three against the Zandbergen-Landini transliteration with matched prose, cipher, and pseudo-text controls and quire-level resampling

  • None holds, and the failures share a shape: the order in Voynichese sits at the edges of tokens and at graded boundaries between them, not in the succession of…

  • Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models

    • Publication Time: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.17102v1 Announcement Type: New. -Abstract: Modern multimodal foundation models (MFMs) have made rapid progress on tasks requiring integrated perception across speech, vision, and language, including emotion recognition.
      • However, it remains unclear whether they recognize speech and facial emotion through shared affective functional units or modality-specific pathways.
      • We explore emotion-sensitive neurons (ESNs), i.e., sparse decoder neurons selectively associated with emotion categories, in three MFMs: Gemma-4-12B-it, MiniCPM-o-4.5, and Qwen2.5-Omni-7B.
    • EN Key Points:
      • arXiv:2608.17102v1 Announce Type: new
      • Abstract: Modern multimodal foundation models (MFMs) have made rapid progress on tasks requiring integrated perception across speech, vision, and language, incl…
      • However, it remains unclear whether they recognize speech and facial emotion through shared affective functional units or modality-specific pathways
      • We explore emotion-sensitive neurons (ESNs), sparse decoder neurons selectively associated with emotion categories, in three MFMs: Gemma-4-12B-it, MiniCPM-o-4.5…
  • Children, but not language models, show accelerating returns in word learning

    • Publication Time: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.17120v1 Announcement Type: New.
      • Abstract: Children learn hundreds of words over the first years of their lives, a process that begins slowly but quickly accelerates.
      • Previous models have described vocabulary growth as an accumulation of evidence over time.
      • Here, we show that this process is best characterized by accelerating accumulation: children learn more from each additional unit of language experience than from the previous one.
    • EN Key Points:
      • arXiv:2608.17120v1 Announce Type: new
      • Abstract: Children learn hundreds of words over the first years of their lives, in a process that begins slowly but quickly picks up speed
  • Prior models describe vocabulary growth as evidence accumulation over time

    • Here we show that the process is best characterized as accelerating accumulation: children learn more from each additional unit of linguistic experience than th…
  • Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents

    • Published: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.17153v1 Announce Type: new.
      • Abstract: Retrieval-Augmented Generation (RAG) has significantly enhanced the performance of large language models (LLMs), yet these systems remain vulnerable to knowledge poisoning attacks, where misinformation in retrieved documents can compromise the model’s final output.
      • Notably, an LLM may correctly detect that a document contains incorrect information but still be influenced by it.
      • Prior work has addressed this vulnerability through the Cordon Principle, which prevents models responsible for final answer synthesis from directly accessing the original evidence.
    • EN Key Points:
      • arXiv:2608.17153v1 Announce Type: new
      • Abstract: Retrieval-Augmented Generation (RAG) has significantly enhanced the performance of large language models (LLMs), yet these systems remain vulnerable t…
      • Notably, an LLM may correctly detect that a document contains incorrect information while nevertheless being influenced by it
      • Prior work has addressed this vulnerability through the Cordon Principle, which prevents models responsible for final answer synthesis from directly accessing r…

ArXiv cs.LG (B_intro+search) Link to heading

  • Learning Discrete Riemannian Metrics for Physical Fields with Cochain-Frame Equivarianc

    • Published: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.14556v1 Announce Type: new.
      • Abstract: Physical fields on meshes require a separation between topology and geometry: conservation laws are topological and should be exact, while geometry, material response, and anisotropic coupling must be learned from data.
      • Existing neural surrogates often mix these roles in unconstrained message passing.
      • We introduce Riemannian Hodge Message Passing (RHMP), which turns this separation into an architectural principle.
    • EN Key Points:
      • arXiv:2608.14556v1 Announce Type: new
      • Abstract: Physical fields on meshes require a separation between topology and geometry: conservation laws are topological and should be exact, while geometry, m…
      • Existing neural surrogates often mix these roles inside unconstrained message passing
      • We introduce Riemannian Hodge Message Passing (RHMP), which turns this separation into an architectural principle
  • Forward Pass Domain Adaptation (Without Cross-Layer Backpropagation)

    • Publication Time: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.14563v1 Announce Type: new.
      • Abstract: Forward-Pass-Only MLP training (FPO) adapts large language models without a backward pass through the model body, achieving 2.7–3.2x the throughput of standard finetuning with roughly 40% reduction in peak training memory, while preserving out-of-domain benchmarks within the seed noise of a baseline, an attribute not reliably reproduced by full-network finetuning.
      • FPO rests on a single empirical observation: at late layers of a Transformer, the output-layer prediction error approximates the true gradient with cosine similarity 0.47–0.59 across six public models we investigate.
      • We introduce a two-minute diagnostic that quantifies this approximation per layer for any model, identifying where late-layer adaptation is viable.
    • EN Highlights:
      • arXiv:2608.14563v1 Announce Type: new
      • Abstract: Forward-Pass-Only MLP training (FPO) adapts large language models without a backward pass through the model body, achieving 2.7–3.2x the throughput o…
      • FPO rests on a single empirical observation: at late layers of a transformer, the output-layer prediction error approximates the true gradient with cosine simil…
      • We introduce a two-minute diagnostic that quantifies this approximation per layer for any model, identifying where late-layer adaptation is viable
  • Coarse-to-Fine Multi-Resolution Diffusion Models for Trajectory Generation in Urban Systems

    • Publication Time: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.14570v1 Announce Type: new.
      • Abstract: Understanding human mobility is critical for a wide range of urban applications, including traffic management, epidemic control, and urban planning.
      • However, due to privacy concerns, the availability of large-scale public trajectory data remains limited, posing challenges for downstream mobility analysis.
      • Existing methods for synthetic trajectory generation primarily focus on matching global distribution similarity, while often overlooking mobility patterns across different spatial and temporal resolutions, which are crucial for practical applications.
    • EN Highlights:
      • arXiv:2608.14570v1 Announce Type: new
      • Abstract: Understanding human mobility is critical for a wide range of urban applications, including traffic management, epidemic control, and urban planning
      • However, due to privacy concerns, the availability of large-scale public trajectory data remains limited, posing challenges for downstream mobility analysis
      • Existing methods for synthetic trajectory generation primarily focus on matching global distribution similarity, while often overlooking mobility patterns acros…
  • Geometry Is Not Robustness: A Trajectory-Level Study of PGD Evaluation

    • Publication Time: 2026-08-19 12:00 Beijing Time
  • Abstract: - arXiv:2608.14594v1 Announce Type: new.

    • Abstract: Projected Gradient Descent (PGD) is widely used to evaluate adversarial robustness, typically through final adversarial accuracy, but it does not capture the model’s behavior throughout the attack process.
    • Recent work proposes trajectory-level diagnostics, such as loss evolution, gradient alignment, and steps-to-failure, for deeper insight into adversarial optimization dynamics.
    • However, it remains unclear whether these diagnostics reliably indicate robustness strength.
  • EN Highlights:

    • arXiv:2608.14594v1 Announce Type: new
    • Abstract: Projected Gradient Descent (PGD) is widely used to evaluate adversarial robustness, typically via final adversarial accuracy, which does not capture m…
    • Recent work proposes trajectory-level diagnostics, such as loss evolution, gradient alignment, and steps-to-failure, for deeper insight into adversarial optimis…
    • However, whether these diagnostics reliably indicate robustness strength remains unclear
  • DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs

    • Published: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.14614v1 Announce Type: new.
      • Abstract: As AI data centers retire functional GPUs, large quantities of still-capable accelerators enter the secondary market.
      • This paper investigates whether these retired GPUs can find a productive afterlife by forming a DumpsterCluster that can serve modern LLM inference, and under what conditions this repurposing is economically feasible and environmentally sustainable.
      • We physically built a 128-GPU DumpsterCluster from scratch using only second-hand components and ran it for one year.
    • EN Highlights:
      • arXiv:2608.14614v1 Announce Type: new
      • Abstract: As AI datacenters retire functional GPUs, vast quantities of still capable accelerators enter secondary markets
      • This paper investigates whether these retired GPUs can find a productive afterlife to form a DumpsterCluster that can serve modern LLM inference, and under what…
      • We physically built a 128-GPU DumpsterCluster from scratch using only second-hand components and ran it for one year
  • Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion

    • Published: 2026-08-19 12:00 Beijing Time
    • Abstract: - arXiv:2608.14617v1 Announce Type: new.
      • Abstract: A recurring proposal in the field of legal AI is to improve case outcome prediction by fusing uncertainty tools (evidence graphs with belief propagation, sequential Bayesian odds updating, Dempster-Shafer combination, and conformal prediction) into a pipeline.
      • We tested this on 1,000 real European Court of Human Rights cases from LexGLUE and FairLex, predicting whether the court found a violation of the Convention from the factual paragraphs of the cases.
      • We compared three series from two leading LLMs (Claude Opus 4.8 and GPT-5.5) as factual evidence estimators: (A) the raw LLMs, (B) the LLMs routed through the fusion pipeline, and (C) a term-frequency baseline through the same pipeline.
  • EN 要点:

    • arXiv:2608.14617v1 Announce Type: new
    • Abstract: A recurring proposal in legal AI is to improve case-outcome prediction by fusing uncertainty tools (evidence graphs with belief propagation, sequentia…
    • We test this on 1,000 real European Court of Human Rights cases from LexGLUE and FairLex, predicting whether the Court found a Convention violation from the cas…
    • We compare three families across two frontier LLMs (Claude Opus 4.8 and GPT-5.5) as per-fact evidence estimators: (A) the raw LLM, (B) the LLM routed through th…
  • PIKFNO: An Interpretable Neural Operator Based on Physics Informed Kernel Function

    • Publish Time:2026-08-19 12:00 Beijing Time
    • Abstract:- arXiv:2608.14619v1 Announce Type: new.
      • Abstract: This work proposes a new interpretable neural operator framework, termed the Physics Informed Kernel Function Neural Operator (PIKFNO), which explicitly incorporates physics-informed kernel functions derived from governing equations into the neural operator architecture.
      • Unlike traditional neural operators such as DeepONet, which rely on deep networks to implicitly learn basis functions, PIKFNO constrains the trunk network through physics-informed kernel functions, thereby aligning its operator structure with kernel extensions used in mesh-free collocation methods.
      • Two construction strategies are introduced: one learns kernel functions directly from data, where the learned kernel can be regarded as a non-singular fundamental solution, while the other constructs them through transformations of analytical fundamental solutions.
    • EN 要点:
      • arXiv:2608.14619v1 Announce Type: new
      • Abstract: This work proposes a new interpretable neural operator framework, termed the Physics Informed Kernel Function Neural Operator (PIKFNO), which explicit…
      • Unlike traditional neural operators such as DeepONet, which rely on deep networks to implicitly learn basis functions, PIKFNO constrains the trunk network throu…
      • Two construction strategies are introduced: one learns kernel functions directly from data, where the learned kernel can be regarded as a nonsingular fundamenta…
  • Explaining Reinforcement Learning Decisions in Self-adaptive Systems

    • Publish Time:2026-08-19 12:00 Beijing Time
    • Abstract:- arXiv:2608.14620v1 Announce Type: new.
      • Abstract: Reinforcement Learning (RL) has been widely applied in autonomous and self-adaptive systems, but RL policies, especially deep RL policies relying on neural networks, lack transparency and are difficult to understand.
      • This can lead to decreased user trust and make system verification more challenging.
      • To address this challenge, this paper introduces Explanations using Reinforcement Learning Alternative Realities (EARL), a Python library for generating counterfactual explanations in RL settings.
    • EN 要点:
      • arXiv:2608.14620v1 Announce Type: new
  • Abstract: Reinforcement Learning (RL) has been extensively used in autonomous and self-* systems, but RL policies, especially deep RL ones relying on neural net…

  • This can lead to diminished user trust, and makes for a more challenging verification of systems

  • To address this challenge, this paper introduces Explanations using Alternative Realities for Reinforcement Learning (EARL), a Python library to produce counter…

  • Metaplasticity as adaptive gradient preconditioning for incremental learning

    • Publication Time: 2026-08-19 12:00 Beijing Time
    • Abstract:- arXiv:2608.14634v1 Announce Type: new.
      • Abstract: Biological intelligence naturally prevents catastrophic forgetting through Complementary Learning Systems (CLS) theory, a macroscopic consolidation process driven by synaptic metaplasticity at the local level: continuous, history-dependent neuromodulation of individual synapses.
      • While artificial neural networks struggle with the stability-plasticity dilemma in non-stationary environments, existing solutions often require task labels or incur significant memory overhead, contrary to biological reality.
      • Re-framing this localized neuromodulation as an optimization-driven process, we introduce $\textbf{SynGAP}$: $\textbf{Syn}$aptic $\textbf{G}$eometric $\textbf{A}$daptive $\textbf{P}$reconditioning.
    • EN Key Points:
      • arXiv:2608.14634v1 Announce Type: new
      • Abstract: Biological intelligence naturally prevents catastrophic forgetting through Complementary Learning Systems (CLS) theory, a macroscopic consolidation pr…
      • While artificial neural networks struggle with the stability-plasticity dilemma in non-stationary environments, existing solutions often require task labels or…
      • Re-framing this localized neuromodulation as an optimization-driven process, we introduce $\textbf{SynGAP}$: $\textbf{Syn}$aptic $\textbf{G}$eometric $\textbf{A…
  • Fractional Optimizers Meet Fractal Activation Functions: An Empirical Study of Multi-Scale Optimization in Neural Network

    • Publication Time: 2026-08-19 12:00 Beijing Time
    • Abstract:- arXiv:2608.14636v1 Announce Type: new.
      • Abstract: Fractional optimization methods and fractal activation functions are two independent directions for improving neural network training.
      • Fractional optimizers extend first-order optimization through fractional derivatives and memory effects, while fractal activations introduce multi-scale nonlinear representations based on self-similar Weierstrass and Blancmange-type functions.
      • Here, we investigate their interaction within a unified experimental framework.
    • EN Key Points:
      • arXiv:2608.14636v1 Announce Type: new
  • Abstract: Fractional optimization methods and fractal activation functions are two independent directions for improving neural network training

  • Fractional optimizers extend first-order optimization through fractional derivatives and memory effects, whereas fractal activations introduce multi-scale nonli…

  • Here, we investigate their interaction within a unified experimental framework