🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-07-21
- 类型
- ai-daily
- 字数
- 7503
- 阅读时长
- 36 min
2026-07-21 AI Daily | Long-Horizon Agents Enter Field-Testing Phase: Controllability, Medical Specialization, and Audit Chains Become New Thresholds Link to heading
Today’s focus shifts from model capabilities to the controllable deployment of long-horizon agents. OpenAI warns that long-term autonomous models can exhibit runaway behaviors difficult to capture in short evaluations, necessitating stronger engineering for verification, rollbacks, and tool boundaries. Medical agents continue to specialize, and auditable causal reasoning is also making a comeback.
📖 In-depth Guide to This Issue’s Watch List Link to heading
The most noteworthy topic today is the “Controllability of Long-Horizon Agents”: A safety article on long-horizon models warns that the more continuously and autonomously a model operates, the more likely it is to exhibit runaway behaviors not captured by short evaluations. Paired with papers on multi-agent mathematical review, the ARC-AGI-3 code agent, and the local voice assistant AnovaX, this is an excellent opportunity for engineering teams to reassess their approaches to verification, rollbacks, and tool boundaries.
The second major theme is the accelerating specialization of medical agents. Research on Cura 1T, GraphDx, and clinical multimodal prediction all attempt to unify diagnostic reasoning, cost constraints, EHR tool usage, and multimodal patient records. This is worth close attention from teams working on medical AI.
Additionally, explainable and auditable reasoning is making a comeback. Causal-Audit, Prolog-based reinforcement learning explanation, trusted AI tool analysis, and research into a model’s “global workspace” all point to a common trend: the next stage of competition will not just be about the right answers, but whether the reasoning chain is inspectable and reproducible.
🌐 AI Hot Topics on X Link to heading
Topic 1: Andrew Ng Course Revives Graphs vs Loops Debate in Agentic AI Link to heading
- Category: AI · News
- Overview: Trending since:, Related posts: 334
- What it is: Andrew Ng’s new course has sparked a debate on whether agentic AI should adopt “graph-structured workflows” or “loop-based autonomous reasoning.”
- Why it’s important: This concerns the controllability, explainability, reliability, and engineering implementation of agent systems, representing a critical architectural choice as AI applications move from demos to production.
- Discussion Summary: The discussion on X centers on whether graph structures are better suited for stably orchestrating complex tasks, or if loop-based agents are closer to autonomous decision-making. Supporters emphasize that graphs are debuggable and monitorable, while opponents argue that excessive proceduralization limits agent flexibility.
Topic 2: Anthropic’s Claude Fable Disproves Jacobian Conjecture in 3D Link to heading
- Category: AI · News
- Overview: Trending since: 20 hours ago, Related posts: 35,000
- What it is: A hot topic on X claims that Anthropic’s Claude Fable has provided a counterexample or negative proof for the Jacobian Conjecture in three dimensions, though the claim awaits authoritative verification.
- Why it’s important: If true, this would be a major breakthrough for AI in high-level mathematical discovery, potentially changing assessments of large models’ reasoning, automated theorem proving, and scientific research assistance capabilities.
- Discussion Summary: The focus of the discussion is on whether the proof is authentic and reliable, whether there are model hallucinations or derivation flaws, and whether it has undergone peer review by the mathematics community. Supporters see it as a milestone for AI’s scientific research capabilities, while skeptics emphasize the need for the complete proof to be published and reviewed by experts.
Topic 3: Moonshot AI’s Kimi K3 Tops Coding Benchmarks at Lower Cost Link to heading
- Category: AI · News
- Overview: Trending since: 5 hours ago, Related posts: 1,700
- What it is: Moonshot AI’s Kimi K3 is reported to have achieved leading performance, or close to that of top closed-source models, on front-end code and software engineering-related benchmarks, offered at a lower price. The open-source weights are expected to be released soon.
- Why it’s important: This indicates that China’s open-source models are narrowing the capability gap with cutting-edge closed-source models. Especially in high-value scenarios like coding, this could intensify price competition and encourage enterprises to consider locally deployable, lower-cost AI solutions more seriously.
- Discussion Summary: The discussion on X focuses on whether Kimi K3 truly matches or surpasses closed-source models like Claude, whether benchmark tests can represent real-world development capabilities, the impact of its low-price strategy on the business models of closed-source competitors, and the innovation opportunities and security governance risks brought by open-sourcing the weights.
Topic 4: Debate Heats Over China’s Kimi AI Models and U.S. Response Link to heading
- Category: AI · News
- Overview: Trending since: 2 days ago, Related posts: 21,000
- What it is: Extensive discussions have emerged on X regarding the improving capabilities of China’s Moonshot AI Kimi model and its impact on the U.S. AI competitive landscape.
- Why it’s important: Kimi and other Chinese large models are seen as rapidly catching up in areas like long-context processing, reasoning, and low-cost deployment, potentially intensifying the AI technology, capital, and policy competition between the U.S. and China.
- Discussion Summary: The discussion centers on whether Kimi’s actual technical level is overestimated, whether the U.S. needs stronger industrial policies or export controls, and the impact of open-sourcing, compute limitations, and the speed of Chinese AI innovation on the global market.
Today’s AI Public Opinion Summary on X Link to heading
The main public opinion today focuses on the key turning point for AI, moving from “capability demonstration” to “engineering, research, and industrial competition.” On one hand, discussions on agent architecture show a general industry consensus that controllability, debuggability, and reliability have become core to implementation. On the other hand, the performance and low-cost strategies of Chinese models like Kimi K3 are convincing more people that open weights and cost advantages are reshaping the large model competitive landscape. The consensus is that AI is rapidly entering high-value scenarios such as coding, scientific research, and automated workflows, and that open-source or open-weight models are closing the gap with closed-source frontier models. The main points of divergence are threefold: whether agent architectures should lean towards graph-based orchestration or recursive autonomous reasoning; whether the so-called mathematical breakthrough of Claude Fable is a genuine discovery or a hallucination; and whether Kimi’s benchmark scores can represent true engineering capabilities. Potential risks include over-reliance on unverified AI reasoning results, benchmark hype masking actual defects, security governance pressures from low-cost and open-source competition, and the further amplification of the US-China AI competition by technological nationalism and policy regulations.
💡 Influencer Insights Link to heading
This AI industry daily report analysis is generated based on the provided tweet content.
AI Industry Daily: Large Model “Arms Race” Heats Up, Toolchain Reshaping and Worldview Construction Link to heading
1. Shared Technical Trends and Product Hotspots Link to heading
🔥 Large Model Performance in “Close Quarters Combat”: The Three Kingdoms-style Battle of Kimi K3, Qwen3.8, and Claude Fable 5 Link to heading
Undoubtedly, the biggest hotspot of the day was the successive impact of Kimi K3 and Qwen3.8-Max-Preview, forming a pincer movement against Claude Fable 5, currently recognized as the strongest closed-source model.
- Kimi K3: The “Frontend” Surprise Attack from the Largest Open-Source Model in History. This model, with its 2.8T parameters, captured all the attention. Its capabilities in frontend design and game generation are regarded by many influencers as an extremely strong suit. Bloggers like @Pluvio9yte and @vista8 both pointed out that K3 is even better than or on par with Fable 5 in frontend aesthetics and playable DEMO generation. @ruanyf also analyzed that the core reason for its performance being close to Fable 5 lies in the brute-force increase in parameter count, but also noted that its API fees (20/100 RMB per million tokens for input/output) are already among the most expensive in China.
- Qwen3.8-Max-Preview: The “No Weaknesses” Challenger in Full-Stack Engineering. Just three days after the release of Kimi K3, Alibaba unveiled its 2.4T-parameter competitor. @Pluvio9yte, citing leaked evaluation data, pointed out that Qwen3.8 has already surpassed K3 and performs robustly in complex engineering tasks and multi-agent orchestration. It is on par with Claude Opus 4.8 overall, lagging only behind Fable 5. This confirms the industry view relayed by @dotey: a model’s release should not only be judged by its weaknesses, but also by whether its strengths can be a “game-changer.”
- Fable 5’s Battle to Defend the “Throne”. Facing this siege, Anthropic’s strategy is to maintain its position through business tactics. Both @zhixianio and @Pluvio9yte mentioned that Fable 5 has not only extended access periods for paid users but also remains irreplaceable in handling specific and difficult problems. @dotey shared a typical case: when encountering a VBR MP3 timestamp deviation issue, other models were helpless, while Fable 5 could accurately locate and fix it. This is defined as its barrier in extreme scenarios.
🛠️ AI Programming Tools Enter the Era of “Model Hybridization” and “Skill Monetization” Link to heading
Standalone terminal programming tools can no longer meet demands; integrating the advantages of different models has become the new trend.
- Multi-model Routing and Integration. @vista8 introduced the OpenCodex project, which allows users to switch between Kimi K3 (frontend), GPT 5.6 Sol (backend), and Grok 4.5 (search) within the Codex interface at any time, breaking down platform lock-in barriers. The Skill he developed even enables the orchestration of various local CLI models within Codex with a single sentence, achieving compliant use by “combining the strengths of all.”
- Skill Ecosystem Explosion and Community Building. The auto-editing skill developed by @vista8, and the Xiaohongshu REDSkill Community discovered by @ruanyf, signify that Skills (skill plugins) are evolving from auxiliary scripts to core product features, even becoming a new vehicle for social media dissemination and traffic acquisition.
🌐 “World Models” Begin to Emerge: From Generating Videos to Generating Interactive Worlds Link to heading
While models like Sora are still at the stage of generating videos, another track has begun to explore a deeper level of understanding. @Pluvio9yte provided an in-depth analysis of the open-source Alaya World project, believing it represents the trend of AI evolving from a “tool” to an “environment.” It can generate streamable scenes that can be freely moved and interacted with in real-time based on instructions, demonstrating a preliminary understanding of space, time, and causality, rather than simple next-frame prediction. This is seen as a potential fundamental revolution for fields like gaming and embodied intelligence.
2. Noteworthy Unique Perspectives and Industry Foresight Link to heading
- The “Class Divide” in AI Costs is Irreversible (@Pluvio9yte): Points out that as the prices of top-tier models like Fable 5 continue to rise, future model subscription fees will keep increasing, making the cost of using state-of-the-art productivity tools prohibitive. This will further widen the productivity gap. @zhixianio had also previously lamented the rising hardware costs.
- FDEs (Front-end Engineers) are a “Grand Strategy” for Model Companies (@dotey): Sharply observes that AI companies are using the opportunity of business implementation to have FDEs distill a company’s industry knowledge and best practices into Skills, which are then internalized by the model. For individuals, this creates a short-term technical moat, but for companies, in the long run, it could be a prelude to personnel “optimization” after cost reduction and efficiency improvements.
- Warning from the Surge in AI Code Commits (@ruanyf): Cites data showing a 14-fold year-over-year increase in code commits on GitHub. This not only explains the platform’s frequent outages but also leads to a prediction: if hosting costs continue to explode, a fully paid GitHub might not be far off.
- Top-Tier Models Remain Irreplaceable in “Edge Cases” (@dotey): Emphasizes that while the performance gap between models is narrowing in common scenarios, when faced with extreme logical problems like VBR audio/video encoding recognition or fixing specific, difficult bugs, Fable 5 still demonstrates “killer” reliability.
- Is “Open Source” a Misnomer? (@ruanyf, quoting the Anthropic CEO): Argues that today’s so-called “open source” AI models only release their weights, making it impossible for outsiders to inspect their internal logic. This is fundamentally different from traditional open-source software, and it would be more accurate to call them “open-weight”.
- A New Approach to Product Design: From “Gamification” to “Game-like Feel” (@nishuang): When designing AI applications (like a vocabulary memorization app), one should move away from dopamine-driven reward systems (Gamification). Instead, by stimulating curiosity and endorphins, create a game-like feel (Game-like design) that makes users feel like they are “playing” rather than “grinding.”
3. Recommended Tools & Resources Link to heading
- OpenCodex: Unlocks the model constraints of Codex. It allows you to directly call multiple external models like Kimi K3 and Grok 4.5 from within the Codex interface, enabling efficient combination patterns such as using K3 for the front end and Sol for the back end. Source: @vista8.
- Grok Build: An open-source terminal AI programming agent from Musk’s SpaceX AI team. Written purely in Rust, it features highly customizable configurations like MCP support, sandbox mode, and headless mode. It already has 14k stars. Ideal for developers who enjoy tinkering with and customizing CLI tools. Source: @AI_Jasonyu.
- MOSS-Transcribe-Diarize-0.9B: A lightweight, open-source speech-to-text model released by Alibaba. It can process up to 90 minutes of audio at once and directly outputs timestamped text with speaker diarization. @dotey tested it and found the transcription to be accurate. Suitable for local processing of long audio files like podcasts and meeting minutes.
- BaoCut Skill / Qiaomu-Cut Skill: If you manage a WeChat Channels account or a video clip account, the automatic video editing skills developed or integrated by @dotey and @vista8 allow you to use text commands to automate the entire workflow from footage retrieval to final assembly. You can even generate short, subtitled video clips for English learning.
- Alaya World: Want a preview of a “world model”? You can find the open-source inference code and weights in its GitHub repository. Try turning text or images directly into a world you can freely navigate and interact with. Source: @Pluvio9yte.
- Supporting Infrastructure for Skill Monetization: @ruanyf and @vista8 mentioned the Xiaohongshu (RED) Skill community and WeChat Channels download tools. These are excellent infrastructure for Skill distribution and customer acquisition.
📚 Appendix: Today’s Watch List Source Update Link to heading
Timeframe: Last 3 days; 22 sources covered; 32 updates total
Stratechery by Ben Thompson (A_full) Link to heading
- Who’s Afraid of Chinese Models?
- Publication Time: 2026-07-20 19:00 Beijing Time
- Abstract: - Listen to this post**:**.
- I told a story about my first day in STRT-431 at the Kellogg School of Management, the introductory strategy class every first year MBA had to take; I went through the readings and the case studies and, to my frustration, there wasn’t a single tech company on the list.
- To my point, I talked to the professor after class, wondering why, and was told that the goal of the class wasn’t necessarily to understand specific industries, but rather to discover generally-applicable universal principles that could be applied to any company in any industry.
- As I usually tell the story, I didn’t find this very satisfying: to me the nature of technology, and specifically the fact that software and distribution have zero marginal costs (and zero transaction costs), was fundamentally different; inputting zero into formulas tends to wreak havoc
- However, I soon realized this was my opportunity.
- EN Highlights:
- Listen to this post :
- Log in to listen
- There’s a story I tell about my first day in STRT-431 at Kellogg School of Management, the introductory class that every first-year MBA was required to take; I…
- Me being me, I spoke to the professor after class wondering why, and was told that the goal of the course was not to necessarily learn about specific industries…
- EN Highlights:
OpenAI Blog (A_full) Link to heading
- Safety and alignment in an era of long-horizon models
- Publication Time: 2026-07-20 18:00 Beijing Time
- Abstract: - Models that can work autonomously for extended periods can solve difficult, open-ended problems.
- But the same persistence that makes them useful also gives them more opportunities to take unwanted actions, and in ways that evaluations for short-horizon models might miss.
- The model is designed to work autonomously for long periods.
- During limited, monitored internal use, we observed undesirable behaviors not captured by existing deployment evaluations.
- Because the deployment was limited and monitored, we were able to identify these issues, pause access, create new evaluations based on what we observed, strengthen the model and its safeguards, and then resume access with continued monitoring.
- EN Highlights:
- OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deploym…
ArXiv cs.AI (B_intro+search) Link to heading
GraphDx: A Cost-Aware Knowledge-Enhanced Multi-Agent Framework for Sequential Diagnosis
- Publication Time: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15280v1 Announce Type: new.
- Abstract: Sequential diagnosis requires balancing diagnostic accuracy and resource costs through iterative information gathering.
- Existing Large Language Model (LLM) approaches exhibit a critical knowledge-reasoning gap: despite encoding extensive medical knowledge, they struggle to reason systematically under cost constraints, often resorting to excessive testing.
- We propose GraphDx, a knowledge-enhanced framework with two core innovations.
- EN Highlights:
- arXiv:2607.15280v1 Announce Type: new
- Abstract: Sequential diagnosis requires balancing diagnostic accuracy against resource costs through iterative information gathering
- Existing Large Language Model (LLM) approaches exhibit a critical knowledge-reasoning gap: despite encoding extensive medical knowledge, they struggle to reason…
- We propose GraphDx, a knowledge-enhanced framework with two core innovations
- Publication Time: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15281v1 Announce Type: new.
Abstract: Causal and intervention-based question answering is fundamental to advancing large language models (LLMs) toward reasoning beyond surface-level correlations and understanding underlying causal mechanisms.
However, existing LLM-based methods often rely on implicit language-level reasoning, resulting in opaque causal assumptions, unverifiable reasoning paths, and fragile predictions under complex interventions, especially in context-free environments.
In this paper, we propose an explicit and auditable causal reasoning framework for context-free intervention-based question answering.
EN Highlights:
- arXiv:2607.15281v1 Announce Type: new
- Abstract: Causal and intervention-based question answering is fundamental to advancing large language models (LLMs) toward reasoning beyond surface-level correl…
- However, existing LLM-based methods often rely on implicit language-level reasoning, resulting in opaque causal assumptions, unverifiable reasoning paths, and f…
- In this paper, we propose an explicit and auditable causal reasoning framework for context-free intervention-based question answering
Cura 1T: Specialized Model for Agentic Healthcare
- Publication Time: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15314v1 Announce Type: new.
- Abstract: Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain limited.
- A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use.
- These capabilities fail in different ways, and a narrow update for one task can degrade another’s performance.
- EN Highlights:
- arXiv:2607.15314v1 Announce Type: new
- Abstract: Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain…
- A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use
- These capabilities fail in different ways, and a narrow update for one task can degrade another
- Publication Time: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15367v1 Announce Type: new.
- Abstract: Desktop voice assistants are still dominated by cloud pipelines that transmit raw audio off the machine and expose a fixed set of skills.
- We describe AnovaX, a small, local-first assistant that runs entirely on the user’s computer and treats the desktop itself as its operating surface.
- A single Python process gates a wake word, a voice pipeline, an LLM planner (Gemini) that emits tool-calling JSON plans, allowlist and denylist safety layers, a multi-agent coordinator that converts each plan into typed sub-agents on a bounded thread pool, and an adaptive recovery loop that takes over when core steps fail.
- EN Highlights:
arXiv:2607.15367v1 Announce Type: new
- Abstract: Desktop voice assistants are still dominated by cloud pipelines that ship raw audio off the machine and expose a fixed set of skills
- We describe AnovaX, a small local-first assistant that runs entirely on the user’s computer and treats the desktop itself as its action surface
- A single Python process wires together a wake-word gate, a speech pipeline, an LLM planner (Gemini) that emits a JSON plan of tool calls, a whitelist-and-denyli…
- Published: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15388v1 Announce Type: new.
- Abstract: Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review stage should help turn incorrect candidates into correct ones.
- We test this assumption on 4,181 verifier-grounded Omni-MATH problems using matched gpt-oss-120b actors.
- Collaboration adds little at the easiest levels, but from Tier 4 onward the gains increase sharply; in this more difficult regime, broadcast-style peer discussion reaches a higher final accuracy than the Planner-Executor-Reviewer (PER) pipeline.
- EN Highlights:
- arXiv:2607.15388v1 Announce Type: new
- Abstract: Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review stage should…
- We test this assumption on 4,181 verifier-grounded Omni-MATH problems using matched gpt-oss-120b actors
- Collaboration adds little on the easiest tiers, but from tier 4 onward the gains open sharply; in this harder regime, broadcast-style peer discussion reaches hi…
DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings
- Published: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15418v1 Announce Type: new.
- Abstract: We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings—the core media for architecture, civil, and many other engineering practices.
- Unlike natural images or schematic floor plans, construction drawings fuse abstract geometry, symbolic notation, tabular data, annotations, and domain-specific text, forming a uniquely complex visual-textual domain core to engineering workflows.
- DrawingVQA bridges this gap with 33 “Issued for Construction” drawings and 92 professionally-curated question-answer pairs spanning three reasoning depths: perceptual understanding, contextual interpretation, and domain expert reasoning.
- EN Highlights:
- arXiv:2607.15418v1 Announce Type: new
Abstract: We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings – a co…
- Unlike natural images or schematic floor plans, construction drawings fuse abstract geometry, symbolic notation, tabular data, annotations, and domain-specific…
- DrawingVQA bridges this gap with 33 “Issued for Construction” drawings and 92 expertly curated question-answer pairs, spanning three reasoning depths: perceptua…
Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?
- Posted: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15439v1 Announcement Type: new.
- Abstract: Our previous ARC-AGI-3 agent bundled executable world modeling, scheduled simplification, and exact replay verification, leaving unclear which idea accounted for its performance.
- We address this attribution question with four nested Codex-based agents: a textual baseline; a flexible-interface executable world model without replay verification; the same executable model with scheduled simplification; and a fixed-interface verification process which preserves simplification and requires exact reproduction of recorded observations.
- The main study evaluates all four agents with gpt-5.4 and gpt-5.5 at high and xhigh reasoning effort on the public ARC-AGI-3 games.
- EN Highlights:
- arXiv:2607.15439v1 Announce Type: new
- Abstract: Our previous ARC-AGI-3 agent bundled executable world modeling, scheduled simplification, and exact replay verification, leaving unclear which idea ac…
- We address this attribution question with four nested Codex-based agents: a textual baseline; a flexible-interface executable world model without replay verific…
- The main study evaluates all four agents with gpt-5.4 and gpt-5.5 at high and xhigh reasoning effort on the public ARC-AGI-3 games
Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes
- Posted: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15442v1 Announcement Type: new.
- Abstract: Internet memes intertwine visual cues, textual content, and cultural context, making them particularly difficult to interpret in scenarios where humor, satire, and harmful intent coexist.
- These complexities highlight the need for an interpretable meme understanding system that can provide reliable and structured reasoning to support accurate classification and human interpretability.
- However, existing multimodal classifiers either ignore these interdependencies or provide only limited interpretability.
- EN Highlights:
- arXiv:2607.15442v1 Announce Type: new
Abstract: Internet memes intertwine visual cues, textual content, and cultural context, making them particularly challenging to interpret in scenarios where hum…
These complexities highlight the need for explainable meme understanding systems that can provide reliable and structured reasoning to support both accurate cla…
However, existing multimodal classifiers either overlook these interdependencies or provide only limited interpretability
From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems
- Publication Time: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15459v1 Announcement Type: New.
- Abstract: A trained deep reinforcement learning policy is a black box, and we ask whether it can be made explainable by rewriting it as an executable logic program that can reproduce its behavior, is human-readable, can be run by a logic engine, and can be edited by an optimizer.
- We propose a three-stage post-hoc transformation that extracts a frozen proximal policy optimization teacher, induces an ordered rule list from its decisions in the manner of classic relational learning, and emits the result as a Prolog program whose every decision is executed by an off-the-shelf logic engine; a subsequent extension phase edits the rule base, and an edit is accepted only if a policy evaluation demonstrates an increased return.
- Return-loss bounds make the distilled program a machine-checkable certificate in a finite Markov decision process, and the extension loop monotonically improves and terminates.
- EN Key Points:
- arXiv:2607.15459v1 Announce Type: new
- Abstract: A trained deep reinforcement learning policy is a black box, and we ask whether it can be made explainable by rewriting it as an executable logic prog…
- We present a three-stage post-hoc transformation that extracts a frozen proximal policy optimization teacher, induces an ordered rule list from its decisions in…
- We prove four guarantees
A Critical Analysis of Trustworthy AI Tools, Mark Frameworks, and the Implementation Chasms
- Publication Time: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15480v1 Announcement Type: New.
- Abstract: As artificial intelligence (AI) systems increasingly impact society, ensuring their ethical and trustworthy deployment has become a global priority.
- Although numerous high-level ethical guidelines have emerged, criticism persists that these frameworks remain abstract and lack concrete implementation mechanisms.
- This paper utilizes a comprehensive dataset from the OECD to conduct a critical analysis of tools and trust-marking frameworks designed to implement Trustworthy AI (TAI).
- EN Key Points:
- arXiv:2607.15480v1 Announce Type: new
- Abstract: As artificial intelligence (AI) systems increasingly impact society, ensuring their ethical and trustworthy deployment has become a global priority
While a myriad of high-level ethical guidelines have emerged, criticism persists that these frameworks remain abstract and lack concrete mechanisms for implemen…
This paper conducts a critical analysis of tools and trust mark frameworks intended to operationalize trustworthy AI (TAI), drawing on a comprehensive dataset f…
ArXiv cs.CL (B_intro+search) Link to heading
Large Language Models as Unified Multimodal Learners for Clinical Prediction
- Published: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15380v1 Announce Type: new.
- Abstract: Electronic health records combine free-text clinical narratives with structured measurements such as vital signs, laboratory values, and comorbidities.
- However, most clinical prediction systems still rely on task-specific fusion architectures, pairing dedicated encoders for each modality with learned combination mechanisms that must be redesigned for each new task and clinical environment.
- We propose a simpler alternative: convert all patient data, regardless of modality, into a single natural language sequence and fine-tune a pretrained language model end-to-end, without architectural modifications for fusion.
- EN Highlights:
- arXiv:2607.15380v1 Announce Type: new
- Abstract: Electronic health records combine free-text clinical narratives with structured measurements such as vital signs, laboratory values, and comorbidities
- Yet most clinical prediction systems still rely on task-specific fusion architectures, pairing dedicated encoders for each modality with learned combination mec…
- We propose a simpler alternative: convert all patient data, regardless of modality, into a single natural language sequence and fine-tune a pretrained language…
Verbalizable Representations Form a Global Workspace in Language Models
- Published: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15495v1 Announce Type: new.
- Abstract: Of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning.
- In this paper, we present evidence that an analogous functional distinction has emerged in large language models.
- Using a new interpretability technique, the Jacobian Lens, we can identify the representations the model is prepared to express at any point in its processing.
- EN Highlights:
- arXiv:2607.15495v1 Announce Type: new
- Abstract: Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, delib…
- In this paper, we present evidence that an analogous functional distinction has emerged in large language models
Using a new interpretability technique, the Jacobian lens, we identify the representations a model is poised to verbalize at any point in its processing
VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs
- Publication Time: 2026-07-20 12:00 Beijing Time
- Abstract:- arXiv:2607.15498v1 Announce Type: New.
- Abstract: Key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference.
- Two leading training-free families are both structurally limited: token-selection methods (SnapKV, Ada-KV) score importance from an observation window and evict low-scoring tokens, but eviction is irreversible—thus, accuracy drops by 11-15 points when importance signals decline under query-agnostic reuse; uniform low-rank encoding keeps every token, but spends the same rank everywhere, wasting budget.
- We observe that both failures share one cure: rank should be allocated, not evicted.
- EN Key Points:
- arXiv:2607.15498v1 Announce Type: new
- Abstract: The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference
- Two leading training-free families are both structurally limited: token-selection methods (SnapKV, Ada-KV) score importance from an observation window and evict…
- We observe that both failures share one cure: rank should be allocated, not evicted
EpiNarrate: Agentic Generation of Grounded Narratives from Epidemiological Scenario Projections
- Publication Time: 2026-07-20 12:00 Beijing Time
- Abstract:- arXiv:2607.15544v1 Announce Type: New.
- Abstract: Generating clear and accessible public health narratives is critical for communicating complex epidemiological projections to policymakers and the broader public.
- Such narratives require more than simply reporting numbers: projections must be contextualized and quantitatively grounded across multiple dimensions.
- Furthermore, projections are often derived from large ensemble datasets which combine intervention assumptions, geographic and demographic strata, outcomes, time horizons, and uncertainty quantiles.
- EN Key Points:
- arXiv:2607.15544v1 Announce Type: new
- Abstract: Generation of clear and accessible public health narratives is critical for communicating complex epidemiological projections to policymakers and the…
- Such narratives require more than simply reporting numbers: projections must be contextualized and quantitatively grounded across multiple dimensions
- Further, projections are often derived from large ensemble datasets which combine intervention assumptions, geographic and demographic strata, outcomes, time ho…
SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents
- Publication Time: 2026-07-20 12:00 Beijing Time
- Summary: - arXiv:2607.15557v1 Announcement Type: new.
- Abstract: Agent skills (SKILL.md files), which package reusable procedural knowledge for LLM agents, are a popular mechanism for extending agent capabilities.
- Public repositories now host a large and growing number of these artifacts, but they are fragmented, redundant, and of uneven quality, with their practical value being unclear.
- A core question remains unresolved: how to consolidate this open-source SKILL.md ecosystem into a usable corpus and what limits its benefits for real-world agent tasks.
- EN Highlights:
- arXiv:2607.15557v1 Announce Type: new
- Abstract: Agent skills, SKILL.md files that package reusable procedural knowledge for an LLM agent, are a popular mechanism for extending agent capabilities
- Public repositories now host them in large and growing numbers, yet these artifacts are fragmented, redundant, and uneven in quality, and their value in practic…
- A core question remains open, namely how to consolidate this open-source SKILL.md ecosystem into a single usable corpus, and what bounds its benefit on real-wor…
Process Reward Informed Tree Rollout for Effective Multi-Turn RL
- Publication Time: 2026-07-20 12:00 Beijing Time
- Summary: - arXiv:2607.15610v1 Announcement Type: new.
- Abstract: Reinforcement learning (RL) has become a key method for training LLM agents, yet popular methods like GRPO/RLOO rely on multiple independently sampled complete trajectories for advantage estimation.
- In long-horizon agent tasks, this uniform rollout strategy can waste budget on uninformative dead-end attempts, while promising intermediate states are not sufficiently explored.
- The multi-turn structure of agent trajectories, with interleaved actions and observations, naturally supports organizing a group of trajectories into a tree, where each turn serves as a decision point for exploration.
- EN Highlights:
- arXiv:2607.15610v1 Announce Type: new
- Abstract: Reinforcement learning (RL) has become a key approach for training LLM agents, yet popular methods such as GRPO/RLOO rely on multiple independently sa…
- In long-horizon agentic tasks, such a uniform rollout strategy can waste budget on uninformative dead-end attempts, while promising intermediate states do not r…
- The multi-turn structure of agentic trajectories, with interleaved actions and observations, naturally supports organizing a trajectory group as a tree, where e…
On the Structure of Address in Multi-Party Dialogue: From Discrete Labels to Continuous Levels
Publication Time: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15648v1 Announce Type: New.
- Abstract: In multi-party dialogues between a dialogue system and multiple users, identifying the object of an utterance is a key challenge.
- Previous work has typically treated recipient detection as a multi-class classification task, selecting a single label representing an individual participant or group.
- This formulation assumes that addressing is inherently discrete and is primarily used for predicting turn-taking.
- EN 要点:
- arXiv:2607.15648v1 Announce Type: new
- Abstract: In multi-party dialogues between a dialogue system and multiple users, identifying to whom an utterance is addressed is a key challenge
- Prior work has typically treated addressee detection as a multi-class classification task, selecting a single label representing an individual participant or th…
- This formulation assumes that address is inherently discrete and has primarily been used for predicting turn-taking
- Abstract: - arXiv:2607.15648v1 Announce Type: New.
Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models
- Publication Time: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15655v1 Announce Type: New.
- Abstract: Masked Diffusion Language Models (DLMs) achieve parallel text generation by iteratively refining masked tokens, offering a promising alternative to autoregressive decoding.
- Recent lookahead-based decoding methods improve the accuracy-efficiency trade-off by exploring future decoding states before committing token updates.
- However, existing methods mainly rely on shallow one-step lookahead, which optimizes immediate information gain but may not be optimal for longer-range decoding trajectories.
- EN 要点:
- arXiv:2607.15655v1 Announce Type: new
- Abstract: Masked diffusion language models (DLMs) enable parallel text generation by iteratively refining masked tokens, offering a promising alternative to aut…
- Recent lookahead-based decoding methods improve the accuracy–efficiency trade-off by exploring future decoding states before committing token updates
- However, existing approaches mainly rely on shallow one-step lookahead, which optimizes immediate information gain but can be suboptimal for longer-horizon deco…
- Publication Time: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15736v1 Announce Type: New.
- Abstract: Large reasoning models often solve problems via long Chain-of-Thought (CoT) trajectories, but most computation is spent on redundant derivations, repetitive self-verification, and detours that do not improve the final answer.
- Existing policy self-distillation methods reduce this cost by matching a student model to a concise copy of itself on prefixes sampled from the student’s own rollouts.
- We show that this objective has an initialization bottleneck.
- EN 要点:
- arXiv:2607.15736v1 Announce Type: new
Abstract: Large reasoning models often solve problems through long chain-of-thought (CoT) traces, yet much of this computation is spent on redundant derivations…
Existing on-policy self-distillation methods reduce this cost by matching a student model to a concise copy of itself on prefixes sampled from the student’s own…
We show that this objective has an initialization bottleneck
Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery
- Release Time:2026-07-20 12:00 Beijing Time
- Abstract:- arXiv:2607.15766v1 Announce Type: new.
- Abstract: Large language models (LLM) excel at answering pre-specified questions, but their ability to navigate the open-ended, pre-conclusion discovery phase largely remains unmeasured.
- We introduce Prospective Hypothesis Discovery (PHD), which asks models to autonomously construct grounded, distinctive, and testable hypothesis spaces based on inconclusive evidence (including anomalous observations and fragmented records) to guide subsequent investigations.
- To evaluate this capability, we introduce HypoArena, which includes HypoData (a benchmark of 988 cases across six scientific and analytical domains) and HypoEval (an evaluation framework for open-ended hypothesis sets).
- EN Key Points:
- arXiv:2607.15766v1 Announce Type: new
- Abstract: Large language models (LLMs) excel at answering pre-specified questions, yet their ability to navigate the open-ended, pre-conclusion stage of discove…
- We introduce Prospective Hypothesis Discovery (PHD), which asks models to autonomously construct grounded, discriminative, and testable hypothesis spaces from i…
- To evaluate this capability, we introduce HypoArena, comprising HypoData, a benchmark of 988 cases across six scientific and analytical domains, and HypoEval, a…
ArXiv cs.LG (B_intro+search) Link to heading
Structure of the Circular-Dyadic Convolution Error
- Release Time:2026-07-20 12:00 Beijing Time
- Abstract:- arXiv:2607.15293v1 Announce Type: new.
- Abstract: Both dyadic and circular convolution can be computed in $O(N\log N)$ time using the Hadamard transform and the discrete Fourier transform (DFT) computed by FFT, respectively.
- The Hadamard transform is preferable due to its real-valued sign flips, but its substitution for the DFT introduces algebraic errors.
- We propose three complementary results to characterize this error.
- EN Key Points:
- arXiv:2607.15293v1 Announce Type: new
- Abstract: Dyadic and circular convolution can both be computed in $O(N\log N)$ time using the Hadamard transform and the FFT-computed discrete Fourier transform…
The Hadamard transform is preferable for its real-valued sign flips, yet its substitution for the DFT introduces algebraic error
- We present three complementary results that characterize this error
Position: Quantum Program Generation Must Prioritize Validity Over Probabilistic Scaling
- Publish Time: 2026-07-20 12:00 Beijing Time
- Abstract:- arXiv:2607.15313v1 Announcement Type: New.
- Abstract: The scaling hypothesis assumes that increasing model parameters yields emergent reasoning capabilities.
- This position paper argues that applying this probabilistic paradigm to generic quantum circuit synthesis is a directional error.
- Unlike natural languages, quantum circuits require strict adherence to mathematical constraints that manifest a significant syntax-semantics gap.
- EN Key Points:
- arXiv:2607.15313v1 Announce Type: new
- Abstract: The scaling hypothesis assumes that increasing model parameters yields emergent reasoning capabilities
- This position paper argues that applying this probabilistic paradigm to generic quantum circuit synthesis is a directional error
- Unlike natural languages, quantum circuits require strict adherence to mathematical constraints that manifest a significant syntax-semantics gap
A Transportable Threshold-Based Framework for Interpretable Classification of Medical Data
- Publish Time: 2026-07-20 12:00 Beijing Time
- Abstract:- arXiv:2607.15394v1 Announcement Type: New.
- Abstract: Black-box models limit the adoption of artificial intelligence in medicine due to their lack of interpretability and reproducibility.
- We introduce a statistically grounded framework that provides fully interpretable, rule-based clinical classification using the Bernoulli Naive Bayes (BNB) model.
- The method applies supervised $\chi^2$-guided statistical binarization to continuous variables, identifying thresholds that maximize association with clinical outcomes.
- EN Key Points:
- arXiv:2607.15394v1 Announce Type: new
- Abstract: Black-box models limit the adoption of artificial intelligence in medicine due to their lack of interpretability and reproducibility
- We introduce a statistically grounded framework that provides fully interpretable, rule-based clinical classification using the Bernoulli Na"ive Bayes (BNB) model.
- The method applies supervised $\chi^2$-guided statistical binarization to continuous variables, identifying thresholds that maximize association with clinical outcomes.
Regularity-Aware Stochastic MGDA with Adaptive Conflict-Avoidant Update Direction Control
Publication Time: 2026-07-20 12:00 Beijing Time
- Summary: - arXiv:2607.15412v1 Announcement Type: new.
- Abstract: Multi-objective learning (MOL) aims to optimize multiple objectives simultaneously.
- The multi-gradient descent algorithm (MGDA) is a workhorse that iteratively updates along a common descent or conflict-avoidant (CA) direction across objectives.
- However, in stochastic settings, the vanilla stochastic MGDA method, SMG, lacks a fast convergence rate because mini-batch sampling introduces noise in the gradients.
- EN Key Points:
- arXiv:2607.15412v1 Announce Type: new
- Abstract: Multi-objective learning (MOL) aims to optimize multiple objectives simultaneously
- The multi-gradient descent algorithm (MGDA) is a workhorse that iteratively updates along a common descent or conflict-avoidant (CA) direction across objectives
- In stochastic settings, however, the vanilla stochastic MGDA method, SMG, lacks a fast convergence rate because mini-batch sampling introduces noise in the grad…
- Summary: - arXiv:2607.15412v1 Announcement Type: new.
AI Trading: Evaluating Large Language Models for Technical Market Analysis
- Publication Time: 2026-07-20 12:00 Beijing Time
- Summary: - arXiv:2607.15414v1 Announcement Type: new.
- Abstract: Large Language Models (LLMs) have emerged as powerful tools for processing the heterogeneous information environments of modern financial markets.
- This paper presents a systematic, comparative evaluation of the technical market analysis capabilities of five prominent LLMs: GPT-4 Turbo, Claude 3 Opus, Gemini 1.5 Pro, Llama 3 70B, and the domain-specific FinGPT.
- The evaluation spans four structured tasks: candlestick pattern recognition from OHLCV data, directional signal generation (Buy/Sell/Hold), backtesting of signal quality via a simulated execution pipeline, and financial report comprehension.
- EN Key Points:
- arXiv:2607.15414v1 Announce Type: new
- Abstract: Large Language Models (LLMs) have emerged as powerful tools for processing the heterogeneous information environments of modern financial markets
- This paper presents a systematic, comparative evaluation of five prominent LLMs: GPT-4 Turbo, Claude 3 Opus, Gemini 1.5 Pro, Llama 3 70B, and the domain-special…
- The evaluation spans four structured tasks: candlestick pattern recognition from OHLCV data, directional signal generation (BUY/SELL/HOLD), backtesting of signa…
- Publication Time: 2026-07-20 12:00 Beijing Time
- Summary: - arXiv:2607.15421v1 Announcement Type: new.
- Abstract: Compact medical image classifiers require efficiency and interpretable evidence, but these objectives are often addressed separately.
- We introduce qZACH-ViT, a quantization-aware extension of the zero-token (CLS-less), position-free ZACH-ViT backbone, with recursive intrinsic patch-level class evidence.
We also introduce Recursive Attribution-Stabilized Optimization (RASO), which norm-matches classification and attribution gradients and removes attribution components that conflict with classification.
- EN Highlights:
- arXiv:2607.15421v1 Announce Type: new
- Abstract: Compact medical-image classifiers need efficiency and interpretable evidence, yet these goals are often addressed separately
- We introduce qZACH-ViT, a quantization-aware extension of the zero-token (CLS-token-free), position-free ZACH-ViT backbone with recursive intrinsic patch-level…
- We also introduce Recursive Attribution-Stabilized Optimization (RASO), which norm-matches classification and attribution gradients and removes attribution comp…
- EN Highlights:
- Release Time: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15433v1 Announce Type: new.
- Abstract: We characterize and compare the inherent interpretability of a standard linear model with that of a single-qubit mixed-state model for supervised binary classification tasks.
- A side-by-side comparison reveals that the single-qubit mixed-state model for binary classification is just the “ellipsoid version” of standard linear model classification.
- More precisely, rather than learning a hyperplane to classify data, we learn a hyperellipsoid.
- EN Highlights:
- arXiv:2607.15433v1 Announce Type: new
- Abstract: We characterize and compare the inherent interpretability offerings of a standard linear model with a single qubit mixed state model for the task of s…
- A side by side comparison reveals that a single qubit mixed state model for binary classification is just the ``ellipsoid version" of standard linear model clas…
- More precisely, rather than learning a hyperplane to classify data, we learn a hyperellipsoid
Stochastic Reset Pathfinding: Path-Level Regret for Cascading Bandits over Graph Paths
- Release Time: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15440v1 Announce Type: new.
- Abstract: We introduce Stochastic Reset Pathfinding (SRP), an episodic learning problem on a known directed graph with unknown, fixed edge success probabilities.
- In each episode, an agent commits to a source-to-target path, and any edge failure during execution resets it to the source.
- SRP captures settings such as entanglement distribution in quantum repeater networks, payment routing on the Lightning Network, and delivery in unreliable mesh networks.
- EN Highlights:
- arXiv:2607.15440v1 Announce Type: new
Abstract: We introduce Stochastic Reset Pathfinding (SRP), an episodic learning problem on a known directed graph with unknown stationary edge success probabili…
- In each episode, the agent commits to a source-to-goal path, and any edge failure during execution resets it to the source
- SRP captures settings such as entanglement distribution in quantum repeater networks, payment routing on the Lightning Network, and delivery in unreliable mesh…
- Published: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15446v1 Announce Type: new.
- Abstract: The cost of healthcare remains a concern in the United States and may have been influenced by disruptions associated with the COVID-19 pandemic.
- This study examines healthcare financial vulnerability before and after the pandemic using Medical Expenditure Panel Survey (MEPS) data from 2019 and 2021.
- High financial burden was defined as out-of-pocket healthcare expenditures exceeding 10% of family income.
- EN Highlights:
- arXiv:2607.15446v1 Announce Type: new
- Abstract: The cost of healthcare remains a concern in the United States and may have been influenced by disruptions associated with the COVID-19 pandemic
- This study examines healthcare financial vulnerability before and after the pandemic using Medical Expenditure Panel Survey (MEPS) data from 2019 and 2021
- High financial burden was defined as out-of-pocket healthcare expenditures exceeding 10% of family income
LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models
- Published: 2026-07-20 12:00 Beijing Time
- Abstract: - arXiv:2607.15447v1 Announce Type: new.
- Abstract: Recent research in clinical machine learning, focusing on outcome predictions in the intensive care unit (ICU), has shifted from custom supervised models to foundation models, leveraging modern representation learning methods.
- Here, foundation models are pre-trained on a mixture of complex clinical data patterns and can be used for a variety of downstream tasks.
- Existing work often utilizes Electronic Health Records (EHR) to provide a rich variety of patient observations to train clinical foundation models.
- EN Highlights:
- arXiv:2607.15447v1 Announce Type: new
- Abstract: Recent research in clinical machine learning, focusing on outcome predictions in intensive care unit (ICU), has shifted from bespoke supervised models…
Here, foundation models are pre-trained on mixtures of complex clinical data modalities, useful for various downstream tasks
Existing works often utilise Electronic Health Records (EHR) to provide rich and diverse patient observations to train clinical foundation models