System translated (Gemini)

🤖 AI 速览

Today’s main theme shifted from model capabilities to ecosystem control. Tech companies call for the protection of open-weight models, and the Anthropic settlement case also brought renewed attention to copyright, regulation, and competitive boundaries. On the product side, ChatGPT desktop …
📋 文章元数据
发布时间
2026-07-25
类型
ai-daily
字数
7749
阅读时长
37 min

2026-07-25 AI Daily | Open-Weight Models Enter the Policy Arena, Voice-Based Multi-Agent Assistants Start to Take Shape Link to heading

Today’s main theme shifts from model capabilities to ecosystem control. Tech companies are calling for the protection of open-weight models, and the Anthropic settlement case brings renewed attention to the boundaries of copyright, regulation, and competition. On the product side, ChatGPT’s desktop voice features, real-time voice in Codex, and multi-agent orchestration indicate that AI assistants are evolving from text-based tools into commandable workflow entry points.

📖 In-depth Guide to This Issue’s Watch List Link to heading

The most important reads today follow three threads. First, podcasts and Stratechery are both probing the regulatory, copyright, and competitive boundaries behind the “Open AI War” and the Anthropic 1.5B settlement. This isn’t just legal news; it’s about the realignment of power in the industry. Second, several arXiv papers focus on deconstructing MoE, hallucinations, compression, and knowledge editing. Topics like whether routing approaches Huffman coding and how to perform fixes without compromising core capabilities are worth close attention from algorithm teams. Third, agent-based AI is moving from concept to vertical implementation. Human-machine collaboration is being explored in clinical review, materials literature, and test management, but evaluation, controllability, and trustworthiness remain key hurdles.

🌐 AI Hot Topics on X Link to heading

Topic 1: Tech Giants Urge U.S. to Protect Open-Weight AI Models Link to heading

  • Category: AI · News
  • Overview: Trending time: 10 hours ago, Related posts: 111,000
  • What it is: Several major tech giants are urging the U.S. government to protect open-weight AI models and avoid policies or regulations that would restrict their release and use.
  • Why it matters: This affects whether AI models can be widely reused, fine-tuned, and deployed, directly impacting the speed of technological innovation, the openness of the ecosystem, and the competitive landscape between the U.S. and China in AI.
  • Discussion summary: Discussions on X are mainly focused on whether open-weight or closed-source models are more beneficial for innovation and safety, whether the U.S. should support open models through policy, and whether the emergence of Chinese models like Kimi K3 signifies a new shift in the U.S.-China AI model race.

Topic 2: Uncle Bob Skips AI Code Reviews with Strict Testing Gauntlet Link to heading

  • Category: AI · Other
  • Overview: Trending time: 1 day ago, Related posts: 11,000
  • What it is: Robert C. Martin (Uncle Bob) stated that he does not use AI for code reviews, instead relying on a rigorous testing pipeline to ensure code quality.
  • Why it matters: This touches on the boundaries of AI in software engineering: whether AI can replace or assist in code reviews, and the central role of automated testing in ensuring code reliability.
  • Discussion summary: The discussion on X is divided into two main camps: one side agrees with the pragmatic “tests first, AI later” approach, believing that code quality should ultimately be determined by verifiable tests; the other side argues that AI can still be used to improve review efficiency, questioning this stance as overly conservative or dismissive of AI’s assistive value.

Topic 3: Anthropic Launches Claude Opus 5 with Frontier Performance at Half the Cost Link to heading

  • Category: AI · News
  • Overview: Trending time: 6 hours ago, Related posts: 39,000
  • What it is: Anthropic has released Claude Opus 5, claiming it achieves near-frontier performance at about half the cost. It also introduces low, medium, and high inference intensity options to control task expenses.
  • Why it matters: This indicates that the competition among large models is shifting from a pure focus on peak performance to a balanced optimization of performance, cost, and controllability. This could influence budget decisions for enterprises deploying AI Agents and high-intensity inference tasks.
  • Discussion summary: Discussions on X center on whether Claude Opus 5 truly delivers “frontier performance,” whether its cost advantages can be realized in practical scenarios, and if the benchmarks are inflated for marketing. Others are focused on the practical value of the inference intensity switch for developers and enterprise users.

Topic 4: xAI Announces Grok 4.6 and 4.7 Releases Weeks Apart Link to heading

  • Category: AI · News
  • Overview: Trending time: 5 hours ago, Related posts: 5,300
  • What it is: xAI announced on X that Grok 4.6 and Grok 4.7 will be released sequentially, just weeks apart.
  • Why it matters: This reflects that large model products are entering a phase of rapid iteration. It also suggests that xAI is accelerating its efforts to catch up and close the gap with major competitors in terms of capabilities, speed, and product cadence.
  • Discussion summary: The discussion on X is focused on two points: first, whether these two updates will bring significant improvements in reasoning, code, and multimodal capabilities; second, whether such a rapid release schedule is a sign of enhanced strength or merely a marketing tactic to capture attention.

Topic 5: Etched Raises $300M at $10.3 Billion Valuation for AI Inference Chips Link to heading

  • Category: AI · News
  • Overview: Trending time: 1 day ago, Related posts: 3,900
  • What it is: AI chip startup Etched has completed a $300 million financing round, reaching a valuation of $10.3 billion, with a focus on AI inference chips.
  • Why it matters: The cost of AI inference and the supply of computing power are becoming critical bottlenecks for the commercialization of large models. Etched’s high-valuation financing shows that capital continues to bet on specialized inference chips to challenge general-purpose GPU solutions like NVIDIA’s.
  • Discussion summary: Discussions on X are focused on whether Etched’s valuation is too high, whether specialized ASICs can maintain an advantage amid rapidly changing model architectures, and whether it has a chance to shake NVIDIA’s dominant position in the inference market.

Topic 6: AI Community Awaits Opus 5 as OpenAI Rolls Out Voice on Desktop Link to heading

  • Category: AI · News
  • Overview: Trending: 2 days ago, Related posts: 18,000
  • What it is: As OpenAI launches its voice feature on desktop, the AI community is closely watching for the potential release of Anthropic’s Claude Opus 5.
  • Why it matters: This reflects how major AI companies are accelerating competition in both model capability upgrades and multimodal interactive experiences. Voice and more powerful models could further push AI assistants into daily office and productivity scenarios.
  • Discussion summary: Discussions on X are centered on whether Opus 5 will significantly surpass existing models, whether OpenAI’s desktop voice experience is practical enough, and the competitive gap between the two companies in multimodality, inference capabilities, and speed of product implementation.

Topic 7: Class of 2027 Prospects Land First Division I Offers Link to heading

  • Category: AI · Other
  • Overview: Trending:, Related posts: 437
  • What it is: On X, some are discussing Vicor Corporation ($VICR) as an under-the-radar “next-generation AI power architecture” stock, debating its opportunities in AI infrastructure.
  • Why it matters: As the computing power of AI chips continues to increase, so do the demands for power supply and cooling. Power architecture has become a critical component in the expansion of AI servers and data centers.
  • Discussion summary: The current discussion focuses on whether Vicor is undervalued by the market, whether it can benefit from upgrades in AI power supply, and whether its valuation and performance realization pace are sufficient to support a bullish outlook.

Topic 8: AI Clip of Green-Eyed Woman at Blue Jays Game Divides Opinions on Beauty Link to heading

  • Category: AI · Entertainment
  • Overview: Trending:, Related posts: 40
  • What it is: A short video, allegedly generated by AI, of a “green-eyed woman at a Blue Jays game” has spread on X, sparking discussions among users about her appearance and its authenticity.
  • Why it matters: This event reflects the increasing realism of generative AI in creating entertainment content and human images. It also highlights the impact of synthetic media on aesthetics, authenticity recognition, and platform distribution.
  • Discussion summary: Discussions on X mainly focus on whether the video is real, whether AI-generated beautiful women reinforce a single standard of beauty, and why people are attracted to virtual figures. Some also believe it’s just harmless entertainment and shouldn’t be over-analyzed.

AI Public Opinion Summary on X Today Link to heading

The main theme of today’s public opinion is that AI competition is shifting from single-point model capabilities to a full-chain battle encompassing “open ecosystems, cost efficiency, product experience, and infrastructure.” Policies on open-sourcing weights, the high-frequency iterations of Claude Opus 5 and Grok, OpenAI’s voice feature, and investments in inference chips and power architecture all indicate that AI is accelerating towards large-scale deployment. A clear consensus is emerging that inference cost, controllability, computing power, and deployment efficiency will be key in the next stage of competition. Companies and developers are no longer just looking at the highest benchmark scores but are also paying more attention to practical usability, cost, and the risk of ecosystem lock-in. Points of disagreement are centered on openness versus security, whether AI should be deeply involved in code review, whether model releases represent genuine performance breakthroughs or just marketing hype, and whether the valuations of specialized chips and AI infrastructure stocks are already overdrawn. Potential risks include innovation-security imbalances caused by excessive or insufficient regulation, investment bubbles and technological misjudgments due to overly rapid model and hardware iterations, and the further blurring of reality and fiction by generative imagery, which amplifies aesthetic homogenization and information credibility issues.

💡 Influencer Insights Link to heading

Okay, based on the tweet content from multiple AI influencers over the past 24 hours, here is an in-depth analysis report combined with recent hot topics.

A. The Comprehensive Outbreak of Multimodality, Multi-Agent Collaboration, and Voice Interaction This is the core theme today, signaling that the interaction dimension of AI assistants is evolving from singular text/code towards a more anthropomorphic and parallelized direction.

  • “Voice Control for Everything” on ChatGPT Desktop: @dotey and @vista8 both highlighted the update to the ChatGPT desktop client (formerly the Codex App). It integrates a voice mode based on a GPT-Live full-duplex architecture. The core breakthrough is that it allows you to use natural language to orchestrate multiple background Agents simultaneously (such as Codex for writing code and ChatGPT Work for running tasks), and can see the screen context via Appshots (as mentioned by @dotey). This is akin to a commander issuing voice commands to a digital team.
  • Codex’s Real-Time Voice Mode: @Pluvio9yte cited a leak claiming that Codex is about to launch Realtime Voice Mode, where the main assistant handles conversation while worker agents processes tasks like Slack, Spotify, and web browsing in the background. This aligns with the strategy for the OpenAI Chat desktop client, moving towards a “voice + multi-Agent collaboration” super personal assistant.
  • “Dictation Programming” and Stream-of-Consciousness Input: @Pluvio9yte relayed @karpathy’s unique workflow: when faced with complex ideas and not wanting to type, he switches to voice mode for a 10-minute “stream-of-consciousness” dump, letting the AI grasp the original intent. This indicates that voice is not just for commands but is also becoming a “high-bandwidth” input channel for unstructured thoughts, solving the information loss problem associated with typing.

B. The “Warring States Period” of the Programming Agent Ecosystem and Paradigm Shifts Programming Agents remain the most competitive field, and today’s discussion has shifted from model evaluation to toolchains, cost control, and a new philosophy of code supervision.

  • The Model Battle: Claude Opus 5’s “Cost-Effectiveness” Positioning: @dotey provided a detailed analysis of Anthropic’s release of Claude Opus 5. It is positioned to “provide near-Fable 5 frontier intelligence at half the price,” performing impressively across multiple benchmarks and particularly excelling at long-range tasks like autonomously building testing frameworks. This indicates that model vendors are starting to differentiate their product lines, balancing peak performance with cost.
  • Open-Source Tool Competition Heats Up:
    • @Pluvio9yte shared “This week’s top 10 trending AI open-source projects on GitHub,” the vast majority of which are Agent-related. Examples include mattpocock/skills (composable Agent engineering patterns), orca (managing multiple coding Agents in parallel), and code-review-graph (parsing codebases into knowledge graphs to reduce token consumption).
    • @AI_Jasonyu mentioned SpaceX’s open-source terminal AI coding Agent, grok-build, which is feature-complete, has 14k stars, and competes directly with tools like Claude Code.
    • @Pluvio9yte also discovered OpenCodex, which allows Codex applications to connect to other large models like Kimi, Grok, and GLM, reflecting the developer demand to avoid single-model lock-in.
  • The New “Don’t Read the Code” Programming Philosophy: @dotey quoted Uncle Bob (@unclebobmartin), author of Clean Code, who said: “I don’t read AI-written code because humans read code too slowly, which defeats the purpose of using AI.” His new method involves setting up a series of gates for the Agent (tests, quality metrics, mutation testing) to manage code quality through metrics rather than manual inspection. This signals that the core competency in programming is shifting from “reading and writing code” to “defining constraints, writing tests, and interpreting metrics.”

C. Practical Application of On-Device Models and Hardware Choices @zhixianio and @ruanyf continue to focus on on-device models, with today’s topic delving into specific applications and hardware comparisons. @ruanyf suggested that for running AI locally, a mini PC with an on-board chipset like the AMD Strix Halo is often a better choice than a high-end dedicated graphics card, thanks to its 128GB unified memory advantage. @zhixianio, on the other hand, considers Google’s Gemma 4 Quantization-Aware Training (QAT) to be an important optimization approach for on-device models that will accelerate their deployment on Android devices.

2. Notable Unique Perspectives or Industry Foresight Link to heading

  • “The Pharmaceutical Business with a Ten-Month Patent Period”: @dotey forwarded a view from @xleaps that compares the large model industry to the pharmaceutical business but with extremely short patent periods, starkly revealing the harsh reality of rapid model iterations and brief windows for monetization.
  • AI Hasn’t Brought Leisure, But Stronger Shackles: In a podcast, @vista8 reflected that many people are chained to the credit reset times of their AI programming tools, fostering a “scarcity mindset” akin to farmers being domesticated by wheat. AI was supposed to bring abundance, but instead, it has trapped people in a higher-frequency production rhythm—a profound warning about the relationship between technology and humanity.
  • Product Design from ‘Gamification’ to ‘Game Sense’: @nishuang, through a language learning App CapWords, incisively differentiated between dopamine-driven ‘gamification’ (e.g., Duolingo’s reward mechanism) and endorphin-driven ‘game sense’ (e.g., the joy of collecting in Pokémon-style games), pointing the direction for AI product interaction design.
  • Etymological Research on the Term ‘TikTok’: @ruanyf discovered that the term TikTok is actually the name of a robot in ‘The Wizard of Oz’ series of novels, and not coined by ByteDance.
  • New Professions in the AI Era: @gefei55 observed that by rapidly learning cutting-edge knowledge with AI and combining it with deliberate practice, becoming an offline conference speaker is emerging as a free, high-income side hustle.
  • Controversy over the Global Competitiveness of Chinese Open-Source Models: @vista8 mentioned in an article that multiple US AI startups jointly called for not banning Chinese models, as a ban would only protect the high pricing of US frontier models, indirectly confirming the huge competitiveness of Chinese models in terms of cost-effectiveness.

Open Source / Tools:

  • Skills and Workflows:
    • Agent Skills Collection (@mattpocock): Compiles engineering experience such as TDD and debugging into composable skill packs, recommended by @Pluvio9yte.
    • Xiangyang Qiaomu (@vista8)’s Skill Series: Includes video editing/download, frontend design, AI PRD generation, server deployment, etc., which can be installed with one click via npx skills add, highly practical.
    • Topview MCP (@TopviewAIhq): A full-stack marketing MCP integrating Amazon, YouTube, and TikTok Shop data, enabling automation from data analysis to content generation. Recommended by @AI_Jasonyu.
    • claude-tap (recommended by @seekjourney, retweeted by @dotey): A local observability platform for Claude Code and other Agents, making it convenient for developers to gain insight into the actual running process of Agents.
  • Download Tools:
    • vista8 developed a Video Account download Skill, solving the problem of difficulty in downloading Video Account content.
    • Flclash (recommended by @AI_Jasonyu): An Android open-source VPN tool based on Clash, with a user-friendly interface and card-style layout.
    • OfficeCLI: A rising star project from @Pluvio9yte’s weekly report, allowing Agents to directly read and write Word, Excel, and PPT without installing Office.

Platforms / Resources:

  • AIHOT (@Khazix0918): An AI news aggregation platform with over 600,000 monthly active users, jointly recommended by @dotey and @vista8, who called it a ‘crystallization of taste and experience’.
  • Xiaohongshu REDSkill Community: @ruanyf discovered that Xiaohongshu is building a social media-based Skill Hub that supports uploading and sharing Agent Skills, considered the Github of the Skill domain, providing a new channel for programmers to reach a massive number of C-end users.
  • API Relay Services: @ruanyf and @Pluvio9yte respectively mentioned @fennoAI and self-built relay stations, which have become a choice for many developers to solve service stability issues amidst increased risk of overseas model account bans.
  • Bolivian Exchange Rate Difference Loophole: @Pluvio9yte discovered that the sharp drop in the Bolivian exchange rate could be exploited to subscribe to Codex 20x services at extremely low prices, but this carries high risks.

📚 Appendix: Today’s Watch List Update Sources Link to heading

Time Window: Recent 3 days; Covering 22 sources; Total 32 updates

All-In Podcast (A_full) Link to heading

  • The Fight Over Open Source AI, Anthropic’s $1.5B Payout, NYC Socialists: Evictions = Violence?
    • Published: 2026-07-25 04:46 Beijing Time
    • Summary:
      • (0:00) Bestie intros.
      • (0:18) The Fight to Save Open Source AI: Kimi K3 panic, Anthropic/OpenAI regulatory capture.
      • (27:38) Anthropic/OpenAI historical growth rates, China’s protracted war.
      • (48:29) Anthropic’s $1.5B piracy settlement and massive IP theft hypocrisy.
    • EN Highlights:
      • (0:00) Bestie intros
  • (0:18) The fight to save open source AI: Kimi K3 panic, Anthropic/OpenAI regulatory capture
  • (27:38) Anthropic/OpenAI historic growth rates, China’s long game
  • (48:29) Anthropic’s $1.5B piracy settlement and the great IP theft hypocrisy

Stratechery by Ben Thompson (A_full) Link to heading

  • 2026.30: The Copium Wars
    • Published: 2026-07-25 01:00 Beijing Time
    • Abstract: - (Photo by Ng Hanguan-Pool/Getty Images).
      • Welcome back to This Week in Stratechery!
      • As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone.
      • Additionally, you have complete control over what we send to you.
      • With that said, here are some of our favorites from this week.
    • EN Key Points:
      • (Photo by Ng Han Guan-Pool/Getty Images)
      • Welcome back to This Week in Stratechery
      • As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone
      • Additionally, you have complete control over what we send to you

ArXiv cs.AI (B_intro+search) Link to heading

  • AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics

    • Published: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20452v1 Announce Type: new.
      • Abstract: Modern software quality assurance demands intelligent, autonomous systems capable of adaptive decision-making across distributed cloud environments.
      • This paper presents AINTMA (Agentic Intelligent Test Management Architecture), a multi-agent agentic AI system that transforms traditional test management into an autonomous quality intelligence ecosystem.
      • AINTMA deploys six specialized AI agents (Test Discovery, Risk Assessment, Reinforcement Learning Prioritization, Execution Orchestration, Generative Quality Intelligence, and Cloud Security Monitor), coordinated through a secure multi-agent communication framework on a cloud-native microservices infrastructure.
    • EN Key Points:
      • arXiv:2607.20452v1 Announce Type: new
      • Abstract: Modern software quality assurance demands intelligent, autonomous systems capable of adaptive decision-making across distributed cloud environments
      • This paper presents AINTMA (Agentic Intelligent Test Management Architecture), a multi-agent agentic AI system that transforms traditional test management into…
      • AINTMA deploys six specialized AI agents (Test Discovery, Risk Assessment, Reinforcement Learning Prioritization, Execution Orchestration, Generative Quality In…
  • Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Abstract:
      • arXiv:2607.20462v1 Announcement Type: New.
      • Abstract: Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need for reliable traceability of watermarked model-generated outputs.
      • However, most watermarks are evaluated on general-purpose benchmarks, leaving domains like medicine, where small token-level perturbations can result in significant semantic changes, largely unexplored.
      • In this work, we present the first rigorous study of how LLM watermarks affect medical performance, benchmarking 5 watermarking schemes across 11 LLMs and 7 VLMs, encompassing various tasks in unimodal and multimodal clinical reasoning.
    • Key Points:
      • arXiv:2607.20462v1 Announce Type: new
      • Abstract: Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need for reliable traceability of model-generated outp…
      • Yet, most watermarks are evaluated on general-purpose benchmarks, leaving domains like medicine, where small token-level perturbations can result in significant…
      • In this work, we present the first rigorous study of how LLM watermarks affect medical performance, benchmarking 5 watermarking schemes across 11 LLMs and 7 VLM…
  • ClickGuard: Detecting and Spoiling Clickbait News with Informativeness Measures and Large Language Models

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Abstract:
      • arXiv:2607.20463v1 Announcement Type: New.
      • Abstract: This paper proposes an AI-driven browser extension that identifies clickbait to help users avoid misleading internet articles.
      • The application goes beyond traditional detection, employing a hybrid machine learning architecture that combines transformer-based embeddings with linguistically-driven features and a custom ‘bait’ score.
      • After evaluating various natural language processing techniques (from classic vectorizers to large language model (LLM) embeddings), an XGBoost-based model was developed, achieving a 91% F1 score on an open composite dataset.
    • Key Points:
      • arXiv:2607.20463v1 Announce Type: new
      • Abstract: This paper presents an AI-driven browser extension that identifies clickbait to help users avoid misleading Internet articles
      • Moving beyond traditional detection, the application employs a hybrid machine learning architecture that combines transformer-based embeddings with linguistical…
      • After evaluating various natural language processing techniques – from classic vectorizers to large language model (LLM) embeddings – an XGBoost-based model w…
  • Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Summary: - arXiv:2607.20464v1 Announcement Type: New.
      • Abstract: When a language model gives different answers on repeated runs, does that variation reveal what it does not know?
      • Self-consistency turns the variation into a per-question uncertainty estimate via majority voting.
      • But does the same variation reveal cross-question structure – related questions flipping together, the way a diverse ensemble does?
    • EN Key Points:
      • arXiv:2607.20464v1 Announce Type: new
      • Abstract: When a language model gives different answers on repeated runs, does that variation reveal what it does not know
      • Self-consistency turns the variation into a per-question uncertainty estimate via majority voting
      • But does the same variation reveal cross-question structure – related questions flipping together, the way a diverse ensemble does
  • JAXBench: Benchmarking Autonomous TPU Kernel Optimization

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Summary: - arXiv:2607.20466v1 Announcement Type: New.
      • Abstract: Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equivalent target exists for TPUs.
      • We introduce JAXBench, a TPU-native benchmark suite for optimizing AI-generated kernels on Google Cloud TPUs.
      • JAXBench includes 50 JAX workloads that are both relevant and offer room for optimization.
    • EN Key Points:
      • arXiv:2607.20466v1 Announce Type: new
      • Abstract: Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equ…
      • We present JAXBench, a TPU-native benchmark suite for AI-generated kernel optimization on Google Cloud TPUs
      • JAXBench comprises 50 JAX workloads that are both relevant and provide headroom for optimization
  • DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Summary: - arXiv:2607.20467v1 Announcement Type: New.
      • Abstract: While parallel decoding is crucial for the efficiency of diffusion large language models (dLLMs), current strategies are often hindered by overly conservative confidence thresholds.
      • These thresholds, required by the Joint Probability Discrepancy Error (JPDE), lead to redundant denoising iterations and suboptimal inference speeds.
      • To overcome this problem, we propose DC-Leap, a training-free framework that can reliably accelerate dLLMs under moderate confidence.
    • EN Key Points:
      • arXiv:2607.20467v1 Announce Type: new
  • Abstract: While parallel decoding is central to the efficiency of Diffusion Large Language Models (dLLMs), current strategies are often hindered by overly conse…

    • These thresholds, necessitated by the Joint Probability Dependence Error (JPDE), result in redundant denoising iterations and suboptimal inference speeds
    • To overcome this, we propose DC-Leap, a training-free framework that enables reliable acceleration of dLLMs in the moderate-confidence regime
  • InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20468v1 Announcement Type: New.
      • Abstract: AI agents are increasingly used to automate research and development tasks, but existing benchmarks typically evaluate them on prescribed workflows or narrow action spaces.
      • Even nominally open-ended tasks can often be solved by retrieving well-known recipes and adjusting some hyperparameters, making it unclear whether strong results reflect true optimization or memorized solutions.
      • We introduce InferenceBench, where an agent must deploy an OpenAI-compatible inference server and optimize the speed of LLM inference.
    • EN 要点:
      • arXiv:2607.20468v1 Announce Type: new
      • Abstract: AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or…
      • Even nominally open-ended tasks can often be solved by retrieving a well-known recipe and tuning a few hyperparameters, making it unclear whether strong results…
      • We introduce InferenceBench, where an agent must deploy an OpenAI-compatible inference server and optimize the speed of LLM inference
  • DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20469v1 Announcement Type: New.
      • Abstract: Large Language Models (LLMs) use a set of parameters to handle many tasks, but under KV cache inference, it is unclear what task-general structure, if any, is used during decoding rather than during prefill.
      • We propose DecodeShare, a protocol that identifies a consistently shared low-dimensional subspace among tasks in the hidden states at decode time, and then tests its causal role by ablating only this subspace during decoding.
      • In our experiments, under the same intervention budget, interfering with the discovered shared subspace degrades decision performance more than interfering with prefill-derived subspaces or random subspaces.
    • EN 要点:
      • arXiv:2607.20469v1 Announce Type: new
  • Abstract: Large language models (LLMs) handle many tasks with one set of parameters, but under KV-cached inference it is unclear what task-general structure, if…

  • We propose DecodeShare, a protocol that identifies a low-dimensional subspace consistently shared across tasks in decode-time hidden states, and then tests its…

  • In our experiments, disturbing the discovered shared subspace degrades decision performance far more than disturbing either a prefill-derived or random subspace…

  • PlanE: Meta Planning of Data, Tuning, and Inference for Extractive-based LLMs

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20470v1 Announcement Type: New.
      • Abstract: Enhancing the task-specific capabilities of Large Language Models (LLMs) primarily requires substantial instruction-tuning datasets.
      • However, the sheer volume of such data imposes a considerable annotation cost, and there is a lack of optimization methods for tailoring LLMs for specific tasks.
      • To address the above issues, we propose a framework for building extractive-based LLMs called \textbf{PlanE}, which includes data decomposition, instruction tuning, and prompt inference.
    • EN Key Points:
      • arXiv:2607.20470v1 Announce Type: new
      • Abstract: Enhancing the task-specific capabilities of Large Language Models (LLMs) primarily requires substantial instruction-tuning datasets
      • However, the sheer volume of such data imposes a considerable annotation cost, and a lack of optimization methods for tailoring LLMs to specific tasks
      • To address the above issues, we propose a \textbf{Plan}ning framework for constructing \textbf{E}xtractive-based LLMs called \textbf{PlanE}, which includes data…
  • Benchmarking the Personalization Capabilities of Large Language Models

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20471v1 Announcement Type: New.
      • Abstract: Personalization is the act of altering a message to induce action in a specific recipient while keeping the sender, channel, and time fixed. It has a long tradition in psychology and marketing as a two-party problem where the sender and receiver have independent goals.
      • Large language models remove the finite inventory constraints of classic retrieval and ranking methods by generating a continuum of message variants conditioned on the inferred state of the recipient, raising the question of how effectively current models can perform personalization in the classic sense.
      • Existing LLM personalization benchmarks measure the adaptability of the sender, where the recipient is the same user being served by the model.
    • EN Key Points:
      • arXiv:2607.20471v1 Announce Type: new
  • Abstract: Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long trad…

  • Large language models remove the bounded-inventory constraint of classical retrieval-and-ranking approaches by generating a continuum of message variants condit…

  • Existing LLM personalization benchmarks measure sender-side adaptation, in which the receiver is the same user the model is serving

ArXiv cs.CL (B_intro+search) Link to heading

  • What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces

    • Release Time: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20425v1 Announcement Type: New.
      • Abstract: What makes writing “good” remains a long-standing problem in literary studies and computational linguistics.
      • We present a two-study investigation into how reasoning-enabled LLMs evaluate literary quality.
      • In Study 1, we construct a benchmark of 30 real texts spanning six quality tiers (from canonical literature to anonymous forum posts) and extract the model’s implicit theories of quality from its reasoning traces.
    • EN Highlights:
      • arXiv:2607.20425v1 Announce Type: new
      • Abstract: What makes writing “good” remains a persistent question in literary studies and computational linguistics
      • We present a two-study investigation of how reasoning-enabled LLMs evaluate literary quality
      • In Study 1, we construct a benchmark of 30 real texts spanning six quality tiers, from canonical literature to anonymous forum posts, and extract the model’s im…
  • Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs’Hallucinations

    • Release Time: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20426v1 Announcement Type: New.
      • Abstract: Existing LLM hallucination mitigation methods, including prompt engineering and model optimization, either find it difficult to alter the model’s internal knowledge or have poor cross-domain generalization.
      • Contrastive decoding mitigates hallucinations by using layer-wise differences in LLMs.
      • However, previous research has only explored Transformer-based models (e.g., GPT), neglecting other effective frameworks like Mixture-of-Experts (MoE) models.
    • EN Highlights:
      • arXiv:2607.20426v1 Announce Type: new
      • Abstract: Existing LLM hallucination mitigation methods, including prompt engineering and model optimization, either hardly alter models’internal knowledge or h…
      • Contrastive decoding mitigates hallucinations by using layer-wise differences in LLMs
  • However, prior studies only explore transformer-based models (e.g., GPT), ignoring other effective frameworks like mixture-of-experts (MoE) models

  • Is MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20427v1 Announcement Type: New.
      • Abstract: Mixture-of-Experts architectures have revolutionized scaling, yet the underlying logic of their routing remains a black box.
      • In this paper, we uncover a fundamental governing principle: MoE routing is not merely selection, but a manifestation of Huffman Coding.
      • We introduce the Frequency-Diversity Law, revealing that state-of-the-art models, such as Phi-3.5-MoE and Gemma-4-27B-A4B, spontaneously act as information-theoretic engines.
    • EN Key Points:
      • arXiv:2607.20427v1 Announce Type: new
      • Abstract: Mixture-of-Experts architectures have revolutionized scaling, yet the underlying logic of their routing remains a black box
      • In this paper, we uncover a fundamental governing principle: MoE routing is not merely selection, but a manifestation of Huffman Coding
      • We introduce the Frequency-Diversity Law, revealing that state-of-the-art models, such as Phi-3.5-MoE and Gemma-4-27B-A4B, spontaneously act as information-theo…
  • Human-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20428v1 Announcement Type: New.
      • Abstract: This study evaluated a retrieval-augmented, multi-agent large language model (LLM)-driven, human-in-the-loop framework for detecting cutaneous immune-related adverse events (cirAE) from clinical records.
      • Compared with unassisted manual review, the LLM-assisted workflow improved accuracy (F1 = 0.88 vs 0.77), inter-rater agreement as measured by Cohen’s kappa (kappa = 0.82 vs 0.50), and reduced the average review time by approximately half.
      • This framework pilots the application of LLMs to identify immune-related toxicities across organ systems and, more broadly, to achieve accurate, scalable, and transparent extraction of adverse event data.
    • EN Key Points:
      • arXiv:2607.20428v1 Announce Type: new
      • Abstract: This study evaluated a retrieval-augmented, multi-agent large language model (LLM)-driven, human-in-the-loop framework for detecting cutaneous immune-…
      • Compared with unassisted manual review, the LLM-assisted workflow improved accuracy (F1 = 0.88 vs 0.77), inter-rater agreement measured by Cohen’s kappa (kappa…
  • This framework pilots how LLMs can be applied to identify immune-related toxicities across organ systems and, more broadly, enable accurate, scalable, and trans…

  • More Is Not More: What Matters for Diversity in LLM Opinions?

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20429v1 Announcement Type: new.
      • Abstract: Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group modeling, and public opinion forecasting.
      • However, LLM outputs exhibit systematic opinion homogenization.
      • Practitioners have explored various interventions to increase diversity, but the landscape remains fragmented: different methods are evaluated in isolation with incomparable metrics, and in practice they are often deployed and scaled concurrently, making it difficult to attribute gains to specific components.
    • EN Key Points:
      • arXiv:2607.20429v1 Announce Type: new
      • Abstract: Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group modeling, an…
      • However, LLM outputs exhibit systematic opinion homogenization
      • Practitioners have explored various interventions to increase diversity, but the landscape remains fragmented: different methods are evaluated in isolation with…
  • LLM-INSTRUCT at UZH Shared Task 2026: Constraint-Aware Retrieval and Selective Debate for Paragraph-Level Argument Mining

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20430v1 Announcement Type: new.
      • Abstract: We introduce LLM-INSTRUCT, the winning system for the UZH Shared Task at ArgMining 2026, which involves paragraph-level argument mining in UN and UNESCO resolutions.
      • The task requires paragraph-type classification, prediction of a subset of 141 official tags, and directed relation prediction under a strict JSON schema setting using only open-weight models with up to 8B parameters.
      • We define the task as constrained structured prediction.
    • EN Key Points:
      • arXiv:2607.20430v1 Announce Type: new
      • Abstract: We present LLM-INSTRUCT, the winning system for the UZH Shared Task at ArgMining 2026 on paragraph-level argument mining in UN and UNESCO resolutions
      • The task requires paragraph-type classification, prediction of a subset of 141 official tags, and directed relation prediction under a strict JSON schema settin…
      • We frame the task as constrained structured prediction
  • Skill-Contracted Agents for Evidence-Aware Materials Literature Analysis

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20431v1 Announcement Type: new.
  • Abstract: Materials science literature analysis requires simultaneous attention to composition, processing, characterization, and property relationships, but traditional retrieval-augmented generation pipelines struggle to coordinate heterogeneous tasks in a single retrieve-then-generate architecture.

  • Here, we introduce AlphaAgent, a skill-driven agent framework that decouples retrieval-based question answering from paper-level report generation through an explicit skill contract.

  • A dedicated retrieval skill rewrites user requests into material-specific search intents, queries a curated index of over 300,000 papers from the Journal Citation Reports metallurgy and metallurgical engineering category, and reformulates queries when initial evidence is insufficient.

  • EN Highlights:

    • arXiv:2607.20431v1 Announce Type: new
    • Abstract: Materials science literature analysis requires simultaneous attention to composition, processing, characterization, and property relationships, yet co…
    • Here we present AlphaAgent, a skill-driven agent framework that decouples retrieval-based question answering from paper-level report generation through explicit…
    • A dedicated retrieval skill rewrites user requests into material-specific search intents, queries a curated index of more than 300,000 papers from the Journal C…
  • Position: Natural Language Should Not Fully Replace Formal Languages

    • Published: 2026-07-24 12:00 Beijing Time
    • Abstract:- arXiv:2607.20432v1 Announce Type: new.
      • Abstract: Recent advances in large language models and their widespread adoption have prompted claims that natural language could entirely replace formal languages, such as programming languages for software design.
      • In this position paper, we argue that this perspective overlooks fundamental linguistic properties of natural language, specifically that it is optimized for irregularities in open-ended contexts.
      • We introduce a formal framework centered on “task specificity,” defining it as the information-theoretic reduction of uncertainty in an output space (e.g., all possible images) according to the user’s specific requirements.
    • EN Highlights:
      • arXiv:2607.20432v1 Announce Type: new
      • Abstract: Recent advances in large language models and their widespread adoption have prompted claims that natural language could entirely replace formal langua…
      • In this position paper, we argue that this perspective overlooks fundamental linguistic properties of natural language, specifically that it is optimized for un…
      • We introduce a formal framework centered on task specificity, defining it as the information-theoretic reduction of uncertainty in an output space – such as…
  • Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing

    • Published: 2026-07-24 12:00 Beijing Time
    • Abstract:- arXiv:2607.20433v1 Announce Type: new.
      • Abstract: While language models remain in a static training state, the world continues to evolve.
      • Knowledge editing has emerged as a key alternative to full retraining, but its deployment is bottlenecked by the erosion of core capabilities: mathematical and procedural reasoning collapse, while encyclopedic recall remains intact.
  • We trace this asymmetric degradation to a distributional mismatch.

    • EN 要点:
      • arXiv:2607.20433v1 Announce Type: new
      • Abstract: While language models remain frozen at their training state, the world evolves continuously
      • Knowledge editing has emerged as a key alternative to full retraining, but its deployment is bottlenecked by the erosion of core capabilities: mathematical and…
      • We trace this asymmetric degradation to a distributional mismatch
  • Break Through the Compression Bottleneck: From Theory to Practice

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Abstract:- arXiv:2607.20434v1 Announce Type: new.
      • Abstract: As the parameter size of language models continues to grow, effective model compression is required to reduce their computational and memory overhead.
      • Existing compression methods suffer from bottleneck issues: when the compression ratio is increased, performance degrades significantly.
      • Low-rank decomposition and quantization are two important compression methods that have been proven to significantly reduce the computational and memory requirements of large language models (LLMs) while maintaining model accuracy.
    • EN 要点:
      • arXiv:2607.20434v1 Announce Type: new
      • Abstract: As the parameter size of language models continues to grow, effective model compression is required to reduce their computational and memory overhead
      • Existing compression methods suffer from bottleneck issues: when the compression ratio is increased, performance degrades significantly
      • Low-rank decomposition and quantization are two prominent compression methods that have been proven to significantly reduce the computational and memory require…

ArXiv cs.LG (B_intro+search) Link to heading

  • DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Abstract:- arXiv:2607.20465v1 Announce Type: new.
      • Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), but there is no unified benchmark to measure how LLMs, agents, and data-centric workflows actually prepare training data end-to-end.
      • We believe LLM-driven data preparation includes two complementary capabilities: data construction, which transforms raw sources into supervised training data, and data quality evaluation, which predicts the training value of candidate datasets before downstream training; throughout, “quality” refers to downstream training utility, not superficial text attributes.
      • We introduce DataPrep-Bench, the first unified benchmark to jointly evaluate these two capabilities under a shared downstream foundation protocol across six domains and multiple foundation models.
    • EN 要点:
      • arXiv:2607.20465v1 Announce Type: new
      • Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how…
  • We view LLM-driven data preparation as comprising two complementary capabilities: data construction, which transforms raw sources into supervised training data,…

  • We introduce DataPrep-Bench, the first unified benchmark that jointly evaluates both capabilities under a shared downstream-grounded protocol over six domains a…

  • PhantomFill: When the Form Demands an Answer, Language Models Invent One

    • Published: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20492v1 Announce Type: new.
      • Abstract: Language models in production do not write prose.
      • They fill forms: JSON fields, function arguments, extraction templates.
      • We show that the form itself causes hallucination.
    • EN Key Points:
      • arXiv:2607.20492v1 Announce Type: new
      • Abstract: Language models in production do not write prose
      • They fill forms: JSON fields, function arguments, extraction templates
      • We show that the form itself causes hallucination
  • The Active Ingredient in Muon’s Grokking

    • Published: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20512v1 Announce Type: new.
      • Abstract: The Muon optimizer reaches the grokking threshold on modular arithmetic faster than AdamW.
      • Prior work attributes this to “spectral-norm constraints plus orthogonalized momentum” but does not isolate which mechanism matters.
      • To better understand Moun’s behavior, we run multi-seed and multi-learning-rate sweeps to decompose and stress-test the effect.
    • EN Key Points:
      • arXiv:2607.20512v1 Announce Type: new
      • Abstract: The Muon optimizer reaches the grokking threshold on modular arithmetic faster than AdamW
      • Prior work attributes this to “spectral-norm constraints plus orthogonalized momentum” but does not isolate which mechanism matters
      • To better understand Moun’s behavior, we run multi-seed and multi-learning-rate sweeps to decompose and stress-test the effect
  • Scaling Closed-Loop Feature Channel Configuration with LLMs

    • Published: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20516v1 Announce Type: new.
      • Abstract: Initial results from closed-loop large language model-based channel configuration search have shown that neural network width can be directly optimized via executable code generation and accuracy feedback.
      • However, these results were obtained from a relatively sparse set of valid evaluations, and it remains unknown whether the observed optimization behavior transfers to a more dense sampling regime and if additional architectural regularities emerge when more generated networks are evaluated.
      • To test this, the same search setup is extended to 250 candidate networks per fine-tuning cycle.
    • EN Key Points:
      • arXiv:2607.20516v1 Announce Type: new
  • Abstract: Promising initial results in closed-loop large-language-model-based channel-configuration search demonstrated that neural-network widths can be optimi…

  • However, those results were obtained from a relatively sparse set of valid evaluations, leaving open whether the observed optimization behavior transfers to a d…

  • To test this, the same search setting is scaled to 250 candidate networks per fine-tuning cycle

  • Multimodal CoLRAG-TF: Triple-Filtered Retrieval for Complex PDFs

    • Published: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20517v1 Announcement Type: new.
      • Abstract: Retrieval-augmented generation (RAG) over heterogeneous PDF collections remains challenging due to multimodal content, domain-specific terminology, and the need for multi-hop reasoning across disparate evidence.
      • We propose Multimodal CoLRAG-TF, a four-axis fusion architecture that integrates dense text embeddings, BM25 keyword matching, knowledge graph triple filtering, and image-based similarity for robust retrieval on complex documents.
      • Our system builds a multimodal index of 2,403 chunks extracted from 43 Japanese disaster lesson PDFs, supported by a hybrid OCR pipeline and LLM-based caption generation.
    • EN Key Points:
      • arXiv:2607.20517v1 Announce Type: new
      • Abstract: Retrieval-augmented generation (RAG) over heterogeneous PDF collections remains challenging due to multimodal content, domain-specific terminology, an…
      • We present Multimodal CoLRAG-TF, a four-axis fusion architecture that integrates dense text embeddings, BM25 keyword matching, knowledge-graph triple filtering,…
      • Our system constructs a multimodal index of 2,403 blocks extracted from 43 Japanese disaster lesson PDFs, supported by a hybrid OCR pipeline and LLM-based capti…
  • Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts

    • Published: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20519v1 Announcement Type: new.
      • Abstract: Looped Transformers increase test-time computation by repeatedly applying a shared recurrent block.
      • The learned halting objective in a Looped Transformer typically uses a single exit distribution to both serve as an inference-time stopping rule and a training-time weighting over per-depth losses.
      • This entangles exit selection with trajectory shaping: the gate not only chooses which recurrent state to use, but also determines the supervision strength on each intermediate state.
    • EN Key Points:
      • arXiv:2607.20519v1 Announce Type: new
      • Abstract: Looped Transformers increase test-time computation by repeatedly applying a shared recurrent block
  • Learned halting objectives in looped Transformers typically use a single exit distribution both as the inference-time stopping rule and as the training-time wei…

  • This entangles exit selection with trajectory formation: the gate not only chooses which recurrent state to use, but also determines how strongly each intermedi…

  • Generative Bayesian Filtering for State Estimation

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20521v1 Announcement Type: New.
      • Abstract: The state of a dynamic system evolves over time, switching among several latent modes that govern its observable behavior.
      • Filtering methods infer the latent state from observations.
      • Classical filtering approaches, including Kalman filters, typically rely on simple observation models, such as linear-Gaussian models, that are incapable of characterizing the increasingly nonlinear and heterogeneous patterns in high-dimensional sensor signals.
    • EN Highlights:
      • arXiv:2607.20521v1 Announce Type: new
      • Abstract: The state of a dynamic system evolves over time, switching among several latent modes that govern its observable behavior
      • Filtering methods infer the latent state from observations
      • Classical filtering approaches, including Kalman filters, typically rely on simple observation models, such as linear-Gaussian models, that are incapable of cha…
  • Do Active SAE Feature Planes Carry More Holonomy? A Preregistered Reversal in Gemma

    • Publication Time: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20522v1 Announcement Type: New.
      • Abstract: This paper tests whether holonomy concentrates on active sparse-autoencoder (SAE) feature planes in Gemma 2 2B, a concrete operationalization of the broader semantic concentration prediction.
      • Holonomy is measured at the final-token layer-12 to layer-13 residual-stream readout by carrying a local frame around small loops using the instrument’s restricted Jacobian transport rule, then standardizing the resulting rotation by the enclosed area.
      • The design, materiality threshold, analysis, and verdict rules were preregistered and frozen before the analysed measurements were inspected.
    • EN Highlights:
      • arXiv:2607.20522v1 Announce Type: new
      • Abstract: This paper tests whether holonomy concentrates on active sparse-autoencoder (SAE) feature planes in Gemma 2 2B, a concrete operationalization of the b…
      • Holonomy is measured at the final-token layer-12 to layer-13 residual-stream readout by carrying a local frame around small loops using the instrument’s restric…
      • The design, materiality threshold, analysis, and verdict rules were preregistered and frozen before the analysed measurements were inspected
  • Uncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement

    • Published: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20529v1 Announcement Type: new.
      • Abstract: Large Language Model (LLM) ensembles are increasingly used to improve reliability by combining predictions from multiple LLMs.
      • However, existing aggregation methods typically assume that all models are equally trustworthy, overlooking differences in uncertainty quality.
      • This assumption is poorly suited to heterogeneous LLMs, whose reliability and capability vary significantly, making naive aggregation vulnerable to unreliable or adversarial experts.
    • EN Key Points:
      • arXiv:2607.20529v1 Announce Type: new
      • Abstract: Large Language Model (LLM) ensembles are increasingly used to improve reliability by combining predictions from multiple LLMs
      • However, existing aggregation methods typically assume that all models are equally trustworthy, overlooking differences in uncertainty quality
      • This assumption is poorly suited to heterogeneous LLMs, whose reliability and capability vary significantly, making naive aggregation vulnerable to unreliable o…
  • CLOE: Christoffel Loss Autoencoder for Anomaly Detection

    • Published: 2026-07-24 12:00 Beijing Time
    • Abstract: - arXiv:2607.20530v1 Announcement Type: new.
      • Abstract: Semi-supervised anomaly detection plays a key role in diverse fields such as process monitoring, healthcare, and finance.
      • However, lightweight methods often struggle with high-dimensional data and typically require careful tuning of multiple hyperparameters.
      • Among existing approaches, Christoffel Function–based methods are attractive due to their simplicity, requiring at most a single hyperparameter.
    • EN Key Points:
      • arXiv:2607.20530v1 Announce Type: new
      • Abstract: Semi-supervised anomaly detection plays a key role in diverse fields such as process monitoring, healthcare, and finance
      • However, lightweight methods often struggle with high-dimensional data and typically require careful tuning of multiple hyperparameters
      • Among existing approaches, Christoffel Function–based methods are attractive due to their simplicity, requiring at most a single hyperparameter