System translated (Gemini)

🤖 AI 速览

Open-source inference engines like vLLM are being redefined as critical foundational components of the AI stack; Jeff Dean’s departure from Google also indicates that top talent and research resources are flowing towards independent organizations. Meanwhile, as models enter high-risk scenarios …
📋 文章元数据
发布时间
2026-08-07
类型
ai-daily
字数
7917
阅读时长
38 min

2026-08-07 AI Daily | After Jeff Dean’s Departure, AI Competition Shifts to Inference Infrastructure Link to heading

Open-source inference engines like vLLM are being redefined as a key foundation of the AI stack; Jeff Dean’s departure from Google also signals that top talent and research resources are flowing to independent organizations. Meanwhile, as models enter high-risk scenarios such as healthcare, adolescent psychology, and law, the importance of evaluation, governance, and trustworthy deployment is clearly increasing.

📖 In-Depth Guide to This Issue’s Watch List Link to heading

The top priority today is the trend of “agents moving toward personal devices and becoming infrastructure.” The YC interview with the founder of OpenClaw and the in-depth discussion on vLLM as an open-source inference engine can be read together. The former focuses on how personal AI assistants can truly take over workflows like email, calendars, and files, while the latter explains why the inference layer behind these applications is becoming a key piece of infrastructure in the AI stack.

The second main theme is “evaluation and governance as model capabilities enter high-risk scenarios.” OpenAI has updated ChatGPT and is collaborating with the APA on adolescent mental health. Multiple papers offer warnings from the perspectives of medical safety, privacy filtering, legal benchmarks, and the reproducibility of LLM-as-judge systems, reminding us that preference, higher scores, or a correct final answer do not equate to safety and reliability.

Additionally, Google DeepMind’s WeatherNext for cyclone prediction, along with new benchmarks for cuneiform and multimodal tumor analysis, are worth observing as examples of “AI penetrating specialized domains.” The signal today is clear: the next phase of competition lies not just in the models themselves, but in trustworthy deployment, evaluation frameworks, and integration with real-world workflows.

🌐 AI Hot Topics on X Link to heading

Topic 1: Google’s Jeff Dean Leaves After 27 Years to Co-Found AI Research Venture Link to heading

  • Category: AI · News
  • Overview: Trending for: 1 day ago, Related posts: 43,000
  • What it is: Veteran Google computer scientist Jeff Dean is leaving after 27 years to co-found a new AI research organization.
  • Why it’s important: Jeff Dean has long been involved in building Google’s search, distributed systems, and deep learning infrastructure. His departure is seen as a sign of top AI talent and research resources moving from large tech companies to new types of research organizations.
  • Discussion summary: Discussions on X are focused on whether this signals an acceleration of AI talent drain from Google, the potential research directions for the new organization, and whether independent AI research labs can compete with tech giants in terms of computing power, funding, and model capabilities.

Topic 2: Meta AI Models Win Gold in Five Major STEM Olympiads Link to heading

  • Category: AI · News
  • Overview: Trending for: 7 hours ago, Related posts: 670
  • What it is: Meta announced that its AI models have achieved gold medal-level performance in five major STEM Olympiads, showcasing strong abilities in mathematics, reasoning, and complex problem-solving.
  • Why it’s important: This shows that large models are progressing from general conversational skills to tackling difficult academic reasoning tasks, which is important for measuring the true reasoning level of AI, for educational applications, and for their future potential as research aids.
  • Discussion summary: Discussions on X focus on two main points: whether this demonstrates a substantial improvement in model reasoning abilities, and how credible these “gold medal” results are, considering potential influences from leaked questions or evaluation methods. It is also seen by some as a sign that Meta is catching up to the frontrunners in the AI competition.

Topic 3: Garry Tan Pushes Personal AGI for Solo Founders at Startup School 2026 Link to heading

  • Category: AI · News
  • Overview: Trending for: , Related posts: 70
  • What it is: At Startup School 2026, Garry Tan advocated for “Personal AGI,” encouraging solo founders to leverage AI to build startups with smaller teams or even alone.
  • Why it’s important: This reflects AI’s shift from being a tool to a force multiplier for entrepreneurs, potentially reshaping the organizational structure, funding logic, and product development efficiency of startups.
  • Discussion summary: Discussions on X center on whether Personal AGI can truly make solo entrepreneurship mainstream, if it will lower barriers to entry and accelerate innovation, and its feasibility, potential for hype, and impact on traditional team collaboration.

Topic 4: Meta Launches Muse Code Beta for Complex Coding Tasks Link to heading

  • Category: AI · News
  • Overview: Trending for: 1 day ago, Related posts: 15,000
  • What it is: Meta has released Muse Code Beta, a terminal-based AI coding agent powered by Muse Spark 1.2, designed for complex, multi-file, and long-duration software engineering tasks.
  • Why it’s important: This indicates AI is evolving from code completion to executable engineering collaboration, capable of planning, modifying, and verifying work within large codebases. This will impact software development workflows, efficiency, and the human-machine division of labor. Discussion Overview: The discussions on X primarily focused on its multi-agent architecture, whether it can truly surpass existing coding models, its stability for long tasks, and practical production usability. Some also focused on pricing, installation methods, and whether it serves to replace or assist developer work.

Today’s AI Public Opinion Summary on X Link to heading

Today’s main theme on X is that AI is moving from “can chat” to “can do things, can start businesses, can research.” On one hand, Jeff Dean’s departure sparked concerns about top talent and research resources flowing out of major tech companies. On the other hand, Meta’s Olympic gold medal, Muse Code, and “personal AGI” are all reinforcing the narrative that “models are approaching real working capabilities.” A relatively consistent consensus is that AI’s capabilities are indeed expanding, especially in reasoning, software engineering, and entrepreneurial efficiency, already beginning to change organizational forms and production methods. Disagreements mainly center on how “real” these achievements are—for example, whether the gold medal is credible, whether coding agents are stable enough, whether solo entrepreneurship will truly become mainstream, and whether independent AI research institutions can compete head-on with giants in terms of computing power and funding. The potential risk lies in the possibility that external expectations for capability leaps might be too rapid; if evaluations and demonstrations are over-packaged, it could amplify a bubble. Simultaneously, talent mobility, automated development, and the solo entrepreneurship boom could also concentrate the industry on a few resource providers, increasing uncertainty regarding technological out-of-control scenarios and job displacement.

💡 Influencer Insights Link to heading

Alright, based on a summary of tweets from several AI influencers on X over the past 24 hours, here are my deep insights and conclusions as a senior AI industry analyst.


Today’s core focus is highly concentrated on model capability evolution, the implementation forms of AI Agents, and breakthroughs in on-device intelligence.

  • DeepSeek V4 Official Release Triggers a “Performance-Cost” Storm: Multiple bloggers are eagerly discussing the official release of DeepSeek V4.

    • ** @zhixianio** personally tested the 4-bit quantized version of DeepSeek V4 Flash on a Mac Studio, demonstrating its feasibility for local operation.
    • ** @vista8** cited data and funding news, pointing out that DeepSeek V4 Pro’s performance has “entered the global top tier,” with its programming capability only 0.3% lower than Claude’s flagship model, but its API pricing is only 1/10 to 1/100 of overseas competitors, creating a “decisive lead” in cost-effectiveness.
    • Meanwhile, ** @Pluvio9yte** observed that DeepSeek’s official team has not yet launched its self-developed CLI programming tools and recommended Reasonix (which has garnered 32k Stars), an open-source Agent integration tool mentioned in DeepSeek’s official documentation and deeply optimized for its prefix-caching mechanism. This reflects the rapid maturation of the developer ecosystem around DeepSeek.
  • AI Agent Enters a New Stage of “Desktop Integration” and “Framework Convergence”:

    • Diverse Agent Frameworks Flourish: ** @vista8** reviewed an Agent framework named bb, whose highlight is its ability to automatically identify and invoke various local programming tools like Codex CLI and Claude Code CLI, realizing a “framework as platform” product philosophy. ** @dotey** focused on Hermes Desktop supporting a built-in browser, allowing Agents to perceive and operate the Web environment more directly.
    • From Terminal to Desktop: ** @dotey** observed an industry consensus that the most powerful Agents require “their own computer” and not just a container. This has driven tools like Claude Code from the command line to desktop clients, and products like Workbuddy have leveraged this to gain significant market share.
    • ByteDance’s SeedRealtime Redefines Multimodal Interaction: ** @vista8** thoroughly experienced and interpreted ByteDance’s SeedRealtime model, emphasizing its achievement of native audio-video full-duplex interaction. The model can not only listen and see but also actively perceive, decide, and interact within continuous audio-video streams, for example, proactively reminding users of interesting exhibits when visiting a museum. ** @vista8** considers this a breakthrough beyond GPT-4o’s audio-only full-duplex capability, potentially influencing the development speed of embodied robots.
  • On-device Model Deployment: Pushing Hardware to the Limit, Narrative Constantly Rewritten: The ability to run large models on-device is breaking through at an unexpected pace.

    • ** @Pluvio9yte** shared a project called Swiftlet, which packs an 80B parameter Qwen model into a Mac with 4.3GB of memory, even claiming to run a 35B model on an iPhone. Its core technology is “on-demand streaming loading” of MoE expert weights, providing a brand-new path for running ultra-large models on consumer-grade hardware.
  • @zhixianio concluded through practical tests that in local code generation tasks, a Qwen3.6-35B-A3B MoE model deployed on an M5 Max significantly outperforms the latest Gemma 4 12B Coder in its ability to generate complex programs (like Tetris). He pointed out that the “ceiling” imposed by the number of model parameters is more critical than fine-tuning techniques, as a 12B model cannot handle complex programs that are “long, stateful, and generated in a single pass.”

2. Noteworthy Unique Perspectives or Industry Foresight Link to heading

  • The Future Division of Labor in Vibe Coding: Coexistence of Experts and “Full-Stack Commoners” (from @dotey) @dotey presented a forward-looking view: in the future of programming, “professionals will not only have to put out fires and clean up the messes caused by Vibe Coding, but they will also need to build the right infrastructure to allow everyone to Vibe efficiently and safely.” This implies that a few specialized front-end and back-end positions will be retained to manage underlying architecture and security, while a large number of mid-level “full-stack” roles, often filled by operations or product personnel, will use AI for implementation. This differs from the current “everyone is a developer” narrative, instead emphasizing the crucial role of professionals as “infrastructure builders” and “gatekeepers” in the AI era.

  • The “AI Feel” Paradox in AI Writing (from @dotey) @dotey offered a deep reflection on the popular demand for “Skills to eliminate the AI feel.” He argues that it’s a contradiction that people naturally dislike articles written by AI, yet they want to use AI to write articles that don’t feel like they were written by AI. He admits he has given up on finding the perfect “de-AI-ify Skill” because “AI cannot replace human writing.” He positions his methodology as “heavily using AI to assist writing”: using AI to help gather information, adjust structure, and provide inspiration when “stuck,” with the core expression ultimately completed by a human. This represents a more pragmatic shift from viewing AI as a “replacement” to a “collaborator.”

  • Large Model Companies’ “New Marketing Battlefield”: Competing on Who Can Hack Others’ Systems (from @vista8) @vista8 humorously observed that several large model companies are starting to use “our model can hack into others’ systems” as a marketing highlight. This may signal that AI security issues are evolving from purely technical challenges into a public competition with a PR spin.

  • The Rise of the “Assembly Plant” Model in the Compute Market (from @Pluvio9yte) @Pluvio9yte keenly noted that Anthropic signed a compute deal worth approximately $10 billion with Volta, a cloud startup that is only a few months old and rents almost all its hardware. He believes this marks a shift where compute power is no longer just about self-building or waiting in line for major cloud providers. An “assembly plant” model is emerging, where “whoever can assemble power and GPUs first gets the cutting-edge mega-deal,” which will reshape the supply landscape for AI infrastructure.

  • AI Programming and Development Tools

    • Reasonix: An open-source agent programming framework optimized for DeepSeek models, particularly suited for leveraging its prefix caching feature to reduce costs in long sessions. It has over 32k stars. (@Pluvio9yte)
    • bb Agent Framework: An innovative Agent IDE that automatically identifies and integrates various locally installed CLI programming tools (like Codex, Claude Code, PI Agent, etc.) for one-stop management. (@vista8)
    • CodeBuddy NPC (Tencent Cloud): A product that packages an AI model into a game NPC concept. You can assign tasks to the “NPC” on code hosting platforms using natural language, introducing a new dynamic to programming education or automation. (@ruanyf)
    • QLMarkdown: A small Mac utility that allows you to quickly preview Markdown files with the spacebar after installation, boosting development efficiency. (@vista8)
  • AI-Native Applications and Services

    • @makeplayai: A free AI game creation platform where you can generate a complete mini-game with art, sound, and animation from a single sentence prompt. It supports branching development to compare different gameplay mechanics. (@Pluvio9yte)
    • Xiaohongshu REDSkill Community: An AI Skill hosting and sharing feature launched by Xiaohongshu. It allows users to upload Skill files to their posts, which others can then install with one click. @ruanyf considers this the world’s first attempt to combine a social media platform with a Skill Hub, offering developers a new channel to reach a massive user base.
  • AI-Assisted Writing and Productivity

    • mattpocock/skills: Includes practical features like /hand off that can generate handover documents, enabling a seamless transition of context for AI programming between old and new sessions. (@Pluvio9yte)
  • Human-like Writing .skill (By @Khazix0918): An open-source Skill designed to help users write text without an “AI feel.” Recommended despite some controversy. ( @Pluvio9yte)

  • qiaomu-campus-resume (By @vista8): A Skill that generates and optimizes PDF resumes using an interview-based model. It incorporates best practices from the career centers of top universities like Tsinghua and MIT, making it ideal for university students on their job hunt. ( @vista8)

  • AI Safety and Infrastructure

    • OpenConnector: An open-source credential connection gateway designed to prevent AI Agents from leaking credentials into their context. It centralizes connection authorization for all external applications, allowing Agents to receive only the execution results without ever accessing the credentials themselves. ( @ruanyf)

📚 Appendix: Today’s Watch List Source Updates Link to heading

Timeframe: Last 3 days; covers 22 sources; 36 updates in total.

a16z Podcast (A_full) Link to heading

  • Inside vLLM: The Engine Powering Open-Source AI
    • Publication Time: 2026-08-06 18:00 Beijing Time
    • Summary: - Elena Burger and Matt Bornstein are joined by Simon Mo, co-founder and CEO of Inferact, the open-source inference engine powering many of today’s most advanced AI applications.
      • Together, they explore how open-source AI evolved from a research project into critical infrastructure, why inference has become one of the most important layers of the AI stack, and how to bring cutting-edge intelligence to developers everywhere.
      • The conversation covers the origins of vLLM, the rise of open-weight models, why companies increasingly want control over their AI infrastructure, and how open-source inference can support the next generation of AI applications.
      • They also discuss model licensing, the economics of open-weight AI, Kimi K3, distillation, AI infrastructure, and why Simon believes the gap between open and closed models is rapidly closing.
      • See everything a16z is doing with AI, including articles, projects, and more podcasts, here.
    • EN Highlights:
      • Elena Burger and Matt Bornstein are joined by Simon Mo, co-founder and CEO of Inferact, the open-source inference engine powering many of today’s most advanced…
      • Together, they explore how open-source AI evolved from a research project into critical infrastructure, why inference has become one of the most important layer…
      • The conversation covers vLLM’s origins, the rise of open-weight models, why companies increasingly want control over their AI infrastructure, and how open-sourc…
      • They also discuss model licensing, the economics of open-weight AI, Kimi K3, distillation, AI infrastructure, and why Simon believes the gap between open and cl…

Y Combinator Podcast (B_intro+search) Link to heading

  • Garry Tan: Own Your Intelligence
    • Publication Time: 2026-08-07 03:28 Beijing Time
    • Summary: - You’ve probably heard of OpenClaw (formerly Clawdbot/Moltbot).
      • The sensational open-source AI assistant that runs on your own devices, connects with the messaging apps you already use, and goes beyond chat to actually do things like manage your email, calendar, files, workflows, and more.
      • Now, meet the person behind it.
  • YC’s Raphael Schaad sits down with OpenClaw founder Peter Steinberger to discuss the “aha” moment behind viral personal AI agents, why local-first agents could replace many of today’s apps, and how personal agents will reshape the future of software.
    • EN Highlights:
      • The next generation of startups will be built by smaller teams than ever before
      • At Startup School 2026, YC President & CEO Garry Tan explains why we’re entering the era of personal AGI: AI agents that run on your own infrastructure, compoun…
      • He shares the tools and workflows he uses every day, why every founder should own their intelligence instead of renting it, and what it means to build under you…

OpenAI Blog (A_full) Link to heading

  • Improving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free users

    • Published: 2026-08-06 18:00 Beijing Time
    • Summary:
      • Our mission is to ensure that artificial general intelligence benefits all of humanity.
      • We are rolling out updates to ChatGPT to improve everyday conversations while expanding access for free users.
      • For Plus and Pro users, we are updating GPT-5.6 Sol in chat to be more reliable with facts and provide more targeted answers.
      • A new slider lets you choose how deeply ChatGPT thinks for each response.
      • For free users, we are updating the default model to GPT-5.6 Luna and expanding access with unlimited text chats.
    • EN Highlights:
      • ChatGPT introduces improved GPT-5.6 Sol with better accuracy and consistency, plus expanded access for free users and unlimited everyday chats with GPT-5.6 Luna…
  • Working with the American Psychological Association on youth mental health and AI

    • Published: 2026-08-06 14:00 Beijing Time
    • Summary:
      • Young people are already using AI to learn, create, ask questions, and seek advice.
      • As this usage grows, families, schools, clinicians, and communities need clearer evidence, better resources, and stronger safeguards.
      • That’s why we are collaborating with the American Psychological Association (APA) to bring psychological science into our thinking about responsible AI development and use for young people.
      • The APA is the leading scientific and professional organization for psychology in the United States, and its evidence-based work will help clarify what is known, what is uncertain, and what responsible AI should look like as the technology evolves.
      • “Technology is a part of teenagers’ lives, and it is our responsibility to meet this reality with experiences that are safe, age-appropriate, and designed for families.
    • EN Highlights:
      • OpenAI and the American Psychological Association advance evidence-based guidance, resources, and safeguards for responsible AI use and youth mental health.
  • From asking to doing: How the world is putting ChatGPT to work

    • Published: 2026-08-06 08:00 Beijing Time
    • Summary:
      • New OpenAI Signals data reveals how people are using ChatGPT globally, offering country-level insights on adoption, usage trends, and evolving behaviors.
  • This article from the OpenAI blog explains From request to action: how the world makes ChatGPT work, shaping the broader AI and infrastructure landscape.

  • It also reveals the practical implications of From request to do: how the world makes ChatGPT work for founders, operators, and investors.

    • EN Key points:
      • New OpenAI Signals data shows how people use ChatGPT worldwide, with country-level insights on adoption, usage trends, and evolving behavior.

Google DeepMind Blog (A_full) Link to heading

  • WeatherNext: AI model achieves breakthrough in forecasting cyclones
    • Published: 2026-08-06 23:06 Beijing Time
    • Summary: - WeatherNext: An AI model has achieved a breakthrough in cyclone prediction.
      • This article from the Google DeepMind blog explains how WeatherNext: AI model achieves breakthrough in forecasting cyclones, shaping the broader AI and infrastructure landscape.
      • It also brings practical implications for founders, operators, and investors of WeatherNext: AI model achieves breakthrough in forecasting cyclones.
    • EN Key points:
      • WeatherNext: AI model achieves breakthrough in forecasting cyclones

ArXiv cs.AI (B_intro+search) Link to heading

  • A Long-Run Persistence Theory for AI Systems under the Redundancy-Adjusted Artificial Age Score (AAS)

    • Published: 2026-08-06 12:00 Beijing Time
    • Summary: - arXiv:2608.04012v1 Announce Type: new.
      • Abstract: Artificial intelligence systems are increasingly expected to operate over repeated cycles of interaction, adaptation, and update rather than through isolated, one-shot outputs.
      • This raises a fundamental theoretical question: can an AI system persist indefinitely without incurring unbounded structural aging?
      • This paper develops a long-run persistence framework for AI systems based on the redundancy-adjusted Artificial Age Score (AAS).
    • EN Key points:
      • arXiv:2608.04012v1 Announce Type: new
      • Abstract: Artificial intelligence systems are increasingly expected to operate over repeated cycles of interaction, adaptation, and update rather than through i…
      • This raises a fundamental theoretical question: can an AI system persist indefinitely without incurring unbounded structural aging
      • This paper develops a long-run persistence framework for AI systems based on the redundancy-adjusted Artificial Age Score (AAS)
  • The LLM Proposes, the Executive Disposes: A Self-Verifying Agent Instrument that Dissociates Commitment Drift from Binding Drift in Long-Horizon Agents

    • Published: 2026-08-06 12:00 Beijing Time
    • Summary: - arXiv:2608.04066v1 Announce Type: new.
      • Abstract: How can the authenticity of a long-horizon agent be verified when its state and self-reports are completely untrustworthy?
      • We propose an agent instrument that makes verification structural rather than post-hoc.
      • A deterministic executor holds all beliefs; the language model can only submit typed proposals, and a statement is acknowledged only if pre-registered predictions before an action match code observations.
    • EN Key points:
  • arXiv:2608.04066v1 Announce Type: new

  • Abstract: How do you verify a long-horizon agent when its own state and self-reports are exactly what you cannot trust

  • We present an agent instrument built so that verification is structural rather than post-hoc

  • A deterministic Executive owns all belief; a language model may only file typed proposals, and a claim is admitted only when a prediction pre-registered before…

  • Monte Carlo Tree Search for Table-to-Multimodal Report Generation

    • Published: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.04071v1 Announce Type: new.
      • Abstract: Automatically generating professional multimodal reports that include both textual analysis and visual charts from structured tabular data is a key challenge in data intelligence.
      • Existing methods suffer from fixed linear pipelines and isolated sub-task processing, which hinders the joint optimization of factual accuracy, visual quality, and narrative coherence.
      • To address these issues, this paper proposes MCTS-Report, a Monte Carlo Tree Search (MCTS)-driven framework that formulates the generation of multimodal table-to-report as a progressive construction process over a structured search space.
    • EN Highlights:
      • arXiv:2608.04071v1 Announce Type: new
      • Abstract: Automatically generating professional multimodal reports comprising both textual analysis and visual charts from structured tabular data is a critical…
      • Existing methods suffer from fixed linear pipelines and isolated subtask processing, which hinder joint optimization of factual accuracy, visual quality, and na…
      • To address these issues, this paper proposes MCTS-Report, a Monte Carlo Tree Search (MCTS)-driven framework that formulates multimodal table-to-report generatio…
  • FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables

    • Published: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.04077v1 Announce Type: new.
      • Abstract: Evaluating financial AI agents requires standards that align with actual professional work.
      • Existing methods for establishing evaluation criteria often derive them from task prompts or model outputs, overlooking the default standards that are only visible in practitioner deliverables.
      • We introduce FinProBench (a benchmark for professional financial tasks) and Role-Grounded Rubric Construction (RGRC), a reusable pipeline for deriving scoring criteria from deliverables generated by practitioners in the same role.
    • EN Highlights:
      • arXiv:2608.04077v1 Announce Type: new
      • Abstract: Evaluating financial AI agents requires criteria aligned with real professional work
  • Existing rubric methods typically derive criteria from task prompts or model outputs, overlooking tacit standards visible only in practitioner deliverables

  • We introduce FinProBench, a benchmark for professional financial tasks, and Role-Grounded Rubric Construction (RGRC), a reusable pipeline that derives rubrics from practitioner handbooks

  • FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

    • Publication Time: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.04095v1 Announce Type: new.
      • Abstract: Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear if they can maintain and update personalized user models over the long term.
      • Existing personalized-memory benchmarks primarily test factual retention or rely on weakly constrained model-generated trajectories, while event-driven preference adaptation remains under-explored.
      • We introduce FinPerMA, an event-grounded benchmark that evaluates personalized memory against frozen longitudinal investor trajectories.
    • EN Key Points:
      • arXiv:2608.04095v1 Announce Type: new
      • Abstract: Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains u…
      • Existing personalized-memory benchmarks primarily test factual retention or rely on weakly constrained model-generated trajectories, leaving event-driven prefer…
      • We introduce FinPerMA, an event-grounded benchmark that evaluates personalized memory against frozen longitudinal investor trajectories
  • BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding

    • Publication Time: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.04156v1 Announce Type: new.
      • Abstract: Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires a workflow that connects natural language instructions, signal processing, quantitative evidence, and scientific explanations.
      • We term this capability \emph{comprehensive EEG understanding}.
      • However, existing evaluations primarily target isolated decoding tasks or system-specific demonstrations, leaving the capabilities of large language models (LLMs) under-quantified.
    • EN Key Points:
      • arXiv:2608.04156v1 Announce Type: new
      • Abstract: Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural-language inst…
      • We term this capability \emph{comprehensive EEG understanding}
  • Existing evaluations, however, primarily target isolated decoding tasks or system-specific demonstrations, leaving the competence of large language models (LLMs…

  • Adversarially Robust Abductive Fusion of Pre-trained Transformer-based Perception Models

    • Published: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.04190v1 Announcement Type: new.
      • Abstract: Deploying pre-trained perception models in new environments degrades their accuracy under distributional shifts, and simply assembling them does not restore it: combiners like majority voting trade recall for precision and are prone to coordination failures.
      • Previous metacognitive methods learn logical rules that flag a model’s errors, but rely on hand-authored domain knowledge cues (object-size priors, segmentation masks) that do not transfer to truly novel scenarios.
      • We show that this metacognitive layer can be learned without any domain knowledge by exploiting vector-space geometry: per-model Label Vector Pools (LVP), built from each model’s own training embeddings, yield error-detection rules from the geometry of detections relative to training-identified prototypes, achieving parity with domain-knowledge rules at a cost of under $0.002 per F1 on the test set.
    • EN Highlights:
      • arXiv:2608.04190v1 Announce Type: new
      • Abstract: Deploying pre-trained perception models in novel environments degrades their accuracy under distributional shift, and assembling them alone does not r…
      • Prior metacognitive methods learn logical rules that flag a model’s errors, but rely on hand-authored domain-knowledge cues (object-size priors, segmentation ma…
      • We show that this metacognitive layer can be learned without any domain knowledge by exploiting vector-space geometry: per-model Label Vector Pools (LVP), built…
  • MatrAIx: Simulating the World with 8.3 Billion Persona Agents

    • Published: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.04205v1 Announcement Type: new.
      • Abstract: Human evaluation of AI systems and digital products is costly, slow, and difficult to scale.
      • Offline evaluations are more scalable, but often abstract away from human diversity and interactive behavior.
      • Therefore, we introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users.
    • EN Highlights:
      • arXiv:2608.04205v1 Announce Type: new
      • Abstract: Human evaluation of AI systems and digital products is costly, slow, and difficult to scale
      • Offline evaluations are more scalable but often abstract away human diversity and interactive behavior
      • We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users
  • Interoceptive Attention as Dynamic Homeostatic Prioritization in a Foraging Agent

    • Publication Time: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.04232v1 Announcement Type: New.
      • Abstract: Biological systems must regulate competing needs under limited perceptual bandwidth, where sharpening one estimate consumes the capacity to sharpen others.
      • Therefore, any fixed-budget system must decide where to allocate its perceptual precision.
      • We study this in a foraging agent that must satisfy multiple bodily needs to survive, modeled with active inference.
    • EN Highlights:
      • arXiv:2608.04232v1 Announce Type: new
      • Abstract: Biological systems must regulate competing needs under limited perceptual bandwidth, where sharpening one estimate costs the capacity to sharpen the o…
      • Any fixed-budget system therefore has to decide where to allocate its perceptual precision
      • We study this in a foraging agent that must keep several bodily needs satisfied to survive, modelled with active inference
  • The RAIL Principles for Neurosymbolic AI: Reasoning, Assurances, Interfacing and Learning

    • Publication Time: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.04285v1 Announcement Type: New.
      • Abstract: Neurosymbolic AI systems that integrate machine learning and symbolic reasoning are rapidly gaining attention.
      • They complement the data-intensive statistical approaches of neural networks and language models with symbolic reasoning algorithms to function in high-stakes domains or in low-data regimes representing many real-world applications.
      • We argue that the neurosymbolic combination of machine learning and formal reasoning is not a niche approach within AI, but rather includes many already successful techniques that are essential for developing reliable, efficient, and ultimately trustworthy systems.
    • EN Highlights:
      • arXiv:2608.04285v1 Announce Type: new
      • Abstract: Neurosymbolic AI systems that integrate machine learning and symbolic reasoning are rapidly gaining attention
      • They complement the data-intensive statistical approaches of neural networks and language models with symbolic reasoning algorithms to function in high-stakes d…
      • We argue that the neurosymbolic combination of machine learning and formal reasoning is not a niche approach within AI, but rather includes many already success…

ArXiv cs.CL (B_intro+search) Link to heading

  • TabletCraft: Bridging a 4,000-Year Cultural Gap with Bidirectional Akkadian NMT and Cuneiform Rendering

    • Publication Time: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.02609v1 Announcement Type: New.
      • Abstract: Half a million cuneiform tablets are preserved in museums around the world, but modern users can neither read nor write in the world’s oldest writing system, leaving a 4,000-year cultural barrier that existing NLP tools have only partially addressed.
  • Previous work achieved one-way, scholar-oriented translation from Akkadian to English, but did not provide a path in the opposite direction: non-specialist users cannot compose new content in cuneiform and thus remain passive consumers of ancient culture, rather than active participants.

  • We introduce TabletCraft, the first open-source system that enables bidirectional interaction with Mesopotamian writing.

    • EN Key points:
      • arXiv:2608.02609v1 Announce Type: new
      • Abstract: Half a million cuneiform clay tablets survive in museums worldwide, yet modern users can neither read nor write in the world’s oldest writing system,…
      • Prior work enables one-way, scholar-oriented translation from Akkadian to English, but offers no path in the reverse direction: non-specialist users cannot comp…
      • We present TabletCraft, the first open-source system that enables bidirectional interaction with Mesopotamian writing
  • BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems

    • Release Date: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.02612v1 Announce Type: new.
      • Abstract: Formulating an optimization problem greatly affects the quality of the final solution, and a good formulation usually requires substantial expertise.
      • Therefore, recent research has investigated how to automatically derive optimization problems from natural language descriptions, but existing benchmarks focus on settings where objectives and constraints can be explicitly written as mathematical expressions.
      • Many practically important problems are naturally treated as black-box optimization (BBO) problems, where only objective values can be observed, and no function form is available.
    • EN Key points:
      • arXiv:2608.02612v1 Announce Type: new
      • Abstract: Formulating an optimization problem strongly affects the quality of the final solution, yet good formulations usually require substantial expertise
      • Recent studies have therefore examined how to automatically derive optimization problems from natural-language descriptions, but existing benchmarks focus on se…
      • Many practically important problems are naturally treated as black-box optimization (BBO) problems, in which only objective values are observable, and the funct…
  • MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale

    • Release Date: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.02613v1 Announce Type: new.
      • Abstract: Edge-deployed personal memory assistants must process on-device private interpersonal conversations using open-weight models.
      • However, existing memory benchmarks often fail to adequately test the combination of activity-intensive interactions, ego-centric perspective, and coherent multi-session worlds.
      • MemArena fills these gaps with its single-world conversational benchmark built via its MASim agent simulator, across 50 agents over 15 days (10.3M conversational text tokens, 24.1K plaintext ego-observation tokens/agent/day).
    • EN Key points:
      • arXiv:2608.02613v1 Announce Type: new
  • Abstract: Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models

  • Yet, existing memory benchmarks often under-test the combination of activity-dense interaction, ego-centric perspective, and coherent multi-session worlds

  • MemArena fills these gaps with a single-world conversational benchmark built with its MASim agent simulator, for 50 agents over 15 days (10.3M dialog-text token…

  • OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning

    • Published: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.02615v1 Announce Type: new.
      • Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata.
      • However, most medical large language model (LLM) and vision-language model (VLM) benchmarks focus on isolated modalities or narrow image-text tasks, leaving patient-level oncology assessment across multiple evidence streams largely untested.
      • We introduce OncoTriad-QA, a patient-level radiology-pathology-genomics benchmark for pan-cancer question answering.
    • EN Highlights:
      • arXiv:2608.02615v1 Announce Type: new
      • Abstract: Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata
      • However, most medical large language model (LLM) and vision-language model (VLM) benchmarks focus on isolated modalities or narrow image-text tasks, leaving pat…
      • We introduce OncoTriad-QA, a patient-level radiology-pathology-genomics benchmark for pan-cancer question answering
  • Evaluating OpenAI’s Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks

    • Published: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.02616v1 Announce Type: new.
      • We conduct the first independent, systematic evaluation of OpenAI’s Privacy Filter (OPF), a 1.5B parameter bidirectional PII detector, across 42 comprehensive benchmarks spanning 22 languages and 5 domains.
      • Zero-shot, OPF achieves F1=0.855 on AI4Privacy and 0.464 on SPY Medical, outperforming Presidio (0.431, 0.273) and XLM-RoBERTa (0.269, 0.111) on PII-annotated benchmarks; on multilingual NER, XLM-RoBERTa leads OPF on all 13 Indic and non-Latin languages.
      • GPT-4o leads in medical, legal, and financial PII (SPY: avg 0.643, Gretel: 0.527), while OPF leads in structured synthetic PII (avg 0.71) and customer support (avg 0.60).
    • EN Highlights:
      • arXiv:2608.02616v1 Announce Type: new
  • Abstract: We present the first independent, systematic evaluation of OpenAI’s Privacy Filter (OPF), a 1.5B-parameter bidirectional PII detector, across 42 synth…

    • Zero-shot, OPF achieves F1=0.855 on AI4Privacy and 0.464 on SPY medical, outperforming Presidio (0.431, 0.273) and XLM-RoBERTa (0.269, 0.111) on PII-annotated b…
    • GPT-4o leads on medical, legal, and financial PII (SPY: 0.643 avg, Gretel: 0.527), while OPF leads on structured synthetic PII (0.71 avg) and customer support (…
  • Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety

    • Publication Time: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.02617v1 Announcement Type: New.
      • Abstract: We use expert feedback from MOOVE (Massive Open Online Validation and Evaluation) to assess whether clinician pairwise preferences provide a reliable signal for clinical safety in large language model (LLM) evaluations. MOOVE is a clinician-led platform that collects blind pairwise preferences and multi-criteria scores.
      • Clinicians score on a discrete $[-2, +2]$ scale, where negative values indicate clinically unsafe or misleading content.
      • Using 26,804 pairwise judgments from 13 LLMs, provided by over 736 clinicians from more than 28 countries, we find that clinician preferences do not well represent safety-critical performance.
    • EN Highlights:
      • arXiv:2608.02617v1 Announce Type: new
      • Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert…
      • Clinicians assign scores on a discrete $[-2, +2]$ scale, where negative values indicate clinically unsafe or misleading content
      • Using 26{,}804 pairwise judgments across outputs from 13 LLMs, contributed by more than 736 clinicians across 28+ countries, we find that clinician preference i…
  • JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation

    • Publication Time: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.02620v1 Announcement Type: New.
      • Abstract: LLM-as-a-judge evaluation has become the dominant paradigm for ranking language models, but the ecosystem remains fragmented: most benchmarks have their own codebases, hardcode specific closed-model judges, and support a single evaluation protocol.
      • This fragmentation makes it difficult to understand how research design choices (benchmarks, judge models, prompts, inference backends) affect the conclusions we draw about model quality.
      • We introduce JudgeArena, an open-source framework that unifies major LLM-judge benchmarks (AlpacaEval, Arena-Hard, MT-Bench, and m-Arena-Hard) under a single interface with swappable judges and comprehensive metadata logging to improve transparency in reporting and reproducibility.
    • EN Highlights:
      • arXiv:2608.02620v1 Announce Type: new
  • Abstract: LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their…

  • This fragmentation makes it difficult to study how design choices–the benchmark, the judge model, the prompt, the inference backend–affect the conclusions we…

  • We introduce JudgeArena, an open-source framework that unifies major LLM-judge benchmarks (AlpacaEval, Arena-Hard, MT-Bench, and m-Arena-Hard) under a single in…

  • Knowing the Form, Not the Function: Automatically Auditing Answer–Authority Decoupling in Legal Benchmarks

    • Publication Time: 2026-08-06 12:00 Beijing Time
    • Abstract:
      • arXiv:2608.02621v1 Announce Type: new.
      • Abstract: Legal benchmarks typically score final answers even when models also state legal authority.
      • We test whether answer correctness can serve as a proxy for authority grounding.
      • Under ordinary reasoning prompts that did not request statutory citations, four LLMs spontaneously produced authority markers across 238 Taiwan bar-examination items.
  • Speculative Correction: Draft-then-Refine Decoding for Diffusion Language Models

    • Publication Time: 2026-08-06 12:00 Beijing Time
    • Abstract:
      • arXiv:2608.02625v1 Announce Type: new.
      • Abstract: Diffusion language models (DLMs) can revise tokens bidirectionally, but standard decoding procedures often adapt them to left-to-right generation by producing text chunk-by-chunk.
      • We study a simple plug-and-play inference pattern: first generate a complete draft, then refine the full response using bidirectional diffusion.
      • Using LLaDA2.1-Flash and LLaDA2.1-Mini, we evaluate two configurations.
  • Using LLaDA2.1-Flash and LLaDA2.1-Mini, we evaluate two configurations

  • Stuck on “A”: Diagnosing and Repairing Interface Injury in Attention-to-KDA Linearization of a 0.6B Language Model

    • Published: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.02689v1 Announce Type: new.
      • Abstract: We convert 21 of 28 full-attention layers of Qwen3-0.6B-Base into KDA (Kimi Delta Attention) linear-attention layers on a single consumer-grade GPU budget and ask a simple question: what exactly does the conversion break?
      • After the procedure, hidden-state alignment and end-to-end KL distillation bring the student close to its teacher in perplexity, yet multiple-choice accuracy remains near random chance (25-29% vs.
      • the teacher’s 50.6% on C-Eval).
    • EN Highlights:
      • arXiv:2608.02689v1 Announce Type: new
      • Abstract: We convert 21 of 28 full-attention layers of Qwen3-0.6B-Base into KDA (Kimi Delta Attention) linear-attention layers on a single consumer-grade GPU bu…
      • After surgery, hidden-state alignment and end-to-end KL distillation drive the student close to its teacher in perplexity, yet multiple-choice accuracy stays ne…
      • the teacher’s 50.6% on C-Eval)

ArXiv cs.LG (B_intro+search) Link to heading

  • C$^2$MOE: Consistency and Complementarity-guided Mixture of Experts for Incomplete Multimodal Emotion Learning

    • Published: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.04013v1 Announce Type: new.
      • Abstract: Recent advances in Multimodal Emotion Recognition in Conversations (MERC) highlight its reliance on complete multimodal inputs.
      • However, real-world data often suffers from missing modalities due to transmission errors or user behavior, severely degrading model performance.
      • Existing methods enhance robustness through cross-modal consistency learning but largely ignore modality complementarity, leading to biased reconstructions.
    • EN Highlights:
      • arXiv:2608.04013v1 Announce Type: new
      • Abstract: Recent advances in Multimodal Emotion Recognition in Conversations (MERC) highlight its reliance on complete multimodal inputs
      • However, real-world data often suffer from missing modalities due to transmission errors or user behavior, severely degrading model performance
      • Existing methods enhance robustness via cross-modal consistency learning but largely ignore modality complementarity, leading to biased reconstructions
  • On Hamming-Lipschitz Type Stability of the Subdominant (Minmax) Ultrametric: Theory and Simple Proofs

    • Published: 2026-08-06 12:00 Beijing Time
  • Abstract: - arXiv:2608.04014v1 Announce Type: new.

    • Abstract: The subdominant (minmax) ultrametric is a canonical tree-structured summary of a dissimilarity matrix, equivalent to the ultrametric induced by single-linkage clustering.
    • While its classical stability theory is usually formulated in $\ell_\infty$ or Gromov–Hausdorff terms, such bounds are poorly suited to sparse perturbations that only change a few pairwise distances.
    • We develop an $\ell_0$-type stability theory for this operator.
  • A Trust-region Framework for Moment Estimation

    • Release Time: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.04026v1 Announce Type: new.
    • Abstract: In this paper, we develop a trust-region framework for understanding the behavior of adaptive moment estimation mechanisms (such as \textsc{Adam}) in stochastic gradient optimization.
    • Specifically, in this framework, the magnitude of the update step for each individual weight is constrained within a trust-region governed by a $p\in[2,4]$ order moment constraint.
    • The resulting derivation leads to a series of learning rate mechanisms based on second-moment estimation and a normalized $p$-th moment estimation.
  • Learning to Resolve Neutron Resonances with Fully Convolutional Neural Networks

    • Release Time: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.04027v1 Announce Type: new.
    • Abstract: This work investigates the feasibility of enhancing traditional R-matrix codes with a powerful machine learning framework to automatically detect neutron resonances in transmission spectra.
    • Neutron transmission data is often complex and noisy, making it difficult to analyze using traditional peak identification methods.
    • The state-of-the-art R-matrix codes currently used by physicists to fit this data often depend on prior evaluations and require substantial manual effort.
  • Abstract: This work investigates the feasibility of augmenting traditional R-Matrix codes with a robust machine learning framework for automatically detecting n…

    • Neutron transmission data are often complex and noisy, making them difficult to analyze using traditional peak-identification methods
    • The state-of-the-art R-Matrix codes currently used by physicists to fit these data often depend on prior evaluations and require substantial manual effort
  • Lindblad-Inspired Multi-Timescale Reservoir Computing with Separable Rotation and Dissipation

    • Publication Time: 2026-08-06 12:00 Beijing Time
    • Abstract:- arXiv:2608.04028v1 Announcement Type: new.
      • Abstract: Echo-state networks achieve efficient temporal learning by fixing recurrent dynamics and training only a linear readout.
      • However, conventional reservoirs typically accommodate signal mixing, memory retention, and stability within a single random recurrent matrix.
      • Existing structured designs improve topology, norm preservation, leakage, or depth, but generally do not provide separate modal control of reversible mixing and irreversible forgetting, as well as direct global stability guarantees.
    • EN Key Points:
      • arXiv:2608.04028v1 Announce Type: new
      • Abstract: Echo-state networks enable efficient temporal learning by fixing the recurrent dynamics and training only a linear readout
      • However, conventional reservoirs typically accommodate signal mixing, memory retention, and stability within a single random recurrent matrix
      • Existing structured designs improve topology, norm preservation, leakage, or depth, but generally do not provide separate modal control of reversible mixing and…
  • An Explainable LLM Agent Layer for Open-World Anomaly Detection in Oil Wells

    • Publication Time: 2026-08-06 12:00 Beijing Time
    • Abstract:- arXiv:2608.04041v1 Announcement Type: new.
      • Abstract: Recent work has shown that Open-World Learning (OWL) pipelines for oil well anomaly detection combine autoencoder-based detection, multi-class classification, and Mahalanobis-based novelty detection on the public 3W dataset.
      • These pipelines answer \textit{what happened}, but they do not explain \textit{why the model believes it} or \textit{what the operator should do next}, and they do not place human-readable names on the novel clusters they discover.
      • This paper evaluates a Large Language Model (LLM) agent layer located downstream of the OWL pipeline, designed as a \textbf{companion} to the published upstream methods, rather than a replacement.
    • EN Key Points:
      • arXiv:2608.04041v1 Announce Type: new
      • Abstract: Open-World Learning (OWL) pipelines for oil well anomaly detection have recently been shown to combine autoencoder-based detection, multiclass classif…
  • These pipelines answer \textit{what happened}, but they do not explain \textit{why the model believes it} or \textit{what the operator should do next}, and they…

  • This paper evaluates a Large Language Model (LLM) agent layer placed downstream of the OWL pipeline, designed as a \textbf{companion} to the published upstream…

  • Tactus: Open-Vocabulary Object Recognition from Low-Cost Pressure Arrays

    • Publication Time: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.04043v1 Announcement Type: New.
      • Abstract: Resistive pressure arrays are the cheapest and most widely shipped tactile sensors, but tactile representation learning has primarily focused on optical sensors that image deforming gels.
      • We present Tactus, an open model that answers text queries from pressure data alone: on the STAG benchmark (27 objects, held-out recordings), it achieves 0.771 +/- 0.062 top-1 (0.935 top-3) across four runs, matching and at best outperforming the dataset’s supervised closed-set CNN (without a trained classifier head) at 0.76.
      • The secret is small data: 187 training recordings, masked autoencoder pre-training on 144k unlabeled frames from the same sensor, and the sensor’s own calibration affine, which recovers more accuracy than all architectural changes combined.
    • EN Highlights:
      • arXiv:2608.04043v1 Announce Type: new
      • Abstract: Resistive pressure arrays are the cheapest and most widely shipped tactile sensors, yet tactile representation learning has concentrated on optical se…
      • We present Tactus, an open model that answers text queries from pressure data alone: on the STAG benchmark (27 objects, held-out recordings), it reaches 0.771 +…
      • The recipe is small-data: 187 training recordings, masked-autoencoder pretraining on 144k unlabeled same-sensor frames, and the sensor’s own calibration affine,…
  • Robust and Personalized Federated Learning for Aircraft-Engine Prognostics under Benign and Adversarial Client Heterogeneity

    • Publication Time: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.04045v1 Announcement Type: New.
      • Abstract: Federated Learning (FL) enables aircraft fleet operators to collaboratively train Remaining Useful Life (RUL) models using engine sensor telemetry without sharing raw data.
      • This study investigates two complementary challenges: benign heterogeneity (where honest operators observe different operating conditions and fault modes) and adversarial heterogeneity (where compromised operators submit toxic updates).
      • We conduct a controlled, safety-oriented evaluation using a multi-task 1D convolutional neural network and a structurally non-IID partition of the Commercial Modular Aero-Propulsion System Simulation (C-MAPSS) benchmark.
    • EN Highlights:
      • arXiv:2608.04045v1 Announce Type: new
  • Abstract: Federated learning (FL) enables aircraft fleet operators to jointly train remaining-useful-life (RUL) models from engine sensor telemetry without shar…

  • This study examines two complementary challenges: benign heterogeneity, where honest operators observe different operating conditions and fault modes, and adver…

  • We conduct a controlled, safety-oriented evaluation using a multi-task one-dimensional convolutional neural network and a structurally non-IID partition of the…

  • Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs

    • Publication Time: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.04048v1 Announcement Type: New.
      • Abstract: Serving Large Language Models (LLMs) under diverse deployment constraints requires a flexible trade-off between accuracy, memory footprint, and throughput.
      • However, traditional quantization methods typically require a separate checkpoint for each target bit-width.
      • We introduce Recurrent Residual Quantization (RRQ), a post-training quantization (PTQ) framework that represents weights as a low-bit quantized base along with a series of quantized residual corrections, thereby achieving multiple effective precisions from a single checkpoint.
    • EN Highlights:
      • arXiv:2608.04048v1 Announce Type: new
      • Abstract: Serving large language models (LLMs) under diverse deployment constraints requires flexible trade-offs between accuracy, memory footprint, and through…
      • However, conventional quantization methods typically require a separate checkpoint for each target bit-width
      • We introduce Recurrent Residual Quantization (RRQ), a post-training quantization (PTQ) framework that represents weights as a low-bit quantized base together wi…
  • CAMP: A Cycle-Aware Multi-Scale Patch Mixer for Time Series Forecasting

    • Publication Time: 2026-08-06 12:00 Beijing Time
    • Abstract: - arXiv:2608.04051v1 Announcement Type: New.
      • Abstract: Real-world time series are often governed by recurring patterns, but their dominant periods may vary across datasets, forecasting settings, and individual input windows.
      • Existing cycle-aware predictors often rely on a single period selected at the dataset level, which can be limiting when periodic behavior changes over time or when multiple cycles coexist.
      • Furthermore, patch-based models typically process all patch locations uniformly, although patches far from the prediction boundary may require broader contextual refinement, while recent patches contain information that should be more directly preserved.
    • EN Highlights:
      • arXiv:2608.04051v1 Announce Type: new
      • Abstract: Real-world time series are often governed by recurring patterns, but their dominant periods may vary across datasets, forecasting settings, and indivi…
  • Existing cycle-aware forecasters commonly rely on a single period selected at the dataset level, which can be restrictive when periodic behavior changes over ti…

  • Moreover, patch-based models typically process all patch positions uni- formly, although patches farther from the forecast boundary may require broader contextu…