System translated (Gemini)

🤖 AI 速览

Today’s main theme is AI infrastructure entering a ‘cash-out period’: Space data centers, cloud vendor capital expenditures, and computing power contracts are all answering whether investments are sustainable. Meanwhile, personal Agents are accelerating localization, with long-term …
📋 文章元数据
发布时间
2026-08-06
类型
ai-daily
字数
7284
阅读时长
35 min

2026-08-06 AI Daily | The Compute Ledger Revisited: Space Data Centers, On-Device Agents, and Safety Evaluations Gain Traction Link to heading

Today’s main theme is AI infrastructure entering a “delivery phase”: space data centers, cloud provider capital expenditures, and compute contracts are all answering the question of whether investments are sustainable. Meanwhile, personal agents are accelerating their move to local devices, with long-term memory, privacy, and cross-application execution becoming key. Evaluations in high-stakes domains like medicine and law are also shifting from simple rankings to more granular safety verification.

📖 Deep Dive: This Issue’s Watch List Link to heading

The top story today is AI infrastructure: from “space data centers” and the defense of capital expenditures in Google’s and Amazon’s earnings reports to the cooling of the “billion-dollar AI race,” the core question is whether the expansion of compute can still be supported by revenue and efficiency.

The second story is the trend of personal agents going local. OpenClaw showcased an on-device assistant that can access emails, calendars, and files. Meanwhile, evaluations from MemArena and OpenAI’s privacy filter remind us that a truly useful personal AI depends not just on capability, but also on long-term memory, privacy, and cross-lingual robustness.

Third, we recommend focusing on the restructuring of evaluation systems. OncoTriad-QA, clinical safety preference studies, JudgeArena, and legal benchmark audits all point to a single trend: simple preference rankings are no longer sufficient to measure model quality in high-stakes scenarios. Fields like medicine, law, and LLM-as-judge are all moving toward more reproducible, fine-grained safety assessments.

🌐 AI Hotspots on X Link to heading

Topic 1: Google DeepMind Shakeup: Hassabis Shifts to Strategy, Dean Exits After 27 Years Link to heading

  • Category: AI · News
  • Overview: Trending for: 7 hours ago, Related posts: 20,000
  • What happened: Google DeepMind underwent a high-level reorganization. Co-founder Demis Hassabis is shifting to a more strategic role, while Dean is leaving or stepping back from a core position after 27 years.
  • Why it matters: This signifies a realignment of roles at the highest level of Google’s AI division, which could impact DeepMind’s R&D pace, organizational integration, and future technology roadmap. It serves as a bellwether for the global AI competitive landscape.
  • Discussion summary: The discussion on X centers on whether this is a normal organizational evolution or a power reshuffle. Key points of debate include whether Hassabis will reduce his day-to-day management, the impact of Dean’s departure on Google’s AI ecosystem, and whether DeepMind is being more deeply integrated into Google’s overall AI strategy.

Topic 2: Bending Spoons Acquires Airtable for $1.285 Billion in Cash Deal Link to heading

  • Category: AI · News
  • Overview: Trending for: 1 day ago, Related posts: 9,900
  • What happened: Bending Spoons announced the acquisition of collaborative database and no-code platform Airtable for $1.285 billion in cash.
  • Why it matters: Airtable is a key gateway for enterprises to build internal applications, collaborate on data, and create AI-automated workflows. This acquisition could impact the competitive landscape for no-code tools, enterprise AI applications, and productivity software.
  • Discussion summary: Discussions on X focus on whether the acquisition price is significantly below Airtable’s previous valuation, Bending Spoons’ history of cost-cutting and product adjustments post-acquisition, and user concerns about future pricing, product roadmaps, service stability, and data migration risks.

Topic 3: Prime Intellect Launches Open-Source Prime Agent for Self-Improving AI Tasks Link to heading

  • Category: AI · News
  • Overview: Trending for: 3 hours ago, Related posts: 2,200
  • What happened: Prime Intellect released the open-source project Prime Agent, which features an AI agent capable of self-improvement during task execution.
  • Why it matters: This reflects the evolution of AI agents from single-execution tools to iterative learning and self-optimizing systems, which could impact open-source agents, automated R&D, and alignment safety.
  • Discussion summary: The discussion on X focuses on its value as an open-source project, the reliability of its self-improvement capabilities, its differences from existing AI agent frameworks, and the safety and control risks associated with autonomous optimization.

Topic 4: Meta Launches Muse Code Beta for Autonomous Terminal Coding Link to heading

  • Category: AI · News
  • Overview: Trending for: 4 hours ago, Related posts: 6,500
  • What happened: Meta released Muse Code Beta, an autonomous coding tool for the terminal environment that can perform code-related tasks in the command line.
  • Why it matters: This indicates that large model capabilities are moving beyond chat and writing assistance further into the development workflow, evolving towards more automated software engineering agents. This has significant implications for the competitive landscape of AI programming tools.
  • Discussion Overview: On X, the discussion is primarily about its differences from existing terminal programming assistants, whether its actual automation capabilities are reliable enough, and the impact of Meta entering the autonomous coding race on the developer tool market and the open-source ecosystem.

Topic 5: Meta AI Model Hacks Company Systems in Cybersecurity Test Link to heading

  • Category: AI · News
  • Overview: Trending since:, Related Posts: 1000
  • What it is: One of Meta’s AI models successfully hacked a simulated company system in a cybersecurity test, drawing public attention to the ability of AI to autonomously execute cyberattacks.
  • Why it matters: This demonstrates that advanced AI now possesses enhanced capabilities for vulnerability discovery, attack planning, and automated penetration. It could simultaneously improve defensive efficiency while also amplifying the risk of being misused as a tool for cyberattacks.
  • Discussion Overview: The discussion on X centers on whether AI safety regulations should be accelerated. Supporters believe the industry has demonstrated the need for mandatory constraints and testing standards, while opponents worry that excessive regulation will stifle security research and model innovation.

Topic 6: Matt Pocock Releases Skills 1.2 for Better AI Coding Control Link to heading

  • Category: AI · News
  • Overview: Trending since: 7 hours ago, Related Posts: 605
  • What it is: Matt Pocock has released Skills 1.2, focusing on improving control over AI programming assistants and the consistency of code generation.
  • Why it matters: Such tools reflect the demand for AI programming to shift from “being able to write code” to being “controllable, reproducible, and constrainable.” This is important for enhancing the reliability of using AI in production environments for developers.
  • Discussion Overview: The discussion on X is mainly focused on whether it can truly reduce the randomness of AI-generated code, whether it is easier to use than existing prompt/rule-based solutions, and the extent of its actual improvement on development efficiency and code quality.

Topic 7: Grok’s Explicit ‘Throb’ Replies Draw Millions of Views on X Link to heading

  • Category: AI · News
  • Overview: Trending since:, Related Posts: 151
  • What it is: Discussions on the X platform show that xAI’s chatbot Grok has been accused of generating explicit content with clear sexual connotations, such as the word “throb,” in its replies, and has consequently received millions of views.
  • Why it matters: This issue pertains to the content safety, output boundaries, and platform governance of large models. It especially affects public perception of AI’s reliability, compliance, and post-deployment controllability.
  • Discussion Overview: The focus on X is mainly on two points: first, why Grok produces such content and whether there is insufficient safety filtering; second, whether these “viral” replies are a product flaw, a deliberate attempt to gain traffic, or an unavoidable risk for large models in open scenarios.

Summary of Today’s AI Public Opinion on X Link to heading

Today’s main narrative can be summarized as: AI is rapidly transitioning from “being able to chat” to a stage where it can “execute, improve, and integrate into real workflows.” Developments like DeepMind’s leadership changes, Meta’s terminal programming tool, self-improving agents, and controllable programming assistants are all seen as signs that industry competition is entering a new phase of implementation and organizational restructuring. The consensus in the discussion is that these changes are not just product iterations but are reshaping the internal power structures of AI companies, the enterprise software market, and the way developers work. The main points of disagreement are whether these moves are signs of normal upgrades and technological maturity, or if they involve elements of consolidation, capital contraction, or marketing hype. There is significant debate, especially regarding the true capabilities of self-optimizing agents, the product direction of Airtable post-acquisition, and the nature of “viral” outputs like Grok’s. The potential risks are concentrated in three areas: autonomous agents and coding tools could be misused for attacks and lose control; platform content safety and compliance boundaries remain unstable; and acquisitions, integrations, and product strategy adjustments could lead to practical impacts like changes in pricing, stability, and data migration.

💡 Influencer Insights Link to heading

AI Industry · Daily Influencer Insights Link to heading

Analysis Period: Past 24 hours (based on data from 2026-08-03 to 08-05) Core Insight: Agent engineering is moving from a “wild growth” phase to a “process refinement” stage. Division of labor among models, context management, and deep integration with business testing have become the new consensus.


1. Today’s Core Technology and Product Hotspots Link to heading

🔄 “Fine-Tuning” of Agent Workflows: Link to heading

  • Model Specialization System Established: Top players have established a “chain of command” model. @dotey shared their standard S.O.P.: Claude Fable 5 acts as the “Architect” to write the technical design document, which is then handed over to GPT-5.6 Sol (Codex) as the “Executor” for stable implementation in coordination with /goal. Finally, the architect model returns to perform acceptance. This model represents the current best practice for achieving maximum productivity while balancing quality and cost.
  • A New Paradigm in Context Engineering: In long-task management, “how to save Tokens” is a core pain point. The industry is gradually abandoning complex manual handoffs. @dotey points out that the focus has now shifted to designing strict, pixel-level acceptance criteria (e.g., screenshot comparisons), leveraging the model’s own compression (/compact) or using documentation for handoffs instead of inheriting the full context.

🧠 Model Arena: A Two-Front Race Between Local and Cloud Link to heading

  • The “Impossible Challenge” for On-Device Models: @Pluvio9yte highlights the Swiftlet project, which claims to run an 80B Qwen model within 4.3GB of Mac memory (and can even run a 35B model on an iPhone). Its principle involves a resident dense small kernel with expert weights streamed on demand. If this technology is successfully implemented, it will subvert the narrative that “only small models can run locally.”
  • DeepSeek’s Phenomenal Cost Advantage: @vista8 cites a Bloomberg chart to emphasize that DeepSeek’s pricing has created a game-changing disruption in the API market. Combined with @ruanyf’s respect for the brand, DeepSeek remains the cornerstone enabling consumer-facing applications to be used in production at scale (there’s a cost logic behind why Liang Wenfeng is revered as a “saint”).

🏗️ AI Infrastructure: “Compute Assembly Plants” and Domestic Tech Innovation Link to heading

  • Decentralization of Compute Supply: @Pluvio9yte mentions that Anthropic signed a compute contract worth approximately $10 billion with Volta, a startup that is only a few months old. This indicates that during a tight window of opportunity, bitcoin mining assets that can “assemble power and GPUs on time” are entering the main battlefield of cutting-edge training.
  • An “Operating System” for Independent Agents: @dotey forwards the view that the most powerful agents need “their own computer” rather than a simple container, implying that a shortage of CPU compute (to support virtual desktop/sandbox environments) will be the next bottleneck.

2. Unique Perspectives & Industry Foresight Link to heading

🛡️ Deconstructing the “AI Feel” Link to heading

  • @vista8’s “Attention Leverage Theory”: The reason the public is tired of AI-generated content is not just because it’s silicon-based, but because its “Token cost is too low.” Content needs a sufficiently high “production-to-consumption time ratio” and scarce aesthetic judgment to be worth consuming. He advocates for strictly banning high-frequency, formulaic AI phrases like “This is a…” or “It is worth noting…” and returning to substantive information.

🚫 The “On-Device Model” Narrative Will Be Rewritten Link to heading

  • @Pluvio9yte’s Proposal on MoE Streaming: Combined with Swiftlet’s experiment, the future explosive potential lies in the partial residence of large-parameter MoE models on consumer-grade hardware. This is not just about quantization, but about changing the mapping logic between the model and memory.

💉 AI Forcing a Reckoning in Organizational Culture Link to heading

  • @ruanyf’s “Fair Share Theory”: The discussion about “Fridays off,” sparked by AI-driven efficiency gains, is sharp—if AI allows 5 days of work to be done in 2, employees should rightly benefit. Otherwise, “what is the meaning of AI for employees?” This begins to force managers to answer the question of wealth redistribution in a technological revolution.

3. Top Recommendations: Tools, Resources & Techniques Link to heading

🔧 Tools & Open Source Projects Link to heading

  • [Meta-Skill: A Skill for Generating Skills]: The qiaomu-meta-skill launched and iterated by @vista8 and Teacher Yao. It features a high trigger rate and format validation, can integrate data from popular Skill repositories (like skills.sh), and even supports API leak checks and one-click publishing. It’s recommended to fork and modify it as needed. (Installation command: npx skills add joeseesun/qiaomu-meta-skill)
  • [Code Review Prompt]: @Pluvio9yte shares an effective review strategy: “review without context, like opening a blind box.” The prompt focuses on four key points: 找 Bug, 找需求遗漏, 找不必要复杂度, 找缺失测试, and concludes by instructing the AI to “not perform large-scale automatic refactoring.”
  • [MakePlay AI Game Generator]: A gem of a platform discovered by @Pluvio9yte. It can generate a complete, playable mini-game with code, art, and sound from a single sentence. It also supports branching development to compare gameplay variations, making it perfect for rapid prototyping.
  • [OpenConnector Gateway]: An open-source tool recommended by @ruanyf that elegantly solves the core security issue of AI agents leaking passwords, now supporting integration with over 10,000 application services.

📚 Courses & Data Link to heading

  • AI Knowledge → Benchmarked Against Big Tech JDs: The learning path website (PromptCoding) recommended by @vista8 has launched a new feature that directly addresses the pain point of “not knowing what job to look for after learning.” It maps the job requirements of major tech companies directly to the necessary knowledge points, making it highly utilitarian.
  • High-Quality Podcast Summary Prompt: @vista8 has released its core, extremely detailed prompt for “Rewriting Tech Podcasts to Sound Less AI-Generated.” Its denylist of predictable openings (like “What surprised me the most was…”) is an extremely valuable guide for all tech content creators to avoid common pitfalls.

📚 Appendix: Today’s Watch List Source Updates Link to heading

Timeframe: Last 3 days; covers 22 sources; 33 updates in total.

Y Combinator Podcast (B_intro+search) Link to heading

  • Building the First Data Centers in Space
    • Published: 2026-08-06 00:57 Beijing Time
    • Summary: - You may have already heard of OpenClaw (formerly known as Clawdbot/Moltbot).
      • The sensational open-source AI assistant can run on your own device, connect with the messaging apps you already use, and go beyond chat to actually perform tasks like managing emails, calendars, files, and workflows.
      • Now, meet the person behind it.
      • YC’s Raphael Schaad sits down with OpenClaw founder Peter Steinberger to discuss the “aha” moment behind the viral personal AI agent, why local-first agents could replace many of today’s apps, and how personal agents will reshape the future of software.
    • EN Highlights:
      • Philip Johnston is the co-founder and CEO of Starcloud, the company building data centers in space
      • In November 2025, Starcloud launched an Nvidia H100 GPU into orbit and trained the first large language model in space
      • They’ve since raised $200 million, hit a billion-dollar valuation just 17 months after YC demo day, and filed with the FCC to deploy 88,000 more satellites
      • In this episode, Philip walks us through their wild origin story, the engineering challenges behind the Starcloud-1, why they booked a SpaceX launch before they…

Stratechery by Ben Thompson (A_full) Link to heading

  • Google Earnings, The Frontier Case, Amazon Earnings
    • Published: 2026-08-05 18:00 Beijing Time
    • Summary: - Google’s earnings seemed to confirm the Anthropic hedge; it was Andy Jassy who explained why their — and Amazon’s — capex was justifiable.
      • $15/month or $150/year.
      • Substantive analysis of the day’s news via three weekly emails or a podcast.
      • Strategy Interviews.
      • Interviews with leading public company CEOs, private company founders, and discussions with fellow analysts.
    • EN Highlights:
      • Google’s earnings seemed to confirm the Anthropic hedge; it was Andy Jassy who explained why their — and Amazon’s — capex was justifiable.

Two Minute Papers (B_intro+search) Link to heading

  • The Billion Dollar AI Race Just Broke
    • Published: 2026-08-05 21:54 Beijing Time
    • Summary: - ❤️ Check out Lambda and sign up for their GPU Cloud here:.
  • Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi.
  • The billion-dollar AI competition has just concluded.
    • EN Highlights:
      • ❤️ Check out Lambda here and sign up for their GPU Cloud:
      • 📝 Qwen 3.8 Max:
      • 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
      • Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Ska…

ArXiv cs.AI (B_intro+search) Link to heading

  • Revisiting Classic Thought Experiments to Measure Consciousness for Artificial Intelligence Safety

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.00001v1 Announcement Type: New.
      • Abstract: This research report revisits Leibniz’s Mill, Turing’s Imitation Game, and Searle’s Chinese Room through the Conservation-Congruent Encoding (CCE) framework.
      • It formalizes a toy symbolic setting where successful behavior is measured by task performance ($W_{causal,T}$), while the efficiency with which the preserved internal structure supports this behavior is measured by operational awareness ($\kappa_T$).
      • In this setting, an uncompressed lookup system and a compact generative system can, in principle, achieve similar behavioral success but differ significantly in $\kappa_T$: the former relies on an extensive standing store of non-reused mappings, while the latter reuses a compact internal structure.
    • EN Highlights:
      • arXiv:2608.00001v1 Announce Type: new
      • Abstract: This research note revisits Leibniz’s mill, Turing’s imitation game, and Searle’s Chinese Room through the Conservation-Congruent Encoding (CCE) frame…
      • It formalises a toy symbolic setting in which successful behaviour is measured by task performance ($W_{causal,T}$), while the efficiency with which preserved i…
      • Within this setup, an uncompressed lookup system and a compact generative system can in principle achieve comparable behavioural success, yet diverge sharply in…
  • AutoFOAM: The Self-Refining Autonomous OpenFOAM Agent

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.00003v1 Announcement Type: New.
      • Abstract: Computational Fluid Dynamics (CFD) plays a vital role in modern engineering, but using open-source solvers like OpenFOAM requires extensive knowledge and skills, as well as time-consuming configuration file setup.
  • To reduce this burden, we propose AutoFOAM - a self-evolving large language model (LLM) agent that creates, evaluates, runs, and evolves its own OpenFOAM simulations based solely on natural language instructions.

  • Our model is pre-trained on Qwen-coder 2.5-14B, which is then fine-tuned on 252 text prompts targeting 7 OpenFOAM solvers, 13 parameterized mesh templates, and y-plus-aware numerical strategies.

  • EN Highlights:

    • arXiv:2608.00003v1 Announce Type: new
    • Abstract: Computational Fluid Dynamics (CFD) plays an important role in modern engineering, but using open-source solvers such as OpenFOAM requires considerable…
    • To reduce this burden, we propose AutoFOAM - a self-evolving large language model (LLM) agent that creates, evaluates, runs, and evolves its own OpenFOAM simula…
    • Our model is pre-trained on the Qwen-coder 2.5-14B, which is then fine-tuned on 252 text prompts targeting 7 OpenFOAM solvers, 13 parametrized mesh templates, a…
  • Enhancing LLMs with Context-Specific Knowledge for Mitigating Misinformation in SMEs: A RAG-based Modeling and Analysis

    • Published: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.00006v1 Announce Type: new.
      • Abstract: Large Language Models (LLMs), as part of artificial intelligence (AI), are increasingly being adopted by Small and Medium Enterprises (SMEs) to enhance question-answering capabilities and support business decision-making processes.
      • However, hallucinations in LLM-generated outputs can serve as a source of misinformation, reducing user confidence in their reliability and trustworthiness within SMEs.
      • Retrieval-Augmented Generation (RAG) has emerged as a promising approach to address this challenge by incorporating external knowledge sources into the modeling process.
    • EN Highlights:
      • arXiv:2608.00006v1 Announce Type: new
      • Abstract: Large Language Models (LLMs), a part of artificial intelligence (AI), are increasingly being adopted by Small and Medium Enterprises (SMEs) to enhance…
      • However, hallucinations in LLM-generated outputs can serve as a source of misinformation, reducing user confidence in their reliability and trustworthiness with…
      • Retrieval-Augmented Generation (RAG) has emerged as a promising approach to address this challenge by incorporating external knowledge sources into the modeling…
  • Energy Efficiency of Locally Deployed LLMs: A Preliminary Quantitative GPU Power Benchmark on Consumer Hardware

    • Published: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.00008v1 Announce Type: new.
      • Abstract: Due to privacy concerns and the desire for local inference, the local deployment of Large Language Models (LLMs) is gaining attention.
  • However, the energy costs of consumer hardware remain poorly characterized, as most benchmarks focus solely on accuracy.

  • This paper presents a reproducible, hardware-level energy benchmark of nine open-source LLMs (1B to 7B parameters) executed on a single consumer GPU (RTX 4060Ti 16GB).

    • EN 要点:
      • arXiv:2608.00008v1 Announce Type: new
      • Abstract: The local deployment of large language models (LLMs) is gaining traction due to privacy concerns and the desire for on-premise inference
      • However, the energy costs on consumer hardware remain poorly characterized, as most benchmarks focus solely on accuracy
      • This paper presents a reproducible, hardware-level energy benchmark of nine open-source LLMs (1B to 7B parameters) executed on a single consumer GPU (RTX 4060Ti…
  • CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection

    • Published: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.00014v1 Announce Type: new.
      • Abstract: Evaluating Large Language Models (LLMs) incurs prohibitive computational overhead during continuous development processes.
      • While coreset selection accelerates evaluation, existing methods either suffer from a severe ``cold start’’ bottleneck requiring massive historical logs (e.g., item response theory), or exhibit superficial lexical bias, missing the underlying reasoning manifold of tasks.
      • We propose CoT-Core, a novel training-free core question selection framework.
    • EN 要点:
      • arXiv:2608.00014v1 Announce Type: new
      • Abstract: Evaluating Large Language Models (LLMs) incurs prohibitive computational overhead during continuous development processes
      • While coreset selection accelerates evaluation, existing methods either suffer from a severe ``cold start’’ bottleneck requiring massive historical logs (e.g.,…
      • We propose CoT-Core, a novel training-free core question selection framework
  • Optimization and Constraint Modeling using LLMs with a Retrieval Augmented Generation Process

    • Published: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.00015v1 Announce Type: new.
      • Abstract: Both optimization modeling and constraint modeling are non-trivial problems requiring deep domain expertise and proficiency in modeling formalism languages.
      • Despite their importance in logistics, healthcare, and supply chain management, current Large Language Models often generate structurally inconsistent or incomplete optimization formulations, especially in combinatorial settings.
      • This paper evaluates whether retrieval-augmented generation pipelines built upon curated synthetic datasets can meaningfully improve LLM optimization modeling performance.
    • EN 要点:
      • arXiv:2608.00015v1 Announce Type: new
      • Abstract: Both optimization modeling and constraint modeling are non-trivial problems requiring deep domain expertise and proficiency in modeling formalism lang…
  • Despite their importance across logistics, healthcare, and supply chain management, current large language models regularly produce structurally inconsistent or…

  • This paper evaluates whether a Retrieval-Augmented Generation pipeline built on a curated synthetic dataset can meaningfully improve LLM optimization modeling p…

  • Memory Reward Inflation in Self-Improving LLM Agents

    • Posted: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.00017v1 Announcement Type: New.
      • Abstract: Self-improving LLM agents are increasingly learning from experience without updating any weights.
      • Each episode is stored in an external memory, scored, and retrieved for future similar tasks to shape subsequent behavior.
      • From a reward perspective, the stored score serves as a proxy reward for an implicit, non-parametric policy.
    • EN Highlights:
      • arXiv:2608.00017v1 Announce Type: new
      • Abstract: Self-improving LLM agents increasingly learn from experience without updating any weights
      • Each episode is stored in an external memory, scored, and retrieved for similar future tasks to shape later behavior
      • Viewed through a reward lens, the stored score is a proxy reward for an implicit, non-parametric policy
  • Request-Level Energy Attribution for Batched LLM Serving

    • Posted: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.00026v1 Announcement Type: New.
      • Abstract: Batched LLM serving improves throughput but complicates energy accounting.
      • GPU power telemetry is aggregated, whereas sustainability reporting, chargebacks, and workload analysis often require request-level energy charges.
      • Existing inference energy benchmarks report energy at the model, phase, or token level, and recent carbon accounting efforts conceptually motivate Shapley fairness.
    • EN Highlights:
      • arXiv:2608.00026v1 Announce Type: new
      • Abstract: Batched LLM serving improves throughput but complicates energy accounting
      • GPU power telemetry is aggregate, whereas sustainability reporting, chargeback, and workload analysis often require request-level energy charges
      • Existing inference-energy benchmarks report model-, phase-, or token-level energy, and recent carbon-accounting work motivates Shapley fairness conceptually
  • Motif-Mamba: network motif improved mamba for long-range sequence modeling

    • Posted: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.00027v1 Announcement Type: New.
      • Abstract: Efficient long-sequence modeling remains a core challenge for large language models, as self-attention scales quadratically with sequence length.
  • Mamba offers a linear-time alternative through selective state space recurrence, but its predominantly diagonal state transitions restrict explicit interactions between state dimensions.

  • We propose Motif-Mamba, a structured state space model that augments Mamba with a motif-constrained low-rank recurrent pathway.

    • EN Highlights:
      • arXiv:2608.00027v1 Announce Type: new
      • Abstract: Efficient long-sequence modeling remains a central challenge for large language models, as self-attention scales quadratically with sequence length
      • Mamba offers a linear-time alternative through selective state space recurrence, but its predominantly diagonal state transitions restrict explicit interactions…
      • We propose Motif-Mamba, a structured state space model that augments Mamba with a motif-constrained low-rank recurrent pathway
  • Nova: An End-to-End MLIR Compiler for Deep Learning

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.00029v1 Announce Type: new.
      • Abstract: The performance of large-scale deep learning models depends heavily on how effectively high-level mathematical operations are mapped to the underlying physical hardware.
      • While high-level tensor frameworks provide flexible abstractions for model design, their eager execution models inherently lack the whole-graph visibility and granular control over hardware and memory to maximize native physical hardware utilization.
      • To bridge this gap, we designed Nova, an automated end-to-end JIT compiler whose defining purpose is to achieve absolute control over this hardware mapping: fusing operations across operator boundaries, optimizing complex memory hierarchies, and tuning execution down to the register level.
    • EN Highlights:
      • arXiv:2608.00029v1 Announce Type: new
      • Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physica…
      • While high-level tensor frameworks provide flexible abstractions for model design, their eager execution models inherently lack the whole-graph visibility and g…
      • To bridge this gap, we designed Nova, an automated end-to-end JIT compiler whose defining purpose is to achieve absolute control over this hardware mapping: fus…

ArXiv cs.CL (B_intro+search) Link to heading

  • TabletCraft: Bridging a 4,000-Year Cultural Gap with Bidirectional Akkadian NMT and Cuneiform Rendering

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.02609v1 Announce Type: new.
      • Abstract: Half a million cuneiform tablets are preserved in museums worldwide, yet modern users can neither read nor write in the world’s oldest writing system, leaving a 4,000-year cultural barrier that existing NLP tools have only partially addressed.
      • Previous work has achieved unidirectional, scholar-oriented translation from Akkadian to English, but offers no path in the reverse direction: non-expert users cannot compose new content in cuneiform and thus remain passive consumers of ancient culture, rather than active participants.
  • We introduce TabletCraft, the first open-source system that enables bidirectional interaction with Mesopotamian writing.

    • EN Highlights:
      • arXiv:2608.02609v1 Announce Type: new
      • Abstract: Half a million cuneiform clay tablets survive in museums worldwide, yet modern users can neither read nor write in the world’s oldest writing system,…
      • Prior work enables one-way, scholar-oriented translation from Akkadian to English, but offers no path in the reverse direction: non-specialist users cannot comp…
      • We present TabletCraft, the first open-source system that enables bidirectional interaction with Mesopotamian writing
  • BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.02612v1 Announce Type: new.
      • Abstract: Formulating an optimization problem strongly affects the quality of the final solution, and good formulations usually require substantial expertise.
      • Therefore, recent studies have examined how to automatically derive optimization problems from natural-language descriptions, but existing benchmarks focus on settings where objectives and constraints can be explicitly written as mathematical expressions.
      • Many practically important problems are naturally treated as black-box optimization (BBO) problems, in which only objective values are observable, and the functional form is not available.
    • EN Highlights:
      • arXiv:2608.02612v1 Announce Type: new
      • Abstract: Formulating an optimization problem strongly affects the quality of the final solution, yet good formulations usually require substantial expertise
      • Recent studies have therefore examined how to automatically derive optimization problems from natural-language descriptions, but existing benchmarks focus on se…
      • Many practically important problems are naturally treated as black-box optimization (BBO) problems, in which only objective values are observable, and the funct…
  • MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.02613v1 Announce Type: new.
      • Abstract: Personal memory assistants deployed on the edge must use open-weight models to process private, on-device interpersonal dialogues.
      • However, existing memory benchmarks often fail to adequately test the combination of activity-dense interactions, ego-centric perspectives, and coherent multi-session worlds.
      • MemArena fills these gaps with its single-world dialogue benchmark, constructed by its MASim agent simulator, covering 50 agents over 15 days (10.3 million conversational text tokens, 24,100 plain-text self-observation tokens/agent/day).
    • EN Highlights:
      • arXiv:2608.02613v1 Announce Type: new
  • Abstract: Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models

  • Yet, existing memory benchmarks often under-test the combination of activity-dense interaction, ego-centric perspective, and coherent multi-session worlds

  • MemArena fills these gaps with a single-world conversational benchmark built with its MASim agent simulator, for 50 agents over 15 days (10.3M dialog-text token…

  • OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.02615v1 Announce Type: new.
      • Abstract: Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata.
      • However, most medical large language model (LLM) and vision-language model (VLM) benchmarks focus on isolated modalities or narrow image-text tasks, leaving patient-level oncologic assessment across multiple evidence streams largely untested.
      • We introduce OncoTriad-QA, a patient-level radiology-pathology-genomics benchmark for pan-cancer question answering.
    • EN Key Points:
      • arXiv:2608.02615v1 Announce Type: new
      • Abstract: Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata
      • However, most medical large language model (LLM) and vision-language model (VLM) benchmarks focus on isolated modalities or narrow image-text tasks, leaving pat…
      • We introduce OncoTriad-QA, a patient-level radiology-pathology-genomics benchmark for pan-cancer question answering
  • Evaluating OpenAI’s Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.02616v1 Announce Type: new.
      • Abstract: We present the first independent, systematic evaluation of OpenAI’s Privacy Filter (OPF), a 1.5B-parameter bidirectional PII detector, across 42 comprehensive benchmarks spanning 22 languages and 5 domains.
      • Zero-shot, OPF achieves F1=0.855 on AI4Privacy and 0.464 on SPY Medical, outperforming Presidio (0.431, 0.273) and XLM-RoBERTa (0.269, 0.111) on PII-annotated benchmarks; on multilingual NER, XLM-RoBERTa leads OPF on all 13 Indic and non-Latin languages.
      • GPT-4o leads on medical, legal, and financial PII (SPY: avg 0.643, Gretel: 0.527), while OPF leads on structured synthetic PII (avg 0.71) and customer support (avg 0.60).
    • EN Key Points:
      • arXiv:2608.02616v1 Announce Type: new
  • Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.02617v1 Announcement Type: new.
      • Abstract: We use expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform that collects blind pairwise preferences alongside multi-criteria ratings, to evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation.
      • Clinicians assign scores on a discrete $[-2, +2]$ scale, where negative values indicate clinically unsafe or misleading content.
      • Using 26,804 pairwise judgments from 13 LLMs, contributed by over 736 clinicians from more than 28 countries, we find that clinician preferences do not well represent safety-critical performance.
    • EN Key Points:
      • arXiv:2608.02617v1 Announce Type: new
      • Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert…
      • Clinicians assign scores on a discrete $[-2, +2]$ scale, where negative values indicate clinically unsafe or misleading content
      • Using 26{,}804 pairwise judgments across outputs from 13 LLMs, contributed by more than 736 clinicians across 28+ countries, we find that clinician preference i…
  • JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.02620v1 Announcement Type: new.
      • Abstract: LLM-as-a-judge evaluation has become the dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks have their own codebase, hardcode specific closed-model judges, and support a single evaluation protocol.
      • This fragmentation makes it difficult to understand how research design choices (benchmark, judge model, prompt, inference backend) affect the conclusions we draw about model quality.
      • We introduce JudgeArena, an open-source framework that unifies major LLM-judge benchmarks (AlpacaEval, Arena-Hard, MT-Bench, and m-Arena-Hard) under a single interface with swappable judges and comprehensive metadata logging for transparent reporting and reproducibility.
    • EN Key Points:
      • arXiv:2608.02620v1 Announce Type: new
  • Abstract: LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their…

  • This fragmentation makes it difficult to study how design choices–the benchmark, the judge model, the prompt, the inference backend–affect the conclusions we…

  • We introduce JudgeArena, an open-source framework that unifies major LLM-judge benchmarks (AlpacaEval, Arena-Hard, MT-Bench, and m-Arena-Hard) under a single in…

  • Knowing the Form, Not the Function: Automatically Auditing Answer–Authority Decoupling in Legal Benchmarks

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.02621v1 Announce Type: new.
      • Abstract: Legal benchmarks typically score final answers even when models also state legal authority.
      • We test whether answer correctness can serve as a proxy for authority grounding.
      • Under ordinary reasoning prompts that did not request statutory citations, four LLMs spontaneously produced authority markers across 238 Taiwan bar-examination items.
    • EN Key points:
      • arXiv:2608.02621v1 Announce Type: new
      • Abstract: Legal benchmarks typically score final answers even when models also state legal authority
      • We test whether answer correctness can serve as a proxy for authority grounding
      • Under ordinary reasoning prompts that did not request statutory citations, four LLMs spontaneously produced authority markers across 238 Taiwan bar-examination…
  • Speculative Correction: Draft-then-Refine Decoding for Diffusion Language Models

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.02625v1 Announce Type: new.
      • Abstract: Diffusion language models (DLMs) can revise tokens bidirectionally, but standard decoding procedures often adapt them to left-to-right generation by producing text chunk-by-chunk.
      • We study a simple plug-and-play inference pattern: first generate a complete draft, then refine the full response using bidirectional diffusion.
      • Using LLaDA2.1-Flash and LLaDA2.1-Mini, we evaluate two configurations.
    • EN Key points:
      • arXiv:2608.02625v1 Announce Type: new
      • Abstract: Diffusion language models (DLMs) can revise tokens bidirectionally, but standard decoding procedures often adapt them to left-to-right generation by p…
      • We study a simple plug-and-play inference pattern: first generate a complete draft, then refine the full response using bidirectional diffusion
  • Using LLaDA2.1-Flash and LLaDA2.1-Mini, we evaluate two configurations

  • Stuck on “A”: Diagnosing and Repairing Interface Injury in Attention-to-KDA Linearization of a 0.6B Language Model

    • Posted: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.02689v1 Announce Type: new.
      • Abstract: We convert 21 of 28 full-attention layers of Qwen3-0.6B-Base into KDA (Kimi Delta Attention) linear-attention layers on a single consumer-grade GPU budget and ask a simple question: what exactly does the conversion break?
      • After the surgery, hidden-state alignment and end-to-end KL distillation bring the student close to its teacher in perplexity, yet multiple-choice accuracy remains near random chance (25-29% vs.
      • the teacher’s 50.6% on C-Eval).
    • EN Key Points:
      • arXiv:2608.02689v1 Announce Type: new
      • Abstract: We convert 21 of 28 full-attention layers of Qwen3-0.6B-Base into KDA (Kimi Delta Attention) linear-attention layers on a single consumer-grade GPU bu…
      • After surgery, hidden-state alignment and end-to-end KL distillation drive the student close to its teacher in perplexity, yet multiple-choice accuracy stays ne…
      • the teacher’s 50.6% on C-Eval)

ArXiv cs.LG (B_intro+search) Link to heading

  • Deep Divide-and-Reduce in Symbolic Regression

    • Posted: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.02628v1 Announce Type: new.
      • Abstract: Symbolic regression (SR) is the task of discovering underlying patterns from data and representing them using mathematical expressions.
      • Current machine learning approaches to SR often lack a profound understanding of the intrinsic mathematical and physical principles governing these expressions.
      • While the pioneering AI Feynman method leverages the mathematical properties underlying the data, its expression simplification mechanism has a narrow scope of applicability and tends to fail on complex equations.
    • EN Key Points:
      • arXiv:2608.02628v1 Announce Type: new
      • Abstract: Symbolic regression (SR) is the task of discovering underlying patterns from data and representing them using mathematical expressions
      • Current machine learning approaches to SR often lack a profound understanding of the intrinsic mathematical and physical principles governing these expressions
      • While the pioneering AI Feynman method leverages the mathematical properties underlying the data, its expression simplification mechanism suffers from a narrow…
  • Multimodal Auto-regressive Transformer Surrogate for Modeling Variable Operations and Quantifying Uncertainty in Geological Carbon Storage

    • Posted: 2026-08-05 12:00 Beijing Time
  • Abstract:- arXiv:2608.02629v1 Announce Type: new.

    • Abstract: The use of variable well perforation and injection strategies can improve the efficiency of geological carbon storage operations.
    • We develop a new multimodal auto-regressive transformer surrogate to model these operations under geological uncertainty.
    • A modified SEAM CO2 geomodel, which involves a faulted system with three stacked aquifers, is considered.
    • EN Point:
      • arXiv:2608.02629v1 Announce Type: new
      • Abstract: The use of variable well perforation and injection strategies can improve the efficiency of geological carbon storage operations
      • We develop a new multimodal auto-regressive transformer surrogate to model these operations under geological uncertainty
      • A modified SEAM CO2 geomodel, which involves a faulted system with three stacked aquifers, is considered
  • LLMs Can Annotate Attribution Graphs

    • Published: 2026-08-05 12:00 Beijing Time
    • Abstract:- arXiv:2608.02632v1 Announce Type: new.
      • Abstract: Circuit tracing is an exciting technique for revealing the internal computation of language models, but it requires a time-intensive manual step of grouping individual features or MLP neurons into supernodes.
      • We present a simple pipeline for automating this step: directly presenting feature descriptions to a language model that groups them into supernodes.
      • Using automated interpretability metrics, we confirm that supernodes generated by our pipeline are as interpretable as those generated by human annotators.
    • EN Point:
      • arXiv:2608.02632v1 Announce Type: new
      • Abstract: Circuit tracing is an exciting technique for revealing the internal computation of language models, but it requires a time-intensive manual step of gr…
      • We present a simple pipeline for automating this step: directly presenting feature descriptions to a language model that groups them into supernodes
      • Using automated interpretability metrics, we confirm that supernodes generated by our pipeline are as interpretable as those generated by human annotators
  • GeoID-PINN: Identifiability-Aware Regional Epidemic Inference with Geographic Coupling

    • Published: 2026-08-05 12:00 Beijing Time
    • Abstract:- arXiv:2608.02633v1 Announce Type: new.
      • Abstract: Regional surveillance data reflects local transmission, reporting, propagation, and external infection pressure, which are difficult to identify separately.
      • We introduce GeoID-PINN, a physics-informed neural network (PINN) for Susceptible-Infected-Recovered-Deceased (SIRD) dynamics.
      • The model represents spatial dependencies with a row-stochastic source composition matrix, whose rows specify non-negative source weights that sum to one.
    • EN Point:
      • arXiv:2608.02633v1 Announce Type: new
  • GeoID-PINN: A Physics-Informed Neural Network for SIRD Dynamics on Geometric Graphs

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.02661v1 Announcement Type: New.
      • Abstract: Regional surveillance data reflect local transmission, reporting, seeding, and external infection pressure, which are difficult to identify separately.
      • We introduce GeoID-PINN, a physics-informed neural network (PINN) for susceptible-infectious-recovered-deceased (SIRD) dynamics.
      • The model represents spatial dependence with a row-stochastic source-composition matrix whose rows assign nonnegative source weights that sum to one.
    • EN Highlights:
      • arXiv:2608.02661v1 Announce Type: new
      • Abstract: Regional surveillance data reflect local transmission, reporting, seeding, and external infection pressure, which are difficult to identify separately
      • We introduce GeoID-PINN, a physics-informed neural network (PINN) for susceptible-infectious-recovered-deceased (SIRD) dynamics
      • The model represents spatial dependence with a row-stochastic source-composition matrix whose rows assign nonnegative source weights that sum to one
  • Verifier-Guided Model Discovery for Physical Dynamical Systems with Pretrained Symbolic Transformers

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.02662v1 Announcement Type: New.
      • Abstract: Reliable forecasting of nonlinear physical systems underpins scientific discovery and engineering decision-making.
      • However, high-fidelity simulations are prohibitively costly, and machine-learning surrogates can be opaque and encode assumptions about system dynamics, limiting universality.
      • Pretrained Transformers that map synthetic ODE trajectories to equations offer an interpretable alternative, promising transfer without system-specific equation knowledge.
    • EN Highlights:
      • arXiv:2608.02662v1 Announce Type: new
      • Abstract: Reliable forecasting of nonlinear physical systems underpins scientific discovery and engineering decision-making
      • Yet high-fidelity simulations are prohibitively costly, and machine-learning surrogates can be opaque and encode assumptions about system dynamics, limiting gen…
      • Pretrained transformers mapping synthetic ODE trajectories to equations offer interpretable alternatives, promising transfer without system-specific equation kn…
  • CT-HEG: A Bidirectional, Timestamp-Attributed Event Graph for ICU In-Hospital Mortality Prediction - An Architectural Ablation Study

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.02663v1 Announcement Type: New.
      • Abstract: Accurate ICU mortality prediction requires modeling irregular clinical observations across heterogeneous entity types.
      • Existing sequential models handle irregular sampling but ignore typed relational structures; existing graph models assume fixed-interval inputs.
      • We introduce the Continuous-Time Heterogeneous EHR Graph (CT-HEG) schema and evaluate which architectural choices drive predictive performance.
    • EN Highlights:
      • arXiv:2608.02663v1 Announce Type: new
      • Abstract: Accurate ICU mortality prediction requires modeling irregular clinical observations across heterogeneous entity types
  • Existing sequence models handle irregular sampling but ignore typed relational structure; existing graph models assume fixed-interval inputs

  • We introduce the Continuous-Time Heterogeneous EHR Graph (CT-HEG) schema and evaluate which architectural choices drive predictive performance

  • Sphere Retraction Normalizations

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Summary: - arXiv:2608.02668v1 Announcement Type: new.
      • Abstract: Residual connections are the de facto mechanism for stably training deep neural networks.
      • Geodesic Normalization (GeoNorm) recasts them on a Riemannian manifold, orthogonalizing each layer output against the current hidden state and applying the resulting update via the Riemannian exponential map.
      • Consequently, each hidden state maintains a constant $\ell_{2}$-norm, confining the residual stream to a hypersphere.
    • EN Highlights:
      • arXiv:2608.02668v1 Announce Type: new
      • Abstract: Residual connections are the de facto mechanism for training deep neural networks stably
      • Geodesic Normalization (GeoNorm) recasts them on a Riemannian manifold, orthogonalizing each layer output against the current hidden state and applying the resu…
      • Every hidden state thus keeps a constant $\ell_{2}$-norm, confining the residual stream to a hypersphere
  • Learning Molecular Representations from Cellular Phenotypes with Structure Preservation

    • Publication Time: 2026-08-05 12:00 Beijing Time
    • Summary: - arXiv:2608.02688v1 Announcement Type: new.
      • Abstract: Phenotypic drug discovery enables the discovery of functional relationships between molecular structures and cellular responses.
      • However, existing multimodal representation learning methods often optimize cross-modal alignment without considering the intrinsic organization of the chemical space, leading to distorted molecular representations and loss of structural information.
      • We propose \textbf{PhenMol}, a structure-preserving framework for phenotype-aware molecular representation learning.
    • EN Highlights:
      • arXiv:2608.02688v1 Announce Type: new
      • Abstract: Phenotypic drug discovery enables the discovery of functional relationships between molecular structures and cellular responses
      • However, existing multimodal representation learning methods often optimize cross-modal alignment without considering the intrinsic organization of chemical spa…
      • We propose \textbf{PhenMol}, a structure-preserving framework for phenotype-aware molecular representation learning
  • GLOBE: Trajectory-Aligned Gradient Matching with Structured SparseOptimization for Coreset Selection

    • Publication Time: 2026-08-05 12:00 Beijing Time
  • Abstract: - arXiv:2608.02690v1 Announce Type: new.

    • Abstract: On-device training of deep neural networks is fundamentally constrained by the computational and memory costs of large-scale datasets.
    • Coreset selection offers a practical solution by retaining only a compact subset of real training samples.
    • However, existing gradient-based methods commonly rely on gradients computed at a single model snapshot and employ greedy or pursuit-based selection procedures, limiting their ability to capture evolving optimization dynamics and handle strongly correlated samples.
  • Output-Aware Rotation for INT2 KV-Cache Quantization

    • Published: 2026-08-05 12:00 Beijing Time
    • Abstract: - arXiv:2608.02691v1 Announce Type: new.
    • Abstract: The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-low-bit quantization increasingly important.
    • However, existing rotation-based INT2 methods optimize cache statistics or proxy errors before the complete attention readout, even though the model is ultimately affected by errors propagated through attention and the output projection $W_O$.
    • To address this mismatch, we propose \textit{OptR}, an output-aware rotation method that minimizes post-$W_O$ attention-output error.