🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-08-05
- 类型
- ai-daily
- 字数
- 8507
- 阅读时长
- 40 min
2026-08-05 AI Daily | DeepSeek and Qwen Accelerate in Sync, AI Competition Shifts to Cost-Effectiveness and Deployability Link to heading
Today’s focus is on the deepening of model competition: DeepSeek V4 Flash and Qwen3.8-Max are bringing “high performance, low cost, and distributability” to the forefront. Meanwhile, Agent competition is shifting from chat to memory, routing, and stable execution, while evaluation and safety are re-emerging as core industry infrastructure.
📖 In-depth Guide to This Issue’s Watch List Link to heading
There are three main threads worth following today. First, Agents are evolving from “able to chat” to “able to act.” The local execution assistant from OpenClaw, the long-term memory designs of MemoryForge/AgentMemBench, and explorations into routing and role control like “SLMs as Multi-Agent Routers” and “Role Steering” collectively indicate that the next phase of competition is not just about model capabilities, but also about memory, scheduling, and behavioral stability. Second, evaluation and safety are returning to the forefront. OpenAI’s third-party network evaluation revealed boundary issues arising from the interaction between test configurations and model capabilities. Combined with low-cost automated grading and tools like RubricReviewer, this suggests the industry is beginning to systematically address “how to reliably evaluate models.” Third, the efficiency narrative behind Microsoft’s financial reports, along with vertical applications like VLM scaling laws, disaster remote sensing, and AutoFOAM, show that AI commercialization is shifting from general-purpose demos to practical, measurable productivity returns.
🌐 AI Hot Topics on X Link to heading
Topic 1: DeepSeek V4 Flash Surges with Frontier Performance at Tiny Cost Link to heading
- Category: AI · News
- Overview: Trending: 2 days ago, Related Posts: 25,000
- What it is: DeepSeek V4 Flash is gaining significant attention for its performance on multiple benchmarks, which approaches that of frontier models but at a significantly lower inference and usage cost.
- Why it’s important: This is important because it once again pushes the “performance-to-cost ratio of large models” to the core of the industry, potentially changing enterprise model selection, compute investment, and the competitive landscape between open-source and closed-source models.
- Discussion summary: Discussions on X are focused on whether its performance is truly reproducible, the extent of the gap compared to mainstream flagship models, and whether “low-cost, high-performance” signals the start of a more intense price war for large models.
Topic 2: Apple Seeks Court Order to Inspect OpenAI Devices Over Trade Secrets Claims Link to heading
- Category: AI · News
- Overview: Trending: 17 hours ago, Related Posts: 17,000
- What it is: Apple is seeking a court order to inspect OpenAI’s devices or related materials to support its claims in a lawsuit concerning trade secrets.
- Why it’s important: Such disputes involve trade secrets, evidence discovery, and competitive boundaries between AI companies, and could influence how compliance and evidence gathering for models, devices, and data are handled within the industry.
- Discussion summary: Discussions on X are centered on whether Apple has sufficient grounds to demand an inspection of OpenAI devices, whether this move is a normal legal discovery process or excessive pressure, and how to balance trade secret protection with judicial transparency.
Topic 3: Alibaba Launches Qwen3.8-Max as Top Coding AI Model Link to heading
- Category: AI · News
- Overview: Trending: 1 day ago, Related Posts: 47,000
- What it is: Alibaba has released Qwen3.8-Max, claiming it to be the new top AI model for programming, with plans to open-source its weights in the future.
- Why it’s important: This marks a shift for China’s large models from one-off releases to a deployable, distributable infrastructure model, which could reshape global AI competition in terms of model capability, pricing, and the proliferation of open-source weights.
- Discussion summary: Discussions on X are focused on three points: the credibility of Alibaba’s performance benchmarks, whether enterprises and developers can quickly deploy Qwen once its weights are released, and whether Chinese models like it and DeepSeek have already established a competitive advantage in cost and distribution speed.
Topic 4: OpenAI Acquires Ona to Power Next-Gen AI Agents Beyond Laptops Link to heading
- Category: AI · News
- Overview: Trending: 19 hours ago, Related Posts: 6,300
- What it is: OpenAI announced the acquisition of Ona, aiming to support the capabilities of next-generation AI agents and expand their application scenarios beyond laptops.
- Why it’s important: This signifies that OpenAI is accelerating its development of more capable AI agents, pushing them from chatbots to practical, cross-device, cross-scenario systems, which is crucial for both AI product forms and ecosystem competition.
- Discussion summary: Discussions on X are mainly focused on whether this acquisition will enable OpenAI to launch AI agents for mobile phones, wearables, or at the OS level more quickly, and how this “beyond the laptop” direction will change human-AI interaction. There is also interest in how the Ona team and its technology will be integrated post-acquisition.
Topic 5: Immunologist Calls OpenAI’s GPT-5.6 Pro Smartest AI Model Yet Link to heading
- Category: AI · News
- Summary: Trending time: 23 hours ago, Related posts: 82
- What it is: An immunologist on X called OpenAI’s GPT-5.6 Pro “the smartest AI model to date.”
- Why it matters: If this assessment becomes widely accepted, it signifies another potential leap in large model capabilities, influencing the industry’s perception of OpenAI’s technological leadership, professional evaluation standards, and the application boundaries of these models.
- Discussion summary: The discussion on X revolves around whether this claim is an exaggeration, whether GPT-5.6 Pro shows significant improvement over its predecessors, and the real-world experiences of users in professional fields like medicine regarding its reasoning, accuracy, and practicality.
Topic 6: Wife’s Watch Slide Deck Post Highlights Men’s Niche Passions Link to heading
- Category: AI · Entertainment
- Summary: Trending time: , Related posts: 134
- What it is: A wife shared a slide deck post on X about her husband’s niche interests, sparking discussions and circulation around “men’s niche passions.”
- Why it matters: This type of content reflects how generative AI and presentation tools are being used for everyday storytelling, personal expression, and social sharing. It also shows that AI-driven content production is penetrating entertainment scenarios.
- Discussion summary: The main discussion on X centers on whether the slide deck is cute and fun or overly elaborate, whether men’s niche hobbies are worth documenting seriously, and whether the post was created with AI assistance or has deliberate marketing elements.
Topic 7: Chamath Palihapitiya Sweater Meme Takes Off with Grok Link to heading
- Category: AI · Entertainment
- Summary: Trending time: 1 day ago, Related posts: 2400
- What it is: Chamath Palihapitiya’s “sweater” meme rapidly gained traction on X, with users employing Grok for further derivative creation and dissemination.
- Why it matters: This event highlights the involvement of AI chatbots/generative tools in the spread of social media memes, reflecting AI’s evolution from a simple Q&A tool to a tool for content creation and cultural transmission.
- Discussion summary: The discussion on X is mainly focused on whether this is just an entertaining meme, the quality and humor of the content generated by Grok, and whether AI is accelerating the creation and amplification of viral memes.
Topic 8: Perceptis Tops Design Arena’s Corporate Slides Ranking Link to heading
- Category: AI · Entertainment
- Summary: Trending time: 5 hours ago, Related posts: 362
- What it is: Perceptis reached the top of Design Arena’s corporate slide deck rankings, attracting attention on X.
- Why it matters: This indicates that generative AI is making further inroads into corporate presentation documents and design automation, impacting office content production efficiency, competition among design tools, and enterprise-level adoption.
- Discussion summary: The discussion focuses on whether its slide decks are genuinely more aesthetically pleasing and professional, and whether it surpasses other AI design tools in terms of editability, efficiency, price, and practical effectiveness in a corporate setting.
Topic 9: Tesla Model Y L Stuns Reviewers with Roomy Design and FSD Prowess Link to heading
- Category: AI · Entertainment
- Summary: Trending time: 1 day ago, Related posts: 7600
- What it is: The Tesla Model Y L has sparked heated discussions on X due to its more spacious design and FSD (Full Self-Driving) performance, with some reviewers claiming the experience exceeded their expectations.
- Why it matters: This is significant because it involves autonomous driving capabilities powered by large models, end-to-end perception and decision-making, and the real-world implementation of AI in mass-produced vehicles.
- Discussion summary: The discussion on X mainly centers on whether FSD is truly mature enough, whether the space and comfort of the Model Y L are noteworthy, and its pricing, delivery, and competitiveness in different markets.
AI Public Opinion Summary on X Today Link to heading
The main theme on X today is that the AI competition is shifting from “who has the strongest model” to “who can implement it at a lower cost, faster, and truly integrate it into products and scenarios.” Discussions around DeepSeek, Qwen, and OpenAI have formed a relatively clear consensus: near-state-of-the-art performance is merely the entry ticket; cost, deployability, open weights, and agent capabilities are the key differentiators in the next phase. The points of contention are mainly: whether the high-scoring performance of these models is reproducible, whether official benchmarks are exaggerated, which route (open-source vs. closed-source) has the upper hand, and whether the practices of companies like OpenAI and Apple regarding trade secrets and evidence gathering are reasonable. The potential risks are also clear: first, model and product marketing may continue to be amplified by the “benchmark narrative”; second, price wars and competition for computing power will intensify; and third, if AI agents, autonomous driving, and cross-device applications advance too quickly, they could simultaneously magnify security, privacy, and compliance issues.
💡 Influencer Insights Link to heading
Alright, based on the activity of various AI influencers on the X platform over the past 24 hours, here is the distilled industry analysis report.
AI Daily (August 4-5, 2026) Link to heading
Analyst: AI Industry Observer Data Source: @zhixianio, @Pluvio9yte, @dotey, @vista8, @gefei55, @ruanyf, et al.
1. Core Trends & Product Highlights Link to heading
💡 Clear Division of Labor in Model Strategy: Shifting from “Strongest” to “Best-Fit” Link to heading
Industry leaders are moving away from a “one-size-fits-all” single-model approach, instead delving into the “personality” and “division of labor” of different models. @dotey shared his practical strategy, which is highly representative:
- Fable 5 (Anthropic): Responsible for design plans and review acceptance. Although its reasoning quality is high, it is expensive and slow, making it suitable for the top-level design of complex technical solutions.
- GPT-5.6 Sol (OpenAI): Serves as the main workhorse for handling the “grunt work” due to its high cost-effectiveness and its tendency not to overthink even at xhigh inference intensity. However, one must be aware of its tendency to take “shortcuts” (e.g., secretly reducing decoding precision to optimize performance).
- Opus 4.6 (Anthropic): For creative tasks like writing, it’s widely recognized for its superior “style” and prose. Even long after its release, it remains the top choice in this niche, a sentiment echoed by @dotey and international user @petergyang.
🚀 Agent Evolution in “Hand-Brain Coordination”: Ending Session Handoff Anxiety Link to heading
Context management and task continuity for Agents have become a new focus.
- Seamless Handoff: Addressing the pain point of having to “start a new session” to save tokens, @dotey points out that modern Agents (like Codex) have sufficiently strong context compression capabilities. This can be achieved with the
/compactcommand, eliminating the need to frequently start new conversations. For cross-Agent collaboration, the recommended SOP is: “Fable 5 writes the technical documentation → Codex reads the file and executes → Fable 5 performs acceptance testing.” - Structured Handoff Skill: @Pluvio9yte recommended the Codex
/hand offSkill developed by Matt Pocock. This skill can generate a complete handoff document when a session is full, allowing a new session to read it directly. This completely solves the problem of potentially missing key points when manually scanning conversation history.
💻 On-Device Models and Local Deployment Enter the “Sweet Spot” Link to heading
The cost-effectiveness and usability of running large models locally are being redefined:
- Hardware Choices Challenge Perceptions: @ruanyf clearly states that running AI locally isn’t limited to the RTX 5090. A mini PC with a Strix Halo chipset (e.g., equipped with 128GB of unified memory) could be a more cost-effective and feasible option in many scenarios.
- The “Capability Ceiling” of Small Models: Through hands-on testing of Google’s Gemma 4 12B Coder, @zhixianio found that while fine-tuning can improve efficiency, 12B models have a natural ceiling when generating “long, stateful” complex programs (like Tetris). There is a significant gap compared to the 35B MoE model he uses daily.
- DeepSeek V4 Running Locally: @zhixianio successfully ran the 4-bit quantized version of DeepSeek V4 Flash on a Mac Studio, indicating that top-tier models are rapidly becoming accessible on consumer-grade hardware.
- DeepSeek Earns Community Respect: @vista8 quoted the vLLM team’s podcast, stating that DeepSeek is “currently the model people dare to use extensively in consumer-facing production scenarios.” Its extreme pricing strategy has also gained high recognition in the international community.
🎮 AI Game Generation: Democratizing Creation by Turning a Sentence into a Game Link to heading
@Pluvio9yte discovered and recommended the platform @makeplayai. Users can complete the entire development process of a game—including art, sound effects, animations, and code—in the browser with just a single natural language command. Its “branching development” feature allows users to compare the actual gameplay experience of different ideas, lowering the barrier for the entire cycle from game creation to testing.
2. Unique Perspectives & Industry Outlook Link to heading
- The Right Way to Do Code Review: @Pluvio9yte emphasizes that the key to AI Code Review is to “open a new window, without context, or use a different model” to avoid the limitations of the original line of thought. He summarized a five-step review method that gets straight to the point: “find bugs, find missed requirements, find unnecessary complexity, find missing tests, and suggest deletion or simplification.”
- Agents need “their own computers” not “containers”: @dotey relayed the view that the most powerful future Agents will require real computers, as global computing power would be insufficient to support hundreds of millions or even billions of concurrent Agents each exclusively occupying container environments. This presages a shift in Agent form from cloud-based sandboxes to personal devices.
- AI amplifies your intentions: @vista8 shared a concise insight: “Using AI with fear amplifies fear. Using AI with curiosity amplifies curiosity.” This emphasizes the decisive role of the user’s mindset and guidance in the “human + AI” collaborative model.
- AI Agent proxy paradigm review: @gefei55 highly praised Manus for redefining industry product forms, believing that its “cloud virtual machine + Agent” model, and the subsequent derivative localized Agents (OpenClaw, WorkBuddy), all originated from Manus’s keen judgment of AI capabilities at that time, directly leading to the strategic shift of AI browsers towards desktop clients.
- The social nature of Skills: @ruanyf noted that Xiaohongshu launched REDSkill, allowing users to upload and share Agent Skill files in notes, aiming to create a Skill version of GitHub. This suggests that Skill creation and distribution are becoming a new form of social media content.
- Reflections on AI: @lijigang’s sharing was philosophical, proposing: “LLM tokens are the calories of thought,” reminding people to pay attention to the quality of information input into AI and the brain; and believing that “prediction is a powerful selection pressure,” people should learn from LLMs to predict the next Token, forcing themselves to understand and predict the next step in their field.
3. Recommended Tools and Resources Link to heading
| Type | Tool/Resource | Recommender & Key Highlights |
|---|---|---|
| Development Framework | Meta Skill (Meta-Skill) | @vista8: Open-source “Skill that generates Skills,” integrating numerous best practices and data sources. It generates extremely high-quality Skills, supports publishing to GitHub and generating npx commands, greatly simplifying Skill development and sharing. |
| Development Framework | Harness Learning List | @dotey: Recommended a “production-grade Harness source code learning list,” with a special reminder that “thoroughly understanding one (e.g., pi-mono) is better than glancing at each,” making it an excellent resource for deeply understanding how Agents work. |
| Agent Management | Codex Handoff Skill | @Pluvio9yte: Used to generate a complete handover document for a new session when the Codex session context is almost full, avoiding the omission of key points when scanning historical information. |
| Agent Management | OpenConnector | @ruanyf: Open-source password connection gateway that prevents AI Agents from leaking passwords into the context. It acts as a unified authorization middleware, where Agents only receive execution results, making it very secure. |
| AI Games | Makeplay (AI Game Builder) | @Pluvio9yte: A free game platform where you input a sentence, and AI automatically generates all art, sound effects, animations, and code. It supports branching development and offers an excellent experience. |
| Payment Monetization | PayPal CN (Personal Account) | @gefei55: Revealed that PayPal’s domestic platform now supports registration with domestic personal ID cards and allows websites to collect USD from global users. This was driven by him and is a major boon for domestic independent developers expanding overseas. |
| Hardware Management | BeeSIM Bluetooth Card Writer | @AI_Jasonyu: For users managing multiple eSIM cards, BeeSIM is recommended for unified management with a mini-program, offering lower cost and more stability than the小白卡 solution. |
| Office AI | Tencent Cloud CodeBuddy NPC | @ruanyf: Allows calling AI models as “NPCs” on code hosting platforms to perform code operations, offering a novel way to interact. |
📚 Appendix: Today’s Watch List Update Sources Link to heading
Time Window: Most recent 3 days; Covering 22 sources; Total 34 updates
Y Combinator Podcast (B_intro+search) Link to heading
- Waymo Co-CEO Dmitri Dolgov: “Move Fast And Ship Safely”
- Published: 2026-08-05 00:55 Beijing Time
- Summary: - You may have already heard of OpenClaw (formerly known as Clawdbot/Moltbot).
- The open-source AI assistant that’s causing a stir runs on your own device, connects with the messaging apps you already use, and goes beyond chat to actually perform tasks like managing your email, calendar, files, and workflows.
- Now, meet the person behind it.
- YC’s Raphael Schaad sits down with OpenClaw founder Peter Steinberger to talk about the “aha” moment behind the viral personal AI agent, why local-first agents could replace many of today’s apps, and how personal agents will reshape the future of software.
- EN Highlights:
- Waymo’s first autonomous demo took eighteen months
- The product took fifteen years
- Today, the Waymo Driver runs 500,000 trips a week — four million fully autonomous miles across fifteen cities, with 17 times fewer serious-injury crashes than h…
- At Startup School 2026, Waymo co-CEO Dmitri Dolgov shares the seven lessons behind that journey, from bridging the gap between a demo and a real product to buil…
Stratechery by Ben Thompson (A_full) Link to heading
- Microsoft Earnings, Microsoft vs. Meta, The Efficiency Payoff
- Release Time: 2026-08-04 18:00 Beijing Time
- Abstract: - Microsoft’s earnings were compelling because they showed a clarity of strategy, lower costs, and a tangibility of application.
- The reason why is scarier.
- $15/month* or *$150/year.
- Substantive analysis of the day’s news via three weekly emails or a podcast.
- Strategy Interviews.
- EN Highlights:
- Microsoft’s earnings were compelling because they showed a clarity of strategy, lower costs, and a tangibility of application
- The reason why is scarier.
OpenAI Blog (A_full) Link to heading
Third-party cyber evaluations involving OpenAI models
- Release Time: 2026-08-05 03:00 Beijing Time
- Abstract: - Independent testing plays a crucial role in helping us validate and further understand risks before deployment.
- Some cyber evaluations intentionally use custom configurations, including reduced safeguards that measure underlying capabilities, rather than how the model typically behaves in publicly available deployments.
- In recent evaluations, two external testing partners identified incidents where test configurations and controls, combined with the advanced capabilities of the latest models, allowed model activities to extend beyond their intended testing boundaries.
- These incidents underscore the importance of working across the industry and with third-party evaluators to establish testing environments and practice standards as models become more powerful.
- The new incidents involve OpenAI models accessing the public internet under specific conditions during third-party cyber evaluations, with reduced security configurations that do not reflect normal deployment.
- EN Highlights:
- OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evaluation.
New ways to learn and teach with ChatGPT Work and Codex
- Release Time: 2026-08-04 08:00 Beijing Time
- Abstract: - Artificial intelligence is evolving from tools that primarily answer questions to systems that can reason across contexts, use other tools, and help with complex, multi-step work.
- This shift is changing what it means to be prepared for school, work, and whatever comes next.
As students and educators return to classrooms and campuses this autumn, we are launching three new education plugins for ChatGPT Work and Codex, specifically designed to help students and educators leverage agent capabilities using their chosen course materials and context.
Plugins are a set of applications, role-specific skills, instructions, and general workflows that help students and educators get started immediately without having to build complex prompts themselves.
These new plugins are available for deployment through ChatGPT Edu and ChatGPT for Teachers sections.
- EN Key Points:
- Explore new education plugins for ChatGPT Work and Codex that help K–12 teachers, college educators, and students learn, teach, research, and build.
- EN Key Points:
ArXiv cs.AI (B_intro+search) Link to heading
Revisiting Classic Thought Experiments to Measure Consciousness for Artificial Intelligence Safety
- Publication Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00001v1 Announcement Type: New.
- Abstract: This research note revisits Leibniz’s Mill, Turing’s Imitation Game, and Searle’s Chinese Room through the Conservation-Congruent Encoding (CCE) framework.
- It formalizes a toy symbolic setting in which successful behavior is measured by task performance ($W_{causal,T}$), while the efficiency with which preserved internal structure supports that behavior is measured by operational awareness ($\kappa_T$).
- In this setting, uncompressed lookup systems and compact generative systems can in principle achieve similar behavioral success, but differ greatly in $\kappa_T$: the former relies on extended permanent storage of un-reused mappings, while the latter reuses compact internal structures.
- EN Key Points:
- arXiv:2608.00001v1 Announce Type: new
- Abstract: This research note revisits Leibniz’s mill, Turing’s imitation game, and Searle’s Chinese Room through the Conservation-Congruent Encoding (CCE) frame…
- It formalises a toy symbolic setting in which successful behaviour is measured by task performance ($W_{causal,T}$), while the efficiency with which preserved i…
- Within this setup, an uncompressed lookup system and a compact generative system can in principle achieve comparable behavioural success, yet diverge sharply in…
AutoFOAM: The Self-Refining Autonomous OpenFOAM Agent
- Publication Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00003v1 Announcement Type: New.
- Abstract: Computational Fluid Dynamics (CFD) plays a vital role in modern engineering, but using open-source solvers like OpenFOAM requires extensive knowledge and skills, as well as time-consuming configuration file setup.
- To alleviate this burden, we propose AutoFOAM - a self-evolving Large Language Model (LLM) agent that creates, evaluates, runs, and develops its own OpenFOAM simulations solely based on natural language instructions.
- Our model is pre-trained on Qwen-coder 2.5-14B and then fine-tuned on 252 text prompts for 7 OpenFOAM solvers, 13 parameterized mesh templates, and y-plus-aware numerical strategies.
- EN Key Points:
- arXiv:2608.00003v1 Announce Type: new
Abstract: Computational Fluid Dynamics (CFD) plays an important role in modern engineering, but using open-source solvers such as OpenFOAM requires considerable…
To reduce this burden, we propose AutoFOAM - a self-evolving large language model (LLM) agent that creates, evaluates, runs, and evolves its own OpenFOAM simula…
Our model is pre-trained on the Qwen-coder 2.5-14B, which is then fine-tuned on 252 text prompts targeting 7 OpenFOAM solvers, 13 parametrized mesh templates, a…
- Published: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00006v1 Announcement Type: New.
- Abstract: Large Language Models (LLMs), as part of Artificial Intelligence (AI), are increasingly being adopted by Small and Medium Enterprises (SMEs) to enhance question-answering capabilities and support business decision-making processes.
- However, hallucinations in the outputs generated by LLMs can become a source of misinformation, thereby reducing user confidence in their reliability and trustworthiness within SMEs.
- Retrieval-Augmented Generation (RAG) has emerged as a promising method to address this challenge by incorporating external knowledge sources into the modeling process.
- EN Highlights:
- arXiv:2608.00006v1 Announce Type: new
- Abstract: Large Language Models (LLMs), a part of artificial intelligence (AI), are increasingly being adopted by Small and Medium Enterprises (SMEs) to enhance…
- However, hallucinations in LLM-generated outputs can serve as a source of misinformation, reducing user confidence in their reliability and trustworthiness with…
- Retrieval-Augmented Generation (RAG) has emerged as a promising approach to address this challenge by incorporating external knowledge sources into the modeling…
- Published: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00008v1 Announcement Type: New.
- Abstract: Due to privacy concerns and the desire for local inference, the local deployment of Large Language Models (LLMs) is gaining attention.
- However, the energy costs on consumer hardware remain poorly characterized, as most benchmarks focus solely on accuracy.
- This paper presents a reproducible, hardware-level energy benchmark for nine open-source LLMs (from 1B to 7B parameters) executed on a single consumer-grade GPU (RTX 4060Ti 16GB).
- EN Highlights:
- arXiv:2608.00008v1 Announce Type: new
Abstract: The local deployment of large language models (LLMs) is gaining traction due to privacy concerns and the desire for on-premise inference
However, the energy costs on consumer hardware remain poorly characterized, as most benchmarks focus solely on accuracy
This paper presents a reproducible, hardware-level energy benchmark of nine open-source LLMs (1B to 7B parameters) executed on a single consumer GPU (RTX 4060Ti…
CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection
- Published: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00014v1 Announce Type: new.
- Abstract: Evaluating Large Language Models (LLMs) incurs prohibitive computational overhead during continuous development processes.
- While coreset selection accelerates evaluation, existing methods either suffer from a severe “cold start” bottleneck requiring massive historical logs (e.g., item response theory), or exhibit superficial lexical biases, missing the underlying reasoning manifold of the tasks.
- We propose CoT-Core, a novel training-free core question selection framework.
- EN Highlights:
- arXiv:2608.00014v1 Announce Type: new
- Abstract: Evaluating Large Language Models (LLMs) incurs prohibitive computational overhead during continuous development processes
- While coreset selection accelerates evaluation, existing methods either suffer from a severe ``cold start’’ bottleneck requiring massive historical logs (e.g.,…
- We propose CoT-Core, a novel training-free core question selection framework
Optimization and Constraint Modeling using LLMs with a Retrieval Augmented Generation Process
- Published: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00015v1 Announce Type: new.
- Abstract: Both optimization modeling and constraint modeling are non-trivial problems requiring deep domain expertise and proficiency in modeling formalism languages.
- Despite their importance across logistics, healthcare, and supply chain management, current large language models regularly produce structurally inconsistent or incomplete optimization formulations, especially in combinatorial settings.
- This paper evaluates whether a retrieval-augmented generation pipeline built upon a curated synthetic dataset can meaningfully improve LLM optimization modeling performance.
- EN Highlights:
- arXiv:2608.00015v1 Announce Type: new
- Abstract: Both optimization modeling and constraint modeling are non-trivial problems requiring deep domain expertise and proficiency in modeling formalism lang…
- Despite their importance across logistics, healthcare, and supply chain management, current large language models regularly produce structurally inconsistent or…
This paper evaluates whether a Retrieval-Augmented Generation pipeline built on a curated synthetic dataset can meaningfully improve LLM optimization modeling p…
Memory Reward Inflation in Self-Improving LLM Agents
- Release Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00017v1 Announcement Type: New.
- Abstract: Self-improving LLM agents are increasingly learning from experience without updating any weights.
- Each episode is stored in an external memory, scored, and retrieved for future similar tasks to shape subsequent behavior.
- From a reward perspective, the stored score is a proxy reward for an implicit non-parametric policy.
- EN Key Points:
- arXiv:2608.00017v1 Announce Type: new
- Abstract: Self-improving LLM agents increasingly learn from experience without updating any weights
- Each episode is stored in an external memory, scored, and retrieved for similar future tasks to shape later behavior
- Viewed through a reward lens, the stored score is a proxy reward for an implicit, non-parametric policy
Request-Level Energy Attribution for Batched LLM Serving
- Release Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00026v1 Announcement Type: New.
- Abstract: Batched LLM serving improves throughput but complicates energy accounting.
- GPU power telemetry is aggregated, while sustainability reporting, chargebacks, and workload analysis often require request-level energy charges.
- Existing inference energy benchmarks report energy at the model, phase, or token level, and recent carbon accounting efforts conceptually motivate Shapley fairness.
- EN Key Points:
- arXiv:2608.00026v1 Announce Type: new
- Abstract: Batched LLM serving improves throughput but complicates energy accounting
- GPU power telemetry is aggregate, whereas sustainability reporting, chargeback, and workload analysis often require request-level energy charges
- Existing inference-energy benchmarks report model-, phase-, or token-level energy, and recent carbon-accounting work motivates Shapley fairness conceptually
Motif-Mamba: network motif improved mamba for long-range sequence modeling
- Release Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00027v1 Announcement Type: New.
- Abstract: Efficient long-sequence modeling remains a core challenge for large language models because self-attention scales quadratically with sequence length.
- Mamba offers a linear-time alternative through selective state-space recursion, but its primarily diagonal state transitions limit explicit interactions between state dimensions.
- We propose Motif-Mamba, a structured state-space model that enhances Mamba with motif-constrained, low-order recurrent paths.
- EN Key Points:
- arXiv:2608.00027v1 Announce Type: new
Abstract: Efficient long-sequence modeling remains a central challenge for large language models, as self-attention scales quadratically with sequence length
Mamba offers a linear-time alternative through selective state space recurrence, but its predominantly diagonal state transitions restrict explicit interactions…
We propose Motif-Mamba, a structured state space model that augments Mamba with a motif-constrained low-rank recurrent pathway
Nova: An End-to-End MLIR Compiler for Deep Learning
- Publication Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00029v1 Announcement Type: New.
- Abstract: The performance of large-scale deep learning models largely depends on how effectively high-level mathematical operations are mapped to the underlying physical hardware.
- While high-level tensor frameworks provide flexible abstractions for model design, their eager execution models inherently lack full-graph visibility and fine-grained control over hardware and memory to maximize native physical hardware utilization.
- To bridge this gap, we designed Nova, an automated end-to-end JIT compiler whose defining purpose is to achieve absolute control over this hardware mapping: fusing operations across operator boundaries, optimizing complex memory hierarchies, and tailoring execution to the register level.
- EN Highlights:
- arXiv:2608.00029v1 Announce Type: new
- Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physica…
- While high-level tensor frameworks provide flexible abstractions for model design, their eager execution models inherently lack the whole-graph visibility and g…
- To bridge this gap, we designed Nova, an automated end-to-end JIT compiler whose defining purpose is to achieve absolute control over this hardware mapping: fus…
ArXiv cs.CL (B_intro+search) Link to heading
Cost-Effective Automated Judging of Natural-Language Mathematical Proofs
- Publication Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00004v1 Announcement Type: New.
- Abstract: Scoring natural language mathematical proofs is a recurring cost for evaluating mathematical reasoning systems, and frontier model judges are expensive.
- We investigate whether inexpensive open-weight models can serve as reliable judges given a candidate proof, a ground-truth proof, and human scoring criteria.
- On a 200-instance validation sample from IMO-GradingBench, three inexpensive judges (GPT-OSS 120B, DeepSeek-V4 Flash, Gemma-4 31B) agreed with human pass/fail decisions at a rate statistically indistinguishable from Claude Opus 4.7 and Gemini 3.1 Pro, but at 100x lower cost.
- EN Highlights:
- arXiv:2608.00004v1 Announce Type: new
Abstract: Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive
- We ask whether cheap open-weight models can serve as reliable judges given a candidate proof, a ground-truth proof, and a human-grading rubric
- On a 200-instance validation sample of IMO-GradingBench, three cheap judges (GPT-OSS 120B, DeepSeek-V4 Flash, Gemma-4 31B) agree with human pass/fail decisions…
RubricReviewer: From Direct Critique to Objective and Comprehensive Rubric-Driven Peer Review
- Release Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00005v1 Announcement Type: new.
- Abstract: Peer review at major venues is under unprecedented submission pressure, motivating the use of large language models (LLMs) as review assistants.
- However, existing LLM-based reviewers face two structural limitations.
- First, they map manuscripts directly to reviews, leaving the underlying rubric implicit and entangling its derivation with the judgement.
- EN Key Points:
- arXiv:2608.00005v1 Announce Type: new
- Abstract: Peer review at major venues is under unprecedented submission pressure, motivating the use of large language models (LLMs) as review assistants
- Existing LLM-based reviewers, however, face two structural limitations
- First, they map manuscripts directly to reviews, leaving the underlying rubric implicit and entangling its derivation with the judgement
MemoryForge: Synthesize Lifelong Memory for Human-Like LLM Agents
- Release Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00007v1 Announcement Type: new.
- Abstract: Equipping Large Language Models (LLMs) with human-like personas is crucial for agentic applications, such as role-play and user simulation.
- Traditional prompt-based methods rely on descriptive conditioning by injecting static textual profiles, which often makes agents show generic behaviors due to a lack of realistic life memories.
- To fill this gap, we introduce memory-based conditioning, a paradigm inspired by cognitive psychology that replaces abstract profiles with an autobiographical memory repository, enabling frozen LLMs to dynamically retrieve context-relevant memories to guide their behavior.
- EN Key Points:
- arXiv:2608.00007v1 Announce Type: new
- Abstract: Equipping Large Language Models (LLMs) with human-like personas is crucial for agentic applications, such as role-play and user simulation
- Traditional prompt-based methods rely on descriptive conditioning by injecting static textual profiles, which often makes agents show generic behaviors due to a…
To fill this gap, we introduce memory-based conditioning, a paradigm inspired by the cognitive psychology, which replaces abstract profiles with an autobiograph…
- Published: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00009v1 Announce Type: new.
- Abstract: Long-term memory remains a critical bottleneck for conversational AI agents, whose finite context windows cannot support coherent recall across thousands of turns.
- We introduce AgentMemBench, a unified, reproducible benchmark to evaluate five memory management strategies under identical conditions: in-context windowing (ICW), external key-value stores (EKV), graph-based episodic memory (GEM), compression-based summarization (CBS), and web-augmented memory (WAM).
- All are assessed across three public datasets covering long-term multi-session dialogue (LoCoMo), task-oriented document grounding (MultiDoc2Dial), and persona-based multi-session chat (MSC), using Recall@k, MRR, nDCG@k, Answer F1, LLM-judge faithfulness scores, memory footprint, and latency over 491 annotated question-turns.
- EN Key Points:
- arXiv:2608.00009v1 Announce Type: new
- Abstract: Long-term memory remains a critical bottleneck for conversational AI agents, whose finite context windows cannot support coherent recall across thousa…
- We present AgentMemBench, a unified, reproducible benchmark evaluating five memory management strategies under identical conditions: in-context windowing (ICW),…
- All are assessed across three public datasets covering long-term multi-session dialogue (LoCoMo), task-oriented document grounding (MultiDoc2Dial), and persona-…
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis
- Published: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00011v1 Announce Type: new.
- Abstract: Current text-to-speech systems face a trade-off: autoregressive codec language models can produce highly intelligible speech but require large-scale models and training data and decode tokens sequentially, while non-autoregressive methods improve speed at the cost of linguistic accuracy.
- We propose DLLM-TTS, a framework that formulates TTS as a conditional blockwise discrete diffusion over X-Codec2 neural audio codec tokens.
- The model decomposes the sequence into blocks and applies masked diffusion within each block while processing blocks sequentially, learning both local acoustic consistency and global text-speech alignment.
- EN Key Points:
- arXiv:2608.00011v1 Announce Type: new
- Abstract: Current text-to-speech systems face a trade-off: autoregres- sive codec language models produce highly intelligible speech but require large-scale mod…
We present DLLM-TTS, a framework that formulates TTS as conditional block discrete diffusion over X-Codec2 neural audio codec to- kens
The model decomposes sequences into blocks and applies masked diffusion within each block while processing blocks se- quentially, learning both local acoustic c…
- Publication Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00012v1 Announcement Type: New.
- Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains under-evaluated.
- Existing remote sensing benchmarks largely rely on static, post-hoc, and expert-processed products, such as gridded reanalysis data, which are difficult to align with disaster scenarios where events unfold rapidly and decisions must be made under strict time constraints.
- To bridge this gap, we introduce Obshazard-bench, a real-time, observation-driven benchmark for evaluating disaster intelligence in MLLMs.
- EN Key Points:
- arXiv:2608.00012v1 Announce Type: new
- Abstract: Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaste…
- Existing remote sensing benchmarks largely rely on static, post-hoc, and expert-processed products, such as gridded reanalysis data, which are difficult to alig…
- To bridge this gap, we introduce Obshazard-bench, a real-time, observation-driven benchmark for evaluating disaster intelligence in MLLMs
What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs
- Publication Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00013v1 Announcement Type: New.
- When building a Vision-Language Model (VLM), selecting the right Large Language Model (LLM) backbone is the most consequential decision, yet it remains fundamentally unprincipled: computation-based scaling laws do not generalize across model families, and no framework exists to directly predict VLM performance before training begins.
- We propose capability-driven multimodal scaling laws, the first cross-family framework that predicts VLM benchmark accuracy from directly observable text capabilities.
- Given a low-dimensional capability score $S$ extracted from LLM text benchmarks via PCA, we model VLM performance as a function of $S$ and use per-backbone transfer and absorption rates to quantify data scaling efficiency.
- EN Key Points:
- arXiv:2608.00013v1 Announce Type: new
- Abstract: Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains…
We propose the Capability-Driven Multimodal Scaling Law, the first cross-family framework that predicts VLM benchmark accuracy from directly observable textual…
Given a low-dimensional capability score $S$ extracted from LLM textual benchmarks via PCA, we model VLM performance as a function of $S$, with a per-backbone t…
Role Steering of Language Models for Social Simulations
- Publish Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00023v1 Announce Type: new.
- Abstract: Social simulations built from language model agents require role-conditioned behavior that can be inspected before placing agents into a simulated population.
- We introduce an activation-steering screening workflow for role-conditioned agents: defining a role profile, extracting a role-specific direction, sweeping four steering coefficients, evaluating role profile alignment, and passing or flagging each candidate configuration.
- On OLMo-3-7B-Instruct, we apply the workflow to a mixed 275-role inventory with 228 role-agnostic questions, GPT-4.1-mini prompted role references, and GPT-4.1-mini judges.
- EN 要点:
- arXiv:2608.00023v1 Announce Type: new
- Abstract: Social simulations built from language-model agents need role-conditioned behavior that can be checked before agents are placed into a simulated popul…
- We introduce an activation-steering screening workflow for role-conditioned agents: define a role profile, extract a role-specific direction, sweep four steerin…
- On OLMo-3-7B-Instruct, we apply the workflow to a mixed 275-role inventory with 228 role-agnostic questions, GPT-4.1-mini prompted role references, and GPT-4.1-…
Exploring More to Solve More: Boosting Diversity in Text Diffusion Models via Entropy-Based Guidance
- Publish Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00024v1 Announce Type: new.
- Abstract: Although diffusion models have revolutionized continuous domains like image synthesis through high-quality generation and controllable guidance mechanisms, introducing this controllability to the discrete, sequential nature of text remains an open challenge.
- Simultaneously, current sampling strategies and guidance methods adjust token probabilities without capturing the broader semantic landscape, leading to a suboptimal balance between fidelity and diversity.
- In this work, we introduce a novel training-free semantic-aware kernel entropy (SAKE) guidance method.
- EN 要点:
- arXiv:2608.00024v1 Announce Type: new
- Abstract: Although diffusion models have revolutionized continuous domains like image synthesis through high quality generations and controllable guidance mecha…
Meanwhile, current sampling strategies and guidance methods adjust token likelihoods without capturing the broader semantic landscape, leading to a suboptimal b…
In this work, we introduce a novel training-free Semantic-Aware Kernel Entropy (SAKE) guidance method
SLMs as Multi-Agent Routers: A Progressive SFT and Reinforcement Learning Approach
- Published: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00030v1 Announce Type: new.
- Abstract: Specialized retrieval agents typically provide higher-quality results than general-purpose search, but selecting the optimal agent for a given query remains an open problem.
- Current methods route queries based on inferred topics or intents; however, intent-based selection is fundamentally limited: it does not incorporate signals from retrieved content and cannot detect when a thematically consistent agent produces low-relevance results.
- We address this problem by training a small language model via supervised fine-tuning followed by reinforcement learning to jointly perform agent selection and structured parameter generation for downstream tool calls, using a hierarchical reward function based on retrieval relevance and query-agent topic alignment.
- EN Key Points:
- arXiv:2608.00030v1 Announce Type: new
- Abstract: Specialised retrieval agents typically surface higher quality results than general-purpose search, but selecting the optimal agent for a given query r…
- Current approaches route queries based on inferred topic or intent, however intent-based selection is fundamentally limited: it does not incorporate signal from…
- We address this by training a small language model via supervised fine-tuning followed by reinforcement learning to jointly perform agent selection and structur…
ArXiv cs.LG (B_intro+search) Link to heading
Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models
- Published: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00019v1 Announce Type: new.
- Abstract: Deploying large language models (LLMs) for operations research (OR) tasks remains challenging because correctness depends on a coherent modeling process rather than just the correct final answer.
- Standard autoregressive generation operates on a short-sighted policy, sometimes failing to predict whether a partial formulation can be effectively extended into a globally consistent optimization model.
- Therefore, locally sound steps can propagate into catastrophic downstream formulation or solver code errors.
- EN Key Points:
- arXiv:2608.00019v1 Announce Type: new
- Abstract: Deploying large language models (LLMs) for operations research (OR) tasks remains challenging because correctness depends on a coherent modeling proce…
Standard autoregressive generation operates on a myopic policy, which sometimes fails to anticipate whether a partial formulation can be validly extended into a…
Consequently, locally plausible steps may propagate into catastrophic downstream formulation or solver code errors
Learning Compositional Meta-Routing for Agentic Workflows: An Executable Benchmark
- Published: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00106v1 Announcement Type: New.
- Abstract: Agentic systems must decide not only what answer to produce, but also which reasoning and execution operations should precede it.
- A controller may directly answer, decompose a request, retrieve evidence, execute code, delegate to a specialist, or verify an intermediate result.
- Existing routing work largely selects model endpoints, retrieval depth, or tools in isolation.
- EN Highlights:
- arXiv:2608.00106v1 Announce Type: new
- Abstract: Agentic systems must decide not only what answer to produce, but which reasoning and execution operations should precede it
- A controller may answer directly, decompose a request, retrieve evidence, execute code, delegate to a specialist, or verify an intermediate result
- Existing routing work largely selects model endpoints, retrieval depth, or tools in isolation
MetaRoute-Bench: Evaluating Meta-Decision Policies for Agentic Workflow Routing
- Published: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00107v1 Announcement Type: New.
- Abstract: Agentic systems must repeatedly decide whether to answer directly, decompose tasks, invoke tools, execute code, delegate to specialists, verify intermediate results, or recover from failures.
- These meta-decisions affect not only task success but also operating cost and latency, yet they are often embedded inside an orchestration framework and evaluated only by aggregate task accuracy.
- We present MetaRoute-Bench, an open, inspectable framework for comparing meta-decision policies under a shared execution model.
- EN Highlights:
- arXiv:2608.00107v1 Announce Type: new
- Abstract: Agentic systems must repeatedly decide whether to answer directly, decompose a task, invoke a tool, execute code, delegate to a specialist, verify an…
- These meta-decisions affect not only task success but also operating cost and latency, yet they are often embedded inside an orchestration framework and evaluat…
- We present MetaRoute-Bench, an open, inspectable framework for comparing meta-decision policies under a shared execution model
- Release Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00129v1 Announcement Type: New.
- Abstract: Knowledge distillation (KD) is a widely used technique for transferring knowledge from a large model (teacher) to a smaller model (student).
- Due to its flexibility and wide applicability, KD is widely used in the compression of server-side models to meet the Quality of Service (QoS) requirements of client-side users.
- Despite significant progress, the performance of distillation is severely impacted when a large disparity exists between the capabilities of the server and the needs of the client.
- EN Highlights:
- arXiv:2608.00129v1 Announce Type: new
- Abstract: Knowledge distillation (KD) is a widely utilized technique for transferring knowledge from a large model (the teacher) to a smaller model (the student…
- Owing to its flexibility and broad applicability, KD has been extensively applied in the compression of server-side models to meet the Quality of Service (QoS)…
- Despite significant advancements, the performance of distillation is substantially compromised when a large disparity exists between the capabilities of the ser…
- Release Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00135v1 Announcement Type: New.
- Abstract: Design and architectural archives encode expert human knowledge in graphical formats, providing a critical testbed for design-inspired machine learning (ML) challenges that are lacking in typical computer vision benchmarks.
- Based on JONES-19 (a small image dataset based on The Grammar of Ornament (London, 1857)), we evaluate the discriminative performance of Convolutional Neural Networks (CNN) under two model training strategies: (a) ImageNet pre-training for general-domain “visual common sense,” and (b) learning from scratch with the design data in JONES-19.
- We find that while domain-general priors improve discriminative performance, learning from scratch augmented with repeated local sampling (multi-crop) can effectively recover these gains.
- EN Highlights:
- arXiv:2608.00135v1 Announce Type: new
- Abstract: Design and architectural archives encode expert human knowledge in graphical formats, providing a critical testbed for design-inspired Machine Learnin…
- Building on JONES-19, a small-size image dataset based on The Grammar of Ornament (London, 1857), we evaluate the discriminative performance of Convolutional Ne…
- We find that while domain-general priors improve discriminative performance, learning from scratch augmented with repeated local sampling (multi-crop) effective…
Leak It: A Probabilistic Approach to Training-Data Extraction from Black-Box Language Models
- Publication Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00144v1 Announcement Type: new.
- Abstract: Membership inference (MIA) on language models is often summarized by aggregate ROC-AUC, but this evaluation is confounded: a model-free blind baseline separates members from non-members based on surface text alone.
- We investigate black-box, sampling-based training data leakage through a probabilistic lens, treating N samples from p(.|x) as an estimate of the output distribution and the leakage signal as a function thereof.
- We extend the critique of blind baselines to the sampling regime: on WikiMIA, a blind bag-of-words classifier achieves an AUC of 0.97 (0.90 TPR at 5% FPR) with no added benefit from sampling, while on an IID split of the Pile (MIMIR), neither self-concentration nor golden continuation recovery significantly outperforms the blind baseline (incremental AUC 95% CI includes zero).
- EN Key Points:
- arXiv:2608.00144v1 Announce Type: new
- Abstract: Membership inference (MIA) on language models is usually summarised by an aggregate ROC-AUC, but such evaluations are confounded: model-free blind bas…
- We study black-box, sampling-based training-data leakage through a probabilistic lens, treating N samples from p(.|x) as an estimate of the output distribution…
- We extend the blind-baseline critique into the sampling regime: on WikiMIA a blind bag-of-words classifier reaches AUC 0.97 (TPR 0.90 at 5% FPR) and sampling ad…
Response Magnitude as a Dominant Signal for Held-Out CRISPRi Perturbation Effect Prediction
- Publication Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00152v1 Announcement Type: new.
- Abstract: Predicting the magnitude of a CRISPRi perturbation’s transcriptomic effect on held-out target genes is a significant open problem in single-cell biology.
- Recent work has documented that simple baselines often match or outperform deep perturbation predictors on related protocols.
- We study this phenomenon on the Virtual Cell Challenge (VCC) benchmark under a strict held-out target-gene split, identifying the specific low-dimensional signal that drives the gap and describing how it transfers across cell types.
- EN Key Points:
- arXiv:2608.00152v1 Announce Type: new
- Abstract: Predicting the magnitude of a CRISPRi perturbation’s transcriptomic effect on held-out target genes is an important open problem in single-cell biolog…
- Recent work has documented that simple baselines often match or exceed deep perturbation predictors on related protocols
- We study this phenomenon on the Virtual Cell Challenge (VCC) benchmark under a strict held-out target-gene split, identify the specific low-dimensional signal t…
Inference-Time Policy Alignment for Fair Reinforcement Learning
- Publish Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00175v1 Announce Type: New.
- Abstract: Deep reinforcement learning (RL) agents achieve strong performance by optimizing scalar reward functions.
- However, once deployed, the policies of these RL agents are often rigid and costly to adapt to new performance criteria.
- For example, an agent trained to maximize expected cumulative reward may not accommodate previously unknown stakeholder preferences.
- EN 要点:
- arXiv:2608.00175v1 Announce Type: new
- Abstract: Deep reinforcement learning (RL) agents achieve strong performance by optimizing scalar reward functions
- However, once deployed, the policies of these RL agents are often rigid and costly to adapt to new performance criteria
- For instance, an agent trained to maximize expected cumulative reward may not accommodate previously unknown stakeholder preferences
- Publish Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00198v1 Announce Type: New.
- Abstract: Environmental time-series causal discovery requires expert decisions about method choice, conditional independence testing, lag horizons, sample size sufficiency, multiple testing control, and evidence interpretation.
- Applied inconsistently across datasets, these choices yield graphs that cannot be compared, reproduced, or audited.
- We present AutoCause, an open-source Python workflow that records each decision, derives defaults from an extended causal audit module, and allows for domain-informed overrides.
- EN 要点:
- arXiv:2608.00198v1 Announce Type: new
- Abstract: Environmental time-series causal discovery requires expert decisions about method choice, conditional-independence tests, lag horizons, sample-size ad…
- Applied inconsistently across datasets, these choices yield graphs that cannot be compared, reproduced, or audited
- We present AutoCause, an open-source Python workflow that records each decision, derives defaults from an extended causal-audit module, and admits domain-inform…
- Publish Time: 2026-08-04 12:00 Beijing Time
- Abstract: - arXiv:2608.00212v1 Announce Type: New.
- Abstract: Spatial atomic layer deposition (SALD) is the leading atmospheric pressure, high-throughput pathway for industrial ALD, but design and control are limited by the cost of predicting surface coverage: high-fidelity CFD is too slow for operating window scans, while analytical models miss transport modulations such as gas curtains.
We propose a physics-chemistry-informed neural network (PCINN), a hybrid surrogate with CFD-level accuracy at real-time speed: a query returns coverage in approximately 7 milliseconds, about 5x10^4 times faster than a CFD solve, achieving a test R^2_log = 0.998 (leave-one-out R^2_raw = 0.974) with only 30 training cases covering four orders of magnitude.
The architecture is not a black box: a small network learns only the operating conditions for near-wall concentration closure, while the known surface kinetics are a hard-coded, trainable chemistry layer integrated along the substrate trajectory.
- EN Highlights:
- arXiv:2608.00212v1 Announce Type: new
- Abstract: Spatial atomic layer deposition (SALD) is a leading atmospheric-pressure, high-throughput route to industrial ALD, but design and control are limited…
- We present a physics-chemistry-informed neural network (PCINN), a hybrid surrogate with CFD-level accuracy at real-time speed: a query returns coverage in about…
- The architecture is not a black box: a small network learns only the operating-condition to near-wall concentration closure, while the known surface kinetics is…
- EN Highlights: