System translated (Gemini)

🤖 AI 速览

Today’s focus shifts from “stronger models” to “higher throughput.” OpenAI previews GPT-5.6 Ultrafast, pushing frontier models into low-latency scenarios; Google releases Gemini 3.7 Flash, continuing to enhance its coding capabilities and cost-effectiveness. Meanwhile, …
📋 文章元数据
发布时间
2026-08-14
类型
ai-daily
字数
6863
阅读时长
33 min

2026-08-14 AI Daily Update | Highly Intelligent Models Begin Competing on Throughput: GPT-5.6 Speeds Up, Gemini Flash Fills the Gap Link to heading

Today’s main theme shifts from “stronger models” to “more output per unit of time.” OpenAI previews GPT-5.6 Ultrafast, pushing frontier models into low-latency scenarios; Google releases Gemini 3.7 Flash, continuing to enhance coding capabilities and cost-effectiveness. Meanwhile, Agent engineering is extending from skill orchestration and context compression to production monitoring and real-world evaluation.

📖 In-depth Guide to This Issue’s Watch List Link to heading

The most important trend to follow today is that “model capabilities continue to advance, but the real watershed lies in cost and speed.” Google DeepMind has launched Gemini 3.7 Flash, while OpenAI has released the developer guide for GPT-5.6 and previewed its Ultrafast mode. They are bringing high intelligence, low latency, and low cost to the forefront, which deserves close attention from product and infrastructure teams. The second major theme is Agent engineering: from skill learning and orchestration to context compression, several papers are addressing the same question—how to make agents cheaper, more stable, and more controllable. The third theme covers progress in evaluation and vertical applications. Areas like trading, role-playing, robotics, and sign language translation are all bridging the final gap from “functional” to “trustworthy.”

🌐 AI Hot Topics on X Link to heading

Topic 1: DeepSeek Launches V4-Pro with Agent Upgrades and Open-Source Framework Link to heading

  • Category: AI · News
  • Overview: Trending time: 1 day ago, Related posts: 24,000
  • What it is: DeepSeek released V4-Pro, focusing on upgrading agent capabilities and launching a supporting open-source framework.
  • Why it matters: This shows that open-source large models are shifting from simple text generation to agent systems capable of executing tasks, which could further lower the barrier for developers to build AI applications and automated workflows.
  • Discussion summary: Discussions on X focus on the actual performance of V4-Pro, the reliability of its agent capabilities, the impact of the open-source framework on the development ecosystem, and how the competitive landscape between DeepSeek and closed-source model providers like Google and OpenAI might change.

Topic 2: Cursor Welcomes Firetiger Team to Build AI Agents for Production Monitoring Link to heading

  • Category: AI · News
  • Overview: Trending time: , Related posts: 69
  • What it is: AI coding tool Cursor announced it has welcomed the Firetiger team, and together they will build AI Agents for production monitoring scenarios.
  • Why it matters: This indicates that AI Agents are extending beyond code generation into production environment aspects like software operation, observability, and incident response. This could enhance DevOps efficiency and help close the loop in the AI toolchain.
  • Discussion summary: Discussions on X are centered on whether Cursor will evolve from an AI coding assistant into a complete platform covering both development and operations. Users are also discussing the reliability, security, and false-positive handling of AI Agents in production monitoring. Some are also curious about what specific product features the Firetiger team’s addition will bring.

Topic 3: Google Launches Gemini 3.7 Flash with Major Coding Gains Link to heading

  • Category: AI · News
  • Overview: Trending time: 14 hours ago, Related posts: 15,000
  • What it is: Google released Gemini 3.7 Flash, highlighting significant improvements in its coding capabilities.
  • Why it matters: This type of model update directly impacts code generation, agent-based development, and product integration capabilities, reflecting the accelerating competition among foundational AI models in practical applications.
  • Discussion summary: Discussions on X focus on whether its coding performance is sufficient to close the gap with other top-tier models. Some also mention the improvements in Sonnet 5, Google’s reduction in media generation costs, and Andrew Ng’s views on agentic loops.

Topic 4: Rails ‘Agents on Rails’ Benchmark Ranks Top AI Coding Models Link to heading

  • Category: AI · News
  • Overview: Trending time: 5 hours ago, Related posts: 255
  • What it is: Rails released the “Agents on Rails” benchmark, which ranks major AI coding models on their agent-based development capabilities within real Rails projects.
  • Why it matters: The benchmark emphasizes model performance under framework constraints, code modification, passing tests, and multi-step task execution. It helps measure the ability of AI coding assistants to transition from generating code to automating actual software engineering.
  • Discussion summary: Discussions on X focus on whether the rankings accurately reflect the developer experience, which models are most reliable in complex codebases, and whether the benchmark is biased towards the Rails ecosystem or specific agent workflows. Others are interested in the performance and cost differences between open-source and closed-source models.

Topic 5: MiniMax Launches MiniMax-Music3 for Open AI Song Generation Link to heading

  • Category: AI · Entertainment
  • Overview: Trending time: 7 hours ago, Related posts: 1700
  • What it is: MiniMax has released the MiniMax-Music3 model for open AI song generation, drawing attention in the music generation field.
  • Why it’s important: This model shows that generative AI is further entering the music creation scene, which could impact content production, copyright licensing, creator tools, and the entertainment industry’s workflow.
  • Discussion summary: Discussions on X are focused on the quality, controllability, and openness of the songs generated by the model. Supporters believe it lowers the barrier to music creation, while critics are concerned about copyright ownership, training data sources, the income of original musicians, and the potential flood of AI-generated music.

Topic 6: MiniMax Releases MiniMax-Music3 for Open Song Generation Link to heading

  • Category: AI · News
  • Overview: Trending time: 2 hours ago, Related posts: 837
  • What it is: MiniMax has released MiniMax-Music3, a model for open song generation that can create complete songs including melody, vocals, and accompaniment.
  • Why it’s important: This indicates that AI music generation is moving from fragmented audio synthesis to more complete songwriting capabilities, potentially driving changes in content production, music tools, and copyright governance.
  • Discussion summary: Discussions on X mainly focus on the sound quality and controllability of the generated songs, comparisons with products like Suno, whether the weights or API will be open-sourced, its commercial potential, and risks related to training data and music copyrights.

AI Public Opinion Summary on X Today Link to heading

The main narrative on X today is that AI is rapidly transitioning from “being able to chat and write code” to the next stage of “being able to execute tasks, enter production systems, and directly generate content.” DeepSeek, Google, Cursor, and the Rails benchmark are all seen as different facets of this trend. The general consensus is that agent capabilities and the open-source ecosystem are lowering development barriers and accelerating the implementation of programming, operations, and automated workflows. However, whether the various models can be stable and reliable in real-world, complex scenarios remains a key point of discussion. The main disagreement centers on the gap between “on-paper capabilities” and “real-world usability”: some are optimistic about the ecosystem impact from open-source and benchmark rankings, while others question whether these leaderboards and demos are biased towards specific frameworks or task setups and may not represent the daily experience of developers. Potential risks are more focused on false positives in production monitoring, agent operational errors, the impact of music generation on copyrights and creator revenues, and the pressure on industry order and governance rules as generative content becomes more widespread.

💡 Influencer Insights Link to heading

No influencer insights for today. We recommend reading the in-depth content from the Watch List.

📚 Appendix: Today’s Watch List Update Source List Link to heading

Time window: Last 3 days; covers 22 sources; 36 updates in total

Y Combinator Podcast (B_intro+search) Link to heading

  • Chelsea Finn: This is the State of the Art in Robotics
    • Release Time: 2026-08-14 00:53 Beijing Time
    • Abstract: - You’ve probably heard of OpenClaw (formerly Clawdbot/Moltbot).
      • The viral open-source AI assistant that can run on your own devices, connects with the messaging apps you already use, and goes beyond chat to actually do things like manage your email, calendar, files, workflows, and more.
      • Now meet the person behind it.
      • YC’s Raphael Schaad sits down with Peter Steinberger, founder of OpenClaw, to talk about the “aha” moment behind the viral personal AI agent, why a local-first agent could replace many of today’s apps, and how personal agents will reshape the future of software.
    • EN Key Points:
      • Robots can already fold laundry, make espresso, clean kitchens, and assemble things
      • The harder problem is getting them to do those tasks reliably, for long periods of time, without a human babysitting them
      • At Startup School 2026, Physical Intelligence cofounder Chelsea Finn explains what it takes to build general-purpose robots that work in the real world
  • She shares how reinforcement learning pushed robot throughput up 2x, how their systems can run autonomously for hours, and why she believes robotics is entering…

All-In Podcast (A_full) Link to heading

  • Rahm Emanuel: Trump’s Foreign Policy, China, Europe’s Decline, Immigration & DSA vs Democrats
    • Publication time: 2026-08-13 10:10 Beijing Time
    • Summary:- (0:00) Jason and Friedberg welcome Rahm Emanuel.
      • (1:16) Lessons learned from working with Clinton and Obama.
      • (7:45) China: How to approach conflict, new economic blocs, isolation, and Taiwan.
      • (28:49) Europe: Reasons for the decline of the West.
    • EN Highlights:
      • (0:00) Jason and Friedberg welcome Rahm Emanuel
      • (1:16) Lessons from working with Clinton and Obama
      • (7:45) China: How to approach conflict, new economic bloc, isolation, and Taiwan
      • (28:49) Europe: Reasons for the decay of the West

OpenAI Blog (A_full) Link to heading

  • The builder’s guide to GPT‑5.6

    • Publication time: 2026-08-13 19:00 Beijing Time
    • Summary:- ## GPT‑5.6 sets a new standard for price-performance.
      • The GPT-5.6 model series makes frontier-level agent performance significantly more affordable, while also advancing the frontier of what’s possible.
      • In this guide, we show how startups can build faster, more powerful agents at extremely low costs using smarter model selection and new API controls to aid reasoning continuity, multi-agent orchestration, and programmatic tool calling.
      • Better out-of-the-box experience. Link to heading

      • Since GPT-5, each model generation has sought to solve long-term tasks with fewer tokens.
    • EN Highlights:
      • Learn how startups use GPT-5.6 to build faster, more cost-efficient AI agents with smarter model selection and new Responses API capabilities.
  • Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed

    • Publication time: 2026-08-13 18:00 Beijing Time
    • Summary:- Today, we’re sharing an early look at Ultrafast, a new service tier that runs GPT‑5.6 Sol up to 14x faster than standard processing, launching first in the OpenAI API.
      • Powered by Cerebras, Ultrafast generates up to 750 output tokens per second, bringing our smartest models to products and workflows where every second counts.
      • Until now, achieving real-time speed often meant choosing smaller or more specialized models.
      • Ultrafast points to progress in a new direction: doing more useful work per second.
      • When speed no longer requires sacrificing intelligence, AI can enter the most time-sensitive parts of enterprises, and new types of work become possible.
    • EN Highlights:
      • Preview Ultrafast, a new OpenAI API service tier that runs GPT-5.6 Sol up to 14× faster
      • Powered by Cerebras, it delivers up to 750 output tokens per second.
  • OpenAI appoints Dali Rajic as Chief Revenue Officer

    • Publication time: 2026-08-13 17:00 Beijing Time
  • Summary: - OpenAI appoints Dali Rajic as Chief Revenue Officer to lead its global revenue organization, helping businesses realize the full value of artificial intelligence.

    • This article from the OpenAI blog explains how OpenAI’s appointment of Dali Rajic as Chief Revenue Officer shapes the broader AI and infrastructure landscape.
    • Following OpenAI’s appointment of Dali Rajic as Chief Revenue Officer, this also has practical implications for founders, operators, and investors.
    • EN Highlights:
      • OpenAI appoints Dali Rajic as Chief Revenue Officer to lead its global revenue organization and help businesses realize the full value of AI.

Google DeepMind Blog (A_full) Link to heading

  • Introducing Gemini 3.7 Flash
    • Published: 2026-08-14 01:04 Beijing Time
    • Summary: - Introducing Gemini 3.7 Flash.
      • This article from the Google DeepMind blog explains how the launch of Gemini 3.7 Flash shapes the broader AI and infrastructure landscape.
      • Following the launch of Gemini 3.7 Flash, it also has practical implications for founders, operators, and investors.
    • EN Highlights:
      • Introducing Gemini 3.7 Flash

ArXiv cs.AI (B_intro+search) Link to heading

  • Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes

    • Published: 2026-08-13 12:00 Beijing Time
    • Summary: - arXiv:2608.11207v1 Announcement Type: New.
      • Abstract: When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition, but collapse: the visitor surrenders, the site agent stops changing its approach, and the dialogue terminates without achieving the stated goals of either agent.
      • This paper asks whether a control-theoretic governance layer can substitute for the missing goal function.
      • The Experience Orchestrator (EO) addresses this problem in a simulated financial services environment where a site agent guides a visitor toward contacting an advisor while the visitor maintains psychologically realistic resistance.
    • EN Highlights:
      • arXiv:2608.11207v1 Announce Type: new
      • Abstract: When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competitio…
      • This paper asks whether a control-theoretic governance layer can substitute for that missing goal function
      • The Experience Orchestrator (EO) addresses this in a simulated financial services environment where a site agent guides a visitor toward advisor contact while t…
  • Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration

    • Published: 2026-08-13 12:00 Beijing Time
    • Summary: - arXiv:2608.11210v1 Announcement Type: New.
      • Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter.
      • Despite decades of methodological work, researchers almost always rely on uniform priors.
      • The main reason is that constructing informative priors from the scientific literature is slow and requires both domain and statistical expertise.
    • EN Highlights:
      • arXiv:2608.11210v1 Announce Type: new
  • Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter

  • Despite decades of methodological work, researchers almost always fall back on uniform priors

  • The main reason is that building informative priors from scientific literature is slow and needs both domain and statistical expertise

  • A Forced-Structure Reduction and Verifiable Bounds for Conway’s 99-Graph

    • Publication Time: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11211v1 Announcement Type: New.
      • Abstract: The Conway 99-graph problem asks whether a strongly regular graph with parameters $\mathrm{srg}(99,14,1,2)$ exists.
      • We report a systematic, fully reproducible attack initiated by an autonomous AI research agent, scored according to the track’s partial credit metrics.
      • Our verifiable contributions are: (1) an exhaustive proof that no circulant graph on $\mathbb{Z}/99$ satisfies more than $3366/4950=68.0%$ of the constraints (33 of 49 difference classes), with the same upper bound for other abelian groups of order 99; (2) a forced-structure reduction: $\lambda=1$ makes each neighborhood a perfect matching and $\mu=2$ bijects external vertices to unmatched neighbor-pairs, collapsing existence into a 12-regular graph on 84 vertices, encoded for CP-SAT and validated by recovering the unique $\mathrm{srg}(9,4,1,2)$; (3) a verified framework for the existence of prescribed automorphism orbits (fixed-point-free and single-fixed-point actions, checked on $\mathrm{srg}(9,4,1,2)$ and the Paley graph $\mathrm{srg}(13,6,2,3)$), and (4) a best-verified artifact at $69.43%$ with evidence this is a robust frontier (fourteen distinct methods, none exceeding it) entangled with the open problem, since any provable bound less than 4950 is a proof of non-existence.
    • EN Key Points:
      • arXiv:2608.11211v1 Announce Type: new
      • Abstract: Conway’s 99-graph problem asks whether a strongly regular graph with parameters $\mathrm{srg}(99,14,1,2)$ exists
      • We report a systematic, fully reproducible attack by an autonomous AI research agent, scored under the track’s partial-credit metric
      • Our verifiable contributions are: (1) an exhaustive proof that no circulant graph on $\mathbb{Z}/99$ satisfies more than $3366/4950=68.0%$ of the constraints (…
  • Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-Experts

    • Publication Time: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11212v1 Announcement Type: New.
      • Abstract: Top-k Mixture-of-Experts (MoE) routing is discontinuous, so deployment-driven numerical disturbances (simulated 4-bit KV cache quantization read by a protected BF16 gate) push tokens across decision boundaries and flip the experts they trigger.
      • This paper does not propose new mitigations; it contributes causal machinery, empirical findings, and a detection-limit result.
      • Four-wheel machinery prices out the Route-Mediated Fraction (RMF) of quantization damage, token-level attributions decompose it by mechanism, and pre-registered probes carry the results across three architectures.
  • EN Highlights:

    • arXiv:2608.11212v1 Announce Type: new
    • Abstract: Top-k Mixture-of-Experts (MoE) routing is discontinuous, so a deployment-motivated numerical disturbance – simulated 4-bit KV-cache quantization read…
    • This paper proposes no new mitigation; it supplies a causal apparatus, empirical findings, and a detection-limit result
    • A four-run apparatus prices the route-mediated fraction (RMF) of quantization damage, a token-level attribution decomposes it by mechanism, and pre-registered p…
  • Poor Man’s Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop

    • Release Time: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11215v1 Announce Type: new.
      • Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behavior, stylized facts, and scaling with the number of agents $N$, rather than the cognition of any single agent.
      • We turn a statistical-physics observation into a method: replace each LLM agent by a low-parameter model fitted from a few hundred to a few thousand cheap queries, then run the society on a laptop at arbitrary $N$.
      • Whether this works is decided before the simulation runs, chiefly by what each agent perceives.
    • EN Highlights:
      • arXiv:2608.11215v1 Announce Type: new
      • Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phas…
      • We turn a statistical-physics observation into a method: replace each LLM agent by a low-parameter model fitted from a few hundred to a few thousand cheap queri…
      • Whether this works is decided before the simulation runs, chiefly by what each agent perceives
  • AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

    • Release Time: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11216v1 Announce Type: new.
      • Abstract: World-modeling is a volatile field: architectures, training objectives, and state representations interact in complex ways, and no single method dominates across environments.
      • This makes it an ideal testbed for AI coding agents as autonomous researchers – where, in this case, the direction of improvement is not specified in advance, unlike the engineering-spec tasks that dominate current agent benchmarks.
      • We introduce AutoWorldModel-Bench, a closed-loop benchmark where frontier coding agents autonomously improve a provided world-model starter under a fixed compute budget.
    • EN Highlights:
      • arXiv:2608.11216v1 Announce Type: new
  • Abstract: World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dom…

  • This makes it an ideal testbed for AI coding agents acting as autonomous researchers–a setting in which the improvement direction is not specified in advance,…

  • We introduce AutoWorldModel-Bench, a closed-loop benchmark in which frontier coding agents autonomously improve a provided world-model starter under a fixed com…

  • MaSRead: Content-Addressed Reading of Replicated Latent Stores

    • Publication Time: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11218v1 Announcement Type: New.
      • Abstract: Independent agents that reason in latent space can share computed state as key-value cache fragments rather than text.
      • Merged by a conflict-free replicated data type, these fragments form a store that converges under any delivery order or duplication.
      • Yet a later query, unknown at encode time, cannot reliably read the merged cache: colocated fragments interfere, so colocation is not addressability.
    • EN Key Points:
      • arXiv:2608.11218v1 Announce Type: new
      • Abstract: Independent agents that reason in latent space can share computed state as key-value cache fragments rather than text
      • Merged by a conflict-free replicated data type, these fragments form a store that converges under any delivery order or duplication
      • Yet a later query, unknown at encode time, cannot reliably read the merged cache: colocated fragments interfere, so colocation is not addressability
  • From Monolithic to Modular: Segment-level Automatic Prompt Optimization

    • Publication Time: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11219v1 Announcement Type: New.
      • Abstract: Automatic Prompt Optimization (APO) typically rewrites prompts monolithically, which can improve one behavior while degrading others.
      • We propose SAPO, a segment-level APO method that decomposes the prompt into role, context, task, and output format, then applies targeted improvements based on the top 5 and bottom 5 examples.
      • The optimization loop uses an LLM with a static meta-prompt and structured output for segmentation, weakness analysis, and candidate generation.
    • EN Key Points:
      • arXiv:2608.11219v1 Announce Type: new
      • Abstract: Automatic Prompt Optimization (APO) often rewrites prompts monolithically, which can improve one behavior while degrading others
      • We present SAPO, a segment-level APO method that decomposes prompts into role, context, tasks, and output format, then applies targeted improvements based on to…
  • The optimization loop uses one LLM with static meta-prompts and structured outputs for segmentation, weakness analysis, and candidate generation

  • LLMs in Process Diagram Engineering: From Optimal PFDs to Validated P&IDs

    • Published: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11220v1 Announce Type: new.
      • Abstract: Nowadays, the creation of a process flow diagram (PFD) and its subsequent transformation into a piping and instrumentation diagram (P&ID) is predominantly performed manually.
      • Applying artificial intelligence in the task could potentially lead not only to process automation and time savings, but also to financial gains by exploring numerous topological options for the diagrams and reducing manual labor.
      • This research presents P&ID Pilot - a practical end-to-end AI pipeline capable of handling flowsheet development for both stages.
    • EN Highlights:
      • arXiv:2608.11220v1 Announce Type: new
      • Abstract: Nowadays, the creation of a process flow diagram (PFD) and its subsequent transformation into a piping and instrumentation diagram (P&ID) is predomina…
      • Applying artificial intelligence in the task could potentially lead not only to process automation and time savings, but also to financial gains by exploring nu…
      • This research presents P&ID Pilot - a practical end-to-end AI pipeline capable of handling flowsheet developing for both stages
  • A Conceptual Framework for Refining Influence Knowledge from Simulation Evidence in Cyber-Physical Systems

    • Published: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11221v1 Announce Type: new.
      • Abstract: Cyber-physical systems (CPS) are typically developed by multiple stakeholders who produce artifacts tailored to their specific domains of expertise.
      • The behavior of these systems emerges from the interaction between those artifacts and their operational environment.
      • Simulation and co-simulation have become essential approaches for analyzing CPS behavior, and through simulation campaigns, developers can explore system responses under changing conditions, including interactions with the environment.
    • EN Highlights:
      • arXiv:2608.11221v1 Announce Type: new
      • Abstract: Cyber-physical systems (CPS) are typically developed by multiple stakeholders who produce artefacts tailored to their specific domains of expertise
      • The behaviour of these systems emerges from the interaction between those artefacts and their operational environment
      • Simulation and co-simulation have become essential approaches for analysing CPS behaviour and, through simulation campaigns, developers can explore system respo…

ArXiv cs.CL (B_intro+search) Link to heading

  • Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs

    • Publication Time: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11232v1 Announcement Type: New.
      • Abstract: Evaluating LLM coding agents in algorithmic trading is difficult, as static benchmarks risk data contamination, and numerical backtest outputs require ground truth from actual code execution.
      • We introduce Backtrader-Bench, a framework with two complementary pipelines.
      • The deterministic multiple-choice question (MCQ) pipeline generates questions from backtest configurations across five trading strategies, 33 templates, and three difficulty levels, with each answer re-derived by an independent checker.
    • EN Highlights:
      • arXiv:2608.11232v1 Announce Type: new
      • Abstract: Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs requ…
      • We present Backtrader-Bench, a framework with two complementary pipelines
      • A deterministic multiple-choice question (MCQ) pipeline generates questions from backtest configurations across five trading strategies, 33 templates, and three…
  • Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets

    • Publication Time: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11233v1 Announcement Type: New.
      • Abstract: Dense pretrained language models can be retrofitted with recurrent depth and learn iterative latent transitions that persist after outcome-only annealing.
      • Qwen2.5-0.5B-Instruct is divided into a prelude, a weight-tied recurrent block, and a coda, with an identity-preserving single-loop path and a re-entry bridge on later loops.
      • In loop 1, the retrofit remains non-inferior to its base on a pre-registered ARC battery.
    • EN Highlights:
      • arXiv:2608.11233v1 Announce Type: new
      • Abstract: A dense, pretrained language model can be retrofitted with recurrent depth and learn an iterative latent transition that persists after outcome-only a…
      • Qwen2.5-0.5B-Instruct is split into a Prelude, a weight-tied Recurrent Block, and a Coda, with an identity-preserving one-loop path and a re-entry bridge on lat…
      • At loop 1 the retrofit remains non-inferior to its base on a preregistered ARC battery
  • TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation

    • Publication Time: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11236v1 Announcement Type: New.
      • Abstract: Roleplay evaluation should not just yield a single score: it should reveal which role requirements were tested, which were failed, and what conversational evidence supports the judgment.
      • We propose TRACE Bench, a task-driven agentic checklist evaluation framework.
  • It decomposes each role profile offline into a fixed checklist, then uses a User Agent to converse naturally with the target roleplay model while privately updating the checklist status based on model responses.

    • EN Key Points:
      • arXiv:2608.11236v1 Announce Type: new
      • Abstract: Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue…
      • We propose TRACE Bench, a task-driven agentic checklist evaluation framework
      • It decomposes each role profile offline into a fixed checklist, then uses a User Agent to converse naturally with the target roleplay model while privately upda…
  • Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction

    • Publication Date: 2026-08-13 12:00 Beijing Time
    • Abstract:- arXiv:2608.11242v1 Announce Type: new.
      • Abstract: When the context window is under pressure, LLM systems compact prior context to continue ongoing tasks.
      • We identify a class of user-issued instructions, namely Session Constraints (SCs), such as “do not delete any emails until I confirm,” which are intended to constrain the LLM’s behavior for the remainder of the session but are silently dropped during compaction.
      • To quantify this loss, we introduce COMPINT, an evaluation suite that evaluates compactors across three long-context scenarios: multi-turn chat, agentic trajectories, and long-horizon research.
    • EN Key Points:
      • arXiv:2608.11242v1 Announce Type: new
      • Abstract: When the context window is under pressure, LLM systems compact prior context to continue ongoing tasks
      • We identify a class of user-issued instructions, Session Constraints (SCs), such as “do not delete any emails until I confirm,” that are meant to constrain LLM’…
      • To quantify this loss, we introduce COMPINT, an evaluation suite that evaluates compactors across three long-context scenarios: multi-turn chat, agentic traject…
  • Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression

    • Publication Date: 2026-08-13 12:00 Beijing Time
    • Abstract:- arXiv:2608.11249v1 Announce Type: new.
      • Abstract: We investigate the problem of lossless text compression, motivated by the rapid growth in collection and storage of digital text data (including plain text, source code, and structured formats like XML), and recent advances in compression based on neural language models.
      • In particular, recent LLM-based approaches, whether built on symbolic ranking pipelines or used in conjunction with statistical compressors, have demonstrated significantly better compression ratios than general-purpose compressors such as zstd, gzip, or bzip on text and code.
      • However, these neural methods suffer from severe throughput limitations, making them not yet practical for real-world use.
    • EN Key Points:
      • arXiv:2608.11249v1 Announce Type: new
  • Abstract: We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including pla…

  • In particular, recent LLM-based approaches, whether built on symbol-ranking pipelines or paired with a statistical compressor, have demonstrated compression rat…

  • However, these neural approaches suffer from severe throughput limitations, making them not yet practically usable

  • Gloss-Free Representation Learning for Cross-Dataset Sign Spotting

    • Published: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11332v1 Announce Type: new.
      • Abstract: Research on sign languages for resource-constrained languages is often limited by the cost of dense linguistic labels (e.g., glosses, temporal boundaries, and sign order).
      • Broadcast news offers a practical alternative by pairing continuous signing with spoken-language transcripts, but this supervision is weak as the text and signing are loosely aligned.
      • Morphologically rich languages (such as Turkish) add further difficulty, as the same lexical meaning can appear in many inflected forms, while some derived forms should remain distinct.
    • EN Highlights:
      • arXiv:2608.11332v1 Announce Type: new
      • Abstract: Sign-language research for resource-constrained languages is often limited by the cost of dense linguistic labels such as glosses, temporal boundaries…
      • Broadcast news offers a practical alternative by pairing continuous signing with spoken-language transcripts, but this supervision is weak since text and signin…
      • Morphologically rich languages such as Turkish add further difficulty, as the same lexical meaning can appear in many inflected forms while some derived forms s…
  • Better, Faster, Stronger: Programmatic Skill Learning Best Reduces Agent Cost

    • Published: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11338v1 Announce Type: new.
      • Abstract: Recently, the practice of augmenting LLM agent capabilities with skills has become prevalent.
      • We explore the cost-effective adaptation of agents to new domains by learning skills.
      • Existing work focuses on performance gains rather than cost-effectiveness.
    • EN Highlights:
      • arXiv:2608.11338v1 Announce Type: new
      • Abstract: Recently, the practice of augmenting LLM agent capability with skills has gained prevalence
      • We explore the cost effective adaptation of agents to novel domains by means of learning skills
      • Existing works focus on performance gain over cost effectiveness
  • Self-Evolving Embodied Agents via Skill-Harness Evolution

    • Publication Time: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11350v1 Announce Type: new.
      • Embodied agents are increasingly built as systems around foundation models, where performance depends not only on model weights but also on the skills, context, operational interfaces, and execution tools surrounding the model.
      • While supervised fine-tuning and reinforcement learning can adapt agents to new environments, they require additional data, rewards, and training runs; meanwhile, many train-free, code-centric methods rely on programmable robot APIs that may not be available in fixed-interface settings.
      • We propose SHAPER, a self-evolving framework for train-free embodied adaptation that keeps model parameters frozen and improves the non-parametric agent system by evolving reusable skills and context code utilization through target environment rollouts.
    • EN Key Points:
      • arXiv:2608.11350v1 Announce Type: new
      • Abstract: Embodied agents are increasingly built as systems around foundation models, where performance depends not only on model weights but also on the skills…
      • While supervised fine-tuning and reinforcement learning can adapt agents to new environments, they require additional data, rewards, and training runs; meanwhil…
      • We propose SHAPER, a self-evolving framework for train-free embodied adaptation that keeps model parameters frozen and improves the non-parametric agent system…
  • ODE-Based Transformer Decoders for Iterative Sign Language Translation

    • Publication Time: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11352v1 Announce Type: new.
      • Sign language translation has achieved strong results with Transformer architectures, yet recent improvements largely rely on scaling model capacity at the cost of increased computation.
      • We propose a parameter-efficient alternative that improves expressiveness without increasing model size.
      • Rather than scaling capacity, we focus on enhancing the update dynamics of iterative refinement decoders, where each refinement step corresponds to one internal decoder iteration that progressively improves the latent representation before the translation is generated.
    • EN Key Points:
      • arXiv:2608.11352v1 Announce Type: new
      • Abstract: Sign language translation has achieved strong results with Transformer architectures, yet recent improvements largely rely on scaling model capacity a…
      • We propose a parameter-efficient alternative that improves expressiveness without increasing model size
      • Rather than scaling capacity, we focus on enhancing the update dynamics of iterative refinement decoders, where each refinement step corresponds to one internal…
  • Measure, Don’t Optimize: Forecasting Recovery in LLM Unlearning

    • Publication Time: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11408v1 Announce Type: new.
  • Abstract: Prior white-box studies show that large language models can retain latent traces of target knowledge after unlearning, even when this knowledge is no longer expressed in their outputs.

  • However, existing audits remain limited to one-off diagnostics: it is unclear whether these residual signals can predict future recovery under continued training or serve as reliable optimization targets.

  • Resolving this gap is essential to determine whether internal auditing can move beyond post-hoc evaluation toward proactive risk monitoring and safer unlearning.

  • EN Highlights:

    • arXiv:2608.11408v1 Announce Type: new
    • Abstract: Prior white-box studies show that large language models can retain latent traces of target knowledge after unlearning, even when the knowledge is no l…
    • However, existing audits remain limited to one-off diagnostics: it is unclear whether these residual signals can predict future recovery under continued trainin…
    • Resolving this gap is essential to determine whether internal auditing can move beyond post-hoc evaluation toward proactive risk monitoring and safer unlearning

ArXiv cs.LG (B_intro+search) Link to heading

  • FarSky: Task-Aware Latent-Space Coupling for Generative Intra-Hour Solar Forecasting

    • Published: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11254v1 Announce Type: new.
      • Abstract: Accurate solar irradiance forecasting is essential for the reliable integration of photovoltaic power into modern electricity grids.
      • All-sky imagers (ASI) provide high-resolution observations of clouds, making them well suited for intra-hour forecasting.
      • Recent deep learning approaches have substantially improved forecast accuracy but are often limited by deterministic predictions and a reduced capability to anticipate ramp events.
    • EN Highlights:
      • arXiv:2608.11254v1 Announce Type: new
      • Abstract: Accurate solar irradiance forecasting is essential for the reliable integration of photovoltaic power into modern electricity grids
      • All-sky imagers (ASI) provide high-resolution observations of clouds, making them well suited for intra-hour forecasting
      • Recent deep learning approaches have substantially improved forecast accuracy but are often limited by deterministic predictions and a reduced capability to ant…
  • Why AI Detection Fails for Academic Integrity

    • Published: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11256v1 Announce Type: new.
      • Abstract: Institutions use commercial AI detectors to ensure academic integrity, but the detectors cannot distinguish between AI editing and full LLM drafts, and may treat both as misconduct.
      • A controlled study of published English abstracts (four fields; 2013-2015 vs.
      • 2023-2025), we quantify this policy failure under proxy human/AI labels with tau=0.50.
    • EN Highlights:
      • arXiv:2608.11256v1 Announce Type: new
  • Abstract: Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both a…

  • In a controlled study of published English abstracts (four domains; 2013 to 2015 vs

  • 2023 to 2025), we quantify this policy failure under proxy human/AI labels at tau=0.50

  • Basin: Efficient and Extensible Numerical Optimization in Rust

    • Published time:2026-08-13 12:00 Beijing Time
    • Abstract:- arXiv:2608.11279v1 Announce Type: new.
      • Abstract: Basin is a numerical optimization library for the Rust programming language.
      • Numerical optimization is the task of finding the inputs that minimize a function, and it is a fundamental element across the sciences: fitting a model to data, calibrating simulations, training machine learning models, or selecting engineering parameters that minimize cost.
      • Basin gives users a single, consistent way to both state and solve such problems, with a broad catalog of solvers and first-class support for constraints.
    • EN Highlights:
      • arXiv:2608.11279v1 Announce Type: new
      • Abstract: Basin is a numerical optimization library for the Rust programming language
      • Numerical optimization is the task of finding the inputs that minimize a function, and it is a fundamental element across the sciences: fitting a model to data,…
      • Basin gives users a single, consistent way to both state and solve such problems, with a broad catalog of solvers and first-class support for constraints.
  • Federated Learning for Distributed CNC Tool Wear Prediction

    • Published time:2026-08-13 12:00 Beijing Time
    • Abstract:- arXiv:2608.11281v1 Announce Type: new.
      • Abstract: Tool wear prediction is an important task in CNC machining, where accurate monitoring of tool condition supports product quality and process reliability.
      • Machine learning methods have shown potential for this task, but their use in industrial environments is limited by the distributed nature of machining data and restrictions on data sharing between machines, sites, or organizations.
      • Federated learning provides a suitable framework for this setting by enabling collaborative model training without transmitting raw operational data.
    • EN Highlights:
      • arXiv:2608.11281v1 Announce Type: new
      • Abstract: Tool wear prediction is an important task in CNC machining, where accurate monitoring of tool condition supports product quality and process reliabili…
      • Machine learning methods have shown potential for this task, but their use in industrial environments is limited by the distributed nature of machining data and…
  • Federated learning offers a suitable framework for this setting by enabling collaborative model training without transferring raw operational data

  • Terminal Symmetry as a Decision Resource: Statewise Refinement for Anytime Verified Construction

    • Publication Date: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11318v1 Announce Type: new.
      • Abstract: Many sequential construction tasks exhibit exact symmetry at completion, while their execution remains directed and history-dependent.
      • We develop a decision-resource view of terminal symmetry: process evidence supplies directionality, terminal correspondence transfers this structure between equivalent results, enabling state evidence to refine its current decision relevance after transformation, and a fixed verifier validates execution.
      • This decomposition yields transport—refine—certify.
    • EN Highlights:
      • arXiv:2608.11318v1 Announce Type: new
      • Abstract: Many sequential construction tasks exhibit exact symmetry at completion while their execution remains directed and history-dependent
      • We develop a decision-resource view of terminal symmetry: process evidence supplies directionality, terminal correspondence transports that structure across equ…
      • This decomposition yields transport–refine–certify
  • Contextual Quality-Diversity Evolutionary Reinforcement Learning for HVAC Control in Tropical Commercial Buildings

    • Publication Date: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11324v1 Announce Type: new.
      • Abstract: This paper proposes a contextual quality-diversity evolutionary reinforcement learning controller, CQD-ERL, for the supervisory control of tropical water-cooled chillers and their associated air-side.
      • Instead of converging to a single scalarized policy, the controller maintains a product archive of specialized policies jointly indexed by a data-driven operational context, a set of daily weather and load conditions, and context-invariant behavior descriptors, populated by a gradient-free evolutionary operator and a soft actor-critic policy gradient operator sharing a replay buffer.
      • Every action is filtered through a deterministic safety shield before execution.
    • EN Highlights:
      • arXiv:2608.11324v1 Announce Type: new
      • Abstract: This paper proposes a contextual quality-diversity evolutionary reinforcement-learning controller, CQD-ERL, for the supervisory control of a tropical,…
      • Rather than converging to a single scalarised policy, the controller maintains a product archive of specialised policies indexed jointly by a data- driven opera…
      • Every action is filtered through a deterministic safety shield before execution
  • Long-Horizon Forecasting of Complete Financial Statements with Forma

    • Publication Date: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11327v1 Announce Type: new.
  • Abstract: Specialist training beats generalist scale when forecasting financial statements.

  • To our knowledge, no prior work jointly forecasts complete financial statements beyond one year, yet in a discounted-cash-flow valuation most firm value sits beyond that window.

  • We release ProForma-20Q, a reproducible benchmark for forecasting 78 statement line items 1-20 quarters ahead, for anonymized firms, from past statements and an industry code, scored by variance-space $R^2$.

  • EN Highlights:

    • arXiv:2608.11327v1 Announce Type: new
    • Abstract: Specialist training beats generalist scale when forecasting financial statements
    • To our knowledge, no prior work jointly forecasts complete financial statements beyond one year, yet in a discounted-cash-flow valuation most firm value sits pa…
    • We release ProForma-20Q, a reproducible benchmark for forecasting 78 statement line items 1-20 quarters ahead, for anonymized firms, from past statements and an…
  • Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport

    • Publication Time: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11342v1 Announce Type: new.
      • Abstract: Supervised fine-tuning (SFT) is a standard approach for adapting LLMs to a target distribution, but its cost becomes prohibitive in settings such as personalization, where each author requires separate weight access, optimization, storage, and retraining.
      • We propose Weightless Fine-Tuning (WFT), a training-free decoding-time method that approximates the distributional effect of SFT without weight updates.
      • WFT computes supervised residuals on an author’s training sequence and transports them to the current prompt through a cross-prefix transport operator estimated from dropout-induced inter-covariance.
    • EN Highlights:
      • arXiv:2608.11342v1 Announce Type: new
      • Abstract: Supervised fine-tuning (SFT) is a standard approach for adapting LLMs to a target distribution, but in settings such as personalization, where each au…
      • We propose Weightless Fine-Tuning (WFT), a training-free decoding-time method that approximates the distributional effect of SFT without weight updates
      • WFT computes supervised residuals on an author’s training sequence and transports them to the current prompt through a cross-prefix transport operator estimated…
  • Dynamics Models for Offline Hyperparameter Selection in Real-World RL

    • Publication Time: 2026-08-13 12:00 Beijing Time
    • Abstract: - arXiv:2608.11349v1 Announce Type: new.
      • Abstract: A key obstacle to deploying Reinforcement Learning in real-world systems is hyperparameter selection, especially when simulators are unavailable and online experimentation is costly.
      • Prior work has proposed calibrated models trained on offline data to approximate environment dynamics and enable offline hyperparameter selection, but to date, these methods have only been evaluated in simple simulated settings.
      • In this paper, we demonstrate the first application of calibrated models in a real-world industrial setting: a municipal water treatment plant.
    • EN Highlights:
  • arXiv:2608.11349v1 Announce Type: new

  • Abstract: A key obstacle to deploying reinforcement learning in real-world systems is hyperparameter selection, particularly when simulators are unavailable and…

  • Prior work has proposed calibration models trained on offline data to approximate environment dynamics and enable offline hyperparameter selection, but these me…

  • In this paper, we present the first application of calibration models in a real-world industrial setting: a municipal water treatment plant

  • Market-Information-Aware Gated-LoRA of Foundation Models for Transferable Day-Ahead Electricity Price Forecasting

    • Publication Time: 2026-08-13 12:00 Beijing Time
    • Abstract:
      • arXiv:2608.11359v1 Announce Type: new
      • Abstract: Electricity price forecasting is crucial for market participants but remains difficult because prices are volatile, market-specific, and closely tied to expected system conditions.
      • Existing supervised methods depend largely on market-specific historical data, limiting their use in newly established or data-scarce markets.
      • This paper proposes a market-information-aware adaptation framework that transfers the Chronos-2 time-series foundation model to day-ahead electricity price forecasting.
    • EN Key Points:
      • arXiv:2608.11359v1 Announce Type: new
      • Abstract: Electricity price forecasting is crucial for market participants but remains difficult because prices are volatile, market-specific, and closely tied…
      • Existing supervised methods depend largely on market-specific historical data, limiting their use in newly established or data-scarce markets
      • This paper proposes a market-information-aware adaptation framework that transfers the Chronos-2 time-series foundation model to day-ahead electricity price for…