🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-07-10
- 类型
- ai-daily
- 字数
- 7734
- 阅读时长
- 37 min
2026-07-10 AI Daily | OpenAI Pushes GPT-5.6 into Office: The Enterprise AI Race Enters the Workflow Gateway Link to heading
Today’s main story is OpenAI’s integration of GPT-5.6 into Microsoft 365 Copilot and the launch of ChatGPT Work, embedding model capabilities directly into enterprise workflows like documents, spreadsheets, and slides. Meanwhile, the focus of agent implementation is shifting from “task completion” to execution trajectories, permission boundaries, orchestration costs, and verifiability. The role of developers is also evolving: they are becoming more like engineering managers than just coders.
📖 In-depth Guide to This Issue’s Watch List Link to heading
The top story today is OpenAI’s set of updates: GPT-5.6 is being integrated into Microsoft 365 Copilot, and ChatGPT Work is positioned as a work agent that can deliver documents, spreadsheets, slides, and web applications across apps. This means “model upgrades” are rapidly reaching enterprise productivity gateways. Product and engineering teams should pay close attention to its workflow boundaries, permissions, and delivery quality.
The second main theme is agent evaluation and cost control. AgentLens emphasizes reviewing the complete execution trajectory, not just task success. Meanwhile, the “Harness Effect” suggests that the true cost lever for enterprise agents is at the orchestration layer, not simply waiting for token prices to fall. These insights are crucial for teams implementing agents.
On the research front, key areas include reasoning and tool augmentation: contextual search theory, ARC low-cost agents, and SageMath-enhanced mathematical agents all address how to achieve verifiable capabilities with less trial-and-error and more reliable feedback. Also noteworthy is the GPT-5.5 Bio Bug Bounty, which security and governance teams should follow.
🌐 AI Hot Takes from X Link to heading
Topic 1: SpaceXAI Launches Grok 4.5 for Coding and Engineering Tasks Link to heading
- Category: AI · News
- Overview: Trending for: 1 day ago, Related posts: 159,000
- What it is: Buzz on X claims SpaceXAI has released its flagship model, Grok 4.5, for coding and engineering tasks, highlighting high-speed inference, complex software development capabilities, and lower API call prices.
- Why it matters: If true, this would intensify competition in the AI coding assistant and enterprise model markets, especially in terms of cost, speed, engineering task automation, and multi-model platform integration, putting pressure on competitors like OpenAI and Anthropic.
- Discussion summary: Discussions focus on whether Grok 4.5’s real-world coding capabilities live up to the claims, if its low-price strategy will attract developers, whether the investment in AI infrastructure will hurt financial performance, and the validity of rumors about related acquisitions and corporate restructuring.
Topic 2: Anthropic Adds /checkup Command to Clean Up Claude Code Link to heading
- Category: AI · News
- Overview: Trending for: 22 hours ago, Related posts: 2,300
- What it is: Anthropic has added a “/checkup” command to Claude Code, designed to check project status, clean up context, and help developers organize their coding workflows.
- Why it matters: This reflects a shift in AI coding tools from one-off code generation towards continuous collaboration and engineering maintenance, emphasizing context management, code quality, and long-term project viability.
- Discussion summary: The discussion on X focuses on whether this feature can reduce context confusion in Claude Code and boost efficiency on large projects. Some users question its practical effectiveness, wonder if the automated cleanup could omit important information, and debate whether it holds a clear advantage over competing AI coding tools.
Topic 3: OpenAI Launches GPT-Live for Natural Voice Conversations Link to heading
- Category: AI · News
- Overview: Trending for: 1 day ago, Related posts: 55,000
- What it is: OpenAI has launched the GPT-Live voice model, enabling ChatGPT to support more natural, real-time, full-duplex voice conversations. It can listen and speak at the same time, handle interruptions, and offer features like real-time translation.
- Why it matters: This signals a major shift for AI assistants from text-based Q&A toward natural voice interaction, which could transform how users interact with AI and push voice models, real-time inference, and multimodal experiences to the forefront of competition.
- Discussion summary: The discussion on X centers on whether GPT-Live delivers a truly human-like conversational experience, the impact of full-duplex voice on customer service and personal assistant applications, the capability differences between free and paid tiers, and whether OpenAI will further extend its lead in the voice AI space.
Topic 4: OpenClaw Foundation Launches to Secure Open-Source AI Forever Link to heading
- Category: AI · News
- Overview: Trending for: 20 hours ago, Related posts: 691
- What it is: The OpenClaw Foundation has been established with the goal of ensuring open-source AI projects remain permanently open and accessible through foundation-based governance and long-term resource support.
- Why it matters: As foundational AI models and toolchains become increasingly centralized among a few companies, the governance, funding, security, and long-term maintenance of open-source AI have become critical industry issues. The emergence of this foundation is seen as an exploration of the sustainability of the open ecosystem.
- Discussion overview: Discussions on X primarily focus on whether the foundation can truly prevent open-source AI from being commercialized or closed off, whether its governance structure is transparent and trustworthy, and how to balance the security risks and innovative freedom of open models.
Topic 5: ICML 2026 Spotlights Agentic AI Breakthroughs in Seoul Link to heading
- Category: AI · News
- Overview: Trending: 15 hours ago, Related posts: 88
- What it is: ICML 2026 will be held in Seoul and will focus on showcasing research and breakthroughs in Agentic AI.
- Why it matters: This indicates that the focus of AI research is shifting from the capabilities of single models to systems capable of planning, tool use, collaboration, and autonomous decision-making. This has significant implications for next-generation AI applications and security governance.
- Discussion overview: Discussions on X are mainly centered on whether Agentic AI will become the core focus of ICML 2026, the significance of Seoul hosting the event for the Asian AI ecosystem, and the ongoing challenges for agentic systems in terms of reliability, evaluation standards, and security risks.
AI Public Opinion Summary on X Today Link to heading
The main narrative today is that the AI competition is shifting from “single model capabilities” to system-level abilities that are closer to real-world use cases: code engineering, real-time voice, agent collaboration, context maintenance, and open-source governance have all become focal points. The consensus is that developers and users are increasingly prioritizing low cost, high speed, long-term availability, and natural interaction experiences. AI tools are evolving from one-off generation to continuous collaboration and autonomous execution. Disagreements mainly center on whether the announced or rumored capabilities of various companies live up to their claims, such as the authenticity and actual programming proficiency of Grok 4.5, the engineering value of new Claude Code features, whether GPT-Live is truly close to human conversation, and whether open-source foundations can maintain transparency and independence. Potential risks include over-marketing leading to an expectation bubble, insufficient reliability and security evaluation for agentic systems, automated tools missing critical context, and the difficulty for the open-source ecosystem to balance commercialization, security regulation, and long-term funding.
💡 Influencer Insights Link to heading
AI Industry Daily Briefing (07/09-07/10) Link to heading
1. Today’s Top Story: OpenAI’s Full-Suite Release and the Escalating Model Arms Race Link to heading
Today’s discussion was almost entirely dominated by OpenAI, sparked by the official public launch of the full GPT-5.6 model series (Sol/Terra/Luna), accompanied by the grand unification of the ChatGPT and Codex applications.
Model Matrix and Platform Unification: According to a deep dive by @dotey, among the three models released by OpenAI, Sol is the flagship, specializing in complex reasoning and autonomous work; Terra focuses on cost-effectiveness; and Luna is the lightweight, high-speed version. Accompanying the model release is the ChatGPT Work feature, which transforms the AI from a chat assistant into an agent capable of executing tasks across applications, connecting to tools like Google Drive and Slack. This move is seen by @dotey as a key step in OpenAI’s strategy towards becoming a “super app,” aiming to compete with Google Workspace and Microsoft 365 and build a strong enterprise narrative for its upcoming IPO. Additionally, the launch of the GPT-Live full-duplex voice mode upgrades voice interaction from a “walkie-talkie” style to more natural interruptions and parallel processing. However, users, including @dotey, have reported that its real-world tests show responses like “mhmm” are too frequent, making it slightly annoying.
Top-Tier Competition Enters the “Ultimate Move” Phase: @dotey broke the news and summarized information that GPT-6 will be released within a month, pointing out that this is OpenAI’s move to skip minor version iterations in direct response to Anthropic’s Mythos model. Meanwhile, Meta’s Muse Spark 1.1 and xAI’s Grok 4.5 have also been making waves.
- Grok 4.5 Hands-on: @vista8 ( @vista8) provided an initial review, finding it not comprehensive enough for independent CLI development but superior to Codex in front-end aesthetics. This further enhances its value as a perk for Premium+ subscribers.
- The Sudden Reset of Fable 5: @Pluvio9yte ( @Pluvio9yte) and several other bloggers noticed that with the launch of GPT-5.6, the weekly quota for Claude Code’s Fable 5 was suddenly reset, jokingly referred to as a “defensive maneuver” by Anthropic.
Undercurrents in On-Device and Open-Source Models: While the giants are waging a nuclear war, edge models are still making breakthroughs. @zhixianio (@zhixianio) conducted in-depth tests on Gemma 4 12B Coder and compared it with Qwen 3.6 35B, concluding that although the 12B small model is highly efficient after fine-tuning, it hits a ceiling when handling complex programs that are “long, stateful, and require one-shot generation” due to its limited parameter count. Additionally, China’s Tencent Hunyuan Hy3 model (295B MoE) has garnered significant attention. @ruanyf and @vista8 both pointed out that it achieves a level close to GLM 5.1 with a smaller parameter size, making it suitable for frequent daily use due to its cost-effectiveness.
2. Unique Perspectives & Industry Foresight Link to heading
The New Role in Vibe Coding: Engineering Manager: @dotey (@dotey) proposes that when developing with a Coding Agent, the developer’s role shifts from a programmer to an Engineering Manager (EM). The developer is responsible for breaking down requirements, assigning tasks, and accepting the results. If you don’t review the code and only focus on functionality, you’re like an incompetent EM. He also suggests adopting a “Continuous Integration” mindset for Vibe Coding, having the AI work on one small feature at a time for easier verification and debugging, rather than generating a mountain of unmaintainable code at once.
The Science and Non-Science of Skill Management: @dotey summarized a set of rules for Skill management, citing data from SkillsBench: AI-generated Skills can perform even worse than using no Skills at all, and only Skills produced under the guidance of human experts are valuable. Also, large, comprehensive Skills are less effective than small, focused ones. Software engineering-related Skills provide minimal improvement to models because they are already saturated with such data, whereas Skills in long-tail domains like healthcare show significant boosts.
Effort Must Be Applied in the Right Place: Targeting indie developers in the AI era, @gefei55 (@gefei55) went full-throttle, pointing out that while one used to write one unused app a month, with AI assistance, one can now write 37, but this is just a “token-burning trap.” What’s truly lacking is market research, marketing, and the courage to leave one’s comfort zone.
Xiaohongshu and Skill Distribution: @ruanyf discovered that Xiaohongshu is beta-testing a REDSkill community, allowing users to distribute AI Agent Skill files on the platform, attempting to combine social media with a Skill Hub to create a “GitHub for Skills.” This provides a new channel for developers to reach a massive non-technical user base.
3. Recommended Tools & Resources Link to heading
Open Source Projects & Releases:
- RN-Skill Collection (Open-sourced by @Pluvio9yte): A set of AI Agent Skills covering writing, video production, and quality inspection, especially suitable for refining AI-generated articles to remove the “AI flavor” and for directing motion graphics videos.
- Qiaomu RSS Reader (Open-sourced by @vista8): Features AI-powered automatic translation and rewriting of newsletters, integrates quality sources like Hacker News, and is ideal for alleviating information overload.
- Topview 3D Shot Composer (Recommended by @AI_Jasonyu): A new tool for solving composition challenges in AI video creation. It allows you to first set up character positions and camera angles in a 3D space, and then have the AI generate the scene.
Development Aids & Plugins:
- Obsidian → X Long-form Article Plugin (Developed by @kaitoxhacker, recommended by @AI_Jasonyu): Solves the pain point of publishing long articles from local Markdown notes to the X platform with one click.
- WeChat Official Account Batch Downloader Skill: A tool based on Python’s standard library that automatically converts WeChat Official Account articles to Markdown, downloads images, and creates an index.
Design Aesthetics & References:
- Apple-Design Skill (Published by @emilkowalski, recommended by @vista8): Summarizes 17 design and motion principles from Apple’s WWDC videos, which can effectively improve the aesthetics of AI-generated front-end interfaces.
📚 Appendix: Today’s Watch List Source Updates Link to heading
Time window: Last 3 days; 22 sources covered; 36 updates in total
Y Combinator Podcast (B_intro+search) Link to heading
- How To Better Understand Your Users
- Publication Time: 2026-07-10 04:25 Beijing Time
- Summary: - You may have already heard of OpenClaw (formerly known as Clawdbot/Moltbot).
- The sensational open-source AI assistant that runs on your own device, connects with the messaging apps you already use, and goes beyond chat to actually perform tasks like managing email, calendars, files, workflows, and more.
- Now meet the person behind it.
- YC’s Raphael Schaad sits down with Peter Steinberger, founder of OpenClaw, to discuss the “aha” moment behind the viral personal AI agent, why a local-first agent could replace many of today’s apps, and how personal agents will reshape the future of software.
- EN Highlights:
- Most founders obsess over dashboards and aggregate metrics, but some of the best product insights come from understanding how individual users actually use thei…
- In this episode of Startup School, YC’s David Lieb walks through one of his favorite tools for better understanding your users, the dot plot
- It’s a simple two-dimensional grid that reveals usage patterns no aggregate chart can show you
- He’ll cover why it gives founders a better sense of product health, what patterns to look for, and real-world exam
- EN Highlights:
Stratechery by Ben Thompson (A_full) Link to heading
- Muse Image, Grok 4.5, Alex Karp on CNBC
- Published: 2026-07-09 18:00 Beijing Time
- Summary: - The battle for verifiable data is increasingly defining the AI race, from Meta to Grok to the frontier labs.
- $15/month* or *$150/year.
- Substantial analysis of the day’s news via three weekly emails or podcasts.
- Strategy Interviews.
- Interviews with leading public company CEOs, private company founders, and discussions with fellow analysts.
- EN Highlights:
- The batter for verifiable data is increasingly defining the AI race, from Meta to Grok to the frontier labs.
OpenAI Blog (A_full) Link to heading
GPT-5.6 is now the preferred model in Microsoft 365 Copilot
- Published: 2026-07-09 21:00 Beijing Time
- Summary: - Today, OpenAI released GPT‑5.6, which will become the new preferred model in Microsoft 365 Copilot (Word, Excel, PowerPoint, Chat, and Cowork).
- For Microsoft 365 customers, this update brings OpenAI’s latest flagship model series into the productivity tools people use every day, helping them leverage more powerful AI assistance to create, analyze, and collaborate within their workflows.
- GPT-5.6 is OpenAI’s latest flagship model series, delivering more useful work from every token, with stronger price-performance and on-demand capabilities for the most complex tasks.
- With GPT‑5.6, Microsoft 365 users will be able to create higher-quality work products with less effort in the applications they already rely on:
- In Word, GPT-5.6 can help people draft, edit, and refine documents with fewer prompts.
- EN Highlights:
- Learn how GPT-5.6 powers Microsoft 365 Copilot with stronger AI capabilities across Word, Excel, PowerPoint, Chat, and Cowork for faster, higher-quality work.
ChatGPT is now a partner for your most ambitious work
- Publication Time: 2026-07-09 18:00 Beijing Time
- Abstract: - Introducing ChatGPT Work, an agent within ChatGPT that helps you take on more demanding tasks.
- It can gather information across applications and workflows to create finished materials like worksheets, slides, documents, and web applications, handling them in hours by breaking down complex projects into smaller steps and completing them independently.
- With built-in Codex technology, ChatGPT can now not only answer questions but also complete actual work across web, mobile, and desktop devices.
- Over 5 million people use Codex every week.
- Although it was originally created as a coding agent for developers, over 1 million people now use it for work outside of software development, demonstrating how its capabilities support a broader range of tasks.
- EN Key points:
- ChatGPT Work is an agent that can take action across your apps and files, stay with a project for hours if needed, and turn a goal into finished work.
- Publication Time: 2026-07-09 18:00 Beijing Time
- Abstract: - Detailed information about the OpenAI Bio Bounty program.
- This article from the OpenAI blog explains how the GPT-5.5 Bio Bug Bounty is shaping the broader AI and infrastructure landscape.
- The GPT-5.5 Bio Bug Bounty also has practical implications for founders, operators, and investors.
- EN Key points:
- Details about the OpenAI Bio Bounty program
GPT-5.6: Frontier intelligence that scales with your ambition
- Publication Time: 2026-07-09 18:00 Beijing Time
- Abstract: - More intelligence from every token, stronger performance per dollar, and more of the capabilities you need for your toughest work.
- This article from the OpenAI blog explains how GPT-5.6: Frontier intelligence that scales with your ambition is shaping the broader AI and infrastructure landscape.
- It also offers practical implications for founders, operators, and investors following GPT-5.6: Frontier intelligence that scales with your ambition.
- EN Key points:
- More intelligence from every token, stronger performance per dollar, and more capability on demand for your hardest work.
ArXiv cs.AI (B_intro+search) Link to heading
AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06624v1 Announcement Type: new.
- Abstract: We introduce AgentLens, a production-assessed benchmark for interactive code agents.
- Most code agent benchmarks reduce a run to a single bit - did the task pass?
- But people who actually use these agents experience the whole trajectory: how the agent follows instructions, uses its tools, verifies its own work, recovers from errors, and talks to them along the way.
- EN Key points:
- arXiv:2607.06624v1 Announce Type: new
- Abstract: We present AgentLens, a production-assessed benchmark for interactive code agents
- Most code-agent benchmarks reduce a run to a single bit – did the task pass
– but the people who actually use these agents experience the entire trajectory: how the agent follows instructions, uses its tools, verifies its own work, rec…
When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06720v1 Announce Type: new.
- Abstract: Training large language models (LLMs) with extended reasoning has enabled in-context search, where models iteratively generate, critique, and modify solution attempts.
- We provide a theoretical analysis of in-context search by modeling it as approximate inference over reasoning trajectories, where the base model defines a prior and self-reflection provides feedback for posterior updates, and study the resulting inference-time sampling complexity—the number of sequential attempts required to achieve a high success probability.
- We show that when reflection reliably localizes early mistakes, in-context search can yield exponential improvements over the base model, solving problems with exponentially small zero-shot pass rates using only a polynomial number of sequential attempts, whereas when this property fails, conditioning on past attempts provides no asymptotic advantage over parallel sampling.
- EN 要点:
- arXiv:2607.06720v1 Announce Type: new
- Abstract: Training large language models (LLMs) with extended reasoning has enabled in-context search, in which models iteratively generate, critique, and revis…
- We provide a theoretical analysis of in-context search by modeling it as approximate inference over reasoning traces, where the base model defines a prior and s…
- We show that when reflections reliably localize early mistakes, in-context search can yield exponential improvements over the base model, solving problems with…
LLM-powered reasoning in agent-based modeling
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06757v1 Announce Type: new.
- Abstract: Agent-based modeling (ABM) enables the modeling of millions of individuals and their interactions, which is highly useful for policymaking.
- However, ABM has traditionally relied on static priors, which prevents models from adapting to real-time changes.
- Our research provides a novel approach to addressing this information gap.
- EN 要点:
- arXiv:2607.06757v1 Announce Type: new
- Abstract: Agent-based modeling (ABM) has the capability to model millions of individuals and their interactions, which is useful for policy making
- However, ABMs have traditionally relied on static prior, which prevents the models from adapting to real-time changes
- Our research provides a novel approach to addressing this information gap
QANTIS: Hardware-Calibrated Sequential POMDP Belief Updates on IBM Heron
- Publication Time: 2026-07-09 12:00 Beijing Time
摘要:
- arXiv:2607.06760v1 Announce Type: new.
- Abstract: Autonomous systems under partial observability act on beliefs, not raw sensor events.
- QANTIS treats the quantum processor as a calibrated belief-update service in that loop: it receives a prior and an observation model, estimates rare event evidence terms, and returns the ordinary posterior to the classical planner.
- This paper asks whether that service can be reused across a sequential Tiger POMDP horizon on present IBM Heron hardware without corrupting the planner-facing posterior.
- EN 要点:
- arXiv:2607.06760v1 Announce Type: new
- Abstract: Autonomous systems under partial observability act on beliefs, not raw sensor events
- QANTIS treats the quantum processor as a calibrated belief-update service in that loop: it receives a prior and an observation model, estimates the rare-event e…
- This paper asks whether that service can be reused across a sequential Tiger POMDP horizon on present IBM Heron hardware without corrupting the planner-facing p…
Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1
- Release Time: 2026-07-09 12:00 Beijing Time
- Abstract:
- arXiv:2607.06764v1 Announce Type: new.
- Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or training specific to the benchmark where small models are fine-tuned on ARC data, often with task-specific architectures.
- We study a third regime: an open-weight model in non-thinking mode (DeepSeek V3.2) under a strict budget, with no ARC-specific fine-tuning.
- We study what is recoverable through architecture alone, building agentic harnesses that explicitly decompose pattern-discovery and program-synthesis stages.
- EN 要点:
- arXiv:2607.06764v1 Announce Type: new
- Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionar…
- We study a third regime: an open-weight model in non-thinking mode (DeepSeek V3.2) under a strict budget, with no ARC-specific fine-tuning
- We study what is recoverable through architecture alone, building agentic harnesses that decompose pattern-discovery and program-synthesis stages explicitly
Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics
- Release Time: 2026-07-09 12:00 Beijing Time
- Abstract:
- arXiv:2607.06820v1 Announce Type: new.
- Abstract: Recent progress in mathematical AI has largely focused on automated formalization and theorem proving, while the role of Computer Algebra Systems (CAS) in agent LLM workflows remains underexplored.
- We propose a ReAct-style agent setup combining LLM reasoning with verifiable feedback from SageMath, and Context7 for up-to-date documentation.
We evaluate this agentic setup across frontier models for solving research-level mathematical problems from the RealMath benchmark in a setting that emulates a computational mathematics research loop.
- EN Highlights:
- arXiv:2607.06820v1 Announce Type: new
- Abstract: Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS…
- We propose a ReAct-style agentic setup that combines LLM reasoning with verifiable feedback from SageMath, together with Context7 for the up-to-date documentati…
- We evaluate this agentic setup across frontier models for solving research-level mathematical problems from the RealMath benchmark in a setting that emulates a…
- EN Highlights:
The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06906v1 Announce Type: new.
- Abstract: Agentic AI development today runs on token maxing: buying capability with tokens—longer reasoning traces, more turns, wider tool payloads, bigger replay contexts—so per-task tokens grow faster than task value.
- Falling per-token prices mask the pattern; total spend rises anyway.
- We argue the decisive lever against token maxing is the harness: the orchestration layer that assembles context, exposes tools, sequences turns, delegates work, and hosts enterprise observability and governance.
- EN Highlights:
- arXiv:2607.06906v1 Announce Type: new
- Abstract: Agentic AI development today runs on token maxing: buying capability with tokens – longer reasoning traces, more turns, wider tool payloads, bigger r…
- Falling per-token prices mask the pattern; total spend rises anyway
- We argue the decisive lever against token maxing is the harness: the orchestration layer that assembles context, exposes tools, sequences turns, delegates work,…
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06925v1 Announce Type: new.
- Abstract: Compact world models conditioned on language goals promise to ground relations like “put the red block left of the blue one” using a sparse set of explicit \emph{referent anchors}.
- We ask when such referents truly ground relations, and identify a pitfall: a goal-conditioned predictor achieves a surprising $0.90 relational readout accuracy, but this is only \emph{instruction transcription}, not perception.
- Withholding the goal collapses it ($0.90!\to!0.27$, three seeds), and counterfactual instructions make predicted anchors follow the \emph{false} instruction $94.5%$ of the time (vs. the true scene $2.3%$; $N{=}256$).
- EN Highlights:
- arXiv:2607.06925v1 Announce Type: new
Abstract: Compact world models that condition on a language goal promise to ground relations such as ``put the red block left of the blue block’’ using a sparse…
We ask when such references actually ground a relation, and identify a trap: a goal-conditioned predictor reaches a striking $0.90$ relation-readout accuracy, y…
Withholding the goal collapses it to chance ($0.90\
Large Behavior Model: A Promptable Digital Twin of the Retail Customer
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06993v1 Announcement Type: new.
- Abstract: Customer behavior modeling is the foundation for recommendation, marketing, and decision support, but existing methods either optimize for predictive accuracy without explaining the decisions, or simulate users without being grounded in real behavioral data.
- We present the Large Behavioral Model (LBM), which learns customer decision-making directly from large-scale retail transactions through a unified person-environment formula.
- Customer state is represented by a behavioral profile derived from historical purchases, while product context is incorporated through retrieval-augmented generation.
- EN Points:
- arXiv:2607.06993v1 Announce Type: new
- Abstract: Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive accuracy with…
- We present the Large Behavioral Model (LBM) that learns customer decision making directly from large-scale retail transactions through a unified Person-Environm…
- Customer state is represented by a behavioral profile derived from historical purchases, while product context is incorporated through retrieval-augmented gener…
Learning social norms enhances compatibility in dynamic human-AI coordination
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.07021v1 Announcement Type: new.
- Abstract: Humans continuously coordinate with others in dynamic interactions, often through implicit, hard-to-quantify social norms that act as shared tacit expectations between interacting agents.
- As AI agents, including Large Language Models (LLMs), are integrated into daily life, they increasingly participate in such interactions and reshape the structure of social interactions.
- However, they often fail to coordinate with humans in an effective, considerate, and natural way.
- EN Points:
- arXiv:2607.07021v1 Announce Type: new
- Abstract: Humans continuously coordinate with others in dynamic interactions, often through implicit, hard-to-quantify social norms that act as shared tacit exp…
As AI agents, including large language models (LLMs), become embedded in daily life, they increasingly participate in such interactions and reshape social inter…
Yet they often fail to coordinate with humans in an effective, considerate, and natural manner
ArXiv cs.CL (B_intro+search) Link to heading
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06611v1 Announcement Type: New.
- Abstract: Automatically identifying positive or negative sentiment in speech is a challenging task, requiring the analysis of vocal changes and the interpretation of spoken words.
- Recent solutions rely on audio foundation models to address the task, but it is unclear whether such models can consider all aspects.
- To this end, we propose a multimodal solution that integrates audio and text information through a cross-modal transformer, where text transcriptions are automatically generated by an Automatic Speech Recognition (ASR) tool.
- EN Key Points:
- arXiv:2607.06611v1 Announce Type: new
- Abstract: Automatically recognizing the sentiment, positive or negative, from speech is a challenging task, requiring both the analysis of vocal inflections and…
- Recent solutions rely on audio foundation models to solve the task, but it remains unclear if such models can take all aspects into account
- To this end, we propose a multimodal solution that integrates audio and text information via cross-modal transformers, where text transcripts are automatically…
Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06641v1 Announcement Type: New.
- Abstract: Large Language Models (LLMs) have achieved promising results on medical question-answering benchmarks, but their use in public health is limited by hallucinations and the rapid evolution of official guidance.
- Retrieval-Augmented Generation (RAG) mitigates these risks by grounding responses in an explicitly maintained corpus, but end-to-end performance primarily depends on retrieval configuration and evaluation beyond multiple-choice formats.
- We extend PubHealthBench (a Question Answering (QA) benchmark of 7,929 questions derived from UK government public health guidance) to a retrieval-augmented environment and systematically evaluate retrieval and generation choices.
- EN Key Points:
- arXiv:2607.06641v1 Announce Type: new
- Abstract: Large language models (LLMs) achieve promising results on medical question answering benchmarks, yet their use in public health is constrained by hall…
Retrieval-Augmented Generation (RAG) mitigates these risks by grounding responses in an explicitly maintained corpus, but end-to-end performance depends critica…
We extend PubHealthBench, a question answering (QA) benchmark of 7,929 questions derived from UK Government public health guidance, into a retrieval-augmented s…
Ad Headline Generation using Self-Critical Masked Language Model
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract:
- arXiv:2607.06818v1 Announce Type: new
- For any E-commerce website it is a nontrivial problem to build enduring advertisements that attract shoppers.
- It is hard to pass the creative quality bar of the website, especially at a large scale.
- We thus propose a programmatic solution to generate product advertising headlines using retail content.
Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract:
- arXiv:2607.06831v1 Announce Type: new
- Speech-to-text alignment means finding the temporal boundaries of each word in the audio.
- Some models provide such an alignment directly and others do not.
- Connectionist temporal classification (CTC) and transducer models have an alignment by construction, whereas attention-based encoder-decoders (AED) and speech l…
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06845v1 Announce Type: new.
Abstract: African American English (AAE), a rule-governed dialect spoken by over 30 million people, is routinely misinterpreted and “corrected” by large language models (LLMs).
Across six instruction-tuned LLMs (14B to 70B), we show that state-of-the-art models systematically prefer Standard American English (SAE) continuations even when the preceding context is in AAE, effectively rewriting AAE to SAE.
We present an end-to-end framework to audit and mitigate this bias.
- EN Highlights:
- arXiv:2607.06845v1 Announce Type: new
- Abstract: African American English (AAE), a rule-governed dialect spoken by over 30 million people, is routinely misinterpreted and “corrected” by large languag…
- Across six instruction-tuned LLMs (14B to 70B), we show that state-of-the-art models systematically prefer Standard American English (SAE) continuations even wh…
- We present an end-to-end framework to audit and mitigate this bias
- EN Highlights:
Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06940v1 Announce Type: new.
- Abstract: The remarkable performance of large language models (LLMs) in linguistic tasks underscores the urgent need for a comprehensive evaluation of their response quality.
- Prevailing methods are often confined to a single dimension, failing to capture the full spectrum of model capabilities.
- This study introduces a multi-factor scoring paradigm that integrates accuracy, conciseness, factual consistency, readability, and coherence, supplemented by a graphical user interface (GUI) for visualizing results.
- EN Highlights:
- arXiv:2607.06940v1 Announce Type: new
- Abstract: The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their respon…
- Prevailing methods, often confined to singular dimensions, fall short of capturing the full spectrum of model capabilities
- This study introduces a multifactor scoring paradigm, integrating accuracy, conciseness, factual consistency, readability, and coherence, complemented by a grap…
MILES: Modular Instruction Memory with Learnable Selection for Self-Improving LLM Reasoning
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06974v1 Announce Type: new.
- Abstract: Large language models (LLMs) are increasingly improving their reasoning abilities at test time through additional computation, but most existing work handles each problem in isolation.
- When problems are presented sequentially, accumulating reusable experience can further enhance performance.
- Existing memory-based methods either store holistic solution templates that generalize poorly to new problems or use heuristic, step-level selection that is not optimized for final answer correctness.
- EN Highlights:
- arXiv:2607.06974v1 Announce Type: new
Abstract: Large language models (LLMs) increasingly improve their reasoning at test time via additional computation, yet most existing works treat each problem…
When problems arrive sequentially, accumulating reusable experience across them can further improve performance
Existing memory-based methods either store whole-solution templates that generalize poorly to novel problems or use heuristic step-level selection that is not o…
Riemannian Geometry for Pre-trained Language Model Embeddings
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.07047v1 Announcement Type: New.
- Abstract: Understanding the geometric structure of pre-trained language model embeddings is crucial for interpretability and safety.
- We ask whether sentence-level classification signal exists in the Riemannian geometry of contextual token embeddings, and we probe it by extracting per-token pullback metrics from the analytic Jacobian of the learned encoder and aggregating them with the Fr’echet mean on the Symmetric Positive Definite (SPD) manifold; we call this process Riemannian Mean Pooling (RMP).
- In three datasets with significant linguistic structure (CoLA, CREAK, RTE), RMP outperforms Euclidean mean pooling, while on FEVER-Symmetric (a benchmark constructed to eliminate annotation-driven lexical artifacts), the method correctly maintains chance performance.
- EN Highlights:
- arXiv:2607.07047v1 Announce Type: new
- Abstract: Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety
- We ask whether sentence-level classification signal lives in the Riemannian geometry of contextual token embeddings, and probe it by extracting per-token pullba…
- Across three datasets with non-trivial linguistic structure (CoLA, CREAK, RTE), RMP outperforms Euclidean mean pooling, while on FEVER-Symmetric, a benchmark co…
Behavior Leverage Imbalance in Multi-Teacher On-Policy Distillation
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.07050v1 Announcement Type: New.
- Abstract: Agentic language models must learn when to call tools, when to consume tool responses, and when to answer directly.
- This makes multi-teacher on-policy distillation a natural training strategy: one teacher can specialize in tool-calling, another can specialize in direct responses, and the student can learn from both on
- its own generated distribution.
- EN Highlights:
- arXiv:2607.07050v1 Announce Type: new
- Abstract: Agentic language models must learn when to call tools, when to consume tool responses, and when to answer directly
This makes multi-teacher on-policy distillation a natural training strategy: one teacher can specialize in tool calls, another in direct responses, and the stud…
its own generated distribution
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.07141v1 Announcement Type: new.
- Abstract: Newly developed items must typically be field-tested before their psychometric properties are known, which creates a cold-start problem for item calibration.
- Predicting item parameters from features is a long-standing measurement problem dating back to the Linear Logistic Test Model; modern text embeddings can now automate the design matrices that were traditionally specified by hand.
- We propose an evaluation framework that combines regularized regression on item text embeddings, reporting of repeated cross-validated R-squared with its resampling standard deviation, and two performance ceilings: a reliability ceiling derived from parameter standard errors, and a design ceiling derived from simulation-based power calibration.
- EN Highlights:
- arXiv:2607.07141v1 Announce Type: new
- Abstract: Newly developed items must ordinarily be field tested before their psychometric properties are known, creating a cold start problem for item calibrati…
- Predicting item parameters from features is a long standing measurement problem dating back to the Linear Logistic Test Model; modern text embeddings now automa…
- We propose an evaluation framework combining regularized regression on item text embeddings, repeated cross validated R squared reported with its resampling sta…
ArXiv cs.LG (B_intro+search) Link to heading
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06601v1 Announcement Type: new.
- Abstract: Conditional computation can decouple language model quality from per-token inference cost, but leading techniques operate on a single axis in isolation: Mixture-of-Experts (MoE) sparsifies the FFN, Mixture-of-Depths (MoD) skips entire transformer blocks, and KV-cache quantization compresses attention memory.
- We argue that these three decisions (attention resolution, expert choice, and cache bit-width) are strongly coupled and should be made jointly: a token rare enough to warrant full attention may also require high-precision caching, regardless of which expert processes it.
- We introduce TriRoute, a single lightweight controller shared across all three axes that, for each token at each layer, emits a coordinated policy for: (i) attention mode (skip/local/full), (ii) a sparse set of FFN experts (with a null expert to recover MoD), and (iii) KV-cache bit-width.
- EN Highlights:
- arXiv:2607.06601v1 Announce Type: new
- Abstract: Conditional computation can decouple language model quality from per-token inference cost, yet leading techniques act on a single axis in isolation: M…
We argue these three decisions (attention resolution, expert selection, and cache bit-width) are strongly coupled and should be made jointly: a token rare enoug…
We introduce TriRoute, a single lightweight controller shared across all three axes that, for every token at every layer, emits a coordinated policy: (i) an att…
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06605v1 Announcement Type: New.
- Abstract: Conformal prediction is adopted in drug discovery to provide an honest number on model reliability: by selecting an error rate alpha, the method returns a prediction set containing the true label with a probability of at least 1 - alpha.
- We demonstrate that this guarantee can be dangerous for imbalanced datasets.
- Across four datasets, standard (marginal) conformal prediction achieved its global 90% coverage target while severely exposing the minority class: the realized minority class coverage decreased to 64.8% for blood-brain barrier permeability and 4.2% for clinical trial toxicity, with the rare class being almost abandoned.
- EN Key Points:
- arXiv:2607.06605v1 Announce Type: new
- Abstract: Conformal prediction is being adopted in drug discovery to put an honest number on model reliability: pick an error rate alpha, and the method returns…
- We show this guarantee can be dangerous on imbalanced datasets
- Across four datasets, standard (marginal) conformal prediction hits its global 90% coverage target while leaving the minority class badly exposed: realized mino…
NEST: Tackling Dataset-Level Distribution Shifts via Regime-Oriented Mixture-of-Experts
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06607v1 Announcement Type: New.
- Abstract: Accurate long-term forecasting in complex systems is frequently affected by dataset-level distribution shifts, where different underlying behavioral patterns and evolving system states drive dynamic multivariate time series.
- While existing methods primarily focus on local temporal variations, they fail to explicitly model global structural challenges where datasets are combinations of different operational mechanisms.
- In this paper, we propose NEST, a specialized framework designed to model and reconstruct these evolving structures through a two-phase dense MoE architecture.
- EN Key Points:
- arXiv:2607.06607v1 Announce Type: new
- Abstract: Accurate long-term forecasting in complex systems is frequently compromised by dataset-level distribution shifts, where diverse underlying behavioral…
While existing methods predominantly focus on local temporal shifts, they fail to explicitly model the global structural challenge where datasets are composites…
In this paper, we propose NEST, a specialized framework designed to model and recompose these evolving structures through a two-phase dense MoE architecture
D2PO: Optimizing Diffusion Samplers via Dynamic Preference
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06609v1 Announcement Type: New.
- Abstract: We propose D2PO (Dynamic Direct Preference Optimization), a principled framework for optimizing diffusion sampling policies with respect to timestep schedules and classifier-free guidance (CFG) weights.
- Our work is motivated by a fundamental limitation of existing student-teacher regression frameworks; low-NFE student samplers trained to mimic high-NFE teachers often sacrifice high-frequency texture fidelity while preserving coarse global structures, thereby misaligning the sampler with perceptual quality.
- D2PO addresses this challenge by reformulating sampler optimization as a preference-based alignment problem, leveraging the Direct Preference Optimization (DPO) framework.
- EN Highlights:
- arXiv:2607.06609v1 Announce Type: new
- Abstract: We propose D2PO (Dynamic Direct Preference Optimization), a principled framework for optimizing diffusion sampling policies with respect to timestep s…
- Our work is motivated by a fundamental limitation of existing student-teacher regression frameworks; low-NFE student samplers are trained to mimic high-NFEteach…
- D2PO addresses this challenge by reformulating sampler optimization as a preference-based alignment problem, leveraging the Direct Preference Optimization (DPO)…
Deep Reinforcement Learning for Reliability Based Bi-Objective Portfolio Optimization
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06610v1 Announcement Type: New.
- Abstract: Portfolio optimization under uncertainty is inherently a multi-objective decision problem involving complex interactions among return, risk, market dynamics, and practical investment constraints.
- Existing reliability-based portfolio optimization methods primarily rely on static optimization frameworks, which often fail to capture sequential decision-making, tail risks, and market frictions such as transaction costs.
- To address these limitations, we propose a deep reinforcement learning framework for multi-objective reliability-based portfolio optimization (MORP-DRL).
- EN Highlights:
- arXiv:2607.06610v1 Announce Type: new
- Abstract: Portfolio optimization under uncertainty is inherently a multi-objective decision problem involving complex interactions among return, risk, market dy…
Existing reliability based portfolio optimization approaches primarily rely on static optimization frameworks and often fail to capture sequential decision maki…
To address these limitations, we propose a deep reinforcement learning framework for multi-objective reliability based portfolio optimization (MORP-DRL)
STAGformer: A Spatio-temporal Agent Graph Transformer for Micro Mobility Demand Forecasting
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06614v1 Announcement Type: New.
- Abstract: Accurate station-level demand forecasting is crucial for the efficient operation of bike-sharing systems, but it remains challenging due to complex spatiotemporal dependencies and large-scale urban networks.
- This paper proposes STAGformer, a spatiotemporal agent graph transformer that enables efficient global modeling with linear computational complexity.
- The model introduces a two-step agent attention mechanism, where a small set of learnable spatial and temporal agent tokens first aggregate global information and then broadcast it back to individual stations and timesteps, effectively capturing long-range interactions while reducing the quadratic cost of standard self-attention O(NT).
- EN Highlights:
- arXiv:2607.06614v1 Announce Type: new
- Abstract: Accurate station-level demand forecasting is essential for the efficient operation of bike-sharing systems, yet it remains challenging due to complex…
- This paper presents STAGformer, a Spatio-Temporal Agent Graph Transformer that achieves efficient global modeling with linear computational complexity
- The model introduces a two-step agent attention mechanism, where a small set of learnable spatial and temporal agent tokens first aggregate global information a…
WHERE to Generate Matters: Budget-Aware Synthetic Augmentation for Label Skewed Federated Learning
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06616v1 Announcement Type: New.
- Abstract: Label skew in Federated Learning (FL) leads to client drift and reduces global accuracy.
- Synthetic data augmentation can reduce this imbalance; however, full-class balancing requires significant computational cost.
- We propose FedEAS, a strategy that assigns each client an entropy-adaptive per-class generation budget calculated based on its local label distribution.
- EN Highlights:
- arXiv:2607.06616v1 Announce Type: new
- Abstract: Label skew in federated learning (FL) causes client drift and degrades global accuracy
- Synthetic data augmentation can reduce this imbalance; however, full class balancing requires substantial computation cost
We propose FedEAS, a policy that assigns each client an entropy-adaptive per-class generation budget computed from its local label distribution
Inertia-1: An Open Exploration of Wearable Motion Foundation Models
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06617v1 Announcement Type: New.
- Abstract: Wearable motion sensing provides a continuous and scalable window into human behavior and health, making it highly suitable for foundation models, yet its pre-training and scaling principles remain largely unknown.
- Prior work has studied isolated design choices, such as sensor placement or sampling frequency, typically under fixed settings and for narrow downstream tasks, failing to capture real-world sensing diversity.
- We introduce Inertia-1, a fully open exploration of wearable motion foundation models.
- EN Highlights:
- arXiv:2607.06617v1 Announce Type: new
- Abstract: Wearable motion sensing provides a continuous and scalable window into human behavior and health, making it a natural fit for foundation models, yet i…
- Prior work studies isolated design choices, such as sensor placement or sampling frequency, often under fixed settings and narrow downstream tasks that fail to…
- We introduce Inertia-1, a fully open exploration of wearable motion foundation models
Fingerprint, Not Blueprint: How Positional Schemes Set the Default Spectral Algebra of Attention
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06621v1 Announcement Type: New.
- Abstract: The pre-softmax score of an attention head is a bilinear form $score(i,j) = x_i^T M x_j$ in a learned operator $M = W_q^T W_k$.
- Because M is generally non-symmetric, and therefore non-normal, it has a complex eigenspectrum and non-orthogonal eigenvectors—the regime where non-Hermitian and random matrix tools apply.
- We ask what this spectrum encodes at three levels for previous-token and induction circuits.
- EN Highlights:
- arXiv:2607.06621v1 Announce Type: new
- Abstract: The pre-softmax score of an attention head is a bilinear form $score(i,j) = x_i^T M x_j$ in a learned operator $M = W_q^T W_k$
- Because M is generally non-symmetric, hence non-normal, it has a complex eigenspectrum and non-orthogonal eigenvectors, the regime where non-Hermitian and rando…
- We ask what this spectrum encodes, at three levels for previous-token and induction circuits
LLM-Guided Task-Semantic Field Factorization for Industrial Process Forecasting
- Publication Time: 2026-07-09 12:00 Beijing Time
- Abstract: - arXiv:2607.06623v1 Announcement Type: New.
Abstract: Process industries rely on time-series forecasting and soft sensing to estimate quality variables that are hard to measure online.
Labeled data are scarce, operating regimes change frequently, and retraining models or rebuilding alignment pipelines for each scenario is costly.
Such settings often provide variable tables and process documents that record variable names, units, physical meanings, and process roles.
EN Key Points:
- arXiv:2607.06623v1 Announce Type: new
- Abstract: Process industries rely on time-series forecasting and soft sensing to estimate quality variables that are hard to measure online
- Labeled data are scarce, operating regimes change frequently, and retraining models or rebuilding alignment pipelines for each scenario is costly
- Such settings often provide variable tables and process documents that record variable names, units, physical meanings, and process roles