🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-07-11
- 类型
- ai-daily
- 字数
- 7560
- 阅读时长
- 36 min
2026-07-11 AI Daily Update | Deutsche Telekom Samples Emerge: AI Agents Enter Operational Ledgers Link to heading
Today’s focus isn’t just on new model releases, but on how AI Agents are truly entering enterprise operations. Deutsche Telekom showcases an AI-native transformation from customer service and networks to decision-making processes; OpenAI’s ChatGPT Work pushes Agents to the office entry point. Meanwhile, the industry is increasingly focusing on trajectory evaluation, orchestration efficiency, and token costs, as enterprise adoption enters the accounting phase.
📖 In-Depth Guide to This Issue’s Watch List Link to heading
Today’s most important read is about the implementation of enterprise AI agents: Deutsche Telekom’s “AI-native telecommunications company” case, advancing generative AI from customer service efficiency to network, decision-making, and customer journey re-engineering. At the same time, “The Harness Effect” reminds us that enterprise Agent costs depend not only on model price, but also on how the orchestration layer organizes context, tools, and turns, which determines long-term token economics.
The second main thread is that Agent evaluation is shifting from “task completion” to “process quality.” AgentLens focuses on the complete trajectory of code agents, and papers like ARC-AGI and SageMath enhanced mathematical agents are also discussing: how reflection, tool calling, and verification feedback can truly improve generalization within a limited budget.
Furthermore, application research on long-tail fairness in healthcare, large behavioral models in retail, and temporal graph interpretability are worth a quick read, showing that AI evaluation is further aligning with real deployment risks.
🌐 X Platform AI Hot News Briefs Link to heading
Topic 1: Anthropic Resets Claude Rate Limits After Rival AI Launches Link to heading
- Category: AI · News
- Overview: Hot for: 1 day ago, Related Posts: 21000
- What happened: Anthropic adjusted and reset Claude’s usage rate limits after a competitor launched new AI products.
- Why it’s important: This reflects intensifying competition among leading AI companies in model releases, user growth, and compute resource allocation, and also highlights the significant impact of throttling strategies on user experience and product adoption.
- Discussion overview: Discussions on X focus on whether Anthropic relaxed limits due to competitive pressure, whether Claude users will get a more stable experience, and how AI platforms balance performance, price, and availability.
Topic 2: 1X Unveils Most Advanced Robotic Hands for NEO Humanoid Link to heading
- Category: AI · News
- Overview: Hot for: 2 days ago, Related Posts: 37000
- What happened: Robotics company 1X released a new generation of highly dexterous robotic hands for the NEO humanoid robot, demonstrating its capabilities in grasping, manipulation, and human-like hand movements.
- Why it’s important: Dexterous hands are critical components for humanoid robots to move from demonstrations to real home and work scenarios, directly affecting whether robots can complete complex, unstructured daily tasks, and also reflecting the progress of embodied AI and hardware collaborative development.
- Discussion overview: Discussions on X focus on whether its hand’s degrees of freedom, load capacity, durability, and cost are sufficient for commercialization; supporters believe this is an important step for household humanoid robots, while skeptics argue that the demonstrations are still far from stable mass production and generalization in real-world scenarios.
Topic 3: GPT-5.6 Sol Challenges Claude Fable 5 in AI Coding Debate Link to heading
- Category: AI · News
- Overview: Hot for: 1 day ago, Related Posts: 7500
- What happened: A popular discussion emerged on X around “whether GPT-5.6 Sol challenges Claude Fable 5 in programming capability,” and was widely circulated by AI weekly accounts.
- Why it’s important: AI programming capability has become one of the core indicators of large model competition, and relevant comparisons will influence developer tool selection, enterprise procurement, and judgments on the frontier capabilities of advanced models.
- Discussion overview: The discussion centers on the code generation, debugging, long-context understanding, and practical engineering usability of both types of models; the divergence lies in some users believing GPT-5.6 Sol has greater generality and speed advantages, while others emphasize that Claude Fable 5 performs better in complex code understanding and stability. Some also question the lack of public, reproducible benchmark tests for such comparisons.
Topic 4: OpenAI Rolls Out GPT-Live Voice and GPT-5.6 Models for ChatGPT Link to heading
- Category: AI · News
- Overview: Hot for: 2 days ago, Related Posts: 120000
- What happened: X is abuzz with reports that OpenAI is rolling out the real-time audible GPT-Live voice model to ChatGPT, and the GPT-5.6 series is entering public release after restrictions were lifted.
- Why it’s important: If true, this means AI assistants are moving from text-based conversations towards low-latency, two-way voice interaction, and may enhance reasoning, agency, and multimodal application capabilities through new-generation models.
- Discussion Summary: The discussion focuses on whether the real-time conversation experience of GPT-Live is natural enough, the extent of GPT-5.6’s capability improvement and its scope of availability, and the impact of lifting regulatory restrictions on the release cadence and competitive landscape of frontier models. Some also questioned the authenticity and practical availability of the related news.
Topic 5: SpaceXAI Launches Grok 4.5, Frontier AI Model with Top Efficiency Link to heading
- Category: AI · News
- Overview: Trending Time: 2 days ago, Related Posts: 239,000
- What it is: SpaceXAI released its new-generation frontier AI model, Grok 4.5, featuring higher reasoning capabilities and efficiency.
- Why it matters: If its efficiency and performance metrics are true, Grok 4.5 could intensify competition among frontier large models in terms of computing cost, inference speed, and commercial deployment.
- Discussion Summary: Discussions on X mainly focus on whether Grok 4.5 truly achieves “top-tier efficiency,” its performance compared to models from OpenAI, Google, Anthropic, etc., and whether the related benchmarks are transparent and credible.
Summary of AI Public Opinion on X Today Link to heading
The main theme of today’s public opinion is the significant heating up of competition in frontier AI: from Claude adjusting its rate limits, rumors of OpenAI’s new voice features and GPT-5.6 release, to Grok 4.5 and comparisons of various coding models, market attention is focused on capability improvements, availability, cost-efficiency, and release cadence. The broad consensus is that AI products are no longer just competing on “model scores,” but also on real user experience, including whether rate limits are stable, voice interactions are natural, coding abilities are applicable in engineering scenarios, and whether robot hardware can support practical tasks. The main point of disagreement lies in the credibility of the claimed leadership in various model or hardware demonstrations: supporters believe the new models and dexterous hands represent key progress toward general-purpose assistants and home robots, while skeptics emphasize the lack of publicly reproducible benchmarks, insufficient generalization in real-world scenarios, and the possibility that marketing narratives may outweigh actual capabilities. Potential risks lie in the fact that unverified release announcements and opaque evaluations can easily inflate market expectations, rate limits and computing constraints may affect user trust, and advancements in embodied intelligence and real-time voice assistants will also bring pressures related to security, privacy, regulation, and commercialization.
💡 Influencer Insights Link to heading
AI Industry Daily Insights Report (July 10) Link to heading
1. Today’s Core Focus: The Release of GPT-5.6 Sol and the Wave of Agentification Link to heading
OpenAI’s Super-App Strategy Takes Shape Link to heading
The most significant industry event today is OpenAI’s official release of GPT-5.6, along with the simultaneous launch of ChatGPT Work.
- Product Matrix Reorganization: OpenAI has integrated ChatGPT, Codex, and Work into a single desktop application (@dotey). It is shifting from the original “conversational AI” to an all-in-one super-app covering Chat, Work (productivity Agent), and Codex (programming Agent).
- GPT-5.6 Model Tiers: The new model is divided into Sol (flagship/complex reasoning), Terra (cost-effective/daily), and Luna (lightweight/high-speed). The Sol model adds an Ultra mode that can call multiple sub-agents to process complex tasks in parallel (@dotey).
- Competitive Response: Interestingly, right after GPT-5.6 was released, the Claude team immediately reset user quotas (@Pluvio9yte citing @bourneliu66). Furthermore, paid access to Claude Fable 5 has been extended to July 12 (@zhixianio), making the competition extremely intense.
Agent Workflows Move from Code to Office Link to heading
- ChatGPT Work Defines a New Paradigm: No longer just for simple Q&A or code generation, Work can connect to business applications like Gmail, Slack, and Google Drive, autonomously break down complex projects, and has capabilities for long-term execution, scheduled tasks, and Computer Use. This marks the official transition of AI Agents from developer tools to daily enterprise office scenarios (@dotey).
- Local Agent Desktop Showdown: ChatGPT Work and Claude Cowork are now in direct competition on the desktop. Both have computer control capabilities, but they differ in their technical approaches to isolation mechanisms (Seatbelt vs. virtual machine) and data synchronization strategies (@dotey).
2. Unique Perspectives and Industry Foresight Link to heading
2.1 The Developer’s Shifting Identity and Institutional Thinking Link to heading
- From Programmer to “Engineering Manager”: @dotey accurately points out that in the context of high token consumption with Vibe Coding, directing an AI Agent to work is more like being an Engineering Manager (EM). The core tasks become requirement decomposition, task allocation, and results acceptance (Review). AI can replace programmers, but human value lies in controlling the overall architecture and security boundaries, reviewing AI code through a “continuous integration” approach of small, rapid iterations (@dotey).
- The Disappearance of Code Moats: @ruanyf recounted the case of a Cloudflare engineer replicating Next.js for $1100, raising the point that in the AI era, code moats have vanished. The key to preventing large software from being rapidly replicated lies in test cases, which are the new barrier.
2.2 Model Capability Involution and Hallucination Link to heading
- The “Whoever’s Guilty Resets” Race: @Pluvio9yte mocked the current industry trend where models, once at a disadvantage, try every means to reset quotas to retain users, implying that the current competition has shifted from mere benchmark scores to a battle for traffic and stickiness.
- The Ceiling of Small Model Limitations: @zhixianio thoroughly reviewed Gemma 4 12B Coder, concluding that fine-tuning can accelerate convergence but cannot raise the ceiling for small-scale models in complex tasks that are “long-form, stateful, and one-shot.” This proves that in some scenarios, large-parameter MoE is still the sweet spot.
- Token Traps and Cost Black Holes: @Pluvio9yte provided real feedback that GPT-5.6 Sol consumes several times more tokens than Fable 5, with Pro 20 accounts hitting their window limit in tens of minutes. @ruanyf also estimated that if top-tier models were made unlimited, heavy users could face annual costs in the tens of millions. The stronger the model, the more expensive it is; if not restrained, “unlimited Token could become an unlimited bill.”
2.3 General Ecosystem Trends Link to heading
- Heavyweight Competition in the Chinese Market: @ruanyf tested Tencent Hunyuan Hy3, noting that despite its smaller parameters, it approached GLM 5.1 levels; @dotey broke the news that DeepSeek V4 is launching soon and plans to increase prices, while MiniMax M3 Pro will become China’s largest open-source model.
- The Double-Edged Sword of GPT-Live Voice: OpenAI’s release of full-duplex voice capabilities makes human-machine interaction more natural, but in actual experience, the AI’s “American accent” and overly intrusive response words (e.g., mhmm) are distracting (@dotey), indicating that the restraint of emotional interaction is a current challenge for voice assistant deployment.
3. Emerging Tools, Resources, and Inspiration Link to heading
3.1 Hardcore Development and Productivity Link to heading
- Vercel Native SDK: A newly released desktop application framework, using the Zig language, with its own declarative UI markup language (.native) and self-rendering engine, attempting to solve the problems of electronic software size and memory footprint, and natively considering support for AI Agent automation (@dotey).
- Qiaomu Design Skill & IBM Carbon: @vista8 recommended IBM’s Carbon design system and compiled it into an AI-absorbable Skill to standardize the level of AI-generated interactive design, offering a new approach to “constraining AI hallucination with design systems.”
- Obsidian to X Long-Form Post Plugin: The tool recommended by @AI_Jasonyu solves the pain point of converting long-form posts from local Markdown editing to the X platform (developed by @kaitoxhacker).
3.2 Open-Source Information Acquisition Suite Link to heading
- Hackernews Speed Reading Site: @vista8 open-sourced an HN news site based on AI translation and summarization, using AI to extract key comments from popular English posts, lowering the barrier to accessing tech news.
- Qiaomu RSS Reader: An integrated reader of 35+ high-quality overseas Newsletters, supporting AI translation, rewriting, and sidebar conversations.
3.3 Creativity and Design Link to heading
- Topview 3D Shot Composer: An AI video new paradigm recommended by @AI_Jasonyu, allowing direct placement of characters and camera positions in 3D space before generating video, solving the problem of difficulty in precisely controlling composition in text-to-video generation, pushing AI tools from “gacha” to “director.”
📚 Appendix: Today’s Watch List Update Source List Link to heading
Time Window: Most recent 3 days; Covering 22 sources; Total 32 updates
All-In Podcast (A_full) Link to heading
- Open Source Wins, AGI Is Here, and Scorsese’s AI Toolkit with CEOs of Cerebras & Black Forest Labs
- Published Time: 2026-07-10 09:26 Beijing Time
- Summary: - AppLovin Ads - AppLovin’s AI advertising platform covers over 1 billion daily active users in the mobile gaming sector.
- Full-screen video ads with a median watch time of 35 seconds.
- Advertisers spend hundreds of thousands of dollars daily for profit, and advertiser access is still in closed beta.
- Nasdaq - Industry, capital, and intelligence are merging into a single, interconnected system, and the infrastructure behind it needs to evolve just as quickly.
- Nasdaq was built for this moment: powering over 135 markets and regulators globally and connecting capital with companies shaping the future.
- EN Key Points:
- (0:00) The AI Buildout: Datacenters Bigger Than Cities (Andrew Feldman)
- (1:50) Reasoning, Inference, and Breaking Moore’s Law
- (16:28) Open Source, AI Sovereignty, and the Road to AGI
- (40:54) The Innovation Behind Generative Video (Robin Rombach)
Stratechery by Ben Thompson (A_full) Link to heading
- 2026.28: XBOX On the Rocks
- Publication Time: 2026-07-11 01:00 Beijing Time
- Summary: - (Photo by Jeff Christensen/Liaison).
- Welcome back to This Week in Stratechery!
- As a reminder, each week, every Friday, we send out an overview of the content in the Stratechery bundle; highlighted links are free for everyone.
- Additionally, you have complete control over what we send to you.
- With that in mind, here are some of our favorites from this week.
- EN Key Points:
- (Photo by Jeff Christensen/Liaison)
- Welcome back to This Week in Stratechery
- As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone
- Additionally, you have complete control over what we send to you
OpenAI Blog (A_full) Link to heading
- How Deutsche Telekom is rewiring telecommunications with AI
- Publication Time: 2026-07-10 15:00 Beijing Time
- Summary: - Operating at such a scale means managing a huge customer service operation, a complex network infrastructure, and the millions of daily interactions that keep people connected.
- With the accelerating capabilities of generative AI, Deutsche Telekom saw an opportunity that went beyond productivity gains.
- The company set an ambitious goal: to become the world’s first AI-native telecommunications company.
- Instead of viewing AI as just another software rollout, the leadership team saw it as a fundamental shift in how decisions are made, how customer journeys are designed, and how telecommunication services are delivered.
- We sat down with Jonathan Abrahamson, Chief Product and Digital Officer at Deutsche Telekom, to discuss how the company is redesigning its operating model across the entire organization - from customer service and employee workflows to network operations and the future of voice communications.
- EN Key Points:
- How Deutsche Telekom is becoming an AI-native telco with OpenAI-transforming customer service, employee workflows, network operations, and the future of voice.
ArXiv cs.AI (B_intro+search) Link to heading
AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
- Publication Time: 2026-07-10 12:00 Beijing Time
- Summary: - arXiv:2607.06624v1 Announce Type: new.
- Summary: We introduce AgentLens, a production evaluation benchmark for interactive code agents.
- Most code agent benchmarks reduce a run to a single bit - did the task pass?
- But people who actually use these agents experience the whole trajectory: how the agent follows instructions, uses its tools, verifies its own work, recovers from errors, and talks to them along the way.
- EN Key Points:
- arXiv:2607.06624v1 Announce Type: new
Abstract: We present AgentLens, a production-assessed benchmark for interactive code agents
Most code-agent benchmarks reduce a run to a single bit – did the task pass
– but the people who actually use these agents experience the entire trajectory: how the agent follows instructions, uses its tools, verifies its own work, rec…
When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning
- Published: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.06720v1 Announce Type: new.
- Abstract: Training large language models (LLMs) with extended reasoning has enabled in-context search, in which models iteratively generate, critique, and revise solution attempts.
- We provide a theoretical analysis of in-context search by modeling it as approximate inference over reasoning traces, where the base model defines a prior, self-reflection provides feedback for posterior updates, and we study the resulting reasoning-time sampling complexity—the number of sequential attempts required to achieve a high success probability.
- We show that when reflections reliably localize early mistakes, in-context search can yield exponential improvements over the base model, solving problems with exponentially small zero-shot pass rates using only a polynomial number of sequential attempts, whereas when this property fails, conditioning on past attempts offers no asymptotic advantage over parallel sampling.
- EN Highlights:
- arXiv:2607.06720v1 Announce Type: new
- Abstract: Training large language models (LLMs) with extended reasoning has enabled in-context search, in which models iteratively generate, critique, and revis…
- We provide a theoretical analysis of in-context search by modeling it as approximate inference over reasoning traces, where the base model defines a prior and s…
- We show that when reflections reliably localize early mistakes, in-context search can yield exponential improvements over the base model, solving problems with…
LLM-powered reasoning in agent-based modeling
- Published: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.06757v1 Announce Type: new.
- Abstract: Agent-based modeling (ABM) can model millions of individuals and their interactions, which is useful for policymaking.
- However, ABMs have traditionally relied on static priors, which prevent the models from adapting to real-time changes.
- Our research offers a new approach to address this information gap.
- EN Highlights:
- arXiv:2607.06757v1 Announce Type: new
- Abstract: Agent-based modeling (ABM) has the capability to model millions of individuals and their interactions, which is useful for policy making
- However, ABMs have traditionally relied on static prior, which prevents the models from adapting to real-time changes
Our research provides a novel approach to addressing this information gap
QANTIS: Hardware-Calibrated Sequential POMDP Belief Updates on IBM Heron
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.06760v1 Announcement Type: New.
- Abstract: Autonomous systems under partial observability act on beliefs, not raw sensor events.
- QANTIS treats the quantum processor as a calibrated belief-update service in that loop: it receives a prior and an observation model, estimates the rare-event evidence term, and returns the ordinary posterior to the classical planner.
- This paper asks whether that service can be reused across a sequential Tiger POMDP horizon on present IBM Heron hardware without corrupting the planner-facing posterior.
- EN Highlights:
- arXiv:2607.06760v1 Announce Type: new
- Abstract: Autonomous systems under partial observability act on beliefs, not raw sensor events
- QANTIS treats the quantum processor as a calibrated belief-update service in that loop: it receives a prior and an observation model, estimates the rare-event e…
- This paper asks whether that service can be reused across a sequential Tiger POMDP horizon on present IBM Heron hardware without corrupting the planner-facing p…
Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.06764v1 Announcement Type: New.
- Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or training for specific benchmarks where small models are fine-tuned on ARC data, often with task-specific architectures.
- We study a third regime: an open-weight model in non-thinking mode (DeepSeek V3.2) under a strict budget, with no ARC-specific fine-tuning.
- We study what is recoverable through architecture alone, building agentic harnesses that explicitly decompose the pattern-discovery and program-synthesis stages.
- EN Highlights:
- arXiv:2607.06764v1 Announce Type: new
- Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionar…
- We study a third regime: an open-weight model in non-thinking mode (DeepSeek V3.2) under a strict budget, with no ARC-specific fine-tuning
- We study what is recoverable through architecture alone, building agentic harnesses that decompose pattern-discovery and program-synthesis stages explicitly
Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics
Published: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.06820v1 Announce Type: new.
- Abstract: Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS) in agentic LLM workflows underexplored.
- We propose a ReAct-style agentic setup that combines LLM reasoning with verifiable feedback from SageMath, together with Context7 for the up-to-date documentation.
- We evaluate this agentic setup across frontier models for solving research-level mathematical problems from the RealMath benchmark in a setting that emulates a computational mathematics research loop.
- EN Key Points:
- arXiv:2607.06820v1 Announce Type: new
- Abstract: Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS…
- We propose a ReAct-style agentic setup that combines LLM reasoning with verifiable feedback from SageMath, together with Context7 for the up-to-date documentati…
- We evaluate this agentic setup across frontier models for solving research-level mathematical problems from the RealMath benchmark in a setting that emulates a…
- Abstract: - arXiv:2607.06820v1 Announce Type: new.
The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
- Published: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.06906v1 Announce Type: new.
- Abstract: Agentic AI development today runs on token maxing: buying capability with tokens – longer reasoning traces, more turns, wider tool payloads, bigger replay contexts – so that per-task tokens grow faster than task value.
- Falling per-token prices mask the pattern; total spend rises anyway.
- We argue the decisive lever against token maxing is the harness: the orchestration layer that assembles context, exposes tools, sequences turns, delegates work and hosts enterprise observability and governance.
- EN Key Points:
- arXiv:2607.06906v1 Announce Type: new
- Abstract: Agentic AI development today runs on token maxing: buying capability with tokens – longer reasoning traces, more turns, wider tool payloads, bigger r…
- Falling per-token prices mask the pattern; total spend rises anyway
- We argue the decisive lever against token maxing is the harness: the orchestration layer that assembles context, exposes tools, sequences turns, delegates work,…
- Published: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.06925v1 Announce Type: new.
- Abstract: Compact world models conditioned on language goals promise to ground relations like “put the red block to the left of the blue block” using a sparse set of explicit \emph{reference anchors}.
We ask when such references actually ground a relation, and identify a trap: a goal-conditioned predictor reaches a striking $0.90 relation-readout accuracy, but this is merely \emph{instruction transcription}, not perception.
- Withholding the goal collapses it ($0.90!\to!0.27$, three seeds), and counterfactual instructions cause the predicted anchor to follow the \emph{false} instruction $94.5%$ of the time (true scene $2.3%$; $N{=}256$).
- EN Key Points:
- arXiv:2607.06925v1 Announce Type: new
- Abstract: Compact world models that condition on a language goal promise to ground relations such as ``put the red block left of the blue block’’ using a sparse…
- We ask when such references actually ground a relation, and identify a trap: a goal-conditioned predictor reaches a striking $0.90$ relation-readout accuracy, y…
- Withholding the goal collapses it to chance ($0.90\
Large Behavior Model: A Promptable Digital Twin of the Retail Customer
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract:
- arXiv:2607.06993v1 Announcement Type: New.
- Customer behavior modeling is foundational to recommendation, marketing, and decision support, but existing methods either optimize for predictive accuracy without explaining decisions or simulate users without grounding in real behavioral data.
- We present the Large Behavioral Model (LBM), which learns customer decision-making directly from large-scale retail transactions through a unified person-environment formulation.
- Customer state is represented by a behavioral profile derived from historical purchases, while product context is incorporated through retrieval-augmented generation.
- EN Key Points:
- arXiv:2607.06993v1 Announce Type: new
- Abstract: Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive accuracy with…
- We present the Large Behavioral Model (LBM) that learns customer decision making directly from large-scale retail transactions through a unified Person-Environm…
- Customer state is represented by a behavioral profile derived from historical purchases, while product context is incorporated through retrieval-augmented gener…
ArXiv cs.CL (B_intro+search) Link to heading
Unveiling Public Opinion: A Study of Sentiment Analysis Using LSTM and Traditional Models
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract:
- arXiv:2607.07772v1 Announcement Type: New.
- In this era of social media, sites like Twitter have become gathering places for people to share their opinions and feelings on various issues and current events in real-time.
- Sentiment analysis, a key application of NLP, has become indispensable due to the massive influx of user-generated content, enabling the extraction of meaningful insights from the opinions and emotions expressed in text data.
- Sentiment analysis on Twitter employs complex computational techniques to classify tweets into positive, negative, or neutral sentiments.
- EN Key Points:
arXiv:2607.07772v1 Announce Type: new
- Abstract: In this age of social media, sites like Twitter have become meeting places for people to share their views and feelings on a wide range of issues and…
- Sentiment analysis, a critical application of NLP, has become indispensable due to the massive influx of user-generated content, enabling the extraction of mean…
- Sentiment analysis on Twitter employs sophisticated computational techniques to categorize tweets into positive, negative, or neutral sentiments
From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract:- arXiv:2607.07779v1 Announce Type: new.
- Abstract: Recent developments in AI for Mathematics (AI4Math), especially Large Language Model (LLM)-driven theorem provers, have achieved remarkable success in generating formal proofs for well-defined mathematical problems using interactive theorem proving (ITP) languages.
- However, current systems remain fundamentally limited in tackling frontier research mathematics, such as discovering new theorems or resolving open conjectures, which are often open-ended, ill-defined, and involve multiple layers of abstraction.
- We argue that the next leap in AI4Math systems requires a decisive shift from predefined problem-solvers to research agents that can address frontier mathematical challenges through rigorous formal mathematical reasoning.
- EN Key Points:
- arXiv:2607.07779v1 Announce Type: new
- Abstract: Recent developments in AI for Mathematics (AI4Math), especially Large Language Model (LLM)-driven theorem provers, has achieved remarkable success in…
- However, current systems remain fundamentally limited in tackling frontier research mathematics, such as discovering new theorems or resolving open conjectures,…
- We argue that the next leap in AI4Math systems requires a decisive shift from predefined problem-solvers to research agents that can address frontier mathematic…
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract:- arXiv:2607.07820v1 Announce Type: new.
- Abstract: Training tool-using agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher distillation trajectories, and sparse reward reinforcement learning provides weak supervision for long-horizon interactions.
- We introduce DeepSearch-Evolve, a web agent self-distillation framework built on DeepSearch-World, a deterministic and verifiable environment with reproducible search and page-reading tools.
- DeepSearch-World incorporates 420K multi-hop QA tasks constructed from entity-level random walks and supports key agent cognitive behaviors useful for self-evolution, including progress verification, grounded reflection, and failure recovery.
EN Key Points:
- arXiv:2607.07820v1 Announce Type: new
- Abstract: Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher-distilled traject…
- We present DeepSearch-Evolve, a self-distillation framework for web agents built on DeepSearch-World, a deterministic and verifiable environment with reproducib…
- DeepSearch-World contains 420K multi-hop QA tasks constructed from entity-level random walks and supports key agentic cognitive behaviors useful for self-evolvi…
- Published: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.07891v1 Announce Type: new.
- Abstract: Roy Harris’s Integrationist linguistics offers a compelling critique of the referentialist tradition embedded deep at the heart of computational approaches to language, arguing that language is not a code that maps onto a pre-given world, but a context-situated, two-sided activity geared toward prospective joint action.
- Yet Integrationism leaves certain explanatory gaps: it does not fully account for the structural mechanism by which signs sustain prospective openness, it undertheorizes the continuity between linguistic and non-linguistic semiotic activities, and it does not provide a detailed account of the structural properties of the archive of past integrations built up over time.
- This paper argues that Elan Barenholtz’s autogenerative theory of language, developed in response to the behaviour of Large Language Models (LLMs), can fill precisely these gaps, enriching Integrationism without compromising any of its core commitments.
- EN Key Points:
- arXiv:2607.07891v1 Announce Type: new
- Abstract: Roy Harris’s Integrationist linguistics offers a compelling critique of the referentialist tradition embedded deep at the heart of computational appro…
- Yet Integrationism leaves certain explanatory gaps: it does not fully account for the structural mechanism by which signs sustain prospective openness, it under…
- This paper argues that Elan Barenholtz’s autogenerative theory of language, developed in response to the behaviour of Large Language Models (LLMs), can fill pre…
Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration
- Published: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.07895v1 Announce Type: new.
- Abstract: Due to the lack of datasets in other languages and the high cost of manual annotation in underrepresented cultures, research on stereotypes in Large Language Models (LLMs) has primarily focused on English-speaking contexts.
- To address this gap, we introduce a cost-effective human-LLM collaborative annotation framework and apply it to build EspanStereo, a Spanish stereotype dataset spanning multiple Spanish-speaking countries in Europe and Latin America.
EspanStereo captures both well-documented stereotypes from prior literature and culturally specific biases absent from English-centric resources.
- EN Key Points:
- arXiv:2607.07895v1 Announce Type: new
- Abstract: Research on stereotypes in large language models (LLMs) has largely focused on English-speaking contexts, due to the lack of datasets in other languag…
- To address this gap, we introduce a cost-efficient human-LLM collaborative annotation framework and apply it to construct EspanStereo, a Spanish-language stereo…
- EspanStereo captures both well-documented stereotypes from prior literature and culturally specific biases absent from English-centric resources
- EN Key Points:
When Debiasing Backfires: Counterintuitive Side Effects of Preprocessing-Based Stereotype Mitigation
- Release Time: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.07937v1 Announce Type: new.
- Abstract: Preprocessing-based stereotype mitigation methods, such as pre-/post-training on debiased corpora, are widely used in NLP.
- While these methods reduce measurable stereotypes for targeted groups, we find they often induce unintended shift side-effects, where stereotyping or counter-stereotyping may increase relative to a neutral baseline for other demographics, including across unrelated demographic categories.
- We demonstrate these side effects across two model families (encoder-only and decoder-only), multiple preprocessing strategies (removing stereotypical sentences, removing group mentions, and swapping group references), and in pre-training and post-training on Wikipedia with different data scales.
- EN Key Points:
- arXiv:2607.07937v1 Announce Type: new
- Abstract: Preprocessing-based methods for stereotype mitigation, such as pre-/post-training on debiased corpora, are widely used in NLP
- While these approaches reduce measurable stereotypes for targeted groups, we find they often induce unintended shifts-side effects, where stereotyping or counte…
- We demonstrate these side effects across two model families (encoder-only and decoder-only), multiple preprocessing strategies (removing stereotypical sentences…
A Multi-cluster Boundary Learning Method for Out-of-Scope Intent Detection via MiniLM Embedding
- Release Time: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.07974v1 Announce Type: new.
- Abstract: Intent detection is a critical task in human-computer interaction systems that connects human intent with system operations.
- However, detecting out-of-scope (OOS) intents remains a challenge.
- (i) Traditional methods treat OOS intent detection as multi-class classification, where the detection accuracy then decreases as the number of known intent classes increases; (ii) LLM embedding methods require a large number of parameters, which makes them difficult to train and deploy in practice.
- EN Key Points:
- arXiv:2607.07974v1 Announce Type: new
Abstract: Intent detection is a critical task that bridges human intents and system actions in human-machine interaction systems
- However, there still exist challenges for detecting out-of-scope (OOS) intents
- (i) The traditional methods view the OOS intent detection as a multi-class classification, then the detection accuracy decreases as the class number of the know…
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract:- arXiv:2607.07976v1 Announce Type: new.
- Abstract: Reinforcement learning (RL) has achieved significant success in enhancing the reasoning capabilities of large language models (LLMs).
- However, widely used critic-free reinforcement learning methods rely on uniform credit assignment, broadcasting the same advantage to all tokens regardless of their differences.
- We identify a critical failure mode of this design, which we refer to as positive credit contamination: low-probability tail tokens that are contextually erroneous receive the same positive credit within the same trajectory as plausible ones, leading to indiscriminate reinforcement of flawed reasoning behaviors.
- EN 要点:
- arXiv:2607.07976v1 Announce Type: new
- Abstract: Reinforcement learning (RL) has achieved remarkable success in enhancing the reasoning capabilities of large language models (LLMs)
- However, widely used critic-free RL methods rely on uniform credit assignment, broadcasting the same advantage to all tokens regardless of their differences
- We identify a critical failure mode of this design, which we refer to as Positive-Credit Contamination: low-probability tail tokens that are contextually errone…
A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract:- arXiv:2607.07985v1 Announce Type: new.
- Abstract: We report the empirical reliability of Gemini models as audio judges, scoring full-duplex agent conversations directly from raw stereo waveforms, tested across three models in the Gemini family: 2.5 Flash, 3.5 Flash, and 3.1 Pro.
- Our primary evidence base used Gemini 2.5 Flash as the ground truth model, validated against three calibrated human evaluators across 209 stereo sessions, scoring on 8 production dimensions: 152 full-duplex conversations spanning 13 accent and condition layers, and 57 adversarial defect injection clips.
- Evidence from Gemini 2.5 Flash was consistent across three tests.
- EN 要点:
- arXiv:2607.07985v1 Announce Type: new
- Abstract: We report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stereo waveform,…
Our primary evidence base uses Gemini 2.5 Flash as the ground-truth model, validated against three calibrated human raters on 209 stereo sessions, scored on 8 p…
The evidence for Gemini 2.5 Flash is consistent across three tests
Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.07993v1 Announcement Type: new.
- Abstract: Identifying faithfulness hallucinations in LLM-generated outputs remains challenging due to the scarcity of high-quality annotated data.
- Recent work relies on advanced LLMs to synthesize training data, including rationales, labels, and hallucinated claims.
- However, these methods treat the generator as a static component, limiting the iterative improvement of the detector.
- EN Key Points:
- arXiv:2607.07993v1 Announce Type: new
- Abstract: Identifying faithfulness hallucinations in LLM-generated outputs remains challenging due to the scarcity of high-quality annotated data
- Recent work relies on advanced LLMs to synthesize training data, including rationales, labels, and hallucinated claims
- However, these methods treat the generator as a static component, limiting iterative improvement of the detector
ArXiv cs.LG (B_intro+search) Link to heading
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.07716v1 Announcement Type: new.
- Abstract: Temporal graphs are ubiquitous in real-world applications, and Temporal Graph Networks (TGNs) have achieved superior predictive accuracy.
- Understanding which historical events drive model predictions can enhance the trustworthiness of TGNs.
- Existing explanation methods overlook the memory module, the core component that records and updates node histories, without exploring the influence of past events.
- EN Key Points:
- arXiv:2607.07716v1 Announce Type: new
- Abstract: Temporal graphs are ubiquitous in real-world applications and Temporal Graph Networks (TGNs) have achieved superior predictive accuracy
- Understanding which historical events drive model predictions can enhance trustworthiness of TGNs
- Existing explanation methods overlook the memory module, the core component that records and updates node histories, leaving the influence of past events unexpl…
Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.07717v1 Announcement Type: new.
- Abstract: In chest X-ray (CXR) classification, acceptable ranking performance can still place rare positive patients below the threshold, especially within subgroups.
- We study this pre-deployment fairness issue as an audit question: after a long-tail multi-label CXR model converts scores into decisions, who is missed?
- In VinDr-CXR and MIMIC-CXR/CXR-LT, we use a diagnostic ladder to separate class-level long-tail losses, subgroup-aware weighting, group robustness, and threshold selection.
- EN Highlights:
- arXiv:2607.07717v1 Announce Type: new
- Abstract: In chest X-ray (CXR) classification, acceptable ranking performance can still leave rare-positive patients below threshold, especially within subgroup…
- We study this pre-deployment fairness problem as an audit question: after a long-tailed multi-label CXR model is converted from scores into decisions, who is mi…
- Across VinDr-CXR and MIMIC-CXR/CXR-LT, we use a diagnostic ladder to separate class-level long-tail losses, subgroup-aware weighting, group robustness, and thre…
- Abstract: - arXiv:2607.07717v1 Announcement Type: new.
LLT: Local Linear Transformer for PDE Operator Learning
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.07718v1 Announcement Type: new.
- Abstract: Neural operators have become a common method for learning PDE solution maps and accelerating numerical simulations.
- Transformer-based neural operators are particularly interesting because attention can learn long-range dependencies in the computational domain.
- However, standard attention has two main limitations when applied to PDEs: it scales quadratically with the number of computational nodes, and it lacks an explicit bias for local interactions.
- EN Highlights:
- arXiv:2607.07718v1 Announce Type: new
- Abstract: Neural operators have become a common approach for learning PDE solution maps and accelerating numerical simulations
- Transformer-based neural operators are of particular interest, since attention can learn long-range dependencies in the computational domain
- However, standard attention has two major limitations when applied to PDEs: it scales quadratically with the number of computational nodes, and it lacks an expl…
ReCoLoRA: Spectrum-Aware Recursive Consolidation for Continual LLM Fine-Tuning
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.07719v1 Announcement Type: new.
- Abstract: Parameter-efficient fine-tuning can inexpensively adapt large language models to a task, but across a sequence of tasks, LoRA-style methods continuously stack low-rank updates on the same frozen weights, so each new task tends to overwrite previous ones.
We present ReCoLoRA (Recursive Consolidation of Low-Rank Adapters), a spectrum-aware framework for continual fine-tuning: adapters are initialized from a random SVD of pre-trained weights, effective ranks are selected per layer via the elbow criterion, and the principal subspace is adjusted before opening up residual capacity.
Before each new task, ReCoLoRA re-decomposes the current effective weight (rather than the original one) into a frozen residual, slowly updated principal components, and new adapters (recursive consolidation), so each task begins with a model that has already absorbed its predecessors.
EN Highlights:
- arXiv:2607.07719v1 Announce Type: new
- Abstract: Parameter-efficient fine-tuning adapts a large language model to one task cheaply, but across a task sequence LoRA-style methods keep stacking low-ran…
- We present ReCoLoRA (Recursive Consolidation of Low-Rank Adapters), a spectrum-aware framework for continual fine-tuning: adapters are initialized from a random…
- Before each new task, ReCoLoRA re-decomposes the current effective weight, rather than the original one, into a frozen residual, a slowly updated principal comp…
Omni-Sleep: A Sleep Foundation Model via Hierarchical Contrastive Learning of CNS–ANS Dynamic
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract:- arXiv:2607.07720v1 Announce Type: new.
- Abstract: Sleep physiology arises from the coordinated dynamics of the central nervous system (CNS) and autonomic nervous system (ANS), as reflected by multimodal polysomnography signals such as electroencephalogram (EEG), electrooculogram (EOG), electromyogram (EMG), electrocardiogram (ECG), and respiration.
- However, existing sleep foundation models often fuse heterogeneous biosignals in a topology-agnostic manner, overlooking their physiological organization.
- We introduce Omni-Sleep, a sleep foundation model that uses the CNS/ANS partition as a physiological prior for topology-constrained representation learning.
- EN Highlights:
- arXiv:2607.07720v1 Announce Type: new
- Abstract: Sleep physiology arises from the coordinated dynamics of the central nervous system (CNS) and autonomic nervous system (ANS), as reflected by multimod…
- However, existing sleep foundation models often fuse heterogeneous biosignals in a topology-agnostic manner, overlooking their physiological organization
- We introduce Omni-Sleep, a sleep foundation model that uses the CNS/ANS partition as a physiological prior for topology-constrained representation learning
Uncertainty-gated selection for block-sparse attention
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract:- arXiv:2607.07724v1 Announce Type: new.
- Abstract: Block-sparse attention extends long-context language models by replacing O(N^2) softmax with a top-k selection per query on key blocks.
This truncation is myopic: when the k-th and (k+1)-th blocks are nearly tied in score, the selector commits without spending extra budget, and a dropped block carrying answer evidence is unrecoverable downstream.
We propose a value-of-information router that measures, for each query, how decisively the top-k cut was made, and doubles the kept set for the queries where the gap is smallest; this rule is backbone-agnostic and stacks with existing block scoring methods (e.g., Quest).
- EN Highlights:
- arXiv:2607.07724v1 Announce Type: new
- Abstract: Block-sparse attention scales long-context language models by replacing the O(N^2) softmax with a per-query top-k selection over key blocks
- This cutoff is myopic: when the k-th and (k+1)-th blocks are nearly tied in score, the selector commits without spending extra budget, and a dropped block carry…
- We propose a value-of-information router that measures, for each query, how decisively the top-k cut was made, and doubles the kept set for the queries where th…
- EN Highlights:
SHIFT: Survival Prediction from Incomplete and Heterogeneous Genomic Data
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.07725v1 Announce Type: new.
- Abstract: Genomic prediction models often fail to transfer across institutions because sequencing panels differ across sites, leading to structural feature missingness at deployment.
- Existing approaches to this challenge typically restrict analysis to genes shared across cohorts, exclude patients with incomplete profiles, or rely on test-time imputation, all of which reduce robustness and limit the use of multi-center data.
- We propose Survival prediction Handling Incomplete Features using Transformer (SHIFT), a missingness-aware survival model that can directly predict from incomplete genomic inputs without test-time imputation.
- EN Highlights:
- arXiv:2607.07725v1 Announce Type: new
- Abstract: Genomic prediction models often fail to transfer across institutions because sequencing panels differ across sites, creating structural feature missin…
- Existing approaches to this challenge typically restrict analysis to genes shared across cohorts, exclude patients with incomplete profiles, or rely on test-tim…
- We propose Survival prediction Handling Incomplete Features using Transformer (SHIFT), a missingness-aware survival model that directly predicts from incomplete…
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.07740v1 Announce Type: new.
- Abstract: Modern LLMs are increasingly deployed in long-context applications, such as retrieval-augmented generation, repository-level coding, and agentic workflows, where their accumulated reasoning and tool traces often push the input an order of magnitude beyond the pre-training window, making zero-shot context extension a primary deployment path for open-weight checkpoints.
Most existing zero-shot methods pre-fix a single scaling factor, so an aggressive factor sacrifices short-context fidelity, while a conservative one collapses in long contexts.
- We propose Jet-Long, a tuning-free zero-shot method that pairs a local RoPE-faithful window with a long-range window whose scaling factor dynamically adapts to the current sequence length, precisely recovering the base model on short inputs while cleanly extrapolating on long inputs.
- EN Highlights:
- arXiv:2607.07740v1 Announce Type: new
- Abstract: Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workfl…
- Most existing zero-shot methods fix a single rescaling factor up front, so an aggressive factor sacrifices short-context fidelity while a conservative one break…
- We propose Jet-Long, a tuning-free zero-shot method that pairs a local RoPE-faithful window with a long-range window whose rescaling factor adapts dynamically t…
Architecture Generalization with MetaNCA
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.07743v1 Announce Type: new.
- Abstract: Self-organization is an emergent property of life, driven by the collective behavior of individual components acting on local information.
- Biological neurons, through local interactions transmitted via synapses, can learn efficiently and adapt their connections throughout an organism’s lifespan.
- Motivated by these desirable properties of adaptability and local interaction, Neural Cellular Automata (NCA) models have successfully learned morphogenesis using only local update rules, demonstrating stability over multiple updates and robustness to perturbations.
- EN Highlights:
- arXiv:2607.07743v1 Announce Type: new
- Abstract: Self-organization is an emergent property of life, driven by the collective behavior of individual components acting on local information
- Biological neurons, through local interactions transmitted through synapses, are able to learn efficiently and can adapt their connections over an organism’s li…
- Motivated by these desirable properties of adaptability and local interaction, neural cellular automata (NCA) models have been successful at learning morphogene…
LiST: Lipschitz Scaling Training for Robust and Calibrated Neural Networks
- Publication Time: 2026-07-10 12:00 Beijing Time
- Abstract: - arXiv:2607.07745v1 Announce Type: new.
- Abstract: While accuracy, robustness, and calibration are all crucial for reliable neural networks, they are often studied separately; developing models that satisfy all three requirements remains a core challenge.
- Lipschitz-constrained models guarantee robustness by design, but the manual selection of the Lipschitz constant L controls the resulting accuracy-robustness trade-off, and their calibration properties remain largely under-explored.
- In this work, we highlight the theoretical and empirical connections between enforcing a Lipschitz constraint and temperature scaling, a state-of-the-art calibration method.
- EN Highlights:
arXiv:2607.07745v1 Announce Type: new
Abstract: While accuracy, robustness, and calibration are all essential for reliable neural networks, they are often studied separately; developing models that…
Lipschitz-constrained models guarantee robustness by design, yet the manual selection of the Lipschitz constraint L governs the resulting accuracy-robustness tr…
In this work, we highlight a theoretical and empirical link between the enforced Lipschitz constraint and Temperature Scaling, a state-of-the-art calibration me…