🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-06-17
- 类型
- ai-daily
- 字数
- 7284
- 阅读时长
- 35 min
2026-06-17 AI Daily | AI Agents Begin to Enter Real Workflows: From Mobile Operations to Physical Experiments Link to heading
Today’s main theme is the transition of AI Agents from demos to real-world workflows: mobile agent evaluations, end-to-end development plugins, and hosted intelligent agents are all strengthening execution capabilities. At the same time, the industry is placing a greater emphasis on reliable reasoning, testing, and contracts, focusing on how model answers can be verified. On-device models and specialized small models also continue to gain traction, indicating that efficiency, cost, and local deployment are becoming new competitive battlegrounds.
📖 In-depth Guide to This Issue’s Watch List Link to heading
The most noteworthy topic today is “AI Agents transitioning from demos to real-world workflows.” PhoneHarness redefines mobile agent evaluation, emphasizing the mixed use of GUI, CLI, and tools. Meanwhile, research on Web Agents reminds us that memory and skill modules are not free—they must be re-evaluated within the token budget.
The second main theme is “reliable reasoning and interpretability.” CoRA investigates whether confidence scores are genuinely supported by the reasoning process, while papers on Lean 4 automatic formalization and the definition of LLM explanations are all asking the same question: How can model answers be verified and trusted?
Additionally, the AI planning prototype from DeepMind and the UK government is worth reading for product and policy teams, as it shows AI entering high-friction public infrastructure scenarios like housing approvals. Context Compression, multilingual tokenizers, and Nemotron 3 Ultra can serve as technical supplements for improving model efficiency.
🌐 AI Hot Topics on X Link to heading
Topic 1: SpaceX Acquires Cursor Maker Anysphere in $60 Billion Stock Deal Link to heading
- Category: AI · News
- Summary: Trending Time: 13 hours ago, Related Posts: 78,000
- What happened: According to trending news on X, SpaceX is acquiring Anysphere, the developer of the AI coding tool Cursor, in a $60 billion stock deal.
- Why it’s important: If true, this would indicate that aerospace and hard-tech companies are accelerating the integration of AI programming capabilities, highlighting the strategic value of code generation tools in engineering R&D, automated development, and enterprise productivity.
- Discussion summary: Discussions on X are focused on the authenticity of the deal, whether the $60 billion valuation is excessive, SpaceX’s strategic intentions for acquiring an AI development tool, and whether Cursor will be deeply integrated into aerospace software and internal engineering workflows.
Topic 2: Matt Shumer Seeks Mobile Control for Claude Code AI Agent Link to heading
- Category: AI · Other
- Summary: Trending Time: , Related Posts: 33
- Abstract: Matt Shumer Seeks Mobile Control for Claude Code AI Agent:
Topic 3: OpenAI Fixes Codex Outage and Promises Rate Limit Resets Link to heading
- Category: AI · News
- Summary: Trending Time: 11 hours ago, Related Posts: 2,400
- What happened: OpenAI has fixed the Codex service outage and stated that it will reset the relevant rate limits for affected users.
- Why it’s important: Codex is a critical gateway for developers using AI programming tools. This outage highlights the crucial impact of AI infrastructure stability, availability, and quota management on the adoption of productivity tools.
- Discussion summary: Discussions on X are mainly focused on the impact of the service interruption on development workflows, whether OpenAI’s compensation measures are adequate, and concerns from heavy users about the transparency and reliability of rate limits.
Topic 4: OpenAI Launches Developers Plugin for Codex Coding App Link to heading
- Category: AI · News
- Summary: Trending Time: 23 hours ago, Related Posts: 294
- What happened: OpenAI has launched a developer plugin for its Codex coding application, which allows users to build, preview, and test iOS apps within Codex, optimizing the development process with SwiftUI Preview and hot reloading.
- Why it’s important: This shows that AI programming tools are evolving from code generation to end-to-end development environments, capable of handling writing, running, debugging, and feedback in a single workflow. This enhances the ability of AI agents to participate in real software development.
- Discussion summary: Discussions on X center on the competition between Codex and other tools like Claude Code and Cursor, and whether a plugin ecosystem will become a key moat for AI programming products. Supporters believe it streamlines the development process, while critics are concerned about stability, controllability, and practical efficiency in complex projects.
Topic 5: AI Builders Embrace Agentic Loops for Self-Reliant Tasks Link to heading
- Category: AI · News
- Summary: Trending Time: 6 hours ago, Related Posts: 105
- What happened: AI developers are increasingly adopting “agentic loops” to build AI systems that can autonomously plan, execute, check, and iterate on tasks.
- Why it matters: This marks a shift in AI applications from single Q&A sessions to more autonomous workflows, promising to enhance efficiency in complex task processing, software development, data analysis, and automated operations.
- Discussion landscape: The discussion on X centers on whether agentic loops are truly reliable. Supporters believe they can reduce manual intervention and make AI assistants more practical, while skeptics worry about error accumulation, runaway costs, security boundaries, and the immaturity of evaluation standards.
Topic 6: Anthropic’s Claude Managed Agents Speed Up AI Production Deployment Link to heading
- Category: AI · News
- Overview: Trending since: 2 hours ago, Related posts: 217
- What it is: Anthropic has launched managed agents and a toolchain for enterprises and developers, built around Claude, to accelerate the transition of AI applications from prototype to production.
- Why it matters: This indicates that the competition among large models is shifting from pure model capabilities to enterprise-grade integration, automated deployment, and developer ecosystems. This could impact the adoption speed of AI in e-commerce, engineering, and office scenarios.
- Discussion landscape: The main topics on X are whether Anthropic is catching up to or even surpassing OpenAI, the practical value of Claude’s managed agents for enterprise AI deployment, and whether ecosystem collaborations with partners like Shopify will create new development workflows. The point of disagreement is whether these tools represent a productivity leap or are being overhyped by marketing and the “vibe coding” trend.
AI Public Opinion Summary on X Today Link to heading
Today’s main narrative focuses on the rapid evolution of AI programming tools from “code assistants” to “end-to-end development and agentic execution platforms.” Whether it’s the rumor of Cursor’s high-priced acquisition by SpaceX, Codex’s launch of an iOS development plugin, or Anthropic’s release of managed agents, all signs point to the developer ecosystem and engineering workflows becoming the new battleground for large model competition. The consensus is that capabilities for code generation, preview, testing, deployment, and autonomous iteration are now considered strategic infrastructure for enterprise productivity and hard-tech R&D, making toolchain integration as important as the models themselves. The main disagreement lies in valuation versus actual utility: supporters believe these products will significantly shorten development cycles and drive AI assistants into production environments. Skeptics, however, think that acquisition rumors, plugin ecosystems, and “agentic loops” might be over-inflated by capital narratives and the “vibe coding” trend. Potential risks are concentrated in three areas: infrastructure stability and rate limits directly impact development workflows; autonomous agent execution could lead to error accumulation, runaway costs, and security boundary issues; and once a platform ecosystem becomes highly locked-in, it could diminish developers’ control and freedom to migrate their toolchains.
💡 Influencer Insights Link to heading
Analysis of AI Industry Dynamics (Mid-June 2026) Link to heading
1. Today’s Hot Topics: Local On-Device Models and the Expansion of Programming Agents into the Physical World Link to heading
The focus of today’s discussion among influencers has undergone a significant “gravity shift”: from cloud-based super models to local on-device deployment, and from purely digital programming to physical world manipulation.
On-device models hit the “sweet spot”:
- After intensive testing, @zhixianio concluded that Qwen3.6-35B-A3B running locally on a Mac has secured the “sweet spot” throne, surpassing remote LLMs in speed and intelligence, with a native multimodal experience that is even better than cloud-based large models.
- Regarding Google’s newly released Gemma 4 12B Coder, @zhixianio noted that despite optimizations, its 12B size remains a bottleneck for complex, “long-form, stateful, single-shot” generation tasks (like writing a complete game), showing a significant gap compared to 35B MoE models. He also tested Gemma 4’s audio capabilities, pointing out poor Chinese recognition but good performance in Japanese and English, and suggested that Quantization Aware Training (QAT) is a new approach to improving on-device efficiency.
- @AI_Jasonyu corroborated the “on-device intelligence” trend from another angle: Baidu’s PP-OCRv6, with extremely few parameters (1.5MB), achieves higher OCR accuracy in the browser than large models like GPT-5.5, proving the huge advantage of “small models mastering vertical scenarios.”
AI Agents Enter the Physical World (AutoResearch):
- @dotey highlighted NVIDIA GEAR lab’s ENPIRE project. This is the first time a fully autonomous research loop for an AI programming agent (design, experiment, failure analysis, code iteration) has been implemented in a real physical environment. The agent can autonomously control robots to perform high-precision tasks and discovered a “physical scaling law”: parallel robots can accelerate research. This marks a leap for AI capabilities from digital code generation to physical-world productivity.
2. Unique Perspectives and Industry Outlook Link to heading
The Real Moat in AI Programming: “Contracts” and “Tests,” Not Code Logic
- @Pluvio9yte proposed that the secret to Vibe Coding lies in “Contract First.” Based on practical experience, he concluded that only by externalizing and clearly defining the contracts for APIs and data models can human-AI collaboration avoid context drift. This framework is more critical than just requirements or code alone.
@ruanyf cites the case of a Cloudflare engineer replicating Next.js, sharply pointing out: “Code itself no longer has a moat; testing is the new moat.” This is because AI can easily replicate large projects, but the core barrier is the ability to pass high-quality tests and ensure stable operation.
A Rational Look and “Contrarian View” on Claude Fable 5
- @Pluvio9yte provided a deep-dive “contrarian” experience, differing from the hype: Fable 5 is extremely slow, token consumption isn’t as outrageous as imagined (about 1.5 times that of Opus), and while its capability boundaries are wider, it hasn’t reached a “stunning” level. It feels more like a hybrid of Opus 4.6++ and GPT-5.5++, prompting a call for rational use.
- @dotey pointed out the token consumption black hole in Claude Code’s new Dynamic Workflows, where a simple task consumed 1.3 million tokens.
Production Relations and Economics in the AI Era
- @ruanyf raised a sharp sociological question: After AI significantly boosts efficiency, will employees get time off? If there are no raises and no breaks, what is the point of AI for employees? He also calculated that the cost of unlimitedly using top-tier models for AI programming already far exceeds human programmer salaries, suggesting that businesses will need to weigh the ROI of AI in the future.
- @vista8 shared a forward-looking perspective from the CEO of Factory AI: In the future, the most valuable people will be engineers who can deliver end-to-end business results, not just those who write code. Furthermore, within three years, the median token expenditure per employee will equal their salary.
AI Product Design and Traffic Strategies
- @Pluvio9yte suggested a way to eliminate the “AI feel” from UIs: use DESIGN.md files from famous brands as a constraint for AI generation to enhance the quality.
- @gefei55 shared practical SEO experience, stressing the importance of evolving with Google’s algorithm and warning that low-quality AIGC content will eventually face a backlash. He also shared a highly profitable domain investment story and the feasibility of a high-pricing strategy for SaaS overseas.
- @AI_Jasonyu observed that the paywall competition in the AI video space has shifted from competing on features to competing on “the explanation of the credit system.” The “vertical scenario + long-form to short-form video” logic, like that of OpusClip, is best suited for independent developers.
3. Recommended Tools & Resources Link to heading
| Tool/Resource | Recommended by | Core Highlights & Use Case |
|---|---|---|
| baoyu-design Skill | @dotey | Local design-to-code. Supports importing Figma files, generating design systems and PPTs locally, and can even export to editable PPTX files. |
| info-digest Skill | @dotey | AI News Digesting Assistant. Baoyu’s public daily writing Skill, containing adaptable strategies like reader perspective, fact-checking, and refined formatting. |
| getdesign.md | @Pluvio9yte | The ultimate tool to remove the “AI feel” from UI design. A collection of DESIGN.md system files from real brands like Linear, Vercel, and Apple, which can be fed to an AI to generate high-quality UIs. |
| Papr | @vista8 | A lightweight, open-source RSS client. Supports connecting your own API key for AI summaries and Q&A. |
| Figma Chrome Extension | @vista8 | Game-changer for website cloning. One-click conversion of any webpage element into editable layers and import into Figma. |
| PP-OCRv6 | @AI_Jasonyu | Ultimate on-device OCR. A 1.5MB model that can run in the browser, surpassing large models like GPT-5.5 in speed and accuracy. Fully open-source. |
| App Store Review Analysis Tool | @vista8 | Open-source user feedback miner. Can scrape reviews for any app and use an LLM to analyze pain points and opportunities. |
| GPT Image Prompt | @dotey (from @Ciri_ai) | Photo to doodle illustration. A specific prompt to transform photos into a “decorative folk flat doodle style.” |
| Maccy / Mos | @Pluvio9yte | Mac productivity duo. Maccy is an open-source clipboard tool, and Mos solves the unnatural scrolling direction issue with external mice. |
📚 Appendix: Today’s Watch List Source Update Link to heading
Timeframe: Last 3 days; 22 sources covered; 34 updates in total
Stratechery by Ben Thompson (A_full) Link to heading
- Fox Buys Roku, The Problem With Fox’s Smart Strategy, Streaming That Works
- Published: 2026-06-16 18:00 Beijing Time
- Abstract: - The market hates Fox’s acquisition of Roku, but the company is trading extraction from rights holders for leverage as a renter.
- $15/month or $150/year.
- Substantive analysis of the day’s news via three emails or a podcast per week.
- Strategy Interviews.
- Interviews with leading public company CEOs, private company founders, and discussions with fellow analysts.
- EN Key Points:
- The market hates Fox’s acquisition of Roku, but the company is trading extraction from rights holders for leverage as a renter.
OpenAI Blog (A_full) Link to heading
- Predicting model behavior before release by simulating deployment
- Published: 2026-06-16 08:00 Beijing Time
- Abstract: - Before releasing new models, labs need to understand not only what they can do, but also how they will behave in real-world use, including where they might introduce new risks.
- This becomes even more important as capabilities increase.
- As part of our pre-deployment safety review, we use targeted evaluations, red teaming, and other checks to understand model behavior.
- We are now beginning to use a method to simulate model deployments before they happen, which adds a complementary signal: a deployment-like preview of a candidate model’s behavior before it reaches users.
- Deployment simulation is a method of simulating future deployments before they occur.
- EN Key Points:
- OpenAI introduces Deployment Simulation, a method to predict AI model behavior before deployment using real conversation data to improve safety and evaluation a…
Google DeepMind Blog (A_full) Link to heading
- Unlocking UK house-building with AI-accelerated planning
- Published: 2026-06-17 05:29 Beijing Time
- Abstract: - The UK government is partnering with Google DeepMind to build a new AI prototype aimed at making faster housing decisions.
- This article from the Google DeepMind Blog explains how to unlock UK house-building with AI-accelerated planning, shaping the broader AI and infrastructure landscape.
- After unlocking UK house-building through AI-accelerated planning, it also brings practical implications for founders, operators, and investors.
- EN Key Points:
- UK government partners with Google DeepMind to build a new AI-powered prototype aimed at faster housing decisions.
Two Minute Papers (B_intro+search) Link to heading
- They Looked Inside Claude’s AI’s Mind. It Got Weird
- Published: 2026-06-16 23:53 Beijing Time
- Abstract: - ❤️ Check out Lambda and sign up for their GPU Cloud here:.
- 📝 The paper is available here:.
- Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi.
- They delved deep into the mind of Claude AI.
- EN Key Points:
- ❤️ Check out Lambda here and sign up for their GPU Cloud:
- 📝 The paper is available here:
- 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
- Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Ska…
- EN Key Points:
ArXiv cs.AI (B_intro+search) Link to heading
A Definition of Good Explanations and the Challenges Explaining LLM Outputs
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14838v1 Announcement Type: New.
- Abstract: How to define a good explanation is a long-standing philosophical debate that has recently seen renewed interest in the context of AI outputs.
- Explainability is crucial for the adoption of AI in many contexts, but in order to produce good explanations for AI systems, we must first have an understanding of what constitutes a good explanation.
- In this paper, we propose a definition inspired by the concept of counterfactual explanations, but we argue that one must also consider the interlocutor’s prior beliefs about each fact that might be provided in the explanation.
- EN Key Points:
- arXiv:2606.14838v1 Announce Type: new
- Abstract: How to define a good explanation is a long-standing philosophical debate which has found recent renewed interest in the context of AI outputs
- Explainability is crucial for AI adoption in many contexts, but in order to produce good explanations of AI systems, we must first have an understanding of what…
- In this paper we propose a definition inspired by the notion of counterfactual explanations, however we argue that one must also take into account the interlocu…
Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14885v1 Announcement Type: New.
- Abstract: Agentic search over large corpora relies on retriever-mediated interfaces (e.g., BM25 or ColBERT) for scalable candidate discovery.
- While effective at ranking relevant documents, these interfaces only expose evidence as ranked results or bounded document views, limiting an agent’s ability to reorganize materials and validate cross-document constraints.
- Direct Corpus Interaction (DCI) addresses this limitation by exposing shell-executable corpus operations for flexible search, filtering, comparison, and validation.
- EN Key Points:
- arXiv:2606.14885v1 Announce Type: new
- Abstract: Agentic search over large corpora relies on retriever-mediated interfaces (e.g., BM25 or ColBERT) for scalable candidate discovery
While effective at ranking relevant documents, these interfaces expose evidence only as ranked results or bounded document views, limiting agents’ ability to re…
Direct Corpus Interaction (DCI) addresses this limitation by exposing shell-executable corpus operations for flexible search, filtering, comparison, and verific…
Relational Structural Causal Models
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14892v1 Announce Type: new.
- Abstract: An artificial intelligence must have a causal model of its environment, supporting reasoning about interventions and counterfactuals, and also a compositional model of its environment, supporting generalization to unseen combinations of objects.
- In this work, we formally study when and how such a model can be learned.
- We develop relational structural causal models, extending structural causal models (Pearl 2009) to settings where objects and their relations vary.
- EN 要点:
- arXiv:2606.14892v1 Announce Type: new
- Abstract: An artificial intelligence must have a model of its environment that is causal, supporting reasoning about interventions and counterfactuals, and also…
- In this work, we formally study when and how such a model can be learned
- We develop relational structural causal models, extending structural causal models (Pearl 2009) to settings where objects and their relations vary
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14923v1 Announce Type: new.
- Abstract: As language-model agents increasingly work in teams, each agent must decide how much to trust its teammates.
- However, we lack a standard way to measure trust between AI agents.
- We propose a behavioral measure based on costly verification.
- EN 要点:
- arXiv:2606.14923v1 Announce Type: new
- Abstract: As language-model agents increasingly work in teams, each agent must decide how much to trust its teammates
- Yet we lack a standard way to measure trust between AI agents
- We propose a behavioral measure based on costly verification
PrologMCP: A Standardized Prolog Tool Interface for LLM Agents
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14935v1 Announce Type: new.
- Abstract: State-of-the-art language models fine-tuned for reasoning still fail on deep deductive tasks, and the cost of improving performance by expanding internal reasoning is also poor.
- Symbolic delegation offers a complementary path: the language model translates the problem, while the solver performs the reasoning.
However, current autoformalization pipelines for logic programming are typically bespoke integrations tied to particular tasks or agents.
- EN Highlights:
- arXiv:2606.14935v1 Announce Type: new
- Abstract: Frontier reasoning-tuned language models still fail on deductive tasks at depth, and the cost of improved performance through extended internal reason…
- Symbolic delegation offers a complementary route: a language model translates the problem, while a solver performs the inference
- However, current autoformalization pipelines for logic programming are typically bespoke integrations tied to particular tasks or agents
- EN Highlights:
Semantics-Enhanced Retrieval-Augmented Time Series Forecasting
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14941v1 Announce Type: new.
- Abstract: Time series forecasting models often benefit from historical patterns.
- Inspired by Retrieval-Augmented Generation (RAG), recent research explored retrieving relevant historical time series segments to enhance forecasting.
- However, relying solely on time series similarity is often insufficient for retrieval under non-stationarity.
- EN Highlights:
- arXiv:2606.14941v1 Announce Type: new
- Abstract: Time series forecasting models often benefit from historical patterns
- Inspired by Retrieval-Augmented Generation (RAG), recent research explored retrieving relevant historical time series segments to enhance forecasting
- However, relying solely on time series similarity is often insufficient for retrieval under non-stationarity
AI Engram: In Search of Memory Traces in Artificial Intelligence
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14997v1 Announce Type: new.
- Abstract: Memory formation is fundamental to intelligence, yet whether deep neural networks preserve identifiable memory traces analogous to biological memory units remains an open question.
- This work introduces a geometric framework to identify such “AI engrams” by formalizing the neuroscientific criteria of specificity, reactivation, sufficiency, and necessity as a constrained inverse problem.
- We derive a closed-form estimator that isolates individual memory traces from globally entangled parameters and show that this biologically-derived solution corresponds to a natural gradient update on the parameter manifold.
- EN Highlights:
- arXiv:2606.14997v1 Announce Type: new
- Abstract: Memory formation is fundamental to intelligence, yet whether deep neural networks preserve identifiable memory traces analogous to biological memory u…
- This work introduces a geometric framework to identify such “AI engrams” by formalizing the neuroscientific criteria of specificity, reactivation, sufficiency,…
We derive a closed-form estimator that isolates individual memory traces from globally entangled parameters, and show that this biologically-derived solution co…
Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.15029v1 Announcement Type: New.
- Abstract: LLM judges are used to reduce the need for expensive human labor when evaluating open-ended text generation.
- However, the reliability of these judges largely depends on their consistency with human raters—a property that itself relies on costly human annotations.
- In this work, we develop a method (Metric Match) for estimating correlation-based reliability metrics of LLM judges from limited annotations.
- EN 要点:
- arXiv:2606.15029v1 Announce Type: new
- Abstract: LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation
- However, the reliability of these judges depends critically on their alignment with human raters – a property that itself depends on costly human annotations
- In this work, we develop a method (Metric Match) for estimating correlation-based reliability metrics of LLM judges from limited annotations
OSGuard: A Benchmark for Safety in Computer-Use Agents
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.15034v1 Announcement Type: New.
- Abstract: Computer-use agents are increasingly evaluated on whether they complete realistic desktop and Web tasks.
- However, task success alone can miss failures where an agent reaches the nominal goal through an unsafe shortcut.
- We introduce OSGuard, a dual-granularity benchmark suite for evaluating safety in computer-use agents under benign, unchanged user instructions.
- EN 要点:
- arXiv:2606.15034v1 Announce Type: new
- Abstract: Computer-use agents are increasingly evaluated by whether they complete realistic desktop and web tasks
- However, task success alone can miss failures in which an agent reaches the nominal goal through an unsafe shortcut
- We introduce OSGuard, a dual-granularity benchmark suite for evaluating safety in computer-use agents under benign, unchanged user instructions
Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.15038v1 Announcement Type: New.
- Abstract: Accurately predicting Time-to-Event (TTE) from multimodal clinical data remains challenging due to modal imbalances and distribution shifts.
We introduce a foundation model-driven framework for cross-modal representation alignment between CT imaging and longitudinal EHR data, designed to generalize across tasks and institutions.
CT and EHR modalities are encoded independently using domain-specific foundation models and aligned in a shared latent space through four principled fusion strategies: late fusion, contrastive alignment, cross-attention, and co-attention.
EN Highlights:
- arXiv:2606.15038v1 Announce Type: new
- Abstract: Accurate time-to-event (TTE) prediction from multimodal clinical data remains challenging due to modality imbalance and distribution shift
- We introduce a foundation model-driven framework for cross-modal representation alignment between CT imaging and longitudinal EHR data, designed to generalize a…
- CT and EHR modalities are encoded independently using domain-specific foundation models and aligned in a shared latent space through four principled fusion stra…
ArXiv cs.CL (B_intro+search) Link to heading
PhoneHarness: Harnessing Phone-Use Agents through Mixed GUI, CLI, and Tool Actions
- Published: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14832v1 Announce Type: new.
- Abstract: Phone agents are increasingly expected to complete real mobile workflows rather than merely predict the next screen action.
- However, much of the current mobile-agent literature still evaluates agents primarily as GUI controllers that observe a screen, emit taps and swipes, and are scored based on the target application state.
- Real phone-use tasks are broader: they require deciding when to use app GUIs, device-side commands, or structured tools, while leaving evidence that the intended side effects have indeed occurred.
- EN Highlights:
- arXiv:2606.14832v1 Announce Type: new
- Abstract: Phone agents are increasingly expected to complete real mobile workflows rather than merely predict the next screen action
- However, much of the current mobile-agent literature still evaluates agents primarily as GUI controllers that observe a screen, emit taps and swipes, and are sc…
- Real phone-use tasks are broader: they require deciding when to use app GUIs, device-side commands, or structured tools, while leaving evidence that the intende…
Evaluating the Robustness of Proof Autoformalization in Lean 4
- Published: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14867v1 Announce Type: new.
- Abstract: Proof autoformalization aims to translate informal mathematical proofs written in natural language into formal proofs in a formal language (e.g., Lean 4).
- Several works have developed LLM-based models for proof autoformalization.
- However, existing evaluations often focus on translating well-formed informal proofs from curated datasets.
- EN Highlights:
- arXiv:2606.14867v1 Announce Type: new
Abstract: Proof autoformalization aims to translate a mathematical informal proof written in natural language into a formal proof in a formal language such as L…
Several works have developed LLM-based models for proof autoformalization
However, existing evaluations have typically focused on translating well-formed informal proofs from curated datasets
- Published: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14875v1 Announce Type: new.
- Abstract: We study context compression for multi-hop question answering with small language models.
- We propose Telegraph English, a readable symbolic format that rewrites retrieved passages into structured entity-relation statements, preserving reasoning evidence at a lower token cost.
- In controlled experiments on MuSiQue, TwoWiki, and HotpotQA, Telegraph English outperforms three matched-budget compression baselines (character-level deletion, truncation, and random subsampling) with gains of 13 to 20 F1 percentage points.
- EN Key Points:
- arXiv:2606.14875v1 Announce Type: new
- Abstract: We study context compression for multi-hop question answering with small language models
- We propose Telegraph English, a readable symbolic format that rewrites retrieved passages into structured entity-relation statements, preserving reasoning evide…
- In controlled experiments on MuSiQue, TwoWiki, and HotpotQA, Telegraph English outperforms three matched-budget compression baselines (character-level deletion,…
Simplifying the Modeling of Arbitrary Conditionals in Natural Language
- Published: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14943v1 Announce Type: new.
- Abstract: Causal Transformers model sequences through an autoregressive factorization of the joint distribution, which enables efficient left-to-right decoding and conditional likelihood computation.
- However, they cannot tractably sample from or evaluate arbitrary conditionals – e.g., a block of text conditioned on past and future tokens.
- Recent work has aimed to solve this with novel architectures, but they often result in suboptimal modeling of such conditionals and degenerate generation.
- EN Key Points:
- arXiv:2606.14943v1 Announce Type: new
- Abstract: Causal Transformers model sequences through an autoregressive factorization of the joint distribution, which enables efficient left-to-right decoding…
- However, they cannot tractably sample from or evaluate arbitrary conditionals – e.g., a block of text conditioned on past and future tokens
CoRA: Confidence-Rationale Alignment for Reliable Chain-of-Thought Reasoning
- Release Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14961v1 Announce Type: new.
- Abstract: Chain-of-thought (CoT) reasoning can improve LLM performance, but high answer confidence can be misleading when the accompanying CoT rationale is seemingly plausible but incomplete or unsupported.
- We study confidence-rationale alignment: whether a model’s confidence in its committed answer is justified by its generated rationale.
- We introduce a GRPO-based reinforcement learning framework that jointly rewards answer correctness, committed-answer probability, and rubric-based rationale support, where the rubric assesses grounding, coherence, task-matching, and connection to the chosen answer without revealing the gold answer to the judge.
- EN Key Points:
- arXiv:2606.14961v1 Announce Type: new
- Abstract: Chain-of-thought (CoT) reasoning can improve LLM performance, but high answer confidence may be misleading when the accompanying CoT rationale is plau…
- We study confidence–rationale alignment: whether a model’s confidence in its committed answer is justified by its generated rationale
- We introduce a GRPO-based reinforcement learning framework that jointly rewards answer correctness, committed-answer probability, and rubric-based rationale sup…
- Release Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.15007v1 Announce Type: new.
- Abstract: We introduce Nemotron 3 Ultra, a Mixture-of-Experts hybrid Mamba-Attention language model with a total of 550 billion and 55 billion active parameters.
- We pre-trained Nemotron 3 Ultra on 20 trillion text tokens, then extended the context length to 1M tokens, and post-trained using Supervised Fine-Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation (MOPD).
- Nemotron 3 Ultra is our most powerful model to date, employing several key technologies - LatentMoE, Multi-Token Prediction (MTP), NVFP4 pre-training, multi-environment RLVR, MOPD, and inference budget control.
- EN Key Points:
- arXiv:2606.15007v1 Announce Type: new
- Abstract: We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model
- We pre-trained Nemotron 3 Ultra on 20 trillion text tokens, then extended the context length to 1M tokens, and post-trained using Supervised Fine Tuning (SFT),…
Nemotron 3 Ultra is our most capable model yet, employing multiple key technologies - LatentMoE, Multi Token Prediction (MTP), NVFP4 pre-training, multi-environ…
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.15017v1 Announce Type: new.
- Abstract: Online web agents often augment a base actor with memory, workflow, or skill modules.
- These modules can improve performance, but they also consume test-time tokens, a cost rarely reported alongside the actor’s inference cost.
- We study online augmentation, where this overhead is paid on every task, and re-evaluate its benefits under a fixed total inference budget.
- EN Highlights:
- arXiv:2606.15017v1 Announce Type: new
- Abstract: Online web agents often augment a base actor with memory, workflow, or skill modules
- These modules can improve performance, but they also consume test-time tokens, a cost rarely reported alongside the actor’s inference cost
- We study online augmentation, where this overhead is paid on every task, and re-evaluate its benefits under a fixed total inference budget
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.15026v1 Announce Type: new.
- Abstract: Physiological stress and emotion recognition are important for health monitoring and affective computing.
- In this work, we present a comprehensive evaluation of deep learning models such as Long Short-Term Memory (LSTM), Temporal Convolutional Networks (TCN), and Transformer on the WESAD dataset for multimodal emotion recognition using wrist and chest sensor signals.
- We perform ablation studies to assess the individual contributions of each modality by training models on wrist-only and chest-only inputs.
- EN Highlights:
- arXiv:2606.15026v1 Announce Type: new
- Abstract: Physiological stress and emotion recognition are important for health monitoring and affective computing
- In this work, we present a comprehensive evaluation of deep learning models such as Long Short-Term Memory (LSTM), Temporal Convolutional Networks (TCN), and Tr…
- We perform ablation studies to assess the individual contributions of each modality by training models on wrist-only and chest-only inputs
ReportQA: QA-Based Radiology Report Evaluation
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.15037v1 Announce Type: new.
Abstract: Radiology report evaluation is essential for advancing automated report generation.
Natural language generation metrics have limited clinical relevance.
Clinical efficacy (CE) metrics evaluate important medical findings, but focus mainly on presence and cover only a limited set of entities.
EN Highlights:
- arXiv:2606.15037v1 Announce Type: new
- Abstract: Radiology report evaluation is essential for advancing automated report generation
- Natural language generation metrics have limited clinical relevance
- Clinical efficacy (CE) metrics evaluate important medical findings, but focus mainly on presence and cover only a limited set of entities
Equity with Efficiency: An Empirical Study of Tokenizers for Multilingual Large Language Models
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.15044v1 Announcement Type: New.
- Abstract: Multilingual Large Language Models (LLMs) depend on subword tokenization to bridge discrete text and continuous neural representation.
- State-of-the-art multilingual LLMs often use Byte-level Byte-Pair Encoding (BPE) tokenizers that structurally favor high-resource languages and Latin scripts.
- For speakers of underrepresented languages, particularly those across Southeast Asia, this bias inflates inference costs and widens the cross-lingual capability gap.
- EN Highlights:
- arXiv:2606.15044v1 Announce Type: new
- Abstract: Multilingual large language models (LLMs) depend on subword tokenization to bridge discrete text and continuous neural representation
- State-of-the-art multilingual LLMs often use Byte-level Byte-Pair Encoding (BPE) tokenizers that structurally favor high-resource languages and Latin scripts
- For speakers of underrepresented languages, particularly those across Southeast Asia, this bias inflates inference costs and widens cross-lingual capability gap…
ArXiv cs.LG (B_intro+search) Link to heading
QPILOTS: Efficient Test-Time Q-Steering for Flow Policies
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14801v1 Announcement Type: New.
- Abstract: Flow-matching and diffusion policies are expressive action generators, but optimizing them with temporal-difference reinforcement learning (RL) remains difficult.
- Effective policy extraction requires leveraging the critic’s action-gradients, but directly backpropagating this signal through a multi-step denoising process can be numerically unstable.
- Existing methods address this issue by discarding gradient information, distilling the policy into a simpler single-step actor, or repeatedly fine-tuning the denoising policy as the critic improves.
- EN Highlights:
- arXiv:2606.14801v1 Announce Type: new
- Abstract: Flow-matching and diffusion policies are expressive action generators, but optimizing them with temporal-difference reinforcement learning (RL) remain…
Effective policy extraction requires exploiting the critic’s action gradient, yet directly backpropagating this signal through a multi-step denoising process ca…
Existing methods work around this either by discarding gradient information, distilling the policy into a simpler one-step actor, or repeatedly fine-tuning the…
GRAPE: Guided Parameter-Space Evolution for Compact Adversarial Robustness
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14865v1 Announce Type: new.
- Abstract: Adversarial Training (AT) improves neural network robustness, but most methods train a fixed parameter space from the start.
- This paper asks whether the order in which parameters become optimizable can affect the final robust solution, even when the final architecture or computation budget is controlled.
- We propose GRAPE (Guided Parameter-Space Evolution), a training framework for compact adversarial robustness.
- EN Highlights:
- arXiv:2606.14865v1 Announce Type: new
- Abstract: Adversarial Training (AT) improves neural network robustness, but most methods train a fixed parameter space from the start
- This paper asks whether the order in which parameters become optimizable can affect the final robust solution, even when the final architecture or computation b…
- We propose GRAPE, Guided Parameter-Space Evolution, a training framework for compact adversarial robustness
{\alpha}-Fair Insurance Pricing: A Fairness Continuum
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14898v1 Announce Type: new.
- Abstract: Fairness in insurance pricing remains a long-standing and deeply debated puzzle.
- On one hand, insurers, driven by profitability considerations, set premiums that differentiate across individual risks to achieve actuarial fairness.
- On the other hand, insurance serves a critical societal function by pooling risks across a population, incentivizing cross-subsidization among groups to promote solidarity fairness.
- EN Highlights:
- arXiv:2606.14898v1 Announce Type: new
- Abstract: Fairness in insurance pricing remains a long-standing and deeply debated puzzle
- On one hand, insurers, driven by profitability considerations, set premiums that differentiate across individual risks to achieve actuarial fairness
- On the other hand, insurance serves a critical societal function by pooling risks across a population, motivating cross-subsidization among groups to promote so…
GRASP: Gradient-Aligned Sequential Parameter Transfer for Memory-Efficient Multi-Source Learning
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14900v1 Announcement Type: new.
- Abstract: Multi-source transfer learning faces a fundamental scalability bottleneck: existing approaches either require loading all K source models into memory simultaneously during parameter fusion, which needs O(K) memory, or deploying all models at inference, making production deployment infeasible.
- We propose GRASP (Gradient-Aligned Sequential Parameter Transfer), which achieves superior knowledge integration while maintaining O(1) memory consumption through three key innovations: (1) sequential processing, merging one source at a time into an evolving target model; (2) parameter-gradient alignment, selectively transferring only parameters whose optimization direction aligns with the target domain to avoid negative transfer; and (3) iterative fine-tuning to adapt the transferred knowledge before integrating the next source.
- Extensive experiments across three continual learning benchmarks (Yearbook, CLEAR-10, CLEAR-100), spanning 10 to 108-year temporal distribution shifts and four architectures (1.3M to 25.6M parameters), show that GRASP achieves an average accuracy of 93.5% across all datasets and architectures, compared to 71.7% for ensemble methods, while requiring only constant memory, whereas standard multi-source fusion for K models requires memory.
- EN Highlights:
- arXiv:2606.14900v1 Announce Type: new
- Abstract: Multi-source transfer learning faces a fundamental scalability bottleneck: existing approaches require either loading all K source models into memory…
- We propose GRASP (Gradient-Aligned Sequential Parameter Transfer), which achieves superior knowledge integration while maintaining O(1) memory consumption throu…
- Extensive experiments across three continual learning benchmarks (Yearbook, CLEAR-10, CLEAR-100) spanning 10 to 108-year temporal distribution shifts and four a…
Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14929v1 Announcement Type: new.
- Abstract: Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models.
- Despite its practical significance, this problem remains poorly understood under realistic conditions like adversarial queries, bandit feedback, and limited observability of the models.
- We formalize embedding model routing as an adversarial contextual linear bandit with low-rank experts, where contexts are queries, actions are items, and experts are the embedding models that operate on a low-rank latent representation space.
- EN Highlights:
- arXiv:2606.14929v1 Announce Type: new
- Abstract: Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models
- Despite its practical significance, this problem remains poorly understood under realistic conditions like adversarial queries, bandit feedback, and limited obs…
We formalize embedding model routing as an adversarial contextual linear bandit with low-rank experts, where contexts are queries, actions are items, and expert…
Separable Neural Architectures as Physical World Models: from Mathematical Theory to Applications
- Posted: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14934v1 Announce Type: new.
- Abstract: This work introduces the Separable Neural Architecture (SNA), a function representational class combining neural approximation with tensor decomposition.
- The SNA decouples localized coordinate functions (atoms) from global interactions governed by a sparse, low-rank interaction object.
- This architecture possesses a compact and smooth inductive bias well-suited for solving partial differential equations (PDEs).
- EN Highlights:
- arXiv:2606.14934v1 Announce Type: new
- Abstract: This work introduces the Separable Neural Architecture (SNA), a function representational class combining neural approximation with tensor decompositi…
- The SNA decouples localized coordinate functions (atoms) from global interactions governed by a sparse, low-rank interaction object
- This architecture possesses a compact and smooth inductive bias well-suited for solving partial differential equations (PDEs)
Remember, Don’t Re-read: Stateful ReAct Agents for Token-Efficient Autonomous Experimentation
- Posted: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14945v1 Announce Type: new.
- Abstract: The autoresearch pattern enables autonomous experimentation by having a large language model (LLM) iteratively modify code to optimize a target metric.
- Its stateless design, however, reconstructs experimental context from scratch at every iteration, incurring $O(n)$ token cost per iteration and $O(n^{2})$ total.
- This work reformulates the pattern as a stateful ReAct agent using LangGraph, where typed persistent state carries experimental history across iterations via a tool-call interface.
- EN Highlights:
- arXiv:2606.14945v1 Announce Type: new
- Abstract: The autoresearch pattern enables autonomous experimentation by having a large language model (LLM) iteratively modify code to optimize a target metric
- Its stateless design, however, reconstructs experimental context from scratch at every iteration, incurring $O(n)$ token cost per iteration and $O(n^{2})$ total
- This work reformulates the pattern as a stateful ReAct agent using LangGraph, where typed persistent state carries experimental history across iterations via a…
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14956v1 Announcement Type: New.
- Abstract: Autonomous driving systems rely on precise trajectory prediction to plan safe and efficient movement.
- Graph Neural Networks (GNNs) have become a promising approach for modeling the spatiotemporal interactions between road agents.
- However, designing GNN architectures for trajectory prediction remains non-standardized, with little guidance on which layers effectively capture spatial interactions and temporal dynamics.
- EN Key Points:
- arXiv:2606.14956v1 Announce Type: new
- Abstract: Autonomous driving systems rely on precise trajectory prediction to plan safe and efficient movement
- Graph Neural Networks (GNNs) have become a promising approach for modelling spatiotemporal interactions among road agents
- However, designing GNN architectures for trajectory prediction remains non-standardized, with little guidance on which graph layers effectively capture spatial…
Leveraging Physiological Signals to Predict Exam Outcomes with Machine Learning
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14960v1 Announcement Type: New.
- Abstract: This study investigates the application of machine learning models to predict exam outcomes by utilizing physiological data collected during exams.
- Physiological stress indicators, including electrodermal activity, heart rate, and skin temperature, are analyzed to reveal their relationship with academic performance.
- A variety of machine learning methods were employed, from standard models like logistic regression, random forests, and support vector machines to more advanced architectures, including transformers, Long Short-Term Memory (LSTM), and Gated Recurrent Unit (GRU) models.
- EN Key Points:
- arXiv:2606.14960v1 Announce Type: new
- Abstract: This study investigates the application of machine learning models to predict exam outcomes using physiological data collected during examination sess…
- Physiological stress indicators, including electrodermal activity, heart rate, and skin temperature, were analyzed to uncover their association with academic pe…
- A variety of machine learning approaches were employed, ranging from standard models like logistic regression, random forest, and support vector machines to mor…
Benchmarking Instance-Dependent Label Noise with Controlled Corruptions
- Publication Time: 2026-06-16 12:00 Beijing Time
- Abstract: - arXiv:2606.14965v1 Announcement Type: New.
- Abstract: Synthetic instance-dependent label noise (IDN) benchmarks are widely used to evaluate noisy label learning methods, but existing methods often generate noise through imperfect annotators or classifier evaluators, thereby obscuring the source of ambiguity.
We introduce CILN, a benchmark generation framework that creates IDN through controlled input corruptions.
A diverse voter pool labels corrupted instances, producing benchmark datasets in which both the source and severity of ambiguity are explicit and controllable.
- EN Key Points:
- arXiv:2606.14965v1 Announce Type: new
- Abstract: Synthetic instance-dependent label noise (IDN) benchmarks are widely used to evaluate noisy-label learning methods, yet existing approaches typically…
- We introduce CILN, a benchmark generation framework that creates IDN through controlled input corruptions
- A diverse voter pool labels corrupted instances, producing benchmark datasets in which both the source and severity of ambiguity are explicit and controllable
- EN Key Points: