🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-07-17
- 类型
- ai-daily
- 字数
- 7907
- 阅读时长
- 38 min
2026-07-17 AI Daily | Kimi K3 Escalates Open-Source Scaling War, AI Governance Shifts to “Traceability and Auditability” Link to heading
Today’s main narratives are one explicit and one implicit: Kimi K3’s massive scale is reheating the open-source model competition, but its true capabilities and costs are yet to be validated. Meanwhile, secure access, bioresilience, data traceability, and premise auditing are emerging as key governance priorities. Agents are also enhancing their infrastructure for long-term memory, self-improvement, and collaboration.
📖 In-depth Guide to This Issue’s Watch List Link to heading
The first major theme worth exploring today is the governance issue of “making AI safe and usable.” From ensuring protected AI access for minors to DeepMind’s bioresilience, OriginBlame’s training data traceability, and Grounding Audits’ premise dependency auditing, several articles address the same fundamental question: can AI remain controllable, deletable, and verifiable while being openly accessible? The second theme is the development of agent infrastructure. When considered together, Oracle Agent Memory, the Self-Improvements survey, Networked Intelligence, and SPINE clearly show that long-term memory, self-improvement, multi-agent collaboration, and robotics implementation are transitioning from prototypes to systematic engineering. The third theme is methodological, with new papers on safe behavior learning, neuro-symbolic reasoning, and automatic differentiation, which are perfect for R&D teams to study.
🌐 AI Hot Topics on X Link to heading
Topic 1: Moonshot AI Launches Kimi K3, Largest Open Frontier Model at 2.8 Trillion Parameters Link to heading
- Category: AI · News
- Overview: Trending for 10 hours, 11,000 related posts
- What it is: Moonshot AI has released Kimi K3, claiming it to be an open-source frontier large model with 2.8 trillion parameters.
- Why it matters: This indicates that open-source large models are continuing to close the gap with top-tier closed-source models in scale, capability, and competitiveness. It could also lead to a reassessment of ultra-large model architectures, training costs, and the open-source ecosystem within the industry.
- Discussion summary: The discussion on X is centered around whether its actual capabilities live up to its “frontier model” status, the technical implications and training/inference costs of 2.8 trillion parameters, and its real-world advantages compared to other leading models and open-source alternatives.
Topic 2: Cerebras Unveils Blueprint for Internal Knowledge Base Handling 15,000 Daily Queries Link to heading
- Category: AI · News
- Overview: Trending for 2 hours, 124 related posts
- What it is: Cerebras has announced an internal enterprise knowledge base solution that it claims can handle around 15,000 queries per day.
- Why it matters: This move shows that AI infrastructure companies are expanding from training and inference hardware into enterprise knowledge management, reflecting the demand for large models in internal search, Q&A, and workflow automation.
- Discussion summary: The discussion on X focuses on whether its performance and cost are better than existing RAG solutions, its reliability for enterprise data security and access control, and whether Cerebras is leveraging this to enhance its AI inference ecosystem.
Topic 3: Moonshot AI Launches Kimi K3, Matching Top U.S. Models Link to heading
- Category: AI · News
- Overview: Trending for 2 days, 15,000 related posts
- What it is: Moonshot AI has launched its next-generation model, Kimi K3, claiming its performance now approaches or matches top-tier U.S. AI models.
- Why it matters: This is viewed as a signal that Chinese large models are continuing to close the gap with leading U.S. models in reasoning, coding, and general capabilities. It also heightens global AI competition and the debate between open-source and closed-source strategies.
- Discussion summary: Discussions on X are focused on whether Kimi K3’s actual benchmark performance, cost, and usability are sufficient to challenge OpenAI, Anthropic, and Google. Supporters believe it’s evidence of Chinese models catching up quickly, while skeptics question the transparency of its evaluations, its stability in practical applications, and whether its claims are exaggerated.
Topic 4: OpenAI Launches $230 Codex Micro Macro Pad, Sells Out in Hours Link to heading
- Category: AI · News
- Overview: Trending for 1 day, 24,000 related posts
- What it is: OpenAI released a Codex Micro Macro Pad priced at $230, and it sold out within hours.
- Why it matters: This demonstrates the market power of AI brands and developer-oriented hardware. It also points to new commercialization strategies centered on AI tool ecosystems, developer workflows, and brand extensions.
- Discussion summary: Discussions on X are focused on whether the product is merely high-priced merchandise, its practical utility, and the success of OpenAI’s marketing tactic of using a limited release to generate hype and scarcity.
Topic 5: Researcher Adds Vision to Top Open AI Model with Just 50 Million Parameters Link to heading
- Category: AI · News
- Overview: Trending Time: 17 hours ago, Related Posts: 191
- What it is: A researcher has added vision capabilities to a top open-source AI model by adding only about 50 million parameters.
- Why it matters: This suggests that multimodal capabilities do not necessarily rely on significant parameter expansion or retraining. If the method proves effective, it will help reduce model upgrade costs, enhance the scalability of open-source models, and promote more lightweight vision-language integration solutions.
- Discussion Summary: Discussions on X primarily focus on how efficient this “low-parameter-increment for vision capabilities” approach is, whether it’s sufficient to approach native multimodal models, and what such methods imply for the open-source model ecosystem and computational barriers.
Topic 6: Basis AI’s Ode to Accounting Film Traces 8000-Year Legacy Link to heading
- Category: AI · Entertainment
- Overview: Trending Time: , Related Posts: 41
- What it is: Basis AI released a film themed around accounting, tracing its evolution over approximately 8,000 years, from ancient bookkeeping to modern intelligent tools.
- Why it matters: This reflects how AI companies are attempting to package vertical industry applications with generative content and narrative storytelling, highlighting the potential value of AI in finance, auditing, and business operations automation.
- Discussion Summary: Discussions on X are centered on whether the film’s creative concept is novel, whether the accounting industry will be profoundly transformed by AI, and whether this brand narrative is effective education or just marketing hype.
Topic 7: MoonPay and Venice AI Launch Lumara Film Festival for AI-Generated Shorts Link to heading
- Category: AI · Other
- Overview: Trending Time: , Related Posts: 157
- What it is: MoonPay and Venice AI have jointly launched the Lumara Film Festival, calling for and showcasing short films generated by AI.
- Why it matters: The event demonstrates that AI-generated video is moving from technical demos to creative industry applications, and it also shows a further convergence of crypto payments, the creator economy, and generative AI.
- Discussion Summary: Discussions on X focus on whether an AI film festival can lower the barrier to creation and promote exposure for independent creators. At the same time, some question the originality and copyright ownership of AI-generated content, as well as its impact on traditional film and television professionals.
Summary of AI Public Opinion on X Today Link to heading
The main theme of discussion on X today is that “AI capabilities continue to advance, but the market is more concerned with real-world performance and practical value.” Many have reached a consensus: whether it’s ultra-large open-source models like Kimi K3, enterprise knowledge bases, lightweight multimodal enhancements, or generative content events, all indicate that AI is shifting from single-point tech demos to large-scale applications. Furthermore, the open-source camp and Chinese models are rapidly closing the gap with leading closed-source models. The main point of disagreement is about what is “truly cutting-edge”—whether Kimi K3’s parameters, benchmarks, and costs live up to the hype, and whether enterprise solutions and related products are practical innovations or just marketing ploys. Some also question the actual substance of cases like low-parameter increments, multimodal patches, and AI film festivals. Potential risks lie in non-transparent evaluations, over-marketing that inflates expectations, and issues like enterprise data security, copyright ownership, and the displacement of creators/professionals, which may be underestimated amid the excitement.
💡 Influencer Insights Link to heading
Daily AI Influencer Insights Briefing Link to heading
Based on the X posts from several AI influencers over the past 24 hours, we have summarized the most-watched industry trends from yesterday.
1. Technical Trends & Product Hotspots Link to heading
🧠 New Model Battle: Kimi K3 Emerges as a Domestic Dark Horse Link to heading
One of yesterday’s hottest topics was the release of Kimi K3.
- @Pluvio9yte first broke the news: “Kimi K3 is out, rumored to surpass Opus…”
- @vista8 followed up with high praise, calling it the “number one domestic model” and showcasing its stunning front-end code generation capabilities, such as replicating high-quality websites in multiple styles with a single prompt, admiring its “aesthetic sense.” He also announced that a detailed review would be released soon.
💻 AI Programming Ecosystem: The Continued Rise of Codex and the Open-Sourcing of Grok Build Link to heading
OpenAI Codex and Grok Build were two major hotspots in programming tools.
- @dotey provided frequent updates on Codex: its user count surpassed 8 million, and usage quotas were reset again. Its companion hardware keyboard,
kbd-1.0-codex-micro, was officially unveiled with a cool design. He shared his “design-develop” closed-loop workflow based on Claude Code and Codex in detail and recommended a video editing Skill he developed,BaoCut. - ** @vista8** pointed out that Musk has officially open-sourced Grok Build, and shared the complete documentation after AI analyzed its source code. ** @vista8** interpreted this move as “Is this a slap in the face for OpenAI?”, highlighting the route dispute between open-sourcing and closed-sourcing model tools.
🌍 “World Model” Becomes a New Focus Link to heading
** @Pluvio9yte** published an in-depth long article, proposing that “AI is evolving from generating videos to generating worlds”. He elaborated on the concept of “World Model” and highlighted the open-source project Alaya World: it allows users to generate freely explorable and interactive 3D environments in real-time through text, images, or videos, supporting 720P 24FPS streaming generation and stable exploration for over one minute. He believes this is a key step for AI to transform from a tool into an environment.
🏠 On-Device Models and Lightweighting Link to heading
- ** @zhixianio** continues to focus on on-device models and released a podcast special on the topic. He previously reviewed the full-duplex audio and video capabilities of MiniCPM-o 4.5, expressing satisfaction with the performance of the 9B model. At the same time, he agrees with @geekbb’s vision of “Model-Pak” (model cartridge-ization), foreshadowing the future form of on-device model distribution and usage.
- The industry is paying attention to Quantization-Aware Training (QAT). ** @zhixianio** forwarded Google’s QAT model, believing it is “specialized” for quantization during the training phase, which enables models to run better on consumer-grade hardware.
2. Unique Perspectives and Industry Foresight Link to heading
- The Traps and Solutions of “Vibe Coding” ( @gefei55): Addressing the growing phenomenon of developers indulging in using AI to quickly generate Apps, ** @gefei55** sharply pointed out: “Before, you would spend a month writing an App that nobody used; now, you spend ten thousand yuan in tokens per month to write 37 Apps that nobody uses”. The increase in production efficiency does not mean making money. He emphasized that developers must step out of their comfort zone and learn to identify user needs and do promotional marketing, otherwise it’s just a waste of tokens.
- Best Practices for AI Programming Workflow ( @dotey): @dotey shared his efficient development loop: first, use AI to generate and refine design prototypes (Claude Opus/Fable is better than GPT), then let AI recreate the implementation 1:1 based on the design draft, and finally, Agents automatically complete testing and publishing. He emphasized that the core role of “humans” is to propose ideas and ultimately verify them, not to replace AI.
- New Interview Questions in the AI Era ( @ruanyf): He posed a profound question: “If in the future all code is written by AI, how do we hire programmers?” He pointed out that future interviewers will need to assess candidates’ ability to harness AI, rather than handwriting code itself, but how to assess this is still undecided.
- Learning to “Predict” to Understand the Essence ( @lijigang): @lijigang offered an insight: LLMs learn language structures by predicting the next token, and people should also force themselves to see the “generative mechanism” of their field by predicting the next step in that field.
- “Reverse Consumption” under AI Experience ( @AI_Jasonyu & @zhixianio): @AI_Jasonyu shared highly praised prompt techniques seen on Douyin, including asking AI “What are you most unsure about right now?” and “What is my biggest oversight?”, to encourage AI to perform more thorough “self-scrutiny” and improve output quality.
3. Recommended Tools and Resources Link to heading
🛠️ AI Agent Tools & Skills Link to heading
grillme(Skill): ** @Pluvio9yte** highly recommends it. Through intense, exhaustive questioning, it helps you extract all demand details during the development planning phase, avoiding rework later.rn-wechat-extract(Skill): ** @Pluvio9yte** open-sourced. It solves the pain point of Agents being unable to read WeChat official account articles by simulating WeChat UA to bypass blockades, converting articles to Markdown with one click.AnySearch(AI Search Infrastructure): ** @Pluvio9yte** shared. A search tool designed specifically for Agents, supporting vertical domain searches like finance and academia, and can directly convert web pages into structured Markdown, greatly improving Agent research efficiency.qiaomu-tiny-gif(Skill): ** @vista8** open-sourced. Used for GIF compression, solving the GIF size limit problem for official accounts.REDSkillCommunity: ** @ruanyf** observed that Xiaohongshu has started internal testing “REDSkill”, allowing users to upload and share Skill files, attempting to combine social media platforms with a Skill Hub to create a “GitHub for Skills”.- Glossary of Frontend Interaction Design: @vista8 recommended a website that compiles common components, animations, and design style names for Web/App, facilitating communication with AI using precise terminology during Vibe Coding.
📚 Opinions & Reports Link to heading
- Grok Build Source Code Analysis Document: @vista8 has made public an AI’s analysis document of the Grok Build source code, providing valuable material for developers to study.
- Ben’s GPT 5.6 Model Combination Guide: @vista8 relayed the experience of a well-known blogger, Ben: use Sol for creative development, Ultra for complex tasks, and Luna for daily conversations.
The above content was compiled by AI industry analysts based on the views of bloggers such as @zhixianio, @Pluvio9yte, @dotey, @vista8, @ruanyf, @gefei55, @lijigang, and @AI_Jasonyu.
📚 Appendix: Today’s Watch List Update Source List Link to heading
Time window: Last 3 days; 22 sources covered; 34 updates in total.
OpenAI Blog (A_full) Link to heading
Why teens deserve access to safe AI
- Publication Time: 2026-07-17 00:00 Beijing Time
- Summary: - Teenagers are the first generation to grow up with artificial intelligence, and this technology will largely shape their future.
- Today, nearly nine out of ten teens on ChatGPT use it within a week to learn, access information, develop skills, or enhance productivity.
- This is why we believe it is crucial for teens to have access to AI.
- Preventing teens from using it until they are adults would be like asking the previous generation to avoid the internet or search engines before they turned 18, leaving them less prepared for one of the defining technologies of their era.
- But access must be paired with safeguards specifically designed for teens.
- EN Key Points:
- Learn how OpenAI is making ChatGPT safer for teens with age-appropriate protections, learning tools, parental controls, and expert partnerships.
How Cars24 scales conversations and builds faster with OpenAI
- Publication Time: 2026-07-16 08:00 Beijing Time
- Summary: - Cars24 operates one of the world’s largest AI-native automotive ecosystems in India for buying and selling cars, with additional operations in the UAE and Australia.
- The company supports the entire car ownership journey, from discovery and financing to resale and after-sales service. It helps extend the lifecycle of vehicles through a more efficient and convenient used-car ecosystem, in a market where most transactions are still manual, regulated, and fragmented.
Expanding a complex, conversation-driven market. Link to heading
- Unlike traditional e-commerce, buying and selling cars in India rarely happens in a single transaction.
- Much of the process occurs outside of an app, involving phone calls, document checks, and follow-ups that can take days or weeks.
- EN Key Points:
- Cars24 uses OpenAI-powered voice and chat agents to handle 1M+ monthly conversation minutes, recover 12% of lost leads, and bring agentic workflows to teams acr…
Google DeepMind Blog (A_full) Link to heading
- Our approach to bioresilience
- Publication Time: 2026-07-16 17:30 Beijing Time
- Summary: - Google DeepMind and Isomorphic Labs are sharing our joint approach on bioresilience and artificial intelligence models.
- This article from the Google DeepMind blog explains how our approach to bioresilience is shaping the broader AI and infrastructure landscape.
- It also brings practical implications for founders, operators, and investors following our bioresilience approach.
- EN Key Points:
- Google DeepMind and Isomorphic Labs are sharing our joint approach to bioresilience and AI models.
Two Minute Papers (B_intro+search) Link to heading
- Claude Just Revealed AI’s Biggest Problem
- Release Time: 2026-07-16 23:39 Beijing Time
- Abstract: - ❤️ Check out Lambda here and sign up for their GPU Cloud:.
- 📝 The paper is available here:
- Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi.
- Claude just revealed AI’s biggest problem.
- EN Key Points:
- ❤️ Check out Lambda here and sign up for their GPU Cloud:
- 📝 The paper is available here:
- 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
- Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Ska…
ArXiv cs.AI (B_intro+search) Link to heading
OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets
- Release Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13037v1 Announce Type: new.
- Abstract: When a data contributor requests removal, model trainers face a practical gap: unlearning algorithms require a forget set, yet no tool can locate which training records belong to a given author.
- Existing provenance systems operate at the file or dataset level, leading to catastrophic over-deletion.
- We propose ob, a record- and token-level data provenance system that propagates author identity through data processing pipelines and resolves revocation requests into precise forget sets through deterministic queries.
- EN Key Points:
- arXiv:2607.13037v1 Announce Type: new
- Abstract: When a data contributor requests removal, model trainers face a practical gap: unlearning algorithms require a forget set, yet no tool can locate whic…
- Existing provenance systems operate at file or dataset level, forcing catastrophic over-deletion
- We present ob, a record- and token-level data provenance system that propagates author identity through data processing pipelines and resolves revocation reques…
SPINE: Bridging the Cyber-Physical Gap with Agentic AI
- Release Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13049v1 Announce Type: new.
- Abstract: Foundation models provide robots with complex brains for sophisticated decision-making, but deploying this intelligence into physical platforms still requires tedious, expert-driven calibration.
This deployment gap, the robot’s spinal cord, remains a primary bottleneck for scalable Embodied AI.
Therefore, we propose SPINE (Scalable Physical Integration with ageNtic Expertise): an agentic framework for systematically debugging and deploying bimanual robots with minimal robotics expertise.
- EN Key Points:
- arXiv:2607.13049v1 Announce Type: new
- Abstract: Foundation models have given robots a sophisticated brain for complex decision-making, yet deploying that intelligence into a physical platform still…
- This deployment gap, the robot’s spinal cord, remains a primary bottleneck to scalable Embodied AI
- Hence, we propose SPINE (Scalable Physical Integration with ageNtic Expertise): an agentic framework for systematically debugging and deploying bimanual robots…
- EN Key Points:
- Publication Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13069v1 Announce Type: new.
- Abstract: The Chain-of-Thought (CoT) reasoning produced by large language models may appear logically sound, but it may not genuinely depend on its stated premises.
- We introduce interventional grounding audits, a black-box, step-level test of premise dependency: we intervene on a single premise by substituting its target predicate with a new symbol, re-run the model, and check if the normalized conclusion (canonical predicate form) of each reasoning step changes.
- We evaluate on ProntoQA, a synthetic multi-hop deductive reasoning benchmark with gold proof trees, where step-level premise dependencies are known.
- EN Key Points:
- arXiv:2607.13069v1 Announce Type: new
- Abstract: Large language models produce chain-of-thought (CoT) reasoning that appears logically sound yet may not genuinely depend on its stated premises
- We introduce interventional grounding audits, a black-box, step-level test of premise dependency: we intervene on a single premise by substituting its target pr…
- We evaluate on ProntoQA, a synthetic multi-hop deductive reasoning benchmark with gold proof trees, where step-level premise dependencies are known
Probabilistic Extension of Neuro-Symbolic AGI Robots based on Belnap’s Typed Intensional FOL
- Publication Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13073v1 Announce Type: new.
- Abstract: Neuro-symbolic AGI based on $IFOL_B$ is an approach that combines neural learning and symbolic reasoning to overcome the limitations of purely neural systems (e.g., lack of interpretability and logical structure), and features a formal logic mechanism for self-reference.
- In this paper, we extend the cognitive abilities of $IFOL_B$ by performing probabilistic calculations on currently unknown sentences, based on Nilsson’s probabilistic structure for $IFOL_B$.
We introduce global symmetry transformations that preserve the current knowledge database and logical deduction, and local ones for real-time decision-making on concrete (sub)problems involving only a very strict subset of $IFOL_B$ predicates.
- EN Highlights:
- arXiv:2607.13073v1 Announce Type: new
- Abstract: Neuro-symbolic AI based on $IFOL_B$ is a way to combine neural learning and symbolic reasoning to overcome limitations of purely neural systems (like…
- In this paper we expand the cognitive power of $IFOL_B$ by using the probability computation for the currently unknown sentences, based on Nilsson’s probability…
- We introduce the global symmetry transformation that preserves the current knowledge database and logical deduction, and the local one used for real-time decisi…
- EN Highlights:
Self-Improvements in Modern Agentic Systems: A Survey
- Published: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13104v1 Announce Type: new.
- Abstract: Self-improving autonomous agents are moving from research prototypes to deployed systems.
- The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input.
- This survey frames modern self-improving agents as adaptive systems that convert experience into accumulated capability gains.
- EN Highlights:
- arXiv:2607.13104v1 Announce Type: new
- Abstract: Self-improving autonomous agents are moving from research prototypes to deployed systems
- The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input
- This survey frames modern self-improving agents as adaptive systems that convert experience into accumulated capability gains
Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools
- Published: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13115v1 Announce Type: new.
- Abstract: Small language models (SLMs) have shown promise for zero-shot molecular property prediction from SMILES strings, yet they often suffer from structural blindness as their sequence representations do not specify key graph topological cues.
- We propose a modular, context-augmented prompting framework that uses agentic tools at inference time: a trained GNN expert model can confidently provide prediction hints, and the GNN extracts instance-specific explanatory subgraphs (e.g., subgraph SMILES and accompanying explanatory paragraphs).
- We evaluate three popular SLMs on MUTAG and Tox21 under five prompting configurations, ranging from SMILES-only to using all available tools.
- EN Highlights:
- arXiv:2607.13115v1 Announce Type: new
- Abstract: Small language models (SLMs) have shown promise for zero-shot molecular property prediction from SMILES strings, yet they often suffer from structural…
We propose a modular Context-Augmented Prompting framework that enables agentic tool use at inference time: a trained GNN expert model provides a predictive hin…
We evaluate three commonly used SLMs on MUTAG and Tox21 under five prompting configurations ranging from SMILES-only to using all available tools at hand
Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents
- Publication Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13157v1 Announcement Type: New.
- Abstract: Agent memory is a systems problem for long-horizon agents.
- Practical deployments require retaining task state in extended dialogues, restoring user-specific facts and preferences during sessions, and accumulating procedural knowledge from previous results.
- These requirements exceed the scope of document retrieval: the memory layer must determine which interactions become persistent state, how to scope that state, how to retrieve it under latency constraints, and how to modify or delete it over time.
- EN Highlights:
- arXiv:2607.13157v1 Announce Type: new
- Abstract: Agent memory is a systems problem for long-horizon agents
- Practical deployments require retention of task state across extended conversations, recovery of user-specific facts and preferences across sessions, and accumu…
- These requirements extend beyond document retrieval: a memory layer must determine which interactions become durable state, how that state is scoped, how it is…
Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models
- Publication Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13172v1 Announcement Type: New.
- Abstract: We address the problem of safely training an agent policy and deploying a good and safe policy in situations where the environmental dynamics are unknown and no suitable reward function is available.
- In the context of safety-critical environments, we consider traditional reinforcement learning to be impractical and resort to human input resources.
- We introduce DROPJ, a human-centric method for safe training and deployment.
- EN Highlights:
- arXiv:2607.13172v1 Announce Type: new
- Abstract: We address the problem of safely training an agent policy and deploying a good and safe policy, in settings where the environment dynamics are unknown…
- In the context of safety-critical environments, we consider traditional reinforcement learning impractical and resort to the resource of human input
- We introduce DROPJ, a human-centred method for both safe training and deployment
CayleyR: Solving the TopSpin puzzle via cycle intersection
- Publication Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13219v1 Announcement Type: New.
- Abstract: We present cayleyR, an R package for solving permutation puzzles by detecting cycle intersections in Cayley graphs.
- The core algorithm performs an iterative bidirectional search: from both the initial and target permutation states, random operation sequences generate cycles in the Cayley graph of the symmetric group Sn; their intersection yields a connecting path.
- When no direct intersection is found, a distance-guided bridge selection narrows the gap, and the process repeats.
- EN Highlights:
- arXiv:2607.13219v1 Announce Type: new
- Abstract: We present cayleyR, an R package for solving permutation puzzles by detecting cycle intersections in Cayley graphs
- The core algorithm performs an iterative bidirectional search: from both the initial and target permutation states, random operation sequences generate cycles i…
- When no direct intersection is found, a distance-guided bridge selection narrows the gap, and the process repeats
Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science
- Publication Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13220v1 Announcement Type: New.
- Abstract: Most AI-for-science systems focus on scaling a single reasoning process through better models, larger context windows, long-horizon agentic execution, or a digital co-scientist that partners with one primary user.
- However, challenging scientific problems are rarely solved by one reasoner alone.
- These problems are solved by teams whose members bring different priors, experimental backgrounds, tacit knowledge, and domain-trained intuitions.
- EN Highlights:
- arXiv:2607.13220v1 Announce Type: new
- Abstract: Most AI-for-science systems focus on scaling a single reasoning process through better models, larger context windows, long-horizon agentic execution,…
- However, challenging scientific problems are rarely solved by one reasoner alone
- They are solved by teams whose members bring different priors, experimental backgrounds, tacit knowledge, and domain-trained intuitions
ArXiv cs.CL (B_intro+search) Link to heading
FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents
- Publication Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13035v1 Announcement Type: New.
- Abstract: Cloud services frequently experience incidents that require rapid diagnosis and resolution.
- Troubleshooting guides help engineers respond consistently, but creating them manually is labor-intensive, leading to incomplete coverage and outdated documentation.
- We introduce FixItFlow, an automated system that generates troubleshooting guides from historical incident data using large language models.
- EN Highlights:
- arXiv:2607.13035v1 Announce Type: new
Abstract: Cloud services experience frequent incidents that require rapid diagnosis and resolution
Troubleshooting guides help engineers respond consistently, but creating them manually is labor-intensive, resulting in incomplete coverage and outdated documen…
We present FixItFlow, an automated system that generates troubleshooting guides from historical incident data using large language models
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry
- Published: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13036v1 Announcement Type: New.
- Abstract: Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving.
- When the available information is insufficient to support a reliable answer, models should request clarification or abstain rather than provide an unsupported answer.
- However, existing medical benchmarks typically assume that complete information is available upfront.
- EN Highlights:
- arXiv:2607.13036v1 Announce Type: new
- Abstract: Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving
- When the available information is insufficient to support a reliable answer, models should request clarification or abstain rather than provide unsupported resp…
- Existing medical benchmarks, however, typically assume that complete information is available upfront
The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI
- Published: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13044v1 Announcement Type: New.
- Abstract: The European Patent Office (EPO) reported record filings in 2025, and the 2026 EPO Guidelines hold applicants strictly responsible for LLM-assisted content under Articles 83 and 42, which creates pressure to classify suspected AI-generated patent texts.
- Two constraints make this difficult.
- First, realistic prosecution settings often only have consumer-grade GPUs with about 8 GB of VRAM, not data-center-class scoring stacks.
- EN Highlights:
- arXiv:2607.13044v1 Announce Type: new
- Abstract: The European Patent Office (EPO) reported record filings in 2025, and the 2026 EPO Guidelines hold applicants strictly responsible for LLM-assisted co…
- Two constraints make this hard
- First, realistic prosecution settings often have only consumer GPUs with about 8 GB VRAM, not datacenter-class scoring stacks
- Publication Time: 2026-07-16 12:00 Beijing Time
- Abstract:
- arXiv:2607.13158v1 Announce Type: new.
- Abstract: Simultaneous speech translation (SimulST) requires incremental translation under strict latency constraints, but remains challenging for decoder-only LLM systems due to limited context and cross-lingual reordering.
- Recent approaches often introduce architectural changes or explicit read/write policies to control output timing, which can be brittle in conversational speech where segment boundaries are ambiguous.
- We present a simple data-driven alternative: fixed-length chunks for cumulative streaming decoding with a rewind-based committed prefix, and teacher-labeled prefix-to-prefix (P2P) targets with limited-wait fine-tuning, yielding CSSEL-P2P, where CSSEL is our proposed chunked streaming speech encoder LLM.
- EN Highlights:
- arXiv:2607.13158v1 Announce Type: new
- Abstract: Simultaneous speech translation (SimulST) requires incremental translation under strict latency constraints, yet remains challenging for decoder-only…
- Recent approaches often introduce architectural changes or explicit read/write policies to control output timing, which can be brittle in conversational speech…
- We present a simple data-driven alternative: fixed-length chunks for cumulative streaming decoding with a rewind-based committed prefix, and teacher-labeled pre…
What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
- Publication Time: 2026-07-16 12:00 Beijing Time
- Abstract:
- arXiv:2607.13162v1 Announce Type: new.
- Abstract: What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone.
- Persona vectors, behavioral directions in activation space, can probe this organization, but prior work has covered only a few traits.
- We present the first systematic application of persona vectors at this scale, compiling a 53-trait inventory across four behaviorally distinct domains and labeling each trait in two open-weight models as natural (expressed on baseline), manipulable (latent but amplifiable), or intractable (difficult to extract conventionally).
- EN Highlights:
- arXiv:2607.13162v1 Announce Type: new
- Abstract: What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by…
- Persona vectors, behavioral directions in activation space, can probe this organization, but prior work covers only a handful of traits
- We present the first systematic application of persona vectors at this scale, compiling a 53-trait inventory across four behaviorally distinct domains and label…
Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation
- Release Time: 2026-07-16 12:00 Beijing Time
- Abstract:
- arXiv:2607.13164v1 Announcement Type: new.
- Abstract: Sign language is the primary communication channel for millions of Deaf and hard-of-hearing people, yet text-to-signer video generation remains costly due to the high training and evaluation costs of video diffusion models.
- This paper introduces Text2Sign, a text-conditioned diffusion model for short sign language clips that runs on a single NVIDIA L4 GPU.
- It combines a frozen visual-language text encoder with a 3D encoder-decoder and factorized spatiotemporal attention to reduce the cost of full-video attention while maintaining motion coherence.
- EN 要点:
- arXiv:2607.13164v1 Announce Type: new
- Abstract: Sign language is a primary communication channel for millions of Deaf and hard-of-hearing people, yet text-to-signer video generation remains costly b…
- This paper presents Text2Sign, a text-conditioned diffusion model for short sign-language clips that runs on a single NVIDIA L4 GPU
- It combines a frozen vision-language text encoder with a 3D encoder-decoder and factorized spatiotemporal attention to reduce the cost of full-video attention w…
- Release Time: 2026-07-16 12:00 Beijing Time
- Abstract:
- arXiv:2607.13189v1 Announcement Type: new.
- Abstract: We introduce RAGthoven, our system for SemEval-2026 Task 1 (MWAHAHA), Subtask A (multilingual constrained humor generation in English, Spanish, and Chinese).
- RAGthoven decomposes creative text generation into a multi-stage large language model (LLM) pipeline (Planner, Best-of-N Writer, Self-Critique Reflector, LLM-as-a-judge Judge), which is grounded in computational humor theories (benign violation theory, script-based semantic theory of humor) and refined through ten experiments.
- In our final configuration, we augment the Planner with retrieval-augmented generation (RAG) from a curated joke corpus, and seed generation with diverse joke mechanisms.
- EN 要点:
- arXiv:2607.13189v1 Announce Type: new
- Abstract: We present RAGthoven, our system for SemEval-2026 Task 1 (MWAHAHA), Subtask A (multilingual constrained humor generation in English, Spanish, and Chin…
- RAGthoven decomposes creative text generation into a multi-stage large language model (LLM) pipeline (Planner, Best-of-N Writer, Reflector for self-critique, LL…
- In our final configuration, we augment the Planner with retrieval-augmented generation (RAG) from a curated joke corpus, seeding generation with diverse joke me…
Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference
- Publish Time: 2026-07-16 12:00 Beijing Time
- Abstract:
- arXiv:2607.13205v1 Announcement Type: new.
- Abstract: Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of long-context models by ranking tokens based on accumulated attention quality (considered here as signal energy) and retaining the heaviest tokens.
- On schema-dense input streams (e.g., nested JSON), this score acts as a non-stationary filter that disproportionately retains noise: non-content sink roles (delimiters or whitespace) carry an order of magnitude more energy than any content role, and structural KEY tokens are over-retained at approximately 1.8 times the rate of VALUE tokens carrying answers, reducing exact match accuracy from 88% to 0% at a 5% budget as the signal-to-noise ratio of the retained state decreases.
- Counterfactual experiments demonstrate that suppressing KEY tokens is the best deployable filter.
- EN Key Points:
- arXiv:2607.13205v1 Announce Type: new
- Abstract: Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking tokens on accum…
- On schema-dense input streams such as nested JSON, this score acts as a non-stationary filter that disproportionately retains noise: a non-content sink role (de…
- A counterfactual experiment establishes that suppressing KEY tokens is the best deployable filter
- Publish Time: 2026-07-16 12:00 Beijing Time
- Abstract:
- arXiv:2607.13248v1 Announcement Type: new.
- Abstract: The evaluation of mathematical reasoning in large language models (LLMs) has predominantly focused on high-resource languages like English.
- This poses a significant barrier to the equitable development and deployment of AI in linguistically diverse regions such as Bangladesh, where over 230 million people speak Bengali.
- Despite its global significance, there has been little prior work on mathematical reasoning in Bengali, and no existing research systematically benchmarks perturbed Bengali mathematical datasets, leaving critical gaps in evaluating model robustness and genuine understanding beyond pattern recognition.
- EN Key Points:
- arXiv:2607.13248v1 Announce Type: new
- Abstract: The evaluation of mathematical reasoning in large language models (LLMs) has predominantly focused on high-resource languages like English
- This has created a significant barrier to the equitable development and deployment of AI in linguistically diverse regions such as Bangladesh, where over 230 mi…
- Despite this global significance, there has been minimal prior work on mathematical reasoning in Bengali and no existing research that systematically benchmarks…
- Publish Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13260v1 Announce Type: new.
- Abstract: Policy documents shape governance outcomes, but their reasoning is often implicit.
- Participatory commitments and managerial control routinely coexist in the same text, and the tensions between them are rarely stated directly.
- Existing computational approaches to policy discourse cannot express the frame-mediated relations that drive these tensions, where one argument narrows or instrumentalizes another rather than rejecting it.
- EN 要点:
- arXiv:2607.13260v1 Announce Type: new
- Abstract: Policy documents shape governance outcomes, but their reasoning is often implicit
- Participatory commitments and managerial control routinely coexist in the same text, and the tensions between them are rarely stated directly
- Existing computational approaches to policy discourse cannot express the frame-mediated relations that drive these tensions, where one argument narrows or instrumentalizes another rather than rejecting it.
ArXiv cs.LG (B_intro+search) Link to heading
- Publish Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13042v1 Announce Type: new.
- Abstract: This paper traces, with explicit numerical values, how PyTorch’s automatic differentiation (AD) engine computes gradients for Physics-Informed Neural Networks (PINN) training—a setting requiring two levels of differentiation: computing physical derivatives through the network $\hat{y}’(t)=d\hat{y}/dt$, and computing the parameter gradients of the loss $\nabla_\theta L$ which itself depends on $\hat{y}’(t)$.
- Using a 1-3-3-1 multilayer perceptron and the initial value problem $y’(t)+y(t)=0$, $y(0)=1$, we trace the complete pipeline at every node: the computational graph built during the forward pass, the reverse-mode back-traversal that computes all 22 parameter gradients in a single pass, and \texttt{create_graph=True} enabling correct differentiation of physics-informed residuals via a graph-on-graph mechanism.
- Each adjoint value is validated against hand derivations from Tahimi (2026), connecting the $P/Q$ sensitivity framework to the vector Jacobian products used by PyTorch’s autograd engine.
- EN 要点:
- arXiv:2607.13042v1 Announce Type: new
- Abstract: This paper traces, with explicit numerical values, how PyTorch’s automatic differentiation (AD) engine computes gradients for Physics-Informed Neural Networks (PINN) training—a setting requiring two levels of differentiation: computing physical derivatives through the network $\hat{y}’(t)=d\hat{y}/dt$, and computing the parameter gradients of the loss $\nabla_\theta L$ which itself depends on $\hat{y}’(t)$.
- Using a 1-3-3-1 multilayer perceptron and the initial value problem $y’(t)+y(t)=0$, $y(0)=1$, we trace the complete pipeline at every node: the computational graph built during the forward pass, the reverse-mode back-traversal that computes all 22 parameter gradients in a single pass, and \texttt{create_graph=True} enabling correct differentiation of physics-informed residuals via a graph-on-graph mechanism.
Every adjoint value is verified against the hand derivations of Tahimi (2026), connecting the $P/Q$ sensitivity framework to the vector–Jacobian products used…
Beyond Backbone Backpropagation: A Decoupled Strategy for Efficient Transfer Learning
- Published: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13043v1 Announce Type: new.
- Abstract: Deep learning models achieve state-of-the-art image classification but face deployment challenges due to computational costs and energy demands.
- We propose a lightweight training strategy that adapts the model’s normalization layers to the new domain and decouples feature extraction from classifier optimization, reducing overhead by pre-calculating features only once.
- A redesigned classifier head with a margin-based weighted loss further minimizes ambiguity without end-to-end backpropagation.
- EN Highlights:
- arXiv:2607.13043v1 Announce Type: new
- Abstract: Deep learning models achieve state-of-the-art image classification but face deployment challenges due to computational costs and energy demands
- We propose a lightweight training strategy that adapts normalization layers of the model to the new domain and decouples feature extraction from classifier opti…
- A redesigned classifier head with margin-based weighted loss further minimizes ambiguity without end-to-end backpropagation
Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges
- Published: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13045v1 Announce Type: new.
- Abstract: Federated Learning (FL) has become a key paradigm for privacy-preserving collaborative model training across distributed and heterogeneous data sources.
- By keeping raw data local, FL addresses data confidentiality concerns, yet it does not resolve the opacity of modern machine learning models.
- Meanwhile, eXplainable Artificial Intelligence (XAI) has gained attention for improving transparency, trust, and accountability, particularly in high-stakes domains.
- EN Highlights:
- arXiv:2607.13045v1 Announce Type: new
- Abstract: Federated Learning (FL) has emerged as a key paradigm for privacy-preserving collaborative model training across distributed and heterogeneous data so…
- By keeping raw data local, FL addresses data confidentiality concerns, yet it does not resolve the opacity of modern machine learning models
- In parallel, Explainable Artificial Intelligence (XAI) has gained attention for improving transparency, trust, and accountability, particularly in high-stakes d…
- Publication Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13046v1 Announcement Type: New.
- Abstract: We develop a framework for the information discarded by machine learning models whose inputs carry a Lie group action.
- Given a representation $\pi$ of a Lie group $G$ on a space $V$ and a learned function $f\colon V \to \mathbb{R}$, we define two objects measuring the symmetry invisible to $f$.
- The null fiber at a point $x \in V$ is the set of group elements $N_G(f,x) = {g \in G : f(\pi(g^{-1}) \cdot x) = f(x)}$, whose inverse action on $x$ is undetectable by $f$.
- EN Highlights:
- arXiv:2607.13046v1 Announce Type: new
- Abstract: We develop a framework for the information discarded by machine learning models whose inputs carry a Lie group action
- Given a representation $\pi$ of a Lie group $G$ on a space $V$ and a learned function $f\colon V \to \mathbb{R}$, we define two objects measuring the symmetry i…
- The null fiber at a point $x \in V$ is the set $N_G(f,x) = {g \in G : f(\pi(g^{-1}) \cdot x) = f(x)}$ of group elements whose inverse action on $x$ is undetec…
Targeted Recovery of Weight-Space Mechanisms From Neural Networks
- Publication Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13047v1 Announcement Type: New.
- Abstract: Parameter decomposition (PD) decomposes neural networks into interpretable computational components that faithfully reflect the original network’s operations.
- However, scaling PD to large models requires vast computation, making it a costly and risky endeavor.
- Here, we propose targeted PD (tPD), which identifies only the components that process specific inputs of interest—from isolated prompts to large subtasks—by introducing a high-level, catch-all component to handle all non-targeted data.
- EN Highlights:
- arXiv:2607.13047v1 Announce Type: new
- Abstract: Parameter decomposition (PD) decomposes neural networks into interpretable computational components that faithfully reflect the original network’s ope…
- However, scaling PD to large models requires vast compute, making it a costly and risky endeavor
- Here we propose targeted PD (tPD), which identifies only the components that process specific inputs of interest – from isolated prompts to large subtasks – b…
Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems
- Publication Time: 2026-07-16 12:00 Beijing Time
Abstract: - arXiv:2607.13048v1 Announce Type: new.
- Abstract: Streaming inference pipelines increasingly pair lightweight fast models with Large Language Models (LLMs) to provide rich semantic understanding at a significant cost.
- The central question of when to invoke the LLM has received limited formal treatment.
- We cast this as a risk-based sequential stopping problem, where a trigger policy fires when a risk functional over the observation history exceeds a threshold.
- EN Highlights:
- arXiv:2607.13048v1 Announce Type: new
- Abstract: Streaming inference pipelines increasingly pair lightweight fast models with Large Language Models (LLMs) that provide rich semantic understanding at…
- The central question of when to invoke the LLM has received limited formal treatment
- We cast this as a risk-based sequential stopping problem, where a trigger policy fires when a risk functional over the observation history exceeds a threshold
- Publication Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13101v1 Announce Type: new.
- Abstract: Global Station Weather Forecasting (GSWF) is pivotal for localized and extreme weather prediction over key regions.
- Despite efforts to exploit look-back windows, existing methods show limited accuracy gains and struggle with extreme events and error accumulation.
- These limitations stem from overreliance on short-term patterns, which are insufficient to capture chaotic weather dynamics, especially under partial observation.
- EN Highlights:
- arXiv:2607.13101v1 Announce Type: new
- Abstract: Global Station Weather Forecasting (GSWF) is pivotal for localized and extreme weather prediction over key regions
- Despite efforts to exploit look-back windows, existing methods show limited accuracy gains and struggle with extreme events and error accumulation
- These limitations stem from overreliance on short-term patterns, which are insufficient to capture chaotic weather dynamics, especially under partial observatio…
Disentangling Knowledge States with Ability and Proficiency Modeling for Knowledge Tracing
- Publication Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13103v1 Announce Type: new.
- Abstract: Knowledge Tracing (KT) aims to predict students’ future performance by modeling their changing knowledge states throughout historical interactions.
- Existing KT methods often treat the raw interaction sequence as a unified behavioral process, overlooking the stage-specific nature of learning behaviors.
- Our preliminary observations indicate that after sufficient practice, students are more likely to correctly answer previously failed knowledge concepts, suggesting a student’s transition from ability cultivation to proficiency-based learning.
- EN Highlights:
- arXiv:2607.13103v1 Announce Type: new
Abstract: Knowledge tracing (KT) aims to predict students’ future performance by modeling their evolving knowledge states from historical interactions
- Existing KT methods usually treat the raw interaction sequence as a unified behavioral process, overlooking the phase-specific nature of learning behaviors
- Our preliminary observations show that students are more likely to correctly answer previously failed knowledge concepts after sufficient practice, suggesting a…
STKAN: Kolmogorov-Arnold Networks for Spatio-Temporal Forecasting
- Release Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13108v1 Announcement Type: new.
- Abstract: Real-world traffic data exhibits heterogeneous spatial correlations and non-linear temporal dynamics, posing significant challenges for accurate spatio-temporal forecasting.
- Existing methods have developed increasingly complex graph, attention, and decomposition architectures, while the influence of the underlying non-linear function approximators has received relatively little attention.
- In this work, we propose STKAN, a spatio-temporal forecasting architecture that introduces Taylor-polynomial Kolmogorov-Arnold Network modules into spatio-temporal token mixing.
- EN Highlights:
- arXiv:2607.13108v1 Announce Type: new
- Abstract: Real-world traffic data exhibit heterogeneous spatial correlations and nonlinear temporal dynamics, posing substantial challenges for accurate spatio-…
- Existing approaches have developed increasingly sophisticated graph, attention, and decomposition architectures, while the influence of the underlying nonlinear…
- In this work, we propose STKAN, a spatio-temporal forecasting architecture that introduces Taylor-polynomial Kolmogorov–Arnold Network modules into spatial and…
A Hybrid Mamba for Audio-Visual Navigation
- Release Time: 2026-07-16 12:00 Beijing Time
- Abstract: - arXiv:2607.13110v1 Announcement Type: new.
- Abstract: Since the paradigm centered on convolutional neural networks and recurrent architectures was established in 2020, the fundamental backbone networks for audio-visual navigation have not fundamentally changed in over five years, making them insufficient to support the effective representation of dynamic multi-modal sequences.
- This paper proposes Samba (a hybrid Mamba for audio-visual navigation).
- It replaces the traditional GRU for temporal aggregation with a Mamba State Encoder (M-SE) that supports adaptive selection, and builds an Audio Mamba Encoder (AME) to compensate for the limitations of convolutional operators in capturing global time-frequency dependencies in spectrograms.
- EN Highlights:
- arXiv:2607.13110v1 Announce Type: new
- Abstract: Since the paradigm centered on convolutional neural networks and recurrent architectures was established in 2020, the fundamental backbone networks fo…
This paper proposes Samba(A Hybrid Mamba for Audio-Visual Navigation)
It uses the adaptive selection-enabled Mamba State Encoder (M-SE) to replace conventional GRUs for temporal aggregation, and constructs an Audio Mamba Encoder (…