🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-07-28
- 类型
- ai-daily
- 字数
- 9012
- 阅读时长
- 43 min
2026-07-28 AI Daily | Kimi K3 Open Weights Reach 2.8 Trillion Parameters, Agents Begin Shift from Usable to Auditable Link to heading
Today’s two main highlights: Moonshot’s release of Kimi K3’s open weights, with 2.8 trillion parameters, elevates both the capabilities and the deployment challenges for open-source models. Concurrently, the focus for AI Agents is shifting from demos to production, with the industry prioritizing auditing, traceability, compliance, and evaluations based on real-world tasks over raw scores.
📖 In-Depth Guide from This Issue’s Watch List Link to heading
The most critical trend to watch today is the shift of “Agents from usable to archivable, auditable, and compliant.” FlowEvo compiles successful workflows into reusable skills, LoRA directly bakes document knowledge into model parameters, and Humanly ensures the human-machine co-writing process is traceable. Coupled with evaluations for fidelity and copyright compliance in document-to-podcast generation, this indicates that implementation challenges have moved from demos to production. The second trend is the evolution of evaluation paradigms, shifting from static accuracy to preference, consensus, and operational consistency. LLM benchmarks, MeSH retrieval, and wildfire risk assessments all underscore that the alignment of metrics with real-world tasks is more crucial than the scores themselves. Finally, multimodal security continues to gain importance, with VLM jailbreaking and cross-modal protection demanding close attention from security teams.
🌐 Trending AI News from the X Platform Link to heading
Topic 1: ChatGPT Work Surpasses Codex in Active Users Amid Rapid Growth Link to heading
- Category: AI · News
- Overview: Trending for: 2 days ago, Related posts: 7,400
- What happened: The number of active users for ChatGPT Work has surpassed that of Codex amid rapid growth.
- Why it matters: This reflects stronger user adoption of AI products designed for workplace scenarios, indicating that the value of AI tools is expanding from code generation to broader office and collaboration use cases.
- Discussion summary: Discussions on X are focused on whether the growth of ChatGPT Work signals stronger enterprise demand, if Codex is becoming marginalized, and whether this shift heralds the transition of AI products from developer tools to general-purpose productivity platforms.
Topic 2: Neurosurgeon Claims Muscle Building Shortens Life, Sparks Fitness Debate Link to heading
- Category: AI · News
- Overview: Trending for: N/A, Related posts: 33
- What happened: A neurosurgeon claimed that intentional muscle building could shorten one’s lifespan, sparking a debate on X about fitness, muscle, and health risks.
- Why it matters: Such controversies touch upon the dissemination of medical knowledge and the credibility of health advice. When generating, retrieving, and recommending health content, AI needs to more accurately distinguish between opinions, evidence, and misinformation.
- Discussion summary: The debate centers on whether muscle building is genuinely harmful, the sufficiency of the evidence, the impact of different training methods on lifespan and metabolic health, and whether the doctor’s views were taken out of context.
Topic 3: Coffee Shop Guy Sips in Peace, No Screens in Sight Link to heading
- Category: AI · News
- Overview: Trending for: N/A, Related posts: 50
- What happened: A picture of a man quietly sipping coffee in a cafe without using any screens has drawn attention on X.
- Why it matters: This type of content is relevant to the AI field because it reflects public contemplation on the over-intrusion of digital devices, algorithmic recommendations, and smart interfaces into daily life. It also highlights a demand for more restrained and less intrusive technological experiences.
- Discussion summary: Discussions on X focus on three points: some see it as an ideal example of a “digital detox” or slow living; others believe it’s just a glamorized moment that doesn’t reflect daily reality; and some use it to discuss whether AI and mobile devices are making it harder for people to find focus and tranquility.
Topic 4: Moonshot AI Releases Kimi K3 Open Weights with 2.8 Trillion Parameters Link to heading
- Category: AI · News
- Overview: Trending for: 21 hours ago, Related posts: 27,000
- What happened: Moonshot AI has released the open-weights version of Kimi K3, uploading approximately 1.4TB of model files to Hugging Face under a permissive license.
- Why it matters: This marks the entry of ultra-large-scale frontier models into the market as open-weights, which could drive changes in the open-source ecosystem, model accessibility, and the US-China AI competition landscape. It also brings more attention to the issues of large model capability proliferation and deployment barriers.
- Discussion summary: Discussions on X center on three key points: first, Kimi K3 sets a new scale for open-weights models with its 2.8 trillion total parameters, million-token context, and multimodal capabilities. Second, whether its performance on coding and long-context tasks is sufficient to rival closed-source frontier models. Third, the strategic advantages, compliance, and security risks of open-weighting, as well as the practical infrastructure requirements for a 1.4TB model.
Topic 5: NVIDIA Launches Open Secure AI Alliance After Major Breach Link to heading
- Category: AI · News
- Overview: Trending for: 14 hours ago, Related posts: 19,000
- What it is:Following a major security incident, NVIDIA announced the launch of an AI alliance focused on openness and security.
- Why it’s important:This is significant for the AI field because it elevates the security of models, infrastructure, and the supply chain to the level of industry-wide collaboration, potentially influencing future security standards and compliance directions for AI products.
- Discussion summary:The discussion on X is mainly focused on whether the alliance can genuinely improve AI security, if it’s merely a public relations response to the crisis, and how to balance open collaboration with closed security.
Topic 6:Why No $2000 AI Subscriptions for Power Users? Link to heading
- Category:AI · News
- Summary:Hotness time:18 hours ago,Related posts:1600
- What it is:A discussion on X questions why major AI companies have not yet launched high-end subscription services, priced at around $2000 per month, for power users.
- Why it’s important:This relates to the commercialization path of AI products, the recovery of computing costs, and the willingness of professional users to pay for higher performance, larger quotas, and more stable services.
- Discussion summary:The discussion focuses on whether there is real market demand for high-priced subscriptions. Supporters believe developers, enterprises, and content creators are willing to pay for more powerful models and higher limits. Skeptics argue the price is too high, existing enterprise APIs already meet the needs, and it could increase the usage barrier and resource divide for AI tools.
Topic 7:Caitlin Clark Sets WNBA Record with 38.3 Points-plus-Assists Average Link to heading
- Category:AI · Sports
- Summary:Hotness time:,Related posts:328
- What it is:WNBA player Caitlin Clark set a new record with an average of 38.3 points plus assists, sparking heated discussion on X.
- Why it’s important:While the event itself is in the sports domain, its high popularity reflects the influence of sports data analysis, real-time content recommendation, and AI-driven event broadcasting in public discourse.
- Discussion summary:The discussion on X primarily centers on Clark’s historical standing, the significance of her stats, the boost in attention for the WNBA, and whether she should be considered the league’s most influential player right now.
Topic 8:Vukic Crushes Svajda in DC Open Upset Link to heading
- Category:AI · Sports
- Summary:Hotness time:,Related posts:55
- What it is:In the DC Open, Vukic scored an upset victory over Svajda, becoming one of the day’s surprising results.
- Why it’s important:The significance of such sports hotspots for the AI field lies in their ability to demonstrate a model’s capacity for rapid understanding and summarization of real-time events, unexpected outcomes, and public opinion shifts. They can also be used for training in sports news generation and trend analysis.
- Discussion summary:The discussion on X mainly focuses on whether this upset was unexpected, whether Vukic was in better form, and the reasons for Svajda’s defeat. Some discussions also compare the players’ rankings, pre-match predictions, and key points in the match.
Topic 9:LeBron James Signs with 76ers for Two-Year Deal at Age 41 Link to heading
- Category:AI · Sports
- Summary:Hotness time:1 day ago,Related posts:172000
- What it is:A trending topic on X claims LeBron James is joining the Philadelphia 76ers on a two-year contract, with related discussions escalating quickly.
- Why it’s important:This type of high-interest, breaking sports news serves as a classic test case for AI’s capabilities in news comprehension, rumor detection, topic clustering, and real-time summarization. It also reflects a model’s ability to process rapidly spreading content on social media.
- Discussion summary:The discussion mainly centers on the authenticity of the news, the potential impact of LeBron’s arrival on the 76ers’ championship prospects, whether his age and physical condition can still support high-level contributions, and speculative discussions and fan banter about his “best next team.”
Topic 10:Lando Norris Wins Hungarian Grand Prix in McLaren Masterclass Link to heading
- Category:AI · Sports
- Summary:Hotness time:1 day ago,Related posts:378000
- What it is:Lando Norris won the Hungarian Grand Prix, with McLaren delivering a dominant race performance.
- Why it’s important:Such high-interest sports events are important to the AI field as they are classic scenarios for testing a model’s ability in real-time topic identification, cross-domain summarization, and public opinion aggregation.
- Discussion summary:The discussion on X is mainly focused on Norris’s individual performance, McLaren’s overall competitiveness and tactical execution. There are also debates on whether the victory was more due to the driver’s skill or the team’s strategy, and the reasons for other teams’ poor performance.
Topic 11:Arsenal Eye Vinicius Junior After Premier League Title Win Link to heading
- Category:AI · Sports
- Summary:Hotness time:2 days ago,Related posts:607000
- What it is:Reports suggest that Arsenal is interested in signing Real Madrid forward Vinícius Júnior after their Premier League title win.
- Why it’s important:High-interest sports transfer rumors like this are often used for AI news summarization, public opinion analysis, and content generation. They also test an AI’s ability to identify and de-bias unverified information.
- Discussion Overview: Discussions on X are mainly focused on whether the rumor is credible, whether Arsenal can afford the transfer and salary costs, whether Vinícius Júnior will leave Real Madrid, and how the deal, if it happens, would change the competitive landscape of the Premier League. Some also believe it’s just hype.
Topic 12: Mizkif Runs YouTube Ads on Asmongold’s Videos to Promote Defamation Video Link to heading
- Category: AI · Entertainment
- Overview: Hot Since:, Related Posts: 124
- What Happened: Twitch/YouTube streamer Mizkif is accused of running YouTube ads before Asmongold’s videos to promote a video involving defamation allegations.
- Why It Matters: While not core AI technology news, this incident involves platform ad placement, content recommendation, and reputational disputes, reflecting how algorithmic distribution and advertising systems can be used to amplify creator conflicts.
- Discussion Overview: The discussion on X focuses on whether Mizkif’s actions constitute malicious marketing or harassment, whether the video content constitutes defamation, and whether YouTube should strengthen its review of similar targeted placements and the promotion of controversial content.
Topic 13: UK Takeaway Order Delivers Handful of Limp Fries Instead of Loaded Bowl Link to heading
- Category: AI · Entertainment
- Overview: Hot Since:, Related Posts: 27
- What Happened: A customer in the UK received a handful of limp fries instead of the expected loaded bowl they ordered, sparking discussion on social media.
- Why It Matters: Such delivery fails are noteworthy because they reflect reliability issues in platform delivery, order recognition, and automated customer service. If more of these processes involve AI in the future, similar errors will directly impact user experience and trust.
- Discussion Overview: Discussions on X are mainly about whether this was the merchant’s mistake, a platform error, or a delivery mishap. Some also joked about the stark contrast between the food received and the order photo, using it as an opportunity to complain about inconsistent takeaway quality.
Topic 14: Once Human Players Share Stunning Post-Apocalyptic Photos for Second Anniversary Link to heading
- Category: AI · Entertainment
- Overview: Hot Since:, Related Posts: 48
- What Happened: On the second anniversary of “Once Human,” players shared numerous beautiful post-apocalyptic screenshots and photo-mode works on X.
- Why It Matters: This topic shows how high-quality visual generation, photo modes, and user-created ecosystems in games of the AI era can amplify community engagement. It also reflects how AI-driven content production is impacting game marketing and the barrier to creation.
- Discussion Overview: Discussions on X focus on the visual quality of the screenshots, the post-apocalyptic art style, and whether the second-anniversary content is sufficient. There is also discussion about whether these works were created using AI-assisted tools or were simply captured in-game.
Topic 15: Tesla Robotaxis Go Silent on Turn Signals and Hazards Link to heading
- Category: AI · News
- Overview: Hot Since:, Related Posts: 233
- What Happened: A discussion on X has gained traction regarding Tesla Robotaxis allegedly failing to use turn signals and hazard lights correctly while driving, drawing attention to their on-road performance.
- Why It Matters: This issue pertains to the autonomous driving system’s compliance with basic traffic regulations and its safety credibility. It also affects public perception of the commercial maturity and regulatory acceptability of Robotaxis.
- Discussion Overview: The debate centers on whether this is a systemic flaw or an isolated situational error. Supporters argue that it is still in the testing and iteration phase, while opponents believe that the inability to handle basic signals indicates that autonomous driving is still far from being truly ready for use.
AI Public Opinion Summary on X Today Link to heading
The main narrative about AI on X today has clearly shifted from “which model is superior” to “how AI is becoming productized, platformized, and infrastructural.” The growth of ChatGPT Work, the open-sourcing of Kimi K3’s weights, and discussions about high-priced subscriptions all indicate that users and the market are more concerned with whether AI can truly integrate into office, collaboration, and heavy production workflows. A relative consensus is emerging that the value boundary of AI is expanding from developer tools to general productivity platforms, and that open models, long context, and greater computing power will continue to drive ecosystem proliferation. Disagreements are focused on two points: first, whether open-sourcing weights is a strategic advantage or a security and compliance risk; and second, whether high-end subscriptions can create a real paid market or will only exacerbate resource disparity and barriers to entry. The potential risks have therefore become more prominent, including security and abuse issues from the rapid spread of model capabilities, whether the AI alliance will devolve into a PR crisis management tool, and the real-world consequences of misjudgment or misinformation in autonomous driving and health-related content.
💡 Influencer Insights Link to heading
Okay, based on the summary of tweets from various AI influencers over the past 24 hours that you provided, I have compiled the following analysis report for you:
Daily Observations from AI Influencers: Technology, Trends, and Resource Insights Link to heading
1. Commonly Watched Technology Trends Link to heading
Anthropic’s full suite of models has become the absolute focus: As
@claudeaimade a series of announcements, discussions among prominent figures highly concentrated on Anthropic’s new models and pricing strategy.Claude Opus 5 Release:
@doteyprovided a detailed analysis of the Opus 5 release, emphasizing its positioning as “offering near cutting-edge intelligence at half the price of Fable 5,” and performing remarkably well in key benchmarks across coding, automation, and more, making it especially suited for Agents handling multiple complex tasks simultaneously.Fable 5 Access Democratized:
@zhixianiohighlighted that Claude Fable 5 will be included in premium subscription plans such as Max and Team starting July 20th, marking a rapid popularization of Anthropic’s most powerful model to a wider range of paying users.Open-source Model ‘Size Race’ vs. Practicality Debate:
- Kimi K3 Release Shakes the Open-Source World: Both
@doteyand@ruanyfquickly noted Moonshot AI’s open-sourcing of the 2.8T parameter MoE model, Kimi K3.@ruanyfaffirmed its performance through personal testing, stating it “is truly close to Fable 5,” and pointing out that the surge in parameters is the primary reason for its leap in capability, while also noting that its API pricing (20 RMB / 100 RMB per million tokens domestically) is already the most expensive in China. - The Surprising Performance and Limitations of Small Models:
@zhixianiothoroughly tested MiniCPM-o 4.5 (9B)’s audio and video full-duplex capabilities, frankly stating “it’s hard to imagine this is the effect a 9B model can achieve.” However, his practical coding tests with Gemma 4 12B Coder showed that when handling “long, stateful, one-shot” complex programs, the “ceiling” of 12B parameters is still evident, not as good as his daily-used Qwen 35B MoE.
- Kimi K3 Release Shakes the Open-Source World: Both
On-device Models Emerge:
@zhixianiohas become the standard-bearer in this area, with his tweets indicating that on-device models are no longer just a concept. He not only officially discussed this topic on the podcast “Cognition Has Bounds” but also actively tested Google’s Gemma 4 QAT (Quantization-Aware Training) model series (including E4B and 12B Coder), praising their performance on specific tasks and Google’s emphasis on on-device solutions. He even envisions “model cartridges” as a highly forward-looking future interaction paradigm.AI Video and Automation Tools Stepping into a ‘Blue Ocean’:
@Pluvio9yte’s tweet thread entirely revolved around “AI Video + Self-media Monetization,” positioning this as a “blue ocean.” He open-sourced 55 AI video skills and deeply integrated tools like Codex, Hyperframes, HeyGen, building a fully automated workflow from content generation to digital human cloning. This is a vivid microcosm of current AI application layer exploration.
2. Noteworthy Unique Perspectives Link to heading
The ‘Paradox’ of Model Collaboration and Best Practices:
@doteyraised sharp doubts about the then-popular/advisorpattern (where weaker models design, and stronger models act as consultants). He believes this approach is logically unsustainable and proposed his own best practice: “Let the smart models design, the weaker models execute, and finally the smart models perform acceptance testing,” which he deems the most cost-effective combination.The ‘Creativity Trap’ of Structured Output:
@vista8shared the conclusion of an important paper: requiring large models to output in JSON or XML format significantly reduces the diversity of their responses. In tests, when asked to “say any word,” the probability of outputting “serendipity” in JSON format was as high as 64%. This reminds developers that in pursuing the convenience of structured data, they might inadvertently stifle the model’s creativity.AI is ‘Tenfold Efficiency,’ Not ‘Four-Day Work Week’:
@ruanyfparaphrased an article titled “Can I Take a Holiday Today?”, raising a profound societal question: When AI doubles white-collar work efficiency, allowing a week’s work to be completed in a few hours, should employees enjoy more holidays? He believes that beyond individual skill improvement, the overall societal productivity increase brought by AI should ultimately manifest in increased average salaries or benefits, rather than no change.The ‘Sense of Play’ and ‘Gamification’ in the AI Era:
@nishuangprecisely distinguished between “sense of play” and “gamification” in product design by comparingCapWordsand多邻国. He believes the former is the sense of play that stimulates curiosity and endorphins, while the latter is gamification that drives arduous tasks through dopamine rewards. In an era of increasingly shrewd users, designs that bring genuine intrinsic joy through a “sense of play” will surpass simple point reward systems.The Strategic “Convergence” of Claude Code: According to data cited by
@Pluvio9yte, Anthropic has cut 80% of Claude Code’s system prompts for Opus 5. This move not only suggests a deep and highly efficient integration with its own most powerful model but could also signal a reduction in compatibility with other models to achieve peak performance and interoperability.“Free Puppies” and the Burden of Open-Source Maintenance:
@ruanyfshared the philosophy of SQLite creator Richard Hipp on rejecting external PRs. He likens each PR to a “free puppy,” implying that accepting a contribution means shouldering the moral responsibility for its maintenance, documentation, and testing for the next 25 years or more. This serves as a sober reflection on sustainability for the booming open-source AI community.
3. Recommended Tools & Resources Link to heading
AI Audio and Video Creation
- Suno:
@vista8recommends a music generation workflow: hum and upload a melody, then let the AI expand it into a full song, which can help bypass copyright restrictions. - ChatCut: After comparative testing by
@Pluvio9yteand@teach_fireworks, it was found to be superior to video-use in finely-tuned motion graphics effects and rhythmic variations, making it more suitable for production. - HeyGen + Hyperframes + Codex Workflow:
@Pluvio9yteoffers a fully open-source workflow for low-cost, zero-entry monetization in content creation, particularly through digital human and voice cloning technologies.
- Suno:
Development & Productivity Tools
- Bento PPT / Skill:
@vista8developed an HTML PPT Skill (npx skills add joeseesun/qiaomu-bento-ppt) based on this, which can one-click generate collaborative and dynamic web-based PPTs from content. - BaoCut (v0.8.2): The personally iterated video screen translation feature by
@doteyis now live. It supports automated operations via an Agent and allows exporting results to Jianying for secondary editing. - Mole: A Mac cleaning and optimization tool recommended by
@gkxspace. Its CLI can be installed for free viabrew install mole. - eSIM Wallet:
@AI_Jasonyurecommends this free mobile card management software, suitable for users who manage multiple eSIM cards.
- Bento PPT / Skill:
Content Creation & Monetization Tools
- Xiaohongshu Viral Cover Skill: Open-sourced by
@pyang1235005and highly recommended by prominent figures like@AI_Jasonyu. It uses an Agent to guide users through rounds of prompts, offering 10 composition styles to generate covers, significantly boosting click-through rates. - PayPal China:
@gefei55shared an “information gap,” revealing that it now supports registration with domestic personal IDs and can be integrated into websites to send and receive USD payments from global users. A complete, compliant path from registration to withdrawal was also shared. - GEO Knowledge Base:
@yaojingangreleased a collection of materials from their GEO open course, including a data warehouse, tools, and Skills, suitable for users providing SEO services to international businesses.
- Xiaohongshu Viral Cover Skill: Open-sourced by
AI Frontiers & Learning
- Cutting-Edge AI Papers Selection (Newsletter): A highly recommended weekly resource from
@vista8and one of the best ways to keep pace with AI model development. - “AI Also Has a Subconscious”: An article recommended by
@vista8covering Anthropic’s latest research, which explores the deeper internal states of models beyond their outputs. - Hardcore AI Concepts Series: Educational videos produced by
@vista8on topics like RLHF and Multi-Head Attention, suitable for foundational knowledge in deep learning.
- Cutting-Edge AI Papers Selection (Newsletter): A highly recommended weekly resource from
📚 Appendix: Today’s Watch List Source Updates Link to heading
Timeframe: Last 3 days; 22 sources covered; 33 updates in total
Y Combinator Podcast (B_intro+search) Link to heading
- Jensen Huang: The Mindset That Built NVIDIA
- Published: 2026-07-28 05:29 Beijing Time
- Summary: - You may have already heard of OpenClaw (formerly known as Clawdbot/Moltbot).
- The sensational open-source AI assistant that runs on your own device, connects with the messaging apps you already use, and goes beyond chat to actually perform tasks like managing email, calendars, files, workflows, and more.
- Now meet the person behind it.
- YC’s Raphael Schaad sat down with Peter Steinberger, founder of OpenClaw, to discuss the “aha” moment behind the viral personal AI agent, why local-first agents could replace many of today’s apps, and how personal agents will reshape the future of software.
- EN Points:
- NVIDIA started with the wrong technology, learned the right one from three textbooks bought at Fry’s, and went on to invent most of the major breakthroughs in m…
- EN Points:
Stratechery by Ben Thompson (A_full) Link to heading
- Vacation: Week of July 27
- Published: 2026-07-27 18:00 Beijing Time
- Summary: - Stratechery is on vacation the week of July 27.
- There will be no Weekly Article or Updates.
- The next Update will be on Monday, August 3.
- Sharp Tech and Greatest of All Talk will also return the week of August 3.
- EN Points:
- Stratechery is on vacation the week of July 27
- There will be no Weekly Article or Updates
- The next Update will be on Monday, August 3
- Sharp Tech , and Greatest of All Talk will also return the week of August 3
OpenAI Blog (A_full) Link to heading
- How AI is expanding what people do at work
- Published: 2026-07-27 11:30 Beijing Time
- Summary: - AI is changing the work people do.
- An analysis of over 800,000 messages from the US
- ChatGPT users, our new research shows that 16.8% of work-related messages and 43.5% of occupation-specific messages relate to tasks associated with another occupation.
- Small business owners can independently draft copy, review contracts, or conduct basic financial analysis.
- Salespeople can use AI to explore customer datasets that were once handed off to analysts.
- EN Points:
- New OpenAI research shows how AI is expanding what workers do, with ChatGPT users taking on tasks across roles and reshaping job boundaries.
ArXiv cs.AI (B_intro+search) Link to heading
FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills
- Published: 2026-07-27 12:00 Beijing Time
- Summary: - arXiv:2607.21596v1 Announce Type: new.
- Abstract: Large language model agents are increasingly solving complex tasks by constructing inference-time workflows that combine reasoning, tool use, and code execution.
- While such workflows enable flexible problem-solving, useful processes discovered during execution are often ephemeral: they help solve the current task but are not preserved in a form that can systematically benefit future tasks.
- We introduce FlowEvo, a training-free framework that compiles successful traces into reusable skill records.
- EN Points:
- arXiv:2607.21596v1 Announce Type: new
Abstract: Large language model agents increasingly solve complex tasks by constructing inference-time workflows that combine reasoning, tool use, and code execu…
While such workflows enable flexible problem solving, the useful procedures discovered during execution are often transient: they help solve the current task bu…
We present FlowEvo, a training-free framework that compiles successful traces into reusable skill records
Risk Is Not the Target: A Monotonic Framework for Evaluating Wildfire Operational Risk Signals
- Published: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21597v1 Announce Type: new.
- Abstract: Evaluating wildfire risk systems using standard machine-learning metrics such as F1-score or IoU is fundamentally flawed: these metrics assess event prediction accuracy, not the operational consistency of a continuous risk signal.
- This work proposes a novel monotonic evaluation framework that measures whether increases in a predicted risk score consistently correspond to increases in observed operational load, such as the number of fires, intervention times, and deployed resources.
- Moreover, we compare three structurally different approaches on the French Alpes-Maritimes department: the expert-based DFE index, GRU-based predictive models, and FARS (a hybrid multi-agent system combining predictive AI with LLM-based reasoning).
- EN Highlights:
- arXiv:2607.21597v1 Announce Type: new
- Abstract: Evaluating wildfire risk systems using standard machine-learning metrics such as F1-score or IoU is fundamentally flawed: these metrics assess event p…
- This work proposes a novel monotonic evaluation framework that measures whether increases in a predicted risk score consistently correspond to increases in obse…
- Moreover, we compare three structurally different approaches on the French Alpes-Maritimes department: the expert-based DFE index, GRU- based predictive models,…
Securing Multimodal AI through Internal Information Decomposition
- Published: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21600v1 Announce Type: new.
- Abstract: Multimodal large language models introduce attack surfaces not present in unimodal systems: adversaries can distribute malicious intent across modalities to evade unimodal safeguards.
- This motivates the use of cross-modal consistency as a detection signal, rather than inspecting each modality in isolation.
- Our key observation is that benign inputs induce compatible predictive behaviors from text-only and vision-only reasoning, which stabilize upon fusion, whereas adversarial manipulations disrupt this agreement, leading to anomalous multimodal behavior.
- EN Highlights:
- arXiv:2607.21600v1 Announce Type: new
Abstract: Multimodal large language models introduce attack surfaces absent in unimodal systems: adversaries can distribute malicious intent across modalities t…
This motivates using cross-modal consistency as a detection signal rather than inspecting each modality in isolation
Our key observation is that benign inputs induce compatible predictive behavior from text-only and vision-only reasoning that stabilizes when fused, whereas adv…
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21601v1 Announcement Type: New.
- Abstract: Public-space gesture interaction is often evaluated as a frame-level recognition problem, but deployed systems expose different failure boundaries.
- In scenic kiosks, exhibition halls, and service terminals, the user’s experience is whether an intentional action becomes a stable interaction event, not whether a single hand bounding box is correct.
- We call this the gap between cognition and interaction.
- EN Highlights:
- arXiv:2607.21601v1 Announce Type: new
- Abstract: Public-space gesture interaction is often evaluated as a frame-level recognition problem, but deployed systems expose a different failure boundary
- In scenic kiosks, exhibition halls, and service terminals, users experience whether an intended action becomes a stable interaction event, not whether individua…
- We call this the recognition-to-interaction gap
Transferable Latency Prediction for Fast LLM Screening on Heterogeneous Edge Devices
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21602v1 Announcement Type: New.
- Abstract: Accurate latency prediction is crucial for deploying Large Language Models (LLMs) on heterogeneous edge devices, where inference latency is affected by model architecture, prompt behavior, runtime backend, hardware utilization, dynamic voltage and frequency scaling (DVFS), and thermal variations.
- This paper proposes a runtime-aware latency prediction framework for deployment-oriented LLM selection.
- The framework represents each inference request as a hardware-runtime-model-prompt configuration, divides inference into prefill and decoding stages, and adaptively fuses static descriptors with dynamic hardware telemetry through a gated prediction model.
- EN Highlights:
- arXiv:2607.21602v1 Announce Type: new
- Abstract: Accurate latency prediction is critical for deploying large language models (LLMs) on heterogeneous edge devices, where inference latency is affected…
- This paper presents a runtime-aware latency prediction framework for deployment-oriented LLM selection
The framework represents each inference request as a hardware-runtime-model-prompt configuration, separates inference into prefill and decode phases, and adaptively adjusts the KV cache based on the prompt’s structure.
AgentKVShift: Efficient KV Cache Reuse for Agentic Memory Systems
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21604v1 Announcement Type: New.
- Abstract: Memory-augmented LLM agents maintain context across hundreds of interactions through agentic memory systems that actively manage retrieved content using LLM-generated metadata (such as summaries, keywords, and tags).
- From an inference cost perspective, each retrieval triggers a full re-encoding of these structured memory units into Key-Value (KV) states, which determines the prefill latency.
- Existing training-free KV reuse methods mitigate this issue by selectively recomputing a small fraction of tokens, but they are designed for RAG-style raw passages and degrade the performance of structured agentic memory.
- EN Highlights:
- arXiv:2607.21604v1 Announce Type: new
- Abstract: Memory-augmented LLM agents maintain context across hundreds of interactions through agentic memory systems that actively curate retrieved content wit…
- From an inference cost standpoint, every retrieval triggers a full re-encoding of these structured memory units into Key-Value (KV) states, which dominates pref…
- Existing training-free KV reuse methods mitigate this by selectively recomputing a small fraction of tokens, but were designed for RAG-style raw passages and de…
TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21606v1 Announcement Type: New.
- Abstract: Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that can modify the sampling trajectory to generate images more faithful to complex compositional prompts.
- We propose TILT, a training-free framework for compositional text-to-image generation through test-time reward alignment.
- We interpret compositional failures as overlapping modes between the joint distribution and single-concept distributions, and define a reward that favors samples where all concepts co-exist.
- EN Highlights:
- arXiv:2607.21606v1 Announce Type: new
- Abstract: Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling…
- We present TILT, a training-free framework for compositional text-to-image generation via test-time reward alignment
- We interpret compositional failures as overlap modes between joint and single-concept distributions, and define a reward that favors samples where all concepts…
Spectral Flow Certificates for Depth-Aware Long-Range Propagation in Graph Neural Networks
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21607v1 Announcement Type: new.
- Abstract: Graph Neural Networks propagate information through local message passing, but the graph topology itself can silently prevent any amount of training from solving long-range tasks.
- When we deploy GNNs on new graphs, there is currently no inexpensive way to know, before training begins, whether the graph’s structure will allow information to be transmitted sufficiently far between distant nodes.
- We address this gap by proposing Spectral Flow Certificates (SFCs), a single scalar computed from the graph’s normalized Laplacian in seconds, requiring no model training or labeled data.
- EN Highlights:
- arXiv:2607.21607v1 Announce Type: new
- Abstract: Graph Neural Networks propagate information through local message passing, but the graph topologies themselves can silently prevent any amount of trai…
- When we deploy GNNs on new graphs, there is currently no inexpensive way to know, before training begins, whether the graphs’ structures will allow information…
- We address this gap by proposing Spectral Flow Certificates (SFCs), single scalars computed from the graphs’ normalised Laplacians in seconds, requiring no mode…
Coupled Hierarchical Search over Topology and Execution for Agentic Workflow Synthesis
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21609v1 Announcement Type: new.
- Abstract: Although structured workflows empower Large Language Models (LLMs) to solve complex problems, the automation of their creation is severely hindered by a vast combinatorial search space, often leading to inflexible and resource-intensive offline training dependencies.
- To address this, we conceptualize workflow generation as an intertwined topology-and-execution search paradigm, where the broader topological layer dictates sub-task boundaries, while lower-level execution results actively reshape the topology itself.
- Building on this foundation, we introduce HierFlow, a training-free, test-time hierarchical search architecture that automates agentic workflow design by merging feedback-guided topological adjustments with a rapid, MCTS-inspired tree search for sub-workflow optimization.
- EN Highlights:
- arXiv:2607.21609v1 Announce Type: new
- Abstract: Although structured workflows empower Large Language Models (LLMs) to tackle complex problems, automating their creation is severely hindered by a vas…
- To address this, we conceptualize workflow generation as an intertwined topology-and-execution search paradigm, where the broader topological layer dictates sub…
- Building on this foundation, we introduce HierFlow, a training-free, test-time hierarchical search architecture that automates agentic workflow design by mergin…
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21610v1 Announcement Type: New.
- Abstract: Schema graphs are an upstream bottleneck for schema-based information extraction and knowledge graph construction, yet most extraction systems assume that a schema is already available.
- We introduce SCOPE (Schema Construction and Ontology-induction Pipeline Evaluation), a train-text-only benchmark for corpus-to-schema induction and optional schema fusion from raw text, constructed from 24 public information extraction sources (15 RE and 9 EE), and standardized to an evaluation-only gold schema graph; its core event extraction targets cover event types and intra-event argument roles, with inter-event links reported separately.
- We present SCION (Schema Construction and Induction with Ontology Normalization), an auditable reference pipeline rather than a new extraction architecture; it builds a candidate space from training text and constrains naming, merging, filtering, validation, and conservative fusion to candidate-relative evidence under a strict JSON contract.
- EN Highlights:
- arXiv:2607.21610v1 Announce Type: new
- Abstract: Schema graphs are an upstream bottleneck of schema-grounded information extraction and knowledge graph construction, yet most extraction systems assum…
- We introduce SCOPE (Schema Construction and Ontology-induction Pipeline Evaluation), a train-text-only benchmark for corpus-to-schema induction and optional sch…
- We present SCION (Schema Construction and Induction with Ontology Normalization), an auditable reference pipeline rather than a new extraction architecture; it…
ArXiv cs.CL (B_intro+search) Link to heading
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21619v1 Announcement Type: New.
- Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive performance, but their safety remains vulnerable to jailbreak attacks.
- Existing content-based jailbreaks are often inconsistent and exhibit unsatisfactory performance against rapidly evolving MLLMs, failing to exploit non-content-based vulnerabilities.
- Unlike previous research, we empirically discover a stylistic inconsistency between the understanding and safety capabilities of MLLMs: MLLMs can robustly comprehend content regardless of visual style, yet their defense mechanisms can be easily bypassed by specific stylistic triggers.
- EN Highlights:
- arXiv:2607.21619v1 Announce Type: new
- Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive performance, but their safety alignment remains vulnerable to jailbreak attacks
- Existing content-based jailbreaks are often inconsistent and show unsatisfying performance against the rapidly evolving MLLMs, failing to exploit non-content-ba…
Unlike previous research, we empirically find that MLLMs exhibit a Stylistic Inconsistency between their comprehension ability and safety ability: MLLMs can rob…
A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21632v1 Announcement Type: new.
- Abstract: Traditional benchmarks for LLMs primarily rely on static datasets and objective scoring metrics, which often fail to capture differences in response quality when multiple answers are acceptable.
- In this context, correctness alone is insufficient to distinguish between responses that vary in clarity, completeness, and usefulness.
- This paper introduces a consensus-based evaluation framework that measures the relative preference between model-generated responses, rather than absolute correctness.
- EN Highlights:
- arXiv:2607.21632v1 Announce Type: new
- Abstract: Traditional benchmarks for LLMs primarily rely on static datasets and objective scoring metrics, which often fail to capture differences in response q…
- In such settings, correctness alone is insufficient to distinguish between responses that vary in clarity, completeness, and usefulness
- This paper introduces a consensus-based evaluation framework that measures relative preference among model-generated responses rather than absolute correctness
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21685v1 Announcement Type: new.
- Abstract: Systematic reviews begin by reading thousands of abstracts to identify a few relevant ones, then using classifiers to prioritize the reading.
- Their inputs are often enhanced with Medical Subject Headings (MeSH), which are assigned either by expert indexers weeks or months after publication, or immediately by automated tools.
- To our knowledge, these two have not been directly compared as classifier features, and previous work has not investigated whether the outcome of the comparison depends on how the classifier is evaluated.
- EN Highlights:
- arXiv:2607.21685v1 Announce Type: new
- Abstract: A systematic review begins with someone reading thousands of abstracts to identify the few that are relevant, and classifiers are used to prioritise t…
- Their inputs are often augmented with Medical Subject Headings (MeSH), assigned either by expert indexers weeks or months after publication or by automatic tool…
To our knowledge the two have not been compared directly as classifier features, and no previous work has asked whether that comparison’s outcome depends on how…
Humanly: A Configurable and Traceable Environment for Human-AI Collaborative Writing
- Publish Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21758v1 Announcement Type: new.
- Abstract: Teachers, conference chairs, and public readers all judge writing from limited evidence, seeing only a finished document and not the process that generated it.
- Final text alone cannot reveal whether a document was produced through human typing, AI generation, or mixed human-AI collaboration.
- Existing process-tracking tools help, but many are tied to host-document histories, provide coarse activity records, and offer limited control over the writing environment.
- EN Key Points:
- arXiv:2607.21758v1 Announce Type: new
- Abstract: Teachers, conference chairs, and public readers all judge writing from limited evidence, seeing only a finished document and not the process that prod…
- Final text alone cannot reveal whether a document was produced through human typing, AI generation, or mixed human-AI collaboration
- Existing process-tracking tools help, but many are tied to host-document histories, provide coarse activity records, and offer limited control over the writing…
Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders
- Publish Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21774v1 Announcement Type: new.
- Abstract: Large language models may infer demographic attributes from subtle linguistic cues even when those attributes are not explicitly stated.
- This pilot study examines whether Qwen2.5-7B-Instruct internally represents Colombian identity, socioeconomic status, or stereotype-related information when processing Colombian-Spanish and English prompts.
- We use Natural Language Autoencoders (NLA) to describe residual-stream activations from layer 20 across four positional quartiles per prompt.
- EN Key Points:
- arXiv:2607.21774v1 Announce Type: new
- Abstract: Large language models may infer demographic attributes from subtle linguistic cues even when those attributes are not explicitly stated
- This pilot study examines whether Qwen2.5-7B-Instruct internally represents Colombian identity, socioeconomic status, or stereotype-related information when pro…
- We use Natural Language Autoencoders (NLA) to verbalize residual-stream activations from layer 20 across four positional quartiles per prompt
Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21780v1 Announcement Type: New.
- Abstract: Document packets, i.e., multiple documents concatenated into a single file, are common in government and administrative workflows, but splitting them into their constituent documents is difficult, especially for under-resourced languages.
- We introduce Khondo (Bengali for split/segment), the first benchmark for document packet splitting on Bangladeshi government forms.
- Unlike previous English and OCR text-based datasets, Khondo is bilingual (Bangla-English) and vision-native, where models operate directly on page images.
- EN Key Points:
- arXiv:2607.21780v1 Announce Type: new
- Abstract: Document packets, multiple documents concatenated into a single file, are common in government and administrative workflows, yet splitting them into t…
- We introduce Khondo (Bangla for split/segment), the first benchmark for document packet splitting on Bangladeshi government forms
- Unlike prior English and OCR-text-based datasets, Khondo is bilingual (Bangla–English) and vision-native; where models operate directly on page images
Agentic Evaluation of Copyright Law Compliance
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21799v1 Announcement Type: New.
- Abstract: Large language model (LLM) agents are increasingly performing commercial tasks that involve retrieving external content, such as images, and, where appropriate, reproducing that content.
- LLM agents should comply with the law, including copyright law.
- However, we currently lack adequate frameworks to assess whether they do so in practice.
- EN Key Points:
- arXiv:2607.21799v1 Announce Type: new
- Abstract: Large language model (LLM) agents increasingly perform commercial tasks that involve retrieving external content such as images and, where appropriate…
- LLM agents should comply with the law, including copyright law
- Presently, however, we lack adequate frameworks to assess whether they do so in practice
Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21861v1 Announcement Type: New.
- Abstract: We investigate baking documents directly into the weights of a 4-bit Gemma-4-e4b model via LoRA, so the system can answer questions about the corpus closed-book: without retrieval, and without a context window budget.
- In approximately 100 training runs on corpora ranging from a single document to 99 documents, we find that once adapter capacity is sufficient, training data quality becomes the dominant lever for closed-book accuracy, outweighing the combined effects of LoRA rank, learning rate, and two alternative architectures. Capacity itself is a hard gate, below which no data intervention is effective.
A single curation pass (shortening gold answers to canonical 1-6 word spans and dropping trivia) moved closed-book accuracy from 57.7% to 85.7% on a 15-document corpus, more than any architectural change.
- EN Highlights:
- arXiv:2607.21861v1 Announce Type: new
- Abstract: We study baking documents directly into the weights of a 4-bit Gemma-4-e4b model via LoRA, so a system can answer questions about a corpus closed-book…
- Across roughly 100 training runs from single documents to a 99-document corpus, we find that once adapter capacity is adequate, training-data quality is the dom…
- A single curation pass (shortening gold answers to canonical 1-6 word spans and dropping trivia) moved closed-book accuracy from 57.7% to 85.7% on a 15-document…
- EN Highlights:
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21936v1 Announce Type: new.
- Abstract: Historical documents are invaluable knowledge archives but often suffer from illegibility due to physical deterioration and damage.
- While existing restoration methods based on masked language modeling effectively utilize local context, they struggle to restore named entities that require external historical knowledge.
- To address this limitation, we introduce a novel framework for historical document restoration that leverages large language models with retrieval-augmented generation (RAG).
- EN Highlights:
- arXiv:2607.21936v1 Announce Type: new
- Abstract: Historical documents act as invaluable knowledge archives but often suffer from illegibility due to physical deterioration and damage
- While existing restoration methods based on masked language modeling effectively utilize local context, they struggle to restore named entities that require ext…
- To address this limitation, we introduce a novel framework for historical document restoration that leverages large language models with retrieval-augmented gen…
On Improving Faithfulness of Podcasts from Documents
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21961v1 Announce Type: new.
- Abstract: Large Language Models (LLMs) are increasingly used to generate long-form conversational content, such as podcasts from text sources.
- While these systems produce fluent and engaging narratives, they often introduce unsubstantiated information.
- In this work, we present the first systematic study of faithfulness in document-based podcast generation, where grounding must be maintained across conversational turns in a long-form, multi-speaker transcript.
- EN Highlights:
- arXiv:2607.21961v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to generate long-form conversational content such as podcasts from textual sources
While these systems produce fluent and engaging narratives, they often introduce ungrounded information
In this work, we present the first systematic study of faithfulness in document-grounded podcast generation, where grounding must be maintained across conversat…
ArXiv cs.LG (B_intro+search) Link to heading
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21623v1 Announcement Type: New.
- Abstract: We present EaaS, a cloud-native reference architecture that operationalizes AI evaluation methods as six stateless Kubernetes microservices: conformal prediction with finite-sample correction for adaptive prediction sets, calibration evaluation, drift detection via RFF approximation of maximum mean discrepancy, fairness monitoring using bootstrap confidence intervals, a DAG-based pipeline orchestrator, and a results storage API.
- We validate four key methodological concerns.
- First, empirical coverage is consistent with the marginal conformal guarantee across K=50 random calibration/test splits, with mean coverage within 1.4 percentage points of the nominal target.
- EN Highlights:
- arXiv:2607.21623v1 Announce Type: new
- Abstract: We present EaaS, a cloud-native reference architecture that operationalizes AI evaluation methods as six stateless Kubernetes microservices: conformal…
- We validate four key methodological concerns
- First, empirical coverage is consistent with the marginal conformal guarantee across K=50 random calibration/test splits, with mean coverage within 1.4 percenta…
On the Depth Scalability of Logic Gate Networks
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21633v1 Announcement Type: New.
- Abstract: Logic Gate Networks (LGNs) implement computation through compositions of Boolean operations, yet unlike classical Boolean circuits, existing LGNs do not reliably benefit from increased depth.
- We identify two distinct reasons: optimization collapse in deep relaxed LGNs, and a topology-induced limitation that persists even when skip-bias initialization and straight-through estimation stabilize training.
- Therefore, trainability alone is not sufficient; deeper layers must also receive information that supports useful computation.
- EN Highlights:
- arXiv:2607.21633v1 Announce Type: new
- Abstract: Logic Gate Networks (LGNs) implement computation through compositions of Boolean operations, yet unlike classical Boolean circuits, existing LGNs do n…
We identify two distinct causes: optimization collapse in deep relaxed LGNs and a topology-induced limitation that persists even when skip-biased initialization…
Thus, trainability alone is insufficient; deeper layers must also receive information that supports useful computation
MotifRole-Diff: Risk-Optimal Role-Aware Corruption for Masked Molecular Graph Diffusion
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21634v1 Announcement Type: new.
- Abstract: Masked discrete diffusion for molecular graph generation typically applies a uniform corruption schedule to all tokens in a lossless graph-to-sequence representation, implicitly treating structurally heterogeneous molecular components as equally difficult and important to reconstruct.
- However, different molecular graph token roles exhibit substantial variation in denoising difficulty and their influence on the decoded molecule, motivating role-specific corruption strategies.
- We introduce MotifRole-Diff, a role-aware corruption process that allocates masking rates according to empirically measured denoising difficulty and graph-level perturbation impact, while preserving the model architecture, clean sequence space, and lossless molecular graph decoder.
- EN Highlights:
- arXiv:2607.21634v1 Announce Type: new
- Abstract: Masked discrete diffusion for molecular graph generation typically applies a uniform corruption schedule to all tokens in a lossless graph-to-sequence…
- However, different molecular graph token roles exhibit substantial variation in denoising difficulty and their influence on the decoded molecule, motivating rol…
- We introduce MotifRole-Diff, a role-aware corruption process that allocates masking rates according to empirically measured denoising difficulty and graph-level…
Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21635v1 Announcement Type: new.
- Abstract: Personal agents maintain memory, learned skills, tool configurations, and policy states that evolve with each user.
- Existing agent benchmarks typically evaluate these capabilities in isolation: tool benchmarks test invocation under fixed APIs, memory benchmarks test recall or forgetting, and security benchmarks test static policy compliance.
- We argue that personal agent evaluation requires a different protocol: replaying the same temporal interventions under different persistent user-conditioned states and measuring how failures propagate between agent components.
- EN Highlights:
- arXiv:2607.21635v1 Announce Type: new
- Abstract: Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user
- Existing agent benchmarks often evaluate these capabilities in isolation: tool benchmarks test invocation under fixed APIs, memory benchmarks test recall or for…
We argue that personal-agent evaluation requires a different protocol: replaying the same temporal intervention across different persistent user-conditioned sta…
Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21636v1 Announcement Type: New.
- Abstract: The value of synthetic tabular data lies not only in preserving the marginal distribution of each column but also in preserving the dependencies between columns, a structure that carries many discriminative signals for minority classes in imbalanced domains such as fraud and clinical risk.
- However, we show that the metrics most commonly used to certify synthetic tabular data are largely blind to inter-column dependencies: a baseline that models each column independently (thereby destroying all dependencies) is judged indistinguishable from real data by a logistic regression C2ST, and pairwise trend scores are only partially sensitive.
- We introduce a dependency-aware fidelity diagnostic that decomposes a strong classifier two-sample test (XGB-C2ST) into marginal, dependency, and numerical-categorical cross-components, anchored between a worst-case fully-factorized reference (where all dependencies are destroyed) and a best-case true data oracle.
- EN Highlights:
- arXiv:2607.21636v1 Announce Type: new
- Abstract: Synthetic tabular data is valued for preserving not only each column’s marginal distribution but the dependencies between columns – structure that ca…
- Yet the metrics most commonly used to certify synthetic tabular data are, we show, largely blind to inter-column dependency: a baseline that models every column…
- We introduce a dependency-aware fidelity diagnostic that decomposes a strong classifier two-sample test (XGB-C2ST) into marginal, dependency, and numerical-cate…
Quasi-Monte Carlo Initialization for Meta-Reinforcement Learning
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21637v1 Announcement Type: New.
- Abstract: This paper explores the efficacy of quasi-Monte Carlo (QMC) weight initialization for meta-reinforcement learning in modern benchmark environments.
- Various sampling methods are used to constrain a population-based search and aggregate an optimal prior from a set of baseline tasks.
- When extrapolating to similar, unseen continuous control environments, the QMC meta-prior shows improved training convergence compared to modern orthogonal (SB3) defaults.
- EN Highlights:
- arXiv:2607.21637v1 Announce Type: new
- Abstract: This paper explores the efficacy of quasi-Monte Carlo (QMC) weight initialization for meta-reinforcement learning within modern benchmark environments
- Various sampling methods are used to bound a population-based search and aggregate an optimal prior from a baseline set of tasks
The QMC meta-priors show improvements in training convergence compared to modern orthogonal (SB3) defaults when extrapolated to similar unseen continuous contro…
Toward Goal-Agnostic Joint-Embedding Predictive Control of Partial Differential Equations
- Publish Time: 2026-07-27 12:00 Beijing Time
- Abstract:
- arXiv:2607.21644v1 Announce Type: new.
- Abstract: We present a goal-agnostic control framework for partial differential equations (PDEs) built around a joint-embedding predictive architecture (JEPA).
- The small 2D ViT encoder and action-conditioned latent dynamics are trained offline without a reward or downstream goal, frozen, and reused by a Model Predictive Path Integral (MPPI) controller.
- We find that when available, the control objective is better applied to an explicit physical observable (providing injectivity) than to minimizing raw Euclidean distance ($L^2$) in the learned latent space.
Multi-Horizon Consistency as Geometry: When Latent Dynamics Contract, and When They Do Not
- Publish Time: 2026-07-27 12:00 Beijing Time
- Abstract:
- arXiv:2607.21645v1 Announce Type: new.
- Abstract: Multi-horizon latent consistency is a common training knob in video predictors and world models, but practitioners rarely know what it does to transition geometry.
- We treat lambda (the weight on multi-step latent consistency) as a diagnostic control and measure an empirical expansion proxy L20,q95 together with Horizon-20 prediction error E20.
- On Moving-MNIST (n=6 seeds on key pairs), increasing lambda from 0 to 0.8 cut L20 from 4.96 +/- 2.01 to 1.01 +/- 0.06 (paired t p=0.005, Wilcoxon p=0.031) and halved E20 (0.365 to 0.177, paired t p=1.1e-13).
On Moving-MNIST (n=6 seeds at the critical pair), raising lambda from 0 to 0.8 cuts L20 from 4.96 +/- 2.01 to 1.01 +/- 0.06 (paired t p=0.005, Wilcoxon p=0.031)…
Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21646v1 Announce Type: new.
- Abstract: Ensuring safety in reinforcement learning under nonstationarity requires determining whether a learning system can safely adapt to forecasted environmental changes within the required recovery scope.
- Existing safe reinforcement learning methods typically assume stationary environments and do not explicitly consider adaptation speed as a safety concern.
- However, when environments evolve over time, delayed adaptation may result in transient unsafe behavior.
- EN Key Points:
- arXiv:2607.21646v1 Announce Type: new
- Abstract: Ensuring safety in reinforcement learning under nonstationarity requires determining whether a learning system can safely adapt to forecasted environm…
- Existing safe reinforcement learning methods typically assume stationary environments and do not explicitly consider adaptation speed as a safety concern
- However, when environments evolve over time, delayed adaptation may result in transient unsafe behavior
A Drift Stable Quantum Federated Learning for Intelligent Services
- Publication Time: 2026-07-27 12:00 Beijing Time
- Abstract: - arXiv:2607.21647v1 Announce Type: new.
- Abstract: Quantum federated learning enables distributed clients to train quantum neural networks without sharing local data, making it promising for privacy-aware intelligent services.
- Intelligent services in this context refer to privacy-sensitive distributed decision systems, such as fraud detection and genomic classification, where reliable and fair client-level learning is as important as the accuracy of the aggregated model.
- However, heterogeneous client data and noisy quantum optimization often cause unstable local updates, client drift, and unfair performance between clients.
- EN Key Points:
- arXiv:2607.21647v1 Announce Type: new
- Abstract: Quantum federated learning enables distributed clients to train quantum neural networks without sharing local data, making it promising for privacy-aw…
- Intelligent services in this context refer to privacy-sensitive distributed decision systems, such as fraud detection and genomic classification, where reliable…
- However, heterogeneous client data and noisy quantum optimization often cause unstable local updates, client drift, and unfair performance between clients