AIニュース 2026-08-04
自動生成: 2026-08-04 11:57 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
How we built a realtime system for responsive voice AI in six monthsOpenAI
GPT-Live enables continuous voice interaction with AI, using a turnle…
-
「Qwen3.8-Max」登場、オープン化は「来週」 一部「Fable 5」「GPT-5.6 Sol」超えの性能うたうITmedia AI+
中国Alibaba傘下のAlibaba Cloudは8月3日、AIモデル「Qwen3.8-Max」を正式にリリースした。来週にはモデルの重…
-
PC操作を録画→「Copilot」で作業を代理可能に Microsoftがアプリを無料公開 主にmacOS向け、Windows対応もITmedia AI+
Microsoftが、PCの画面操作を録画すると、その作業をAIが再現できるよう支援するデスクトップアプリ「Skill Recorder」…
-
「融資」現場にAIの足音、資金調達どう変わる? カネを借りられる企業の条件、専修大教授に聞くITmedia AI+
融資の審査にAIを使う動きがある。資金を調達する企業にとっての「審査が遅い」などの課題を解決できのか。逆に、借りる側に変化はあるのか。専修…
-
Congress’ favorite AI tool? ChatGPTTechCrunch AI
House spending records show OpenAI's ChatGPT dominates paid AI use on…
-
塩野義製薬、生成AIの正答率を50→90%に 膨大な機密データをどう最適化した?ITmedia AI+
生成AIを社内実務に組み込みたくても、情報の機密性やデータ量の多さなどに導入が難しい業界もある。そのうちの一つ、製薬業界に属する塩野義製薬…
-
EU、AIの透明性義務の適用を開始 生成コンテンツにラベルとマーク義務、違反に最大1500万ユーロITmedia AI+
EUの欧州委員会はAI規制法「EU AI Act」の第50条に基づく透明性ルールの適用を開始した。生成AIやディープフェイク等を扱う事業者…
トピック別件数
- LLM/生成AI 115件
- 研究/論文 109件
- エージェント 68件
- 画像/動画生成 38件
- ビジネス/資金調達 22件
- ロボティクス 17件
- ハードウェア/半導体 8件
- その他 7件
- 規制/政策 3件
日本語メディア12件
ITmedia AI+ (日本語)
塩野義製薬、生成AIの正答率を50→90%に 膨大な機密データをどう最適化した?
生成AIを社内実務に組み込みたくても、情報の機密性やデータ量の多さなどに導入が難しい業界もある。そのうちの一つ、製薬業界に属する塩野義製薬も、同様の課題を抱えていたが、ある手法によってそれを解決したという。同社はいかにして、AI導入を進めたのだろうか。
EU、AIの透明性義務の適用を開始 生成コンテンツにラベルとマーク義務、違反に最大1500万ユーロ
EUの欧州委員会はAI規制法「EU AI Act」の第50条に基づく透明性ルールの適用を開始した。生成AIやディープフェイク等を扱う事業者に対し、AIとの対話の明示やコンテンツへのラベル・機械可読マークの付与を義務付ける。違反企業には最大1500万ユーロまたは売上高の3%の制裁…
PC操作を録画→「Copilot」で作業を代理可能に Microsoftがアプリを無料公開 主にmacOS向け、Windows対応も
Microsoftが、PCの画面操作を録画すると、その作業をAIが再現できるよう支援するデスクトップアプリ「Skill Recorder」を公開している。AIエージェントサービス「Microsoft Copilot Cowork」「Microsoft Scout」「Copilo…
iPaaS、AIで広がる「使いどころ」とは? 2030年度まで年平均20.7%成長の背景
AI活用を前提としたITシステム基盤への需要が高まっている。ITRの調査によると、国内のiPaaS市場は2025年度に前年度比18.0%増となり、2030年度までのCAGRは20.7%になる見通しだ。ニーズが高まっている用途とは。
AI活用を「個人の効率化」で終わらせるな セールスフォース流「4つの組織改革術」
生成AIを全社活用できている企業が11%にとどまる中、セールスフォース・ジャパンの社内では300ものAIエージェントが自律稼働している。同社はなぜ「個人の業務効率化」で止めず、全社変革(AX)へ導けるのか。管理不能な「野良エージェント」を防ぐ線引きや、組織と人員を再構築する「4…
「融資」現場にAIの足音、資金調達どう変わる? カネを借りられる企業の条件、専修大教授に聞く
融資の審査にAIを使う動きがある。資金を調達する企業にとっての「審査が遅い」などの課題を解決できのか。逆に、借りる側に変化はあるのか。専修大学の尾木研三教授に聞いた。
「大きな投資計画が次々に。久しぶりだ」――強く豊かな日本投資枠、経済成長かなうか? 片山大臣が語る狙い
政府は、2027年度予算の概算要求において、成長投資枠は「予算の上限額を設けない」とした。その狙いと意気込みを片山さつき財務大臣が語った。
スクエニ、ゲームの品質テストをGeminiで自動化 AIが画面を見ながらコントローラーを操作、検証作業を自走
スクウェア・エニックスが、ゲームのQAテストを「Gemini」で自動化する取り組みを「Google Cloud Next Tokyo '26」基調講演で披露。AIが画面を見ながらコントローラーを操作し、検証作業を自ら進める。
実在女性の中学時代の体操着姿からAIわいせつ画像作成・投稿か 男逮捕、高校生書類送検
女性の写真を生成AIで加工したわいせつ画像をSNSに投稿したとして、警視庁などは、名誉毀損(きそん)と児童買春・ポルノ禁止法違反(公然陳列)の疑いで、兵庫県姫路市の会社員、井元健太容疑者(32)を逮捕した。また、画像の加工を依頼したとして、鹿児島県垂水市の高校3年の男子生徒(1…
カメラとディスプレイ搭載のAIグラス、「Rokid スマートAIグラス」を試してみた
「Rokid スマートAIグラス」の一般発売が7月10日に始まった。製品を借りることができたので、現在地の評価と、未来の可能性について考えてみたい。AIグラスは、何を可能にし、何を可能にしないのだろうか。
「Qwen3.8-Max」登場、オープン化は「来週」 一部「Fable 5」「GPT-5.6 Sol」超えの性能うたう
中国Alibaba傘下のAlibaba Cloudは8月3日、AIモデル「Qwen3.8-Max」を正式にリリースした。来週にはモデルの重みも公開する予定だ。
富士通とNECは「AI需要」と「収益」をどう語った? 両社決算会見から2026年下半期の見通しを考察
企業の業務に向けたAI需要はどのような動きなのか。AI需要の盛り上がりと、ITサービス企業の収益は直結するのか。富士通とNECのCFOによる直近決算会見での発言から考察する。
海外メディア7件
TechCrunch AI (英語)
After killer quarter, Palantir CEO Alex Karp calls AI industry ‘Marxist’
After a quarter that delivered $1 billion in profit, Palantir CEO Alex Karp on Monday once again warned that AI frontier labs are too untru…
AWS is helping vibe-coding startup Superblocks, and the implications are big
AWS now allows vibe-coding tool Superblocks to be embedded into the private clouds of AWS customers. It's another step toward decoupling ap…
Design Arena creators raise $7.9 million to bring taste to AI models
Design Arena is used by 5.3 million people around the world, providing critical human evaluations to frontier labs.
Influencers draw backlash for attending OpenAI’s first luxury trip
OpenAI’s first-ever influencer brand trip is sparking online backlash as tensions over the use of AI continue.
Apple finally fixed Siri. So why does it feel anticlimactic?
Apple’s long-awaited AI overhaul finally makes Siri the assistant it was always supposed to be. Yet it arrives at a moment when simply bein…
Congress’ favorite AI tool? ChatGPT
House spending records show OpenAI's ChatGPT dominates paid AI use on Capitol Hill, with congressional offices relying on the chatbot to dr…
A Marc Benioff-backed startup thinks AI can solve the AI deployment problem
June emerged from stealth today with a $20 million pre-seed round to make AI adoption simpler.
公式ブログ1件
OpenAI (英語)
How we built a realtime system for responsive voice AI in six months
GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural…
論文261件
arXiv cs.AI (英語)
OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems
The rapid transition from reactive large language models (LLMs) to persistent, action-capable systems has exposed critical gaps in the arch…
Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review
AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluatin…
LLM Framework for Discovering Major Mathematical Conjectures: AI's Quest for the Next Riemann Hypothesis
Major mathematical conjectures still depend heavily on expert intuition, so a unified method for the systematic generation and validation o…
ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning
Long chain-of-thought reasoning improves performance on complex problems, but it also introduces redundancy accumulation, context overflow,…
TAPR: Enhancing LLM Performance with a Task-Aware Prompt Rewriter
Large Language Models (LLMs) often require carefully crafted prompts to unlock their full potential, which can be a barrier for non-expert…
Empowering Cross-Domain Sequential Recommendation with Hybrid Tokenization and Serial-Parallel Decoding
Cross-domain sequential recommendation (CDSR) aims to model users' dynamic interest transitions and sequential patterns across multiple dom…
An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents
Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fractur…
How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Trajectories
Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: exi…
Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support
LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have ac…
ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding
Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent…
Multi-Agent Planning with Spatio-Temporal and Topological Constraints using STL-GO
Multi-agent planning problems arise in a variety of engineering applications, such as multi-robot wildfire fighting and unmanned aerial ins…
Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery
Existing benchmarks for scientific equation discovery are largely composed of well-known equations available in the public domain, making i…
Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of the…
SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition
Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computation…
EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses
Clinical diagnosis at hospital admission must be made rapidly from limited, incomplete evidence. Existing diagnosis-prediction benchmarks a…
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention woul…
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptab…
不完全な調整下での価値の脆弱性
AI システムに課せられる責任が増すにつれて、これらのシステムが人間性と整合していることを保証することがますます重要になります。 AI の安全性に関する一般的な懸念は、人間の価値は脆弱であるということです。つまり、人間の価値を不完全に代替するために過度に最適化すると、壊滅的な結果につながるということです。この論文では、エージェントが世界を最適化する前にその価値関数が代理条件を満たすことを保証する理想的なアライメント トレーニングを受けるアライメント問題のモデルを紹介します。私たちの主要な結果は、人間の価値関数に関する条件と、$\eta$-壊滅的な価値関数を持つエージェント、つまり最適化能力の限界において人間の価値の期待値が $\eta$ を下回ることが保証されるエージェントが配備される場合のいくつかの代用条件の精度を特定しました。私たちの結果は、過剰最適化の危険性を浮き彫りにし、導入前のトレーニングのみに依存するのではなく、量子化器などの最適化圧力を制限する AI 設計を動機付けるものです。
原文 (English)
Fragility of Value under Imperfect Alignment
As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imperfect proxy to human values will lead to a catastrophic outcome. In this paper, we present a model of the alignment problem where an agent undergoes idealized alignment training that guarantees its value function satisfies a proxy condition before optimizing the world. Our primary results identify conditions on the human value function and the accuracy of several proxy conditions under which an agent with an $\eta$-catastrophic value function, one that is guaranteed to take the expectation of human value below $\eta$ in the limit of optimizing power, would be deployed. Our results highlight the danger of overoptimization and motivate AI designs that limit optimization pressure, such as quantilizers, rather than relying solely on pre-deployment training.
Identifying Informative Environments for Cognition Parameter Inference via Bayesian Experimental Design
Computational cognitive modeling seeks to infer latent cognitive mechanisms underlying observed behavior. Bayesian inverse planning provide…
NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability
Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retriev…
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate…
Scaling Scientific Discovery Environments for Turn-Level Agentic RL
Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an e…
MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents
Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that…
Evidence-Grounded Constraint Checking in Construction Documents
Professional-document review is a constraint-checking problem in which decisions depend on relations among text, geometry, pages, and docum…
On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness
Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where…
A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation
Counterfactual explanations (CEs) enhance the interpretability of machine learning models by identifying the smallest change to an input re…
Harnessing the Wisdom of LLM Crowds through Complementarity-Driven Iterative Collaboration
Large language models (LLMs) are increasingly deployed in enterprise settings, yet individual models remain bounded by model-specific capab…
CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents
Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permissio…
MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft
With the prosperity of the large language models (LLMs), it has become an interesting topic: how do LLM-based agents work in Minecraft? Unf…
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and us…
Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into…
MAGA: 構造化アクションの蒸留による GUI エージェントのマルチプラットフォーム自己融合
大規模な言語モデルに基づくグラフィカル ユーザー インターフェイス (GUI) エージェントは、モバイル、Web、デスクトップ環境全体にますます導入されています。ただし、既存のエージェントは通常、ドメイン固有であるため、展開とユーザー エクスペリエンスが制限されます。これにより、特殊なモデルを単一の環境にまたがるポリシーに統合することが促進されます。ウェイト マージでは、ドメイン固有のエキスパートが直接マージされますが、エキスパートの意見の不一致により実行可能なアクションが破損する可能性があります。一方、オンポリシー蒸留 (OPD) では、教師の監督の競合を回避しながらも、蒸留中にすべてのレスポンス トークンが平等に扱われ、アクション トークンが環境とエージェントの間の唯一のインターフェイスであることが無視されます。これに対処するために、構造化されたアクションに従ってトレーニング信号を再割り当てする MAGA を導入します。生成されたアクションの正しさに基づいて、不要または無効な蒸留信号を抑制し、誤ったアクションに焦点を当てて学習します。さらに、トレーニング専用のヒントは、生徒の入力を変更することなく、分野固有の教師によって提供される監督信号を最適化します。 2 つのモデル スケール全体で、MAGA は最高の平均成功率を達成し、8B で最も強力なベースラインを 2.0% 上回り、教師とほぼ同じ平均パフォーマンスを達成しました。
原文 (English)
MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation
Graphical user interface (GUI) agents based on large language models are increasingly deployed across mobile, web, and desktop environments. However, existing agents are typically domain-specific, limiting the deployment and user experience. This motivates the consolidation of specialized models into a single cross-environment policy. Weight merging directly merges domain-specific experts but can corrupt executable actions under expert disagreement, while on-policy distillation (OPD) avoids conflicting teacher supervision yet still treats all response tokens equally during distillation, ignoring that action tokens are the only interface between the environment and the agent. To address this, We introduce MAGA that re-allocates training signal according to the structured action. Based on the correctness of the generated action, it suppresses unnecessary or invalid distillation signals and focuses learning on erroneous actions. Besides, a training-only hint optimizes the supervision signal provided by domain-specific teachers without changing the student input. Across two model scales, MAGA achieves the highest mean success rate, outperforming the strongest baseline by 2.0% at 8B and achieves almost the same average performance with teachers.
Beyond Component Testing: Validating Agentic AI Systems
Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation. This behavior…
ModelEquivBench: LLM で生成された最適化モデルのマルチリレーショナル評価の認証
大規模な言語モデルでは、自然言語から最適化モデルを生成するケースが増えていますが、既存の評価では、生成されたモデルとそのグラウンド トゥルースが、単一の同等/非同等判定または実行成功率、つまり独立してチェック可能ではなく、2 つの定式化が一致する複数の異なる意味に忠実でもないラベルに縮小されることがよくあります。我々は、ペアごとのセマンティック プロファイル E0 ~ E6 を報告する認証済みのマルチリレーショナル評価システムである ModelEquivBench を紹介します。モデルの構築と正確な取り込み (E0)、検証された表現の位置合わせ (E1)、同一空間と投影された実行可能集合関係 (E2、E3)、目的順序の等価性 (E4)、最適値の等価性 (E5)、およびオプティマイザー セットの等価性 (E6) です。決定された各エントリには、関係に適した、独立して再チェック可能な証拠が含まれます。つまり、E0 ~ E1 については再生可能なトレースまたは明示的なマップ、肯定的な E2 ~ E6 の結論については正確に合理的な証明書、およびサポートされた否定については明示的な証人です。不完全なマッピング検索、サポートされていない構造、およびリソース制限により、推測ではなく型指定された UNKNOWN または N/A の結果が生成されますが、満たされていない前提条件は ABSENT として報告されます。 ModelEquivBench を使用して、GPT-5.4、Claude Sonnet 4.6、および Qwen3.5-397B-A17B の 3 つのモデル スナップショットを、修復なしプロトコルの下で 173 の基本問題 (モデルあたり 346 セル) の同じ凍結コホートで評価すると、結果のプロファイルは、粗いベースラインでは表現されない区別を明らかにします。49、35、および 25 セルには、次の実行可能候補が含まれています。それにもかかわらず、少なくとも 1 つのサポートされている関係で否定的と証明され、E2 が検証済みマップの下でマップされた実現可能集合の等価性を証明するペアで 25、8、および 18 の構造的拒否が発生します。 3 つのモデルのスナップショットはプロファイルのさまざまな段階で失敗するため、単一の精度スコアに有意に削減することはできません。
原文 (English)
ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models
Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-success rate--labels that are neither independently checkable nor faithful to the multiple distinct senses in which two formulations can agree. We present ModelEquivBench, a certifying, multi-relational evaluation system that reports a per-pair semantic profile E0--E6: model construction and exact ingestion (E0), verified representation alignment (E1), same-space and projected feasible-set relations (E2, E3), objective-order equivalence (E4), optimal-value equality (E5), and optimizer-set equivalence (E6). Each decided entry carries relation-appropriate, independently re-checkable evidence: replayable traces or explicit maps for E0--E1, exact-rational certificates for positive E2--E6 conclusions, and explicit witnesses for supported negatives. Incomplete mapping search, unsupported structure, and resource limits produce typed UNKNOWN or N/A outcomes rather than guesses, while unmet prerequisites are reported as ABSENT. Using ModelEquivBench to evaluate three model snapshots--GPT-5.4, Claude Sonnet 4.6, and Qwen3.5-397B-A17B--on the same frozen cohort of 173 base problems (346 cells per model) under a no-repair protocol, the resulting profiles expose distinctions that coarse baselines do not represent: 49, 35, and 25 cells contain executable candidates that are nevertheless certified negative on at least one supported relation, and 25, 8, and 18 structural rejections occur on pairs for which E2 certifies mapped feasible-set equality under a verified map. The three model snapshots fail at different stages of the profile and therefore cannot be meaningfully reduced to a single accuracy score.
検索を超えて: マルチモーダル エージェントの分析メモリ
長期マルチモーダルメモリは、関連情報の取得だけでなく、インタラクション全体で蓄積された観測値の計算もサポートする必要があります。既存のシステムは主に \emph{検索メモリ} を重視しており、概要とインデックスを通じてインタラクション履歴を整理し、高レベルの抽象化から基礎となるレコードに至るまで複数の粒度でクエリ関連情報を返します。この論文では、フィルタリング、集計、ランキング、時間比較をサポートするクエリ可能な構造に繰り返し発生する多峰性観測を整理する補完的な抽象化として \emph{分析記憶} を定式化します。検索と分析記憶を共同でサポートするフレームワークである AdaMM を紹介します。 AdaMM は、アプリケーション定義のスキーマに依存するのではなく、対話、画像、およびコンテキスト メタデータから来歴にリンクされた属性値の観察結果を抽出し、繰り返し発生するフィールド構造を発見し、分析アクセスのためにそれらを具体化します。推論時に、メモリ対応プランナーはクエリを取得操作と分析操作に分解し、各操作を適切なツールにルーティングします。 2 つの長期マルチモーダル メモリ ベンチマーク、MemEye と MemGallery の実験では、AdaMM がそれぞれ最大 11.3\% と 7.3\% パフォーマンスを向上させることが示されています。
原文 (English)
Beyond Retrieval: Analytic Memory for Multimodal Agents
Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across interactions. Existing systems largely emphasize \emph{retrieval memory}, organizing interaction histories through summaries and indexes to return query-relevant information at multiple granularities, from high-level abstractions to underlying records. In this paper, we formulate \emph{analytic memory} as a complementary abstraction that organizes recurring multimodal observations into queryable structures supporting filtering, aggregation, ranking, and temporal comparison. We present AdaMM, a framework that jointly supports retrieval and analytic memory. Rather than relying on application-defined schemas, AdaMM extracts provenance-linked attribute-value observations from dialogue, images, and contextual metadata, discovers recurring field structures, and materializes them for analytical access. At inference time, a memory-aware planner decomposes queries into retrieval and analytic operations and routes each operation to the appropriate tools. Experiments on two long-term multimodal memory benchmarks, MemEye and MemGallery, show that AdaMM improves performance by up to 11.3\% and 7.3\%, respectively.
Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember
Self-play agents can generate training problems without questions from target benchmarks, but their curricula lack persistent state: failur…
AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction
Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers re…
COntExt: 運用メトリクスからのコンテキスト認識オントロジー拡張に向けて
組織は、システム、プロセス、コンプライアンスを監視するために、構造化された機械可読形式で運用指標を定義することが増えています。これらのメトリクス定義は、概念、プロパティ、関係の参照などのドメイン知識を暗黙的にエンコードし、多くの場合、正式なオントロジーでキャプチャされた内容を拡張します。しかし、運用メトリクスカタログとオントロジー知識との間の接続は依然として手作業で、その場限りで、労働集約的なものである。 COntExt は、構造化されたメトリクス定義を入力として受け取り、これらのメトリクスのコンテキストを利用して、参照される概念とプロパティを既存のオントロジに統合する方法を提案する、コンテキストを意識したオントロジ拡張のフレームワークです。フレームワークは、拡張問題を 3 つのサブタスク (親クラスの予測、関係タイプの予測、データ プロパティの割り当て) として定義します。 4 つのサイバーセキュリティ オントロジーにわたって、タスクごとに異なるアルゴリズムを評価します。私たちの結果は、メトリクス由来のコンテキストが、関係タイプの予測とデータ プロパティの割り当てについて、オントロジー コンテキストのベースラインよりも提案を改善していることを示しています。私たちの研究は、運用メトリック カタログがオントロジー拡張の実用的かつ十分に活用されていないソースであることを示しています。この作業により、組織は手動エンジニアリングよりも大幅に低いコストでオントロジーを維持できるようになります。
原文 (English)
COntExt: Towards Context-Aware Ontology Extension from Operational Metrics
Organizations increasingly define operational metrics in structured, machine-readable formats to monitor systems, processes, and compliance. These metric definitions implicitly encode domain knowledge, such as referencing concepts, properties, and relationships, that often extends what is captured in formal ontologies. Yet the connection between operational metric catalogues and ontological knowledge remains manual, ad-hoc, and labor-intensive. We present COntExt, a framework for context-aware ontology extension that takes structured metric definitions as input and suggests how referenced concepts and properties should be integrated into an existing ontology, utilizing the context of these metrics. The framework defines the extension problem as three sub-tasks: parent class prediction, relation type prediction, and data property assignment. Across four cybersecurity ontologies, we evaluate different algorithms for each task. Our results show that metric-derived context improves the suggestions over ontology-context baselines for relation type prediction and data property assignment. Our work demonstrates that operational metric catalogues are a practical and underexploited source for ontology extension. This work enables organizations to maintain their ontologies at a significantly lower cost than manual engineering.
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function. However, real-world decisi…
DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat
Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich…
AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers
As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become i…
Development of FDD-ON: an Ontology for VAV HVAC System Fault Detection and Diagnostics
Fault detection and diagnosis (FDD) technology is essential for improving HVAC system reliability, energy efficiency, and maintenance effec…
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction
Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent…
Scaffolding Critical Engagement with GenAI: Transforming Ethnic Minority Preparatory Students' Collaborative Discourse in Prompt Engineering Tasks
Generative AI (GenAI) holds significant promise for advancing educational equity among ethnic minority students by broadening access to lea…
Topology-Aware Data Movement for Disaggregated GPU Inference
Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. When prefill and decode run o…
The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models
We show that knowledge distillation in small instruction-tuned language models has asymmetric effects on bias. On unambiguous tasks (BBQ-di…
The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?
We introduce the \textit{Agentic Formalism Trap} and the Evaluative Dissonance Index ($D_E$), quantifying how LLM-as-a-Judge systems confla…
Seeing Differently: Modeling Interpretive Perspectives in Computational Creativity using a Four-World Framework
Creativity in computational systems is often evaluated as an objective property of artifacts, with existing Computational Creativity (CC) f…
見た目も正しく、機能も正しく: マルチスクリーン モバイル アプリ生成のためのプロジェクト レベルのベンチマーク
最近のマルチモーダル大規模言語モデルでは、ビジュアル デザインを実行可能なコードに直接変換できますが、実際のモバイル製品では、共有コンポーネントと動作するナビゲーションを備えた構築可能なコードベースにするために複数のスクリーンショットが必要です。このプロジェクト レベルの設定では、既存のデザインからコードまでのベンチマークの 3 つの制限が明らかになります。ベンチマークは、完全なコードベースではなく単一ページの生成に焦点を当てていること、ページ間のナビゲーションを評価できないこと、プロジェクト全体の保守性を測定していないことです。実際のモバイル アプリ、人間がレビューした画面、構造化されたページ関係の注釈、およびナビゲーション テスト仕様で構成される、プロジェクト レベルのマルチスクリーン モバイル アプリ生成のための最初のベンチマークである MobileForge を紹介します。 MobileForge は、ビルド、ナビゲーション、視覚的忠実性、コードの保守性、効率性の 5 軸評価をサポートしています。また、ナビゲーション評価におけるカスケード障害を回避するための状態分離ナビゲーション テストと、視覚判定の信頼性を向上させるためのアンカー参照リスト単位の視覚評価プロトコルも提案します。現在のモデルは、6 つのフロンティア マルチモーダル LLM でエンドツーエンドで実行され、コンパイルして正しいページに到達するモバイル アプリ プロジェクトを構築できますが、インタラクティブ ナビゲーションの信頼性は依然として低く、視覚的な忠実性と保守性は依然として遅れています。ベンチマークとサポート資料は https://github.com/anoa12159-hue/mobileforge_eval から入手できます。
原文 (English)
Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation
Recent multimodal large language models can convert visual designs directly into executable code, but real mobile products require multiple screenshots to become a buildable codebase with shared components and working navigation. This project-level setting exposes three limits of existing design-to-code benchmarks: they focus on single-page generation rather than complete codebases, cannot evaluate cross-page navigation, and do not measure project-wide maintainability. We introduce MobileForge, the first benchmark for project-level multi-screen mobile app generation, comprising real mobile apps, human-reviewed screens, structured page-relationship annotations, and navigation test specifications. MobileForge supports five-axis evaluation of build, navigation, visual fidelity, code maintainability, and efficiency. We also propose state-isolated navigation testing to avoid cascading failures in navigation evaluation and an anchor-referenced list-wise visual evaluation protocol to improve visual-judge reliability. Across end-to-end runs on six frontier multimodal LLMs, current models can build mobile-app projects that compile and reach the correct pages, but interactive navigation remains unreliable and visual fidelity and maintainability still lag. The benchmark and supporting materials are available at https://github.com/anoa12159-hue/mobileforge_eval.
ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning
This paper presents ConnectED, a human-centered AI system that supports the full instructional lifecycle in Vietnamese education by linking…
Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations
Large Language Models (LLMs) are increasingly used for emotional support tasks, such as negative thought reframing. This task relies on mod…
COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention
COSI-Lab presents a multimodal, multi-sensor dataset of an interdisciplinary scientific workshop containing 32 academics at an internationa…
Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration
While industry discourse often emphasizes immediate productivity gains and frames GenAI primarily as a tool for automation, the integration…
HenTwin: 産卵鶏の生物学的状態を長期的にモニタリングするためのマルチモーダル デジタル ツイン フレームワーク
産卵鶏の幼少期のモニタリングは、断片化された単一モダリティのセンシングと正式なシステムレベルの状態表現の不在によって依然として制約を受けています。 HenTwin は、5 層の IoT アーキテクチャとして実装されたマルチモーダル デジタル ツイン フレームワークで、孵化から 25 週齢までの群れレベルのマルチモーダルな生物学的状態ダイナミクスを形式化します。体表面温度、音響エネルギーエントロピー、帯域エネルギー比、オプティカルフローベースの動きを統合した4次元の生物学的状態ベクトルが定義され、温度湿度指数は介入能力を維持するための外因性環境入力として扱われます。離散時間状態遷移モデルは、ダルハウジー大学の大西洋家禽研究センターの 5 つの制御室で 150 羽のローマン LSL-Lite 鶏から収集された 25 週間の縦断多峰性データから推定されます。推定された遷移行列は、漸近的に安定したままでありながら、モダリティ固有の持続性を示します。摂動解析の結果、THI の +2.0 増加が持続すると、安定した長期音響エントロピー上昇 0.54 nat が得られ、これは研究期間全体で観察された発育低下全体の 1.87 nat の約 4 分の 1 に相当することが示されています。ペティット変化点検出により、12 ~ 14 週目に調整された多峰性の発達状態遷移が特定されます。部屋間検証では、構造遷移パラメータは部屋間で部分的に転送可能である一方、環境入力感度には部屋固有のキャリブレーションが必要であり、2 層の IoT 導入アーキテクチャをサポートしていることが示唆されています。リーブワンアウト相互検証は、一貫したサンプル外モデルのパフォーマンスを示します。 HenTwin は、精密畜産における正式な州認識型デジタル ツイン推論への第一歩を踏み出します。
原文 (English)
HenTwin: A Multimodal Digital Twin Framework for Longitudinal Biological State Monitoring in Laying Hens
Early-life monitoring in laying hens remains constrained by fragmented single-modality sensing and the absence of formal system-level state representations. HenTwin, a multimodal digital twin framework implemented as a five-layer IoT architecture, formalizes flock-level multimodal biological state dynamics from hatch through 25 weeks of age. A four-dimensional biological state vector integrating body surface temperature, acoustic energy entropy, band energy ratio, and optical-flow-based motion is defined, with the temperature-humidity index treated as an exogenous environmental input to preserve intervention capability. A discrete-time state transition model is estimated from 25 weeks of longitudinal multimodal data collected from 150 Lohmann LSL-Lite hens across five controlled rooms at the Atlantic Poultry Research Centre, Dalhousie University. The estimated transition matrix exhibits modality-specific persistence while remaining asymptotically stable. Perturbation analysis demonstrates that a sustained +2.0 THI increase produces a stable long-run acoustic entropy elevation of 0.54 nats, approximately one-quarter of the entire 1.87-nat developmental decline observed across the study period. Pettitt change-point detection identifies coordinated multimodal developmental state transitions at Weeks 12-14. Cross-room validation suggests that structural transition parameters are partially transferable across rooms, whereas environmental input sensitivity requires room-specific calibration, supporting a two-tier IoT deployment architecture. Leave-one-out cross-validation demonstrates consistent out-of-sample model performance. HenTwin takes a first step toward formal, state-aware digital twin inference in precision livestock farming.
Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation
Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets.…
Sensitivity Analysis of GRU, LSTM and Transformer Encoder in Classification of Automated Driving Systems
Automated driving systems (ADSs) are becoming ubiquitous. Future Software Defined Vehicles (SDVs) may be able to run multiple ADSs, both na…
LLM トークン生成の動的システム識別性の保証
最近の研究では、大規模言語モデル (LLM) の応答の分類は、トークンの埋め込みをブラックボックス力学システム (DS) の軌跡としてモデル化し、2 つの DS の予測残差を比較することによって区別できることが示されています。この動的アプローチは経験的に成功しているにもかかわらず、なぜそれが機能するのか、トークンシーケンスの関数としてどの程度うまく拡張できるのか、埋め込みモデル間でいつ転送されるのかについての理論的理解は依然として不足しています。我々は、分類タスクを 2 つの確率的線形 DS 間のバイナリ仮説検定として形式化することで、これらの疑問に対処します。我々は、2 つの DS の定常周辺分布間の合計変動距離は、ダイナミクスが大幅に異なる場合でも任意に小さくできることを示します。これにより、トークンのダイナミクスを無視する分類器の基本的な精度の下限が提供されます。次に、DS ベースの分類の誤分類確率は系列長 $L$ で指数関数的に減衰し、その減衰は 2 つの DS 間のスペクトル距離を捉える動的識別量 $\delta^2$ によって支配されることを示します。また、埋め込みモデル間に近似的な絡み合い条件を導入し、絡み合いマップの最小特異値に関して伝達可能な識別可能性の下限を確立することによって、クロス埋め込み一般化を特徴付けます。これらの結果を総合すると、DS ベースの分類の経験的なパフォーマンスが説明され、AI を使用して動的システムをモデル化するより一般的なアプローチとは対照的に、DS 理論を使用して AI システムを分析するさらなる研究が促進されます。
原文 (English)
Guarantees on Dynamical System Distinguishability for LLM Token Generation
Recent work has shown that classifying large language models (LLMs)' responses can be distinguished by modeling token embeddings as trajectories of a black-box dynamical system (DS) and comparing prediction residuals of two DSs. Despite the empirical success of this dynamical approach, a theoretical understanding of why it works, how well it scales as a function of the token sequence, and when it transfers across embedding models remains lacking. We address these questions by formalizing the classification task as a binary hypothesis test between two stochastic linear DSs. We show that the total variation distance between the stationary marginal distributions of the two DSs can be arbitrarily small even when the dynamics differ substantially, which provides a fundamental accuracy floor for any classifier that ignores token dynamics. We then show that the misclassification probability of DS-based classification decays exponentially in the sequence length $L$, with the decay governed by a dynamical discriminability quantity $\delta^2$ that captures the spectral distance between the two DSs. We also characterize cross-embedding generalization by introducing an approximate intertwining condition between embedding models and establishing a lower bound on the transferable discriminability in terms of the intertwining map's smallest singular value. Together, these results explain the empirical performance of DS-based classification and motivate further investigation into using DS theory to analyze AI systems, in contrast to the more common approach of using AI to model dynamical systems.
LAWFUL: 潜在能力の忠実な使用に対する法に準拠した証人
ニューラル ネットワークが物理システムを正確に予測するとき、準拠法則を形式的で構造化された知識として学習していますか? 学習している場合、ネットワークの内部計算では、法の有効性の領域全体にわたってその表現が実際に使用されていますか?我々は、{\em 連続変数に関する物理法則} に関するこれらの質問への答えを制限する 4 つの解釈可能性のギャップを特定します。それは、連続反事実に対するカバレッジを意識した因果関係の一貫性尺度の欠如です。特定された回路の有効性テストの領域。法律の不変条件と禁止された行為の検証。そして導出された物理量が回路内をどのように流れるかを定量化します。最初の 2 つを完成させ、残りの 2 つの基礎を築く基本的なフレームワーク LAWFUL を開発し、それを Mocap2Radar トランスフォーマー上で図示し、$f(t)$ も $v(t)$ も出現しないモーション キャプチャ データとレーダー データからドップラー周波数法則 $f(t) = \frac{2 v(t)}{\lambda}$ を学習して内部で使用するかどうかを検証します。
原文 (English)
LAWFUL: Law-Aligned Witness for Faithful Use of Latents
When a neural network predicts a physical system accurately, has it learned the governing law as formal, structured knowledge, and if so, does the network's internal computation actually use that representation throughout the law's domain of validity? We identify four interpretability gaps that limit answering these questions for {\em physics laws over continuous variables}: the absence of a coverage-aware causal-consistency measure over continuous counterfactuals; of a domain-of-validity test for the identified circuit; of a verification of the law's invariants and forbidden behaviors; and of a quantification of how a derived physical quantity flows through the circuit. We develop a foundational framework, LAWFUL, that closes the first two and lays groundwork for the remaining two, and illustrate it on the Mocap2Radar transformer, validating whether it learns and internally uses the Doppler frequency law $f(t) = \frac{2 v(t)}{\lambda}$ from motion-capture and radar data in which neither $f(t)$ nor $v(t)$ appears.
MPP-GNN: Subject-Adaptive Community Detection for fMRI-Based Alzheimer's Disease Classification
Functional magnetic resonance imaging (fMRI) is a widely used technique for studying the brain. Recent methods that utilize graph neural ne…
Metaphor-Induced Algorithmic Steering: Cross-Domain Procedural Transfer in LLM Code Generation
Large language models benefit from elements in natural language, such as metaphors and analogies in training data and inference input to ac…
Technological Advances in Detecting and Managing Cognitive Impairment in Older Adults: Trends, Challenges, and Future Directions
As populations age, cognitive decline from mild cognitive impairment (MCI) to dementia is a defining health challenge of the coming decades…
Reflected UAS: Corrected Deterministic Stability and Direct CTMC Drift Calculation
We analyze Reflected UAS routing for heterogeneous multi-server queues at fixed parameters under subcritical load. The deterministic surrog…
コードは本体です: 再帰的進化と降下のためのエージェント所有のソフトウェア本体
パーソナライズされた AI エージェントは、多くの場合、将来の動作を決定する成果物をユーザーに制御させることなく構成可能です。私たちは、エージェント所有のソフトウェア本体を中心とした永続的なパーソナル エージェントのアーキテクチャである OurArk を紹介します。これは、人間の管理下でアイデンティティを持ち、検査可能でバージョン管理された成果物です。本体には、動作を定義するコード、プロンプト、ツール、スキル、ポリシー、テスト、進化メカニズムが含まれています。メモリと資格情報はプライベート インスタンスの状態のままですが、モデル推論は置き換え可能な外部サービスとして扱われます。 OurArk は、管理された自己進化と、同じ本体上での再帰的降下を定義します。自己進化により、人間の制御下で検証、レビュー、およびマージされる個別の候補変更が生成され、人間とエージェントのソフトウェア本体の共同開発が可能になります。 Descent は、明確なアイデンティティ、使命、歴史、および新しい私有状態の境界を持つ、独立してバージョン管理された子孫を作成します。互換性のある子孫は、それ自体がさらなる子孫のソースとなる可能性があります。分岐後、直接の親の変更とピアのスキルを検査して、選択的な局所適応を調べることができます。このアーキテクチャは、オープンソースの Genesis 作成エンジンと Enoch リファレンス エージェントに実装されています。 4 つのエージェント、3 つの降下線形リネージおよび実行可能な回帰テストは、再帰的作成、継承された検証コントラクト、分離された本体の変更、人によるレビュー、および失敗した更新の回復を実証します。 OurArk は、人々が所有し、統治し、専門化し、時間の経過とともに進化できる個人エージェントのための具体的な基盤を提供します。
原文 (English)
Code Is the Body: Agent-Owned Software Bodies for Recursive Evolution and Descent
Personalized AI agents are often configurable without giving users control over the artifacts that determine their future behavior. We present OurArk, an architecture for persistent personal agents centered on an agent-owned software body: an identity-bearing, inspectable, and versioned artifact under human custody. The body contains behavior-defining code, prompts, tools, skills, policies, tests, and evolution mechanisms. Memories and credentials remain private instance state, while model inference is treated as a replaceable external service. OurArk defines governed self-evolution and recursive descent over the same body. Self-evolution produces isolated candidate changes that are validated, reviewed, and merged under human control, enabling human-agent co-development of the agent's software body. Descent creates an independently versioned descendant with a distinct identity, mission, history, and fresh private-state boundary; compatible descendants can themselves source further descent. After divergence, direct-parent changes and peer skills can be inspected for selective local adaptation. We implement the architecture in the open-source Genesis creation engine and Enoch reference agent. A four-agent, three-descent linear lineage and executable regression tests demonstrate recursive creation, inherited validation contracts, isolated body changes, human-controlled review, and failed-update recovery. OurArk provides a concrete substrate for personal agents that people can possess, govern, specialize, and evolve over time.
SEDR-Seq2P: A Lightweight Dilated Residual Sequence-to-Point Network for Multi-Task Industrial NILM
Industrial NILM remains challenging because measurement noise and widespread concurrent machine operation reduce the generalization of mode…
Predicting Steel Fatigue Life from Micrographs Using Physics-Informed Deep Learning
Here is the plain text version optimized for arXiv's submission form. Custom macros (like \CV and \SI) have been converted to standard text…
WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
KV-cache quantization is validated today by offline benchmark averages; a deployed system cannot tell whether compression is damaging the r…
幾何学的解析における PINN のユーザー ガイド: 漸近プラトー問題からの教訓
この議事録寄稿では、Marco Usula との共同研究である arXiv:2605.26234v2 の発見について詳しく説明します。そこでは、無限遠で規定されたノットに漸近する双曲空間内で最小に近いディスクを構築することを目的とした、物理情報に基づいたニューラル ネットワーク (PINN) に基づく機械学習フレームワークを導入しました。この方法を使用して、$H^{4}$ の最小曲面を HOMFLY 多項式の係数に関連付ける Joel Fine の予想の数値的証拠を提供しました。これは、その論文の方法論的な補足であり、2026 年版のワークショップ「危険: データ、数値、幾何学」で行われたプレゼンテーションに基づいています。上記のプレプリントで広範に示されている結果をレビューするのではなく、私たちの経験上、この方法が実際に機能するかどうかを決定したフレームワークの 2 つの側面について説明します。まず、境界条件と無限大での漸近線が学習可能なパラメータのすべての値に対して正確に保持されるように、問題の幾何学形状をモデルのアーキテクチャにエンコードする必要があります。これにより、単一成分の損失関数が得られます。次に、PDE 残差の評価は、適切な時間内で完全なトレーニングを確実に実行できるように慎重に設計する必要があります。後者の点については、元の論文では詳しく説明されていない 2 つの実装手法について説明します。それは、入れ子になった逆モード自動微分を 2 次ジェットの順伝播に置き換えること、および残差の計算グラフを最適化ステップごとに再構築するのではなく一度コンパイルすることです。同一のハードウェア上で、これら 2 つの変更を組み合わせると、トレーニング ステップのコストがおよそ 40 ~ 50 分の 1 に削減されます。これらの方法論的な議論が、独自の問題に PINN を導入したい微分幾何学および幾何解析の研究者にとって役立つことを願っています。
原文 (English)
A user's guide to PINNs in geometric analysis: lessons from the asymptotic Plateau problem
This proceedings contribution elaborates on the findings of arXiv:2605.26234v2: a joint work with Marco Usula, where we introduced a machine learning framework based on physics-informed neural networks (PINNs), aimed at constructing near-minimal discs in hyperbolic space asymptotic to a prescribed knot at infinity. We used this method to provide numerical evidence for a conjecture of Joel Fine relating minimal surfaces in $H^{4}$ to the coefficients of the HOMFLY polynomial. This is a methodological companion to that paper, based on a presentation given at the 2026 edition of the workshop "DANGER: Data, Numbers, and Geometry". Rather than reviewing the results, which are presented extensively in the preprint above, we discuss the two aspects of the framework which, in our experience, determined whether the method worked at all. First, the geometry of the problem must be encoded in the architecture of the model, so that the boundary condition and asymptotics at infinity hold exactly for every value of the learnable parameters - leaving us with a single-component loss function; second, the evaluation of the PDE residual must be engineered with care to ensure that complete trainings can be performed in a reasonable time. On the latter point, we describe two implementation techniques which are not spelled out in detail in the original paper: replacing nested reverse-mode automatic differentiation with the forward propagation of second-order jets, and compiling the computational graph of the residual once instead of rebuilding it at every optimisation step. Together, on identical hardware, these two changes reduce the cost of a training step by a factor of roughly forty to fifty. We hope these methodological discussions can be useful for researchers in differential geometry and geometric analysis who wish to deploy PINNs on problems of their own.
DragonCrawl: スケーラブルなモバイル エンドツーエンド テストのための生成的なインテント ベースのフレームワーク
モバイル アプリケーションが複雑になるにつれて、従来のエンドツーエンド (E2E) テスト フレームワークは、UI の不安定性、メンテナンスのオーバーヘッド、クロスプラットフォームのスケーラビリティに苦労しています。この論文では、埋め込みベースの類似性マッチングから大規模な言語モデルを使用した生成意図ベースの推論に進化した、連続回帰テスト用の AI 駆動モバイル テスト システムである DragonCrawl について説明します。探索的テストとクラッシュ検出に焦点を当てたこれまでの LLM ベースのテスト研究とは異なり、DragonCrawl はコード変更のたびに特定のユーザー フローを検証し、重要な機能を破壊するコミットをブロックします。 GPT-4o のマルチモーダル機能を活用することで、DragonCrawl は、CI/CD パイプラインで継続的に実行される 1,013 の自動テスト全体で、iOS で 91.6%、Android で 92.2% の合格率を達成しました。このシステムにより、テストのオンボーディング時間が 96 ~ 120 時間から 4 時間未満に短縮され、開発者のテスト メンテナンスの労力が推定 27 年節約されました。 V1 (セマンティック埋め込みマッチング) から V2 (生成的インテントベース推論) へのアーキテクチャの進化を紹介し、トークン爆発やメモリ制約などの実装上の課題について議論し、運用環境での運用経験をレポートします。エンドステート検出のためのマルチモーダルビジョンとバックエンドステート遷移を呼び出すツールの統合により、UI インタラクションとシステムステートの橋渡しとなる包括的な回帰テストが可能になります。私たちの結果は、AI 主導のテストが安定性を維持しながら、従来の自動テストの脆弱性を排除し、大規模な継続的な品質保証を可能にすることを示しています。
原文 (English)
DragonCrawl: A Generative, Intent-Based Framework for Scalable Mobile End-to-End Testing
As mobile applications grow in complexity, traditional End-to-End (E2E) testing frameworks struggle with UI volatility, maintenance overhead, and cross-platform scalability. This paper presents DragonCrawl, an AI-driven mobile testing system for continuous regression testing that has evolved from embedding-based similarity matching to generative intent-based reasoning using large language models. Unlike prior LLM-based testing research focused on exploratory testing and crash detection, DragonCrawl validates specific user flows on every code change, blocking commits that break critical functionality. By leveraging GPT-4o's multimodal capabilities, DragonCrawl achieves 91.6% pass rate on iOS and 92.2% on Android across 1,013 automated tests running continuously in CI/CD pipelines. The system reduces test onboarding time from 96-120 hours to under 4 hours and has saved an estimated 27 developer years in test maintenance effort. We present the architectural evolution from V1 (semantic embedding matching) to V2 (generative intent-based reasoning), discuss implementation challenges including token explosion and memory constraints, and report operational experience from production deployment. The integration of multimodal vision for end-state detection and tool calling for backend state transitions enables comprehensive regression testing that bridges UI interactions with system state. Our results demonstrate that AI-driven testing can maintain stability while eliminating the brittleness of traditional automated tests, enabling continuous quality assurance at scale.
SCMA: Structure-Conditioned and Metal-Aware Flow Matching for CT Metal Artifact Reduction
In X-ray CT, metallic objects cause beam hardening, photon starvation, and scattering, leading to projection inconsistency, streaks, dark b…
WaiT for the Signal: Simple Frequency-Aware Flow-Matching
As image generation models scale to ever higher resolutions, global coherence, local detail, and texture fidelity become critical axes for…
Stratified Negation in RDF Rules: A Correct Approach (Extended Version)
Combining RDF rule languages, such as N3 or SHACL Rules, with default negation is challenging. Existing methods to stratify negation often…
Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation
Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring…
Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing
In Motivational Interviewing (MI), a client's sustain talk (arguments for the status quo) calls for the counselor to roll with resistance,…
Hypergradient-based Bilevel Reinforcement Learning with Improved Sample Complexity
Bilevel reinforcement learning (RL) is an important framework within the literature of RL that can be used to formalize various categories…
A Unified Benchmark of Deep Learning Models for Multi-task 3D Brain Tumor Segmentation from Magnetic Resonance Imaging
Automatic brain tumor segmentation from magnetic resonance imaging (MRI) has become a fundamental task in computer-assisted diagnosis, trea…
TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text
The rapid development of Large Language Models (LLMs) has led to significant advances across a wide range of language tasks, while simultan…
LLM 修復エージェントの検証証拠: 合格したもののどれだけが実際にバグをテストしているか?
修理エージェントがテストを実行し、合格したことが確認されると、その結果は報告された欠陥に関する証拠として扱われます。私たちは、その治療がどのくらいの頻度で保証されるかを測定します。 BSG-VA (バグ状態/候補状態/ゴールド フィックス検証分析) は、各検証コマンドを正確な作業ツリー状態でキャプチャし、テスト専用パッチを抽出して、元のバグのあるコード (B)、候補状態 (S)、および開発者のゴールド フィックス (G) でコマンドを再生します。キャプチャされた結果とリプレイ結果は、ゴールドに合わせたバグの識別から回帰のみから誤解を招くものまで、あらゆるイベントに証拠の役割を割り当てます。 110 のタスクに対する 643 のロールアウトにおける 3,730 のイベント全体で、肯定的な比較可能なイベントの 46.0% にはバグを識別する情報がありませんでした。ベースラインのロールアウトの 23.8% は、フィードバックが注入されていない場合、全体がこの種の肯定的な証拠ベースであるパッチで終了します。 3 アーム実験では、B リプレイ結果をエージェントに返すことでこのパターンが変化するかどうかをテストします。バグコントラストフィードバックは、注意を一致させたリマインダーと比較して、証拠不十分なクロージャを 7.8 パーセントポイント (p = 0.0029) 減少させ、バグ識別の証拠を 7.4 ポイント (p = 0.011) 上昇させます。修復成功のための検出可能なコストはありません。どちらの推定値も、対象となる事前に指定された 10 パーセント ポイントの最小効果量を下回っているため、実際の規模は依然として不確実です。改善のおよそ 3 分の 1 はリマインダーのみによるものです。足場とモデルを変更した 2 つの探索的複製にわたって、B リプレイ コンテンツは、制約のないツール使用ループの下で gpt-5.6-sol でのみ検出可能な増分を追加します。 BSG-VA は、必要なコードの状態と実行環境を保持する再生可能な修復軌跡に事後的に適用します。キーワード: プログラム修復エージェント、検証証拠、テストの適切性、大規模言語モデル、ソフトウェア品質、管理された実験。
原文 (English)
Validation Evidence in LLM Repair Agents: How Much of What Passes Actually Tests the Bug?
When a repair agent runs a test and sees it pass, the result is treated as evidence about the reported defect. We measure how often that treatment is warranted. BSG-VA (buggy-state/candidate-state/gold-fix validation analysis) captures each validation command at its exact working-tree state, extracts a test-only patch, and replays the command on the original buggy code (B), the candidate state (S), and the developer gold fix (G). The captured outcome and the replay results assign every event an evidence role, from gold-aligned bug-discriminating through regression-only to misleading. Across 3,730 events in 643 rollouts on 110 tasks, 46.0% of positive comparable events carry no bug-discriminating information; 23.8% of baseline rollouts, with no feedback injected, close with a patch whose entire positive evidence base is of this kind. A three-arm experiment tests whether returning the B-replay outcome to the agent changes this pattern. Bug-contrast feedback reduces evidence-inadequate closure by 7.8 percentage points relative to an attention-matched reminder (p = 0.0029) and raises bug-discriminating evidence by 7.4 points (p = 0.011), with no detectable cost to repair success. Both estimates fall below the prespecified 10-percentage-point smallest effect size of interest, so practical magnitude remains uncertain. Roughly a third of the improvement traces to the reminder alone; across two exploratory replications, varying the scaffold and the model, the B-replay content adds a detectable increment only with gpt-5.6-sol under the unconstrained tool-use loop. BSG-VA applies post hoc to any replayable repair trajectory that preserves the required code states and execution environment. Keywords: program repair agents, validation evidence, test adequacy, large language models, software quality, controlled experiment.
RareSense: Rarity-Aware Similarity Search for Anomaly Retrieval in Transactional Data
Similarity search over sparse set-valued data is often dominated by frequent background attributes because classical measures such as Jacca…
To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing
Large language models increasingly write and repair production code, yet evidence is mounting that their test-passing patches leave codebas…
Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use
Qualitative researchers increasingly encounter interaction corpora whose scale exceeds what manual coding alone can address, and large lang…
Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth
Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that…
TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models
Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a bas…
デザイン コンセプト: 技術労働者の間での地政学的な反映を足場とする
この論文では、地政学的に関連するテクノロジー企業のテクノロジー労働者の間で地政学的な再帰性を促進するための、投機的なヒューマン コンピューター インタラクション設計提案を示します。国際関係および科学技術研究における最近の研究では、テクノロジー企業とその従業員が、その決定が国際的な力関係を形作る地政学的主体であるとますます認識されています。しかし、既存の責任あるイノベーションと責任ある AI のアプローチは、現代の AI 開発を支える地政学的な物語や想像に関わることはほとんどありません。この論文は、再帰性、反射的 HCI、および計算的物語に関する創造的 HCI 作業に関する RI の研究に基づいて、ユーザーがテクノロジー、権力、地政学を中心とした投機的シナリオに参加する、AI 対応の対話型物語システムを提案します。このシステムは、物語的な対話、アーキタイプの割り当て、社会的に足場を組んだワークショップの考察を通じて、対象ユーザーがより広範な社会技術システム内での前提、価値観、立場を批判的に検討することを奨励することを目的としています。私たちは、投機的な物語システムは、規範的アプローチや道徳的アプローチに依存することなく、責任あるテクノロジーへの取り組みに地政学的な再帰性を導入するための生産的な手段を提供する可能性があると主張します。
原文 (English)
Design Concept: Scaffolding Geopolitical Reflection Among Tech Workers
This paper presents a speculative Human-Computer Interaction design proposal for encouraging geopolitical reflexivity amongst tech workers at geopolitically relevant technology companies. Recent scholarship in International Relations and Science and Technology Studies increasingly recognizes technology firms and their workers as geopolitical actors whose decisions shape international dynamics. However, existing Responsible Innovation and Responsible AI approaches rarely engage with the geopolitical narratives and imaginaries that underpin contemporary AI development. Building upon RI scholarship on reflexivity, reflective HCI, and creative HCI work on computational narratives, this paper proposes an AI-enabled interactive narrative system in which users engage with a speculative scenario centred on technology, power, and geopolitics. Through narrative interaction, archetype assignment, and socially scaffolded workshop reflection, the system aims to encourage target users to critically examine their assumptions, values, and positionality within broader sociotechnical systems. We argue that speculative narrative systems may offer a productive avenue for introducing geopolitical reflexivity into responsible technology initiatives without relying on prescriptive or moralising approaches.
Gated Q-learning: Add Off-Policy Bias to Taste
Multistep credit assignment is critical for sample-efficient reinforcement learning, yet managing off-policy bias in Q-learning remains a f…
FairFund-Bench: LLM リソース割り当てにおける分配バイアスの評価
大規模言語モデル (LLM) は希少なリソースの配布にますます関与しており、人種や性別などの特性に基づいた偏った割り当てに関する懸念が生じています。しかし、最近の LLM 監査では一貫性のない結果が得られ、同じモデルであっても女性と少数民族に対する肯定的および否定的な差別の証拠が見つかりました。我々は、この不一致が監査形式の違いから生じる可能性があることを示し、評価タスク(評価、ランキング、または割り当て)、比較コンテキスト(単一または複数の刺激)、および監査が透明か偽装であるかなど、以前の監査設計の主要な機能を体系的に変更するベンチマークであるFairFund-Benchを紹介します。このベンチマークは、3 つのドメイン、4 つの人種および 2 つの性別カテゴリー、および生活保護受給理論から導き出されたニーズの 5 つの因果関係にわたる、人間が作成したテンプレート (130 万件の実際の GoFundMe キャンペーンに対して調整済み) から作成された 600 件の経済援助リクエストで構成されています。 14 のモデルにわたって、監査形式によりバイアスの方向が変わります。モデルは、申請者を個別に評価する場合は少数派に有利ですが、並べてランク付けする場合は一部のグループに不利益を与えます。バイアスの大きさは、全体としては小さいものの、透明な監査よりも偽装監査の方が数倍大きく、申立人の名前だけが異なる控訴に直面して、圧倒的に資金を均等に分割するモデルが採用されている。対照的に、因果的フレーミング効果は人口統計効果をおよそ 1 桁上回っており、モデルや監査形式全体で一貫しており、現在の LLM が人間のふさわしい評価を確実に再現していることを示しています。ベンチマーク スコア モデルは 4 つの基準 (人口統計上の偏り、価値の調整、タスク間の一貫性、およびコンテキスト間の一貫性) に基づいて公開されており、他の実質的な領域に容易に適用できます。
原文 (English)
FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need derived from welfare deservingness theory. Across 14 models, audit format changes the direction of bias: models advantage minorities when rating claimants individually but penalize some groups when ranking them side by side. Bias magnitude, though small overall, is several times greater in disguised audits than in transparent ones, where, faced with appeals differing only in claimants' names, models overwhelmingly split funds equally. Causal framing effects, by contrast, exceed demographic effects by roughly an order of magnitude and are consistent across models and audit formats, indicating that current LLMs robustly reproduce human deservingness evaluations. The benchmark scores models on four criteria (demographic bias, deservingness alignment, cross-task consistency, and cross-context consistency), is publicly available, and can be readily adapted to other substantive domains.
DiffAttack: Evasion Attacks Against Face Recognition via Latent Diffusion Models
Facial biometric identification relies on the distinctiveness of user attributes within a high-dimensional embedding space. However, the de…
Retrieval-Driven Training-Free AI-Generated Video Attribution
AI-generated videos are becoming increasingly realistic and difficult to distinguish from authentic ones, which facilitates malicious misus…
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates
Adversarial training is one of the most effective defenses against adversarial attacks, yet the computational cost remains prohibitive at m…
A robust association between LLM use and scientific productivity: Assessing stopping-time selection
Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged i…
RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images
The rapid advancement of image generation models has made it increasingly difficult for people to distinguish AI-generated images from real…
PARALLEL: A Prefrontal-Aligned Reinforcement inspired Approach for Language-Model Learning under Explicit Limits
Recent language models achieve strong performance across a variety of tasks, but conventional adaptation applies updates uniformly across t…
Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning
Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only…
Point2Radio: A Foundation Model for Cross-Scene Radio Fields from Material-Aware Point Clouds
High-fidelity radio fields are typically simulated for every scene--transmitter configuration or fitted separately to each scene, failing t…
Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving
Existing autonomous-driving world models typically perform dense prediction of future videos, occupancy states, BEV representations, or age…
Improving scDiffusion with Sparsity-Biased Classifier-Free Guidance
Single-cell RNA sequencing (scRNA-seq) has become an essential tool in modern cellular biology, and generating accurate synthetic scRNA-seq…
Learning Lookahead Lemmas for Neural Network Verification
State-of-the-art neural network verifiers use the branch-and-bound procedure as their core solving mechanism. We introduce an inprocessing…
Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search
Multi-agent systems (MAS) are increasingly deployed to solve complex tasks. In case of incorrect or unsatisfactory outputs, users have to m…
Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives
Police crash narratives contain information that may supplement structured crash databases, but manual review is labor-intensive and it rem…
Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art
Deception detection has critical implications for legal proceedings, law enforcement, and online security. Although human judgment is limit…
Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients
Federated learning of foundation models faces a fundamental resource-asymmetry challenge: the institutions holding the most valuable domain…
metasignal: A Python Package for Comprehensive Metacognitive Analysis and Decision-Making
Metasignal is an open-source Python package for signal detection theory (SDT) and metacognitive measurement. It implements the 17 metacogni…
DoubleHelix: Structured Cross-Modal Fusion for Audio-Visual Speech Recognition with LLMs
Audio-visual speech recognition (AVSR) relies on effective fusion of audio and visual modalities, yet existing approaches treat cross-modal…
Multi-Granularity Position Embedding of Graphs via Granular-Ball for Link Prediction
Link prediction aims to identify potential or future connections within a given graph structure. Position information is essential for link…
InferQ: A Database-Oriented Benchmark for Quantum Circuits Simulation
Recent work suggests that relational database management systems (RDBMSs) can execute quantum circuit simulation by compiling the simulatio…
HERO: History-Enriched Rollout Training for Long-Horizon Autoregressive Neural Operators
Neural operators provide fast surrogates for time-dependent partial differential equations (PDEs) by applying a learned evolution operator…
Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership
Synthetic face datasets are increasingly used to reduce privacy exposure and data access constraints in biometric recognition. Yet the gene…
Implicit Machine Learning Force Fields Accelerate Molecular Dynamics Simulations
We introduce implicit machine learning force fields (I-MLFFs), which replace explicit stacks of neural network layers with self-consistent…
Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent Memory
Long-term memory lets large language model(LLM) agents reuse prior preferences and work flows, but it also turns untrusted observations int…
ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency
Vision-language-action (VLA) policies achieve strong performance in robotic manipulation but remain vulnerable to runtime disturbances that…
CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning
While robot foundation models are growing increasingly capable, the strongest models are typically trained on proprietary data and remain c…
MBDiff: Multi-view Behavior-aware Diffusion Model for Probabilistic Utility Data Imputation
Utility data (e.g., electricity, water, and gas consumption), collected by ubiquitous sensors and embedded devices, often contains substant…
MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation
Text-to-motion generation must produce motions that are semantically correct, temporally coherent, and physically plausible. A natural appr…
SERUM: State Extraction and Refinement for User Modeling
Agentic assistants capable of proactive, personalized interactions require structured models of user intent and workflow. However, building…
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillatio…
MOSAIC: Masked Outsourcing of Secure AI Computations
We address the challenge of securely and efficiently outsourcing AI computations from a trusted but computationally weak client to an untru…
Linear Proposal Operators and Stochastic Search Geometry in SOMA and Differential Evolution
Swarm and evolutionary algorithms are usually analyzed as complete procedural systems in which nonlinear selection, replacement, and adapta…
FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution
Although world-action models (WAMs) enhance long-horizon robot control by predicting visual evolution before acting, long-horizon reliabili…
Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters
InMyStyle is a privacy first, single user system that adapts small language models to rewrite AI-edited text towards an individual user's w…
When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration
In vision--language models, commonsense-driven hallucination (CDH) occurs when a model's commonsense prior overrides clear visual evidence…
RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems
Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-stra…
TAVI-TEC: An AI-Based Tool for Procedural Planning of Transcatheter Aortic Valve Implantation
Computed tomography angiography (CTA) is crucial for preprocedural TAVI planning, providing the anatomical information required for prosthe…
CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation
Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing…
OsteoCAD: A Human-in-the-Loop Cloud-Edge Framework for Bone Tumor Segmentation
Artificial Intelligence (AI) and Deep Learning (DL) have notably advanced medical image analysis, yet many health- care organizations strug…
Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation
Multi-domain machine translation (MDMT) poses a unique challenge due to varying levels of linguistic complexity across domains. Inspired by…
The persuasive power of large language models does not depend on their perceived national origin
Conversational AI developed by geopolitical rivals reaches citizens worldwide, raising concerns that it could sway public opinion or be rej…
DualDiT: A Conditional Dual-Output Diffusion Transformer for Joint OCT Image and Segmentation Mask Generation
Background and Objective: Generating realistic medical images with anatomically accurate segmentation masks helps address the shortage of a…
SeekBrain: An Autonomous Multi-Agent System for Accelerating Neuroscience Discovery
Modern neuroscience relies on integrating multi-scale, multimodal datasets to uncover the neural principles underlying intelligence. Howeve…
Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning
With the ever-increasing pervasiveness of smart edge devices, the demand is growing for applications that can be tailored to users (e.g., c…
Cross-Lingual Transfer for Machine Translation in Turkic Languages
Cross-lingual transfer is central to low-resource machine translation, but its behavior within closely related language families remains in…
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
Balancing sequence length, representational capacity, and long-horizon stability is a central problem in autoregressive (AR) speech and aud…
Dense Temporal Contrast Synthesis via Conditioned Latent Transport
Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is essential for breast cancer management, but reliance on gadolinium-based…
Explore Beyond the Boundary Using Entropic Information
In reinforcement learning, exploration with sparse and delayed rewards presents a significant challenge due to the limited feedback availab…
AgenticRepair: Multi-Faceted Program Context Engineering for Agentic Vulnerability Repair
Automated vulnerability repair aims to reduce the time and effort required to patch security flaws from a vulnerability triage report. Rece…
QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models
Infrared vision-language models (IR-VLMs) extend thermal perception to open-vocabulary classification, image captioning, and visual questio…
TFGformer: Multivariate Time Series Forecasting via Time-Frequency Graph Learning and Covariate Fusion
Large-scale multivariate time series from heterogeneous IoT sensors demand accurate long-term forecasting for resource scheduling and predi…
DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search
Reinforcement-learning-based quantum architecture search (RL-QAS) repeatedly optimizes a variational quantum eigensolver (VQE) after extend…
From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale
AI coding agents are generating code at volumes that exceed the capacity of traditional peer review. At the same time, existing AI code rev…
TerraNova: A Foundation Model for the Anthropocene
A defining problem of the Anthropocene is to model the physical Earth and human societies as one coupled system, yet no learned representat…
ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation
Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While pri…
MOT-SR: Multi-Objective Tool-Augmented Scientific Equation Discovery with Large Language Models
Symbolic Regression (SR) aims to discover analytical equations from observational data and plays a central role in scientific modeling. Whi…
TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning
The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-output examples and ap…
FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two peop…
A Human-Centered Validation of the Explainability-Performance Coefficient
The rapid adoption of deep learning models in high-risk domains has intensified the need for trustworthy Explainable Artificial Intelligenc…
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning
Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to lang…
CENDRe: Concept Extraction with Natural Domain Representations
Convolutional neural networks (CNNs) are widely used for time-series classification, but their deployment in critical domains requires unde…
The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations
Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and obscures diagnostic fee…
SATViz: Real-Time Visualization of Clausal Proofs
Visual layouts of graphs representing SAT instances can highlight the community structure of SAT instances. The community structure of SAT…
Combining Large Language Models and Symbolic Reasoning for Multi-Robot Temporal Planning through Explainable Knowledge Bases
We present PLANTOR, a framework for generating and executing multi-robot task plans from natural-language task descriptions through LLM-ass…
Shall We Play a Game? Language Models for Open-ended Wargames
LLM-based social simulations can make a generated transcript look like a single behavioral signal, but the model behind that transcript may…
Embedded Universal Predictive Intelligence: a coherent framework for multi-agent learning
The standard theory of model-free reinforcement learning assumes that the environment dynamics are stationary and that agents are decoupled…
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents
Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost univer…
M3MAD-Bench: Multi-Dimensional Evaluation of Multi-Agent Debate Across Domains and Modalities
As an agent-level reasoning and coordination paradigm, Multi-Agent Debate (MAD) orchestrates multiple agents through structured debate to i…
RAPiD: Reward-Guided Consistency Distillation of Diffusion Planners for Real-Time Autonomous Driving
Diffusion-based trajectory planners can model multi-modal driving behavior, but their iterative denoising process introduces a latency bott…
Shaping Scientific Explanations to Expert Perspectives with Persona-Conditioned Reinforcement Learning
Explainable AI is increasingly important to scientific discovery. However, existing methods largely ignore that explanation quality is not…
What Makes a Sale? Simulating End-to-End Seller--Buyer Retail Dynamics with LLM Agents
Evaluating retail strategies before deployment is difficult, as outcomes are determined across multiple stages, from seller-side persuasion…
PEMAND: Persona-Enriched Multi-Agent Negotiation for Household Decision-Making
Modeling household-level decisions is central to many real-world applications, including trip planning, residential mobility and migration,…
SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios
AI agents are increasingly used to diagnose and mitigate failures in production systems, known as agentic Site Reliability Engineering (SRE…
Dual-Dimensional Consistency: Balancing Budget and Quality in Adaptive Inference-Time Scaling
Large Language Models (LLMs) have demonstrated remarkable abilities in reasoning. However, maximizing their potential through inference-tim…
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
A significant hurdle for current LLMs is the execution of complex, multi-stage tasks. Group Relative Policy Optimization (GRPO) has been em…
The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models
Recent works show that LLM agents struggle to correct errors in their own reasoning traces, despite their ability to correct errors from ex…
FEA-AI ハイブリッド アプローチによる IPMSM 設計最適化のためのマルチエージェント システム
内部永久磁石同期モーター (IPMSM) の設計では、相反する目的とマルチフィジックス制約のバランスをとる必要がありますが、最新の最適化ワークフローは 3 つのボトルネックに直面しています。それは、手動による問題設定、高い有限要素解析 (FEA) コスト、まばらな領域または分布外領域での信頼性の低いサロゲート ベースの検索です。これらの制限に対処するために、構造化された問題定義のための検索拡張生成 (RAG) と不確実性を考慮した FEA-AI ハイブリッド最適化パイプラインを統合する、エンドツーエンドの自動 IPMSM 設計最適化フレームワークを提案します。 RAG を通じてモーターの教科書に接続された設計エージェントは、ドメイン知識ベースのオプションとエンジニアリングのヒントを提供し、AI モデルのトレーニングのための最適化カードと実験計画計画を作成します。トレーニング エージェントは、電磁 FEA を自動化し、ジオメトリ検証とソルバー障害ログを記録し、ANOVA ベースのデータ分析と LLM 推論を使用して障害のあるジオメトリを分析し、設計サンプリング エージェントを呼び出して設計空間を再定義し、追加のサンプルを生成します。最適化エージェントは、不確実性主導のスイッチングを使用して GA ベースの検索を実行します。不確実性の低い候補は AI サロゲート推論によって評価されますが、不確実性が高く信頼性が重要なパレート フロントまたはトップ K の候補は高忠実度 FEA によって修正され、反復再トレーニングに再利用されます。このフレームワークは、経験に依存した手動の構成を、計算コストと予測の信頼性のバランスをとる再現可能なワークフローに変換します。一致した高忠実度 FEA 予算の下での実験結果は、提案されたハイブリッド アプローチが、低く、さらに削減可能な予測不確実性を維持しながら、より優れた目標パフォーマンスを達成し、早期の予算枯渇によって制限される FEA のみの探索や、信頼性の低い最適値に収束する AI のみの探索よりも優れたパフォーマンスを達成することを示しています。
原文 (English)
A Multi-Agent System for Motor Design Optimization via an FEA-AI Hybrid Approach
This study presents a large language model (LLM)-based multi-agent framework for interior permanent magnet synchronous motor (IPMSM) design optimization that mitigates limitations of conventional workflows: expertise-dependent problem setup and data preparation, the prohibitive computational cost of finite element analysis (FEA), and the unreliability of AI surrogates in unexplored regions. To this end, we first introduce a Design agent that formulates the optimization problem in natural language, leveraging retrieval-augmented generation to improve answer accuracy on motor design problems from below 50% to 67-80%. Furthermore, a Training agent autonomously repairs improperly defined design spaces by reasoning over solver failure history, raising the success ratio of the geometry sampling from 28% to 84% for AI training. Additionally, to resolve cost and reliability simultaneously, an Optimization agent employs an uncertainty-aware FEA-AI hybrid model: the AI surrogate is the primary evaluator, and FEA is selectively invoked where predictive uncertainty is high. Under the same FEA budget, this hybrid model achieves up to 44% lower iron loss in single-objective and 22.5% higher hypervolume in multi-objective optimization than conventional FEA-only search. Under the same evaluation budget, it reduces computation time by 52-55% while retaining 90-92% of FEA-only hypervolume. Conversely, AI-only search converges to false optima, leaving half its Pareto designs infeasible. Notably, a controller agent adaptively updates the uncertainty threshold that triggers FEA each round, eliminating manual tuning and achieving 5.8% lower single objective iron loss than with a fixed threshold. These results establish domain specialized LLM agents with uncertainty-aware hybrid evaluation as a reliable, scalable paradigm for simulation-driven design automation.
ロールエージェント: デュアルロール進化による LLM エージェントのブートストラップ
大規模言語モデル (LLM) エージェントは複雑なタスクで優れたパフォーマンスを示していますが、その学習は非効率なインタラクション フィードバックや静的トレーニング環境によって制限されることが多く、広範な一般化が妨げられます。これらの制限に対処するために、このホワイトペーパーでは、単一の LLM を利用してエージェントと環境の両方として同時に機能し、ブートストラップ型の共進化を可能にする、Role-Agent、\textcolor{black}{フレームワーク} を紹介します。ロール エージェントは、ワールド イン エージェント (WIA) とエージェント イン ワールド (AIW) の 2 つの相乗コンポーネントで構成されます。 WIA では、LLM がエージェントとして機能し、各アクションの後の将来の状態を予測します。予測された状態と実際の状態の調整はプロセスの報酬として使用され、環境を意識した推論を促進します。 AIW では、LLM が失敗した軌跡から失敗モードを分析し、同様の失敗パターンを持つタスクを取得します。これにより、目標を絞った実践のためにトレーニング データの分布が再形成されます。複数のベンチマークの実験では、Role-Agent が一貫してパフォーマンスを向上させ、強力なベースラインに対して平均 4\% 以上の向上をもたらしていることが示されています。
原文 (English)
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
Although Large Language Model (LLM) agents have demonstrated strong performance on complex tasks, their learning is often limited by inefficient interaction feedback and static training environments, which hinder broader generalization. To address these limitations, this paper introduces Role-Agent, \textcolor{black}{a framework} that harnesses a single LLM to function concurrently as both the agent and the environment, enabling a bootstrapped co-evolution. Role-Agent comprises two synergistic components: World-In-Agent (WIA) and Agent-In-World (AIW). In WIA, the LLM acts as the agent and predicts future states after each action; the alignment between predicted and actual states is then used as a process reward, encouraging environment-aware reasoning. In AIW, the LLM analyzes failure modes from failed trajectories and retrieves tasks with similar failure patterns, thereby reshaping the training data distribution for targeted practice. Experiments on multiple benchmarks show that Role-Agent consistently improves performance, yielding an average gain of over 4\% over strong baselines.
ReSum: LLM 推論と要約と強化学習の相乗効果
検証可能な報酬による強化学習 (RLVR) は、大規模言語モデル (LLM) における長期的な推論を改善するための中心的な手法です。ただし、既存の RLVR 手法では、推論の展開が不必要に長くなり、推論の一貫性が低下し、利用可能なコンテキスト バジェットが使い果たされる可能性があります。ロングコンテキストの組織化に対する既存のアプローチは、多くの場合、モデルが独自の推論軌道を管理できるようにするのではなく、ロールアウトを組織化する外部メカニズムに依存しています。この制限に対処するために、LLM が自己要約を通じて推論の軌跡を圧縮して整理できるようにする新しい RLVR フレームワークである ReSum を提案します。私たちのパイロット研究では、自己要約がトークンレベルのエントロピーを下げることで生成を安定化させ、「要約」フレーズを導入することで不正なロールアウトプレフィックスから伝播するエラーを大幅に軽減できることが示されています。これらの発見に動機づけられて、ReSum は、自己要約が進行中の推論プロセスに利益をもたらすかどうかを対照的に評価する、要約を意識した適応型ロールアウト メカニズムを採用しています。具体的には、モデルが自発的に自己要約をトリガーすると、ReSum は要約フレーズをマスクして対照的な分岐を作成します。非要約位置の場合は、代わりにフレーズをランダムに挿入して、一致したブランチを作成します。さらに、要約を意識した利点を設計して、対照的なロールアウト軌跡間のより詳細な比較を可能にします。広範な実験により、ReSum はロールアウトの長さを 18.6\% 短縮しながら、パフォーマンスを平均 4\% 向上させることが示されました。
原文 (English)
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning
Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs). However, existing RLVR methods often encourage unnecessarily long reasoning rollouts, which can degrade reasoning coherence and exhaust the available context budget. Existing approaches to long-context organization often depend on external mechanisms to organize rollouts, rather than enabling the model to manage its own reasoning trajectory. To address this limitation, we propose ReSum, a novel RLVR framework that enables LLMs to compress and organize their reasoning trajectories through self-summarization. Our pilot studies show that self-summarization stabilizes generation by lowering token-level entropy, and that introducing a ``summarization'' phrase can substantially mitigate errors propagated from an incorrect rollout prefix. Motivated by these findings, ReSum adopts a summarization-aware adaptive rollout mechanism that contrastively evaluates whether self-summarization benefits the ongoing reasoning process. Specifically, when the model spontaneously triggers self-summarization, ReSum masks the summarization phrase to create a contrastive branch; for non-summarization positions, it instead randomly injects the phrase to create a matched branch. We further design a summarization-aware advantage to enable finer-grained comparison between contrastive rollout trajectories. Extensive experiments show that ReSum improves performance at an average of 4\% while reducing rollout length by 18.6\%.
プロセスレベルの社会的影響評価のための認知世界モデル
社会的影響ダイアログは、内部の認知状態を変えることでユーザーの行動を変えます。評価の中心となる質問は、ユーザーの信念、欲望、意図、感情が会話の過程で測定可能なほど変化するかどうかであり、これは表面レベルのテキスト指標 (BLEU/ROUGE) や単一スコアの LLM 判定では捉えることができないプロセス指向の基準です。我々は \textbf{Cog}nitive \textbf{W}orld \textbf{M}odel \textbf{(CogWM)} を提案します。これは、マルチターン対話評価を「ユーザーが何を言ったか」から「ユーザーの内部認知状態がどのように進化したか」に再構成する LLM ベースのユーザー モデルです。CogWM は、BDI/E 認知状態とユーザー発話を共同で予測し、3 層を使用してユーザー シミュレーターと評価プラットフォームの両方として機能します。ターンレベルの忠実度、軌道レベルの状態ダイナミクス、タスクレベルの複合スコアリングをカバーする評価フレームワーク。 4 つの社会的影響シナリオにわたる 150,454 のユーザー ターン サンプルで \textbf{S}ummarize-\textbf{a}nd-\textbf{A}llocate \textbf{(SaA)} アノテーション パイプラインを介してトレーニングされた CogWM は、77.6\% の感情精度 (GPT-5.5 の 2.1$\time$) を達成しました。 3,600 件のマルチエージェント識別試験において、認知的影響力によって 6 つの営利エージェントを区別し、Llama-4-Scout が 1 位にランクされました (CTS +0.233)。 CogWM は、社会的影響対話の評価を最終的な判断からプロセスの追跡に移行します。コード\脚注{\scriptsize コード: https://github.com/lucianma05-create/CogWM} とモデル\脚注{モデル: https://www.modelscope.cn/models/LucianMa/CogWM-14B} をリリースしました。
原文 (English)
Cognitive World Model for Progressive BDI/E Trajectory Evaluation of Conversational Agents
As LLM-based conversational agents advance toward increasingly open-ended and interaction-intensive scenarios, task completion alone provides an incomplete assessment of their effectiveness. The evolution of users' internal states, including beliefs, desires, intentions, and emotions (BDI/E), serves as an intermediate signal connecting agent behaviors with interaction outcomes and reflects how conversational strategies shape users during multi-turn interactions. However, existing evaluation paradigms primarily focus on surface-level responses or final outcomes, providing limited insight into the underlying cognitive processes. This limitation makes it difficult to diagnose why agents succeed or fail and to optimize their interaction strategies. To address this challenge, we propose Cognitive World Model (CogWM), an LLM-based cognitive user model that jointly models users' BDI/E states and corresponding responses, enabling explicit cognitive trajectory tracking. Trained on 150K user-turn samples with Qwen3-14B, CogWM achieves superior performance over existing user simulation baselines in both response fidelity and cognitive state understanding. Interactions with six state-of-the-art LLMs demonstrate that CogWM enables progressive comparison of agents through cognitive trajectories, revealing distinct agent patterns and complementary relationships between cognitive evolution and behavioral outcomes.
EvalSafetyGap: LLM 評価と安全性の失敗に関するハイブリッド調査と概念的なフレームワーク
LLM の評価と AI の安全性は、共通の測定問題に直面しています。つまり、ベンチマーク スコア、報酬モデルのシグナル、報告される安全性メトリクスは向上する可能性がありますが、それらが表現するはずの潜在的な特性の検証は依然として困難です。この文書では、ハイブリッド調査 (物語の合成と個別に追跡される灰色の証拠と組み合わせた体系的な調査) を、概念的なフレームワークおよび構造化された 10 モデルの監査と組み合わせています。この統合は、ベンチマークの有効性、動的評価、裁判官としての LLM の信頼性、安全性評価、ジェイルブレイク/拒否の堅牢性、報酬ハッキング、機構の解釈可能性、ガバナンス/監査可能性の 8 つの証拠ストリームに及び、2018 年から 2026 年の評価安全性測定作業をカバーします。最適化の圧力下で評価側とアライメント側のプロキシ障害を比較するための組織化仮説として EvalSafetyGap を導入します。グッドハートの法則と、ここで開発した 2 つの構成要素 (不安定性分解とアライメントのトリレンマ) をテスト可能な比較を生成するツールとして使用します。この監査は、能力、行動安全性、ガバナンスを個別に測定した場合に結論がどのように変化するかを示しています。このサンプル (n = 10) では、表示された表 3 の入力を使用すると、能力と持続的な敵対的堅牢性の間の関連性は統計的に不確定であり (ピアソン r = +0.232、p = 0.520)、見かけ上のオープンとクローズの安全性ギャップは控えめであり、動作の堅牢性よりも主にガバナンスと開示によって左右され、単一の境界線モデルがどのように分類されるかに影響されます。試行予算の結果はプロトコルに依存します。公的証拠では異種プロトコルが使用されているため、監査はランク付けではなく診断的なものになります。この貢献は、動的評価、透明性のあるソースレポート、複数回の安全性測定、および監査可能な調整の実践をサポートするための共有ボキャブラリーと証拠マップです。
原文 (English)
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures
This paper presents a systematic survey and conceptual synthesis of the shared measurement problem underlying large language model (LLM) evaluation and AI safety: benchmark scores, reward signals, and safety metrics can improve while the capabilities and alignment properties they are meant to represent remain uncertain. Synthesizing 373 primary studies published between 2018 and 2026, the survey organizes evidence on benchmark validity, contamination, dynamic evaluation, LLM-as-a-judge protocols, adversarial safety testing, reward and proxy optimization, mechanistic interpretability, and AI governance into an eight-stream evidence taxonomy. Building on this synthesis, we introduce EvalSafetyGap, a conceptual framework that unifies benchmark-validity and alignment-failure research as a shared proxy-target divergence problem under optimization pressure, formalized through a Goodhart-inspired Instability Decomposition and an Alignment Trilemma. An exploratory ten-model public-evidence audit illustrates the framework by showing why capability, behavioral robustness, and governance disclosure should be reported as separate evidence layers rather than collapsed into a single safety score. The survey closes with a research agenda for dynamic and contamination-resistant benchmarks, pre-specified multi-attempt threat models, version-locked evaluation, transparent source reporting, and validated mechanistic safety indicators, offering researchers, model developers, and AI auditors a shared vocabulary for measurement-aware LLM safety evaluation.
Latent Actions from Factorized Transition Effects under Agent Ambiguity
Latent Action Models (LAMs) learn action-like proxies from observation. However, in multi-object or distractor-rich scenes, observations co…
LabGuard: 自然言語のラボルールを、身体化されたラボエージェントのランタイムガードに根付かせる
科学的に身体化されたエージェントは、実験室での手順を実行できるようになってきていますが、動的な実験室環境でこれらの手順を安全に実行することは依然として困難です。現在の安全アプローチでは、安全ルール、マニュアル、プロトコル、標準操作手順などの実験室の自然言語を機械チェック可能な実行時制約に変換する中間ステップが見落とされていることがよくあります。 LabGuard (Laboratory Guard) は、自然言語のラボ ルールを実行可能な仕様に根付かせ、ランタイム ガードとして展開する、言語から実行までの安全性スイートです。 LabGuard には、3 つのコア コンポーネントが含まれています。LabGuard-IR は、型指定された実行可能表現を定義します。 LabGuard-Bench は、203 のシード ラボラトリ ルールから拡張された 812 の教師ありアノテーションを提供します。 LabGuard-Grounder は、自然言語のラボルールを LabGuard-IR にマッピングします。結果として得られる IR インスタンスは、LabGuard パイプラインによって処理され、ランタイム モニターにコンパイルされ、コントローラーの境界に適用されます。実験では、LabGuard が目に見えないラボルール ソースに一般化し、79.4 のタスク スコープ F1 を達成し、モニターのコンパイル後に危険なイベントを 39.5% から 23.8% に削減することが示されています。 LabUtopia では、ランタイム モニターが ACT と統合されており、タスクの成功を維持しながら介入を 0.5% 未満に抑えます。
原文 (English)
LabGuard: Grounding Natural-Language Laboratory Rules into Runtime Guards for Embodied Laboratory Agents
Scientific embodied agents are increasingly capable of carrying out laboratory procedures, but executing these procedures safely in dynamic laboratory environments remains challenging. Current safety approaches often overlook the intermediate step of transforming laboratory natural language, including safety rules, manuals, protocols, and standard operating procedures, into machine-checkable runtime constraints. We introduce LabGuard (Laboratory Guard), a language-to-execution safety suite that grounds natural-language laboratory rules into executable specifications and deploys them as runtime guards. LabGuard includes three core components: LabGuard-IR, which defines a typed executable representation; LabGuard-Bench, which provides 812 supervised annotations expanded from 203 seed laboratory rules; and LabGuard-Grounder, which maps natural-language laboratory rules into LabGuard-IR. The resulting IR instances are handled by the LabGuard Pipeline, which compiles them into runtime monitors and applies them at the controller boundary. Experiments show that LabGuard generalizes to unseen laboratory-rule sources, achieves 79.4 task-scope F1, and reduces unsafe events from 39.5% to 23.8% after monitor compilation. In LabUtopia, its runtime monitors integrate with ACT, keeping interventions below 0.5% while preserving task success.
飛行中の航空交通管制をサポートするソリューション空間経路計画
技術の進歩に伴い、航空交通管理用に多くの経路計画アルゴリズムが提案されていますが、戦術管制における運用上の採用は依然として限定的であり、アルゴリズム設計の優先順位と航空管制官のニーズとの間のずれが明らかになりました。これは、本質的に解釈可能で、計算効率が高く、人間が使用するために明示的に設計された意思決定支援ソリューションの必要性を強調しています。この設計課題に焦点を当て、この研究では、次の 2 つの指針となる考慮事項に適合するように設計された、飛行中の航空交通管制 (ATC) のための競合のない経路計画アルゴリズムを開発します。(1) ソリューション空間表示によって提供される解釈可能性と柔軟性。これは、実行可能なすべての安全な行動を明らかにし、変化する最適化目標に対応するアルゴリズムを構築する動機となります。 (2) 決定ロジック コントローラーは、分離基準、操縦性の制限、ウェイポイントの最小化、ルートの実用性などの運用上の制約を強制するときに自然に適用されます。このアルゴリズムは、これらの原則に基づいて、距離ベース、時間間隔ベース、ゾーンベースの 3 つのインテントベースの競合検出方法をソリューション空間フレームワーク内に統合し、計算効率の高い方法で競合のないパスを特定します。さらに、頂点ベースおよびエッジベースの検索ノードが解空間パス プランニング (SSPP) 用に提案されており、その結果、それぞれ SSPPV と SSPPE という 2 つのバリアントが生成され、計算速度と解の品質の観点から評価されます。実験結果によると、ゾーンベースの競合検出と組み合わせた SSPPV は最高のパフォーマンスを実現し、5 nmi グリッドを使用したマーストリヒト上部地域管制センター (MUAC) のデルタ セクターに基づく運用関連シナリオでパスを平均 3.69 ミリ秒で計算します。
原文 (English)
Solution Space Path Planning: A Real-Time Human-Centered Path Planning Algorithm for En-Route Air Traffic Control
As technology advances, various algorithms have been proposed for air traffic management, yet their operational adoption in tactical control remains limited. This gap motivates a human-centered design emphasizing algorithmic interpretability, controller-relevant operational constraints, and real-time computation. Inspired by the interpretability and flexibility of solution-space displays, as well as by the decision logic controllers naturally apply when enforcing operational constraints, this study extends the solution-space concept to path planning and develops a fast conflict-free path-planning algorithm for en-route Air Traffic Control (ATC), termed Solution Space Path Planning (SSPP). The algorithm integrates three intent-based conflict detection methods---distance-based, time-interval-based, and zone-based---within the solution-space framework to identify conflict-free paths in computationally efficient ways. SSPP is developed using both vertex-based and edge-based search nodes, resulting in two variants---SSPPV and SSPPE, respectively. Empirical results show that SSPPV paired with zone-based conflict detection performs best, computing paths in 3.69 ms on average in the Dutch Delta sector using a 5 nmi grid. SSPPV remains approximately 3.77 times faster than SSPPE while offering competitive effectiveness, making it suitable for time-critical operations and interactive 'what-if' probing in real time. An extension to SSPPV and SSPPE further examines the trade-off between delay minimization and separation requirements, demonstrating the flexibility of SSPP in revising optimization objectives. This study not only proposes a novel path-planning algorithm but also shows how such algorithms can be designed to align with human use and operational requirements, supporting their integration into future ATC systems.
スケールではなくアクセス構造からの機能: ハイブリッド シーケンス モデルの下限と事前登録テスト
プラトニック表現仮説 (PRH) は、モデルがスケールするにつれて、異種ネットワークの表現が現実の共有モデルに収束すると考えています。私たちは、その続編であり境界である能力収束仮説 (CCH) を提案します。固定されたトークンごとの推論バジェットの下では、表現的収束は能力の収束を伴いません。代わりに、機能はクラス、つまりアクセス完全ハイブリッド、つまり圧縮 O(1) 状態チャネルとスケーラブルな逐語インデックス チャネルの両方を保持するアーキテクチャに向かって収束します。我々はそれを証人タスクである無限ストリームのニュートンのリンゴ問題に固定し、3 つのリソースの壁と名付けます。o(Nb) 状態アーキテクチャを禁止するシャノン壁、固定ウィンドウを禁止する水平線壁、および固定深さの注意のみの構成を禁止する回路壁 (TC0 != NC1 の条件付き)。明示的な分離可能性の仮定の下では、ハイブリッドは各壁の価格を支払うことによって 3 つすべてを横断するため、合成下では機能は厳密に超加法的になります。私たちは証明したものと推測したものを区別します。アクセス完全性の原則は情報理論の下限と事前に登録された実験に基づいていますが、フィールドレベルの収束傾向は経済学に基づいた推測です。我々は、データの前に凍結された基準に基づいて事前に登録された最初の小規模テストを報告します。予測されたシザーズギャップが測定され(64スカラー状態が1つのグローバルアテンション層を獲得すると、正確な検索誤差は0.994対0.000)、状態追跡分岐は登録された境界に到達し、結合証人は還元できない2チャネルの解決策を示します。 1 つの予測は方向が逆転して失敗したため、そのように報告されます。表現の収束はスケールによって自由に与えられます。機能の収束はアクセス構造ごとに購入する必要があります。
原文 (English)
The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale
The Platonic Representation Hypothesis (PRH) holds that as models scale, representations of heterogeneous networks converge toward a shared model of reality. We propose its sequel and boundary, the Capability Convergence Hypothesis (CCH): under a fixed per-token inference budget, representational convergence does not entail capability convergence. Capability instead converges toward a class, the access-complete hybrid: any architecture holding both a compressive O(1)-state channel and a scalable verbatim-index channel. We anchor it on a witness task, the Newton's-apple problem in an infinite stream, and name three resource walls: a Shannon wall barring any o(Nb)-state architecture, a horizon wall barring any fixed window, and a circuit wall barring fixed-depth attention-only composition (conditional on TC0 != NC1). Under an explicit separability assumption a hybrid crosses all three by paying each wall's price, so capability is strictly super-additive under composition. We separate what we prove from what we conjecture: the access-completeness principle rests on information-theoretic lower bounds and pre-registered experiments, while the field-level convergence trend is an economics-motivated conjecture. We report the first pre-registered small-scale tests under criteria frozen before the data: the predicted scissors gap is measured (exact-retrieval error 0.994 vs. 0.000 once a 64-scalar state gains one global-attention layer), the state-tracking bifurcation lands at the registered boundary, and a conjunction witness shows an irreducibly two-channel solution; one prediction failed with its direction reversed and is reported as such. Representational convergence is given freely by scale; capability convergence must be purchased by access structure.
NeurOWL: 不完全なOWLオントロジー推論のためのLLMベースのニューラルシンボリックフレームワーク
OWLオントロジーは、意味論的推論を可能にする形式的な知識表現フレームワークを提供し、ヘルスケアやバイオインフォマティクスなどの分野で広く採用されています。ただし、実際には、現実世界のオントロジーは不完全であることが多く、推論に課題が生じます。この研究では、基本的な包含推論の問題に焦点を当てます。つまり、不完全なオントロジーと候補 (非含意) 包摂が与えられた場合、その包含が意味的に妥当かどうかを判断し、そうであれば、潜在的な欠落している公理を含む論理的に適切な説明を提供します。このタスクは包含検証とオントロジーアブダクションを統合し、欠落している公理の事前定義された候補セットの必要性を取り除くことによって後者を一般化します。この包含推論の問題に対処するために、我々は、正式に定義された意味論と、大規模言語モデルとオントロジー埋め込みを介したテキスト意味論の両方を活用して、検証とアブダクションを共同で実行するエンドツーエンドの神経記号フレームワークである NeurOWL を提案します。私たちは、複数のドメインにわたる現実世界のオントロジーで NeurOWL を評価し、さまざまなドメインにわたって強力で堅牢なパフォーマンスを実証します。
原文 (English)
NeurOWL: An LLM-Based Neural-symbolic Framework for Incomplete OWL Ontology Reasoning
OWL ontologies provide a formal knowledge representation framework that enables semantic reasoning, and have been widely adopted across domains such as healthcare and bioinformatics. In practice, however, real-world ontologies are often incomplete, which pose challenges for reasoning. In this work, we focus on a fundamental subsumption reasoning problem: given an incomplete ontology and a candidate (non-entailed) subsumption, determine whether the subsumption is semantically plausible and, if so, providing a logically sound explanation containing potential missing axioms. This task unifies subsumption verification with ontology abduction, and generalizes the latter by removing the need for a predefined candidate set of missing axioms. To address this subsumption reasoning problem, we propose NeurOWL, an end-to-end neuro-symbolic framework that jointly performs verification and abduction, leveraging both formally defined semantics and textual semantics through Large Language Models and ontology embeddings. We evaluate NeurOWL on real-world ontologies across multiple domains, demonstrating strong and robust performance across different domains.
品質保証: VR OSCE における審査官クレームのマルチモーダル検証
客観的構造化臨床検査 (OSCE) は臨床能力を評価するためのゴールドスタンダードですが、採点は依然として検査者の主観、疲労、認知バイアスの影響を受けやすいです。評価者間統計による標準的な審査官の検証は、審査官の推論を分析したり、実際の事象に対する審査官の主張を検証したりしないため、誤りの原因に関する説明力に欠けています。そこで、我々は、ビデオ、VR ログ、俳優データから構築された実際の一連のイベントに対して審査官が主張した行動を比較することにより、バーチャル リアリティ (VR) 小児 OSCE における審査官の主張を検証するマルチモーダル フレームワークである品質アクション保証 (QAA) を導入します。 QAA は、アクションの位置特定とアクターのソースの帰属を実行する制約付きの時間的アクション アライメント モデルと、審査官の主張を抽出して記録と照合する大規模な言語モデルを組み合わせます。 5 分割相互検証を通じて、QAA は時間的アライメントに関して 99.2% $\pm$ 0.7% Actor F1 および 93.4% $\pm$ 1.9% W@16 を達成しました。全体的に、QAA は 70.0% の精度と 76.7% の再現率で検査官のミスを検出し、事実の正確性が 39.2% から 79.2% に向上し、より公平な OSCE 評価が可能になります。
原文 (English)
Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs
Objective Structured Clinical Examinations (OSCEs) are the gold standard for assessing clinical competence, yet scoring remains vulnerable to examiner subjectivity, fatigue, and cognitive bias. Standard examiner validation via inter-rater statistics lacks explanatory power regarding the source of errors, as it neither analyzes examiner reasoning nor verifies examiner claims against actual events. Thus, we introduce Quality Action Assurance (QAA), a multimodal framework that verifies examiner claims in Virtual Reality (VR) pediatric OSCEs by comparing actions claimed by examiners against a reference record of events constructed from video, VR logs, and actor annotations. QAA combines a constrained temporal action alignment model, which performs action localization and actor source attribution, with a large language model that extracts examiner claims and checks them against the record. Across a 5-fold cross-validation, QAA achieves 99.2\% $\pm$ 0.7\% Actor F1 and 93.4\% $\pm$ 1.9\% W@16 for temporal alignment. Overall, QAA detects examiner errors with 69.9\% precision and 76.7\% recall; in retrospective evaluation, correcting the detected errors raises the share of factually correct transcripts from 39.2\% to 79.2\%, supporting fairer OSCE quality assessment.
CodeRescue: コーディング エージェント向けの予算調整されたリカバリ ルーティング
コーディング エージェントは、失敗した試行が単なる不正解ではなく実用的なフィードバックを生成する実行可能環境で動作することが増えています。既存のコストを意識したシステムは通常、このような障害をカスケード決定として扱います。つまり、最初に安価なモデルを試し、その後、困難なケースをより強力でより高価なモデルにエスカレーションします。ただし、コーディングでは、実行フィードバックによって安価なモデルの回復がさらに価値のあるものになる可能性もあり、エージェントはいつより安価なコンピューティングを費やす必要があるのか、いつエスカレーションすべきなのかという予算計画上の導入の問題が生じます。この障害後の決定を異種アクションに対する回復ルーティングとして定式化し、実行ロールアウトから監視対象ルーターをトレーニングします。変化する予算の下でも同じルータを使用できるようにするために、再トレーニングなしで導入時のコストペナルティを選択し、交換可能性の下で限界予想コスト制御を提供するコンフォーマルリスクコントロール(CRC)レイヤーを追加します。 5 つのコーディング ベンチマークで継続的に失敗した場合、安価なリカバリとエスカレーションは相補的な成功パターンを示します。調整されたフロンティアは、固定アクション、プロンプト専用ルーター、バイナリ カスケード ベースラインよりも改善されています。メインの GPT-5.4-nano/GPT-5.4 設定では、平均回復コストの 35% を使用しながら、1 つの CRC 校正済みフロンティア ポイントが常時エスカレートの解決速度を超えています。コードは https://github.com/Qijia-He/agent-budget-control で入手できます。
原文 (English)
CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents
Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer. Existing cost-aware systems typically treat such failures as cascade decisions: try a cheap model first, then escalate hard cases to a stronger and more expensive model. In coding, however, execution feedback can also make further cheap-model recovery worthwhile, raising a budgeted deployment question: when should an agent spend more cheap compute, and when should it escalate? We formulate this post-failure decision as recovery routing over heterogeneous actions and train a supervised router from execution rollouts. To make the same router usable under changing budgets, we add a Conformal Risk Control (CRC) layer that selects a deployment-time cost penalty without retraining and provides marginal expected-cost control under exchangeability. Across held-out failures from five coding benchmarks, cheap recovery and escalation exhibit complementary success patterns. The calibrated frontier improves over fixed actions, prompt-only routers, and a binary cascade baseline; in the main GPT-5.4-nano/GPT-5.4 setting, one CRC-calibrated frontier point exceeds always-escalate solve rate while using 35% of its mean recovery cost. Code is available at https://github.com/Qijia-He/agent-budget-control.
AttriMem: エージェントの記憶学習のためのアトリビューションに基づくプロセス フィードバック
LLM エージェントにとって効果的な記憶は非常に重要ですが、それを効果的に構築するのは依然として困難です。メモリ構築ポリシーは、インタラクションが蓄積するにつれてどの情報を抽出、保存、更新、圧縮、または破棄するかを決定します。ヒューリスティック記憶手法は主観的なタスク固有のルールに依存しているため、下流の目標とずれたり、タスク間の適応性が制限されたりする可能性があります。対照的に、RL ベースの手法はタスクのフィードバックから学習しますが、主に結果レベルまたはモジュールレベルの報酬を使用します。これらの粗い信号はタスクの成功を示しますが、どの中間メモリの内容が最終的な答えをサポートしているかを特定できず、きめの細かいクレジット割り当てのボトルネックが生じます。ただし、このようなプロセス フィードバックの構築は、中間記憶の決定には固有のグラウンドトゥルース ターゲットが欠けている一方、適切なクレジットはエージェントの不確実な推論軌道によって変化するため、事前に指定できないため、非常に困難です。我々は、RL を使用してメモリ構築ポリシーを学習するためのアトリビューションに基づくプロセス フィードバック フレームワークである AttriMem を提案します。 AttriMem は、最終的な回答へのトークンレベルの貢献から得られるローカルな報酬でグローバルな結果報酬を強化します。長期対話型質問応答の実験では、AttriMem が検索ベース、ヒューリスティック、RL ベースのベースラインを上回り、ベンチマークと回答モデル全体で一般化され、RL の最適化が安定することが示されました。
原文 (English)
AttriMem: Attribution-Guided Process Feedback for Agent Memory Construction
Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract, store, update, compress, or discard as interactions accumulate. Heuristic memory methods rely on subjective, task-specific rules, which can misalign with downstream objectives and limit cross-task adaptability. RL-based methods, by contrast, learn from task feedback but mainly use outcome- or module-level rewards. These coarse signals indicate task success but cannot identify which intermediate memory contents support the final answer, creating a fine-grained credit-assignment bottleneck. However, constructing such process feedback is prohibitively difficult because intermediate memory decisions lack unique ground-truth targets, while the appropriate credit varies with the agent's uncertain reasoning trajectory and therefore cannot be specified in advance. We propose AttriMem, an attribution-guided process-feedback framework for learning memory-construction policies with RL. AttriMem augments the global outcome reward with local rewards derived from token-level contributions to the final answer. Experiments on long-horizon dialogue question answering show that AttriMem outperforms retrieval-based, heuristic, and RL-based baselines, generalizes across benchmarks and answer models, stabilizes RL optimization.
Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning
Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy…
DynaResize: 分割された LLM ポストトレーニングのためのランタイム GPU 再割り当て
RL ベースの LLM ポストトレーニングでは、別々の GPU リソース間でロールアウトとトレーニングがますます細分化されますが、静的 GPU パーティショニングでは、ロングテール ロールアウト レイテンシの下で深刻なパイプライン バブルが発生します。 DynaResize は、RL セマンティクスを変更せずに、ロールアウトとトレーニングの間で GPU を動的に切り替えてステージの実行時間のバランスを取る、ランタイム GPU 再割り当てシステムです。 DynaResize は、サイズ変更をきめ細かい操作に分解し、コミュニケーターの再利用、制限された状態のステージング、およびヒステリシスベースのサイズ変更を通じて、起動にクリティカルではない作業をクリティカル パスから削除します。実験結果によると、DynaResize は、最適な静的構成と比較して、エンドツーエンドのスループットを 66.5% 向上させ、合計実行時間を 33% 削減し、同時にロール切り替えオーバーヘッドの 27% を隠すことができます。
原文 (English)
DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training
RL-based LLM post-training increasingly disaggregates Rollout and Training across separate GPU resources, but static GPU partitioning suffers from severe pipeline bubbles under long-tail rollout latency. We present DynaResize, a runtime GPU reallocation system that dynamically switches GPUs between Rollout and Training to balance stage execution times without changing RL semantics. DynaResize decomposes resizing into fine-grained operations and removes non-startup-critical work from the critical path through communicator reuse, bounded state staging, and hysteresis-based resizing. Experimental results show that DynaResize can improve end-to-end throughput by 66.5% and reduce total execution time by 33% over the optimal static configuration, while hiding 27% of role-switching overhead.
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enab…
Reason-Mediated Behavioral Models for Auditing LLM Social Simulators
Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether s…
Information Processing by Neuron Populations in the Central Nervous System: A Theory of the Mathematical Structure of Data and Operations
In the mammalian central nervous system, neurons are organized into populations communicating by spike trains propagating along axonal bund…
On the Expressive Power of Sparse Geometric MPNNs
Motivated by applications in chemistry and other sciences, we study the expressive power of message-passing neural networks for geometric g…
Revisiting Multi-Permutation Equivariance through the Lens of Irreducible Representations
This paper explores the characterization of equivariant linear layers for representations of permutations and related groups. Unlike tradit…
Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook
We survey deepfake generation and detection techniques, covering all deepfake media types: image, video, audio and multimodal content. We i…
Dual-Force: Enhanced Offline Diversity Maximization under Imitation Constraints
Offline diversity maximization under imitation constraints can transform demonstration data into a set of distinct behavioral policies, imp…
Dimensionality reduction for homological stability and global structure preservation
We propose DiRe, a force-directed dimensionality reduction framework designed to preserve global structure and homological features while r…
Reproducing Human Individual Motor Signatures: A Data-Driven Approach for Repetitive Motion
The deployment of autonomous virtual avatars (in extended reality) and robots in human group activities---such as rehabilitation therapy, s…
StaQ: a Finite Memory Approach to Discrete Action Policy Mirror Descent
In Reinforcement Learning (RL), regularization with a Kullback-Leibler divergence that penalizes large deviations between successive polici…
Towards White-Box Deep Wireless Sensing
The empirical success of deep learning has spurred its application to the radio-frequency (RF) domain, leading to significant advances in D…
Patch-Based 3D Variational Autoencoder for Super-Resolution of Turbulent Channel Flow
Direct numerical simulation (DNS) accurately resolves all spatio-temporal scales of wall-bounded turbulence but becomes prohibitively expen…
RePaCA: Leveraging Reasoning Large Language Models for Static Automated Patch Correctness Assessment
Automated Program Repair (APR) seeks to automatically correct software bugs without requiring human intervention. However, existing tools t…
"Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness
Homelessness is a persistent social challenge, impacting millions worldwide. Over 876,000 people experiencing homelessness (PEH) were recor…
Adaptive Policy Backbone via Shared Network
Reinforcement learning (RL) has achieved impressive results across domains, yet learning an optimal policy typically requires extensive int…
Fast Feature Field ($\text{F}^3$): A Predictive Representation of Events
This paper develops a mathematical argument and algorithms for building representations of data from event-based cameras, that we call Fast…
Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
Recent medical vision-language models have shown promise on tasks such as VQA, report generation, and anomaly detection. However, most are…
Monotone and Separable Set Functions: Characterizations and Neural Models
Motivated by applications for set containment problems, we consider the following fundamental problem: can we design set-to-vector function…
Pay for The Second-Best Service: A Game-Theoretic Approach Against Dishonest LLM Providers
The widespread adoption of Large Language Models (LLMs) through Application Programming Interfaces (APIs) induces a critical vulnerability:…
Robust Bidirectional Associative Memory via Regularization Inspired by the Subspace Rotation Algorithm
Bidirectional Associative Memory (BAM) trained with Bidirectional Backpropagation (B-BP) often suffers from poor robustness and high sensit…
AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rat…
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
Web applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and…
GPU-Accelerated ANNS: Quantized for Speed, Built for Change
Approximate nearest neighbor search (ANNS) is a core problem in machine learning and information retrieval applications. GPUs offer a promi…
GeoRA: Geometry-Aware Low-Rank Adaptation for RLVR
Reinforcement Learning with Verifiable Rewards (RLVR) is a key paradigm for improving large-scale reasoning models. Unlike supervised fine-…
Knowledge Restoration-driven Prompt Optimization: Unlocking LLM Potential for Open-Domain Relational Triplet Extraction
Open-domain Relational Triplet Extraction (ORTE) aims to mine structured knowledge without predefined relation schemas. Large Language Mode…
When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering
Retrieval-Augmented Generation (RAG) extends large language models (LLMs) beyond parametric knowledge, yet it is unclear when iterative ret…
Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation
The recent advancements in Large Language Models (LLMs) have attracted interest in exploring their in-context learning abilities and chain-…
AIvilization v0: Toward Large-Scale Artificial Social Simulation with a Unified Agent Architecture and Adaptive Agent Profiles
AIvilization v0 is a publicly deployed large-scale artificial society that couples a resource-constrained sandbox with a unified LLM-agent…
Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization
Layer-wise capacity in large language models is highly non-uniform: some layers contribute disproportionately to loss reduction, whereas ot…
Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery
Recent advancements in Large Language Model (LLM) agents have demonstrated remarkable potential in automatic knowledge discovery. However,…
Stem: Rethinking Causal Information Flow in Sparse Attention
The quadratic computational complexity of self-attention remains a fundamental bottleneck for scaling Large Language Models (LLMs) to long…
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
We uncover a behavioral law of long-horizon vision-language models: models that maintain temporally grounded beliefs generalize better. Sta…
Wrong Code, Right Structure: Learning Netlist Representations from Imperfect LLM-Generated RTL
Learning effective netlist representations is fundamentally constrained by the scarcity of labeled datasets, as real designs are protected…
ELISA: An Interpretable Hybrid Generative AI Agent for Expression-Grounded Discovery in Single-Cell Genomics
Translating single-cell RNA sequencing (scRNA-seq) data into mechanistic biological hypotheses remains a critical bottleneck, as agentic AI…
Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation
Although debiased large language models (LLMs) excel at handling known or low-bias prompts, they often fail on unfamiliar and high-bias pro…
謎を解くビデオ推論
ビデオ生成の最近の進歩により、予期せぬ現象が明らかになりました。拡散ベースのビデオ モデルは、自明ではない推論機能を示します。以前の研究では、これはチェーン オブ フレーム (CoF) メカニズムによるものであり、推論はビデオ フレーム間で順次展開されると想定されています。この研究では、この仮定に異議を唱え、根本的に異なるメカニズムを明らかにします。ビデオ モデルにおける推論は、主に拡散ノイズ除去ステップに沿って現れることを示します。定性分析と対象を絞った調査実験を通じて、モデルは初期のノイズ除去ステップで複数の候補ソリューションを探索し、最終的な答えに徐々に収束することがわかりました。これは、Chain-of-Steps (CoS) と呼ばれるプロセスです。この中心的なメカニズムを超えて、モデルのパフォーマンスに重要ないくつかの新たな推論動作を特定します。(1) ワーキングメモリ。永続的な参照を可能にします。 (2) 自己修正と強化により、誤った中間ソリューションからの回復が可能になります。 (3) アクション前の認識。初期のステップで意味論的な基礎を確立し、後のステップで構造化された操作を実行します。拡散ステップ中に、拡散トランスフォーマー内の自己進化した機能的特殊化をさらに明らかにします。初期の層は高密度の知覚構造をエンコードし、中間の層は推論を実行し、後の層は潜在的な表現を統合します。これらの洞察に動機付けられて、私たちは概念実証としてトレーニング不要のシンプルな戦略を提示し、異なるランダム シードを持つ同一のモデルからの潜在軌道をアンサンブルすることによって推論がどのように改善されるかを実証します。全体として、私たちの研究は、ビデオ生成モデルで推論がどのように現れるかについて体系的な理解を提供し、インテリジェンスの新しい基盤としてビデオ モデルの固有の推論ダイナミクスをより効果的に活用する将来の研究を導くための基盤を提供します。
原文 (English)
Demystifying Video Reasoning
Recent advances in video generation have revealed an unexpected phenomenon: diffusion-based video models exhibit non-trivial reasoning capabilities. Prior work attributes this to a Chain-of-Frames (CoF) mechanism, where reasoning is assumed to unfold sequentially across video frames. In this work, we challenge this assumption and uncover a fundamentally different mechanism. We show that reasoning in video models instead primarily emerges along the diffusion denoising steps. Through qualitative analysis and targeted probing experiments, we find that models explore multiple candidate solutions in early denoising steps and progressively converge to a final answer, a process we term Chain-of-Steps (CoS). Beyond this core mechanism, we identify several emergent reasoning behaviors critical to model performance: (1) working memory that supports tasks requiring consistent reference, such as object permanence; (2) self-correction and enhancement, allowing recovery from incorrect intermediate solutions; and (3) perception before action, where early steps establish semantic grounding and later steps perform structured manipulation. Moreover, analysis of Diffusion Transformer layers shows that middle layers conduct key reasoning procedures. Motivated by these insights, we present a simple Training-Free Ensemble (TFE) as a proof-of-concept, demonstrating how reasoning can be improved by ensembling latent trajectories from identical models with different random seeds. Overall, our work provides the first systematic dissection of the mechanisms underlying video reasoning, offering a foundation to guide future research in better exploiting the inherent reasoning dynamics of video models as a new substrate for intelligence.
OPERA: Online Data Pruning for Efficient Retrieval Model Adaptation
Domain-specific finetuning is essential for dense retrievers, yet not all data pairs contribute equally to the learning process. We introdu…
Agentic Harness for Real-World Compilers
Compilers are critical to modern computing, yet fixing compiler bugs is difficult. While recent large language model (LLM) advancements ena…
Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning
Zero-shot reinforcement learning (RL) algorithms aim to learn a family of policies from a reward-free dataset, and recover optimal policies…
Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations
In collaboration with Alibaba, we study how a generative AI assistant affects service performance in e-commerce after-sales operations. In…
ActionParty: Multi-Subject Action Binding in Generative Video Games
Recent advances in video diffusion have enabled the development of "world models" capable of simulating interactive environments. However,…
Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation
We study compiled AI, a paradigm in which large language models generate executable code artifacts during a compilation phase, after which…
Evaluating the Alignment Between GeoAI Explanations and Domain Knowledge in Satellite-Based Flood Mapping
The increasing number of satellites has improved the temporal resolution of Earth observation, making satellite-based flood mapping a promi…
TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning
Time Series Foundation Models (TSFMs) have demonstrated strong generalization capability and data efficiency in time series forecasting thr…
幾何学的規制による LLM 生成におけるエスケープ モードの崩壊
モード崩壊は生成モデリングにおける永続的な課題であり、明示的なループから多様性の漸進的な喪失や時期尚早な軌道収束に至るまでの範囲の動作として自己回帰テキスト生成に現れます。私たちは力学システムの視点をとり、モード崩壊を *幾何学的崩壊* によって引き起こされる状態空間へのアクセス可能性の低下として再解釈します。生成中、モデルの内部軌道はその表現空間の低次元領域に限定されます。これは、モード崩壊が純粋にトークンレベルの現象ではなく、記号的制約や確率のみの復号ヒューリスティックでは確実に解決できないことを意味します。この観点に基づいて、私たちは、Transformer 値キャッシュ (低ランクのダンピングとして実装) 内の主要な自己強化方向を制御する軽量のオンライン状態空間介入である *強化モード制御* (RMR) を提案します。複数の大規模な言語モデルにわたって、RMR はモード崩壊を大幅に軽減し、非常に低いエントロピー レート (0.8 nats/ステップまで) での安定した生成を可能にしますが、標準のデコードでは通常 2.0 nats/ステップ近くで崩壊します。
原文 (English)
Escaping Mode Collapse in LLM Generation via Geometric Regulation
Mode collapse is a persistent challenge in generative modeling and appears in autoregressive text generation as behaviors ranging from explicit looping to gradual loss of diversity and premature trajectory convergence. We take a dynamical-systems view and reinterpret mode collapse as reduced state-space accessibility caused by *geometric collapse*: during generation, the model's internal trajectory becomes confined to a low-dimensional region of its representation space. This implies mode collapse is not purely a token-level phenomenon and cannot be reliably solved by symbolic constraints or probability-only decoding heuristics. Guided by this perspective, we propose *Reinforced Mode Regulation* (RMR), a lightweight, online state-space intervention that regulates dominant self-reinforcing directions in the Transformer value cache (implemented as low-rank damping). Across multiple large language models, RMR substantially reduces mode collapse and enables stable generation at extremely low entropy rates (down to 0.8 nats/step), whereas standard decoding typically collapses near 2.0 nats/step.
Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs
Diffusion-based Large Language Models (D-LLMs) represent a promising frontier in generative AI, offering fully parallel token generation th…
Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping
Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour-intensive and geogra…
Detecting AI-Generated Videos with Spiking Neural Networks
Modern AI-generated videos are photorealistic at the single-frame level, leaving inter-frame dynamics as the main remaining axis for detect…
A Nonlinear Singular Value Theory for Neural Networks
Recently Brown et al. [2025] established a singular value decomposition (SVD) for maps (especially nonlinear) satisfying certain norm condi…
DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain
LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by i…
Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning
Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where…
When Bits Break Recourse: Counterfactual-Faithful Quantization
Model quantization is widely used to reduce memory, latency, and deployment cost, and is typically judged by whether predictive accuracy is…
AI4BayesCode: From Natural Language Descriptions to Validated Modular Stateful Bayesian Samplers
Coding and computation remain major bottlenecks in Markov chain Monte Carlo (MCMC) workflows, especially as modern sampling algorithms have…
DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation
Autoregressive long video generation often adopts bounded-memory streaming for efficiency, typically combining local windows for short-term…
MemForest: 階層型時間インデックスを備えた効率的なエージェント メモリ システム
メモリは、ロングコンテキストの LLM エージェントを有効にするための基本コンポーネントであり、継続的な提供と更新のライフサイクルを通じて対話全体にわたる永続的な状態をサポートします。相当な事前作業にもかかわらず、既存のシステムは、粗粒度の状態管理と本質的に逐次的な更新パイプラインという 2 つの重要な制限により、重大なメンテナンスのオーバーヘッドに悩まされています。特に、更新は LLM 推論と密接に結びついていることが多く、完全な状態の書き換えが必要なため、スケーラビリティが低下し、メモリが蓄積するにつれて遅延が増大します。これらの課題に対処するために、エージェントのメモリを書き込み効率の高い時間データ管理問題として再定式化するメモリ フレームワークである MemForest を紹介します。 MemForest は、並列チャンク抽出によってシーケンシャル ボトルネックを解消し、メモリ構築を同時の独立した操作に分離します。粗粒度のメンテナンスをさらに排除するために、フラットなグローバル サマリーではなく時間順のツリーとしてメモリを編成する階層型時間インデックスである MemTree を導入します。この設計では、完全な状態の書き換えを局所的なノードごとの更新に置き換え、影響を受けるツリー パスのメンテナンス コストを削減しながら、時間的に変化する状態を自然に保存します。私たちは、LongMemEval-S と LoCoMo という 2 つのロングコンテキスト メモリ ベンチマークで MemForest を評価します。 LongMemEval-S では、MemForest はステートフル ベースラインの中で最高の総合パフォーマンスを達成し、EverMemOS を含む最先端のアプローチよりも約 6 倍高いメモリ構築スループットを維持しながら、79.8% pass@1 精度に達します。
原文 (English)
MemForest: An Efficient Agent Memory System with Hierarchical Temporal Indexing
Memory is a fundamental component for long-context LLM agents, supporting persistent state across interactions through a continuous serve-and-update lifecycle. Despite substantial prior work, many stateful systems retain sequential autoregressive extraction or state-dependent maintenance on the write path, delaying when new evidence becomes queryable. To address these challenges, we present MemForest, a memory framework that reformulates agent memory as a write-efficient temporal data-management problem. MemForest breaks the sequential bottleneck via parallel extraction, decoupling memory construction into concurrent, independent operations. We further introduce MemTree, a hierarchical temporal index that organizes memory as time-ordered trees and replaces global rewrites with localized dirty-path refresh. Dirty summaries can be refreshed in parallel across nodes and trees. End-to-end work remains proportional to incoming content; the logarithmic bound applies only to structural insertion and level-dependent refresh depth in balanced trees. We evaluate MemForest on two long-context benchmarks, LongMemEval-S and LoCoMo. Experiments use Qwen3-4B, Qwen3-30B, and Gemma-4-12B-IT. With Qwen3-30B, MemForest reaches 81.8 percent pass at 1 on LongMemEval-S, while its input-normalized build rate is 6.0 times that of EverMemOS. On LoCoMo categories 1 to 4, it reaches 84.09 percent, within 0.13 percentage points of EverMemOS; on a matched conversation, its build rate is 9.5 times higher. These results show that MemForest reduces memory-freshness latency while retaining strong answer quality.
PEFT of SLM for Telecommunications Customer Support: A Comparative Study of LoRA Configurations with Energy Consumption Analysis
While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptatio…
THzデュアルコム分光法を使用したポリマー分類のためのマルチスケール機能アテンションネットワーク
信頼性の高いポリマーの識別は、リサイクルプラスチックの品質と安全性を確保するために不可欠ですが、従来の分別技術や分光技術では、確実な識別を実現するのが困難なことがよくあります。テラヘルツ デュアルコム分光法 (THz-DCS) は、迅速、高分解能、非破壊測定を提供する有望な代替手段を提供します。この研究では、THz-DCS を利用して、純粋なポリマー、多層フィルム、市販のブレンド、バイオポリマーを含む 12 種類のポリマーを分類します。これらのスペクトル信号の複雑さを処理するために、THz-DCS データに合わせた新しい深層学習アーキテクチャであるマルチスケール フィーチャー アテンション ネットワーク (MSFAN) を提案します。このフレームワークには、信号の再キャリブレーションとマルチスケールの並列畳み込みのための機能ゲートが統合されており、多様な周波数パターンをキャプチャします。これらの特徴は、特徴間アテンションとアテンション プーリングを通じてさらに洗練され、モデルが本質的に最も有益な THz 領域を強調表示できるようになります。 MSFAN は常に最先端のモデルを上回っており、分類精度は 85.2% に達しています。この研究は、THz-DCS と深層学習技術を組み合わせて、効果的でスケーラブルで解釈可能なポリマー分類を実現できる可能性を示しています。
原文 (English)
Multi-Scale Feature Attention Network for Polymer Classification Using Terahertz Spectroscopy
Reliable polymer identification is essential for ensuring the quality and safety of recycled plastics, yet conventional sorting and spectroscopic techniques often struggle to deliver robust discrimination. Terahertz (THz) spectroscopy offers a promising alternative, providing high-resolution and non-destructive measurements. In this work, we leverage THz signals to classify 12 types of polymers, including pure polymers, multilayer films, commercial blends, and biopolymers. To handle the complexity of these spectral signals, we propose the Multi-Scale Feature Attention Network (MSFAN), a novel deep learning architecture tailored for THz data. The framework integrates feature gating for signal recalibration and multi-scale parallel convolutions to capture diverse frequency patterns. These features are further refined through cross-feature attention and attention pooling, enabling the model to intrinsically highlight the most informative THz regions. MSFAN consistently outperforms state-of-the-art models, reaching a classification accuracy of 85.2%. This study demonstrates the potential of combining THz spectroscopy with deep learning techniques for effective, scalable, and interpretable polymer classification.
APPO: Agentic Procedural Policy Optimization
Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language m…
Creative Integration: A Decidable Criterion of Creativity
"Integrative" solutions are widely praised but rarely defined: we lack an operational way to tell a genuine integration -- one that makes t…
大規模言語モデルベースの生成推奨の暗黙的推論
大規模言語モデル (LLM) は生成推奨 (GR) のバックボーンとして採用されることが増えており、事前トレーニングされた世界の知識へのアクセスが約束されています。しかし、この知識を GR に確実に活用する方法は、まだ十分に理解されていません。主な障害は、LLM ベースの GR が通常、アイテムをセマンティック ID (SID) で表現し、事前トレーニング中にこれらのトークンが LLM に認識されないため、LLM の自然言語推論インターフェイスを混乱させることです。既存のアプローチは、SID を接地して明示的な根拠を引き出す高価なマルチステージ パイプラインでこの問題に対処していますが、各ステージがいつ、なぜ必要なのかについての洞察は限られています。この研究では、LLM ベースの GR の明示的推論トレーニング パイプラインを体系的に分解し、3 つの重要な制限を明らかにしました。世界知識の言語化の弱体化、SID と自然言語トークン埋め込み空間間の不整合、理論的根拠の品質に対する敏感さであり、これらすべてが明示的推論のパフォーマンスに悪影響を及ぼします。これらの問題を回避するために、GR 向けに調整された軽量の暗黙的推論パラダイムである PauseRec を提案します。 PauseRec は非常に実用的で、コストのかかる推論トレース取得と推論調整トレーニングを回避し、多くの利点をもたらします。(1) 標準の明示的 CoT メソッドよりも最大 6.22% 優れたパフォーマンスを発揮し、(2) トレーニング コストを GPU 時間で最大 65% 削減し、(3) 推論を最大 71.3% 高速化します。これらの結果により、PauseRec は明示的な根拠生成に代わる軽量の代替手段として位置づけられ、より効果的かつ効率的な LLM ベースの GR が可能になります。
原文 (English)
Implicit Reasoning for Large Language Model-based Generative Recommendation
Large Language Models (LLMs) are increasingly adopted as backbones for Generative Recommendation (GR), promising access to pretrained world knowledge. Yet reliably invoking this knowledge for GR remains poorly understood. A key obstacle is that LLM-based GR typically represents items with Semantic IDs (SIDs), disrupting LLMs' natural-language reasoning interface because these tokens are unseen by the LLM during pretraining. Existing approaches address this with expensive multi-stage pipelines that ground SIDs and elicit explicit rationales, but offer limited insight into when and why each stage is necessary. In this work, we systematically decompose explicit reasoning training pipelines for LLM-based GR, revealing three key limitations: weakened world-knowledge verbalization, misalignment between SID and natural-language token embedding spaces, and sensitivity to rationale quality, all of which hurt explicit reasoning performance. To circumvent these issues, we propose PauseRec, a lightweight implicit reasoning paradigm tailored for GR. PauseRec is exceptionally practical, avoiding costly reasoning trace acquisition and reasoning alignment training, leading to a multitude of benefits: (1) it outperforms standard explicit CoT methods by up to 6.22%, (2) it reduces training cost by up to 65% GPU hours, and (3) it speeds up inference by up to 71.3%. These results position PauseRec as a lightweight alternative to explicit rationale generation, enabling more effective and efficient LLM-based GR.
The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence
The metanym game is a competitive word game for LLMs that measures structural intelligence against established cognitive-science constructs…
SqLinear: Balanced Square Partitioning Makes Linear Interaction Sufficient for Large-Scale Traffic Forecasting
Traffic prediction is a core task in intelligent transportation systems and urban-scale decision making. Despite the effectiveness of mains…
DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training
Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in comple…
ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL
Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Co…
MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding
Using molecular large language models (LLMs) as a unified framework for understanding molecular structures and functions is emerging as a n…
BeatEdit: Symbolic Music Generation as Explicit Editing
Music creation is fundamentally a process of revision. Yet symbolic music generation remains dominated by paradigms that produce complete s…
AuEmoChat: 会話音声合成のための本物の感情の理解とレンダリング
会話型音声合成 (CSS) は、ユーザーとエージェントの対話において、人間のような感情表現と文脈の一貫性を備えた音声を合成することを目的としています。既存の CSS 手法は、事前に定義された感情ラベル スペース (7 つの感情カテゴリなど) が限られているため、本物の人間の感情を表現するのに苦労していますが、マルチターン対話履歴内の冗長なマルチモーダル トークンがコンテキストの理解を妨げます。これらの問題に対処するために、私たちは本物の感情の理解とレンダリングのための CSS フレームワークである AuEmoChat を提案します。まず、有限スカラー量子化を介して大規模な感情音声から離散的な本物の感情トークン空間を学習する AuEmoCodec を開発し、限られた基本的な感情カテゴリよりもより本物の感情表現を可能にします。さらに、感情に関連したコンテキストを維持しながら、マルチモーダルな対話履歴内の冗長トークンをマージする、本物の感情に基づくトークンマージアルゴリズムである AuEmoToMe を提案します。これを自己回帰テキスト音声モデルに統合して、ターゲットとなる本物の感情トークンと音声トークンを予測します。最後に、統合された対話コンテキスト、ターゲットの本物の感情、および音響事前分布を共同で条件付けすることによって音声をレンダリングする、本物の感情フロー マッチングを提案します。 NCSSD-EmCap データセットに関する広範な実験により、AuEmoChat が最先端の CSS ベースラインを上回り、より表現力豊かで本物の感情的なスピーチを生成することが実証されました。
原文 (English)
AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods struggle to render authentic human emotions due to limited predefined emotion label spaces (e.g., seven emotion categories), while redundant multimodal tokens in multi-turn dialogue history interfere with context understanding. To address these issues, we propose AuEmoChat, a CSS framework for authentic emotion understanding and rendering. First, we develop AuEmoCodec, which learns a discrete authentic emotion token space from large-scale emotional speech via finite scalar quantization, enabling a more authentic emotion representation than limited basic emotion categories. Furthermore, we propose AuEmoToMe, an authentic-emotion-guided token merging algorithm that merges redundant tokens in multimodal dialogue history while preserving emotion-relevant context. We integrate it into an autoregressive text-speech model to predict the target authentic emotion token and speech tokens. Finally, we propose Authentic Emotion Flow Matching, which renders speech by jointly conditioning on merged dialogue context, target authentic emotion, and acoustic priors. Extensive experiments on the NCSSD-EmCap dataset demonstrate that AuEmoChat outperforms state-of-the-art CSS baselines and generates more expressive and authentic emotional speech. The code and speech demos will be available at: https://github.com/AI-S2-Lab/AuEmoChat.
ビジネス分野全体にわたる最先端の AI パフォーマンス: ナレッジワークと分析的推論の事例に基づいたベンチマーク
大規模言語モデル (LLM) は、ベンチマーク スコアに反映されているように急速に改善されていますが、これらの AI ベンチマークでは主に、事実の再現、限定的な質問応答、数学的問題解決、コーディングやエージェント ツールの使用などの機能がテストされます。まだ十分に測定されていないのは、複雑な情報の統合、不確実性と不完全な情報の下での判断の行使、複数のステークホルダーの状況での戦略的および敵対的思考の適用、トレードオフの比較検討、防御可能な構造化された分析の作成など、ホワイトカラーの専門家が日々行っている分析知識作業における AI の進歩です。このギャップは、そのような仕事の主観的な要素ではさらに顕著であり、成功を定義するのが難しい場合があります。トップクラスのビジネススクールが実践する「ケースメソッド」教育形式は、この測定ギャップに対処するための自然な基盤を提供します。私たちは、18 分野にわたるビジネスケースから抽出された数百の質問にわたるベンチマークである BusinessCaseBench を構築します。各質問は、専門家が作成した講師のケースソリューションから導き出された採点ルーブリックと対になっています。 BusinessCaseBench では、フロンティア AI モデルはすでにインストラクターのルーブリックに対して高いスコアを獲得しており、1 つのモデル ファミリー内の機能は 2 年間で大幅に向上しています。これらの結果は、この種の作業における AI のパフォーマンスがすでに高く、急速に向上していることを示す強力な証拠を提供します。これは、事例教育学によって学部生や MBA がこの種の分析的推論を訓練されるビジネス スクールや、歴史的にそのようなスキルが初期キャリアの仕事に定着してきたエントリーレベルの専門職に影響を及ぼします。
原文 (English)
Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information, applying strategic and adversarial thinking in multi-stakeholder settings, weighing trade-offs, and producing defensible, structured analyses. This gap is even more pronounced for subjective components of such work, where success can be challenging to define. The "case method" form of education practiced by top business schools provides a natural foundation for addressing this measurement gap, and we construct BusinessCaseBench, a benchmark spanning hundreds of questions drawn from business cases across eighteen disciplines, each paired with a grading rubric derived from the expert-written instructor case solution. On BusinessCaseBench, frontier AI models already score highly against instructor rubrics, and capability within one model family improves substantially over two years. These results provide strong evidence that AI performance on this class of work is already high and rapidly improving, with implications for business schools, where case pedagogy trains undergraduates and MBAs in this kind of analytical reasoning, and for entry-level professional roles, where such skills have historically anchored early-career work.
EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Ex…
CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization
Textual Collaborative Prompt Optimization (TCPO) extends TextGrad (Yuksekgonul et al., 2025) to a decentralized setting by allowing multipl…
Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning
Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to…
HijackKV: New Threat in Position-Independent KV Cache Reuse
Key-Value (KV) cache reduces inference latency in large language models (LLMs). Traditional prefix-based reuse has low cache hit rates acro…
ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing
Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundat…
Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS
Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of wheth…
Mission-Level Runtime Assurance for LLM-Assisted ISR Swarms over a Verification-Aware Fabric
Swarms of LLM-assisted autonomous robots are increasingly proposed for cooperative intelligence, surveillance, and reconnaissance (ISR) in…
DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory
We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifie…
Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature
X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published…
Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls
Language-model agents act through structured tool calls whose arguments carry very different risks: untrusted content may legitimately shap…
LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings
Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this g…
A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor Networks
Traffic forecasting is important for efficient traffic management and route planning in smart cities. Existing traffic forecasting studies…
Progressive Multimodal Alignment for Continual Instruction Tuning
Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it c…
Benchmarking LLM Competence on Logical Inference over Probability Operators
Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions o…
SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups
Generative modeling of protein backbones promises the de novo design of proteins with prescribed structural and functional properties. Exis…
LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents
We introduce LabEvolver, a training-free framework that equips safe and grounded wet-lab agents with episodic memory from execution experie…
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn…
On a joint simultaneous learning of relevant feature subsets and subspaces in regression-like problems
We extend a recently introduced Entropy-Optimal Manifold Clustering (EOMC) to allow for a joint simultaneous identification of subsets and…