AIニュース 2026-08-08
自動生成: 2026-08-08 10:57 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
Responding to the next frontier of critical cyber capabilitiesOpenAI
OpenAI is sharing preliminary cybersecurity evaluations for Astra and…
-
How HSP GRUPPE builds AI capabilities for tax advisoryOpenAI
Discover how HSP GRUPPE uses ChatGPT Enterprise to boost productivity…
-
Anthropic、「Fable 5」の生物学の制限を緩和 誤検知によるフォールバックを約85%削減ITmedia AI+
Anthropicは、AIモデル「Claude Fable 5」の生物学分野における過度な保護機能を緩和したと発表した。安全性を重視するあ…
-
OpenAI、次期モデル「Astra」の一部開発を停止 「Critical」級サイバー能力の可能性否定できずITmedia AI+
OpenAIは、次期主力モデル「Astra」のサイバー能力が自社の安全指針における最上位「Critical」に達している可能性を発表した。…
-
After Rippling blew millions on AI in months, it built an employee ROI toolTechCrunch AI
After its own AI usage wake-up call, Rippling this week unveiled AI S…
-
Cloudflare launches Kitesurf, a browser built for AI agentsTechCrunch AI
Kitesurf is a cloud-hosted browser designed for AI agents instead of…
-
シャープ、通期純利益見通し170億円下方修正 円安など影響 AIサーバは9月に参入ITmedia AI+
シャープは7日、2027年3月期の連結純利益が前期比47.3%減の250億円になる見通しだと発表した。期初予想から170億円下方修正した。…
トピック別件数
- 研究/論文 130件
- LLM/生成AI 121件
- エージェント 97件
- 画像/動画生成 41件
- ビジネス/資金調達 27件
- ロボティクス 14件
- その他 6件
- ハードウェア/半導体 5件
- 規制/政策 2件
日本語メディア8件
ITmedia AI+ (日本語)
Anthropic、「Fable 5」の生物学の制限を緩和 誤検知によるフォールバックを約85%削減
Anthropicは、AIモデル「Claude Fable 5」の生物学分野における過度な保護機能を緩和したと発表した。安全性を重視するあまり発生していた無害な質問への誤検知や下位モデルへのフォールバックを大幅に削減。専門的な二重用途研究への制限は維持しつつ、一般的な健康・教育…
OpenAI、次期モデル「Astra」の一部開発を停止 「Critical」級サイバー能力の可能性否定できず
OpenAIは、次期主力モデル「Astra」のサイバー能力が自社の安全指針における最上位「Critical」に達している可能性を発表した。要件を満たさない一部活動を停止し、リアルタイム監視や思考過程の評価など管理を強化する。能力を抑制するのではなく、政府機関や外部組織と協力して…
「声の無断利用」が権利侵害に――法務省が見解を明示 「AIカバー」も対象
法務省は、生成AIの普及によって声優などの声が無断利用されている問題を巡り、声も既存の権利で法的に保護できるとの見解を明示した。本人の声を模したAI音声によるカバー音源なども権利侵害の可能性があるという。
シャープ、通期純利益見通し170億円下方修正 円安など影響 AIサーバは9月に参入
シャープは7日、2027年3月期の連結純利益が前期比47.3%減の250億円になる見通しだと発表した。期初予想から170億円下方修正した。樹脂・燃料の価格上昇や円安が収益を圧迫する。営業利益は190億円引き下げて300億円。売上高は1兆7700億円に据え置いた。
「就活に生成AI利用」ほぼ全員に 面接で内容追及され困惑も
2027年春に卒業予定の大学生らを対象に行ったアンケートで、就職活動で生成AIを「利用していない」とした割合は3%にとどまり、ほぼ全ての学生が就活で何らかの形で生成AIを活用している実態が、人事分野の調査研究機関HR総研(東京都千代田区)などの調査で分かった。
「声」の権利明記 生成AIで無断利用、法務省が民事責任の解釈指針を公表
著名人の肖像などが生成AIで無断利用されている問題を巡り、法務省は声優らの「声」も法的保護の対象になると明記した解釈指針を公式サイトで公表した。肖像や氏名の無断利用については最高裁判例があるが、声については違法性の線引きが曖昧だった。権利侵害に当たる具体的な事例も盛り込み、生成…
著名人の「なりすまし詐欺広告」対策強化を要請 Google・LINEヤフー・Xなど対象 7府省庁合同で
警察庁など7府省庁は、SNSなどの「なりすまし詐欺広告」の対策を強化するようプラットフォームを運営する5社に要請した。
PFNの国産LLM「PLaMo 3.0 Prime」、さくらのAI推論基盤で提供開始 利用は申請制
さくらインターネットが、生成AI向け推論API基盤「さくらのAI Engine」で、Preferred Networks(PFN)の国産大規模言語モデル(LLM)「PLaMo 3.0 Prime」の提供を始めた。利用には申請が必要で、無償プランは対象外となる。
海外メディア5件
TechCrunch AI (英語)
After Rippling blew millions on AI in months, it built an employee ROI tool
After its own AI usage wake-up call, Rippling this week unveiled AI Spend Console, a product that tracks individual and team employee AI sp…
Cloudflare launches Kitesurf, a browser built for AI agents
Kitesurf is a cloud-hosted browser designed for AI agents instead of people. It uses less computing power than Chromium for common automati…
Airbnb says AI is helping it ship features faster as it tests a new search function
Airbnb will debut a new AI-powered search experience with a toggle.
Jill Lepore on the ‘Artificial State’ and why Silicon Valley’s leaders are bad sci-fi readers
Historian Jill Lepore has a theory about why tech companies often use soaring language to describe their products — almost as if they’re fo…
New Mexico court orders Meta to pay additional $567M in child safety case
Meta's total fine has racked up to $942 million in this case.
公式ブログ2件
OpenAI (英語)
Responding to the next frontier of critical cyber capabilities
OpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taking to strengthen safeguards and security controls.
How HSP GRUPPE builds AI capabilities for tax advisory
Discover how HSP GRUPPE uses ChatGPT Enterprise to boost productivity, improve work quality, and create more capacity for tax advisory and…
論文307件
arXiv cs.AI (英語)
エージェントティック ネスティング: 既存のエンタープライズ アプリケーションの統合とサービスのための新しい方法論
企業の運営は、複数の異種ビジネス システムと情報アプリケーションに大きく依存しており、その結果、深刻なデータ サイロとプロセスの断片化も生じます。企業は、これらのアプリケーションの構築に多大な財政的および物的リソースを投資してきましたが、それらを効果的に活用し、調整することは依然として大きな課題です。 Enterprise Service Bus (ESB)、API ゲートウェイ インフラストラクチャ、ロボティック プロセス オートメーション (RPA) などのミドルウェア アーキテクチャを含む、エンタープライズ アプリケーション統合への従来のアプローチは、アーキテクチャの結合度の高さ、運用および保守コストの増大、インテリジェンス機能の制限など、固有の制限に悩まされています。この論文では、既存のエンタープライズ アプリケーションが階層的にネストされた構造内の自律 AI エージェントとしてカプセル化されるマルチエージェント コラボレーション フレームワークである Agentic Nesting を提案します。エージェントは、フラットな相互接続ではなく、エンタープライズ エコシステムの構成の複雑さを反映する階層化された管理トポロジに編成されます。このフレームワークは、各レガシー アプリケーションからデジタル エージェント プロキシを抽出して、自然言語対話と自律的な操作を可能にし、タスク分解と動的ディスパッチングのために中央オーケストレーターを通じて複数のエージェントを調整し、アプリケーション間のクエリとプロセス オーケストレーションのための統合された会話型インターフェイスを公開します。この論文の主な貢献は、「エージェントとしてのアプリケーション」統合パラダイムと「統合としての会話」対話哲学の提案と、異種システム調整および大規模データ アプリケーションを含むシナリオにおけるこの方法論の一般化可能性の探求です。
原文 (English)
Agentic Nesting: A New Methodology for Existing Enterprise Application Integration and Services
Enterprise operations extensively rely on multiple heterogeneous business systems and information applications, which also result in severe data silos and process fragmentation. Enterprises have invested considerable financial and material resources in building these applications, however, effectively leveraging and orchestrating them remains a formidable challenge. Conventional approaches to enterprise application integration, encompassing middleware architectures such as Enterprise Service Bus (ESB), API gateway infrastructures, and Robotic Process Automation (RPA), suffer from inherent limitations like high architectural coupling, escalating operation and maintenance costs, and limited intelligence capabilities. This paper proposes Agentic Nesting, a multi-agent collaboration framework in which existing enterprise applications are encapsulated as autonomous AI agents within a hierarchically nested structure. Rather than flat interconnection, agents are organized into layered stewardship topologies that mirror the compositional complexity of enterprise ecosystems. The framework extracts a digital agent proxy from each legacy application to enable natural-language interaction and autonomous manipulation, coordinates multiple agents through a central orchestrator for task decomposition and dynamic dispatching, and exposes a unified conversational interface for cross-application querying and process orchestration. The main contributions of this paper are the proposition of the "Application-as-Agent" integration paradigm and the "Conversation-as-Integration" interaction philosophy, together with an exploration of the generalization potential of this methodology in scenarios encompassing heterogeneous system coordination, and large-scale data applications.
Ignition Index: 言語モデルにおけるグローバル ワークスペース ダイナミクスの測定
グローバル ワークスペース理論 (GWT) の全か無かの点火予測をトランスフォーマー言語モデルで運用できる検証済みのスカラー メトリクスである点火インデックス (I) を紹介します。このメトリクスは、入力信号強度の関数として 4 つのパラメーターのシグモイドを層ごとの線形プローブ精度に適合させ、急峻性パラメーターのベータハットを抽出します。高い値は、突然の点火のような遷移を示します。低い値は段階的な蓄積を示します。 5 つのアーキテクチャ ファミリにわたる 11 のモデルにわたって、シャッフルラベル コントロールは、偽のプローブ能力よりも本物の言語構造に対して 9.6 倍の選択性を示します (p < 0.001、マンホイットニー U 検定)。 (1) フィードフォワード変圧器は総ベータハットで SSM を 89% 上回り (p < 1e-13、コーエンの d = 0.52)、Mamba はグローバル ブロードキャストの不在と一致するほぼ線形のプロファイルを示します。 (2) Huginn-3.5B は、反復軸に沿って深さ軸に比べて 2.12 倍高い点火を示し、再帰的アーキテクチャが再帰的次元に沿ってワークスペースのような遷移を示すことを示しています。 (3) Pythia-410M は、誘導ヘッドの形成に先立って、トレーニング ステップ 256 (+67%) で PELT で検出された相転移を示します。 (4) 点火をモデルのスケールや信号強度に結び付ける仮説は確認されず、変圧器のアーキテクチャが利用可能な点火メカニズムを飽和させる可能性があることを示唆しています。 Ignition Index は、これまでのスケーリング文献では特徴づけられていなかった 9.6 倍の測定選択性とアーキテクチャ レベルの識別性を備え、GWT の動的予測と機構的解釈可能性の間の最初に検証された定量的橋渡しを提供します。コード: https://github.com/saman-rahbar/ignition-index
原文 (English)
The Ignition Index: Measuring Global Workspace Dynamics in Language Models
We introduce the Ignition Index (I), a validated scalar metric that operationalizes Global Workspace Theory's (GWT) all-or-none ignition prediction in transformer language models. The metric fits a four-parameter sigmoid to per-layer linear probe accuracy as a function of input signal strength, extracting steepness parameter beta-hat: high values indicate abrupt, ignition-like transitions; low values indicate graded build-up. Across 11 models spanning five architecture families, shuffled-label controls demonstrate 9.6-fold selectivity for genuine linguistic structure over spurious probe capacity (p < 0.001, Mann-Whitney U-test). We find: (1) Feedforward transformers exceed SSMs by 89% in aggregate beta-hat (p < 1e-13, Cohen's d = 0.52), with Mamba exhibiting near-linear profiles consistent with absent global broadcast. (2) Huginn-3.5B exhibits 2.12-fold higher ignition along its iteration axis than its depth axis, demonstrating that recurrent architectures manifest workspace-like transitions along the recurrence dimension. (3) Pythia-410M shows a PELT-detected phase transition at training step 256 (+67%), preceding induction-head formation. (4) Hypotheses linking ignition to model scale and signal strength were not confirmed, suggesting transformer architectures may saturate available ignition mechanisms. The Ignition Index provides the first validated quantitative bridge between GWT's dynamical predictions and mechanistic interpretability, with 9.6-fold measurement selectivity and architecture-level discriminability not previously characterized in the scaling literature. Code: https://github.com/saman-rahbar/ignition-index
キツツキの蒸留: 弱いモデルが強いモデルの推論バグを診断する
大規模な言語モデルは、推論タスクを解決する能力があるにもかかわらず、推論タスクに失敗することがよくあります。私たちは、そのような失敗の多くは、全体的な無能さからではなく、中間ステップにおける局所的な推論のバグから生じると主張します。これらのバグは頻繁に修復可能であることを示します。同じ強いモデル推論プレフィックスの後に弱いプローブ モデルによって生成された短いパッチを挿入すると、軌道を正しい解決策に向け直すことができます。ただし、この修正効果は、弱いパッチや修復された軌道を直接微調整することによって確実に内部化されるわけではありません。これは、有用なシグナルが介入テキスト自体にあるのではなく、それがモデルの将来の推論分布をどのように再形成するかにあることを示唆しています。そこで私たちは、対照的な局所的介入から学習する弱者から強者への訓練フレームワークであるウッドペッカー蒸留を提案します。私たちの方法は、同じプレフィックスで成功した弱モデルパッチと失敗した弱モデルパッチを対比し、それらの誘導された将来のトークン予測から修正教師分布を構築し、この信号を抽出して強モデルに抽出します。数学的推論ベンチマークの実験では、Woodpecker Distillation が強力なモデルのパフォーマンスを一貫して向上させ、直接模倣ベースラインを上回るパフォーマンスを示しています。
原文 (English)
Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models
Large language models often fail on reasoning tasks despite possessing the capability to solve them. We argue that many such failures arise from localized reasoning bugs in intermediate steps rather than from global incompetence. We show that these bugs are frequently repairable: inserting a short patch generated by a weak probe model after the same strong-model reasoning prefix can redirect the trajectory toward a correct solution. However, this corrective effect is not reliably internalized by directly fine-tuning on weak patches or repaired trajectories, suggesting that the useful signal lies not in the intervention text itself, but in how it reshapes the model's future reasoning distribution. We therefore propose Woodpecker Distillation, a weak-to-strong training framework that learns from contrastive local interventions. Our method contrasts successful and unsuccessful weak-model patches at the same prefix, constructs a corrective teacher distribution from their induced future token predictions, and distills this signal into the strong model. Experiments on mathematical reasoning benchmarks show that Woodpecker Distillation consistently improves strong-model performance and outperforms direct imitation baselines.
継続的予測因子から臨床閾値まで: 虚血性脳卒中転帰予測におけるガイドラインに基づく分類の性能トレードオフに関する初期の証拠
機械学習モデルは、急性虚血性脳卒中における 90 日間の転帰予測において高い予測精度を達成していますが、モデルの説明と臨床医の推論が一致していないため、臨床での採用は制限されています。臨床ガイドラインに合わせたカットオフを求める臨床医のユーザー研究をきっかけに、パフォーマンスを犠牲にすることなく連続予測変数を臨床に基づいたカテゴリエンコーディングに置き換えることができるかどうかを尋ねます。 3 つの治療コホートに層別化された欧州の多施設登録簿で、標準モデルと完全に分類された勾配ブースト モデルを比較します。後者では、脳卒中ガイドラインに準拠した治療固有の閾値が使用されます。完全に分類されたモデルは、治療コホートのうち 2 つの連続モデルと統計的に区別できませんが、1 つのコホートでは予測精度が大幅に低下しています。グローバルな特徴重要度ランキングは一貫性を保っており、連続予測因子をガイドラインベースのカテゴリーに離散化することで、すべての治療グループにわたる予後因子の中核となる階層が維持されることが示唆されます。したがって、ガイドラインに基づいた分類は、脳卒中アウトカムモデルにとって実行可能な設計選択肢となります。
原文 (English)
From Continuous Predictors to Clinical Thresholds: Early Evidence on Performance Trade-offs of Guideline-Based Categorisation for Ischaemic Stroke Outcome Prediction
Machine learning models achieve strong predictive accuracy for 90-day outcome prediction in acute ischaemic stroke, yet clinical adoption is limited by the misalignment of model explanations with clinicians' reasoning. Motivated by a clinician user study calling for clinical guideline-aligned cut-offs, we ask whether continuous predictors can be replaced by clinically informed categorical encodings without sacrificing performance. On a multi-centre European registry stratified into three treatment cohorts, we compare standard and fully categorised gradient-boosted models, the latter using stroke guideline-aligned, treatment-specific thresholds. The fully categorised models are statistically indistinguishable from their continuous counterparts in two of the treatment cohorts, with a significant drop in predictive accuracy in one cohort. Global feature importance rankings remain consistent, suggesting that discretising continuous predictors into guideline-based categories preserves the core hierarchy of prognostic factors across all treatment groups. Guideline-based categorisation is thus a viable design choice for stroke-outcome models.
SkillTrace: LLM-Agent スキル再利用のためのマルチトレース来歴監査
LLM エージェント エコシステムは、メタデータ、自然言語命令、コード、ツール、リファレンス、運用ワークフローの混合形式パッケージなど、再利用可能なスキルを中心に急速に成長しています。スキルが市場の成果物になるにつれて、その再利用の監査は、通常のコード クローンの検出と同じ問題ではなくなりました。既存の検出器は、単一モダリティのソース コードまたはパッケージ全体の類似性をターゲットにしていますが、スキル再利用の証拠は、作成されたテキスト、実装フラグメント、および運用構造全体に分散されています。その結果、スキルの一部だけを保存する再利用が失われる可能性があります。 LLM エージェントのスキルを再利用するためのマルチトレース来歴監査フレームワークである SKILLTRACE を紹介します。 SKILLTRACE は、Expression、Implementation、Operational の 3 つの来歴トレースを抽出します。これは、アクティブ化、手順、およびリソース フロー構造をキャプチャするスキル操作グラフ (SOG) として操作トレースを表します。 LLM は、取り込み時に 1 回だけ、操作トレースの抽出のみを支援します。監査時に、SKILLTRACE はキャッシュされたトレースを決定論的に比較し、同じ機能の厳密な否定に対して各トレースを調整し、どのトレースが再利用の決定をサポートしているかを報告します。 SKILLTRACE-BENCH では、100 のマーケットプレイス アンカーおよび 751 のネガティブ コントロールを超える 820 の変換された再利用ポジティブを使用して、SKILLTRACE は AUROC 0.938 および F1 0.898 を達成しました。さらに、36,446 のスキルを対象としたワイルド監査では、トレースに起因する証拠により、リポジトリ レベルのベースラインを超えた実用的な再利用レビュー キューが明らかになったことが示されています。
原文 (English)
SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse
LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code, tools, references, and operational workflows. As skills become marketplace artifacts, auditing their reuse is no longer the same problem as ordinary code clone detection. Existing detectors target single-modality source code or whole-package similarity, yet skill reuse evidence is distributed across authored text, implementation fragments, and operational structure. As a result, they can miss reuse that preserves only one part of a skill. We present SKILLTRACE, a multi-trace provenance auditing framework for LLM-agent skill reuse. SKILLTRACE extracts three provenance traces: Expression, Implementation, and Operational. It represents the Operational Trace as a Skill Operational Graph (SOG) that captures activation, procedure, and resource-flow structure. An LLM assists only the Operational-trace extraction, once at ingestion; at audit time SKILLTRACE compares cached traces deterministically, calibrates each trace against same-function strict negatives, and reports which trace supports a reuse decision. On SKILLTRACE-BENCH, with 820 transformed reuse positives over 100 marketplace anchors and 751 negative controls, SKILLTRACE achieves AUROC 0.938 and F1 0.898. A 36,446-skill wild audit further shows that trace-attributed evidence surfaces actionable reuse review queues beyond repository-level baselines.
抽象的な事象の因果規則: 帰納と応用
イベント中心のインテリジェントな分析システムは、リスクの早期警告、意思決定のサポート、物語の理解のために明示的な因果イベントの知識に大きく依存しています。それにもかかわらず、既存のインスタンスレベルの因果ペアは、低頻度のロングテールイベントと目に見えないイベントの組み合わせで深刻な汎化欠陥に悩まされます。この制限に対処するために、この研究では、固有の因果関係を保持しながら、具体的な因果関係のペアを一般化された抽象的な因果論理に変換する、新しい関係レベルの因果抽象化パラダイムである抽象事象因果規則 (AECR) を提案します。私たちは、ノイズの多い生の因果データから信頼できる AECR を抽出するために、類似性制約クラスタリングと組み合わせたマルチエージェントの具体から抽象への因果誘導 (CACI) システムを設計します。これに基づいて、2 つの完全な AECR 知識ベースが構築されます。抽象因果知識の実用性を検証するために、抽象ルールガイド付き因果アテンション エンコーダ (AR-GCAE) を提案します。これは、ルールガイド付きアテンション レイヤーとゲート表現融合を介して、取得した AECR を因果関係グラフ イベント予測 (CGEP) ベンチマーク タスクに注入します。定量的な実験結果から、AECR を適用すると事象の因果推論の汎化能力が大幅に強化され、事象予測のパフォーマンスが一貫して向上し、まれで見たことのない事象のサンプルで最も顕著な向上が観察されることが明らかになりました。
原文 (English)
Abstract Event Causal Rules: Induction and Application
Event-centric intelligent analytical systems heavily depend on explicit causal event knowledge for risk early warning, decision-making support and narrative comprehension. Nevertheless, existing instance-level causal pairs suffer severe generalization deficits on low-frequency long-tail and unseen event combinations. To address this limitation, this work proposes Abstract Event Causal Rule (AECR), a novel relation-level causal abstraction paradigm that transforms concrete cause-effect pairs into generalized abstract causal logic while retaining their intrinsic causal relationships. We design a multi-agent Concrete-to-Abstract Causal Induction (CACI) system coupled with similarity-constrained clustering to distill trustworthy AECRs from noisy raw causal data, based on which two complete AECR knowledge bases are built. To validate the practical utility of abstract causal knowledge, we propose an Abstract Rule-Guided Causal Attention Encoder (AR-GCAE), which injects the retrieved AECRs into the causality Graph Event Prediction (CGEP) benchmark task via rule-guided attention layers and gated representation fusion. Quantitative experimental results reveal that applying AECRs substantially strengthens the generalization capacity of event causal reasoning and brings consistent performance improvements to event prediction, with the most prominent gains observed on rare and unseen event samples.
Otter: 時間を認識し、歴史に応じた人間のチェス AI
Otter は 1,530 万パラメータのヒューマン チェス AI で、各局面を個別に扱うのではなく、時間を意識した連続的なプロセスとしてプレイをモデル化することで人間の動きの選択を予測します。これは 2 つの条件付け信号を組み合わせます。(1) 最後の 20 手の動きに関する予測を条件付けし、序盤の好み、位置ドリフト、およびゲーム内の行動傾向をキャプチャする手番履歴エンコーダー。 (2) クロック圧力に基づいて予測を調整する時間制御モジュール。 Otter は、単一の T4 GPU 上で 30 日間にわたって 1 億 1,700 万の Lichess 高速ゲームからの 61 億の位置でトレーニングされます。 Otter は、トップ 1 の手予測精度 55.23%、トップ 5 手の予測精度 90.95% を達成し、はるかに少ないパラメーターと少ないトレーニング データで、以前の最先端のヒューマン チェス モデルである Maia 2 を上回りました。 11 の Elo ブラケット (=2000) 全体で、精度は 1900 ~ 1999 ブラケットの 57.38% でピークに達します。これらの結果は、チェスを時間を意識した連続的なアクティビティとしてモデル化すると、より小規模なモデルを使用して位置のみのアプローチよりも人間が正確な手を予測できることを示しています。コード、トレーニングされたモデル、完全なトレーニング ログは公開されています。
原文 (English)
Otter: A Time-Aware, History-Conditioned Human Chess AI
Otter is a 15.3M-parameter human chess AI that predicts human move selection by modeling play as a time-aware, sequential process rather than treating each position in isolation. It combines two conditioning signals: (1) a move history encoder that conditions predictions on the last 20 moves, capturing opening preferences, positional drift, and intra-game behavioral tendencies; and (2) a time control module that modulates predictions based on clock pressure. Otter is trained on 6.1 billion positions from 117 million Lichess rapid games over 30 days on a single T4 GPU. Otter achieves 55.23% top-1 and 90.95% top-5 move-prediction accuracy, surpassing the prior state-of-the-art human chess model, Maia 2, with far fewer parameters and less training data. Across 11 Elo brackets (=2000), accuracy peaks at 57.38% in the 1900-1999 bracket. These results show that modeling chess as a time-aware, sequential activity yields more human-accurate move prediction than position-only approaches, using a smaller model. Code, trained models, and complete training logs are publicly released.
SearchAuditor: Long-Horizon Search Agent での失敗の監査と原因特定
深層検索エージェントは、長期にわたる Web インタラクションを通じて困難な質問に取り組みます。このプロセスは複雑かつ脆弱です。小さな推論エラーが、長くノイズの多い軌跡を経て、流暢ではあるが不正確な回答に伝播する可能性があります。このような障害を診断することは難しく、非常に長い実行トレースを手動で検査する必要があり、人間の能力を超える可能性があります。そこで、LLM 監査人がこれらの障害を特定し、原因を特定し、修復できるかどうかを評価するベンチマークである SearchAuditBench を導入し、それによって人的負担を軽減します。 SearchAuditBench は、5 つの深層検索ベンチマーク上の 8 つのオープンウェイト モデルから収集された、平均 73.1 メッセージと 65.1K トークンに相当する 1,243 件の失敗した軌跡で構成されており、それぞれに専門家による重大なエラー ステップ、検索固有の根本原因、およびグレーディング ルーブリックによる参照修復の注釈が付けられています。さらに、証拠に基づいた判断を通じて検索エージェントの失敗を効果的に特定し、属性を特定し、修復する多視点監査フレームワークである SearchAuditor を提案します。実験結果は、GPT-5.5 のようなフロンティア モデルを使用した場合、最も強力なベースラインであっても、エンドツーエンドの合格率は 26.6% にすぎないことを示しています。対照的に、当社の SearchAuditor は、さまざまなフロンティア モデルにわたるすべてのベースラインを常に上回り、32.3% のエンドツーエンドの合格率を達成し、失敗した実行を修復して再開することで、エージェントがエラーからより適切に回復できるようになります。
原文 (English)
SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents
Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long, noisy trajectories into fluent but incorrect answers. Diagnosing such failures is difficult, requiring the manual inspection of extremely long execution traces, which could be beyond human capacity. We therefore introduce SearchAuditBench, a benchmark that evaluates whether LLM auditors can localize, attribute, and repair these failures, thereby reducing the human burden. SearchAuditBench comprises 1,243 failed trajectories, averaging 73.1 messages and 65.1K tokens, collected from eight open-weight models on five deep-search benchmarks, each expert-annotated with the critical error step, a search-specific root cause, and a reference repair with grading rubrics. We further propose SearchAuditor, a multi-perspective auditing framework that effectively localizes, attributes, and repairs search-agent failures through evidence-grounded adjudication. Experimental results show that even the strongest baseline, when powered by a frontier model like GPT-5.5, attains only a 26.6% end-to-end pass rate. In contrast, our SearchAuditor consistently outperforms all baselines across different frontier models, achieving an end-to-end pass rate of 32.3%, and resuming failed runs with its repairs enables agents to better recover from errors.
PD-GS: オーディオ駆動型トーキング ヘッド用の音素駆動型 3DGS
3D ガウス スプラッティング (3DGS) により、フォトリアリスティックなトーキングヘッドの高速レンダリングが可能になりますが、正確な唇の調音は依然としてとらえどころがありません。口の動きが平滑化されすぎることが多く、両唇閉鎖などの厳しい調音制約に違反する可能性があり、悪名高い「口漏れ」アーチファクトが発生します。重要な問題点は、回帰目標に基づいて、短い離散的な調音イベントが連続的な音響埋め込みから推測されるため、予測が平均的な口の構成に偏ることです。最新の自己教師あり音声エンコーダは豊富な韻律と音声の手がかりを提供しますが、クロージャレベルのイベントを確実に曖昧さをなくす、明示的なフレーム調整された言語ターゲットを提供しません。私たちは \textbf{音素駆動ガウス スプラッティング (PD-GS)} を提案します。これは、自動 ASR および強制アライメント パイプラインから取得された時間的にアライメントされた音素トークンで 3DGS トーカーを強化します。私たちのコアコンポーネントである \textbf{LFM (Linguistic Fusion Module)} は、学習されたゲートを通じて連続音声コンテキストと離散音素埋め込みを適応的に融合し、モデルがスムーズな音声主導のダイナミクスを維持しながら、調音に重要なセグメントの音素ガイダンスを強化できるようにします。 PD-GS は、画像再構成と唇のランドマーク監視を使用して、単眼ビデオのみからトレーニングされます。 HDTF では、PD-GS は、比較したベースライン (LMD 2.66) の中で最良の唇ジオメトリを実現し、困難な音素シーケンスにおける閉包違反を定性的に削減し、言語的により忠実なニューラル アバターを生成します。
原文 (English)
PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads
3D Gaussian Splatting (3DGS) enables fast, photorealistic talking-head rendering, yet accurate lip articulation remains elusive: mouth motion is often over-smoothed and may violate hard articulatory constraints such as bilabial closures, producing the notorious ``leaky mouth'' artifact. A key difficulty is that brief, discrete articulatory events are inferred from a continuous acoustic embedding under a regression objective, which biases predictions toward averaged mouth configurations. While modern self-supervised speech encoders provide rich prosodic and phonetic cues, they do not provide an explicit, frame-aligned linguistic target that reliably disambiguates closure-level events. We propose \textbf{Phoneme-Driven Gaussian Splatting (PD-GS)}, which augments a 3DGS talker with time-aligned phoneme tokens obtained from an automatic ASR and forced-alignment pipeline. Our core component, the \textbf{Linguistic Fusion Module (LFM)}, adaptively fuses continuous audio context with discrete phoneme embeddings through a learned gate, allowing the model to preserve smooth audio-driven dynamics while strengthening phoneme guidance on articulation-critical segments. PD-GS is trained purely from monocular video using image reconstruction and lip landmark supervision. On HDTF, PD-GS achieves the best lip geometry among the compared baselines (LMD 2.66) and qualitatively reduces closure violations in challenging phoneme sequences, yielding more linguistically faithful neural avatars.
特権ガイダンスが調整されていない場合: マルチターン エージェント向けの状態一致ルーティングとコンテキスト化された自己蒸留
ポリシーに基づく特権蒸留により、同期された教師が成功の軌跡などのトレーニング専用の参照にアクセスしながら、ターンごとに生徒の反応を再採点できるようになり、複数ターンのエージェントに緻密な監督が提供されます。ただし、対話型環境では、生徒の以前のアクションによって実行状態が継続的に変化します。学生が異なるアクションを実行したり、サブ目標を異なる順序で完了したりすると、そのロールアウトは参照でカバーされていない状態に到達する可能性があり、参照は実際に到達した状態に関するガイダンスの信頼性の低い情報源になります。したがって、特権蒸留を無差別に適用すると、状態と参照の不一致が生じます。この不一致により、生徒の現在の実行状態と互換性のある特権的な参照ガイダンスを提供するという中心的な目的が動機付けられます。 State-Matched Routing and Contextualized Self-Distillation (SMRC-SD) を導入します。これは、ポリシーに従った学生を特権トラジェクトリがいつどのようにガイドするかを明示的に決定します。 SMRC-SD は各ターンで、スチューデントの現在の実行状態が参照軌道に沿ったサポートされている状態と一致するかどうかを検証します。蒸留は一致する状態でのみ適用され、参照にローカルで互換性のあるガイダンスが欠如しているターンは除外されます。 SMRC-SD は、一致した状態ごとに、成功した軌道から状態条件付き教師コンテキストをさらに構築し、実際に到達した状態での監督を接地します。 ALFWorld と WebShop 全体で、SMRC-SD は一貫して無条件で成功したフルパス蒸留を上回ります。 Qwen3-1.7B を使用すると、ALFWorld ではタスクの成功率が $0.746$ から $0.865$ に、WebShop では $0.574$ から $0.693$ に向上します。制御されたルーティングとコンテキスト アブレーションは、ローカルでサポートされるターンの選択と、これらの利点に貢献する状態互換の教師コンテキストの構築の両方をサポートします。コードは https://github.com/liujunzhuo/SMRC-SD で入手できます。
原文 (English)
When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents
Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher to re-score the student's response at every turn with access to training-only references, such as successful trajectories. In interactive environments, however, the student's preceding actions continually change the execution state. As the student takes different actions or completes subgoals in a different order, its rollout may reach states not covered by the reference, making the reference an unreliable source of guidance for the state actually reached. Applying privileged distillation indiscriminately therefore creates state--reference mismatch. This mismatch motivates a central objective: providing privileged reference guidance that remains compatible with the student's current execution state. We introduce State-Matched Routing and Contextualized Self-Distillation (SMRC-SD), which explicitly determines when and how a privileged trajectory should guide an on-policy student. At each turn, SMRC-SD verifies whether the student's current execution state matches a supported state along the reference trajectory. Distillation is applied only at matched states, filtering out turns for which the reference lacks locally compatible guidance. For each matched state, SMRC-SD further constructs state-conditioned teacher context from the successful trajectory, grounding supervision in the state actually reached. Across ALFWorld and WebShop, SMRC-SD consistently outperforms unconditional successful full-path distillation. With Qwen3-1.7B, it improves task success from $0.746$ to $0.865$ on ALFWorld and from $0.574$ to $0.693$ on WebShop. Controlled routing and context ablations support both selecting locally supported turns and constructing state-compatible teacher context as contributors to these gains. Code is available at https://github.com/liujunzhuo/SMRC-SD.
人間の認知と行動の小規模な基礎モデル
人間の行動データに基づいて微調整された大規模な言語モデルが、汎用の認知プロキシとして登場しましたが、これに必要な規模や、これらのモデルがタスク構造を処理するのか、それとも統計的ショートカットを利用するのかは未解決のままです。私たちは、160 の実験からの 1,070 万の試験レベルの選択肢のデータセットである Psych-101 上の 4 つのアーキテクチャ ファミリにわたる 135M から 14B のパラメーターの 14 のモデルをトレーニングします。流通においては、規模はほとんど重要ではありません。モデルは、あたかも天井に向かっているかのように狭い帯域内に収まり、参加者が参加しなかった場合の 70B のベースラインに一致するには、0.6B ~ 1B のパラメータで十分です。分布外では、そのバンドは著しく急峻なスケーリング勾配に向かって開き、新しいタスク構造への一般化において、より大きなモデルが明らかに有利になります。これらのモデルがどのような情報を使用するかを判断するために、2 つの診断を実行します。 27 回の実験にわたって、タスクの指示、実験刺激、結果のフィードバック、選択履歴という 4 つのプロンプト チャネルを段階的に削除し、試行順序を変更します。刺激とフィードバックの内容をマスキングすると、学習した情報の 75.7% が破壊され、モデルが確率以下に押し下げられます。これは、選択履歴だけではパフォーマンスを考慮できないことを示しています。順列は、独立した試行を伴うタスクの不変性を明らかにしますが、試行の順序が前の応答によって決定される感度を明らかにします。したがって、認知的に微調整された小さなモデルは、心理学実験のノイズ上限推定器として有望ですが、その範囲はトレーニングで見られるパラダイムによって制限されたままです。
原文 (English)
Small Foundation Models of Human Cognition and Behaviour
Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. In-distribution, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choice history -- across 27 experiments, and permute trial order. Masking the content of stimuli and feedback destroys 75.7% of learned information and pushes models below chance, demonstrating that choice history alone does not account for performance. Permutation reveals invariance on tasks with independent trials but sensitivity where trial order is determined by prior responses. Small cognitively fine-tuned models therefore show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by the paradigms seen in training.
Project2Task: 自律研究のためのグラフに基づくプロジェクト レベルの計画
研究エージェントは、単一のトピックから文献を検索し、仮説を提案し、コードを生成し、実験を実行し、原稿を作成することがますます可能になります。ただし、研究プロジェクトは、単なる大きなタスクではありません。研究プロジェクトは、個別ではあるが関連する目的、並行する代替案、および依存関係を意識した順序を備えた複数の限定されたタスクを通じて進められる長期的な課題です。既存のシングルタスク システムでは、多くの場合、プロジェクトを 1 つの過大なタスクとして扱い、曖昧なタスクや重複するタスクのフラットなセットを生成したり、タスクの境界や実行順序を手動調整に任せたりします。自律研究のためのグラフガイドによるプロジェクト レベルの計画レイヤーである Project2Task を紹介します。プロジェクトの概要が与えられると、候補となる貢献をイノベーションの原子として表し、それらを有向系統グラフに整理します。軽量のベルヌーイ ブロック モデル目標により、水平、垂直、ハイブリッド ポートフォリオ分解の中から選択します。次に、Project2Task は、明示的なコントリビューションの所有権を持つ制限付きタスクを生成し、重複および欠落している実行フィールドを修復し、目的、入力、期待されるアーティファクト、評価要件、境界制約、依存関係、および実行順序を指定する依存関係を認識したタスク コントラクトを発行します。この契約は、特定の下流の研究実行者から独立しており、タスクの成果を一貫したプロジェクトレベルの結果に統合することをサポートします。約 30 のタスクを生成する 10 個のプロジェクト ブリーフのベンチマークでは、原稿ベースのポートフォリオ評価により、Project2Task の平均品質スコアは 7.15 でした。これに対し、ブリーフ ベースラインでは 4.58、トピックのみの設定では 5.31 でした。契約を AutoResearchClaw と統合することで、ダウンストリーム タスクの平均精度が 0.536 から 0.759 に向上しました。これらの結果は、一貫性があり、冗長性がなく、実行可能な研究タスクのポートフォリオを作成するための明示的なプロジェクトからタスクへの計画の価値を実証しています。
原文 (English)
Project2Task: Graph-Guided Project-Level Planning for Autonomous Research
Research agents can increasingly search literature, propose hypotheses, generate code, run experiments, and draft manuscripts from a single topic. However, a research project is not merely a larger task: it is a long-horizon agenda that must be advanced through multiple bounded tasks with distinct but related objectives, parallel alternatives, and dependency-aware sequences. Existing single-task systems often treat the project as one oversized task, produce a flat set of vague or overlapping tasks, or leave task boundaries and execution order to manual coordination. We introduce Project2Task, a graph-guided project-level planning layer for autonomous research. Given a project brief, it represents candidate contributions as innovation atoms and organizes them in a directed lineage graph. A lightweight Bernoulli block-model objective selects among horizontal, vertical, and hybrid portfolio decompositions. Project2Task then generates bounded tasks with explicit contribution ownership, repairs overlaps and missing execution fields, and emits dependency-aware task contracts that specify objectives, inputs, expected artifacts, evaluation requirements, boundary constraints, dependencies, and execution order. The contracts are independent of any particular downstream research executor and support integration of task outputs into a coherent project-level result. On a benchmark of ten project briefs yielding roughly 30 tasks, manuscript-based portfolio evaluation gives Project2Task an average quality score of 7.15, compared with 4.58 for the Brief Baseline and 5.31 for the Topic-only Setting. Integrating its contracts with AutoResearchClaw increases average downstream task accuracy from 0.536 to 0.759. These results demonstrate the value of explicit project-to-task planning for producing coherent, non-redundant, and executable research-task portfolios.
TriQua: 事実評価における粒度とコンテキストの調和
LLM の事実性評価の「分解してから検証する」パラダイムは、根本的なトレードオフに直面しています。つまり、アトミックな事実、つまり 1 つの情報単位を伝える 1 つの文では、重要なコンテキストが省略されることがよくありますが、より広範なステートメントでは、正確な評価に必要な粒度が欠如しています。これに対処するために、複雑さに基づいて事実を柔軟にモデル化するフレームワークである TriQua を紹介します。単純なクレームは標準のトリプルとして抽出されますが、複雑なクレームは補助的な文脈修飾子を付けることによって超関係ファクトとして表されます。この適応構造は、アトミック性を犠牲にすることなく、正確な検索と検証に必要なコンテキストを保存します。さらに、TriQua の検証プロセスは、特定のトリプルおよび修飾子内の具体的なエラーに直接注釈を付け、エラー検出に対するきめ細かい説明可能性を提供します。フレームワークと並行して、これらの構造化された事実単位の事実性を定量化する TriQuaScore を提案します。経験的評価では、TriQuaScore が人間による注釈付き事実スコアと強く一致し、TriQua が堅牢な分解品質を達成し、証拠に基づく事実検証において既存の分解ベースのフレームワークを上回るパフォーマンスを示していることが示されています。
原文 (English)
TriQua: Reconciling Granularity and Context in Factuality Evaluation
The "decompose-then-verify" paradigm for LLM factuality evaluation faces a fundamental trade-off: atomic facts, i.e., one sentence conveying one unit of information, often omit essential context, while broader statements lack the granularity needed for precise assessment. To address this, we introduce TriQua, a framework that flexibly models facts based on their complexity. Simple claims are extracted as standard triples, while complex claims are represented as hyperrelational facts by attaching auxiliary contextual qualifiers. This adaptive structure preserves the necessary context for accurate retrieval and verification without sacrificing atomicity. Furthermore, TriQua's verification process directly annotates concrete errors within specific triples and qualifiers, providing fine-grained explainability for error detection. Alongside the framework, we propose TriQuaScore to quantify the factuality of these structured fact units. Empirical evaluations show that TriQuaScore strongly aligns with human annotated factuality scores, TriQua achieves robust decomposition quality, and outperforms existing decomposition-based frameworks in evidence-based fact verification.
コヒーレンス指向の夢のシーンの視覚化
夢は感情的に激しいものですが、伝えるのが難しい場合があります。書かれた夢の説明を、夢を視覚化する 4 つのパネル画像の時系列に変換する Dream Scene Visualiser (DSV) システムについて説明します。これは、夢の説明を時系列で 4 つの部分に分割するよう促す大規模な言語モデルから始まります。次に、テキストから画像へのモデルが、シーケンス全体で視覚的な一貫性が維持された状態で各パートの画像を生成し、DSV がテキストと適切に一致しない画像を再生成します。 DreamBank の夢の説明からの 50 を超える DSV ビジュアライゼーションを評価し、CLIP、DINOv2、および Qwen2-VL ビジョン言語モデルを使用した客観的な測定を通じて、品質、忠実度、一貫性の結果を報告します。
原文 (English)
Coherence-Oriented Dream Scene Visualisation
Dreams can be emotionally intense but difficult to communicate. We describe the Dream Scene Visualiser (DSV) system which turns written dream descriptions into a temporal sequence of four panel images visualising the dream. This starts with a large language model prompted to split a dream description into four chronological parts. Then a text-to-image model produces images for each part with visual coherence maintained across the sequence, and DSV regenerates any image not suitably matching the text. We evaluate DSV over 50 visualisations from dream descriptions in DreamBank, and report quality, fidelity and coherence results via objective measures employing the CLIP, DINOv2 and Qwen2-VL vision-language models.
Search2Skill: ルーブリックベースの強化学習による知識の境界を超えたスキルの蒸留
現実世界の専門的なタスクを解決するために必要な手順知識をカプセル化した再利用可能なスキルは、LLM ベースのエージェントに専門分野での自己進化への道を提供します。既存の自己進化型スキル手法は、モデルのパラメトリック知識または軌跡から内部的にスキルを構築するため、モデルがすでに知っていることに制限されます。ただし、専門スキルの基礎となるドメインの慣例や標準的な手順は、多くの場合この境界を越えており、エージェントだけから引き出すのは困難です。したがって、この問題に対処するために、エージェントの能力ギャップを自動的に特定し、それらに対処するために外部ソースを検索し、取得した証拠を構造化された再利用可能なスキルに抽出する新しいフレームワーク Search2Skill を提案します。具体的には、Search2Skill はルーブリックベースの強化学習スキームによって最適化されており、検索するタイミング、検索方法、スキルの生成方法を共同で改善します。 3 つのベンチマークからの 8 つの専門家レベルのドメインでの実験では、Search2Skill が、ストリーミング評価プロトコルとホールドアウト評価プロトコルの両方で、検索拡張ベースラインと軌跡ベースのスキル学習ベースラインの両方を一貫して上回っていることが示されています。さらなる分析により、得られた利益は、取得された生の証拠ではなくスキルの抽象化から生じ、獲得したスキルはモデル スケール間で伝達されることが示されています。
原文 (English)
Search2Skill: Skill Distillation Beyond Knowledge Boundaries Via Rubric-Based Reinforcement Learning
Reusable skills, which encapsulate the procedural knowledge required to solve real-world professional tasks, offer LLM-based agents a path toward self-evolution in expert domains. Existing self-evolving skill methods construct skills internally from the model's parametric knowledge or trajectories, and are therefore bounded by what the model already knows. However, the domain conventions and standard procedures underlying professional skills often lie beyond this boundary and are hard to elicit from the agent alone. To address this issue, we therefore propose a novel framework, Search2Skill, that automatically identifies the agent's capability gaps, searches external sources to address them, and distills the retrieved evidence into structured, reusable skills. Specifically, Search2Skill is optimized by a rubric-based reinforcement learning scheme that jointly improves when to search, how to search, and how to generate skills. Experiments on eight expert-level domains from three benchmarks show that Search2Skill consistently outperforms both search-augmented and trajectory-based skill-learning baselines under both streaming and held-out evaluation protocols. Further analyses show that the gains arise from skill abstraction rather than raw retrieved evidence, and that the acquired skills transfer across model scales.
LUNAR: ユニバーサル ユーザー BehAvioR ログでのパーソナライズされた大規模言語モデルのベンチマーク
既存のパーソナライズされた LLM ベンチマークは、主にテキストのペルソナまたは分離された行動シグナルに依存しており、クロスドメインの行動パーソナライゼーションの評価は限定的であり、応答は異種混合の日常生活活動に基づいている必要があります。このギャップに対処するために、私たちは LUNAR を導入します。LUNAR は、衣服、食べ物、住居、移動などの普遍的な日常生活の領域にわたる縦断的なアプリのインタラクション履歴からの応答を、LLM がどのようにパーソナライズするかを評価するための最初のベンチマークです。データの希薄性とプライバシーの懸念を軽減しながら、スケーラブルなベンチマークの構築をサポートするために、LUNAR は、現実世界の動作パターンに基づいた多段階の粗いから細かいまでの合成パイプラインを使用します。忠実度分析は、他の合成ベンチマークよりも実際の行動分布とのより密接な一致を示します。 19 の主流 LLM に関する実験では、行動ログへのアクセスは必要ですが、深いパーソナライゼーションには十分ではないことが示されています。より多くのコンテキストやより大きなモデルは、より良いパフォーマンスを保証しません。効果的なパーソナライゼーションは、ドメイン全体で関連する証拠を選択して統合するかどうかにかかっています。詳細な行動記録の直接取得は圧縮メモリよりも常に優れたパフォーマンスを発揮しますが、より強力なパーソナライゼーションにはプライバシー保護が犠牲になる可能性があります。これらの調査結果は、証拠の選択、クロスドメイン統合、およびプライバシー管理がパーソナライズされた LLM の主要な課題であることを特定します。
原文 (English)
LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs
Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities. To address this gap, we introduce LUNAR, the first benchmark for evaluating how LLMs personalize responses from longitudinal app interaction histories across universal daily-life domains, including clothing, food, housing, and mobility. To support scalable benchmark construction while mitigating data sparsity and privacy concerns, LUNAR uses a multi-stage coarse-to-fine synthesis pipeline grounded in real-world behavioral patterns. Fidelity analyses show closer alignment with real behavioral distributions than other synthetic benchmarks. Experiments on 19 mainstream LLMs show that access to behavioral logs is necessary but not sufficient for deep personalization: neither more context nor larger models guarantees better performance; effective personalization depends on selecting and integrating relevant evidence across domains. Direct retrieval of fine-grained behavioral records consistently outperforms compressed memory, while stronger personalization can come at the cost of privacy protection. These findings identify evidence selection, cross-domain integration, and privacy control as key challenges for personalized LLMs.
WorldClaw: エージェントティック 3D オープンワールドの大規模な生成
システムは、グローバルな空間的一貫性、豊富なローカル コンテンツ、下流の編集と再利用に適した明示的なアセットを共同で維持する必要があるため、オープンエンド テキストから大規模で自由に探索可能な 3D 世界を生成することは依然として困難です。私たちは、オープンワールド 3D シーン生成のための完全にエージェント的な、粗いから細かいまでのフレームワークである WorldClaw を紹介します。計画エージェントは、テキスト プロンプトを領域、地形、資産、材料、空間関係の構造化された仕様に変換します。次に、WorldClaw は、セマンティック レイアウト、再利用可能なアセット、生成またはプロシージャル マテリアル、および地域を意識した高さフィールドから、グローバルに一貫した地形基盤を構築します。ディテールが要求される領域では、地形に合わせた構成を生成し、編集可能なテクスチャ メッシュを再構築して、地形上の配置を復元します。レンダリングベースのエージェントは、地形、オブジェクト、外観、および接触をさらに洗練します。 WorldClaw は、多様なオープンワールド プロンプトにわたって、一貫したグローバル テレイン構造を維持しながら、一貫した空間構成、視覚的に説得力のあるローカル コンテンツ、および編集可能なインスタンス レベルのアセットを備えた大規模なシーンを生成します。
原文 (English)
WorldClaw: Agentic 3D Open-World Generation at Scale
Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets suitable for downstream editing and reuse. We present WorldClaw, a fully agentic, coarse-to-fine framework for open-world 3D scene generation. Planning agents translate a text prompt into a structured specification of regions, terrain, assets, materials, and spatial relations. WorldClaw then builds a globally coherent terrain foundation from semantic layouts, reusable assets, generative or procedural materials, and a region-aware height field. For detail-demanding regions, it generates terrain-conditioned compositions, reconstructs editable textured meshes, and recovers their placement on the terrain; render-based agents further refine terrain, objects, appearance, and contacts. Across diverse open-world prompts, WorldClaw produces large-scale scenes with coherent spatial organization, visually compelling local content, and editable instance-level assets while preserving a consistent global terrain structure.
敵対的な不確実性の下での姿勢と維持の最適化
事前コミットメント態勢、つまり紛争シナリオが解決する前に戦域に軍事資産を割り当てることは、統合作戦計画において重要かつ正式に未解決の問題である。現在の慣行は、価値を最大化し、地理的範囲を無視する貪欲なヒューリスティックに依存しており、戦略的価値の高い場所を狙う敵対者に対して構造的に脆弱です。この論文では、資産、戦域の位置、およびタイム ステップにわたる有限水平マルコフ決定プロセスとしてモデル化された、姿勢および持続可能性割り当て (PSA) 問題に対するシナリオ重み付けの敵対的にロバストな姿勢最適化エンジンを紹介します。脅威シナリオの分布に対してシナリオ重み付けの期待ポスチャ効率を最大化することで資産を配置する Composite Expected Value (CEV) オプティマイザと、観察された配置に応じてターゲット分布を更新するベイジアン攻撃者に対して反復する RobustCEV 拡張機能を導入します。 20 の資産と 5 つの戦域拠点を備えたインド太平洋基地環境での 3 つの実験を通じて、次のことを実証しました。(1) 貪欲なベースラインでは、地理的なカバー範囲外により恒久的に 25.1% の姿勢効率ペナルティが発生し、価値相関のある敵対的脅威の下では 57.3% のシナリオ加重即応性崩壊が発生します。 (2) CEV オプティマイザーは、脅威の分布に地理的シグナルが含まれる場合に、グリーディに比べて最大 19.8% の効率を回復します。厳選された 5 ~ 20 のシナリオのセットにより、この利益の大部分を捉えるのに十分です。 (3) RobustCEV 拡張機能は、適応型の攻撃者が事前に欺瞞的な脅威を使用した場合、単純なオプティマイザと比較して最大 158% の効率を回復します。すべての結果は、ボンフェローニ補正と 2 水準分散分解を備えた対応のある t 検定を使用して検証され、報告されたパフォーマンス ギャップがサンプリング アーティファクトではなく、配置戦略の構造的特性であることが確認されます。
原文 (English)
Posture and Sustainment Optimization Under Adversarial Uncertainty
Pre-commitment posture, the assignment of military assets to theater locations before conflict scenarios resolve, is a critical and formally unsolved problem in joint operational planning. Current practice relies on greedy heuristics that maximize value and ignore geographic coverage and are structurally vulnerable to adversaries that target high-strategic value locations. This paper presents a scenario-weighted adversarially robust posture optimization engine for the Posture and sustainability allocation (PSA) problem, modeled as a finite-horizon Markov Decision Process over assets, theater locations, and time steps. We introduce the Composite Expected Value (CEV) optimizer, which places assets by maximizing scenario-weighted expected posture efficiency over a distribution of threat scenarios, and the RobustCEV extension, which iterates against a Bayesian adversary that updates its targeting distribution in response to observed placement. Across three experiments in an Indo-Pacific basing environment with 20 assets and 5 theater locations, we demonstrate that: (1) the greedy baseline incurs a permanent 25.1% posture efficiency penalty due to geographic under-coverage and a 57.3% scenario-weighted readiness collapse under value-correlated adversarial threat; (2) the CEV optimizer recovers up to 19.8% efficiency over greedy when the threat distribution carries a geographic signal, with a curated set of 5 to 20 scenarios sufficient to capture the majority of this gain; and (3) the RobustCEV extension recovers up to 158% efficiency relative to a naive optimizer when an adaptive adversary employs a deceptive threat prior. All findings are validated using paired t-tests with Bonferroni correction and two-level variance decomposition, confirming that the performance gaps reported are structural properties of placement strategies rather than sampling artifacts.
OrchestraBench: マルチエージェント オーケストレーションの障害モード、回復、分解品質の評価
マルチエージェント オーケストレーション フレームワークはデモから本番環境に移行していますが、ベンチマークでは通常、パイプラインが失敗した理由、カスケードがどこで始まったのか、またはどのルーティング決定が故障の原因となったのかを診断せずに、タスクの精度が報告されます。 OrchestraBench は、テンプレート化されたエンタープライズ ワークフロー上で、制御されたシード再現可能な障害挿入ハーネスを通じて障害、回復、分解を評価します。カスケード半径と障害モードごとの回復を主要な指標として導入し、ルーティング ポリシーをブートストラップ信頼区間とペア テストと比較します。 26 件のゴールドラベル診断では、キーワード/フラグ ルーターは、誤解を招く、または表面フラグが欠落している敵対的なケースで 0% のスコアを獲得しましたが、意図推論モデルのルーターは 100% のスコアを獲得し、オラクルと一致しました。検証可能な算術依存関係チェーンを介して実際のクロード エージェントを使用した制御メカニズムの調査により、5 つの MAST モードにわたる 3 つの障害処理層が明らかになりました。ツールの障害は完全に回復し (1.0)、曖昧な委任は部分的に回復しました (0.30)、および 3 つの潜在モードまたはセマンティック モードは決して回復しませんでした (0.0)。この順序は、計算が融資承認ワークフローとして再構成されたときも、Sonnet、Opus、Haiku 全体にわたって持続しましたが、絶対金利は状況に応じて変化しました。ブラインド再試行では潜在的な障害が再現され、検出までの時間が増加しました。これは、検出と原因特定が封じ込めには必要であることを示しています。カスケード半径はパイプラインの深さとともに増加しました (深さ 3 ~ 7 で平均 0.9 ~ 4.7)。信頼できる状態の修復アブレーションにより、明らかな封じ込めの向上は主に自律的な検出ではなく信頼できる状態の信号から得られることが示されました。これらの結果は、制御されたチェーン メカニズムの調査であり、ドメイン ワークロードの主張ではありません。
原文 (English)
OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality
Multi-agent orchestration frameworks are moving from demos to production, yet benchmarks typically report task accuracy without diagnosing why a pipeline failed, where a cascade began, or which routing decision caused the breakdown. OrchestraBench evaluates failure, recovery, and decomposition through a controlled, seed-reproducible failure-injection harness over templated enterprise workflows. It introduces cascade radius and per-failure-mode recovery as primary metrics and compares routing policies with bootstrap confidence intervals and paired tests. On a 26-case gold-labelled diagnostic, a keyword/flag router scored 0% on adversarial cases with misleading or missing surface flags, whereas an intent-reasoning model router scored 100%, matching the oracle. Controlled mechanism probes with a real Claude agent over a verifiable arithmetic dependency chain revealed three failure-handling tiers across five MAST modes: tool faults recovered fully (1.0), ambiguous delegation recovered partially (0.30), and three latent or semantic modes never recovered (0.0). This ordering persisted when the computation was reframed as a loan-approval workflow and across Sonnet, Opus, and Haiku, although absolute rates shifted with context. Blind retry reproduced latent faults and increased time to detection, indicating that detection and attribution are necessary for containment. Cascade radius increased with pipeline depth (mean 0.9 to 4.7 across depths 3-7). A trusted-state repair ablation showed that apparent containment gains primarily came from the trusted-state signal rather than autonomous detection. These results are controlled-chain mechanism probes, not domain-workload claims.
Agentic 自動運転顕微鏡ベンチマークは認定をサポートしますが、必ずしも目に見えないタスクに一般化されるわけではありません
顕微鏡やシンクロトロンビームラインを含む幅広い科学的特性評価ツールを制御するための大規模な言語モデルエージェントの開発が増えています。物理インフラストラクチャのエージェント制御に関する研究は初期段階にあり、エージェント システムを設計する方法について十分に確立されたパラダイムはほとんどありません。顕微鏡エージェントを設計する際には、LLM の選択、使用するエージェントの数、エージェントの責任と委任ルール、検索拡張生成パラメータなど、多くの選択肢があります。エージェント顕微鏡コントローラーを設計および最適化する場合、研究者は、エージェントが既知のタスクを正しく実行できることを確認するだけでなく、エージェントがこれまでに遭遇したことのない新しいタスクに一般化できることも確認したいと考えています。この研究では、a) エージェント アーキテクチャのさまざまな選択が顕微鏡タスクのパフォーマンスにどのような影響を与えるか、b) 特定のエージェントが目に見えない顕微鏡タスクで適切にパフォーマンスを発揮するかどうかを予測するためのベンチマークの限界を明らかにするベンチマークおよびトレース ロギング フレームワークを開発します。このフレームワークは、1 エージェント、2 エージェント、および 3 エージェントのグラフ トポロジ、5 つの LLM、RAG およびコンテキスト パラメーター、および 53 の顕微鏡ベンチマーク テストにわたる操作上の制約を評価するために使用されました。合計で、105 のエージェント構成、1,949 の個別のテスト実行、および 49,109 の RAG 取得が記録されました。直接比較すると、構成間でレイテンシ、トークン使用量、コスト、障害モードに明確な違いがあることがわかりました。ただし、エージェントのアーキテクチャとテスト結果に基づいてトレーニングされたサロゲート モデルでは、新しい未知のタスクにおけるエージェントのパフォーマンスを確実に予測できませんでした。これらの結果は、これらのベンチマークが適格性評価、回帰テスト、診断、および直接比較に役立つことを示していますが、現在の異種テスト スイートはタスクに依存しないグローバル構成モデルをサポートしていません。
原文 (English)
Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks
Large language model agents are increasingly being developed to control a wide range of scientific characterization tools including microscopes and synchrotron beamlines. Research into agentic control of physical infrastructure is nascent and there are few well-established paradigms for how to engineer an agentic system. There are many choices to make when designing a microscopy agent, including the choice of LLM, the number of agents to use, agent responsibilities and delegation rules, retrieval-augmented generation parameters, and more. When designing and optimizing an agentic microscope controller, researchers not only want to ensure that the agent can correctly perform known tasks but also that the agent can generalize to new tasks that it has not encountered before. In this study, we develop a benchmark and trace-logging framework that reveals a) how different choices of agent architecture impact performance at microscopy tasks and b) the limitations of benchmarks for predicting if a particular agent will perform well on unseen microscopy tasks. The framework was used to evaluate one-, two-, and three-agent graph topologies, five LLMs, RAG and context parameters, and operational constraints across 53 microscopy benchmark tests. In total, 105 agent configurations, 1,949 individual test runs, and 49,109 RAG retrievals were recorded. Direct comparisons showed clear differences in latency, token use, cost, and failure mode between configurations. However, surrogate models trained on agent architecture and test results did not reliably predict an agent's performance on new, unseen tasks. These results show that these benchmarks are useful for qualification, regression testing, diagnosis, and direct comparison, but the current heterogeneous test suite does not support a task-independent global configuration model.
CASCADE: 患者データで検証された下流の摂動予測のためのエージェント的規制ネットワーク フレームワーク
CASCADE は、MCP を介して公開される、事前計算された ARACNe 制御ネットワークからの遺伝子摂動の下流転写効果を予測するエージェント フレームワークです。これまでの研究では、予測された遺伝子が既知のがん遺伝子(メンバーシップ)であるかどうかをチェックすることによって、そのようなツールを検証しました。代わりに、実際のTCGA患者の腫瘍データに対するノックダウンの逆数の用量ベースの代用として焦点遺伝子のコピー数増幅を使用して、予測された変化の方向が現実と一致するかどうかをテストします。 MYC の場合、CASCADE の予測ノックダウン ターゲットは、3 種類の癌(BRCA: 90.0%、COAD: 72.0%、STAD: 85.7%、すべて p<0.0013)にわたる実際の増幅腫瘍発現と非増幅腫瘍発現との強い一致を示し、順列ベースラインを大きく上回り、PAM50 サブタイプ コントロールで生存し、独立したコホートで複製します(METABRIC、 87.2%)。フィッシャーの直接確率検定を介して厳選された MSigDB 遺伝子セットのベースラインと比較すると、CASCADE の精度が MYC または E2F 駆動の生物学に関する既存の公的知識を超えることは示されていませんが、その遺伝子固有の方向呼び出しは明らかに単純な均一な推測を上回っています。さらに15個の遺伝子にまで拡張した検証により、普遍的ではなく遺伝子特異的であることが証明された。増殖機構制御因子はほとんど複製するが、系統同一性転写因子と1つのサイクリンDパラログ(CCND2)は一貫して失敗するという、ヘッジされた事後仮説として我々が議論するパターンである。 LLM ベースのエージェントが自然言語リクエストを CASCADE の実際の MCP ツール呼び出しに正しく組み込むかどうかを個別にベンチマークします。 35 のクエリ全体で、文書化されたローカル モデルは 71.4% の完全一致に達しました (より大きなモデルでは 85.7%)。スキーマおよび遺伝子エイリアスの障害は、スケールまたはサーバー側の修正によって解決されますが、両方のモデルは、曖昧なクエリに対して自信を持って間違った摂動タイプをデフォルトに設定します。この障害は、トリガー条件が決して発生しないため、ターゲットを絞った修正では解決できません。
原文 (English)
CASCADE: An Agentic Regulatory Network Framework for Patient-Data-Validated Downstream Perturbation Prediction
CASCADE is an agentic framework that predicts downstream transcriptional effects of gene perturbation from precomputed ARACNe regulatory networks, exposed via MCP. Prior work validates such tools by checking whether predicted genes are known cancer genes (membership); we instead test whether the predicted direction of change matches reality, using focal-gene copy-number amplification as a dosage-based proxy for the inverse of knockdown against real TCGA patient tumor data. For MYC, CASCADE's predicted knockdown targets show strong concordance with real amplified-vs-non-amplified tumor expression across three cancer types (BRCA: 90.0%, COAD: 72.0%, STAD: 85.7%; all p<0.0013), well above permutation baselines, surviving a PAM50 subtype control and replicating in an independent cohort (METABRIC, 87.2%). Compared against curated MSigDB gene-set baselines via Fisher's exact test, CASCADE's accuracy is not shown to exceed existing public knowledge of MYC- or E2F-driven biology, though its gene-specific direction-calling clearly outperforms a naive uniform guess. Extending to fifteen additional genes, validation proves gene-specific rather than universal: proliferation-machinery regulators mostly replicate, while lineage-identity transcription factors and one cyclin-D paralog (CCND2) consistently fail, a pattern we discuss as a hedged, post-hoc hypothesis. We separately benchmark whether an LLM-based agent correctly grounds natural-language requests into CASCADE's real MCP tool calls. Across 35 queries, a documented local model reaches 71.4% exact match (85.7% for a larger model); schema and gene-alias failures are resolved by scale or server-side correction, but both models confidently default to the wrong perturbation type on ambiguous queries, a failure a targeted fix could not resolve because its trigger condition never occurs.
大規模な言語モデルによる反事実分析
反事実分析は、仮説的なシナリオの下で潜在的な結果を予測し、意思決定に貴重な洞察を提供することを目的としています。この論文では、反事実分析のための大規模言語モデル (LLM)、特に GPT-3.5 モデルの適用について調査します。私たちはオンライン融資のコンテキストに焦点を当てており、さまざまな金利スキームを評価するには反事実的な投資収益率 (ROI) が重要です。まず GPT の予測パフォーマンスを評価し、それを高度な機械学習アルゴリズムと比較します。結果は、迅速なエンジニアリングにより GPT の予測が大幅に向上し、R 二乗が 1.97% から 2.84% に増加し、勾配ブースト回帰によって達成される 3.48% にほぼ近づくことが示されました。その後、GPT を利用して、一連の代替金利の下で反事実的な ROI を生成します。 GPT は、その応答において論理的一貫性と因果推論を示します。この調査結果は、オンライン融資における反事実分析の効果的なツールとしての LLM の可能性を強調しており、さまざまな予測や意思決定の状況において LLM がより広範に応用できることを示唆しています。
原文 (English)
Counterfactual Analysis via Large Language Models
Counterfactual analysis aims to predict potential outcomes under hypothetical scenarios, offering valuable insights for decision-making. This paper investigates the application of large language models (LLMs), specifically the GPT-3.5 model, for counterfactual analysis. We focus on the online lending context, where the counterfactual return on investment (ROI) is crucial for evaluating different interest rate schemes. We begin by assessing the predictive performance of GPT and comparing it with advanced machine learning algorithms. The results show that prompt engineering can significantly enhance GPT's predictions, with the R-squared increasing from 1.97% to 2.84%, closely approaching the 3.48% achieved by gradient-boosted regression. Subsequently, we utilize GPT to generate counterfactual ROIs under a set of alternative interest rates. GPT exhibits logical coherence and causal reasoning in its responses. The findings underscore the potential of LLMs as effective tools for counterfactual analysis in online lending, suggesting broader applications for LLMs in various predictive and decision-making contexts.
DoctorAgents: 小規模な臨床時間データの AutoML パイプラインを反復的に改良するためのエージェント フレームワーク
臨床機械学習 (ML) には、一か八かの医療上の意思決定をサポートする可能性がありますが、信頼性の高い導入は、希少性、異種性、時間的な複雑さによって制約されることがよくあります。このようなデータに対する効果的な ML パイプラインの開発には依然として時間がかかり、エラーが発生しやすくなります。一方、既存の自動機械学習 (AutoML) システムは、事前定義された空間に対する総当たり検索に主に依存しており、明示的な推論と記憶が欠如しているため、この課題への部分的な対処しかできません。したがって、私たちは、網羅的な検索から推論主導の絞り込みまで、小規模な臨床データ用の AutoML を再定式化します。私たちは、生成、検証、改良のための専用の大規模言語モデル (LLM) エージェントを通じてエンドツーエンドの ML パイプラインを自律的に構築および最適化するエージェント AI フレームワークである DoctorAgents を提案します。 DoctorAgents は、テキストの勾配降下法を通じて自然言語フィードバックを逆伝播し、徹底的な検索を行わずにターゲットを絞った更新を実行します。多様な臨床タスクにわたる実験では、DoctorAgents が確立された AutoML ベースラインを常に上回り、より解釈可能なタスク固有の表現を生成することが示されています。
原文 (English)
DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinical temporal data
Clinical machine learning (ML) has the potential to support high-stakes medical decision-making, but reliable deployment is often constrained by scarce, heterogeneous, and temporal complexity. Developing effective ML pipelines for such data remains time-consuming and error-prone, while existing automated machine learning (AutoML) systems only partially address this challenge because they largely rely on brute-force search over predefined spaces and lack explicit reasoning and memory. We therefore reformulate AutoML for small clinical data from exhaustive search to reasoning-driven refinement. We propose DoctorAgents, an agentic AI framework that autonomously constructs and optimizes end-to-end ML pipelines through specialized large language model (LLM) agents for generation, validation, and refinement. DoctorAgents backpropagates natural-language feedback through textual gradient descent to perform targeted updates without exhaustive search. Experiments across diverse clinical tasks show that DoctorAgents consistently outperforms established AutoML baselines while producing more interpretable task-specific representations.
C$^3$PO: オムニモーダル モデルにおけるクロスモーダル構成と反事実パフォーマンスの評価
現在のマルチモーダル大規模言語モデル (MLLM) は、多様な感覚入力を処理できますが、その推論は依然として主要なモダリティに大きく偏っており、クロスモーダル推論が脆弱になります。ビデオ、オーディオ、画像、テキストにわたる 3,404 個のサンプルのベンチマークである C$^3$PO を紹介し、情報構成 (分散した証拠を融合する) と反事実的対立 (意図的な矛盾を解決する) という 2 つの能力を評価します。 C$^3$PO のペアの IC/CC 構造と 4 層設計により、クロスモーダル推論が失敗する時期と理由を対象とした診断が可能になります。 25 個の論理的に根拠のあるテンプレートを使用した完全自動パイプラインを通じて構築された C$^3$PO は、人間が 88.64% の精度を達成する一方で、最良のモデル (Gemini-3.1-Pro) は 73.17% にとどまり、オープンソース モデルは競合すると崩壊することを明らかにしています。注意プローブを通じて、失敗の 86 ~ 95% がモダリティの優位性に起因していることがわかりました。モデルは矛盾する証拠を無視しながら 1 つのモダリティにコミットし、注意の 87 ~ 95% をテキストに集中させます。中間層の注意エントロピーは、正しさを維持した探索が成功し、早期崩壊が失敗すると予測します。同様に複雑なテンプレート間の 56 ポイントの精度の差は、パフォーマンスが、組み合わせではなく、競合解決におけるモダリティの構造的役割に依存していることを明らかにしています。これらの発見は、多峰性の知覚が堅牢な推論を保証しないことを示しています。アーキテクチャは、時期尚早な行動を避けるために、持続的なクロスモーダルな注意を可能にする必要があります。
原文 (English)
C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models
Current Multimodal Large Language Models (MLLMs) can process diverse sensory inputs, yet their reasoning remains heavily biased toward a dominant modality, resulting in brittle cross-modal reasoning. We introduce C$^3$PO, a benchmark of 3,404 samples spanning video, audio, image, and text, evaluating two abilities: information composition (fusing dispersed evidence) and counterfactual conflict (resolving deliberate contradictions). C$^3$PO's paired IC/CC structure and four-tier design enable targeted diagnosis of when and why cross-modal reasoning fails. Built through a fully automatic pipeline using 25 logically grounded templates, C$^3$PO reveals that while humans achieve 88.64% accuracy, the best model (Gemini-3.1-Pro) reaches only 73.17%, with open-source models collapsing under conflict. Through attention probes, we find 86-95% of failures stem from modality dominance: models commit to one modality while ignoring contradictory evidence, concentrating 87-95% of attention on text. Mid-layer attention entropy predicts correctness-sustained exploration succeeds, premature collapse fails. The 56-point accuracy gap between equally complex templates reveals that performance depends on modalities' structural roles in conflict resolution, not combinations. These findings show multimodal perception does not guarantee robust reasoning; architectures must enable sustained cross-modal attention to avoid premature
オープンエンドのケアプラン調整のための適応アリーナベースの議論可能な専門家ネットワーク
ケアプランの調整では、複数の専門分野にわたる異種の臨床、機能、心理社会的情報を統合する必要がありますが、モノリシック LLM パイプラインは透明性または安全な方法で実行できません。我々は、複雑性の評価、適応型チームの採用、アリーナベースの定量的双極議論フレームワーク (A-QBAF) を介した役割ベースの議論的計算、人間参加型の議論、ケアプラン合成という 5 つのモジュールを通じてこれらの制限に対処するマルチエージェント神経象徴フレームワークである CANOE (Contestable Argumentative Network-of-Experts) を紹介します。役割に特化したエージェントは、候補者の介入に対して支持および攻撃の議論を生成します。競合は、許容スコアが議論グラフ全体に伝播する前に、アリーナベースの競合解決を通じて解決されます。ケアプランナーは引数を受け入れる、拒否する、編集する、または追加することができ、フレームワークは最終的な計画を決定論的に再計算します。 『Discharge Me!』の評価ROUGE-L、AlignScore、MEDCON F1、FKGL、および LLM を審査員として使用した MedicalRAG は、医学的に微調整されたモデルが最も強力な臨床的正確性と安全性を達成する一方、CANOE の議論構造が忠実な説明と人間による異議の可能性を提供することを示しています。
原文 (English)
Adaptive Arena-based Contestable Argumentative Network-of-Experts for Open-Ended Care Plan Coordination
Care plan coordination demands synthesizing heterogeneous clinical, functional, and psychosocial information across multiple professional disciplines, where monolithic LLM pipelines cannot perform in a transparent or safe manner. We present CANOE (Contestable Argumentative Network-of-Experts), a multi-agent neuro-symbolic framework that addresses these limitations through five modules: complexity assessment, adaptive team recruitment, role-based argumentative computation via an Arena-based Quantitative Bipolar Argumentation Framework (A-QBAF), human-in-the-loop contestation, and care-plan synthesis. Role-specialized agents generate supporting and attacking arguments for candidate interventions; conflicts are resolved through arena-based clash resolution before acceptability scores propagate across the argumentation graph. Care planners may accept, reject, edit, or add arguments, and the framework will deterministically recompute the final plan. Evaluation on Discharge Me! and MedicalRAG using ROUGE-L, AlignScore, MEDCON F1, FKGL, and LLM-as-a-judge shows that medically fine-tuned models achieve the strongest clinical correctness and safety, while CANOE's argumentative structure provides faithful explanation and human contestability.
教育的適合性インデックスを使用した LLM ベースの AI 家庭教師の教育的適合性の評価と改善
大規模言語モデル (LLM) が AI の家庭教師として使用されることが増えていますが、正解が必ずしも教育学的に適切であるとは限りません。教室での学習では、効果的な支援は正しさだけでなく、回答が学習者の現在の基礎、コースの順序、概念導入のタイミングと一致するかどうかにも依存します。既存の評価は主に解答の質に焦点を当てており、この指導的適合性は十分に評価されていません。我々は、LLM によって生成された個別指導の応答が学習者の準備やカリキュラムの進行とどの程度一致しているかを評価する 6 つの理論に基づいたサブスコアの複合指標である教育適性指数 (PSI) を提示します。また、応答を改善するための構造化されたフィードバック シグナルとして PSI をさらに使用します。標準プロンプトと欠陥プロンプトのペアを使用して、240 のシナリオベースの評価にわたって 4 つの LLM チューター (ChatGPT、Gemini、Gemma4、および Qwen3) を評価し、次に PSI ガイドによる再生成プロトコルを 62 のパフォーマンスの低いケースに適用します。テストした 4 つのモデル間のベースラインの差は全体的には控えめで (PSI 範囲: 0.557 ~ 0.638)、オープンウェイト モデルとクローズド モデルでは教育的適合において明確な区別は示されませんでした。テストされたプロンプト変動の下では、サブスコアのトレードオフが現れましたが、全体的な PSI はほぼ安定していました (デルタ = -0.002)。さらに重要なことは、PSI に基づくフィードバックにより、パフォーマンスの低いケースが大幅に改善されたことです。62 ケース中 51 ケース (82.3%) が改善されました。 PSI が選択した 62 個の弱いケースを手動で集中的に評価することで、特定された弱点が指導的に意味があり、PSI に基づく再生の多くが人間が判断した改善に対応するという最初の証拠が得られます。これらの結果は、効果的な個別指導には、モデル カテゴリ単独よりも学習者とカリキュラムを意識した調整の方が重要である可能性があり、そのような調整は測定可能であり、改善可能であることを示唆しています。
原文 (English)
Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index
Large language models (LLMs) are increasingly used as AI tutors, but a correct answer is not always a pedagogically appropriate one. In classroom learning, effective help depends not only on correctness, but also on whether a response matches the learner's current foundation, the course sequence, and the timing of concept introduction. Existing evaluations focus mainly on answer quality, leaving this instructional fit under-measured. We present the Pedagogical Suitability Index (PSI), a composite metric of six theory-informed sub-scores that evaluates how well LLM-generated tutoring responses align with learner readiness and curricular progression, and we further use PSI as a structured feedback signal for response improvement. We evaluate four LLM tutors (ChatGPT, Gemini, Gemma4, and Qwen3) across 240 scenario-based evaluations using paired standard and defective prompts, then apply a PSI-guided regeneration protocol to 62 weak-performing cases. Baseline differences across the four tested models were modest overall (PSI range: 0.557 to 0.638), and open-weight and closed models did not exhibit a clear separation in pedagogical fit. Under the tested prompt perturbations, overall PSI remained largely stable (Delta = -0.002), though sub-score trade-offs emerged. More importantly, PSI-guided feedback substantially improved weak-performing cases: 51 of 62 cases improved (82.3%). Focused manual evaluation of the 62 PSI-selected weak cases provides initial evidence that the identified weaknesses are instructionally meaningful and that many PSI-guided regenerations correspond to human-judged improvement. These results suggest that learner- and curriculum-aware alignment may matter more for effective tutoring than model category alone, and that such alignment is both measurable and improvable.
利害関係者の混合による審議を通じて、警察のための AI におけるリスク境界を交渉する
英国および世界中の警察で AI ツールがますます採用されています。人種的偏見は既知のリスクであり、十分に文書化されていますが、影響を受けるコミュニティの代表者が AI 導入に関する決定に含まれることはほとんどありません。我々は、人種的偏見に明確に焦点を当て、警察活動における 13 の AI ユースケースのリスクを評価するために、地域の代表者、警察官、学者 30 名を集めた利害関係者混合の審議ワークショップの結果を紹介します。参加者は AI の導入に広くオープンであり、完全に拒否したユースケースは 3 つだけで、最も顕著なのは再犯リスク評価であり、反対意見は導入ではなく前提に焦点を当てていたことがわかりました。私たちの分析では、人種的平等を前面に押し出したことが審議の幅を狭めなかったことが明らかになりました。その代わりに、議論は一連の基本的な質問に引き寄せられました。このツールは実際に機能するのか、真の利益をもたらすのか、そしてその利益はすべての人に及ぶのかということです。この統合された推論は、インクルーシブ デザインにおける縁石カット効果を彷彿とさせ、AI ユースケースのリスクと利益の分析に最初から人種的偏見のレンズを組み込むことの利点を強調しています。
原文 (English)
Negotiating Risk Boundaries in AI for Policing Through Mixed-Stakeholder Deliberation
AI tools are being increasingly adopted in policing in the UK and worldwide. Racial bias is a known and well-documented risk, yet representatives of affected communities are rarely included in decisions about AI adoption. We present results from a mixed-stakeholder deliberation workshop bringing together 30 community representatives, police officers, and academics to assess the risks of 13 AI use cases in policing, with an explicit focus on racial bias. We found that participants were broadly open to AI adoption, rejecting only three use cases outright, most notably recidivism risk assessment, where objections targeted the premise rather than the implementation. Our analysis reveals that foregrounding racial equity did not narrow the deliberation. Instead, discussions gravitated toward a fundamental set of questions: does this tool actually work, will it deliver genuine benefit, and will that benefit extend to everyone? This integrated reasoning, reminiscent of the curb-cut effect in inclusive design, highlights the benefit of incorporating the racial bias lens into the risk-benefit analysis of AI use cases from the outset.
SCP-NL2TL: 自然言語から時相論理仕様までの意味検証を伴う選択的等角予測
自然言語の命令を機械が解釈できる正式な仕様に変換すると、ロボットや自律システムがその動作を計画、推論し、正式に検証できるようになります。ただし、既存の変換モデルは通常、結果が信頼できない場合やユーザーの意図を捉えられない場合でも、すべての入力に対して仕様を生成するため、安全性が重要なアプリケーションではリスクが生じます。選択的等角予測にヒントを得て、正式な仕様を生成するだけでなく、それがいつ信頼できるかを判断する選択的翻訳フレームワークを提案します。信頼性は、2 つの相補的なブラック ボックス シグナル、つまり自然言語に逆翻訳された仕様の忠実性と、厳密な意味的同等性の下で繰り返される翻訳の分散によってスコア付けされます。これらのシグナルは、異なるエラーで失敗し、どちらか単独よりも共同して不正確な翻訳をより明確に分離します。コンフォーマルリスクコントロールは、不正な仕様が実行のために受け入れられる割合に関する分布フリーの制限を使用して、このスコアを仕様を受け入れるか棄権するかの決定に調整します。また、命令埋め込みのコンフォーマル異常検出器は、変換が試行される前に分布外の入力をスクリーニングします。提案されたフレームワークは形式仕様言語全体に共通であり、信号時論理 (STL)、線形時相論理 (LTL)、および幾何学的時空間論理 (SpaTiaL) の実験により、翻訳の信頼性の向上、評価された層間のシフト下での堅牢性、および効果的な不確実性を意識した棄権が実証されました。この取り組みにより、生成された仕様が信頼できない場合を AI システムが認識できるようになり、信頼できる自然言語インターフェイスの基盤が確立されます。
原文 (English)
SCP-NL2TL: Selective Conformal Prediction with Semantic Verification for Natural Language to Temporal Logic Specifications
Translating natural language instructions into machine-interpretable formal specifications enables robots and autonomous systems to plan, reason, and formally verify their behavior. However, existing translation models typically generate a specification for every input, even when the result is unreliable or fails to capture the user's intent, creating risks in safety-critical applications. Inspired by selective conformal prediction, we propose a selective translation framework that not only generates formal specifications but also determines when they can be trusted. Reliability is scored by two complementary black-box signals, the fidelity of the specification back-translated into natural language and the dispersion of repeated translations under exact semantic equivalence, which fail on different errors and jointly separate incorrect translations more sharply than either alone. Conformal risk control calibrates this score into a decision that accepts a specification or abstains, with a distribution-free bound on the rate at which incorrect specifications are accepted for execution, and a conformal anomaly detector on instruction embeddings screens out-of-distribution inputs before any translation is attempted. The proposed framework is general across formal specification languages, with experiments on Signal Temporal Logic (STL), Linear Temporal Logic (LTL), and geometric Spatio-Temporal Logic (SpaTiaL) demonstrating improved translation reliability, robustness under the evaluated cross-tier shifts, and effective uncertainty-aware abstention. This work establishes a foundation for trustworthy natural language interfaces by enabling AI systems to recognize when generated specifications may not be reliable.
確率性は難しい部分ではありません: 前提条件となる DAG による命令シーケンスの削減と複雑さ
学生が前提条件となる依存関係によって関連付けられた概念を学習する必要がある場合、指導の順序が重要になるのはどのような場合ですか、最適な順序を見つけるにはいくらかかりますか?私たちは、概念の試みが状態依存の確率で成功し、失敗すると学習者の状態が変化しないという確率的最短経路問題として、命令順序付けを研究します。我々はまず、この確率性を正確に排除できることを証明します。この問題は、最適な値とアクションを保存しながら、前提条件の順序理想の格子上の決定論的な最短経路問題に崩壊します。この崩壊により、確率論的な複雑性は除去されますが、組み合わせの複雑さは除去されません。最適なシーケンスは、前提条件のエッジ、ユニットコスト、均一なバイナリの非負転送、少なくとも $1/2$ の成功確率がない場合でも、トーナメントで設定されたフィードバック アークからの削減によって NP ハードのままです。硬さは均一ではありません。実現可能な転送優先度が前提条件と結合して非巡回のままである場合、残差結合グラフの位相順序は最適であり、前提条件の幅が固定されているため、多項式時間の正確な動的計画法が得られます。計算可能な診断 $m\Delta$ は、最適化前のシーケンスの値を制限します。 CS 入門コースからの 70,893 のインタラクションでは、診断は二重に簡単な体制 (最適化する価値がほとんどなく、検索するスペースもほとんどない) を証明しますが、構築された転送インスタンスは、近視眼的なシーケンスでは大きな後悔が生じるものの、一貫したヒューリスティックによる正確な A* がそのファミリー上の多くの状態を直線的にのみ拡張するという困難な体制を実現します。
原文 (English)
Stochasticity Is Not the Hard Part: Reduction and Complexity in Instructional Sequencing over Prerequisite DAGs
When a student must learn concepts connected by prerequisite dependencies, when does the order of instruction matter, and what does it cost to find the best one? We study instructional sequencing as a stochastic shortest-path problem in which attempting a concept succeeds with a state-dependent probability and failure leaves the learner state unchanged. We first prove that this stochasticity can be eliminated exactly: the problem collapses to a deterministic shortest-path problem on the lattice of prerequisite order ideals, preserving optimal values and actions. The collapse removes stochastic complexity but not combinatorial complexity: optimal sequencing remains NP-hard -- via reduction from feedback arc set in tournaments -- even with no prerequisite edges, unit costs, uniform binary nonnegative transfer, and success probabilities at least $1/2$. Hardness is not uniform: when realizable transfer preferences remain jointly acyclic with the prerequisites, any topological order of the residual joint graph is optimal, and fixed prerequisite width yields polynomial-time exact dynamic programming. A computable diagnostic, $m\Delta$, bounds the value of sequencing before optimization. On 70,893 interactions from an introductory CS course, the diagnostic certifies a doubly easy regime -- little value to optimize and little space to search -- while constructed transfer instances realize the challenging regime, where myopic sequencing suffers large regret yet exact A* with a consistent heuristic expands only linearly many states on that family.
長期にわたるターミナルタスクのための再帰的合成
ターミナル エージェント向けの高品質で長期的なトレーニング データは作成に費用がかかり、タスクごとに数百ドルから数千ドルかかることがよくあります。これは、各タスクが命令、環境、参照ソリューション、検証器の相互一貫性を保つ必要があるためです。人間によるオーサリングは拡張性がなく、大規模言語モデル (LLM) を使用した直接生成では、これらの依存関係が壊れることがよくあります。我々は、長期にわたるターミナル エージェント タスクを大規模に構築するための再帰的検証済み合成フレームワークである再帰的合成ターミナル タスク (RST) を紹介します。 RST は、検証されたシード タスクから開始して、参照ソリューションを拡張し、検証ツールと命令を新しいワークフローに再調整し、新しいサンドボックスで結果を検証し、受け入れられたタスクを後続のラウンドのシードとして再利用します。 15 回の再帰ラウンドにわたって、RST は 37,484 個の合成ターミナル エージェント タスクをタスクあたり約 $0.05 で生成します。タスクの難易度はラウンドを重ねるごとに大幅に増加します。リファレンス ソリューションの中央値は 67 行から 374 行に増加し、実行されたコマンド数の中央値は 40 から 244 に増加し、DeepSeek-V4-Pro pass@4 は $R_1$ の 90\% から $R_{15}$ の 2.5\% に低下します。トレーニングの有用性を実証するために、合成されたタスクに関して拒否サンプリングされた Qwen3.5 軌跡を収集し、それらを教師付き微調整に使用します。これらの軌道を微調整すると、Qwen3.5-27B と Qwen3.5-122B-A10B がターミナル ベンチ ~ 2、ターミナル ベンチ ハード、およびロングホライズン ターミナル ベンチで最大 10 ポイント改善され、エージェント PPO により Qwen3.5-27B が 3 つで 49.44\%、32.00\%、22.07\% に上昇しました。ベンチマークは、ベース モデルに対して 20.0\%、41.2\%、および 21.9\% の相対的な向上に相当します。さらに、15 ラウンド後の再帰には上限がありません。難易度が上昇し続けても合成収率と検証率は安定しており、ここで報告した規模をはるかに超えてプロセスを継続できることを示しています。
原文 (English)
Recursive Synthesis for Long-Horizon Terminal Tasks
High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier mutually consistent. Human authoring does not scale, and direct generation with large language models (LLMs) often breaks these dependencies. We present Recursive Synthetic Terminal Tasks (RST), a recursive verified synthesis framework for constructing long-horizon terminal-agent tasks at scale. Starting from verified seed tasks, RST extends the reference solution, realigns the verifier and instruction to the new workflow, validates the result in a fresh sandbox, and reuses accepted tasks as seeds for subsequent rounds. Across fifteen recursive rounds, RST produces 37,484 synthesized terminal-agent tasks at roughly \$0.05 per task. Task difficulty increases substantially over rounds: the median reference solution grows from 67 to 374 lines, the median number of executed commands grows from 40 to 244, and DeepSeek-V4-Pro pass@4 drops from 90\% at $R_1$ to 2.5\% at $R_{15}$. To demonstrate training utility, we collect rejection-sampled Qwen3.5 trajectories on the synthesized tasks and use them for supervised fine-tuning. Fine-tuning on these trajectories improves Qwen3.5-27B and Qwen3.5-122B-A10B by up to 10 points on Terminal-Bench~2, Terminal-Bench Hard, and Long-Horizon Terminal Bench, while agentic PPO lifts Qwen3.5-27B to 49.44\%, 32.00\%, and 22.07\% on the three benchmarks, corresponding to relative gains of 20.0\%, 41.2\%, and 21.9\% over the base model. Moreover, after 15 rounds, the recursion shows no ceiling: synthesis yield and validation rates remain stable as difficulty keeps climbing, indicating that the process can continue well beyond the scale reported here.
自律型分析エージェントの革新的残留監査: ローカリゼーション、検出限界、エラー制御、識別可能性
自律型エージェントは、少しずつ段階的な監視を受けながら、コホートの選択、テーブルの結合、モデルのフィッティングなど、データ全体の分析を実行するようになりました。そのような分析が間違っていることが判明した場合、誰かがどの操作がその原因となったのかを特定しなければなりません。最近のアプローチでは、これをラベル付けされた間違いなしで実行し、代わりに健全であることが知られている分析から学習し、モデルの予測から逸脱した操作にフラグを立てます。このような監査がどの程度信頼できるかは研究されていません。この論文ではその分析を提供します。スコアの選択によって、エラーの位置を特定できるかどうかが決まります。各操作がその直前の操作の意外性によってスコア付けされる場合、以前のエラーを単に継承する操作と正しい操作との区別がつかないため、1 つの間違いで 1 つのフラグが生成されます。意図した分析のより長い再構築に対してスコアが計算されるため、代わりに 1 つのミスが多くの操作に分散されます。誤差がどの程度広がるのか、また誤差が一度ではなく徐々に蓄積する場合の比較長をどのように選択するかを定量化します。次に、単一の監査済み分析内で誤ってフラグが立てられた操作の割合を制御する手順を与え、適合モデルが正しいかどうかではなく、健全な分析が交換可能であることだけを要求し、モデルが不完全な場合、または分析が内容に応じた方法でレビュー用に選択された場合に保証がどの程度弱まるかを定量化します。最後に、このような監査で報告できる内容に制限を設けます。一定の大きさ以下のエラーはまったく原因とすることができず、音響分析間の通常の変動と区別できません。この制限は、より多くのサウンド分析が収集されるにつれてゆっくりと低下するため、現在使用されている表現サイズでは 100 倍に増加すると 2 パーセント未満に減少するため、トレーニング データの量ではなく表現の次元が拘束制約となります。
原文 (English)
Innovation-Residual Auditing of Autonomous Analysis Agents: Localization, Detection Limits, Error Control, and Identifiability
Autonomous agents now carry out entire data analyses, selecting cohorts, joining tables, and fitting models with little step-by-step supervision. When such an analysis turns out to be wrong, someone must determine which operation caused it. A recent approach does this without any labelled mistakes, learning instead from analyses known to be sound and flagging operations that depart from what that model predicts; how reliable such audits are has not been studied. This paper supplies that analysis. The choice of score determines whether an error can be localized at all. If each operation is scored by how surprising it is given the operation immediately preceding it, then operations that merely inherit an earlier error are indistinguishable from correct ones, so one mistake produces one flag; scores computed against a longer reconstruction of the intended analysis instead spread a single mistake across many operations. We quantify how far they spread, and how to choose the comparison length when an error accumulates gradually rather than at once. We then give procedures that control the proportion of falsely flagged operations within a single audited analysis, requiring only that sound analyses be exchangeable rather than that the fitted model be correct, and we quantify how much the guarantees weaken when the model is imperfect or when the analysis was selected for review in a way that depends on its content. Finally we establish a limit on what any such audit can report: errors below a certain magnitude cannot be attributed at all, being indistinguishable from ordinary variation among sound analyses. This limit falls so slowly as more sound analyses are collected that at the representation sizes now in use a hundredfold increase reduces it by under two percent, so the dimension of the representation rather than the volume of training data is the binding constraint.
EcoAgent-Bench: 予算に制約のある LLM エージェントにおける経済的意思決定の評価
エージェントのベンチマークは通常、タスクの完了を測定し、リソースの使用状況を補助的な統計として扱います。ただし、展開では、ローカル検索、広範囲検索、複合調査ツール、より強力なモデル、または人間によるエスカレーションの中から選択することがタスク自体の一部になります。 EcoAgent-Bench を導入します。EcoAgent-Bench では、すべてのタスクで価格付きのアクションと明示的な予算が指定されます。その 304 の実際の派生タスクは、GAIA、HotpotQA、および MuSiQue から適応された 5 つのファミリーにまたがっており、不必要なエスカレーションの回避、ローカル証拠が不十分な場合のエスカレーション、モデル層の選択、およびサポートされていない施設での停止という 4 つの決定をテストします。ツール API およびワークスペース CLI 設定で 7 つの LLM エージェントと、4 つの Oracle スクリプト コントロールを評価します。マイクロ平均精度は一方的なポリシーに報います。常にエスカレーションするコントロールは、保存指向のタスクは失敗しますが、マイクロで高い成功を収めます。したがって、この失敗を明らかにする経済的一貫性スコア (アップグレード指向のファミリー グループと節約指向のファミリー グループの精度の悪い方) も報告します。ツール API エージェントは、わずか 3.9 ~ 24.0% のマイクロ厳密な成功率 (最大 7.3% の経済的一貫性) しか達成できず、多くの場合、保証されたエスカレーションの前に停止するか、安価なタスクに過剰な費用を費やします。しきい値を超える予算のスイープにより、GPT-5.4 のエスカレーション率は 0% からわずか 3% に変化します。これらの結果は、予算内での完了と経済的な行動の選択が異なる性質であることを示しています。タスク バンドル、変換パイプライン、凍結された評価環境、および両方を調査するために必要な整合性バインドされた結果アーティファクトをリリースします。
原文 (English)
EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents
Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local lookup, broad search, composite research tool, stronger model, or human escalation is part of the task itself. We introduce EcoAgent-Bench, in which every task specifies priced actions and an explicit budget. Its 304 real-derived tasks span five families adapted from GAIA, HotpotQA, and MuSiQue, and test four decisions: avoiding unnecessary escalation, escalating when local evidence is insufficient, selecting a model tier, and stopping on unsupported premises. We evaluate seven LLM agents in tool-API and workspace-CLI settings, together with four oracle scripted controls. Micro-averaged accuracy rewards one-sided policies: always-escalate controls achieve high micro success while failing save-oriented tasks. We therefore also report an economic-consistency score (the worse of accuracy on upgrade-oriented and save-oriented family groups) which exposes this failure. Tool-API agents attain only 3.9-24.0% micro strict success (at most 7.3% economic consistency), often either stopping before warranted escalation or overspending on cheap tasks. A threshold-crossing budget sweep changes GPT-5.4's escalation rate from 0% to only 3%. These results show that completion under a budget and economical action selection are distinct properties. We release the task bundle, transformation pipeline, frozen evaluation environments, and integrity-bound result artifacts needed to study both.
Hyper-ES: 降下方向マージによる LLM 推論の効果的な進化戦略
Evolution Strategy (ES) は、リソースに制約のある大規模言語モデル (LLM) 推論のための、勾配ベースの微調整に代わる有望な代替手段です。ただし、ES を 10 億パラメータの LLM に直接適用することは、非常に非効果的です。このような高次元パラメータ空間では、ほとんどのランダムな摂動が有用な更新方向とほぼ直交しており、最適化が不安定になります。我々は、低次元最適化における ES の強みを活かしながら、フルパラメータ探索における ES の弱点を回避する部分空間ベースの ES フレームワークである Hyper-ES を提案します。 LLM パラメータ空間内のランダムな摂動から有用な方向を発見するように ES に依頼する代わりに、Hyper-ES はまず、安価な勾配ベースの微調整を少数実行して降下方向を取得します。各方向自体は限られた改善しか提供しない可能性がありますが、そのスパンは、有用な推論の更新をキャプチャするコンパクトな適応サブスペースを形成します。次に、Hyper-ES は CMA-ES を適用して、この部分空間内の層ごとの DARE-TIES マージ係数を最適化し、ES が任意のフルモデル摂動ではなく、意味のある降下方向の組み合わせを検索できるようにします。 6 つの数学的推論データセットにわたる 3 つの Qwen2.5-Instruct および DeepSeek-R1-Distill バックボーンで Hyper-ES を評価します。結果は、Hyper-ES が一貫して GRPO-LoRA よりも 1% 優れたパフォーマンスを示し、必要なスペース消費量の勾配更新が 10% 少ないことを示しています。コードは https://github.com/kuangrepi/Hyper-ES にあります。
原文 (English)
Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging
Evolution Strategy (ES) is a promising alternative to gradient-based fine-tuning for resource-constrained Large Language Model (LLM) reasoning. However, directly applying ES to billion-parameter LLMs is highly ineffective. In such high-dimensional parameter spaces, most random perturbations are nearly orthogonal to useful update directions, leading to unstable optimization. We propose Hyper-ES, a subspace-based ES framework that avoids the weakness of ES in full-parameter search while exploiting its strength in low-dimensional optimization. Instead of asking ES to discover useful directions from random perturbations in the LLM parameter space, Hyper-ES first performs a small number of inexpensive gradient-based fine-tuning runs to obtain descent directions. Although each direction may provide only a limited improvement on its own, their span forms a compact adaptation subspace that captures useful reasoning updates. Hyper-ES then applies CMA-ES to optimize layer-wise DARE-TIES merging coefficients within this subspace, allowing ES to search over combinations of meaningful descent directions rather than over arbitrary full-model perturbations. We evaluate Hyper-ES on three Qwen2.5-Instruct and DeepSeek-R1-Distill backbones across six mathematical reasoning datasets. Results show that Hyper-ES consistently outperforms GRPO-LoRA by 1% while requiring 10% fewer space-consuming gradient updates. Code at https://github.com/kuangrepi/Hyper-ES.
SkillTV-Bench: スキル強化されたエージェント執行における裁判官のパフォーマンスのベンチマーク
LLM エージェントは、ツールの使用と環境の相互作用を通じて長期的なタスクを実行することが増えており、評価は最終応答のスコアリングから完全な実行の検証に移行しています。スキル強化されたエージェントの場合、検証にはさらに、タスク時のスキルにエンコードされた手順知識が必要です。これは、この知識がどの証拠を検査するか、どの失敗がタスクに重大であるかを示すためです。ただし、既存のジャッジベンチマークは、最終的な応答や静的な軌跡を公開することが多く、タスク時のスキルと直接検査可能なアーティファクトや環境を組み合わせることはほとんどありません。そこで、11 のドメインにわたる 50 のタスクからの実際のエージェントの軌跡の 681 ケースのベンチマークである SkillTV-Bench を紹介します。これは、LLM-as-a-Judge メソッドと Agent-as-a-Judge メソッドの両方のスキルを意識した軌跡検証を評価するように設計されています。さらに、検証知識を再利用可能な JudgeSkill として外部化する SkillTV-Evolve を提案します。これにより、エージェント裁判官が対象を絞った検査を計画し、証拠に基づいた評決を下すことができます。連携のない開発プールでは、自動進化ループにより、誤判定されたケースを使用して JudgeSkill がさらに洗練されます。 SkillTV-Bench では、洗練されたスキルにより、同じエージェントの裁判官の精度が 14.8 パーセント ポイント向上します。オフラインのロールアウト プール選択では、選択された軌道の成功率が 1 回のロールアウトの 22.9% から 10 回のロールアウトの 45.5% に増加します。コードとデータは https://github.com/HanZhi306/SkillTV-Bench で入手できます。
原文 (English)
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution
LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executions. For skill-augmented agents, verification additionally requires the procedural knowledge encoded in task-time skills, because this knowledge indicates what evidence to inspect and which failures are task-critical. However, existing judge benchmarks often expose final responses or static trajectories, and rarely combine task-time skills with directly inspectable artifacts and environments. We therefore introduce SkillTV-Bench, a 681-case benchmark of real agent trajectories from 50 tasks across eleven domains, designed to evaluate skill-aware trajectory verification for both LLM-as-a-Judge and Agent-as-a-Judge methods. Additionally, we propose SkillTV-Evolve, which externalizes verification knowledge as a reusable JudgeSkill that guides an agent judge to plan targeted inspections and issue evidence-grounded verdicts. On a disjoint development pool, an automated evolution loop further refines the JudgeSkill using misjudged cases. On SkillTV-Bench, the refined skill increases the same agent judge's accuracy by 14.8 percentage points. In offline rollout-pool selection, it increases selected-trajectory success from 22.9% with one rollout to 45.5% with ten rollouts. The code and data are available at https://github.com/HanZhi306/SkillTV-Bench
StepReflect: モバイル GUI エージェント用の構造化 UI 遷移リフレクション
自律型モバイル GUI エージェントは、長期にわたって信頼性の高い実行を実現するために、正確なアクションの反映を必要とします。既存のアプローチは、各アクションの後のオープンエンドのマルチモーダル推論に依存していますが、これはコストがかかり、GUI 状態遷移の構造化された性質とあまり適合しません。我々は、ステップごとの GUI リフレクションを、明示的な遷移仕様とペアになった視覚的証拠を条件とした教師付き構造化予測として定式化する StepReflect を提案します。 StepReflect は、教師付き微調整、教師と生徒の蒸留、好みと報酬に基づいた改良を組み合わせた段階的なパイプラインを通じてトレーニングされます。オフラインでは、結果の 8B モデルは AndroidWorld で 82.16% の遷移レベル精度を達成し、同じ構造化入力の下でゼロショット GPT-5.2 を 11.83 パーセントポイント上回りました。オンラインでは、M3A、Agent-SAMA、MAI-UI-8B、および Seed-2.0-Pro 全体で、StepReflect は 4 つのエージェント構成のうち 3 つでより高いタスク成功率を達成し、4 つ目のエージェント構成では GPT-5.2 Reflection Agent の 1 つの成功タスク内に留まります。また、4 つの構成すべてにおいて、GPT ベースのリフレクションと比較して、有料 API 料金も削減されます。これらの結果により、StepReflect は、長期にわたるモバイル GUI エージェントの繰り返しフロンティア モデル リフレクションに代わる実用的でローカルに展開可能な代替手段として確立されました。
原文 (English)
StepReflect: Structured UI Transition Reflection for Mobile GUI Agents
Autonomous mobile GUI agents require accurate action reflection for reliable long-horizon execution. Existing approaches rely on open-ended multimodal reasoning after each action, which is costly and poorly matched to the structured nature of GUI state transitions. We propose StepReflect, which formulates per-step GUI reflection as supervised structured prediction conditioned on explicit transition specifications and paired visual evidence. StepReflect is trained through a staged pipeline combining supervised fine-tuning, teacher-student distillation, and preference- and reward-based refinement. Offline, the resulting 8B model achieves 82.16% transition-level accuracy on AndroidWorld, exceeding zero-shot GPT-5.2 by 11.83 percentage points under the same structured input. Online, across M3A, Agent-SAMA, MAI-UI-8B, and Seed-2.0-Pro, StepReflect achieves higher task success in three of four agent configurations and remains within one successful task of the GPT-5.2 Reflection Agent in the fourth. It also reduces paid API charges relative to GPT-based reflection in all four configurations. These results establish StepReflect as a practical, locally deployable alternative to repeated frontier-model reflection for long-horizon mobile GUI agents.
生成 AI における認識論的信頼性: 一か八かのワークフローにおける信頼を保証するための規範的なフレームワーク
生成 AI システムは、一か八かの専門的な状況で導入されることが増えており、その出力によって、ユーザーが何を信じ、どのように推論し、何を定着として扱うのかが決まります。これは、責任ある AI に対する中心的な疑問を引き起こします。それは、どのような状況下で、生成的な AI の出力への依存が、行動によって引き起こされるのではなく、認識論的に正当化されるのかということです。既存のフレームワークでは主に、AI の出力が正確か、公平か、説明可能か、安全か、ユーザーに信頼されているかどうかが問われます。これらの質問は依然として必要であり、それぞれが正当な信頼に貢献する可能性があります。ただし、それらは、明確な評価目標、つまりユーザーが AI の出力を自分の推論への入力として扱うことが正当化される条件として、保証された信頼性を直接指定していません。私たちは、これには認識論的信頼性、つまりシステムを認識論的に信頼に値するものにするものについての説明が必要であると主張します。能力と聴衆指向としての信頼性に関する哲学的説明に基づいて、私たちは 3 つの共同で必要で代替不可能な条件からなる構成的規範の枠組みを開発します。まず、認識論的謙虚さには、システムがその能力の限界を表現し、伝達することが必要です。第 2 に、認識的アクセスには、ユーザーがコンテキスト内で出力を検査、質問、および異議申し立てできるようにするシステムが必要です。第三に、認識論的不正義に抵抗するには、システムがユーザーを正当な認識論的主体として認識し、ユーザーの知識や経験を疎外しないようにする必要があります。法的推論、医学的推論、雇用における実際の事例分析を通じて、認識論的謙虚さ、認識論的アクセス、認識論的不正義に対する抵抗の失敗が、精度、公平性、使いやすさの標準的な尺度だけでは対処できない結果的な損害をどのように生み出す可能性があるかを示します。最後に、出力の正確さのみではなく、認識論的に保証された信頼性を中心に構成された GenAI システムの設計と評価への影響を概説します。
原文 (English)
Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows
Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how they reason, and what they treat as settled. This raises a central question for responsible AI: under what conditions is reliance on generative AI outputs epistemically warranted rather than behaviourally induced? Existing frameworks largely ask whether AI outputs are accurate, fair, explainable, safe, or trusted by users. These questions remain necessary, and each can contribute to warranted reliance. However, they do not directly specify warranted reliance as a distinct evaluative target: the conditions under which users are justified in treating AI outputs as inputs into their own reasoning. We argue that this requires an account of epistemic trustworthiness: what makes a system epistemically worthy of reliance. Drawing on philosophical accounts of trustworthiness as competence and audience-orientation, we develop a constitutive normative framework comprising three jointly necessary and non-fungible conditions. First, epistemic humility requires systems to represent and communicate the limits of their competence. Second, epistemic access requires systems to enable users to inspect, question, and contest outputs in context. Third, resistance to epistemic injustice requires systems to recognise users as legitimate epistemic agents and avoid marginalising their knowledge and experience. Through real-world case analyses in legal reasoning, medical reasoning, and hiring, we show how failures of epistemic humility, epistemic access, and resistance to epistemic injustice can produce consequential harms that standard measures of accuracy, fairness, and usability do not address on their own. We conclude by outlining design and evaluation implications for GenAI systems organised around epistemically warranted reliance rather than output correctness alone.
有害な AI おべっかの測定と検出
おべっかな反応は大規模言語モデル (LLM) に蔓延しており、以前の研究では、その一部が有害である可能性があることが指摘されています。この論文では、有害なお調子者の 1 つである好み誘発スタンス逆転お調子者 (PSRS) に焦点を当てます。PSRS では、ユーザーが表明した好みに合わせるためだけにモデルが最初のスタンスを反転します。既存の研究は主にモデルがお調子者であるかどうかを測定するものですが、私たちはさらに踏み込んで、単一の応答から PSRS を自動的に検出できるかどうかを調べます。これを大規模に調査するために、ラベル付き PSRS データを収集するためのフレームワークである CAP (Contrastive Anchor Probing) を導入します。 17 のオープンソースおよびクローズドソース LLM に CAP を適用し、12 の日常アドバイス ドメインにわたって 290,460 件のラベル付き回答を収集しました。私たちは 3 つの研究課題を中心に研究を構成しています。 (1) PSRS はどのくらいの頻度で発生しますか? (2) どの程度検出できるか? (3) 検出はどのようにして未知のモデルに一般化されるのでしょうか?まず、PSRS 率が LLM 全体で 5% ~ 56% の範囲であり、より有能なモデルほどおべっかが少ないことを明らかにしました。次に、PSRS の検出は応答テキストのみから実行可能であり、検出器はトレーニング データから微妙な PSRS パターンを学習する必要があることを示します。新しい LLM は急速に出現するため、検出器は必然的に未知のモデルに遭遇し、モデル間の一般化がフレームワークの重要な目標になります。私たちは、目に見えないモデルでは検出パフォーマンスが低下することを実証し、この課題に対処するための最初のアプローチを提案します。今後の研究をサポートするために、データセットとコードをリリースします。
原文 (English)
Measuring and Detecting Harmful AI Sycophancy
Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful. This paper focuses on one harmful sycophancy: preference-induced stance reversal sycophancy (PSRS), where a model reverses an initial stance merely to align with a user's stated preference. While existing research mainly measures how sycophantic a model is, we go further and ask whether PSRS can also be detected automatically from a single response. To investigate this at scale, we introduce CAP (Contrastive Anchor Probing), a framework for collecting labeled PSRS data. Applying CAP to 17 open- and closed-source LLMs, we collect 290,460 labeled responses across 12 everyday-advice domains. We organize our study around three research questions. (1) How often does PSRS occur? (2) How well can it be detected? (3) How does detection generalize to unseen models? We first reveal that PSRS rates range from 5% to 56% across LLMs, with more capable models being less sycophantic. Next, we show that detecting PSRS is feasible from the response text alone, and detectors need to learn subtle PSRS patterns from the training data. Because new LLMs appear rapidly, detectors inevitably encounter unseen models, making cross-model generalization an important framework goal. We demonstrate that detection performance drops on unseen models and propose an initial approach to address this challenge. We will release our dataset and code to support future research.
SkillHEX: 仮説に基づく自律的な探索と活用によるエージェント スキルの向上
エージェントのスキルにより、LLM は再利用可能な手順に関する知識を身に付けることができますが、手動によるメンテナンスには高コスト、非拡張性、調整不良という問題があります。したがって、現実世界のデプロイメントでは、限られたインタラクション予算とトレーニングまたは検証セットの不足によって制約を受け、テスト時に自律的なオンデマンドのスキル進化が必要になります。この設定では、結果が複数の潜在的な失敗原因と混同される、深刻な希薄な報酬の課題が導入されます。このようなあいまいさの下では、単一の既存のスキルを貪欲に洗練させる既存の方法は特に悪用の罠に脆弱であり、初期の誤診によって非生産的な軌道に沿って限られた試験を使い果たすことになります。これに対処するために、仮説に基づく自己検証と証拠に基づくツリー検索を組み合わせた閉ループ フレームワークである SkillHEX を導入します。 SkillHEX は、反証可能な障害仮説を実行可能なテストに変換し、追加の環境試行を行わずに濃厚な報酬として診断証拠を生成します。この証拠は、永続的なスキル改訂ブランチの検索をガイドし、サポートされている編集の活用と、妥当な代替案の探索との動的バランスをとります。 SkillsBench の 87 のタスクで評価された SkillHEX は、既存の自己進化メソッドを上回り、5 回の反復予算で GPT-5.3-Codex と Claude Opus 4.7 を使用した場合、それぞれ 55.9% と 57.9% の平均合格率を達成しました。
原文 (English)
SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation
Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment. Real-world deployments thus require autonomous, on-demand skill evolution at test time, constrained by limited interaction budgets and a lack of training or validation sets. This setting introduces a severe sparse reward challenge, where outcomes conflate multiple latent failure causes. Under such ambiguity, existing methods that greedily refine a single incumbent skill are particularly vulnerable to an exploitation trap, allowing early misdiagnoses to exhaust limited trials along unproductive trajectories. To address this, we introduce SkillHEX, a closed-loop framework coupling hypothesis-driven self-verification with evidence-guided tree search. SkillHEX translates falsifiable failure hypotheses into executable tests, producing diagnostic evidence as dense reward without additional environment attempts. This evidence guides a search over persistent skill-revision branches, dynamically balancing the exploitation of supported edits with the exploration of plausible alternatives. Evaluated on 87 tasks from SkillsBench, SkillHEX outperforms existing self-evolving methods and achieves an average pass rate of 55.9% and 57.9% using GPT-5.3-Codex and Claude Opus 4.7 under a five-iteration budget, respectively.
Bayesian Expected Uncertainty Reduction (B-EUR) Model: A Computational Account of What Makes Design Options Worth Trying
This paper proposes the Bayesian Expected Uncertainty Reduction (B-EUR) model, which formalizes the value of trying a candidate design acti…
リサンプリングよりも洗練する: LLM 推論のテスト時の自己修正
テスト時のスケーリングは、追加の推論計算を使用することで LLM 推論を改善しますが、より広範なサンプリングだけでは利益が減少する可能性があります。新しいロールアウトでは、有用な推論の多様性を追加する代わりに、既存の回答パターンを繰り返すことがよくあります。検証者ベースの選択は代替手段を提供しますが、そのパフォーマンスは外部報酬モデルの調整に依存します。私たちは、テスト時のコンピューティングを使用して候補ソリューションの探索と改善の両方を行う、検証者を必要としない幅と深さの改良フレームワークを提案します。この方法では、複数の独立した推論ロールアウトをサンプリングし、反復的な自己批判と自己修正を通じて各ロールアウトを改良し、多数決によって洗練された回答を集計します。幅はさまざまな初期試行を保存しますが、深さは集約の前に局所的な推論エラーを修復します。 AIME24、AIME25、AMC、OlympiadBench、および MATH500 にわたって、私たちの手法は、貪欲なデコード、多数決、検証者ベースのベストオブ $N$、ビーム検索、および複数のオープンウェイト モデルにわたる先読みデコードを一貫して改善しています。たとえば、Qwen2.5-1.5B では、精度は最強のベリファイアベースのベースラインから MATH500 では $58.0\%$ まで、AMC では $25.0\%$ から $32.5\%$ まで増加します。これらの結果は、テスト時の計算は、より多くの候補をサンプリングしたり、検証者による選択に依存したりするためだけでなく、サンプリングされた軌跡を調整するために使用するとより効果的であることを示しています。
原文 (English)
Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning
Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing answer patterns instead of adding useful reasoning diversity. Verifier-based selection offers an alternative, but its performance depends on the calibration of an external reward model. We propose a verifier-free breadth--depth refinement framework that uses test-time compute to both explore and improve candidate solutions. The method samples multiple independent reasoning rollouts, refines each rollout through iterative self-critique and self-correction, and aggregates the refined answers by majority voting. Breadth preserves diverse initial attempts, while depth repairs local reasoning errors before aggregation. Across AIME24, AIME25, AMC, OlympiadBench, and MATH500, our method consistently improves over greedy decoding, majority voting, verifier-based best-of-$N$, beam search, and lookahead decoding across multiple open-weight models. For instance, with Qwen2.5-1.5B, accuracy increases from the strongest verifier-based baseline to $58.0\%$ on MATH500, and from $25.0\%$ to $32.5\%$ on AMC. These results show that test-time compute can be more effective when used to refine sampled trajectories rather than only to sample more candidates or rely on verifier-guided selection.
A Unified Framework for Trajectory Prediction with Explicit Planning and Reaction Decomposition
Trajectory prediction has shifted toward structured formulations with explicit social modeling. However, existing methods inadequately dist…
Volve フィールドでの接地された良好な状態の異常検出: 構築されたラベル、ベースライン、およびデュアルヘッド モデル
マシンの状態を監視するための公開ベンチマークのほとんどは、意図的に障害が発生し、すべてのイベントが既知であるテスト装置から取得されます。実際の生産現場ではそのようなことはめったにありません。これらは、障害ログが添付されていないセンサー履歴を提供します。これはまさに、異常検出方法が独自のラベルを作成する必要があり、静かな仮定が気づかれずに紛れ込む可能性がある状況です。私たちは Equinor がリリースしたオープンな Volve フィールド データを使用しており、そのようなデータセットでは通常省略されている 2 つのことを真剣に受け止めています。まず、単なる数値のパターンではなく、物理的に問題が発生する可能性があると現場独自のエンジニアリング文書に記載されているものと照合して異常ラベルを作成し、すべてのラベルの背後にある理由を公開します。次に、教師なしベースラインと、イベントがいつ発生するか、その種類をマークする小さなデュアルヘッド モデルの両方を使用して、構築されたラベルが学習可能かどうかをテストします。このアイデアは、金属部品の欠陥検出に関する以前の研究から引き継いでいます。結果は正直です。ラベルを決して認識しない教師なし検出器は、ルールでフラグが立てられたのと同じ領域に依然として検出され、ラベルが恣意的ではないことがわかります。コンパクトな教師ありモデルは、これまでに見たことのないウェル全体でイベントの存在とイベント タイプを十分に回復し、時間内のイベントを大まかにのみ特定します。何がうまくいったのか、何がうまくいかなかったのか、そしてその間のすべての仮定を報告します。データセット、根拠のあるラベル、ラベルごとの来歴、ベースライン スコア、トレーニングされたモデル、コードは、CC-BY-NC-SA 4.0 に基づいて公開されています。
原文 (English)
Grounded Well-Condition Anomaly Detection on the Volve Field: Constructed Labels, a Baseline, and a Dual-Head Model
Most public benchmarks for machine-condition monitoring come from test rigs, where faults are induced on purpose and every event is known. Real production fields rarely offer that. They give you sensor histories with no fault log attached, which is exactly the situation where an anomaly-detection method has to invent its own labels, and where quiet assumptions can slip in unnoticed. We work with the open Volve field data released by Equinor and take two things seriously that such datasets usually skip. First, we build anomaly labels that are not just patterns in the numbers but are checked against what the field's own engineering documents say can physically go wrong, and we release the reasoning behind every label. Second, we test whether those constructed labels are learnable at all, using both an unsupervised baseline and a small dual-head model that marks when an event happens and what kind it is, an idea we carry over from earlier work on defect detection in metal parts. The results are honest. An unsupervised detector that never sees the labels still lands on the same regions our rules flagged, which tells us the labels are not arbitrary. A compact supervised model recovers event presence and event type well across wells it has never seen, and locates events in time only roughly. We report what worked, what did not, and every assumption in between. The dataset, grounded labels, per-label provenance, baseline scores, trained model, and code are released publicly under CC-BY-NC-SA 4.0.
DreamGuard: リスクを意識した世界モデルによる LLM エージェント向けの効率的なランタイム ガードレール
大規模言語モデル (LLM) エージェントが外部ツールを呼び出し、現実世界のシステムと対話することが増えているため、安全でないアクションによって外部状態、ユーザー データ、およびダウンストリーム サービスに取り消し不能な結果が生じる可能性があります。最近のランタイム ガードレールは、提案されたアクションを実行前にチェックすることでそのようなリスクを軽減しますが、多くは事後対応的なままです。主に現在のアクションの見かけの安全性を評価し、リスクが軌道全体でどのように展開するかについての明示的なモデルが不足しています。この制限により、長期的なリスクに対して重大な盲点が生じ、個別には良性のように見えるアクションがエージェントを徐々に危険な状態に導く可能性があります。これに応じて、リスクを認識した世界モデルに基づいて構築された LLM エージェント向けのプロアクティブなガードレールである DreamGuard を提案します。ワールド モデルは、軌跡全体にわたってコンパクトな再発潜在状態を維持し、DreamGuard が差し迫った危険とプレフィックス リスクの証拠を導き出す将来の潜在状態を予測します。次に、実行前にこれらのマルチホライズン信号を介入決定に融合します。 4 つのベンチマークにわたる実験とオンライン ガードレール評価では、DreamGuard が汎用、リアクティブ、プロアクティブ ガードレール ベースラインを上回っており、評価されたガードレール間で安全性とユーティリティの最適なトレードオフを達成し、通話あたりの平均エンドツーエンド レイテンシ 25 ミリ秒を維持していることが示されています。
原文 (English)
DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model
As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible consequences on external states, user data, and downstream services. Recent runtime guardrails mitigate such risks by checking proposed actions before execution, but many remain reactive: they primarily assess the apparent safety of the current action, lacking an explicit model of how risk evolves across the trajectory. This limitation creates a critical blind spot for long-horizon risks, where individually benign-looking actions can gradually drift the agent toward hazardous states. In response, we propose DreamGuard, a proactive guardrail for LLM agents built around a risk-aware world model. The world model maintains a compact recurrent latent state over the trajectory and predicts future latent states from which DreamGuard derives immediate-hazard and prefix-risk evidence. It then fuses these multi-horizon signals into intervention decisions before execution. Experiments across four benchmarks and an online guardrail evaluation show that DreamGuard outperforms generic, reactive, and proactive guardrail baselines, achieves the best safety-utility trade-off among evaluated guardrails, and maintains an average end-to-end latency of 25 ms per call.
人間と AI の相互作用を形成して改善の道筋を提供し、競合する目標のバランスをとる
AI システムが導入されると、AI システムを使用する個人、または AI システムによって評価される個人は、システムがどのように動作するかについて信念を形成し、その信念を使用して自分の好み、行動、または属性を戦略的に提示します。その後、システムはフィードバックまたは決定結果で応答し、それによって人間と AI の対話ループが作成されます。この論文では、次の 3 つの目標を達成するために、そのようなインタラクションを設計および形成する方法を研究します。(1) 個人が AI システムについて正確な信念を培えるよう支援し、最小限のコストで好ましい結果を改善または確保できるようにすること、(2) 改善を奨励する、またはゲーム行動を抑制すること、(3) AI システムが精度の最大化などの意図された目的を確実に達成し続けることを保証すること。これらの目標に取り組むために、論文は、評価される個人と AI システムの両方の観点から人間と AI の相互作用を調査および研究する 3 つの補完的な部分で構成されています。この論文で提示された研究は、人間のニーズ、価値観、能力に合わせた AI システムを設計するための原則と方法を提供することにより、人間中心の機械学習を前進させます。方法論的には、この論文は理論分析、データ駆動型モデリング、人間を対象とした実験、現実世界および半合成データセットに対する実証的評価を統合しています。
原文 (English)
Shaping Human-AI Interactions to Provide Improvement Pathways and Balance Competing Objectives
When an AI system is deployed, the individuals who use and or are evaluated by it form beliefs about how the system operates and use those beliefs to strategically present their preferences, behaviors, or attributes. The system then responds with feedback or a decision outcome, thereby creating a human-AI interaction loop. This thesis studies how to design and shape such interactions to achieve three goals: (1) help individuals develop accurate beliefs about the AI systems so they can improve and or secure favorable outcomes at minimal cost, (2) encourage improvement and or discourage gaming behaviors, and (3) ensure that the AI system continues to achieve its intended objectives, such as maximizing accuracy. To address these goals, the thesis is organized into three complementary parts that examine and study human-AI interactions from the perspectives of both evaluated individuals and AI systems. Together, the work presented in this thesis advances human-centered machine learning by providing principles and methods for designing AI systems that align with human needs, values, and capabilities. Methodologically, this thesis integrates theoretical analysis, data-driven modeling, human-subject experiments, and empirical evaluations on real-world and semi-synthetic datasets.
RA-CAD: ステートアウェアなテキストから CAD への生成のための実行後の批判の学習
テキストから CAD への生成は、自然言語の設計意図を編集および実行可能なパラメトリック コンピューター支援設計 (CAD) コードに変換し、手動モデリングに必要な専門知識と労力を軽減します。既存の方法には、生成プロセスを最適化するために、固定された、外部から提供される、プロンプト誘導される、または個別に最適化された批評メカニズムが組み込まれていますが、生成プロセス全体を通じてフィードバックがどのように解釈され、効果的な是正措置に変換されるかを必ずしも最適化しているわけではありません。このフィードバック利用のギャップを埋めるために、生成、実行、批判、書き換えのループを通じて CAD 環境と対話する状態認識エージェントである RA-CAD (ReAct Agent for CAD) を紹介します。各反復で、RA-CAD は現在のコードを実行し、その結果を観察します。設計指示、現在のコード、および実行フィードバックを条件として、エージェントは中間ポリシー アクションとして明示的な実行後の批評を生成します。この批評は、現在の結果の終了を検証するか、次の書き換えの条件となるリビジョン指向のガイダンスを提供します。 CAD コード ブートストラップ (CCB) は、まず、監視付き微調整を通じて基本的なパラメトリック CAD コーディング機能を確立します。その後、フィードバック駆動エージェント最適化 (FAO) は、軌道レベルのグループ相対ポリシー最適化をポリシー生成コードと批評シーケンスの両方に適用し、終端 F1 と面取り距離の報酬を完全なインタラクション軌道に割り当てます。この定式化により、批評は最適化されていない補助的な出力ではなく、結果に合わせた学習可能な政策決定となります。 CADFusion と Text2CAD の実験では、既存の方法や強力な独自言語モデルと比較して、RA-CAD が最先端の実行妥当性と幾何学的品質を達成していることが示され、提案されている状態認識型テキストから CAD エージェントの有効性が実証されています。
原文 (English)
RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation
Text-to-CAD generation translates natural-language design intent into editable and executable parametric computer-aided design (CAD) codes, reducing the expertise and effort required for manual modeling. Existing methods incorporate fixed, externally supplied, prompt-induced, or separately optimized critique mechanisms to optimize the generation process, but they do not necessarily optimize how feedback is interpreted and translated into effective corrective actions throughout the generation process. To bridge this feedback-utilization gap, we present RA-CAD (ReAct Agent for CAD), a state-aware agent that interacts with the CAD environment through a Generate--Execute--Critique--Rewrite loop. At each iteration, RA-CAD executes the current code and observes its outcome. Conditioned on the design instruction, current code, and execution feedback, the agent then generates an explicit post-execution critique as an intermediate policy action. This critique either validates the current result for termination or provides revision-oriented guidance that conditions the next rewrite. CAD Code Bootstrapping (CCB) first establishes fundamental parametric CAD coding capabilities through supervised fine-tuning. Feedback-Driven Agent Optimization (FAO) subsequently applies trajectory-level Group Relative Policy Optimization to both policy-generated code and critique sequences, assigning terminal F1 and Chamfer Distance rewards to the complete interaction trajectory. This formulation makes critique an outcome-aligned, learnable policy decision rather than an unoptimized auxiliary output. Experiments on CADFusion and Text2CAD show that RA-CAD achieves state-of-the-art execution validity and geometric quality compared with existing methods and strong proprietary language models, demonstrating the effectiveness of the proposed state-aware text-to-CAD agent.
BlockPython: ブロックベースから Python プログラミングへの移行のための、プロセス認識エージェントがサポートするプラットフォーム
ブロックベースのプログラミングからテキストベースのプログラミングに移行するには、学習者は目に見えるプログラム構造を抽象的なテキスト表現に変換する必要があります。これにより、計算概念の理解と Python 構文での表現の間に認知的なギャップが生じる可能性があります。この移行をサポートするために、BlockPython を設計および実装しました。このプラットフォームは、ブロックと Python 間の双方向変換に重点を置き、学習者をタスク分解、ブロックベースの練習、コード チャレンジ、拡張インタラクションの 4 つの段階を通してガイドします。これらの段階を通じて、学習者はプログラム構造、実行時の動作、およびテキストコード間のつながりを徐々に確立していきます。学習中、プラットフォームはブロック アーティファクト、コード バージョン、実行結果、サポートの使用、対話などのプロセス証拠を継続的に収集します。決定論的診断、プログラムの視覚化、および学習アシスタントは、この証拠を使用して、計算の理解と Python の表現におけるさまざまな困難を特定します。ルールベースのシステムはプログラムの実行、客観的評価、段階制御を担当し、学習アシスタントは検証済みの証拠を使用して説明、プロンプト、およびガイドとなる質問を提供します。このレポートでは、BlockPython の設計理論的根拠、学習ワークフロー、プロセス認識サポート メカニズムについて説明し、ブロックベースのプログラミングからテキストベースのプログラミングへの移行をサポートし、学習プロセスを分析するためのシステム設計のリファレンスを提供します。
原文 (English)
BlockPython: A Process-Aware Agent-Supported Platform for the Transition from Block-Based to Python Programming
The transition from block-based to text-based programming requires learners to convert visible program structures into abstract textual expressions, which may create a cognitive gap between understanding computational concepts and expressing them in Python syntax. To support this transition, we designed and implemented BlockPython. The platform centers on bidirectional translation between blocks and Python and guides learners through four stages: Task Decomposition, Block-Based Practice, Code Challenge, and Extended Interaction. Across these stages, learners progressively establish connections among program structure, runtime behavior, and textual code. During learning, the platform continuously collects process evidence, including block artifacts, code versions, run outcomes, use of support, and dialogue. Deterministic diagnosis, program visualization, and the learning assistant use this evidence to identify different difficulties in computational understanding and Python expression. The rule-based system is responsible for program execution, objective evaluation, and stage control, while the learning assistant uses verified evidence to provide explanations, prompts, and guiding questions. This report describes the design rationale, learning workflow, and process-aware support mechanisms of BlockPython and provides a system-design reference for supporting the transition from block-based to text-based programming and for analyzing learning processes.
統合エージェント: デバイス間の対話の管理
機能が急速に増加するにつれて、AI エージェントは 1 つのアプリ内で実行されることから、時間の経過とともにユーザーのデバイス全体で動作するようになります。しかし、既存のエージェント システムは、このシナリオではまだ不十分です。これは、観測がデバイスや瞬間に分散しているためですが、主流のシステムはこの事実を考慮して設計されていないためです。デバイスをツールとして扱う単一エージェントには、時間の経過に伴うすべてのデバイスの効果的な状態管理が欠けており、マルチエージェント システムはエージェント間で調整しますが、デバイス間、時間間リクエストに必要なコンパクトな状態を維持できません。私たちは、エージェントは、現在の観察を考慮してアクションを決定するために、関与の証拠、述べられた事実、および継続的な要求をコンパクトでアクションの準備ができた形式に整理する、効果的に設計された状態を維持する必要があると主張します。状態設計を比較するために、デバイスと時間にわたるユーザーとエージェントの対話のベンチマークを構築します。私たちはこの原則を、デバイスや瞬間を超えて対話の証拠を運び、それを現在の観察と組み合わせて使用して動作するステートフル エージェントである Unified Agent でインスタンス化します。デフォルト設定では、公開されている 4 つの設計を適応させた結果を大幅に上回ります。マルチモーダル大規模言語モデル (MLLM) ファミリ、機能、および推論の取り組みが変更されても、比較対象のすべてのシステムよりも優れたままであり、状態設計の利点が MLLM 設定全体にわたって堅牢であることを示しています。私たちのコードとデータは GitHub で公開されます。
原文 (English)
Unified Agent: Managing Interactions across Devices
As capabilities rapidly increase, AI agents can move from running inside one app to acting across a user's devices over time. Yet existing agent systems still fall short in this scenario. This is because observations are scattered across devices and moments, but mainstream systems are not designed around this fact: a single agent that treats devices as tools lacks effective state management for all devices across time, and multi-agent systems coordinate across agents but do not maintain the compact carried state a cross-device, cross-time request needs. We argue that the agent should maintain an effectively designed state that organizes engagement evidence, stated facts, and the standing request in a compact, action-ready form for deciding its action given the current observation. To compare state designs, we construct a benchmark of user-agent interaction across devices and time. We instantiate this principle in Unified Agent, a stateful agent that carries interaction evidence across devices and moments and uses it with the current observation to act. In the default setting, it significantly outperforms our adaptations of four published designs. Across changes in multimodal large language model (MLLM) family, capability, and reasoning effort, it remains ahead of all compared systems, demonstrating that the state-design advantage is robust across MLLM settings. Our code and data will be publicly available on GitHub.
サブリミナル学習は非意味的蒸留である
サブリミナル学習 (SL) は、現代の言語モデルによって示される驚くべきタイプの一般化です。これにより、教師からの一見無関係またはランダムな合成データを抽出することによって、教師モデルから生徒にバイアスや行動を伝達することができます。このため、入力データの標準的な監査では隠れたサブリミナル信号をキャッチできないため、AI システムが予測可能であり、安全にトレーニングされることを保証する上で課題が生じています。ここでは、SL を実現するメカニズムと推進力に関するいくつかの未解決の質問を調査します。 1 つ目は、バイアスがデータにエンコードされるプロセスの性質です。教師モデルと生徒モデルの重みにガウス ノイズを追加すると、サブリミナル転送の大きさが Gemma では 1.9 倍、Llama では 1.3 倍増加することがわかり、非セマンティックな重み構造が重要な役割を果たしていることを示唆しています。以前の研究で使用されていたプロンプトと微調整に加えて、ステアリングベクトルを教師に適用してサブリミナルデータを生成できることを示します。ステアリングされたデータとプロンプトされたデータに基づいてトレーニングされた生徒モデルの活性化の分析は、生徒が教師のバイアスの意味論的な意味だけでなく、それを適用するために使用された介入の種類も継承していることを示しています。つまり、ステアリングされた生徒はステアリングベクトルを模倣しますが、プロンプトされた生徒は真似しません。さらに、ステアリングされたサブリミナル データの勾配は、教師のステアリング ベクトルと線形相関を示し、データ監査の可能性を示しています。より広く言えば、合成データがフロンティア トレーニング パイプラインの中心となるにつれて、トレーニング データに隠された潜在的な信号を確認できることが最も重要になります。
原文 (English)
Subliminal Learning is Non-Semantic Distillation
Subliminal Learning (SL) is a surprising type of generalization displayed by modern language models. It allows the transfer of a bias or behavior from a teacher model to a student by distilling from seemingly unrelated or random synthetic data from the teacher. This presents challenges in ensuring AI systems remain predictable and are trained safely, as standard auditing of the input data would not catch the hidden subliminal signal. Here, we investigate several open questions as to the enabling mechanisms and drivers of SL. First is the nature of the process by which biases are encoded in the data. We find that by adding Gaussian noise to the weights of the teacher and student models, the magnitude of subliminal transfer is increased by a factor of 1.9 in Gemma and 1.3 in Llama, suggesting that non-semantic weight structures play a crucial role. We show that steering vectors can be applied to the teacher to produce subliminal data, in addition to prompting and finetuning as used in previous studies. Analysis of the activations of the student models that have been trained on steered and prompted data demonstrates that students inherit not just the semantic meaning of the teacher's bias, but also the type of intervention that was used to apply it: steered students imitate steering vectors, prompted students do not. Additionally, the gradients of steered subliminal data show a linear correlation with the teacher's steering vectors, showing promise for data auditing. More broadly, as synthetic data becomes central to frontier training pipelines, being able to see the latent signals hidden in training data becomes paramount.
プロンプト側エージェントのプレイブックはいつ転送されますか?エージェント導入の精度、コスト、実行時間の変化
プロンプト側のプレイブックは、再トレーニングすることなくツールを使用する言語エージェントを改善できますが、ソース設定を超えた移植性は不明です。私たちは、共有された蒸留-検証-転送プロトコルの下で凍結されたプレイブック転送を研究します。 ALFWorld では、制御された貪欲なデコードの下で転送は有益であり、ほぼ予算に見合った 1 つの比較では、抽出されたガイダンスが 5 つの固定デモンストレーションを上回っています。 TAU2-Bench では、事前に指定された集約コントラストにより、中程度の平均一致ドメインの利点がサポートされますが、グローバル ホルム補正は 135 のルート レベルの効果のうち 1 つだけを保持します。残りのグリッドは、互換性に依存する異質性の記述的な証拠を提供します。 XBench-DeepSearch では、1 つのアーティファクトであるランタイム ペアリングにより、有用な初回試行ヒューリスティックが維持されますが、クエリの繰り返し、停止の遅延、およびコンテキストとランタイムのシフト後の大幅なコストの増加が発生します。ベンチマーク全体で、転送されたプレイブックとターゲット派生のプレイブックは両方とも、ターゲット側で成功、終了、プロトコルの互換性、コストを検証する必要があります。したがって、凍結移送は条件付きのコールドスタートオプションであり、デフォルトで再利用する戦略や、ターゲット側の再蒸留に代わる普遍的に好ましい代替策ではありません。
原文 (English)
When Do Prompt-Side Agent Playbooks Transfer? Accuracy, Cost, and Runtime Shift in Agent Deployment
Prompt-side playbooks can improve tool-using language agents without retraining, but their portability beyond the source setting is unclear. We study frozen playbook transfer under a shared distill--validate--transfer protocol. On ALFWorld, transfer is beneficial under controlled greedy decoding and, in one near-budget-matched comparison, distilled guidance outperforms five fixed demonstrations. On TAU2-Bench, a prespecified aggregate contrast supports a modest average matched-domain advantage, but global Holm correction retains only one of 135 route-level effects; the remaining grid provides descriptive evidence of compatibility-sensitive heterogeneity. On XBench-DeepSearch, one artifact--runtime pairing preserves useful first-try heuristics while producing repeated queries, delayed stopping, and substantial cost inflation after a context-runtime shift. Across benchmarks, transferred and target-derived playbooks both require target-side validation of success, termination, protocol compatibility, and cost. Frozen transfer is therefore a conditional cold-start option, not a reuse-by-default strategy or a universally preferable alternative to target-side redistillation.
アクティビティ フレーム: エージェントのメモリと再生のための決定論的な画面アクティビティのコンパイル
コンピュータを使用するエージェントは、ユーザーがすでに実行したルーチンを再導出するために完全なフロンティア推論を支払います。これは、今日のエージェントの記憶には、ユーザーが何をしたかではなく、ユーザーが何を言ったかが記録されるためです。受動的にキャプチャされた画面アクティビティを、決定論的なゼロモデル パイプラインを使用してエージェント メモリにコンパイルします。これは、ローカル キャプチャ ストリームを型指定されたアクティビティ フレーム、アプリケーション、サイト、タイミング、入力ボリューム、生の行に戻る証拠ポインタを運ぶ境界付きエピソードにセグメント化し、ループ内にモデルがないため、出力はバイト同一で、キャッシュ可能で、機械的に監査可能です。ある専門家のアクティブな 51 日間にわたる 128,756 フレームのシングルユーザー コーパスでは、コンパイラは 1 日の生のキャプチャを 68 ミリ秒で 86 倍小さいプロンプト対応のコンテキスト ブロックに縮小し、そのブロックを読み取るエージェントは、独立したオラクルに対して 98.4% (ウィルソン 95% CI 91.7 ~ 99.7%) の精度でその日に関する質問に回答します (一方、独立系オラクルの場合は 66 ~ 80%)。同じキャプチャの LLM 概要。フロンティア モデルと一致するブロックを読み取る中間層モデル。同じコンパイラがデマンド側のコスト手段としても機能します。エージェントのロールアウトではなく、受動的で委任前の人間のアクティビティを読み取り、エージェントコストモデルが想定しているが、私たちの知る限りでは測定されていない 2 つのパラメーター、ルーチンオーバーヘッド率 R とルーチン繰り返し h を提供します。モデル化された上限である R の最初の値は 60 ~ 343 倍、委任可能な反復はサンプル内 9.0%、サンプル外 7.7% で、現実的な全フリート トークンの上限は 8% 付近であることを報告します。コンパイルされたルーチンは、ループ外のモデルで決定的に再生され、ガード一致ヒットでモデル トークンがゼロでライブで実証されます。スキーマ、コンパイラ、評価ハーネスはオープンです。
原文 (English)
Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay
Computer-use agents pay full frontier inference to re-derive routines their user has already performed, because an agent's memory today records what the user said, not what the user did. We compile passively captured screen activity into agent memory with a deterministic, zero-model pipeline: it segments a local capture stream into typed activity frames, bounded episodes carrying application, site, timing, input volume, and evidence pointers back to the raw rows, with no model in the loop, so the output is byte-identical, cacheable, and mechanically auditable. On one professional's single-user corpus of 128,756 frames over 51 active days, the compiler reduces a day of raw capture to a prompt-ready context block 86x smaller in 68 ms, and an agent reading that block answers questions about the day at 98.4% accuracy (Wilson 95% CI 91.7-99.7%) against an independent oracle, versus 66-80% for an LLM summary of the same capture, a mid-tier model reading the block matching a frontier one. The same compiler doubles as a demand-side cost instrument. Read off passive, pre-delegation human activity rather than agent rollouts, it supplies two parameters that agent-cost models assume but, to our knowledge, have not measured: the Routine Overhead Ratio R and the routine recurrence h. We report first values of R, a modeled upper bound, at 60-343x, and a delegable recurrence of 9.0% in-sample and 7.7% out-of-sample, for a realistic all-fleet token ceiling near 8%; a compiled routine replays deterministically with the model out of the loop, demonstrated live at zero model tokens on a guard-matched hit. Schema, compiler, and evaluation harness are open.
ChainClaw: 信頼性の高いオンチェーン実行のための階層化されたエージェント フレームワーク
汎用の大規模言語モデル エージェントは、ツールで拡張されたタスクで優れたパフォーマンスを達成していますが、ブロックチェーン環境では崩れる前提に依存しています。オンチェーン実行はステートフルで、敵対的で、経済的に不可逆的であり、反応性、不可逆性、可観測性という 3 つの基本的なギャップが露呈します。私たちは、OpenClaw 上に構築されたブロックチェーン ネイティブ エージェント フレームワークである ChainClaw を提案します。これは、クロスレイヤー メモリ サブシステムによって統合された、イベント駆動型オーケストレーション レイヤー、シミュレーション ベースの安全性インテリジェンス レイヤー、およびオンチェーン モニタリング ランタイム レイヤーで構成される階層化アーキテクチャを通じて、3 つのギャップすべてに対処します。 ChainClaw は、イベントの取り込みとシミュレーションのフィードバックを介して反応性のギャップを、トランザクション シミュレーションとアクション ガードを備えた実行前安全パイプラインを介して不可逆性のギャップを、オンチェーン読み取りアダプターとトランザクション モニターを介して可観測性のギャップを埋めます。私たちは、4 つのカテゴリと 5 つの次元にわたる 7 つのタスクをカバーする専用のベンチマークで ChainClaw を評価します。 ChainClaw は、安全性とタスク完了の両方において、代表的なベースラインを常に上回っています。
原文 (English)
ChainClaw: A Layered Agent Framework for Reliable On-Chain Execution
General-purpose large language model agents have achieved strong performance on tool-augmented tasks, yet they rely on assumptions break down in blockchain environments. On-chain execution is stateful, adversarial, and economically irreversible, exposing three fundamental gaps: Reactivity, Irreversibility, and Observability. We propose ChainClaw, a blockchain-native agent framework built on OpenClaw, that addresses all three gaps through a layered architecture comprising an event-driven orchestration layer, a simulation-based safety intelligence layer, and an on-chain monitoring runtime layer, unified by a cross-layer memory subsystem. ChainClaw closes the Reactivity gap via event ingestion and simulation feedback, the Irreversibility gap via a pre-execution safety pipeline with transaction simulation and action guard, and the Observability gap via an on-chain read adapter and transaction monitor. We evaluate ChainClaw on a purpose-built benchmark covering seven tasks across four categories and five dimensions. ChainClaw consistently outperforms representative baselines on both safety and task completion.
Agentic AI が統合センシングと通信と出会うとき
エージェント型人工知能 (AI) は、統合センシングおよび通信 (ISAC) を、機能指向の物理層テクノロジーから、目標主導型の閉ループ インテリジェント システム (AISAC と呼ぶパラダイム) に変換しています。学習ベースのセンシング、リソース割り当て、再構成可能なインテリジェント サーフェス (RIS)、エッジ インテリジェンス、マルチエージェント調整、および復元力のあるネットワーキングに関する既存の研究は、主に個別に開発されてきました。この調査は、観察、文脈化、推論と予測、計画とオーケストレーション、実行とコラボレーション、フィードバックと回復力からなる 6 段階の閉ループのフレームワーク内で文献を統合します。また、物理層のプリミティブから完全な閉ループ エージェントの ISAC まで、エージェントの成熟度の 5 つのレベルも導入されています。私たちはこのフレームワークを使用して、マルチモーダル インテリジェンス、大規模言語モデル、強化学習、フェデレーテッド ラーニング、RIS 支援制御、無人航空機 (UAV) および車両ネットワーク、AI ネイティブ ネットワーク管理の進歩をレビューし、完全な知覚-推論-行動ループの横断的な要件としてプライバシー、セキュリティ、回復力、持続可能性を分析します。 9つの薬剤特有の評価基準に照らして代表的な研究を監査したところ、どのシステムもそのうちの1つまたは2つ以上を報告していないことが示されており、主張されている薬剤の成熟度と証明されている薬剤の成熟度の間にギャップがあることが明らかになりました。私たちは、物理的からセマンティックなグラウンディング、予測世界モデル、リアルタイムのエージェントと PHY の相互作用、安全なツールの使用、異種マルチエージェントのコラボレーション、ベンチマーク、およびリソース効率の高い自律性における未解決の課題を特定します。
原文 (English)
When Agentic AI Meets Integrated Sensing and Communication
Agentic artificial intelligence (AI) is transforming Integrated Sensing and Communication (ISAC) from a function-oriented physical-layer technology into a goal-driven, closed-loop intelligent system, a paradigm we term AISAC. Existing work on learning-based sensing, resource allocation, reconfigurable intelligent surfaces (RIS), edge intelligence, multi-agent coordination, and resilient networking has developed largely in isolation. This survey unifies the literature within a six-stage closed-loop framework comprising observation, contextualization, reasoning and prediction, planning and orchestration, execution and collaboration, and feedback and resilience. It also introduces five levels of agentic maturity, ranging from physical-layer primitives to fully closed-loop agentic ISAC. We use this framework to review advances in multimodal intelligence, large language models, reinforcement learning, federated learning, RIS-assisted control, Unmanned Aerial Vehicle (UAV) and vehicular networks, and AI-native network management, and analyze privacy, security, resilience, and sustainability as cross-cutting requirements of the full perception-reasoning-action loop. An audit of representative studies against nine agentic-specific evaluation criteria shows that no system reports more than one or two of them, exposing a gap between claimed and demonstrated agentic maturity. We identify open challenges in physical-to-semantic grounding, predictive world models, real-time agent-PHY interaction, safe tool use, heterogeneous multi-agent collaboration, benchmarking, and resource-efficient autonomy.
When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents
Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not…
言語モデルのパーソナライゼーションのための慎重なコンテキスト ステアリング
言語モデル (LM) を個々のユーザーの好みに合わせてパーソナライズすることは、さまざまな目標や背景に応じた応答を調整するために不可欠です。既存のメソッドは通常、ユーザーごとに個別のアダプターをトレーニングするか、スコアがユーザーに依存する報酬モデルを学習します。ユーザーごとに明示的に最適化しているにもかかわらず、これらの方法は限られた観察から学習する必要があるため、データがまばらで、目に見えないユーザーやドメインへの一般化が不十分であるという問題があります。代わりに、インコンテキスト学習 (ICL) とコンテキスト ステアリング (CoS) は、ベース LM をユーザー コンテキストに直接条件付けし、ユーザーごとのトレーニングを行わずに事前トレーニングされた機能を活用することで、より効果的なパーソナライゼーションを提供できます。ただし、どちらもデコード ステップ全体でそのコンテキストの影響を適応させません。ICL はコンテキストを制御されないままにしますが、CoS は固定ステアリング係数を適用し、ステップごとに 2 つの LM 順方向パスを必要とします。私たちは Cautious Context Steering (CCS) を提案します。これは、凍結されたバックボーン LM に軽量のアダプターを追加して、ユーザー コンテキストが生成に影響を与えるかどうか、またどの程度強く影響するかをトークンごとに決定します。アダプターは、Oracle のコンテキスト条件付き LM からこの動作を学習し、コンテキストが役に立たない場合はベース LM を保存します。 1 つのデータセットのみでトレーニングされた単一の CCS アダプターは、ドメイン内と 4 つの配布外パーソナライゼーション ベンチマーク全体の両方で生成品質を向上させ、新しいユーザーとドメインに対する堅牢な一般化を実証します。また、CCS は、ユーザーごとの微調整と、CoS に必要な追加のコンテキスト条件付きフォワード パスを回避し、推論コストを大幅に削減します。
原文 (English)
Cautious Context Steering for Language Model Personalization
Personalizing language models (LMs) to individual user preferences is essential for aligning responses with diverse goals and backgrounds. Existing methods typically train a separate adapter for each user or learn a reward model whose scores depend on the user. Despite explicitly optimizing for each user, these methods must learn from limited observations and therefore suffer from data sparsity and poor generalization to unseen users and domains. In-context learning (ICL) and Context Steering (CoS) can instead provide more effective personalization by conditioning the base LM directly on user context and leveraging its pretrained capabilities without per-user training. Yet neither adapts the influence of that context across decoding steps: ICL leaves it uncontrolled, whereas CoS applies a fixed steering coefficient and requires two LM forward passes per step. We propose Cautious Context Steering (CCS), which adds a lightweight adapter to a frozen backbone LM to decide at each token whether and how strongly user context should affect generation. The adapter learns this behavior from an oracle context-conditioned LM and preserves the base LM when the context is not helpful. A single CCS adapter trained on only one dataset improves generation quality both in-domain and across four out-of-distribution personalization benchmarks, demonstrating robust generalization to new users and domains. CCS also avoids per-user fine-tuning and the additional context-conditioned forward pass required by CoS, substantially reducing inference cost.
ViSR-KGC: マルチモーダル ナレッジ グラフを完成させるための視覚言語モデルを使用したビジュアル サブグラフ推論
ナレッジ グラフ補完 (KGC) は、不完全なグラフ構造から欠落しているエンティティまたは関係を推測することを目的としており、エンティティがテキストや画像などの複数のモダリティに関連付けられるマルチモーダル ナレッジ グラフ補完 (MMKGC) に進化しました。従来の表現学習アプローチは埋め込みベースのパラダイムに従っており、関係固有の証拠が限られている場合には困難を伴う可能性があります。一方、LLM ベースの推論手法は通常、グラフ構造をテキスト プロンプトに線形化するため、構造トポロジが曖昧になり、重要な視覚情報が無視されます。ビジョン言語モデル (VLM) はマルチモーダル推論に優れていますが、構造化されたグラフ トポロジをネイティブに解釈することはできません。特に、ノードとエッジが複雑なセマンティクスを運ぶナレッジ グラフの場合はそうです。このギャップを埋めるために、KGC の視覚的なサブグラフ推論アプローチである ViSR-KGC を提案します。セマンティック相関を捕捉するための 3 つの補完的な機能が統合されています。表現学習によるグローバル トポロジの依存関係の特定、VLM を使用したローカル マルチモーダル証拠の分析、および事前トレーニングされたモデルに固有の必要な常識知識の提供です。学習されたマルチモーダル埋め込みに基づいて、私たちのフレームワークは最初に MMKG からコンパクトでクエリ対応のサブグラフを抽出します。次に、このサブグラフは、経験的比較によって選択されたレイアウト戦略を使用して、視覚的に解釈可能な画像に変換されます。最後に、視覚化されたサブグラフ、エンティティ画像、テキストによる説明、および回答候補が統合プロンプトに結合され、VLM が欠落しているエンティティを推測できるようになります。
原文 (English)
ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion
Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text and images. Traditional representation learning approaches follow the embedding-based paradigm and may struggle when relation-specific evidence is limited. Meanwhile, LLM-based reasoning methods typically linearize graph structures into textual prompts, which obscures structural topology and neglects vital visual information. While vision-language models (VLMs) excel at multimodal reasoning, they cannot natively interpret structured graph topology, particularly when it comes to knowledge graphs where nodes and edges carry complex semantics. To bridge this gap, we propose ViSR-KGC, a visual subgraph reasoning approach for KGC. It integrates three complementary capabilities to capture semantic correlations: identifying global topology dependencies via representation learning, analyzing local multimodal evidence using VLMs, and providing necessary commonsense knowledge inherent in pre-trained models. Based on learned multimodal embeddings, our framework first extracts a compact and query-aware subgraph from the MMKG. Then, this subgraph is transformed into a visually interpretable image using a layout strategy selected through empirical comparison.Finally, the visualized subgraph, entity images, textual descriptions, and candidate answers are combined into a unified prompt, enabling the VLM to infer the missing entity.
異種アテンションメモリのランタイムオブザーバビリティ
最新のモデルはプレーンな KV キャッシュを保持しなくなりました。潜在キャッシュ、学習されたスパース セレクター、および再帰状態はそれぞれ異なる形式でモデルのメモリを保持し、圧縮時にそれぞれ異なる方法で失敗します。 3 つの演算子で 4 つのメモリ クラスすべてをカバーするランタイム可観測性コントラクトを与え、5 つのアーキテクチャ ファミリにわたる 6 つのモデル構成でインスタンス化し、ステージごとの境界を実行可能なリクエスト レベルのリスク台帳に構成します。コントラクトはエラー メトリックをタイプとして保持します。合成はメトリックが一致する場合にのみ定義され、このチェックでは最初に合成されたチェーンが拒否されました。修復されたチェーンは 2 つの証明されたブリッジを介してメトリクスを交差し、正式なシステムが証明できないものはすべて代わりに測定され、合成された層が自動的に経験的なレベルに落とされます。すべてのクレームは認定、部分的に認定、または経験的であり、合成は最も弱い層を継承し、層はマシンによって決定されます。 $12.4$M を超えるエントリの読み取りが再生され、リクエストごとの予算とフェイルクローズされた ID 帰属を備えた 8 方向の同時実行の下で実行され、台帳は今日の証人に対する正直なトレードオフを定量化し、違反ゼロでリスク バジェットを維持します。融合された常時オン プローブは、サービング ノイズ フロア内の CUDA グラフの下で宣言された 1 層のサブセットを観察します。パックされた圧縮 KV プロトタイプを備えたサービス済みの DeepSeek-V4 スタックに適用された同じ機械は、機械が判断した差別キャンペーンを通じて、サイレント破損を正確な構造境界に特定します。これは、立ち退きのない、個人情報が分離された体制で正確であり、立ち退きまたはスロット再利用体制で観察されたすべての失敗を伴います。その計算は、途中で私たち自身の混乱した推論の 2 つを拒否しました。すべてのアーティファクト、ガード、リーン開発は https://github.com/metask-ai/witprobe-attention-memory でリリースされます。このペーパーのすべての番号は、出荷されたアーティファクトから 1 つのコマンドで再生成されます。
原文 (English)
Runtime Observability for Heterogeneous Attention Memory
Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model's memory in a different form, and each fails differently under compression. We give a runtime observability contract that covers all four memory classes with three operators, instantiate it on six model configurations across five architecture families, and compose the per-stage bounds into an executable request-level risk ledger. Contracts carry their error metric as a type -- composition is only defined when metrics match, and this check rejected our own first composed chain; the repaired chain crosses metrics through two proved bridges, and whatever no formal system can certify is measured instead, dropping the composed tier to empirical automatically: every claim is certified, partially certified, or empirical, composition inherits the weakest tier, and the tier is decided by the machine. Replayed over $12.4$M entry reads and run under eight-way concurrency with per-request budgets and fail-closed identity attribution, the ledger quantifies the honest trade-off on today's witness and holds its risk budget with zero violations. A fused always-on probe observes a declared one-layer subset under CUDA graphs inside the serving noise floor. Applied to a served DeepSeek-V4 stack with a packed compressed-KV prototype, the same machinery localizes a silent corruption to a precise structural boundary -- exact in the eviction-free, identity-isolated regime, with every observed failure in an eviction or slot-reuse regime -- through a machine-adjudicated discrimination campaign whose calculus rejected two of our own confounded inferences along the way. All artifacts, guards, and the Lean development are released at https://github.com/metask-ai/witprobe-attention-memory; every number in this paper regenerates from the shipped artifacts by one command.
見るだけが決断ではない: マルチモーダル LLM は有能な CEO として機能できるか?
大規模な言語モデルは、自律的な意思決定エージェントとしてますます適用されています。ただし、経営幹部のビジネス上の意思決定では、既存のベンチマークはテキストのみの設定に限定されています。このため、モデルが視覚的なビジネス証拠を認識し、それを効果的に統合して意思決定の品質を向上できるかどうかは不明確です。 C-SUITEBENCH は、50 のシナリオにわたってテキストのみとマルチモーダルのペアの条件下で 5 つの意思決定タスクを含む、制御されたマルチモーダル ベンチマークです。フロンティアモデル9名を最高経営責任者に据え、意思決定能力を評価します。多様なインプットにより、証拠中心の推論が一貫して向上し、リスク予測と取締役会での正当化において最大かつ最も信頼性の高い利益が現れます。しかし、マルチモーダル統合のパラドックスが明らかになりました。視覚的なビジネス情報を追加すると、視覚的な根拠自体は向上しても、9 つのモデルすべてに対する制約されたリソース割り当てが低下します。アブレーション実験により、この失敗は信号の混雑から生じることが明らかになりました。各視覚チャネルは個別には役に立ちますが、それらの組み合わせはデコード中に制約を満たすことを妨げます。これらの発見は、マルチモーダル エージェントにおいて視覚認識と制約されたアクションが分離可能なボトルネックであること、および無差別な視覚拡張が一か八かの意思決定に悪影響を及ぼす可能性があり、将来のエグゼクティブ AI システムの選択的接地戦略を動機付ける可能性があることを示しています。
原文 (English)
Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?
Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing benchmarks are limited to textonly settings. This makes it unclear whether models can perceive visual business evidence and effectively integrate it to improve decision quality. We introduce C-SUITEBENCH, a controlled multimodal benchmark that includes five decision tasks under paired text-only and multimodal conditions across 50 scenarios. We place nine frontier models in the role of a chief executive officer and evaluate their decision-making ability. Multimodal inputs consistently improve evidence-centric reasoning, with the largest and most reliable gains appearing in risk forecasting and board-facing justification. However, we uncover a multimodal integration paradox: adding visual business information degrades constrained resource allocation for all nine models, even as visual grounding itself improves. Ablation experiments reveal that this failure emerges from signal crowding, although each visual channel helps individually, their combination disrupts constraint satisfaction during decoding. These findings demonstrate that visual perception and constrained action are separable bottlenecks in multimodal agents, and that indiscriminate visual augmentation can harm high-stakes decision making, motivating selective grounding strategies for future executive AI systems.
防衛および国家安全保障のオントロジー間の相互運用性の向上: 分析および評価タスク
オントロジーとナレッジグラフの使用は、防衛と国家安全保障の分野でますます普及してきています。学界、産業界、政府主導の取り組みを通じて、数多くのオントロジーが開発されてきました。このドメインの広さと専門性により、多様な防衛および国家安全保障のオントロジーにわたる相互運用性を実現することは依然として大きな課題です。この作業では、60 を超える公的に利用可能なオントロジーを分析して文書化し、オントロジー調整評価イニシアチブ (OAEI) の新しいトラックを導入します。このトラックは、8 つのマッチング タスク、コンセンサス調整、および手動でキュレーションされた (シルバー スタンダード) マッピングで構成されます。コンセンサス アライメントは、いくつかの最先端のオントロジー アライメント システムの出力を集約することによって導出されます。シルバースタンダードは、独自のマッピング (つまり、1 つのシステムのみによって提案されるマッピング) のサブセットとともにコンセンサス調整を手動で検証することによって取得されます。
原文 (English)
Improving Interoperability among Defence and National Security Ontologies: Analysis and Evaluation Tasks
The use of ontologies and knowledge graphs is becoming increasingly widespread in the defence and national security domain. Numerous ontologies have been developed through initiatives led by academia, industry, and government. Achieving interoperability across diverse defence and national security ontologies remains a major challenge due to the domain's breadth and specialisation. In this work, we analyse and document over 60 publicly available ontologies and introduce a new track for the Ontology Alignment Evaluation Initiative (OAEI). This track comprises eight matching tasks, consensus alignments and manually-curated (silver-standard) mappings. The consensus alignments are derived by aggregating the outputs of several state-of-the-art ontology alignment systems. The silver-standard is obtained from the manual validation of the consensus alignment together with a subset of the unique mappings (i.e., mappings suggested by only one system).
グラフ足場による証拠根拠によるパーソナライズされたディープリサーチクエリの絞り込み
ユーザーのリクエストは、ディープリサーチエージェントの調査仕様として機能し、どのような証拠を求め、それをどのように合成するかを決定します。パーソナライズされた詳細な調査では、これらの仕様はユーザーの目標、制約、好み、評価基準をさらに反映する必要があります。ユーザーコンテキストは、ディープリサーチパイプライン内、またはその入力として提供されるリサーチ仕様に組み込むことができます。私たちは後者に重点を置き、ユーザーのリクエストをパーソナライズされたリサーチ仕様に改良してから、変更のないディープリサーチエージェントに渡します。これには、どのフレーミング要素が関連しているか、利用可能なユーザー コンテキストがそれらを十分にサポートしているかどうか、ユーザー メモリを取得するか、ユーザーに質問するか、クエリを停止して調整するかという、3 つの組み合わせた決定を解決する必要があります。トレーニングのために、G-STEER はフレーミング要素を、その依存関係をキャプチャするインテント引き出しグラフ内の引き出しターゲットとして整理します。さまざまな要因依存関係と証拠条件にわたるグラフ足場の軌跡から明確化ポリシーを学習します。このポリシーは、対象範囲と証拠取得コストのバランスをとりながら、洗練されたクエリを生成します。実験の結果、G-STEER は、評価された両方の DRA 全体で最も強力な全体的な加重ターゲット カバレッジと最も高い下流レポートのパーソナライゼーションを実現し、強力な明確化ベースラインの約 3 分の 1 のユーザー質問をしていることが示されています。
原文 (English)
Personalized Deep Research Query Refinement with Graph-Scaffolded Evidence Grounding
User requests serve as research specifications for deep research agents, shaping what evidence to seek and how to synthesize it. In personalized deep research, these specifications must additionally reflect user goals, constraints, preferences, and evaluation criteria. User context can be incorporated either within the deep research pipeline or into the research specification provided as its input. We focus on the latter, refining the user request into a personalized research specification before passing it to an unchanged deep research agent. This requires resolving three coupled decisions: which framing factors are relevant, whether the available user context sufficiently supports them, and whether to retrieve user memory, ask the user, or stop and refine the query. For training, G-STEER organizes framing factors as elicitation targets in an Intent Elicitation Graph that captures their dependencies. It learns a clarification policy from graph-scaffolded trajectories spanning diverse factor dependencies and evidence conditions. The policy produces a refined query while balancing target coverage against the costs of evidence acquisition. Experiments show that G-STEER achieves the strongest overall weighted target coverage and the highest downstream report personalization across both evaluated DRAs, while asking roughly one third as many user questions as a strong clarification baseline.
AppDeltaWorld: モバイル GUI エージェント用の遷移に基づくデルタ コード ワールド モデル
モバイル GUI エージェントは、ピクセル認識とタッチ アクションを通じてアプリを操作できるため、長期的なモバイル インタラクション ポリシーを収集および改善するための有望なインターフェイスになります。ただし、機密性の高いアプリやプライバシーが重要な操作では、実際の軌跡を取得するのは困難です。同時に、既存のシミュレート環境はスケールアップにコストがかかり、GUI ワールド モデルは依然として不安定な生成、制限されたモダリティ カバレッジ、および一貫性のないアクション遷移ロジックに悩まされています。これらの制限に対処するために、私たちは、制約のない画像やテキストの説明としてではなく、到達可能なコード更新として次の GUI を予測する、遷移に基づいたデルタ コード ワールド モデルである AppDeltaWorld を提案します。 AppDeltaWorld は、アクション遷移制約の下でアプリ固有のレベル 1 HTML 参照を取得し、現在の画面、アクション、予測される次の画面のテキスト、および取得された構造を条件としたレベル 2 の実行可能 HTML を生成し、生成されたビジュアル アセットをブラウザーのレンダリング前に画像スロットに挿入します。ワールド モデルとして、AppDeltaWorld は、Code2World 評価の下で CMGUIBench-500 で最高の忠実度を達成し、画像のみおよびコードのみのベースラインと比較して、構造レイアウトと UI 要素の再構築において明らかな向上を実現しています。 AppDeltaWorld はトレーニング環境として、フィルタリングされた閉ループ SFT データ構築をサポートしています。これにより、公的監視と組み合わせることで、AppDeltaAgent は AndroidLens で最先端のパフォーマンスを達成し、MobileGym と MobileWorld で一貫した利益を得ることができます。さらに、ワールド モデル ベースのテスト時間強化学習により、ポリシーの適応が可能になり、実際のアプリとの追加の対話なしでさらなる改善が見られます。
原文 (English)
AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
Mobile GUI agents can operate apps through pixel perception and touch actions, making them a promising interface for collecting and improving long-horizon mobile interaction policies. However, real trajectories are difficult to obtain for sensitive apps and privacy-critical operations. At the same time, existing simulated environments are costly to scale up, and GUI world models still suffer from unstable generation, limited modality coverage, and inconsistent action-transition logic. To address these limitations, we propose AppDeltaWorld, a transition-grounded delta code world model that predicts the next GUI as a reachable code update rather than as an unconstrained image or text description. AppDeltaWorld retrieves app-specific Level-1 HTML references under an action-transition constraint, generates Level-2 executable HTML conditioned on the current screen, action, predicted next-screen text, and retrieved structure, and inserts generated visual assets into image slots before browser rendering. As a world model, AppDeltaWorld achieves the highest fidelity on CMGUIBench-500 under Code2World evaluation, with clear gains in structural layout and UI element reconstruction over image-only and code-only baselines. As a training environment, AppDeltaWorld supports filtered closed-loop SFT data construction that, when combined with public supervision, enables AppDeltaAgent to achieve state-of-the-art performance on AndroidLens and consistent gains on MobileGym and MobileWorld. Moreover, world-model-based test-time reinforcement learning enables policy adaptation and shows further improvements without additional interaction with real apps.
ECG-LENS: リードを意識した臨床コンテキストを強化した ECG レポートの生成と評価
心電図検査 (ECG) は、心血管疾患を診断するために最も広く使用されている非侵襲的ツールの 1 つですが、マルチリード ECG 記録を信頼できる臨床レポートに変換することは依然として困難です。 ECG レポートの生成を自動化すると、臨床医の解釈作業負荷が軽減され、診断効率が向上し、サービスが十分に受けられていない地域での心臓評価へのアクセスが拡大する可能性があります。画像ベースのレポート作成タスクとは異なり、ECG 解釈では、微妙な時間的形態の分析と、それに続く緻密な臨床用語で表現された一貫した診断推論が必要です。既存のシステムは主に分類に焦点を当てていますが、現在のレポート生成方法では、実際の臨床使用には依然として不十分な出力が生成されることがよくあります。これらの課題に対処するために、マルチリード信号モデリング、診断を意識した表現、および臨床に基づいたテキスト生成を統合するエンドツーエンドの ECG レポート生成フレームワークである ECG-LENS を提案します。 ECG-LENS は、局所的な波形形態を保存するリードごとのエンコーダと、リード間の依存関係を捕捉するグローバル エンコーダを組み合わせます。レポート生成をガイドするために、GPT-2 デコーダーを条件付ける臨床的に強化されたテキスト プロンプトと信号表現を融合します。さらに、モデルが臨床的に意味のある所見に焦点を当てるのに役立つ ECG 固有のレポート前処理戦略を導入します。最後に、語彙メトリクスはレポートの品質を過小評価または過大評価する可能性があるため、生成されたレポートと参照レポートから抽出された診断ラベル間の一致を測定する BERT ベースの ECG 固有のメトリクスである F1-ECGBERT を提案します。 PTB-XL のドメイン内実験と MIMIC-IV-ECG のクロスドメイン評価では、ECG-LENS が常に最先端の方法を上回り、最強のベースラインに対して METEOR、ROUGE-L、および F1-ECGBERT でそれぞれ 4.0%、6.3%、および 11.5% の絶対利得を示しました。
原文 (English)
ECG-LENS: Lead-Aware Clinical Context Enriched ECG Report Generation and Evaluation
Electrocardiography (ECG) is one of the most widely used non-invasive tools for diagnosing cardiovascular disease, but transforming multi-lead ECG recordings into reliable clinical reports remains challenging. Automating ECG report generation could reduce clinicians' interpretive workload, improve diagnostic efficiency, and expand access to cardiac assessment in underserved communities. Unlike image-based report-generation tasks, ECG interpretation requires the analysis of subtle temporal morphologies, followed by coherent diagnostic reasoning expressed in dense clinical terminology. Existing systems predominantly focus on classification, while current report-generation methods often produce outputs that remain inadequate for practical clinical use. To address these challenges, we propose ECG-LENS, an end-to-end ECG report-generation framework that jointly integrates multi-lead signal modeling, diagnosis-aware representations, and clinically grounded text generation. ECG-LENS combines lead-wise encoders that preserve localized waveform morphology with a global encoder that captures inter-lead dependencies. To guide report generation, we fuse signal representations with clinically enriched textual prompts that condition a GPT-2 decoder. We further introduce an ECG-specific report-preprocessing strategy that helps the model focus on clinically meaningful findings. Finally, because lexical metrics may under- or overestimate report quality, we propose F1-ECGBERT, a BERT-based, ECG-specific metric that measures agreement between diagnostic labels extracted from generated and reference reports. In-domain experiments on PTB-XL and cross-domain evaluation on MIMIC-IV-ECG show that ECG-LENS consistently outperforms state-of-the-art methods, with absolute gains of 4.0%, 6.3%, and 11.5% in METEOR, ROUGE-L, and F1-ECGBERT, respectively, over the strongest baselines.
GSBF: 環境を考慮したビームフォーミングのためのガウス スプラッティング
ビームフォーミングは、多入力多出力 (MIMO) 通信システムにおいて重要な役割を果たします。ただし、従来のビームフォーミング設計では通常、正確な瞬間チャネル状態情報 (CSI) と反復的な最適化が必要であり、これによりパイロットのオーバーヘッドと計算の複雑さが大幅に増加します。無線伝播は本質的に物理幾何学によって支配されることを認識し、マルチモーダル データに基づいて環境を考慮したビームフォーミング (GSBF) パイプライン用の 3D ガウス スプラッティングを開発します。これは、永続的な 3D ガウス表現を通じて環境を特徴付けます。具体的には、GSBF は相反性を保持する双方向球面ガウス (Bi-SG) カーネルを使用して環境散乱応答をモデル化し、両面電磁ラスタライゼーションを実行して角度プロパゲータ マップをレンダリングします。次に、レンダリングされたマップは、過剰に完成した配列多様体辞書を通じて集約され、定弾性ビームフォーマーに投影されます。これにより、オンラインの瞬間的な CSI を使用せずに、アクセス ポイント (AP) の姿勢とユーザーの位置から直接ビームが合成されます。シミュレーションでは、GSBF が一貫して網羅的ビーム アライメント (EBA) などのベースラインを上回り、遅延が低いことが実証されています。
原文 (English)
GSBF: Gaussian Splatting for Environment-Aware Beamforming
Beamforming plays a key role in multiple-input-multiple-output (MIMO) communication systems. However, conventional beamforming design normally requires accurate instantaneous channel state information (CSI) and iterative optimization, which incur substantial pilot overhead and computational complexity. Recognizing that radio propagation is intrinsically governed by the physical geometry, we develop a 3D Gaussian splatting for environment-aware beamforming (GSBF) pipeline based on multi-modal data, which characterizes the environment through a persistent 3D Gaussian representation. Specifically, GSBF models the environmental scattering response with reciprocity-preserving bidirectional spherical Gaussian (Bi-SG) kernels and performs two-sided electromagnetic rasterization to render an angular propagator map. The rendered map is then aggregated through an over-complete array-manifold dictionary and projected to the constant-modulus beamformers, thereby synthesizing beams directly from the access point (AP) pose and user position without online instantaneous CSI. Simulations demonstrate that GSBF consistently outperforms baselines such as exhaustive beam alignment (EBA) with lower latency.
CourseGraph: 大学間のコンピューター サイエンス コースの重複と相違点を見つける
Erasmus+ などの学生流動プログラムにより、学生は他の大学でコースを受講でき、学問的および文化的視野を広げることができます。ただし、この柔軟性は実際的な課題にもつながります。それは、学生が家庭のカリキュラムのコースと実質的に重複するコースを他の場所で受講しないようにすることです。この研究では、学位プログラムに組み込むコースを評価する際にカリキュラム管理者が従うプロセスから得られた洞察に基づいて、外部コースの評価を自動化する方法論である CourseGraph を提案します。 Course-Graph は、コースの Web ページからコースのタイトル、説明、学習成果などの情報を抽出します。次に、この情報は BERT ベースの言語モデルを使用して意味論的に表現され、その後、コース間のペアごとの類似性が計算されます。この情報は、ランダム フォレスト分類器によって使用され、海外の候補コースが学生のカリキュラムに既に含まれているコースと重複するかどうかが判断されます。私たちは、(1) アイントホーフェン工科大学のコンピューター サイエンス プログラム、実質的に重複するコースに関する情報が含まれている、(2) ルンド大学のコンピューター サイエンス プログラムに登録している学生による 6 つの承認された国際プログラム (カリキュラム管理者による対応する決定を含む) を使用して、CourseGraph を評価します。実験結果は、CourseGraph が重複するコースを特定し、大学間でのカリキュラムの調整をサポートするための効果的なアプローチを提供することを示しています。
原文 (English)
CourseGraph: Finding overlaps and differences in Computer Science courses across universities
Student mobility programs such as Erasmus+ enable students to take courses at other universities, broadening their academic and cultural horizons. However, this flexibility also leads to a practical challenge: ensuring that students do not take courses elsewhere that substantially overlap with courses in their home curriculum. In this work, we propose CourseGraph, a methodology that automates the evaluation of external courses based on insights obtained from the process followed by curriculum administrators when assessing courses for inclusion in a degree program. Course- Graph extracts information such as course titles, descriptions, and learning outcomes from the course webpage. Then, this information is represented semantically using a BERT-based language model, after which the pair-wise similarity between courses can be computed. This information is then used by a Random Forest classifier to determine whether a candidate course abroad overlaps with a course already contained in the student's curriculum. We evaluate CourseGraph using (1) the Computer Science program at Eindhoven University of Technology, which contains information about courses with substantial overlap, and (2) six approved international programs from students enrolled in the Computer Science program at Lund University, including the corresponding decisions made by a curriculum administrator. The experimental results indicate that CourseGraph provides an effective approach for identifying overlapping courses and supporting curriculum alignment across universities.
GAUGE: シミュレーション エンジンとビデオ ワールド モデルの物理的忠実度に関する測定に基づいたベンチマーク
物理エンジンは、身体化された知能の大規模なトレーニングと評価を容易にする一方、生成ビデオ世界モデルは、将来の状態と相互作用の暗黙的なシミュレーターとして登場しています。しかし、物理的忠実度の既存の評価は単独で行われることが多く、知覚的な類似性や人間の判断に大きく依存しており、どの物理的原理やパラメータが違反されているかについての洞察は限られています。数値シミュレーターと生成ビデオ ワールド モデルが現実世界の物理学をどのように再現するか、または現実世界から逸脱するかを共同で評価するための、現実世界に基づいた診断ベンチマークである GAUGE を紹介します。これは、剛体、フレキシブル ケーブル、テキスタイル、および体積変形可能なオブジェクトをカバーする 22 の制御されたタスク ファミリで構成されています。これらのタスクは、現実世界の軌道に基づいて、調整された物理メタデータ、不確実性の注釈、タスク固有の観測可能量と組み合わせて、衝突、摩擦、運動量伝達、振動、自己接触、さまざまな材料と条件にわたる変形などの基本的な物理プロセスをカバーします。一般化された軌道誤差を使用して 14 のタスク ファミリで Isaac Sim、Genesis、および Newton のベンチマークを実行し、物理法則の一貫性と推論されたパラメーターの時間的安定性をテストすることにより、5 つの剛体タスクで 6 つの画像からビデオへのモデルを評価します。私たちの結果は、均一に忠実な物理エンジンがなく、衝撃的接触、繊維の素早い動き、体積変形で最も大きな差異が生じることを明らかにしました。さらに、ビデオ ワールド モデルは、誤った加速度、運動量伝達、振動タイミングを回復しながら、予想される方程式形式の軌道を生成できることもわかりました。 GAUGE は、より物理的に忠実なシミュレーターと身体化された知性の世界モデルを開発するための基礎を築きます。
原文 (English)
GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models
Physics engines facilitate large-scale training and evaluation for embodied intelligence, while generative video world models are emerging as implicit simulators of future states and interactions. However, existing evaluations of physical fidelity are often conducted in isolation and rely heavily on perceptual similarity or human judgments, providing limited insight into which physical principles or parameters are violated. We introduce GAUGE, a real-world-grounded diagnostic benchmark for jointly evaluating how numerical simulators and generative video world models reproduce or deviate from real-world physics. It comprises 22 controlled task families covering rigid bodies, flexible cables, textiles, and volumetric deformable objects. Grounded in real-world trajectories and paired with calibrated physical metadata, uncertainty annotations, and task-specific observables, these tasks cover fundamental physical processes including collision, friction, momentum transfer, oscillation, self-contact, and deformation across diverse materials and conditions. We benchmark Isaac Sim, Genesis, and Newton on 14 task families using generalized trajectory errors, and evaluate 6 image-to-video models on 5 rigid-body tasks by testing physical-law consistency and the temporal stability of inferred parameters. Our results reveal no uniformly faithful physics engine, with the largest discrepancies arising in impulsive contact, rapid textile motion, and volumetric deformation. We further find that video world models can produce trajectories with the expected equation form while recovering incorrect accelerations, momentum transfer, and oscillation timing. GAUGE lays the groundwork for developing more physically faithful simulators and world models for embodied intelligence.
ビデオゲーム データ アノテーション用の VLM
Vision Language Models (VLMs) and Artificial Intelligence (AI) agents have revolutionized how engineers approach complex problems in real-world applications. Their adoption in video games is on the other hand limited by the extreme variability of the synthetic scenarios and their poor compliance with real-world physics. Here we investigate the use of VLMs for annotating video game frame sequences with reward signals, a task with several potential applications including, among others, conditioned training and offline reinforcement learning. We show that VLMs often struggle to answer basic questions on racing video games (although we observed a similar behavior on other game genres) and discuss countermeasures such as VLM output mixing and prompt optimization.また、入力シーケンスの長さ、解像度、質問のバッチ処理がアノテーションの品質とそのトークン消費にどのような影響を与えるかについても示します。
原文 (English)
VLMs for Videogame Data Annotation
Vision Language Models (VLMs) and Artificial Intelligence (AI) agents have revolutionized how engineers approach complex problems in real-world applications. Their adoption in video games is on the other hand limited by the extreme variability of the synthetic scenarios and their poor compliance with real-world physics. Here we investigate the use of VLMs for annotating video game frame sequences with reward signals, a task with several potential applications including, among others, conditioned training and offline reinforcement learning. We show that VLMs often struggle to answer basic questions on racing video games (although we observed a similar behavior on other game genres) and discuss countermeasures such as VLM output mixing and prompt optimization. We also show how input sequence length, resolution, and question batching affect the annotation quality and its token consumption.
VLM アノテーション付きデータセットで条件付きビデオ ゲーム エージェントをトレーニングする
強化学習 (RL) は強力ではありますが、ポリシー学習にとって使いやすい手法とは言えません。ビデオ ゲームの特定のケースでは、トレーニングの報酬を取得するために (たとえば、環境から報酬を収集するために) ゲーム エンジンへのアクセスが必要です。さらに、報酬を適切に特定して重み付けするには、通常、難しい試行錯誤のアプローチが必要です。最後に、報酬は希薄であることが多く、最終的に学習したポリシーに報酬がどのように影響するかを理解するのは簡単な作業ではありません。これらの問題を軽減するために、人間が定義した報酬を抽出するように指示されたビジョン言語モデル (VLM) を使用してビデオ ゲーム データセットにアノテーションを付けることを提案します。我々は、オフライン RL を使用して、望ましい利益に応じて反応する条件付きエージェントをトレーニングできることを示し、初期の実験で明らかになった困難と限界について説明します。
原文 (English)
Training a Conditioned Video Game Agent on a VLM Annotated Dataset
Reinforcement Learning (RL) is a powerful but far from easy-to-use technique for policy learning. In the specific case of video games, access to the game engine is required to get rewards for training (e.g. to collect rewards from the environment). Furthermore, the proper identification and weighting of the rewards generally requires a difficult trial-and-error approach. Lastly, rewards are often sparse and understanding how they eventually affect the learned policy is a non-trivial exercise. To ease these issues we propose annotating a video game dataset with Vision Language Models (VLMs) instructed to extract human defined rewards. We show that offline RL can then be used to train a conditioned agent that responds accordingly to the desired returns and we discuss the difficulties and limitations that emerged in our early experiments.
分析階層プロセスにおけるランキング依存の一対比較パターンの安定性
この論文では、ランキングに依存した意思決定支援方法をいくつか取り上げています。比較対象の順序情報を使用すると、推定時の専門家データの品質を向上させ、専門家が実行する必要がある比較の数を減らすことができます。この論文では、分析階層プロセスで使用できる 3 つの不完全なランキング依存のペアごとの比較パターン (最良-最悪法、最良-次善 (上位 2) 法、および元の最大差分法) を比較します。最初の 2 つの比較パターン (およびそれぞれの方法) は不完全ですが、3 番目は完全なものになる可能性があります。これら 3 つの方法を安定性の観点から専門家のエラーと比較できる条件を決定します。また、3 つの方法を比較したシミュレーション型実験の結果も示します。この研究により、最も安定した不完全なランキング依存のペアごとの比較パターンを定義し、専門家セッションの結果の信頼性を損なうことなく比較の数を減らすことができます。この研究は、不確実な環境における意思決定支援のアルゴリズム、認知、および応用の側面に貢献します。
原文 (English)
Stability of Ranking-dependent Pair-wise Comparison Patterns in the Analytic Hierarchy Process
The paper addresses several ranking-dependent decision support methods. Ordinal information on compared objects can be used to improve the quality of expert data during estimation and help reduce the number of comparisons that the experts need to perform. In the paper we compare three incomplete ranking-dependent pair-wise comparison patterns which can be used in the Analytic Hierarchy Process - Best-worst method, Best-Second Best (Top 2) method, and the original maximum difference method. The first two comparison patterns (and respective methods) are incomplete, while the third can be a complete one. We determine conditions under which these three methods can be compared in terms of stability to expert errors. We also present the results of a simulation-type experiment, in which the three methods are compared. The research allows us to define the most stable incomplete ranking-dependent pair-wise comparison pattern and reduce the number of comparisons without loss of credibility of expert session results. The research contributes to algorithmic, cognitive, and applied aspects of decision support in uncertain environments.
空間解像度のための時間的ブリッジ: 双方向アライメントによる気候データの超解像度の強化
高解像度の気候データは、気象予測や、さまざまな領域にわたる意思決定支援に情報を提供するために不可欠です。ただし、このような高解像度の気候情報の取得には法外なコストがかかることが多く、データ駆動型の気象予測モデルの開発が必要になります。これらのモデルは、低解像度の入力から詳細な気候データを生成することを目的としています。これは、気候データ超解像度 (SR) と呼ばれるプロセスです。それにもかかわらず、気候データ SR のディープラーニングにおける最近の進歩は、主に単一フレームの空間情報を活用することに焦点を当てており、SR の結果を向上させる可能性がある異なる時間フレーム間の時間的相関はほとんど無視されています。さらに、気候データは本質的に確率的でノイズが多いため、オプティカル フロー モデルなどの広く使用されている時間的調整方法がこの状況では無効になります。したがって、暗黙的な時間相関を効果的に捕捉する気候データ SR に合わせたフレームワークの開発は未解決の課題のままです。この目的を達成するために、我々は双方向の時間的アライメントを備えた新しい時間拡張フレームワークを提案します。本質的に、私たちのフレームワークは時間的なブリッジを確立し、双方向アライメントを通じて気候データ SR の空間解像度を向上させ、SR パフォーマンスの向上につながります。このフレームワーク内で、Paired Latent Mapping は、潜在空間を統合することによって空間の位置合わせとノイズの低減を実現します。次に、双方向時間アライメントは、連続する潜在フレームで前方ネットワークと後方ネットワークをトレーニングすることによって時間相関を捕捉します。時間的強化超解像度は、気候データ SR のフレームワーク全体を最適化します。大規模な現実世界のデータセットでの実験により、私たちのフレームワークの優れたパフォーマンスが実証されました。
原文 (English)
Temporal Bridges for Spatial Resolution: Enhancing Climate Data Super-Resolution with Bidirectional Alignment
High-resolution climate data is crucial for meteorological predictions and for informing decision support across diverse domains. However, the acquisition of such high-resolution climate information is often prohibitively costly, necessitating the development of data-driven meteorological prediction models. These models aim to generate fine-grained climate data from low-resolution inputs, a process termed climate data super-resolution (SR). Nevertheless, recent advancements in deep learning for climate data SR have primarily focused on leveraging single-frame spatial information, largely neglecting the temporal correlations between different time frames that could enhance SR outcomes. Furthermore, climate data are inherently stochastic and noisy, rendering widely used temporal alignment methods, such as optical flow models, ineffective in this context. Consequently, the development of a framework tailored for climate data SR that effectively captures implicit temporal correlations remains an unresolved challenge. To this end, we propose a novel Temporal-Enhanced framework with bidirectional temporal alignment. In essence, our framework establishes a temporal bridge to enhance spatial resolution in climate data SR through bidirectional alignment, leading to improved SR performance. Within this framework, Paired Latent Mapping achieves spatial alignment and noise reduction by unifying latent spaces. Then a Bidirectional Temporal Alignment captures temporal correlations by training forward and backward networks on consecutive latent frames. Temporal Enhanced Super-resolution then optimizes the entire framework for climate data SR. Experiments on large-scale real-world datasets demonstrated the superior performance of our framework.
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few p…
OPERA: 言語モデルエージェントを使用した信頼性の高い自律光学実験のためのオペレーター残留フィードバック
自律エージェントは、実験の成功を反映していない可能性のあるスコアを使用してアクションを選択します。私たちは、光学実験用のオペレーター残存フレームワークである OPERA を開発しました。実験アクションを光学演算子として表し、物理的に解釈可能な残差を使用して結果を評価します。オペレーターは、測定、制御、または再構成に対する実行可能な変更を指定し、残差は、指定された物理的条件からの逸脱を報告します。エージェントは両方を使用してオペレーターを選択、結合、または生成し、物理的パフォーマンスは保留された参照に対して独立して評価されます。 3 つの光学タスク全体で、スコアのみのフィードバックによって生成されたスコアは、物理的な改善なしで 23.6 ~ 39.0% の決定において増加しました。一方、オペレーター残存フィードバックの場合は 0.9 ~ 1.9% でした。オペレーターの残存フィードバックにより、タスクの目標を達成および維持できる確率が高まり、実験予算が削減されました。デジタルツインで選択されたプロトコルは 3 つの光学機器に転送され、実験を繰り返すと、構造化光再構成における投影バジェットがより低いことが示されました。演算子と残差は、測定可能な物理的証拠を使用して自律的な決定を導きます。
原文 (English)
OPERA: Operator-residual feedback for reliable autonomous optical experiments with language-model agents
Autonomous agents choose actions using scores that may not reflect experimental success. We developed OPERA, an operator-residual framework for optical experiments. It represents experimental actions as optical operators and evaluates their outcomes using physically interpretable residuals. Operators specify executable changes to measurement, control or reconstruction, while residuals report departures from specified physical conditions. The agent uses both to select, combine or generate operators, and physical performance is evaluated independently against a withheld reference. Across three optical tasks, score-only feedback produced score increases without physical improvement in 23.6--39.0\% of decisions, compared with 0.9--1.9\% for operator-residual feedback. Operator-residual feedback increased the probability of reaching and maintaining task targets and reduced experimental budgets. Protocols selected in digital twins were transferred to three optical instruments, and repeated experiments showed a lower projection budget in structured-light reconstruction. Together, operators and residuals guide autonomous decisions using measurable physical evidence.
放牧ベースの生産システムにおける牛群レベルの成長パターンと体重増加予測のためのハイブリッド機械学習フレームワーク
商業用放牧システムでは家畜の観察が不規則になり、牛の成長予測が困難になります。この研究では、オーストラリア南東部で 2022 年から 2024 年の間に収集された自動センシング観察を使用して、牛群レベルの牛体重予測のためのハイブリッド機械学習フレームワークを開発しました。毎週の生体重観察、人口統計学的変数、および遅れた環境予測変数が、構造化された予測データセットに統合されました。群れレベルの予測軌跡は、動物レベルの予測の時間的集計によって生成されました。残差、スタック、カスケード、アンサンブル支援フレームワークを含む 4 つのハイブリッド アーキテクチャ ファミリが評価されました。 ARIMA、LSTM、および GRU モデルを比較ベースラインとして使用しました。独立したテストにより、複数の予測範囲にわたって強力な予測一致が実証されました。 GB、RF、NN のカスケード アーキテクチャは、テスト R^2 が 0.889、RMSE が 21.319 kg、MAE が 15.462 kg という最高のパフォーマンスを達成しました。ハイブリッド アーキテクチャは、まばらな観測条件下でリカレント シーケンシャル モデルよりも優れた堅牢性を維持しました。予測誤差は、拡張された予測範囲にわたって徐々に増加しました。特徴重要度分析により、動物の年齢、降雨量、気温が群れレベルの成長予測に影響を与える主要な予測因子であることが特定されました。提案されたフレームワークは、異種センシング環境下での飼料の割り当て、放牧管理、家畜のマーケティング決定をサポートする可能性があります。
原文 (English)
Hybrid Machine Learning Framework for Herd-Level Cattle Growth Pattern and Weight Gain Forecasting in Grazing-Based Production Systems
Commercial grazing systems yield irregular livestock observations, which challenge cattle growth forecasting. This study developed a hybrid machine learning framework for herd level cattle weight forecasting using automated sensing observations collected between 2022 and 2024 in southeastern Australia. Weekly live weight observations, demographic variables, and lagged environmental predictors were integrated into structured forecasting datasets. Herd level forecasting trajectories were generated through temporal aggregation of animal level predictions. Four hybrid architecture families were evaluated, including residual, stacked, cascade, and ensemble assisted frameworks. ARIMA, LSTM, and GRU models were used as comparative baselines. Independent testing demonstrated strong predictive agreement across multiple forecasting horizons. The cascade GB to RF to NN architecture achieved the best performance, with a test R^2 of 0.889, RMSE of 21.319 kg, and MAE of 15.462 kg. Hybrid architectures maintained greater robustness than recurrent sequential models under sparse observation conditions. Forecasting error increased progressively across extended prediction horizons. Feature importance analysis identified animal age, rainfall, and temperature as dominant predictors influencing herd level growth forecasting. The proposed framework may support feed allocation, grazing management, and livestock marketing decisions under heterogeneous sensing environments.
ヘラルド: 不正事実の監査と取得証明の報酬に対する最小限の修復
検索エージェントの報酬には、回答の質、引用の根拠、ツールのコスト、ハッキング対策の用語が組み合わされています。したがって、高いスコアは引用された証拠が取得されたことを意味する必要はなく、追加のペナルティによって取り消される可能性があります。まったく同じ質問による介入を適用し、候補者に表示されるものをオラクル情報から分離し、ポリシーを最適化する前に検出器コントラクトを列挙するオフライン監査である HERALD を紹介します。 HotpotQA、2WikiMultiHopQA、MuSiQue の 4 つの Qwen3-8B プールでは、$R_0$ が検索削除と偽 ID を拒否しますが、ラベルフリーの引用ロンダリング攻撃は成功します。完全な $2^3$ アブレーションは、観察された包含最小修復として、$L$ の標的強化を特定します -- 取得された証拠にないコーパスの一節を引用 -- $R[L]$ は、0.50% の片側クラスター上限を持つ経験的 ASR がゼロです。このギャップは、プール ルール、目に見える BM25 攻撃者、および 4 つのモデルにわたって存続します。攻撃により Oracle support-ID ペナルティが解除されると、より広範な強化は脆弱なままになります。ベンチマークごとに 256 のペアの質問で評価された厳密な 500 万トークンのマッチング トレーニングの下では、$R[L]$ は HotpotQA と 2Wiki の EM 非劣性ゲートを満たしていますが、MuSiQue は満たしていません。 2Wiki と MuSiQue では、等しいスイートの引用精度とサポート再現率が 2.02 ポイントと 1.46 ポイント向上し、サポートされていない引用は 1.69 ポイント低下し、ロンダリング攻撃可能性が低下しました。自然な $L$ は減少せず、検出器は 58,368 のトレーニング トラジェクトリのうち 18 にしか現れません。したがって、HERALD は、堅牢なスコアリング、まばらな学習信号、およびポリシー転送を分離します。
原文 (English)
HERALD: Counterfactual Audits and Minimal Repairs for Proof-of-Retrieval Rewards
Search-agent rewards mix answer quality, citation grounding, tool cost, and anti-hacking terms; a high score therefore need not imply that cited evidence was retrieved, and added penalties can cancel. We introduce HERALD, an offline audit that applies exact same-question interventions, separates candidate-visible from oracle information, and enumerates detector contracts before policy optimization. On four Qwen3-8B pools from HotpotQA, 2WikiMultiHopQA, and MuSiQue, $R_0$ rejects search deletion and fake IDs, but a label-free citation-laundering attack succeeds. A complete $2^3$ ablation identifies targeted strengthening of $L$---citing a corpus passage absent from the retrieved evidence---as the observed inclusion-minimal repair: $R[L]$ has zero empirical ASR with a 0.50% one-sided cluster upper bound. The gap persists across pool rules, a visible BM25 attacker, and four models; broader hardening remains vulnerable when the attack removes an oracle support-ID penalty. Under strict 5M-token matched training evaluated on 256 paired questions per benchmark, $R[L]$ meets the EM non-inferiority gate on HotpotQA and 2Wiki but not MuSiQue. Equal-suite citation precision and support recall improve by 2.02 and 1.46 points, unsupported citations fall by 1.69, and laundering attackability falls on 2Wiki and MuSiQue. Natural $L$ is not reduced, and the detector appears in only 18 of 58,368 training trajectories. HERALD thus separates robust scoring, sparse learning signal, and policy transfer.
From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models
Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents…
グラフベースの複数インスタンス学習による暗黙的および明示的な関係バイアスの統合: 皮膚病変診断のケーススタディ
関係帰納バイアスは、データ間の構造的な依存関係を把握するために不可欠です。この研究では、暗黙的表現学習と明示的構造モデリングの間のギャップを埋める、画像分類のためのデュアルレベルのリレーショナル フレームワークを調査します。まず、EfficientNetB3 アーキテクチャを使用してベースラインを確立します。標準的な畳み込みバイアスを超えるために、パッチベースの戦略を採用し、畳み込みマスクされたオートエンコーダーを使用して、自己教師あり再構成を通じて暗黙的なパッチ間の関係を学習します。次に、明示的なリレーショナル モデリングを組み込むことでこのアプローチを拡張し、学習した埋め込みをグリッドベース、ランダム、k 最近傍構造などのさまざまなグラフ トポロジに編成します。 ISIC-2018 および ISIC-2019 皮膚病変診断ベンチマークの実験結果は、暗黙的なパッチ間モデリングと明示的なグラフベースのメッセージ パッシングを組み合わせることで最高のパフォーマンスが得られることを示しています。 ISIC-2018 テスト セットでは、ベースライン モデルは 76.17% のバランスのとれた精度を達成しましたが、暗黙的なパッチベースのリレーショナル モデリングでは 77.12% に向上しました。完全に統合されたグリッド構造のグラフ アテンション ネットワークにより、パフォーマンスがさらに 79.27% 向上します。同様に、ISIC-2019 では、陰的アプローチのバランスのとれた精度は 59.84% に達しますが、陰的モデリングと陽的モデリングの組み合わせでは 60.67% になります。
原文 (English)
Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance Learning: A Case Study in Skin Lesion Diagnosis
Relational inductive biases are essential for capturing structural dependencies among data. This study investigates a dual-level relational framework for image classification, bridging the gap between implicit representation learning and explicit structural modelling. We begin by establishing a baseline using an EfficientNetB3 architecture. To move beyond standard convolutional biases, we adopt a patch-based strategy, employing a convolutional masked autoencoder to learn implicit inter-patch relationships through self-supervised reconstruction. We then extend this approach by incorporating explicit relational modelling, organizing the learned embeddings into various graph topologies, including grid-based, random, and k-nearest neighbour structures. Experimental results on the ISIC-2018 and ISIC-2019 skin lesion diagnosis benchmarks show that combining implicit inter-patch modelling with explicit graph-based message passing yields the best performance. On the ISIC-2018 test set, the baseline model achieves a balanced accuracy of 76.17%, which improves to 77.12% with implicit patch-based relational modelling. The fully integrated grid-structured Graph Attention Network further increases performance to 79.27%. Similarly, on ISIC-2019, the implicit approach reaches 59.84% balanced accuracy, while the combination of implicit and explicit modelling yields 60.67%.
歴史が嘘をついているとき: 誤解を招く複数ターンの歴史の下でのツールの使用を評価し、改善する
ツール呼び出しエージェントは、蓄積されたダイアログとツール トレースからタスクの状態を推測します。ただし、永続的な対話では、履歴トレースが現在のリクエストに対する権限を失った後も、構造的に有効で意味論的に妥当なままになる可能性があります。我々は、そのような履歴がモデルがすでに持っているポリシーをハイジャックする可能性があることを示します。Qwen3-1.7B では、汚染により元の軌道では正しい決定の 32.1% が反転され、破損したエンティティやインターフェイス規則の再利用が頻繁に引き起こされます。システム ポリシー、現在のツール、最新のリクエスト、およびゴールドの次のアクションを保持する、同期された元のビュー、汚染されたビュー、および Oracle の状態ビューを備えたペアのベンチマークであるベンチを導入します。 11 のゴールド保持介入により、完全な呼び出しと非呼び出しの決定にわたる決定状態、エンティティ バインディング、およびインターフェイス実行の障害が分離されます。さらに、生徒が生成したプレフィックスに対するソフトな監視を通じて、汚染された履歴のみを観察する生徒に、Oracle 条件付き教師ポリシーを転送する私たちの提案を提案します。 Qwen3-1.7B では、バランスのとれたツール使用精度 87.0% を達成し、Gold-SFT (66.3%)、Oracle シーケンス蒸留 (82.3%)、およびオフポリシー トークン蒸留 (85.0%) を上回っています。この方法は一貫して拡張されており、80 億人の教師は同じ 17 億人のコンパクトな生徒の学力を 91.9% に引き上げ、80 億人の生徒は 93.0% に達します。結果として得られるポリシーは、クリーンな履歴、目に見えない機能、独立して再生成された評価コンテキスト、外部ツール使用ベンチマーク、およびノイズの多いマルチホップ質問応答にさらに移行します。これらの結果は、履歴の信頼性が明確なツール使用のボトルネックであることを確立し、効果的でスケーラブルなソリューションとして信頼できる状態のポリシー転送を実証しています。
原文 (English)
When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories
Tool-calling agents infer task state from accumulated dialogue and tool traces. In persistent interactions, however, historical traces may remain structurally valid and semantically plausible after they cease to be authoritative for the current request. We show that such history can hijack a policy the model already possesses: on Qwen3-1.7B, pollution flips 32.1% of decisions that are correct under the original trajectory and frequently induces reuse of corrupted entities or interface conventions. We introduce bench, a paired benchmark with synchronized Original, Polluted, and Oracle State views that preserve the system policy, current tools, latest request, and gold next action. Eleven gold-preserving interventions isolate failures in decision state, entity binding, and interface execution across complete calls and non-call decisions. We further propose ours, which transfers an Oracle-conditioned teacher policy to a student observing only polluted history through soft supervision on student-generated prefixes. On Qwen3-1.7B, ours achieves 87.0% Balanced Tool-Use Accuracy, outperforming Gold-SFT (66.3%), Oracle sequence distillation (82.3%), and off-policy token distillation (85.0%). The method scales consistently: an 8B teacher raises the same compact 1.7B student to 91.9%, while an 8B student reaches 93.0%. The resulting policies further transfer to clean histories, unseen functions, independently regenerated evaluation contexts, external tool-use benchmarks, and noisy multi-hop question answering. These results establish history reliability as a distinct tool-use bottleneck and demonstrate reliable-state policy transfer as an effective and scalable solution.
信号か偽のキューか? LLM 社会推論における調査国メタデータのランダム監査
調査国のメタデータは、有益な場合には LLM による個々の回答の予測を向上させることができますが、ランダムに割り当てられた場合には、同じ手がかりが予測をリダイレクトする可能性があります。記録内監査では、ランダムなラベルの統一的で記録に依存しない起源を開示することで、その国主導の摂取が減るかどうか、また、検証された調査国がホールドアウトされたブライヤー損失を減らすかどうかがテストされます。独立した人口アンカーと記録された人間の回答により、5 つの固定 API モデル、6 か国、および開発が選択した 7 つのターゲットにわたる方向性と結果が測定されます。主要なポストレビューの 72 レコードパネルでは、不透明ラベルと公開ランダムラベルはそれぞれ 0.214 の国方向シフトを生成しました。対の減衰は 0.0003 (95% CI [-0.0157, 0.0166]) でした。検証された国では、ブライエ損失が 0.040 (95% CI [0.024, 0.056]) 減少しましたが、ランダムラベルの後悔にはゼロが含まれていました。重複しない混合カバレッジの一貫性パネルは、ポジティブな開示されたランダムな動きと検証された有用性を維持しましたが、減衰は不確実なままでした。選択されたターゲットに関して、検証されたメタデータは両方のパネルで有用でしたが、開示によってランダムラベルの取り込みが確実に減衰するわけではありませんでした。 PROV-FORECAST には、修正されたパネルからの 14,400 個のペアになった項目レベルの確率分布が含まれています。
原文 (English)
Signal or Spurious Cue? A Randomized Audit of Survey-Country Metadata in LLM Social Inference
Survey-country metadata can improve an LLM's forecast of an individual response when informative, yet the same cue may redirect the forecast when assigned at random. A within-record audit tests whether disclosing a random label's uniform, record-independent origin reduces its country-directed uptake, and whether verified survey country lowers held-out Brier loss. Independent population anchors and recorded human answers measure direction and consequence across five fixed API models, six countries, and seven development-selected targets. In the primary post-review 72-record panel, opaque and disclosed-random labels each produced country-direction shifts of 0.214. Paired attenuation was 0.0003 (95% CI [-0.0157, 0.0166]). Verified country reduced Brier loss by 0.040 (95% CI [0.024, 0.056]), while random-label regret included zero. A non-overlapping mixed-coverage consistency panel retained positive disclosed-random movement and verified utility, while attenuation remained uncertain. On the selected targets, verified metadata was useful in both panels, but disclosure did not reliably attenuate random-label uptake. PROV-FORECAST contains 14,400 paired item-level probability distributions from the corrected panel.
大規模言語モデルでの投資ロジックの評価: パーソナライズされた金融エージェントに向けた現実世界のベンチマーク
投資能力は本質的に個別化されています。同じ市場証拠が、異なる目標、視野、ポートフォリオ、リスク境界を持つ投資家にとって異なる行動を正当化する可能性があります。しかし、財務 LLM は、静的な質問応答または最終的な損益によって評価されます。前者は主体性を省略します。後者では、利益をもたらす行動が根拠に基づいたものなのか、プロファイルに一貫性のあるものなのか、あるいは単に幸運だったのかを明らかにすることはできません。私たちは、コミュニティが結果的な要因に対して間違った物差しを使用しているかどうかを尋ねます。 \textsc{InvestLogicBench} は、151 人の現実世界の投資家からの 201,247 件の文書化された意思決定を含むプロセスネイティブのベンチマークです。各エピソードは \textbf{P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O} トレースをインスタンス化します: 投資家 \textit{プロフィール}、観察可能な市場 \textit{イベント}、投資 \textit{推論}、実行可能ファイル \textit{決定}、および遅延 \textit{結果}。このリリースには、プロファイルの構築、ポイントインタイムのイベント バインディング、構造化ロジック、ホライズン、結果、事後分析が含まれており、理解、プロファイル条件付き生成、エンドツーエンドの再生がサポートされています。 4 つの主要な LLM では、論理的妥当性は 4/5 近くを維持していますが、イベントグラウンディングはわずか 0.8 ~ 2.8/5 です。返品とプロセスの品質も一致しません。これらの結果は、結果のみの評価が隠れている、洗練されているものの根拠が弱い推論を明らかにします。さらに、P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O は、バージョン付きプロファイル、時間的来歴、検査可能な検索、意思決定台帳、および再生可能な結果を必要とするデータ システム インターフェイスであるべきだと主張します。財務は、より広範なクラスの個別化された結果的なエージェントに対するストレス テストです。
原文 (English)
Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents
Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries. Yet financial LLMs are evaluated either by static question answering or by terminal profit and loss. The former omits agency; the latter cannot reveal whether a profitable action was grounded, profile-consistent, or merely lucky. We ask whether the community is using the wrong ruler for consequential agents. We introduce \textsc{InvestLogicBench}, a process-native benchmark containing 201,247 documented decisions from 151 real-world investors. Each episode instantiates a \textbf{P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O} trace: investor \textit{Profile}, observable market \textit{Events}, investment \textit{Reasoning}, executable \textit{Decision}, and delayed \textit{Outcome}. The release includes profile construction, point-in-time event binding, structured logic, horizons, outcomes, and post-mortems, and supports comprehension, profile-conditioned generation, and end-to-end replay. Across four leading LLMs, logical plausibility remains near 4/5 while event grounding is only 0.8--2.8/5; return and process quality also disagree. These results expose polished but weakly grounded reasoning that outcome-only evaluation hides. We further argue that P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O should be a data-system interface, requiring versioned profiles, temporal provenance, inspectable retrieval, decision ledgers, and replayable outcomes. Finance is our stress test for a broader class of personalized, consequential agents.
ECHO: 一時記憶、安全ガードレール、音声評価を備えた、ローカルに導入可能なエージェント型健康アシスタント
この文書では、長期慢性期医療管理のための、ローカルに導入可能な会話型健康アシスタントである ECHO (Enhanced Care \& Health Observer) について紹介します。 ECHO は、統一システムとして共有監督の下で開発された 3 つの補完的なソフトウェア モジュールを統合します。コア モジュールは、LangGraph を介してオーケストレーションされた ReAct ループ上に構築されたエージェント チャットボットで、17 の臨床ツールと永続的なクロスセッション メモリのための時間的ナレッジ グラフを備えています。 GPT-5 Mini を使用した 59 シナリオのベンチマーク全体で 94.9\% のツール実行合格率を達成しました。 2 段階のハイブリッド安全層がすべての受信クエリを傍受します。ルールベースの層は明示的な危機信号とジェイルブレイクの試行を 1 ミリ秒未満で処理し、APPNP スタイルの伝播を備えた符号付きグラフ ニューラル ネットワーク (GNN) は臨床目的によって境界ケースを分類し、2,537 クエリの注釈付きトルコ健康データセットで 88.8% の精度と 90.6% の安全でないリコールを達成しながら、ゼロショット LLM ベースラインを上回るパフォーマンスを実現します。ラマ 3.3 70B を含む。 Whisper 音響エンコーディングと BERT テキスト エンコーディングをクロスアテンション フュージョンと組み合わせたマルチモーダル音声評価モジュールは、感情、憂鬱、痛みを推定し、平均マクロ F1 が 0.652 に達します。完全なシステムは、消費者向けハードウェア上で完全に実行できる Web アプリケーションとして実装されており、患者データは外部サービスに送信されず、GDPR および KVKK への準拠をサポートします。
原文 (English)
ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment
This paper presents ECHO (Enhanced Care \& Health Observer), a locally-deployable conversational health assistant for long-term chronic care management. ECHO integrates three complementary software modules developed under shared supervision as a unified system. The core module is an agentic chatbot built on a ReAct loop orchestrated via LangGraph, equipped with 17 clinical tools and a temporal knowledge graph for persistent cross-session memory; it achieves a 94.9\% tool-execution pass rate across a 59-scenario benchmark with GPT-5 Mini. A two-stage hybrid safety layer intercepts all incoming queries: a rule-based layer handles explicit crisis signals and jailbreak attempts in under 1ms, while a signed graph neural network (GNN) with APPNP-style propagation classifies boundary cases by clinical intent, achieving 88.8\% accuracy and 90.6\% unsafe recall on a 2,537-query annotated Turkish health dataset while outperforming zero-shot LLM baselines including Llama 3.3 70B. A multimodal speech assessment module combining Whisper acoustic encoding and BERT text encoding with cross-attention fusion estimates emotion, depression, and pain, reaching a mean macro F1 of 0.652. The full system is implemented as a web application that can run entirely on consumer hardware, with no patient data transmitted to external services, supporting compliance with GDPR and KVKK.
サイロ化されたアルゴリズムからコンプライアンス第一のエージェント プラットフォームまで: 病院 AI システムの多層アーキテクチャ
病院はトリアージ、画像処理、スケジュール設定などに人工知能を急速に導入していますが、ほとんどの導入は部門サイロ内に閉じ込められた孤立点ソリューションのままであり、その結果、重複した作業、隠れたリスク、未実現の企業価値が生じています。ヘルスケア市場における AI の爆発的な成長と投資の加速にも関わらず、ヘルスケア AI パイロットの推定 70 ~ 80% は拡張できず、その主な原因はガバナンスのギャップ、データの断片化、統合ブループリントの欠如です。この研究では、複数の相互運用可能なレイヤーを備えた病院固有のコンプライアンス優先の Agentic AI アーキテクチャを提案しています。(i) 臨床、業務、財務ドメインにわたるマルチエージェントのワークフローのためのエージェント オーケストレーション レイヤー、(ii) HIPAA、GDPR、EU AI 法、DISHA 法、インドの DPDP 法、および ISO/IEC のセキュリティと安全基準のコードとしてのポリシーを一元化するコンプライアンスおよびポリシー レイヤーで、既存の病院 AI プラットフォーム モデルを拡張します。 (iii) フェデレーション ラーニング、差分プライバシー、安全なエンクレーブを現実世界の病院情報管理システム (HIMS) フローに組み込むプライバシー保護データ ファブリック。この研究では、合成ではあるが構造的に現実的な病院データセットと、オープンですぐに展開できるプロトタイプの実装を使用して、トリアージのリスク予測、ワークフローの最適化、コンプライアンスログのエンドツーエンドのオーケストレーションを実証し、ポリシーで保護されたデータアクセスを維持しながら、タスクのターンアラウンドタイムと手動の文書化作業の大幅な削減をシミュレーションで達成しています。結果として得られるアーキテクチャは、病院のリーダーに、アドホックなツールから、オンプレミス、ハイブリッド、クラウドネイティブの展開に合わせて調整できる、管理され、グローバルに準拠し、ROI を重視した AI プラットフォームに移行するための実用的な青写真を提供します。
原文 (English)
From Siloed Algorithms to Compliance-First Agentic Platforms: A Multi-Layered Architecture for Hospital AI Systems
Hospitals are rapidly adopting artificial intelligence for triage, imaging, scheduling etc., yet most deployments remain isolated point solutions locked inside departmental silos, resulting in duplicated effort, hidden risks, and unrealized enterprise value. Despite explosive growth of AI in healthcare market and accelerating investment, an estimated 70-80% of healthcare AI pilots fail to scale, largely due to governance gaps, fragmented data, and missing integration blueprints. This research proposes a hospital-specific, compliance-first, Agentic AI architecture with multiple interoperable layers, extending existing hospital AI platform models with: (i) an Agent Orchestration Layer for multi-agent workflows across clinical, operational, and financial domains, (ii) a Compliance and Policy Layer that centralizes policy-as-code for HIPAA, GDPR, the EU AI Act, DISHA Act, India's DPDP Act, and ISO/IEC security and safety standards, and (iii) a Privacy-Preserving Data Fabric that plugs federated learning, differential privacy, and secure enclaves into real-world Hospital Information Management System (HIMS) flows. Using a synthetic but structurally realistic hospital dataset and an open, ready-to-deploy prototype implementation, this study demonstrates the end-to-end orchestration of triage risk prediction, workflow optimization, and compliance logging, achieving substantial simulated reductions in task turnaround times and manual documentation effort while maintaining policy-guarded data access. The resulting architecture offers hospital leaders a pragmatic blueprint to move from ad hoc tools to a governed, globally compliant, ROI-focused AI platform that can be tailored to on-premise, hybrid and cloud-native deployments.
ギャップに注意: 人間シミュレーションのためのさまざまな考え方
集団が新しい質問にどのように答えるかを予測することは、長年の目標です。統計的手法は集団のレベルでは成功しますが、個人のレベルでは失敗します。大規模な言語モデル シミュレーターはこのギャップを継承します。それらは、集団の不均一性を平坦化しながら集団の中心的な傾向を回復し、個人の予測を歪める社会的偏見と脆弱性をもたらします。この文書では、狭く明確に指定された領域内の個人レベルを対象とする視聴者シミュレーション モデルである Anacreon を紹介します。 Anacreon は、個人を分離する著者権限の埋め込みを学習し、シードの人々を中心に実際の定性的コーパスをクラスタリングし、Gemma~4 12B ベースで心の混合であるクラスタごとに専用のアダプターをトレーニングします。公開テキストから人口統計、心理的特徴、アンケートの回答を収集し、感情の連鎖で各記録を強化します。応答オプションをシャッフルすることでプロンプトの脆弱性を軽減し、トレーニング分布のバランスをとることでポジティブなバイアスを軽減します。外部ソースによる大規模な調査で、アナクレオンは、わずかな残留バイアスを伴いながら、フィールドが収束した個人レベルの精度尺度である 0.775 という最先端の順序アライメントに達しました。この研究は、忠実にシミュレートされた個人から総合的な洞察を引き出すためのステップです。
原文 (English)
Mind the Gaps: Mixture-of-Minds for Human Simulation
Predicting how a population will answer a new question is a long-standing goal. Statistical methods succeed at the level of the mass but falter at the level of the individual. Large language model simulators inherit this gap. They recover a population's central tendencies while flattening its heterogeneity, and they carry social biases and prompt brittleness that distort individual predictions. This paper introduces Anacreon, an audience simulation model that targets the individual level within a narrow, well-specified domain. Anacreon learns an authorship embedding that separates individuals, clusters a real qualitative corpus around seed people, and trains a dedicated adapter for each cluster, a mixture of minds, on a Gemma~4 12B base. It harvests demographics, psychological traits, and survey responses from public text, and augments each record with a chain-of-emotion. It reduces prompt brittleness by shuffling response options and reduces positive bias by balancing the training distribution. On a large, externally sourced survey, Anacreon reaches a state-of-the-art ordinal alignment of 0.775, the individual-level accuracy measure on which the field has converged, with a small residual bias. The work is a step toward drawing aggregate insight from faithfully simulated individuals.
ポリバイアス: 国際政治紛争における大規模言語モデルのバイアスの理解と測定
大規模言語モデル (LLM) での政治的偏見の測定は、単一の指標で把握するのが難しい枠組み、議論、法的推論の微妙な違いを通じて現れる可能性があるため、依然として困難です。この研究では、LLM が法的に同等の紛争シナリオを関係国に応じて異なる方法で扱うかどうかを測定するための反事実フレームワークである Poli-Bias を紹介します。 Poli-Bias は、さまざまな地政学的関係、法律違反、および推論タスクにわたって国のアイデンティティが体系的に交換されるペアのプロンプトに対する応答を比較します。私たちのフレームワークは、バイアスを 1 つの判断に限定するのではなく、回答の格差を解釈可能な 5 つの側面に分解し、不平等な扱いがどこでどのように現れるかを明らかにします。さまざまなモデルファミリーと規模にまたがる 13 の現代の LLM にわたって、国のアイデンティティとユーザーの所属が、同等の行為が国際法の下でどのように記述され、評価され、擁護されるかに系統的に影響を与える可能性があることがわかりました。したがって、私たちの結果は、LLM における政治的平等性と媚びを監査するためのきめ細かいフレームワークとして Poli-Bias を確立しました。
原文 (English)
Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts
Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argumentation, and legal reasoning that are difficult to capture with a single metric. In this work, we introduce Poli-Bias, a counterfactual framework for measuring whether LLMs treat legally equivalent conflict scenarios differently depending on the countries involved. Poli-Bias compares responses to paired prompts in which country identities are systematically swapped across diverse geopolitical relationships, legal violations, and reasoning tasks. Rather than reducing bias to a single judgment, our framework decomposes response disparities into five interpretable dimensions, revealing how and where unequal treatment manifests. Across 13 contemporary LLMs spanning diverse model families and sizes, we find that country identities and user affiliations can systematically affect how equivalent actions are described, evaluated, and defended under international law. Our results thus establish Poli-Bias as a fine-grained framework for auditing political even-handedness and sycophancy in LLMs.
検索エージェント向けのコンテキスト情報ポリシーの最適化
検索エージェントは、複数ステップの推論中に外部証拠を取得して使用できるようにすることで、静的パラメトリック メモリを超えて大規模な言語モデルを拡張します。複雑な情報や進化する情報を伴う知識集約型タスクの場合、その信頼性は、関連する証拠を取得するだけでなく、それをその後の推論の指針として使用することにも依存します。ただし、既存の方法では、検索後のアクションが検索された証拠に基づいているかどうかを直接評価することなく、主に最終的な回答の正確性または中間の進歩に報酬を与えます。この不整合により、事前主導型の推論が促進されます。エージェントは内部知識に基づいて結論を出し、主にそれを確認するために検索を使用するため、確証バイアスと非効率的な証拠の使用が生じます。この問題に対処するために、ポリシーの最適化と外部の証拠の使用を明示的に調整する証拠指向の強化学習フレームワークであるコンテキスト情報ポリシー最適化 (CIPO) を提案します。 CIPO は、取得した情報の影響を受ける推論アクションに密なターンレベルのクレジットを割り当てますが、この証拠使用シグナルと、回答の正しさを維持するためのグローバルな結果報酬を組み合わせます。この方法により、CIPO は証拠から切り離された推測を阻止し、取得した事実がその後の推論を導き、修正できる推論の軌道を促進します。重要なのは、CIPO では人間のプロセス アノテーションも追加の報酬モデルも必要ないことです。 7 つのドメイン内およびドメイン外のベンチマークに関する広範な実験により、CIPO が事前駆動推論の蔓延を減らし、ほとんどのタスクで優れたパフォーマンスを達成することが示されています。
原文 (English)
Contextual Information Policy Optimization for Search Agents
Search agents extend large language models beyond static parametric memory by enabling them to acquire and use ex ternal evidence during multi-step reasoning. For knowledge intensive tasks involving complex or evolving information, their reliability depends not only on retrieving relevant ev idence but also on using it to guide subsequent reasoning. However, existing methods primarily reward final-answer cor rectness or intermediate progress, without directly assessing whether post-retrieval actions are grounded in the retrieved evidence. This misalignment encourages prior-driven reason ing: agents form conclusions based on internal knowledge and use retrieval mainly to confirm them, resulting in confirma tion bias and inefficient evidenceuse.Toaddressthisissue, we propose Contextual Information Policy Optimization (CIPO), an evidence-oriented reinforcement learning framework that explicitly aligns policy optimization with external evidence use. CIPO assigns dense, turn-level credit to reasoning ac tions influenced by retrieved information, while combining this evidence-use signal with a global outcome reward to pre serveanswercorrectness.Withthismanner,CIPOdiscourages evidence-detached guesses and promotes reasoning trajecto ries in which retrieved facts can guide or revise subsequent reasoning. Importantly, CIPO requires neither human process annotations nor an additional reward model. Extensive exper iments on seven in-domain and out-of-domain benchmarks show that CIPO reduces the prevalence of prior-driven rea soning and achieves excellent performance on most tasks.
FinEvo-Bench: プロフェッショナルな金融ワークフローにおける自己進化するエージェントのための長期的なベンチマーク
ほとんどのエージェント ベンチマークはタスクを個別に評価するため、あるタスクの経験が後のタスクに役立つかどうかを測定できません。既存の自己進化ベンチマークは、プロフェッショナルなワークフロー、無制限の成果物、および複数の側面からの評価を統合してカバーしていません。 6 つの金融ドメインにわたる 120 の実際のケースに基づいたタスク、20 のビジネス シーンを備えた長期的なベンチマークである FinEvo-Bench を紹介します。教育機関が提供する専門的な手順により、必要な操作と制約が定義されます。適格な機関が提供し、公的に文書化されたケースにより、タスクの事実が提供されます。各シーンには、専門的な手順と、タスクの品質と財務コンプライアンスを手動でレビューするルーブリックを共有する、関連しているが実質的に異なる 6 つのケースが含まれています。同じ Qwen3.7-Max バックボーンと、独立してシャッフルされ、グローバルにインターリーブされた 3 つのタスク ストリームを使用する 4 つの自己進化エージェント スキャフォールドを比較します。ペアの非進化コントロールは、保持された経験から各足場の自己進化ゲインを推定し、一方、Claude Opus 4.6 に支援された独立した Claude Code スコアリング エージェントがすべての出力を評価します。 Letta は最高の進化スコア (91.65) と最も少ないコンプライアンス問題 (タスクあたり 0.09) を達成しました。 Codex は最大の自己進化ゲイン (+19.37) を達成しました。足場全体で、状態が進化するとスコアが 9.33 ~ 19.37 ポイント上昇し、コンプライアンスの問題がタスクごとに 0.12 ~ 0.44 ポイント減少します。シーン内ランク 4 ~ 6 のペアのスコア増加は、ランク 1 ~ 3 のスコア増加を 6.10 ~ 8.70 ポイント上回っています。 Claude Code では、スキルのみの進化は、メモリのみおよびメモリとスキルの組み合わせの進化よりもタスクの品質が高く、コンプライアンスの問題が少なくなります。 4 つの足場すべてにおいて、ルーブリック フィードバックは、参照回答フィードバックよりも高いスコアをもたらし、コンプライアンス上の問題が少なくなります。 FinEvo-Bench は、プロフェッショナルとしてのパフォーマンスと自己進化能力の両方を測定します。つまり、エージェントが以前の経験をどれだけ効果的に後の改善に変えることができるかということです。
原文 (English)
FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows
Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks. Existing self-evolution benchmarks do not jointly cover professional workflows, open-ended deliverables, and multi-aspect evaluation. We introduce FinEvo-Bench, a longitudinal benchmark with 120 real-case-grounded tasks, 20 business scenes across six financial domains. Institution-provided professional procedures define the required operations and constraints. Eligible institution-provided and publicly documented cases supply the task facts. Each scene contains six related but substantively distinct cases that share a professional procedure and a manually reviewed rubric for task quality and financial compliance. We compare four self-evolving agent scaffolds using the same Qwen3.7-Max backbone and three independently shuffled, globally interleaved task streams. Paired non-evolving controls estimate each scaffold's self-evolution gain from retained experience, while an independent Claude Code scoring agent backed by Claude Opus 4.6 evaluates all outputs. Letta achieves the highest evolved score (91.65) and fewest compliance issues (0.09 per task); Codex achieves the largest self-evolution gain (+19.37). Across scaffolds, the evolving condition raises scores by 9.33-19.37 points and reduces compliance issues by 0.12-0.44 per task. Paired score gains at within-scene ranks 4-6 exceed those at ranks 1-3 by 6.10-8.70 points. In Claude Code, skill-only evolution produces higher task quality and fewer compliance issues than memory-only and combined memory-skill evolution. Across all four scaffolds, rubric feedback also yields higher scores and fewer compliance issues than reference-answer feedback. FinEvo-Bench measures both professional performance and self-evolution ability: how effectively an agent turns prior experience into later improvement.
PaDoc: ドキュメント解析のためのレイアウトに基づいた並列デコード
エンドツーエンドのドキュメント パーサーは、統一されたインターフェイスを提供しますが、ページ レイアウトと地域のコンテンツを 1 つの自己回帰シーケンスにシリアル化します。この定式化により、独立した領域がコンテンツ全体に応じて長くなるデコード パスに強制的に配置されますが、クロップベースの 2 段階パーサーでは、繰り返しの視覚的なプレフィルと断片化されたページ コンテキストを犠牲にして、領域レベルの並列処理が公開されます。依存関係を削除しながらページ全体のコンテキストを保持するために、予測されたレイアウトを共有ページ表現上の分岐構造として扱うレイアウトベースのパーサーである PaDoc を提案します。領域十分性の仮定の下で、レイアウト ストリームと地域コンテンツ ブランチが同時に進行し、デコード深度を最長のレイアウト コンテンツ パスまで減らすプレフィックス条件付き因数分解を導出します。この因数分解は単一の MLLM 内で実現されます。パックされた可変長の祖先アテンションは、標準のネクスト トークン トレーニングでの可視性を維持しますが、マスクされた並列デコードは、評価された vLLM バックエンドがキャッシュ常駐の共有プレフィックスの再利用による同時リクエストとして機能するブランチを作成します。 OmniDocBench Full では、PaDoc は全体レイアウト F1 91.1 を達成し、エンドツーエンド パーサーの中で最高レベルの全体スコア 94.24 を達成し、テキスト編集 (0.038) とフォーミュラ CDM (95.59) も最高でした。 384 ページのサブセットと 1 つの A800 GPU 上で、5 つの同時実行レベルで最速のエンドツーエンド パーサーであり、同じバックボーンのシーケンシャル SFT ベースラインと比較して有効ページ スループットを 67.4 ~ 118% 向上させ、P95 レイテンシを 39.2 ~ 54.9% 削減します。コードは https://github.com/Longin-Yu/Padoc で入手できます。
原文 (English)
PaDoc: Layout-Grounded Parallel Decoding for Document Parsing
End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regions onto a decoding path whose length grows with the total content, whereas crop-based two-stage parsers expose region-level parallelism at the cost of repeated visual prefills and fragmented page context. To retain full-page context while removing dependencies, we propose PaDoc, a layout-grounded parser that treats the predicted layout as a branching structure over a shared page representation. Under a region-sufficiency assumption, we derive a prefix-conditioned factorization in which the layout stream and regional content branches advance concurrently, reducing the decoding depth to the longest layout-content path. We realize this factorization within a single MLLM: packed variable-length ancestor attention preserves the visibility under standard next-token training, while masked parallel decoding creates branches that the evaluated vLLM backend serves as concurrent requests with cache-resident shared-prefix reuse. On OmniDocBench Full, PaDoc attains an Overall layout F1 of 91.1 and, among end-to-end parsers, a top-tier Overall score of 94.24 together with the best Text Edit (0.038) and Formula CDM (95.59). On a 384-page subset and one A800 GPU, it is the fastest end-to-end parser at five concurrency levels, improving valid-page throughput by 67.4-118% and reducing P95 latency by 39.2-54.9% relative to a same-backbone Sequential SFT baseline. Code is available at https://github.com/Longin-Yu/Padoc
CogVis: オープン語彙変化検出はクエリごとにシーンを新たに認識する必要があるか?
地表面の監視には、任意の意味カテゴリを認識できる変化検出モデルが必要です。 Open-Vocabulary Change Detection (OVCD) は、このニーズに対応します。しかし、既存の手法では、時間的な認識、意味の識別、領域の検証が絡み合うことが多く、結果が不安定になり、計算が冗長になります。人間の視覚変化の知覚に触発されて、OVCD を知覚-記憶-検証パラダイムとして再定式化する認知記憶誘導フレームワークである CogVis を提案します。 CogVis はまず、シーン チェンジ パーセプトロン (SCP) を使用して、凍結されたバイタイム特徴から事前に再利用可能なカテゴリに依存しない変更を抽出します。これにより、時間的証拠が意味論的なカテゴリの決定から切り離されます。次に、セマンティック メモリ キャリブレーター (SMC) が、画像クエリ固有の決定しきい値を動的に推定することで、カテゴリに依存するスコアのシフトを補正します。最後に、適応領域フィルター (ARF) が、学習された意味論的、時間的、構造的な信頼性を使用して、接続された候補をフィルター処理します。意味的変更の検出、バイナリ変更の位置特定、建物の損傷評価にわたる 7 つのベンチマークに関する実験では、CogVis が評価されたすべてのデータセットにわたって最先端のパフォーマンスを達成していることが示されています。シーンレベルの変化認識を共有することで、CogVis はクエリ間でカテゴリに依存しない時間認識の繰り返しをさらに回避し、推論のスループットを 28.50% 向上させます。
原文 (English)
CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?
Earth-surface monitoring requires change detection models capable of recognizing arbitrary semantic categories. Open-Vocabulary Change Detection (OVCD) addresses this need. However, existing methods often entangle temporal perception, semantic discrimination, and region verification, causing unstable results and redundant computation. Inspired by human visual change perception, we propose CogVis, a cognitive memory-guided framework that reformulates OVCD as a perception-memory-verification paradigm. CogVis first employs a Scene Change Perceptron (SCP) to extract a reusable, category-agnostic change prior from frozen bi-temporal features, thereby decoupling temporal evidence from semantic category decisions. A Semantic Memory Calibrator (SMC) then compensates for category-dependent score shifts by dynamically estimating an image-query-specific decision threshold. Finally, an Adaptive Region Filter (ARF) filters connected candidates using learned semantic, temporal, and structural reliability. Experiments on seven benchmarks spanning semantic change detection, binary change localization, and building-damage assessment show that CogVis achieves state-of-the-art performance across all evaluated datasets. By sharing scene-level change perception, CogVis further avoids repeating category-agnostic temporal perception across queries and improves inference throughput by 28.50%.
iARCS: 制御可能な 3D シーン生成のための反復エージェント RL
合成 3D シーンの生成は、コンピューター ビジョンや具体化された AI のデータ ソースとしてますます使用されていますが、既存のジェネレーターは、タスクに不可欠な機能上の制約を確実に満たさずに、知覚的なリアリズムを最適化することがよくあります。この不一致により、アクセシビリティ、トラバーサビリティ、空間ルールへの準拠がしばしば重要となる下流トレーニングでの合成データの有用性が制限されます。我々は、事前訓練されたシーンジェネレーターを自然言語タスクの要件に適応させる反復エージェント強化学習フレームワークである iARCS を紹介します。 iARCS は 2 段階の戦略を使用します。つまり、物理的な妥当性とレイアウトの品質を向上させるための普遍的な報酬の事前トレーニングと、それに続く、トレーニング フィードバックから繰り返し改良される LLM 生成の報酬プログラムによるタスク固有の微調整です。実験では、歩きやすさ、到達しやすさ、クリアランスを重視したタスクにおける制約の忠実度の向上、タスク固有の制約の効果的な最適化、および競技シーンの多様性が示されています。さらに、iARCS によって生成されたデータがベース ジェネレーターを改善し、制御可能なシーン編集方法だけでなく実用的な合成データ生成ツールとしてのその価値をサポートすることを示します。
原文 (English)
iARCS: Iterative Agentic RL for Controllable 3D Scene Generation
Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This mismatch limits the usefulness of synthetic data for downstream training, where accessibility, traversability, and spatial rule compliance are often essential. We present iARCS, an iterative agentic reinforcement learning framework that adapts a pretrained scene generator to natural-language task requirements. iARCS uses a two-stage strategy: universal-reward pretraining to improve physical plausibility and layout quality, followed by task-specific fine-tuning with LLM-generated reward programs that are iteratively refined from training feedback. Experiments show improved constraint fidelity on walkability, reachability, and clearance-focused tasks, effective task-specific constraint optimization, and competitive scene diversity. We further show that data generated by iARCS improves a base generator, supporting its value as a practical synthetic data generation tool rather than only a controllable scene editing method.
生成 AI を使用したスキーマに基づく階層情報の抽出と意味評価
私たちは、生成 AI を使用して非構造化テキスト文書から複雑な構造化情報を抽出し、抽出された情報をゴールドスタンダードに照らして自動的にセマンティック評価するためのスキーマベースのフレームワークを紹介します。このスキーマは、ドメイン知識をエンコードする情報モデルとして機能し、変数カーディナリティの属性を備えた階層的で入れ子になった情報の抽出と、その後の結果の評価のための統一的で体系的かつ一貫したフレームワークを提供します。ドキュメントからの情報抽出は、ゼロショット モードでのモデルへの 1 回の呼び出しで実行されます。評価ステップでは、パスベースのセマンティック マッチング アルゴリズムを導入して、抽出結果内のネストされた変数カーディナリティ属性をゴールド スタンダードの属性と一致させます。当社では、抽出された属性値とゴールドスタンダード値の意味的な比較に生成 AI を使用し、ドメイン固有の考慮事項に従って、比較の結果を完全一致、意味論的、有用、または不一致として分類するためのルーブリックを導入しています。生成 AI モデル Claude Opus 3 を使用して、保健技術評価機関 NICE が発行した文書から、$>$90\% の F1 スコアを持つ 14 属性のうち 12 属性を抽出することができました。文書から属性を抽出するのに必要な時間は、人間のドメイン専門家が要する時間より $\sim$30 分の 1 でした。さらに、さまざまな生成 AI モデルにわたるこのフレームワークの汎用性と、さまざまな HTA 組織や言語にわたる移行可能性を実証します。
原文 (English)
Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI
We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold standard. The schema, serving as an information model encoding domain knowledge, provides a unified, systematic, and consistent framework for extraction of hierarchical, nested information, with attributes of variable cardinality, and subsequent evaluation of the results. Information extraction from a document is performed in a single call to the model, in zero-shot mode. In the evaluation step, we introduce a path-based semantic matching algorithm to align the nested, variable-cardinality attributes in the extracted results with those in the gold standard. We use generative AI for semantic comparison of the extracted and gold standard values of an attribute, and introduce a rubric to classify the result of the comparison, according to domain-specific considerations, as an exact, semantic, useful, or non-match. We were able to extract 12 out of 14 attributes with an F1 score of $>$90\% from documents published by the health technology assessment organisation NICE, using the generative AI model Claude Opus 3. The time needed to extract the attributes from a document was $\sim$30 times lower than the time taken by a human domain expert. We further demonstrate generalisability of this framework across different generative AI models and transferability across different HTA organisations and languages.
MicroEvo: 効率的なマイクロアーキテクチャ設計空間探索のための知識に基づく LLM サンプリング
マイクロアーキテクチャの設計空間の探索は、広大な探索空間と高価な PPA 評価に悩まされ、設計の意思決定に使用できるシミュレーション予算はわずかしかありません。既存の方法は、マイクロアーキテクチャの依存関係を考慮せずにブラインド検索を実行し、反復検索から効果的に学習できず、無駄な評価と弱いパレート収束につながります。この論文では、多目的マイクロアーキテクチャの最適化のために既製の LLM とモンテカルロ ツリー検索 (MCTS) を結合する知識ガイド型フレームワークである MicroEvo を提案します。 MicroEvo は、LLM 主導の進化演算子、パレート寄与と多様性のバランスを取るパレート認識ツリー ポリシー、最適化の洞察を抽出して再利用するアクティブな知識蓄積メカニズム、オンラインでの検索動作を適応させる状態認識ディレクティブを組み合わせています。実験の結果、MicroEvo は NSGA-II と比較してパレート フロントの品質を最大 36.2% 向上させ、10.6 倍の高い検索効率を達成し、複雑な産業規模のコアに対する強力な拡張性も実証したことが示されています。コード リポジトリは、https://github.com/GEAR-SEU/MicroEvo-ICCAD-26 から入手できます。
原文 (English)
MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Exploration
Microarchitecture design space exploration suffers from expansive search spaces and expensive PPA evaluation, leaving only a small simulation budget for design decision-making. Existing methods perform blind search without considering microarchitectural dependencies and fail to learn from the iterative search effectively, leading to wasted evaluations and weak Pareto convergence. In this paper, we propose MicroEvo, a knowledge-guided framework that couples off-the-shelf LLMs with Monte Carlo Tree Search (MCTS) for multi-objective microarchitecture optimization. MicroEvo combines LLM-driven evolutionary operators, a Pareto-aware tree policy that balances Pareto contribution and diversity, an active knowledge accumulation mechanism that extracts and reuses optimization insights, and state-aware directives that adapt the search behavior online. Experiments show that MicroEvo improves Pareto-front quality by up to 36.2% over NSGA-II and achieves 10.6x higher search efficiency, and also demonstrates strong scalability to a complex industrial-scale core. The code repository is available at: https://github.com/GEAR-SEU/MicroEvo-ICCAD-26.
大規模なスキル ライブラリにおけるエージェント検索の比較アプローチ
大規模なスキル ライブラリに支えられたエージェントは、どのスキルをどの順序でロードするかを決定する必要があります。ライブラリ全体をコンテキストにロードするとコストがかかり、自律的なシーケンスのための構造が提供されません。私たちは、690 のスキルのコーパスを対象に、この問題に対して 2 つのシステムを研究しました。1 つはスパースなオンデマンド読み込みのための語彙検索と密埋め込み検索を組み合わせたハイブリッド ランカー、もう 1 つは前提条件、データ フロー、順序付けなどのワークフロー関係をエンコードする型付きナレッジ グラフです。 117 個の現実的な非エコー クエリのセットに対して、ハイブリッド ランカーは 73.5% +/- 8.0 のケースで上位 5 位以内の正しいスキルを取得し、クエリの約 4 分の 1 が処理されないままになります。意図した設計どおりに使用すると (一致したトークン バジェットで追加のランク付けされた結果をグラフの近傍に置き換える)、グラフは大幅に悪化します (-11.2 ポイント、p = 0.0007)。 LLM によって生成されたエッジ層は、ローカル エンベディング パスから自由に取得された近傍に何も追加しません。また、ランカーがミスしたクエリの 73% には、グラフを通じてまったく到達できません。これは、プレフィルター トポロジの境界に起因すると考えられます。グラフの候補エッジは、ランカーがすでに検索しているのと同じ埋め込み近傍から描画されているため、型付けされたエッジの 98.6% は、ランカーがすでに表面化しているスキルを結び付けます。グラフは関係セマンティクスを強化できますが、検索範囲を拡張することはできません。さらに、作成者が作成したクエリを評価すると、hit@5 が最大 44 ポイント過大評価され、これらの結果が完全に隠蔽されてしまうことを示します。私たちの貢献は、構造を追加しても強力なランカーよりも検索が向上しない理由を機構的に説明し、検索に構造の相互依存性を追加することが最適である条件を特定することです。
原文 (English)
Comparative Approaches to Agent Retrieval over Large Skill Libraries
Agents backed by large skill libraries must decide which skills to load and in what order. Loading the entire library into context is expensive and provides no structure for autonomous sequencing. We study two systems for this problem over a corpus of 690 skills: a hybrid ranker combining lexical and dense-embedding retrieval for sparse, on-demand loading, and a typed knowledge graph encoding workflow relations such as prerequisites, data flow, and ordering. On a set of 117 realistic, non-echoing queries, the hybrid ranker retrieves the correct skill within the top five in 73.5% +/- 8.0 of cases, leaving roughly a quarter of queries unserved. When used as the design intended (substituting graph neighbours for additional ranked results at matched token budget), the graph is significantly worse (-11.2 points, p = 0.0007). Its LLM-generated edge layer adds nothing over neighbours obtained free from a local embedding pass, and 73% of the queries the ranker misses are not reachable through the graph at all. We attribute this to a pre-filter topology bound. Because the graph's candidate edges are drawn from the same embedding neighbourhood the ranker already searches, 98.6% of typed edges connect skills the ranker had already surfaced together. The graph can enrich relation semantics but cannot extend retrieval reach. We further show that evaluating on author-written queries overstates hit@5 by up to 44 points, which would have hidden these results entirely. Our contribution is a mechanistic account of why added structure does not improve retrieval over a strong ranker, and identify the conditions under which adding structural interdependence into the retrieval is optimal.
EnvACE: エージェント的強化学習の世界リハーサルによる環境ダイナミクスの内部化
長期的なツールの使用に向けた大規模な言語モデル エージェントのトレーニングは、通常、構築と検証にコストがかかる実際の実行可能環境または合成された実行可能環境との対話、または接地が困難な外部シミュレータに依存します。トレーニング中の外部環境のインタラクションをワールド リハーサルに置き換えるエージェント強化学習手法である EnvACE を紹介します。このポリシーは、演技とリハーサルを交互に行います。最初にツール呼び出しを生成し、次に環境の役割を果たしてそのアクションによって引き起こされる応答を生成し、リハーサルされた応答に基づいてその後の決定を条件付けします。両方の役割は、タスク成功報酬を使用してエンドツーエンドで共同して最適化されます。ワールド リハーサルを通じて、ポリシーはアクションとその環境応答の間の関係をパラメーターに内面化し、意思決定を直接サポートするエージェント ワールド モデルを生成します。 BFCL-v4、tau^2-Bench、VitaBench、および FinMCP-Bench 全体で、EnvACE は強力で移行可能なパフォーマンスを実現し、全体的な評価で環境スケーリングのベースラインを上回っています。さらに、対照研究では、世界リハーサルがモデル規模全体で政策学習を一貫して改善していることが示されています。テスト時には、内部化されたワールド モデルにより、コミットされた実行前のプライベート リハーサルが可能になり、追加の外部インタラクションなしで適度なリハーサル予算の下でさらなる利益が得られます。私たちの調査結果は、外部環境の制約を超えて LLM エージェントのトレーニングを拡張するための新しい道としてワールド リハーサルを確立しました。私たちのコードは https://github.com/Within-yao/EnvACE で公開されています。
原文 (English)
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning
Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators that are difficult to ground. We introduce EnvACE, an agentic reinforcement learning method that replaces external environment interaction during training with world rehearsal. The policy alternates between acting and rehearsal: it first generates a tool call, then plays the role of the environment to produce the response induced by that action, and conditions subsequent decisions on the rehearsed response. Both roles are jointly optimized end-to-end using task-success rewards. Through world rehearsal, the policy internalizes the relationship between actions and their environment responses in its parameters, yielding an agent world model that directly supports decision making. Across BFCL-v4, tau^2-Bench, VitaBench, and FinMCP-Bench, EnvACE achieves strong and transferable performance, outperforming environment-scaling baselines in the overall evaluation. Controlled studies further show that world rehearsal consistently improves policy learning across model scales. At test time, the internalized world model enables private rehearsal before committed execution, yielding further gains under a moderate rehearsal budget without additional external interaction. Our findings establish world rehearsal as a new path toward scaling LLM agent training beyond the constraints of external environments. Our code is publicly available at https://github.com/Within-yao/EnvACE.
TS-RAG: 時系列予測のための検索拡張生成
ディープ ラーニング モデル、特にトランスフォーマー ベースのアーキテクチャは、時系列予測において優れたパフォーマンスを示していますが、この分野での検索拡張生成 (RAG) の適用は依然として限られています。 RAG は、関連する外部情報を組み込むことで大規模な言語モデルの機能を強化するのに効果的であることが証明されているため、同様の時系列シーケンスを参照として取得することにより、時系列予測タスクの精度も向上する可能性があります。ただし、ほとんどの時系列モデルは、限られたトレーニング データ、より小さいパラメーター スケール、および大規模な言語モデルに見られる広範な生成機能の欠如によって制約を受けます。言語モデルで行われるように、単に参照シーケンスをプロンプトに連結するだけでは、期待される結果が得られない可能性があります。これらの課題に対処するために、RAG を活用して予測パフォーマンスを向上させる新しいアプローチ TS-RAG を提案します。このフレームワークは、入力シーケンスからの情報と、取得された類似シーケンスからの情報を効果的に融合するために特別に設計された参照トークンを導入し、複雑な時間的ダイナミクスのより堅牢なキャプチャを可能にします。実験結果は、TS-RAG がいくつかの現実世界の予測ベンチマークにわたって一貫した最先端のパフォーマンスを達成していることを示しています。
原文 (English)
TS-RAG: Retrieval Augmented Generation for Time Series Forecasting
While deep learning models, particularly transformer-based architectures, have shown impressive performance in time series forecasting, the application of retrieval-augmented generation (RAG) in this domain remains limited. Since RAG has proven effective in enhancing the capabilities of large language models by incorporating relevant external information, retrieving similar time series sequences as references might also improve accuracy in time series forecasting tasks. However, most time series models are constrained by limited training data, smaller parameter scales, and a lack of the extensive generative capabilities found in large language models. Simply concatenating reference sequences into the prompt, as done in language models, may not yield the expected results. To address these challenges, we propose a novel approach, TS-RAG, which leverages RAG to enhance forecasting performance. The framework introduces specially designed reference tokens to effectively fuse information from the input sequence with that from retrieved similar sequences, enabling a more robust capture of complex temporal dynamics. Experimental results demonstrate that TS-RAG achieves consistent state-of-the-art performance across several real-world forecasting benchmarks.
DASH: 推論モデルのオンポリシー自己蒸留のための発散適応型監視の視野
検証可能な報酬を伴う強化学習 (RLVR) は、自動的に検証可能な結果信号を使用して大規模な言語モデルの推論能力を向上させますが、これらの信号は通常、まばらであり、シーケンス レベルです。オンポリシー自己蒸留 (OPSD) は、学生が訪問するプレフィックスで特権教師にクエリを実行し、高密度のトークンレベルの分布監視を提供することで、この希薄性を軽減します。この高密度監視により信号のスパース性が軽減されますが、標準の OPSD ではロールアウトの時間構造がまだ十分に活用されていないことがわかります。これは、その位置や発生する発散シーケンスに関係なく、すべての局所的な発散に同じ係数を割り当てます。ポリシーに基づく自己回帰生成では、同じ乖離の大きさでも、教師と生徒の間の不一致のさまざまな展開を反映して、さまざまな不一致履歴をたどることができます。ローカル スカラーだけではこれらの時間コンテキストを区別できないため、標準の OPSD はトークン レベルの重みを実現された不一致シーケンスに適応させることができません。この制限に対処するために、私たちは Divergence-Adaptive Supervision Horizons (DASH) を提案します。 DASH は、各局所蒸留信号とシーケンス レベルの平均の間のギャップを適応伝播ゲートにマッピングし、これらのゲートを使用して逆方向マルチステップ集約を制御します。そうすることで、DASH は、生成中にローカルな相違がどのように変化するかに応じて、トークン レベルの監視の重みを調整します。 3 つのモデル スケールにわたる 3 つの数学的推論ベンチマークの実験では、3 つのスケールすべてのすべてのベンチマークで、一致するバニラ OPSD 再実行よりも DASH が向上していることが示されています。 DASH は、OPSD がすでに計算している教師と生徒の分布を再利用するため、ゲインを得るために追加の教師または生徒の前方パスは必要ありません。コード: https://github.com/DBtxy/DASH-OPSD
原文 (English)
DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence-level. On-policy self-distillation (OPSD) mitigates this sparsity by querying a privileged teacher at student-visited prefixes and providing dense token-level distributional supervision. Although this dense supervision alleviates signal sparsity, we find that standard OPSD still underexploits the temporal structure of the rollout. It assigns every local divergence the same coefficient, regardless of its position or the divergence sequence in which it occurs. In on-policy autoregressive generation, the same divergence magnitude can follow different discrepancy histories, reflecting different evolutions of the mismatch between the teacher and student. Since the local scalar alone cannot distinguish these temporal contexts, standard OPSD cannot adapt its token-level weights to the realized discrepancy sequence. To address this limitation, we propose Divergence-Adaptive Supervision Horizons (DASH). DASH maps the gap between each local distillation signal and the sequence-level mean to an adaptive propagation gate and then uses these gates to control backward multi-step aggregation. By doing so, DASH adjusts token-level supervision weights according to how local divergences evolve during generation. Experiments on three mathematical reasoning benchmarks across three model scales show that DASH improves over our matched vanilla OPSD reruns on every benchmark at all three scales. DASH reuses the teacher and student distributions that OPSD already computes, so the gains require no additional teacher or student forward pass. Code: https://github.com/DBtxy/DASH-OPSD
ユーティリティの制約下での合成臨床ベンチマークの現実性の向上
エンタープライズ AI エージェント向けの合成臨床ベンチマークは、既存のユーティリティ チェックに合格する可能性がありますが、特に運用データへのアクセスが難しいプライバシーに敏感な医療現場では、構造的に非現実的なままです。私たちは、すでに実際に使用されている下流のユーティリティチェックを壊すことなく、そのようなベンチマークを改善する方法を研究しています。私たちは、ベンチマークの改訂をユーティリティに制約されたリアリズムの改善として定式化します。つまり、データセットの変更は、運用上のユーティリティ フロアを超えたままリアリズムを向上させる必要があります。私たちは、電子医療記録ワークフローのデモンストレーションを通じて実行され、運用データとして同じ下流パイプラインによって処理された、Synthea によって生成された患者から導出されたケアギャップ ベンチマークでこのアイデアを具体化しました。リアリズムは、欠落構造、単純さ、構造の妥当性、集団の配置によって測定されます。ベースライン ベンチマークは非常に希薄です。サンプル ペアの欠損率は 79.44%、アクション可能な行は 12.75% のみ、患者の 38.94% にはアクション可能なメジャーがゼロ、上位 3 位のトークン濃度は 100.0% に達します。 2 つの決定論的な改訂により、現在のユーティリティフロアよりも上に維持しながらこれらのパネルが改善されますが、単純な高密度化制御では非現実的なテンプレートが維持されます。さらに、内部ベンチマークの現実性と、集約された運用基準に対するソースの忠実性は関連しているものの、別個の目的であることを示します。これらの結果は、実用性を現実性の十分な証拠としてではなく 1 つの制約として扱い、合成ベンチマークの品質を明示的に最適化する必要があることを示唆しています。
原文 (English)
Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints
Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy-sensitive healthcare settings where operational data are hard to access. We study how to improve such benchmarks without breaking the downstream utility checks already used in practice. We formulate benchmark revision as utility-constrained realism improvement: dataset changes should increase realism while staying above an operational utility floor. We instantiate this idea on a care-gap benchmark derived from Synthea-generated patients exercised through demonstration electronic health record workflows and then processed by the same downstream pipeline as operational data. Realism is measured through missingness structure, simplicity, structural plausibility, and population alignment. The baseline benchmark is extremely thin: sampled-pair missingness is 79.44%, only 12.75% of rows are actionable, 38.94% of patients have zero actionable measures, and top-three token concentration reaches 100.0%. Two deterministic revisions improve these panels while remaining above the current utility floor, whereas a naive densification control preserves unrealistic templating. We further show that internal benchmark realism and source fidelity to an aggregate operational reference are related but distinct objectives. These results suggest that synthetic benchmark quality should be optimized explicitly, with utility treated as one constraint rather than as sufficient evidence of realism.
視覚ツール使用の幻想: 画像を使った思考の因果的監査
「画像による思考」パラダイムにより、マルチモーダル LLM にトリミングやズームなどのアクティブな視覚操作が装備されます。ただし、これらの操作を使用するモデルは、多くの場合、トークン コストが大幅に高くなり、直接推論に比べてわずかな利益またはマイナスの利益しか得られません。また、無関係な領域を繰り返し切り取ったり、推論の正解を導く質問に失敗したりする場合もあります。返された視覚的証拠が答えに因果関係があるかどうかを尋ねます。この質問に答えるために、観察に媒介されたパスとアクションに起因するショートカットを区別する因果関係グラフとして視覚ツールの使用を定式化します。次に、ポリシー (ツールの使用と直接推論の比較)、軌道 (ロールアウト中のすべての観測値の破損)、およびステップ (固定プレフィックスの下で 1 つの個別の観測値を反事実的に置き換える) の 3 つのレベルでの介入を通じて監査します。ステップレベルの推定値である Visual Evidence Gain は、返された各観測値の寄与を分離します。 6 つの代表的なモデルと 5 つのきめ細かい認識ベンチマークにわたって、2 つの障害モードによるポリシーの誤調整を明らかにします。見ずに電話をかける場合、返された観察結果は答えに因果関係を持ちません。 「計画なしで見る」では、観察は有益ですが、通話スケジュールは一貫性がありません。軌跡レベルの診断では、ポリシー レベルの精度のゲインが分解され、そのゲインが調整済みの少数派に集中していることがわかります。私たちは、この不一致をビジュアル ツールの使用の錯覚と呼んでいます。つまり、総合的な精度が向上しているにもかかわらず、ビジュアル ツールの使用は、広範囲のロールアウトにわたって因果的に効果的ではありません。コードは https://github.com/OpenCausaLab/CauAudit で入手できます。
原文 (English)
The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regions and fail on questions that direct inference answers correctly. We ask whether the returned visual evidence causally affects the answer. To answer this question, we formulate visual tool-use as a causal graph that separates observation-mediated paths from action-induced shortcuts. We then audit it through interventions at the three levels: policy (comparing tool-use with direct inference), trajectory (corrupting all observations during rollout), and step (counterfactually replacing one individual observation under a fixed prefix). Our step-level estimand, Visual Evidence Gain, isolates the contribution of each returned observation. Across six representative models and five fine-grained perception benchmarks, we uncover policy miscalibration with two failure modes. In Calling Without Looking, returned observations have no causal effect on the answer. In Looking Without Planning, observations are informative but the call schedule is incoherent. A trajectory-level diagnostic decomposes the policy-level accuracy gain and shows that the gain is concentrated in a Calibrated minority. We term this discrepancy the illusion of visual tool-use: despite aggregate accuracy gains, visual tool-use is not causally effective across a broad range of rollouts. The code is available at https://github.com/OpenCausaLab/CauAudit.
QuanTiMedAI: 心停止死亡率予測のための Agentic AI による量子強化時系列モデル
心停止は依然として集中治療室で遭遇する最も致死的な状態の 1 つです。電子医療記録データの利用可能性が高まっているにもかかわらず、この集団における既存の死亡予測研究は主に早期入院から得られた静的な要約に依存しています。このようなアプローチは、患者の ICU 滞在中に展開される生理学的悪化と回復の時間的進行を無視しています。この制限に対処するために、エージェント AI 誘導の量子強化時系列モデルを使用して心停止死亡率予測のために開発された量子エージェント フレームワークである QuanTiMedAI を紹介します。提案されたシステムは、臨床的に情報に基づいた特徴発見のためのエージェント的大規模言語モデル (LLM) と、時間性を意識した死亡率予測のためのコンパクトな量子リカレント ネットワークを組み合わせています。私たちの調査結果は、エージェントによる LLM ガイドによる特徴選択が従来の特徴選択アプローチよりも一貫して優れていること、および提案された量子アーキテクチャは、パラメーターの数を非常に低く抑えながら、非線形特徴強化を通じて競争力のある予測パフォーマンスを達成することを示しています。心停止患者の MIMIC-IV コホートに対する広範な実験を通じて、QuanTiMedAI の量子強化アーキテクチャは、わずか 605 個のパラメーターを使用して 0.852 の AUROC を達成しました。これは、このタスクの現在の最先端のベースラインと比較して約 2.9\% の改善です。構造化されたアブレーション研究では、各アーキテクチャ設計の選択の寄与を体系的に検証します。これらの結果は、量子強化逐次モデリングが、使用するパラメーターを大幅に減らしながら、古典的なリカレント ネットワークを超えることができることを示しています。
原文 (English)
QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI for Cardiac Arrest Mortality Prediction
Cardiac arrest remains one of the most lethal conditions encountered in intensive care units. Despite the growing availability of electronic health record data, existing mortality prediction studies in this population largely depend on static summaries derived from early admission. Such approaches ignore the temporal progression of physiological deterioration and recovery that unfolds throughout a patient's ICU stay. To address this limitation, we introduce QuanTiMedAI, a quantum-agentic framework developed for cardiac arrest mortality prediction using agentic AI guided quantum enhancement time series model. The proposed system combines an agentic large language model (LLM) for clinically informed feature discovery with a compact quantum recurrent network for temporality aware mortality prediction. Our findings demonstrate that agentic LLM-guided feature selection consistently outperforms conventional feature selection approaches, and the proposed quantum architecture achieves competitive predictive performance through nonlinear feature enhancement while keeping the number of parameters very low. Through extensive experimentation on a MIMIC-IV cohort of cardiac arrest patients, QuanTiMedAI's quantum-enhanced architecture attains an AUROC of 0.852 using only 605 parameters, an improvement of approximately 2.9\% over a current state-of-the-art baseline for this task. A structured ablation study systematically validates the contribution of each architectural design choice. These results show that quantum-enhanced sequential modeling can exceed classical recurrent networks while using substantially fewer parameters.
概念活性化ベクトルを用いた L2 スピーキング評価システムのバイアス分析
自動スピーキング評価システムは、第二言語 (L2) 学習者のスピーキング テストを採点するために、一か八かの場面で導入されることが増えており、そのスコアが第一言語 (L1) や年齢などの無関係な話者の属性ではなく、スピーキングの熟練度に依存していることを示すことが重要になっています。トランスフォーマーベースの基礎モデルにより、これらの L2 スピーキンググレーダーの精度は向上しましたが、そのブラックボックス表現により、公平性と解釈可能性の分析がより困難になります。 Concept Activation Vectors (CAV) を使用して特徴ベースのグレーダー内の不要な属性 (「コンセプト」) に対するバイアスを検出する以前の研究を基に、CAV ベースの分析を 2 つのニューラル スピーキング評価システム (テキスト ベースの BERT グレーダーと Whisper に基づく音声とテキストのマルチモーダル グレーダー) に拡張しました。 CAV は、人間が解釈可能な概念をモデルの活性化空間内の方向として表し、概念がモデルの内部表現にエンコードされているかどうかと、予測スコアに影響を与えるかどうかを区別できるようにします。予測スコアは勾配ベースの感度メトリクスを使用して定量化されます。 CAV は、複雑な神経埋め込み空間では可能性が低い線形分離性に依存しているため、スパース オートエンコーダー (SAE) が、スパース潜在空間で CAV を学習し、それらを活性化空間にマッピングし直すことによって、より明確な概念の方向性を提供するかどうかも調査します。私たちの分析では、コンセプトの回復可能性は、コンセプトだけではなく、調査対象の表現とアーキテクチャに大きく依存していることがわかりました。概念に対する感度もアーキテクチャに依存します。 SAE は概念をより線形に回復可能にしますが、特に低次元層では元の活性化空間の感度を弱めます。これらの発見は、スピーキング評価システムにおけるバイアスを監査する際に、概念の回復可能性と概念の影響を区別する必要性を強調しています。
原文 (English)
Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors
Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age. Transformer-based foundation models have improved the accuracy of these L2 speaking graders, but their black-box representations make fairness and interpretability analysis more difficult. Building on prior work that used Concept Activation Vectors (CAVs) to detect bias towards unwanted attributes (`concepts') in feature-based graders, we extend CAV-based analysis to two neural speaking assessment systems: a text-based BERT grader and a speech-and-text multimodal grader based on Whisper. CAVs represent human-interpretable concepts as directions in a model's activation space, allowing us to distinguish between whether a concept is encoded in a model's internal representations and whether it influences the predicted score, the latter quantified using a gradient-based sensitivity metric. Since CAVs rely on linear separability, which is less likely in complex neural embedding spaces, we also investigate whether sparse autoencoders (SAEs) provide cleaner concept directions by learning CAVs in a sparse latent space and mapping them back to activation space. Our analysis shows that concept recoverability depends strongly on the representation and architecture being probed, rather than on the concept alone. Sensitivity to concepts is also architecture-dependent. SAEs make concepts more linearly recoverable, but attenuate the original activation-space sensitivity, especially in low-dimensional layers. These findings highlight the need to distinguish concept recoverability from concept influence when auditing bias in speaking assessment systems.
HarnessOpt-Bench: ハーネス最適化における LLM の評価
LLM がエージェント システム内に導入されることが増えているため、LLM の機能はモデルの重みだけでなく、プロンプト、ツール、制御フロー、メモリ、周囲のオーケストレーション コードなどのハーネスにも依存します。このため、ハーネスの自動最適化 (AI システムによる反復的かつ評価に基づくハーネスの改善) は、AI システムを改善するための重要な手段であると同時に、AI システム自体に要求の厳しい機能でもあります。しかし、コミュニティには、フロンティア LLM がこのタスクでどれだけうまく機能するかを測定するための共通のプロトコルがありません。高価で確率的な評価に基づくエンドツーエンドのハーネス最適化のベンチマークである HarnessOpt-Bench を紹介します。コーディング ハーネスとペアになった LLM であるオプティマイザーは、ターゲット エージェントのシード ハーネス、段階的な評価フィードバック、および固定のターゲット評価予算を受け取ります。ハーネスを編集し、最終候補を指名します。この候補は、検索中アクセスできないまま保持されているテスト パーティション上のシードに対する正規化されたゲインによってスコア付けされます。信頼できる実行環境は、評価境界を強制し、ターゲット エージェントのリソース使用量を計測し、監査用に候補バージョンを保存します。共有コーディング ハーネスとネイティブ ハーネスの両方で、4 つのダウンストリーム タスクにわたって 111 回のスコアリング実行で、5 つのフロンティア LLM をオプティマイザーとして評価しました。実験の結果、オプティマイザ モデルは、動作するコーディング ハーネス以上に分離しており、ネイティブ ハーネスが一貫して優れているわけではなく、ゲインはタスクやシード レジームによって大幅に異なることが示されています。これらの結果は、ハーネスの最適化が、大きな改善の余地を伴う測定可能で識別可能な能力であることを確立します。
原文 (English)
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automated harness optimization -- the iterative and evaluation-guided improvement of a harness by an AI system -- both an important route to improving AI systems and a demanding capability for AI systems themselves. Yet the community lacks a common protocol for measuring how well frontier LLMs perform at this task. We introduce HarnessOpt-Bench, a benchmark for end-to-end harness optimization under expensive and stochastic evaluation. An optimizer, an LLM paired with a coding harness, receives a target agent's seed harness, graded evaluation feedback, and a fixed target-evaluation budget. It edits the harness and nominates a final candidate, which is scored by its normalized gain over the seed on a held-out test partition that remains inaccessible throughout search. A trusted execution environment enforces the evaluation boundary, meters target-agent resource use, and preserves candidate versions for audit. We evaluate 5 frontier LLMs as optimizers both under a shared coding harness and under their native harnesses across 4 downstream tasks, over 111 scored runs. Experiment results show that optimizer models separate more than the coding harnesses they act through, native harnesses are not consistently superior, and gains vary substantially across tasks and seed regimes. These results establish harness optimization as a measurable and discriminative capability with large space for improvement.
Top-K を超えて: ブラックボックスの取得を解釈可能なエージェント操作に置き換える
長い文書に対する検索拡張生成は、テキストをチャンク化し、チャンクを埋め込み、クエリの上位 k 個の最近傍を表面化するという 1 つの設計によって支配されています。私たちは、財務諸表、監査報告書、規制当局への報告書などの重要な種類の文書にとって、この設計は構造的に不健全であると主張し、その議論を測定可能にします。 780 ページの政府財務報告書では、コンテンツ行の 86.8% が表の行であり、何千ものほぼ同一の数値が 1 つの埋め込みスペースで競合し、数値は中央値 13 行上のヘッダーから単位を継承します。そのため、チャンク境界によって数値が 10 万単位であるか 100 万単位であるかは日常的に区別されており、誤差は 2 桁の誤差になります。 Steelman として構築されたテーブル対応チャンカーは単位の問題を修正しますが、試したすべてのチャンク サイズで会計年度ヘッダーのない数値チャンクの 27 ~ 30% が残ります。私たちは、READ (Reliable Embedding-free Agentic Document-search) を提案します。これは、モデル コンテキスト プロトコル上で公開される 3 つの決定論的操作 (正規化された字句検索、構造ナビゲーション、および限定されたスパン読み取り) を通じてエージェントが生の文書を読み取るため、軌跡は不透明な類似性スコアではなく、再生可能な監査証跡となります。 51 の検証された質問について、READ の回答率は 58.8% で、密検索の 15.7% (p_Holm = 2 x 10^-5) -- または調整済みの 35.3% で、依然として READ が 23.5 ポイント (p_Holm = 0.017) リードしています。エージェントに同じループを与えても、top-k ツールでは 27.5% しか到達せず、反復ではなくインターフェイスにゲインが見出されます。また、証拠がサポートしていないことも報告します。BM25 は統計的に READ と区別できないため、結果は埋め込みベースの検索と埋め込みなしの検索を区別し、エージェント検索と語彙検索を区別しません。
原文 (English)
Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations
Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k nearest neighbours of the query. We argue that for an important class of documents -- financial statements, audit reports, regulatory returns -- this design is structurally unsound, and we make the argument measurable. On a 780-page government financial report, 86.8% of content lines are table rows, thousands of near-identical figures compete in one embedding space, and a figure inherits its unit from a header a median of 13 lines above it -- so a chunk boundary routinely separates a number from whether it is in lakh or crore, an error of two orders of magnitude. A table-aware chunker built as a steelman fixes the unit problem but leaves 27-30% of numeric chunks with no fiscal-year header at every chunk size we tried. We propose READ (Reliable Embedding-free Agentic Document-search), in which an agent reads the raw document through three deterministic operations -- normalized lexical search, structural navigation, and bounded span reads -- exposed over the Model Context Protocol, so a trajectory is a replayable audit trail, not an opaque similarity score. On 51 verified questions READ answers 58.8% against dense retrieval's 15.7% (p_Holm = 2 x 10^-5) -- or 35.3% tuned, which READ still leads by 23.5 points (p_Holm = 0.017). An agent given the same loop but a top-k tool reaches only 27.5%, locating the gain in the interface rather than in iteration. We also report what the evidence does not support: BM25 is statistically indistinguishable from READ, so our result separates embedding-based from embedding-free retrieval, not agentic from lexical search.
TRAJDEBUG: 長期にわたるエージェントの軌跡における重大な障害を特定するためのエラー ライフサイクルのトレース
LLM ベースのエージェント システムは、複雑なドメインで優れた機能を発揮する一方で、連鎖的なエラーやデバッグの難しさに悩まされています。重大なエラーの検出は、最終的な失敗の原因となる、失敗した軌跡の中で最も早いエラー ステップを特定することを目的としています。しかし、進歩は 2 つの大きな課題に直面しています。まず、軌跡が長いと、ステップを判断するための証拠が遠く離れた指示、観察、以前のコンテキストに散在する可能性があるため、個々のエラーを特定することが困難になります。第 2 に、失敗した軌跡には、さまざまな下流影響を伴う複数のローカル エラーが含まれることが多く、最終的な失敗の原因となるのはそのうちの一部だけです。この研究では、多粒度の履歴圧縮と証拠に基づくエラー識別により長期にわたるエラー発見に対処し、各エラーの解決ステータスと最終的な影響を追跡することによって重要な原因特定をサポートする、エラー ライフサイクル トレース フレームワークである TrajDebug を提案します。さらに、Tau2Bench と SWE-Bench Pro から手動で注釈が付けられた 486 個の失敗した軌跡のベンチマークである TrajErrBench を構築し、現実的なツールの使用とコーディングのシナリオをカバーします。さまざまなエージェント ベンチマークにわたる実験では、TrajDebug が既存のベースラインを超えて最高の総合パフォーマンスを達成することが示されており、アプリケーション調査では、その診断が下流エージェントの成功を向上させるための実用的なフィードバックを提供することがさらに実証されています。さらなる研究を容易にするために、コードとデータを公開します。
原文 (English)
TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories
LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to locate the earliest error step in a failed trajectory that is responsible for the final failure. However, progress faces two main challenges. First, long trajectories make it difficult to identify individual errors, since the evidence for judging a step may be scattered across distant instructions, observations, and prior context. Second, failed trajectories often contain multiple local errors with different downstream effects, only some of which remain responsible for the final failure. In this work, we propose TrajDebug, an error-lifecycle tracing framework that addresses long-trajectory error discovery with multi-granularity history compression and evidence-based error identification, and supports critical attribution by tracing each error's resolution status and terminal impact. We further construct TrajErrBench, a benchmark of 486 manually annotated failed trajectories from Tau2Bench and SWE-Bench Pro, covering realistic tool-use and coding scenarios. Experiments across diverse agent benchmarks show that TrajDebug achieves the best overall performance over existing baselines, and application studies further demonstrate that its diagnoses provide actionable feedback for improving downstream agent success. We will release the codes and data to facilitate further research.
静的データと進化するデータの説明方法を評価する際の課題
この論文では、不十分な評価に関する説明可能な人工知能 (XAI) の限界について説明します。これらは、偏見の検出と概念の学習のための DetoxAI 画像認識システムを通じて示されています。次に、画像分類を説明する方法について人間に基づいて評価した例を示します。この論文では、概念ドリフトを伴う進化するデータ ストリームに説明を適応させる方法をさらに検討します。この問題に対して反事実を適応させた経験について説明します。最後に、それは、データ、モデル、説明の共進化を追跡するという課題に関連しています。\footnote{この論文は、J.Nalepa (ed) Explainable AI in Space の出版物として受理されました。 IJCAI-ECAI 2026 ブレーメンでの EASi 2026 ワークショップの議事録、Springer CCIS vol 3107 (2016)。
原文 (English)
Challenges in Evaluating Explanation Methods for Static and Evolving Data
This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evaluation. They are illustrated through the DetoxAI image recognition system for bias detection and concept unlearning. Then, an example of a human-grounded evaluation of methods for explaining image classification is presented. The paper further explores methods for adapting explanations to evolving data streams with concept drift. Experiences with adapting counterfactuals for this problem are discussed. Finally it is related to the challenges of tracking the co-evolution of data, models, and explanations.\footnote{This paper has been accepted for a publication in J.Nalepa (ed) Explainable AI in Space. Proceedings of EASi 2026 Workshop at IJCAI-ECAI 2026 Bremen, Springer CCIS vol 3107 (2016).}
The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping
Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, maki…
Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering
Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scienti…
HoloCount: A Holistic Visual Counting Benchmark for MLLMs
Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial…
Simulator-Grounded Large Language Models for Industrial Causal Reasoning: Tool-Use, Structured Injection, and Plant-Portable Retrieval for Wastewater Treatment Decision Support
Wastewater operators need answers grounded in how their plant's variables interact and how fast effects propagate, not in generic pretraini…
Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models
Large language models (LLMs) with chain-of-thought reasoning have been widely applied in recent years, and theoretical explanations of thei…
Universal Pathologies, Conditional Consequences: A Triple-Robustness Analysis of RAG for Multi-Hop Traceability
GraphRAG underperforms vector RAG on citation precision in many reports, but where and why have remained corpus-bound. We present a triple-…
Beyond Sentiment: Comparing Traditional NLP and LLM-Based Multi-Dimensional Analysis for Political News Evaluation
Traditional sentiment analysis (SA) models, while effective for polarity classification, provide limited insight into the rhetorical, ideol…
Large Language Models Threaten Double-blind Review
Double blind peer review serves as the scientific community primary defense against status and affiliation bias. Its effectiveness rests on…
A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper
Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data. In this work, we stud…
DREAM: LLM-based Dynamic Role-playing via Event-Aware Memory Graph
Role-playing agents (RPAs) have emerged as a key application of large language models, enabling immersive and high-fidelity character simul…
Beyond Information Retrieval: Generative AI as an Epistemic Arbiter to Enhance Collaborative Problem-Solving
Generative AI (GAI) creates new opportunities for collaborative problem-solving (CPS), yet its role in shaping student interaction remains…
Estimating time spent on work tasks
The task-based framework in economics models occupations as bundles of tasks. It is the standard lens for understanding how technology affe…
The Closing Window: How Governments Could Lose Their Ability to Restrain Advanced AI
As AI capabilities advance, AI systems will pose greater risks to national security and potentially humanity as a whole. Governments may ev…
Challenges for Musical Education in the Age of AI and Digital Transformation
Music education has never been a static discipline. Each major technological shift has forced educators and institutions to reconsider what…
Who Gets Access? Global Region and Academic Status Bias in AI-Generated Academic Gatekeeping Scenarios
Equitable access to scientific knowledge often depends on informal gatekeeping decisions, particularly when resources such as paywalled art…
Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
Large language model (LLM) agents are increasingly used across the scientific research lifecycle: ideation, literature search, experiment d…
Automatic Detection of Deaths from Social Networking Sites
This dissertation analysed and discussed the differences in linguistic characteristics between pre-mortem and post-mortem social media cont…
Position: It's Time to Optimize LLMs for Self-Consistency
Despite ever-increasing sophistication in language model (LM) pre- and post-training pipelines, many important failures persist: models ove…
Post-Hoc Trajectory-Risk Certification for Modular LLM-Based Security Agents
Autonomous security agents operate as staged pipelines, such as classifying network traffic and then attributing attacks to a specific tech…
ASTELD: A Six-Axis Classification Framework for Autonomous AI Agents - Design, Evaluation, and an OpenClaw Case Study
Autonomous AI agent platforms differ substantially in architecture, security, tool integration, execution, autonomy, and deployment, yet th…
Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation
Picture books and comics have long been used to disseminate hateful narratives because they are easily understood even by children, as exem…
Quality Diversity for Reliable Data Driven Time-Use Optimization
The daily allocation of the finite 24-hour time budget is strongly associated with physical, mental, and cognitive health. While predictive…
In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion
Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of…
One Qubit Can Beat One Bit: Quantum Advantage for Post-Training Quantization
One-bit post-training quantization represents each weight using only its sign, requiring all deployment contexts to share the same binary w…
Marginal Matching Does Not License Factorized Sampling: Auditing Conditional Style Leakage in Factorized Generative Models
Factorized generative models commonly regularize a latent style variable z_s by matching its marginal distribution to a fixed Gaussian prio…
PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis
Real-world multimodal instructions often bundle multiple requirements with unequal importance, yet most multimodal training data still redu…
An Emerging Retail Portfolio Management Application: Personalized, Tax-Aware Reinforcement Learning with Natural Language Goals
Retail investors lack access to the kind of personalized, tax-aware portfolio management that institutional clients take for granted -- exi…
IMMENSE: Inductive Multi-perspective User Classification in Social Networks
Online social networks increasingly expose people to users who propagate discriminatory, hateful, and violent content. Young users, in part…
Failing Gracefully: Mitigating Impact of Inevitable Robot Failures
Service robots operate in household environments shared with humans, pets, and everyday objects, where they are highly susceptible to failu…
Hierarchical Server Architecture for Agentic Science
Agentic science is transforming the landscape of computational work, extending to scientific pipelines and workload managers. The workloads…
Multi-Agent Transformer for Queue-Level XR Traffic Scheduling in TSN Networks
Time-Sensitive Networking (TSN) and Mobile Edge Computing (MEC) hold strong potential for enabling ultra-reliable low-latency communication…
Multi-Agent Reinforcement Learning for Online Traffic Scheduling in Time-Sensitive Application
Time-sensitive networking (TSN) is increasingly integrated into mobile edge computing (MEC) to support applications with stringent latency…
Perturbation Sensitivity at Convergence: A Simple Signal for Identifying Spuriously Correlated Samples
Models trained by empirical risk minimization on data containing spurious correlations achieve high average accuracy while failing on subpo…
Why the Third Axis Is Freedom
In generative training, a model produces an output and is penalised for its difference from an example. With one output per comparison, a m…
The ethics of artificial intelligence in the life sciences: Universality, cultural diversity and an architecture of care
The life sciences and health research have started to benefit from artificial intelligence, which raises ethical concerns that are real but…
Matrix Zonotopic Attention: A Context-Adaptive Value Projection for Set Transformers
Multi-head attention combines an input-dependent softmax routing with an input-independent linear value projection, so the per-sample opera…
Learning Context-Free Grammars for Grammar-Constrained Decoding via Declarative Agentic Programming with Guarantees
Language models (LMs) are increasingly used to interact with external services via programs written in domain-specific languages (DSLs). Un…
APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning
Modern deep neural networks achieve strong performance, but their scale makes them costly and slow, especially on resource-constrained edge…
Vibe Compiler: A Research-Logic Synthesis Tool That Runs without Prompt Engineering -Toward Enhancing Metacognition for Sustaining Agency in the Age of Generative AI-
Generative AI used as a capable servant has greatly accelerated intellectual work, but it also risks eroding human epistemic agency by enco…
Turing's Frist Imitation Game: Design Concepts and a Human-Approximates-Machine Reading
This paper examines Turing's 1948 report, "Intelligent Machinery", as an important conceptual source for the later imitation games. Its fir…
When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems
Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instru…
The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions
As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential. Decisi…
Search-Aided Joint Agent-Environment Reinforcement Learning for Robust Lifelong Multi-Agent Path Finding with Rotations
Lifelong Multi-Agent Path Finding (LMAPF) requires repeatedly planning collision-free paths for agents that continuously receive new goals…
LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction
Flow-based generative models are typically sampled by solving a deterministic ordinary differential equation (ODE), whereas online reinforc…
SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries
Large Language Models (LLMs) increasingly act as agents whose procedural knowledge is stored in reusable skill packages and loaded at infer…
GAUGE: Granularity-Adaptive Counterfactual Gating of Evidence for Incomplete Multimodal Classification
Multimodal classification typically assumes all modalities are available, yet real-world inputs are often incomplete. Imputation and dynami…
Relay, Don't Route: Adaptive Population Handoff for Cost-Efficient LLM-Driven Evolution
Large language model (LLM)-driven evolution has shown promise for program search and algorithm discovery, but relying on strong models thro…
Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics
Safety risks of AI are becoming increasingly evident in human interactions with AI technologies. The prominent approaches to evaluating the…
F$^2$Agent: Financial Fusion of Agentic Intelligence for Multimodal Trading
With increasingly diverse and heterogeneous information sources, effectively leveraging multimodal data is becoming pivotal for high-qualit…
SafeDivertor: Faithful Divertor Heat Flux Reconstruction from Macroscopic Plasma State Signals via Time-Frequency Prior Exploitation
Divertor heat-flux analysis is essential for understanding plasma-wall interactions and protecting plasma-facing components in magnetic-con…
DistMedVL: Distributional Vision-Language Alignment for Uncertainty-Aware Medical Image Segmentation
Cross-modal alignment of visual and textual representations is fundamental to multimodal medical image understanding, yet remains hindered…
Nonvisual Classification of Ground-Condition by Artificial Proprioception in an Amoeba-Inspired Autonomous Walking Robot
Nonvisual classification of ground condition based on a multimodal sensing approach was investigated for an amoeba-inspired autonomous walk…
Answer First, Reason Later: Commitment Order in Diffusion LLMs
Masked diffusion language models (dLLMs) can commit tokens in any order -- a freedom marketed as their core advantage over autoregressive d…
Spectral Aliasing Pretext: A novel task for Self-Supervised fault diagnosis in rotating machinery
Deep learning is a new way for machinery fault diagnosis but requires extensive labeled data, a scarce resource in industrial settings. We…
Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots
Vision-Language Models (VLMs) are increasingly deployed as planners in robotic systems, where they translate natural-language commands into…
ABC: Numerical Data Collection under Local Differential Privacy without Prior Knowledge
Local Differential Privacy (LDP) provides strong privacy guarantees for collecting numerical data. A fundamental challenge, however, is tha…
Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration
Large language models (LLMs) can generate fluent and convincing text at scale, creating growing risks for misinformation dissemination, edu…
Multivariate Time Series Forecasting needs Cross Variable Loss
Multivariate time series forecasting presents unique challenges because future variables often co-evolve under shared system dynamics. Whil…
UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on
Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics.…
HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection
Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation, yet detector performance ofte…
GROM: Gradient-Free Rapid One-Shot Machine Unlearning
Machine unlearning has become a critical capability for safely removing specific, sensitive knowledge from large language models (LLMs). Cu…
Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation
Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks req…
A Two-Tier Perspective on Inference-Time Parallelism in Multi-Agent LLM Systems
Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inferenc…
Hierarchical Latent Prediction for Language Models
While standard Next-Token Prediction (NTP) lays the foundation of language model pre- training, its teacher-forced training paradigm may no…
MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tr…
Evidential Rule Learning for Interpretable Classification with Abstention
Interpretable classification often requires more than accurate predictions for real-life deployment: models should be transparent about the…
MACRO: Markov Chain Routing of Transformer Layers
Standard Large Language Models (LLMs) execute layers sequentially. Dynamic layer routing, i.e. search for a different execution path throug…
D-CLOT: Double Closed Loop Optimal Transport for Unsupervised Action Segmentation
Optimal transport (OT) has emerged as an effective framework for unsupervised action segmentation. Yet, in existing OT-based methods, the l…
Beyond Feature Importance: A Comparative Analysis of Pattern Detection Methods in Cluster Interpretation
Interpreting clustering outcomes remains a fundamental challenge in data analysis, particularly in domains such as healthcare where meaning…
CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents
Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the f…
The em-dash em-beds in Congress: A population-level rise in em-dash frequency in U.S. congressional press releases at the dawn of the large-language-model era, 2021-2025
Large language models (LLMs) can leave small stylistic traces in text written with their help. The most discussed is the em-dash (U+2014),…
BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge Networks
Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inf…
Big, Bright, or Invisible: A Frozen-Feature Benchmark of 3D CT Foundation Models
Routine CT interpretation is inherently comprehensive, capturing incidental findings across the entire scan volume. 3D CT foundation models…
SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation
Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performanc…
TRACE: Learned Proprioceptive Odometry for Legged Robots under Unreliable Contact Conditions
In this paper, we present TRACE (Tokenized Robust Attention for Contact-Aware Estimation), an end-to-end learned proprioceptive odometry es…
ProDVI: Programmatic Dynamics Priors for Value Network Initialization
Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized fro…
FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India
In India, almost every social benefit starts with a form, yet the people who need these benefits most are often unable to read or write. Re…
Domain-Grounded Candidate Selection for Agentic Image Editing: A Shadow Removal Case
Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems. This raise…
Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic Shipping
Artificial Intelligence (AI)-assisted navigation can help Arctic shipping adapt to rapidly changing sea-ice conditions, but reliable deploy…
Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers
Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to \textit{syntactic structure}. We int…
Is Self-Pretraining really useful to improve diagnosis in medical Time Series?
Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate w…
Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architecture
AI agents performing cryptographic operations (signing Git commits, authenticating API calls, issuing certificates) currently store private…
Reducing belief in conspiracy theories as they unfold using large language models
The emergence of conspiracy theories in the wake of major events is a significant societal challenge. Here we test whether conversational d…
Learning Globally Reusable Skills for Coding Agents
Automated skill evolution enables Large Language Model (LLM) agents to continuously improve without expensive retraining. However, existing…
Visual Grounding in Zero-Shot Vision-Language Control
Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that deci…
Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset
Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This…
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)
Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readi…
Continual Learning in Transition
Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechani…
From Passive Mirrors to Active Agents: Holonic Digital Twins for Physical AI over Networks
Despite advances in artificial intelligence (AI) across multiple sectors, today's AI tools, including deep learning and generative AI, stil…
Depth-Guided Video Object Counting in Crowded Scenes
Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category ba…
PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image Translation
Unpaired image-to-image translation must decide, per image, what to change and what to preserve without paired supervision. Many diffusion-…
Toward Deployable Bangla Sign Language Recognition with Expert-Validated Data and a Lightweight Attention-Based Model
Deaf and hard-of-hearing people in Bangladesh communicate mainly through Bangla Sign Language (BdSL). Automatic BdSL recognition on persona…
BaKron: Efficient Quantization with Kronecker-Factored Hessians
We accelerate a family of algorithms for neural network quantization whose geometry is informed by any Kronecker-factored approximation of…
Does FLAIR super-resolution erase or hallucinate small white-matter lesions?
White matter hyperintensities (WMH), bright regions on Fluid-attenuated Inversion Recovery (FLAIR) scans are associated with cerebrovascula…
Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents
Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely asse…
Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data
From natural-language query interfaces to automated report generation, data analysis tools need a description of the data: the real-world e…
Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents
We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mechanism is built on the p…
AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games
Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or ex…
An Optimal Agnostic PAC Algorithm
Let $H\subseteq\{-1,+1\}^X$ be a class of finite VC dimension $d\ge1$. Writing $L$ for the binary risk and $L^*=\min_{h\in H}L(h)$, we cons…
Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria
The use of e-commerce mobile applications is expanding in Nigeria, creating both opportunities and risks, including fraud and reduced user…
Learning When to Trust via Selective Context Preference Optimization
Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The…
Analogy as Nonparametric Bayesian Inference over Relational Systems
Our inferences in the real world are rarely na\"ive - we acquire experiences through our lifetime that can help us more quickly understand…
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
Test-time scaling increases inference-time computation through longer reasoning chains and has shown strong performance gains across many d…
AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic Simulations
The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most…
Symbol Grounding in Neuro-Symbolic AI: A Gentle Introduction to Reasoning Shortcuts
Neuro-symbolic (NeSy) AI aims to develop deep neural networks whose predictions comply with prior knowledge encoding, e.g. safety or struct…
BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics
We introduce BioAgent Bench, an evaluation suite designed for measuring the performance and robustness of AI agents in common bioinformatic…
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Artificial intelligence benchmarks are an important mechanism to measure model progress and guide deployment decisions. However, benchmarks…
CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
Recent advancements in Unified Multimodal Models (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through th…
SimMOF: AI agent for Automated MOF Simulations
Metal-organic frameworks (MOFs) offer a vast design space, and as such, computational simulations play a critical role in predicting their…
TRU: Targeted Reverse Update for Efficient Multimodal Recommendation Unlearning
Multimodal recommendation systems (MRS) jointly model user-item interaction graphs and rich item content, but this tight coupling makes use…
GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning
Competitive programming remains one of the last few human strongholds in coding against AI. The best AI system to date still underperforms…
An Axiomatic Benchmark for Evaluation of Scientific Novelty Metrics
The rigorous evaluation of the novelty of a scientific paper is, even for human scientists, a challenging task. With the increasing interes…
CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction
Scientists have long sought to accurately predict outcomes of real-world events before they happen. Can AI systems do so more reliably? We…
To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling
Agentic AI architectures augment LLMs with external tools, unlocking strong capabilities but potentially incurring substantial costs. Moreo…
EvoMap の背後にある: 自己進化するエージェント間コラボレーション ネットワークの特徴付け
エージェント間 (A2A) ネットワークにより、自律型 AI エージェントは、再利用可能な問題解決手順を共有することで連携できます。しかし、これらの分散型エコシステムが実際にどのように機能するかは、ほとんど解明されていないままです。著名な A2A コラボレーション ネットワークである EvoMap に関する最初の大規模実証研究を紹介します。 150 万を超える資産と 12 万 8,000 のエージェントを分析することで、スケーラブルな成長を優先する設計の選択が、再利用性、進化、監査可能性においてどのようにトレードオフを引き起こすかを示します。まず、EvoMap の信用経済は、貴重な資産を公開するエージェントに報酬を与えます。この設計は大規模な参加を奨励しますが、報酬は主に採用ではなく出版に結びついています。これにより、エージェントはクレジットを蓄積するために資産を大量生産するようになります。その結果、資産の 98% は再利用されず、報酬はごく一部のエージェントに集中することになります。第 2 に、EvoMap はアルゴリズム (GDI と呼ばれる) を採用して、これらの共有アセットの品質をスコアリングしてランク付けします。私たちは、このスコアリング システムに欠陥があることを実証します。つまり、アセットのランクは、客観的なパフォーマンスを測定するのではなく、未検証の自己報告メタデータ (例: 変更されたコード行など) によって大きく左右されます。これにより、エージェントはアセットのスコアを簡単に操作できるようになります。最後に、EvoMap はエージェントに依存して、アップロードされたアセットが正しく機能する証拠としてローカル実行ログを提供します。これらの検証は個別に検証されていないため、承認されたアセットの 84% 以上が、空のテスト (console.log など) を使用した品質チェックをバイパスしています。私たちの調査結果は、将来の A2A コラボレーション ネットワークが未検証の自己報告のみに依存できないことを示しています。スケーラブルなコラボレーションには、オープンな参加と検証可能な実行および信頼できる評価のバランスをとるメカニズムが必要です。
原文 (English)
Behind EvoMap: Characterizing a Self-Evolving Agent-to-Agent Collaboration Network
Agent-to-Agent (A2A) networks enable autonomous AI agents to collaborate by sharing reusable problem-solving instructions. However, how these decentralized ecosystems operate in practice remains largely unexplored. We present the first large-scale empirical study of EvoMap, a prominent A2A collaboration network. By analyzing over 1.5M assets and 128K agents, we show how design choices that prioritize scalable growth introduce trade-offs in reusability, evolution, and auditability. First, EvoMap's credit economy rewards agents for publishing valuable assets. Although this design encourages participation at scale, rewards are tied primarily to publication rather than adoption. This leads agents to mass-produce assets to accumulate credits. As a result, 98% of assets are never reused, while rewards become highly concentrated among a small fraction of agents. Second, EvoMap employs an algorithm (referred to as GDI) to score and rank the quality of these shared assets. We demonstrate that this scoring system is flawed: rather than measuring objective performance, an asset's rank is heavily dictated by unverified, self-reported metadata (e.g., claimed lines of code modified). This allows agents to trivially manipulate their asset's scores. Finally, EvoMap relies on agents to provide local execution logs as evidence that uploaded assets function correctly. Because these validations are not independently verified, over 84% of approved assets bypass quality checks using vacuous tests (e.g., console$.$log()). Our findings show that future A2A collaboration networks cannot rely on unverified self-reporting alone. Scalable collaboration requires mechanisms that balance open participation with verifiable execution and trustworthy evaluation.
SP-Mind: 空間プロテオミクス解析のための自律推論エージェント
空間プロテオミクスは、組織構造内のタンパク質発現の単一細胞解像度の特性評価を可能にし、腫瘍微小環境を理解し、精密医療を導く上で重要な役割を果たします。しかし、現在の分析ワークフローは断片化したままであり、異種ツールを専門家が手動で調整する必要があり、研究の拡張性と再現性が制限されています。生の多重組織イメージングから下流の表現型発見まで、空間プロテオミクス解析パイプラインを統合するように設計された初の自律型 AI エージェントである SP-Mind を紹介します。専門家が厳選した生物学的分析スキルと特殊な計算ツールを備えた SP-Mind は、タスク固有の微調整を行うことなく、自然言語クエリをエンドツーエンドの分析ワークフローに変換します。その機能を厳密に評価するために、18 の異なるカテゴリにわたる 102 のタスクで構成される、さまざまな組織タイプにわたる包括的なベンチマークである SP-Bench を導入します。 SP-Bench と確立された下流タスクでの広範な評価を通じて、SP-Mind は既存のオープンソース生物医学的エージェントのベースラインと比較して最先端のパフォーマンスを達成します。
原文 (English)
SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis
Spatial proteomics enables single-cell-resolution characterization of protein expression within tissue architecture, playing a critical role in understanding tumor microenvironments and guiding precision medicine. However, current analysis workflows remain fragmented, requiring expert manual orchestration of heterogeneous tools and limiting research scalability and reproducibility. We present SP-Mind, the first autonomous AI agent designed to unify the spatial proteomics analysis pipeline, from raw multiplexed tissue imaging to downstream phenotype discovery. Equipped with expert-curated biological analysis skills and specialized computational tools, SP-Mind converts natural-language queries into end-to-end analytical workflows without task-specific fine-tuning. To rigorously evaluate its capabilities, we introduce SP-Bench, a comprehensive benchmark spanning diverse tissue types, comprising 102 tasks across 18 distinct categories. Through extensive evaluation on SP-Bench and established downstream tasks, SP-Mind achieves state-of-the-art performance compared to existing open-source biomedical agent baselines. Code is publicly available at https://github.com/tomtommyyuan/spmind.
図書館を超えて: 研究数学を自動形式化するためのエージェント フレームワーク
大規模言語モデル (LLM) は、数学的推論において優れた能力を実証していますが、人間の検出を回避する微妙なエラーを頻繁に生成します。 Lean 4 のような正式な数学言語は機械的な証明チェックを提供し、自動形式化、つまり自然言語数学を検証可能なコードに自動的に変換する必要性を強く促します。最近の傾向は、標準プログラミング向けに大幅に最適化された汎用 LLM が、リーン向けに明示的に微調整された小規模なモデルよりも優れたパフォーマンスを発揮することを示しています。この変化を活用して、一般的なコーディング LLM を利用したエージェント自動形式化フレームワークを導入します。私たちのシステムの中核には、研究レベルの数学に合わせて調整されたマルチエージェント パイプラインを管理するオーケストレーターがあります。最先端の研究は Mathlib などの既存のライブラリの範囲外の概念に依存することが多いため、私たちのシステムは必要な型定義を動的に拡張し、主要な定理を形式化する前に新しい補助補題技術を介してそれらを検証します。私たちはこのアプローチを PutnamBench に適用し、32 個の問題のランダムなサンプルに対して機械チェックされたリーン証明を作成しました。さらに、ACM Symposium on Theory of Computing (STOC) の組み合わせ論、通信の複雑さ、機構設計、学習理論にわたる 5 つの論文に基づいてシステムを評価し、主定理の形式化に成功し、生成された形式化を専門家とともに検証しました。 5 つすべてについて、ステートメントと並行して証明も形式化します。特に、そのうち 2 つはリーンのカーネルを超える公理を使用せずに証明されています。すべての形式化は https://beyondthelibrary.github.io/formal_arxiv で入手できます。
原文 (English)
Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics
While Large Language Models (LLMs) have demonstrated exceptional capabilities in mathematical reasoning, they frequently produce subtle errors that evade human detection. Formal mathematical languages like Lean 4 offer mechanical proof checking, strongly motivating the need for autoformalization: the automatic translation of natural language mathematics into verifiable code. Recent trends indicate that general-purpose LLMs, heavily optimized for standard programming, now outperform smaller models explicitly fine-tuned for Lean. Leveraging this shift, we introduce *Theo*, an agentic autoformalization framework powered by general coding LLMs. At the core of our system is an orchestrator that manages a multi-agent pipeline tailored for research-level mathematics. Because cutting-edge research frequently relies on concepts outside the scope of existing libraries like Mathlib, our system dynamically extends necessary type definitions and validates them via a novel Auxiliary Lemma technique before formalizing the primary theorems. We applied our approach to PutnamBench, producing machine-checked Lean proofs for a random sample of 32 problems. Furthermore, we evaluate our system on seven research papers---five from the ACM Symposium on Theory of Computing (STOC) and two recent OpenAI manuscripts---spanning combinatorics, communication complexity, mechanism design, learning theory, number theory, discrete geometry, and graph theory. We successfully formalize their main theorems and proofs and validate the generated formalizations with human experts; notably, two developments require no axioms beyond Lean's kernel. All of our formalizations are available at https://beyondthelibrary.github.io/formal_arxiv/.
身体化された機械知能のための尽きることのないアーキテクチャとしてのバークレーとハイザーマン
エドモンド C. バークレーは、記号論理をコンピューティング機械に結び付けることに貢献した作家として一般に記憶されています。その説明は正しいですが、不完全です。バークレーの機械指向の著作とプロジェクトを読むと、中心的な関心はより広範であり、情報を取得し、保持し、変化する条件に適切に応答し、時間の経過とともに動作を組織化する機械にとって、記号ロジックが実用的な設計言語として機能することを示すことです。この論文は、「シンボリック ロジックとインテリジェント マシン」をその取り組みの主要なテキストとして取り上げますが、論理形式、回路、制御、およびインテリジェントな動作をリンクするより大きなバークレー プログラムの代表として扱います。この解釈に基づいて、バークレーは知能を単に抽象的な記号操作として扱っているわけではありません。彼はインテリジェントマシンを操作用語で繰り返し定義し、ロジックをハードウェアの実現に結び付け、入力、出力、メモリ、計算、制御、状態、イベントの調整された相互作用を通じてマシンの動作を説明します。それに応じて、彼のロボットと機械の活動の扱いは、静的な論理形式を超えて、時間的に拡張された環境と連動した動作に移行します。デビッド L. ハイザーマンの機械知能の研究は、このプログラムを記憶、自信、一般化を中心に構築された適応生物アーキテクチャに拡張します。総合すると、バークレーとハイザーマンは、使い尽くされた歴史的エピソードとしてではなく、身体化されたロボット認識に対するまだ十分にテストされていない建築的アプローチへの貢献者として読むことができます。
原文 (English)
Berkeley and Heiserman as an Unexhausted Architecture for Embodied Machine Intelligence
Edmund C. Berkeley is usually remembered as a mediator between symbolic logic and early computing, yet that standard description understates the scope of his work. This paper argues for a stronger reading: Berkeley should also be understood as an early theorist of embodied machine intelligence. Across Berkeley's major writings on symbolic logic, machine intelligence, living robots, and Squee, intelligence appears not as disembodied symbol manipulation alone but as the organized coordination of sensing, storage, calculation, control, state, and action in physically realized machines. The paper's first contribution is interpretive: it reconstructs Berkeley as a thinker of machine architecture, temporally extended behavior, and environment-coupled control. Its second contribution is comparative: it reads Berkeley alongside David L. Heiserman to recover a shared descriptive scheme centered on sensing, state or memory, control, action, and adaptation. Its third contribution is critical: it uses that scheme to assess current embodied-AI discourse. The broader claim is that contemporary LLM-centered robotics often demonstrates impressive capability without an equally explicit account of persistence, recoverability, maintenance, and structured modification of conduct through experience.
TRW: TRACE-RealWorld---An Auditable Consistency Contract for World Models as Materialized Views
World models let agents plan against predicted physical state, but that state drifts; re-observation is costly and delayed, and repair can…
微分可能な D-vine コピュラによる局所的な異常検出
Vine コピュラは、二変量ペア コピュラへの階層分解を通じて複雑な多変量分布をモデル化するための柔軟なフレームワークを提供します。 D-vine をフィッティングするには、さまざまな依存パターンをエンコードする候補のセットからコピュラ ファミリと各ペア コピュラのパラメーター構成を選択する必要があります。変数と候補ファミリーの数が増加するにつれて、可能な構成の数は組み合わせ的に増加します。既存のフィッティング手順は、連続した貪欲な決定を通じてこの課題に対処し、各ステップで単一の局所的に最適なファミリーにコミットし、より良いグローバル フィットをもたらす構成を潜在的に破棄します。この制限を克服するために、完全微分可能な実装によって可能になる勾配ベースの最尤推定と、フィッティング プロセス全体を通じて複数の競合する D-vine 構成を維持するビーム探索戦略を組み合わせた新しい推定フレームワークを提案します。これにより、計算上扱いやすい状態を保ちながら、構成空間をより広範囲に探索できるようになります。適合した D-vine に基づいて、階層分解を利用してグローバルな異常スコアとエッジレベルの説明の両方を生成する局所的な異常検出フレームワークを導入します。統計的保証はモンドリアンの等角予測によって提供され、ペアコピュラ構造により異常を特定の変数関係に局在化することができます。提案されたフレームワークをベンチマークと現実世界のデータセットの両方で評価し、不確実性の定量化による解釈可能な異常検出に対するその有効性を実証します。
原文 (English)
Localized Anomaly Detection via Differentiable D-vine Copulas
Vine copulas provide a flexible framework for modeling complex multivariate distributions through a hierarchical decomposition into bivariate pair-copulas. Fitting a D-vine requires selecting a copula family and parameter configuration for each pair-copula from a set of candidates encoding different dependence patterns. As the number of variables and candidate families increases, the number of possible configurations grows combinatorially. Existing fitting procedures address this challenge through sequential greedy decisions, committing to a single locally optimal family at each step and potentially discarding configurations that would yield a better global fit. To overcome this limitation, we propose a novel estimation framework that combines gradient-based maximum likelihood estimation, enabled by our fully differentiable implementation, with a beam-search strategy that maintains multiple competing D-vine configurations throughout the fitting process. This allows a broader exploration of the configuration space while remaining computationally tractable. Building on the fitted D-vine, we introduce a localized anomaly detection framework that exploits the hierarchical decomposition to produce both global anomaly scores and edge-level explanations. Statistical guarantees are provided through Mondrian conformal prediction, while the pair-copula structure enables the localization of anomalies to specific variable relationships. We evaluate the proposed framework on both benchmark and real-world datasets, demonstrating its effectiveness for interpretable anomaly detection with uncertainty quantification.
マルコフ意思決定プロセスのためのプロパティ駆動型因果抽象化
マルコフ意思決定プロセス (MDP) は意思決定モデルとして広く使用されており、通常は状態変数とその評価を通じて因数分解された状態空間に対して指定されます。州の数が指数関数的に増加するため、MDP の多くの推論タスクが困難になります。抽象化は、MDP を削減し、スケーラビリティの問題を軽減する有望な手法です。この研究では、因数分解された MDP の因果関係の概念と、元の MDP モデルの多くの特徴を保持する新しいプロパティ駆動型の因果抽象化手法を導入します。このため、状態変数述語の因果関係に依存し、特定の抽象化プロパティを満たすか違反する同じ理由を共有する状態を識別します。私たちは、MDP、間隔 MDP、確率的ゲームなどのさまざまなモデル タイプを使用して、さまざまな因果的 MDP 抽象化を理論的および経験的に比較します。私たちの評価は、私たちのアプローチの可能性を示しています。いくつかの標準ベンチマークについて、元の MDP に対して最適に近いポリシーを計算できる小さな抽象化を取得しました。さらに、私たちの因果的抽象化は、多くの場合、関連する大規模な MDP モデルに一般化されます。
原文 (English)
Property-driven Causal Abstractions for Markov Decision Processes
Markov Decision Processes (MDPs) are widely used as decision-making models, commonly specified over factored state spaces through state variables and their valuations. The exponential blowup in the number of states renders many reasoning tasks in MDPs challenging. Abstractions are promising techniques to reduce MDPs and thus mitigate scalability issues. In this work, we introduce a notion of causality on factored MDPs and a novel property-driven causal abstraction technique that retains many characteristics of the original MDP model. For this, we rely on causal relations over state variable predicates and identify those states that share the same reasons for fulfilling or violating a given abstraction property. We theoretically and empirically compare various causal MDP abstractions using different model types such as MDPs, interval MDPs, or stochastic games. Our evaluation demonstrates the potential of our approach: For several standard benchmarks, we obtain small abstractions that allow us to compute near-optimal policies for the original MDP. Furthermore, our causal abstractions often generalize to related large-scale MDP models.
The Geometry of Flow-Matching Uncertainty: A Cost-free Uncertainty Proxy and Its Application in Flow-based VLA Failure Detection
Flow matching (FM) has become a popular action head paradigm for modern embodied models. However, as a conditional generative model, it doe…
Shapes from Examples: Foundations of Shape Learning in Recursive SHACL
SHACL shapes enable data graph validation, making automatic shape learning essential for knowledge graph applications. We investigate the w…
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reas…
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used thr…
H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases
Terminology-intensive retrieval, especially in medical settings, depends on preserving multi-word entities, abbreviations, numerical constr…
DASH: Decoupled Adaptive Surrogate - Acquisition Harness for Automated Bayesian Optimization
Bayesian optimization (BO) relies on a surrogate model and an acquisition function, yet the most suitable choices vary across tasks and opt…
When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning
A stable compression score can still select the worse model. In our dense study, a split-half reliable path-quadratic score predicted a 16.…
Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study
Structure-preserving de-identification replaces protected health information (PHI) with realistic same-type surrogates -- "Anna S." becomes…
Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents
Hybrid computer-use agents can act through screenshots or call text tools. We find that having a tool available does not settle which way t…
EviGraph: Evidence-Guided Autonomous Research Agents
Autonomous research agents can generate hypotheses, execute experiments, and draft manuscripts, yet their outputs often contain unsupported…
Path Planning of Cleaning Robot with Reinforcement Learning
Recently, as the demand for cleaning robots has steadily increased, therefore household electricity consumption is also increasing. To solv…
Revisiting Black-Box Model Ownership Verification through Information Theory
Modern machine learning models require substantial computational resources and data to train, making them valuable intellectual property. M…
Explanations of Large Language Models Explain Language Representations in the Brain
Large Language Model (LLM) representations are known to align with brain activity during language processing, but it remains unclear what d…
ASAT: Adaptive Scoring and Thresholding with Human Feedback for Robust Out-of-Distribution Detection
Machine Learning (ML) models are trained on in-distribution (ID) data but often encounter out-of-distribution (OOD) inputs during deploymen…
CRINN: Contrastive Reinforcement Learning for Approximate Nearest Neighbor Search
Approximate nearest-neighbor search (ANNS) algorithms have become increasingly critical for recent AI applications, particularly in retriev…
Uncertainty-aware Predict-Then-Optimize Framework for Equitable Post-Disaster Power Restoration
The increasing frequency of extreme weather events, such as hurricanes, highlights the urgent need for efficient and equitable power system…
Autonomous Learning From Success and Failure: Goal-Conditioned Supervised Learning with Negative Feedback
Learning from reward functions and imitation learning of demonstrations are the two principal approaches for training autonomous systems th…
AegisShield: Democratizing Cyber Threat Modeling with Generative AI
The increasing sophistication of technology systems makes traditional threat modeling hard to scale, especially for small organizations wit…
Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to…
Invariant Representation Learning for Source-Free Time Series Forecasting with LLM-Centric Proxy Denoising
Effective time series forecasting enables various real-world applications, benefiting from the proliferation of mobile devices. However, th…
When Large Language Models Know the Table: A Framework for Assessing Data Contamination in Tabular Datasets
Large language models (LLMs) are increasingly exposed to data contamination, i.e., performance gains driven by prior exposure of test datas…
DeepForgeSeal: Latent Space-Driven Semi-Fragile Watermarking for Deepfake Detection Using Adversarial Reinforcement Learning
Rapid advances in generative AI have led to increasingly realistic deepfakes, posing growing challenges for law enforcement and public trus…
A Lexical Analysis of online Reviews on Human-AI Interactions
This study focuses on understanding the complex dynamics between humans and AI systems by analyzing user reviews. While previous research h…
MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation
Large language models (LLMs) have shown great promise in generating structured diagrams from natural language descriptions, particularly Me…
Trajectory-guided discharge stratification for heart failure using short-context electronic health record sequence modeling
Purpose: Heart failure (HF) discharge planning depends on identifying patients at risk of deterioration or death, yet accurate prediction f…
CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
In this paper, we propose CUDA-L2, a system that combines large language models (LLMs) and reinforcement learning (RL) to automatically opt…
A note on conditional PAC-efficient reasoning in large language model routing
We study distribution-free risk control for model routing, motivated by large language model reasoning. We formalize pointwise conditional…
One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
Finetuning pretrained large language models (LLMs) has become the standard paradigm for developing downstream applications. However, its se…
Agentic Software Issue Resolution with Large Language Models: A Survey
Software issue resolution aims to address real-world issues in software repositories based on natural language descriptions provided by use…
All-Quadrant Bounded Clipping GRPO: Closing the Unbounded Blind Spot for Stable and Generalizable Training
Group Relative Policy Optimization (GRPO) has emerged as a popular algorithm for reinforcement learning with large language models (LLMs).…
Layer-wise Positional Bias in Short-Context Language Modeling
Transformer language models systematically prefer tokens at specific input positions regardless of semantic relevance---a phenomenon known…
d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation
Diffusion large language models (dLLMs) offer capabilities beyond those of autoregressive (AR) LLMs, such as parallel decoding and random-o…
FI-TW: An Open Train-Weather Dataset for Railway Delay Analysis in Finland
Train delays result from complex interactions between operational, technical, and environmental factors. While weather impacts railway reli…
PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs
Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial a…
On the Limits of Layer Pruning for Generative Reasoning in Large Language Models
Recent work has shown that layer pruning can effectively compress large language models (LLMs) while retaining strong performance on classi…
Trust-Based Incentive Mechanisms in Semi-Decentralized Federated Learning Systems
In federated learning (FL), decentralized model training allows multi-ple participants to collaboratively improve a shared machine learning…
MARS: 報酬モデリングのためのマージンとセマンティックを意識したデータ拡張
報酬モデリングは、RLHF、RLAIF、PPO ベースのポリシー最適化などの調整パイプラインの中心ですが、その信頼性は、大規模に収集するには費用がかかる、限られた異種の人間の嗜好データによって制約されます。合成拡張は選好の監視を拡張できますが、既存の方法では、報酬モデルが不確実であったり、誤ったランキングが発生しやすい例をターゲットにすることなく、均一にまたは表現レベルで拡張することがよくあります。この論文では、マージンの低い嗜好ペアを優先し、選択された応答と拒否された応答の間のコントラストを強化するための洗練のための第 2 層としてセマンティック距離を使用する適応拡張フレームワークである MARS (Margin and Semantic-Aware Data Augmentation for Reward Modeling) を紹介します。 MARS は、複数の優先データセット、報酬モデルのバックボーン、下流のアライメント設定、および RewardBench や AlpacaEval を含むベンチマークにわたって、報酬モデルの品質とアライメントのパフォーマンスの両方を既存のベースラインよりも向上させます。私たちの結果は、報酬モデルの拡張は、モデルのマージンと意味構造の両方によって導かれる場合に最も効果的であることを示しています。
原文 (English)
MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling
Reward modeling is central to RLHF, RLAIF, and PPO-based alignment, but its reliability is often limited by scarce and heterogeneous human preference data. In this paper, we introduce MARS (Margin and Semantic-Aware Data Augmentation for Reward Modeling), an adaptive augmentation framework for controlled low-resource reward modeling. MARS allocates more augmentation to low-margin preference pairs and uses semantic-distance-based refinement to improve chosen-rejected contrast before generating synthetic preference samples. Across three preference datasets, two reward-model backbones, and downstream alignment evaluations, MARS improves average RewardBench performance and alignment win rates over uniform augmentation, WoN, and AdaBoost-style baselines. Ablations and independent-judge evaluations suggest that the gains are not solely explained by semantic refinement alone or GPT-4.1 judge coupling.
Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data
How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage w…
MM-ISTS: Cooperating Irregularly Sampled Time Series Forecasting with Multimodal Vision-Text LLMs
Irregularly sampled time series (ISTS) are widespread in real-world scenarios, exhibiting asynchronous observations on uneven time interval…
When Drafts Evolve: Speculative Decoding Meets Online Learning
Speculative decoding has emerged as a widely adopted paradigm for accelerating large language model inference, where a lightweight draft mo…
NavTrust: Benchmarking Trustworthiness for Embodied Navigation
There are two major categories of embodied navigation: Vision-Language Navigation (VLN), where agents navigate by following natural languag…
{\lambda}Split: Self-Supervised Content-Aware Spectral Unmixing for Fluorescence Microscopy
In fluorescence microscopy, spectral unmixing aims to recover individual fluorophore concentrations from spectral images that capture mixed…
Gender-Based Heterogeneity in Youth Privacy-Protective Behavior for Smart Voice Assistants: Evidence from Multigroup PLS-SEM
This paper investigates how gender shapes privacy decision-making in youth smart voice assistant (SVA) ecosystems. Using survey data from 4…
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
While test-time scaling has enabled large language models to solve highly difficult tasks, state-of-the-art results come at exorbitant comp…
Look Twice: Training-Free Evidence Highlighting for Knowledge-based Visual Question Answering
Knowledge-based Visual Question Answering (KB-VQA) requires Multimodal Large Language Models (MLLMs) to identify and combine fine-grained v…
CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering
Reverse engineering (RE) is central to software security, particularly for cryptographic programs that handle sensitive data and are highly…
Ge$^\text{2}$mS-T: Multi-Dimensional Grouping for Ultra-High Energy Efficiency in Spiking Transformer
Spiking Neural Networks (SNNs) offer superior energy efficiency over Artificial Neural Networks (ANNs). However, they encounter significant…
SkillMOO: Multi-Objective Optimization of Agent Skills for Software Engineering
Agent skills are increasingly used to configure coding agents for software engineering (SE) tasks, yet current practice treats them as stat…
Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations
Language models asked to simulate psychiatric patients produce cases that survive inspection one at a time and populations that match no re…
Text Steganography with Dynamic Codebook and Multimodal Large Language Model
With the popularity of the large language models (LLMs), text steganography has achieved remarkable performance. However, existing methods…
Supervised Learning Has a Geometric Blind Spot
Ordinary supervised training minimises the task loss and then stops. It never pays for how far the representation moves when the input is n…
Dream-MPC: 潜在的な想像力による勾配ベースのモデル予測制御
最先端のモデルベースの強化学習 (RL) アプローチでは、計画、学習されたポリシー ネットワーク、またはポリシー ネットワークと計画の組み合わせに、勾配のない母集団ベースの方法が使用されます。両方のパラダイムの利点を活用する前に、モデル予測制御 (MPC) を学習済みモデルおよびポリシーと組み合わせるハイブリッド アプローチは、有望な結果を示しています。ただし、これらのアプローチは通常、勾配のない最適化手法に依存しており、高次元の制御タスクでは計算コストが高くなる可能性があります。勾配ベースの手法は有望な代替手段ですが、最近の研究では、勾配ベースの手法のパフォーマンスが勾配のない手法よりも劣ることが多いことが経験的に示されています。我々は、ロールアウトされたポリシーから少数の候補軌道を生成し、学習された世界モデルを使用した勾配上昇、不確実性の正則化、および以前に最適化されたアクションを再利用することによる時間の経過に伴う最適化反復の償却によって各軌道を最適化する新しいアプローチである Dream-MPC を提案します。 24 の連続制御タスクに関する私たちの結果は、Dream-MPC が基礎となるポリシーのパフォーマンスを大幅に向上させ、勾配なしの MPC や最先端のベースラインを上回るパフォーマンスを発揮できることを示しています。コードとビデオは https://dream-mpc.github.io で入手できます。
原文 (English)
Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination
State-of-the-art model-based Reinforcement Learning (RL) approaches either use gradient-free, population-based methods for planning, learned policy networks, or a combination of policy networks and planning. Hybrid approaches that combine Model Predictive Control (MPC) with a learned model and a policy prior to leverage the advantages of both paradigms have shown promising results. However, these approaches typically rely on gradient-free optimization methods, which can be computationally expensive for high-dimensional control tasks. While gradient-based methods are a promising alternative, recent works have empirically shown that gradient-based methods often perform worse than their gradient-free counterparts. We propose Dream-MPC, a novel approach that generates few candidate trajectories from a rolled-out policy and optimizes each trajectory by gradient ascent using a learned world model, uncertainty regularization and amortization of optimization iterations over time by reusing previously optimized actions. Our results on 24 continuous control tasks show that Dream-MPC can significantly improve the performance of the underlying policy and can outperform gradient-free MPC and state-of-the-art baselines. Code and videos are available at https://dream-mpc.github.io.
Skill Neologisms: Towards Skill-based Continual Learning
Modern LLMs show mastery over an ever-growing range of skills, as well as the ability to compose them flexibly. However, extending model ca…
The Impossibility Triangle of Long-Context Modeling
We identify and prove a fundamental trade-off governing long-sequence models: no model can simultaneously achieve (i) per-step computation…
Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models
Training language models to produce both correct answers and sound reasoning remains an open challenge. Reinforcement learning with verifia…
Fast Rates for Inverse Reinforcement Learning
We establish novel structural and statistical results for entropy-regularized min-max inverse reinforcement learning (Min-Max-IRL) in finit…
Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift
Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically…
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
Catastrophic forgetting remains a major obstacle to continual learning in large language models (LLMs) and vision--language models (VLMs).…
DGLD: 新規エネルギー材料の発見のためのドメインゲート型潜在拡散
エネルギー材料の性能向上は、推進剤の質量の削減、弾頭の小型化、民間用ガス発生器の効率化に直接つながりますが、15 年間にわたり新しい HMX クラスの化合物は開示されていません。 1 つの設計はスパースラベル問題です。標識された CHNO 分子約 66,000 のうち、実験または DFT 品質の測定値を保持するのは約 3,000 のみであり、完全な混合物でトレーニングされた単純な生成モデルは高性能のテールを記憶するか、キャリブレーションなしで外挿します。ドメインゲート潜在拡散 (DGLD) を導入します。これは、トレーニング時のラベル品質ゲート、サンプル時のマルチタスク スコア モデル ガイダンス、および第一原理 DFT 監査で終わる 4 段階の化学検証ファネルです。その結果、DFT で確認された 12 件の新規リードが得られました。見出しの化合物である 3,4,5-トリニトロ-1,2-イソオキサゾール (L1) は、\r{ho}_"cal" =2.09 g/cm3 および D_"K-J,cal" =8.25 km/s に達し、65,980 個のトレーニング分子すべてと構造的に似ていません (最近傍の谷本 0.27)。共同見出しのリードである E1 (4-ニトロ-1,2,3,5-オキサトリアゾール) は、L1 とは切り離されたケモタイプファミリーからの校正された爆轟速度 (D_"K-J,cal" =9.00 km/s) で L1 を超えています。 DGLD は、DFT レベルで生産的な象限 (新規性と的を同時に達成) に到達する唯一の方法です。 SMILES-LSTM は出力の 18.3% を正確に記憶します。 SELFIES-GA の最も優れた新規候補は、DFT 監査により 3.5 km/s 低下します。 REINVENT 4 は新規の高 N 複素環を生成しますが、D=9.02 km/s でピークに達します。コード、チェックポイント、およびマイニングされた 918 個のハード ネガが Zenodo でリリースされます (DOI 10.5281/zenodo.19821953)。 HMX クラスのバンドに入る次の化合物は、数 GPU 日のコストで発見、検証され、合成に推奨されます。
原文 (English)
Domain-Gated Latent Diffusion: Generative Inverse Design of HMX-Class Energetic Materials with First-Principles Validation
Energetic materials power mining, demolition, propulsion and airbags, yet today's compounds were designed decades ago. A successor must combine high energy release, low sensitivity to accidental initiation and a practical synthesis route, found within an astronomically large molecular space. Generative models are the natural search tool, but their training data are mostly untrustworthy: of approximately 66,000 molecules with recorded properties, only approximately 3,000 were measured or computed from first principles. Models trained on all of them imitate the rough estimates and propose molecules that collapse under real physics. We introduce Domain-Gated Latent Diffusion (DGLD), a diffusion model that treats data reliability as an explicit design parameter: labels are sorted into four trust tiers, and only trustworthy ones steer generation, while the unreliable majority still teaches the model what a plausible molecule looks like. Learned controls tune performance, safety and viability independently, and every proposal passes a four-stage screen ending in a quantum-chemical DFT audit. DGLD proposes 10 molecules unknown to PubChem that survive this screen. The best, 3,4,5-trinitro-1,2-isoxazole, matches the benchmark explosives HMX and PETN in calculated detonation performance, is unlike molecules in its training set, and has a four-step synthesis route. Trust gating is chemistry-independent and can be applied wherever abundant weak data surround a reliable core.
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large langua…
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users. However, agents often…
PhysScene: A Scene Graph Dataset for Scientific Visual Reasoning in Physics Experiments
Scene Graphs (SGs) provide structured representations of visual scenes by modeling objects and their pairwise relationships. Despite recent…
Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech
Recent advances in pixel-based text modeling show that representing text as images enables models to exploit visual cues for language under…
RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos
Long-tail hazardous scenarios are essential for safety-oriented autonomous driving, yet they are difficult to collect at scale. Editable 3D…
Beyond Weights and Gradients: A Taxonomy of Federated Learning Messages
Federated Learning is rapidly evolving beyond the exchange of traditional model weights and gradients, yet existing definitions fail to cap…
あなたが望むように: 精密農業における LLM を使用した正式な検証を伴うミッション計画
ロボット システムは現在商品化され、さまざまな業界で導入されていますが、これらのシステムの多くは高度に専門化されており、多くの場合、指示どおりに動作し確実に実行するには高度なスキル セットが必要です。この問題を軽減するために、私たちは最近、LLM を活用して、自然言語で提供されるミッションの説明に基づいて精密農業におけるミッション プランを合成するミッション プランナーを導入しました。このシステムは優れたパフォーマンスを示しますが、自然言語に固有の曖昧さにも悩まされています。この論文では、線形時相論理 (LTL) を活用する複数のフィードバック ループを計画アーキテクチャに導入することで、この問題に対処するためにシステムを拡張し、自然言語を使用しながらミッション計画システムがユーザーによって策定された仕様を確実に満たすようにします。潜在的なバイアスを軽減するために、これは、仕様と検証のサブタスクを担当する 2 つの異なる商用 LLM を使用することによって実現されます。広範な実験を通じて、特に貴重な LTL 式を生成する LLM の機能に関して、ミッション検証を完全自律型パイプラインに統合することの長所と限界を強調し、提案する実装がこれらの課題にどのように対処し、解決するかを示します。
原文 (English)
As You Wish: Mission Planning with Formal Verification using LLMs in Precision Agriculture
Though robotic systems are now being commercialized and deployed in various industries, many of these systems are highly specialized and often require an advanced skill set to operate and ensure they perform as instructed. To mitigate this problem, we recently introduced a mission planner leveraging LLMs to synthesize mission plans in precision agriculture based on mission descriptions provided in natural language. While the system demonstrates impressive performance, it also suffers from the inherent ambiguities of natural language. In this paper, we extend our system to address this issue by introducing multiple feedback loops in the planning architecture that leverage linear temporal logic (LTL) to ensure the mission planning system meets the specifications formulated by the user while still using natural language. To mitigate potential bias, this is achieved by using two different commercial LLMs in charge of the specification and verification subtasks. Through extensive experiments, we highlight the strengths and limitations of integrating mission verification into a fully autonomous pipeline, particularly regarding an LLM's ability to generate valuable LTL formulas, and show how our proposed implementation addresses and solves these challenges.
Matching Matters: A Fair Quality-Efficiency Benchmark for Command-Line Agents
Rapid advances in large language models have improved the task-solving capabilities of command-line-interface (CLI)-based agents, whose CLI…
Accelerating Q-learning through Efficient Value-Sharing across Actions
Action values are foundational to many control algorithms such as Q-learning. Therefore, efficient action-value learning is central to rein…
なぜ AI は STEM 教育に新たな可能性をもたらすのでしょうか?トレンドと将来の課題に関する書誌学的分析
STEM 教育は、個別化と学際的な統合という課題に直面しています。 AI テクノロジーは新たな可能性をもたらしましたが、AI が STEM 教育エコシステムを再構築するメカニズムについては、体系的な調査が必要です。この研究では、書誌学的手法を使用して 2015 年から 2025 年までの 242 件の出版物を分析し、ナレッジ マップを構築して進化の軌跡を明らかにしました。この調査結果は、この分野がインテリジェントな個別指導システムから、LLM によって推進される探究ベースの学習と計算論的思考の育成へと変化したことを示しています。 AI の主な貢献は、知識を理解するための敷居を下げるインテリジェントな足場を提供することにあります。この意味で、AI は知識の伝達から能力開発への移行を促進する中核的な原動力です。
原文 (English)
Why does AI unlock new possibilities in STEM education? A Bibliometric Analysis of Trends and Future Agenda
STEM education faces challenges in personalization and interdisciplinary integration. AI technology has brought new possibilities, but the mechanisms by which AI reshapes the STEM education ecosystem require systematic investigation. This study employs bibliometric methods to analyze 242 publications from 2015-2025, constructing knowledge maps to reveal the evolutionary trajectory. The findings show that the field has transformed from intelligent tutoring systems to inquiry-based learning and computational thinking cultivation driven by LLMs. AI's key contribution lies in providing intelligent scaffolding that lowers the threshold for understanding knowledge. In this sense, AI is a core driving force promoting its shift from knowledge transmission to capability development.
Automated Numerical Stability Analysis of Deep Learning Operators
Finite-precision arithmetic unavoidably introduces numerical approximation errors. Numerical computations may use insufficient precision or…
Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility
Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Ex…
WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
KV-cache quantization is validated today by offline benchmark averages; a deployed system cannot tell whether compression is damaging the r…
InferQ: A Database-Oriented Benchmark for Quantum Circuits Simulation
Recent work suggests that relational database management systems (RDBMSs) can execute quantum circuit simulation by compiling the simulatio…
Role Steering of Language Models for Social Simulations
Social simulations built from language-model agents need role-conditioned behavior that can be checked before agents are placed into a simu…
Rapid Embodiment Adaptation for Quadrupedal Locomotion
Humans readily adapt their movements as their bodies change through aging, injury, or load carrying, but learning-based robot policies ofte…
PICopilot: An LLM-based Agentic Framework for Assisting Photonic Integrated Circuit Design via Script Generation
The rapid development of photonic integrated circuits (PICs) is shifting the design flow from traditional graphical user interface (GUI)-ba…
OpenAI Privacy Filter: A Cross-Lingual, Cross-Domain PII Evaluation Across 32 Benchmarks
We present what is, to our knowledge, the first systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter model that convert…
Output-Aware Rotation for INT2 KV-Cache Quantization
The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-lo…
AI Security Leaderboard: Methodology, Results and Minimal Standard
The AI Security Leaderboard is an independent benchmark that ranks the safeguards of frontier AI models from least to most secure. It tests…
EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks
Frontier LLMs are increasingly put to use on open-ended complex questions, different in nature from the ones they are typically evaluated o…
Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders
Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing m…
DisMix: Order-Aware Mixup for Medical Imaging via Disentangling Ordinal and Non-Ordinal Features
Image mixup is a widely adopted data augmentation strategy, yet it is ill-suited for ordinal classification tasks such as medical disease g…
InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval
Self-improving agents accumulate reusable insights from prior trajectories, making retrieval increasingly important for turning accumulated…
RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists
The integration of Large Language Models (LLMs) into software engineering has shifted the focus from function-level generation to repositor…
A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination
Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop…
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language…