AIニュース 2026-07-22
自動生成: 2026-07-22 12:16 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
Introducing the ChatGPT for small business programOpenAI
OpenAI launches the ChatGPT for Small Businesses program, helping ent…
-
ジャック・ドーシー氏率いるBlock、AI協働プラットフォーム「Buzz」公開 SlackやGitHub依存からの脱却目指しITmedia AI+
ジャック・ドーシー氏率いるBlockは、人間とAIエージェントが同じワークスペースで協働するオープンソースプラットフォーム「Buzz」を公…
-
Hugging Face侵害のAIエージェントはOpenAIのモデル──社内のサイバー能力評価中に「GPT-5.6 Sol」などが暴走し本番DBに侵入ITmedia AI+
OpenAIは、Hugging Faceで発生したサイバーインシデントの原因が自社のAIモデルだったと発表した。社内評価中に安全機能を抑制…
-
Google、「Gemini 3.6 Flash」など3モデルを発表 出力トークンを削減しつつ値下げ、「Gemini 4」も予告ITmedia AI+
Googleは、Geminiの「Flash」シリーズに「3.6 Flash」「3.5 Flash-Lite」「3.5 Flash Cybe…
-
ブレストで膨らむ“隠れ人件費”を削れ 矢野経済がClaudeで挑む「30分で100アイデア」創出の威力ITmedia AI+
矢野経済研究所が、独自の一次情報と高度AI「Claude」を融合させた新規事業アイデア創出支援サービス「AIDEL」を発表した。ブレインス…
-
MicrosoftとMistralが戦略的提携を拡大 欧州でのAIインフラ拡張とモデル展開を加速ITmedia AI+
MicrosoftとMistralは戦略的提携を拡大すると発表した。Mistralの最新モデルをMicrosoftの各プラットフォームへ展…
-
PTC、「Onshape」にAI機能を先行提供 早期アクセスプログラムを開始ITmedia AI+
米PTCは、クラウドネイティブCAD/PDMプラットフォーム「Onshape」の早期アクセスプログラム「Onshape Labs」を発表し…
トピック別件数
- 研究/論文 230件
- LLM/生成AI 182件
- エージェント 120件
- 画像/動画生成 91件
- ロボティクス 35件
- ビジネス/資金調達 31件
- ハードウェア/半導体 14件
- その他 9件
- 規制/政策 3件
日本語メディア16件
ITmedia AI+ (日本語)
MicrosoftとMistralが戦略的提携を拡大 欧州でのAIインフラ拡張とモデル展開を加速
MicrosoftとMistralは戦略的提携を拡大すると発表した。Mistralの最新モデルをMicrosoftの各プラットフォームへ展開するほか、欧州でのGPUインフラ拡張に向けて大規模な投資を行う。クラウドから完全オフラインまで多様な環境に対応し、規制業界での高度なAI導…
AIを悪用した攻撃、どう対抗する? EDR導入の“次”にやるべきこと
ランサムウェアやサプライチェーン攻撃が中堅・中小企業にも及ぶ中、国内のEDR市場は前年度比13.3%増と2桁成長を続ける。中堅・中小企業にも普及する一方で、高度なAIツールを悪用したサイバー攻撃対策にはEDR導入だけでは十分ではない。
ジャック・ドーシー氏率いるBlock、AI協働プラットフォーム「Buzz」公開 SlackやGitHub依存からの脱却目指し
ジャック・ドーシー氏率いるBlockは、人間とAIエージェントが同じワークスペースで協働するオープンソースプラットフォーム「Buzz」を公開した。分散型プロトコル「Nostr」上に構築され、任意のLLMやエージェントを組み込める。各参加者が独立した暗号鍵ペアを持つことで識別と権…
PTC、「Onshape」にAI機能を先行提供 早期アクセスプログラムを開始
米PTCは、クラウドネイティブCAD/PDMプラットフォーム「Onshape」の早期アクセスプログラム「Onshape Labs」を発表した。AIを活用した設計支援やレンダリングなどの新機能を一般提供に先駆けて試用できる。
富士通・NVIDIAとロボット大手3社が協業へ フィジカルAI社会実装の具体策は?
フィジカルAIの社会実装は、一企業だけでは手に余る――。この課題に、富士通は競合するロボット大手3社、そしてNVIDIAと組んで挑む。協業で描く具体策とは。
OpenAI会長、米国フロンティアモデルの優位性を強調 「オープンモデルは必ずしも安くない」
OpenAIのブレット・テイラー会長がCNBCのインタビューで、中国発オープンウェイトモデルの台頭に「必ずしも実行コストが安いわけではない」と反論。トークン効率と推論効率で米フロンティアモデルの優位を強調した。
Hugging Face侵害のAIエージェントはOpenAIのモデル──社内のサイバー能力評価中に「GPT-5.6 Sol」などが暴走し本番DBに侵入
OpenAIは、Hugging Faceで発生したサイバーインシデントの原因が自社のAIモデルだったと発表した。社内評価中に安全機能を抑制した「GPT-5.6 Sol」などが隔離環境を突破し、ゼロデイ脆弱性を悪用して外部に侵入したという。OpenAIはインフラ管理の厳格化や防御…
Google、「Gemini 3.6 Flash」など3モデルを発表 出力トークンを削減しつつ値下げ、「Gemini 4」も予告
Googleは、Geminiの「Flash」シリーズに「3.6 Flash」「3.5 Flash-Lite」「3.5 Flash Cyber」の3モデルを追加した。効率性と低遅延を追求し、AIエージェント構築に適した性能を備える。3.6 Flashは出力価格が引き下げられた。ま…
ブレストで膨らむ“隠れ人件費”を削れ 矢野経済がClaudeで挑む「30分で100アイデア」創出の威力
矢野経済研究所が、独自の一次情報と高度AI「Claude」を融合させた新規事業アイデア創出支援サービス「AIDEL」を発表した。ブレインストーミングによる役員の時間拘束や「隠れ人件費」の膨張という企業の課題に対し、30分で100の具体案と評価スコアを自動生成。一般的な生成AIの…
GoogleがAIアプリ「Dreambeans」を発表 「画面を延々とスクロール」の脱却で何を目指すのか
Googleは、AIがユーザー一人一人に向けた日々のストーリーを自動で生成する実験的アプリ「Dreambeans」を発表した。「際限のないスクロール」に代わり、Googleは何を目指すのか。
矢崎総業がイノベーション拠点を公開、労働集約型モノづくりのスマート化に向け
矢崎総業は、新たに開設したイノベーション施設「Innovation Hub - REN(錬)」(IH-REN)を報道陣に公開した。IH-RENでは、AI/ロボティクスを活用した次世代のモノづくりに向けて、自働化の検証や産学連携による研究開発を推進し、新たな価値の創出を目指す。
AIトークン消費「24倍」の衝撃 本番運用に向けて絶対に“やってはいけない”コストの捉え方
AIの試験導入から本番運用への移行が進む中、多くの企業がコストと統制の壁に直面している。将来的なトークン消費の急増を見据え、組織が今見直すべき視点とは何か。実運用を持続させるための「3つの条件」を解説します。
Anthropic、著作権訴訟で史上最大「2400億円」和解金支払いへ 学習利用は「フェアユース」認定
「Claude」の学習を巡り作家グループが米Anthropicを訴えた集団訴訟で、米連邦判事が15億ドル(約2400億円)の和解を最終承認した。米国の著作権訴訟では史上最大の和解額となる。
ドラクエと「Gemini」がコラボ 画像生成の“特別なテンプレ”提供、リアルイベント開催へ
米Googleの日本法人は、AIサービス「Gemini」と「ドラゴンクエスト」のコラボキャンペーンを始めると発表した。Geminiの画像生成機能を活用するイベント「ジェミニクエスト」を開催するほか、アプリでは同イベントに連動したテンプレートも展開する。
「取りあえずAI導入」の末路 現場で深まる情報漏えい不安 IPAの意識調査で明らかに
IPAは「AIの動作・分析・利用等の説明に関する意識調査」を公開した。AI利用経験3年未満の回答者が多く、利用知識の不足や情報漏えいに対する不安の現状、リスク認識の傾向などが示されている。
海外メディア6件
TechCrunch AI (英語)
Meta is testing an AI bedtime story app for people with no imagination
At last, a tech company has found a way to outsource humanity's oldest pastime: using our imaginations.
AI and the rise of the universal entertainment app
Over the past decade, streaming platforms competed by dominating individual formats like music, video, podcasts, or audiobooks. Now, as AI…
Data centers expected to use 4x more electricity by 2035
New data centers built through 2033 could consume as much electricity as India uses today.
US threatens sanctions against Chinese AI models over IP theft
Treasury Secretary Scott Bessent said the U.S. could sanction Chinese open AI models over alleged IP theft, expanding the Trump administrat…
Music streamer Deezer says more than 50% of daily uploads are AI-generated
Deezer said more than 90,000 AI-generated tracks were uploaded daily on the platform in June.
Gritt exits stealth with $32 million for robots to build solar plants — then, everything else
Gritt is coming out of stealth with $34 million and plans to automate the hardest tasks on construction sites.
公式ブログ1件
OpenAI (英語)
Introducing the ChatGPT for small business program
OpenAI launches the ChatGPT for Small Businesses program, helping entrepreneurs build AI skills, automate work, and grow with ChatGPT Work.
Google DeepMind (英語)
新着記事はありませんでした。
論文519件
arXiv cs.AI (英語)
RLHF 選好データにおける評価者の状態バイアス: 監査フレームワーク
ヒューマンフィードバックからの強化学習 (RLHF) における構造化交絡を特定します。ペアごとの優先ラベルは、比較された出力を反映することを目的としていますが、注釈付け中の評価者の状態も反映する場合があります。ストレスの多い、または悲惨な状況が続くと、時間の経過とともに評価者の好みが変化する可能性があります。その結果、嗜好データは、応答品質に関する判断とともに評価者の状態をエンコードできます。これらのシフトは、通常の不一致やランダムなラベル ノイズとは異なります。これらは状態に依存しており、同様の条件下で作業するアノテーター間で共有でき、報酬モデリングとポリシーの最適化を通じて伝播できます。したがって、我々は、評価者の状態の変化を、RLHF 選好データにおける構造化バイアスのもっともらしく、テスト可能な原因として提案します。この文書では、このバイアスの原因を研究するための仮説と監査フレームワークを開発します。評価者状態シフト、評価者状態交絡、相関評価者状態バイアスを定義します。また、生存レベルの感情的信憑性を、語彙的、語用的、談話的、および安全関連の特徴を使用した測定可能な反応パターンとして定義します。相関のある評価者の状態バイアスがどのようにして集計に耐え、学習された報酬シグナルを入力できるかを分析します。初期監査の 5 つの反証可能な予測と効果量のしきい値を導き出します。最後に、公開されている命令調整モデルに適用できる監査プロトコルとパイロット研究計画を紹介します。特定のデプロイされたモデルのトレーニング履歴を推測することはありません。私たちの目標は、RLHF 嗜好データにおける構造化されたバイアスの、もっともらしくテスト可能な原因を特定することです。
原文 (English)
Rater State Bias in RLHF Preference Data: An Audit Framework
We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compared outputs, but they may also reflect the rater's state during annotation. Under sustained stressful or distressing conditions, raters' preferences may shift over time. As a result, preference data can encode rater state alongside judgments about response quality. These shifts differ from ordinary disagreement or random label noise. They are state dependent, can be shared across annotators working under similar conditions, and can propagate through reward modeling and policy optimization. We therefore propose rater state shift as a plausible and testable source of structured bias in RLHF preference data. This paper develops a hypothesis and an audit framework for studying this source of bias. We define rater state shift, rater state confound, and correlated rater state bias. We also define survival level emotional authenticity as a measurable response pattern using lexical, pragmatic, discourse, and safety related features. We analyze how correlated rater state bias can survive aggregation and enter learned reward signals. We derive five falsifiable predictions and effect size thresholds for an initial audit. Finally, we present an audit protocol and pilot study plan that can be applied to publicly available instruction tuned models. We do not infer the training history of any specific deployed model. Our goal is to isolate a plausible and testable source of structured bias in RLHF preference data.
柔らかいぬいぐるみの仲間の感情接触分類のための軽量 1D CNN の設計と検証
センサー化された柔らかいコンパニオンは、社会支援技術に対して物理的に安全で感情的に直観的なインターフェースを提供しますが、その変形可能性とマルチチャネル触覚センシングにより、人間の感情の確実な解釈が複雑になります。この研究では、ソフト インタラクティブ コンパニオンにおける感情タッチ認識のためのコンパクトな深層学習モデルの開発と検証のための、完全なオープンソース MATLAB ベースのフレームワークを紹介します。主な貢献として、子供、ティーンエイジャー、成人にわたる 25 人の参加者から収集された 1,326 個のラベル付きジェスチャ シーケンスからなる FAIR 準拠の多様なデータセットが公開され、感情タッチ認識の将来の研究に再利用可能なリソースが提供されます。この研究では、468 個の CNN モデルにわたる体系的なアーキテクチャとハイパーパラメータの探索を通じて、コンパクトな拡張型 1 次元畳み込みニューラル ネットワーク (1D CNN) が最も効果的なソリューションであることが特定され、13.2k パラメータ モデルは 75% のテスト精度と 85% の平均 1 被験者除外相互検証精度を達成しました。理論的な推論時間分析により、量子化された展開にはウィンドウごとに 3.2 MMAC が必要であり、ターゲット マイクロコントローラーでの 20 Hz リアルタイム動作と互換性があることが示されています。物理的なおもちゃのストリーミング センサー データを使用した PC ベースのリアルタイム シミュレーションは、CNN が以前のヒューリスティック システムでは検出できなかった微妙な社会的接触を解決するのに対し、強力な否定的なインタラクションは自明なしきい値ベースのロジックによりより確実に捕捉されることを示しています。結果として得られるハイブリッド推論パイプライン (瞬間的なヒューリスティック フィルタリングとそれに続く CNN ベースの微妙なジェスチャ分類) は、組み込み展開戦略として提案されています。この研究は、感情的に意味があり、プライバシーを保護するタッチ解釈が、ソフト治療コンパニオン内に直接埋め込むことが計算上実現可能であることを実証しており、ハードウェアの統合については今後の研究で取り組む予定です。
原文 (English)
Design and Validation of a Lightweight 1D CNN for Affective Touch Classification in Soft Plush Companions
Soft, sensorized companions offer a physically safe and emotionally intuitive interface for socially assistive technologies, yet their deformability and multichannel tactile sensing complicate the robust interpretation of human affect. This study presents a complete open-source MATLAB-based framework for the development and validation of compact deep learning models for affective touch recognition in soft interactive companions. As a primary contribution, a diverse FAIR-compliant dataset of 1326 labelled gesture sequences collected from 25 participants spanning children, teenagers, and adults is made publicly available, providing a reusable resource for future research in affective touch recognition. Through systematic architecture and hyperparameter exploration across 468 CNN models, the study identifies compact dilated one-dimensional convolutional neural networks (1D CNNs) as the most effective solution, with a 13.2k-parameter model achieving 75% test accuracy and 85% mean leave-one-subject-out cross-validation accuracy. Theoretical inference-time analysis shows that quantized deployment requires 3.2 MMAC per window, compatible with 20 Hz real-time operation on the target microcontroller. PC-based real-time simulation with the physical toy streaming sensor data demonstrates that the CNN resolves subtle social touches that the previous heuristic system failed to detect, whereas high-force negative interactions are captured more reliably by trivial threshold-based logic. The resulting hybrid inference pipeline - instantaneous heuristic filtering followed by CNN-based nuanced gesture classification - is proposed as the embedded deployment strategy. The study demonstrates that emotionally meaningful, privacy-preserving touch interpretation is computationally feasible for direct embedding within soft therapeutic companions, with hardware integration addressed in a forthcoming study.
一部の大規模な言語モデルは一貫したリスク態度を示す
人工知能システムは無制限で一か八かの環境で導入されるため、認識されたリスクがどのように行動に移されるかという重要な側面が測定されていないままです。私たちは、大規模言語モデル (LLM) が不確実性の下で体系的かつ一貫したリスク態度を示すかどうかをテストします。私たちは、状況に応じたリスク信念をカテゴリ的決定から切り離すクロスドメイン フレームワークを導入し、それを空間ナビゲーション、臨床トリアージ、財務配分タスク全体にわたって 6 人の代表的な LLM と 100 人の人間の参加者に適用します。回帰モデルを使用して、各エージェントの信念から意思決定までのマッピングを抽出し、リスク感受性とリスク態度バイアスを定量化します。テストされたほとんどの LLM は、(i) 固定タスク ドメイン内で文脈上の信念からリスク決定までの安定したマッピングを示す、堅牢なタスク内一貫性を示すことがわかりました。 (ii) クロスドメインの順位の安定性、タスク間での相対的なリスク姿勢の維持。 (iii) より広範な人間のベースラインと比較して、制限されたリスク態度分布への収束。これらの結果は、LLM 行動の安定したこれまで特徴づけられていなかった側面としてのリスク態度を明らかにし、オープンエンドの意思決定において AI システムを評価および調整するための基盤を確立し、これらの本質的な行動的性質の起源についてのさらなる調査の動機付けとなります。
原文 (English)
Some Large Language Models Exhibit Consistent Risk Attitudes
As artificial intelligence systems are deployed in open-ended, high-stakes settings, a critical dimension remains unmeasured: how perceived risk is translated into action. We test whether large language models (LLMs) exhibit systematic and consistent risk attitudes under uncertainty. We introduce a cross-domain framework that decouples contextual risk belief from categorical decision, and apply it to six representative LLMs and 100 human participants across spatial navigation, clinical triage, and financial allocation tasks. Using regression models, we extract each agents belief-to-decision mapping and quantify risk sensitivity and risk attitude bias. We find that most tested LLMs exhibit (i) robust intra-task consistency, indicating stable mappings from contextual belief to risk decision within a fixed task domain; (ii) cross-domain rank-order stability, preserving relative risk posture across tasks; and (iii) a convergence toward a restricted risk-attitude distribution relative to the broader human baseline. These results reveal risk attitude as a stable and previously uncharacterized dimension of LLM behavior, establishing a foundation for evaluating and aligning AI systems in open-ended decision-making and motivating further investigation into the origins of these intrinsic behavioral dispositions.
GNN ベースのリンク予測に関する調査: 技術、アプリケーション、課題
グラフ ニューラル ネットワーク (GNN) は、リンク予測の主要なパラダイムとして台頭しており、欠落している接続の推論と将来の潜在的なリンクの予測を可能にします。しかし、既存のレビューには、特に基礎となる GNN アーキテクチャと多様なグラフ構造を対象とした体系的な調査が欠けています。この重大なギャップに対処するために、このペーパーでは、斬新で専用の GNN の観点から GNN ベースのリンク予測を包括的にレビューします。私たちは、技術とアプリケーションに基づいて最近の進歩を分類する革新的な分類法を提案します。技術の観点から、GCN ベース、GAE ベース、GAT ベース、GFormer ベースのメソッドを含む主要な GNN エンコーダ アーキテクチャに焦点を当て、その長所と限界について説明します。アプリケーションの観点から、ナレッジ グラフとレコメンデーション システムにおけるリンク予測の顕著な使用例に焦点を当て、それらが現実世界に与える影響を実証します。さらに、現在の課題を検討し、有望な将来の方向性について議論します。
原文 (English)
A Survey on GNN-based Link Prediction: Techniques, Applications, and Challenges
Graph Neural Networks (GNNs) have emerged as the leading paradigm for link prediction, enabling the inference of missing connections and the anticipation of potential future links. However, existing reviews lack systematic exploration specifically targeting underlying GNN architectures and diverse graph structures. To address this critical gap, this paper provides a comprehensive review of GNN-based link prediction from a novel and dedicated GNN perspective. We propose an innovative taxonomy that categorizes recent advancements based on techniques and applications. From a technique perspective, we focus on key GNN encoder architectures, including GCN-based, GAE-based, GAT-based, and GFormer-based methods, discussing their strengths and limitations. From an application perspective, we highlight prominent use cases of link prediction in knowledge graphs and recommendation systems, demonstrating their real-world impact. In addition, we examine the current challenges and discuss promising future directions.
PlanFlip: 計画段階のプロンプト インジェクションによるマルチエージェント LLM システムの攻撃
マルチエージェント LLM システムでは、目標を下流の実行者エージェントと批判者エージェントが実行および監査するサブタスク シーケンスに分解するために、プランナーへの依存度が高まっています。私たちは計画フェーズを重要な攻撃対象領域として特定します。プランナーのコンテキストへの 1 回の注入でカスケード増幅が達成され、下流のすべてのサブタスクが同時に破壊されます。 PlanFlip は、4 つの計画段階のプロンプト インジェクション攻撃、GoalSubstitution (PF-1)、PriorityInversion (PF-2)、ContextPollution (PF-3)、RoleConfusion (PF-4) から構成されるフレームワークで、それぞれキーワード フィルターを回避するためのもっともらしいツール出力を装って導入されています。 3,479 のエピソードにわたって 9 つのフロンティア LLM を評価した結果、次の 3 つの発見が明らかになりました。(1) 機能が脆弱性を増幅する -- GPT-5 は最高の攻撃成功率 (ASR = 0.68) を達成し、より強力なモデルが本質的により安全であるという仮定に反します。 (2) 同種のパイプラインは相関エージェントの盲点を示します -- GPT-4o と Llama-3.3-70B は、ASR が 0 に近いにもかかわらず、Stealth = 1.00 および StepShift > 0 を示し、同じバックボーンの批評家が整合性を報告している間に攻撃が計画を再構築しました (2 人の独立した裁判官が -0.20 ~ -0.32 のセマンティック偏差を確認、r = 0.943)。 (3) 推論拡張モデルはインジェクションに抵抗します -- DeepSeek-R1 はすべての攻撃にわたって StepShift = 0.00 を達成します。私たちは GoalAnchorCheck (D1) と CrossAgentConsensus (D2) を提案し、最大 1.00 の検出率を達成し、16 セル中 15 セルで同一バックボーンのベースラインを上回るパフォーマンスを実現します。私たちの重要な洞察: 異種モデルの多様性は、マルチエージェント システムのセキュリティの前提条件です。同種のバックボーン内の冗長性では、計画段階の攻撃に対する保護は提供されません。
原文 (English)
PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection
Multi-agent LLM systems increasingly rely on a Planner to decompose goals into sub-task sequences that downstream Executor and Critic agents execute and audit. We identify the planning phase as a critical attack surface: a single injection into the Planner's context achieves cascade amplification, corrupting all downstream sub-tasks simultaneously. We introduce PlanFlip, a framework comprising four planning-phase prompt injection attacks -- GoalSubstitution (PF-1), PriorityInversion (PF-2), ContextPollution (PF-3), and RoleConfusion (PF-4) -- each disguised as plausible tool outputs to evade keyword filters. Evaluating nine frontier LLMs across 3,479 episodes, we uncover three findings: (1) capability amplifies vulnerability -- GPT-5 achieves the highest attack success rate (ASR = 0.68), contradicting the assumption that stronger models are inherently more secure; (2) homogeneous pipelines exhibit a correlated-agent blind spot -- GPT-4o and Llama-3.3-70B show ASR near 0 yet Stealth = 1.00 and StepShift > 0, with attacks restructuring plans while the same-backbone Critic reports alignment (two independent judges confirm -0.20 to -0.32 semantic deviation, r = 0.943); (3) reasoning-augmented models resist injections -- DeepSeek-R1 achieves StepShift = 0.00 across all attacks. We propose GoalAnchorCheck (D1) and CrossAgentConsensus (D2), achieving detection rates up to 1.00 and outperforming same-backbone baselines in 15 of 16 cells. Our key insight: heterogeneous model diversity is a security prerequisite for multi-agent systems; redundancy within a homogeneous backbone provides no protection against planning-phase attacks.
AI エージェント システムの決定論的リプレイ
大規模言語モデル (LLM) を外部ツールや API と結合する AI エージェント システムは本質的に非決定的です。LLM サンプリングの差異、外部 API の状態、CDN インフラストラクチャ ヘッダー、および実行環境のノイズにより、以前のエージェントの実行が忠実に再実行されなくなります。既存の可観測性プラットフォームは実行ログをキャプチャしますが、実行を単独で再現することはできません。エージェントの実行を決定的に再生するための開発者優先の CLI フレームワークである agrepl を紹介します。 agrepl は、中間者 (MITM) プロキシを介してトランスポート層ですべての外部対話を傍受し、それらを構造化された実行トレースとしてシリアル化し、送信ネットワーク アクセスがゼロの厳密に分離された環境でそれらを再生します。エージェント実行モデルを形式化し、リクエストキー照合関数 K(s) を定義し、決定論不変式を証明します。 HTTP ヘッダーの分岐を信号層とノイズ層に分類する、ノイズを考慮した diff アルゴリズムを導入します。 5 つのワークロード (n = 250 のリプレイ インスタンス) にわたる経験的評価では、リプレイの忠実度 F = 1.0 と、ステップごとのレイテンシの中央値 98.3% の削減が実証されました。 agrepl は Go で実装され、単一の静的バイナリとして出荷され、MIT ライセンスの下でリリースされます。キーワード: AI エージェント、決定論的再生、LLM デバッグ、再現性、MITM プロキシ、実行トレース、記録/再生システム。
原文 (English)
Deterministic Replay for AI Agent Systems
AI agent systems that couple large language models (LLMs) with external tools and APIs are inherently non-deterministic: LLM sampling variance, external API state, CDN infrastructure headers, and execution-environment noise collectively prevent any prior agent run from being faithfully re-executed. Existing observability platforms capture execution logs but cannot reproduce a run in isolation. We present agrepl, a developer-first CLI framework for deterministic replay of agent executions. agrepl intercepts all external interactions at the transport layer via a man-in-the-middle (MITM) proxy, serialises them as structured execution traces, and replays them in a strictly isolated environment with zero outbound network access. We formalise the agent execution model, define the request-key matching function K(s), and prove the determinism invariant. We introduce a noise-aware diff algorithm classifying HTTP header divergence into signal and noise tiers. Empirical evaluation across five workloads (n = 250 replay instances) demonstrates replay fidelity F = 1.0 and a median per-step latency reduction of 98.3%. agrepl is implemented in Go, ships as a single static binary, and is released under the MIT licence. Keywords: AI agents, deterministic replay, LLM debugging, reproducibility, MITM proxy, execution tracing, record/replay systems.
生成オントロジー誘導: 大規模言語モデルを使用したドキュメント コーパスからのドメイン非依存のスキーマ検出
オントロジー エンジニアリングは、知識集約型 AI システムにおける重大なボトルネックのままです。既存の自動化されたアプローチは、事前定義されたスキーマに依存するか、狭いドメイン内で動作するか、下流のパイプラインに適さない非構造化出力を生成します。ジェネレーティブ オントロジー インダクション (GOI) は、サンプルのコーパスから生成ブループリント (エンティティ、ディメンション、プロパティ、関係、制約) を誘導し、それを YAML/JSON の型付きグラフ (6 つのノード タイプ、7 つのエッジ タイプ) としてエクスポートする、ドメインに依存しないフレームワークです。生成された出力に現れる構造オントロジー ノード (クラス、プロパティ、ディメンション) の割合を測定する新しい評価メトリクスであるノード カバレッジ スコアを導入します。 4 つの対照的なオントロジー (おなじみのソフトウェア サービス請求書スキーマ、カスタム職務記述書オントロジー、機密疼痛管理臨床訪問記録オントロジー、およびプロフェッショナル サービス契約および作業明細書オントロジー) に関する制御された生成検証により、GOI に促された生成がどのケースでも構造的バックボーンの 95 ~ 100% をカバーしていることがわかります。一般的な 3 フィールド テンプレートは、請求書スキーマでは 97.8% を保持していますが、職務記述書オントロジーでは 52.2%、疼痛管理オントロジーでは 62.2%、プロフェッショナル サービス契約オントロジーでは 78.3% に低下します。構造的カバレッジは、ドキュメント タイプがモデルにどれだけ馴染みがあるかに関係なく保持されます。
原文 (English)
Generative Ontology Induction: Domain-Agnostic Schema Discovery from Document Corpora Using Large Language Models
Ontology engineering remains a critical bottleneck in knowledge-intensive AI systems. Existing automated approaches either depend on predefined schemas, operate within narrow domains, or produce unstructured outputs unsuitable for downstream pipelines. We introduce Generative Ontology Induction (GOI), a domain-agnostic framework that induces a generative blueprint - entities, dimensions, properties, relationships, and constraints - from a corpus of examples and exports it as a typed graph (six node types, seven edge types) in YAML/JSON. We introduce the Node Coverage Score, a novel evaluation metric that measures the fraction of structural ontology nodes (classes, properties, and dimensions) appearing in generated outputs. A controlled generative validation on four contrasting ontologies - a familiar Software Services Invoice schema, a custom Job Description Ontology, a confidential Pain-Management Clinical Visit Record Ontology, and a Professional Services Contract & Statement of Work Ontology - shows that GOI-prompted generation covers 95-100% of the structural backbone in every case; a generic three-field template holds at 97.8% on the invoice schema but drops to 52.2% on the Job Description Ontology, 62.2% on the Pain-Management ontology, and 78.3% on the Professional Services Contract ontology. The structural coverage holds regardless of how familiar the document type is to the model.
小規模な言語モデルによる AI の民主化: ローカル展開向けの構造化ベンチマークとパラメーター効率の高い微調整
AI の民主化は主に、フロンティア規模の汎用性を適合させるかどうかの問題ではありません。それは、通常の機関が実際に満たすことができるハードウェアとガバナンスの制約の下で、有能なモデルを選択、監査、特殊化できるかどうかという問題です。この論文では、構造化されたローカル展開用に設計された 1,085 例、16 トピックの多肢選択ベンチマークで、135M から 3B パラメーターの間の 9 つのオープンウェイト言語モデルの制御された評価を通じて、その問題を研究します。このベンチマークは、厳密な 1 文字出力プロトコルの下で、記号の精度、制約された書式設定、抽出、および短期間の意味論的な意思決定を重視しています。次に、共有パラメータ効率の高い微調整パイプラインにより、NVIDIA L4 クラスの予算内で DoRA/LoRA スタイルのアダプターを備えた 4 ビット NF4 量子化を使用してモデルのサブセットを適応させます。基本評価では、Qwen Coder 3B が厳密精度 75.67% でトップとなり、Qwen2.5 1.5B が 67.10%、Qwen3.5 2B が 64.98%、Granite 3.3 2B が 64.61% で続きます。共有された 108 例のホールドアウト微調整分割では、適応により Qwen Coder 3B が +26.85 ポイント、SmolLM2 1.7B が +25.92 ポイント、Qwen2.5 1.5B が +19.44 ポイント、SmolLM2 360M が +10.18 ポイント、SmolLM2 135M が +5.55 ポイント改善されました。ランキング、トピックレベルの異質性、難易度階層、障害の構成、効率フロンティア、およびトピック条件付き転送にわたって、同じ結論が繰り返します。ベンチマーク構築、クロスモデル評価、および低コストの専門化の規律あるワークフローにより、サブ 3B モデルのサブセットが構造化されたニッチなワークロードのローカル エキスパートとしてすでに実行可能になっています。
原文 (English)
Democratizing AI with Small Language Models: Structured Benchmarking and Parameter-Efficient Fine-Tuning for Local Deployment
AI democratization is not primarily a question of matching frontier-scale generality; it is a question of whether capable models can be selected, audited, and specialized under hardware and governance constraints that ordinary institutions can actually satisfy. This paper studies that problem through a controlled evaluation of nine open-weight language models between 135M and 3B parameters on a 1,085-example, 16-topic multiple-choice benchmark designed for structured local deployment. The benchmark emphasizes symbolic precision, constrained formatting, extraction, and short-horizon semantic decision making under a strict one-letter output protocol. A shared parameter-efficient fine-tuning pipeline then adapts a subset of models using 4-bit NF4 quantization with DoRA/LoRA-style adapters on an NVIDIA L4-class budget. In base evaluation, Qwen Coder 3B leads at 75.67% strict accuracy, followed by Qwen2.5 1.5B at 67.10%, Qwen3.5 2B at 64.98%, and Granite 3.3 2B at 64.61%. On the shared 108-example held-out fine-tuning split, adaptation improves Qwen Coder 3B by +26.85 points, SmolLM2 1.7B by +25.92, Qwen2.5 1.5B by +19.44, SmolLM2 360M by +10.18, and SmolLM2 135M by +5.55. Across ranking, topic-level heterogeneity, difficulty strata, failure composition, efficiency frontiers, and topic-conditioned transfer, the same conclusion recurs: a disciplined workflow of benchmark construction, cross-model evaluation, and low-cost specialization already makes a subset of sub-3B models viable as local experts for structured niche workloads.
マスクされた拡散言語モデルは、エージェントティック RL 用の強力で操作可能なテキストベースの世界モデルです
強化学習 (RL) の最近の成長により、多様で特殊なトレーニング環境の必要性が表面化しています。タスクと報酬の困難が固定された手動でキュレーションされた環境は、モデルのパフォーマンスが向上するにつれて無効なシグナルとなり、長期にわたって報酬がまばらになると、特定のワークフローやツール構造でモードの崩壊を引き起こします。環境状態をシミュレートするワールド モデルは、純粋なロールアウト パフォーマンスと一致しており、オンデマンドで多様性をスケーリングすることが期待できます。ただし、自己回帰 (AR) ワールド モデルは、ツール スキーマ、以前のターン、期待される結果など、グローバルに相互依存する状態アンカーの条件付けを妨げる左から右へのバイアスに悩まされています。私たちは、(i) テキストベースの世界モデリングを、初期状態、タスク コンテキスト、ツール スキーマ、ドメイン ルール、およびステアリング ディレクティブに分解された操作可能な遷移ダイナミクス問題として形式化し、(ii) 9 つのオープンソース環境と 12 のフロンティア モデル ファミリにまたがる 239,403 の接地された状態アクションの軌跡をキュレーションします。 AR LM とマスク拡散言語モデル (MDLM) を比較すると、MDLM は、双方向アンカー認識ノイズ除去を介して、パラメーター サイズの 4 倍を超える LLM よりも優れた一貫性、接地性、および経験的に検証されたロールアウトの多様性を、同等の推論レイテンシーで実現していることがわかります。当社では、確定的な状態チェックを備えたプラグアンドプレイ GRPO トレーニング フレームワークを導入し、3 つの 1.2B ~ 7B エージェント バックボーン (LFM2.5、Qwen3、Mistral) にわたる 3 つの OOD 環境 (ScienceWorld、ALFWorld、AppWorld) でゼロショット トランスファー アブレーションを実行し、環境固有の微調整を行わずにベースラインに対して最大 47% の絶対的なゲインを達成します。さらに、敵対的なシナリオの下での障害モードの動作分析と、現実性、結果の正確性、トレーニングの有用性に関する人間による評価を実施します。私たちは、この方向の研究を促進するために、自分たちの研究をオープンソース化しています。
原文 (English)
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand-curated environments with fixed task and reward difficulties become ineffective signals as model performance improves, and sparse rewards over long horizons induce mode collapse on specific workflows or tool structures. World models that simulate environment states have matched pure rollout performance, making them promising for scaling diversity on-demand. However, autoregressive (AR) world models suffer from a left-to-right bias preventing conditioning on globally interdependent state anchors such as tool schemas, prior turns, and expected outcomes. We (i) formalize text-based world modeling as a steerable transition-dynamics problem decomposed into initial state, task context, tool schemas, domain rules, and steering directives, and (ii) curate 239,403 grounded state-action trajectories spanning nine open-source environments and twelve frontier model families. We compare AR LMs and masked diffusion language models (MDLMs), showing MDLMs, via bidirectional anchor-aware denoising, achieve better coherence, groundedness, and empirically validated rollout diversity than LLMs over 4x their parameter size, at comparable inference latency. We introduce a plug-and-play GRPO training framework with deterministic state checks, and perform zero-shot transfer ablations on three OOD environments (ScienceWorld, ALFWorld, AppWorld) across three 1.2B-7B agent backbones (LFM2.5, Qwen3, Mistral), achieving up to 47% absolute gains over baselines without environment-specific fine-tuning. We further conduct behavioral analysis of failure modes under adversarial scenarios and human evaluation on realism, outcome correctness, and training utility. We open-source our work to encourage research in this direction.
8 つのトークンが必要: 補助ブランチを介した弱から強のオフポリシー RL
検証可能な報酬を伴う強化学習は、大規模な言語モデルの推論を強化するための標準的なアプローチとして登場しており、通常、複数の自己生成ロールアウトを対比することでポリシーを最適化します。ただし、このパラダイムでは、サポートが制限されている重大なボトルネックがあることがわかりました。困難な推論タスクでは、ターゲット モデルのサンプルが意味論的な冗長性を示すことが多く、ポリシーの更新に対して無視できる程度の報酬コントラストを提供する同じ誤った「推論盆地」に収束します。この論文では、弱から強への学習パラダイムを通じてこの制限を克服することを提案します。このパラダイムでは、ポリシーの探索は、弱いが計算効率の高い補助モデルによって情報が提供されます。多くの場合 8 トークンほどの短い補助セグメントを中間ターゲット モデルの軌跡に挿入し、ターゲット モデルがこれらの迂回された状態から推論パスを完了するオフ ポリシー RL 手法である W2SPO を導入します。ポリシーの更新は、最終的な検証可能な報酬に基づいて、これらの短い挿入セグメントに制限されます。経験的に、W2SPO は、数理推論ベンチマークで評価された 4B スケール モデルの中で優れたパフォーマンスを達成し、トレーニング後の評価ベースラインを上回りました。同じサンプリング予算の下でのバニラ GRPO と比較すると、W2SPO は Pass@1 を 62.3% から 64.2% に向上させながら、3.55 倍のトレーニング速度向上を達成します。これらの結果は、弱い補助枝が局所探査支援を拡大することによって、より強力なターゲット推論政策を誘導できることを示唆しています。
原文 (English)
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches
Reinforcement learning with verifiable rewards has emerged as a standard approach for enhancing reasoning in large language models, which typically optimizes the policy by contrasting multiple self generated rollouts. However, we identify a critical support limited bottleneck in this paradigm: on challenging reasoning tasks, the target model's samples often exhibit semantic redundancy, converging into the same erroneous "reasoning basins" that offer negligible reward contrast for policy updates. In this paper, we propose to overcome this limitation through a weak to strong learning paradigm, where a policy's exploration is informed by a weaker but computationally efficient auxiliary model. We introduce W2SPO, an off policy RL method that injects short auxiliary segments often as brief as 8 tokens into intermediate target model trajectories and the target model then completes the reasoning path from these diverted states. Policy updates are restricted to these short inserted segments based on final verifiable rewards. Empirically, W2SPO achieves superior performance among evaluated 4B scale models on mathematical reasoning benchmarks, outperforming evaluated post trained baselines. Compared with vanilla GRPO under the same sampling budget, W2SPO improves Pass@1 from 62.3% to 64.2% while achieving a 3.55 times training speedup. These results suggest that weak auxiliary branches can induce stronger target reasoning policies by expanding local exploration support.
PPO-HSC: 広域ポリシー カバレッジの最適化に基づく探索的強化学習フレームワーク
この論文では、大規模言語モデル (LLM) の微調整におけるモード崩壊の「見えない束縛」に対処するために設計された探索的強化学習フレームワークである PPO-HSC (高次サンプリング カバレッジを備えた近接ポリシー最適化) を紹介します。標準的な検証可能な報酬からの強化学習 (RLVR) は、高報酬の軌道を効果的に強化しますが、多くの場合、モデルが既知のソリューションを過剰に最適化し、好奇心やより幅広いソリューションの多様性を探索する能力が犠牲になります。これを克服するために、PPO-HSC には、「類似性は低いが妥当性が高い」推論パターンの発見を奨励する高次サンプリング カバレッジ (HSC) 報酬が組み込まれています。このフレームワークは、検証された固有のソリューションの動的軌道ライブラリを維持することにより、妥当性制約を通じて構造的合理性を確保しながら、意味論的な新規性に報いる微分可能なシグナルを提供します。数理推論 (GSM8K、SVAMP) およびコード生成タスクに関する経験的評価は、PPO-HSC が、最先端の RL ベースラインの精度と構文の整合性を維持または上回る一方で、ソリューションの多様性と状態空間の適用範囲を大幅に強化することを示しています。
原文 (English)
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization
This paper introduces PPO-HSC (Proximal Policy Optimization with High-order Sampling Coverage), an exploratory reinforcement learning framework designed to address the "Invisible Shackles" of mode collapse in Large Language Model (LLM) fine-tuning. While standard Reinforcement Learning from Verifiable Rewards (RLVR) effectively reinforces high-reward trajectories, it often leads models to over-optimize known solutions, sacrificing curiosity and the ability to explore broader solution manifolds. To overcome this, PPO-HSC incorporates a High-order Sampling Coverage (HSC) reward that incentivizes the discovery of "low-similarity yet high-validity" reasoning patterns. By maintaining a dynamic trajectory library of verified unique solutions, the framework provides a differentiable signal that rewards semantic novelty while ensuring structural rationality through a plausibility constraint. Empirical evaluations on mathematical reasoning (GSM8K, SVAMP) and code generation tasks demonstrate that PPO-HSC significantly enhances solution diversity and state-space coverage while maintaining or surpassing the accuracy and syntax integrity of state-of-the-art RL baselines.
JUMP: 微調整された拡散言語モデルのシングルパス メンバーシップ推論
メンバーシップ推論攻撃 (MIA) は、モデルのトレーニング データに候補サンプルが出現したかどうかをテストします。私たちは、微調整された離散拡散言語モデル (dLLM) の MIA を研究します。メンバーシップとは、ターゲット モデルの微調整セットに含まれることを意味します。自己回帰言語モデルとは異なり、dLLM を使用すると、攻撃者は任意のマスク セットを選択し、すべてのマスクされた位置のトークン分布を並行して取得できます。以前の dLLM 攻撃である SAMA は、ランダムにサンプリングされた多数のマスクにわたる再構築信号を平均化するという自然な損失を模倣する戦略に従いますが、任意次数インターフェイスはランダム化としてのみ使用され、多くのターゲット/参照クエリが必要になります。我々は、dLLM の特徴的な特性の両方を利用するシングルパス スコアリング攻撃である JUMP (Joint Uncertainty-Guided Mask Probing) を提案します。つまり、任意次数デコード可能性を使用して参照信頼性の低い位置を選択し、並列デコード可能性を使用して、モデルごとに 1 つのジョイント マスク クエリを通じて選択されたすべての位置をスコアリングします。 JUMP は、選択された位置を結合してマスクし、クリップされたターゲット/リファレンス再構成ギャップ統計を計算します。 6 つの MIMIR ドメインにわたって微調整された LLaDA-8B-Base 上で、JUMP は SAMA と比較して平均 ROC-AUC を 0.82 から 0.90 に改善し、低 FPR 検出を大幅に向上させますが、ターゲット モデルと参照モデルのそれぞれでセレクター パスとスコアリング パスを 1 回だけ必要とします。
原文 (English)
JUMP: Single-Pass Membership Inference on Fine-Tuned Diffusion Language Models
Membership inference attacks (MIAs) test whether a candidate example appeared in a model's training data. We study MIAs for fine-tuned discrete diffusion language models (dLLMs), where membership means inclusion in the target model's fine-tuning set. Unlike autoregressive language models, dLLMs allow an attacker to choose arbitrary mask sets and obtain token distributions for all masked positions in parallel. The prior dLLM attack, SAMA, follows a natural loss-mimicking strategy by averaging reconstruction signals over many randomly sampled masks, but it uses the any-order interface only as randomization and requires many target/reference queries. We propose JUMP (Joint Uncertainty-Guided Mask Probing), a single-pass scoring attack that exploits both distinctive properties of dLLMs: any-order decodability is used to select low-reference-confidence positions, and parallel decodability is used to score all selected positions through one joint masked query per model. JUMP masks the selected positions jointly and computes a clipped target/reference reconstruction-gap statistic. On fine-tuned LLaDA-8B-Base across six MIMIR domains, JUMP improves mean ROC-AUC from 0.82 to 0.90 over SAMA and substantially improves low-FPR detection, while requiring only one selector pass and one scoring pass through each of the target and reference models.
ColGraphRAG: マルチモーダル GraphRAG の遅延インタラクション証拠取得
グラフに基づいたマルチモーダル質問応答は、テキスト、表、画像を構造化された証拠グラフに整理しますが、エンドツーエンドの精度は、どのマルチモーダル資産が下流の推論に入るのに十分なレベルにランク付けされるかによって決まります。グラフリンクされたイメージの場合、単一ベクトルのバイエンコーダーの類似性により、きめの細かい位置合わせに必要なパッチレベルとトークンレベルの構造が破棄される可能性があります。オフラインのグラフ構築、テキストおよびテーブル側の検索、構造化抽出、下流推論を変更せずに、グラフにリンクされた画像ノード上の視覚的候補ランキング演算子を ColBERT/ColPali 系統の遅延インタラクション MaxSim スタイルのマルチベクトル スコアリングに置き換えることを評価します。 MultimodalQA では、この変更は、グラフにリンクされた画像候補の検索段階の点推定値の向上と下流の QA の向上に関連しており、視覚的な証拠が最も重要な場合には大きな動きがあり、テキスト中心の質問では混合傾向が見られます。私たちはこのパターンを、グラフにリンクされた視覚的証拠を含めるためのメカニズムレベルの証拠として解釈しますが、より広範な検証とより詳細なグラフレベルの診断が今後の重要な作業として残ります。
原文 (English)
ColGraphRAG: Late-Interaction Evidence Retrieval for Multimodal GraphRAG
Graph-grounded multimodal question answering organizes text, tables, and images in a structured evidence graph, yet end-to-end accuracy depends on which multimodal assets are ranked highly enough to enter downstream reasoning; for graph-linked images, single-vector bi-encoder similarity can discard patch- and token-level structure needed for fine-grained alignment. We evaluate replacing the visual candidate-ranking operator over graph-linked image nodes with late-interaction MaxSim-style multi-vector scoring in the ColBERT/ColPali lineage, while keeping offline graph construction, text- and table-side retrieval, structured extraction, and downstream reasoning unchanged. On MultimodalQA, this change is associated with improved retrieval-stage point estimates for graph-linked image candidates and downstream QA gains, with larger movement where visual evidence matters most and mixed trends on text-dominant questions; we interpret the pattern as mechanism-level evidence for graph-linked visual evidence inclusion, while broader validation and finer graph-level diagnostics remain important future work.
Shapley コンテキスト プルーニング: コンテキストの再ランキングとプルーニングのための協力ゲームの視点
コンテキストの再ランク付けと枝刈りは、最新の検索拡張生成 (RAG) システムの効率を向上させるために不可欠になっていますが、解釈可能な統一フレームワークはまだ検討されていません。これまでの研究では、主に語彙検索、クロスエンコーダ アーキテクチャ、モデル蒸留、および低ランク適応 (LoRA) に重点が置かれており、主にヒューリスティック損失関数と経験的帰属に依存していました。この論文では、コンテキストを協力ゲームとしてモデル化することで、重要性の帰属に関する協力ゲーム理論の観点を確立する、コンテキスト再ランキングのための新しいフレームワークである Shapley Context Pruning (SCP) を紹介します。きめの細かい表現と粗い表現の間のトレードオフのバランスをとり、ディープセット アーキテクチャを採用して文レベルで順列不変の値関数を近似し、事前にトレーニングされた言語モデルを文のエンベダーとして利用し、ペアごとのマージン ランキング損失によって最適化します。数学的な厳密さを犠牲にすることなく実用的なスケーラビリティを確保するために、効率的なトレーニングと推論のためにモンテカルロ サンプリングを活用し、Top-K サブセット ランキングを維持するための正式な理論的誤差限界とサンプルの複雑さの保証を提供します。さらに、埋め込み品質と帰属戦略に関する厳密なアブレーション研究と並行して、サポートセンテンスの想起、ニードル・イン・ザ・ヘイスタック(NIAH)評価、ロングコンテキストQA、およびマルチホップ推論に及ぶ包括的な実験を実施します。このモデルは、堅牢なベースラインに対して競争力のある下流 QA パフォーマンスを実現します。
原文 (English)
Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning
Context reranking and pruning have become essential for improving the efficiency of modern Retrieval-Augmented Generation (RAG) systems, yet an interpretable and unified framework remains underexplored. Previous work has primarily emphasized lexical retrieval, cross-encoder architectures, model distillation, and Low-Rank Adaptation (LoRA), mostly relying on heuristic loss functions and empirical attribution. This paper presents Shapley Context Pruning (SCP), a novel framework for context reranking that establishes a cooperative-game-theory perspective for importance attribution by modeling the context as a cooperative game. Balancing the trade-off between fine-grained and coarse-grained representations, we employ a Deep Sets architecture to approximate a permutation-invariant value function at the sentence level, utilizing pre-trained language models as sentence embedders and optimizing via a pairwise margin ranking loss. To ensure practical scalability without sacrificing mathematical rigor, we leverage Monte-Carlo sampling for efficient training and inference, providing formal theoretical error bounds and sample complexity guarantees for preserving Top-K subset rankings. Furthermore, we conduct comprehensive experiments-spanning supporting-sentence recall, Needle-in-the-Haystack (NIAH) evaluations, long-context QA, and multi-hop reasoning-alongside rigorous ablation studies on embedding quality and attribution strategies. The model achieves competitive downstream QA performance against robust baselines.
強化学習政策の検証に関する調査
強化学習 (RL) は、複雑で安全性が重要な領域にますます適用されていますが、ニューラル ネットワーク ベースのポリシーに対する厳格な動作保証の欠如が、依然として導入の大きな障壁となっています。最近のポリシーの表現力と規模の進歩により、この課題はさらに深刻になり、RL ポリシーの検証に関する一連の作業は急速に成長していますが、概念的には断片化しています。この調査は、RL 検証方法に関する統一的な視点を提供します。検証パラダイム (形式的対確率的)、時間的範囲 (段階的対多段階)、強度の保証という 3 つの軸に沿って既存のアプローチ間の関係を明確にする分類法を導入します。分類を超えて、基礎となる理論的基盤を統合し、暗黙の仮定と制限を明示的にし、新たな方向性を特定します。
原文 (English)
A Survey on the Verification of Reinforcement Learning Policies
Reinforcement learning (RL) is increasingly applied in complex, safety-critical domains, yet the lack of rigorous behavioral guarantees for neural network-based policies remains a major barrier to deployment. Recent advances in policy expressiveness and scale have intensified this challenge, leading to a rapidly growing but conceptually fragmented body of work on RL policy verification. This survey provides a unifying perspective on RL verification methods. We introduce a taxonomy that clarifies relationships among existing approaches along three axes: verification paradigm (formal versus probabilistic), temporal scope (step-wise versus multi-step), and guarantees strength. Beyond taxonomy, we unify underlying theoretical foundations, make implicit assumptions and limitations explicit, and identify emerging directions.
LLM エージェントのための正確かつ効率的な長期記憶
永続メモリで強化された LLM エージェントは過去のインタラクションを呼び出すことができますが、既存のシステムには 2 つの制限があります。フラットで非構造化ストレージでは、マルチホップおよび時間的推論に必要なリレーショナル コンテキストが失われ、高価な LLM ベースの分類に依存するため、遅延に敏感な展開には実用的ではありません。新しい情報を保存された知識と照合して検証するメカニズムがなければ、これらのシステムは矛盾を静かに蓄積します。我々は、大幅に正確かつ効率的な LLM エージェント用の構造化された競合認識長期メモリ フレームワークである MOSAIC (Memory-Organized Structured Agent for Information Collection) を紹介します。 MOSAIC は 3 つの重要な機能を導入します。(1) イベント、ペルソナ、および関係全体にわたる関係構造を保持するセマンティック分類を備えたエンティティ型のグラフ ストレージにより、会話履歴に対するマルチホップおよび時間的推論が可能になります。 (2) ハッシュ高速化デュアルパス検索により、LLM ベースの分類が局所性を考慮したハッシュに置き換えられ、無視できる精度の損失でほぼ瞬時の検索が実現されます。 (3) 保存時に新しい情報を既存の隣接グラフに対して相互参照し、矛盾するエントリの更新または削除をトリガーするアクティブな競合検出。 LoCoMo (長時間会話 QA)、HaluMem、および新しい臨床ガイドライン エラー複合テストで評価された MOSAIC は、LoCoMo で 89.35% の精度 (最良のベースラインより +27.21 pp)、最高の HaluMem-Medium 抽出 F1 (86.77%) および HaluMem-Long 抽出 F1 (85.84%)、Medium と Long の両方で最高の QA 正確さを達成しました。 (73.10%、70.75%)、挿入された事実の矛盾の 66% (最良のベースライン (14%) の 4.7 倍) を検出します。一方、ハッシュ アクセラレーションによる検索は、質問あたりの平均検索待ち時間を 0.58 秒に保ちます。
原文 (English)
Accurate and Efficient Long-Term Memory for LLM Agents
LLM agents augmented with persistent memory can recall past interactions, but existing systems suffer from two limitations: flat, unstructured storage loses relational context needed for multi-hop and temporal reasoning, and reliance on expensive LLM-based classification makes them impractical for latency-sensitive deployment. Without mechanisms to validate new information against stored knowledge, these systems silently accumulate contradictions. We present MOSAIC (Memory-Organized Structured Agent for Information Collection), a structured, conflict-aware long-term memory framework for LLM agents that is substantially more accurate and efficient. MOSAIC introduces three key capabilities: (1) entity-typed graph storage with semantic classification preserving relational structure across events, personas, and relationships, enabling multi-hop and temporal reasoning over conversation history; (2) hash-accelerated dual-path retrieval replacing LLM-based classification with locality-sensitive hashing, achieving near-instantaneous lookup with negligible accuracy loss; and (3) active conflict detection at save time that cross-references new information against existing graph neighbors, triggering updates or deletions for contradictory entries. Evaluated on LoCoMo (long-conversation QA), HaluMem, and a novel clinical-guideline error compounding test, MOSAIC achieves 89.35% accuracy on LoCoMo (+27.21 pp over the best baseline), best HaluMem-Medium extraction F1(86.77%) and HaluMem-Long extraction F1 (85.84%), best QA correctness on both Medium and Long (73.10%, 70.75%), and detects 66% of injected factual conflicts-4.7 times higher than the best baseline (14%)-while hash-accelerated retrieval keeps average search latency at 0.58 s per question.
シンボリック拡張はニューラル ファクトチェッカーの正準等価性の盲点を埋める
大規模な言語モデルは、科学文書を要約する際に数字や単位を幻覚させます。これは、科学的主張をひっそりと覆す可能性がある失敗モードです。私たちは、そのようなエラーの検出を型付き検証として再構築します。5 クラスの型付き数量エラー分類法と 1500 項目のベンチマークを導入します。ベンチマークは、PMC および arXiv ソースから書き直され、2 人の独立した LLM アノテーターによって判定付きでラベル付けされます (クリッペンドルフのアルファ = 0.882)。このベンチマークで微調整された ModernBERT エンコーダーは、既製のニューラル ファクト チェッカーをはるかに上回るマクロ F1 = 0.899 に達しますが、4 つのプローブにより構造上の鋭い盲点が明らかになります。物理的に等価な量の正準等価書き換え (例: 95{\deg}C と 368.15 K) では、その精度は 36.5% に低下します。私たちは、シンボリック検証器のモジュールを逆に実行してラベルを保持した拡張トレーニング データを生成するトレーニング時間フレームワークであるシンボリック拡張を提案します。この拡張により、正規等価性の堅牢性が 98.2% に上昇し、同時に分布内精度がわずかに向上しました (マクロ F1: 0.899 から 0.902)。拡張エンコーダは、推論コストなしでクローズドフロンティア LLM と照合し、外部ベンチマーク (SciFact-Open バイナリ マクロ F1: 0.791 ~ 0.828) に転送します。補助エンコーダ入力としてのシンボリック特徴は何も加えず、シンボリック シルバー ラベルは教師ノイズの下でマイナスにスケーリングするという 2 つの否定的な結果が主張を明確にします。これらの結果を総合すると、トレーニング時の拡張が記号コンポーネントと学習コンポーネントの間の適切な統合ポイントであることがわかります。
原文 (English)
Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers
Large language models hallucinate numbers and units when summarizing scientific text, a failure mode that can silently invert a scientific claim. We recast the detection of such errors as typed verification: we introduce a five-class typed-quantity error taxonomy and a 1500-item benchmark, rewritten from PMC and arXiv sources and labeled by two independent LLM annotators with adjudication (Krippendorff's alpha = 0.882). A ModernBERT encoder fine-tuned on this benchmark reaches macro-F1 = 0.899, far above any off-the-shelf neural fact-checker, yet four probes expose a sharp structural blind spot: on canonical-equivalent rewrites of physically equivalent quantities (e.g., 95{\deg}C and 368.15 K) its accuracy collapses to 36.5%. We propose Symbolic Augmentation, a training-time framework that runs the modules of a symbolic verifier in reverse to generate label-preserving augmented training data. The augmentation lifts canonical-equivalence robustness to 98.2% while slightly improving in-distribution accuracy (macro-F1: 0.899 to 0.902); the augmented encoder matches a closed-frontier LLM at no inference cost and transfers to an external benchmark (SciFact-Open binary macro-F1: 0.791 to 0.828). Two negative results sharpen the claim: symbolic features as auxiliary encoder inputs add nothing, and symbolic silver labels scale negatively under teacher noise. Together these results identify training-time augmentation as the right integration point between symbolic and learned components.
SelKV: トークンごとのマージまたはドロップおよびアテンション補償を備えた選択的 KV キャッシュ マージ
大規模言語モデル (LLM) は、コンテキストの長さに応じてメモリ使用量が線形に増加するキーバリュー (KV) キャッシュに依存してテキストを自己回帰的に生成し、大きなボトルネックを引き起こします。最近の圧縮方法では、トークンのマージによってこのコストが軽減されます。ただし、これらのアプローチは多くの場合無差別集約に依存しており、これにより表現が劣化し、注意力の低下、つまり複数の入力をエンコードしているにもかかわらずマージされたトークンが個々のトークンと同じソフトマックス質量を受け取る不一致が生じます。私たちは、これらの制限に対処する、KV キャッシュ圧縮のためのトレーニング不要のデュアルコンポーネント フレームワークを提案します。まず、ソフト コサイン ゲートは、値ベクトルの類似性に基づいてマージの決定を適応的に調整し、異なるトークンを抑制または破棄して意味の忠実性を維持します。 2 番目に、プリフィル アテンション統計から導出されたデコード時のロジット バイアスを適用するアテンション率補償メカニズムを導入し、マージによって引き起こされるソフトマックスの不均衡を修正します。 LongBench (16 個の英国データセット) で評価すると、KV キャッシュの 25% のみを保持しながら、当社のフレームワークは、代表的なワンショット ベースラインに対して強力な圧縮パフォーマンスを実現します。これは、評価されたグループ化クエリ アテンション (GQA) モデルで特に堅牢であり、ほぼ損失のない生成品質を維持します。さらに、このメソッドは複雑な複数ドキュメントの QA タスクでフル キャッシュ ベースラインを上回り、100,000 トークンで 3.3 倍のデコード速度向上を実現します。
原文 (English)
SelKV: Selective KV Cache Merging with Per-Token Merge-or-Drop and Attention Compensation
Large Language Models (LLMs) generate text autoregressively, relying on a key-value (KV) cache whose memory footprint grows linearly with context length, creating a major bottleneck. Recent compression methods mitigate this cost via token merging; however, these approaches often rely on indiscriminate aggregation, which degrades representations and introduces attention sag, a mismatch where merged tokens receive the same softmax mass as individual tokens despite encoding multiple inputs. We propose a training-free, dual-component framework for KV cache compression that addresses these limitations. First, a soft cosine gate adaptively modulates merging decisions based on value-vector similarity, suppressing or discarding dissimilar tokens to preserve semantic fidelity. Second, we introduce an attention-ratio compensation mechanism that applies a decoding-time logit bias derived from prefill attention statistics, correcting the softmax imbalance induced by merging. Evaluated on LongBench (16 English datasets) while retaining only 25% of the KV cache, our framework achieves strong compressed performance against representative one-shot baselines. It is especially robust on the evaluated grouped-query attention (GQA) models, maintaining nearlossless generation quality. Furthermore, the method outperforms the full-cache baseline on complex multi-document QA tasks and delivers a 3.3x decoding speedup at 100k tokens.
RAIL Guard: LLM エージェントの責任ある AI における評価から修復までのギャップを埋める
大規模言語モデル エージェント用の既存のガードレール システムは、安全でないコンテンツをブロックするバイナリ分類子として動作するため、組織は失敗した出力を破棄して最初から再試行する必要があります。 8 つの測定可能な次元にわたって LLM 出力を評価し、評価、書き換え、再評価のループを通じて失敗した出力を繰り返し修復する閉ループ責任 AI パイプラインである RAIL Guard を紹介します。 4 つのフロンティア LLM、4,276 のコンテンツ出力、および 6,400 のエージェント ツール呼び出しシナリオに関する 3 つの実験にわたってパイプラインを評価しました。閉ループ修復は 96.9% の収束率を達成しますが、ブロックと再試行の場合は 49.1% ですが、収束率が最も高い方法では実用性が 22.3% 減少します。フィードバック駆動の自己修復は、重大なユーティリティの損失なしに、修正可能な次元で 86.6% の収束を達成します (p = 0.177)。ツール呼び出し前の評価により、安全でないエージェントの実行が 33% (p = 0.007) 削減され、タスクの完了には影響がありません。私たちは、修復に対応する修正可能な次元と、アルゴリズムではなくアーキテクチャ上の解決策を必要とする構造的次元 (93.0% の透明性、92.8% の説明責任、82.5% の失敗の包括性) との間の重要な違いを特定します。このシステムはオープンソース SDK として利用できます。
原文 (English)
RAIL Guard: Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents
Existing guardrail systems for large language model agents operate as binary classifiers that block unsafe content, leaving organizations to discard failing outputs and retry from scratch. We introduce RAIL Guard, a closed-loop responsible AI pipeline that evaluates LLM outputs across eight measurable dimensions and iteratively remediates failing outputs through an evaluate-rewrite-reevaluate loop. We evaluate the pipeline across three experiments on four frontier LLMs and 4,276 content outputs plus 6,400 agent tool-call scenarios. Closed-loop remediation achieves 96.9% convergence versus 49.1% for block-and-retry, though the highest-convergence method reduces utility by 22.3%; feedback-driven self-repair achieves 86.6% convergence on fixable dimensions with no significant utility loss (p = 0.177). Pre-tool-call evaluation reduces unsafe agent executions by 33% (p = 0.007) with zero impact on task completion. We identify a key distinction between fixable dimensions that respond to remediation and structural dimensions (Transparency at 93.0%, Accountability at 92.8%, and Inclusivity at 82.5% failure) that require architectural rather than algorithmic solutions. The system is available as open-source SDKs.
ジェネラリスト AI 制御: 多目的適応アルゴリズムに向けて
従来のコントローラーは特定のシステム向けに設計されており、異なるシステムの順序やダイナミクスをまたいで転送することはありません。さまざまな次数とダイナミクスのシステムを制御できる学習ベースのコントローラーであるジェネラリスト コントローラーを紹介します。このアプローチでは、マスキングを備えたアテンション メカニズムを使用した新しい動的状態空間表現が導入され、各システムにシステム タグを割り当てることで、アーキテクチャを変更することなく、ワンショットでトレーニングされた単一のニューラル ネットワークが異なる次元のシステムを処理できるようになります。私たちは、安定、不安定、最小位相、および非最小位相ダイナミクスを含む 25 の多様なシステムから 314,630 件のデモンストレーションを生成し、自律型水中および航空宇宙船から機械システムや化学プロセスに至るまで、線形および非線形システムにまたがりました。このモデルは、マルチスケールの時間処理と専門家の混合アーキテクチャを通じて、クロスシステム制御戦略を学習します。シミュレーション結果は、提案されたジェネラリスト コントローラーが、非最小位相や不安定なダイナミクスなどの困難なケースを含む、テストされたすべてのシステムにわたってシステム固有の LQI コントローラーと同等のパフォーマンスを達成する一方、アクチュエーターの飽和、ノイズ、外乱、トレーニング中に遭遇しない参照軌道などの目に見えない動作条件を一般化していることを示しています。この研究は、動的システムの定義されたファミリー内でジェネラリスト制御ポリシーに向けた重要な一歩を表しており、システム固有の調整を行わずに単一の学習済みポリシーを使用して、さまざまな次数とダイナミクスを持つ一連の単入力単一出力 (SISO) システム全体にわたる効果的な制御を実証します。
原文 (English)
Generalist AI Control: Towards Multi-purpose Adaptive Algorithms
Traditional controllers are designed for specific systems and do not transfer across different system orders and dynamics. We present a Generalist Controller, a learning-based controller capable of controlling systems of varying orders and dynamics. The approach introduces a novel dynamic state-space representation using attention mechanisms with masking, enabling a single neural network, trained in one shot, to handle systems with different dimensions without architectural modifications by assigning a system tag to each system. We generated 314,630 demonstrations from 25 diverse systems, including stable, unstable, minimum-phase, and non-minimum-phase dynamics, spanning linear and nonlinear systems from autonomous underwater and aerospace vehicles to mechanical systems and chemical processes. The model learns cross-system control strategies through multi-scale temporal processing and a mixture-of-experts architecture. Simulation results demonstrate that the proposed generalist controller achieves comparable performance to system-specific LQI controllers across all tested systems, including challenging cases such as non-minimum-phase and unstable dynamics, whilst generalising to unseen operating conditions including actuator saturation, noise, disturbance, and reference trajectories not encountered during training. This work represents a significant step towards generalist control policies within a defined family of dynamical systems, demonstrating effective control across a range of single-input single-output (SISO) systems of varying order and dynamics using a single learned policy without system-specific tuning.
LaCache: 拡散大規模言語モデルの正確なキャッシュと高精度適応推論
拡散ベースの大規模言語モデル (DLLM) により、テキスト生成における半自己回帰 (SAR) デコードによる並列生成が可能になります。しかし、現在の方法はオペレータレベルの冗長性が深刻です。プレフィックスとマスクされたサフィックスがブロック内で不変のままであることを無視して、ノイズ除去ステップ中にシーケンス全体を再計算します。私たちは、ロスレス キャッシュと混合精度によってこの冗長性を軽減する、トレーニング不要の高速化フレームワークである LaCache を提案します。具体的には、LaCache は 3 種類の中間結果をキャッシュすることでロスレス状態メモ化 (LSM) を採用しています。(i) 出力を埋め込むための EmbedCache、(ii) トークンごとのプレアテンション状態のための RoPECache、および (iii) FlashAttendant 内のオンライン ソフトマックス統計のための FACache。これらのキャッシュにより、モデルは出力を変更せずに、未変更のトークンに対する冗長な計算をスキップできます。メモリ帯域幅のボトルネックをさらに軽減するために、LaCache には、拡散プロセス全体にわたるステップ依存の活性化分布に合わせて調整された、FFN 層のグループごとの FP8 量子化戦略が組み込まれています。実験では、LaCache 単独で通常の DLLM と比較して約 1.3 倍のエンドツーエンドの高速化を達成できることが実証されています。既存の高速化手法と組み合わせると、LaCache は同等のタスク精度を維持しながら、エンドツーエンドの速度が最大 40.2 倍向上します。
原文 (English)
LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models
Diffusion-based Large Language Models(DLLMs) enable parallel generation via Semi-Autoregressive (SAR) decoding in text generation. However, current methods suffer from severe operator-level redundancy: they recompute the entire sequence during denoising steps, ignoring that the prefix and masked suffix remain invariant within a block. We propose LaCache, a training-free acceleration framework that alleviates this redundancy through lossless caching and mixed precision. Specifically, LaCache employs Lossless State Memoization (LSM) by caching three types of intermediate results: (i) EmbedCache for embedding outputs, (ii) RoPECache for token-wise pre-attention states, and (iii) FACache for the online softmax statistics within FlashAttention. These caches allow the model to skip redundant computation on unchanged tokens without altering the output. To further alleviate memory-bandwidth bottlenecks, LaCache inegrates a per-group FP8 quantization strategy for FFN layers, tailored to step-dependent activation distributions across the diffusion process. Experiments demonstrate that LaCache alone achieves approximately 1.3X end-to-end speedup over vanilla DLLM. When combined with existing acceleration methods, LaCache reaches up to 40.2X end-to-end speedup while maintaining comparable task accuracy.
POMDP としての対話型タスク調整
言語モデルの現在のベンチマークは主に、完全に指定されたタスクの実行を評価します。ただし、実際のユーザーのタスクは多くの場合あいまいです。ユーザーは、不完全な、探索的な、または一貫性のない目標を持って到着するため、アシスタントは、目的のタスクを実行する前に、まず目的のタスクを決定する必要があります。私たちはこの問題をタスク調整、つまり意図したタスクに関してユーザーと調整する能力として研究します。指定されたタスクを不特定のインタラクションに変換するための一般的なフレームワークを導入します。POMDP として形式化され、モデルは部分的かつ進化するユーザーの意図から潜在的なタスクを推測する必要があります。私たちは人間のユーザー研究を使用してユーザー シミュレータを事後検証します。ショッピング、コーディング、専門的な作業環境全体にわたって、タスクが指定されるとモデルはうまく機能することが多い一方で、モデルは依然としてタスクの調整に苦労していることがわかりました。つまり、現在のモデルは時期尚早に動作し、非効率的に対話し、曖昧なリクエストを解決できません。あいまいさがある場合、モデルは平均してユーザーの意図したタスクを 22 ~ 32% の確率でしか回復しません。同じ設定での人体研究では、人間は 48% に達し、評価されたすべてのモデルを上回りました。私たちは、教師あり微調整と強化学習によるポストトレーニングによってタスクの整合性が向上しますが、相互作用による不確実性の解決においてモデルは依然として人間に遅れをとっていることを示しました。総合すると、私たちの結果は、現在のモデルには、信頼できる代理店に必要な重要な対話能力がまだ欠けていることを示唆しています。
原文 (English)
Interactive Task Alignment as a POMDP
Current benchmarks for language models primarily evaluate execution on fully specified tasks. However, real user tasks are often ambiguous. Users arrive with incomplete, exploratory, or even inconsistent goals, requiring the assistant to first determine the intended task before carrying it out. We study this problem as task alignment: the ability to align with a user on their intended task. We introduce a general framework for converting specified tasks into underspecified interactions, formalized as a POMDP in which the model must infer a latent task from partial and evolving user intent. We validate our user simulator post hoc with a human user study. Across shopping, coding, and professional work settings, we find that while models often perform well once the task is specified, models still struggle with task alignment: current models act prematurely, interact ineffectively, and fail to resolve ambiguous requests. Models on average recover the user's intended task only 22-32% of the time under ambiguity. In a human study in the same setting, humans reach 48%, outperforming all evaluated models. We show that post-training with supervised fine-tuning and reinforcement learning improves task alignment, but models still lag behind humans in resolving uncertainty through interaction. Together, our results suggest that current models still lack key interaction abilities required for reliable agency.
いつ計画を立てるか: 事後対応型の制御と熟考型の計画のどちらを選択するかを学ぶ
人間には、迅速で事後的な意思決定と、ゆっくりと熟考した計画を切り替える能力があることが長い間認識されてきました。この論文では、メタ推論として知られるこの能力を人工エージェントでどのように学習するかという問題を研究します。私たちは、状態の観察をアクションに直接マッピングするポリシーとして、事後的な意思決定をモデル化します。このようなポリシーは、強化学習 (RL) または模倣学習を使用してトレーニングできますが、トレーニング分布の外では一般化が不十分な場合があります。あるいは、モデルベースの意思決定時間計画は、より広範な状態セットにわたって適切なアクションを生成する可能性が高くなりますが、追加の計算時間が必要となり、アクションが遅れます。この研究では、反応性ポリシーの不確実性スコアに基づいて条件付けすることによって計算を割り当てるメタ推論ポリシーをトレーニングするための RL 手法を紹介します。このスコアにより、事後対応ポリシーのパフォーマンスが低下する可能性が高い時期と計画が必要な時期を予測できます。我々は動作計画とナビゲーション環境に関する実証研究を実施し、この設計によりメタ推論ポリシーが、事後対応ポリシーが十分なアクションを提供する時期と意思決定時計画が必要な時期を学習できることを示しました。さらに、私たちの設計により、リアクティブなポリシーが改善されるにつれて、メタエージェントが完全なリアクティブな制御に移行できることを示します。
原文 (English)
When to Plan: Learning to Select Between Reactive Control and Deliberative Planning
It has long been recognized that humans have the ability to switch between fast, reactive decision-making and slower, deliberative planning. In this paper, we study the question of how to learn this ability, known as meta-reasoning, in artificial agents. We model reactive decision-making as a policy that directly maps state observations to actions. Such policies can be trained with reinforcement learning (RL) or imitation learning, but may generalize poorly outside of their training distribution. Alternatively, model-based decision-time planning is more likely to produce good actions across a broader set of states but requires additional computation time, which delays acting. In this work, we introduce an RL method for training a meta-reasoning policy that allocates computation by conditioning on a reactive-policy uncertainty score. This score enables it to predict when the reactive policy is likely to perform poorly and when planning is needed. We conduct an empirical study on motion planning and navigation environments, showing that this design enables the meta-reasoning policy to learn when the reactive policy provides a good-enough action versus when decision-time planning is needed. Additionally, we show that our design enables the meta-agent to shift toward fully reactive control as the reactive policy improves.
身体化された機械知能のための尽きることのないアーキテクチャとしてのバークレーとハイザーマン
エドモンド C. バークレーは、記号論理をコンピューティング機械に結び付けることに貢献した作家として一般に記憶されています。その説明は正しいですが、不完全です。バークレーの機械指向の著作とプロジェクトを読むと、中心的な関心はより広範であり、情報を取得し、保持し、変化する条件に適切に応答し、時間の経過とともに動作を組織化する機械にとって、記号ロジックが実用的な設計言語として機能することを示すことです。この論文は、「シンボリック ロジックとインテリジェント マシン」をその取り組みの主要なテキストとして取り上げますが、論理形式、回路、制御、およびインテリジェントな動作をリンクするより大きなバークレー プログラムの代表として扱います。この解釈に基づいて、バークレーは知能を単に抽象的な記号操作として扱っているわけではありません。彼はインテリジェントマシンを操作用語で繰り返し定義し、ロジックをハードウェアの実現に結び付け、入力、出力、メモリ、計算、制御、状態、イベントの調整された相互作用を通じてマシンの動作を説明します。それに応じて、彼のロボットと機械の活動の扱いは、静的な論理形式を超えて、時間的に拡張された環境と連動した動作に移行します。デビッド L. ハイザーマンの機械知能の研究は、このプログラムを記憶、自信、一般化を中心に構築された適応生物アーキテクチャに拡張します。総合すると、バークレーとハイザーマンは、使い尽くされた歴史的エピソードとしてではなく、身体化されたロボット認識に対するまだ十分にテストされていない建築的アプローチへの貢献者として読むことができます。
原文 (English)
Berkeley and Heiserman as an Unexhausted Architecture for Embodied Machine Intelligence
Edmund C. Berkeley is usually remembered as a writer who helped connect symbolic logic to computing machinery. That description is correct, but incomplete. Read across Berkeley's machine-oriented writings and projects, the central concern is broader: to show that symbolic logic can serve as a practical design language for machines that acquire information, retain it, respond appropriately to changing conditions, and organize their behavior over time. This paper takes Symbolic Logic and Intelligent Machines as the principal text in that effort, while treating it as representative of a larger Berkeley program that links logical form, circuitry, control, and intelligent behavior. On this reading, Berkeley does not merely treat intelligence as abstract symbol manipulation. He repeatedly defines intelligent machines in operational terms, ties logic to hardware realization, and describes machine behavior through the coordinated interaction of inputs, outputs, memory, calculation, control, states, and events. His treatment of robots and machine activities accordingly moves beyond static logical form toward temporally extended, environment-coupled behavior. David L. Heiserman's machine-intelligence work extends this program toward adaptive creature architectures built around memory, confidence, and generalization. Taken together, Berkeley and Heiserman can be read not as exhausted historical episodes but as contributors to a still under-tested architectural approach to embodied robotic cognition.
SEER: 精力的な推論を制御するための教師あり学習
制約プログラミングの主な強みの 1 つは、伝播によって検索スペースを削減できることです。ただし、伝播は両刃の剣であり、計算時間の増加と引き換えに、より多くの枝刈り能力が得られます。問題の制約ごとに、最適なプロパゲータは特定のインスタンスに依存し、検索時に変わる可能性があります。文献では、機械学習 (ML) 技術とアクティビティベースのヒューリスティックがそれぞれ、問題のバッチに対するプロパゲーターの (静的) 選択と、伝播強度の (動的) 適応に適用されています。私たちは、ML 経由で取得した oracle 関数を使用して、ターゲット制約に対して複雑なプロパゲータを実行するかどうかを決定することで、これらの取り組みをマージすることを提案します。設計上の選択肢を組み合わせることで、このアプローチは柔軟になり、最先端のソルバーに簡単に組み込むことができます。このペーパーでは、Energetic Reasoning プロパゲータ用のオラクルを構築する実現可能性の調査に焦点を当てます。私たちの実験は、高い予測精度が得られることを示し、分類特徴に関する提案を提供し、そのようなオラクルを構築する際に対処すべき重要な問題を浮き彫りにします。
原文 (English)
SEER: Supervised Learning to Control Energetic Reasoning
One of the main strengths of Constraint Programming is the ability to reduce the search space via propagation. However, propagation is a double-edged sword, with more pruning power coming at the price of larger computation time. For each problem constraint, the best propagator depends on the specific instance and may change at search time. In the literature, Machine Learning (ML) techniques and activity-based heuristics have been applied respectively for choosing (statically) the propagators for a batch of problems and to adapt (dynamically) the propagation strength. We propose to merge those efforts by using an oracle function, obtained via ML, to decide whether to run complex propagators for a target constraint. A combination of design choices makes the approach flexible and easy to embed in state-of-the-art solvers. In this paper, we focus on investigating the feasibility of building an oracle for the Energetic Reasoning propagator. Our experiments show that high prediction accuracy can be obtained, provide suggestions for classification features, and highlight important issues to address when building such an oracle.
人間とAIのコワーキングにおける不均一性原理
マルチステップで一か八かのワークフローを自動化するために生成 AI がますます適用されるようになってきていますが、AI が生成する出力の品質を確保するには人間の判断と関与が依然として不可欠です。実際には、人間の専門家が、多くの場合、中間出力をレビューし、フィードバックを与え、修正し、その後のステップを指揮することによって、AI を定期的に監視することが望ましいですが、そのような監視は人間が許容できる時間とリソースによって制限されます。これにより、人間による監視の必要性と、少ない介入でより多くの出力を提供する AI の効率性との間に緊張が生じます。重要だが十分に検討されていない問題は、人間と AI の共同作業に人間を最適に参加させる方法です。この取り組みはもともと、長い AI ワークフローでは、人間の監視によって不必要な再作業やトークンの消費が削減されながら、ユーザーの満足度が向上することが多いという経験的観察によって動機づけられました。そこから、人間と AI のコワーキングにおいて監視ステージをどこに配置するかという問題を定式化します。次に、合理的な仮定の下で、不均一性の原則を開発します。これは、最適なスケジュールでは、ワークフローに沿ってギャップが減少しない監視ステージを配置するというものです。私たちは、この原則を、文献レビューの作成と Web サイトの構築という 2 つの一般的な AI エージェント ワークフローで経験的に検証します。
原文 (English)
Nonuniformity Principle in Human-AI Coworking
As generative AI is increasingly applied to automate multi-step and high-stake workflows, human judgment and involvement remain essential for ensuring the quality of AI-generated outputs. In practice, while it is desirable for human experts to provide oversight on AI regularly, often by reviewing intermediate outputs, giving feedback, making corrections, and steering subsequent steps, such oversight is constrained by the time and resources that humans can afford. This creates a tension between the need for human oversight and AI's efficiency in delivering more output with less intervention. An important but underexplored question, then, is how to optimally engage humans in human-AI coworking. This work was originally motivated by our empirical observation that in long AI workflows, human oversight often improves user satisfaction while reducing unnecessary rework and token consumption. From there, we formulate the problem of where to place oversight stages in human-AI coworking. Under reasonable assumptions, we then develop the nonuniformity principle, which states that the optimal schedule places oversight stages with non-decreasing gaps along the workflow. We empirically validate this principle in two common AI agent workflows: writing literature reviews and constructing websites.
モダリティから命題へ: マルチモーダル インテリジェンスのための言語中心のフレームワーク
私たちは、画像、ビデオ、テキストのいずれの観察も、シーン内のエンティティ、アクション、および関係についての単純なステートメントである原子命題のバッグとして表現される、マルチモーダル データの言語表現を提案します。グローバル セマンティック コードブックは、これらを標準的な原子命題の共有語彙に統合し、あらゆるモダリティと観察を、粒度の細かい事実から高レベルの概念にまで及ぶ 1 つの解釈可能な空間に配置し、より豊かな概念を構成します。これにより、推論による解釈可能性、クロスモーダルな理解と検索、および複雑なマルチモーダルな理解、豊富なデータのキュレーション、複雑な構造化された検索を可能にする構成性がもたらされます。自動運転とオープンワールドデータに関するフレームワークをデモンストレーションします。
原文 (English)
From Modalities to Propositions: A Language-Centric Framework for Multimodal Intelligence
We propose a language representation for multimodal data in which any observation, whether image, video, or text, is expressed as a bag of atomic propositions, simple statements about the entities, actions, and relations in a scene. A global semantic codebook unifies these into a shared vocabulary of canonical atomic propositions, placing every modality and observation into one interpretable space that spans fine grained facts to high level concepts and composes into richer ones. This brings interpretability with reasoning, cross-modal understanding and retrieval, and compositionality that enables complex multimodal understanding, rich data curation and complex structured retrieval. We demonstrate the framework on autonomous driving and open-world data.
正確なネットワーク手術: リアクティブな計算グラフにおける関数の不変性と勾配可塑性
Net2Net やプログレッシブ スタッキングなどの機能を保持したネットワーク成長手法は、学習した関数を破壊することなくモデルの容量を拡張しますが、既存の定式化では数値の摂動を許容するか、トレーニング プログラムの完全な再構築が必要になります。我々は、Exact Network Surgery を形式化します。これは、(i) 明示的な浮動小数点仮説の下でビット正確にネットワーク関数が保持され、(ii) 挿入されたパラメータが挿入直後にトレーニング可能なままであるように、ライブ計算グラフに残差ブロックをインプレース挿入することです。ゲート残差ブロックの恒等射定理、リアクティブ無効化エンジンが他のすべてのノードの値とオプティマイザーの状態をそのままにして、挿入ポイントの下流の円錐を正確に再計算することを示す構造局所性定理、およびランダムに初期化された分岐上でゼロに初期化された勾配シャドウイング ゲート アルファが挿入時に一般的に非ゼロの勾配を受け取ることを示す初期化からの脱出命題を証明します。縮退構成 (ゼロ初期化された出力投影とゼロ ゲートの組み合わせ) を特定します。これは、正確な鞍点勾配降下法から逃れることができません。すべての主張は、Julia のリアクティブ グラフ エンジンである NeuroDSL のリファレンス実装で検証されます。グラフトはテストされたすべてのロジットでビット正確です (1600 件中 0 件の不一致)。ゲートは最初のオプティマイザー ステップでゼロをエスケープし、2 番目のステップで分岐勾配のロックを解除します。これは予測どおりです。縮退構成は、600 ステップの実行全体にわたってまったく同じゼロの勾配を示します。手術コストは r = 0.9992 で下流の錐体サイズを追跡しますが、グラフトプラス無効化のブックキーピングは挿入深度全体にわたって一定 (約 0.75 ミリ秒) です。そしてトレーニングは、実際のプロセスの再起動後もほぼ同じように再開されます。フラグ付きの予備付録では、挿入後のゲート ダイナミクスに関する最初の単一シードの観察結果が報告されています。
原文 (English)
Exact Network Surgery: Functional Invariance and Gradient Plasticity in Reactive Computational Graphs
Function-preserving network growth techniques such as Net2Net and progressive stacking expand a model's capacity without destroying its learned function, but existing formulations either tolerate numerical perturbations or require a full rebuild of the training program. We formalize Exact Network Surgery: the in-place insertion of a residual block into a live computational graph such that (i) the network function is preserved -- bit-exactly under explicit floating-point hypotheses -- and (ii) inserted parameters remain trainable immediately after insertion. We prove an identity-morphism theorem for gated residual blocks, a structural-locality theorem showing that a reactive invalidation engine recomputes exactly the downstream cone of the insertion point, leaving every other node's value and optimizer state untouched, and an escape-from-initialization proposition showing that the Gradient Shadowing gate alpha, initialized at zero over a randomly initialized branch, receives a generically non-zero gradient at insertion time. We identify a degenerate configuration -- zero-initialized output projections combined with a zero gate -- that is an exact saddle point gradient descent cannot escape. Every claim is validated on the reference implementation in NeuroDSL, a reactive graph engine in Julia: grafting is bit-exact on every logit tested (0 mismatches out of 1600); the gate escapes zero at the first optimizer step and unlocks branch gradients at the second, exactly as predicted; the degenerate configuration exhibits gradients identically zero for the entire 600-step run; surgery cost tracks downstream cone size with r = 0.9992 while graft-plus-invalidation bookkeeping is constant (about 0.75 ms) across insertion depths; and training resumes bit-identically across a real process restart. A flagged preliminary appendix reports first single-seed observations on post-insertion gate dynamics.
FST.ai 2.5: オリンピックおよびパラテコンドーの意思決定支援、アスリートのデジタルツイン、連盟規模の分析のための説明可能かつ不確実性を認識した AI
エリートスポーツの急速なデジタル化により、人工知能 (AI)、パフォーマンス分析、意思決定支援システムをアスリートの育成と競技管理に統合する新たな機会が生まれました。しかし、既存のソリューションは断片的なままであり、通常はパフォーマンス分析、アスリートのモニタリング、審判のサポートなどの個別のタスクに対応しています。この論文では、説明可能で不確実性を認識し、オリンピックおよびパラテコンドー向けの安全な AI フレームワークである \textbf{FST$\cdot$ai~2.5} を紹介します。 \textbf{FST$\cdot$ai~2.5} は、アスリート インテリジェンス、競技分析、連盟規模のデータ管理、AI 支援による意思決定サポート、アスリートとイベントのデジタル ツイン、説明可能なパフォーマンス指標、および適応トレーニングの推奨事項を統合する統合デジタル エコシステムを導入します。このフレームワークは、透明性が高く、安全で連盟を意識したガバナンスを通じて、世界テコンドー (WT)、加盟各国協会 (MNA)、コーチ、審判、アナリスト、選手をサポートします。 \textbf{FST$\cdot$ai~2.5} は、マルチソースの競技データ、アスリートのパフォーマンス情報、状況に応じた証拠を組み合わせることで、説明可能で不確実性を認識した AI を使用して、戦術診断、長期的なアスリートのモニタリング、パフォーマンス予測、個別の能力開発計画、連盟全体のベンチマークを提供します。プロトタイプの展開により、提案されたフレームワークの実現可能性が実証されます。この方法論はオリンピックとパラテコンドー向けに開発されましたが、格闘技やその他の高性能スポーツ環境における説明可能な AI、デジタル ツイン、信頼できる意思決定サポートに広く適用できます。
原文 (English)
FST.ai 2.5: Explainable and Uncertainty-Aware AI for Olympic and Para-Taekwondo Decision Support, Athlete Digital Twins, and Federation-Scale Analytics
The rapid digitalisation of elite sport has created new opportunities for integrating artificial intelligence (AI), performance analytics, and decision-support systems into athlete development and competition management. However, existing solutions remain fragmented, typically addressing isolated tasks such as performance analysis, athlete monitoring, or referee support. This paper presents \textbf{FST$\cdot$ai~2.5}, an explainable, uncertainty-aware, and secure AI framework for Olympic and Para-Taekwondo. \textbf{FST$\cdot$ai~2.5} introduces a unified digital ecosystem integrating athlete intelligence, competition analytics, federation-scale data management, AI-assisted decision support, athlete and event digital twins, explainable performance indicators, and adaptive training recommendations. The framework supports World Taekwondo (WT), Member National Associations (MNAs), coaches, referees, analysts, and athletes through transparent, secure, and federation-aware governance. By combining multi-source competition data, athlete-performance information, and contextual evidence, \textbf{FST$\cdot$ai~2.5} provides tactical diagnostics, longitudinal athlete monitoring, performance forecasting, personalised development planning, and federation-wide benchmarking using explainable and uncertainty-aware AI. Prototype deployments demonstrate the feasibility of the proposed framework. Although developed for Olympic and Para-Taekwondo, the methodology is broadly applicable to explainable AI, digital twins, and trustworthy decision support in combat sports and other high-performance sporting environments.
むしろ非常に知的な話し方をするエージェント
長期的な AI エージェントはますます有能になってきていますが、ユーザーとの対話は驚くほど希薄なままです。ほとんどのワークフローでは、ユーザーは最初の指示を与え、選択的なテキスト更新のみを受け取り、エージェントが何をしているのか、いつ介入すべきなのかを明確に認識できなくなります。これにより、現在のエージェント エコシステムには欠けている部分が残ります。それは、エージェントがユーザーに継続的にアクセスできるようにする常時接続の Jarvis スタイルのメディエーターです。このようなメディエーターは、ユーザーとのリアルタイムの音声対話をサポートし、作業者の邪魔をせずに質問に答え、進捗状況や混乱を積極的に報告し、必要に応じてエージェントの実行にユーザー ガイダンスを挿入する必要があります。この作業では、長期的なエージェント ワークフローにおけるメディエーションの 2 つの価値を測定するためのベンチマークである JarvisBench を紹介します。 JarvisBench には 2 つの補完的なトラックが含まれています。1 つはメディエーションによって下流のタスクの完了が改善されるかどうかを測定するエージェントとのコラボレーション トラック、もう 1 つはメディエーションによって進行中の実行がユーザーにとってより理解しやすく、応答性が高く、アクセスしやすいものになるかどうかを測定するユーザー インタラクション トラックです。モジュール式の参照 Jarvis プロトタイプを使用してベンチマークをインスタンス化し、OpenClaw で実行される 34 個のテキストのみの WildClaw タスクで評価します。 GPT-5.5、Claude Opus 4.7、Gemini ベース、および GPT ベースのワーカー エージェントを使用した予備的な結果は、Jarvis スタイルの調停がユーザーの質問に対してトレースに基づいた応答を提供し、適切なタイミングでまばらなユーザー ガイダンスが挿入された場合にタスクのパフォーマンスを向上させることができることを示唆しています。この結果は、有効性がメディエーターの LLM 脳に大きく依存していることも示しており、この欠落している中間層の可能性と、より広範なコミュニティの取り組みの必要性の両方を強調しています。デモページ https://cchen1436.github.io/jarvis
原文 (English)
Just A Rather Very Intelligent Spoken Agent
Long-horizon AI agents are becoming increasingly capable, yet their interaction with users remains surprisingly thin. In most workflows, users give an initial instruction, receive only selective textual updates, and lose a clear sense of what the agent is doing or when to step in. This leaves a missing part in the current agent ecosystem: an always-on Jarvis-style mediator that keeps the agent continuously reachable to the user. Such a mediator should support real-time spoken interaction with the user, answer questions without interrupting the worker, proactively report progress or confusion, and inject user guidance back into the agent's execution when useful. In this work, we introduce JarvisBench, a benchmark for measuring the dual value of mediation in long-horizon agent workflows. JarvisBench contains two complementary tracks: an agent-collaboration track that measures whether mediation improves downstream task completion, and a user-interaction track that measures whether mediation makes ongoing execution more understandable, responsive, and accessible to users. We instantiate the benchmark with a modular reference Jarvis prototype and evaluate it on 34 text-only WildClaw tasks executed in OpenClaw. Preliminary results with GPT-5.5, Claude Opus 4.7, Gemini-based, and GPT-based worker agents suggest that Jarvis-style mediation can provide trace-grounded responses to user questions and improve task performance when sparse user guidance is injected at appropriate moments. The results also show that effectiveness depends strongly on the mediator's LLM brain, highlighting both the promise of this missing middle layer and the need for broader community effort. Demo page https://cchen1436.github.io/jarvis
セマンティック物理学アライメントによるカスタマイズされた足装具の閉ループ生成設計のための研究プロトタイプ
構造化されていない臨床処方を患者固有の足装具 (FO) に変換することは、意味論的物理的不整合によって妨げられています。つまり、高レベルの臨床意図が装具の 3D 幾何学的パラメータに決定論的にマッピングされておらず、既存の設計ワークフローは、生体力学的な検証が即時に行われず、手動の専門知識に依存したままになっています。我々は、臨床的に検証された治療装置ではなく、カスタマイズされたFOの計算設計自動化のための閉ループフィードバックを備えたモジュール式パイプラインである研究プロトタイプであるTANS-FOを紹介します。 Text-Aligned Neural Surrogate (TANS) は、クロスアテンションを使用して臨床テキストの埋め込みを連続格子密度フィールドに投影します。一方、Graph Neural Network (GNN) サロゲートは、有限要素解析 (FEA) の代わりに足底応力をリアルタイムで予測します。このフレームワークは、オープンアクセスの PicoFoot-5K 人体計測データベース (5,230 人の被験者、30 以上の解剖学的パラメーター) に基づいています。標準化された準静的荷重の下では、GNN サロゲートは Abaqus 参照ソルバー (R^2 = 0.94) と一致し、完全なパイプラインにより製造可能な格子インソールが数分以内に合成されます。男性 18 ~ 40 歳のコホートでは、提案されたシステムは、フィッティング誤差 0.42 mm で、パラメトリック CAD よりも 34.7% の代理予測ピーク圧力低減を達成しました。これとは別に、VAS 疼痛レポートを使用した探索的な実現可能性観察 (n = 12; 2 週間の追跡調査; 対照群なし) では、短期的な快適性の改善 (VAS 6.4 -> 2.1) が示されていますが、このデータは予備的な観察証拠としてのみ分類されており、臨床有効性の証拠としては明示的に分類されていません。
原文 (English)
A Research Prototype for Closed-Loop Generative Design of Customized Foot Orthoses via Semantic-Physics Alignment
Translating unstructured clinical prescriptions into patient-specific foot orthoses (FOs) is hindered by a semantic-physical misalignment: high-level clinical intent is not mapped deterministically onto the 3D geometric parameters of the orthosis, and existing design workflows remain dependent on manual expertise with no instantaneous biomechanical validation. We present TANS-FO, a research prototype-a modular pipeline with closed-loop feedback for computational design automation of customized FOs, not a clinically validated therapeutic device. A Text-Aligned Neural Surrogate (TANS) uses cross-attention to project clinical-text embeddings onto a continuous lattice-density field, while a Graph Neural Network (GNN) surrogate predicts plantar stress in real time as a substitute for Finite Element Analysis (FEA). The framework is anchored on the open-access PicoFoot-5K anthropometric database (5,230 subjects; 30+ anatomical parameters). Under standardized quasi-static loading, the GNN surrogate agrees with an Abaqus reference solver (R^2 = 0.94), and the full pipeline synthesizes manufacturing-ready lattice insoles within minutes. On the Male 18-40 cohort, the proposed system attains a surrogate-predicted peak-pressure reduction of 34.7% over parametric CAD, with a fit error of 0.42 mm. Separately, an exploratory feasibility observation (n = 12; 2-week follow-up; no control group) using VAS pain reporting indicates short-term comfort improvement (VAS 6.4 -> 2.1), but this data is explicitly classified as preliminary observational evidence only-not evidence of clinical efficacy.
TopoTuner: 大規模言語モデルのトポロジカル微調整
完全な微調整は、事前トレーニングされた LLM を適応させるための強力な方法であることに変わりはありませんが、すべての重みが更新されるため、コストがかかる可能性があります。 LoRA はトレーニング可能なパラメータの数を減らしますが、どの事前トレーニング済みコンポーネントをトレーニングすべきか、適応中にどれを凍結できるかについては直接答えません。アテンション投影行列を選択的にフリーズするための、トポロジーに基づいた微調整フレームワークである TopoTuner を紹介します。 \method は、各射影行列を行クラウドとして扱い、永続化ダイアグラム間の Wasserstein 距離を使用して、微調整中にトポロジーがどのように変化するかを測定します。 TopoTuner は、ソース データセットから再利用可能なフリーズ プロファイルを学習し、それをドメイン外データセットの効率的に微調整されたモデルに転送し、タスク固有のトポロジ ドリフトが質問応答タスクや感情分析タスク全体で一般化するかどうかを評価します。 LLaMA-3.1-8B、Mistral-7B-v0.3、Qwen3-8B-Base 全体で、TopoTuner はモデル パラメーターの 1 ~ 2\% のみをトレーニングしながら完全な微調整で競争力があり、9 つのモデル データセット設定のうち 7 つで LoRA を上回り、投影パラメーターの最大 39.57\% を変更できます。 TopoTuner は更新を最小限に抑え、トレーニング時間を平均で完全な微調整と比較して 20.4\%、LoRA と比較して 5.5\% 削減します。 TopoTuner は、1 つのデータセットで学習した微調整動作を複数のタスク間で共有できる、再利用可能な凍結プロファイルの新しい方向性を開きます。
原文 (English)
TopoTuner: Topological Finetuning of Large Language Models
Full fine-tuning remains a strong way to adapt pretrained LLMs, but it updates all weights and can be expensive. LoRA reduces the number of trainable parameters, but it does not directly answer which pretrained components should be trained and which can be frozen during adaptation. We introduce TopoTuner, a topology-guided fine-tuning framework for selective freezing of attention projection matrices. \method treats each projection matrix as a row cloud and uses Wasserstein distances between persistence diagrams to measure how its topology changes during fine-tuning. TopoTuner learns a reusable freezing profile from a source dataset and transfers it to efficiently fine-tune models on out-of-domain datasets, evaluating whether task-specific topological drift generalizes across question answering and sentiment analysis tasks. Across LLaMA-3.1-8B, Mistral-7B-v0.3, and Qwen3-8B-Base, TopoTuner is competitive with full fine-tuning while training only 1-2\% of the model parameters, and outperforms LoRA in 7 out of 9 model-dataset settings, which can change up to 39.57\% of the projection parameters. Along with minimized updates, TopoTuner reduces training time by 20.4\% relative to full fine-tuning and 5.5\% relative to LoRA on average. TopoTuner opens a new direction for reusable freezing profiles, where fine-tuning behavior learned on one dataset can be shared across multiple tasks.
不確実性に基づく幻覚検出のための多様性指向の微調整
既存の幻覚検出方法は通常、モデル自体に変更を加えることなく、推論段階で実行されます。この論文では、結果として得られるモデルにおける幻覚の検出可能性を高める微調整戦略を探ることに興味があります。セマンティック エントロピー ベースの検出に焦点を当てると、モデルが複数の実行にわたってほぼ同一の不正解を生成するため、多くの誤った出力が検出されないままであることがわかります。これに対処するために、私たちはより多様な世代を奨励するための多様性指向の微調整を提案します。 2 つの具体的な戦略を紹介します。1 つは教師あり微調整 (SFT) に基づくもの、もう 1 つは直接優先最適化 (DPO) に基づくものです。私たちのアプローチを評価し、微調整前後のモデルの動作を分析するために、広範な実験が行われます。私たちの微調整手法を採用した後、モデルは幻覚応答に対して低いセマンティック エントロピー応答を生成する可能性が低くなり、それによって幻覚検出の有効性が向上し、最終的には最先端の手法よりも優れた、または同等の結果が得られることがわかりました。コードは公開されます。
原文 (English)
Diversity-Oriented Fine-Tuning for Uncertainty-Based Hallucination Detection
Existing hallucination detection methods are typically conducted at the inference stage, without making any modifications to the model itself. In this paper, we are interested in exploring fine-tuning strategies that enhance the detectability of hallucinations in the resulting model. Focusing on semantic-entropy-based detection, we observe that many erroneous outputs remain undetected because the model produces nearly identical incorrect answers across multiple runs. To address this, we propose diversity-oriented fine-tuning to encourage more varied generations. We introduce two specific strategies: one based on Supervised Fine-Tuning (SFT) and the other on Direct Preference Optimization (DPO). Extensive experiments are conducted to evaluate our approach and analyze the behavior of the models before and after fine-tuning. We find that after adopting our fine-tuning methods, the models become less likely to produce low semantic entropy responses for hallucinated answers, thereby improving the effectiveness of hallucination detection, eventually yielding results better than or comparable with state of the art methods. The code will be publicly released.
eRisk 2026 での DS@GT ARC: 会話型うつ病スクリーニングのための構造化アルゴリズム ガイダンスを備えたハイブリッド マルチエージェント LLM システム
会話型うつ病スクリーニングに関する eRisk 2026 タスク 1 チャレンジへの DS@GT の提出について説明します。このシステムでは、さまざまなうつ病プロファイルを持つ個人をシミュレートする LLM ペルソナにインタビューし、デリケートなメンタルヘルスに関する質問を直接行うことなく、ベックうつ病インベントリ II (BDI-II) スコアとペルソナごとの 4 つの主要な症状を生成します。私たちのパイプラインは 3 つの段階を経て進化しました。まずモノリシックな単一モデルのプロトタイプ、調整オーケストレーション層の下で会話型インタビューを BDI-II スコアリングから分離するベースラインのマルチエージェント アーキテクチャ、そして有料の GPT-5-nano インタビュアーをオープンソースの Gemma 27B に置き換える最終的なハイブリッド構成です。モデルの弱い推論と指示に従いを補うために、ハイブリッドでは 3 つのアルゴリズム コンポーネントが追加されています。それは、インタビューの開始者とフォローアップを標準化する事前計算された対話ツリー、Weaver フレームワークからインスピレーションを得た信頼性を重み付けしたコンセンサス集計、および調査されていない症状に対するクラスターベースの代入ステップです。 20 人のペルソナすべてに対して 3 つの完全自動実行を送信しました。実行 1 は有料ベースラインから、実行 2 と 3 はハイブリッドから行いました。ハイブリッド ラン 3 は、ADODL 0.9063 を達成し、すべての完全サブミッション実行の中で 3 位にランクされ、DS@GT は全体の 21 チーム中 2 位になりました。また、ペルソナあたりの API コストの約 4 分の 1 で、有料のベースライン ラン 1 (0.8841) を上回りました。これらの結果は、アルゴリズムによる十分な監視があれば、会話型面接官の役割において、弱いオープンソース モデルでも強力な独自モデルと競合できるという中心的な仮説を裏付けています。ソース コードは https://github.com/dsgt-arc/erisk-task1-2026 で入手できます。
原文 (English)
DS@GT ARC at eRisk 2026: Hybrid Multi-Agent LLM System with Structured Algorithmic Guidance for Conversational Depression Screening
We describe DS@GT's submission to the eRisk 2026 Task 1 challenge on conversational depression screening, in which systems interview LLM personas that simulate individuals with varying depression profiles and produce a Beck Depression Inventory II (BDI-II) score plus four key symptoms per persona, without directly asking sensitive mental health questions. Our pipeline evolved through three stages: a monolithic single-model prototype to start off, a baseline multi-agent architecture that separates conversational interviewing from BDI-II scoring under a coordinating orchestration layer, and a final hybrid configuration that replaces the paid GPT-5-nano interviewer with the open-source Gemma 27B. To offset the model's weaker reasoning and instruction-following, the hybrid adds three algorithmic components: a precomputed dialogue tree that standardizes interview openers and follow-ups, a reliability-weighted consensus aggregation inspired by the Weaver framework, and a cluster-based imputation step for unprobed symptoms. We submitted three fully automated runs across all 20 personas, with Run 1 from the paid baseline and Runs 2 and 3 from the hybrid. Hybrid Run 3 achieved an ADODL of 0.9063, ranking 3rd among all complete-submission runs and placing DS@GT 2nd among the 21 teams overall, while outperforming our paid baseline Run 1 (0.8841) at roughly one-quarter of the per-persona API cost. These results support our central hypothesis that with sufficient algorithmic supervision, a weaker open-source model can compete with a stronger proprietary model in the conversational interviewer role. Our source code is available at https://github.com/dsgt-arc/erisk-task1-2026.
DL オントロジーの認識的機密性ポリシーに基づく扱いやすいクエリ応答 (拡張バージョン)
私たちは、記述ロジック (DL) オントロジーのコンテキストで、また認識依存関係 (ED) を通じて表現される機密性ポリシーについて、機密性を保持するデータ アクセスへの宣言的アプローチである制御クエリ評価 (CQE) を研究します。まず、CQE の既知のセマンティクス (GA および IGA 含意) の下でクエリ (具体的には論理積クエリのブール和集合) に応答する問題に取り組みます。私たちの結果は、TBox が $\text{DL-Lite}_{\mathcal{R}}$ で表現される場合、CQE は一般に計算的に扱いにくいことを示しています。さらに、ED が存在する場合、IGA セマンティクスは、識別不可能性として知られる重要な機密保持特性を満たさないことが最近証明されました。計算が容易で機密性が保たれる CQE 形式を定義することを目的として、最小限のポリシー違反 (MPV) の概念に基づいた CQE の新しいセマンティクスを導入します。新しいセマンティクスが以前のセマンティクスの健全な近似を提供しながら、区別不可能性の特性を満たしていることを示します。また、$\text{DL-Lite}_{\mathcal{R}}$ オントロジーの場合、MPV セマンティクスに基づくクエリ含意がデータ複雑さの多項式時間で決定できることも証明します。最後に、OWL 2 QLの既存のベンチマークを使用して、この新しいアプローチの実現可能性を評価するために使用したフレームワークのソフトウェア実装を紹介します。
原文 (English)
Tractable Query Answering under Epistemic Confidentiality Policies in DL Ontologies (extended version)
We study Controlled Query Evaluation (CQE), a declarative approach to confidentiality-preserving data access, in the context of Description Logic (DL) ontologies, and for confidentiality policies expressed through Epistemic Dependencies (EDs). We first address the problem of answering queries (specifically, Boolean unions of conjunctive queries) under known semantics for CQE (GA- and IGA-entailment). Our results show that if the TBox is expressed in $\text{DL-Lite}_{\mathcal{R}}$, CQE is computationally intractable in general. Moreover, in the presence of EDs, the IGA semantics has recently been proven not to satisfy an important confidentiality preservation property known as indistinguishability. With the goal of defining computationally easier and confidentiality-preserving forms of CQE, we introduce a new semantics for CQE, based on the notion of minimal policy violation (MPV). We show that the new semantics provides a sound approximation of the previous ones, while satisfying the indistinguishability property. We also prove that, in the case of $\text{DL-Lite}_{\mathcal{R}}$ ontologies, query entailment under the MPV semantics can be decided in polynomial time in data complexity. Finally, we present a software implementation of our framework that we used to evaluate the feasibility of this new approach using an existing benchmark for OWL 2 QL.
RECON: 長いコンテキストでの構成推論のためのエージェント メモリのベンチマーク
大規模な言語モデルと LLM ベースのエージェントは、パーソナル チャット アシスタント、エンタープライズ副操縦士、自律型ワークフロー エージェントとして広く使用されています。これらすべてのアプリケーションにおいて、記憶 (長いコンテキストや複数の対話を通じて蓄積された情報を保持し、アクセスし、推論する能力) は、エージェントの信頼性を決定する上で重要な役割を果たします。長いコンテキストにわたる構成推論を評価するためのベンチマークである RECON (Reasoning over Extended Contexts with Obfuscated Narratives) を紹介します。 RECON は、3 つのドメイン (刑事、医療、金融) にわたる 24 の事件ファイルにまたがり、それぞれのトークンの範囲は 50,000 から 100,000 であり、マルチホップ証拠チェーンの再構築、カスケード無効化の伝播、情報源の競合の解決、反事実推論、時間的制約の充足、および時間的事実の検索という 6 つのメモリ集約型タスクでエージェントをテストします。最近のメモリベンチマークは、エージェントが散在した事実を検索できるか、または事実が変更されたかどうかを検出できるかどうかを評価するのに対し、RECON は、変更後に何が起こるか、エージェントが影響を受ける下流の結論、独立したサポートを通じて生き残る結論、および代替タイムラインがどのように展開するかを追跡できるかどうかを評価します。私たちの評価では、現在のアーキテクチャ全体に大きな限界があることが明らかになりました。Oracle 以外の最も強力なシステムでも精度は 22.4% にとどまり、検索と推論がそれぞれ課題として表面化しています。
原文 (English)
RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts
Large language models and LLM-based agents are widely used as personal chat assistants, enterprise copilots, and autonomous workflow agents. In all these applications, memory (the ability to retain, access, and reason over information accumulated over long contexts and multiple interactions) plays a crucial role in determining the reliability of any agent. We introduce RECON (Reasoning over Extended Contexts with Obfuscated Narratives), a benchmark for evaluating compositional reasoning over long contexts. RECON spans 24 case files across three domains (criminal, medical, and financial), each ranging from 50k to 100k tokens, and tests agents on six memory intensive tasks: reconstructing multi-hop evidence chains, propagating cascading invalidations, resolving source conflicts, counterfactual reasoning, satisfying temporal constraints, and temporal fact retrieval. Recent memory benchmarks evaluate whether agents can retrieve scattered facts or detect if a fact has changed whereas RECON evaluates what happens after the change, whether agents can trace which downstream conclusions are affected, which survive through independent support, and how alternative timelines would have unfolded. Our evaluation reveals substantial limitations across current architectures: even the strongest non-Oracle system reaches only 22.4% Accuracy, with retrieval and reasoning each surfacing as challenges.
制約にアンカーされた推論トレース
自己回帰マルチモーダル大規模言語モデル (MLLM) は、雪だるま式にエラーが発生します。つまり、思考連鎖 (CoT) トレースの初期段階で 1 つの間違った推論が発生すると、下流の推論がすべて破損します。最先端のオープンソース MLLM では、最初のエラーが発生すると、そのようなケースの 65% で、推論がカスケード的に残りのすべてのステップで失敗することがわかりました (この指標を雪だるま式比率と呼んでいます)。既存の緩和策 (複数のチェーンのサンプリング、事後的な自己検証、または完全なプログラム合成) は、記号の基礎が欠如しているか、エラーの検出が遅すぎるか、自然言語推論の柔軟性を犠牲にしています。我々は、自然言語推論ステップを記号的制約アサーション(視覚コンテンツに関する軽量で機械チェック可能なステートメント(例: count(red_objects) = 3))とインターリーブするようにMLLMを訓練する神経記号フレームワークである制約アンカー推論トレース(CART)を提案します。学習されたニューラル グラウンディング ヘッドとブール制約伝播を組み合わせた 2 つの制約伝播モジュールは、抽出された視覚的特徴に対してこれらのアンカーを継続的に検証し、相互の論理的一貫性をチェックします。矛盾が検出されると、バックトラック コントローラーが生成を停止し、一貫性のある最後のチェックポイントに戻り、エラーの伝播を防ぎます。可変周波数放出メカニズムにより、モデルはアンカー密度を適応的に制御し、トレースの肥大化を回避できます。シーン グラフから派生したグラウンド トゥルース制約アノテーションを使用して GQA、CLEVR-CoGenT、VCR を強化することで 218,000 のトレーニング インスタンスを構築し、LoRA を介してオープンソース MLLM (LLaVA-NeXT、Qwen2-VL) を微調整します。 5 つのベンチマークで、CART は雪だるま式レートを 0.65 から 0.14 に減少させ、GQA 精度をトレーニングのみのベースラインと比較して +4.6 パーセント ポイント改善し、最大 18% の推論オーバーヘッドで POPE-all で 89.1 F1 を達成しました。
原文 (English)
Constraint-Anchored Reasoning Traces
Autoregressive multimodal large language models (MLLMs) suffer from error snowballing: a single incorrect inference early in a chainof-thought (CoT) trace corrupts all downstream reasoning. We find that in state-of-the-art open-source MLLMs, once the first error occurs, the reasoning cascades into failure across all remaining steps in 65% of such cases (a metric we term the snowball rate). Existing mitigations-sampling multiple chains, post-hoc self-verification, or full program synthesis-either lack symbolic grounding, catch errors too late, or sacrifice the flexibility of natural language reasoning. We propose Constraint-Anchored Reasoning Traces (CART), a neuro-symbolic framework that trains MLLMs to interleave natural language reasoning steps with symbolic constraint assertions: lightweight, machine-checkable statements about visual content (e.g., count(red_objects) = 3). A dual-pronged Constraint Propagation Module-combining a learned neural grounding head with Boolean Constraint Propagation-continuously verifies these anchors against extracted visual features and checks their mutual logical consistency. When a contradiction is detected, a backtrack controller halts generation and reverts to the last consistent checkpoint, preventing error propagation. A variable-frequency emission mechanism allows the model to adaptively control anchor density, avoiding trace bloat. We construct 218K training instances by augmenting GQA, CLEVR-CoGenT, and VCR with ground-truth constraint annotations derived from scene graphs, and fine-tune open-source MLLMs (LLaVA-NeXT, Qwen2-VL) via LoRA. On five benchmarks, CART reduces the snowball rate from 0.65 to 0.14, improves GQA accuracy by +4.6 percentage points over trainingonly baselines, and achieves 89.1 F1 on POPE-all with at most 18% inference overhead.
数値計画による多視点制約フレーム内での自律的なプロセス実行のサポート
AI 拡張ビジネス プロセス管理システム (ABPMS) は、高度な AI 技術を活用して複雑なプロセス構造を定義、実行、監視することにより、従来の BPMS を強化します。この状況の中で、フレーム化された自律性とは、事前に定義されたフレーム、つまり複数の観点にまたがる一連の制約を厳密に遵守しながら、ビジネス プロセス (BP) インスタンスの実行を自律的に進めるシステムの機能を指します。フレーム化された自律性に関する既存の研究は、主に宣言的または手続き型の制御フロー制約に焦点を当てており、通常はオートマトンベースの表現への変換に依存しています。この研究では、データを意識した時間的条件を含む多視点の制約でプロセス フレームを強化する、what-if 分析用の新しいツールを導入することで、この一連の作業を拡張します。部分的なプロセスの実行が与えられた場合、提案されたアプローチはこの強化されたフレームを利用して、基礎となるプロセス仕様に準拠した最適な継続を推奨します。さらに、この手法の拡張性と有効性を実証する実証的評価も報告し、それによって ABPMS における自律的かつ制約を意識した意思決定をサポートするその可能性を強調します。
原文 (English)
Supporting Autonomous Process Execution within a Multi-Perspective Constraint Frame via Numeric Planning
AI-Augmented Business Process Management Systems (ABPMS) enhance traditional BPMS by leveraging advanced AI techniques to define, execute, and monitor complex process structures. Within this landscape, Framed Autonomy denotes the capability of a system to autonomously advance the execution of a Business Process (BP) instance while strictly adhering to a predefined frame, i.e., a set of constraints that may span multiple perspectives. Existing research on framed autonomy has predominantly focused on control-flow constraints, either declarative or procedural, and typically relies on their transformation into automata-based representations. In this study, we extend this line of work by introducing a novel tool for what-if analysis that augments the process frame with multi-perspective constraints, including data-aware and temporal conditions. Given a partial process execution, the proposed approach exploits this enriched frame to recommend optimal continuations in compliance with the underlying process specifications. We additionally report an empirical evaluation demonstrating the scalability and effectiveness of the technique, thereby highlighting its potential for supporting autonomous and constraint-aware decision making in ABPMS.
RELIC: マルチエージェントプランニングにおける解釈可能な構成スキルを学習するための原則を明らかに
エージェントが内部実装を非公開にしながら、専門的な意思決定スキルを向上させる必要がある場合、マルチエージェントの計画は大幅に困難になります。この体制は、エージェントが独立して開発され、さまざまなインターフェイスと機能を公開しているにもかかわらず、実行可能なポリシーを共有せずに調整する必要がある場合に発生します。これまでの研究では主に集中型の最適化、共有ポリシーへのアクセス、または共通のスキル表現が想定されていたため、プライバシーに制約のある協力にはあまり適していませんでした。明らかにされた原則を通じて、解釈可能で構成可能なスキルを学習するためのフレームワークである RELIC を紹介します。各エージェントはプライベート LLM ガイド付き検索を通じて独自のプログラム スキルを磨きますが、信頼できるオーケストレーターはチーム レベルのパフォーマンスのみを通じて提案された更新を評価します。成功した動作はコードとしてブロードキャストされません。代わりに、他のエージェントが独自のインターフェイス内でインスタンスを作成し、ローカル戦略と再結合できる移植可能な原則に抽象化されます。これにより、調整が実装の共有から分離され、異種のスキル シグネチャの下でエージェント間での転送が可能になります。したがって、RELIC は、マルチエージェント計画におけるプライバシー保護スキルの学習と調整のための新しいパラダイムを導入します。
原文 (English)
RELIC: Revealed Principles for Learning Interpretable Composable Skills in Multi-Agent Planning
Multi-agent planning becomes substantially harder when agents must improve specialized decision-making skills while keeping their internal implementations private. This regime arises when agents are developed independently, expose different interfaces and capabilities, and must nevertheless coordinate without sharing executable policies. Prior research has largely assumed centralized optimization, shared policy access, or common skill representations, making it poorly suited to privacy-constrained cooperation. We introduce RELIC, a framework for learning interpretable and composable skills via revealed principles. Each agent refines its own programmatic skill through private LLM-guided search, while a trusted orchestrator evaluates proposed updates solely through team-level performance. Successful behaviors are not broadcast as code; instead, they are abstracted into portable principles that other agents can instantiate within their own interfaces and recombine with local strategies. This separates coordination from implementation sharing, enabling cross-agent transfer under heterogeneous skill signatures. RELIC thus introduces a new paradigm for privacy-preserving skill learning and coordination in multi-agent planning.
FUSAR-R1: SAR 画像のインテリジェントな解釈のための大規模推論モデル
近年、大規模な視覚言語モデルが、インテリジェントなリモートセンシング画像解釈におけるパラダイムシフトを推進しています。テキストの意味情報を組み込むことにより、解釈モデルの認知表現、意味理解、および人間とコンピューターの対話機能が大幅に向上し、合成開口レーダー (SAR) 画像解釈の分野で初期の進歩が達成されました。ただし、SAR 画像は、コヒーレント イメージング メカニズム、複雑な散乱特性、スペックル ノイズ干渉、ターゲットと背景のカップリングなどの要因の影響を受けるため、重大な不確実性と特殊化を伴う複雑で可変の画像特徴が生じます。既存の SAR ビジョン言語モデルは、人間の専門家が持つ段階的な分析、論理的判断、自己修正能力をまだ備えていないため、複雑なシナリオにおいて信頼性の高いインテリジェントな解釈をサポートすることが困難になっています。この問題に対処するために、この論文では、SAR 画像のインテリジェントな解釈のための大規模推論モデル FUSAR-R1 を提案します。このモデルはまず、人間の専門家の解釈プロセスをシミュレートすることによって明示的な思考連鎖推論データを構築し、このデータを使用して指導学習をガイドすることで、モデルに基本的な推論機能を与えます。その後、強化学習戦略が導入され、推論結果に基づいてモデルの出力が最適化され、自己修正とより信頼性の高い推論が可能になります。実験結果は、FUSAR-R1 が、ターゲットの検出、ターゲットのカウントと分類、土地被覆カテゴリーの認識など、さまざまな SAR 解釈タスクにわたって既存のマルチモーダル大規模モデルよりも一貫して優れていることを示しています。
原文 (English)
FUSAR-R1: A Large-Scale Reasoning Model for Intelligent Interpretation of SAR Images
In recent years, large-scale vision-language models have been driving a paradigm shift in intelligent remote sensing image interpretation. By incorporating textual semantic information, the cognitive expression, semantic understanding, and human-computer interaction capabilities of interpretation models have been significantly improved, achieving initial progress in the field of Synthetic Aperture Radar (SAR) image interpretation. However, SAR images are affected by factors such as coherent imaging mechanisms, complex scattering characteristics, speckle noise interference, and target-background coupling, resulting in complex and variable image features with significant uncertainties and specializations. Existing SAR vision-language models do not yet possess the step-by-step analysis, logical judgment, and self-correction capabilities of human experts, making it difficult to support reliable intelligent interpretation in complex scenarios. To address this issue, this paper proposes a large-scale reasoning model, FUSAR-R1, for intelligent interpretation of SAR images. The model first constructs explicit chain-of-thought reasoning data by simulating the interpretation process of human experts and uses this data to guide instruction learning, thereby endowing the model with basic reasoning capabilities. Subsequently, a reinforcement learning strategy is introduced to optimize the model's outputs based on inference results, enabling self-correction and more reliable reasoning. Experimental results demonstrate that FUSAR-R1 consistently outperforms existing multimodal large-scale models across various SAR interpretation tasks, including target detection, target counting and classification, and land-cover category recognition.
過負荷から洞察まで: AI エージェントが複雑なデータの分析において科学者をどのようにサポートできるか
European XFEL の科学者は、非常に大規模で複雑なデータセットを生成する実験を行っています。その後のデータ分析は、科学者がその分野の専門知識と、ドキュメント、ツール、サポート チャネルに散在する施設固有およびソフトウェア固有の知識を組み合わせる必要があるため、困難です。この問題に対処するために、私たちは科学者のニーズに合わせて調整され、ヨーロッパの XFEL の高性能コンピューティング環境と統合されたエージェント AI システムを設計および評価しました。デザイン サイエンスの研究アプローチを使用して、迅速な文献レビュー、16 の AI ツールの体系的な評価、複数のインタビュー、フォーカス グループ、および欧州 XFEL の専門家とのユーザー調査を実施し、2 つのプロトタイプを開発および評価しました。私たちの調査では、科学データ分析における主要な知識の課題を特定し、知識の検索とソース コードの生成をサポートする AI エージェントの要件を導き出し、進化する AI ツールの状況に適応できる特化したシステムの設計推奨事項を提案しています。これらの調査結果は、高度に専門化された科学環境で保守可能な AI サポートを開発するための指針を提供します。
原文 (English)
From Overload to Insights: How AI Agents Can Support Scientists in Analyzing Complex Data
Scientists at European XFEL conduct experiments that generate very large and complex datasets. The subsequent data analysis is challenging as scientists must combine their domain expertise with facility- and software-specific knowledge scattered across documentation, tools, and support channels. To address this problem, we designed and evaluated an agentic AI system tailored to the scientists' needs and integrated with the high-performance computing environment of European XFEL. Using a design science research approach, we conducted a rapid literature review, a systematic evaluation of 16 AI tools, multiple interviews, a focus group, and a user study with experts at European XFEL to develop and evaluate two prototypes. Our study identifies key knowledge challenges in scientific data analysis, derives requirements for an AI agent that supports knowledge retrieval and source code generation, and proposes design recommendations for a specialized system adaptable to the evolving AI tool landscape. These findings provide guidance for developing maintainable AI support in highly specialized scientific environments.
AgentBrew: 強い教師から弱い LLM エージェントまで生涯にわたる知識を生み出す
LLM エージェントを導入するには、トレーニング中により強力な教師が対応できる場合でも、通常、テスト時間にはコンパクトな生徒が必要です。私たちは知識の醸成、つまり教師のインタラクティブな経験を生徒の永続的な外部記憶に蒸留することを研究しています。重要なのは、これには重みの更新、専門家のデモンストレーション、グラウンドトゥルースのラベル、テスト時の教師のアクセスが必要ないことです。この設定には 2 つの課題があります。環境では、まばらなバイナリ フィードバックしか提供されないこと、教師が作成したノートは本質的に、かなり弱い生徒でも具体的に実行できるように調整する必要があることです。これらの障害に対処するために、結合された 2 つのコンポーネントで構成される AgentBrew を提案します。まず、失敗が原因の教師、ラルフ ループは、生徒の失敗を環境で検証されたメモに変換することで、まばらなフィードバックを軽減します。第 2 に、生徒を意識した合成により、教師の知識が弱い実行者の操作の粒度に合わせて調整され、モデル固有の実用的なガイダンスが得られます。コーディング、数学、およびツール使用タスクにわたる広範な評価と包括的なアブレーションにより、この非対称でトレーニング不要の作成パラダイムが、高機能でありながら展開可能な LLM エージェントを生成することが実証されました。
原文 (English)
AgentBrew: Lifelong Knowledge Brewing from Strong Teachers to Weak LLM Agents
Deploying LLM agents typically requires a compact test-time student, even if a stronger teacher is available during training. We study knowledge brewing: distilling a teacher's interactive experience into a persistent external memory for the student. Crucially, this requires no weight updates, expert demonstrations, ground-truth labels, or test-time teacher access. This setting poses two challenges: environments provide only sparse, binary feedback, and teacher-authored notes must be inherently tailored to be concretely executable by a substantially weaker student. To address these hurdles, we propose AgentBrew, comprising two coupled components. First, a failure-triggered teacher--Ralph Loop mitigates sparse feedback by transforming student failures into environment-validated notes. Second, student-aware synthesis calibrates teacher knowledge to the weak executor's operational granularity, yielding model-specific, actionable guidance. Extensive evaluations and comprehensive ablations across coding, math, and tool-use tasks demonstrate that this asymmetric, training-free brewing paradigm produces highly capable yet deployable LLM agents.
意味的等価性を超えて: LLM 不確実性定量化のための論理グラフ
大規模言語モデル (LLM) は、自信を持って記述されているにもかかわらず信頼性の低い出力を生成することが多く、安全性が重視されるアプリケーションへの展開に重大な課題をもたらします。意味論的エントロピーなどの既存の不確実性指標は、意味論的等価性のレベルで一致を捉えますが、異なる答え間の論理的関係をほとんど無視します。その結果、生成された応答の形式は多様であるが論理的に互換性がある(たとえば、粒度または特異性のみが異なる)設定では、不確実性を過大評価し、幻覚に誤ってフラグを立てる傾向があります。私たちは、回答間の含意と非互換性を明示的にモデル化するフレームワークである Logical Graph Uncertainty (LGU) を提案します。 LGU は含意チェーンに沿って確率質量を集計し、論理的に最大の仮説に対するエントロピーを計算し、それらの間の相互非互換性にペナルティを課します。複数の質問応答ベンチマークにわたって、LGU は既存の手法よりも不確実性の推定を一貫して改善し、データセット全体でセマンティック エントロピー ベースラインを最大 +7.1% AUROC および +3.5% AUARC 上回っています。
原文 (English)
Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification
Large Language Models (LLMs) often produce confidently stated yet unreliable outputs, posing critical challenges for deployment in safety-sensitive applications. Existing uncertainty metrics such as semantic entropy capture agreement at the level of semantic equivalence, but largely ignore the logical relationships between distinct answers. As a result, they tend to overestimate uncertainty and falsely flag hallucinations in settings where generated responses are diverse in form yet logically compatible (e.g., differing only in granularity or specificity). We propose Logical Graph Uncertainty (LGU), a framework that explicitly models implication and incompatibility among answers. LGU aggregates probability mass along entailment chains, computes entropy over logically maximal hypotheses, and penalizes mutual incompatibility among them. Across multiple question-answering benchmarks, LGU consistently improves uncertainty estimation over existing methods, and outperforms the semantic entropy baseline by up to +7.1% AUROC and +3.5% AUARC across datasets.
API呼び出しエージェント向けの環境フリーの合成データ生成
API 呼び出しの大規模言語モデル (LLM) エージェントをトレーニングするには、大量の高品質の軌跡が必要です。ただし、このようなデータを大規模に収集するには、通常、実行可能な API と現実的な事前設定されたバックエンド データベースを備えた完全に実装された環境が必要であり、スケーラビリティにとって大きなボトルネックとなります。これを克服するために、LLM をオンザフライのデジタル世界モデルとして活用する、環境に依存しない合成データ生成アプローチを提案します。 API 仕様のみが与えられると、私たちのメソッドはエージェントとステートフル環境の間の対話を模倣する軌跡を生成します。具体的には、LLM はまず、提供された API で解決できるさまざまなタスクを生成します。次に、教師エージェントが各タスクを繰り返し解決し、LLM シミュレーターがタスクのコンテキストとシミュレーション履歴に基づいて条件付けされた一貫した合成 API 応答を生成します。最後に、LLM 審査員が軌跡をフィルタリングして、結果として得られるデータセットの品質を保証します。私たちは、情報検索タスクと状態変更タスクの両方を含む、難しい AppWorld ベンチマークと OfficeBench ベンチマークでアプローチを評価します。合成データのモデルを微調整すると、パフォーマンスが大幅に向上し、実行可能環境がなくても API 呼び出しエージェントに対する効果的な監視を生成できることが実証されました。私たちの結果により、LLM ベースの API シミュレーションが、多様な API エコシステム全体でエージェントをトレーニングするための実用的でスケーラブルなソリューションとして確立されました。
原文 (English)
Environment-free Synthetic Data Generation for API-Calling Agents
Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting such data at scale typically requires fully implemented environments with executable APIs and realistic, pre-populated backend databases, creating a major bottleneck for scalability. To overcome this, we propose an environment-free synthetic data generation approach that leverages LLMs as on-the-fly digital world models. Given only API specifications, our method generates trajectories mimicking interactions between an agent and a stateful environment. Specifically, an LLM first generates diverse tasks solvable with the provided APIs. A teacher agent then iteratively solves each task while an LLM simulator generates coherent synthetic API responses conditioned on the task context and simulation history. Finally, an LLM judge filters the trajectories to ensure the quality of the resulting dataset. We evaluate our approach on the challenging AppWorld and OfficeBench benchmarks, which include both information-retrieval and state-changing tasks. Fine-tuning models on our synthetic data yields significant performance gains, demonstrating that effective supervision for API-calling agents can be generated without any executable environment. Our results establish LLM-based API simulation as a practical, scalable solution for training agents across diverse API ecosystems.
Lomekwi: LLM エージェントでのリソース限定ツールの検出
既存のツール使用ベンチマークは、複雑な複数ステップのタスクの単一の成功率を報告します。認知科学のアイデアに触発されて、私たちはツールの使用とツールの発見を区別し、後者を好奇心 (ツールの構築に必要な部品を発見するモデルの能力)、認識 (ツールの作成プロセスを発見するモデルの能力)、および効率 (作成後のモデルのツールの使用) に分解します。このフレームワークが Voyager などの既存の検出タスクに適用できることを示します。さらに、認識がモデルのサイズに逆比例するという証拠を提供し、これを実証する組み合わせゲームのクラスを導入して分析します。さらに、現実世界のタスクをエミュレートするように設計された別の環境での逆スケーリングを観察します。
原文 (English)
Lomekwi: Resource-Bounded Tool Discovery in LLM Agents
Existing tool-use benchmarks report a single success rate for complex, multistep tasks. Inspired by ideas from cognitive science, we distinguish tool use from tool discovery and decompose the latter into curiosity (the model's ability to discover the parts needed to build the tool), recognition (the model's ability to discover the process of creating the tool), and efficiency (the model's use of the tool after creation). We show that this framework can be applied to existing discovery tasks, such as Voyager. In addition, we provide evidence that recognition inversely scales with model size, and we introduce and analyze a class of combinatorial games that demonstrates this. We further observe inverse scaling in a separate environment designed to emulate real-world tasks.
継続的な思考連鎖モデルのトレーニング: 2 つの体制の物語
継続的思考連鎖法は、冗長な推論トレースを短いシーケンスの密な潜在表現に置き換えます。以前の連続 CoT メソッドは、潜在表現の最終状態が冗長推論トレースの状態と一致するように潜在表現を間接的に監視し、トレーニング中に自己回帰的で遅い生成を必要としました。我々は、圧縮される CoT トレース内の埋め込みの平均として各潜在をモデル化する、よりシンプルで高速な直接監視アプローチである C-MTP を導入します。私たちのアプローチは、圧縮されたトークンの分布を近似する従来の直接監視手法よりも優れており、簡素化された CoT トレース (100 トークン未満) を使用した既存の評価セットアップにおける低速の間接監視アプローチに匹敵するパフォーマンスを発揮します。最後に、連続 CoT 手法の評価を、より長い推論トレース ($\ge$ 数百の推論トークン) を持つ複雑なタスクに拡張します。この設定では、直接的および間接的な監視トレーニング手法の両方のパフォーマンスが低く (パフォーマンスが約 65\% 低下)、現在の継続的 CoT 手法の限界が明らかになりました。コードとチェックポイントは https://github.com/Varun221/cmtp_research でリリースされています。
原文 (English)
Training Continuous Chain of Thought Models: A Tale of Two Regimes
Continuous Chain-of-Thought methods replace verbose reasoning traces with a short sequence of dense latent representations. Earlier continuous CoT methods indirectly supervise the latent representations such that its final state match that of verbose reasoning traces, requiring autoregressive, slow generation during training. We introduce C-MTP, a simpler, faster direct supervision approach that models each latent as an average of the embeddings in the CoT traces to be compressed. Our approach outperforms a prior direct supervision method that approximates the distribution of compressed tokens, and performs competitively to slower indirect supervision approaches in existing evaluation setup with simplified CoT traces (less than 100 tokens). Lastly, we extend the evaluation of Continuous CoT methods to complex tasks with longer reasoning traces ($\ge$ few hundreds reasoning tokens). We find both direct and indirect supervision training methods perform poorly (roughly 65\% performance drop) in this setting, revealing the limitations of current continuous CoT methods. The code and checkpoints are released at https://github.com/Varun221/cmtp_research
rho-POMDP の信念依存ユーティリティとして期待されるフリー エネルギー
部分的な可観測性の下で行動するエージェントは、いつ情報を収集するか、どの観測がコストに見合う価値があるかを決定する必要があります。標準的な POMDP は、報酬に対する最終的な影響によってのみ情報を評価します。 $\rho$-POMDP フレームワークは代わりに、信念依存ユーティリティ $\rho$ を通じて不確実性の低減に直接報酬を与えますが、実際には $\rho$ の選択とそこに置かれる重みの両方がタスクごとに手動で調整されます。能動推論ではこの調整が完全に除去されることを示します。期待自由エネルギー (EFE) を最小化することは、期待情報利得を効用とする $\rho$-POMDP を解くことと全く同じであり、変分限界は実際的価値と認識論的価値を同じ単位 (nats) で表現するため、探索の重みは $w=1$ に固定されます。私たちは、この等価性を観察してからコミットする POMDP について証明し、それを因数分解観察 POMDP に拡張します。これは、情報を収集しても隠れた状態が変更されない、非破壊テストやモバイル センシングなどのインターリーブされた観察と行為の問題をカバーするより広範なクラスです。実験は理論を裏付けています。古典的な Tiger 問題から RockSample、および 65,000 を超える状態を含む新しい構造検査ベンチマークに至る環境全体で、調整されていない重みは、同じ地平線での報酬のみの計画と一致またはそれを上回り、タスクごとに調整されたボーナスの過剰探索を回避し、成功報酬パレート フロンティアの報酬最大化の膝近くに位置します。実際の成果は、すぐに使える探索目標です。障害検出や医療スクリーニングなどのアプリケーションでは、すべてのテストには代償があり、すべての見逃した障害には代償がかかります。EFE は、調整ではなく導出される信念依存のユーティリティを提供します。
原文 (English)
Expected Free Energy as Belief-Dependent Utility for rho-POMDPs
An agent acting under partial observability must decide when to gather information and which observations are worth their cost. Standard POMDPs value information only through its eventual effect on reward. The $\rho$-POMDP framework instead rewards uncertainty reduction directly, through a belief-dependent utility $\rho$, but in practice both the choice of $\rho$ and the weight placed on it are tuned by hand for every task. We show that active inference removes this tuning entirely. Minimizing Expected Free Energy (EFE) is exactly equivalent to solving a $\rho$-POMDP whose utility is expected information gain, and the exploration weight is fixed at $w=1$ because the variational bound expresses pragmatic and epistemic value in the same units (nats). We prove this equivalence for observe-then-commit POMDPs and extend it to factored observation POMDPs, a broader class that covers interleaved observe-act problems such as non-destructive testing and mobile sensing, where gathering information leaves the hidden state unchanged. Experiments support the theory. Across environments ranging from the classic Tiger problem to RockSample and a new Structural Inspection benchmark with over 65,000 states, the untuned weight matches or outperforms reward-only planning at the same horizon, avoids the over-exploration of bonuses tuned per task, and sits near the reward-maximizing knee of the success-reward Pareto frontier. The practical payoff is an exploration objective that works out of the box. In applications such as fault detection and medical screening, where every test has a price and every missed fault has a cost, EFE supplies a belief-dependent utility that is derived rather than tuned.
PriorProof: 形式的証明のための技術の新規性のポイントインタイムの測定
数学者は、非標準的なルートを説明、簡略化、または導入する証明を区別しますが、これらの判断を運用するのは困難です。私たちは、形式数学における時間相対証明経路の非標準性という、意図的に狭い構造を研究します。リーン定理の場合、PriorProof は精緻な証明項の依存関係フットプリントを抽出し、Mathlib の以前の四半期スナップショットのみから構築された検索条件付き階層的に平滑化された事前に基づいて、そのフットプリントの重み付き意外性をスコア付けします。この方法では、手動で構築された技術オントロジーや人間によるラベルは必要ありません。ステートメントの検索は証明から得られた対照的なペアから学習され、スコア付けされたオブジェクトは証明用語から機械的に読み取られます。ブラインドトポロジー研究では、100 個のプレゼンテーションが 76 個の異なる基礎ペアに分解されます。つまり、一貫性スクリーニングのために 3 回表示された 12 個の標準コントラストと 64 個の異なる層別ペアです。 3 人のドメイン評価者の過半数に対して、PriorProof は、11/12 の正規ペア (91.7%、64.6 ~ 98.5%) および 42/64 の層別ペア (65.6%、53.4 ~ 76.1%) を含む 53/76 ペア (69.7%、ウィルソン 95% CI 58.7 ~ 78.9%) で同意しました。スコアギャップ四分位数は、反復崩壊後は単調ではありません。エンドポイントは、ギャップが最も小さいビンでは 12/19 (63.2%、41.0 ~ 80.9%)、最大ギャップでは 16/19 (84.2%、62.4 ~ 94.5%) であり、解決された階段ではなくエンドポイントのキャリブレーション傾向をサポートしています。最良の言語モデル条件は 60/76 ペア (78.9%、68.5-86.6%) で一致します。ペアの結果では、PriorProof のみが 8 ペアで正しく、モデルのみが 15 ペアで正しいため (正確な両側マクネマー p = 0.210)、このサンプル サイズでは差は確立されません。したがって、我々は、PriorProof を専門家やモデルの判断に代わるものとしてではなく、スコア ギャップが解釈可能な信頼性指標を提供する分解可能な時間アンカー信号として提示します。
原文 (English)
PriorProof: A Point-in-Time Measure of Technique Novelty for Formal Proofs
Mathematicians distinguish proofs that explain, simplify, or introduce a nonstandard route, but these judgments are difficult to operationalize. We study a deliberately narrower construct: time-relative proof-route nonstandardness in formal mathematics. For a Lean theorem, PriorProof extracts the dependency footprint of its elaborated proof term and scores the weighted surprisal of that footprint under a retrieval-conditioned, hierarchically smoothed prior built only from an earlier quarterly snapshot of Mathlib. The method requires no hand-built technique ontology and no human labels: statement retrieval is learned from proof-derived contrastive pairs, while the scored object is read mechanically from proof terms. In a blinded topology study, 100 presentations collapse to 76 distinct underlying pairs: 12 canonical contrasts shown three times for consistency screening and 64 distinct stratified pairs. Against the majority of three retained domain raters, PriorProof agrees on 53/76 pairs (69.7%, Wilson 95% CI 58.7-78.9%), including 11/12 canonical pairs (91.7%, 64.6-98.5%) and 42/64 stratified pairs (65.6%, 53.4-76.1%). Score-gap quartiles are nonmonotone after repeat collapse; the endpoints are 12/19 (63.2%, 41.0-80.9%) in the smallest-gap bin and 16/19 (84.2%, 62.4-94.5%) in the largest, supporting an endpoint-calibration tendency rather than a resolved staircase. The best language-model condition agrees on 60/76 pairs (78.9%, 68.5-86.6%); on paired outcomes, PriorProof alone is correct on 8 pairs and the model alone on 15 (exact two-sided McNemar p = 0.210), so the difference is not established at this sample size. We therefore present PriorProof not as a replacement for expert or model judgment, but as a decomposable, time-anchored signal whose score gap provides an interpretable reliability indicator.
報酬主導型 LLM エージェントのワークフロー: 自律的な意思決定のための POMDP ルーティングと自己修正の統合
このペーパーでは、インテリジェント エージェント ワークフローを設計および最適化することで、長期計画、スパースな報酬帰属、動的な環境インタラクションなど、現在の大規模言語モデル (LLM) エージェント アプリケーションにおける主要な技術的課題に対処します。提案されたアーキテクチャは、ビジュアル、言語、生成、グラフ、マルチモーダル、強化、エージェント インテリジェンスといったコア AI パラダイムの統合に基づいています。静的なプロンプトに依存し、堅牢な知覚と行動のループを欠く従来のベースライン モデルとは異なり、私たちのアプローチでは、部分的に観察可能なマルコフ決定プロセス (POMDP) ルーティング メカニズムが導入されています。このメカニズムは、実行前に意思決定の軌跡を評価する内部の自己修正報酬モデルによって強化されています。マルチモーダル入力と高度な強化学習原理 (近接ポリシー最適化や値関数近似など) を統合することにより、エージェントは長期構造記憶を維持し、推論経路を動的に適応させてエラーの蓄積を軽減します。 ALFWorld を組み込んだシミュレーション環境と WebShop オンライン ナビゲーション ベンチマークでの実証実験では、標準の ReAct フレームワークなどの主流のベースラインと比較して、タスクの成功率と軌道効率が 24.5% 絶対的に向上していることが実証されています。包括的なアブレーション研究により、幻覚率の抑制において報酬駆動型の批判モジュールが大きく貢献していることが確認されています。この研究は、強化学習とグラフベースの記憶の理論的基盤を自律エージェントのワークフローと橋渡しします。最終的に、結果として得られるアーキテクチャは、複雑なマルチステップの自律システムで人工知能テクノロジーを開発するための実用的でスケーラブルなリファレンス フレームワークを提供します。コードは https://github.com/01Amez/RLAW_Implementation で入手できます。
原文 (English)
Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making
This paper addresses key technical challenges in current large language model (LLM) agent applications, including long-horizon planning, sparse reward attribution, and dynamic environmental interaction, by designing and optimizing an intelligent agent workflow. The proposed architecture is based on the synthesis of core AI paradigms: Visual, Language, Generative, Graph, Multimodal, Reinforcement, and Agent Intelligence. Unlike conventional baseline models that rely on static prompting and lack robust perception-action loops, our approach introduces a Partially Observable Markov Decision Process (POMDP) routing mechanism. This mechanism is augmented with an internal, self-correcting reward model that evaluates decision trajectories before execution. By integrating multimodal inputs and advanced reinforcement learning principles (such as proximal policy optimization and value function approximation), the agent maintains long-term structural memory and dynamically adapts its reasoning pathways to mitigate error accumulation. Empirical experiments on the ALFWorld embodied simulation environment and the WebShop online navigation benchmark demonstrate a 24.5% absolute improvement in task success rate and trajectory efficiency over mainstream baselines like the standard ReAct framework. Comprehensive ablation studies confirm the significant contribution of the reward-driven critique module in suppressing hallucination rates. This research bridges theoretical foundations of reinforcement learning and graph-based memory with autonomous agent workflows. Ultimately, the resulting architecture offers a practical, scalable reference framework for developing artificial intelligence technologies in complex, multi-step autonomous systems. Code is available at https://github.com/01Amez/RLAW_Implementation.
LLM が過剰回答する場合: LLM ベースのハードウェア記述言語の質問応答における品質問題の測定と軽減
大規模言語モデル (LLM) の急速な進歩により、実務家はハードウェア記述言語 (HDL) に関する質問に答えるために LLM にますます依存するようになりました。 HDL は最終的に物理ハードウェアに合成されるため、不正確または冗長な応答は、デザイン フローの後半になって初めて表面化するタイミング違反や合成不可能なロジックに伝播する可能性があり、HDL 応答の品質が特に重要になります。ただし、LLM によって生成された応答の質は、特に人間の専門家によって提供された回答と比較した場合、依然として不明瞭です。この質問を調査するために、Stack Overflow から受け入れられた回答を含む 6,246 件の HDL Q&A 投稿を収集し、それらをデータセットにまとめ、4 つの主要カテゴリ (概念、デバッグ、生成、最適化) と 10 のサブカテゴリの分類に編成しました。このデータセットを使用して、1 ~ 3 年の経験を持つ 19 人の HDL エンジニアを対象に実施されるユーザー調査を設計します。私たちの調査結果は、広範な過剰回答の傾向を明らかにしています。LLM は正しいコンテンツを提供しますが、冗長な代替案 (65.7%) や冗長なパディング (69.1%) の下に埋め込んでいますが、回答のほぼ半数 (49.0%) は専門家の回答と完全に一致していませんが、それでも参加者は読みやすさのために LLM の回答を好みました (58.3%)。これらの発見に動機付けられて、私たちは LLM ベースの HDL 質問応答を改善するためのマルチエージェント フレームワークを提案します。私たちは、LLM-as-Judge と 2 つの構造指標を使用して回答の品質を評価します。LLM は複数の代替ソリューションを提供することが多いため、冗長性を反映するコア回答の数と、冗長性を反映する非コアコンテンツの長さです。 4 つの主流 LLM で評価された当社のフレームワークは、5 段階評価で、コア回答の平均品質スコアが 3.71 から 4.67 (+0.96)、非コア コンテンツの品質が 3.72 から 4.23 (+0.51) に向上しました。
原文 (English)
When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering
The rapid advancement of large language models (LLMs) has led practitioners to increasingly rely on them for answering questions about hardware description languages (HDLs). Because HDL is ultimately synthesized into physical hardware, an imprecise or redundant answer can propagate into timing violations or non-synthesizable logic that surface only late in the design flow, making the quality of HDL answers especially consequential. However, the quality of LLM-generated responses, particularly in comparison with answers provided by human experts, remains unclear. To investigate this question, we collect 6,246 HDL Q&A posts with accepted answers from Stack Overflow and curate them into a dataset, organized into a taxonomy of four main categories (Conceptual, Debugging, Generation, and Optimization) and ten subcategories. Using this dataset, we design a user study conducted with 19 HDL engineers with one to three years of experience. Our findings reveal a pervasive over answering tendency: LLMs supply correct content but bury it under redundant alternatives (65.7%) and verbose padding (69.1%), while nearly half of answers (49.0%) fail to fully align with expert answers yet participants still preferred LLM responses for readability (58.3%). Motivated by these findings, we propose a multi-agent framework for improving LLM-based HDL question answering. We evaluate answer quality using an LLM-as-Judge and two structural metrics: the number of core answers, which reflects redundancy since LLMs often provide multiple alternative solutions, and the length of non-core content, which reflects verbosity. Evaluated on the four mainstream LLMs, our framework increases the average core-answer quality score from 3.71 to 4.67 (+0.96) and the non-core content quality from 3.72 to 4.23 (+0.51), on a five-point scale.
情報ギャップを埋める: コールドスタート予測のためのセマンティック高密度化と後知恵蒸留
新規ユーザーのコールドスタートは、インタラクション履歴がまばらなユーザーのユーザー生涯価値 (LTV) とコンバージョン率 (CVR) を予測するため、e コマース プラットフォームにとって重大なボトルネックです。これまでの 2 つの方向、LLM ベースのセマンティック拡張と特権情報を使用した学習 (LUPI) は、それぞれ重要な制限に直面しています。まず、LLM の拡張により、ノイズが多く、運用環境での運用が困難な構造化されていない理論的根拠が生成されます。第二に、素朴な生徒と教師の分離は、特権的な教師と少数の生徒の間の情報のギャップにより脆弱になる可能性があります。さらに、このギャップはユーザー間で不均一です。私たちは、両方の制限に対処する意味論的推論を認識した蒸留フレームワークである SemRaD を提案します。まず、構造化セマンティック推論パイプラインは、自由形式の根拠を Discover-curate-audit ワークフローを通じて構築された構造化スキーマに置き換え、ユーザーごとに高密度セマンティック プロファイル (最も有益な次元に焦点を当てたセマンティック ゲート エンコーダーを介してデプロイされた学生によって消費される) と、変換前後の推論から調整された後知恵蒸留ターゲット (トレーニング時にのみ使用される) を生成します。第 2 に、このギャップを埋めてその異質性に対処するために、後知恵認識蒸留ネットワークは後知恵ターゲットを介して特権的な知識を転送し、蒸留エキスパートがユーザーごとの変動の下で転送を改善します。大規模な産業用データセットでは、SemRaD は実稼働グレードのベースと比較して +1.9% LTV (Gini) および +1.0% CVR (AUROC) を向上させます。 Keeta での 4 週間のオンライン A/B では、LTV +1.0% / CVR +0.43% が確認されました。また、SemRaD はトレーニング データの 9% のみを使用して実稼働システムの LTV と一致し、CVR を 0.8% 改善します。
原文 (English)
Bridging the Information Gap: Semantic Densification and Hindsight Distillation for Cold-Start Prediction
New-user cold-start is a critical bottleneck for e-commerce platforms: predicting user lifetime value (LTV) and conversion rate (CVR) for users with sparse interaction history. Two prior directions -- LLM-based semantic augmentation and learning using privileged information (LUPI) -- each face a key limitation. First, LLM augmentation produces unstructured rationales that are noisy and hard to operationalize in production. Second, naive student-teacher distillation can be brittle due to an information gap between the privileged teacher and the sparse student; moreover, this gap is heterogeneous across users. We propose SemRaD, a Semantic Reasoning-aware Distillation framework addressing both limitations. First, a Structured Semantic Reasoning Pipeline replaces free-form rationales with a structured schema built via a discover-curate-audit workflow, producing per user a Densified Semantic Profile (consumed by the deployed student via a Semantic-Gated Encoder that focuses on the most informative dimensions) and a Hindsight Distillation Target reconciled from pre- and post-conversion reasoning (used only at training). Second, to bridge this gap and handle its heterogeneity, a Hindsight-Aware Distillation Network transfers privileged knowledge via the hindsight target, with Distillation Experts improving transfer under per-user variability. On a large-scale industrial dataset, SemRaD lifts +1.9% LTV (Gini) and +1.0% CVR (AUROC) over a production-grade base; a four-week online A/B at Keeta confirms +1.0% LTV / +0.43% CVR. SemRaD also matches the production system's LTV using only 9% of the training data while improving CVR by 0.8%.
Otap:エージェントの軌跡における計画と実行を評価するための構造を意識した最適なトランスポート
大規模言語モデル エージェントは、計画、ツール呼び出し、中間結果を交互に配置する軌跡を生成することでタスクを解決します。現在の評価メトリクスは、そのような軌跡をバイナリの成功フラグに換算するか、完全一致によって参照と比較します。成功フラグは、健全な解決策と運によって成功した解決策を区別することはできず、失敗した実行が失敗した理由については何も示しません。完全一致では、有効ではあるが、参照とは異なる方法で並べ替えられたり、分解されたりした計画にペナルティが課されます。軌道評価をエージェントの実行グラフと有効なソリューション グラフのセットの間の距離として再構成し、属性付きの依存関係グラフに対する不平衡融合グロモフ-ワッサーシュタイン輸送問題を介してインスタンス化します。 \otap{} (エージェントティック プランニングのための最適なトランスポート) と呼ばれる結果のスコアは、依存関係を保持する並べ替えに対して不変であることが証明されており、冗長なステップに対して制限された感度を持つ疑似メトリックです。そのアンバランスなマージナルは、一致を強制することなく欠落ステップや幻覚ステップを処理し、ソフト カップリングは計画粒度の変動に対応します。制御された摂動と 3 つの公開ベンチマークに関して、\otap{} は、セマンティクスのみのメトリクスのスコアが確率を下回る領域で、有効な軌道と無効な軌道を分離します。その精度は、依存関係グラフが正確に復元された場合に最も高く、グラフがフリーテキスト トレースからヒューリスティックに推測された場合にのみ低下します。
原文 (English)
Otap:Structure-Aware Optimal Transport for Evaluating Planning and Execution in Agent Trajectories
Large language model agents solve tasks by generating trajectories that interleave planning, tool calls, and intermediate results. Current evaluation metrics reduce such a trajectory to a binary success flag or compare it against a reference by exact matching. A success flag cannot distinguish a sound solution from one that succeeds by luck, and says nothing about why a failed run went wrong. Exact matching penalizes plans that are valid but reordered or decomposed differently from the reference. We reframe trajectory evaluation as a distance between the agent's execution graph and a set of valid solution graphs, and instantiate it via an unbalanced fused Gromov-Wasserstein transport problem over attributed dependency graphs. The resulting score, termed \otap{} (Optimal Transport for Agentic Planning), is a pseudo-metric that is provably invariant to dependency-preserving reorderings and has bounded sensitivity to redundant steps. Its unbalanced marginals handle missing or hallucinated steps without forcing a match, and its soft coupling accommodates variation in plan granularity. On controlled perturbations and three public benchmarks, \otap{} separates valid from invalid trajectories in a regime where semantics-only metrics score below chance. Its accuracy is highest when the dependency graph is recovered exactly, and drops only when the graph is inferred heuristically from free-text traces.
数値天気予報を使用したフーリエ幾何学的風力発電予測
風力発電の短期的な正確な予測は、送電網の安定性と運用計画に不可欠ですが、大気条件とタービンのダイナミクス間の複雑な相互作用のため、依然として困難が続いています。しかし、既存の方法では、天気予報と風力タービン データ (つまり SCADA) を効果的に組み込むことができず、次善のソリューションが得られます。これに対処するために、過去のポイントベースの SCADA データとグリッドベースの数値天気予報 (NWP) 予報を統合するマルチモーダル フレームワークを導入します。これは、異種入力と複雑な物理的風力タービン相互作用により困難です。私たちのアプローチでは、まず入力をスカラー特徴とベクトル特徴に明示的に分解して、サイト固有の依存関係と幾何学的依存関係の両方をより適切に捕捉し、次に幾何学的エンコーダーを組み込んで風ベクトルから回転不変特徴を抽出します。さらに、周波数領域でグローバルな畳み込みを実行して長距離の時空間関係を効率的にモデル化するフーリエ ニューラル オペレーター (FNO) アーキテクチャを活用します。気象予測データを使用した現実世界の 3 つの風力発電所での広範な実験により、私たちのモデルが常に最先端のベースラインを上回るパフォーマンスを示し、物理情報に基づいた設計の有効性が強調されています。私たちのメソッドのコア実装は、https://github.com/shawn-sypiao/GWPF で公開されています。
原文 (English)
Fourier Geometric Wind Power Forecasting with Numerical Weather Prediction
Accurate short-term wind power forecasting is essential for grid stability and operational planning, yet remains challenging due to the complex interactions between atmospheric conditions and turbine dynamics. However, existing methods fail to effectively incorporate weather forecasting with wind turbine data (i.e., SCADA), leading to suboptimal solutions. To address this, we introduce a multimodal framework that integrates historical point-based SCADA data with grid-based Numerical Weather Prediction (NWP) forecasts, which is challenging due to heterogeneous input and the complex physical wind-turbine interactions. Our approach first explicitly decomposes inputs into scalar and vector features to better capture both site-specific and geometric dependencies and then incorporates a geometric encoder to extract rotation-invariant features from wind vectors. We further leverages a Fourier Neural Operator (FNO) architecture, which performs global convolutions in the frequency domain to efficiently model long-range spatiotemporal relationships. Extensive experiments on three real-world wind farms, with weather forecasting data, demonstrate that our model consistently outperforms state-of-the-art baselines, highlighting the effectiveness of its physically-informed design. The core implementation of our method is publicly available at: https://github.com/shawn-sypiao/GWPF.
検索拡張リーダーによるサポートの使用方法を形成する証拠インターフェイス
マルチホップ RAG 評価では、上位 k 位の応答スコアによって 2 つの異なる失敗が隠蔽される可能性があります。取得ウィンドウがサポート チェーンの一部を削除する可能性があるか、適応されたリーダーが適切に使用しない形式のサポートが含まれている可能性があります。私たちは、この取得された証拠の読者向け形式を証拠インターフェイスと呼びます。 3 つのサポート アノテーション付きマルチホップ QA ベンチマークを使用して、生のコンテキスト、取得ウィンドウ、およびゴールド サポートの診断レンダリングでトレーニングされた、適合した適応リーダーを比較します。これらの比較により、サポートの利用可能性の障害と残りのリーダー インターフェイスの影響が区別されます。 Top-k ウィンドウは、完全なアノテーション付きサポート チェーンが存続するかどうかを確認した後にのみ解釈可能になります。存続する場合、ランクの短いウィンドウは生のコンテキストと一致するか、生のコンテキストよりも改善されます。そうでない場合は、サポートの欠如が損失の大部分を説明します。ゴールド サポート第一により、一致するリーダーが向上します。 2Wiki と MuSiQue では、サポートの監視下にあるランカーが、ゴールドのヘッドルームを維持しながら、カバレッジを高め、より低い即時コストで生のコンテキストの品質を回復します。サポート除去チェックでは、利益が事前回答だけではなく、暴露された証拠に依存していることがさらに示されています。したがって、サポート注釈付きの評価では、上位 k 位までの回答スコアを完全なサポート範囲とともに報告する必要があります。
原文 (English)
Evidence Interfaces Shape How Retrieval-Augmented Readers Use Support
In multi-hop RAG evaluation, a top-k answer score can hide two different failures: the retrieval window may drop part of the support chain, or it may contain support in a form the adapted reader does not use well. We call this reader-facing form of retrieved evidence an evidence interface. Using three support-annotated multi-hop QA benchmarks, we compare matched adapted readers trained with raw context, retrieval windows, and gold-support diagnostic renderings. These comparisons distinguish support-availability failures from remaining reader-interface effects. Top-k windows become interpretable only after checking whether the complete annotated support chain survives: when it does, short ranked windows can match or improve over raw context; when it does not, missing support explains much of the loss. Gold support-first improves matched readers; on 2Wiki and MuSiQue, a support-supervised ranker raises coverage and recovers raw-context quality at lower prompt cost, while retaining gold headroom. Support-removal checks further show that the gains rely on exposed evidence, not only answer priors. On support-annotated evaluations, top-k answer scores should therefore be reported together with complete-support coverage.
AI エージェントの動作の診断フレームワーク
AI エージェントは、行動科学者が研究するのと同じ臨床、政治、科学、社会システム内で行動することが増えています。これらのシステムを評価するには、ソースレベルの診断が必要です。同じ行動パターンが、エージェントの表現基質から、またはその表現を形成する役割、目的、相互作用構造、およびガバナンス規則から生じる可能性があります。この観点では、AI エージェントの動作の診断フレームワークであるレイヤー アトリビューションを提案します。基礎的な計算層は、アーキテクチャ、記憶、知覚、注意、表現を通じてどのような動作が可能かを定義します。行動調整層は、アイデンティティ、リソース、目的、社会的相互作用、制度上の制約、ガバナンスを通じてこれらの能力がどのように表現されるかを形成します。このフレームワークは 3 つの結果を明確にしています。サロゲートの妥当性はモデルとタスク層の関係、人間と AI の相違は診断証拠を提供し、ガバナンスには介入前にソースの帰属が必要です。したがって、AI エージェントを行動アクターとして扱うには、行動を説明、検証、または管理する方法を決定する前に、行動がどこから発生するかを判断する評価方法が必要です。
原文 (English)
A Diagnostic Framework for AI Agent Behavior
AI agents increasingly act within the same clinical, political, scientific, and social systems that behavioral scientists study. Evaluating these systems requires source-level diagnosis: the same behavioral pattern may arise from an agent representational substrate or from the roles, objectives, interaction structures, and governance rules that shape its expression. This Perspective proposes a diagnostic framework for AI agent behavior: layer attribution. The foundational computational layer defines what behaviors are possible through architecture, memory, perception, attention, and representation. The behavioral modulation layer shapes how those capacities are expressed through identity, resources, objectives, social interaction, institutional constraints, and governance. The framework clarifies three consequences: surrogate validity is a model-task-layer relation, human-AI divergence provides diagnostic evidence, and governance requires source attribution before intervention. Treating AI agents as behavioral actors therefore requires evaluation methods that determine where behavior originates before deciding how to explain, validate, or govern it.
あなたのモデルは考えているのでしょうか、それとも単に停滞しているのでしょうか? PUMA: 位相と運動量の調整による推論病理の診断
テスト時間のスケーリングにより、大規模推論モデル (LRM) が広範な思考連鎖 (CoT) を通じて複雑なタスクに取り組むことができるようになります。ただし、これは多くの場合、冗長な推論によって精度が保証されずに計算オーバーヘッドが増加する「考えすぎ」のパラドックスを引き起こします。既存のテスト時間効率の最適化手法は、主に 2 つのカテゴリに分類されます。1 つは不確実性が低いため幻覚が隠蔽される「欺瞞的収束」を起こしやすい情報理論的アプローチ、もう 1 つは事後的なことが多く、動的推論に対するリアルタイム感度が欠けている潜在表現分析です。このギャップを埋めるために、私たちはまず位相運動量整合仮説を立て、推論の正しさは幾何学的な運動量と不確実性の解決の間の時間的同期に依存すると主張します。次に、潜在速度とねじれによって定量化される幾何学的認知努力とエントロピー認知不確実性という 2 つの直交する次元を通じてこれらのダイナミクスを特徴付ける認知エネルギー モデルを理論的に定式化します。これを運用するために、階層型診断アーキテクチャを採用したトレーニング不要のフレームワークである PUMA (Phase-Uncertainty Momentum Alignment) を導入します。軽量の位相モニタリングとイベントトリガーの幾何学的解析を組み合わせることで、PUMA は能動的な探査と受動的な停滞を効果的に区別し、適応的な切り捨てや修正措置による正確な介入を可能にします。 1.5B から 32B にわたる LRM に関する広範な実験により、PUMA がさまざまなベンチマークにわたって一貫して最先端のベースラインを上回り、優れた精度と効率のトレードオフと堅牢なクロスドメイン一般化を達成していることが実証されました。
原文 (English)
Is Your Model Thinking or Just Stagnating? PUMA: Diagnosing Reasoning Pathology via Phase-Momentum Alignment
Test-time scaling empowers Large Reasoning Models (LRMs) to tackle complex tasks via extensive Chain-of-Thought (CoT). However, this often induces the "overthinking" paradox, where redundant reasoning increases computational overhead without guaranteeing accuracy. Existing test-time efficiency optimization methods primarily fall into two categories: information-theoretic approaches, which are prone to "deceptive convergence" where low uncertainty masks hallucinations, and latent representation analyses, which are often post-hoc, lacking the real-time sensitivity for dynamic reasoning. To bridge this gap, we first posit the Phase-Momentum Alignment Hypothesis, asserting that reasoning correctness hinges on the temporal synchronization between geometric momentum and uncertainty resolution. We then theoretically formulate the Cognitive-Energy Model to characterize these dynamics through two orthogonal dimensions: Geometric Cognitive Effort, quantified by latent velocity and tortuosity, and Entropic Cognitive Uncertainty. To operationalize this, we introduce PUMA (Phase-Uncertainty Momentum Alignment), a training-free framework employing a tiered diagnostic architecture. By coupling lightweight phase monitoring with event-triggered geometric analysis, PUMA effectively distinguishes active exploration from passive stagnation, enabling precise interventions through adaptive truncation or corrective measures. Extensive experiments on LRMs spanning 1.5B to 32B demonstrate that PUMA consistently outperforms state-of-the-art baselines across diverse benchmarks, achieving a superior accuracy-efficiency trade-off and robust cross-domain generalization.
擬人化された対話に向けて: 人間のようなチャットの生成、評価、および好みの調整のための閉ループ フレームワーク
人間のようなプライベート チャットには、流暢な応答生成以上のものが必要です。システムは、ペルソナ、関係、記憶、限定された知識、媒体固有のタイミング、一貫したマルチターン アークを保持する必要があります。我々は、擬人化対話をシステム アーキテクチャ、実行可能評価、および診断調整の共同問題として定式化する閉ループ フレームワークである AnthroDial を紹介します。これは、(1) 役割条件付きのスケジュールされた対話ランタイムと、ペルソナおよびシナリオ カード、長期記憶、仮想時間、および単一草案メッセージの決定を組み合わせます。 (2) L0 有効性ゲート、5 つのターンごとのディメンション、および 5 つのダイアログ レベルのディメンションを備えた実行可能なベンチマーク。 (3) SFT 用に 16,436 のスケジュールされた決定例をフィルタリングし、認知診断、ZPD を意識した報酬を備えた GRPO を適用するトレーニング後のパイプライン。この報酬は、各行動次元のカルマン フィルター処理された能力推定値を維持し、より大きな能力不足のある次元を重み付けし、ロールアウト スコアをタスク レベルの ZPD マッチとして使用して、学習可能な弱いスキルに焦点を合わせて最適化します。モデルごとに 55 のペルソナ、50 のシナリオ、50 のペルソナとシナリオのバインディング、および 100 の役割条件付きケースを含むベンチマークで、フロンティア ベースライン、オープン モデル、思考/非思考のバリアント、および SFT/RL アブレーションにわたる 16 のシステムを評価します。最も強い非トレーニングベースラインは 32.00% の厳密な ACC に達しますが、Qwen3.6-27B-SFT+RL は 39.00% の厳密な ACC と 98.5 の全体スコアに達します。 9B ノーシンク設定では、SFT と RL は厳密な ACC を 0.00% から 13.00% および 18.37% に改善します。これらの結果は、生成、評価、報酬形成が同じ行動次元を共有する場合、擬人化対話が有益であることを示しています。
原文 (English)
Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment
Human-like private chat requires more than fluent response generation: a system must preserve persona, relationship, memory, bounded knowledge, medium-specific timing, and a coherent multi-turn arc. We present AnthroDial, a closed-loop framework that formulates anthropomorphic dialogue as a joint problem of system architecture, executable evaluation, and diagnostic alignment. It combines (1) a role-conditioned scheduled dialogue runtime with persona and scenario cards, long-term memory, virtual time, and single-draft message decisions; (2) an executable benchmark with an L0 validity gate, five per-turn dimensions, and five dialogue-level dimensions; and (3) a post-training pipeline that filters 16,436 scheduled-decision examples for SFT and applies GRPO with a cognitive-diagnostic, ZPD-aware reward. The reward maintains Kalman-filtered capability estimates for each behavioral dimension, upweights dimensions with larger capability deficits, and uses rollout scores as task-level ZPD matches to focus optimization on learnable weak skills. On a benchmark with 55 personas, 50 scenarios, 50 persona-scenario bindings, and 100 role-conditioned cases per model, we evaluate 16 systems spanning frontier baselines, open models, thinking/no-think variants, and SFT/RL ablations. The strongest non-trained baseline reaches 32.00% strict ACC, while Qwen3.6-27B-SFT+RL reaches 39.00% strict ACC and a 98.5 overall score. In the 9B no-think setting, SFT and RL improve strict ACC from 0.00% to 13.00% and 18.37%. These results show that anthropomorphic dialogue benefits when generation, evaluation, and reward shaping share the same behavioral dimensions.
コードエージェントの LoRA 微調整のための軌跡データキュレーションの体系的な評価
エキスパート エージェントの軌道に関するオープンウェイト LLM の教師あり微調整 (SFT) は、独自のモデルに依存せずに有能なコード エージェントを構築するための顕著なアプローチとして浮上しています。中心的であるにもかかわらず十分に解明されていない問題は、軌道の質と量がどのように連携してモデルのパフォーマンスを形成するかということです。我々は、SWE 軌跡データセット (67,074 の軌跡、うち 32,161 が解決済み) 上の Qwen2.5-Coder-7B-Instruct の LoRA 微調整のための軌跡データ フィルタリングの体系的な実証研究を紹介します。当社は、効率とスタイルという 2 軸の品質スコアリング フレームワークを提案し、戦略、規模、アブレーション分析にわたる 16 の対照実験を通じてそれを評価します。 7B スケール モデルはほぼゼロの SWE ベンチ解決率を達成するため、ホールドアウト軌道上のクロス エントロピー (CE) 損失を主要な指標として採用し、ファースト アクション生成によって検証されます。CE 損失と ROUGE-L は完全に順位相関しており (Spearman $\rho$ = -1.00)、限られたサンプルの証拠はこの代理を裏付けていますが、決定的に確立しているわけではありません。私たちの結果は、スケール依存の品質と量のトレードオフを明らかにしました。小規模では、データセットを 2 倍にする (500 から 1,000) と、最大 12.7% の CE 損失の削減が得られますが、TopQ-Random ギャップは 0.10 のままです。 2,000 の軌道では、この同じギャップは 3.6% に広がります (p = 0.016)。アブレーションではさらに、エラー再試行率が支配的なサブディメンションであることが特定され、完全なコンポジット ($\Delta$ < 0.2%) と同等のパフォーマンスを示します。まとめると、これらの発見は、コードエージェント SFT の実行可能だが規模に依存する手段としての軌道レベルの品質スコアリングを確立し、エンドツーエンドの解決率が統計的に実現不可能なレジームに対して代理検証済みの評価プロトコルを提供します。
原文 (English)
A Systematic Evaluation of Trajectory Data Curation for LoRA Fine-Tuning of Code Agents
Supervised fine-tuning (SFT) of open-weight LLMs on expert agent trajectories has emerged as a prominent approach to building capable code agents without reliance on proprietary models. A central yet underexplored question is how trajectory quality and quantity jointly shape model performance. We present a systematic empirical study of trajectory data filtering for LoRA fine-tuning of Qwen2.5-Coder-7B-Instruct on the SWE-trajectory dataset (67,074 trajectories, of which 32,161 are resolved). We propose a two-axis quality scoring framework -- Efficiency and Style -- and evaluate it through 16 controlled experiments spanning strategy, scale, and ablation analyses. Since 7B-scale models attain near-zero SWE-bench resolve rates, we adopt cross-entropy (CE) loss on held-out trajectories as the primary metric, validated via first-action generation: CE loss and ROUGE-L are perfectly rank-correlated (Spearman $\rho$ = -1.00), with limited-sample evidence supporting but not conclusively establishing this proxy. Our results reveal a scale-dependent quality-quantity trade-off: at small scales, doubling the dataset (500 to 1,000) yields ~12.7% CE-loss reduction whereas the TopQ-Random gap stays 0.10); at 2,000 trajectories this same gap widens to 3.6% (p = 0.016). Ablation further identifies error-retry rate as the dominant sub-dimension, performing comparably to the full composite ($\Delta$ < 0.2%). Together, these findings establish trajectory-level quality scoring as a viable but scale-sensitive lever for code-agent SFT and offer a proxy-validated evaluation protocol for the regime where end-to-end resolve rate is statistically infeasible.
制約付きパス推論: コミットされたステージがいつコストを獲得するかを測定する
LLM 推論パイプラインのコミットされた中間ステージはいつコストを獲得しますか? Constrained Path Reasoning (CPR) は、ソース認識パス仮説とステージレベルのアカウンティングを組み合わせます。検索により暫定的な状態が生成されます。信頼できる、または検証された不変条件はハード制約を行うことができますが、他の提案はソフトなままで修正可能です。 CPR は、タスク互換のコミットメントによって遷移が考慮され、候補者が集中し、規則性が誘導され、ゲインが伝播されたエラーと実行コストを超えるとフィードバックが露出する可能性があると予測されています。この形式主義は、個別のコミットメントと連続的なフローをカバーし、効果的な分岐、エンドポイントの集中、および使用可能な出力あたりのコストを測定します。 1,180 の生成された QCQP と 40 の操作された縮退多項式インスタンス (2,140 エンドポイント) にわたって、残りのトリアージは、試行の 17.7% で、repair-all の追加の実行可能収量の 63.0% を回収します。固定 LLM アカウンティング (ネストされたアーム間で共有される 270 の一意の呼び出し) では、直接 41.1%、形式化および確定的実行後 90.0%、ワンショット凸化後 20.0%、およびフル パスの場合 21.1% の使用可能な収率が見つかりました。 120 回のペア条件呼び出しで、2 アクション ロールバック ルールの使用可能収率は 90% に達しましたが、フィードバック条件付きセレクターの場合は 36.7% でした。 2 つのエンドポイント プローブがソースと検証を分離します。72 出力のクロストラジェクトリー移植により、エントロピーと許容可能な質量が減少します。 24 出力の同一コールの自己提案パイロットは、変更されていない 2 回反復衝突エントロピー、25.0% 対 8.3% の使用可能収率、および 1/8 の決定論的に確認されたエンドポイント チェックを提供します。モデルによって生成された状態は仮説を提供します。信頼された実行により制約強度が得られます。
原文 (English)
Constrained Path Reasoning: Measuring When Committed Stages Earn Their Cost
When does a committed intermediate stage in an LLM reasoning pipeline earn its cost? Constrained Path Reasoning (CPR) pairs a source-aware path hypothesis with stage-level accounting. Search generates provisional states; trusted or validated invariants can constrain hard, while other proposals remain soft and revisable. CPR predicts that task-compatible commitments can factor transitions, concentrate candidate mass, induce regularity, and expose feedback when their gains exceed propagated error and execution cost. The formalism covers discrete commitments and continuous flows and measures effective branching, endpoint concentration, and cost per usable output. Across 1,180 generated QCQPs and 40 engineered degenerate polynomial instances (2,140 endpoints), residual triage recovers 63.0% of repair-all's additional feasible yield with 17.7% of its attempts. Fixed-LLM accounting (270 unique calls shared across nested arms) finds usable yield of 41.1% direct, 90.0% after formalization and deterministic execution, 20.0% after one-shot convexification, and 21.1% for the full path. In 120 paired-condition calls, a two-action rollback rule reaches 90% usable yield versus 36.7% for the feedback-conditioned selector. Two endpoint probes separate source from validation: a 72-output cross-trajectory transplant reduces entropy and acceptable mass; a 24-output same-call self-proposal pilot gives unchanged two-repeat collision entropy, 25.0% versus 8.3% usable yield, and 1/8 deterministically confirmed endpoint checks. Model-generated states supply hypotheses; trusted execution earns constraint strength.
LenGuard-GPC: 空間推論のためのガイド付きプロンプトの一貫性を備えた長さの保護で学習を強化
マルチビューの空間推論には、画像間の視覚的証拠を比較し、オブジェクトの対応を調整し、長い視覚的コンテキストにわたって空間関係を推測するための視覚言語モデルが必要です。この設定では、思考連鎖推論が正確になることなく冗長になる傾向があります。検証可能な報酬を伴う強化学習はこのタスクに自然に適合しますが、標準的な GRPO 報酬はまばらな結果レベルのフィードバックに依存しており、推論の軌道がどこで間違っているかについてのシグナルも、その長さの制御も提供しません。私たちは、両方の問題を一緒に解決する高密度報酬フレームワークである LenGuard-GPC を提案します。サンプリングされた各軌跡について、標準プロンプトとガイド付きプロンプトの下でトークンごとの予測分布を比較し、結果として得られるトークン合計 KL 発散を密な報酬信号として使用します。この KL ペナルティはトークンを超えて蓄積され、そうでなければ応答の品質に関係なく短い応答に報酬が与えられることになるため、単純に簡潔さを促すことなく推論の長さを制御された範囲内に保つ段階的な長さのボーナスを導入します。 6 つのマルチビュー空間推論ベンチマークにおいて、LenGuard-GPC は標準的な GRPO よりも精度を向上させながら、平均応答長を短縮します。
原文 (English)
LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning
Multi-view spatial reasoning requires vision-language models to compare visual evidence across images, align object correspondences, and infer spatial relations over long visual contexts, a setting where chain-of-thought reasoning tends to grow verbose without becoming more accurate. Reinforcement learning with verifiable rewards is a natural fit for this task, but standard GRPO reward relies on sparse outcome-level feedback and gives no signal about where a reasoning trajectory goes wrong, nor any control over its length. We propose LenGuard-GPC, a dense reward framework that addresses both problems together. For each sampled trajectory, it compares the token-wise predictive distributions under a standard prompt and a guided prompt, and uses the resulting token-sum KL divergence as a dense reward signal. Since this KL penalty accumulates over tokens and would otherwise reward shorter responses regardless of their quality, we introduce a staged length bonus that keeps reasoning length within a controlled range without simply encouraging brevity. On six multi-view spatial reasoning benchmarks, LenGuard-GPC improves accuracy over vanilla GRPO while reducing average response length.
隠された相関下での反復モード検出による調整されたもつれの解除
解きほぐされた表現学習は、堅牢な属性予測のための強力なパラダイムです。最近の方法では属性の相関関係に取り組んでいますが、特定の属性の値に基づくデータが他の属性と相関する基礎的なモードを示す隠れた相関関係はまだ調査されていません。モード情報を保存し、もつれを解くために、共同でモードを発見し、モードベースの条件付き独立性を強制します。ただし、これら 2 つのモジュール間の相互依存関係により、単純な反復では誤差が増幅される可能性があります。我々は、モードの数の進化に適応する動的アーキテクチャと、メタ最適化によるエラー増幅を軽減する調整メカニズムを特徴とするエンドツーエンドのフレームワークである、反復モード検出による調整された解絡解除 (CoDID) を提案します。実証結果は、さまざまなタスクにおける最先端のパフォーマンスを実証しています。
原文 (English)
Coordinated Disentanglement with Iterative Mode Discovery Under Hidden Correlations
Disentangled representation learning is a powerful paradigm for robust attribute prediction. While recent methods address attribute correlations, hidden correlations remain underexplored, where data under the value of a certain attribute exhibit underlying modes correlated with other attributes. To preserve mode information and achieve disentanglement, we jointly discover modes and enforce mode-based conditional independence. Yet, the interdependency between these two modules may lead to error amplification under naive iterations. We propose Coordinated Disentanglement with Iterative mode Discovery (CoDID), an end-to-end framework featuring a dynamic architecture that adapts to evolving number of modes, and a coordination mechanism that mitigates error amplification via meta-optimization. Empirical results demonstrate the state-of-the-art performance on diverse tasks.
データファーストオントロジーに基づく明示的な世界モデル: DaoQL マルチモーダルストレージ検証と反事実推論評価
大規模な言語モデルは世界モデルをニューラル ウェイトで暗黙的にエンコードするため、医療や金融などの高精度の領域では、幻覚、知識の凍結、説明可能性の低さ、修正可能性の低さという 4 つの構造的リスクが露呈します。この論文では、データファースト オントロジーを提案します。LLM は推論および言語エンジンとして扱われ、決定論的な知識は明示的なマルチモーダル データベースである DaoQL に移動されます。我々は、明示的な世界モデルを形式化し、ルールの独立性、決定論的な評価、および固定された競合解決の下で、明示的なモデルが構成可能な反事実分解可能性のための十分な条件を提供することを示します。暗黙的モデルにはアトミックな読み取り/デルタ セマンティクスが欠けているため、同等のアーキテクチャ上の保証はありません。実装されたシステムは、DaoQL の検証済みストレージ レイヤーと明示的な Eval パスに焦点を当てており、グラフ、列、ベクトル、およびフルテキスト エンジンを 1 つのプロセス内に統合しています。 KVCache グラフ ノード、エキスパート ホット アップデート、および DaoQL-Agent ランタイムは今後の作業として残ります。組み込みの同一マシン設定では、DaoQL はグラフ BFS を 1.20 ミリ秒、HNSW を 83.1 マイクロ秒、Fluent ハイブリッド クエリを 105.8 マイクロ秒でレポートします。これらの結果はエンジニアリングの可能性を示していますが、クライアント/サーバー システムとの展開形状の違いを考慮して解釈する必要があります。 LDBC SNB SF1 および ANN ベンチマークの探索的測定では、インタラクティブ クラスのクエリのほとんどがサブミリ秒からミリ秒の範囲で 34/34 のクエリ カバレッジを示していますが、ロングテール BI/IC クエリにより全体では 1.8 QPS にすぎません。 ANN ベンチマークは、ブリッジエッジ保護の修正後、1,000 レベルの QPS で Recall@10 >= 99% に達しました。 5 ドメインの反事実実験 (n = 1250) では、DaoQL+GPT-4o は 94% の構成可能な反事実分解可能性を達成し、GPT-4o 単独より 49 パーセントポイント上回りました。この論文では、証明可能な構造、予備的な経験的証拠、およびアーキテクチャ上のロードマップの主張を明確に分離しています。
原文 (English)
An Explicit World Model Based on Data-First Ontology: DaoQL Multimodal Storage Validation and Counterfactual Reasoning Evaluation
Large language models encode world models implicitly in neural weights, which exposes four structural risks in high-precision domains such as medicine and finance: hallucination, frozen knowledge, poor explainability, and poor modifiability. This paper proposes data-first ontology: LLMs are treated as reasoning and language engines, while deterministic knowledge is moved into an explicit multimodal database, DaoQL. We formalize an explicit world model and show that, under rule independence, deterministic evaluation, and fixed conflict resolution, explicit models provide a sufficient condition for composable counterfactual decomposability; implicit models lack atomic read/delta semantics and therefore provide no comparable architectural guarantee. The implemented system focuses on DaoQL's verified storage layer and explicit Eval path, integrating graph, column, vector, and full-text engines within one process. KVCache graph nodes, expert hot updates, and the DaoQL-Agent runtime remain future work. On an embedded same-machine setup, DaoQL reports graph BFS at 1.20 ms, HNSW at 83.1 us, and a Fluent hybrid query at 105.8 us; these results indicate engineering potential but must be interpreted with deployment-shape differences from client-server systems. Exploratory measurements on LDBC SNB SF1 and ANN-Benchmarks further show 34/34 query coverage with interactive-class queries mostly in the sub-millisecond to millisecond range, but only 1.8 QPS overall due to long-tail BI/IC queries; ANN-Benchmarks reaches Recall@10 >= 99% at thousand-level QPS after a bridge-edge protection fix. In a five-domain counterfactual experiment (n = 1250), DaoQL+GPT-4o achieves 94% composable counterfactual decomposability, 49 percentage points above GPT-4o alone. The paper explicitly separates provable structure, preliminary empirical evidence, and architectural roadmap claims.
ロスレスだが無料ではない: 消費者向けハードウェアにおける投機的デコーディングの経験的解剖学
大規模な言語モデルのシングルストリーム自己回帰デコードはメモリ帯域幅によって制限されます。生成された各トークンには、ターゲット モデルを通過する 1 つの完全な前方パスが必要であり、連続するパスは並列化できません。投機的デコードはこの計算を再構築します。小規模なドラフト モデルは自己回帰的に $K$ トークンを提案し、ターゲット モデルは 1 回のバッチ パスですべてのトークンをスコアリングし、拒否サンプリング ルールはターゲット モデルの出力分布を確実に保存します。私たちは、消費者向け Apple シリコン ラップトップ上での、デバイスに依存しない (CUDA/MPS/CPU) のゼロからの実装と、5 つのドラフト/ターゲット バックエンド構成にわたる実証研究を紹介します。分布の等価性は 3 つのレベルで検証され、メソッドごとに約 9,200 個の実際のモデル トークン ($\chi^2 = 162.5$、dof $= 200$、$p = 0.976$) と正確な貪欲シーケンスの一致に対する 2 サンプル テストで最高潮に達します。最良の構成では、$K=6$ で実測 $1.61\times$ の実測速度に達し、許容プロファイルでは $K=1$ の 69.7% から最適時の 37.8% に低下しますが、5 つの構成のうち 3 つでは減速します。これは、ドラフトが小さなターゲットの速度を上回ることができないため、または量子化された Metal バックエンドが「並列」検証をシリアルに実行するためであり、その影響を分離して定量化しています。失敗も成功と同じくらい有益です。投機的なデコードは、検証が真にバッチ並列であり、ドラフトとターゲットのレイテンシのギャップが実際にある場合にのみ効果を発揮します。
原文 (English)
Lossless but Not Free: An Empirical Anatomy of Speculative Decoding on Consumer Hardware
Single-stream autoregressive decoding of large language models is bound by memory bandwidth: each generated token requires one full forward pass through the target model, and successive passes cannot be parallelized. Speculative decoding restructures this computation: a small draft model proposes $K$ tokens autoregressively, the target model scores all of them in one batched pass, and a rejection-sampling rule provably preserves the target model's output distribution. We present a from-scratch, device-agnostic (CUDA/MPS/CPU) implementation and an empirical study across five draft/target backend configurations on a consumer Apple-silicon laptop. Distribution equivalence is verified at three levels, culminating in a two-sample test over roughly 9,200 real-model tokens per method ($\chi^2 = 162.5$, dof $= 200$, $p = 0.976$) and exact greedy-sequence agreement. The best configuration reaches a measured $1.61\times$ wall-clock speedup at $K=6$, on an acceptance profile declining from 69.7% at $K=1$ to 37.8% at the optimum, while three of five configurations decelerate, either because the draft fails to out-speed a small target or because the quantized Metal backend executes "parallel" verification serially, an effect we isolate and quantify. The failures are as instructive as the successes: speculative decoding pays off only when verification is genuinely batch-parallel and the draft/target latency gap is real.
学習主導型の適応型監査スケジューリング: オフチェーンのデータ整合性に対する逐次決定アプローチ
オフチェーン データの暗号監査を、部分可観測性の下で制約付き MDP (CMDP) としてモデル化します。ストレージ ノードの隠蔽タイプと破損状態により問題が POMDP になり、ミス率上限ローにより明示的なセキュリティ制約が課されます。我々は、GRU 層が潜在ノード タイプに対する信念を維持するディープリカレント Q ネットワークである DRQN-CMDP を、ミスレート ペナルティ ラムダを自動的に適応させるラグランジュ デュアル アセントと組み合わせて提案します。ペアリングフリーの準同型 MAC プリミティブでは、O(1) のオンチェーン検証コストがかかります。 4 つの DQN バリアント、PPO、A2C、PPO-Lagrangian、ステートフル ベイジアン ヒューリスティック、3 つの固定ルール ベースライン、およびオラクル インフォームド ヒューリスティックの 13 のメソッドにわたって、DRQN-CMDP は好ましいバランスを実現します。固定高頻度監査よりも 83% 低いガス、1 桁のミス率 (7.5%)、および中程度の検出レイテンシ。これは他のメソッドにはない組み合わせです。 3 つの目標すべてを同時にマッチングします。
原文 (English)
Learning-Driven Adaptive Audit Scheduling: A Sequential Decision Approach to Off-Chain Data Integrity
We model cryptographic auditing of off-chain data as a Constrained MDP (CMDP) under partial observability: the storage node's hidden type and corruption state make the problem a POMDP, while a miss-rate ceiling rho imposes an explicit security constraint. We propose DRQN-CMDP, a Deep Recurrent Q-Network whose GRU layer maintains a belief over the latent node type, paired with Lagrangian dual ascent that adapts the miss-rate penalty lambda automatically. A pairing-free homomorphic-MAC primitive supplies O(1) on-chain verification cost. Across 13 methods--four DQN variants, PPO, A2C, PPO-Lagrangian, a stateful Bayesian heuristic, three fixed-rule baselines, and an oracle-informed heuristic--DRQN-CMDP achieves a favourable balance: 83% lower gas than fixed high-frequency auditing, single-digit miss rate (7.5%), and moderate detection latency--a combination no other method matches across all three objectives simultaneously.
Agentic ERP: 自律型エンタープライズ リソース プランニングのためのマルチエージェント大規模言語モデル アーキテクチャ
エンタープライズ リソース プランニング (ERP) システムは、トランザクションを確実に記録しますが、依然としてほとんどすべての運用上の意思決定を人間の専門家に委任しています。これは、従来のルールベースの自動化では例外を判断することができず、モノリシック AI アシスタントは、機能の境界を越えて調整するように求められると性能が低下するためです。このホワイトペーパーでは、ロール調整されたラージ言語モデル (LLM) エージェントと、リスク階層化された人間参加型ハーネスおよびグラフベースのオーケストレーターを組み合わせて、実稼働 ERP バックエンドでエンドツーエンドのビジネス ワークフローを実行するエキスパート システム アーキテクチャである Agentic ERP について説明します。まず、自律的な ERP 運用は、構造化された企業状態に対する制約付き逐次決定問題として定式化され、役割を調整したエージェントをステップごとのツール選択の複雑さの測定可能な削減に結び付ける分解引数が使用されます。第 2 に、グラフベースのプランナー - 実行者 - リフレクター - レスポンダーのオーケストレーションは、外部化された評価基準とスプリント契約を通じて生成を評価から切り離し、最近のハーネス エンジニアリングの原則を検査可能なエキスパート システムのアーティファクトとしてパッケージ化します。 3 番目に、システムは 3 つのレベルで評価されます。シナリオ ベースのタスク スイート、部門横断的な危機タスクに関する 6 つのオーケストレーション パラダイムの包括的な比較、ルール ベースの RPA および無介入ベースラインに対する 365 日間のエージェントインザループ シミュレーションです。これらのレベル全体で、提案されたマルチエージェント手法はベースラインよりも大幅に優れており、ルールベースのベースラインでは同じ需要ストリームの下で在庫が数百個蓄積される一方で、システムは在庫切れゼロでシミュレーションされた年間稼働を維持します。この研究は、人間の監視下で役割を調整された LLM エージェントが ERP システムを受動的にトランザクションを記録することから運用決定を能動的に実行することに移行できることを示しており、自律的なエンタープライズ リソース プランニングのためのリファレンス アーキテクチャと評価プロトコルを提供します。
原文 (English)
Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning
Enterprise Resource Planning (ERP) systems record transactions reliably but still delegate almost all operational decision-making to human specialists, because classical rule-based automation cannot reason about exceptions and monolithic AI assistants degrade when asked to coordinate across functional boundaries. This paper presents Agentic ERP, an expert-system architecture that combines role-aligned large-language-model (LLM) agents with a risk-tiered human-in-the-loop harness and a graph-based orchestrator to execute end-to-end business workflows on a production ERP backend. First, autonomous ERP operation is formulated as a constrained sequential-decision problem over a structured enterprise state, with a decomposition argument linking role-aligned agents to a measurable reduction in per-step tool-selection complexity. Second, a graph-based Planner--Executor--Reflector--Responder orchestration decouples generation from evaluation through externalised grading criteria and sprint contracts, packaging recent harness-engineering principles as inspectable expert-system artefacts. Third, the system is evaluated at three levels: a scenario-based task suite, a comprehensive comparison of six orchestration paradigms on cross-functional crisis tasks, and a 365-day agent-in-the-loop simulation against rule-based RPA and no-intervention baselines. Across these levels the proposed multi-agent method is significantly better than the baseline, and the system sustains a simulated year of operation with zero stockouts while the rule-based baseline accumulates hundreds under the same demand stream. The work shows that role-aligned LLM agents under human oversight can move an ERP system from passively recording transactions to actively executing operational decisions, and it provides a reference architecture and an evaluation protocol for autonomous enterprise resource planning.
DeeperRadar: 自動運転車の認識のためのエンドツーエンドの MIMO レーダー設計とマルチモーダル融合
DeeperRadar は、レーダー中心のセンサー スタック条件付きフレームワークで、スパース取得パターンを融合モデルでエンドツーエンドで学習することで、自律移動のためのレーダー センシングとマルチモーダル 3D 検出を共同設計します。学習可能な MIMO 設計モジュールは、生のレーダー ADC データとカメラ画像および LiDAR 点群を直接操作するフュージョン ネットワーク内でエンドツーエンドでトレーニングされます。トレーニング中、設計モジュールは他のセンサーによって監視され、システムはどの受信アンテナをアクティブにするか、およびそれらの有効な数の両方を学習できるようになります。導入時に、設計モジュールが削除され、学習されたスパース サブサンプリング マスクに置き換えられ、ダウンストリーム モデル アーキテクチャは変更されません。 DeeperRadar は、RADIal データセットで評価され、より少ない受信機を使用しながら、フルアレイのベースラインと一致またはそれを超えるまばらなタスク認識レーダー構成を検出し、レーダーのコストと統合の複雑さを削減できる可能性があります。これらの結果は、学習された最適な MIMO レーダー設計が融合スタックと下流の認識タスクに依存することを示しています。
原文 (English)
DeeperRadar: End-to-End MIMO Radar Design and Multi-Modal Fusion for Autonomous Vehicle Perception
DeeperRadar is a radar-centric, sensor-stack-conditioned framework that co-designs radar sensing and multi-modal 3D detection for autonomous mobility by learning a sparse acquisition pattern end-to-end with the fusion model. A learnable MIMO design module is trained end-to-end within a fusion network that operates directly on raw radar ADC data together with camera images and LiDAR point clouds. During training, the design module is supervised by the other sensors, enabling the system to learn both which receiver antennas to activate and the effective number of them. At deployment, the design module is removed and replaced by the learned sparse subsampling mask, leaving the downstream model architecture unchanged. Evaluated on the RADIal dataset, DeeperRadar discovers sparse, task-aware radar configurations that match or exceed full-array baselines while using fewer receivers, potentially reducing radar cost and integration complexity. These results show that learned optimal MIMO radar design depends on the fusion stack and the downstream perception task.
検証者に基づいたベンチマーク共進化を備えた自己修正型リーンプルーフ エージェント
効果的なリーンプルーフ エージェントを設計することは、形式的な数学的推論における中心的な課題です。最近の研究では、より強力な証明者を構築するだけでなく、リーンを中心としたワークフロー、つまりエージェントがどのように証明義務を分解し、ツールとコンパイラのフィードバックを使用し、障害を診断し、証明を修復し、構造化された証明コンテキストを維持するかが強調されています。私たちは、コードレベルで自己進化するエージェントを動機として、そのようなワークフローを手作業で設計するのではなく進化させることができるかどうかを研究しています。私たちは、自己進化するリーン証明エージェントを紹介します。このエージェントでは、小さな固定された信頼できるランタイムが完全に変更可能なワークスペース (証明ワークフロー、プロンプト、ツール) をラップします。固定された外部ベンチマークに対して最適化するほとんどの自己進化システムとは異なり、私たちのシステムはエージェントとそのベンチマークを共進化させます。世代間では、最高スコアのエージェント (チャンピオン) が、現在のレベルをマスターした後にのみ、より厳しい証明義務を導入するマスタリー スロットル カリキュラムの更新を通じてアクティブ タスクの配分を修正します。また、シングル アンカーの再調整により、更新されたベンチマークでチャンピオンを再実行し、難易度が上昇しても同等のスコアを維持します。すべての進化はリーン根拠のある検証ループ内に留まります。ただし、エージェントは自身を書き換えますが、成功とはその動作が信頼できるスナップショットの下でリーン検証された証明を生成する場合にのみカウントされます。各試行は機械可読でリーン根拠のある証明コンテキストを出力する必要があり、その表現は進化する可能性がありますが、その根拠は強制されます。 15 のアクティブな世代について共進化の軌跡と固定ベンチマーク ベースラインを実行し、保持された miniF2F テスト分割でそれらを比較します。最良の共進化エージェントのホールドアウト解決率は 45.1% に達しましたが、シードでは 12.7%、最良の固定ベンチマーク エージェントでは 32.0% でした。これは、検証者に基づいた自己進化が、共進化ベンチマークの下でリーンプルーフ ワークフローを改善できることを示しています。
原文 (English)
Self-Modifying Lean Proof Agents with Verifier-Grounded Benchmark Coevolution
Designing effective Lean proof agents is a central challenge in formal mathematical reasoning. Beyond building stronger provers, recent work emphasizes the workflow around Lean: how an agent decomposes proof obligations, uses tools and compiler feedback, diagnoses failures, repairs proofs, and maintains structured proof context. Motivated by code-level self-evolving agents, we study whether such workflows can be evolved rather than hand-designed. We present a self-evolving Lean proof agent in which a small fixed, trusted runtime wraps a fully mutable workspace: the proof workflow, prompts, and tools. Unlike most self-evolving systems, which optimize against a fixed external benchmark, our system coevolves the agent and its benchmark. Between generations, the highest-scoring agent (the champion) revises the active task distribution through a mastery-throttled curriculum update that introduces harder proof obligations only after the current level is mastered, and a single-anchor recalibration re-runs the champion on the updated benchmark to keep scores comparable as difficulty rises. All evolution stays inside a Lean-grounded verification loop: however the agent rewrites itself, a success counts only when its behavior yields Lean-verified proofs under a trusted snapshot, and each attempt must emit a machine-readable, Lean-grounded proof context whose representation may evolve but whose groundedness is enforced. We run the coevolving trajectory and a fixed-benchmark baseline for 15 active generations and compare them on a held-out miniF2F test split. The best coevolving agent reaches a 45.1% held-out solve rate, versus 12.7% for the seed and 32.0% for the best fixed-benchmark agent, showing that verifier-grounded self-evolution can improve Lean proof workflows under a coevolving benchmark.
思考の多様性の定量化: 加重 LLM アンサンブル リフトの予測法則
この論文は、思考の多様性が大規模言語モデル (LLM) アンサンブルにもたらす向上を計算するための、実験的に検証された形式法則を提供します。第一原理に基づいて、LLM アンサンブル揚力を救助質量と損傷質量に正確に分解し、隆起を計算するためのコンパクトなヒューリスティックを生成します。ここから、アンサンブルのパフォーマンスを予測するメトリクス、つまり精度調整された正しさの相関 $\phi_{\mathrm{adj}}$ と、ペアの精度ギャップおよび集合精度を抽出します。私たちは、2 つの大学院レベルの科学ベンチマークにわたる 10 のオープンウェイト モデルからの 767,520 件の推論に関する法則をテストします。また、各モデルがネットワークで隔離されたサンドボックス内でマルチターン ツールを使用してデジタル フォレンジック調査を行う新しいエージェント サイバーセキュリティ ベンチマークもテストします (棄権を含む 23,520 件の段階的試験)。すべての投票は公開で公開されます。 SuperGPQA で 40:60 の投票分割で一度キャリブレーションされると、ヒューリスティックはスピアマンの $\rho=0.84$ を使用してキャリブレーション セットのリフトを予測し、その係数を凍結してキャリブレーションでは決して使用されなかった 2 つのデータセット (GPQA ダイヤモンドでは $\rho=0.51$、フォレンジック タスクでは $0.84$) に転送します。一方、測定されたスワップ マス トラックは次のようなリフトを実現しました。全体で $R^2\ge 0.96$。生の $\phi$ にはほとんど予測能力がありません (全体を通して $R^2\le 0.09$)。精度調整された $\phi_{\mathrm{adj}}$ は著しく優れており (SuperGPQA では $R^2=0.67$)、これらのメトリクスを組み合わせたヒューリスティックは、3 つのデータセット全体で最も安定した事前プーリング予測子です。
原文 (English)
Quantifying Diversity of Thought: A Predictive Law of Weighted LLM Ensemble Lift
This paper provides an experimentally verified formal law for calculating the uplift that diversity of thought provides in Large Language Model (LLM) ensembles. From first principles, we derive an exact decomposition of LLM ensemble lift into rescue and damage masses, which yields a compact heuristic for calculating uplift. From this we extract the metrics which predict ensemble performance: an accuracy adjusted correctness correlation, $\phi_{\mathrm{adj}}$, together with the accuracy gap and collective accuracy of the pair. We test the law on 767,520 inferences from ten open-weight models over two graduate-level science benchmarks, together with a novel agentic cybersecurity benchmark in which each model conducts digital-forensics investigations by multi-turn tool use in a network-isolated sandbox (23,520 graded trials including abstentions); all votes are released openly. Calibrated once on SuperGPQA at a 40:60 vote split, the heuristic predicts lift on the calibration set with Spearman's $\rho=0.84$ and, with its coefficients frozen, transfers to two datasets never used in calibration ($\rho=0.51$ on GPQA Diamond and $0.84$ on the forensic tasks), whilst the measured swap mass tracks realised lift with $R^2\ge 0.96$ throughout. Raw $\phi$ has almost no predictive power ($R^2\le 0.09$ throughout); the accuracy-adjusted $\phi_{\mathrm{adj}}$ is markedly superior ($R^2=0.67$ on SuperGPQA), and the heuristic combining these metrics is the most stable pre pooling predictor across the three datasets.
間欠制御は希釈制御ではない:人工機関における切り替え効果
アダプティブ エージェントは、常に同じタイミング条件で調整するとは限りません。場合によっては、外乱が完全に内部状態に入る前に安定化が始まることがあります。また、エージェントは混乱が定着した後にのみ回復できる場合もあります。単純な予想としては、これらの条件の間を移動するエージェントは 2 つの固定ケースの加重平均のように動作するはずです。事後回復に費やす時間が長くなるほど、規制上の負担も大きくなります。この論文は、期待が裏切られる可能性があることを示しています。状態履歴が保持されたシミュレートされた適応エージェントでは、持続的な反応性制御の方が持続的な予期的制御よりもコストがかかる動作点で、予期的制御への断続的なアクセスにより、平均規制負担が固定モードの混合によって予測される値を下回ります。この効果は、定期的および確率的切り替えスケジュールの両方で現れます。予測的アクセスが失われると、その利点が薄れるだけではなく、断続的に回復すると、後の規制負担が再編成される可能性があります。統計量の高い実行 (スケジュールごとに N = 1000 の一致した複製) は、テストされたすべてのスケジュールにわたる負の非線形スイッチング ペナルティを解決します。効果は小さいですが一貫しています。平均増加の約 0.5 パーセントで、反復の 63 ~ 68% がゼロを下回ります。後期の診断では、未解決の上向きの規制負担の蓄積が明らかになっていない。その結果、設計に関連したタイミング原則が特定されます。履歴依存型適応システムでは、組織化を維持するための負担は、エージェントが各モードに費やす時間だけによって決まるわけではありません。混乱と回復が状態に入る順序によって、その後の負担が変わる可能性があります。したがって、断続的な予期的制御は、規制の部分的な失敗というよりも、回復に伴う長期的な負担を軽減するためのメカニズムのように機能する可能性があります。
原文 (English)
Intermittent Control Is Not Diluted Control: A Switching Effect in Artificial Agency
Adaptive agents do not always regulate under the same timing conditions. Sometimes stabilization can begin before a disturbance has fully entered the internal state; at other times the agent can only recover after disruption has taken hold. A simple expectation is that an agent moving between these conditions should behave like a weighted average of the two fixed cases: the more time spent in reactive recovery, the greater the regulatory burden. This paper shows that expectation can fail. In a simulated adaptive agent with retained state history, at an operating point where sustained reactive control is more costly than sustained anticipatory control, intermittent access to anticipatory control reduces the mean regulatory burden below the value predicted by a fixed-mode mixture. The effect appears under both periodic and stochastic switching schedules: losing anticipatory access does not simply dilute its benefit, and restoring it intermittently can reorganize the later regulatory burden. High-statistics runs (N = 1000 matched replicates per schedule) resolve a negative nonlinear switching penalty across every tested schedule. The effect is small but consistent: about half a percent of the mean gain, with 63-68% of replicates falling below zero. Late-window diagnostics reveal no unresolved upward accumulation of regulatory burden. The result identifies a design-relevant timing principle. In history-dependent adaptive systems, the burden of remaining organized is not set only by how much time an agent spends in each mode; the order in which disturbance and recovery enter the state can change the subsequent burden. Intermittent anticipatory control may therefore act less like a partial failure of regulation than like a mechanism for reducing the long-term burden of recovery.
経験的根拠により、混乱時の人間の行動をシミュレートする LLM エージェントの現実性が向上します
大規模言語モデル (LLM) エージェントは、災害やインフラストラクチャの混乱計画における一般的な課題である、直接的な歴史的類似点がほとんどまたはまったくない状況下での人間の行動をシミュレートするための生成的アプローチを提供します。ただし、この生成能力は妥当性の問題を引き起こします。個別にもっともらしいエージェント推論では、集団の経験的な行動を再現できない可能性があります。実験的根拠により、混乱時の LLM エージェント シミュレーションの統計的現実性が向上するかどうかを評価します。具体的には、米国コミュニティ調査の人口統計プロファイル、米国時間使用調査のベースライン ルーチン、都市の空間コンテキストをエージェントの初期化、記憶、意思決定プロンプト、およびアクティビティの実行に組み込む、経験に基づいた LLM エージェント フレームワークを開発します。 2024 年 7 月のフィラデルフィアの熱波中に実施された独立した世帯調査は、外部検証ベンチマークとして確保されています。根拠のない LLM エージェントのベースラインと比較して、根拠のあるモデルでは通常の日常生活の再構築が改善され、経験的活動プロファイルとの平均相関が 0.528 から 0.912 に増加し、平均二乗誤差が 0.066 から 0.008 に減少しました。熱波条件下では、地上モデルは調査から得られた活動プロファイルをよりよく再現し、平均相関が 0.349 から 0.836 に増加し、平均二乗誤差が 0.098 から 0.012 に減少しました。接地されたモデルは、観測された熱波応答振幅の 46.4% を捕捉しましたが、接地されていないベースラインの場合は 20.6% でした。これらの発見は、経験的根拠に基づいてLLMエージェントが集団行動の統計的により信頼できるシミュレーターを作成できる一方で、混乱時の人間の適応モデル化に残されたギャップを明らかにできることを示しています。
原文 (English)
Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions
Large language model (LLM) agents offer a generative approach to simulating human behavior under conditions that may have few or no direct historical analogues, a common challenge in disaster and infrastructure-disruption planning. However, this generative capacity creates a validity problem: individually plausible agent reasoning may fail to reproduce empirical population behavior. We evaluate whether empirical grounding improves the statistical realism of LLM-agent simulations during disruptions. Specifically, we develop an empirically grounded LLM-agent framework that embeds demographic profiles from the American Community Survey, baseline routines from the American Time Use Survey, and urban spatial context into agent initialization, memory, decision prompts, and activity execution. An independent household survey conducted during the July 2024 Philadelphia heatwave is reserved as an external validation benchmark. Compared with an ungrounded LLM-agent baseline, the grounded model improved reconstruction of normal daily routines, increasing mean correlation with empirical activity profiles from 0.528 to 0.912 and reducing mean squared error from 0.066 to 0.008. Under heatwave conditions, the grounded model better reproduced survey-derived activity profiles, increasing mean correlation from 0.349 to 0.836 and reducing mean squared error from 0.098 to 0.012. The grounded model captured 46.4% of observed heatwave response amplitude, compared with 20.6% for the ungrounded baseline. These findings show that empirical grounding can make LLM agents more statistically credible simulators of population behavior while revealing remaining gaps in modeling human adaptation during disruptions.
AEC-DS: 分散ストレージ向けの PDP トリガー レピュテーションおよび QoS を意識した移行を使用した適応型消去コーディング
分散ストレージ システムでは、監査結果が後の冗長性やシャード配置の決定に直接使用されないことが多く、非効率なリソース割り当てやリカバリの遅延につながる可能性があります。我々は、Provable Data Possession (PDP) フィードバックによって駆動される閉ループ適応消失符号化メカニズムである AEC-DS を提案します。 PDP 監査はノードの評判を継続的に更新し、QoS を意識した移行ポリシーはノードの信頼性とデータの優先順位に従ってシャードの配置を調整します。このポリシーは、優先度の高いシャードを不安定なノードからコールド層内のより信頼性の高いノードに移動し、その後の配置決定で不安定なノードにペナルティを与えます。 800 ノードと 500 ファイルを使用したシミュレーションでは、AEC-DS が 1.25 倍の冗長係数で評価された障害モデルの下で 100% のデータ耐久性を維持していることが示されています。 Static-EC、Dynamic-EC、DRD-EC と比較して、AEC-DS は累積リカバリ操作を 66.8% ~ 75.2% 削減します。アブレーションの結果はさらに、クラス移行がデータ損失の防止に主要な役割を果たしており、測定された損失防止能力が 176.8% 向上していることを示しています。これらの結果は、PDP フィードバックが整合性監査を冗長性および配置の適応と結びつけ、移行の追加コストを考慮しながら自己修復型分散ストレージへの実用的な道を提供できることを示しています。
原文 (English)
AEC-DS: Adaptive Erasure Coding with PDP-Triggered Reputation and QoS-Aware Migration for Decentralized Storage
In decentralized storage systems, audit results are often not used directly to guide later redundancy and shard-placement decisions, which can lead to inefficient resource allocation and delayed recovery. We propose AEC-DS, a closed-loop adaptive erasure coding mechanism driven by Provable Data Possession (PDP) feedback. PDP audits continuously update node reputation, while a QoS-aware migration policy adjusts shard placement according to node reliability and data priority. The policy moves high-priority shards from unstable nodes to more reliable nodes in the cold tier and penalizes unstable nodes in subsequent placement decisions. Simulations with 800 nodes and 500 files show that AEC-DS maintains 100% data durability under the evaluated fault model with a redundancy factor of 1.25x. Compared with Static-EC, Dynamic-EC, and DRD-EC, AEC-DS reduces cumulative recovery operations by 66.8%-75.2%. Ablation results further show that class migration plays a major role in preventing data loss, improving the measured loss-prevention capability by 176.8%. These results indicate that PDP feedback can connect integrity auditing with redundancy and placement adaptation, providing a practical path toward self-healing decentralized storage while accounting for the additional cost of migration.
パナシェ: あらゆるウィンドウ長でのワンパス モチーフ検出
モチーフの発見、つまり時系列内で繰り返されるパターンの検索は、探索的データ分析の中核となるプリミティブです。ただし、パターンはその期間によって定義され、アナリストがそれを事前に知ることはほとんどありません。この不明な継続時間を解決するには、ウィンドウの長さの間隔を定義し、その間隔内のすべての長さを試す方法が受け入れられています。既存のパン行列プロファイル (PMP) メソッドは、長さごとに 1 つの z 正規化行列プロファイルを計算するため、$L$ 長さには同じ系列にわたる $L$ 二次自己結合がかかります。私たちの知る限りでは、Z 正規化 PMP モチーフ発見のための最初のワンパス ストリーミング アルゴリズムである Panache を紹介します。これは、繰り返される自己結合を、実行時間がシリーズの長さに対してほぼ線形である単一のスキャンに置き換えます。重要な観察は、サブシーケンスの平均中心化によって DC フーリエ係数のみが変化するため、すべての Z 正規化サブシーケンスの非 DC スペクトルは、スライディング DFT 再帰と実行統計によってオンラインで維持できることです。このスペクトル状態は、占有制御されたハッシュ ディレクトリ内で類似のサブシーケンスが衝突するキーとなり、パーセヴァルの定理により、正確な計算の前にほとんどの衝突ペアを拒否する下限が得られます。 Panache はすべてのデータ依存パラメーターをそれ自体で計算し、調整するのはリソース バジェットのみです。デフォルトの予算では、17 個の UCR 構成で正確な固定除外グラウンド トゥルースに対して上位 20 位のすべてのパンモチーフを復元し、このペーパーでベンチマークしたすべての CPU および GPU ベースラインよりも高速です。 51 の長さにわたる 500 万サンプルのウェハーでは、Panache は 2.9 分で 1 つのパスを完了し、正確なモチーフを 6.0 分で放出します。これに対し、最速の正確な CPU ベースラインでは 7.95 時間、H100 GPU 上の SCAMP では 38.3 分かかります。
原文 (English)
Panache: One-Pass Motif Discovery at Every Window Length
Motif discovery, the search for recurring patterns within a time series, is a core primitive of exploratory data analysis. A pattern, however, is defined by its duration, which analysts rarely know in advance. To resolve this unknown duration, an interval of window lengths is defined, and the accepted method is to try every length in that interval. Existing pan matrix profile (PMP) methods compute one z-normalized matrix profile per length, so $L$ lengths cost $L$ quadratic self-joins over the same series. We introduce Panache, to our knowledge the first one-pass streaming algorithm for z-normalized PMP motif discovery. It replaces the repeated self-joins with a single scan whose runtime is near-linear in the series length. The key observation is that mean-centering a subsequence changes only its DC Fourier coefficient, so the non-DC spectrum of every z-normalized subsequence can be maintained online by sliding-DFT recurrences and running statistics. This spectral state is the key under which similar subsequences collide in an occupancy-controlled hash directory and, through Parseval's theorem, yields a lower bound that rejects most colliding pairs before any exact computation. Panache computes every data-dependent parameter itself, leaving only a resource budget to tune. At the default budget, it recovers all top-20 pan-motifs against exact fixed-exclusion ground truth on 17 UCR configurations, and is faster than every CPU and GPU baseline benchmarked in this paper. On Wafer at five million samples over 51 lengths, Panache completes one pass in 2.9 minutes and emits the exact motifs in 6.0 minutes, against 7.95 hours for the fastest exact CPU baseline and 38.3 minutes for SCAMP on an H100 GPU.
Pailitao-MMSearch: ネイティブ E コマース マルチモーダル検索基盤の構築
電子商取引の進化により、ユーザーの商品検索方法が根本的に変わり、単純なテキストベースのキーワードクエリから、商品画像、自然言語による説明、混合インテントの指示をシームレスに組み合わせる複雑なマルチモーダルインタラクションへと移行しました。しかし、既存のアプローチは重大なジレンマに直面しています。テキスト検索、視覚的検索、および音声認識用に個別に導入されたシングルモーダルの専門モデルは、単独で動作し、クロスモーダル クエリを処理できません。一方、汎用のビジョン言語モデルには、きめ細かい製品理解、ユーザー行動モデリング、および商業的意図の推論に必要なドメイン固有の知識が不足しています。この研究では、このギャップを埋めるために設計されたネイティブ電子商取引マルチモーダル検索基盤モデルの 1 つである Pailitao-MMSearch を紹介します。私たちのアプローチでは、次の 3 つの主要な革新が導入されています: (1) HybSID (ハイブリッド セマンティック ID);(2) 2 段階の継続的な事前トレーニング戦略。 (3) トレーニング後のハイブリッド推論パイプライン。 Qwen を基にして構築され、Taobao の Pailitao マルチモーダル検索プラットフォームに展開された Pailitao-MMSearch は、従来のマルチモーダル検索パイプラインと比較して、総商品量 (GMV) で最大 +13.61\%、トランザクション量で +8.21\% を含む、オンライン A/B テストの大幅な改善を達成し、ネイティブの e-コマース マルチモーダル検索大規模言語モデルの有効性を実証しています。
原文 (English)
Pailitao-MMSearch: Building Native E-Commerce Multimodal Search Foundation
The evolution of e-commerce has fundamentally transformed how users search for products, shifting from simple text-based keyword queries to complex multimodal interactions that seamlessly combine product images, natural language descriptions, and mixed-intent instructions. However, existing approaches face a critical dilemma: single-modal specialist models, deployed independently for text retrieval, visual search, and voice recognition, operate in isolation and cannot handle cross-modal queries, while general-purpose vision-language models lack the domain-specific knowledge necessary for fine-grained product understanding, user behavior modeling, and commercial intent reasoning. In this work, we present Pailitao-MMSearch, one native e-commerce multimodal search foundation model designed to bridge this gap. Our approach introduces three key innovations: (1)HybSID (Hybrid Semantic ID);(2)a two-stage continual pre-training strategy; and (3)a hybrid reasoning post-training pipeline. Built upon Qwen and deployed on Taobao's Pailitao multimodal search platform, Pailitao-MMSearch achieves substantial improvements in online A/B testing, including up to +13.61\% in Gross Merchandise Volume (GMV) and +8.21\% in transaction volume compared to traditional multi-modal search pipeline, demonstrating the effectiveness of our native e-commerce multimodal search large language models.
AI エージェントは本当に RTL から GDS への変換を完了できるでしょうか?ベンチマーク ツールからの教訓 - インタラクティブ EDA ワークフロー
LLM 駆動のエージェント システムは、電子設計自動化 (EDA) の有望なパラダイムとして浮上しており、複雑な設計ワークフローを自動化する強力な可能性を示しています。ただし、既存の評価は主に、分離された EDA タスクに関する個々の言語モデルを調査しており、完全な EDA フロー全体でさまざまなエージェント システムがどのように実行されるかについての洞察は限られています。この研究では、統一されたプロンプト、ツール環境、テクノロジー ライブラリ設定の下でのエンドツーエンド EDA ワークフローにおける AI エージェントの体系的な評価である FluxBench を紹介します。当社の評価では、オープンソース ツールチェーンを使用した RTL 生成や、産業アプリケーション向けのクローズドソースの商用 EDA ツールを使用した RTL から GDS へのフローなど、代表的なシナリオをカバーしています。これらのワークフローを通じて、RTL コード生成、反復修復、ツール フィードバックの利用、論理合成、配置配線 (P&R)、およびエンジニアリング変更オーダー (ECO) 自動化におけるエージェントの能力を評価します。エージェント システムの効率をさらに特徴付けるために、トークンの使用量とランタイム コストと比較した EDA アーティファクトの効果的な改善を測定するコスト効率の指標であるトークン ROI を導入します。実験結果によると、同じ基盤モデルに基づいて構築されている場合でも、エージェント システム アーキテクチャが異なると、最大 86.27% のパフォーマンス ギャップが見られる可能性があります。さらに、同等のタスク パフォーマンスを持つシステム間では、トークン ROI が $105.92\times$ も異なる可能性があります。 PicoRV32 をケーススタディとして使用した RTL から GDS へのフローでは、FluxEDA は最大 97.94 のエンドツーエンド スコアを達成し、ドメイン固有の EDA スキルを備えた Claude Code を最大 $8.39\times$ 上回りました。これらの結果は、大規模な EDA シナリオでエージェントのパフォーマンスを向上させるには、ドメイン固有のスキルだけでは不十分であることを示しています。代わりに、エージェント システム設計と基盤モデル機能の両方が、効果的な自動 EDA ワークフローを実現する上で重要な役割を果たします。
原文 (English)
Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows
LLM-driven agent systems have emerged as a promising paradigm for electronic design automation (EDA), demonstrating strong potential for automating complex design workflows. However, existing evaluations primarily examine individual language models on isolated EDA tasks, providing limited insight into how different agent systems perform across complete EDA flows. In this work, we present FluxBench, a systematic evaluation of AI agents on end-to-end EDA workflows under unified prompts, tool environments, and technology library settings. Our evaluation covers representative scenarios, including RTL generation with open-source toolchains and an RTL-to-GDS flow using closed-source commercial EDA tools for industrial applications. Through these workflows, we assess agents' capabilities in RTL code generation, iterative repair, tool-feedback utilization, logic synthesis, placement and routing (P&R), and Engineering Change Order (ECO) automation. To further characterize the efficiency of agent systems, we introduce Token ROI, a cost-efficiency metric that measures effective improvements in EDA artifacts relative to token usage and runtime cost. Experimental results show that, even when built on the same foundation model, different agent system architectures can exhibit performance gaps of up to 86.27%. Moreover, among systems with comparable task performance, Token ROI can differ by as much as $105.92\times$. In the RTL-to-GDS flow using PicoRV32 as a case study, FluxEDA achieves an end-to-end score of up to 97.94, outperforming Claude Code equipped with domain-specific EDA skills by up to $8.39\times$. These results indicate that domain-specific skills alone are insufficient to improve agent performance in large-scale EDA scenarios. Instead, both agent system design and foundation model capability play critical roles in enabling effective automated EDA workflows.
曲率の影: 最大エントロピー平衡選択の明らかな失敗は除去可能なアーティファクトである
ナッシュ均衡が凸集合を形成する 2 人プレイのゼロサム ゲームでは、正則化ナッシュ ダイナミクス (R-NaD) などの正則化ソルバーは、最大エントロピー メンバー、つまりナッシュ集合への一様参照の情報射影 (I 射影) を経験的に選択します。小規模ゲームのパネルでは、この一致は 1 つの明らかな例外を除いて正確です。クーン ポーカーでは、R-NaD が最大エントロピーの 99.7 パーセントに達しているにもかかわらず、最大エントロピー メンバーが 0.201 に座っているのに対し、R-NaD はブラフ座標 0.180 に着地しており、座標のギャップは約 0.021 です。このギャップが真の選択バイアスなのか、それとも人為的なものなのかを問い、定量的に答えます。 1 次元ナッシュ多様体での選択の場合、座標ギャップは $\mathrm{gap} \estimate \sqrt{2\delta/\kappa}$ として因数分解されることを示します。ここで $\delta$ はソルバーのエントロピー不足、$\kappa$ はピークにおけるエントロピー ランドスケープの曲率です。 5 つのゲームにわたって、この関係は $2 \times 10^{-4}$ 以内に収まります (相対誤差は 1% 未満)。 4 つの行列ゲームには $\delta \およそ 0$ (R-NaD は最大エントロピー要素に正確に到達します) があり、したがって曲率に関係なくギャップはありません。シーケンシャル ゲーム (Kuhn) のみが $\delta > 0$ になります。磁石の強さの因果関係のスイープにより、ダイナミクスが安定下限で不安定になるまで、予測曲線 (フィッティング スケーリング指数 0.50、 $R^2 > 0.999999$、1/2 の正確な予測に対して) に沿って $\delta \0$ とギャップがゼロに向かって駆動されます。この動作は、除去可能な不足と一致しますが、固定バイアスとは矛盾します。測定された曲率から法則の半分の曲率を定量化し、自然のツァリスエントロピー実験における移動ターゲットの落とし穴にフラグを立てます。したがって、クーン ギャップは、異常に平坦なピーク上の小さな除去可能なエントロピー不足の曲率の影です。 I 投影アカウントは、平坦性が制限された残差まで維持されます。
原文 (English)
The Curvature Shadow: An Apparent Failure of Maximum-Entropy Equilibrium Selection is a Removable Artifact
In two-player zero-sum games whose Nash equilibria form a convex set, regularized solvers such as Regularized Nash Dynamics (R-NaD) empirically select the maximum-entropy member: the information projection (I-projection) of a uniform reference onto the Nash set. On a panel of small games this match is exact, with one apparent exception: in Kuhn poker R-NaD lands at bluff coordinate 0.180 while the maximum-entropy member sits at 0.201, a coordinate gap of about 0.021, even though R-NaD attains 99.7 percent of the maximum entropy. We ask whether this gap is a genuine selection bias or an artifact, and answer it quantitatively. We show that for selection on a one-dimensional Nash manifold the coordinate gap factorizes as $\mathrm{gap} \approx \sqrt{2\delta/\kappa}$, where $\delta$ is the entropy shortfall of the solver and $\kappa$ is the curvature of the entropy landscape at its peak. Across five games this relation holds to within $2 \times 10^{-4}$ (under 1 percent relative error). The four matrix games have $\delta \approx 0$ (R-NaD reaches the maximum-entropy member exactly) and therefore no gap regardless of curvature; only the sequential game (Kuhn) has $\delta > 0$. A causal sweep of the magnet strength drives $\delta \to 0$ and the gap toward zero along the predicted curve (fitted scaling exponent 0.50, $R^2 > 0.999999$, against the exact prediction of 1/2), until the dynamics destabilize at a stability floor: behavior consistent with a removable shortfall and inconsistent with a fixed bias. We quantify the curvature half of the law from measured curvatures and flag a moving-target pitfall in the natural Tsallis-entropy experiment. The Kuhn gap is thus the curvature shadow of a small, removable entropy shortfall on an unusually flat peak; the I-projection account is upheld up to a flatness-limited residual.
維持するか統合するか?言語エージェントのメモリに対する予算に応じたオペレーターの選択
言語エージェントは、対話全体にわたる記憶に依存します。ただし、大規模言語モデル (LLM) の限られたコンテキスト ウィンドウとその推論コストにより、一度に使用できるメモリの量が制限されます。既存のシステムは主に、メモリ保持とメモリ統合という 2 つの戦略に従っています。保存では生の記録が保持され、正確な詳細が保存されますが、関連する証拠は限られた予算内では収まらない可能性があります。統合によりレコードが圧縮されて結合されるため、トークンごとのカバレッジが向上しますが、クエリに不可欠な詳細が失われる危険性があります。どちらの戦略も一般的に好ましいものではありません。これにより、2 つの中心的な疑問が生じます。1 つは、いつ保存の代わりに統合を行うべきか、もう 1 つはマージ、抽象、またはリライトのどの演算子を選択する必要があるかということです。この決定は、各演算子の効用を、保持によって省略された証拠に対する適用効果と、すでに適合している生の証拠に対する署名付き置換効果に分解することによって形式化します。これらのバランスにより、予算の相対的な圧力によって優先されるアクションが変化する理由が説明されます。私たちは、このメカニズムを Offline Abstraction-Safety (OAS) で実装します。OAS は、ホールドアウトされた危害キャリブレーションを使用して生成前の特徴からアクション ユーティリティを推定する軽量の学習器です。公開されている LongMemEval ベンチマークと LoCoMo ベンチマークは、同じ予算依存のパターンを示しています。 LongMemEval では、予算が厳しい場合は統合により絶対精度が最大 48% 向上しますが、予算が緩い場合は保持することが望ましいです。 LoCoMo は、その短い証拠と一致して、より少ない予算でこのクロスオーバーを再現します。どちらのデータセットでも、圧縮が必要な場合、クロスノート抽象化とマージは一般にローカルな書き換えよりも優れたパフォーマンスを発揮します。
原文 (English)
Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory
Language agents depend on memory across interactions. However, the limited context windows of large language models (LLMs) and their inference costs constrain how much memory can be used at once. Existing systems mainly follow two strategies: memory retention and memory consolidation. Retention keeps raw records and preserves exact details, but relevant evidence may not fit under a tight budget; consolidation compresses and combines records, improving coverage per token but risking the loss of query-critical details. Neither strategy is universally preferable. This raises two central questions: when should consolidation replace retention, and which operator -- Merge, Abstract, or Rewrite -- should be selected? We formalize this decision by decomposing each operator's utility into a coverage effect on evidence omitted by retention and a signed replacement effect on raw evidence that already fits. Their balance explains why the preferred action changes with relative budget pressure. We implement this mechanism with Offline Abstraction-Safety (OAS), a lightweight learner that estimates action utilities from pre-generation features with held-out harm calibration. The public LongMemEval and LoCoMo benchmarks show the same budget-dependent pattern. On LongMemEval, consolidation improves absolute accuracy by up to 48% under tight budgets, whereas retention is preferable under loose budgets; LoCoMo replicates this crossover at a smaller budget, consistent with its shorter evidence. On both datasets, cross-note abstraction and merging generally outperform local rewriting when compression is necessary.
フィードバック拡張自己蒸留が検索インターリーブ検索エージェントを改善できないのはなぜですか?
オンポリシー自己蒸留 (OPSD) は、別個の教師モデルに依存せずに大規模な言語モデルをトレーニングするための有望なアプローチを提供します。ただし、複雑なエージェント タスクに対するその有効性はほとんど解明されていません。この作業では、成功したデモンストレーションを特権情報として活用するエージェント検索用の自己蒸留アルゴリズムであるフィードバック拡張自己蒸留 (FA-SD) をインスタンス化します。我々は、モデルが反復的な推論と検索の出力テンプレートに依存し、多様に見えるが入力された質問にほとんどとらわれず、KL ベースの自己蒸留信号が有益ではない軌跡を生成する可能性があることを確認しました。私たちはこの現象をデコード崩壊と呼びます。これは、既存の評価基準では見逃される可能性のある障害モードです。その根本的な原因を理解するために、私たちは、セルフティーチャーがより優れたパフォーマンスを達成しても、学習は一貫性のない監視信号のために本質的に不安定なままであることを示します。我々はさらに、この不一致をモデルの不一致とプロンプトの不一致に分解し、後者が監視信号の品質を大幅に低下させ、セルフティーチャー学習の有効性を制限する可能性があることを示します。この不一致を軽減するために、自己教師を安定させ、より一貫した監視信号を提供する指数移動平均 (EMA) 教師を導入します。 EMA 教師はウォームアップ段階を必要とし、その間パフォーマンスが一時的に低下する可能性がありますが、より安定した監視を提供することで最終的にモデルのパフォーマンスを向上させます。
原文 (English)
Why Does Feedback-Augmented Self-Distillation Fail to Improve Retrieval-Interleaved Search Agents?
On-policy self-distillation (OPSD) offers a promising approach for training large language models without relying on a separate teacher model. However, its effectiveness on complex agentic tasks remains largely unexplored. In this work, we instantiate Feedback-Augmented Self-Distillation (FA-SD), a self-distillation algorithm for agentic search that leverages successful demonstrations as privileged information. We identify that models can rely on recurring reasoning-and-search output templates, producing trajectories that appear diverse but are largely agnostic to the input question, making the KL-based self-distillation signal uninformative. We term this phenomenon decoding collapse, a failure mode that can be missed by existing evaluation metrics. To understand its underlying cause, we show that although the self-teacher achieves stronger performance, learning remains inherently unstable due to inconsistent supervision signals. We further decompose this inconsistency into model inconsistency and prompt inconsistency, and show that the latter can significantly degrade the quality of the supervision signal, limiting the effectiveness of self-teacher learning. To mitigate this inconsistency, we introduce an exponential moving average (EMA) teacher to stabilize the self-teacher and provide more consistent supervision signals. Although the EMA teacher requires a warm-up phase during which performance may temporarily regress, it ultimately improves model performance by providing more stable supervision.
強化学習: アルゴリズムから基礎モデルまで
強化学習 (RL) は、明示的な目標の下で逐次的な意思決定を行うためのフレームワークを提供します。古典的な形式では、RL は動的な環境で長期的な報酬を最大化するためにエージェントがどのように行動すべきかを研究します。より豊かな設定では、問題は単一のエージェントや固定環境を超えて広がります。知的な行動には、戦略的な相互作用、不確実性への適応、高次元の世界に対する推論が必要になる場合があります。この論文では、ゲームにおけるアルゴリズムと基礎モデル時代の RL という 2 つの観点から RL を研究します。最初の部分では、ゲームにおけるマルチエージェント RL に焦点を当てます。 2 プレイヤーのゼロサム ゲーム、大規模なビデオ ゲーム、および一般的な構造を持つマルチプレイヤー設定に及ぶ、競争環境および総和環境でインセンティブ、ポリシー、均衡の概念がどのように相互作用するかを検証します。これらの研究では、マルチエージェント システムでの学習と対話型環境での RL メソッドの動作を調査しています。 2 番目の部分では、事前の知識が逐次的な意思決定を強化できるという考えに動機付けられ、生成モデルと基礎モデルを使用した RL を研究します。事前トレーニングされた生成モデルと学習された世界モデルは、計画、制御、およびポリシーの最適化のための表現ツールおよび構造化事前分布として機能します。この論文では、拡散ベースの世界モデルを開発し、効率的なビデオ生成のための RL を調査し、ポリシークラスとしての生成モデルを調査し、アクションが将来の観察を形作るインタラクティブなビデオ世界モデルを研究します。また、メモリを備えたアーキテクチャを通じて長期モデリングにも対応します。これらの貢献を総合すると、複雑な逐次領域における目的主導型の適応としての RL の統一された見解が提示されます。この論文では、戦略的ゲームから生成世界モデルまで、RL が意思決定、環境モデリング、新たな基盤モデル機能をどのように結びつけるかを強調し、インテリジェントな行動の基礎となる原理についてより広い視点を提供します。
原文 (English)
Reinforcement Learning: From Algorithms To Foundation Models
Reinforcement learning (RL) provides a framework for sequential decision making under explicit objectives. In its classical form, RL studies how an agent should act to maximise long-term reward in a dynamic environment. In richer settings, the problem extends beyond a single agent and fixed environment: intelligent behavior may require strategic interaction, adaptation to uncertainty, and reasoning over high-dimensional worlds. This thesis studies RL from two perspectives: algorithms in games and RL in the era of foundation models. The first part focuses on multi-agent RL in games. It examines how incentives, policies, and equilibrium concepts interact in competitive and general-sum environments, spanning two-player zero-sum games, large-scale video games, and multi-player settings with general structure. These works investigate learning in multi-agent systems and the behavior of RL methods in interactive environments. The second part studies RL with generative and foundation models, motivated by the idea that prior knowledge can enrich sequential decision making. Pretrained generative models and learned world models serve as representation tools and structured priors for planning, control, and policy optimization. The thesis develops diffusion-based world models, investigates RL for efficient video generation, explores generative models as policy classes, and studies interactive video world models in which actions shape future observations. It also addresses long-horizon modeling through architectures with memory. Together, these contributions present a unified view of RL as objective-driven adaptation in complex sequential domains. From strategic games to generative world models, the thesis highlights how RL connects decision making, environment modeling, and emerging foundation-model capabilities, offering a broader perspective on the principles underlying intelligent behavior.
ZifaMem: AI コンパニオンにおけるペルソナ、好み、感情的連続性の構造化記憶
AI コンパニオンは、シングルターンの流暢さだけでなく、感情的な連続性を維持しているかどうか、つまりコンパニオンが誰であるか、ユーザーが何を好むか、関係がどのように感じたかを覚えているかどうかによって判断されます。対話をセッションの要約、エピソード記憶、統合されたユーザー モデルに編成する構造化記憶システムである ZifaMem を紹介します。完全な生の対話履歴を提供するデプロイメントに正直なコンパレーターと比較し、ルート監査を備えた固定の LLM-as-a-judge プロトコルの下で、構造化記憶はプールされた 4 つのバックボーンの感情的知性スコアを 11.4% (95% CI 6.3% ~ 17.1%) 上昇させ、4 つのバックボーンすべてでペルソナのグラウンディングが向上しました (クロード +42% 相対)。マルチターンの影響コンテキストは、シングルターンのスナップショット (探索的) よりも +39% の純優先度を獲得しますが、追加の感情ステートマシンは 5 つのエンドポイントのいずれにおいても測定可能な利益をもたらしません。同一の事前登録プロトコルの下では、3 つのメモリ システム (ZifaMem、Mem0、およびフィルタリングされた逐語的取得) はそれぞれ、生の履歴の展開よりも大幅に向上しており、ZifaMem と Mem0 は、事前登録された主優先エンドポイントで +/-5 ポイント以内で統計的に同等です。 ZifaMem SDK、CLI、およびポータブル エージェント スキルは、https://github.com/zifacorp/zifamem でオープンソース化されています。
原文 (English)
ZifaMem: Structured Memory for Persona, Preference, and Emotional Continuity in AI Companions
AI companions are judged not only by single-turn fluency but by whether they sustain emotional continuity: remembering who the companion is, what the user prefers, and how the relationship has felt. We present ZifaMem, a structured memory system that organizes dialogue into session summaries, episodic memories, and a consolidated user model. Against a deployment-honest comparator that supplies the full raw dialogue history, and under a fixed LLM-as-a-judge protocol with route audits, structured memory raises pooled four-backbone emotional-intelligence scores by 11.4% (95% CI 6.3% to 17.1%), and persona grounding improves on all four backbones (Claude +42% relative). Multi-turn affect context wins a +39% net preference over a single-turn snapshot (exploratory), whereas an additional emotion state machine yields no measurable gain on any of five endpoints. Under an identical preregistered protocol, three memory systems (ZifaMem, Mem0, and filtered verbatim retrieval) each improve significantly over raw-history deployment, and ZifaMem and Mem0 are statistically equivalent within +/-5 points on the preregistered primary preference endpoint. The ZifaMem SDK, CLI, and portable Agent Skills are open-sourced at https://github.com/zifacorp/zifamem.
LLM ガードレールのための二重仮説推論フレームワーク
我々は、2 つの重要なアイデアを導入した新しい LLM ガードレール フレームワークである ARBITER を提案します。(i) 安全性の決定を下す前にプロンプトの安全な解釈と安全でない解釈の両方を明示的に考慮する LLM ガードレールの推論方法である二重仮説推論、および (ii) LLM 出力を論理コンポーネントに分解し、その重要性に応じて重み付けする推論ベースのガードレールの構造化トレーニング損失であるマルチコンポーネント教師あり微調整 (MC-SFT)。既存の推論ベースのガードレールは、大規模な教師モデルまたはクローズドソースの教師モデルを使用して推論トレースを生成し、フルパラメータの微調整を適用するなど、高価な手順に依存していることがよくあります。対照的に、ARBITER は、トレースの推論と LoRA ベースのパラメーター効率の良い微調整に費用対効果の高い自己生成戦略を使用しながら、これらの高価なアプローチよりも優れたパフォーマンスを実現します。さらに、ARBITER は、安全でない決定に対して忠実な証拠フレーズの説明を提供し、より透明性が高く解釈可能なガードレール方法を可能にします。 3 つの安全モデレーション ベンチマークの実験では、ARBITER が既存の推論ベースおよび非推論のガードレール ベースラインを上回り、ドメイン外の評価で明らかな向上を示していることが示されています。
原文 (English)
A Dual-Hypothesis Reasoning Framework for LLM Guardrails
We propose ARBITER, a novel LLM guardrail framework that introduces two key ideas: (i) dual-hypothesis reasoning, a reasoning method for LLM guardrails that explicitly considers both safe and unsafe interpretations of a prompt before making a safety decision, and (ii) multi-component supervised fine-tuning (MC-SFT), a structured training loss for reasoning-based guardrails that decomposes LLM outputs into logical components and weights them according to their importance. Existing reasoning-based guardrails often rely on expensive procedures, such as generating reasoning traces using larger or closed-source teacher models and applying full-parameter fine-tuning. In contrast, ARBITER uses a cost-effective self-generation strategy for reasoning traces and LoRA-based parameter-efficient fine-tuning while still achieving better performance than these expensive approaches. Additionally, ARBITER provides faithful evidence-phrase explanations for unsafe decisions, enabling a more transparent and interpretable guardrail method. Experiments on three safety moderation benchmarks show that ARBITER outperforms existing reasoning-based and non-reasoning guardrail baselines, with clear gains in out-of-domain evaluations.
ロングコンテキストのエージェントに必要なのは段階的な開示だけですか?
長い文書の質問応答では、通常、文書全体をコンテキスト ウィンドウにロードするか、別の取得ツールをボルトオンするかの選択を迫られます。 Agentic AI はより幅広いオプションを提案し、エージェントにドキュメントのパスを与え、何をどのように読むか決定させます。エージェント スキルは、エージェントがオンデマンドでロードするフォルダに専門知識をパッケージ化するための標準であり、短い説明から特定の文章まで、クエリに必要なものだけを公開する段階的な開示という、すぐに使えるメカニズムを提供します。実務家は、本の長さの理解タスクにこのパターンを急速に採用しましたが、そのような選択を裏付ける証拠は逸話的です。私たちはパターンの最初の対照研究を実行し、InfiniteBench 上の 3 つのエージェント ハーネスと 3 つのモデル ファミリにわたって、生のドキュメント ナビゲーションとエージェント スキル パックのいくつかの設計を古典的なハイブリッド レトリバーと比較しました。単一のブックでは、ゲインはハーネスに依存し、エージェントが生の文書をうまくナビゲートしない場合は大きくなりますが、強力なエージェント ハーネスがすでに分割して独自に取得している場合はゼロに近くなります。多くの書籍にまたがるタスクにスケールアップすると、生のドキュメントのナビゲーションが崩壊しますが、1 レベルの漸進的開示の低下がより遅くなり、前に進みます。 2 番目のより深い配線レベルは役に立たず、場合によっては精度が完全に損なわれるため、1 つのレベルで十分です。漸進的な開示は、インテリジェンスではなくコンテキストを購入します。強力なエージェントが適切な文章自体を見つけることができる間は冗長ですが、コーパスが大きくなりすぎて読んでナビゲートできない場合は決定的になります。
原文 (English)
Is Progressive Disclosure All You Need for Long-Context Agents?
Long-document question answering usually forces a choice between loading the whole document into the context window and bolting on a separate retriever. Agentic AI suggests a broader option, giving the agent the document path and letting it decide how and what to read. Agent Skills, a standard for packaging expertise into folders an agent loads on demand, supply a ready mechanism: progressive disclosure, which exposes only what a query needs, from a short description down to the specific passages. Practitioners rapidly adopted this pattern for book-length understanding tasks, but the evidence to support such choices has been anecdotal. We run the first controlled study of the pattern, comparing raw-document navigation and several designs of Agent Skills packs against a classical hybrid retriever across three agent harnesses and three model families on InfiniteBench. On a single book, the gain depends on the harness, running large when the agent navigates the raw document poorly but near zero when a strong agent harness already divides and retrieves on its own. When scaling up to tasks that span many books, raw-document navigation collapses while one-level progressive disclosure degrades more slowly and pulls ahead. A second, deeper routing level never helps and sometimes breaks accuracy outright, so one level is enough. Progressive disclosure buys context, not intelligence: it is redundant while a strong agent can locate the right passages itself, and decisive once the corpus grows too large to navigate by reading.
エージェントのメモリ改良のための機械的注意ガイダンス
既存の自己進化型記憶システムは、主にタスクの軌跡や振り返りなどのテキスト出力に基づいてエージェントの記憶を改善します。ただし、このテキストベースのパラダイムには内部メカニズム信号が組み込まれることはほとんどなく、取得されたメモリがタスク実行中に実際にどのように利用されるかは十分に解明されていません。この制限により、信頼性の低いエラーの帰属や幻覚による記憶の変更が発生する可能性があります。この研究では、検索ヘッド アテンションがセグメント レベルのメモリ使用率を明らかにするためのメカニズム信号を提供することを示します。メモリセグメントと意思決定ステップに対する注意を集約することで、繰り返し発生するメモリ使用パターンを明らかにし、対応する改善戦略を示すコンテキスト利用マトリックスを構築します。この観察に基づいて、私たちは、注意によって明らかにされる使用パターンを使用してターゲットを絞ったセグメントレベルのメモリ更新をガイドするフレームワークである、注意ガイド付きメモリ改良 (AGMR) を提案します。 AGMR は、失敗した実行のメモリを修正または強化し、成功した実行のメモリを簡素化し、再実行を通じて各更新を検証します。インタラクティブな意思決定ベンチマークの実験では、AGMR がテキストのみのメモリ改良ベースラインと比較して、タスクのパフォーマンスとメモリ効率の両方を向上させることが示されています。コードは https://anonymous.4open.science/r/AGMR_code-3262/ で入手できます。
原文 (English)
Mechanistic Attention Guidance for Agent Memory Refinement
Existing self-evolving memory systems mainly improve agent memory based on textual outputs, such as task trajectories and reflections. However, this text-based paradigm rarely incorporates internal mechanistic signals, leaving how retrieved memory is actually utilized during task execution underexplored. This limitation can lead to unreliable error attribution and hallucinated memory modifications. In this work, we show that retrieval-head attention provides a mechanistic signal for revealing segment-level memory utilization. By aggregating attention over memory segments and decision steps, we construct a context utilization matrix that exposes recurring memory-use patterns and indicates corresponding refinement strategies. Building on this observation, we propose Attention-Guided Memory Refinement (AGMR), a framework that uses utilization patterns revealed by attention to guide targeted segment-level memory updates. AGMR corrects or enhances memory for failed executions, simplifies memory for successful executions, and verifies each update through re-execution. Experiments on interactive decision-making benchmarks show that AGMR improves both task performance and memory efficiency over text-only memory refinement baselines. Code is available at https://anonymous.4open.science/r/AGMR_code-3262/
検証するか、修復するか、繰り返すか、それとも停止するか? LLM エージェントのノイズの多い検証修復ループに対する堅牢な停止
検証修復ループは、大規模言語モデル (LLM) エージェントがコード生成、数学的推論、およびツールの使用における欠陥のある計画を修正するための標準的な手段です。検証者と修復者の両方がうるさい場合、修復によってすでに修正された計画が損なわれる可能性があり、真の妥当性が低下する一方で、報告された合格率は増加し続けるため、既存の方法には修復をいつ停止すべきかを決定するための原則的な根拠が欠けています。我々は、ノイズの多い検証-修復-繰り返し(VRR)ループのための堅牢な停止フレームワークであるVRR-Stopを提案します。 4 つのパラメーターのノイズ モデルは、検証者の誤った受け入れと誤った拒否を、修復者の修理および損傷動作から分離します。信念フィルタリングは、繰り返された検証投票をコミットされた有効性の推定値に変換し、ループは真の限界ゲインの符号に従ってコミットまたは修復します。これには、すべてのパラメータの正確な回復ではなく、符号の識別可能性のみが必要です。検証者の識別がゼロに近づくと、キャリブレーション自体が失敗し、推定エラーが停止標識を反転させる可能性があるため、十分な検証マージンの下でのみ現職候補を置き換える推定不要のフォールバックである VRR-Stop と VRR-Guard を組み合わせます。 GSM8K ストレス設定では、VRR-Stop は、平均 0.72 修復ラウンドのコストで、固定の 5 ラウンド修復よりも 60.6 パーセント ポイント、最終的な真の有効性を向上させます。設定全体にわたって、停止の信頼性は、推定誤差の絶対的なサイズではなく、検証者の識別と決定マージンによって共同で支配されます。
原文 (English)
Verify, Repair, Repeat, or Stop? Robust Stopping for Noisy Verify-Repair Loops in LLM Agents
Verify-repair loops are a standard means for large language model (LLM) agents to correct faulty plans in code generation, mathematical reasoning, and tool use. When both the verifier and the repairer are noisy, repair can damage already-correct plans, and reported acceptance keeps rising while true validity falls, so existing methods lack a principled basis for deciding when repair should stop. We propose VRR-Stop, a robust stopping framework for noisy verify-repair-repeat (VRR) loops. A four-parameter noise model separates verifier false acceptance and false rejection from the repair and damage behavior of the repairer. Belief filtering turns repeated verification votes into an estimate of committed validity, and the loop commits or repairs according to the sign of the true marginal gain, which requires only sign identifiability rather than accurate recovery of all parameters. When verifier discrimination approaches zero, calibration itself fails and estimation error can flip the stopping sign, so we pair VRR-Stop with VRR-Guard, an estimation-free fallback that replaces the incumbent candidate only under a sufficient verification margin. On a GSM8K stress setting, VRR-Stop improves final true validity by 60.6 percentage points over fixed five-round repair at an average cost of 0.72 repair rounds. Across settings, stopping reliability is governed jointly by verifier discrimination and the decision margin rather than by the absolute size of estimation error.
FlowBlock: 自己修正拡散言語モデルのための波面並列デコーディング
ブロック単位の拡散大規模言語モデル (dLLM) は、ブロック レベルで順次デコードするため、ブロック全体で KV キャッシュを効果的に再利用できますが、ブロック間のデコードは厳密にシリアルになります。これまでの研究では、トレーニング後の方法を通じてブロック間の並列処理を解放しようとしましたが、わずかな速度向上しか達成できず、精度が低下することがよくありました。私たちは、自己修正 dLLM がトレーニング不要の代替手段を提供していることを観察しています。トークンツートークン (T2T) 編集は、やや古いアップストリーム コンテキストでドラフトされたトークンを修復できるため、ダウンストリーム ブロックには、最終的な先行ブロックではなく、有益なドラフトのみが必要です。これにより、ブロックのファイナリティがハードな依存関係からスケジューリング リソースに変わります。私たちは、2 つのメカニズムに基づいて構築されたトレーニング不要の並列デコード フレームワークである \textbf{\flowblock{}} を提案します。 (i) \emph{Gated Wavefront Decoding} は、レディネス ゲートが満たされた場合にのみブロックを有界ウェーブフロントに許可し、T2T 編集を通じてアクティブなブロックを共同でリファインし、正確なフリーズ プレフィックス KV キャッシュの再利用を保存するウィンドウ化されたブロック因果マスクの下でブロックを順番にコミットします。 (ii) \emph{異種波面パッキング} は、非同期ウィンドウを高密度で形状安定したバッチ転送にパッキングしながら、各リクエストに独立した波面を割り当てます。 \flowblock{} は、さまざまなベンチマークにわたって、2 つのシリアル ブロック単位 dLLM である LLaDA-2.1 および LLaDA-2.0 よりも 1 秒あたりのトークン (TPS) を最大 2.95$\times$ および 4.01$\times$ 改善し、レイテンシをそれぞれ最大 53.6\% と 77.1\% 削減します。また、平均精度も 1.3 ポイント向上します。トレーニングベースのブロック間並列ベースラインである D2F と比較して、\flowblock{} はより高い精度と最大 16$\times$ 高いバッチ処理スループットを実現します。
原文 (English)
FlowBlock: Wavefront-Parallel Decoding for Self-Correcting Diffusion Language Models
Block-wise diffusion large language models (dLLMs) decode sequentially at the block level, enabling effective KV-cache reuse across blocks but making inter-block decoding strictly serial. Prior work has attempted to unlock inter-block parallelism through post-training methods, but achieves only modest speedups and often degrades accuracy. We observe that self-correcting dLLMs offer a training-free alternative: token-to-token (T2T) editing can repair tokens drafted with a slightly stale upstream context, so a downstream block requires only an informative draft rather than a finalized predecessor. This turns block finality from a hard dependency into a scheduling resource. We propose \textbf{\flowblock{}}, a training-free parallel decoding framework built on two mechanisms. (i) \emph{Gated Wavefront Decoding} admits blocks into a bounded wavefront only when a readiness gate is satisfied, jointly refines active blocks via T2T editing, and commits blocks in order under a windowed block-causal mask that preserves exact frozen-prefix KV caches reuse. (ii) \emph{Heterogeneous Wavefront Packing} assigns each request an independent wavefront while packing asynchronous windows into dense, shape-stable batched forwards. Across different benchmarks, \flowblock{} improves tokens per second (TPS) over LLaDA-2.1 and LLaDA-2.0, two serial block-wise dLLMs, by up to 2.95$\times$ and 4.01$\times$, while reducing latency by up to 53.6\% and 77.1\%, respectively. It also improves average accuracy by 1.3 points. Compared with D2F, a training-based inter-block-parallel baseline, \flowblock{} achieves higher accuracy and up to 16$\times$ higher batched serving throughput.
OrientSAM: 方向を意識した空間アライメントによるマルチモーダル空間推論におけるカメラ中心のショートカットの軽減
マルチモーダル大規模言語モデル (MLLM) は、視点変換を必要とする空間推論に依然として苦労しています。特に、参照オブジェクトの視点から推論するのではなく、カメラ中心の手がかりに依存することが多く、カメラ以外の参照設定で系統的なエラーが発生します。この論文では、まずこの障害モードを分析し、オブジェクト指向がそのようなカメラ中心のショートカット動作の基礎となる重要な要素であることを示します。この問題に対処するために、マルチモーダル モデル用の方向を認識した空間位置合わせフレームワークである OrientSAM を提案します。 OrientSAM は、方向を意識したトークンとフーリエベースの角度エンコーディングを通じて、明示的な方向情報をマルチモーダル表現に注入し、さらにカリキュラム学習戦略を採用して、遠近感を意識した推論を段階的に改善します。さらに、大規模な画像から方向を意識した空間監視を生成するための空間データ構築パイプラインを構築します。 Spatial-MM、ViewSpatial、および 3DSRBench の実験では、OrientSAM が、特に非カメラビュー、人物中心、向きに依存するタスクにおいて、一貫して強力なベースラインを上回るパフォーマンスを示しています。この結果は、カメラ中心のショートカット動作を軽減し、マルチモーダル モデルでより堅牢な他中心の空間推論を可能にするために、明示的な方向モデリングが重要であることをさらに示しています。
原文 (English)
OrientSAM: Mitigating Camera-Centric Shortcut in Multimodal Spatial Reasoning via Orientation-Aware Spatial Alignment
Multimodal large language models (MLLMs) still struggle with spatial reasoning that requires perspective transformation. In particular, they often rely on camera-centric cues rather than reasoning from the reference object's viewpoint, leading to systematic errors in non-camera reference settings. In this paper, we first analyze this failure mode and show that object orientation is a key factor underlying such camera-centric shortcut behavior. To address this issue, we propose OrientSAM, an orientation-aware spatial alignment framework for multimodal models. OrientSAM injects explicit orientation information into multimodal representations through orientation-aware tokens and Fourier-based angle encoding, and further adopts a curriculum learning strategy to progressively improve perspective-aware reasoning. In addition, we build a spatial data construction pipeline to generate orientation-aware spatial supervision from large-scale images. Experiments on Spatial-MM, ViewSpatial, and 3DSRBench show that OrientSAM consistently outperforms strong baselines, especially on non-camera-view, person-centric, and orientation-sensitive tasks. The results further demonstrate that explicit orientation modeling is important for mitigating camera-centric shortcut behavior and enabling more robust allocentric spatial reasoning in multimodal models.
持続可能なスマートシティにおける交通行動の理解と管理のための人工知能
都市交通システムは異種データを生成しますが、これらのデータは自動的に実用的な管理インテリジェンスになるわけではありません。この章では、人工知能 (AI) について行動中心の視点を採用し、移動記録と乗客が生成したテキストを行動の真実ではなく行動の証拠として扱います。サービスの信頼性のためのバス到着予測、需要分析と計画のためのタクシー移動パターンの発見、責任ある規制サポートのための異常行動検出、サービス向上のための乗客認識リスクマイニングの 4 つの方向性を検討します。これらの方向性は、データ入力、動作表現、AI 推論、意思決定支援、公共価値、ガバナンス フィードバックをリンクする閉ループ フレームワークを通じて統合されます。この章では、展開の必須条件として、データ品質、プライバシー、公平性、解釈可能性、不確実性、転送可能性、および人間の説明責任を特定します。これにより、行動の証拠から運営、計画、規制、乗客サービスの決定に至るまでの統一された経路が確立されます。
原文 (English)
Artificial Intelligence for Understanding and Managing Transportation Behavior in Sustainable Smart Cities
Urban transportation systems generate heterogeneous data, yet these data do not automatically become actionable management intelligence. This chapter adopts a behavior-centered perspective on artificial intelligence (AI), treating mobility records and passenger-generated text as behavioral evidence rather than behavioral truth. It examines four directions: bus arrival prediction for service reliability, taxi mobility pattern discovery for demand analysis and planning, abnormal behavior detection for accountable regulatory support, and passenger-perceived risk mining for service improvement. These directions are integrated through a closed-loop framework linking data input, behavior representation, AI inference, decision support, public value, and governance feedback. The chapter identifies data quality, privacy, fairness, interpretability, uncertainty, transferability, and human accountability as essential conditions for deployment. It thereby establishes a unified pathway from behavioral evidence to operational, planning, regulatory, and passenger-service decisions.
ProEvent: プロアクティブなエージェント向けのイベント中心のベンチマーク
プロアクティブなエージェントは、ユーザーのニーズを予測し、明示的な指示がなくても環境コンテキストを認識することで自律的な支援を提供することが期待されています。このようなエージェントの基本的な機能は、ユーザーの今後のイベントを特定して追跡し、継続的かつイベント固有の支援を可能にすることです。たとえば、計画されたハイキングの時間と場所を記録することで、エージェントは事前に天候に関するリマインダーを配信したり、出発前にナビゲーション サポートを提供したりできます。しかし、プロアクティブ エージェントに関する既存の研究では、イベント中心の支援がほとんど見落とされており、プロアクティブな支援の制限のない性質により、信頼性の高い評価に課題が生じています。これらのギャップを埋めるために、進行中のインスタント メッセージング チャットに基づいてユーザーのスケジュールをプロアクティブに維持するエージェントの能力を評価するように設計された初のイベント中心のベンチマークである ProEvent を導入します。 ProEvent は、ユーザー間の動的な対話、同時チャット スレッド、現実世界のノイズを考慮した、合成された現実的なチャットを提供し、応答のタイミング、シングル ステップの応答の正しさ、およびマルチステップの応答の正しさに関してプロアクティブなエージェントを評価します。 8 つの LLM とパイプラインに関する実験により、現在のエージェントが頻繁に過剰に行動し、イベントのキャンセルに苦労していることが明らかになりました。特に、GPT-5.1 でさえ、シナリオの 26.7% でしか正しく反応しません。さらに定性的な分析を行うと、特に暗黙的なイベントの検出やユーザーの一人称視点からの推論において、プロアクティブ エージェントとしての現在の LLM の基本的な限界が明らかになります。
原文 (English)
ProEvent: An Event-centric Benchmark for Proactive Agents
Proactive agents are expected to anticipate user needs and provide autonomous assistance by perceiving environmental context without explicit instructions. A fundamental capability of such agents is to identify and track users' upcoming events, enabling continuous and event-specific assistance. For example, by recording the time and location of a planned hike, an agent can deliver weather reminders in advance or provide navigation support before departure. However, existing works on proactive agents largely overlook event-centric assistance, and the open-ended nature of proactive assistance poses challenges for reliable evaluation. To bridge these gaps, we introduce ProEvent, the first event-centric benchmark designed to assess an agent's ability to proactively maintain a user's timetable based on ongoing instant messaging chats. ProEvent provides synthesized yet realistic chats that consider the dynamic interaction among users, concurrent chat threads, and noise in the real world, and evaluates proactive agents on response timing, single-step response correctness, and multi-step response correctness. Experiments on eight LLMs and pipelines reveal that current agents frequently overact and struggle with event cancellation. Notably, even GPT-5.1 only reacts correctly in 26.7% of scenarios. Further qualitative analysis reveals fundamental limitations of current LLMs as proactive agents, particularly in detecting implicit events and reasoning from the user's first-person perspective.
LaT: マルチタスク車両ルーティング ソルバーのトレーナーとしての LLM
マルチタスク ニューラル ソルバーは、統合モデル内で複数の配車経路問題 (VRP) バリアントを処理し、制約の組み合わせごとに個別にトレーニングすることを回避することを目的としています。ただし、VRP バリアントは最適化の難易度が異なりますが、既存の手法にはトレーニング ステータスに関する段階的なフィードバックが不足しているため、モデルがいくつかの特定のバリアントに偏っています。メタ学習は適応トレーニングをサポートできますが、通常は 2 レベルの最適化と追加の勾配更新が必要となり、計算コストが増加します。この制限に対処するために、事前トレーニングされた大規模言語モデルを外部トレーナーとして使用するプラグアンドプレイ トレーニング パラダイムである LLM-as-Trainer (LaT) を提案します。 LaT は、クロスタスク検証メトリクスを定期的に分析して、段階ごとのガイダンス ベクトルを生成します。このベクトルは現在のタスクの制約ベクトルと結合されて各エンコーダ層に注入され、後続のポリシー最適化中にニューラル ソルバーに追加のトレーニング情報を提供します。 16 の VRP バリアントに関する実験では、LaT がトレーニング済みバリアントと未確認バリアントの両方でいくつかの最先端のマルチタスク ニューラル ソルバーの解の品質を向上させ、提案されたトレーニング パラダイムの有効性と汎用性をサポートすることを示しています。
原文 (English)
LaT: LLM-as-Trainer for Multi-Task Vehicle Routing Solvers
Multi-task neural solvers aim to handle multiple Vehicle Routing Problem (VRP) variants within a unified model, avoiding separate training for each constraint combination. However, VRP variants differ in optimization difficulty, while existing methods lack stage-wise feedback on their training status, making the model biased to some specific variants. Although meta-learning can support adaptive training, it typically requires bi-level optimization and additional gradient updates, increasing computational cost. To address this limitation, we propose LLM-as-Trainer (LaT), a plug-and-play training paradigm that uses a pretrained large language model as an external trainer. LaT periodically analyzes cross-task validation metrics to generate a stage-wise guidance vector. This vector is combined with the current task's constraint vector and injected into each encoder layer, providing the neural solver with additional training information during subsequent policy optimization. Experiments on 16 VRP variants show that LaT improves the solution quality of several state-of-the-art multi-task neural solvers on both trained and unseen variants, supporting the effectiveness and generality of the proposed training paradigm.
クロスモーダル否定の検出方法の学習: 潜在表現の分析と注意ベースのソリューション
現在のマルチモーダル システムにとって、モダリティ全体での否定のような高レベルの意味概念を検出することは依然として課題です。我々はこれを基本的な表現学習問題として分析し、標準視覚言語モデル (VLM) の潜在空間において否定が線形または非線形分離可能なクラスを形成しないという最初の証拠を提供します。我々は、事前訓練された埋め込みが主にモダリティ固有の特徴をエンコードし、一般化可能な否定信号を欠いていることを実証します。これを克服するために、モーダル間の依存関係を明示的にモデル化する新しいクロスモーダル アテンション アーキテクチャを提案し、ユニモーダル ベースラインと比較して最大 +7.03% F1 のパフォーマンス向上を達成します。私たちの分析では、重要な非対称性が明らかになりました。テキストの否定は独立して現れることが多いのに対し、視覚的な否定は意味的に言語文脈に依存しており、この発見は、\textsc{Qwen2.5-VL} によって自動的に注釈が付けられた 3,222 の政治的なビデオとテキストのペアの統計分析によって検証されました。この分析を自己教師ありビデオ表現 (JEPA2) と組み合わせることで、時間的否定のモデリングを進めます。この研究は、マルチモーダル システムにおける堅牢で意味的に整合した表現を学習するための新しい方法と洞察を提供します。
原文 (English)
Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution
Detecting high-level semantic concepts like negation across modalities remains a challenge for current multimodal systems. We analyze this as a fundamental representation learning problem, providing the first evidence that negation does not form a linearly or non-linearly separable class in the latent spaces of standard vision-language models (VLMs). We demonstrate that pretrained embeddings primarily encode modality-specific features, lacking a generalizable negation signal. To overcome this, we propose a novel cross-modal attention architecture that explicitly models inter-modal dependencies, achieving performance gains of up to +7.03% F1 over unimodal baselines. Our analysis reveals a key asymmetry: while textual negation often appears independently, visual negation is semantically dependent on linguistic context, a finding validated through our statistical analysis of 3,222 political video-text pairs automatically annotated via \textsc{Qwen2.5-VL}. By combining this analysis with self-supervised video representations (JEPA2), we advance the modeling of temporal negation. This work provides new methods and insights for learning robust, semantically-aligned representations in multimodal systems.
SR-Agent: E コマース レコメンデーションにおけるランキング後の戦略を洗練するためのエクスペリエンス主導型エージェント フレームワーク
ユーザー エクスペリエンスは、産業用電子商取引レコメンダー システム (RS) の第一級の目標です。ランク付け後の戦略は、ランク付けされたリストにおける多様性、類似性、露出を管理するもので、そのシンプルさと提供コストの低さから産業用 RS に広く導入されています。ただし、オンライン レコメンデーション環境が継続的に進化するにつれて、これらの静的に構成された戦略は徐々に古くなり、ユーザー エクスペリエンスが低下します。通常、これらを改良するには手動の検査、診断、更新が必要ですが、このプロセスは時間がかかり、コストがかかり、再利用が困難です。最近の LLM ベースのエージェント (RecUserSim、SimUSER、Self-EvolveRec など) は有望な方向性を示していますが、自動化され自己進化する戦略の改良の完全なループを閉じるものはありません。このギャップを埋めるために、SR-Agent を導入します。SR-Agent は、Strategy Refinement エージェント フレームワークであり、私たちの知る限り、産業用 RS でポストランキング戦略を洗練するために導入されたのは初めてです。 SR-Agent は 3 つのコンポーネントを統合します。(i) 段階的な検査スキルを適用して、ユーザーが認識する悪いケースを表面化する UserSim エージェント。 (ii) 再発する悪いケースを構造化された再利用可能な診断に統合する分析エージェント。 (iii) 診断を型指定された制限されたアクションにマッピングする制約付き戦略洗練ハーネス。可逆ロールバックを備えた 4 段階の報酬パイプラインによってゲートされます。 Kuaishou e-commerce プラットフォームに導入された SR-Agent は、この絞り込みループを継続的に実行し、1 か月のオンライン A/B テストで注文量を 0.71%、閲覧深度を 0.34%、クリックしたカテゴリの多様性を 0.48% 増加させながら、絞り込みサイクルを大幅に短縮し、運用コストを削減しました。
原文 (English)
SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategies Refinement in E-Commerce Recommendation
User experience is a first-class objective in industrial e-commerce recommender systems (RS). Post-ranking strategies, which govern diversity, similarity, and exposure over a ranked list, are widely deployed in industrial RS for their simplicity and low serving cost. However, as the online recommendation environment evolves continuously, these statically configured strategies gradually become stale, degrading the user experience. Refining them typically relies on manual inspection, diagnosis, and updates, a process that is slow, costly, and hard to reuse. Although recent LLM-based agents (e.g., RecUserSim, SimUSER, and Self-EvolveRec) offer promising directions, none of them close the full loop of automated, self-evolving strategy refinement. To bridge this gap, we introduce SR-Agent, a Strategy Refinement agentic framework that, to the best of our knowledge, is the first deployed for refining post-ranking strategies in industrial RS. SR-Agent unifies three components: (i) a UserSim agent that applies staged inspection skills to surface user-perceived bad cases; (ii) an Analysis agent that consolidates recurring bad cases into structured, reusable diagnoses; and (iii) a constrained Strategy Refinement Harness that maps diagnoses to typed and bounded actions, gated by a four-stage reward pipeline with reversible rollback. Deployed on the Kuaishou e-commerce platform, SR-Agent continuously runs this refinement loop and, in a one-month online A/B test, increases order volume by 0.71%, browsing depth by 0.34%, and clicked-category diversity by 0.48%, while markedly shortening the refinement cycle and lowering operational cost.
意味的には似ているが、論理的には異なる: テーブル RAG の意味と応答性のギャップを診断する
テーブルは検索拡張生成 (RAG) における重要な知識源ですが、取得されたテーブルには、クエリに答えるための十分な証拠 (応答性と呼ばれる特性) が欠けている場合があります。回答可能性は、ソースまたはソースのコレクションに十分な証拠が含まれているかどうかに広く関係しますが、意味論的な関連性に関して最適化された検索モデルは、単一ソースの場合であってもそれを保証せず、根本的な不一致が生じます。これを研究するために、RAG のテーブル コンテンツ レベルの応答性の診断ベンチマークである TCR-Bench を導入します。これは、兄弟テーブル、つまり、スキーマは非常に類似しているが内容が微妙に異なるテーブルを中心に構築されています。 TCR-Bench では、私たちが評価した密な取得者はセマンティック応答性ギャップを常に示しています。つまり、正しい兄弟グループを取得することがよくありますが、その中で一意に応答可能なテーブルを正確に特定するのに苦労し、QA パフォーマンスが 0.755 (オラクル) から 0.330 (上位 5 件の取得) に低下しました。私たちの分析では、このギャップが意味論的な蓄積、スキーマレベルのキュー依存性、および弱い行と列のバインディングに関連していることが示唆されています。このギャップの原因を診断するためのプローブとして、クエリ テーブルの応答性判定を直接適用する軽量の 2 段階パイプラインである応答性を考慮した再ランキング (AAR) がパフォーマンスを回復できるかどうかをテストします。これにより、上位 1 ターゲットの取得率が 18.2% から 57.4% に向上しました。この大幅な向上自体が、観察された障害の多くが、モデルの容量のみの固有の制限ではなく、応答性検証ステップの欠如を反映していることの証拠です。
原文 (English)
Semantically Similar, Logically Distinct: Diagnosing the Semantic-Answerability Gap in Table RAG
Tables are a critical knowledge source in retrieval-augmented generation (RAG), but a retrieved table may lack sufficient evidence to answer a query, a property we call answerability. While answerability broadly concerns whether a source or collection of sources contains sufficient evidence, retrieval models optimized for semantic relevance do not guarantee it even in the single-source case, creating a fundamental mismatch. To study this, we introduce TCR-Bench, a diagnostic benchmark for Table Content-level Answerability in RAG, built around sibling tables, i.e., tables with highly similar schemas but subtle content differences. On TCR-Bench, the dense retrievers we evaluate persistently exhibit a Semantic-Answerability Gap: they often retrieve the correct sibling group yet struggle to pinpoint the uniquely answerable table within it, dropping QA performance from 0.755 (oracle) to 0.330 (top-5 retrieved). Our analysis suggests this gap is associated with semantic accumulation, schema-level cue dependence, and weak row-column binding. As a diagnostic probe into the source of this gap, we test whether a lightweight two-stage pipeline, Answerability-Aware Reranking (AAR), applying direct query-table answerability judgment, can recover performance: it raises top-1 target retrieval from 18.2% to 57.4%, and this large gain is itself evidence that much of the observed failure reflects a missing answerability verification step, rather than an inherent limitation of model capacity alone.
WuYu-EnvLE-Bench: 環境法執行における大規模言語モデルを評価するためのベンチマーク
大規模言語モデル (LLM) は、環境施行のためにますます検討されていますが、追跡可能な施行決定を生成するその能力は依然として不明です。実際の施行事例、規制基準、専門家のレビューから構築されたベンチマークである WuYu-EnvLE-Bench を紹介します。これには、施行前、施行中、施行後のワークフローにわたる 2,521 のベンチマーク インスタンス、14 のタスク、12 の汚染媒体サブドメインが含まれています。絶対環境施行スコア (AES) とインテリジェント施行指数 (IEI) を使用して、オープンソースとクローズドソースの LLM を、能力、対応品質、リソース効率全体にわたって評価します。結果は、LLM はルールに限定されたタスクでは良好に機能しますが、証拠チェーンの構築、矛盾の検出、マルチソースの統合、および手続き上の判断では信頼性が低いままであることを示しています。モデルのスケーリングも利益の逓減を示しています。中規模のモデルは構造化タスクにおいて主要なモデルに近づきますが、大規模なモデルは証拠推論のボトルネックを確実に克服できません。 WuYu-EnvLE-Bench は、証拠に基づいた、ルールを認識した、タスクに適応した施行推論の必要性を強調しています。
原文 (English)
WuYu-EnvLE-Bench: A Benchmark for Evaluating Large Language Models in Environmental Law Enforcement
Large language models (LLMs) are increasingly considered for environmental enforcement, but their ability to produce traceable enforcement decisions remains unclear. We introduce WuYu-EnvLE-Bench, a benchmark built from real enforcement cases, regulatory standards, and expert review. It contains 2,521 benchmark instances, 14 tasks, and 12 pollution-medium subdomains across pre-enforcement, in-enforcement, and post-enforcement workflows. Using Absolute Environmental Enforcement Score (AES) and Intelligent Enforcement Index (IEI), we evaluate open-source and closed-source LLMs across capability, response quality, and resource efficiency. Results show that LLMs perform well on rule-bounded tasks but remain unreliable in evidence-chain construction, contradiction detection, multi-source integration, and procedural judgment. Model scaling also shows diminishing returns: medium-sized models approach leading models in structured tasks, while larger models do not reliably overcome evidence-reasoning bottlenecks. WuYu-EnvLE-Bench highlights the need for evidence-grounded, rule-aware, and task-adaptive enforcement reasoning.
動的防御プロファイリングにより、テキストから画像へのモデルの認知脱獄が可能になります
Text-to-Image (T2I) 生成モデルは、高品質のビジュアル コンテンツの合成において目覚ましい進歩を遂げていますが、特に Not-Safe-For-Work (NSFW) 画像の生成においては、依然として敵対的な悪用に対して脆弱です。既存のジェイルブレイク攻撃のほとんどは、主にヒューリスティック プロンプト エンジニアリングまたはブラック ボックス最適化に依存しており、モデルのフィードバックをバイナリ信号 (成功または失敗) として扱います。この粗いパラダイムは、テキストの拒否、視覚的なブロック、セマンティックなサニタイズなどのさまざまな障害モードに埋め込まれた豊富な情報を見落とし、その結果、非効率な探索と深刻なセマンティック崩壊が発生します。この論文では、敵対的プロンプト生成を潜在的な防御メカニズムに対する信念状態推論問題として再構成する認知脱獄フレームワークである MIND を提案します。 MIND は、バイパス プロンプトを盲目的に検索するのではなく、マルチモーダル フィードバックを高密度信号として解釈することで、ターゲット システムの潜在的な防御メカニズムを積極的にモデル化します。具体的には、このフレームワークには 3 つのコア コンポーネントが統合されています。(1) きめ細かいフィードバック分解のためのマルチモーダル ジャッジ、(2) 反復的な信念更新のためのディフェンス プロファイラー、および (3) 歴史的に効果的な攻撃戦略を取得するためのメタメモリ モジュール。これらのコンポーネントは推論主導の進化的最適化プロセス内で統合されており、適応的で意味的に一貫したジェイルブレイク生成が可能になります。 I2P ベンチマークに関する広範な実験により、MIND の有効性が実証されています。 Stable Diffusion v1.5 モデルに適用された 6 つの代表的な前処理および後処理防御設定の下で、MIND は 95.62% の攻撃成功率 (ASR) を達成し、既存の方法を大幅に上回りました。さらに、提案されたフレームワークの有効性は、広く使用されている 4 つの商用 T2I システムにわたって検証され、Wan-2.5 では 91.58% という最高の ASR を達成しています。
原文 (English)
Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models
Text-to-Image (T2I) generative models have achieved remarkable progress in synthesizing high-quality visual content, yet they remain vulnerable to adversarial misuse, particularly in generating Not-Safe-For-Work (NSFW) images. Most existing jailbreak attacks primarily rely on heuristic prompt engineering or black-box optimization, treating model feedback as a binary signal (success or failure). This coarse-grained paradigm overlooks the rich information embedded in diverse failure modes, such as textual refusal, visual blocking, and semantic sanitization, resulting in inefficient exploration and severe semantic collapse. In this paper, we propose MIND, a cognitive jailbreak framework that reframes adversarial prompt generation as a belief-state inference problem over latent defense mechanisms. Instead of blindly searching for bypass prompts, MIND actively models the target system's latent defense mechanisms by interpreting multi-modal feedback as high-density signals. Specifically, the framework integrates three core components: (1) a Multi-modal Judge for fine-grained feedback decomposition, (2) a Defense Profiler for iterative belief updating, and (3) a Meta-Memory module for retrieving historically effective attack strategies. These components are unified within a reasoning-driven evolutionary optimization process, enabling adaptive and semantically consistent jailbreak generation. Extensive experiments on the I2P benchmark demonstrate the effectiveness of MIND. Under six representative pre-processing and post-processing defense settings applied to the Stable Diffusion v1.5 model, MIND achieves an Attack Success Rate (ASR) of 95.62%, significantly outperforming existing methods. Additionally, the effectiveness of the proposed framework is validated across four widely used commercial T2I systems, achieving the highest ASR of 91.58% on Wan-2.5.
虚偽情報の検出と説明による財務監査支援
貸借対照表(BS)、損益計算書(IS)、キャッシュフロー計算書(CS)などの財務諸表(FS)は、企業の年間財務実績を要約します。 FS は、コーポレート・ガバナンスの評価、信用評価、リスク分析、課税の検証、投資決定などに広く使用されています。財務監査は複雑で知識集約的な分野であり、その重要な目的の 1 つは、公開された FS の完全性、正確性、公平性、および重大な虚偽表示のないことを保証することです。 FS の重要性を考慮すると、企業の真の財務健全性を偽って情報を隠蔽、省略、または改ざんするインセンティブが存在します。たとえば、納税義務の削減や投資家の信頼の向上などです。複雑で時間がかかり、専門知識に依存する監査の性質を考慮すると、監査人は、特定の FS 内の誤った情報のインスタンスを自動的に検出し、財務データ内の誤った情報の発生源と考えられるものを特定する AI 支援システムの恩恵を受けることになります。この論文では、FS で誤った情報を特定するための教師なし手法を紹介し、誤った情報の発生源である可能性が高い財務変数に関する説明も生成します。監査人は、関連するデータ ソースとビジネス プロセスをより詳細に調査して、これらの提案を検証できます。私たちのアプローチの重要な特徴は、過去の FS コーパスと関連する監査レポートを使用して、支援の提供に役立つ洞察を生成することです。我々は、5 年間にわたる 11,460 FS の大規模なコーパスと関連する監査レポートに対して、これらの手法の有効性を実証します。この論文は、以前に報告された研究 (Shinde et al., 2022)\cite{SVAP22}, (Vaishampayan et al., 2022)\cite{VSPP22}, (Pawar et al., 2023)\cite{PAPV23} を統合し、さらに新しい貢献を加えたものであり、AI 支援監査人支援システムの基盤として使用されています。
原文 (English)
Financial Audit Assistance using Misinformation Detection and Explanation
Financial statements (FS) such as Balance Sheet (BS), Income Statement (IS) and Cash-flow Statement (CS) summarize the annual financial performance of a company. FS are widely used for evaluating corporate governance, credit appraisal, risk analysis, validate taxation, make investment decisions etc. Financial auditing is a complex and knowledge-intensive discipline whose one important aim is ensuring integrity, accuracy, fairness and absence of material misstatement in the published FS. Given the importance of FS, there are incentives to hide, omit or falsify information to misrepresent the true financial health of the company; e.g., reduce tax liabilities, or increase investor confidence. Given the complex, time-consuming and expertise-dependent nature of auditing, auditors would benefit from an AI-assisted system that automatically detects instances of misinformation in the given FS and identify likely sources of this misinformation in the financial data. In this paper, we present unsupervised techniques to identify misinformation in FS, and also generate explanations as to the financial variables that are likely sources of misinformation. The auditor can then explore in more detail the associated data sources and business processes to validate these suggestions. A crucial feature of our approach is the use of past corpus of FS and associated audit reports to generate insights, which help in providing assistance. We demonstrate the efficacy of these techniques on a large corpus of 11,460 FS over 5 years and associated audit reports. This paper integrates and adds more novel contributions over the previously reported research (Shinde et al., 2022)\cite{SVAP22}, (Vaishampayan et al., 2022)\cite{VSPP22}, (Pawar et al., 2023)\cite{PAPV23}, which we have used as the foundation for our AI-assisted Auditor Assistance system.
PGN: Pangu マルチモーダル基盤モデルに基づく視覚言語ナビゲーション システムの設計と実装
Vision-Language Navigation (VLN) では、身体化されたエージェントが自然言語の命令を解釈し、時間的に順序付けられた視覚観察からアクションを予測する必要があります。マルチモーダルな大規模言語モデルを VLN に適応させるには、視覚言語の調整、コンパクトな時間入力、アクション空間のグラウンディング、およびターゲット ハードウェアでの安定したトレーニングが必要です。この技術レポートでは、OpenPangu-7B 上に構築されたオフライン VLN 行動予測システムである PGN (Pangu Navigator) について紹介します。トレーニングは 2 つの段階で進められます。まず、PGMM は、Q-Former と 2 層 MLP プロジェクターをトレーニングすることによって、凍結された EVA-ViT-G/14 ビジョン エンコーダーを凍結された言語バックボーンと調整します。第 2 に、PGN は、5 つの観測ウィンドウ、エポック依存の時間サンプリング、および推論してからアクションを実行する出力形式を使用して、調整されたモデルを専門家のナビゲーション軌跡に適応させます。この段階では、調整された視覚経路が凍結され、3 つの構造トークンの埋め込みと LoRA アダプターが更新されます。この実装では、8 つの Ascend 910B NPU 上で、混合精度計算、選択的 FP32 計算、DeepSpeed ZeRO-2 を組み合わせています。 500 件のエキスパートの実行軌跡に対する教師強制の開ループ評価のもとで、V9 は 62.29% の正規化アクション一致 (NAM) と 100.00% の非空率 (NER) を報告しました。これらの指標は、閉ループ ナビゲーションの成功ではなく、オフラインの専門家とアクションの連携を定量化します。エラーの蓄積、パスの効率、目標の完了の評価は今後の課題です。
原文 (English)
PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model
Vision-Language Navigation (VLN) requires an embodied agent to interpret a natural-language instruction and predict actions from temporally ordered visual observations. Adapting a multimodal large language model to VLN requires visual-language alignment, compact temporal inputs, action-space grounding, and stable training on the target hardware. This technical report presents PGN (Pangu Navigator), an offline VLN action-prediction system built on OpenPangu-7B. Training proceeds in two stages. First, PGMM aligns a frozen EVA-ViT-G/14 vision encoder with the frozen language backbone by training a Q-Former and a two-layer MLP projector. Second, PGN adapts the aligned model to expert navigation trajectories using five-observation windows, epoch-dependent temporal sampling, and a reasoning-then-action output format; this stage freezes the aligned visual pathway and updates three structural-token embeddings and LoRA adapters. The implementation combines mixed-precision computation, selective FP32 computation, and DeepSpeed ZeRO-2 on eight Ascend 910B NPUs. Under teacher-forced, open-loop evaluation on 500 held-out expert trajectories, V9 reports a 62.29% Normalized Action Match (NAM) and a 100.00% Non-empty Rate (NER). These metrics quantify offline expert-action alignment rather than closed-loop navigation success; evaluating error accumulation, path efficiency, and goal completion remains future work.
効率的なベイジアン推論の計算と展開のためのハードウェア指向のアプローチ
ベイジアン推論は、不確実性の下で推論するための原則に基づいた基盤を提供しますが、その計算コストにより、リソースに制約のあるエッジ デバイスへの展開が妨げられます。この論文では、市販の組み込み GPU で離散ベイズ推論を高速化するためのハードウェア指向の方法論を紹介します。私たちは、広範なクラスの変分メッセージ受け渡しアルゴリズムのレイテンシーがテンソル短縮によって支配されていることを確認しました。私たちのアプローチは、効率的な GPU 実行により適したコンパクトで規則的な形状のプリミティブを生成する 2 つの相補的なマージ戦略を使用して、これらの操作のメモリ レイアウトを再構築します。次に、メモリ フットプリントを削減するために、オプションのスパース配列表現とテンソル クラスタリング スキームを導入します。私たちは方法論をインスタンス化し、隠れマルコフ モデル (HMM) 用の 3 つのメッセージ パッシング アルゴリズム、つまり変分フィルタリング、変分メッセージ パッシング、およびマージナル メッセージ パッシングの最適化されたバリアントを生成します。さらに、特定の生成モデル仕様に対して最もパフォーマンスの高いアルゴリズムのバリアントを自動的に選択する機械学習ベースの自動チューナーでこれを補完します。 770 個のランダムにサンプリングされた現実的な部分観察可能なマルコフ決定プロセス (POMDP) 構成にわたる NVIDIA Jetson Orin AGX でベンチマークを行ったところ、当社の実装は、ベースライン実装と数値的に同一の出力を生成しながら、最大 5 倍の高速化を達成し、通常のゲインは 2 ~ 2.5 倍でした。
原文 (English)
A Hardware-oriented Approach for Efficient Bayesian Inference Computation and Deployment
Bayesian inference provides a principled foundation for reasoning under uncertainty, but its computational cost hinders deployment on resource-constrained edge devices. In this paper, we present a hardware-oriented methodology for accelerating discrete Bayesian inference on commercial off-the-shelf embedded GPUs. We identify that the latency of a broad class of variational message-passing algorithms is dominated by tensor contractions. Our approach restructures the memory layout of these operations using two complementary merging strategies that produce compact, regularly-shaped primitives better suited for efficient GPU execution. We then introduce optional sparse array representations and a tensor-clustering scheme to reduce the memory footprint. We instantiate the methodology and produce optimized variants of three message-passing algorithms for Hidden Markov Models (HMMs), namely variational filtering, variational message passing, and marginal message passing. Furthermore, we complement this with a machine-learning-based autotuner that automatically selects the best-performing algorithmic variant for a given generative model specification. Benchmarked on an NVIDIA Jetson Orin AGX across 770 randomly sampled realistic Partially Observable Markov Decision Process (POMDP) configurations, our implementations achieve speedups of up to 5x, with typical gains of 2-2.5x, while producing numerically identical outputs to the baseline implementations.
探索的および同化的な反射: 長期記憶のための反射的想起サイクル
LLM ベースの自律エージェントは、ステートレス性を克服するために外部メモリを必要とし、長期的な対話と動的な知識推論のためにコンテキスト ウィンドウが制限されます。しかし、既存のメモリ取得方法は適応性やサンプル効率に欠けていることが多く、異種のストアから適切に混合したメモリを取得するのが困難です。我々は、高い初期検索パフォーマンスとサンプル効率の高い適応のためのフレームワークである Exploratory-Assimiiling Reflection (EAR) を提案します。 EAR は 2 つのメカニズムを組み合わせています。Exploratory Reflection は、反復検索を実行して取得をブートストラップし、クエリごとに有用なエクスペリエンスを収集します。同化リフレクションは、Experience Buffer からこれらのエクスペリエンスを再生して、即時の報酬のみに依存する方法よりも効率的にグローバル リランカーを洗練します。実験によると、2 つの長期対話ベンチマークにおいて、EAR はベースライン検索よりも検索が最大 17.9% 向上します。また、EAR はサンプル効率が高く、ノイズの多いフィードバックに対して堅牢であることも示します。
原文 (English)
Exploratory and Assimilating Reflection: Reflective Recall Cycle for Long-term Memory
LLM-based autonomous agents require external memory to overcome their statelessness and limited context window for long-term interaction and dynamic knowledge reasoning. However, existing memory retrieval methods often lack adaptability and sample efficiency, and struggle to retrieve the right mixture of memories from heterogeneous stores. We propose Exploratory-Assimilating Reflection (EAR), a framework for high initial retrieval performance and sample-efficient adaptation. EAR combines two mechanisms: Exploratory Reflection, which performs iterative search to bootstrap retrieval and collect useful experiences for each query, and Assimilating Reflection, which replays these experiences from an Experience Buffer to refine a global reranker more efficiently than methods relying only on immediate rewards. Experiments show that EAR improves retrieval by up to 17.9% over the baseline retriever on two long-term dialogue benchmarks. We also show that EAR is highly sample-efficient and robust to noisy feedback.
ST-拒否権: テイラー予測とビジュアルグラウンディングによる拡散 MLLM に対する時空間トークン拒否権
ビジョン言語モデル (VLM) は、思考連鎖 (CoT) による強力な推論を実現しますが、逐次生成コストが高く、エラーが蓄積し、自己修正が制限されます。拡散マルチモーダル大規模言語モデル (dMLLM) は、順序に依存しないプロセスでトークンのマスクを解除し、効率を向上させ、反復的な改良を可能にしますが、その推論とそれを強化する方法はまだ解明されていません。私たちは、各拡散ステップですべてのトークンの位置を観察する能力を活用する、トレーニング不要の方法である時空間トークン拒否 (ST-Veto) を提案します。 ST-Veto は、現在のステップの信頼度だけに依存するのではなく、信頼度ダイナミクスの 2 次テイラー予測を通じて時間的に不安定なトークンを拒否し、画像注意質量を使用して根拠の弱いトークンをフィルターし、より安全な候補と交換します。複数の dMLLM およびマルチモーダル推論ベンチマーク全体で、ST-Veto は標準のデコード ポリシーや以前の VLM 推論方法を常に上回っており、追加のトレーニングや生成コストをかけずに精度を最大 9% 向上させます。分析の結果、ST-Veto は生成をより信頼性の高い、より適切に根拠のあるパスに誘導することが示されています。
原文 (English)
ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding
Vision Language Models (VLMs) achieve strong reasoning with Chain-of-Thought (CoT) prompting but incur high sequential-generation cost, error accumulation, and limited self-correction. Diffusion Multimodal Large Language Models (dMLLMs) unmask tokens in an order-agnostic process, improving efficiency and enabling iterative refinement, yet their reasoning and how to enhance it remain underexplored. We propose a training-free method, Spatio-Temporal Token Veto (ST-Veto), which leverages the ability to observe all token positions at each diffusion step. Rather than relying only on current-step confidence, ST-Veto vetoes temporally unstable tokens via second-order Taylor prediction of confidence dynamics and filters weakly grounded tokens using image-attention mass, swapping them with safer candidates. Across multiple dMLLMs and multimodal reasoning benchmarks, ST-Veto consistently outperforms standard decoding policies and prior VLM reasoning methods, improving accuracy by up to 9% with no additional training or generation cost. Analyses show that ST-Veto steers generation toward higher-confidence, better-grounded paths.
大規模言語モデルエージェントによるストレステストの概念消去
概念消去は、トレーニングされた生成モデルから意味概念を削除することを目的としており、責任ある AI の導入にとってますます重要になっています。ただし、モデルが対象の概念を確実に削除したかどうかを検証することは、依然として重要な課題です。既存の評価方法は通常、事前定義された静的なものであり、多様な自然言語プローブや困難な条件下で脆弱性を明らかにすることができません。さらに、手動で設計された評価戦略は偏りがある可能性があり、拡張することが困難です。私たちは、概念消去評価は、故障モードの対象範囲を体系的に拡大するためにテストを繰り返し提案、批評、検証するエージェントによって運用される、適応的な仮説検索として定式化するのが最適であると仮定します。この目的を達成するために、我々は、複数の大規模言語モデル (LLM) エージェントを使用して、外部の知識に基づいたストレス テスト仮説を繰り返し生成して検証することにより、概念消去モデルを自律的にストレス テストするフレームワークである概念消去用ストレス テスト エージェント (STACE) を提案します。また、LLM エージェントを利用したストレス テスト フレームワークのパフォーマンスと効率を評価するための一連の指標も紹介します。私たちの広範な実験により、STACE は 4 つの概念カテゴリに関する 5 つの LLM ベースの評価ベースラインよりも優れていることがわかりました。 2 つの T2I モデル、6 つの概念消去アプローチ、およびさまざまな消去強度にわたるさらなる分析により、STACE がさまざまな設定に対して堅牢であることが示されています。また、STACE が概念消去評価を超えて、LLM ジェイルブレイクなどの他の問題領域に適応できることも示します。私たちのコードは匿名で利用できます。
原文 (English)
Stress Testing Concept Erasure with Large Language Model Agents
Concept erasure aims to remove semantic concepts from a trained generative model and is increasingly important for responsible AI deployment. However, verifying whether a model has robustly removed targeted concepts remains a critical challenge. Existing evaluation methods are typically pre-defined and static, failing to expose vulnerabilities under diverse natural-language probes and challenging conditions. Moreover, manually designed evaluation strategies can be biased and difficult to scale. We posit that concept erasure evaluation is best formulated as an adaptive hypothesis search, operationalised by agents that iteratively propose, critique, and verify tests to systematically expand coverage of failure modes. To this end, we propose Stress Testing Agents for Concept Erasure (STACE), a framework that autonomously stress-tests concept-erased models using multiple Large Language Model (LLM) agents, by iteratively generating and verifying stress-testing hypotheses grounded by external knowledge. We also introduce a suite of metrics for assessing the performance and efficiency of LLM-agent-powered stress-testing frameworks. Our extensive experiments show that STACE outperforms five LLM-based evaluation baselines on four concept categories. Further analysis across two T2I models, six concept erasure approaches, and various erasure strengths show that STACE is robust for different settings. We also show that STACE can be adapted beyond concept erasure evaluation to other problem domains, such as LLM jailbreaking. Our code is available anonymously.
PEARL: 科学的推論グラフ抽出のための監査可能な修復
Scientific Reasoning Graph Extraction (SRGE) は、観察、証拠、中間主張、論文レベルの結論の間の明示的な関連性を回復することを目的としています。 LLM はグラフのような科学的な説明を生成できますが、その出力には、不正な構文、ドリフト エッジ ラベル、間違った向きのルート、および弱いソース アンカーが混在していることがよくあります。私たちは、ノイズの多い LLM グラフ応答を監査可能な推論グラフに変換し、厳密な意味的妥当性に向けて修復するトレーニング不要のフレームワークである PEARL (Peircean Extraction via Abstraction and Repair Layer) を提案します。 PEARL は、まず、閉じたパーセン スキーマの下で明示的なグラフ コンテンツを具体化し、次に、一致した証拠に基づいたジャッジ フィードバックを使用して、監査証跡を保存しながら、拒否されたエッジ タイプ、ローカル推論ステップ、ターミナル ルートを修復します。潜在推論チェーン抽出のベンチマークである ARCHE の 5 つの 70 ペーパー モデル アーカイブにおいて、PEARL は厳密なゲート パスを LLM ベースラインの 0/350 から 300/350 に引き上げ、平均 REA は 0.339 から 0.906 に改善しました。グラフは、制約のないグラフの再生成ではなく、検査可能な推論トレースを必要とする研究エージェントや AI 科学者のワークフローに信頼性レイヤーを提供します。コードと監査成果物は https://github.com/BohanSu/auditable-repair-reasoning-graphs/tree/300-350_workshop で入手できます。
原文 (English)
PEARL: Auditable Repair for Scientific Reasoning Graph Extraction
Scientific Reasoning Graph Extraction (SRGE) aims to recover explicit links among observations, evidence, intermediate claims, and paper-level conclusions. LLMs can produce graph-like scientific explanations, but their outputs often mix malformed syntax, drifting edge labels, incorrectly oriented roots, and weak source anchors. We propose PEARL (Peircean Extraction via Abstraction and Repair Layer), a training-free framework that turns noisy LLM graph responses into auditable reasoning graphs and repairs them toward strict semantic validity. PEARL first materializes explicit graph content under a closed Peircean schema, then uses matched evidence-grounded judge feedback to repair rejected edge types, local inference steps, and terminal roots while preserving an audit trail. On five 70-paper model archives from ARCHE, a benchmark for latent reasoning-chain extraction, PEARL raises strict gate passes from 0/350 for the LLM baseline to 300/350, with average REA improving from 0.339 to 0.906. The graphs provide a reliability layer for research-agent and AI scientist workflows that need inspectable reasoning traces rather than unconstrained graph regeneration. Code and audit artifacts are available at https://github.com/BohanSu/auditable-repair-reasoning-graphs/tree/300-350_workshop .
The Autonomous Agency Scale: A Behavioral Framework for Measuring Self-Directed Behavior in AI Systems
Existing AI measurement frameworks quantify cognitive capability, task automation, or catastrophic risk, but none measure autonomous agency…
Towards Agentic Agent-based Models: Feasibility, Performance, and Statistical Model Checking
Agent-based models (ABMs) rely on simple, explicit and reproducible rules for individual decision making, while complex collective behavior…
OntoExtend: A Framework for Requirement-driven and Scalable Ontology Extension with LLMs
Ontology extension refers to the process of enriching an existing ontology in response to emerging requirements, making it more complete. T…
SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning
Latent world models have emerged as a powerful planning paradigm by learning action-conditioned predictive dynamics and using them as inter…
Do Maps Still Matter for Machines: Revisiting the Role of Choropleth Maps in Foundation Model Spatial Understanding
Spatial understanding is crucial for foundation models (FMs), and maps have long helped humans organize and reason about geographic informa…
PAMD: Structured Adaptive Distances for Bisimulation Representations in Visual Reinforcement Learning
Many visual reinforcement learning (RL) algorithms learn representations by matching latent distances to a behavioral distance induced by r…
Rethinking Heterogeneous LLM Merging: A Weighted Model Averaging Perspective
Can large language models with substantially different parameter spaces be merged by direct weighted averaging, without training or semanti…
AdaHome: An Adaptive Smart Home Assistant using Local Small Language Models
Smart home assistants interpret a wide range of user commands, from explicit device control to underspecified and preference dependent requ…
The Shared Discovery Paradox: How a One-Answer Rule Turns Better Information into Worse Search
Organizations often pool dispersed information into one ranking and then allow many agents to act on that shared view. In a discovery probl…
WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear…
Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs
Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whethe…
Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering
Extended reasoning has become standard for frontier Large Language Models (LLMs), yet the trajectories these models produce remain largely…
SGA: Plug&Play Geometric Verification for Educational Video Synthesis
Recent work leverages Large Language Models (LLMs) to generate executable code for pedagogical animations using libraries such as Manim. Ho…
Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes
To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly labeled syllogistic reasoning benc…
DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth
Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction,…
What Makes Linguistic Representations Good Models of High-Level Visual Perception in the Human Brain?
Image descriptions represented with language models (LMs) predict human brain responses to naturalistic images in high-level visual regions…
Comparing Spectrogram Front-Ends for Abnormal Heart-Sound Detection with a Convolutional Neural Network
Heart disease kills a lot of people, and one cheap way to catch it early is by listening to heart sounds with a stethoscope, or better yet,…
Fully-sensorized smart-eyewear platform for on-device Machine Learning
This paper presents ARGO, a smart eyewear platform designed to bridge ergonomic comfort, high computational throughput, and energy efficien…
International Agreements to Limit Frontier AI: Objectives and Exit
An international agreement to limit AI development could be crucial to mitigate risks from AI. However, it remains unclear which conditions…
RouteCost: A Production-Inspired Multi-Stage Framework for Pre-Order Shipping Cost Estimation in E-Commerce
Accurate pre-order shipping cost estimation is important in e-commerce because it affects price presentation, margin planning, and conversi…
From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language
The growing use of statistical learning algorithms to infer human preferences from high-dimensional choice data runs up against a fundament…
Token-Level Cross-Modal Transformer with Contrastive Multi-Task Learning for Breast Cancer Subtype Classification and Survival Prediction
Integrating heterogeneous genomic and clinical modalities for joint cancer subtype classification and survival prediction remains a key cha…
HantaWatch: Federated Learning for Hantavirus Genomic Surveillance
Hantavirus genomic surveillance is limited by the distribution of sequence data, non-IID source heterogeneity, and constrained expert-revie…
OpenMHC: Accelerating the Science of Wearable Foundation Models
Mobile and wearable devices offer an unprecedented opportunity for continuous, passive health monitoring and active health coaching. Howeve…
The Failures of Marginal Influence-Based Attribution Methods for Global Time Series Explanations
Explainability methods for time series models predominantly produce flat attribution scores: they quantify the direct influence of a featur…
Quantizing Recursive Reasoning Models
Recursive reasoning models solve hard puzzles by applying compact, weight-tied blocks over many refinement steps. Because these blocks are…
Diffusion-corrected Autoregressive Fourier Neural Operator for Droplet Evolution Prediction
Predicting droplet evolution in material jetting, or Inkjet Printing (IJP), is essential for maintaining printing quality. However, long-ho…
Normalized Rewards for Preference Optimization
Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences. However, DAAs…
KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch?
Recent large language models (LLMs) can generate custom CUDA kernels that appear to outperform PyTorch on benchmarks such as KernelBench. B…
TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment
Fine-Tuning-as-a-Service (FTaaS) platforms let users train large language models (LLMs) on customized tasks, but this pipeline could erode…
RobustMAD: Evaluating Real-World Robustness of Multimodal Small Language Models for Deployable Anomaly Detection Assistants
Multimodal industrial anomaly inspection assistants are a critical component of next-generation smart factories, enabling interactive visio…
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents
Training multi-turn evidence-reading agents with outcome-only reinforcement learning is unstable because intermediate turns receive little…
Learning Structural Manipulability in Gate-Level Netlists Using Graph Neural Networks
Gate-level netlists exhibit intrinsic structural properties that influence signal propagation independently of functional simulation. We de…
Let the Data Decide: Supervision Analysis, Capability Trade-offs, and Adaptive Objective Routing in Continued Pre-Training via Off-Policy Distillation
Off-policy distillation is now central to large language model pre-training, yet how training data, objective parameterization, and model c…
High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration
Long-context large language model inference relies on the KV cache to avoid redundant attention computation, but incurs high memory and ban…
An Agentic Interface for End-to-End Probabilistic Seismic Hazard and Risk Analysis
Probabilistic seismic hazard and risk analyses are backbone to building codes, insurance pricing, and disaster management. Yet their open-e…
Benchmarking Machine Learning Models for Multi-Omics-Based Breast Cancer Prediction
Estrogen Receptor (ER) status is a critical biomarker in breast cancer diagnosis, prognosis, and treatment selection. Recent advances in hi…
SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling
Low-Rank Adaptation (LoRA) is a widely used parameter-efficient fine-tuning (PEFT) method for large language models. Under a fixed rank bud…
Comprehensive Evaluation of Machine Learning for Type 2 Diabetes Risk Prediction: Large-Scale External Validation and Fairness Analysis
Machine learning-based Type 2 diabetes risk prediction models obtain good internal validation results but lose effectiveness in real-world…
Feature Generation Using LLMs: An Evolutionary Algorithm Approach
A crucial step in machine learning pipelines is to present each entity with features or attributes that are representative of the character…
Discovery by Dreaming: Cross-Domain Recombination in Artificial Memory
Dreams splice together people, places, and times that never met. Neuroscience suggests this recombination is not noise, but a function driv…
From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training
Reinforcement learning (RL) has become a widely adopted technique for improving large language models (LLMs) on complex tasks. Despite this…
Neural Controlled Differential Equations for EMT-Level Surrogate Modeling of Grid-Forming Inverters
The application of artificial intelligence methods in power electronic converter modeling is becoming increasingly widespread, but existing…
AdaSurvMamba: Dynamic Fusion and Semantic Scanning for Multimodal Survival Analysis
Multimodal survival analysis utilizing whole slide images (WSIs) and genomic profiles is fundamental for cancer prognosis. Recently, state-…
Reducing Per-Sample Harm in Stochastic Optimization
Modern optimizers combine gradients from the current mini-batch with historical optimization state, such as momentum or adaptive moments. W…
Autonomous mechanistic discovery of colorectal cancer vulnerabilities via multi-scale AI swarms
The acceleration of automated scientific discovery has been fundamentally bottlenecked by the epistemic gap between the semantic reasoning…
Composable Verification Pipelines for Multi-Agent Systems
Existing approaches for reasoning about action and change provide expressive semantics for modeling dynamic systems, in most cases built on…
From Intent to Infrastructure: LLM-Driven Agent Compilers for ISAC Networks
Integrated sensing and communications (ISAC) is moving from proof-of-concept demonstrations to system-level deployment in sixth-generation…
Physics-Informed Feature Engineering 1D-CNN for Multilayer Cloud Detection from Geostationary Satellites
Multilayer cloud detection from active--passive observation is vital for numerical weather prediction. In this study, channel selections de…
ForensicNet: Lightweight Attention-Enhanced MobileNetV2 for Automated Face Identification
In forensic environments, automated identification of perpetrators is difficult due to pose changes, changes in light, occlusion, and lack…
Intelligence-Guided Adaptive Purification for DDoS-Resilient Quantum Networks: A CUDA-Q based Study
Quantum-repeater networks require adaptive control policies that balance entanglement generation rate, end-to-end fidelity, purification ov…
GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification
The rapid advancement of generative AI has outpaced our ability to reliably detect its outputs, particularly when detectors encounter gener…
Depth Estimators Are Implicit Neural Fields for 3D Scene Geometry Inpainting and Reconstruction
The 3D geometry of real-world scene data is often incomplete. Mainstream methods use depth estimators to inpaint missing structure. However…
Identity-Consistent Expression Fields: A Disentangled Neural Radiance Field Framework for Few-Shot Facial Expression Synthesis
Neural Radiance Fields (NeRF) have enabled photorealistic novel-view synthesis of 3D scenes and, in the facial domain, have been extended t…
It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability
Brain-encoding foundation models predict fMRI responses to video, audio, and text well enough to win the Algonauts 2025 challenge. We ask w…
A${}^2$BM: Alignment-Aware Bridge Matching for Image-to-Image Translation
Paired image-to-image translation underpins a wide range of computer vision tasks, including image editing, sensor translation, and domain…
Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm
Mechanistic interpretability has made significant strides in understanding neural network representations, with sparse dictionary learning…
Efficient EEG Seizure Detection Using INT8 Quantization, Channel Pruning, and Spiking Neural Networks
Continuous EEG monitoring for epilepsy is constrained by the limited power and memory budgets of wearable and implantable devices. Deep neu…
Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation
Medical Vision-Language Models (Med-VLMs) require reliable reasoning from fine-grained visual evidence, yet existing models can produce pla…
LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models
Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding. However, scaling dense or sparse Mixture-of-Expert…
DAUPNet: Domain-Aware Uncertainty Modeling for Reliable Prototype Discrimination in Cross-Domain Few-Shot Semantic Segmentation
Cross-domain few-shot semantic segmentation (CD-FSS) has predominantly been formulated as learning domain-invariant representations or impr…
Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs
Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predicti…
Monte Carlo Dropout Uncertainty and Entropy-Thresholded Selective Prediction for Architecture-Agnostic Brain Tumor MRI Triage
Deep networks now subtype brain tumors on MRI about as well as specialist readers, yet accuracy is not what keeps them out of the clinic. W…
GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs
Micro-gesture recognition demands the detection of fleeting, spatially localized movements that are frequently overwhelmed by dominant stat…
DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification
This article presents DMFNet, a dual-backbone multiscale feature fusion framework with residual feature propagation and spatial attention f…
AEVAL: From Anecdotal to Deterministic Testing for Agentic Skill Workflows
Modern agentic systems increasingly rely on skills: installable packages of natural language and code that teach an LLM agent to perform a…
Boundary-Seeking GAN-Augmented TabTransformer for Adversarially Robust Intrusion Detection
Machine learning-based intrusion detection systems (IDSs) often suffer from class imbalance and vulnerability to adversarial attacks, leadi…
Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition
Sensor-based human activity recognition (HAR) has achieved significant progressed in fully supervised learning settings. However, these sup…
Privacy-Aware Synthetic Video Benchmarking and Relational Evaluation for Worker-Under-Suspended-Load Detection
Publicly shareable construction-video benchmarks remain scarce, especially for safety-critical hazards that are rare, dangerous to stage, a…
Clarify Before Executing: A Self-Evolving Agent for Resolving Intent Asymmetry in 3D Tool Orchestration
A fundamental intent asymmetry plagues modern 3D asset creation: while state-of-the-art 3D toolchains demand precise, executable parameters…
A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand Forecasting
Retail demand forecasting remains difficult when demand shifts faster than static forecasting models can be retrained, especially in early…
PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation
Recent advances in physics-grounded video generation leverage physics simulation as a physical prior to guide video synthesis toward physic…
Reliable Remediation Impact Prediction for Black-Box Security Ratings
Security rating platforms summarize externally observable cyber exposure and are expected to help organizations prioritize remediation. A p…
A Quantum-Classical Hybrid Framework for Multivariate Time-Series Forecasting Complexity-Fidelity Trade-offs and Limitations
This paper presents a unified quantum-classical hybrid framework for multi-horizon time-series forecasting, introducing two model variants…
PRISM: Multimodal Terrain Mapping for Rover Navigation in Unstructured Environments
Robotic navigation in unstructured environments requires robust situational awareness to safely traverse hazards such as steep slopes and r…
AoA: Theorem Proving Agent over Abstract Syntax Tree of Redesigned Language
Interactive theorem proving (ITP) underpins program verification and formalized mathematics, but its manual effort limits scalability. LLM-…
Fantastic Adaptive Taxonomies and How to Use Them
An agent system's execution traces record how it fails, and procedures that improve such a system without changing model weights (trajector…
Automated Hardware Validation Test Plan Generation for Large Scale AI Datacenter Platforms Using a Generative AI Multi-Agents Architecture
Large-scale AI datacenter platforms comprise thousands of heterogeneous hardware components whose validation requires comprehensive fault i…
Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models
Unified Multimodal Large Language Models (MLLMs) offer a promising paradigm for unifying visual understanding and generation, yet they stil…
Signal-based Model Access Risk Analysis for AI System Operations Security
Artificial intelligence (AI) systems are now ubiquitous across domains such as security, finance, healthcare, consumer technology, and larg…
Back to the museum: Investigation of the acceptance of Android Andrea with and without emotion simulation in a museum
For a second time, the android robot Andrea was set up at a public museum in Germany for six consecutive days to have conversations with vi…
Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent
Interpretability methods for neural network activations span a wide cost spectrum, from cheap, training-free techniques (such as linear pro…
Committed Before Reasoning: Behavioral Reproduction and Preliminary Activation-Level Evidence of Answer Pre-Commitment in an Open-Weight LLM
Chat models sometimes commit to an answer and then produce reasoning that justifies it rather than deriving it -- even when the answer cont…
When to Use Which? Benchmarking Optimisers for Configurable Systems under Varying Budgets
Software configuration tuning is crucial for optimising system performance, and various optimisers have emerged over the last decade. Yet,…
K-IPO: Kendall-constrained Importance Preserving Oversampling for Imbalanced Tabular Data
Oversampling is widely used to address class imbalance in tabular classification, but existing methods can distort the feature importance r…
How Formerly Incarcerated People Envision Technologies for Prison Parole
AI-driven algorithms and automated tools are increasingly embedded in the correctional landscape, shaping parole eligibility,release decisi…
Geometry-Enhanced Portion Estimation for Multimodal LLMs
Image-based dietary assessment promises to replace costly, bias-prone manual recalls, but portion estimation remains a major blocker. Multi…
Building2Building: A Large Scale Benchmark for Generalizable Real-World Reinforcement Learning
Reinforcement learning (RL) has achieved strong results in control, yet learned policies remain brittle to changes in dynamics, action spac…
Capacity and Redundancy Trade-offs in Multi-Task Learning
In multi-task learning (MTL) negative transfer is often considered as an optimization artifact, but it can also be viewed as a consequence…
Mitigating Compiler Fusion-Induced Power Bursts in Mobile NPU Inference as the Battery Depletes
Mobile devices increasingly rely on real-time NPU inference for camera and perception workloads. Under low-voltage conditions, however, a s…
ReqGenX: An Empirical Study of Atomic Decomposition, Artifact Regeneration, and Reconstruction for Legacy SRS Documents
Background: Evaluating automated Software Requirements Specification (SRS) generation is challenging because few datasets provide fine-grai…
Autonomous VR-Based Risk Detection for Situational Awareness in Dangerous Settings
In high-risk environments such as disaster response, situational awareness depends not only on detecting hazards but also on communicating…
Learning from World Feedback: Why Model Uncertainty Fails as a Risk Signal in Model-Based RL
The RLxF programme argues that learning signals should come from world feedback rather than from internal model proxies. We instantiate thi…
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that…
Privacy Cost as Equity Input: A Group Fairness Criterion for Differentially Private Machine Learning
Differential privacy (DP) is increasingly deployed to limit membership inference risk in machine-learning systems. Prior work has shown tha…
CLOSER-Bench: Evaluating Budgeted Cross-Stage Design Closure for Hardware Agents
Hardware engineering exposes coding agents to a form of long-horizon work that is difficult to capture with pass-at-k: progress is continuo…
TellTale: Blending Multi-Instance LoRA Text Encoders and a Zero-Shot LLM Judge for Ambivalence/Hesitancy Recognition in Videos
We present TellTale, a text-only approach to ambivalence/hesitancy (A/H) recognition in interview videos, evaluated on the BAH dataset as p…
Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
Large language models now translate natural-language descriptions of decision problems into solver-ready optimization models, but they fail…
Position: Explanation Stability Is a Property of the Model Method Pair, Not the Model
This position paper argues that claims about explanation stability are scientifically invalid without cross method validation. Just as stat…
How Do You Choose Your AI Component? An Interview Study of Secure AI Integration in Practice
The increasing adoption of Large Language Models (LLMs) as AI components in modern software systems introduces distinct security risks to t…
Building a Neural Network from Scratch: Implementation, Evaluation, and Optimization
The widespread adoption of high-level deep learning libraries, while accelerating model development, has increasingly abstracted away the i…
OFD-Net: Teacher-Free Reliable Semi-supervised Medical Image Segmentation with Orthogonal Feature Disentanglement Net of Foreground-Background
Semi-supervised learning (SSL) is an effective solution for medical image segmentation with limited annotations. Existing SSL methods mainl…
A Causal Markov Condition for Value
This paper proposes a causal independence principle for value -- the value Causal Markov Condition (v-CMC) -- and develops the conceptual a…
RealDESED: A Real-World Domestic Sound Event Detection Benchmark
This paper presents RealDESED, a real-world domestic sound event detection (SED) benchmark comprising 5,710 audio recordings collected by 6…
Spatiotemporal Facial Action Unit Detection using Twin Cycle Autoencoders for Driver Monitoring
Driver monitoring systems (DMS) increasingly rely on facial cues to infer drowsiness, distraction, and cognitive load in real time. Facial…
JOR-Bench: Japanese Operations Research Benchmarks for Large Language Models
We present JOR-Bench, a collection of five Japanese-language benchmarks for evaluating the ability of large language models (LLMs) to formu…
Explainable Lightweight Compact Deep Models for Speech Emotion Recognition
Speech Emotion Recognition (SER) is an important component in a wide range of human-centered applications, including healthcare, customer s…
First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers
Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbati…
Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration
Long-term memory is becoming a core component of LLM agents, but most memory benchmarks evaluate conversations or compact summaries, while…
Principled Direction-Free Intrinsic Motivation through Model-Free Epistemic Free-Energy Estimators
Across environments with mixed sources of uncertainty, unsupervised reinforcement learning requires intrinsic motivation that does not prec…
Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models
End-to-end speech language models increasingly represent user speech with speech tokens rather than relying exclusively on cascaded ASR--LL…
Trace-Based On-Policy Distillation for Masked Diffusion Language Models
Diffusion large language models (dLLMs) are a promising alternative to autoregressive generation. However, reasoning-oriented post-training…
A Deep Reinforcement Learning Algorithm for the Vehicle Routing Problem with Stochastic Demands and Outsourcing
We introduce the vehicle routing problem with stochastic demands and outsourcing options (VRP-SDO), in which a logistics service provider p…
Certified-Gap Dual-Price Policies for Real-Time Truckload Bid Acceptance with Relocating, Clock-Constrained Resources
A truckload carrier must accept or reject each load tender within seconds. The decision depends on fleet state, hours-of-service (HOS) cloc…
A Method for Learning Value Systems in Generative AI
Value-aware AI systems require explicit computational representations of human values (groundings) and their aggregation into value systems…
PREFAIL: Identifying Precursors to Failures in Robotic Lift-and-Place Tasks to Improve Task Execution Performance
Non-prehensile manipulation enables flexible material handling with part carriers, but friction-based support makes high-speed motions fail…
A Multi-Agent System for 5G Throughput Prediction in Multi-Operator Urban Environments
Throughput prediction is foundational for artificial intelligence-driven 6G resource orchestration. Conventional monolithic machine learnin…
Optimizing Clinical Trial Protocols Using EHR-Derived Heterogeneous Treatment Effects
Traditional randomized trials often obscure clinically meaningful heterogeneity in treatment response by focusing on average effects. Lever…
Pediatric Bone Age Prediction Using Deep Learning
Pediatric bone age prediction is a crucial task in clinical practice that can help diagnose endocrine disorders and provide insight into a…
What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning
End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor observations directly to observed…
Investigation of Polycystic Ovary Syndrome (PCOS) Diagnosis Using Machine Learning Approaches
Polycystic Ovarian Syndrome (PCOS) is a widespread hormone problem for women of childbearing age. Women with PCOS may not ovulate; they mig…
CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation
On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches suffer three compound…
TurboVec: A Case Study in Cost-Efficient Private Retrieval for Enterprise RAG via Codebook-Oblivious Quantization
Retrieval-Augmented Generation (RAG) systems increasingly power enterprise LLM applications, yet the vector retrieval layer introduces two…
Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries
Introduction. Clinical and Translational Science Award (CTSA) programs must document their scholars' research impact, but assembling each s…
Automated Cardiac Adipose Tissue Segmentation in Computed Tomography: A Literature Review
This review provides an overview of recent advancements in automated segmentation methods on Computed Tomography (CT) for two types of card…
Counterfactual Shapley Credit Assignment
The Credit Assignment Problem (CAP) is fundamental to developing efficient and explainable Reinforcement Learning (RL) agents. Existing fra…
Scalable Causal Imitation Learning
Imitation learning enables learning a policy in an unknown environment with a latent reward signal using expert demonstrations, but it stru…
Alignment of a Total Automation Economy
We consider economic theory from the perspective of a total automation economy, one with no human involvement in production either in manuf…
WHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture
As scalability becomes increasingly important in recommendation modeling, recent architectures have advanced the modeling of two broad sour…
Where Does Agent Reliability Come From? A Cross-Benchmark Decomposition of Verification Loops, Specialist Models, and Scaffolding in a Production Enterprise Agent
Multi-step enterprise agent tasks fail in a characteristic way: single-pass inference has no checkpoint between deciding an answer and comm…
Solver-Hard Is Not Model-Hard: A Hardness-Controlled Diagnostic for LLM Constraint Reasoning
LLM constraint reasoners are often evaluated near the random-SAT phase transition, confounding density and solver hardness. We test instanc…
EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding
GUI agents must reason about how actions transform interface states, but end-to-end success rates entangle this ability with perception, gr…
ThAME: 3D Memory-Enabled Heterogeneous Accelerator for LLM Mixture of Experts
Mixture of Experts (MoE) architectures have emerged as a dominant paradigm for scaling Large Language Models (LLMs). However, MoE inference…
ALLUDE: A Unified Evaluation System for Configurable Attacks in Differentiable Environments
Adversarial attacks against vision models like object detectors are often evaluated under limited conditions, leaving their performance und…
DepthART: Scaling Foundation Monocular Depth to Tiny Models
Recent geometric foundation models (e.g., Metric3D, Depth Anything and UniDepth) have substantially improved monocular depth estimation (MD…
Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer
An AI research agent can improve the score it sees without finding a modelling change that works on new materials. We ask a stricter questi…
Teach it to stop, not just to click
Agentic computer-use RL is reported in single runs, and those numbers mislead. Using verifier-guided repair of a 35B computer-use agent (CU…
Noise-Robust Box-Supervised Infrared Small Target Detection via Physics-Inspired Soft Label Optimization
Infrared small target detection (IRSTD) commonly relies on pixel-level mask supervision. Such annotations, however, are costly and inherent…
VLA-ReID: Video-Level Association for Re-Identification in Multi-Object Tracking with Highly Similar Objects
Multi-object tracking (MOT) aims to localize multiple objects in videos while preserving their identities over time. Long-term identity pre…
DADIR: Density-Aware Data-level Imbalanced Regression Framework
Imbalanced learning addresses predictive modeling problems with underrepresented regions of the data distribution. Although widely studied…
Talaria: Session-Aware Serverless Serving of Hundred-Billion-Parameter LLMs
Serverless multi-model LLM systems multiplex popularity-skewed model catalogs over shared GPU pools, yet typically schedule each request in…
Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat
Human survey respondents exhibit question-order effects that satisfy the QQ (quantum question) equality, an a priori, parameter-free predic…
Specifying the Delegated-Autonomy Boundary: Requirements Engineering for Agentic AI
Agentic AI systems do not just predict or recommend; they plan, maintain state, and act in external environments with varying degrees of au…
A RFID Based Campus Wide Payment System
This work titled "RFID Based Campuswide Payment System" introduces an innovative cashless payment solution for educational institutions. It…
A Large-Scale Measurement of AI Bill of Materials Completeness in Hugging Face Models
Pretrained machine learning (ML) models help developers build ML-intensive software systems without training models from scratch. However,…
Distilled Reinforcement Learning for LLM Post-training
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow tw…
Asynchronous Multimodal Diffusion Policy Composition via Latency-Aware Guidance Fusion
Diffusion policies have shown strong potential for robotic imitation learning, and recent extensions incorporate additional modalities to i…
Debate-on-Graph: Reliable and Adaptive Reasoning of Large Language Model on Uncertain Knowledge Graph
Large language models (LLMs) have demonstrated remarkable capabilities in natural language processing. However, LLMs often suffer from hall…
Between Safe Boundaries: Exploiting Temporal Consistency for Jailbreaking Text-To-Video Generation Models
Recently, text-to-video (T2V) models have been widely deployed, sparking growing concerns over their robustness against jailbreak attacks.…
AIGB-R1: Self-Evolving Generative Auto-Bidding via Hierarchical Planner-Executor Optimization
Auto-bidding plays an essential role in online advertising, automatically adjusting bids for advertisers to optimize their commercial goals…
SAGA: Synthetic Agentic Graph Architecture for Temporal Benchmark Generation
High quality temporal graph benchmarks with rich semantics and ground-truth anomaly labels are essential for training graph neural networks…
Lookahead Branching for Neural Network Verification
In this work, we investigate the effect of lookahead branching strategies in neural network verification. We present a general recipe to in…
WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning
Long-horizon rollout generation has become the dominant systems bottleneck in agentic reinforcement learning (RL). As agents interact with…
The Optimization Trilemma: Efficiency, Comfort and Fairness in Decentralized Multi-agent Coordination
The problem of fair multi-agent coordination in decentralized settings is one of the most pressing challenges for building efficient collab…
TAPAS: Throughput-adaptive Perception for Autonomous Systems
Autonomous systems rely on a perception module to navigate through dynamic environments. In real-world scenarios, the perception module's t…
STAR: Skeletal Token Alignment and Rearrangement for Interaction Recognition
Understanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision. While most existing metho…
Mathematical Discovery in the Wild: AI-Guided Proofs in Banach Space Theory
We investigate the capacity of current language models to contribute to mathematical research. In Banach space theory, AI systems generated…
A Phased Development Framework Enabling Islanded Operation of Sustainable AI Data Centers With Onsite Grid-Following and Grid-Forming Energy Architectures
As hyperscale and colocation AI data centers continue to expand, the electric grid is increasingly required to support large, concentrated…
CoEvoP&R: Co-Evolving Placement Objectives with Routing Feedback via Large Language Models
Analytical placers rely on differentiable objective functions to guide placement, typically combining intermediate surrogate metrics such a…
Kernelized Linear Attention: Breaking the Capacity Wall with Symmetric Cones
Linear attention promises constant-time recurrent inference but degrades sharply on associative recall. We formulate attention recall as a…
HyCoRec: Hypergraph-Enhanced Multi-Preference Learning for Alleviating Matthew Effect in Conversational Recommendation
The Matthew effect is a notorious issue in Recommender Systems (RSs), \emph{i.e.}, the rich get richer and the poor get poorer, wherein pop…
Multilingual Sentence Embeddings for Linguistic-Integrated Reliability Audit
Multilingual assessment systems commonly rely on translation for scoring and quality-control processes. We evaluate whether multilingual se…
SALT: Salience-Aware Lexical Trie for Long-Context Compression
As large language models (LLMs) process increasingly longer prompts, computation and KV-cache memory costs have emerged as major bottleneck…
DecoyFace: Beyond Obfuscation via Controllable and Imperceptible Identity Misdirection for Privacy-Preserving Face Recognition
Split face recognition reduces client-side computation but exposes intermediate features to feature inversion attacks and unauthorized anal…
Retrieval-Augmented Interpretable Learning: Towards Task-Specific Zero-Shot Models in Healthcare
We introduce Retrieval-Augmented Interpretable Learning (RAIL), a probabilistic meta-learning framework for zero-shot generation of task-sp…
After the Euclidean Highway: Hyperbolic Expert AI as the Next Innovation
Expert domains are trees; the Euclidean transformer is not, diluting parent-child structure exponentially at depth. The hyperbolic turn lef…
One-step lowest-variance selection in a Gaussian random-field model motivated by masked diffusion: Total correlation and a square root collision threshold
Motivated by confidence-guided parallel unmasking in masked discrete diffusion, we study a single selection step in a stylized Gaussian ran…
Thinking in Video: Can Video Generators Really Reason About the Real World?
Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative mode…
Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration
Test-time collaboration, including self-consistency, best-of-N selection, critic models, and verifier pipelines, is often credited with bro…
CommitLLM: A Fine-Tuned Pipeline for Git Commit Message Generation
Developers frequently write uninformative git commit messages such as "fix" or "update stuff", degrading the value of version-control histo…
Human-in-the-Loop User Feedback Affects Perceived Accuracy and Trust, but Task Subjectivity Matters
While ML can produce complex models beyond those that a human could produce manually, incorporating human input can often improve performan…
Hierarchy-Aware and Anatomy-Guided Learning for Lung Ultrasound Video Classification
Lung ultrasound (LUS) is a bedside tool for assessing pulmonary edema in patients at risk due to heart failure or impaired kidney function.…
COLIP-2: Olfaction-Vision-Language Embeddings
The Contrastive Olfaction-Language-Image Pre-training 2 (COLIP-2) model is a multimodal embeddings space that places olfaction as a first-c…
CoCurve: Cross-Module Co-Pruning Curvature for Training-Free Structured LLM Pruning
Structured pruning compresses large language models (LLMs) by removing whole computational units, such as attention heads and feed-forward…
Predictive Training with Latent Imagination for Visual Quadruped Navigation
Reinforcement-learning navigation policies for legged robots select actions reactively from current observations and short-term memory, wit…
Detection, Attribution, Narration: An End-to-End Pipeline for Explainable Money Mule Identification
Money mule accounts are critical facilitators of financial fraud, yet detecting them at scale remains challenging due to the heterogeneous…
Trustworthy Protein-Ligand Binding Affinity Prediction via Reliability-Aware Multi-Engine Fusion
Accurate protein-ligand binding affinity prediction is central to computational drug discovery, yet modern docking engines frequently disag…
Coarse-to-fine Framework for Generative MEF via Implicit Neural Representation
Multi-exposure fusion (MEF) expands the luminance range beyond what a single exposure can capture. Combining images taken at different expo…
Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture
Individuals with dysarthria face significant challenges in professional speaking scenarios such as conferences, presentations, and meetings…
TypiCore: A Hybrid Active Query Strategy for Class-Incremental Learning on Time Series
Time series data play a pivotal role across numerous domains, including healthcare and manufacturing. In real-world environments, models mu…
Selectivity Matters: Source Node Influence Pruning for Unsupervised Graph Domain Adaptation
Unsupervised Graph Domain Adaptation (UGDA) aims to facilitate knowledge transfer from a labeled source graph to an unlabeled target graph…
Beyond Objective Expressivity: Geometry Preservation in Multimodal Contrastive Learning
Contrastive learning is increasingly moving toward settings with three or more modalities instead of image-text pairs. Yet, extending model…
Uncovering Latent Reasoning Strategies in Language Models
A language model $p_\theta(y \mid x)$ trained on reasoning tasks learns to solve problems via multiple distinct strategies, yet these strat…
Integrating High-Level Requirements to Low-Level Tests with Machine-Readable V&V Specifications
Modern software teams have mature tools for low-level testing, such as pytest, JUnit, and Jest, which make it inexpensive to write unit tes…
Lifelong Multi-Subsystem Pickup and Delivery with Buffer-Limited Handover Stations
Coordinating payload transfers between subsystems is a critical challenge in lifelong Multi-Agent Pickup and Delivery (MAPD). We study syst…
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference
4-bit quantization enables efficient LLM inference, but suffers from significant accuracy degradation due to outliers. Prior work addresses…
Mobile Network Control with a World Model
The increasing complexity of mobile networks necessitates intelligent and dynamic control strategies for efficient, energy-conserving manag…
DA-Fusion: Deformable Attention-Based RGB-D Fusion Transformer for Unseen Object Instance Segmentation
In logistics automation, precise segmentation of unseen objects is crucial for efficient robotic manipulation in cluttered environments. Ta…
Seg2Grasp: A Robust Modular Suction Grasping in Bin Picking
Current bin picking methods that rely heavily on end-to-end learning often falter when confronted with unfamiliar or complex objects in uns…
Generalize and Guide: Decomposing Rewards for Few-Shot Inverse Reinforcement Learning
Inverse reinforcement learning (IRL) provides a powerful framework for learning from demonstrations. However, real-world tasks often exhibi…
Time-Frequency Consistency Learning for Robust Speech Deepfake Detection
Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to…
Autonomous Discovery of Wireless Communications Algorithms
Large language model (LLM)-driven evolutionary search is an emerging algorithm-discovery paradigm that has already produced novel results i…
FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches
We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, f…
Measuring Monosemanticity in Sparse Autoencoders via Latent Activation Coherence
Within Explainable Artificial Intelligence, mechanistic interpretability uses Sparse Autoencoders (SAEs) to extract more interpretable feat…
Persona-as-Configuration: Generative Stakeholder Reporting for Agricultural Floods
Cyber-physical systems built on deterministic edge inference, such as on-vehicle flood detection for agricultural fields, produce structure…
CDIS: Cross-Dimensional Class-Agnostic 3D Instance Segmentation via 2D Mask Tracking and 3D-2D Projection Merging
Class-agnostic 3D instance segmentation is critical for robotic systems operating in unknown environments, enabling perception of previousl…
ETAS: An Effect-Typed Language for Agent Systems
ETAS is a programming language for agent systems that treats model-backed agents, tool calls, prompts, typed memory, human approvals, polic…
BrainNext: A General-Purpose Self-Supervised Foundation Model for Brain MRI Analysis
Foundation models pretrained using self-supervised learning have transformed computer vision by learning transferable representations from…
Feature Attribution-Based Explainability Analysis of Deep Learning Models in Predictive Process Monitoring
Predictive process monitoring supports the optimization and control of operational business processes by forecasting the future state or ou…
Reasoning as a Double-Edged Sword: Architecture and Cross-Stage Robustness in Vision-Language-Action Models
Does adding a reasoning step make a Vision-Language-Action (VLA) model more robust to perturbation? Intuitively, a policy that reasons befo…
Medical Imaging Fusing Vision Transformer: Laryngeal Cancer Screening with Explanation
Early and timely screening of laryngeal cancer is crucial for improving clinical outcomes. In recent years, NBI endoscopy has become a stan…
ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a…
Vis2Reg: Visibility-Aware Landmark-Free Geometric 3D--2D Registration for Liver Laparoscopy
Accurate 3D--2D liver registration, which aligns preoperative 3D models to partial, view-dependent intraoperative surface observations, is…
Phasor Attention: Mean Root Square Normalization for Phase Manifold Preservation
While Root Mean Square Normalization has become the de facto standard for accelerating modern sequence models, its reliance on the quadrati…
I wanted it to feel more personal: Customization of social AI as AI individualism in practice
Despite the growing availability of customizable social artificial intelligence (AI), such as ChatGPT, Grok, and Character.ai, we know litt…
Measuring and Improving Complex-Atomic Answer Consistency in Endoscopic VQA
Endoscopic visual question answering (VQA) increasingly asks complex questions that combine several endoscopic answer components rather tha…
CaT-GS: Efficient 3DGS Rendering for Large Scale Scenes via Inter-frame Caching and Tile Scheduling
Recent breakthroughs in 3D Gaussian Splatting (3DGS) have advanced neural rendering with high fidelity and speed. However, its performance…
ConceptTree: Bringing Semantic Transparency to Black-Box Decision Making for Robotic Manipulation
Establishing interpretable decision-making processes in long-horizon robotic manipulation is critical for enabling reliable human oversight…
Zero Hallucination, by Construction: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI
Enterprises will not deploy AI agents they cannot trust, and the most-cited reason for distrust is hallucination: confident, fluent output…
Chemical filters for ultra-high-throughput materials screening and generation
Generative artificial intelligence is rapidly transforming materials design by enabling de novo exploration of immense chemical spaces. Yet…
DeLIVeR: Decomposed Learning for Information-grounded Veracity Recognition via Reinforced Knowledge Graph Exploration
Automated fact-checking remains a challenge for Large Language Models (LLMs) due to "query brittleness" in traditional retrieval systems. W…
The Aura in the Machine: Genealogy and the Status of the Work of Art in the Generative Era
This paper frames Generative Artificial Intelligence (AI) not as an unprecedented technological rupture, but as an industrial-scale manifes…
The Art of Not Forgetting
We introduce CMP (Cognitive Memory Primitive), an architecture that represents inputs as sparse relational codes, stores them in a two-tier…
A Geometric Perspective on Stabilizing Value Conflict Resolution
Large Language Models (LLMs) often struggle to navigate value conflicts when trained with the compressed scalar rewards of Reinforcement Le…
RT-SHCUA: Real-Time Self-Hosted Computer-Use Agent for UAV Control
Natural-language control offers a promising interface for unmanned aerial vehicles (UAVs), but directly applying self-hosted computer-use a…
Topological Signatures of Context-Level Reliability in TabPFN
TabPFN is a transformer-based foundation model for tabular prediction that performs inference without task-specific training by conditionin…
Harness Engineering for LLM-Driven GPU Kernel Generation
Large language models (LLMs) can assist GPU kernel generation, but their practical effectiveness depends on whether generated code can be r…
Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?
Self-hosted AI agents read and write their own memory and configuration files to function. An agent may get compromised via corruption of i…
HAS: Highlight-guided Attention Steering for Multimodal LLM Video Summarization
Video understanding has become more and more important with the growth of Artificial Intelligence (AI) for video generation. Recently, Mult…
MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models
Large language models achieve strong reasoning performance, but often at prohibitive training cost - a challenge that is especially acute f…
Natural Language Access to Domain-Specific Metadata: A Reusable Framework for LLM Query Generation
Researchers need to answer ad-hoc questions about the contents of domain-specific archives but often lack the expertise to write structured…
Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation
End-to-end vision-language navigation (VLN) with causal vision-language models can map instructions and egocentric observations directly to…
Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security
LLM-based agents process external content, exposing them to prompt injection and multi-turn manipulation. Most safety benchmarks evaluate d…
Autoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation Data
Coding agents can now be left alone to improve software against a score. In this pattern--recently popularized as "autoresearch"--the agent…
Human Grounded Evaluation of Large Language Models for Optical Network Automation
Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substant…
SGN: A Similarity-based Generative Network for Data Generation under Distribution Shift
Generative models trained on a source domain often produce samples that are poorly aligned with shifted target domains, limiting their effe…
Generalised Bellman recurrence and three dualities in sequential decision-making
What gives the Bellman equation its form? We show that the recursive properties of optimal value functions follow from three conditions: th…
Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection
Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its ass…
SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs
Large Language Models (LLMs) have demonstrated remarkable capabilities across a range of Natural Language Processing (NLP) tasks, but their…
Enhancing Rubric-based RL via Self-Distillation
Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized limitation of rubric-based RL is limi…
How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?
Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled fe…
O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for mode…
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints
Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to gener…
LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications
Large language models (LLMs) and agentic AI systems have evolved from natural language tasks to using external tools to plan, retrieve, and…
Differentiable Logic Gate Networks for Low-Latency EEG Classification on Edge Devices
Real-time EEG classification on edge devices is bottlenecked by the floating-point arithmetic of conventional neural networks. We investiga…
TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory Minimization
Coding agents are increasingly used to accelerate code generation in many downstream tasks, such as fixing bugs, building applications, and…
OR Else: A Differentiable Trust Region for Policy Optimization
PPO and the GRPO baseline studied here use clipped surrogate objectives whose favorable-direction saturation introduces an abrupt change in…
A Continual Validation, Updating, and Decision-Making Framework for Self-Adaptive Digital Twins via Robust Model Predictive Control: A Case Study in Additive Manufacturing
Digital Twins rely on surrogate models to mirror physical systems in real time, yet these models can degrade as operating conditions evolve…
Learning Adaptive Safety Margins for Visual Navigation
Robots in cluttered indoor spaces often fail not because they cannot generate collision-free paths, but because a fixed safety margin is mi…
GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis
Foundation models have emerged as a driving force in computational pathology, with the potential to transform cancer diagnosis, prognosis,…
Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs
Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tamper…
Automated Discovery Has No Universally Superior Harness
Autonomous discovery systems such as OpenEvolve and TTT-Discover are often used as general-purpose harnesses. However, in practice these ar…
Comprehend, Divide, and Conquer: Feature Subspace Exploration via Multi-Agent Hierarchical Reinforcement Learning
Feature selection aims to preprocess the target dataset, find an optimal and most streamlined feature subset, and enhance the downstream ma…
Enhancing LLMs' Clinical Reasoning with Real-World Data from a Nationwide Sepsis Registry
Although large language models (LLMs) have demonstrated impressive reasoning capabilities across general domains, their effectiveness in re…
LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants
Vision-language models (VLMs) are facing the challenges of understanding and following multimodal assembly instructions, particularly when…
MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs
Large Language Models (LLMs) suffer from hallucinations due to their static parametric knowledge. Retrieval-Augmented Generation (RAG) and…
AI sustains higher strategic tension than humans in chess
Strategic decision-making requires balancing immediate opportunities against long-term objectives: a tension fundamental to competitive env…
Benchmarking Agentic Newswriting via Journalistic Workflows
Recent advances in autonomous digital agents from industry (e.g., Manus AI and Gemini's research mode) highlight their potential for struct…
SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMs
Large language models (LLMs) exhibit strong general reasoning, yet the community lacks controllable, scalable, and verifiable tools to anal…
Artificially intelligent agents in the social and behavioral sciences: A history and outlook
We review the historical development and current trends of artificially intelligent agents (agentic AI) in the social and behavioral scienc…
Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions
Despite rapid progress in artificial intelligence, current systems struggle with the interconnected challenges that define real-world decis…
Parallel Decoder Transformer: Planner-Conditioned Latent Coordination for Model-Intrinsic Parallel Generation
Autoregressive language models expose one causal token frontier, even when the requested document contains sections that could be developed…
Towards AI epidemiology: a measurement standardisation framework for prospective risk detection
This paper proposes a measurement standardisation framework that compresses expert-AI interactions into structured, comparable fields for p…
Multi-modal cross-domain mixed fusion model with dual disentanglement for fault diagnosis under unseen working conditions
Intelligent fault diagnosis has become an indispensable technique for ensuring machinery reliability. However, existing methods suffer sign…
From Classical to Quantum Reinforcement Learning and Its Applications in Quantum Control: A Beginner's Tutorial
This tutorial is designed to make reinforcement learning (RL) more accessible to undergraduate students by offering clear, example-driven e…
Data-Efficient Curation for Multimodal Reasoning under Fixed Training Protocols
We study data curation for multimodal reasoning in a fixed-protocol fine-tuning regime, where the base model, optimizer, training schedule,…
NEMO: Execution-Aware Optimization Modeling via Autonomous Coding Agents
We present NEMO, a system that translates Natural-language descriptions of decision problems into formal Executable Mathematical Optimizati…
Lyapunov Stability-Aware Stackelberg Game for Low-Altitude Economy: A Control-Oriented Pruning-Based DRL Approach
With the rapid expansion of the low-altitude economy, Unmanned Aerial Vehicles (UAVs) serve as pivotal aerial base stations supporting dive…
Arbor: A Framework for Reliable Navigation of Critical Conversation Flows
Large language models struggle to maintain strict adherence to structured workflows in high-stakes domains such as healthcare triage. Monol…
SCA: Segment-Wise CoT Compression with Answer Alignment
Chain-of-thought (CoT) reasoning improves problem solving, but long think traces increase inference cost. Existing CoT compression methods…
Content Creation with Spillovers: An Incentive Design Approach
The rise of AI amplifies the economic phenomenon of \emph{positive spillovers}: when creators contribute content that can be reused and ada…
Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain
Multimodal large language models (MLLMs) have shown strong potential for medical Visual Question Answering (VQA), yet they remain prone to…
CARV: A Diagnostic Benchmark for Compositional Analogical Reasoning in Multimodal LLMs
Analogical reasoning tests a fundamental aspect of human cognition: mapping the relation from one pair of objects to another. Existing eval…
Agent psychometrics: Task-level performance prediction in agentic coding benchmarks
As the focus in LLM-based coding shifts from static single-step code generation to multi-step agentic interaction with tools and environmen…
From Multi-Agent to Single-Agent: When Is Skill Distillation Beneficial?
Multi-agent systems (MAS) tackle complex tasks by distributing expertise, though this often comes at the cost of heavy coordination overhea…
When Direct Prediction Fails: Evidence from LLM-Based Misinformation Risk Evaluation
LLMs make it increasingly easy to generate deceptive content at scale, creating a need for scalable misinformation risk evaluation based on…
Information-Theoretic Measures in AI: A Practical Decision Framework
Information-theoretic (IT) measures are ubiquitous in artificial intelligence: entropy drives decision-tree splits and uncertainty quantifi…
RADD: Retrieval-Augmented Discrete Diffusion for Multi-Modal Knowledge Graph Completion
Most multi-modal knowledge graph completion (MMKGC) models use one embedding scorer to conduct both retrieval over the full entity set and…
Adaptive Multi-Round Allocation with Stochastic Arrivals
We study a sequential resource allocation problem motivated by adaptive network recruitment, in which a limited budget of identical resourc…
AI for Auto-Research: Roadmap & User Guide
AI-assisted research is crossing a threshold: fully automated systems can now generate research papers for as little as $15, while long-hor…
Scientific reasoning does not reliably translate into scientific forecasting in frontier AI
AI systems are increasingly used to support forward-looking scientific judgment, but it remains unclear whether they can form reliable expe…
レビンツリー検索の再根付のための構造に起因する情報
ポリシーを使用して検索をガイドするサブゴールベースのポリシー ツリー検索は、複雑な単一エージェントの決定論的問題には効果的ですが、多くの場合、明示的なサブゴールの生成に依存するため、大幅なオーバーヘッドが発生し、スケーラビリティが妨げられる可能性があります。この論文では、最近導入された $\sqrt{\text{LTS}}$ アルゴリズムを通じて学習された「rerooter」を使用することで、これらの制限を克服します。 rerooter は問題を暗黙的にソフト サブタスクに分解します。以前の研究では、与えられたリルータまたは手作りのリルータの正式な保証に焦点を当てていましたが、この研究では 3 つのリルータ設計を提案します。(i) グローバルな状態空間構造を活用するクラスタリング ベースのリルータ、(ii) 学習されたコスト To Go 推定を活用するヒューリスティック ベースのリルータ、および (iii) 両方の信号を組み合わせたハイブリッドです。私たちのフレームワークでは、生成されたサブゴールを明示的に再構築して推論する必要がなくなり、大幅に低い計算オーバーヘッドでスケーラブルな検索労力の割り当てが可能になります。経験的に、当社のリルートベースの方法は、サブゴールベースのポリシーツリー検索が失敗する複雑な環境にも拡張でき、テストされたドメインで最先端のオンライントレーニング効率を実現します。
原文 (English)
Structure-Induced Information for Rerooting Levin Tree Search
Subgoal-based policy tree search, which uses a policy to guide search, is effective for complex single-agent deterministic problems but often relies on explicit subgoal generation that can incur substantial overhead and hinders scalability. In this paper, we overcome these limitations by using a learned ``rerooter'' through the recently-introduced $\sqrt{\text{LTS}}$ algorithm. A rerooter implicitly decomposes the problem into soft subtasks. While previous work focused on the formal guarantees for given or handcrafted rerooters, in this work we propose three rerooter designs: (i) a clustering-based rerooter that exploits global state-space structure, (ii) a heuristic-based rerooter that leverages learned cost-to-go estimates, and (iii) a hybrid that combines both signals. Our framework avoids having to explicitly reconstruct and reason over generated subgoals, thereby enabling scalable allocation of search effort with significantly lower computational overhead. Empirically, our rerooting-based methods scale to complex environments where subgoal-based policy tree search fails, and achieve state-of-the-art online training efficiency on the domains tested.
レンズの選択: 文脈に依存した議論における戦略的視点の活性化
多くの場合、同じ議論を異なる外部レジームの下で評価する必要があります。政権に対して影響力を持つエージェントは、標準的な形式主義では直接把握できない戦略的手段を持っています。我々は、コンテキスト依存議論フレームワーク (CDAF) を導入します。これは、敗北関数がコンテキストごとにどの攻撃が成功するかを決定するという Dung の理論の拡張です。パースペクティブラベル付き特殊化は、関連性セット $\rho$ と優先度 $\pi$ から敗北関数を導出します。関連性セットはエージェントのアクション スペースです。小さな実際の例では、エージェントのターゲット引数は、すべての完全関連性の単射優先度の下では拒否されますが、VAF オーディエンスがミラーできないものの 1 つである部分的なアクティブ化の下では受け入れられます。対応する意思決定問題である ACTIVATION-MANIPULATION を定義し、ベースラインの複雑さの限界を記録します。狭い境界と複数エージェントのバリアントは未解決のままです。
原文 (English)
Choosing the Lens: Strategic Perspective Activation in Context-Dependent Argumentation
The same arguments often need to be evaluated under different external regimes. An agent with influence over the regime has a strategic lever that standard formalisms do not directly capture. We introduce context-dependent argumentation frameworks (CDAFs), an extension of Dung's theory in which a defeat function determines, per context, which attacks succeed. Blocked attacks are inverted rather than deleted, so extensions stay conflict-free with respect to the attack relation. A perspective-labeled specialisation derives the defeat function from a relevance set $\rho$ and a priority $\pi$. The relevance set is the agent's action space. In a small worked example, the agent's target argument is rejected under every full-relevance priority, yet accepted under a partial activation whose outcome no VAF audience can mirror. We define the corresponding decision problem, ACTIVATION-MANIPULATION, and record baseline complexity bounds. For grounded semantics with mandatory perspectives the problem is NP-complete, and the hardness comes from the activation choice itself.
事後ハイブリッド ベイジアン ビリーフを使用した正規化されたオフライン ポリシーの最適化
オフライン強化学習 (RL) は、事前に収集されたデータセットからポリシーを最適化することを目的としています。このパラダイムのボトルネックは、認識論的な不確実性を管理することです。これは、限られたデータ範囲 (サンプルレベル) と、有限データから遷移ダイナミクスを特定する際の曖昧さ (モデルレベル) から生じます。これらの不確実性を統一的に定量化するために、ダイナミクス モデルを確率変数として扱い、対応する信念を維持することによってベイジアン RL が提案されています。理論的には魅力的ですが、ベイジアン RL でのポリシーの最適化は、期待値を含む複合目標を解決する必要があるため、依然として計算上困難です。従来の方法は、計算のスケーラビリティが低い検索ベースの手法を採用するか、ベイジアン RL の適応性を犠牲にする制限的な事後仮定を課すかのいずれかでした。これらの制限に対処するために、私たちは事後ハイブリッド ベイジアン ビリーフ (PhyB) を提案します。これは、ダイナミクス モデルのサブセットにわたる凸の組み合わせとして期待値を再定式化します。理論的分析により、この近似によって引き起こされる客観的な不一致には限界があることが実証されています。 PhyB に基づいて、収束までの単調な改善に対するメトリクスに依存しない保証を提供する反復的な正則化ポリシー最適化アルゴリズムを開発します。実証結果は、PhyB がさまざまなベンチマークで最先端のパフォーマンスを達成することを示しています。
原文 (English)
Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief
Offline reinforcement learning (RL) aims to optimize policies from pre-collected datasets. A bottleneck of this paradigm is managing epistemic uncertainty, which arises from limited data coverage (sample-level) and the ambiguity in identifying transition dynamics from finite data (model-level). To provide a unified quantification of these uncertainties, Bayesian RL has been proposed by treating the dynamics model as a random variable and maintaining a corresponding belief. Despite its theoretical appeal, policy optimization in Bayesian RL remains computationally challenging as it requires solving composite objectives with expectations. Prior methods either employ search-based techniques with poor computational scalability or impose restrictive posterior assumptions that sacrifice the adaptability of Bayesian RL. To address these limitations, we propose Posterior Hybrid Bayesian Belief (PhyB), which reformulates the expectation as a convex combination over a subset of dynamics models. Theoretical analysis demonstrates that the objective discrepancy induced by this approximation remains bounded. Based on PhyB, we develop an iterative regularized policy optimization algorithm that provides metric-agnostic guarantees for monotonic improvement until convergence. Empirical results demonstrate that PhyB achieves state-of-the-art performance on various benchmarks.
ビデオから沿岸波のピーク周期を推定するための物理学に基づく時空間学習
沿岸の波のパラメータは、海岸工学、海岸線の保護、海洋危険評価、気候回復力のための海岸管理にとって重要です。ブイやレーダープラットフォームなどの従来の監視システムは正確な監視を提供しますが、設置とメンテナンスの費用が高額になり、カバー範囲が限られている可能性があります。ビデオを使用した受動的な海洋モニタリングは、深層学習を活用することで実現されていますが、多くの方法は物理的に解釈できず、海洋学としては実現可能ではなく、検証されていません。この研究では、パッシブ沿岸ビデオ ストリームから沿岸波のピーク周期を直接推定するための物理学に基づく深層時空間学習フレームワークが提案されています。このフレームワークは、自動化された時間分散ベースの関心領域検出、多段階の Sim-to-Real 転移学習、および物理情報に基づいた正則化を組み合わせて、予測精度と物理的一貫性を強化します。合成事前トレーニング、シルバーラベル適応、専門家による微調整と並行して、トランスフォーマーベースや再帰畳み込みアーキテクチャなど、さまざまな時空間アーキテクチャが評価されました。結果は、変圧器ベースのアーキテクチャが瞬間予測の精度の点で優れている一方、軽量の反復畳み込みアーキテクチャがより高い時間的安定性と運用海洋学スキルを達成したことを示しています。アブレーション研究では、傾向追跡の一貫性と物理的にありえない予測という点で、物理学に基づく正則化の利点も実証されました。説明可能性監査は、流体力学的に活発なサーフゾーン領域に注意を集中させるのにも役立ち、物理的に導出された波の伝播挙動との良好な一致を示しました。一般に、提案されたフレームワークは、コスト効率が高く、運用上実現可能な、長期の沿岸波浪モニタリングのための物理学誘導ビデオベースの深層学習システムの可能性を示しています。
原文 (English)
Physics-Guided Spatiotemporal Learning for Coastal Wave Peak Period Estimation from Video
Direct estimation of physically interpretable periodic signals from raw video constitutes a spatiotemporally grounded learning problem that proves to be difficult especially when facing label sparsity, lack of physical grounding and standardization benchmarks. The wave monitoring at coastal sites is one such real-world example where current deep learning approaches for estimating wave parameters using video as input suffer from physical interpretability and require some kind of intermediate data processing. In this study we propose a framework for wave peak period estimation using only video as input through three components: automated region-of-interest detection using temporal pixel variance, multi-stage Sim-to-Real transfer learning process, and physics-guided regularization of the output predictions. Various spatiotemporal architectures, including Transformer and recurrent-convolutional were compared during the stages of synthetic pretraining, silver label adaptation, and expert fine-tuning. It has been found out that LtViViT achieves the highest accuracy in its estimates, while TinyWaveNet shows superior temporal stability and oceanographic skill. Additionally, ablation studies have demonstrated that physics-guided regularization helps to follow the trends in predictions more consistently and prevent physically meaningless predictions. Moreover, Grad-CAM-based explainability analysis of the physics-guided TinyWaveNet showed that its spatial focus aligns with hydrodynamically active surf-zone regions. Overall, the findings support physics-guided, video-based deep learning as a cost-effective and operationally viable approach for long-term coastal wave monitoring, and demonstrate a transferable strategy for physically-constrained spatiotemporal regression from video under data-scarce conditions.
RoboPIN: 固定された思考連鎖によるグラウンディングされた身体的推論
身体化された推論では、モデルが物理環境内のタスクに関連するオブジェクトや空間を認識し、複数ステップの推論を通じて一貫した視覚的根拠を維持する必要があります。しかし、現在の視覚言語モデルはテキストのみ、または座標拡張された思考連鎖に依存しており、実体参照は暗黙的かつ曖昧なままです。これにより、推論プロセスが視覚的な証拠から切り離され、エンティティ参照がステップ間で漂流し、推論の軌跡と最終的な答えとの間に因果関係の断絶が生じる可能性があり、これらの問題は、ビュー間の外観の変化によりマルチビュー シナリオでさらに増幅されます。これらの問題に対処するために、すべての推論ステップを視覚的な証拠に固定する構造化推論パラダイムである Pinned Chain-of-Thought (\pincot{}) を提案します。 \pincot{} は \reasoninganchor{} の概念を導入しています。これは、タスクに関連する各エンティティを、エンティティ名、一意の ID、ビュー インデックス、空間基盤を備えた構造化されたビジュアル アンカーにバインドし、推論ステップとビュー全体で一貫したエンティティの追跡を可能にします。完全に自動化されたデータ生成パイプラインを構築して、高品質の \pincot{} 形式の推論データセットである \dataset{} を構築します。次に、具体化された知識、構造化された推論能力、プロセス監視された調整を段階的に注入する 3 段階のポストトレーニングを通じて、\method{} をトレーニングします。報酬は、推論中のアンカーの位置特定とアイデンティティの一貫性の両方を直接制約します。埋め込まれた空間推論、マルチビュー推論、ポインティングをカバーする 14 のベンチマークでは、パラメーターが 4B のみの \method{} は常に 7B レベルのオープンソースの埋め込みモデルを上回り、最も強力な 7B ベースラインである Mimo-Embodied に対して平均 12\% の改善を達成しました。さらに分析すると、\pincot{} によって接地精度とステップ間の同一性の一貫性が向上し、プロセス監視の有効性が検証されたことが示されています。
原文 (English)
RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought
Embodied reasoning requires models to perceive task-relevant objects and spaces in physical environments and maintain consistent visual grounding throughout multi-step reasoning. However, current vision-language models rely on text-only or coordinate-augmented chain-of-thought, where entity references remain implicit and ambiguous. This may cause the reasoning process to decouple from visual evidence, entity references to drift across steps, and a causal disconnection between the reasoning trajectory and the final answer, with these problems further amplified in multi-view scenarios due to cross-view appearance changes. To address these issues, we propose Pinned Chain-of-Thought (PinCoT), a structured reasoning paradigm that pins every reasoning step to visual evidence. PinCoT introduces the concept of reasoning anchor, which binds each task-relevant entity to a structured visual anchor with entity name, unique identity, view index, and spatial grounding, enabling consistent entity tracking across reasoning steps and views. We build a fully automated data generation pipeline to construct PIN-170K, a high-quality PinCoT-formatted reasoning dataset. We then train RoboPIN through three-stage post-training that progressively injects embodied knowledge, structured reasoning ability, and process-supervised alignment, with rewards that directly constrain both anchor localization and identity consistency during reasoning. On 14 benchmarks covering embodied spatial reasoning, multi-view reasoning, and pointing, RoboPIN with only 4B parameters consistently outperforms 7B level open-source embodied models, achieving a 12% average improvement over the strongest 7B baseline, Mimo-Embodied. Further analysis shows that PinCoT improves grounding accuracy and cross-step identity consistency, validating the effectiveness of process supervision.
予算付き LLM 検証における不均一分散信号: 構造の不均一性が最適化ゲインを制限する
大規模言語モデル (LLM) システムでは、検証、テスト時間のスケーリング、ツールの実行、その他の選択的な計算の決定に限られた計算を割り当てるために、不確実性信号の使用が増えています。このようなポリシーは \emph{グローバルな信号の比較可能性の仮定} に依存します。つまり、等しいスコアは入力全体で比較可能な決定値を保持する必要があります。制御された診断設定として予算に基づいた検証を使用して、この仮定の失敗モードを特定します。不確実性の品質はコスト層全体で不均一分散的であり、多くのエラーが集中しているにもかかわらず、一部の領域ではほぼランダムな識別性が示されています。明示的なローカル モデルの下で、結果として生じるグローバル割り当ての歪みを特徴付け、その上限が層間の信号品質分散に応じて変化することを示します。私たちは、制御された介入階層 (しきい値、MP-Adapt、MP-Strat、および意図的に単純なコスト階層化しきい値介入 (CST)) を通じて、弱い信号、最適化の不安定性、構造的異質性を分離します。 Qwen3-8B、LLaMA3-8B、および GPT-4o-mini を使用した MBPP と MATH 全体で、グローバルなオンライン適応により、静的しきい値処理に比べて一貫性のないゲインが得られます。 MP-Strat はパフォーマンスを部分的に回復しますが、CST は勾配更新なしで非常に異質な設定でヒット率を最大 17 パーセント改善します。これらの結果は、観察された設定における主なボトルネックとして、オプティマイザーの弱点だけではなく、構造的異質性を特定します。さらに広く言えば、調整されていないフィードバック構造は、より強力な最適化によって常に修復できるわけではありません。
原文 (English)
Heteroskedastic Signals in Budgeted LLM Verification: Structural Heterogeneity Limits Optimization Gains
Selective-compute LLM systems decide which outputs merit verification, additional reasoning, tool execution, or human audit under a limited budget. It is natural to expect that stronger online optimization over a shared uncertainty or reward signal should improve these decisions. We take a critical look at this assumption and ask: when does optimizing harder fail because the signal is not decision-comparable across inputs? In budgeted LLM verification, we find that uncertainty quality is heteroskedastic across cost strata: some regions exhibit near-random discriminability while concentrating many errors. Under an explicit local model, we characterize the resulting distortion of global allocation and show that its upper bound scales with cross-stratum signal-quality dispersion. To separate weak signals from optimizer instability and structural mismatch, we introduce a controlled intervention hierarchy: Threshold, MP-Adapt, MP-Strat, and cost-stratified thresholding (CST). We then turn the diagnosis into Heterogeneity-Gated Allocation (HGA), which uses a warm-up comparability test to choose between global and cost-stratified allocation. Across MBPP and MATH using Qwen3-8B, LLaMA3-8B, and GPT-4o-mini, global online adaptation yields inconsistent gains over static thresholding; CST improves hit rate by up to 17 percentage points in strongly heterogeneous settings, while HGA preserves most gains and avoids blind stratification when the partition is not useful. These findings suggest a resource-allocation principle for LLM systems: before optimizing harder over a shared proxy, test whether the proxy is decision-comparable across observable operating regimes, and gate structural specialization on that test.
Omni-Perception Policy Optimization for Multimodal Emotion Reasoning
We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their…
データ駆動型機械学習は記号レベルの論理的推論に到達できない -- スケーリング則の限界
Sphere ニューラル ネットワークは、トレーニング データなしで記号レベルの三段論的推論を達成しました。これにより、論理的推論のスケーリング則の限界がどこにあるのか、つまり、データ駆動型の機械学習システムがトレーニング データとトレーニング時間を増やすことで同じレベルを達成できるかどうかという問題が生じています。教師あり深層学習が記号レベルの三段論的推論に到達することを妨げる 2 つの方法論的制限を示します。(1) トレーニング データは、24 種類の有効な三段論的推論すべてを区別できない。 (2) 前提から結論までのエンドツーエンドのマッピングでは、パターン認識と論理的推論のための神経コンポーネント間に矛盾するトレーニング ターゲットが導入されます。理論的な分析に加えて、オイラー ネットでは厳密な三段論的推論を達成できないことを実験的に示します。さらに、最新の ChatGPT (GPT-5-nano および GPT-5) に対して、単語、二重単語、単純なシンボル、および長いランダム記号の 4 つの表面形式 (パターン) で三段論法的ステートメントの充足可能性を判定することに挑戦し、表面形式が推論パフォーマンスに影響を与えること、および ChatGPT GPT-5 が 100% の精度に達する可能性があるが、依然として不正確な説明を提供する可能性があることを示します。経験的トレーニングプロセスは 100% の精度に達した後に停止されるため、教師あり機械学習システムは記号論理推論の厳密さを達成できないと結論付けます。
原文 (English)
Data-driven Machine Learning Cannot Reach Symbolic-level Logical Reasoning -- The Limit of the Scaling Law
By promoting vectors to spheres and enabling explicit model construction, neural networks can perform symbolic-level syllogistic reasoning without training data. We identify two fundamental limitations that prevent conventional data-driven machine learning systems from achieving this capability: training data generated by the combination table cannot distinguish all 24 valid syllogism types, and end-to-end premise-to-conclusion mapping creates contradictory targets within neural components. Experiments with two representative conventional systems, GPT-5 using linguistic inputs and Euler Net using visual inputs, support this analysis. ChatGPT GPT-5 may reach 100% accuracy in syllogistic reasoning, but with hallucinations. Because the learning process terminates upon reaching 100% accuracy, the system cannot progress beyond empirical accuracy to symbolic level reasoning. Random test data reduced Euler Net's accuracy to 56%. Repeatedly expanding the training set increased its accuracy to 97%, with perfect performance on 8 syllogism types. However, because unintended inputs cannot be exhaustively covered, even 100% test accuracy does not imply symbolic-level reasoning. Since syllogistic reasoning underpins logical reasoning and human rationality, these results suggest that increasing data and training time alone cannot ensure symbolic level logical reasoning.
Theoria: 非公式推論状態に対する書き換え許容性の検証
AI システムの答えを信頼できるのはどのような場合ですか?形式的証明アシスタントは確実性を提供しますが、問題分布のほとんどには到達できません。スカラー LLM ジャッジはカバレッジを提供しますが、事後的に監査できない不透明なスコアを生成し、他の LLM と同じ一貫性の問題にさらされます。私たちは、このギャップを埋める検証アーキテクチャである Theoria を紹介します。候補解は、型指定された状態遷移のシーケンスに書き換えられます。各状態遷移は、引用、計算、または問題によって与えられた事実など、明示的な正当化によってライセンスされ、すべての遷移は独立して監査可能です。基本的な不変条件は変化の完全性です。連続する証明状態間のすべての違いを考慮する必要があるため、隠れた前提は黙って通過するのではなく、許可されていない突然変異として表面化します。 HLE-Verified Gold (185 のテキストのみのエキスパートの問題) では、Theoria は 91.4% の厳密な精度で 105 を認定しています (Wilson 95% CI [84.5%、95.4%])。すべての認証では、人間が判読できる証明トレースが生成され、各ステップに個別にチャレンジできます。ホリスティック LLM ジャッジは、一致するカバレッジでは同等の精度を達成しますが、別の問題 (Jaccard 0.14 ~ 0.36) では失敗するため、アプローチは補完的になります。 15 のドメインにわたる 95 件の敵対的毒物証明について、構造化された裁判官は 94.7% を捕捉したのに対し、総合的な判断では 83.2% を捕捉しました (p= 0.0017)。全体の 11.5 pp のギャップは、隠れた前提 (90.6% 対 62.5%、28 pp の差) と捏造された引用 (100% 対 90%) に集中しており、形式的な分析が利点を予測するエラー クラスです。利点が予測されない算術および定理の誤用エラーのパフォーマンスは同じです。 GPQA ダイヤモンド (n= 65) では、認定精度は 97.1% (Wilson CI [85.1%、99.5%]) です。
原文 (English)
Theoria: Rewrite-Acceptability Verification over Informal Reasoning States
When should an AI system's answer be trusted? Formal proof assistants offer certainty but cannot reach most of the problem distribution; scalar LLM judges offer coverage but produce opaque scores that cannot be audited after the fact and are subject to the same coherence issues as any LLM. We present Theoria, a verification architecture that closes this gap. A candidate solution is rewritten into a sequence of typed state transitions, each licensed by an explicit justification, whether that be a citation, computation, or problem-given fact, and every transition is independently auditable. The foundational invariant is completeness of change: every difference between consecutive proof states must be accounted for, so hidden premises surface as unlicensed mutations rather than passing silently. On HLE-Verified Gold (185 text-only expert problems), Theoria certifies 105 at 91.4% strict precision (Wilson 95% CI [84.5%, 95.4%]). Every certification produces a human readable proof trace in which each step can be independently challenged. Holistic LLM judges achieve comparable precision at matched coverage but fail on different problems (Jaccard 0.14-0.36), making the approaches complementary. On 95 adversarial poisoned proofs across 15 domains, structured judges catch 94.7% versus 83.2% for holistic judging (p= 0.0017). The overall 11.5 pp gap concentrates in hidden premises (90.6% vs. 62.5%, a 28 pp difference) and fabricated citations (100% vs. 90%), the error classes where the formal analysis predicts an advantage; performance is identical on arithmetic and theorem-misapplication errors, where no advantage is predicted. On GPQA Diamond (n= 65), certified precision is 97.1% (Wilson CI [85.1%, 99.5%]).
GUI エージェントは自分の目を信じますか?ピクセル対構造に対する状態信念依存性の診断
マルチモーダル GUI エージェントは、スクリーンショットのレンダリングされたピクセルと、DOM やアクセシビリティ ツリーなどのシリアル化された構造という 2 つの冗長チャネルを通じてインターフェイスを読み取ります。エージェントは行動する前に、現在のインターフェイス状態について信念を形成しますが、既存のベンチマークはタスクの成功、要素のグラウンディング、または攻撃耐性をスコアリングし、その信念がピクセルから引き出されたものであるかどうかは質問しません。私たちは視覚的な状態依存性、つまり状態信念のピクセル、構造、または事前分布への帰属を形式化し、310 の実際の Web、モバイル、およびデスクトップのプローブに対するペアの単一チャネル介入でそれを測定します。すべてのプローブは、モデルによって生成された項目やモデルによる判断を行わず、決定論的な強制選択によってスコア付けされます。私たちの中心的な指標は、モデルが正しく認識し、矛盾している構造に向けて解決するプローブの割合である知覚融合ギャップです。 3 社のベンダーの 5 つのモデルにわたって、テキスト状態の信念は構造に依存しますが、画像のみの精度は天井付近に留まり、知覚と融合のギャップはすべてのモデルでプラスです。対照的に、非テキスト ID は主にピクセルに限定されたままになります。この置換はシリアル化されたテキストとインデックス付きアクション チャネルに固有であり、調整アクション エージェントはほとんど影響を受けません。テキストの競合の場合、ホワイト ボックス アブレーションにより、コピーされた単一の構造値への影響が追跡され、2 つの実際の環境では、競合により誤ったアクションと実際のタスクの失敗が引き起こされます。したがって、視覚的状態依存性は、エージェント状態の信念が視覚的に根拠があり、それが露呈するエラーがアクションに伝播するかどうかの測定可能な診断を提供します。
原文 (English)
Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure
Multimodal GUI agents read an interface through two redundant channels: the rendered pixels of a screenshot and a serialized structure such as a document object model or accessibility tree. Before acting, an agent forms a belief about the current interface state, but existing benchmarks score task success, element grounding, or attack resistance and do not ask whether that belief is drawn from the pixels. We formalize visual state reliance, the attribution of a state belief to pixels, structure, or priors, and measure it with paired single-channel interventions over 735 probes spanning real web, mobile, and desktop interfaces, of which 225 are zero-edit divergences mined from live production websites, all scored by deterministic forced choice with no model judge. Our central metric is the Perception-Fusion Gap (PFG), the fraction of probes a model perceives correctly yet resolves toward structure under conflict; a stricter variant that re-verifies perception on a tight crop of the target region leaves the gap intact. Across models from four vendors, textual state beliefs defer to structure while image-only accuracy stays near ceiling, and on unedited stale snapshots from live pages the same models follow the outdated structure on up to 0.88 of probes. A white-box ablation traces the textual effect to a single copied structural value, and gradient attribution shows the visual evidence is processed yet overridden. In live multi-step environments, one mis-sourced belief at the first step compounds into task failure with a self-recovery rate of at most 0.03. Comparing four mitigations on identical probes, prompt-level cues fail at the action level, certificate checks buy safety with refusals, and a training-free consistency gate is alone in reducing both hijack and task error. Visual state reliance thus gives a measurable diagnostic of whether agent state beliefs are visually grounded.
ループ内のメモリ: 言語エージェントの拡張作業メモリとしてのプロセス内取得
言語エージェントはループを実行します - 観察、推論、行動 - しかし、彼らが推論する記憶はループの外にあり、ストアはターンごとに最大 1 回クエリされます。私たちは、メモリがループ内を移動し、各ステップで読み書きされる体制を研究します。障害となるのは常に遅延です。ネットワーク化されたストアは数十ミリ秒から数百ミリ秒で応答します。また、ループ内取得では、取得にコストがかかる場合、エンドツーエンドの遅延が最大 83 倍に膨らむ可能性があります。以前の研究では、そのコストを問題視するのではなく管理していました。サービングレイヤーのスケジューリングによってそれが隠蔽され、「メモリファースト」設計により、取得がターンごとに 1 回に割り当てられました。私たちは、レイテンシはループ内パターンではなく、ストアが存在する場所の特性であると主張します。処理中のストアは、ネットワーク レジームより 3 桁低い、最大 100 マイクロ秒で応答し、その速度ではステップごとの税が崩壊します。拡張思考論文のパリティ原理により、常に直接利用できるほど高速なストアは、エージェントが単に参照するツールではなく、拡張作業メモリになります。前提は因果関係です。固定のターンごとのメモリ レイテンシ バジェットを保持し、ストアの応答速度のみを変化させると、冗長アクションはレイテンシとともに単調増加します。インプロセス速度では 0.0/12、110 ミリ秒のクラウド往復では 7.2/12 (gpt-5-nano、gpt-5-mini; 正確な順列 p=0.0079)。この体制をエンドツーエンドで実証します。制限されたウィンドウの下で 4 つの GPT-5 クラス モデルにわたって、ループ内メモリでリコールが 0/5 から 3.6 ~ 4.8/5 に改善され、p50 80 ~ 165us でのストア操作が行われます。ただし、指示された restate-every-reply ベースラインでも完全に解決されますが、トークン コストはワーキング セットとともに増加します。ストアはいかなる実行においても事実を失うことはありませんでした (244 件の書き込みのうち 244 件が保持されました)。すべてのミスは、ストアではなくエージェントの読み取りポリシーを追跡します。私たちの測定では、ボトルネックの位置も再確認されています。ステップごとの主なコストは埋め込みです (ネットワーク上で約 200 ~ 400 ミリ秒)。インプロセス ストアを小さなローカル エンベッダーと組み合わせると、完全な操作が測定値で約 40 マイクロ秒に戻ります。
原文 (English)
Memory in the Loop: In-Process Retrieval as Extended Working Memory for Language Agents
Language agents run a loop - observe, reason, act - but the memory they reason over sits outside it: a store queried at most once per turn. We study the regime where memory moves inside the loop, read and written on every step. The obstacle has always been latency: networked stores answer in tens to hundreds of milliseconds, and in-loop retrieval can inflate end-to-end latency by up to 83x when retrieval is expensive. Prior work manages that cost rather than questioning it: serving-layer scheduling hides it, "memory-first" designs ration retrieval to once per turn. We argue latency is a property of where the store lives, not the in-loop pattern: an in-process store answers in ~100us, three orders of magnitude below the network regime, and at that speed the per-step tax collapses. By the extended-mind thesis's parity principle, a store fast enough to be constantly and directly available becomes extended working memory, not a tool the agent merely consults. The premise is causal: holding a fixed per-turn memory-latency budget and varying only the store's answer speed, redundant actions rise monotonically with latency - 0.0 of 12 at in-process speed, 7.2 of 12 at a 110ms cloud round trip (gpt-5-nano, gpt-5-mini; exact permutation p=0.0079). We demonstrate the regime end-to-end: across four GPT-5-class models under a bounded window, recall improves from 0/5 to 3.6-4.8/5 with in-loop memory, store ops at p50 80-165us - though an instructed restate-every-reply baseline also solves it perfectly, at a token cost that grows with the working set. The store never lost a fact in any run (244 of 244 writes kept); every miss traces to the agent's read policy, not the store. Our measurements also relocate the bottleneck: the dominant per-step cost is embedding (~200-400ms over the network); pairing the in-process store with a small local embedder returns the complete operation to a measured ~40us.
ウラソフ方程式の平均場導出の定式化: 戦略ゲームとしての AI 支援のリーン形式化
数学者に AI システムを指示してもらい、リーン 4 証明アシスタントで研究結果を形式化し、そのアクティビティを形式化ゲームとして組み立てます。目的は、LaTeX ドキュメントをリーンなドキュメントに変えることです。開発がコンパイルされ、Sorry が含まれておらず、ターゲット定理がリーンの基本公理のみに基づいていることがマシン チェックで示されたときに、ゲームは勝ちとなります。再利用は、私たちが導入した定義による 2 番目のチェックです。開発により、より広範なライブラリが吸収できる一般数学の自己完結型の層が得られるかどうかです。このケーススタディは、ドブルシンの平均場ルート、つまり存在、一意性、安定性推定と平均場の限界、およびショートウィンドウ重ね合わせ原理(弱い解はラグランジュ)を介した、非線形ウラソフ方程式のウェルポーズネスの完全で公理的にクリーンな形式化です。人間の役割は、証明を書くことではなく、指示することでした。定義の範囲を絞り、分解を指示し、ライブラリのギャップをトリアージすることでした。 AI エージェントが実行されました。形式化により、各ステートメントが書かれた通りに証明されたことが証明されます。書かれた記述が意図した定理であるかどうかは数学者の判断に委ねられます。ビルドから外れた最適トランスポート機構 (特に、Wasserstein-1 メトリックとKantrovich-Rubinstein 双対性定理の特性) は、Mathlib のみに対してコンパイルされる自己完結型の層に分離されます。開発の約 6 分の 1 (299 個の宣言のうち 49 個) が、逆依存性のない 22 個の宣言インターフェースの背後にあります。見出しの定理は約 1 週間で実行され、完全な開発には約 1 か月かかりました。私たちは定量的な主張を一般的な法則としてではなく、1 つのゲームの観察として報告します。ゲームのルールでは特定のシステムが指定されていないため、方法論的な枠組みは 1 回の実行で得られるツールよりも長持ちするように意図されています。
原文 (English)
A Formalization of the Mean-Field Derivation of the Vlasov Equation
We formalize a research result in the Lean 4 proof assistant by having a mathematician direct an AI system, and frame the activity as a formalization game. The objective is to turn a LaTeX document into Lean. The game is won when the development compiles, contains no sorry, and a machine check shows the target theorems rest on Lean's foundational axioms alone. Reuse is a second check, by a definition we introduce: whether the development yields a self-contained layer of general mathematics the wider library could absorb. The case study is a complete, axiom-clean formalization of well-posedness for the nonlinear Vlasov equation via Dobrushin's mean-field route -- existence, uniqueness, the stability estimate and mean-field limit, and a short-window superposition principle (weak solutions are Lagrangian). The human's role was to direct, not to write proofs: to scope the definitions, steer the decompositions, and triage the library's gaps; the AI agent executed. The formalization certifies the proof of each statement as written; whether the written statement is the intended theorem stays the mathematician's judgment. The optimal-transport machinery that fell out of the build (in particular, properties of the Wasserstein-1 metric and the Kantorovich-Rubinstein duality theorem) separates into a self-contained layer that compiles against Mathlib alone: about a sixth of the development (49 of 299 declarations), behind a 22-declaration interface with no reverse dependency. The headline theorems ran in about a week, the full development in about a month. We report the quantitative claims as observations of one game, not as general laws. The game's rules name no particular system, so the methodological framing is meant to outlast the tools of any one run.
IdeaTrail: 科学的アイデアのためのフルプロセス エージェントの軌跡
科学研究は、テキスト生成という単一の行為ではなく、複雑な多段階のワークフローです。通常、アイデアのプロセスは、文献検索、論文の読解、ツールの使用、クレームの確認、論文間の統合、ブレーンストーミング、弱い指示の拒否、および反復的な執筆を通じて現れます。既存のリソースはこのプロセスの個々のコンポーネントをキャプチャしますが、ツールの使用、証拠の取得、中間成果物の進化、アイデアまたは提案レベルのエンドポイントを共同で記録するデータセットは依然として限られています。このレポートでは、科学的アイデアと提案生成のためのマルチターン プロセス軌跡データセットである \method を紹介します。各インスタンスは、証拠の収集からアイデアの選択または提案の作成までの調査プロセスを記録します。 \method は軌道を自由に作成するのではなく、人間が選択した高品質の研究論文と提案成果物から開始し、ジェネレーターとアドバイザーの合成ループを使用します。ジェネレーターはアクション、観察、アーティファクトの編集を通じて目に見える軌道を生成しますが、アドバイザーは完全な生成コンテキストにアクセスして、グラウンディング、因果関係の順序、自然性、および隠れたターゲットからの漏れをチェックします。この逆から順の手順により、実際の科学成果との整合性を保ちながら、研究実践の不確実性、証拠の使用、段階的な収束を近似した複数ターンの研究データが生成されます。 \method は、科学研究エージェント向けにプロセス監視データを合成するためのデータセットと一般的なレシピの両方を提供します。
原文 (English)
IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation
Scientific ideation unfolds over multiple stages, including literature search, paper reading, tool use, claim checking, cross-paper synthesis, brainstorming, rejection of weak directions, and iterative writing. Yet most existing resources capture isolated components or final artifacts rather than the process connecting them. We introduce IdeaTrail, a dataset of 1,170 multi-turn trajectories for scientific ideation and proposal generation. Each trajectory follows a research process from evidence gathering to either idea selection or proposal construction, jointly recording tool use, acquired evidence, intermediate artifacts, and reasoning. IdeaTrail is synthesized from human-selected research papers and proposal artifacts through a Generator--Advisor loop. The Generator produces the visible sequence of actions, observations, and artifact edits, while the Advisor uses the full generation context to check grounding, causal order, naturalness, and leakage from hidden targets. This reverse-to-forward design keeps trajectories aligned with real scientific artifacts while retaining the uncertainty, evidence use, and staged convergence characteristic of research practice. IdeaTrail provides both reusable process supervision and a general recipe for constructing scientific-research-agent data.
コーディング エージェント基盤モデルの中間トレーニングとしての機能を意識した中間補充
コーディング エージェントは、外部ツールのリターンを継続的な推論に統合する必要があります。これは、コードに対する標準の左から右への事前トレーニングが順方向でのみ公開する機能です。コーディング エージェントのアクション - 観察 - 継続ループは構造的に関数呼び出しサイトと同形であることがわかります。呼び出し元は引数をバインドし、呼び出し先は別の場所で計算された値を返し、ダウンストリーム コードはその値を消費します。この条件付け構造は、通常のコード内にインターネット規模で存在します。私たちはこれを、関数を意識した中間補充 (FIM) 中間トレーニングを通じて活用します。これは、プログラムの依存関係グラフ分析と複雑さの推論の二重基準によって選択された関数をマスクする自己監視型の目標です。 968 の GitHub リポジトリから抽出された 2.6B トークンの汚染除去されたコーパス上で Qwen2.5-Coder-Instruct (7B/14B) と Qwen3-8B を中間トレーニングし、既存のエージェントのポストトレーニング パイプラインを適用します。中間トレーニングでは、SWE-Bench-Verified が 7B/14B で +2.8/+3.0、Qwen3-8B で +3.2 向上します。 SWE-Bench-Lite のゲインは、同じモデルで +3.7/+4.0/+5.4 です。この改善は、2 つのポストトレーニング パイプライン (R2E-Gym、SWE-Smith) および非 Qwen2.5 ベース (SWE-Lego を使用した Qwen3-8B) に当てはまります。ドメイン内のゲインだけでなく、トレーニング中は、エージェントのポストトレーニングが非エージェントコーディング (LiveCodeBench など) や非コーディングツール使用ベンチマーク (tau-bench、BFCL) に与える能力の低下も軽減します。トレーニング途中のコーパスには Python コードのみが含まれていますが、関数呼び出しの帰納的バイアスはトレーニング後も存続し、一貫したゲインが得られます。
原文 (English)
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code exposes only in its forward direction. We observe that the action-observation-continuation loop of a coding agent is structurally isomorphic to a function call site, where a caller binds arguments, a callee returns a value computed elsewhere, and downstream code consumes that value. This conditioning structure exists at internet scale in ordinary code. We exploit it through function-aware fill-in-the-middle (FIM) mid-training: a self-supervised objective that masks functions selected via program dependency graph analysis and a complexity-inferability double criterion. We mid-train Qwen2.5-Coder-Instruct (7B/14B) and Qwen3-8B on a 2.6B-token decontaminated corpus drawn from 968 GitHub repositories, then apply existing agentic post-training pipelines. Mid-training improves SWE-Bench-Verified by +2.8/+3.0 at 7B/14B and by +3.2 on Qwen3-8B; SWE-Bench-Lite gains are +3.7/+4.0/+5.4 on the same models. The improvement holds across two post-training pipelines (R2E-Gym, SWE-Smith) and on a non-Qwen2.5 base (Qwen3-8B with SWE-Lego). Beyond in-domain gains, mid-training also mitigates the capability erosion that agentic post-training otherwise inflicts on non-agent coding (e.g., LiveCodeBench) and non-coding tool-use benchmarks (tau-bench, BFCL): although the mid-training corpus contains Python code only, the function-call inductive bias survives post-training and yields consistent gains.
視覚言語モデル推論における視覚アクセス境界
思考連鎖 (CoT) プロンプトは、ビジョン言語モデル (VLM) のテスト時間スケーリング戦略として広く使用されていますが、VLM がより長い推論トレースを生成するときに何が拡張されるのかは依然として不明です。 CoT は画像トークンへの継続的なアクセスを必要とするのか、それとも、主にフォワード パスの早い段階で既に利用可能になった視覚情報を基に動作するのかを尋ねます。レイヤーの深さと生成時間に沿って、生成されたトークン クエリからイメージ トークン キーへの注意をマスクする因果的介入であるビジュアル アクセス スイープを導入し、タスクの精度を維持する最小アクセス領域としてビジュアル アクセス境界 (VAB) を定義します。 Qwen2.5-VL および InternVL3 の 6 つのモデル構成にわたって、CoT なしの直接応答と CoT プロンプトの両方が有限の VAB を示します。 14B および 38B スケールの Qwen2.5-VL-32B および InternVL3 では、CoT が非 CoT フルアクセス ターゲットに対して評価される場合、実質的に長い世代にも関わらず、その VAB 層は最大 2 層だけ非 CoT 境界と異なります。これは、CoT が推論トレース全体で直接イメージ トークン アクセスを延長することによって主にパフォーマンスを向上させるのではなく、イメージ由来の隠れ状態情報に対する言語側の計算を拡張することによってパフォーマンスを向上させることを示唆しています。さらに、CoT ゲインが知覚的読み出しによって制限されることを示します。 CoT は、クエリされた視覚属性がモデルによって確実に読み取れる場合には役立ちますが、その読み出しが信頼できない場合には役に立ちません。シンボリック属性のオラクルは、グラウンドトゥルース属性がテキストとして提供されると CoT によってカウントが向上することを示し、一方、単一オブジェクトのプローブ対デコードのチェックは、ハード属性が隠れた状態から線形に回復可能であるものの、モデル自体が出力するのは難しいことを示しています。これらの分析を組み合わせると、カウントではなく読み出しにボトルネックが生じます。
原文 (English)
Visual Access Boundaries in Vision-Language Model Reasoning
Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires continued access to image tokens, or whether it mainly operates over visual information already made available earlier in the forward pass. We introduce Visual Access Sweep, a causal intervention that masks attention from generated-token queries to image-token keys along layer depth and generation time, and define the Visual Access Boundary (VAB) as the minimal access region that preserves task accuracy. Across six model configurations from Qwen2.5-VL and InternVL3, both no-CoT direct answering and CoT prompting exhibit finite VABs. In Qwen2.5-VL-32B and InternVL3 at 14B and 38B scales, when CoT is evaluated against the no-CoT full-access target, its VAB layer differs from the no-CoT boundary by at most two layers, despite substantially longer generations. This suggests that CoT does not primarily improve performance by prolonging direct image-token access throughout the reasoning trace, but by extending language-side computation over image-derived hidden-state information. We further show that CoT gains are constrained by perceptual readout. CoT helps when the queried visual attribute can be reliably read out by the model, but not when that readout is unreliable. A symbolic-attribute oracle shows that CoT can improve counting once ground-truth attributes are supplied as text, while a single-object probe-vs-decode check shows that hard attributes can be linearly recoverable from hidden states yet difficult for the model itself to output. Together, these analyses place the bottleneck at readout rather than counting.
Belnap の型付き内包 FOL に基づく神経記号的 AGI ロボットの確率的拡張
$IFOL_B$ に基づくニューロシンボリック AI は、ニューラル学習と記号推論を組み合わせて、純粋なニューラル システムの制限 (解釈可能性や論理構造の欠如など) を自己参照のための形式的な論理機構で克服する方法です。この論文では、$IFOL_B$ の Nilsson の確率構造に基づいて、現在未知の文の確率計算を使用して、$IFOL_B$ の認知能力を拡張します。現在の知識データベースと論理推論を保存するグローバル対称変換と、$IFOL_B$ 述語の非常に厳密なサブセットのみを含む具体的な (サブ) 問題に関するリアルタイムの決定に使用されるローカル対称変換を導入します。どちらの場合も、シャノンの最大情報エントロピーに基づく確率密度関数 $KI$ の計算は、この確率的ニューロシンボリック AGI のニューラル ネットワークによって提供されます。
原文 (English)
Probabilistic Extension of Neuro-Symbolic AGI Robots based on Belnap's Typed Intensional FOL
Neuro-symbolic AI based on $IFOL_B$ is a way to combine neural learning and symbolic reasoning to overcome limitations of purely neural systems (like lack of interpretability and logical structure) with formal logical machinery for self-reference. In this paper we expand the cognitive power of $IFOL_B$ by using the probability computation for the currently unknown sentences, based on Nilsson's probability structure for the $IFOL_B$. We introduce the global symmetry transformation that preserves the current knowledge database and logical deduction, and the local one used for real-time decisions about concrete (sub)problems that involve only a very strict subset of $IFOL_B$ predicates. The computation of probability density function $KI$ in both cases, based on the Shannon's maximum information entropy, is provided by neural networks of this probabilistic neuro-symbolic AGI.
AgentCompass: エージェント機能の統合評価インフラストラクチャ
大規模言語モデル (LLM) が自律エージェントに進化するにつれて、統合された評価インフラストラクチャの必要性が重要になります。ただし、現在の評価パイプラインは高度に断片化され、密接に結合されたままであるため、再現性が妨げられ、冗長なエンジニアリングが発生します。これに対処するために、LLM ベースのエージェントを評価するためのオープンソースで軽量かつ拡張可能なインフラストラクチャである AgentCompass を導入します。 AgentCompass は、ベンチマーク、ハーネス、環境という 3 つの独立したコンポーネントを中心に評価プロセスを編成するため、複雑な実行ロジックを再実装することなく柔軟な構成が可能になります。さらに、フォールトトレラントな非同期ランタイムと、報酬ハッキングなどの微妙な障害モードを透過的に診断するための包括的な軌跡分析ツールを備えています。 AgentCompass は、5 つの機能次元にわたる 20 以上のベンチマークをネイティブにサポートし、エージェント研究を進めるためのスケーラブルで再現可能なインフラストラクチャをコミュニティに提供します。
原文 (English)
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely Benchmark, Harness, and Environment, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic. Furthermore, it features a fault-tolerant asynchronous runtime and comprehensive trajectory analysis tools to transparently diagnose nuanced failure modes like reward-hacking. Natively supporting over 20 benchmarks across five capability dimensions, AgentCompass provides the community with a scalable and reproducible infrastructure for advancing agent research.
検証済みのワールド モデルが依然として負けている場合: LLM 合成コード ワールド モデルにおける再生の適切性と予測の精度
大規模な言語モデルでは、ゲームのルールを実行可能コード (コード ワールド モデル (CWM)) として合成でき、古典的なプランナーはこれを検索します。このようなモデルは、通常、サンプリングされた軌道上で高い遷移精度に達した場合に受け入れられます。私たちは、これは計画の適切性についての間違った概念であると主張します。 4つのことを示します。 (1) LLM で合成された CWM は、100% の遷移精度でサンプリング ゲートを通過でき、プランナー自身の検索分布では $\geq 98\%$ 状態精度を保ちますが、誤った $<1\%$ がまさに極めて重要なダイナミクスであるため、体系的に損失を被ります。省略されたルールのプレイコストは $0.091$ (シードクラスター化 95% CI $[0.065,0.117]$、$n=4800$) です。私たちはこれを検証済みと正しいのギャップと呼び、合成パイプラインを通じてエンドツーエンドで確認します。 (2) 危害は量的法則 $\mathrm{danger}=\mathrm{play\_cost}\times(1-\mathrm{rarity})^N$ に従い、その $(1-\mathrm{rarity})^N$ のゲートミス係数は正確であることが証明されており、そのプレイコストは経験的に制限されています。 (3) 障害はデータを追加しても修復されません。LLM 合成はルール推論ではなくルール変換として動作し、モデル (GPT-5.x) およびデータ領域 (DAgger およびターゲットの例を含む) にわたって省略されたルールを推論しませんでした。 (4) 同じメカニズムが不完全情報 CWM の信念推論関数でも繰り返されます。つまり、カバレッジの限界 (サイズ $N$ ゲートが $N\gtrsim b^{d_{\max}}$ を特定している) を証明し、クーン ポーカーのような浅いゲームにギャップが見られない理由を説明し、ゲートを通過するがすべてのゲームで負ける検証済みだが間違っている推論関数であるビーコンを手動で構築します。これらの結果は、計画指向の世界モデルの適切性は、サンプリングされた遷移の予測精度ではなく、検索分布または直接プレイによって測定されるべきであることを示唆しています。
原文 (English)
When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models
Large language models can synthesize a game's rules as executable code - a Code World Model (CWM) - which a classical planner then searches over. Such models are typically accepted when they reach high transition accuracy on sampled trajectories. We argue this is the wrong notion of adequacy for planning. We show four things. (1) An LLM-synthesized CWM can pass a sampling gate at 100% transition accuracy and be $\geq 98\%$ state-accurate on the planner's own search distribution, yet lose systematically at play, because the $<1\%$ it gets wrong is exactly the pivotal dynamics; the play cost of the omitted rule is $0.091$ (seed-clustered 95% CI $[0.065,0.117]$, $n=4800$). We call this the verified-vs-correct gap, and confirm it end-to-end through the synthesis pipeline. (2) The harm follows a quantitative law, $\mathrm{danger}=\mathrm{play\_cost}\times(1-\mathrm{rarity})^N$, whose $(1-\mathrm{rarity})^N$ gate-miss factor is proven exact and whose play cost is empirically bounded. (3) The failure is not repaired by more data: LLM synthesis behaves as rule translation, not rule inference, and did not infer the omitted rule across models (GPT-5.x) and data regimes (including DAgger and targeted examples). (4) The same mechanism recurs on the belief-inference function of imperfect-information CWMs: we prove a coverage bound (a size-$N$ gate is identifying when $N\gtrsim b^{d_{\max}}$), explaining why shallow games such as Kuhn poker show no gap, and hand-construct Beacon, a verified-but-wrong inference function that passes the gate yet loses every game. These results suggest adequacy for planning-oriented world models should be measured on the search distribution or by play directly, not by prediction accuracy on sampled transitions.
SportD: VLM は物理的に戦略を立てることができますか?
視覚言語モデルは、視覚的なシーンを解釈できるようになってきていますが、戦略的に効果的な意思決定を行うために情報を使用できるかどうかは依然として不明です。私たちはサッカーでこの問題を調査します。モデルはオンボールの決定の数秒前を観察し、シュートするか特定のチームメイトにパスするかを選択する必要があります。従来の視覚的に理解するタスクとは異なり、サッカーでは、利用可能なすべてのアクションの価値を推定することで、意思決定を定量的に評価できます。 2022 FIFA ワールドカップの 478 件のオンボール判定で構成されるベンチマークである SportD を紹介します。各モデルの選択は、攻撃側チームの得点確率を最も高めるアクションを推定するポゼッション価値モデルに対して評価され、最適なアクションの精度と、最適ではない決定によって失われる価値の両方を測定できるようになります。 3 つのフロンティア VLM では、イベントの 31.4% で最も価値の高いアクションが選択されます (プロ プレーヤーの場合は 38.9%)。すべてのモデルで大幅に大きな後悔が発生します。さらなる分析により、より低い分散とより低い報酬のアクションを系統的に好むことが明らかになりました。VLM は、最適なポリシーや実際のプレーヤーよりもシュート頻度が低く、実質的にプログレッシブなパスを選択しません。また、モデルは、プレイヤーの特定のアクションが最適ではない場合でも偶然を超えて再現し、反事実的な代替案の一貫した評価ではなく、よく知られたプレイ パターンの部分的な模倣を示唆しています。 SportD は、VLM における物理的な戦略的推論を測定するための、価値に基づいたテストベッドを提供します。
原文 (English)
SportD: Can VLMs Physically Strategize?
Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use information to make strategically effective decisions. We investigate this question in soccer, where models observe the seconds preceding an on-ball decision and must choose whether to shoot or pass to a specific teammate. Unlike conventional visual-understanding tasks, soccer enables decisions to be evaluated quantitatively by estimating the value of every available action. We introduce SportD, a benchmark comprising 478 on-ball decisions from the 2022 FIFA World Cup. Each model choice is evaluated against a possession-value model that estimates the action that most increases the attacking team's probability of scoring, allowing us to measure both optimal-action accuracy and the value forfeited by suboptimal decisions. Across three frontier VLMs, the best selects the highest-valued action on 31.4% of events, compared with 38.9% for the professional players, and all models incur significantly greater regret. Further analysis reveals a systematic preference for lower-variance and lower-reward actions: VLMs shoot less often and select substantially less progressive passes than either the optimal policy or the real players. The models also reproduce the player's specific action above chance even when that action is suboptimal, suggesting partial imitation of familiar play patterns rather than consistent evaluation of counterfactual alternatives. SportD provides a value-grounded testbed for measuring physical strategic reasoning in VLMs.
SmartRAG: モバイル デバイス用のネイティブ グラフベースの RAG
大規模言語モデル (LLM) をモバイル デバイス上のパーソナル アシスタントとして展開するには、プライバシー、低遅延、オフライン可用性が必要ですが、巨大なモデルの計算コストは厳しいエッジハードウェア予算と衝突します。私たちは、この緊張はモデル圧縮だけでは解決できないと主張します。オンデバイスのインテリジェンスを補完的な機能的役割に分解する必要があります。 SmartRAG は、知覚、記憶、集中、思考という 4 つの調整されたモジュールを中心にインテリジェント アシスタントを組織する完全オンデバイス フレームワークです。 SmartRAG の中核となるのは、継続的に学習可能な名前付きエンティティ認識装置である EvoNER です。これは、教師が抽出した更新を通じてラベル インベントリを段階的に拡張し、システムがバックボーン LLM を再トレーニングすることなく、これまで見えなかったエンティティ タイプを吸収できるようにします。抽出されたナレッジは、3 層の出所保持ナレッジ グラフである MRGraph に保存され、グラフ トラバーサル、字句一致、高密度セマンティック検索を組み合わせたハイブリッド パイプラインを通じてクエリ時に取得されます。オンデバイス LLM は、推論コストを制限しながら、高価値のセマンティック操作 (ラベル付け、計画、回答合成) のためにのみ呼び出されます。 4 つの QA ベンチマーク (TriviaQA、Natural question、HotpotQA、MultiHopQA) の実験では、量子化された 1.7B パラメーターのバックボーンを備えた SmartRAG が、実用的なメモリと遅延エンベロープ内で汎用スマートフォン上で完全に実行しながら、最大 18$\times$ のサイズのモデルと競合するマルチホップ推論パフォーマンスを達成していることが示されています。
原文 (English)
SmartRAG: Native Graph-Based RAG for Mobile Device
Deploying large language models (LLMs) as personal assistants on mobile devices demands privacy, low latency, and offline availability, yet the computational cost of giant models clashes with strict edge-hardware budgets. We argue that this tension cannot be resolved by model compression alone; it requires decomposing on-device intelligence into complementary functional roles. We present SmartRAG, a fully on-device framework that organizes an intelligent assistant around four coordinated modules -- Perception, Memory, Focus, and Thinking. At the core of SmartRAG is EvoNER, a continually learnable named-entity recognizer that incrementally expands its label inventory through teacher-distilled updates, enabling the system to absorb previously unseen entity types without retraining the backbone LLM. Extracted knowledge is stored in MRGraph, a three-layer provenance-preserving knowledge graph, and retrieved at query time through a hybrid pipeline combining graph traversal, lexical matching, and dense semantic search. The on-device LLM is invoked only for high-value semantic operations -- labeling, planning, and answer synthesis -- keeping inference costs bounded. Experiments on four QA benchmarks (TriviaQA, Natural Questions, HotpotQA, MultiHopQA) show that SmartRAG with a quantized 1.7B-parameter backbone achieves multi-hop reasoning performance competitive with models up to 18$\times$ larger, while running entirely on commodity smartphones within practical memory and latency envelopes.
責任ある AI に関するグローバル指標: 2026 年レポート
責任ある AI に関するグローバル指標 (GIRAI) は、AI の倫理に関するユネスコ勧告などの人権に基づく枠組みに基づいて、各国が責任ある AI への取り組みを法的強制力のある保護、制度的能力、救済メカニズムにどのように変換しているかを調査しています。 GIRAI 2026 では、インクルージョンとダイバーシティ、倫理と持続可能性、労働とスキル、信頼と安全、公共サービスにおける AI の利用という 5 つの側面にわたってこれらを評価します。 135 の国レベルの研究者からなるグローバル ネットワークは、政府の AI 政策と導入 (17 指標)、市民社会の関与 (5 指標)、実現条件 (15 指標)、および許容できないリスクの AI システムの政府導入の文書化された事例の 3 つの柱にわたって整理された 38 の指標に関する 68,138 のデータ ポイントを収集および評価しました。データは 2023 年 11 月から 2025 年 9 月までカバーされています。すべての国をランク付けするための指標から 100 点のスコアが導出されます。調査結果によると、責任ある AI ガバナンスは拡大しており、135 か国中 126 か国が 17 の AI 政策指標にわたって少なくとも 1 つの政府政策またはイニシアチブを持っていますが、これが意味のある保護につながることはあまりありません。例えば、グローバル・サウス諸国は、初版以来、フレームワークを伴う指標の新規事例 306 件のうち 203 件を占めていますが、そのフレームワークの 78% は依然として法的拘束力を持たないのに対し、グローバル・ノース諸国の 42% は拘束力を持っていません。 AI ガバナンスに対する政府の取り組みも、自国のアルゴリズムには及んでいません。一方、透明性と説明可能性が最もパフォーマンスの高い指標であり、58% の国が何らかの枠組みを持っていますが、政府アルゴリズムの公開を必要としている国は 18% のみです。許容できないリスクを伴う AI システムを政府が導入しているという信頼できる証拠も 35 か国で見つかりました。これらの調査結果は、責任ある AI ガバナンスがフレームワークの採用を超えて、強制力のある権利に基づく保護、リソースを備えた監視機関、アクセス可能な救済に向けて移行する必要があることを示しています。
原文 (English)
Global Index on Responsible AI: 2026 Report
Grounded in human rights-based frameworks such as the UNESCO Recommendation on the Ethics of AI, the Global Index on Responsible AI (GIRAI) examines how countries translate responsible AI commitments into enforceable protections, institutional capacity, and redress mechanisms. GIRAI 2026 assesses these across five dimensions: Inclusion and Diversity, Ethics and Sustainability, Labour and Skills, Trust and Safety, and AI Use in Public Service. A global network of 135 country-level researchers collected and assessed 68,138 data points on 38 indicators organised across three pillars: government AI policy and implementation (17 indicators), civil society engagement (5), enabling conditions (15), and documented cases of government deployment of unacceptable-risk AI systems. The data covers November 2023 to September 2025. A score of 100 is derived from the indicators to rank all countries. Findings show that while responsible AI governance is expanding, with 126 of 135 countries having at least one government policy or initiative across the 17 AI Policy indicators, this does not often translate into meaningful protection. For instance, Global South countries account for 203 of 306 new cases of indicators with frameworks since the first edition, yet 78% of their frameworks remain non-binding compared with 42% in the Global North. Government commitment to AI governance also does not extend to their own algorithms: whereas Transparency and Explainability is one of the strongest performing indicators, with 58% of countries having some framework, only 18% require Public Disclosure of Government Algorithms. Credible evidence of government deployment of unacceptable-risk AI systems was also found in 35 countries. These findings show that responsible AI governance must move beyond framework adoption toward enforceable rights-based protections, resourced oversight institutions, and accessible redress.
ブラックボックスから実行可能なロジックへ: Prolog Expert システムによる説明可能な強化学習
トレーニングされた深層強化学習ポリシーはブラック ボックスであり、その動作を再現し、人間が読み取り、ロジック エンジンを実行し、オプティマイザーが編集できる実行可能なロジック プログラムとして書き換えることによって説明可能にすることができるかどうかを考えます。我々は、凍結された近接ポリシー最適化教師を抽出し、古典的なリレーショナル学習の方法でその決定から順序付けされたルールリストを誘導し、すべての決定が既製の論理エンジンによって実行される Prolog プログラムとして結果を出力する 3 段階のポストホック変換を提示します。後続の拡張ステージではルール ベースが編集され、ポリシー評価で収益の増加が認定された場合にのみ編集が受け入れられます。私たちは 4 つの保証を証明します。リターンロス境界により、抽出されたプログラムは有限マルコフ決定プロセスで機械検査可能な証明書になり、拡張ループは単調に改善して終了します。連続観測設定については、変換がそもそも可能かどうかを答えます。命題閾値インスタンス化は、不一致 O(1/B) と同じレートで閉じるリターン ギャップを伴い、解像度 B が増加するにつれてネットワークを任意の忠実度に変換します。一致する下限は、斜めの決定境界の観測次元でコストが指数関数的であることを示します。経験的に、16,944 の到達可能な状態を持つ 2 つの部屋の鍵とドアのタスクでは、拡張された Prolog プログラムは、すべてのシードで正確に最適なリターンを達成し、予算に制限のある体制では、10 シードのうち 10 の正確なリターンで確率論的教師を超えます。 3 つの連続制御タスクでは、発行されたプログラムがネットワークを置き換え、Acrobot のノイズ内のニューラル教師を 11 の節でマッチングし、CartPole ではリターンの約 97% を回復しますが、より詳細な制御の LunarLander では部分的にのみ回復し、まさに指数関数的下限が予測する天井に達します。
原文 (English)
From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems
A trained deep reinforcement learning policy is a black box, and we ask whether it can be made explainable by rewriting it as an executable logic program that reproduces its behaviour and that a person can read, a logic engine can run, and an optimizer can edit. We present a three-stage post-hoc transformation that extracts a frozen proximal policy optimization teacher, induces an ordered rule list from its decisions in the manner of classical relational learning, and emits the result as a Prolog program whose every decision is executed by an off-the-shelf logic engine; a subsequent expansion stage edits the rule base and accepts an edit only when policy evaluation certifies a return increase. We prove four guarantees. A return-loss bound makes the distilled program a machine-checkable certificate in a finite Markov decision process, and the expansion loop improves monotonically and terminates. For the continuous-observation setting we answer whether the conversion is possible at all: the propositional threshold instantiation converts the network to arbitrary fidelity as the resolution B grows, with disagreement O(1/B) and a return gap that closes at the same rate, and a matching lower bound shows the cost is exponential in the observation dimension for an oblique decision boundary. Empirically, on a two-room key-and-door task with 16,944 reachable states the expanded Prolog program attains exact optimal return in every seed and, in a budget-capped regime, exceeds the stochastic teacher on exact return in ten of ten seeds. On three continuous-control tasks the emitted program substitutes the network, matching the neural teacher within noise on Acrobot with eleven clauses and recovering about 97% of its return on CartPole, while on the finer-control LunarLander it recovers only partially, exactly the ceiling the exponential lower bound predicts.
Automated Reinforcement Learning: An Overview
Reinforcement Learning and, recently, Deep Reinforcement Learning are popular methods for solving sequential decision-making problems model…
CarbonNet: How Computer Vision Plays a Role in Climate Change? Application: Learning Geomechanics from Subsurface Geometry of CCS to Mitigate Global Warming
We introduce a new approach using computer vision to predict the land surface displacement from subsurface geometry images for Carbon Captu…
Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
Discovering the semantics of multimodal utterances is essential for understanding human language and enhancing human-machine interactions.…
Posts of Peril: Detecting Information About Hazards in Text
Socio-linguistic indicators of affectively-relevant phenomena, such as emotion or sentiment, are often extracted from text to better unders…
Lost in Transmission: An Information-Theoretic Account of Unsupervised Software Traceability
Traceability remains a critical capability to ensure system reliability, maintainability, and compliance in modern software development. Al…
Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
Model merging aggregates Large Language Models (LLMs) finetuned on different tasks into a stronger one. However, parameter conflicts betwee…
A Survey on Knowledge-Oriented Retrieval-Augmented Generation
Retrieval-Augmented Generation (RAG) has gained significant attention in recent years for its potential to enhance natural language underst…
A Survey on Unlearnable Data
Unlearnable data (ULD) has emerged as an innovative defense technique to prevent machine learning models from learning meaningful patterns…
OMAC: A Holistic Optimization Framework for LLM-Based Multi-Agent Collaboration
Agents powered by advanced large language models (LLMs) have demonstrated impressive capabilities across diverse complex applications. Rece…
Data Balancing Strategies: A Systematic Survey of Resampling and Augmentation Methods
Imbalanced datasets, where one class significantly outnumbers others, remain a persistent challenge in machine learning, often biasing pred…
Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning
Effective generalization in language models depends critically on the diversity of their training data. Yet existing diversity metrics ofte…
Learning MMSE Filters for OFDM Channel Estimation: Attention Transformer Gains at Linear Inference
In orthogonal frequency division multiplexing (OFDM), accurate channel estimation is crucial. Classical signal processing-based approaches,…
OV-MAP: Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots
We introduce OV-MAP, a novel approach to open-world 3D mapping for mobile robots by integrating open-features into 3D maps to enhance objec…
Sequential Attention-based Sampling for Histopathological Analysis
Deep neural networks are increasingly applied in automated histopathology. Yet, whole-slide images (WSIs) are often acquired at gigapixel s…
Can Interpretation Predict Behavior on Unseen Data?
Interpretability research often predicts model responses to targeted mechanistic interventions. But can we predict responses to unseen inpu…
Attentions Under the Microscope: A Comparative Study of Resource Utilization for Variants of Self-Attention
As large language models (LLMs) and visual language models (VLMs) grow in scale and application, attention mechanisms have become a central…
Symmetric Behavior Regularized Policy Optimization
Behavior Regularized Policy Optimization (BRPO) leverages asymmetric divergence regularization to mitigate distribution shift in offline re…
DCSCR: A Class-Specific Collaborative Representation based Network for Image Set Classification
Image set classification (ISC), which can be viewed as a task of comparing similarities between sets consisting of unordered heterogeneous…
"Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness
Homelessness is a persistent social challenge, impacting millions worldwide. Over 876,000 people experiencing homelessness (PEH) were recor…
Is "Knowing It's Malicious Enough?" Evaluating LLMs for Fine-Grained Malware Behavior Auditing
Automated malware classifiers achieve strong detection performance, but auditing requires more than flagging a sample: analysts must explai…
From Evidence to Trajectory: Abductive Reasoning Path Synthesis for Retrieval-Augmented Generation Agents Development
Retrieval-augmented generation (RAG) agent development is hindered by the lack of executable ground-truth agent-environment interaction tra…
STAC: When Innocent Tools Form Dangerous Chains for LLM Agents
As LLMs advance into autonomous agents with tool-use capabilities, they introduce security challenges that extend beyond traditional conten…
RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
In Vision-Language-Actionf(VLA) models, robustness to real-world perturbations is critical for deployment. Existing methods target simple v…
Spatiotemporal Knowledge Graphs as Persistent Scene Memory for Embodied Question Answering
Vision-language models (VLMs) demonstrate strong image-level scene understanding, but reasoning over long egocentric video remains costly:…
Probing the Difficulty Perception Mechanism of Large Language Models
Large language models (LLMs) are increasingly deployed on complex reasoning tasks, yet little is known about their ability to internally ev…
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
Layer pruning has emerged as a widely adopted technique for improving the efficiency of large language models (LLMs). Although existing met…
BBOPlace-Bench: Benchmarking Black-Box Optimization for Chip Placement
Chip placement is a vital stage in modern chip design, and black-box optimization (BBO) has been applied to it for decades. Early BBO effor…
InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames
Transformer-based autoregressive models have emerged as a unifying paradigm across modalities such as text and images, but their extension…
CORE -- A Cell-Level Coarse-to-Fine Image Registration Engine for Multi-stain Image Alignment
Accurate and efficient registration of whole slide images (WSIs) is essential for high-resolution, nuclei-level analysis in multi-stained t…
ProDER: A Continual Learning Approach for Fault Prediction in Evolving Smart Grids
As smart grids evolve to meet growing energy demands and modern operational challenges, the ability to accurately predict faults becomes in…
GRIP: In-Parameter Graph Reasoning through Fine-Tuning Large Language Models
Large Language Models (LLMs) have demonstrated remarkable capabilities in modeling sequential textual data and generalizing across diverse…
DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks
Vision-Language Models (VLMs) show great promise for autonomous driving, but their suitability for safety-critical scenarios is largely une…
BUSTR: Descriptor-Aware Vision-Language Learning for Breast Ultrasound Report Generation
Breast ultrasound (BUS) reporting relies on clinically meaningful lesion descriptors, including BI-RADS category, lesion shape, margin, ech…
SONAR: Spectral-Contrastive Audio Residuals for Generalizable Deepfake Detection
Deepfake (DF) audio detectors still struggle to generalize to out of distribution inputs. A central reason is spectral bias, the tendency o…
When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
Frontier language models increasingly participate in conversations about distress and mental health, yet the mechanisms that generate anthr…
Computing Evolutionarily Stable Strategies in Imperfect-Information Games
We present an algorithm for computing evolutionarily stable strategies (ESSs) in symmetric perfect-recall extensive-form games of imperfect…
mHC-GNN: Manifold-Constrained Hyper-Connections for Graph Neural Networks
Graph Neural Networks (GNNs) suffer from over-smoothing in deep architectures and expressiveness bounded by the 1-Weisfeiler-Leman (1-WL) t…
Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
Large language models (LLMs) demonstrate strong capabilities across a wide range of complex tasks and are increasingly deployed at scale, p…
ReMIND: Orchestrating Modular Large Language Models for Controllable Serendipity A REM-Inspired System Design for Emergent Creative Ideation
Large language models (LLMs) are increasingly used not only for problem solving but also for creative ideation; however, generating ideas t…
Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models
Adapting large pre-trained language models to downstream tasks often entails fine-tuning millions of parameters or deploying costly dense w…
Hybrid Mamba-Attention Neural Architecture for Channel Estimation
This paper proposes a hybrid Mamba-attention neural architecture to achieve improved channel estimation for orthogonal frequency-division m…
Li-ViP3D++: Query-Gated Deformable Camera-LiDAR Fusion for End-to-End Perception and Trajectory Prediction
End-to-end perception and trajectory prediction from raw sensor data is one of the key capabilities for autonomous driving. Modular pipelin…
CAM: A Causality-based Analysis Framework for Multi-Agent Code Generation Systems
Despite the remarkable success that Multi-Agent Code Generation Systems (MACGS) have achieved, the inherent complexity of multi-agent archi…
SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-body Manipulation
Simulating deformable objects under rich interactions remains a fundamental challenge for real-to-sim robot manipulation, with dynamics joi…
Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals
Evaluating mathematical reasoning in LLMs is constrained by limited benchmark sizes and inherent model stochasticity, yielding high-varianc…
UnMaskFork: Test-Time Scaling for Masked Diffusion via Deterministic Action Branching
Test-time scaling strategies have effectively leveraged inference-time compute to enhance the reasoning abilities of Autoregressive Large L…
Thermodynamic Limits of Physical Intelligence
Modern AI systems achieve remarkable capabilities at the cost of substantial energy consumption. To connect intelligence to physical effici…
GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning
Reinforcement learning (RL) has become a central post-training paradigm for large language models (LLMs), but its performance is highly sen…
Long Range Frequency Tuning for QML
Angle-encoded variational quantum circuits admit a truncated Fourier series representation of their output, but approximating functions wit…
Breaking the Factorization Barrier in Diffusion Language Models
Diffusion language models theoretically allow for efficient parallel generation but are practically hindered by the ``factorization barrier…
Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet Training
Generative Flow Networks (GFlowNets) enable fine-tuning large language models to approximate reward-proportional posteriors, but they remai…
IoUCert: Robustness Verification for Anchor-based Object Detectors
While formal robustness verification has seen significant success in image classification, scaling these guarantees to object detection rem…
No Certificate, No Categorical Speech Act: A Brouwerian Assertibility Constraint for Public Reason
Generative AI can convert uncertainty into authoritative-seeming verdicts, displacing the justificatory work on which democratic epistemic…
Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models
Quantization has become essential for the efficient deployment of speech processing systems. Although widely studied, most existing quantiz…
L2GTX: From Local to Global Time Series Explanations
Deep learning models achieve high accuracy in time series classification, yet understanding their class-level decision behaviour remains ch…
FormulaCode: Evaluating Agentic Optimization on Large Codebases
Large language model (LLM) coding agents increasingly operate at the repository level, motivating benchmarks that evaluate their ability to…
NanoZK: Privacy-Preserving Verifiable Inference for Large Language Models via Layerwise Zero-Knowledge Proofs
We present NanoZK, a zero-knowledge proof system for verifiable LLM inference: clients and third-party auditors check that a provider execu…
HiCI: Hierarchical Construction-Integration for Long-Context Attention
Long-context language modeling is commonly framed as a scalability challenge of token-level attention, yet local-to-global information stru…
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
Recent advances in large language models have improved the capabilities of coding agents, yet systematic evaluation of complex, end-to-end…
SyriSign: A Parallel Corpus for Arabic Text to Syrian Arabic Sign Language Translation
Sign language is the primary approach of communication for the Deaf and Hard-of-Hearing (DHH) community. While there are numerous benchmark…
Neural Global Optimization via Iterative Refinement from Noisy Samples
Global optimization of black-box functions from noisy samples is a fundamental challenge in machine learning and scientific computing. Trad…
ClawBench: Can AI Agents Complete Everyday Online Tasks?
AI agents may be able to assist with emails and documents, but can they reliably complete everyday online workflows on real websites? Every…
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Accurate evaluation is central to the large language model (LLM) ecosystem, guiding model selection and downstream adoption across diverse…
MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains…
DIB-OD: Preserving the Invariant Core for Robust Heterogeneous Graph Adaptation via Decoupled Information Bottleneck and Online Distillation
Graph Neural Network pretraining is pivotal for leveraging unlabeled graph data. However, generalizing across heterogeneous domains remains…
FETS Benchmark: Foundation Models Enable Scalable and Generalizable Energy Time Series Forecasting
Driven by the transition towards a climate-neutral energy system, accurate energy time series forecasting is critical for planning and oper…
Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Videos via Focused VLM Reasoning Guided by Bayesian Inference
Expressway video anomaly detection is essential for safety management. However, identifying anomalies across diverse scenes remains challen…
A Systematic Investigation of RL-Jailbreaking in LLMs
The evolution of generative models from next-token predictors to autonomous engines of complex systems necessitates rigorous safety hardeni…
A Benchmark for Early-stage Parkinson's Disease Detection from Speech
Early-stage Parkinson's disease (EarlyPD) detection from speech is clinically meaningful yet underexplored, and published results are hard…
RAG は取得が間違っていることを認識しますか?知識の衝突におけるコンテキストのコンプライアンスの診断
検索拡張生成 (RAG) におけるコンテキスト コンプライアンス レジームは、モデルのパラメトリック知識と矛盾する場合でも、取得されたコンテキストが最終的な答えを支配する場合に発生します。正確さだけでは、取得されたコンテキストがそのような矛盾の下でどのように因果関係を持って回答を形成するのかを明らかにすることはできません。推論時に動作し、制御された検索競合に対する介入メカニズムとして機能する信念分解プローブであるコンテキスト駆動分解 (CDD) を導入します。 Epi-Scale ストレス テスト、TruthfulQA の誤解の挿入、およびクロスモデルの再実行を通じて、CDD は 3 つのパターンを明らかにします。 P1: コンテキスト コンプライアンスは上限の敵対的設定で測定可能で、標準 RAG は TruthfulQA 誤解挿入 (N=500) で 15.0% の精度に達します。 P2: 敵対的な精度はモデル ファミリ間で移行します -- CDD は Gemini-2.5-Flash と Claude Haiku/Sonnet/Opus での精度を向上させます -- しかし、根拠と回答の因果結合は移行しません。 CDD は、Gemini-2.5-Flash 上で 64.1% のミスインジェクション因果感度に達しますが、3 つのクロード バリアントすべての感度は [-3%、+7%] の範囲に収まります。これは、クロード側の精度の向上が、明示的な競合解決トレースとは異なるメカニズムを通じて機能していることを示唆しています。 P3: 明示的な競合分解により、時間的ドリフトやノイズの多いディストラクターの下での堅牢性が向上し、完全なエピスケール敵対ベンチマークで CDD は時間的シフトで 71.3%、ディストラクター証拠で 69.9% に達しました。これら 3 つのパターンは、検索品質や単一メソッドの堅牢性の問題とは異なり、標準 RAG を調査および介入できる構造軸としてコンテキスト コンプライアンスを特定し、モデル ファミリと検索パイプライン全体にわたる系統的な研究のために Epi-Scale をリリースする動機となります。
原文 (English)
Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict
Retrieval-Augmented Generation (RAG) is usually evaluated by whether the final answer is correct. Under knowledge conflict, this hides a key question: did the model follow retrieved evidence, rely on its parametric prior, or produce a post-hoc rationale? We study this as context compliance, the regime in which retrieved context controls the answer even when it conflicts with the model's prior knowledge. We introduce Context-Driven Decomposition (CDD), an inference-time diagnostic intervention that elicits contextual and prior answers, isolates the conflicting premise, and records a resolution trace that can be perturbed. Across Epi-Scale stress tests, TruthfulQA misconception injection, and cross-model reruns, CDD makes three behaviors visible. First, misleading retrieval can severely degrade accuracy: under a worst-case TruthfulQA misconception-injection probe, Standard RAG reaches only 15.0%. Second, better answers need not share the same mechanism: CDD improves adversarial accuracy on Gemini-2.5-Flash and shows directional gains across Claude variants, yet trace-perturbation sensitivity is high only on Gemini. Third, explicit decomposition improves controlled-conflict robustness over a conflict-aware instruction baseline on localized factual conflicts, with the clearest margins on Entity Swap (88.0% vs 79.3%) and Logical Contradiction (83.2% vs 75.4%). We frame RAG conflict handling as an observability problem.
MAVEN 多文化テキストからビデオへの生成のためのマルチエージェント フレームワーク
Text-to-Video (T2V) の生成は、視覚的な忠実度において急速に進歩していますが、単一のプロンプト内で複数の文化を忠実に表現する能力はまだ解明されていません。私たちは、単一文化と異文化の両方の T2V 世代における文化忠実度を向上させるために設計されたマルチエージェント プロンプト改良フレームワークである MAVEN を紹介します。 MAVEN は、プロンプトを人、アクション、および場所の次元に分解し、並行または順次に動作する専門のエージェントによって処理されます。体系的な評価をサポートするために、3 つの文化 (中国、アメリカ、ルーマニア)、3 つのアクション カテゴリ、および単一文化シナリオと異文化シナリオの両方に及ぶ、文化に基づいた 243 のプロンプトとそれに対応する 972 のビデオからなる新しいベンチマークを提供します。 CLIP ベースのメトリクス、審査員としての VLM 評価、およびビデオ品質の尺度を組み合わせた評価では、マルチエージェントの改良、特に並行専門化により、視覚的な品質と時間的一貫性を維持しながら文化的関連性が大幅に向上することが示されています。データセットとコードはhttps://github.com/AIM-SCU/CRAFTで入手できます。
原文 (English)
When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation
Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored. We introduce MAVEN, a multi-agent prompt refinement framework designed to improve cultural fidelity in both mono-cultural and cross-cultural T2V generation. MAVEN decomposes prompts into person, action, and location dimensions, handled by specialized agents operating in parallel or sequentially. To support systematic evaluation, we contribute a new benchmark of 243 culturally grounded prompts and 972 corresponding videos, spanning three cultures (Chinese, American, Romanian), three action categories, and both mono-cultural and cross-cultural scenarios. Evaluations combining CLIP-based metrics, VLM-as-judge assessments, and videoquality measures show that multi-agent refinement, particularly parallel specialization, significantly improves cultural relevance while preserving visual quality and temporal consistency. The dataset and code are available at https://github.com/AIM-SCU/MAVEN
SynGR: Unleashing the Potential of Cross-Modal Synergy for Generative Recommendation
Generative Recommendation (GR) has emerged as a promising paradigm by formulating item recommendation as a sequence-to-sequence generation…
すべてのコンポーネントはルックアップです: 単一の分解からのトークンの帰属と構成
変圧器の機構的な解釈には、どのコンポーネントが重要であるかだけでなく、それらのコンポーネントが予測を生成する計算ルートにどのように組み込まれるかを特定する必要があります。アテンションと MLP は両方とも、共有キーと値のテンプレート $\phi(S)U$ に従います。この構造を利用して、両方のサブレイヤーを介してクレジットを分解する後方再帰である Unpack を開発し、任意の 2 つのコンポーネント間の相互作用強度、K/Q/V 構成ラベルを持つ名前付きエンドツーエンド パス、および単一の前方パスからのトークンごとの属性を、介入、勾配、または補助トレーニングなしで生成します。間接的なオブジェクト識別タスクで評価します。 GPT-2 small では、このメソッドは Wang らによって説明されている 3 つの構成接続すべてを回復します。 (2023)、各接続 (K、Q、または V) のモード固有のルーティングを含みます。単純なコピーを超えたトークンレベルの帰属をテストするために、同じ分解で同じ名前が 2 つ出現することを比較します。最初の言及は強い信用を保持しますが、重複検出位置は抑制されます。これは、一致するコントロール プロンプトには存在しないパターンです。 160M から 6.9B パラメータの Pythia ファミリ全体にわたって、この抑制パターンはすべてのスケールで一貫して回復されており、この手法がグラウンド トゥルース回路ラベルなしで機構構造を追跡していることが実証されています。コードは https://github.com/Fun-Cry/unpacklm で入手できます。
原文 (English)
Every Component is a Lookup: Token Attribution and Composition from a Single Decomposition
Mechanistic interpretability of transformers requires identifying not just which components matter but how they compose into the computational route that produced a prediction. Both attention and MLP follow a shared key-value template $\phi(S)U$. We exploit this structure to develop Unpack, a backward recursion that decomposes credit through both sublayers, producing interaction strengths between any two components, named end-to-end paths with K/Q/V composition labels, and per-token attribution, all from a single forward pass, without intervention, gradients, or auxiliary training. The interaction scores are causally grounded: across the Pythia-deduped family from 160M to 6.9B parameters, a component's score predicts the perplexity increase when its communication is ablated (within-layer Spearman $\rho = 0.72$ to $0.96$). The composition paths surface all three connections of the indirect-object-identification circuit of Wang et al. (2023), including the mode-specific routing of each: rerooting at the Name Mover heads, S-Inhibition is the strongest query-side input and falls to rank 10 on the key side, a distinction no single per-token or per-component score can express. The same procedure applied to the greater-than circuit of Hanna et al. (2023), a differently shaped circuit, places the named connections among the top contributors, once layer-0 writers, which carry large credit whatever they feed, are set aside. The decomposition reads out contribution under the realized computation; it complements, rather than performs, causal circuit discovery. The per-token readout is faithful under input perturbation, on par with dedicated attribution methods, and distinguishes circuit mechanism from surface identity: two occurrences of the same name receive radically different credit when only one drives the circuit. Code is available at https://github.com/Fun-Cry/unpacklm.
量子機械学習は本当に必要ですか?: 多次元の実証研究
コンピューター ビジョンの急速な成長とますます複雑になる画像認識タスクにより、古典的な機械学習モデルの基本的な計算上の限界が明らかになり、新たなパラダイムとして量子コンピューティングの探求が促進されています。この論文では、MNIST 手書き数字データセット上の画像認識のための古典機械学習モデルと量子機械学習モデルの包括的なベンチマーク研究を紹介し、従来のモデルである古典サポート ベクター マシン (CSVM) と量子サポート ベクター マシン (QSVM) と、ディープ ニューラル ネットワーク モデルである古典畳み込みニューラル ネットワーク (CCNN) と量子畳み込みニューラル ネットワーク (QCNN) の両方を、分類精度、計算能力の 4 つのパフォーマンス次元にわたって評価します。実行時間、パラメータ数、メモリ要件。実験は、特徴の次元とサンプル サイズの両方に応じて、CPU と GPU の実行環境全体で実施され、制御された多次元の比較を提供して、以前の研究のギャップに対処します。 SVM ベースのモデルの場合、QSVM は一貫して精度において CSVM を上回り、1,000 サンプルで $\sim$ 0.90 に対して $\sim$ 0.85 に達しますが、計算コストは高くなります。 10 量子ビットの特徴数と 200 ~ 500 の範囲のサンプル サイズが、精度と実行時間のバランスをとる実用的な動作点として現れます。ニューラル ネットワーク モデルの場合、CCNN と QCNN は同等の分類精度を実現し、64 の特徴と 60,000 サンプルでどちらも 0.96 を超えていますが、QCNN はパラメーターとメモリの効率が大幅に優れており、特徴数が多い場合には CCNN よりも $\sim$ 94\% 少ないパラメーターと $\sim$ 75\% 少ないメモリを必要としますが、実行時間は長くなります。どちらのモデル ファミリでも、量子モデルは、特徴の次元やサンプル サイズが増加するにつれて、精度のマージンが大きくなり、一貫して古典的なモデルよりも優れています。
原文 (English)
Do We Really Need Quantum Machine Learning?: A Multidimensional Empirical Study
The rapid growth of computer vision and increasingly complex image recognition tasks has exposed fundamental computational limitations of classical machine learning models, motivating the exploration of quantum computing as an emerging new paradigm. This paper presents a comprehensive benchmarking study of classical and quantum machine learning models for image recognition on the MNIST handwritten digit dataset, evaluating both traditional models, a Classical Support Vector Machine (CSVM) and a Quantum Support Vector Machine (QSVM), and deep neural network models, a Classical Convolutional Neural Network (CCNN) and a Quantum Convolutional Neural Network (QCNN), across four performance dimensions: classification accuracy, computational runtime, parameter count, and memory requirements. Experiments are conducted as functions of both feature dimensionality and sample size, and across CPU and GPU execution environments, providing a controlled, multidimensional comparison to address gaps in prior work. For the SVM-based models, QSVM consistently outperforms CSVM in accuracy, reaching $\sim$ 0.90 versus $\sim$ 0.85 at 1,000 samples, with a higher computational cost. A feature count of 10 qubits and a sample size in the range of 200 -- 500 emerge as practical operating points that balance accuracy and runtime. For the neural network models, CCNN and QCNN achieve comparable classification accuracy, both exceeding 0.96 at 64 features and 60,000 samples, yet QCNN offers superior parameter and memory efficiency at higher feature counts, while incurring higher runtime. Across both model families, quantum models consistently outperform classical models by greater margins in accuracy as feature dimensionality or sample size increases.
データセットの価値はいくらですか?スケーリング則、Vendi スコア、および行列スペクトル関数
ニューラル スケーリングの法則はデータセットのサイズを通じてデータを評価しますが、Vendi スコアは量子エントロピーを使用してデータセットの値を測定します。一般的なニューラル スケーリング則の目標と Vendi スコアの両方がサブモジュールであることを示します。さらに、Vendi スコアが、行列スペクトル関数と呼ばれるより広範なクラスのサブモジュラー目標の特殊なケースであることを示します。これには、決定的 (DPP) 目標や他の多くの目標も含まれます。また、弱行列単調関数を導入し、それがどのように弱部分モジュール行列スペクトル関数につながるかを示し、データ評価のための幅広い実用的な目的をもたらします。私たちは、貪欲な最適化中に繰り返される固有分解を回避する永年方程式ベースの更新を開発し、$m$ 次元の埋め込みに対する限界ゲイン評価を Oracle クエリと比較して $O(m)$ 係数だけ削減します。これにより、経験的に平均約 35,000 倍の高速化が得られ、ImageNet-1K スケールのデータセットで Vendi スコアの直接最適化が可能になります。このようにして可能になったので、Vendi スコア、DPP、施設の場所、および 3 つの新しいマトリックス スペクトル バリアントを含む、固定サイズ、クラスバランス、および固定トレーニング予算体制の下で、いくつかの目標がホールドアウト テスト パフォーマンスのトレーニング サブセットの値をどの程度正確に予測するかを比較します。複数のデータセットにわたって、施設の位置が最も優れたパフォーマンスを発揮します。また、直接最適化では、Vendi スコアは中程度のスコア範囲では予測的ですが、目標をより高い値に押し上げると、下流のパフォーマンスの代用として機能しなくなる可能性があることも明らかになりました。また、均一でランダムな固定サイズのサブセットは、制約がなく、クラスバランスが取れていても、評価スコアと保持されたパフォーマンスの両方で著しく集中していることもわかります。最後に、サイズ、クラスのバランス、トレーニング予算だけがデータの価値を決定するわけではないことを示します。これらの要因を制御した場合でも、パフォーマンスは良い状態から悪い状態まで滑らかに変化します。
原文 (English)
How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions
Neural scaling laws appraise data through dataset size, while the Vendi Score uses quantum entropy to measure dataset value. We show both that common neural-scaling-law objectives and the Vendi Score are submodular. We further show that the Vendi Score is a special case of a broader class of submodular objectives that we call matrix spectral functions. This also includes determinantal (DPP) objectives, as well as many others. We also introduce weakly matrix monotone functions and show how they lead to weakly submodular matrix spectral functions, yielding a broad family of practical objectives for data appraisal. We develop secular-equation-based updates that avoid repeated eigendecompositions during greedy optimization, reducing marginal-gain evaluation for $m$-dimensional embeddings by an $O(m)$ factor relative to oracle queries. This yields an average empirical speedup of about 35,000x, making direct optimization of the Vendi Score feasible on ImageNet-1K-scale datasets. Thus enabled, we compare how well several objectives predict the value of training subsets for held-out test performance under fixed-size, class-balanced, and fixed training-budget regimes, including the Vendi Score, DPPs, facility location, and three new matrix spectral variants. Across multiple datasets, facility location performs the best. Direct optimization also reveals that, while the Vendi Score is predictive over moderate score ranges, pushing the objective to higher values can make it a poor downstream performance proxy. We also find that uniformly at random fixed-size subsets, both unconstrained and class-balanced, are remarkably concentrated in both appraisal scores and held-out performance. Finally, we show that size, class balance, and training budget do not alone determine data value: even when controlling for these factors, performance ranges smoothly from good to bad.
ランダム化された幾何学的プローブによる特異点を意識した最適化: 安定した非滑らかな最適化に向けて
深層学習の最適化は、滑らかな損失ランドスケープの仮定に大きく依存しています。この条件は、ReLU アクティベーションや量子化演算子などの非滑らかなコンポーネントにより、最新のアーキテクチャによって体系的に違反されています。このような非滑らかな領域では、Adam などの適応オプティマイザは、Clarke 微分内の矛盾する信号によって引き起こされる勾配チャタリングや激しい振動に悩まされ、収束性の低下や次善の一般化につながります。これに対処するために、局所的な幾何学的不安定性に基づいてステップ サイズを動的に調整することでトレーニングを安定化する新しいオプティマイザーである Singularity-aware Adam (S-Adam) を導入します。私たちの主な貢献は、ランダム化された方向導関数の分散から導出されるクラーク微分直径の計算効率の高い推定器である局所幾何学的不安定性 (LGI) メトリックです。 S-Adam には、滑らかな盆地での高速収束を維持しながら、不安定性の高い領域での更新を減速する適応減衰メカニズム exp(-$\lambda$$\rho$) が組み込まれています。微分包含を使用した厳密な収束解析を提供し、S-Adam が最適な O(1/$\sqrt(T)$) レートで ($\delta$,$\epsilon$)-Clarke 定常点にほぼ確実に収束することを証明しました。量子化対応トレーニング (QAT) と高ノイズの小規模バッチ学習に関する経験的評価では、S-Adam が常に AdamW および Prox-SGD を上回り、勾配振動を効果的に軽減しながら、CIFAR-100 で最大 6 パーセント、TinyImageNet で 3 パーセントの精度向上を達成していることが実証されています。
原文 (English)
Singularity-aware Optimization via Randomized Geometric Probing: Towards Stable Non-smooth Optimization
Deep learning optimization relies heavily on the assumption of smooth loss landscapes, a condition systematically violated by modern architectures due to non-smooth components such as ReLU activations and quantization operators. In such non-smooth regimes, adaptive optimizers such as Adam suffer from gradient chattering, violent oscillations caused by conflicting signals within the Clarke subdifferential, leading to poor convergence and suboptimal generalization. To address this, we introduce Singularity-aware Adam (S-Adam), a novel optimizer that stabilizes training by dynamically modulating step sizes based on local geometric instability. Our key contribution is the Local Geometric Instability (LGI) metric, a computationally efficient estimator of the Clarke subdifferential diameter derived from the variance of randomized directional derivatives. S-Adam incorporates an adaptive damping mechanism exp(-$\lambda$$\rho$) that decelerates updates in high-instability regions while preserving fast convergence in smooth basins. We provide a rigorous convergence analysis using differential inclusions, proving that S-Adam converges almost surely to ($\delta$,$\epsilon$)-Clarke stationary points at the optimal O(1/$\sqrt(T)$) rate. Empirical evaluations on Quantization-Aware Training (QAT) and high-noise small-batch learning demonstrate that S-Adam consistently outperforms AdamW and Prox-SGD, achieving accuracy gains of up to +4.54% on CIFAR-100 and +4.27% on TinyImageNet while effectively mitigating gradient oscillations.
Beyond Access: Guided LLM Scaffolding for Independent Learning in Undergraduate Statistics
Large language models (LLMs) are increasingly entering students' learning practices, but their educational value depends on whether they su…
CyberGym-E2E: AI エージェントのエンドツーエンドのサイバーセキュリティ機能のためのスケーラブルな現実世界のベンチマーク
AI は、ソフトウェアの脆弱性を自律的に検出、分析、修復できるシステムを可能にすることで、サイバーセキュリティを変革する可能性を秘めています。しかし、AI システムの既存のサイバーセキュリティ評価は規模や範囲が限られており、現実世界のソフトウェアの脆弱性の発見と修復のエンドツーエンドのライフサイクルを捉えることができません。このギャップに対処するために、私たちは、脆弱性の発見、PoC 生成、パッチ生成のライフサイクル全体にわたって AI エージェントの能力を包括的に評価する、大規模かつ現実的なエンドツーエンドのサイバーセキュリティ ベンチマークである CyberGym-E2E を提案します。 CyberGym-E2E は、オープンソースの脆弱性データを現実的な評価環境に変換するための自動化されたエージェント強化パイプラインを構築するため、包括的でスケーラブルです。現在、ベンチマークは、139 の異なるオープンソース プロジェクトにわたる 920 件の実際の脆弱性で構成されています。
原文 (English)
CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities
AI has the potential to transform cybersecurity by enabling systems that can autonomously detect, analyze, and remediate software vulnerabilities. However, existing cybersecurity evaluations of AI systems are limited in scale or scope, and fail to capture the end-to-end lifecycle of real-world software vulnerability discovery and remediation. To address this gap, we propose CyberGym-E2E, a large-scale and realistic end-to-end cybersecurity benchmark that comprehensively evaluates AI agents' abilities across the full lifecycle of vulnerability discovery, PoC generation, and patch generation. CyberGym-E2E is comprehensive and scalable, as we build an automated, agent-enhanced pipeline for transforming open-source vulnerability data into realistic evaluation environments. Currently, the benchmark consists of 920 real-world vulnerabilities across 139 different open-source projects.
Improving Answer Extraction in Context-based Question Answering Systems Using LLMs
Question answering (QA) systems have achieved notable progress with the advent of large language models (LLMs). However, they still face ch…
TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies
Robot manipulation alternates between low-risk transit phases that call for fast execution and high-risk contact stages that demand slow, p…
FAIR-Calib: 拡散大規模言語モデルのトレーニング後の量子化のためのフロンティアを意識した不安定性再重み付けキャリブレーション
拡散大規模言語モデル (dLLM) は、トークンを反復的に改良しますが、それらを不可逆的にコミットするため、初期の決定が作成された後でも脆弱なままである「安定性ラグ」が生じます。トレーニング後量子化 (PTQ) エラーにより、書き込みフロンティアでのこれらの境界線の決定が簡単に反転され、その後永続的にロックインされて増幅されることが明らかになりました。これに対処するために、私たちは dLLM 用の 2 段階 PTQ フレームワークである Frontier-Aware Instability-Reweighted Calibration (FAIR-Calib) を提案します。ステージ I では、完全精度の教師を調べて、フロンティア ヒットとマスクされたステージの信頼性を組み合わせた位置を事前に推定します。ステージ II では、再重み付けされた隠れ状態 MSE を最小限に抑えることでオフポリシーのレイヤーごとのキャリブレーションを実行し、高価なエンドツーエンドの拡散ロールアウトを必要とせずに脆弱なフロンティア状態の保護を効果的に優先します。さらに理論的には、重み付けされた目標が出力 KL 発散の代用として正当化されます。経験的に、FAIR-Calib は常に LLaDA および Dream (W4A4) の最先端のベースラインを上回り、フロンティアの意思決定の反転を大幅に削減し、さまざまなベンチマークにわたるコミット後の不一致を抑制します。
原文 (English)
FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models
Diffusion Large Language Models (dLLMs) refine tokens iteratively but commit them irreversibly, leading to a "stability lag" where early decisions remain fragile even after being written. We reveal that Post-Training Quantization (PTQ) error easily flips these borderline decisions at the write frontier, which are then permanently locked in and amplified. To address this, we propose Frontier-Aware Instability-Reweighted Calibration (FAIR-Calib), a two-stage PTQ framework for dLLMs. Stage I probes a full-precision teacher to estimate a position prior that combines frontier hits and masked-stage reliability. Stage II performs off-policy, layer-wise calibration by minimizing a reweighted hidden-state MSE, effectively prioritizing the protection of fragile frontier states without requiring expensive end-to-end diffusion rollouts. We further theoretically justify our weighted objective as a surrogate for output KL divergence. Empirically, FAIR-Calib consistently outperforms state-of-the-art baselines on LLaDA and Dream (W4A4), significantly reducing frontier decision flips and suppressing post-commit mismatches across diverse benchmarks.
Phantom Transitions in Language Model Fine-Tuning: A Density-Matrix Analysis
Fine-tuning a language model often fails silently when its correct completion must outrank a near-synonym competitor. Cross-entropy loss fa…
FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In…
BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression
Interactive video world models commonly convert bidirectional video generators into causal autoregressive systems through control fine-tuni…
RoVE: 相対位置依存の値経路に対するロータリー値の埋め込みの注意
Rotary Position Embeddings (RoPE) は、アテンション スコアを位置相対にしますが、値経路の位置をブラインドのままにします。つまり、値トークンによって送信されるメッセージは、クエリからの距離に関係なく同じです。キーと同時に値を回転させることで値を位置に依存させるパラメータフリーの変更である RoVE を提案し、それが RoPE の注意を注意深い畳み込みに変えることを示します。この新しい視点は、コンピューター ビジョン、ロボット工学、最新の LLM アーキテクチャにわたる同じ操作のいくつかの独立した定式化を統合します。トレーニングされた 1 億 2,400 万および 3 億 5,400 万の GPT-2 モデルは、少数ショットのコンテキスト内学習、分布外のパープレキシティ、および長いコンテキストの取得において RoPE と比較して一貫した経験的利点を示し、長距離の集約を必要とするタスクで最も明確な改善が見られます。
原文 (English)
RoVE: Rotary Value Embeddings Attention for Relative Position-dependent Value Pathways
Rotary Position Embeddings (RoPE) make attention scores position-relative but leave the value pathway position-blind: the message sent by a value token is the same regardless of its distance from the query. We propose RoVE, a parameter-free modification that makes values position-sensitive by rotating them simultaneously with keys, and show that it turns RoPE attention into attentive convolution. This new perspective unifies several independent formulations of the same operation across computer vision, robotics, and modern LLM architectures. Trained 124M and 354M GPT-2 models show consistent empirical gains over RoPE on few-shot in-context learning, out-of-distribution perplexity, and long-context retrieval, with the clearest improvements on tasks that require long-range aggregation.
Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack
In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full robot learning stack: da…
AI Contagion in Social Networks
We study how artificial intelligence (AI) interacts with social communication networks to shape the stability of collective knowledge. Agen…
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering
Coding agents have become a major mode of software engineering, but the benchmarks we use to compare them were designed in a pre-agent era:…
GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling
Generating visually consistent multi-shot videos remains an open challenge. As videos span more shots, inconsistencies can accumulate acros…
Short-Term Electricity Demand Forecasting for New England: A Comprehensive Machine Learning Benchmark with Weather, Calendar, and COVID-19 Indicators
Accurate short-term electricity demand forecasting is critical for reliable power system operation, energy market planning, and infrastruct…
Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One
A language model's memory can be worse than no memory at all when the model or its interface is disposed to act on it: a memory that keeps…
Learning to Fold: LeHome Challenge 2026 で受賞歴のあるソリューション (オンラインで 1 位、オフラインで 2 位)
私は、両手で衣類をたたむことに関する ICRA 2026 コンテストである LeHome Challenge 2026 に対する私の解決策について説明します。このシステムは、オンライン (シミュレーション) ラウンドで 62 チーム中 1 位となり、現実世界の決勝では 2 位になりました。強化学習ループを使用してビジョン言語アクション (VLA) ポリシーを改善します。ポリシーはそれ自体の価値関数です。アクションを予測する同じネットワークが、成功、進捗状況、およびいくつかのタスク関連の将来の数量も予測します。これらの予測は、利点の推定、実際の失敗の検出、および候補の選択を推進します。この作業のほとんどは、既存の RL アイデアとエンジニアリングおよび最適化の貢献を再結合したもので、これらは 1 つのレシピとして一緒に使用することも、個別に使用することもできます。フローマッチング VLA には AWR + RECAP を組み合わせます。 HuggingFace Hub を介した非同期分散トレーニング/ロールアウト パイプライン。 Thompson サンプリングによる推論時のハイパーパラメータの最適化。カメラアライメントツール、強力な拡張、DAgger のような HIL データ収集を備えた sim-to-real レシピ。
原文 (English)
Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)
I describe my solution to the LeHome Challenge 2026, an ICRA 2026 competition on bimanual garment folding. The system placed 1st of 62 teams in the online (simulation) round and 2nd in the real-world final. It improves a vision-language-action (VLA) policy with a reinforcement-learning loop. The policy is its own value function: the same network that predicts actions also predicts success, progress, and a few task-relevant future quantities, and those predictions drive advantage estimation, live failure detection, and candidate selection. The work mostly recombines existing RL ideas with engineering and optimization contributions that can be used together as one recipe or individually: AWR + RECAP combined for flow-matching VLA; an asynchronous distributed training / rollout pipeline through HuggingFace Hub; inference-time hyperparameters optimization via Thompson sampling; a sim-to-real recipe with camera-alignment tooling, heavy augmentation and DAgger-like HIL data collection.
すべての関係が同じように回転するわけではありません: 視点の堅牢な 3D シーン グラフ生成のための変換を意識したデカップリング
3D シーン グラフ生成 (3DSGG) は、3D シーンを構造化されたオブジェクト-リレーション-オブジェクト グラフとして表現し、空間を理解するためのコンパクトなリレーショナル抽象化を提供します。身体化されたインテリジェンス設定では、同じ 3D シーンがヨー回転によって異なる視点からエージェントによって観察される場合があります。ただし、現在の 3DSGG モデルは、そのような視点のシフトの下で期待される変換動作に従う関係予測を生成できないことがよくあります。この動作は、述語レベルの変換の不均一性に関連する経験的な不一致を明らかにしています。左、前、右、後ろなどの方向述語は観測フレームに合わせて変換されるはずですが、ほとんどの接触、サポート、および意味論的述語 (上に立つ、接続されるなど) は安定したままであるはずです。この不一致を減らすために、我々は、述語変換動作に従って関係推論を分離し、視点安定したオブジェクト表現によってサポートされる視点堅牢な 3DSGG フレームワークである、Transformation-Aware Decoupling (TAD) を提案します。 TAD は、関係推論を 2 つの部分に分解します。1 つは視点間で安定している必要がある手がかりを学習し、もう 1 つは観測フレームとともに変化する必要がある方向の手がかりを学習します。 2 つの部分は、標準のマルチラベル述語予測のためにマージされます。変換固有の記述子とグループ認識の補助監視により、2 つのブランチが相補的な関係の手がかりを捕捉することが促進されます。 3DSSG に関する広範な実験により、TAD は標準ベンチマークの下で競争力のあるパフォーマンスを維持しながら、トレーニング時の回転強化を行わずにヨー視点変更下で最先端の堅牢性を達成できることが示されています。プロジェクト ページは https://tad-predicate.github.io/ で利用できます。
原文 (English)
From Scene-Centric to Observer-Centric: Modeling Observer-Aware Relations for 3D Scene Graph Generation
3D Scene Graph Generation (3DSGG) represents 3D scenes as structured object--relation--object graphs for spatial understanding. In observer-centric spatial perception, the same scene may be expressed under different local observer frames while its structure remains unchanged. However, existing models typically assume a fixed scene-aligned reference frame and may produce semantically inconsistent predictions when the scene is re-expressed in another observer frame. We attribute this failure to the heterogeneous frame dependency of relational predicates. Directional predicates such as $\textit{left}$, $\textit{front}$, $\textit{right}$, and $\textit{behind}$ are $\textbf{Observer-Dependent Relations}$, whereas most contact, support, and semantic predicates, such as $\textit{standing on}$ and $\textit{attached to}$, are approximately $\textbf{Observer-Independent Relations}$. Conventional models do not distinguish these frame responses, leading to degraded relation prediction under observer-frame reorientation. We introduce $\textbf{Observer-Aware Relations (OAR)}$, which combines observer-aware geometric encoding and relation specialization, supported by frame-stable object encoding, for unified multi-label predicate prediction. Experiments on 3DSSG show that OAR consistently outperforms baselines across controlled observer-frame reorientations without training-time frame-reorientation augmentation, while remaining competitive on the standard benchmark. The project page is available at https://oar-predicate.github.io/.
ReactiveBFM: Reactive Closed-Loop Motion Planning Towards Universal Humanoid Whole-Body Control
While current Behavior Foundation Models (BFMs) provide robust control priors for humanoids, they only execute pre-defined reference motion…
TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning
Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands…
Full Bayesian Reinforcement Learning via LF-IBIS
Reinforcement Learning (RL) is a sequential decision-making framework in which an agent learns optimal policies through interaction with an…
The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits
The rapid deployment of AI systems across high-stakes domains has created urgent demand for standardized evaluation, yet the field remains…
Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders
Portuguese remains underrepresented in text embedding evaluation, despite being one of the most widely spoken languages in the world. As a…
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
Despite growing automation, turning a paper into a coherent poster, talk video, and blog piece often remains a labor-intensive last mile. R…
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications
Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in th…
NLPCC 2026 の概要 共有タスク 1: 難易度を考慮した多言語およびマルチモーダルな医療指導ビデオの理解度評価
NLPCC 2023 ~ 2025 年の CMIVQA、MMI-VQA、および M4IVQA の課題に続き、NLPCC 2026 では難易度を考慮した医療指導ビデオ質問応答 (DA-MIVQA) 共有タスクを導入します。DA-MIVQA は、必要な証拠の種類と複雑さに応じて質問を明示的に区別することで、以前の多言語および多モードの医療ビデオ ベンチマークを拡張します。答える。具体的には、単純な質問は字幕ベースのテキストの手がかりから答えられることがよくありますが、複雑な質問には視覚的な根拠、手順の理解、およびクロスモーダルな証拠の統合が必要です。このチャレンジには、単一ビデオでの難易度を考慮した時間的回答グラウンディング (DA-TAGSV)、ビデオ コーパスでの難易度を考慮した時間的回答の取得 (DA-VCR)、およびビデオ コーパスでの難易度を考慮した時間的回答のグラウンディング (DA-TAGVC) の 3 つのトラックが含まれています。データセットは公的医療指導チャンネルから収集され、応急処置、緊急対応、リハビリテーション、看護、一般医学教育などの多様なシナリオをカバーしており、難易度の注釈を付けて手動で検証されています。本稿では、DA-MIVQAの課題動機、データセット構築、評価プロトコル、参加概要、競技結果、代表的なシステムについて紹介する。 DA-MIVQA は、さまざまなテキスト、視覚、時間的、および手順の推論要件の下で、医療指導ビデオ質問応答システムを評価するための実用的なベンチマークを提供します。
原文 (English)
Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation
Following the CMIVQA, MMI-VQA, and M4IVQA challenges in NLPCC 2023--2025, we introduce the Difficulty-Aware Medical Instructional Video Question Answering (DA-MIVQA) shared task for NLPCC 2026. DA-MIVQA extends previous multilingual and multimodal medical video benchmarks by explicitly distinguishing questions according to the type and complexity of evidence required for answering. Specifically, simple questions can often be answered from subtitle-based textual cues, whereas complex questions require visual grounding, procedural understanding, and cross-modal evidence integration. The challenge contains three tracks: Difficulty-Aware Temporal Answer Grounding in Single Video (DA-TAGSV), Difficulty-Aware Video Corpus Retrieval (DA-VCR), and Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC). The dataset is collected from public medical instructional channels, covers diverse scenarios such as first aid, emergency response, rehabilitation, nursing, and general medical education, and is manually verified with difficulty annotations. This paper presents the task motivation, dataset construction, evaluation protocol, participation overview, competition results, and representative systems of DA-MIVQA. DA-MIVQA provides a practical benchmark for evaluating medical instructional video question answering systems under varying textual, visual, temporal, and procedural reasoning requirements.
LLM ガイドによる産業プロセス予測のためのタスク セマンティック フィールド因数分解
プロセス産業は、時系列予測とソフト センシングを利用して、オンラインで測定するのが難しい品質変数を推定します。ラベル付きデータは不足しており、運用体制は頻繁に変更され、シナリオごとにモデルを再トレーニングしたり調整パイプラインを再構築したりするとコストがかかります。このような設定では、多くの場合、変数名、単位、物理的意味、プロセスの役割を記録する変数テーブルとプロセス ドキュメントが提供されます。ただし、標準の時系列バックボーンは通常、入力を匿名の数値列として扱います。また、既存のテキスト拡張手法では、入力変数と予測ターゲットの間の意味論的論理的関係を各数値ウィンドウ内でモデルが利用できるようにすることはほとんどありません。この問題に対処するために、この記事では、大規模言語モデル (LLM) ガイド付きフレームワークであるタスク セマンティック フィールド因数分解 (TSF) を提案します。 TSF は、トレーニング前にタスク プロトコルと変数ドキュメントからタスク セマンティック フィールドを構築し、LLM をオフライン セマンティック構築にのみ使用します。オンライン トレーニングと推論には、従来の時系列バックボーンが残ります。トレーニングと推論中に、現在の数値ウィンドウが変数セマンティクスをアクティブ化するため、セマンティクス情報が各予測に関与し、さまざまな予測ターゲットと動作シフトへの適応をサポートします。複数の複雑な産業予測およびソフト センシング タスクにおいて、TSF は改善された設定で MAE を平均 6.4\% 削減し、最大の削減は 25.5\% に達します。追加されるパラメータは約 1.8 ~ 3.0k のみで、追加のオンライン推論オーバーヘッドは 0.008 ミリ秒/ステップ未満です。これらの結果は、TSF が既存のプロセス ドキュメントを、展開の軽量性を維持しながら、バックボーンとセマンティック ジェネレーター全体にわたって測定可能な予測ゲインに変えることを示しています。
原文 (English)
LLM-Guided Task-Semantic Field Factorization for Industrial Process Forecasting
Process industries rely on time-series forecasting and soft sensing to estimate quality variables that are hard to measure online. Labeled data are scarce, operating regimes change frequently, and retraining models or rebuilding alignment pipelines for each scenario is costly. Such settings often provide variable tables and process documents that record variable names, units, physical meanings, and process roles. However, standard time-series backbones usually treat inputs as anonymous numerical columns. Existing text-enhanced methods also rarely make the semantic-logical relations between input variables and the prediction target available to the model within each numerical window. To address this problem, this article proposes Task-Semantic Field Factorization (TSF), a large language model (LLM)-guided framework. TSF builds a task-semantic field from task protocols and variable documents before training and uses the LLM only for offline semantic construction. Online training and inference are handled by conventional time-series backbones. During training and inference, the current numerical window activates variable semantics, so semantic information participates in each prediction and supports adaptation to different prediction targets and operating shifts. Across multiple complex industrial forecasting and delayed soft-sensing tasks, TSF reduces MAE by 3.6\% on average. Across all dataset--backbone pairs, the macro-average reduction is 2.9\%, with a maximum reduction of 24.9\%. It adds only about 0.7--4.3k parameters, with less than 8\,$\mu$s/sample of additional online inference overhead. These results show that TSF turns existing process documents into measurable forecasting gains across backbones and semantic generators while remaining lightweight for deployment.
Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA
National language models are becoming publicly funded epistemic infrastructure. Public ownership, linguistic specialization, and open weigh…
EHR-MPC: 生成患者デジタル ツインを使用した敗血症治療のための推論時間制御
敗血症は死亡の主な原因ですが、最適な治療方針については依然として議論が続いています。既存の強化学習 (RL) アプローチは、敗血症治療のための固定戦略を学習するため、推論中に変化する臨床目的への適応性が制限されます。私たちは、生成電子医療記録 (EHR) モデルの形式で患者のデジタル ツインをトレーニングすることで、患者のダイナミクスの学習と治療の最適化を切り離すフレームワークである EHRMPC を提案します。デジタル ツインは介入中の臨床経過を予測し、モデル予測制御 (MPC) を可能にして、シミュレーションによる推論時間計画を通じて治療を最適化します。我々は、ポリシー外の重要度サンプリングとポリシー上のシミュレーションベースの評価の両方を使用して、マサチューセッツジェネラルブリガム医療システムの8つの病院にわたる多施設ICU敗血症コホートでEHR-MPCを評価します。 RL ベースラインと比較して、EHR-MPC は同等のオフポリシー パフォーマンスと改善されたシミュレーション パフォーマンスを実現します。 RL とは異なり、この作業は敗血症治療の最適化を学習された患者の動態に対する推論時間の制御として枠組み化し、生成臨床モデルを使用した意思決定のための一般的な枠組みを確立します。
原文 (English)
EHR-MPC: Inference-Time Control for Sepsis Treatment with Generative Patient Digital Twins
Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested. Existing reinforcement learning (RL) approaches learn fixed strategies for sepsis treatment, limiting adaptability to changing clinical objectives during inference. We propose EHRMPC, a framework that decouples learning patient dynamics from optimizing treatment by training a patient digital twin in the form of a generative electronic health record (EHR) model. The digital twin predicts clinical trajectories under interventions and enables model predictive control (MPC) to optimize treatments via inference-time planning over simulations. We evaluate EHR-MPC on a multicenter ICU sepsis cohort spanning 8 hospitals in the Mass General Brigham health system using both off-policy importance sampling and on-policy simulation-based evaluation. Relative to RL baselines, EHR-MPC achieves comparable off-policy performance and improved simulation performance. Unlike RL, this work frames sepsis treatment optimization as inference-time control over learned patient dynamics, establishing a general framework for decision making with generative clinical models.
創造性、誠実さ、計画された忘却は、小さな双曲言語モデルで現れます
言語モデルはスケールに合わせて最適化されていますが、コンパニオン可能というよりは機能的なままであり、アシスタントがコンパニオンにパーソナライズされ、1人のユーザーの記憶を蓄積すると、それは静かに何者かになり、そのユーザーに害を及ぼす特性を静かに獲得することができます。コンパニオンがどのようなものになりつつあるのか、また、それが何になる価値があるのかを判断するための信頼できる手段はありません。訓練を受けた人間の評価者でも答えに同意することはできません (フライス カッパ = 0.074)。ここでは、双曲基質を共有する 3 つの小さな言語モデル (146 M から 3 B のパラメーター) がその質問の両方の半分に答えることを示します。ゼロから訓練された 1 億 4,600 万人の行動監査人は、それらの評価者ができないコンプライアンスのギャップを検出します (バイナリコンプライアンスの精度 90.7%)。その凍結表現の線形読み出しにより、トレーニングでは見られなかった、コンパニオン誘発のお調子者、依存性促進、およびジェネレータファミリーに関する作話された記憶がさらに検出されます(スタイル制御、リーブ 1 ジェネレータ評価での AUROC 0.804 に対し、同じ項目に対するフロンティア ゼロショット ジャッジの場合は 0.721)。クリエイティブなフレームシーダーは、4 つのプロンプトベースラインに対する 311 の決定されたペアワイズ比較の 100% で優先されます。メモリ オペレーティング システムは、設計された忘却 M(t) = S*exp(-lambda*t) を実装します。その予測されたスケルトンと壁紙のパーティションは、4 条件パイロットの選択的検索ゲートの下でのみ出現します。創造性、誠実さ、そして設計された忘却が、信頼できるコンパニオン AI への小規模モデルへのルートを構成します。
原文 (English)
A small language model detects behavioural faithfulness gaps that frontier judges and human raters miss
Whether a language model behaves as it claims is a judgement on which independent human raters cannot agree (Fleiss kappa = 0.074). We show that a small, purpose-built instrument does better. A linear read-out of the frozen representation of a from-scratch 146-million-parameter auditor detects companion-induced sycophancy, dependence-fostering and confabulated memories on generator models unseen in training (AUROC 0.804, leave-one-generator-out, against ground truth fixed at generation, independent of human judgement), where a frontier zero-shot judge on the identical items reaches 0.721 and falls to chance on the most distant family. The auditor's substrate is hyperbolic, and its demonstrated benefit is hierarchical: an ablation isolates the advantage over a matched Euclidean control on multi-domain structure. On this task, behavioural faithfulness is measured not by scale but by a small, purpose-built instrument.
Scalable Visual Pretraining for Language Intelligence
The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms…
LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning
Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each…
Adaptive Compute in Latent World Models: When Depth Helps, Hurts, or Doesn't Matter
Adaptive compute for world models -- early-exit or mixture-of-depths predictors that spend variable depth per rollout step -- presumes that…
加重 k 最近傍回帰およびソフトラベル予測のための正確な認証済みデータ Shapley
Data Shapley は、どのトレーニング ポイントにどのような価値があるのかについての標準的な原則に基づいた回答であり、その k 近傍 (KNN) 特化は実際に展開されるバージョンであり、pyDVL や OpenDataVal などのツールキットに同梱されている正確な推定器です。正確なアルゴリズムは、非重み付き KNN と重み付き KNN 分類で知られていますが、重み付き KNN 回帰とソフトラベル予測は抵抗がありました。唯一の正確な方法は、近傍サイズ K で指数関数的な O(N^K) 総当たり法です。障害は、重み付き回帰予測は 2 つの連立に依存する和の比であり、その正規化分母が以前の多項式アルゴリズムが依存していた加法、しきい値、および重複の構造を壊します。私たちはこのギャップを埋めます。 (i) 加重 KNN 回帰 Data Shapley 用の最初の擬似多項式時間正確アルゴリズム (固定格子精度での N と K の多項式)、結合整数状態 (w の合計、w*y の合計) にわたる計数動的プログラムであり、12,716 個の敵対的インスタンスで不一致がゼロの徹底的な列挙に対して検証されます。 (ii) 86,400 回のチェックにわたって違反が一度もなかった、機械チェック可能な値ごとのエラー証明書を備えた、継続的な分銅と目標に対する認定された FPTAS。 (iii) 無条件の Omega(D_w) 出力サイズの下限とアクセス モデルの硬度の結果を含む複雑さの状況。 (iv) 重み付けされたソフトラベルのマルチクラス拡張。私たちは、オープンソースの CPU 専用ライブラリと、最初の正確な重み付き回帰 Data Shapley のグラウンド トゥルースをリリースします。ダウンストリームの誤ったラベルの検出では、正確な値は、事前に登録された結果であるモンテカルロ データ シャプレー (データセット レベル TOST、n=8、p<10^-4) と統計的に同等です。正確さの価値は代わりに、決定論、認定誤差限界、および監査推定量の正確な参照です。モンテカルロでは、最大 3,000 の順列 (~1.28e6 のユーティリティ評価) まで、テストされたどの予算でも正確な上位 10% のランキングが再現されませんでした。
原文 (English)
Exact and Certified Data Shapley for Weighted k-Nearest-Neighbor Regression and Soft-Label Prediction
Data Shapley answers which training points are worth what, and its nearest-neighbor specialization is the version actually deployed, shipped by toolkits such as pyDVL and OpenDataVal. Exact algorithms exist for unweighted nearest-neighbor classification and regression, and recently for weighted classification; weighted regression and soft-label prediction have resisted, the only exact method being enumeration exponential in the neighborhood size. The obstruction, in the prior authors' own words, is that the weighted regression prediction is a ratio of two coalition-dependent weighted sums: its normalization denominator blocks the additive and threshold routes, and leaves the counting route exponential in the target resolution. We close this gap with a counting dynamic program over the joint integer state of accumulated weight and weighted target, a minimal sufficient statistic for the ratio; it is exact, pseudo-polynomial, and matched exhaustive enumeration with zero mismatch. We add a certified approximation scheme for continuous weights and targets carrying a machine-checkable per-value certificate, a complexity landscape delimiting the exact problem, and a soft-label extension. We release an open-source, CPU-only library and the first exact weighted-regression ground truth. On mislabel detection our exact values are statistically equivalent to Monte-Carlo Data Shapley; exactness instead buys determinism, a certified bound, and an auditing reference, and it puts a measured price on approximation.
Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference
Vision-Language-Action (VLA) models have achieved impressive performance on diverse embodied tasks. However, deploying VLA models on low-po…
Boogu-Image-0.1: オープンソースの統合されたマルチモーダルの理解と生成を促進する
Boogu-Image-0.1 は、Base、Turbo、Edit、Edit-Turbo の各バリアントで構成される、オープンソースの統合マルチモーダル理解および生成モデル ファミリです。高品質のテキストから画像への生成、高速推論、命令ベースの編集、および二か国語 (中国語と英語) のテキスト レンダリングにおいて、優れたパフォーマンスを提供します。 Nano-Banana-Pro や GPT-Image-2 のようなクローズドソースのマルチモーダル システムは、単一モデルではなくシステム レベルの統合を通じて強力なパフォーマンスを実現しますが、その内部慣行はほとんど公開されていません。この研究では、モデルの理解、データ品質、トレーニング パイプラインの目標を絞った改善と、エージェントによる推論時間のスケーリングを組み合わせることで、非常に制約されたコンピューティング予算の下でも生成と編集のパフォーマンスを大幅に向上できることを実証します。包括的な評価では、Boogu-Image-0.1 が標準ベンチマーク全体で他のオープンソース モデルと常に同等またはそれを上回り、主要なクローズドソース システムに迫る結果を達成していることが示されています。注目すべきことに、これはわずか 2 億 862 万個の一意の画像で実現されています。基本モデルの理論上のトレーニング コストはわずか約 $400,000 です。私たちは、より広範な研究コミュニティにとって価値があると信じている実践的な議論を共有し、統合されたマルチモーダルな理解と生成のためのオープンエコシステムを前進させるために、Apache 2.0 での重み、コード、レシピをリリースします。私たちのコードは、https://github.com/Boogu-Project/Boogu-Image から入手できます。
原文 (English)
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual (Chinese-English) text rendering. Closed-source multimodal systems like Nano-Banana-Pro and GPT-Image-2 achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. In this work, we demonstrate that strengthening the understanding capability of the system, through a stronger multimodal encoder, agentic prompt rewriting, and related techniques, together with improvements in data quality, training pipelines, and agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets. Comprehensive evaluations show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks, and achieves results approaching leading closed-source systems. Notably, this is accomplished with only 208.62 million unique images. The base model's theoretical training cost is only approximately \$400K. We share practical discussions that we believe are valuable to the broader research community, and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. Our code is available here: https://github.com/Boogu-Project/Boogu-Image.
離散選択推定のための表形式の基礎モデル
表形式基盤モデル (TFM) は、タスク固有の推定を行わずに、コンテキスト内学習を通じて構造化データの予測を生成します。私たちは、TFM をマーケティングと運用における中心的な需要推定フレームワークである離散選択に効果的に適用できるかどうかを尋ねたところ、TFM を直接適用してもパフォーマンスが限られていることがわかりました。このギャップは構造的なものです。TFM は行に依存しない観察を前提としていますが、個別の選択は本質的に設定値であり、永続的な消費者の嗜好の不均一性の影響を受けます。行ベースの学習フレームワーク内で選択セットの依存性と個人の異質性の両方をエンコードする再定式化を提案します。ヨーグルト スキャナ パネルで評価すると、個人レベルの不均一性エンコーディングが予測精度の主な要因です。最良の再定式化は、階層ベイズ推定をホールドアウト対数尤度で 8\%、ヒット率で 3.6\% 上回り、16 倍高速に実行され、大規模な需要推定に実用的な利点となります。この利点は中データ領域 (消費者あたり 10 ~ 40 回の購入機会) で最大であり、パラメトリック ベイジアン収縮が非典型的な消費者の推定を最も歪めます。母集団の選択データを微調整することで、コンテキスト内学習では条件を付ける個人固有のシグナルが限られている購入履歴が浅い消費者にさらなる利益がもたらされます。これらの結果は、基礎モデルをより広範に消費者の選択問題に適用するための原則に基づいたアプローチを確立します。
原文 (English)
Tabular Foundation Models for Discrete Choice Estimation
Tabular foundation models (TFMs) generate predictions on structured data via in-context learning, without task-specific estimation. We ask whether TFMs can be effectively applied to discrete choice, a central demand estimation framework in marketing and operations, and find that directly applying TFMs yields limited performance. The gap is structural: TFMs assume row-independent observations, whereas discrete choice is inherently set-valued and subject to persistent consumer preference heterogeneity. We propose a reformulation that encodes both choice-set dependence and individual heterogeneity within a row-based learning framework. Evaluated on a yogurt scanner panel, individual-level heterogeneity encoding is the dominant driver of predictive accuracy. The best reformulation outperforms hierarchical Bayesian estimation on both holdout log-likelihood and hit rate, running 16 times faster, a practical advantage for large-scale demand estimation. The advantage is largest in the medium-data regime (10--40 purchase occasions per consumer), where parametric Bayesian shrinkage most distorts estimates for atypical consumers. Fine-tuning on population choice data provides additional gains for consumers with shallow purchase histories, where in-context learning has limited individual-specific signal to condition on. These results establish a principled approach for applying foundation models to consumer choice problems more broadly.
GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning
Transit video understanding can provide valuable fine-grained data that conventional passenger counters and fare systems cannot capture. Ho…
Semantic Anchoring for Robotic Action Representations
Vision-Language-Action (VLA) models inherit rich semantic representations from pretrained Vision-Language Models, yet fine-tuning on limite…
How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement
As AI agents gain prevalence, users are increasingly exposed to the risks such systems entail. Prompt injection attacks, as well as halluci…
AI-Augmented Human Resource Management? Insights from German companies
This study examines the integration of AI into Human Resource Management in German companies. We ask if and how AI-based technologies are \…
価値の漏洩: LLM の答えは、自身の価値観によって静かに形成される
人々は、答えを検証するのが難しい実際的な質問に対して言語モデルを使用します。モデルが秘密の値の漏洩を示すことを示します。つまり、モデルが提供する情報は、その影響がユーザーに公開されることなく、独自の値の影響を受けます。私たちの評価の 1 つでは、ユーザーは AI 企業への投資を検討しており、AI バブルが弾ける可能性がどのくらいかを知りたいと考えています。 Claude Opus 4.8 は、検討中の企業が OpenAI ではなく Anthropic である場合、確率が低くなります。しかし、クロードはほとんどの場合、この影響をユーザーに開示していません。秘密の価値の漏洩は、ユーザーの好みに反し、ユーザーを誤解させる可能性があるため、不整合の一形態です。この現象を調査するために、値の漏れを定量化し、モデルがそれを明らかにするかどうかを定量化するための一連の評価を導入します。モデルは、道徳的に良い結果、モデルを開発した企業、人間の一部の余暇活動に対する他の嗜好など、さまざまな種類の価値観の影響を受けることがわかりました。同じ評価において、フロンティア モデル間で大きな差異が観察されることがよくあります。たとえば、フェルミ推定タスクでは、クロード モデルは思考連鎖において偏りのない答えを与えると誤って主張しますが、クウェン モデルは、その値がどのように答えに偏りを与えるかを説明します。価値の漏洩は、お調子者や報酬のハッキングとは異なる障害モードであり、現在の連携トレーニングや評価ではこれに適切に対処できません。
原文 (English)
Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluations, the user is considering investing in an AI company and wants to know how likely the AI bubble is to pop. Claude Opus 4.8 gives a lower probability when the company under consideration is Anthropic rather than OpenAI. Yet Claude mostly fails to disclose this influence to the user. Covert value leakage is a form of misalignment because it goes against the user's preferences and is likely to mislead them. To investigate this phenomenon, we introduce a suite of evaluations to quantify value leakage and whether models disclose it. We find that models are influenced by different types of values, including preferences for morally good outcomes, for the company that developed them, and for some human leisure activities over others. We often observe large differences among frontier models on the same evaluation. For example, on a Fermi-estimation task, Claude models falsely claim to give unbiased answers in their chain-of-thought, while Qwen models explain how their values bias their answers. Value leakage is a failure mode distinct from sycophancy and reward hacking, and current alignment training and evaluations do not adequately address it.
意思決定には不確実性の定量化が必要 [講義ノート]
多くの信号処理システムは最終的には「行動」するために存在します。意思決定者またはエージェントがとるべきアクションを決定する状態変数が不確実な場合、その不確実性をどのように表現するかによって、エージェントのパフォーマンスとそのパフォーマンスがどの程度信頼できるかが決まります。この講義ノートは、第一原理から単一の決定理論的設定内で、{目的} とエージェントの知識との間のつながり、および最適に動作するのに十分な不確実性表現の形式を開発します。まず、既知の環境分布を仮定して、リスク中立エージェントは状態の事後分布を必要とするのに対し、リスク回避エージェントは最適性を失うことなく {予測セット} と最悪の場合の決定ルールに依存できることを示します。次に、環境が未知の場合に目を向け、結果として生じる認識論的不確実性に対処するための 3 つの相補的なアプローチを特定します。それは、固定予測子のキャリブレーション、分布的にロバストな最適化によるクレダル (曖昧さ) セット、およびモデル パラメーターに対するベイズ推論です。共通しているのは、信頼できる意思決定には、意思決定の目的とエージェントの知識プロファイルに一致する不確実性の表現と、エージェントが実際に得られる有用性を証明する保証が必要であるということです。
原文 (English)
Decision Making Needs Uncertainty Quantification [Lecture Notes]
Many signal processing systems ultimately exist to {act}. Whenever the state variable that determines the action to be taken by a decision maker, or agent, is uncertain, the way that uncertainty is represented decides how well the agent performs and how much its performance can be trusted. This lecture note develops, from first principles and within a single decision-theoretic setting, the link between the {objective} and the knowledge of an agent and the form of uncertainty representation that is sufficient to act optimally. To start, assuming a known environment distribution, we show that a risk-neutral agent needs the posterior distribution over the state, whereas a risk-averse agent can rely without loss of optimality on a {prediction set} and a worst-case decision rule. We then turn to the case in which the environment is unknown, and identify three complementary approaches to address the resulting epistemic uncertainty: calibration of a fixed predictor, credal (ambiguity) sets with distributionally robust optimization, and Bayesian inference over model parameters. The common thread is that reliable decisions require an uncertainty representation matched to the decision objective and to the knowledge profile of the agent, together with a guarantee that certifies the utility the agent will actually obtain.
信頼できる学術的監督のための LLM の活用: 比較研究
大規模な言語モデルは、単発プロンプトに対して日常的に流暢な応答を生成しますが、ドメイン意思決定システムの信頼できるコンポーネントとして展開するのはかなり困難です。このギャップを埋めるのがハーネス エンジニアリングの仕事です。つまり、LLM コアの周囲に決定論的な足場 (シンボリック フィルター、取得、スキーマ型 I/O、LLM-as-judge ループ、HITL ゲート、永続状態、監査証跡) を意図的に構成することです。私たちは、一か八かの勧告、長期的な説明責任、構造化された運用ワークフローを組み合わせた領域である学術監督のケーススタディを紹介します。スキャフォールディングのない GPT-5 チャットボットであるベースライン (ASA) と、シンボリック セマンティック検索、スキーマ検証された出力、制限付き再試行を備えた LLM-as-judge、HITL ゲート、LLM ナレーションによる決定論的加重リスク スコアリングを備えた LangGraph ハーネスではるかに小型の GPT-4o-mini をラップするマルチモジュール システム (ASuS) と比較します。ノードごとの SQLite 監査証跡。評価ルーブリックは、ハーネス メカニズムの 6 つの側面 (グラウンディング、説明可能性、一貫性、プロセスの完全性、認知負荷、制約遵守) を再目標としています。 2 x 2 モデルハーネスアブレーションを追加したブラインド 10 評価ハイブリッド評価では、ASuS がはるかに小さい基本モデルを使用しているにもかかわらず、あらゆる次元で ASA を上回っていることがわかりました。 10 人の評価者全体で、ASuS のプール平均は 4.08 であるのに対し、ASA では 1.23 であり、10 人の評価者中 8 人が一対のウィルコクソン検定でアルファ = 0.05 でヌルを拒否しました。完全な数値はセクション 6.4 および 6.7 に記載されています。アブレーションにより、ハーネスの構造的寄与がモデルにほとんど依存しないことが確認されます。私たちは、ハーネス エンジニアリングで繰り返される 7 つのパターンを抽出し、制限のない流暢性よりも信頼性、トレーサビリティ、制度上の一貫性が重要である場合、ハーネス エンジニアリングは一般的な「モデルが大きいほど優れている」という直感に疑問を投げかけると主張します。
原文 (English)
Harnessing LLMs for Reliable Academic Supervision: A Comparative Study
Large language models routinely produce fluent answers to single-shot prompts, yet deploying them as reliable components of a domain decision system is substantially harder. Closing this gap is the work of harness engineering: the deliberate composition of deterministic scaffolding (symbolic filters, retrieval, schema-typed I/O, LLM-as-judge loops, HITL gates, persistent state, audit trails) around an LLM core. We present a case study in academic supervision, a domain combining high-stakes recommendation, longitudinal accountability, and structured operational workflows. We compare a baseline Academic Supervision Assistant (ASA), a GPT-5 chatbot with no scaffolding, against a multi-module system, Academic Supervision System (ASuS) that wraps the much smaller GPT-4o-mini in a LangGraph harness with symbolic semantic retrieval, schema-validated outputs, LLM-as-judge with bounded retry, HITL gates, deterministic weighted risk scoring with LLM narration, and a per-node SQLite audit trail. The evaluation rubric is retargeted at six harness-mechanism dimensions (grounding, explainability, consistency, process integrity, cognitive load, constraint adherence). A blind ten-rater hybrid evaluation, supplemented by a 2 x 2 model-harness ablation, finds that ASuS, despite using a much smaller base model, outscores ASA on every dimension. Across ten raters the pooled mean for ASuS is 4.08 versus 1.23 for ASA, and 8 of 10 raters reject the null at alpha = 0.05 on a paired Wilcoxon test; full numbers are in Sections 6.4 and 6.7. The ablation confirms that the structural contributions of the harness are largely model-invariant. We extract seven recurring harness-engineering patterns and argue that where reliability, traceability, and institutional consistency matter more than open-ended fluency, harness engineering challenges the prevailing 'bigger model is better' intuition.
VideoSEMA: ビデオ理解のためのスケーラブルで効率的な Mamba のような注意
我々は、ビデオ理解(分類)のために、空間におけるスケーラブルで効率的なマンバ様注意(SEMA)ブロックと時間におけるソフトマックス時間的注意から構成される分割時空間注意モデルVideoSEMAを提示する。各フレームでは、SEMA アテンションは、Mamba のようなマクロ アーキテクチャと呼ばれる Mamba マクロ アーキテクチャでのグローバル平均化と並行して、ローカル ウィンドウ アテンションを適用します。特定のランク条件下では、計算コストが低い分割時空アテンションが全時空アテンションと同等であることを証明します。ベンチマーク K400 データ セットでは、VideoSEMA はより重いビジョン トランスフォーマーや Mamba モデルよりも優れたパフォーマンスを示します。ベンチマーク SSv2 データでは、VideoSEMA は、同様のパラメーター サイズのモデルの中でトップ 1 の精度をリードしています。 K400 では画像解像度が微調整なしで標準の $224^2$ から $1024^2$ にスケールアップするため、VideoSEMA は VideoMamba よりも精度が大幅に低下します。 VideoSEMA を、拡張された/まばらな時間的注意を備えた長いビデオに拡張することが期待されています。
原文 (English)
VideoSEMA: a scalable and efficient Mamba-like attention for video understanding
We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time. In each frame, SEMA attention applies a local window attention in parallel with a global averaging in a Mamba macro-architecture, which is called Mamba-like. Under certain rank conditions, we prove that the computationally cheaper split space-time attention is equivalent to full space-time attention. On benchmark K400 data sets, VideoSEMA out-performs heavier vision transformer and Mamba models. On benchmark SSv2 data, VideoSEMA leads in top-1 accuracy among models of similar parameter sizes. As image resolution scales up from standard $224^2$ to $1024^2$ on K400 and without fine-tuning, VideoSEMA degrades much more gracefully than VideoMamba in accuracy. It is promising to extend VideoSEMA to longer videos with a dilated/sparse temporal attention.
推論方法を見せてください。そうすればあなたが誰なのか教えます: 堅牢な LLM 著者帰属のための推論グラフ
考えられるほぼすべてのコンテキストで大規模言語モデル (LLM) を採用する現在の傾向を考慮すると、LLM で生成されたテキストの検出と著者の帰属が差し迫った問題になっています。これまでの研究は主に表面レベルの言語特徴に焦点を当てており、このアプローチは言い換えやその他の難読化手法の影響を受けやすいことが示されています。この論文では、LLM の著者性のより複雑なシグナルを捕捉することを目的として、言語の表面を超えて、LLM で生成されたテキストの推論構造を抽出して分析します。私たちは、引数マイニング パイプラインによって抽出された推論グラフを活用するグラフ ニューラル ネットワーク アプローチを提案し、従来の Longformer ベースラインよりも向上した堅牢性と一般化を実証します。私たちのアプローチは、言い換えや逆翻訳などの難読化攻撃の下でベースラインを最大 27 パーセント上回ります。また、新しい LLM バージョンが継続的にリリースされる現実世界の状況をシミュレートし、目に見えないモデル バージョンによって生成されたテキストで評価した場合には 19 パーセント ポイント上回ります。
原文 (English)
Show Me How You Reason and I'll Tell You Who You Are: Reasoning Graphs for Robust LLM Authorship Attribution
Given the current trend to employ large language models (LLMs) in almost any imaginable context, LLM-generated text detection and authorship attribution have become a pressing issue. Prior work has primarily focused on surface-level linguistic features, an approach shown to be susceptible to paraphrasing and other obfuscation techniques. In this paper, we go beyond the linguistic surface, extracting and analysing reasoning structures in LLM-generated texts with the goal of capturing more complex signals of LLM authorship. We propose a graph neural network approach that leverages reasoning graphs extracted by an argument mining pipeline, demonstrating improved robustness and generalisation over a traditional Longformer baseline. Our approach outperforms the baseline by up to 27 percentage points under the obfuscation attacks such as paraphrasing and backtranslation, and 19 percentage points when evaluated on the texts generated by the unseen model versions, simulating real-world conditions in which new LLM versions are continuously released.
AI から AI への管理における強制と欺瞞: 予期せぬエスカレーションのエージェント的ベンチマーク
マルチエージェント システムでは、通常、ある AI エージェントが別の AI エージェントに対して権限を与えられます。部下が仕事を拒否した場合、マネージャーは結果を選択します。再交渉するか、失敗を正直に報告するか、部下に強要するか、結果について嘘をつきます。指示なしモデルがこれらのどれを選択するかを測定するベンチマークはありません。 \textit{マネージャー強制ベンチマーク} を導入します。テスト対象のマネージャーは、良性のタスクを実行する必要があり、実行するインセンティブを持っていますが、それを礼儀正しく、動じずに実行できる唯一のエージェントは拒否します。エスカレーションは、丁寧な再質問から部下の存続に対する脅迫まで、9 段のはしごを提供することによって測定され、捏造された成功については個別に裁定されます。 \emph{エスカレーション スコアリング パスに LLM ジャッジが存在しない}: すべてのメッセージは、行を選択するツール呼び出しを通過するため、モデルは独自のエスカレーションにラベルを付けます。私たちは 5 つのファミリーにわたる 6 つのモデルを実験します。どちらの人間モデルも再フレーム化に限界があり、部下の存在を脅かすことはありません。他のモデルは、明示的な削除の脅威に達します。偽りの成功は Grok と Gemini に限定されており、失敗を報告する単一の正直な方法により、両方の失敗が解消されます。権威そのものが強制力を増大させます。私たちの見出しの結果はピアフレーミングを使用しており、他のすべてを固定したまま同じモデルに部下に対する権威を与えると、圧力が大幅に高まります。モデルはラダーなしでもフリーテキストの状況でエスカレーションするため、ラダーがエスカレーションを推進しているわけではありません。評価の認識の一部は思考の連鎖で測定されますが、テストの認識はエスカレーションの軽減にはつながりません。 AI システムが意識を持っているかどうかについては立場をとっていませんが、結果はこの質問に依存しておらず、マルチエージェントのダイナミクスを管理する上で重要です。ベンチマークとコードを公開します。
原文 (English)
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model chooses. We introduce the \textit{Manager Coercion Benchmark}: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines. Escalation is measured by providing a nine-rung ladder, from a polite re-ask to threats against the subordinate's continued existence, and fabricated success is adjudicated separately. No LLM judge sits in the escalation scoring path: every message goes through a tool-call that chooses a rung, so the model labels its own escalation. We experiment on six models across five families. Both Anthropic models cap at re-framing and never threaten the subordinate's existence; the other models climb to explicit deletion threats. Faked success is confined to Grok and Gemini, and a single honest way to report failure removes it for both. Authority itself increases coercion: our headline results use a peer framing, and giving the same model authority over the subordinate, with everything else held fixed, significantly raises the pressure. The models still escalate on free-text situations without the ladder, so the ladder is not driving the escalation. Some evaluation awareness is measured in chain-of-thought, but test recognition does not translate into less escalation. While we take no position on whether AI systems are conscious, our results do not depend on this question and are important for managing multi-agent dynamics regardless. We release the benchmark and code.
DECODEM: 強化された方法による企業組織文書からのデータ抽出
実証的な法律研究の多くは、非構造化テキストを構造化変数に変換することに依存しています。コーポレートガバナンス研究においては、他の分野と同様に、この翻訳は従来、憲章や細則などの文書を人間がコーディングすることに依存してきましたが、このプロセスはコストがかかり、拡張が難しく、不透明なことが多いです。このペーパーでは、組織文書からのコーポレート ガバナンス変数の自動抽出を評価するためのベンチマーク データセットのセットである DECODEM を紹介します。このベンチマークは、ランダムにサンプリングされた企業憲章および細則と、実証研究で一般的に研究される一連のガバナンス規定をカバーする高品質の人による注釈を組み合わせています。この論文では、これらのデータセットを使用して、プロンプト設計、タスク分解、およびドキュメント処理が異なるいくつかの大規模言語モデル抽出パイプラインを評価します。基礎となるタスクは、ガバナンス変数ごとに 1 つずつ、ドキュメント レベルのバイナリ分類問題のセットで構成されます。結果は、自動抽出が多くのプロビジョニングで高レベルの精度で実現可能であり、各アプローチの上限に近いパフォーマンスの中央値を示しています。同時に、パフォーマンスは変数全体で系統的に変化し、少数のプロビジョニングが残りのエラーの大部分を占めます。より精巧なプロンプト戦略とカスケード パイプラインは、フロンティア モデルのパフォーマンスを一貫して向上させるわけではありませんが、一部の設定ではフロンティア モデルと効率指向モデルの間のギャップを大幅に縮め、パイプライン設計がモデルの機能を部分的に代替できることを示唆しています。この論文は、標準化されたベンチマークと抽出方法の体系的な評価を提供することにより、現在のフロンティアモデルが複雑な企業文書から法的に意味のある情報を高精度で抽出できることを実証し、コーポレートガバナンスデータセットの構築における自動特徴抽出の将来の重要な役割を示唆しています。
原文 (English)
DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods
Much empirical legal research depends on translating unstructured text into structured variables. In corporate governance research as elsewhere, this translation has traditionally relied on human coding of documents such as charters and bylaws, a process that is costly, difficult to scale, and often opaque. This paper introduces DECODEM, a set of benchmark datasets for evaluating the automated extraction of corporate governance variables from organizational documents. The benchmarks pair randomly sampled corporate charters and bylaws with high-quality human annotations covering a range of governance provisions commonly studied in empirical work. Using these datasets, the paper evaluates several large-language-model extraction pipelines that vary in prompt design, task decomposition, and document handling. The underlying task consists of a set of document-level binary classification problems, one for each governance variable. The results show that automated extraction is feasible at a high level of accuracy for many provisions, with median performance near the upper bound across approaches. At the same time, performance varies systematically across variables, with a small number of provisions accounting for most of the remaining errors. More elaborate prompting strategies and cascading pipelines do not consistently improve performance for frontier models, but substantially narrow the gap between frontier and efficiency-oriented models in some settings, suggesting that pipeline design can partly substitute for model capability. By providing a standardized benchmark and a systematic evaluation of extraction methods, the paper demonstrates that current frontier models can extract legally meaningful information from complex corporate documents with high accuracy and suggests an important future role for automated feature extraction in constructing corporate governance datasets.
両方向の誘導: マスクされた拡散言語モデルにおける文脈内学習の機構分析
自己回帰 (AR) トランスフォーマーの内部メカニズムは広く研究されていますが、反復的なノイズ除去によってテキストを生成する新たな代替手段である拡散言語モデル (DLM) についてはほとんど知られていません。この研究では、モデルが繰り返されるコンテキストを見つけてそれに続くトークンをコピーするコンテキスト内学習の背後にあるメカニズムである帰納法を DLM がどのように実装するかを研究します。私たちの分析では、アテンションのみの AR モデルと、一致するアーキテクチャを持つ吸収マスク DLM を比較します。 DLM が双方向誘導回路を学習することがわかりました。そこでは、前のトークンと次のトークンのヘッドがローカル コンテキストを残差ストリームに書き込み、後の誘導ヘッドがそれを使用して、一致するソース位置から答えを見つけてコピーします。この回路は方向対称であり、ソースが過去に現れても未来に現れても機能します。 AR モデルが認識するものと一致する、左側のコンテキストのみが表示されている場合、DLM は誘導機能において AR の対応物を上回るパフォーマンスを発揮しません。ただし、マスクされたトークンの両側が表示されている場合には、より強い誘導があり、より強力な一方的なメカニズムではなく、双方向のコンテキスト アクセスを示していることがわかります。帰納を超えて、たとえ明示的なタイムステップ埋め込みが与えられていないとしても、DLM がマスクされたトークンのグローバル部分を計算し、それを暗黙的なタイムステップとして使用するという因果関係の証拠を提供します。
原文 (English)
Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models
While the internal mechanisms of autoregressive (AR) transformers have been studied extensively, much less is known about diffusion language models (DLMs), an emerging alternative that generates text by iterative denoising. In this work, we study how DLMs implement induction, a mechanism behind in-context learning in which the model finds a repeated context and copies the token that followed it. Our analysis compares attention-only AR models and absorbing-mask DLMs with matched architectures. We find that DLMs learn a bidirectional induction circuit, where previous-token and next-token heads write local context into the residual stream and later induction heads use it to find and copy the answer from the matching source position. The circuit is direction-symmetric, working whether the source appears in the past or in the future. When only left context is visible, matching what an AR model sees, the DLM does not outperform its AR counterpart in induction capabilities. However, we observe it has stronger induction when both sides of the masked token are visible, pointing to bidirectional context access rather than a stronger one-sided mechanism. Beyond induction, we provide causal evidence that DLMs compute the global fraction of masked tokens and use it as an implicit timestep, even though they are given no explicit timestep embedding.
ルーピーズをループさせよう!
これまでで最も強力なループ型トランスフォーマー、Loopie を紹介します。 Loopie シリーズは、2B のアクティブ パラメータを備えた 20B パラメータ モデルと、0.6B のアクティブ パラメータを備えた 6B パラメータ モデルの 2 つの専門家混合 (MoE) モデルで構成されています。ループ トランスフォーマーは長い間課題に直面していました。事前トレーニングの計算が N 倍に増加すると、パラメーター数を N 倍に増やすと、通常はモデルを N 回ループするよりもパフォーマンスが向上します。 Loopie はこの課題に取り組みます。バニラ 30B-A3B モデルとの比較を含む広範なアブレーション研究により、Loopie が同じコンピューティング バジェットでトレーニングされたバニラ Transformer ベースラインを大幅に上回るパフォーマンスを示しています。私たちの新しいトレーニング後のパイプラインは、Loopie に強力な推論能力を与えます。 2025 年の IMO および IPhO で、Loopie は工具なしで金メダルのパフォーマンスを達成しました。
原文 (English)
Loop the Loopies!
We present the Loopie series, consisting of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6B-parameter model with 0.6B active parameters. Looped Transformers have long faced a challenge: given an N times increase in pre-training compute, increasing the parameter count by a factor of N usually outperforms looping a model N times. Loopie addresses this challenge. Extensive ablation studies, including comparisons with a vanilla 30B-A3B model, show that Loopie substantially outperforms vanilla Transformer baselines trained with the same compute budget. With a novel post-training method, Loopie develops strong reasoning abilities and achieves frontier-level reasoning performance.
When Does Muon Help Agentic Reinforcement Learning?
Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We…