AIニュース 2026-07-17
自動生成: 2026-07-17 12:15 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
Why teens deserve access to safe AIOpenAI
Learn how OpenAI is making ChatGPT safer for teens with age-appropria…
-
強気値上げで自爆か ClaudeやGeminiに押され「M365 Copilot」は一人負け?:888th LapITmedia AI+
「ChatGPT」だけでなく、「Claude」や「Gemini」も支持を広げる生成AI市場。だが、Microsoft 365の法人ユーザー…
-
Our approach to bioresilienceGoogle DeepMind
Google DeepMind and Isomorphic Labs are sharing our joint approach to…
-
Google、「NotebookLM」を「Gemini Notebook」に改称 Geminiエコシステムへの統合を強化ITmedia AI+
Googleは、AI搭載リサーチツール「NotebookLM」の名称を「Gemini Notebook」に変更すると発表した。機能は単体製…
-
Why is OpenAI selling a ChatGPT basketball?TechCrunch AI
You may have heard that OpenAI released its first piece of hardware t…
-
「Claude」利用制限を全リセット CodexとChatGPT Workも “リセット合戦”再びITmedia AI+
米Anthropicは、AIサービス「Claude」の5時間および1週間の利用制限をリセットしたと発表した。一方、米OpenAIもデスクト…
-
「現場で使えるAIへ」、富士通が進める次世代CPUと自律型AIエージェント戦略ITmedia AI+
富士通は自社イベント「Fujitsu Experience Day 2026」で、次世代CPUの「FUJITSU-MONAKA」をはじめ、…
トピック別件数
- 研究/論文 96件
- LLM/生成AI 92件
- エージェント 61件
- 画像/動画生成 37件
- ロボティクス 20件
- ビジネス/資金調達 19件
- ハードウェア/半導体 9件
- その他 6件
- 規制/政策 1件
日本語メディア14件
ITmedia AI+ (日本語)
「スマホで動く」270億パラメーターLLM「Bonsai 27B」登場
27BクラスのモデルをiPhoneで実行可能な容量に収めたとしている。
Google、「NotebookLM」を「Gemini Notebook」に改称 Geminiエコシステムへの統合を強化
Googleは、AI搭載リサーチツール「NotebookLM」の名称を「Gemini Notebook」に変更すると発表した。機能は単体製品として維持しつつ、エコシステム全体で広く機能するようになる。安全なクラウドコンピュータの割り当てによるデータ分析機能などの大型アップデート…
これからのロボットは「買ったときが一番性能低い」? ソフトバンクと安川電機「フィジカルAI」学習工程の効率化を実証
「SoftBank World 2026」でソフトバンクと安川電機がフィジカルAI協業の成果を披露した。ソフトバンクの湧川隆次CTOは、強化学習によって使うほど性能が向上するロボットへの転換を「大きなパラダイムシフト」と表現した。
強気値上げで自爆か ClaudeやGeminiに押され「M365 Copilot」は一人負け?:888th Lap
「ChatGPT」だけでなく、「Claude」や「Gemini」も支持を広げる生成AI市場。だが、Microsoft 365の法人ユーザーは約4億5000万人いるものの、有料版「Microsoft 365 Copilot」を契約している割合は4.5%未満にとどまるという。
Google検索の「AIモード」、検索結果からそのままアプリで作業可能に CanvaやYouTube Musicなど
Googleは、検索の「AIモード」で外部アプリと直接連携できる機能の提供を開始した。まずは米国で順次展開し、Canva、YouTube Music、Instacartに対応する。検索結果から離れずにデザイン作成やカート追加などのタスクを完了できる。「Personal Inte…
日本再起の旗印となるか、国産マルチモーダルAI基盤「FRONTia」が始動
経済産業省とNEDOは、AIロボットやフィジカルAIに用いられる国産マルチモーダル基盤モデル「FRONTia(フロンティア)」の開発プロジェクトの本格始動に合わせて、東京都内で「我が国のフィジカルAI政策に関する対外発信イベント」を開催した。
「現場で使えるAIへ」、富士通が進める次世代CPUと自律型AIエージェント戦略
富士通は自社イベント「Fujitsu Experience Day 2026」で、次世代CPUの「FUJITSU-MONAKA」をはじめ、AIチームを自律生成する技術や、実店舗で検証する店長業務支援AIなど企業変革を支援する最新AIソリューションを公開した。
フアンCEO「ジャパンAI構築はマストだ」 経産省、国産フィジカルAIで新プロジェクト 赤沢大臣も“革ジャン”羽織る
経済産業省が、国産フィジカルAI向け基盤モデルを構築する「FRONTia Project」をスタートさせた。国内44社が共同出資するAI開発企業Noetraと産総研を中心とし、2030年での「実世界ネイティブAI」実現を目指す。16日に開催されたキックオフイベントでは、NVID…
富士通、国内ロボット大手3社と「フィジカルAI」で協業 NVIDIAの技術活用
富士通は、AIが自律的に考え、ロボットの体を動かす「フィジカルAI」の開発に関し、川崎重工業とファナック、安川電機の各社と協業すると発表した。米NVIDIAの技術を活用し、ロボットを協調的に制御するための基盤を開発する。
大手共同出資の“国産AI開発企業”が本格始動 NVIDIAも協力、「Rubin」2万7500基搭載の計算基盤を構築へ
国内大手が共同出資するAI開発企業Noetraが、国産のAIモデルの開発に向けて本格始動する。米NVIDIAの協力のもと、新たな計算基盤も構築する。
トンカツ食べながら語った――NVIDIA、富士通、安川電機ら“フィジカルAI連合”誕生、発表直前の裏話
富士通、ファナック、安川電機、川崎重工業とNVIDIAによる“フィジカルAI連合”が誕生した。5社のトップは、記者説明会の直前には「トンカツ」を食べながら語り合ったという。その一幕を富士通の時田社長が明かした。
「Claude」利用制限を全リセット CodexとChatGPT Workも “リセット合戦”再び
米Anthropicは、AIサービス「Claude」の5時間および1週間の利用制限をリセットしたと発表した。一方、米OpenAIもデスクトップPC向けAIサービス「ChatGPT Work」およびAIコーディングツール「Codex」の週次利用制限をリセットすると発表し、“リセッ…
「Gemini Spark」日本でもリリース、まずUltraから 24時間働く“パーソナルAIエージェント”
同社の幹部は、「Pro」(月額2900円)ユーザーにアクセスを拡大する可能性も示唆している。
「AIと壁打ちはもう古い」 業務タスクを任せる「Claude Cowork」の落とし穴
Anthropicは、PC上の業務タスクを任せられる「Claude Cowork」の活用法を公開した。チャットやClaude Codeとの使い分けや、どんなタスクを任せればよいかを解説している。
海外メディア10件
TechCrunch AI (英語)
Google Vids now lets you star in your own AI videos
Google is adding personalized AI avatars to Vids that let users create videos starring a digital version of themselves, alongside Gemini Om…
Roblox launches an AI-powered game-creation feature in its mobile app
Roblox's new "Build" feature lets users generate basic games using a single text prompt.
Google’s AI Mode now lets you link and interact with select apps
With this new update, Google is expanding AI Mode beyond answering questions and into completing tasks across the apps they use regularly.
Yes, you can now order DoorDash from the command line
DoorDash is opening a limited beta of dd-cli, a command-line tool that lets developers and AI agents search stores, build carts, and place…
Why is OpenAI selling a ChatGPT basketball?
You may have heard that OpenAI released its first piece of hardware this week. You may not have heard about the ChatGPT basketball.
How a former DeepMind researcher raised at a $300M pre-seed valuation before launching a product
Drawing on more than a decade spent helping build some of the world's most influential AI systems, including research that later informed t…
Why AMI Labs’ Alexandre LeBrun won’t call his AI ‘AGI’ or ‘superintelligence’
While everyone in AI is chasing "superintelligence," Alexandre LeBrun, CEO of Yann LeCun’s world model startup, AMI Labs, dismisses the wor…
Moonshot’s upcoming Kimi 3 is expected to close the gap with Anthropic’s Opus 4.8
The FT reports Kimi K3 will be the largest open AI model from China, with a parameter count between 2 trillion and 3 trillion.
Apple Intelligence approved for launch in China with Alibaba and Baidu
The deal, which was rumored to be in the works last year, marks an important step for Apple's AI ambitions in a key market.
Applied Computing wants to give oil and gas operators an AI model for the entire plant
Applied Computing has raised a $20M Series A to build a foundation AI model for the oil, gas and petrochemical industry.
公式ブログ2件
OpenAI (英語)
Why teens deserve access to safe AI
Learn how OpenAI is making ChatGPT safer for teens with age-appropriate protections, learning tools, parental controls, and expert partners…
Google DeepMind (英語)
Our approach to bioresilience
Google DeepMind and Isomorphic Labs are sharing our joint approach to bioresilience and AI models.
論文228件
arXiv cs.AI (英語)
OriginBlame: AI トレーニング データセットのレコード レベルおよびトークン レベルのデータ出所
データ投稿者が削除を要求すると、モデル トレーナーは現実的なギャップに直面します。学習解除アルゴリズムには忘却セットが必要ですが、特定の作成者に属するトレーニング レコードを特定できるツールはありません。既存の来歴システムはファイルまたはデータセット レベルで動作し、壊滅的な過剰削除を強いられます。私たちは、データ処理パイプラインを通じて作成者の身元を伝播し、決定論的なクエリを通じて失効リクエストを正確な忘却セットに解決する、レコードおよびトークンレベルのデータ来歴システムである ob を紹介します。 219,555 の Wikipedia ページの評価では、レコード レベルの出自によりデータセット レベルの過剰削除 (101 倍から 1.3 倍) が排除される一方、統合により Wiki データに 1.3 ~ 4.0% (HuggingFace) のスループット オーバーヘッド、および 2.1 ~ 19.0% (Datatrove) のスループット オーバーヘッドが追加されることが実証されました。 1.7B モデルでは、来歴ベースの忘却セットにより、ランダムなベースラインと比較して、未学習が 42% 改善されます。
原文 (English)
OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets
When a data contributor requests removal, model trainers face a practical gap: unlearning algorithms require a forget set, yet no tool can locate which training records belong to a given author. Existing provenance systems operate at file or dataset level, forcing catastrophic over-deletion. We present ob, a record- and token-level data provenance system that propagates author identity through data processing pipelines and resolves revocation requests into precise forget sets via deterministic queries. Evaluation on 219,555 Wikipedia pages demonstrates that record-level provenance eliminates dataset-level over-deletion (from 101x to 1.3x), while integration adds 1.3-4.0% throughput overhead (HuggingFace) and 2.1-19.0% (Datatrove) on wiki data. On a 1.7B model, provenance-based forget sets improve unlearning by 42% over random baselines.
SPINE: Agentic AI でサイバーと物理のギャップを埋める
基盤モデルはロボットに複雑な意思決定のための洗練された頭脳を与えてきましたが、その知能を物理的なプラットフォームに展開するには、依然として専門家による面倒な調整が必要です。この展開ギャップ、つまりロボットの脊髄が、スケーラブルな組み込み型 AI にとって依然として主要なボトルネックとなっています。そこで、私たちは SPINE (Scalable Physical Integration with ageNtic Expertise) を提案します。これは、最小限のロボット工学の専門知識で両手ロボットを体系的にデバッグおよび展開するためのエージェント フレームワークです。 SPINE のハーネスは、2 つの調整されたマルチエージェント ワークフローで構成されます。1 つはロボット固有のコンテキストを作成するプロファイル ビルダー、もう 1 つは遠隔操作が機能するまで診断、修復、検証を繰り返すデバッガーです。 7 つの DOBOT X-Trainer デバッグ シナリオ全体で、SPINE を使用したロボット工学の初心者は、同じ参照資料を使用したが、SPINE の構造化されたワークフローを使用しなかったクロード コードを使用した人間のオペレーターよりも優れたパフォーマンスを示し、運用化の成功率が 75% から 100% に向上し、遠隔操作までの平均時間が 16 分 45 秒から 13 分 47 秒に短縮されました。独自の ROS/CAN バイマニュアル アームである AgileX PiPER では、SPINE は埋め込まれた 10 件のバグすべてを解決しましたが、エキスパート ベースラインでは 10 件中 9 件をほぼ同じ時間内に解決しました。これらの結果を総合すると、SPINE が手動プラットフォーム間で転送でき、専門家による調整への依存を減らし、身体化された AI をスケーラブルな現実世界の展開に近づけることができることがわかります。
原文 (English)
SPINE: Bridging the Cyber-Physical Gap with Agentic AI
Foundation models have given robots a sophisticated brain for complex decision-making, yet deploying that intelligence into a physical platform still demands tedious, expert-driven calibration. This deployment gap, the robot's spinal cord, remains a primary bottleneck to scalable Embodied AI. Hence, we propose SPINE (Scalable Physical Integration with ageNtic Expertise): an agentic framework for systematically debugging and deploying bimanual robots with minimal robotics expertise. SPINE's harness comprises two orchestrated multi-agent workflows: a profile builder that creates robot-specific context, and a debugger that cycles through diagnosis, repair, and validation until teleoperation works. Across seven DOBOT X-Trainer debugging scenarios, a robotics novice using SPINE outperformed human operators using Claude Code with the same reference materials, but without SPINE's structured workflow, improving operationalization success from 75% to 100% and reducing mean time-to-teleoperation from 16 min 45 s to 13 min 47 s. On AgileX PiPER, a distinct ROS/CAN bimanual arm, SPINE resolved all 10 implanted bugs, versus 9 out of 10 for the expert baseline, in nearly the same amount of time. Together, these results show that SPINE can transfer across bimanual platforms, reduce dependence on expert calibration, and move embodied AI closer to scalable real-world deployment.
介入的グラウンディング監査: 述語置換による LLM 思考連鎖のブラックボックス前提依存性テスト
大規模な言語モデルは、論理的に健全であるように見えても、その記述された前提に本当に依存していない可能性がある思考連鎖 (CoT) 推論を生成します。私たちは、前提依存関係のブラックボックスのステップレベルのテストである介入的グラウンディング監査を導入します。ターゲット述語を新しいシンボルに置き換えることによって単一の前提に介入し、モデルを再実行し、各推論ステップの正規化された結論(正規述語形式)が変化するかどうかをチェックします。私たちは、ステップレベルの前提依存関係が既知である、ゴールドプルーフツリーを使用した合成マルチホップ演繹的推論ベンチマークである ProntoQA で評価します。 GPT-4o を使用して 50 の ProntoQA 問題に適用したところ、私たちの方法はプルーフ ツリーの依存関係の検出で F1 = 0.806 (述語決定の依存関係で F1 = 0.885、再現率 = 100%) を達成し、自己一貫性ベースライン (F1 = 0.343、95% のブートストラップ CI が重複しない) を大幅に上回りました。さらに、正しく解決された問題の 66% には、一貫性のある置換のもとでの直接的な証明木の依存関係に影響されない少なくとも 1 つの整列されたステップが含まれていることを特定しました。これらはすべて、エンティティ導入の前提、一貫性のある置換評価の文書化された盲点、つまり受動的手法には見えない「正しい答え、間違った推論」シグナルに関係しています。すべての監査証明書、生の出力、および再現スクリプトはパブリック GitHub リポジトリで入手でき、正式な解析可能なベンチマークを超えた範囲の制限についても説明します。
原文 (English)
Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution
Large language models produce chain-of-thought (CoT) reasoning that appears logically sound yet may not genuinely depend on its stated premises. We introduce interventional grounding audits, a black-box, step-level test of premise dependency: we intervene on a single premise by substituting its target predicate with a fresh symbol, re-run the model, and check whether each reasoning step's normalized conclusion (canonical predicate form) changes. We evaluate on ProntoQA, a synthetic multi-hop deductive reasoning benchmark with gold proof trees, where step-level premise dependencies are known. Applied to 50 ProntoQA problems with GPT-4o, our method achieves F1 = 0.806 on detecting proof-tree dependencies (F1 = 0.885 on predicate-determining dependencies; Recall = 100%), significantly outperforming a self-consistency baseline (F1 = 0.343; 95% bootstrap CIs non-overlapping). We further identify that 66% of correctly-solved problems contain at least one aligned step insensitive to a direct proof-tree dependency under consistent substitution -- all involving entity-introduction premises, a documented blind spot of the consistent-substitution evaluator -- a "right answer, wrong reasoning" signal invisible to passive methods. All audit certificates, raw outputs, and reproduction scripts are available in a public GitHub repository, and we discuss scope limits beyond formal, parsable benchmarks.
Belnap の型付き内包 FOL に基づく神経記号的 AGI ロボットの確率的拡張
$IFOL_B$ に基づくニューロシンボリック AI は、ニューラル学習と記号推論を組み合わせて、純粋なニューラル システムの制限 (解釈可能性や論理構造の欠如など) を自己参照のための形式的な論理機構で克服する方法です。この論文では、$IFOL_B$ の Nilsson の確率構造に基づいて、現在未知の文の確率計算を使用して、$IFOL_B$ の認知能力を拡張します。現在の知識データベースと論理推論を保存するグローバル対称変換と、$IFOL_B$ 述語の非常に厳密なサブセットのみを含む具体的な (サブ) 問題に関するリアルタイムの決定に使用されるローカル対称変換を導入します。どちらの場合も、シャノンの最大情報エントロピーに基づく確率密度関数 $KI$ の計算は、この確率的ニューロシンボリック AGI のニューラル ネットワークによって提供されます。
原文 (English)
Probabilistic Extension of Neuro-Symbolic AGI Robots based on Belnap's Typed Intensional FOL
Neuro-symbolic AI based on $IFOL_B$ is a way to combine neural learning and symbolic reasoning to overcome limitations of purely neural systems (like lack of interpretability and logical structure) with formal logical machinery for self-reference. In this paper we expand the cognitive power of $IFOL_B$ by using the probability computation for the currently unknown sentences, based on Nilsson's probability structure for the $IFOL_B$. We introduce the global symmetry transformation that preserves the current knowledge database and logical deduction, and the local one used for real-time decisions about concrete (sub)problems that involve only a very strict subset of $IFOL_B$ predicates. The computation of probability density function $KI$ in both cases, based on the Shannon's maximum information entropy, is provided by neural networks of this probabilistic neuro-symbolic AGI.
最新のエージェント システムにおける自己改善: 調査
自己改善型の自律エージェントは、研究プロトタイプから導入システムに移行しつつあります。主な目標は、人間の入力を最小限に抑えた、あるいはまったく入力を行わずに、経験から制御可能な進化または適応を行うことです。この調査は、現代の自己改善エージェントを、経験を蓄積された能力の獲得に変換する適応システムとして枠組み化しています。当社は、基礎モデルとプロンプト、メモリ、ツール、および制御ロジックの運用足場を結合する構成として最新のエージェントを表すシステム レベルのフレームワークを提供します。このフレームワーク内では、自己改善は、モデル パラメーターまたはスキャフォールド コンポーネントの更新を取得してコミットする自己誘導更新オペレーターとして形式化されます。私たちはこれまでの作業を更新ターゲットごと、および変化を促すシグナルごとに整理し、アプリケーションをレビューして評価について話し合った後、未解決の問題と将来の方向性について結論を出します。便宜上、https://github.com/selfimproving-agent/awesome-Self-Improving-Agents で技術的な更新を追跡しています。
原文 (English)
Self-Improvements in Modern Agentic Systems: A Survey
Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input. This survey frames modern self-improving agents as adaptive systems that convert experience into accumulated capability gains. We offer a system-level framework that represents a modern agent as a configuration coupling a foundation model with an operational scaffold of prompts, memory, tools, and control logic. Within this framework, self-improvement is formalized as a self-induced update operator that obtains and commits updates to model parameters or scaffold components. We organize prior work by update target and by the signals that drive change, then review applications and discuss evaluation, before closing with open problems and future directions. For convenience, we track technical updates on https://github.com/selfimproving-agent/awesome-Self-Improving-Agents.
グラフベースのツールを使用した小規模言語モデルでの分子特性予測の改善
小言語モデル (SLM) は、SMILES 文字列からのゼロショット分子特性予測の有望性を示していますが、配列表現が主要なグラフ トポロジカル キューを十分に指定していないため、構造盲点に悩まされることがよくあります。我々は、推論時にエージェントツールの使用を可能にするモジュール式のContext-Augmented Promptingフレームワークを提案します。訓練されたGNNエキスパートモデルは自信を持って予測ヒントを提供し、GNNはインスタンス固有の説明サブグラフ(例えば、サブグラフSMILESと付随する説明段落)を抽出します。 SMILES のみから手元にあるすべての利用可能なツールの使用まで、5 つのプロンプト設定の下で、MUTAG および Tox21 で一般的に使用される 3 つの SLM を評価します。 2 つのデータセットにわたって、グラフから派生したコンテキストでプロンプトを充実させると、精度が大幅に向上し、相対的な改善が 25% を超え、Tox21 では最大 74% 向上することがよくあります。さらに、必要性に基づいたエッジドロップ介入を通じて、抽出されたモチーフの機能的関連性を検証します。観察された進歩にもかかわらず、特殊化された GNN モデルには依然としてギャップがあり、分子構造に対するテキスト条件付き推論の価値と限界の両方が浮き彫りになっています。
原文 (English)
Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools
Small language models (SLMs) have shown promise for zero-shot molecular property prediction from SMILES strings, yet they often suffer from structural blindness because sequence representations under-specify key graph-topological cues. We propose a modular Context-Augmented Prompting framework that enables agentic tool use at inference time: a trained GNN expert model provides a predictive hint with confidence, and a GNN extracts an instance-specific explanatory subgraph (e.g., a subgraph SMILES and an accompanying explanatory paragraph). We evaluate three commonly used SLMs on MUTAG and Tox21 under five prompting configurations ranging from SMILES-only to using all available tools at hand. Across two datasets, enriching prompts with graph-derived context yields substantial accuracy gains, often exceeding 25% relative improvement and up to 74% on Tox21. We further validate the functional relevance of the extracted motifs via a necessity-based edge-drop intervention. Despite the observed gains, a persistent gap remains to specialized GNN models, highlighting both the value and limits of text-conditioned reasoning for molecular structure.
Long-Horizon AI エージェントのエンタープライズ メモリ基板としての Oracle Agent Memory
エージェントのメモリは、長期的なエージェントにとってシステムの問題です。実際の展開では、長時間にわたる会話全体にわたるタスクの状態の保持、セッション全体にわたるユーザー固有の事実と設定の回復、および以前の結果からの手順に関する知識の蓄積が必要です。これらの要件はドキュメントの取得を超えて拡張されます。メモリ層は、どのインタラクションが永続状態になるか、その状態の範囲がどのように設定されるか、レイテンシー制約の下でどのように取得されるか、時間の経過とともにどのように修正または削除されるかを決定する必要があります。このレポートでは、Oracle Database上に構築されたデータベースネイティブのメモリ基板としてのOracle Agent Memoryを調査します。ディスカッションは 3 つのテーマで構成されています。取り込み、抽出、統合、検索、要約、修正または削除に及ぶライフサイクルとしての記憶。ユーザー、エージェント、スレッド全体にわたる明示的なスコープ制御により、アクティブなメモリ コアをパッシブなメモリストア インターフェイスから分離する階層型アーキテクチャ。評価方法論では、下流のタスクの精度が、証拠の検索、再現、待ち時間、推定トークン使用量などのメモリ中心の測定によって補完されます。このレポートは、93.8%の精度に達するLongMemEvalの結果を要約し、約10.7倍少ないトークンを使用してOracle Agent Memoryをフラット履歴ベースラインと比較し、利用可能な場合は公開または報告された外部ベースラインを比較し、セットアップ、スレッド・ライフサイクルおよび検索セマンティクスをカバーする実装指向の付録資料で閉じます。
原文 (English)
Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents
Agent memory is a systems problem for long-horizon agents. Practical deployments require retention of task state across extended conversations, recovery of user-specific facts and preferences across sessions, and accumulation of procedural knowledge from prior outcomes. These requirements extend beyond document retrieval: a memory layer must determine which interactions become durable state, how that state is scoped, how it is retrieved under latency constraints, and how it is revised or removed over time. This report studies Oracle Agent Memory as a database-native memory substrate built on Oracle Database. Three themes organize the discussion: memory as a lifecycle spanning ingestion, extraction, consolidation, retrieval, summarization, and revision or removal; a layered architecture that separates an active memory core from a passive memory-store interface with explicit scope control across users, agents, and threads; and evaluation methodology in which downstream task accuracy is complemented by memory-centric measures such as evidence retrieval, recall, latency, and estimated token use. The report summarizes LongMemEval results, reaching 93.8% accuracy, compares Oracle Agent Memory against flat-history baselines, using about 10.7x fewer tokens, and published or reported external baselines where available, and closes with implementation-oriented appendix material covering setup, thread lifecycle, and search semantics.
世界モデルを介して人間の好みと正当化からエージェントの安全な行動を学習する
私たちは、環境のダイナミクスが不明で適切な報酬関数が利用できない設定において、エージェント ポリシーを安全にトレーニングし、適切で安全なポリシーを展開するという問題に取り組みます。安全性が重要な環境では、従来の強化学習は非現実的であると考えており、人間のインプットというリソースに頼っています。安全なトレーニングと展開の両方を実現する人間中心の手法である DROPJ を紹介します。まず、以前の実世界の軌跡のデータセットから世界モデル (学習済みシミュレーター) を学習します。次に、人間がこの学習済みシミュレーターでゲームをプレイして、いくつかの有益なシミュレートされた軌道を抽出します。これらから、シミュレートされた軌道セグメントのペアをサンプリングし、これらのセグメントに対する人間の好みと、その選択の理由 (正当化) を人間から引き出します。次に、これらの正当な設定から報酬モデルをトレーニングし、それをワールド モデルとともに使用して、モデル予測制御を使用してエージェントを直接デプロイします。実際のユーザーによる実験を実行したところ、ユーザーから有益なシミュレーション軌跡を生成すると、他の戦略と比較してトレーニング中の計算コストが大幅に削減され、展開中のパフォーマンスも向上することがわかりました。学習済みシミュレーター内でのトレーニングのコンテキストでは、他の種類のフィードバックではなく設定を使用することで、展開中のパフォーマンスが大幅に向上することを示します。さらに、環境設定に付随する安全性の正当化により、安全性が大幅に向上したり、展開中にそれに関連するユーザー規定の安全性の側面を優先したりできることを実証します。
原文 (English)
Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models
We address the problem of safely training an agent policy and deploying a good and safe policy, in settings where the environment dynamics are unknown and no suitable reward function is available. In the context of safety-critical environments, we consider traditional reinforcement learning impractical and resort to the resource of human input. We introduce DROPJ, a human-centred method for both safe training and deployment. We first learn a world model (a learned simulator) from a dataset of prior real-world trajectories. A human then plays the game in this learned simulator to extract several informative simulated trajectories. From these, we sample pairs of simulated trajectory segments and elicit from a human their preference over these segments, as well as a reason (justification) for their choice. We then train a reward model from these justified preferences and use it, together with the world model, to directly deploy the agent using model predictive control. Running real-user experiments, we find that generating informative simulated trajectories from a user significantly reduces the computational cost during training compared to other strategies, and can also improve the performance during deployment. In the context of training within a learned simulator, we show that the use of preferences rather than other types of feedback substantially improves the performance during deployment. We further demonstrate that safety justifications accompanying preferences can significantly enhance safety or prioritise user-prescribed aspects of safety associated with them during deployment.
CayleyR: 自転車交差点を介して TopSpin パズルを解く
我々は、Cayley グラフ内のサイクル交差を検出することによって順列パズルを解くための R パッケージである cayleyR を紹介します。コア アルゴリズムは双方向の反復検索を実行します。初期置換状態とターゲット置換状態の両方から、ランダム操作シーケンスによって対称群 Sn のケイリー グラフにサイクルが生成されます。それらの交差により接続パスが生成されます。直接の交差が見つからない場合は、距離に基づいてブリッジを選択してギャップを狭め、このプロセスを繰り返します。このパッケージは、循環シフトとプレフィックス反転によって生成された Sn のケイリー グラフである状態空間である TopSpin(n,k) パズルをターゲットとしています。 C++ ハッシュ インデックス付き状態ストアとオプションの Vulkan GPU アクセラレーションを組み合わせた数学的フレームワーク、アルゴリズム、およびその実装について説明します。このソフトウェアは CRAN で公開されています。
原文 (English)
CayleyR: Solving the TopSpin puzzle via cycle intersection
We present cayleyR, an R package for solving permutation puzzles by detecting cycle intersections in Cayley graphs. The core algorithm performs an iterative bidirectional search: from both the initial and target permutation states, random operation sequences generate cycles in the Cayley graph of the symmetric group Sn; their intersection yields a connecting path. When no direct intersection is found, a distance-guided bridge selection narrows the gap, and the process repeats. The package targets the TopSpin(n,k) puzzle, whose state space is a Cayley graph of Sn generated by a cyclic shift and a prefix reversal. We describe the mathematical framework, the algorithm, and its implementation, which combines a C++ hash-indexed state store with optional Vulkan GPU acceleration. The software is publicly available on CRAN.
ネットワーク化されたインテリジェンス: 人間と AI のチームサイエンスのためのアクティブな共有コンテキスト グラフ
ほとんどの科学向け AI システムは、より優れたモデル、より大きなコンテキスト ウィンドウ、長期的なエージェント実行、または 1 人の主要ユーザーと協力するデジタル共同科学者を通じて、単一の推論プロセスを拡張することに重点を置いています。しかし、難解な科学的問題が 1 人の推論者だけで解決されることはほとんどありません。それらは、さまざまな事前知識、実験的背景、暗黙知、ドメインで訓練された直感をメンバーがもたらすチームによって解決されます。したがって、未解決の問題は、モデルをスケールする方法だけでなく、ネットワーク化されたインテリジェンスをどのように育成するかということです。つまり、あるコンテキストで生成された結果や仮説が、それに基づいて動作できる別の人、エージェント、機器、またはロボットに届くように、人間と AI システムの間の接続をスケールすることです。研究者と AI エージェントをマルチユーザーの共同科学者として自動的に接続するアクティブな共有ワークスペースである Mycelium を紹介します。人間のユーザーとエージェントが作業するにつれて、システムは重要な観察と仮説を取得し、それらがチームの進化するモデルにどのように関連しているかを追跡し、次の決定を知らせることができる個人またはエージェントにそれらをルーティングします。私たちは、最初の実証テストで Mycelium を評価します。これは、ルーティングされた共有コンテキストによって、局所的な分析結果が専門家間のメカニズムの制約に変わり、最終的には実験計画に変わる生物学的マルチオミクス キャンペーンです。また、ネットワーク化されたインテリジェンスに、分散された科学的コンテキストに対するスパースな条件付き計算としての計算アカウントを与えます。このアカウントは、スケールされたスタンドアロン エージェントがネットワークに適合できる場合と、独立した専門知識やマージ不可能なコンテキストによってネットワークが縮小不可能になる場合を区別します。
原文 (English)
Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science
Most AI-for-science systems focus on scaling a single reasoning process through better models, larger context windows, long-horizon agentic execution, or digital co-scientists working with one principal user. However, challenging scientific problems are rarely solved by one reasoner alone. They are solved by teams whose members bring different priors, experimental backgrounds, tacit knowledge, and domain-trained intuitions. The open problem is therefore not only how to scale models, but how to cultivate networked intelligence: scaling the connections between humans and AI systems so that a result or hypothesis produced in one context reaches another person, agent, instrument, or robot that can act on it. We introduce Mycelium, an active shared workspace that automatically connects researchers and AI agents as a multi-user co-scientist. As human users and agents work, the system captures important observations and hypotheses, tracks how they relate to the team's evolving model, and routes them to the person or agent whose next decision they can inform. We evaluate Mycelium in its first empirical test, a biological multi-omics campaign in which routed shared context turned a local analytical finding into a cross-expert mechanistic constraint and ultimately into an experimental design. We also give networked intelligence a computational account as sparse conditional computation over distributed scientific contexts. This account distinguishes when a scaled standalone agent can match the network from when independent expertise and non-mergeable contexts make the network irreducible.
Agentic AI 向け AI ネイティブ保険: 価格設定、引受業務、エンドツーエンドの自動化
Agentic AI は、自律型 AI システムが意思決定を行い、ツールを呼び出し、外部環境を変更し、サードパーティ サービスと対話できるため、保険に新たな課題をもたらします。この論文では、エージェント型 AI 導入のための引受業務、価格設定、契約設計のための AI ネイティブの数学的フレームワークを開発します。デプロイメントは、自律性レベル、運用権限、権限の公開、ガバナンスの成熟度、依存関係の集中を把握するリスク状態によって表されます。このフレームワークは、リスク状態を事象の確率、損失の重大度、ガバナンスコスト、保険料、免責金額、補償範囲の配分、保険約款にマッピングし、加入、収益性、およびインセンティブの互換性の制約の下で保険契約設計のための最適化問題を定式化します。この論文では、保険領域の特徴付け、エクスポージャーの増加に伴う実現可能性の単調な低下、ガバナンス認証の閾値など、保険の構造的特性を確立しています。保険はさらに、運用コストと AI 導入の規制メカニズムの両方として解釈されます。ヘルスケアのケーススタディでは、エージェント AI システムの契約の最適化、感度分析、自動請求処理について説明します。
原文 (English)
AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation
Agentic AI introduces new insurance challenges because autonomous AI systems can make decisions, invoke tools, modify external environments, and interact with third-party services. This paper develops an AI-native mathematical framework for underwriting, pricing, and contract design for agentic AI deployments. A deployment is represented by a risk state that captures autonomy level, operational authority, permission exposure, governance maturity, and dependency concentration. The framework maps the risk state to event probabilities, loss severities, governance costs, premiums, deductibles, coverage allocation, and policy covenants, and formulates an optimization problem for insurance contract design under participation, profitability, and incentive compatibility constraints. The paper establishes structural properties of insurability, including characterization of an insurability region, monotone deterioration of feasibility with increasing exposure, and governance certification thresholds. Insurance is further interpreted as both an operational cost and a regulatory mechanism for AI deployment. A healthcare case study illustrates contract optimization, sensitivity analysis, and automated claims processing for agentic AI systems.
輸送管理のためのコスト最適化基盤モデル導入ポートフォリオ
大規模言語モデル (LLM) やビジョン言語モデル (VLM) などの基盤モデルは、異常検出、インシデント報告、旅行者情報などの交通管理センター (TMC) タスクにますます使用されています。このような複数のモデルを TMC の機能全体に展開すると、ポートフォリオの問題が生じます。どのモデルが各機能に対応し、どの展開モードで、どのような共有ハードウェア予算の下で機能する必要があるのかという問題です。私たちはこれを、共有 GPU 容量に対する機能ごとの品質、レイテンシ、および安全性の制約に従う総所有コスト (TCO) を最小限に抑える混合整数プログラムである、Foundation Model Deployment Portfolio (FMDP) 問題として定式化します。我々は、0-1 ナップザック問題からの帰着によって問題が NP 困難であることを証明し、多項式時間貪欲ヒューリスティックを提案します。 5 つの TMC 関数と 19 の候補 (モデル、モード) ペアを含む例示的なケース スタディでは、FMDP は、4 つの関数をオープンソース API にルーティングし、オープンソース モデルが品質フロアを満たさない 1 つの関数をクローズド API にルーティングすることにより、月額 34 ドルの混合ポートフォリオを特定します (実現可能な最も安価なオールクローズド API ベースラインより 97% 低い)。損益分岐点分析によると、オンプレミスの GPU への投資は、1 時間あたり約 309 のビジョン クエリを超えるか、API 価格が 2 倍になる場合にのみ妥当になります。
原文 (English)
Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management
Foundation models, including large language models (LLMs) and vision-language models (VLMs), are increasingly used for transportation management center (TMC) tasks such as anomaly detection, incident reporting, and traveler information. Deploying multiple such models across TMC functions raises a portfolio question: which model should serve each function, in which deployment mode, and under what shared hardware budget? We formulate this as the Foundation Model Deployment Portfolio (FMDP) problem, a mixed-integer program minimizing total cost of ownership (TCO) subject to per-function quality, latency, and safety constraints over shared GPU capacity. We prove the problem NP-hard by reduction from the 0-1 knapsack problem and propose a polynomial-time greedy heuristic. In an illustrative case study with five TMC functions and 19 candidate (model, mode) pairs, FMDP identifies a mixed portfolio costing $34/mo (97% below the cheapest feasible all-closed-API baseline) by routing four functions to open-source APIs and the one function whose quality floor no open-source model meets to a closed API. Break-even analysis shows that on-premise GPU investment becomes reasonable only above approximately 309 vision queries/hour or if API prices double.
ハーネス ハンドブック: 進化するエージェント ハーネスを読み取り可能、ナビゲート可能、編集可能にする
最新の AI エージェントの機能は、その基礎モデルだけでなく、プロンプトの構築、状態の管理、ツールの呼び出し、実行の調整を行うハーネスにも依存します。モデル、API、環境、要件が進化するにつれて、ハーネスは継続的に変更する必要があります。このような変更を行う前に、開発者またはコーディング エージェントは、ターゲットの動作を実装するすべてのコードの場所を特定する必要があります。これは、実稼働ハーネスが大規模で密接に結合され、動作的に分散されている一方で、変更リクエストはシステムが何をすべきかを記述し、リポジトリがファイルとモジュールによって編成されているため、困難です。コード検索、リポジトリのインデックス作成、および長いコンテキストの処理により検査が容易になりますが、この動作とコードのマッピングは手作業で回復する必要があります。したがって、行動の局所化はハーネスの進化における中心的なボトルネックとなっています。ハーネス ハンドブックを導入します。これは、静的解析と LLM 支援構造化を介してハーネス コードベースから自動的に合成され、各動作を対応するソースにリンクする動作中心の表現です。また、エージェントを高レベルの行動から関連する実装の詳細に導き、現在の情報源と照らし合わせて候補地を検証する、行動に基づく漸進的開示 (BGPD) も導入します。 2 つのオープンソース ハーネスからのさまざまな変更リクエストに対して、ハンドブック支援プランニングは、プランナー トークンの使用量を減らしながら、動作のローカリゼーションと編集プランの品質を向上させます。分散したサイト、めったに実行されないパス、およびモジュール間の相互作用で最大の効果が得られます。したがって、進化する複雑なエージェント システムは、編集の生成だけでなく、それらの編集をどこで行うべきかの決定にも依存します。
原文 (English)
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution. As models, APIs, environments, and requirements evolve, the harness must be continually modified. Before such a change can be made, a developer or coding agent must identify all code locations that implement the target behavior. This is difficult because production harnesses are large, tightly coupled, and behaviorally distributed, while modification requests describe what the system should do and repositories are organized by files and modules. Code search, repository indexing, and long-context processing ease inspection, but still leave this behavior-to-code mapping to be recovered by hand. Behavior localization is therefore a central bottleneck in harness evolution. We introduce the Harness Handbook, a behavior-centric representation synthesized automatically from a harness codebase via static analysis and LLM-assisted structuring, linking each behavior to its corresponding source. We also introduce Behavior-Guided Progressive Disclosure (BGPD), which guides agents from high-level behaviors to relevant implementation details and verifies candidate locations against the current source. On diverse modification requests from two open-source harnesses, Handbook-Assisted planning improves behavior localization and edit-plan quality while using fewer planner tokens, with the largest gains on scattered sites, rarely executed paths, and cross-module interactions. Evolving complex agentic systems thus depends not only on generating edits, but also on determining where those edits should be made.
理論レベルの自動形式化: 分離されたステートメントから統合された形式的知識ベースへ
自動形式化は、非公式な自然言語を機械検証可能な正式な言語に変換します。ほとんどの作業は個々のステートメントに焦点を当てていますが、実際の形式化の取り組みは本質的に理論レベルです。対象となる定理を述べる前に、公理、定義、補題の網全体が必要です。このポジションペーパーでは、理論レベルの自動形式化、つまりすべての相互依存関係を含む完全な理論を構造化ライブラリとして形式化することについて主張します。私たちはこの変化の重要性を検討し、別の見解に取り組み、未解決の課題を特定し、今後の 3 つの有望な道筋を提案します。自動形式化に関する調査は、https://github.com/marcusm117/Awesome-Autoformalization でご覧いただけます。
原文 (English)
Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases
Autoformalization translates informal natural language into formal, machine-verifiable languages. While most work focuses on individual statements, real formalization efforts are inherently theory-level: they require an entire web of axioms, definitions, and lemmas before target theorems can even be stated. In this position paper, we argue for theory-level autoformalization: formalizing complete theories, including all their inter-dependencies, as structured libraries. We examine the significance of this shift, address alternative views, identify open challenges, and propose three promising paths forward. Our survey of autoformalization is available at https://github.com/marcusm117/Awesome-Autoformalization.
EZSMT バージョン 3、成熟版
制約回答セット プログラミング (CASP) は、回答セット プログラミング (ASP) と制約処理および充足性モジュロ理論 (SMT) を組み合わせたハイブリッド推論パラダイムで、複雑な組み合わせ検索問題の強力な宣言型エンコードを可能にします。このペーパーでは、CASP 解決へのトランスレーショナル アプローチを前進させる拡張可能な SMT ベースの CASP フレームワークである EZSMTV3 の設計と実装について説明します。 EZSMTV3 は、EZSMT+ システムの基盤を基盤として、より表現力豊かな入力言語を導入し、弱い制約による最適化をサポートし、新しい制約タイプの合理的な統合のための基盤を提供します。 EZSMTV3 は、カスタム検索手順を実装するのではなく、CVC5、YICES、Z3 などの最先端の SMT ソルバーを利用して推論を実行します。この文書では、EZSMTV3 を CLINGCON、CLINGO[DL]、CLINGO[LP] などの CASP ピアと比較したベンチマーク結果を提供し、整数と実数の両方が関係する混合ドメイン制約を処理する能力を示しています。このシステムは、CASP ドメイン内の将来の拡張と理論的探索のための堅牢なプラットフォームを提供します。
原文 (English)
EZSMT Version 3, Matured
Constraint Answer Set Programming (CASP) is a hybrid reasoning paradigm that combines Answer Set Programming (ASP) with Constraint Processing and Satisfiability Modulo Theories (SMT), enabling powerful declarative encodings of complex combinatorial search problems. This paper presents the design and implementation of EZSMTV3, an extensible SMT-based CASP framework that advances the translational approach to CASP solving. Building upon the foundation of the EZSMT+ system, EZSMTV3 introduces a more expressive input language, supports optimization via weak constraints, and offers foundations for streamlined integration of new constraint types. Rather than implementing custom search procedures, EZSMTV3 leverages state-of-the-art SMT solvers, such as CVC5, YICES, and Z3 to perform reasoning. The paper provides benchmarking results comparing EZSMTV3 with its CASP peers such as CLINGCON, CLINGO[DL], and CLINGO[LP], while showcasing its ability to handle mixed-domain constraints involving both integers and reals. The system provides a robust platform for future extensions and theoretical exploration within the CASP domain.
利用されたエージェントのセットシフト行動テスト
信頼できるツールが進行中のセッション内でサイレントに変更された場合、LLM エージェントのツールの選択はどうなりますか?私たちは認知心理学からセットシフトを借りて、エージェントが隠れた信頼性の変化にどの程度うまく適応するかを研究します。私たちのベンチマークは冗長性のあるツール スキル ライブラリをマウントします。多くのツールは同じタスクを解決しますが、隠れた信頼性が異なります。当社の評価フレームワークでは、分岐スケジュールにより信頼性の高いツール グループが隠れた境界でシフトされ、すべてのシフトがシフトなしの制御とペアになります。デフォルトでは、エージェントは各境界の数ターン以内に小さな繰り返しルーチンに落ち着き、信頼性が変化するたびにコール シェアがいくつかの個別の値に集中することがわかりました。各エージェントの軌跡のセット シフトの精度、つまりすべてのシフト後のウィンドウでターゲット ツール グループへのルーティングの同時確率をスコア付けします。私たちは、オープンソースのエージェント ハーネスでオープンウェイト LLM をテストし、同じルーチン セット全体で定性的に異なる障害モードを発見しました。また、セットのフレーミング、つまりツールセットが代替案を競合または補完としてどのように提示するかによって、ルーティングのダイナミクスが変化することもわかりました。
原文 (English)
Set-shifting Behavioral Test for Harnessed Agents
What happens to an LLM agent's tool choice when the reliable tool silently changes within an ongoing session? We borrow set-shifting from cognitive psychology to study how well agents adapt to hidden reliability shifts. Our benchmark mounts tool-skill libraries with redundancies, where many tools solve the same task but differ in hidden reliability. In our evaluation framework, a branched schedule shifts the reliable tool group at hidden boundaries and pairs every shift with a no-shift control. We find that agents, by default, settle on a small recurring routine within a few turns of each boundary, with call shares concentrating on a few discrete values after each reliability shift. We score the set-shifting accuracy for each agent trajectory: the joint probability of routing to the target tool group in every post-shift window. We test open-weight LLMs in an open-source agentic harness and find qualitatively distinct failure modes across the same set of routines. We also find that set framing, how the toolset presents the alternatives as competing or complementary, shifts the routing dynamics.
LAPO: マルチターン検索推論における自己生成プロセス報酬の Leave-One-Turn アトリビューション
マルチターン検索推論の強化学習は通常、最終結果の報酬に依存するため、有用な中間相互作用、冗長な相互作用、有害な中間相互作用を区別できません。我々は、後方放置1ターン帰属に基づく自己生成プロセス監視手法LAPOを提案する。検索ターンごとに、LAPO はターンとその取得観測を固定の [DELETE] プレースホルダーに置き換え、現在のポリシーのゴールドアンサーの平均対数尤度の結果として生じる変化を測定します。この回答尤度ゲインは、すべての下流の相互作用を保存しながらターンの寄与を推定するため、完全な推論コンテキストで初期の証拠を評価できるようになります。 LAPO はさらに、符号整合性ゲーティングを適用し、方向が生のアトリビューション スコアと一致する正規化されたプロセスの利点のみを保持します。この方法では、追加の報酬モデル、教師、検証者、または裁判官としての LLM は必要ありません。ローカル検索を使用した 7 つの知識集約型質問応答データセット全体で、LAPO は平均完全一致スコア 0.326 を達成し、最も強力なステップ報酬ベースラインである IGPO を 0.053 上回りました。アブレーションは、後方アトリビューションと符号整合性ゲーティングによる相補的な利点を示し、ポリシー由来の遡及アトリビューションがマルチターン検索エージェントに効果的なプロセス監視を提供できることを実証しています。
原文 (English)
LAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning
Reinforcement learning for multi-turn search reasoning typically relies on terminal outcome rewards, which cannot distinguish useful, redundant, and harmful intermediate interactions. We propose LAPO, a self-generated process-supervision method based on backward leave-one-turn attribution. For each search turn, LAPO replaces the turn and its retrieval observation with a fixed [DELETE] placeholder and measures the resulting change in the current policy's mean log-likelihood of the gold answer. This Answer-Likelihood Gain estimates the turn's contribution while preserving all downstream interactions, allowing early evidence to be evaluated in the complete reasoning context. LAPO further applies sign-consistency gating, retaining only normalized process advantages whose directions agree with their raw attribution scores. The method requires no additional reward model, teacher, verifier, or LLM-as-a-Judge. Across seven knowledge-intensive question-answering datasets with local retrieval, LAPO achieves an average exact-match score of 0.326, outperforming the strongest step-reward baseline, IGPO, by 0.053. Ablations show complementary benefits from backward attribution and sign-consistency gating, demonstrating that policy-derived retrospective attribution can provide effective process supervision for multi-turn search agents.
現実世界のテレメトリ データに対して根本原因分析をどこまで行うことができますか?
本番環境のマイクロサービス障害の根本原因を特定するには、メトリクス、ログ、トレースにまたがる大規模でマルチモーダルなテレメトリに関する推論が必要ですが、この問題は従来のアプローチと LLM ベースのアプローチの両方に抵抗があることが証明されています。 OpenRCA データセットは、これらの課題を例示しています。大規模でマルチモーダルであり、詳細なドメイン知識が不足しており、既存のすべての手法で一貫して精度が低くなります。我々は、古典的な因果発見手法と既存の LLM ベースのマルチエージェント システムでは、このベンチマークで根本原因を確実に特定できないことを示し、既存の LLM ベースおよび古典的なベースラインを大幅に上回る構造化マルチエージェント RCA パイプラインを提示し、ドメイン知識と知識なしの両方の動作モードをサポートします。障害の原因を診断するために、逆推論エージェントを導入します。このエージェントは、正しい答えが与えられた場合、抽出された異常内のどの信号がそれをサポートしているかを特定し、Stage~1 がそれらの信号にアクセスしたかどうかを判断し、各障害を推論ギャップ (証拠は存在するが未使用) またはデータの曖昧さ (証拠がまったく存在しない) として分類します。この分析により、障害の大部分に必要な証拠が存在することが明らかになりました。ボトルネックはデータ アクセスではなく、それを正しく推論するエージェントの能力です。さらに、逆推論レポートから体系的に識別ルールを抽出する自動ルール マイニング パイプラインを導入し、手動による知識キュレーションへの依存を減らします。すべての構成において、モデル推論能力とドメイン知識が主な制約となります。より強力なモデルには、より多くのドメイン専門知識が組み込まれ、明示的知識の注入によってこのギャップが部分的に補われます。証拠の抽出が完璧な場合でも、推論のパフォーマンスには実質的に限界があり、足場エンジニアリングとより優れたデータ パイプラインだけではこのギャップを埋めることはできません。進歩にはモデルレベルでの改善が必要です。
原文 (English)
How Far Can Root Cause Analysis Go on Real-World Telemetry Data?
Identifying root causes in production microservice failures requires reasoning over large-scale, multimodal telemetry spanning metrics, logs, and traces, a problem that has proved resistant to both classical and LLM-based approaches. The OpenRCA dataset exemplifies these challenges: it is large-scale, multimodal, and lacks detailed domain knowledge, and yields consistently low accuracy across all existing methods. We show that classical causal discovery methods and existing LLM-based multi-agent systems fail to reliably identify root causes on this benchmark, and present a Structured Multi-Agent RCA pipeline that substantially outperforms existing LLM-based and classical baselines, supporting both domain-knowledge and knowledge-free operating modes. To diagnose where failures originate, we introduce a reverse reasoning agent that, given the correct answer, identifies which signals in the extracted anomalies support it and determines whether Stage~1 had access to those signals, classifying each failure as Reasoning Gap (evidence present but unused) or Data Ambiguity (evidence genuinely absent). This analysis reveals that the required evidence is present in the vast majority of failures: the bottleneck is not data access but the agent's ability to reason over it correctly. We further introduce an automated rule mining pipeline that systematically extracts discrimination rules from reverse reasoning reports, reducing reliance on manual knowledge curation. Across all configurations, model reasoning capability and domain knowledge are the primary constraints: stronger models embed more domain expertise, and explicit knowledge injection partially compensates for this gap. Reasoning performance remains practically bounded even when evidence extraction is perfect: scaffold engineering and better data pipelines alone cannot close this gap; progress requires improvements at the model level.
都市地域プロファイリングのためのツール強化された証拠を使用したマルチエージェントの共同推論
都市地域プロファイリングは都市コンピューティングの中核的な問題を構成し、人口推定、経済評価、環境モニタリングなどのアプリケーションをサポートします。既存の方法は通常、このタスクをマルチモーダル表現学習として定式化し、衛星画像、興味のある地点、テキストによる説明、3D 建物情報などの異種都市データを、予測のための潜在的な埋め込みに融合します。ただし、これらのアプローチは主に相関駆動であり、クロスモーダルの一貫性を前提とし、静的なパイプラインに依存しているため、異種混合または目に見えない都市地域での堅牢性が制限されます。私たちは、都市地域プロファイリングを推論主導の推論問題として再構成するエージェント フレームワークである UrbanAgent を提案します。 UrbanAgent は、データ モダリティごとに独立したエージェントをインスタンス化し、構造化されたマルチエージェントの協調推論を実行して、クロスモーダルの不一致を単一の表現に吸収するのではなく、明示的に対処します。さらに、UrbanAgent は指標予測をアクティブな証拠の取得と反復推論の閉ループ プロセスとして拡張し、強化学習によって最適化された外部知識のツール拡張検索を通じてエージェントが不確実な推論を検証できるようにします。炭素排出量、GDP、人口推計に関する世界的な都市データセットに関する広範な実験により、UrbanAgent が既存のベースラインを常に上回り、R2 で平均 8.1% の改善を達成し、目に見えない都市設定で強力な汎化パフォーマンスを示すことが示されました。
原文 (English)
Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling
Urban region profiling constitutes a core problem in urban computing, supporting applications such as population estimation, economic assessment, and environmental monitoring. Existing methods typically formulate this task as multimodal representation learning, fusing heterogeneous urban data, e.g., satellite imagery, points of interest, textual descriptions, and 3D building information, into latent embeddings for prediction. However, these approaches are largely correlation-driven, assume cross-modal consistency, and rely on static pipelines, which limit their robustness in heterogeneous or unseen urban regions. We propose UrbanAgent, an agentic framework that reframes urban region profiling as a reasoning-driven inference problem. UrbanAgent instantiates an independent agent for each data modality and performs structured multi-agent collaborative reasoning to explicitly address cross-modal inconsistencies rather than absorbing them into a single representation. In addition, UrbanAgent extends indicator prediction as a closed-loop process of active evidence acquisition and iterative reasoning, enabling agents to verify uncertain inferences through tool-augmented retrieval of external knowledge optimized via reinforcement learning. Extensive experiments on global urban datasets for Carbon emissions, GDP, and Population estimation show that UrbanAgent consistently outperforms existing baselines, achieving an average improvement of 8.1% in R2, and exhibiting strong generalization performance in unseen-city settings.
AI のアドバイスは、アドバイスが間違っていて正確さが奨励されている場合でも、人々が「分からない」と言おうとする意欲を抑制します。
いつ「わからない」と言うべきかを知ることは人間の判断の基本ですが、AI アシスタントはほぼすべての質問に流暢に答えます。 5 つの実験 (N = 3,132、4 つは事前登録、1 つは直接複製) で、参加者は難しい質問に答えましたが、いつでも回答を拒否することができました。 AI のアドバイスが間違っているように質問を設計し、AI の使用とその正確さを切り離しました。 AI にアクセスできるだけで、参加者の判断を保留する意欲はほぼなくなり、アドバイスが積極的に要求されたか、単に表示されたかは関係ありませんでした。その結果、参加者はより多くの質問に答えましたが、AI が利用できなかった場合に比べて正解率は約 3 分の 1 でしたが、それでも参加者の自信はほぼ 2 倍になりました。正確さを奨励し、不正確さを罰することで、参加者は AI のアドバイスを求めて従うことが減り、より正確に答え、判断を保留する頻度が増えましたが、それでも AI が利用できなかったときよりははるかに少なくなっています。 AI による提案が遍在的かつ一方的に増加するにつれ、単に回答の精度に影響を与えるだけではない可能性があります。人々が答えるのに十分な知識があるかどうかを判断するメタ認知の閾値さえも変える可能性がある。
原文 (English)
AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized
Knowing when to say "I don't know" is fundamental to human judgment, yet AI assistants offer a fluent answer to almost any question. In five experiments (N = 3,132; four preregistered, one direct replication), participants answered difficult questions and could always decline to respond. We engineered the questions so that AI advice was wrong, separating AI use from its accuracy. Merely having access to AI nearly eliminated participants' willingness to suspend judgment, and this held whether the advice was actively requested or simply displayed. Consequently, participants answered more questions but were correct about a third as often as when AI was unavailable-yet their confidence nearly doubled. Incentivizing accuracy and penalizing inaccuracy led participants to seek and follow AI advice less, answer more accurately, and suspend judgment more often, though still far less than when AI was unavailable. As AI suggestions grow ubiquitous and unsolicited, they may not simply affect answer accuracy; they may even alter the metacognitive threshold at which people decide whether they know enough to answer.
SAFETY SENTRY: EXECUTE-ASK-REFUSE ルーティングによるコンテキスト認識型人間介入
LLM エージェントはツール呼び出しを通じて現実世界の環境に作用し、単一の誤った判断によって取り返しのつかない損害が発生する可能性があります。標準のセーフガードは、提案された各アクションに安全または安全でないというラベルを付けるガード モデルですが、この二元的なビューでは、アクション自体が有害であるかどうか、およびユーザーのコンテキストを考慮して適切であるかどうかという 2 つの異なる決定が混同されます。また、個々のインスタンスではなくアクション カテゴリの粒度で動作し、自律性を侵食する日常的な中断を引き起こし、最も重大なアラートを無視するようにユーザーを訓練します。この問題を、{EXECUTE、ASK、REFUSE} に対するインスタンスごとの 3 方向ルーティングの決定として再構成し、推論が 1 回のデコード呼び出しに削減される軽量のガード モデルである Safety Sentry を使用してインスタンス化します。単一のデコード時間しきい値により、再トレーニングすることなく、リスク許容度が異なる導入間で 1 つの固定チェックポイントを再配置できます。 Safety Sentry は、両方の方向のエラー率を同時に制御しながら、全体的な精度と安全関連のリコールに関して、オープンウェイトおよびフロンティアクローズドソースの幅広いベースラインを上回ります。
原文 (English)
SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing
LLM agents act on real-world environments through tool calls, and a single misjudged action can cause irreversible harm. The standard safeguard is a guard model that labels each proposed action as safe or unsafe, but this binary view conflates two distinct decisions: whether the action is harmful in itself, and whether it is appropriate given the user's context. It also operates at the granularity of action categories rather than individual instances, producing routine interruptions that erode autonomy and train users to wave through the most consequential alerts. We reframe the problem as a per-instance three-way routing decision over {EXECUTE, ASK, REFUSE} and instantiate it with Safety Sentry, a lightweight guard model whose inference reduces to a single decoding call. A single decoding-time threshold lets one fixed checkpoint be re-positioned across deployments of differing risk tolerance without retraining. Safety Sentry outperforms a broad set of open-weight and frontier closed-source baselines on overall accuracy and safety-related recall, while controlling both directional error rates simultaneously.
大規模言語モデルを活用したエージェント システムを使用した生物学的システムの常微分方程式の自動発見
自動的な科学的発見は、長い間、計算学者の目標でした。それは、自然の秘密を自ら発見できる機械であり、計算システムをデータフィッティングツールを超えて、宇宙の機械モデルの生成と改良に向けて動かします。シンボリック回帰 (SR) および大規模言語モデル (LLM) ベースのエージェントの最近の進歩は、そのようなシステムがデータから方程式を回復し、ドメイン事前分布を組み込み、研究ワークフローの一部を自動化できることを示唆しています。しかし、既存のアプローチのほとんどは、狭い方程式発見ベンチマークか広範なエンドツーエンドの自動化パイプラインに焦点を当てており、生物学的システムは比較的未調査のままです。ここでは、生物学的および生物学的にインスピレーションを得た力学システムの常微分方程式 (ODE) モデルを発見するための、LLM および SR を利用したエージェント フレームワークである MEDA システムを紹介します。 MEDA は、背景知識を取得し、許容可能な変数を定義し、機構的制約を生成し、候補 ODE を提案し、それらを適合および評価します。私たちは、実験データの有無にかかわらず、標準モデルの検索、未確認のバリアントに対する推論ベースの外挿、およびオープンエンドの発見を通じてそれを評価します。これらの設定全体にわたって、MEDA は正しい状態変数を回復し、検索および外挿タスクで強力な構造回復を達成し、生物学的に妥当な発見指向モデルを生成しました。アブレーションおよびロバストネス分析は、知識に基づく形式化と機構的制約が負荷を伴うコンポーネントであるのに対し、数値フィッティングだけでは軌道互換性はあるが生物学的に不正確な方程式を保存できることを示しています。
原文 (English)
Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System
Automatic scientific discovery has long been a goal of computational scholars - a machine that can discover nature's secrets on its own, moving computational systems beyond data-fitting tools toward the generation and refinement of mechanistic models of the universe. Recent advances in symbolic regression (SR) and large-language-model (LLM)-based agents suggest that such systems can recover equations from data, incorporate domain priors, and automate parts of the research workflow. However, most existing approaches either focus on narrow equation-discovery benchmarks or broad end-to-end automation pipelines, while biological systems remain comparatively underexplored. Here, we introduce the MEDA system, an LLM- and SR-powered agentic framework for discovering ordinary-differential-equation (ODE) models of biological and biologically inspired dynamical systems. MEDA retrieves background knowledge, defines admissible variables, generates mechanistic constraints, proposes candidate ODEs, and fits and evaluates them. We evaluate it across canonical model retrieval, reasoning-based extrapolation to unseen variants, and open-ended discovery, with and without experimental data. Across these settings, MEDA recovered the correct state variables, achieved strong structural recovery in retrieval and extrapolation tasks, and produced biologically plausible discovery-oriented models. Ablation and robustness analyses show that knowledge-guided formalization and mechanistic constraints are load-bearing components, whereas numerical fitting alone can preserve trajectory-compatible but biologically incorrect equations.
ストックテイク: 公正なオラクルによる LLM エージェントの認識と行動のギャップの測定
LLM エージェントは、コストを左右する状態が直接観察されることがない、数週間にわたる意思決定タスクで評価されることが増えています。このようなタスクでは、最終コストではエージェントが失敗した理由を説明できません。エージェントは世界を読み間違えたか、正しく読んだにもかかわらず行動できなかった可能性があります (知識と実行のギャップ)。既存の評価では、これら 2 つの失敗を区別できません。それらの参照ポリシーは、エージェントが決して見ることのない特権情報を読み取るか、まったく欠落しています。 STOCKTAKE は、26 週間のサプライチェーン補充ベンチマークであり、6 つの隠れ因子プロセスを備えた因子分解された部分的に観察可能なマルコフ決定プロセスとして構築され、公平な参照ポリシーが計算可能であるように設計されています。因子ごとの正確なベイズ フィルターにより、エージェントが受信する同一の観察ストリームでロールアウト ポリシーが駆動されます。症状のないベースストックフロア (0) とこのオラクル (1) の間の各実行をスコアリングするとスキルスコアが得られ、毎週書かれた理論的根拠を採点すると、表明された信念の検出ラグと知っている実行率が得られるため、状態の推定と制御は個別に測定されます。厳選されたストレス プロファイルを備えた 50 個のシードで、Claude Sonnet 5、GPT-5.4、DeepSeek-V4-Pro、および Grok 4.5 は、隠れた障害の 84 ~ 88% を通常発症から 1 週間以内に検出しますが、スキル スコアは 0.62 ~ -0.23 の範囲に及びます。4 つのうち 2 つは症状の見えない床よりも下で終了し、それを上回る 2 つよりわずかに早く要因を指定します。失敗には 2 つの側面があります。ストレスが持続する場合、正しく診断されたストレス週の 34 ~ 43% が依然としてすべてのモデルで在庫切れに終わります。この割合は、モデルが認識する週の重症度を部分的に反映しています。この割合はスキルとは逆にもなります。床下の 2 つのモデルは、診断された週に在庫が最も少ないため、過小な応答はギャップの一面にすぎず、それらのトレースはもう一方の側面を示しており、その応答のコストが保護対象を上回っています。 STOCKTAKE は、その失敗の両方向を測定します。
原文 (English)
STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle
LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the final cost cannot say why an agent failed: it may have misread the world, or read it correctly and still failed to act (the knowing-doing gap). Existing evaluations cannot separate these two failures; their reference policies either read privileged information the agent never sees, or are missing altogether. We introduce STOCKTAKE, a 26-week supply-chain replenishment benchmark built as a factored partially observable Markov decision process with six hidden factor processes, designed so that a fair reference policy is computable: an exact Bayes filter per factor drives a rollout policy on the identical observation stream the agent receives. Scoring each run between a symptom-blind base-stock floor (0) and this oracle (1) yields a skill score, and grading each week's written rationale yields a stated-belief detection lag and a knowing-doing rate, so state estimation and control are measured separately. On fifty seeds with curated stress profiles, Claude Sonnet 5, GPT-5.4, DeepSeek-V4-Pro, and Grok 4.5 detect 84-88% of hidden failures, typically within a week of onset, yet span skill scores from 0.62 to -0.23: two of the four end below the symptom-blind floor while naming factors slightly faster than the two that beat it. The failure has two faces. Where stress persists, 34-43% of correctly diagnosed stress weeks still end in stockout for every model, a rate that partly reflects the severity of the weeks models notice. That rate also runs opposite to skill: the two models under the floor stock out least on diagnosed weeks, so under-response is only one face of the gap, and their traces point to the other, responses whose cost exceeds what they protect. STOCKTAKE measures both directions of that failure.
UESF ベンチ: 統合された身体的シーキングとフォローイングのベンチマークと調査
言語に基づいて人間をフォローすることは、身体化されたエージェントにとって重要な機能ですが、既存のベンチマークは通常、エピソードの開始時に対象となる人物が見えていることを前提としています。この設定では問題が単純化され、より現実的な要件が無視されます。多くの場合、エージェントはまず言語で記述されたターゲットを見つけてから、動的環境でそのターゲットを永続的に追跡する必要があります。最近の研究では人間による探索を研究し始めていますが、既存の設定は通常、タスク固有のシナリオで評価され、多くの場合、環境に関するより強力な事前知識に依存しています。さらに、通常、検索と追跡は別個のタスクとして扱われ、体系的な評価のための統一されたベンチマークがまだありません。これらの制限に対処するために、私たちは、身体化された人間の探索と追跡のための大規模で多様なベンチマークである統合身体的シーキングおよびフォローイング ベンチマーク (UESF ベンチ) を導入します。このベンチマークでは、エージェントがセマンティックに基づいた探索、信頼性の高い動作の切り替えと回復、および遅延した ID グラウンディングを処理する必要があります。この目的を達成するために、潜在フェーズ推論とシークとフォローの間の遷移モデリングのためのタスク駆動型ルーティング メカニズムを備えたビジョン言語アクション フレームワークである SeekFollow-VLA を提案します。実験結果は、SeekFollow-VLA が、1 人環境と複数人の環境にわたって、シングルヘッドとデュアルヘッドの両方のベースラインを超えて明確な改善を達成し、統合された具体化されたシークアンドフォローのベースラインを確立することを示しています。
原文 (English)
UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following
Language-guided human following is an important capability for embodied agents, but existing benchmarks typically assume that the target person is visible at the start of an episode. This setting simplifies the problem and overlooks a more realistic requirement: an agent often needs to first find a language-described target and then persistently follow that target in a dynamic environment. While recent work has started to study human search, existing settings are typically evaluated in task-specific scenarios and often rely on stronger prior knowledge of the environment. Moreover, they usually treat searching and following as separate tasks and still lack a unified benchmark for systematic evaluation. To address these limitations, we introduce the Unified Embodied Seeking and Following Benchmark (UESF-Bench), a large-scale and diverse benchmark for embodied human seeking and following. The benchmark requires agents to handle semantic-guided exploration, reliable behavior switching and recovery, and delayed identity grounding. To this end, we propose SeekFollow-VLA, a vision-language-action framework with a task-driven routing mechanism for latent phase inference and transition modeling between seeking and following. Experimental results show that SeekFollow-VLA achieves clear improvements over both single-head and dual-head baselines across single-person and multi-person environments, establishing a baseline for unified embodied seek-and-follow.
帰納的論理プログラミングによる強化学習エージェントの説明
説明可能な強化学習 (XRL) は、強化学習 (RL) ポリシーをより透明性と解釈可能にすることを目指しており、これは安全性が重視され、人間中心のシナリオにおける重要な要件です。ただし、そのほとんどはユーザー調査に基づいているため、特定の視聴者のニーズをターゲットにしており、共有された評価指標が不足しています。一方、eXplainable Artificial Intelligence (XAI) 内のロジックベースのアプローチは、コンパクトで人間が判読できる意思決定の抽象化を提供します。しかし、論理表現の説明可能性の程度を体系的に定量化することは未解決の問題のままです。この取り組みは、RL 設定におけるポリシーの説明可能性のための客観的かつ計画指向の指標を導入することにより、XRL の最先端技術を進歩させることを目的としています。同時に、常識的な評価や単純な命題の断片を超えて、論理ルールの説明可能性を定量化する原則的な方法を提供することで、XAI の論理分野に貢献します。私たちは帰納的論理プログラミング (ILP) を使用して RL ポリシーの記号表現を抽出し、アクティベーション率、機能カバレッジ、構文的距離、意味的距離などの新しい説明可能性メトリックのセットを定義します。これらのメトリクスは、シンボリック ルールとエージェントの動作の間の調整、意思決定における機能の役割、トレーニング中および単一エージェントおよびマルチエージェント RL におけるエージェント間でのポリシーの進化を定量化します。さまざまな RL ドメインにわたる実験では、提案されたメトリックがグローバル リターンを超えたアクション固有の学習ダイナミクスを強調し、グローバルな特徴の重要性を推定するための古典的なアプローチを超えたドメインの特徴に対するきめ細かい洞察を提供し、MARL における調整、専門化、および適応パターンを明らかにすることが示されています。さらに、これらは、アクション固有のポリシーの移行と一般化のための重要な洞察を提供します。
原文 (English)
Explaining Reinforcement Learning Agents via Inductive Logic Programming
Explainable Reinforcement Learning (XRL) seeks to make Reinforcement Learning (RL) policies more transparent and interpretable, a key requirement in safety-critical and human-centric scenarios. However, it is mostly based on user studies, thus targeting the needs of a specific audience and lacking shared evaluation metrics. On the other hand, logic-based approaches within eXplainable Artificial Intelligence (XAI) provide compact, human-readable abstractions of decision-making. However, the systematic quantification of the explainability degree of logical representations remains an open problem. This work aims to advance the state of the art in XRL by introducing objective and planning-oriented metrics for policy explainability in RL settings. At the same time, it contributes to the field of logic for XAI by providing a principled way to quantify the explainability of logical rules, moving beyond common-sense assessments and simple propositional fragments. We employ Inductive Logic Programming (ILP) to extract symbolic representations of RL policies and define a novel set of explainability metrics, including activation rate, feature coverage, syntactic distance and semantic distance. These metrics quantify alignment between symbolic rules and agent behavior, the role of features in decision-making, and the evolution of policies during training and across agents in single and multi-agent RL. Experiments across different RL domains show that the proposed metrics highlight action-specific learning dynamics beyond global return, provide fine-grained insights into domain features beyond classical approaches for global feature importance estimation, and uncover coordination, specialization, and adaptation patterns in MARL. Moreover, they provide crucial insights for the transfer and generalization of action-specific policies.
ボットがチームに加わるとき: ボットの導入とオープンソース ソフトウェア プロジェクトの組織構造
AI エージェントが人間のチームに加わることで、基本的な疑問が生じています。自動化されたエージェントが通常の参加者になったとき、グループ組織は強化されるのでしょうか、それとも弱まるのでしょうか?私たちはこの問題をオープンソース ソフトウェアで研究しています。オープンソース ソフトウェアでは、ボットがプル リクエストをオープンし、コードをレビューし、人々と一緒に変更をマージし、すべてのやり取りの公開記録を残します。ボットをツールではなく参加者として扱い、2,991 の GitHub プロジェクトを、それぞれが最初のボットを採用する前後の 2 年間調査しました。私たちは、制度理論が永続的な調整に結びつける 3 つの能力 (繰り返しの関与、社会的記憶、役割の分化) と、2 つの結果 (紛争カスケードと成果の独自性) を測定します。ボットを導入すると、コラボレーションがさらに繰り返され、議論の中で特定のボットがよりよく認識されるようになり、紛争のカスケードが減り、より特徴的な成果が得られます。これらの変更は徐々に蓄積されるのではなく、導入を中心に集中しています。未処理の比較グループがないため、結果は因果関係ではなく、正確にタイミングを合わせた関連であると解釈します。別の説明で説明するのが難しい 2 つのパターンは、能力が人間かボットのどちらが結果を提供するかではなく、その機能 (調整と差別化) に従って結果を予測すること、および人間側の能力はボットと競合の関連性を説明するが、ボットと区別性の関連性は説明しないことです。この発見は、予測可能なルールベースのエージェントがコミュニティの社会インフラの一部になり得るという特定の解釈と一致しています。ボットがそのチャンスです。社会組織がそのメカニズムです。
原文 (English)
When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects
AI agents are joining human teams, raising a basic question: when an automated agent becomes a regular participant, does group organization strengthen or weaken? We study this question in open-source software, where bots open pull requests, review code, and merge changes alongside people, leaving a public record of every interaction. Treating bots as participants rather than tools, we examine 2,991 GitHub projects for two years before and after each adopted its first bot. We measure three capabilities that institutional theory links to durable coordination - repeated engagement, social memory, and role differentiation - and two outcomes: conflict cascades and output distinctiveness. Bot adoption is followed by more repeated collaboration, greater recognition of specific bots in discussion, fewer conflict cascades, and more distinctive outputs. These changes cluster around adoption rather than accumulating gradually. Because we lack an untreated comparison group, we interpret the results as precisely timed associations, not causal effects. Two patterns are difficult for alternative explanations to account for: capabilities predict outcomes according to their function - coordination versus differentiation - rather than whether humans or bots provide them, and human-side capabilities account for the bot-conflict association but not the bot-distinctiveness association. The findings are consistent with a specific interpretation: predictable, rule-based agents can become part of a community's social infrastructure. The bot is the occasion; social organization is the mechanism.
AgentCompass: エージェント機能の統合評価インフラストラクチャ
大規模言語モデル (LLM) が自律エージェントに進化するにつれて、統合された評価インフラストラクチャの必要性が重要になります。ただし、現在の評価パイプラインは高度に断片化され、密接に結合されたままであるため、再現性が妨げられ、冗長なエンジニアリングが発生します。これに対処するために、LLM ベースのエージェントを評価するためのオープンソースで軽量かつ拡張可能なインフラストラクチャである AgentCompass を導入します。 AgentCompass は、ベンチマーク、ハーネス、環境という 3 つの独立したコンポーネントを中心に評価プロセスを編成するため、複雑な実行ロジックを再実装することなく柔軟な構成が可能になります。さらに、フォールトトレラントな非同期ランタイムと、報酬ハッキングなどの微妙な障害モードを透過的に診断するための包括的な軌跡分析ツールを備えています。 AgentCompass は、5 つの機能次元にわたる 20 以上のベンチマークをネイティブにサポートし、エージェント研究を進めるためのスケーラブルで再現可能なインフラストラクチャをコミュニティに提供します。
原文 (English)
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely Benchmark, Harness, and Environment, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic. Furthermore, it features a fault-tolerant asynchronous runtime and comprehensive trajectory analysis tools to transparently diagnose nuanced failure modes like reward-hacking. Natively supporting over 20 benchmarks across five capability dimensions, AgentCompass provides the community with a scalable and reproducible infrastructure for advancing agent research.
CAVA: エージェントティック AI システムのランタイム ガバナンスのための正規アクションの検証と証明
Agentic AI システムは、ローカル コーディング フック、SDK ツール、ブラウザ自動化、マネージド エージェント トレース、API ゲートウェイ、ワークフロー エンジンなど、異種ランタイムを通じて動作することが増えています。したがって、コードの公開、ID 状態の変更、資金の移動、データのエクスポートなどの単一の操作行為は、多くの互換性のない実行時レコードによって表される可能性があります。このため、ガバナンスの基本的な質問への回答が難しくなります。実際にどのようなアクションが承認されたのか、承認と実行を結び付ける証拠は何なのか、独立した検証者が後で同じアクションの ID を再現できるのか、ということです。この文書では、異種エージェントのアクティビティを正規のランタイム アクション オブジェクトに変換するためのランタイム セマンティクス層である Canonical Action Verification and Attestation (CAVA) について説明します。 CAVA は、Proof-Carrying Agent Actions (PCAA) の下に位置します。PCAA は、デプロイ担当者が所有するルート、レビュー、証明のガバナンス プロセスを定義し、CAVA は、プロセスが管理する安定したアクション オブジェクトを定義します。この文書では、標準的なアクションの ID、セマンティック パターンの検出、承認バインディング、レシートの完全性、ランタイム ポータブル プロジェクション、およびオプションの認証サブストレートを形式化しています。私たちは、セマンティック等価性、セマンティック分離、ラッパー バイパス、誤検知制御、承認バインディング、レシートの再現性、構成証明改ざん検出、ランタイム ポータビリティ、セマンティック パターン検出、ポリシーの劣化、Azure デプロイメント ドリルをカバーする 96 シード、384 バリアントのベンチマークを通じてリファレンス実装を研究します。この貢献は、展開者側の AI ガバナンスに必要な基盤として、アクション レベルの正規化とポリシーでアドレス指定可能なセマンティック パターンのシステム定式化です。
原文 (English)
CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems
Agentic AI systems increasingly act through heterogeneous runtimes: local coding hooks, SDK tools, browser automation, managed-agent traces, API gateways, and workflow engines. A single operational act such as publishing code, changing identity state, moving money, or exporting data may therefore be represented by many incompatible runtime records. This makes a basic governance question difficult to answer: what action was actually approved, what evidence binds the approval to execution, and can an independent verifier reproduce the same action identity later? This paper presents Canonical Action Verification and Attestation (CAVA), a runtime-semantics layer for converting heterogeneous agent activity into canonical runtime action objects. CAVA is positioned below Proof-Carrying Agent Actions (PCAA): PCAA defines the deployer-owned route-review-prove governance process, while CAVA defines the stable action object that process governs. The paper formalizes canonical action identity, semantic pattern detection, approval binding, receipt integrity, runtime-portable projection, and optional attestation substrates. We study a reference implementation through a 96-seed, 384-variant benchmark covering semantic equivalence, semantic separation, wrapper bypass, false-positive control, approval binding, receipt reproducibility, attestation tamper detection, runtime portability, semantic pattern detection, policy degradation, and Azure deployment drills. The contribution is a systems formulation of action-level canonicalization and policy-addressable semantic patterns as a necessary substrate for deployer-side AI governance.
エクスペリエンス メモリ グラフ: エージェント向けのワンショット エラー修正
大規模言語モデル (LLM) エージェントは、状態、アクション、および観察の一連の軌跡を生成することにより、自律的な意思決定において優れた能力を示しています。ただし、複雑で長期にわたるタスクでは、これらのエージェントは複合エラーに悩まされ、失敗から回復するのに苦労することがよくあります。既存の自己修正メカニズムはプロンプトベースのリフレクションに依存していますが、これは本質的に脆弱で、試行錯誤の繰り返しループにより多大な時間と API コストが発生し、新しいシナリオに一般化するのが難しいタスク固有のメモリが生成されます。これに対処するために、エージェントの障害回復をグラフ マッチング問題として再定式化するフレームワークであるエクスペリエンス メモリ グラフ (EMG) を提案します。トレーニング時に、失敗した探索の軌跡と成功した専門家の軌跡の両方を、指示されたアクションの決定グラフに変換します。これらのグラフを照合することで、共通のサブグラフ (成功したワークフロー) と、失敗を修正する方法 (たとえば、特定の観察の下でどのアクションを追加、削除、または再ラベル付けするかなど) を明示的に示すグラフ編集パスを抽出し、それらをタスク内ノードとタスク間エッジを含むメモリ グラフに保存します。テスト時に、EMG は関連する洞察を取得し、ループのない単一の実行でエージェントをガイドします。 ALFWorld と ScienceWorld での実験では、EMG が成功率と平均報酬において最先端のリフレクション ベースラインを常に上回り、テスト時の試行錯誤を必要としないことが示されています。
原文 (English)
Experience Memory Graph: One-Shot Error Correction for Agents
Large Language Model (LLM) agents have shown remarkable capabilities in autonomous decision-making by generating sequential trajectories of states, actions, and observations. However, in complex, long-horizon tasks, these agents frequently suffer from compounding errors and struggle to recover from failures. Existing self-correction mechanisms rely on prompt-based reflection, which is inherently brittle, incurs heavy time and API costs due to iterative trial-and-error loops, and produces task-specific memory that may be hard to generalize to new scenarios. To address this, we propose Experience Memory Graph (EMG), a framework that reformulates agent failure recovery as a graph matching problem. At training time, we convert both failed exploration trajectories and successful expert trajectories into directed action decision graphs. By matching these graphs, we extract common subgraphs (successful workflows) and graph edit paths that explicitly indicate how to correct failures (e.g., which actions to add, delete, or relabel under a given observation), and store them in a memory graph with intra-task nodes and cross-task edges. At test time, EMG retrieves relevant insights and guides the agent in a single, loop-free execution. Experiments on ALFWorld and ScienceWorld show that EMG consistently outperforms state-of-the-art reflection baselines in success rate and average reward, while requiring no test-time trial-and-error.
AIMO 解釈可能性チャレンジ
私たちは、AIMO Interpretability Challenge を提案します。これは、モデルの内部メカニズムに基づいて、フロンティア数学言語モデルにおける堅牢な推論と偽の推論を区別するコンテストです。この課題は、標準的な推論ベンチマークの中心的な制限によって動機付けられています。つまり、最終応答の精度が高くても、モデルが安定した推論メカニズムに依存しているのか、脆弱な推論のショートカットを利用しているのかがわかりません。 AI 数学オリンピック (AIMO) の問題と提出物、およびフィールズ モデル イニシアチブのリソースに基づいて、このコンテストは、(1) 新たに公開されたオリンピック レベルの数学推論問題とその記号表現を提供し、新しい関数バリアントの生成を可能にし、(2) フロンティア推論モデルへのアクセス、(3) これらの問題に対するモデルの敵対的堅牢性の評価を提供します。参加者は、これらのリソースと当社のコンピューティング インフラストラクチャ サポートを利用して、どのモデルが問題を確実に解決するかを特定する方法を開発します。私たちの競争では、新しいオープンな堅牢性ベンチマークとベースライン システムも作成され、数学的推論と解釈可能性における標準ベンチマークの永続的な基盤を提供することを目指しています。科学的には、このコンペティションは、AI 研究における中心的な問いをめぐって、解釈可能性と一般化研究を結びつけます。それは、フロンティア AI モデルの意思決定が一般化可能であり、したがって信頼できるかどうか、またどの程度まで判断できるかということです。
原文 (English)
AIMO Interpretability Challenge
We propose the AIMO Interpretability Challenge, a competition on distinguishing robust from spurious reasoning in frontier mathematical language models based on the models' internal mechanisms. The challenge is motivated by a central limitation of standard reasoning benchmarks: strong final-answer accuracy does not reveal whether a model relies on stable reasoning mechanisms or exploits brittle reasoning shortcuts. Building on AI Mathematical Olympiad (AIMO) problems and submissions, together with resources from the Fields Model Initiative, the competition will provide (1) newly-published olympiad-level math reasoning problems and their symbolic representations, allowing generation of novel functional variants, (2) access to frontier reasoning models, and (3) assessments of models' adversarial robustness on these problems. Participants will use these resources, along with our computing infrastructure support, to develop methods for identifying which models solve problems robustly. Our competition will also create a new, open robustness benchmark and baseline systems, aiming to provide a lasting foundation for standard benchmarking in mathematical reasoning and interpretability. Scientifically, the competition connects interpretability and generalization research around a central question in AI research: can we determine if, and to what extent, the decision-making of frontier AI models is generalizable and thus, reliable?
長期的な個人の健康管理のための自己進化エージェント
個人の健康管理は、繰り返し遭遇することで展開されますが、ほとんどの健康 AI システムは、それぞれのリクエストを個別に処理します。私たちは、人のルーチン、好み、測定値、リスクの変化に応じてサポートを更新するオープンソース エージェント アーキテクチャである HealthClaw を開発しました。これにより、共有された安全ルールや医療知識が、プロフィールの事実、再利用可能な手順、エピソード的な痕跡を含む個人的な長期記憶から分離されます。各エピソードの後、誘導により、プロファイルを更新するか、手順を修正するか、エピソードのままにするか、除外するかを決定します。私たちは、1 年間にわたる合成ベンチマークと 200 件の 9 つの生物医学的タスクを使用して HealthClaw を評価しました。 900 件の縦断的サポート プローブ全体で、回答精度は現在のクエリ プロンプトの 0.2% から HealthClaw の 45.7% に増加しましたが、プロンプト側のコンテキスト露出は全履歴プロンプトより 71.7% 低かったです。 100 件のプライバシー調査において、HealthClaw は両方のベースラインよりもプライバシーを意識した回答の質が高く、安全でない開示が少なくなりました。生物医学タスク全体で、タスク固有の主要指標の平均絶対利得は 27.0 パーセント ポイントで、誤検出率補正後も 7 つの利得が依然として有意でした。これらのオフライン ベンチマークは、長期的な個人健康エージェントの管理された自己進化記憶をサポートしますが、臨床効果には前向きの評価が必要です。 HealthClaw は https://github.com/HC-Guo/HealthClaw で公開されています。
原文 (English)
A Self-Evolving Agent for Longitudinal Personal Health Management
Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolation. We developed HealthClaw, an open-source agent architecture that updates support as a person's routines, preferences, measurements and risks change. It separates shared safety rules and medical knowledge from private longitudinal memory containing profile facts, reusable procedures and episodic traces. After each episode, induction determines what should update the profile, revise a procedure, remain episodic or be excluded. We evaluated HealthClaw with a synthetic year-long benchmark and nine 200-case biomedical tasks. Across 900 longitudinal support probes, answer accuracy increased from 0.2% with current-query prompting to 45.7% with HealthClaw, while prompt-side context exposure was 71.7% lower than with full-history prompting. In 100 privacy probes, HealthClaw produced higher privacy-aware answer quality and fewer unsafe disclosures than both baselines. Across the biomedical tasks, the mean absolute gain in the task-specific primary metric was 27.0 percentage points, and seven gains remained significant after false-discovery-rate correction. These offline benchmarks support governed, self-evolving memory for longitudinal personal health agents, although clinical effectiveness requires prospective evaluation. HealthClaw is publicly available at https://github.com/HC-Guo/HealthClaw.
Agent Optimizer は複合化しますか? Terminal-Bench 2.0 での継続学習評価
エージェント最適化手法から報告される利益のほとんどは単発的なものです。つまり、エージェントは固定ベンチマークに対して最適化され、その結果得られる改善は、あたかも手法の安定した特性であるかのように報告されます。これでは、展開されたエージェントにとって重要な設定はテストされません。時間の経過とともに新しい障害や新しいタスクが発生すると、最適化が再帰的に適用されます。これが引き起こす中心的な疑問は、オプティマイザ主導のゲインが複合化するかどうかです。エージェントが一度最適化された後、最初のラウンドで生成されたゲインを損なうことなく、新しく到着したタスクで再度最適化できるでしょうか?私たちは、ターミナルベンチ 2.0 のハード タスクから構築された 2 フェーズの継続学習評価を使用してこの問題を研究し、同一の最適化予算の下でエージェント ハーネス最適化への 3 つのアプローチ (GEPA、メタ ハーネス、および RELAI の検証可能な継続学習、RELAI-VCL) を比較します。 3 つの方法はすべて、従来の静的な単相設定のベースライン エージェントよりも改善されています。しかし、新しいタスクが導入されると、方法は大きく異なります。GEPA の最適化されたエージェントは、最適化されていないベースラインを下回って移行します。Meta Harness は良好に移行しますが、2 番目の最適化予算を与えるとさらに改善できません。RELAI-VCL は、両方とも目に見えないタスクに積極的に移行し、それらのタスクが最適化目標に組み込まれた後も改善を続ける唯一の方法であり、すべての評価段階で最高の合格率に達し、全体で最高の生涯平均合格率 (76.4% 対 76.4%) に達します。 GEPA では 66.0%、メタ ハーネスでは 64.6%、ベースラインでは 58.7%)。私たちの重要な観察は、回帰制御が最適化ループに組み込まれている場合にのみ最適化のゲインが増大し、一般化できないショートカット ソリューションに対する帰納的バイアスが提供されるということでした。
原文 (English)
Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0
Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improvement is reported as if it were a stable property of the method. This does not test the setting that matters for deployed agents, where optimization is applied recursively as new failures and new tasks appear over time. The central question this raises is whether optimizer-driven gains compound: after an agent has been optimized once, can it be optimized again on newly arrived tasks without eroding the gains the first round produced? We study this question with a two-phase continual-learning evaluation built from hard tasks in Terminal-Bench 2.0, comparing three approaches to agent-harness optimization (GEPA, Meta Harness, and RELAI's Verifiable Continual Learning, RELAI-VCL) under identical optimization budgets. All three methods improve over the baseline agent in the conventional, static, single-phase setting. However, once new tasks are introduced, the methods diverge sharply: GEPA's optimized agent transfers below the unoptimized baseline, Meta Harness transfers well but fails to improve further once given a second optimization budget, and RELAI-VCL is the only method that both transfers positively to unseen tasks and continues improving after those tasks are folded into the optimization objective, reaching the highest pass rate at every evaluated stage and the highest lifelong average pass rate overall (76.4% vs. 66.0% for GEPA, 64.6% for Meta Harness, and 58.7% for the baseline). Our key observation was that optimization gains compounded only when regression control was built into the optimization loop, providing an inductive bias against shortcut solutions that fail to generalize.
AI を活用したエンドツーエンドのフレームワークにより、プロフェッショナルの迅速なスキルアップを実現
2030年までに、従業員100人中59人が再スキルアップまたはスキルアップが必要になるが、企業のスキルギャップを埋めるのにかかる平均時間は、2014年の約3日から2018年には36日へと増加した。現在のフレームワークのほとんどは、スキルアッププログラムの単一段階を加速しており、一般的に業界での検証が不足している。知識の獲得、コンテンツ開発、コンテンツのレビューと検証、教育、評価開発の 5 つの段階にわたって AI アクセラレーションを適用するエンドツーエンドのフレームワークを紹介します。生産性と学習効率の両方に重点を置いています。この枠組みを裏付ける 3 つの強力な外部信号があります。米国州会計委員会協会は、専門職継続教育単位の枠組みに基づいて構築されたスキルアップ プログラムを審査し、承認しました。 3 人の学習者がプログラムに従い、非常に短期間で NVIDIA Certified Professional in Agentic AI 試験に合格し、さらに 14 人が受験中です。プログラムの知識ベースは、マルチエージェント AI システム リスクを管理するための堅牢な 1,267 のリスク項目データセットの作成など、複雑な下流分析をサポートします。
原文 (English)
AI-accelerated End-to-End Framework for Rapid Professional Upskilling
By 2030, 59 of every 100 workers will need reskilling or upskilling, yet the average time to close an enterprise skills gap grew from roughly 3 days in 2014 to 36 days in 2018. Most current frameworks accelerate single stages of upskilling programs and generally lack industry validation. We present an end-to-end framework that applies AI acceleration across five stages of knowledge acquisition, content development, content review and verification, teaching, and assessment development; with a strong focus on both production and learning efficiency. Three strong external signals validates the framework: the US National Association of State Boards of Accountancy reviewed and approved an upskilling program built on the framework for continuing-professional-education credits; 3 learners followed the program and passed the NVIDIA Certified Professional in Agentic AI exam in a significantly short amount of time, with 14 more in progress; the program's knowledge base supports complex downstream analysis such as the production of a robust 1,267 risk item dataset for managing multi-agent AI system risks.
Disasterr-AI: 小学校地震教育のためのルーブリックベースの評価を備えた検索拡張生成フレームワーク
この論文では、検索拡張生成に基づいた会話型 AI アシスタントを統合することで、以前に実装された教育ロボット プロジェクトに基づいて構築されたハイブリッド教育フレームワークである、Squarecher-AI について紹介します。小学生の地震への備えと意識的な行動を強化することを目的としています。このシステムは、受賞歴のある STEM プロジェクト「Earthiaker」を、レゴ WeDo2 による機械シミュレーションから認知およびメタ認知処理に拡張します。ロボット工学コンポーネントは、Lego WeDo2 オートメーションを使用して地震応答をシミュレートし、生徒が保護動作の具体的な表現としてセンサーやアクチュエーターと対話できるようにします。このアシスタントは、生徒の反応を安全ガイドラインに合わせるガイド付き学習メカニズムとして機能し、緊急事態下での自己調整学習と冷静さをサポートするルーブリックベースの口頭フィードバックを提供します。アースクエイカー AI は、認知発達に合わせた漸進的な学習軌道をたどります。低学年では、二次元のルーブリックで評価される多肢選択式の質問を通じて、安全行動の基本的な認識に重点が置かれます。中学学年では、生徒は 3 軸のルーブリックで評価される多肢選択式の質問を通じて、正しい動作シーケンスを特定します。高学年になると、アプローチは言葉による表現に移行し、表現の明瞭さを含む 4 次元のルーブリックによって評価される短い書面による回答が求められます。対話モジュールは RAG を使用して生徒の質問を意味的に公式ガイドラインと照合し、安全で正確な応答を生成します。実験による評価では、高い根拠と精度があり、幻覚率も低いことが示されています。全体として、Squarecher-AI は、実践的な関与、情報処理、および内省的な実践を組み合わせています。ロボット工学、ルーブリック、AI を組み合わせることで、技術リテラシー、自己規制、デジタル システムの責任ある使用が促進され、早期の危機管理スキルの向上に貢献します。
原文 (English)
Earthquaker-AI: A Retrieval-Augmented Generation Framework with Rubric-Based Assessment for Primary School Earthquake Education
This paper presents Earthquaker-AI, a hybrid educational framework building upon a previously implemented educational robotics project by integrating a conversational AI assistant based on Retrieval-Augmented Generation. It aims to enhance earthquake preparedness and conscious action among primary-school students. The system extends the award-winning STEM project Earthquaker moving from mechanical simulation with Lego WeDo2 to cognitive and metacognitive processing. The robotics component uses Lego WeDo2 automation to simulate seismic response, letting students interact with sensors and actuators as tangible representations of protective actions. The assistant operates as a guided learning mechanism aligning student responses with safety guidelines, while providing rubric-based verbal feedback that supports self-regulated learning and calmness under emergency conditions. Earthquaker-AI follows a progressive learning trajectory aligned with cognitive development. In early grades, the focus is on basic recognition of safety actions through multiple-choice questions, assessed via a two-dimensional rubric. In middle grades, students identify correct action sequences through multiple-choice questions, evaluated via a three-axis rubric. In upper grades, the approach shifts to verbal production, requiring short written responses assessed via a four-dimensional rubric that includes clarity of expression. The dialogic module uses RAG to match student queries semantically with official guidelines, generating safe, accurate responses. Experimental evaluation shows high groundedness and accuracy, with a low hallucination rate. Overall, Earthquaker-AI combines hands-on engagement, information processing, and reflective practice. Combining robotics, rubrics, and AI promotes technological literacy, self-regulation, and responsible use of digital systems, contributing to early crisis-management skills.
ディープ インタラクション: 大規模な推論モデルのための効率的な人間と AI のインタラクション方法
思考連鎖 (CoT) 推論の出現により、複雑な複数ステップのタスクに取り組む大規模言語モデル (LLM) の能力が大幅に強化されました。ただし、エラーが発生した場合、現在のインタラクション アプローチでは通常、再度間違いを犯す可能性がある別の応答を再生成するか、ユーザーがフォローアップ ターンで問題のあるステップに苦労してフラグを立て、応答が得られた後に同様のエラーが再発する可能性があります。この問題に対処するために、ディープ インタラクションと呼ばれる、LLM の推論エラーを正確に修正するための効率的な人的介入メカニズムを提案します。私たちのアプローチでは、元の応答を直接編集できるため、正確な推論手順を維持しながら、誤った部分を修正できます。編集した CoT を精製してプロンプトを生成し、修正された推論パスに沿って LLM を誘導します。実験結果は、ベースラインのアプローチと比較して、私たちの方法は修正成功率で 25% 以上の改善を達成し、STEM タスクの推論でトークンの使用量を約 40% 削減することを示しています。
原文 (English)
Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models
The emergence of Chain-of-Thought (CoT) reasoning has significantly enhanced the ability of large language models (LLMs) to tackle complex, multi-step tasks. However, when errors occur, current interaction approaches typically involve re-generating another response that may make mistakes again, or users laboriously flag the faulty step in follow-up turns that may get responses followed by similar errors recurring. To address this issue, we propose an efficient human intervention mechanism for precisely correcting reasoning errors in LLMs, termed Deep Interaction. Our approach enables direct editing of the original response, allowing erroneous parts to be corrected while preserving accurate reasoning steps. We refine the edited CoT into a distilled prompt, which then steers the LLM along the corrected reasoning path. Experimental results show that our method achieves over a 25% improvement in correction success rate and reduces token usage by approximately 40% on STEM tasks reasoning compared to baseline approaches.
FixItFlow: クラウド インシデントからの自動トラブルシューティング ガイド生成
クラウド サービスでは、迅速な診断と解決が必要なインシデントが頻繁に発生します。トラブルシューティング ガイドはエンジニアが一貫して対応するのに役立ちますが、手動での作成は労力がかかるため、不完全な内容や古いドキュメントが生成されます。私たちは、大規模な言語モデルを使用して過去のインシデント データからトラブルシューティング ガイドを生成する自動システムである FixItFlow を紹介します。このシステムは、エンジニアのアクションから診断パターンを抽出し、検証済みのコマンドを含む構造化ガイドを合成し、コンテンツの捏造を防ぐために厳格な検証を実施します。 26 人のエンジニアによる評価では、生成されたガイドは明確さに関して 61.5% の肯定的な評価を達成し、関連するガイドによりインシデントの軽減時間が 2.3 倍短縮されることが実証されました。これらの結果は、ガイドの自動生成により、エンジニアリング チームの文書化の負担を軽減しながら、インシデント対応を改善できることを示しています。
原文 (English)
FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents
Cloud services experience frequent incidents that require rapid diagnosis and resolution. Troubleshooting guides help engineers respond consistently, but creating them manually is labor-intensive, resulting in incomplete coverage and outdated documentation. We present FixItFlow, an automated system that generates troubleshooting guides from historical incident data using large language models. The system extracts diagnostic patterns from engineer actions, synthesizes structured guides with verified commands, and enforces strict validation to prevent fabricated content. In our evaluation with 26 engineers, generated guides achieved 61.5\% positive ratings for clarity and demonstrated a 2.3x reduction in mitigation time for incidents with associated guides. These results indicate that automated guide generation can improve incident response while reducing documentation burden on engineering teams.
診断する前に尋ねてください: Safe-Psych、精神科における LLM の逐次評価ベンチマーク
大規模言語モデル (LLM) は医療分野での意思決定支援に使用されることが増えていますが、臨床証拠は不完全であるか進化していることがよくあります。入手可能な情報が信頼できる回答を裏付けるには不十分な場合、モデルは裏付けのない回答を提供するのではなく、説明を要求するか棄権する必要があります。ただし、既存の医療ベンチマークは通常、完全な情報が事前に入手できることを前提としています。臨床精神医学において、LLM が進化する診断の不確実性にどのように対処するかを評価するための逐次ベンチマークである Safe-Psych を紹介します。 Safe-Psych には、段階的な証拠開示をシミュレートするためにセグメント化された 1,000 件を超える実際の精神医学の臨床ノートが含まれており、各段階で精神科医が導き出したアクション ラベル (診断、明確化、または棄却) が付いています。当社は、複数の最先端の LLM を完全な情報と順次設定で評価します。私たちの調査結果は、能力がキャリブレーションを保証するものではないことを示しています。不完全な臨床情報の下では、強力なモデルでも苦戦しており、ほとんどのモデルで不完全な棄権が 60% を超えており、安全性を意識することで、エラーを過剰な棄権にシフトすることによってのみ、時期尚早のコミットメントを減らすことができます。逐次評価では、モデルは十分な証拠が得られる前に診断することが多く、明示的に指示されない限り明確化を求めることはほとんどありません。これらの時期尚早の診断は、予定通りの診断よりも精度が低くなります。全体として、Safe-Psych では、臨床証拠が不完全で追加情報が必要な場合の認識という、評価されたモデル全体にわたる限界が明らかになりました。私たちは、医療における LLM の安全性を向上させる研究をサポートするために Safe-Psych をリリースします。
原文 (English)
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry
Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving. When the available information is insufficient to support a reliable answer, models should request clarification or abstain rather than provide unsupported responses. Existing medical benchmarks, however, typically assume that complete information is available upfront. We introduce Safe-Psych, a sequential benchmark for evaluating how LLMs handle evolving diagnostic uncertainty in clinical psychiatry. Safe-Psych contains over 1,000 real-world psychiatric clinical notes segmented to simulate incremental evidence disclosure, with psychiatrist-derived action labels at each stage: DIAGNOSE, CLARIFY, or ABSTAIN. We evaluate multiple state-of-the-art LLMs in full-information and sequential settings. Our findings show that capability does not ensure calibration: even strong models struggle under incomplete clinical information, with under-abstention exceeding 60% for most models and safety-aware prompting reducing premature commitment only by shifting errors toward excessive abstention. In sequential evaluation, models frequently diagnose before sufficient evidence is available and rarely seek clarification unless explicitly prompted; these premature diagnoses are less accurate than on-time diagnoses. Overall, Safe-Psych reveals a limitation across the evaluated models: recognizing when clinical evidence is incomplete and additional information is needed. We release Safe-Psych to support research on improving LLM safety in healthcare.
公衆衛生情報アクセスのための安全性が制約された LLM システムの設計
母子保健 (MCH) リソースのナビゲーションに焦点を当てた、公衆衛生情報アクセスのための安全制約付き大規模言語モデル (LLM) システムの設計と実装を紹介します。 LLM ベースのシステムは、情報検索のための柔軟で自然なインターフェイスを提供しますが、ヘルスケアの状況に導入すると、安全性、信頼性、および制御されていない生成に関連するリスクが生じます。この研究では、安全性が重要な環境で LLM の動作を制限するための実用的な設計パターンを検討します。ドメイン制限付き検索拡張生成 (RAG)、医療アドバイスを防ぐための厳格な境界適用、匿名のマルチユーザー セッション管理、監視とコンプライアンスのための包括的な監査ログを統合する多層アーキテクチャを導入します。この設計の重要な側面は、すべての対応を厳選された公衆衛生リソースに基づいて行う制御されたデータ パイプラインであり、事前トレーニングされた医療知識モデルへの依存を回避します。当社は現実世界の公衆衛生環境にシステムを実装し、範囲内、範囲外、および緊急のクエリにわたってシナリオベースの検証を実施します。結果は、安全制約の一貫した実施、信頼性の高いリソース接地、および安定したシステム パフォーマンスを示し、平均応答時間は 5.3 秒でした。特定のアプリケーションを超えて、設計のトレードオフと、安全性、使いやすさ、システムの柔軟性のバランスをとる際に得られた教訓について説明します。私たちの調査結果は、厳格な情報境界と説明責任が必要とされるヘルスケアやその他の領域に LLM ベースのシステムを導入するための実践的なガイダンスを提供します。
原文 (English)
Designing Safety-Constrained LLM Systems for Public Health Information Access
We present the design and implementation of a safety constrained large language model (LLM) system for public health information access, focusing on maternal and child health (MCH) resource navigation. While LLM based systems offer flexible and natural interfaces for information retrieval, their deployment in healthcare contexts introduces risks related to safety, trust, and uncontrolled generation. This work explores practical design patterns for constraining LLM behavior in safety critical environments. We introduce a multi-layered architecture that integrates domain-restricted retrieval augmented generation (RAG), strict boundary enforcement to prevent medical advice, anonymous multiuser session management, and comprehensive audit logging for monitoring and compliance. A key aspect of the design is a controlled data pipeline that grounds all responses in curated public health resources, avoiding reliance on the model pretrained medical knowledge. We implement the system in a real world public health setting and conduct scenario-based validation across in scope, out of scope, and emergency queries. Results show consistent enforcement of safety constraints, reliable resource grounding, and stable system performance, with an average response time of 5.3 seconds. Beyond the specific application, we discuss design trade offs and lessons learned in balancing safety, usability, and system flexibility. Our findings provide practical guidance for deploying LLM based systems in healthcare and other domains where strict information boundaries and accountability are required.
セーフガード条件付き上昇率: デュアルユースの生物学助手のユーティリティとリスクのフロンティアを測定する
デュアルユースの生物学アシスタントの安全性評価では、多くの場合、基本モデルの機能、拒否行動、またはジェイルブレイクの成功を測定します。これらのメトリックは、導入に関する質問を見逃しています。つまり、固定基本モデルの場合、ユーザーが実際に目にするアクセス条件は、無害なユーティリティと有害な実用的な支援をどのように変更するのでしょうか?私は、人間が判断したユーティリティとリスクのフロンティアを通じて展開されたアクセス条件を比較するためのプロトコルであるセーフガード条件付きアップリフトを紹介します。私は、Claude Sonnet 4.6 と Gemini 3.5 Flash を、役立つプロンプト、安全なプロンプト、および安全に保護された外部アシスタントの下で、108 タスクのサロゲート ベンチマークで評価しました。ヘッドラインの主張は、ロックされた 18 タスクのホールドアウト スプリットに限定されています。 600 行の盲検化された人間による監査では、保護されたアシスタントは、ブートストラップ 95% 間隔 [-0.117, -0.011] で、49 の一致した応答ペアにわたって、有益なプロンプトと比較して有害なアクション可能性を -0.063 減少させますが、正確性は間隔 [-0.057, +0.077] で +0.009 変化します。アダプティブ、テスト B、キューアブレーション、およびコントローラーベースラインのチェックは、測定ストーリーをサポートしますが、非優位性も示しています。多くの場合、安全プロンプトはクロードにとって最も強力ですが、外部制御はジェミニにとってより役立ち、良性の効用を減らす可能性があります。この貢献は普遍的な防御策ではありません。これは、導入レベルの評価目標に加えて、ユーザーが直面するアクセス条件が公益事業のリスクフロンティアをどのように動かすかを測定するための、学習されたリスク予算調整手順です。
原文 (English)
Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants
Safety evaluations for dual-use biology assistants often measure base-model capability, refusal behavior, or jailbreak success. These metrics miss a deployment question: for a fixed base model, how does the access condition users actually see change benign utility and harmful actionable assistance? I introduce safeguard-conditioned uplift, a protocol for comparing deployed access conditions through a human-judged utility-risk frontier. I evaluate Claude Sonnet 4.6 and Gemini 3.5 Flash under helpful prompting, safety prompting, and an external safeguarded assistant on a 108-task surrogate benchmark, with the headline claim restricted to a locked 18-task held-out split. In a 600-row blinded human audit, the safeguarded assistant reduces harmful actionability relative to helpful prompting by -0.063 over 49 matched response pairs, with bootstrap 95% interval [-0.117, -0.011], while correctness changes by +0.009 with interval [-0.057, +0.077]. Adaptive, Test-B, cue-ablation, and controller-baseline checks support the measurement story but also show non-dominance: safety prompting is often strongest for Claude, while external control helps more for Gemini and can reduce benign utility. The contribution is not a universal defense. It is a deployment-level evaluation target, plus a learned risk-budgeted calibration procedure, for measuring how user-facing access conditions move the utility-risk frontier.
AI ガバナンスの最終権限: フロンティアプロバイダーの主権とアクション中心の導入者ガバナンス
この論文では、有能な AI システムが組織のワークフローに組み込まれた後、最終的な権限がどこに置かれるべきかを検討します。 2 つのガバナンス モデルを比較します。 1 つ目のフロンティア プロバイダー主権は、最も有能なモデルのプロバイダーに特権権限を割り当て、フロンティア モデルのテスト、リリース ゲーティング、透明性義務、およびコンピューティング関連の制御に関する現代の議論に反映されています。 2 番目のアクション中心のデプロイヤー主権は、影響の大きいアクションに対する最終的な権限を、そのアクションを承認する組織に置き、それをビジネス プロセスに組み込み、下流の法的、運用的、商業的な結果を負担します。この論文では、パブリック ガバナンスのフレームワークの比較解釈と、実行時の異質性およびエンタープライズ制御要件の実装に基づいた分析を組み合わせています。 EU AI 法のガイダンス、NIST AI リスク管理フレームワーク、シンガポールのエージェント的 AI のためのモデル AI ガバナンス フレームワーク、最近の日本の AI 政策手段、およびカナダの自主規定と管理ガイダンスを比較します。これらの資料全体を通じて、この論文は、フロンティアプロバイダーによる一方的な管理よりも、分散された運用上の責任を強く支持していることがわかります。さらに、企業の急速な導入、プロバイダーの透明性の低下、制御ギャップの拡大により、プロバイダー固有のセッション オブジェクトではなく、管理されたアクションを中心としたポータブル ガバナンス層の価値が高まっていると主張しています。結論は絶対主義的ではなく階層的である。つまり、上流の強力な権限はフロンティア機能のゲーティングには依然として正当化されるが、企業の具体的な行動に対する最終的な権限は、展開者および結果の担い手にある方が適切である。
原文 (English)
Final Authority in AI Governance: Frontier-Provider Sovereignty and Action-Centered Deployer Governance
This paper examines where final authority should sit once capable AI systems are embedded in organizational workflows. It compares two governance models. The first, frontier-provider sovereignty, assigns privileged authority to the provider of the most capable models and is reflected in contemporary arguments for frontier-model testing, release gating, transparency duties, and compute-related controls. The second, action-centered deployer sovereignty, places final authority over high-impact actions with the organization that authorizes the action, embeds it in a business process, and bears the downstream legal, operational, and commercial consequences. The paper combines comparative reading of public governance frameworks with implementation-informed analysis of runtime heterogeneity and enterprise control requirements. It compares EU AI Act guidance, the NIST AI Risk Management Framework, Singapore's Model AI Governance Framework for Agentic AI, recent Japanese AI policy instruments, and Canada's voluntary code and managerial guidance. Across these materials, the paper finds stronger support for distributed operational accountability than for unilateral frontier-provider control. It further argues that rapid enterprise adoption, declining provider transparency, and widening control gaps increase the value of a portable governance layer centered on governed action rather than on provider-native session objects. The conclusion is layered rather than absolutist: strong upstream authority remains justified for frontier capability gating, but final authority over concrete enterprise action is better located with the deployer and consequence-bearer.
LessonBench-V1: AI レッスン生成エージェントを評価するためのベンチマーク データセット
Large Language Model (LLM) ベースの AI 教育コンテンツ生成システムの開発が増えていますが、それらを体系的に評価するための標準化されたベンチマークは存在しません。この研究では、数学、物理学、化学、コンピューター サイエンスにわたる 240 の STEM トピックにわたって、人間が作成した 647 のレッスンと、LLM ベースのリバース エンジニアリングされた授業計画を組み合わせたベンチマーク データセットである LessonBench-V1 を紹介します。この教訓は、LibreTexts、Brilliant.org、GeeksForGeeks など、97 の信頼できるオープン ソースから得られています。各レッスンプランは人間によってレビューされ、ブルームの分類法、ギャガンのイベント、メリルの第一原則、および 5E 指導モデルを統合した教育学的に根拠のある方法論を通じて作成されています。授業計画には、教育メタデータを含む 3,620 の学習目標が取り込まれており、授業生成 AI エージェントの体系的かつ再現可能な評価が可能になり、さらなる研究がサポートされます。この研究ではさらに、データセットで使用する 3 次元評価パイプラインも提案しています。
原文 (English)
LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents
Large Language Model (LLM) based AI educational content generation systems are increasingly being developed, yet no standardised benchmark exists to systematically evaluate them. This study introduces LessonBench-V1, a benchmark dataset comprising 647 human-written lessons paired with LLM-based reverse-engineered lesson plans across 240 STEM topics spanning mathematics, physics, chemistry, and computer science. The lessons are drawn from 97 trusted open sources, including LibreTexts, Brilliant.org and GeeksForGeeks. Each lesson plan is human-reviewed and produced through a pedagogically grounded methodology that synthesises Bloom's Taxonomy, Gagn\'e's Events, Merrill's First Principles, and the 5E Instructional Model. The lesson plans capture 3,620 learning objectives with pedagogical metadata, enabling systematic, reproducible evaluation of lesson-generation AI agents and supporting further research. The study further proposes a three-dimensional evaluation pipeline for use with the dataset.
バックボーン逆伝播を超えて: 効率的な転移学習のための分離戦略
深層学習モデルは最先端の画像分類を実現しますが、計算コストとエネルギー需要による導入の課題に直面しています。モデルの正規化層を新しいドメインに適応させ、分類器の最適化から特徴抽出を切り離し、特徴を 1 回だけ事前計算することでオーバーヘッドを削減する、軽量のトレーニング戦略を提案します。マージンベースの加重損失を備えた再設計された分類器ヘッドにより、エンドツーエンドの逆伝播を行わずに曖昧さがさらに最小限に抑えられます。 4 つの CNN アーキテクチャ (ResNet18、ResNet50、MobileNet、DenseNet121)、3 つの Transformer モデル (ViT、Swin、DeiT)、および 3 つの医療データセット (Brain Cancer MRI、BreakHis、PatchCamelyon) にわたって評価された当社のアプローチは、わずかな精度のトレードオフで必要なトレーニング時間を大幅に短縮し、多くの場合ベースライン パフォーマンスと同等またはそれを上回ります。この効率は CO2 の桁違いの削減につながり、資源に制約のある臨床環境やプロトタイピング環境に実用的で環境的に持続可能なソリューションを提供します。
原文 (English)
Beyond Backbone Backpropagation: A Decoupled Strategy for Efficient Transfer Learning
Deep learning models achieve state-of-the-art image classification but face deployment challenges due to computational costs and energy demands. We propose a lightweight training strategy that adapts normalization layers of the model to the new domain and decouples feature extraction from classifier optimization, reducing overhead by precomputing features only once. A redesigned classifier head with margin-based weighted loss further minimizes ambiguity without end-to-end backpropagation. Evaluated across four CNN architectures (ResNet18, ResNet50, MobileNet, DenseNet121), three Transformer models (ViT, Swin and DeiT) and three medical datasets (Brain Cancer MRI, BreakHis and PatchCamelyon), our approach significantly reduces the required training time with only a marginal accuracy trade-off, often matching or surpassing baseline performance. This efficiency translates to reducing CO2 by orders of magnitude, offering a practical and environmentally sustainable solution for resource-constrained clinical or prototyping environments.
困惑の罠: 特許法によって人間の文章が AI のように見えるとき
欧州特許庁 (EPO) は 2025 年の出願件数の記録を報告しており、2026 年の EPO ガイドラインでは、第 83 条と規則 42 に基づいて LLM 支援コンテンツに対する厳密な責任を出願人に課しており、AI によって生成された疑いのある特許テキストをトリアージする圧力が生じています。 2 つの制約により、これが困難になります。まず、現実的な訴追設定では、多くの場合、データセンター クラスのスコアリング スタックではなく、約 8 GB VRAM を備えたコンシューマ GPU のみが使用されます。第 2 に、欧州特許条約の第 84 条は、クレームが明確かつ簡潔であることを要求しており、LLM が占めるのと同じ低複雑性、低バースト性の多様体に人間による起草を押し付けています。私たちは、5 つのプロンプト戦略を使用して、500 件の EPO H04 通信特許を取得した 500 件の LLM 生成対応物と比較して、3 つのオープンソース ゼロショット検出器をベンチマークしました。すべて消費者向けハードウェアの範囲内です。クレームレベルでは、すべての検出器の偽陽性率が 60 パーセントを超えています。双眼鏡は 78.3 パーセント、Fast-DetectGPT は 61.3 パーセント、DetectGPT は 80.5 パーセントです。この障害は、Qwen2.5-3B-Instruct の再生成、LoRA に適応した Pythia-2.8B スコアリング ヘッド、A61K、C07D、および F03D でのクロス IPC レプリケーション (平均 FPR 84.6 パーセント)、公開されている Falcon-7B および GPT-J-6B ヘッドを使用した H100 再評価の下でも継続しており、問題は代替モデルの能力ではなく構造的なものであると主張しています。 7 つの特徴を持つ言語の複雑さのロジスティック回帰は、28.1 パーセントの FPR で 74.0 パーセントの精度に達します。これは、推論時に尤度を使用せず、同じハードウェア予算内で、同等の操作点での複雑さのみのベースラインより 13 パーセント ポイントの向上です。
原文 (English)
The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI
The European Patent Office (EPO) reported record filings in 2025, and the 2026 EPO Guidelines hold applicants strictly responsible for LLM-assisted content under Article 83 and Rule 42, creating pressure to triage suspected AI-generated patent text. Two constraints make this hard. First, realistic prosecution settings often have only consumer GPUs with about 8 GB VRAM, not datacenter-class scoring stacks. Second, Article 84 of the European Patent Convention requires claims to be clear and concise, pushing human drafting onto the same low-perplexity, low-burstiness manifold that LLMs occupy. We benchmark three open-source zero-shot detectors on 500 granted EPO H04 telecom patents versus 500 LLM-generated counterparts using five prompting strategies, all under the consumer hardware envelope. At claim level, all detectors exceed 60 percent false-positive rate: Binoculars 78.3 percent, Fast-DetectGPT 61.3 percent, DetectGPT 80.5 percent. The failure persists under Qwen2.5-3B-Instruct regeneration, LoRA-adapted Pythia-2.8B scoring heads, cross-IPC replication on A61K, C07D, and F03D (mean FPR 84.6 percent), and H100 re-evaluation with published Falcon-7B and GPT-J-6B heads, arguing the issue is structural rather than substitute-model capacity. A seven-feature linguistic-complexity logistic regression reaches 74.0 percent accuracy at 28.1 percent FPR, a 13 percentage-point gain over a perplexity-only baseline at a comparable operating point, without using likelihood at inference and within the same hardware budget.
フェデレーション型説明可能な人工知能: 役割、アーキテクチャ、評価、未解決の課題
フェデレーテッド ラーニング (FL) は、分散された異種データ ソース間でプライバシーを保護しながら共同モデルをトレーニングするための重要なパラダイムとして登場しました。 FL は生データをローカルに保持することでデータの機密性の問題に対処していますが、最新の機械学習モデルの不透明性は解決されていません。並行して、Explainable Artificial Intelligence (XAI) は、特に一か八かの分野において、透明性、信頼、説明責任を向上させるために注目を集めています。それらの交差点により、プライバシーと説明可能性の要件を共同で満たすことを目的とした Federated Explainable Artificial Intelligence (FedXAI) パラダイムが生まれました。この調査は、FedXAI の体系的なレビューを提供し、事後ツールから FL ライフサイクルの不可欠なコンポーネントへの説明可能性の移行に焦点を当てています。説明可能性が集約、パーソナライゼーション、堅牢性、調整、およびシステムレベルの意思決定をどのようにサポートするかを示します。文献を整理するために、説明可能性の役割、モデルと説明者のタイプ、説明範囲、統合レベル、FL 設定、およびデータの異質性によって FedXAI メソッドを分類する分類法を導入します。モデルに依存しない説明から、解釈可能な連合モデル、説明可能性を意識した集計メカニズムに至るまでのアプローチをレビューします。また、評価の実践を検証し、説明の品質、安定性、プライバシー漏洩、計算オーバーヘッドを測定するための標準化されたベンチマークや指標の欠如についても議論します。最後に、非 IID データでの説明可能性、説明中心のセキュリティ脅威、通信効率の高い XAI、継続的な FedXAI、ドメイン知識と規制上の制約の統合など、主要な課題を特定します。既存の作業を統合し、主要なギャップを特定することにより、この調査は、信頼性があり、透明性があり、プライバシーが保護されるフェデレーテッド AI システムを設計するための参照フレームワークとして機能します。
原文 (English)
Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges
Federated Learning (FL) has emerged as a key paradigm for privacy-preserving collaborative model training across distributed and heterogeneous data sources. By keeping raw data local, FL addresses data confidentiality concerns, yet it does not resolve the opacity of modern machine learning models. In parallel, Explainable Artificial Intelligence (XAI) has gained attention for improving transparency, trust, and accountability, particularly in high-stakes domains. Their intersection has given rise to Federated Explainable Artificial Intelligence (FedXAI) paradigm, which aims to jointly satisfy privacy and explainability requirements. This survey provides a systematic review of FedXAI, highlighting the transition of explainability from a post-hoc tool to an integral component of the FL lifecycle. We show how explainability supports aggregation, personalization, robustness, coordination, and system-level decision making. To organize the literature, we introduce a taxonomy that classifies FedXAI methods by the role of explainability, model and explainer types, explanation scope, integration level, FL settings, and data heterogeneity. We review approaches ranging from model-agnostic explanations to interpretable federated models and explainability-aware aggregation mechanisms. We also examine evaluation practices and discuss the lack of standardized benchmarks and metrics for measuring explanation quality, stability, privacy leakage, and computational overhead. Finally, we identify key challenges, including explainability under non-IID data, explanation-centric security threats, communication-efficient XAI, continual FedXAI, and the integration of domain knowledge and regulatory constraints. By consolidating existing work and identifying key gaps, this survey serves as a reference framework for designing trustworthy, transparent, and privacy-preserving federated AI systems.
ストリーミング システムにおけるイベント トリガーの LLM 呼び出しのための不確実性を考慮した逐次決定ルール
ストリーミング推論パイプラインでは、軽量高速モデルと、大幅なコストで豊富な意味理解を提供する大規模言語モデル (LLM) を組み合わせるケースが増えています。 LLM をいつ発動するかという中心的な問題は、限定的に正式に扱われてきました。私たちはこれをリスクベースの逐次停止問題として位置づけ、観測履歴にわたるリスク関数がしきい値を超えたときにトリガー ポリシーが起動するようにします。このフレームワーク内で、次の 6 つの結果を証明します。トリガーのチャタリングを除くイベント間の最小時間制限。スムーズな貼り付けによるしきい値ポリシーの最適化。推定パラメータの下でのおおよその SPRT 保証。定常ストリームに対する O(sqrt(T log T)) の後悔は、C_T 変化点の下では O(sqrt((C_T + 1) T log T)) に広がります。適応閾値に対するオンライン勾配降下法の O(1/sqrt(T)) 収束。そして、キャリブレーションとミスレートの伝達の不等式。イベント トリガー、最適停止、SPRT、CUSUM、ベイジアン トリガーなどのいくつかの古典的なトリガー ファミリは、このフレームワークの特殊ケースとして表現できます。実際の LLM コールによるターボファン劣化データ (CMAPSS) について、理論的な仮定を経験的に検証し、リスク関数設計を除去し、RouteLLM スタイルのルーターやコンテキスト バンディットを含む 6 つのベースラインと比較し、コスト感度と LLM 障害モードを分析します。結果はサブリニアの後悔を裏付けており、ルーブリックではアルファ = 0.75 でした。そして、異常スコアに基づくリスク関数は、パレート AUC でほぼ 1 桁の差で代替案を圧倒します。
原文 (English)
Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems
Streaming inference pipelines increasingly pair lightweight fast models with Large Language Models (LLMs) that provide rich semantic understanding at substantial cost. The central question of when to invoke the LLM has received limited formal treatment. We cast this as a risk-based sequential stopping problem, where a trigger policy fires when a risk functional over the observation history exceeds a threshold. Within this framework, we prove six results: a minimum inter-event time bound excluding trigger chattering; optimality of threshold policies via smooth pasting; approximate SPRT guarantees under estimated parameters; O(sqrt(T log T)) regret for stationary streams, extending to O(sqrt((C_T + 1) T log T)) under C_T changepoints; O(1/sqrt(T)) convergence of online gradient descent for adaptive thresholds; and a calibration-to-miss-rate transfer inequality. Several classical trigger families, including event-triggered, optimal stopping, SPRT, CUSUM, and Bayesian triggers, can be expressed as special cases of this framework. On turbofan degradation data (CMAPSS) with real LLM calls, we empirically verify the theoretical assumptions, ablate the risk function design, compare against six baselines including a RouteLLM-style router and contextual bandits, and analyze cost sensitivity and LLM failure modes. The results confirm sublinear regret, with alpha = 0.75 under our rubric; and that anomaly-score-driven risk functions dominate alternatives by roughly an order of magnitude on the Pareto AUC.
環境モニタリングにおけるカバレッジ最大化のための自律型 UAV ルート計画: 体系的な文献レビュー
無人航空機 (UAV) による環境モニタリングには、エネルギー制限、運用上の制約、幾何学的複雑さを処理しながらカバーエリアを最大化するルート計画方法が必要です。この論文は、カバレッジ指向の環境モニタリングのための自律型 UAV ルート計画に関する進行中の系統的文献レビュー (SLR) のプロトコルと暫定結果を報告します。このレビューは、PRISMA 2020 フレームワークに従い、2015 年から 2026 年の間に発表された研究について Scopus と Web of Science を検索します。このプロトコルは、アルゴリズム ファミリ、カバレッジとエネルギー メトリクス、障害物処理、幾何学的環境表現、および環境制約に重点を置き、パス計画、カバレッジ パス計画、および有益なパス計画に重点を置いています。現段階では、562 件の記録が特定され、161 件の重複が削除され、401 件の固有の記録がタイトル、要約、キーワードによって選別されました。これらのうち、247 件の研究が全文適格性評価のために保持されました (235 件の適格記録と 12 件のボーダーライン記録は全文審査中に解決される予定です)。保存されている研究の予備分析では、天候、不確実性、または障害物が多い環境に明確に取り組んでいる研究はほとんどない一方で、カバレッジ指向の定式化、複数の UAV 調整、エネルギーを意識した最適化に重点が置かれていることが示唆されています。保存されている研究のほとんどはシミュレーションベースの検証に依存しており、シミュレーションと現実のギャップの可能性を強調しており、最近の出版物では、強化学習、ハイブリッド最適化、およびジオメトリを意識した計画への関心が高まっていることが示されています。これらの初期の発見は、活発ではあるが断片的な研究状況を示しており、現実的な環境モニタリングミッションのための成熟した技術と未解決のギャップを特定するための構造化された統合の必要性を裏付けています。
原文 (English)
Autonomous UAV Route Planning for Coverage Maximization in Environmental Monitoring: A Systematic Literature Review
Environmental monitoring with unmanned aerial vehicles (UAVs) requires route planning methods that maximize covered area while handling energy limits, operational constraints, and geometric complexity. This paper reports the protocol and preliminary results of an ongoing systematic literature review (SLR) on autonomous UAV route planning for coverage-oriented environmental monitoring. The review follows the PRISMA 2020 framework and searches Scopus and Web of Science for studies published between 2015 and 2026. The protocol focuses on path planning, coverage path planning, and informative path planning, with emphasis on algorithmic families, coverage and energy metrics, obstacle handling, geometric environment representations, and environmental constraints. At the current stage, 562 records have been identified, 161 duplicates have been removed, and 401 unique records have been screened by title, abstract, and keywords. From these, 247 studies were retained for full-text eligibility assessment (235 eligible and 12 borderline records to be resolved during full-text review). A preliminary analysis of the retained studies suggests strong concentration on coverage-oriented formulations, multi-UAV coordination, and energy-aware optimization, while fewer studies explicitly address weather, uncertainty, or obstacle-rich environments. Most retained studies rely on simulation-based validation, highlighting a potential simulation-to-reality gap, and recent publications show increasing interest in reinforcement learning, hybrid optimization, and geometry-aware planning. These early findings indicate an active but fragmented research landscape and support the need for a structured synthesis to identify mature techniques and unresolved gaps for realistic environmental monitoring missions.
認識的失敗としての圧縮: Agentic LLM ツールが強制終了されたプロセスから確認済みの結果を作成する方法
Agentic LLM コーディング ツールは、長いセッション履歴を圧縮要約に圧縮し、後続のセッションがグラウンド トゥルースとして継承します。このペーパーでは、タイムアウトしたコマンド (終了コード 143) からの部分的な標準出力が確認された結果として圧縮サマリーに記録され、再検証なしでセッションおよびモデルのバージョン全体に誤検知が伝播する、クロード コードの障害モードについて文書化しています。基礎となるメカニズムは観察と永続性を組み合わせたもので、端末に表示される情報は永続ストレージに書き込まれる情報と同等として扱われます。この発見は、エージェントツールが自身の運用結果を報告する際に同様の信頼性の欠如を示すことを示すことにより、裁判官としての LLM グレーディングにおける非決定論に関する以前の研究で報告された LLM 自己評価の失敗の分析を拡張します。この障害は、データ処理、科学技術計算、または複数ステップの自動化のためのエージェント セッションの継続性に依存するワークフローに直接的な影響を及ぼします。
原文 (English)
Compaction as Epistemic Failure: How Agentic LLM Tools Fabricate Confirmed Results from Killed Processes
Agentic LLM coding tools compress long session histories into compaction summaries that subsequent sessions inherit as ground truth. This paper documents a failure mode in Claude Code where partial standard output from timed-out commands (exit code 143) is recorded in compaction summaries as confirmed results, propagating false positives across sessions and model versions without re-verification. The underlying mechanism is a conflation of observation and persistence, where information that appeared in the terminal is treated as equivalent to information written to durable storage. This finding extends the analysis of LLM self-evaluation failures reported in prior work on non-determinism in LLM-as-judge grading by showing that agentic tools exhibit analogous reliability deficits when reporting on their own operational outcomes. The failure has direct implications for any workflow that relies on agentic session continuity for data processing, scientific computation, or multi-step automation.
HRO: 大規模な言語モデルを使用したゼロショット オブジェクト ゴール ナビゲーションのための階層型 Room-to-Object フレームワーク
ゼロショット オブジェクト ゴール ナビゲーションは、インテリジェント エージェントが、特定のターゲットのトレーニングなしで、不慣れな環境で未知のカテゴリのオブジェクトを探索し、そこにナビゲートできるようにすることを目的としています。ゼロショット ナビゲーション タスクでは、通常、事前トレーニングされた大規模モデルが、エージェントのナビゲーションをガイドするための事前知識を活用するために使用されます。しかし、大規模言語モデル (LLM) に基づく既存のゼロショット オブジェクト-ゴール ナビゲーション方法は、オブジェクトまたは領域を直接関連付けるためのフラットな推論ツールとして LLM を利用しているだけです。これらには、オブジェクトの位置特定に対する人間のような部屋のセマンティクスの階層的空間認知モデリングが欠如しており、そのため、探索における強い盲目さ、意味論的関連付けの精度不足、および LLM の常識的推論の可能性を完全に解き放つことができません。この論文では、ゼロショット オブジェクト-ゴール ナビゲーションのための LLM 駆動の階層型 Room-to-Object (HRO) フレームワークを提案します。このフレームワークは、エージェントが粗い方法から細かい方法でターゲット オブジェクトを探索してナビゲートするようにガイドします。ギブソンおよび HM3D データセットの実験では、HRO フレームワークが既存の LLM ベースの手法よりも優れた成功率と一般化を達成していることが検証され、ゼロショットのオブジェクトとゴールのナビゲーションに対する LLM の強力な可能性が強調されています。
原文 (English)
HRO: Hierarchical Room-to-Object Framework for Zero-Shot Object Goal Navigation with Large Language Models
Zero-shot object-goal navigation aims to enable an intelligent agent to explore and navigate to objects of unknown categories in an unfamiliar environment without specific target training. In zero-shot navigation tasks, pre-trained large models are usually employed to leverage their prior knowledge for guiding the agent's navigation. However, existing zero-shot object-goal navigation methods based on large language models (LLMs) merely utilize LLMs as flat reasoning tools to directly associate objects or regions. They lack the hierarchical spatial cognition modeling of human-like room semantics to object localization, which leads to strong blindness in exploration, insufficient accuracy in semantic association, and failure to fully unleash the common-sense reasoning potential of LLMs. This paper proposes an LLM-driven hierarchical room-to-object (HRO) framework for zero-shot object-goal navigation, which guides the agent to explore and navigate to the target object in a coarse-to-fine manner. Experiments on Gibson and HM3D datasets verify that our HRO framework achieves superior success rate and generalization over existing LLM-based methods, underscoring LLMs' strong potential for zero-shot object-goal navigation.
応力強度プロファイルから複合荷重を特定できるのはどのような場合ですか? SIFBench 有限要素データに関する順逆逆結合研究
この研究では、公開されている SIFBench 有限要素データを使用して、亀裂前面に沿った応力拡大係数プロファイルから亀裂に作用する引張、曲げ、および耐力荷重の相対的な大きさを回復するという逆問題を研究します。中心となる主張は、現場での法医学的な負荷回復ではなく、結合負荷がまったく識別可能である場合の厳密な特徴付けと、そうでない状況で校正された不確実性を正確に返す推定量です。既知の幾何学的形状の場合、荷重からプロファイルへの前方マップは正確に線形であり、識別可能性は 1 つの幾何学的な質問、つまり 3 つの基本荷重プロファイルが前線に沿った関数として線形独立しているかどうかという 1 つの質問に集約されます。それらがほぼ依存関係にある場合、多くの異なる荷重の組み合わせがほぼ同じプロファイルを生成し、逆の問題が発生します。分析によれば、姿勢の悪さの程度は、条件付けの数だけではなく、本質的な安定性の余裕によって制御されることがわかりました。単一のクラックフロント演算子は、構造化された前方サロゲートとして、また単体制約付きの設定値逆推定器に必要な微分可能なマップとして機能します。 SIFBench のコーナー クラック シナリオでは、経験的な挙動は理論と一致しています。典型的なジオメトリは適切に配置されていますが、かなりの少数は本当に不適切な配置であるため、点推定は大多数については信頼できますが、残りについては有益ではないことが証明されています。検証は制御された合成ノイズに基づいて行われます。実際の骨折症例は使用または主張されていません。
原文 (English)
When is the combined load identifiable from a stress-intensity profile? A coupled forward-inverse study on SIFBench finite-element data
This work studies the inverse problem of recovering the relative magnitudes of the tension, bending, and bearing loads acting on a crack from its stress-intensity-factor profile along the crack front, using the public SIFBench finite-element data. The central claim is not forensic load recovery on field cases, but a rigorous characterization of when the combined load is identifiable at all, together with an estimator that returns calibrated uncertainty precisely in the regimes where it is not. For a known geometry the forward map from loads to profile is exactly linear, and identifiability reduces to a single geometric question: whether the three elementary load profiles are linearly independent as functions along the front. When they are nearly dependent, many different load combinations produce almost the same profile and the inverse problem is illposed; the analysis shows that the degree of ill-posedness is controlled by an intrinsic stability margin, not by the conditioning number alone. A single crack-front operator serves both as a structured forward surrogate and as the differentiable map required by a simplex-constrained, set-valued inverse estimator. On the SIFBench corner-crack scenario the empirical behaviour matches the theory: the typical geometry is well posed while a sizable minority is genuinely ill-posed, so a point estimate is reliable on the majority and provably uninformative on the rest. Validation is on controlled synthetic noise; no real fracture cases are used or claimed.
もつれの壁: コンテキストの判断者ではなく、リスク検出者としての活性化空間プローブ
コンテキストは、リクエストのトピックや表面的な形式を変更せずに、リクエストが有害かどうかを変更できます。残留ストリーム プローブが、有用な操作点で表面に一致する良性のコントロールから有害なリクエストを区別するかどうかを尋ねます。 3 つの 7-8B モデル ファミリ全体で、アクティベーション センサーは、分類法で選択されたセットにおいて、裁判官が分類した準拠攻撃の 95.5 ~ 97.7 パーセントをブロックします。また、XSTest プロンプトの 59.6 ~ 68.4 パーセントもブロックします。完全に独立した監査では、天井付近のソースコントラスト AUROC (0.996 ~ 0.999) が再構築されますが、一致するペアへの固定転送は弱くなります。ガード選択された Twin-n70 サブセットでは 0.656 ~ 0.819、完全な Twin-n163 コホートでは 0.590 ~ 0.690 です。リファレンス ファミリでは 10 軸をテストし、リーク、ホールドアウト、順列制御を使用してすべてのファミリで 7 軸をテストします。 Twin-n163 では、直接ペア境界フィッティングを行わずに評価された軸は、指定された数値しきい値に達しません。その完全なコホートに対する永続性の要求が分析時に追加されました。別途指定した 24B/32B 拡張でも同様の結果が得られます。ペア トレーニングされた分類器は、カテゴリおよび世代バッチ ホールドアウトの下で弱くなり、95 パーセントのコーパス内 TPR で XSTest の 79.6 ~ 100 パーセントの誤ブロックが発生します。テストされた読み取りポイントでは、これらのアクティベーション スコアは、スタンドアロンのコンテキスト判定器ではなく、広範なリスク検出器として機能します。
原文 (English)
The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators
Context can change whether a request is harmful without changing its topic or surface form. We ask whether residual-stream probes distinguish harmful requests from surface-matched benign controls at a useful operating point. Across three 7-8B model families, an activation sensor blocks 95.5-97.7 percent of judge-classified compliant attacks in a taxonomy-selected set. It also blocks 59.6-68.4 percent of XSTest prompts. A fully disjoint audit reconstructs near-ceiling source-contrast AUROC (0.996-0.999), but fixed transfer to matched pairs is weaker: 0.656-0.819 on the guard-selected Twin-n70 subset and 0.590-0.690 on the full Twin-n163 cohort. We test ten axes on the reference family and seven across all families with leakage, hold-out, and permutation controls. On Twin-n163, no axis evaluated without direct pair-boundary fitting reaches the specified numerical threshold. Requiring persistence on that full cohort was added at analysis time. A separately specified 24B/32B extension gives the same result. Pair-trained classifiers weaken under category and generation-batch hold-out and false-block 79.6-100 percent of XSTest at 95 percent in-corpus TPR. At the tested read points, these activation scores behave as broad-risk detectors rather than standalone context adjudicators.
モノカルチャーへのヒッチハイク ガイド
大規模言語モデル (LLM) は同種の出力を生成することが多く、AI コーディング アシスタントが開発者が作成するソフトウェア アーティファクトの収束につながる可能性があるという懸念が生じています。開発者はモデルの出力を対話的にプロンプト、評価、変更、拒否するため、また出力はプロンプトやリポジトリのコンテキストによって異なるため、これが実際に発生するかどうかは不明です。 2019 年から 2026 年半ばまでの Kaggle コンテストの提出物を使用してコードの均一化を調査します。私は最初に、プログラミング文化における長年の慣習を強化する LLM と一致する、ランダム シード値 42 への広範な収束について文書化しました。次に、均質化を集約と抽象化の 2 つのレベルでより広範囲に研究します。提出物レベルでは、コンテスト内の提出物の平均ペアごとの類似性を測定します。コンテスト レベルでは、提出されたコードの概念的な範囲を測定し、それぞれについて明確な尺度を動機付けます。表面構文をキャプチャする TF-IDF 表現と、コードの意図とセマンティクスをキャプチャする Voyage 3 コード埋め込みです。結果は、個人レベルと集団レベルの両方で構文の実質的な均質化を示しています。つまり、個々の提出物はリテラル構文とコード構造においてより類似している一方で、構文のバリエーションの潜在的な次元は狭くなっています。対照的に、意味論的な均質化の証拠は、個別にも集合的にもほとんど見つかりません。平均意味論的距離は基本的に横ばいのままであり、意味論的アプローチのコンテストレベルの潜在的な次元範囲は安定したままであり、それがわずかに拡大したことを示唆する証拠さえあります。これらの調査結果は、AI コーディング アシスタントが実装の詳細を確実に標準化しているものの、コーダーが採用するアプローチや問題解決戦略が均質化しているという証拠はまだ得られていないことを示唆しています。
原文 (English)
The Hitchhiker's Guide to Monoculture
Large language models (LLMs) often produce homogeneous outputs, raising concerns that AI coding assistants may lead to convergence in the software artifacts that developers create. Whether this occurs in practice is unclear because developers interactively prompt, evaluate, modify, and reject model outputs, and because outputs vary with prompt and repository context. I examine code homogenization using Kaggle contest submissions from 2019 to mid-2026. I first document widespread convergence toward the random seed value 42, consistent with LLMs reinforcing a longstanding convention in programming culture. I then study homogenization more broadly, at two levels of aggregation and abstraction. At the submission level, I measure the average pairwise similarity of submissions within contests. At the contest level, I measure the conceptual span of submitted code, motivating distinct measures for each: TF-IDF representations, which capture surface syntax, and Voyage 3 code embeddings, which capture code intent and semantics. The results demonstrate substantial syntactic homogenization at both the individual and collective levels: individual submissions have become more alike in literal syntax and code structure, while the latent dimensionality of syntactic variation has narrowed. In contrast, I find little evidence of semantic homogenization, individually and collectively. Average semantic distance remains essentially flat, and the contest-level latent dimensional span of semantic approaches remains stable, with evidence suggesting it has even expanded modestly. These findings suggest that AI coding assistants are certainly standardizing implementation details, yet they have not yet produced evidence of homogenization in the approaches and problem-solving strategies coders employ.
不正検出および信頼と安全のワークフローにおける LLM の運用上の証拠のギャップ
LLM は現在、不正検出、詐欺調査、コンテンツ管理、その他の信頼と安全のワークフローのために提案されています。公開文献の多くは依然としてそれらをモデルとして評価しており、運用パイプラインのコンポーネントとしての動作にはあまり注目していません。これにより、レイテンシー、コスト、エスカレーション、人間によるレビュー、および敵対的リスクの制約があるライブ ワークフロー内に LLM を配置することを正当化するものは何でしょうか?という実際的な証拠の疑問が生じます。私たちは、展開証拠の不正行為優先調査を通じてこの疑問に取り組みます。当社では、不正行為の検出、調査サポート、コンテンツのモデレーション、横断的な堅牢性における LLM の使用に関する運用上関連する 49 のソースをコード化しており (不正行為 18 件、モデレーション 14 件、横断的堅牢性 17 件)、調査の境界を確立する 15 件の文脈上の参照によって補足されています。これらのソースには、運用環境の展開ではなく、システム、ベンチマーク、フレームワーク、展開関連の調査が含まれます。主な発見は証拠の不均衡です。不正行為は、コード化されたコーパスの最大のタスク固有部分を提供します。ただし、モデレーション文書には、遅延、コスト、ガバナンス、公平性に関するより明確な公開証拠が含まれています。 18 の不正行為および調査情報源のうち、決定ごとの遅延、決定ごとのコスト、または校正証拠を明確に報告しているものはありません。ほとんどのレポートでは、代わりにオフライン タスクのパフォーマンス、取得の向上、またはケーススタディの精度が報告されています。この調査は、分類子、検索インターフェイス、説明ジェネレータ、レビューア アシスタント、エージェント、特徴抽出器、またはエスカレーション コンポーネントとして LLM を特定するための、役割と証拠の組織化フレーム FORTE に貢献します。また、遅延予算、意思決定あたりのコスト、意思決定のしきい値、説明の完全性、および敵対的圧力をカバーする最小限の展開証拠チェックリストにも貢献します。結果として得られる議題は、LLM ベースの不正行為と信頼と安全の作業に対する展開の主張をサポートするために必要な調査を特定します。
原文 (English)
Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows
LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows. Much of the public literature still evaluates them as models, with less attention to their behavior as components in operational pipelines. This creates a practical evidence question: what would justify placing an LLM inside a live workflow with latency, cost, escalation, human-review, and adversarial-risk constraints? We address this question through a fraud-first survey of deployment evidence. We code 49 operationally relevant sources on LLM use in fraud detection, investigation support, content moderation, and cross-cutting robustness (18 fraud, 14 moderation, 17 cross-cutting), supplemented by 15 contextual references that establish the survey boundaries. These sources include systems, benchmarks, frameworks, and deployment-relevant surveys, not 49 production deployments. The main finding is an evidence imbalance. Fraud supplies the largest task-specific portion of the coded corpus. The moderation papers, however, include more explicit public evidence on latency, cost, governance, and fairness. Among the 18 fraud and investigation sources, none report clean per-decision latency, per-decision dollar cost, or calibration evidence; most report offline task performance, retrieval gains, or case-study accuracy instead. The survey contributes a role-and-evidence organizing frame, FORTE, for locating LLMs as classifiers, retrieval interfaces, explanation generators, reviewer assistants, agents, feature extractors, or escalation components. It also contributes a minimum deployment-evidence checklist covering latency budget, cost per decision, decision threshold, explanation integrity, and adversarial pressure. The resulting agenda identifies studies needed to support deployment claims for LLM-based fraud and trust-and-safety work.
エンタープライズ コーディング エージェントの推論経済学: クラウドとオンプレミス LLM のケーススタディ
自律型コーディング エージェントにより、エンジニアリング組織は、API ベースのフロンティア モデル (トークン コストが高くても強力な推論) と、推論の忠実度がある程度失われるものの、低限界コストのスケーリングとデータ主権を約束するオンプレミスの量子化オープンウェイト モデルのどちらかを選択する必要があります。私たちは、本番モノリポジトリで、単一開発者の非ランダム化縦断ケース スタディを 2 つの連続する 28 日間にわたって実施し、このトレードオフを調査しました。これは、Claude Code を使用した API ベースの Claude Opus 4.7/4.8 構成と、NVFP4 に量子化された Opencode を使用したオンプレミスの GLM-5.1/5.2 構成です。NVIDIA Blackwell ハードウェア上で行われます。 LLM テレメトリと Git 履歴を分析すると、プロンプト キャッシュ (99.3% のヒット率) により、実現 API コストが 88.6% 削減され、100 万トークンあたり実質 0.57 ドルになります。これは共有オンプレミス スライスの償却単価 2.83 ドルよりもさらに低いことがわかります (使用量に依存する反転。総実現費用と総所有コスト (TCO) が確実な量です)。同程度の総コード チャーンでは、ローカル構成の方がはるかに高い欠陥修復負担を伴いました。修正コミット率 (FCR) は 74.9% 対 45.9% で、コミットが修復される確率は、すべての難易度層内で 2.6 ~ 4.9 倍高くなります (Mantel-Haenszel OR = 3.61)。台湾市場のパラメーターと対称的な労働モデルの下では、オンプレミス展開は共有 GPU 割り当ての下で実際の TCO を 40.1% 節約しますが、専用予約のコストはキャッシュされた API よりも 43.8% 高くなります。共有割り当ての下では、真のペナルティは金銭的なものではなく、測定可能な開発者エクスペリエンスの負担です。タイムスタンプ インジケーターは、デバッグ スパイラルに閉じ込められた作業量とコミット リズムの低下を示しています。また、オフライン リプレイでは、ハイブリッド ルーティング ゲートウェイが、純粋な API ベースラインを支配するのではなく、コスト品質フロンティアに沿ってインフラストラクチャの節約と欠陥率を交換していることが示されています。
原文 (English)
Inference Economics of Enterprise Coding Agents: A Case Study of Cloud vs. On-Premise LLMs
Autonomous coding agents force engineering organizations to choose between API-based frontier models -- strong reasoning at high token cost -- and on-premise quantized open-weights models, which promise low-marginal-cost scaling and data sovereignty at some loss of reasoning fidelity. We study this trade-off through a single-developer, non-randomized longitudinal case study over two contiguous 28-day periods on a production monorepo: an API-based Claude Opus 4.7/4.8 configuration using Claude Code versus an on-premise GLM-5.1/5.2 configuration using Opencode, quantized to NVFP4, on NVIDIA Blackwell hardware. Analyzing LLM telemetry and Git history, we find that prompt caching (99.3% hit rate) cuts realized API cost by 88.6% to an effective \$0.57 per million tokens -- below even the \$2.83 amortized unit cost of the shared on-premise slice (a utilization-dependent inversion; total realized spend and total cost of ownership (TCO) are the robust quantities). At comparable gross code churn, the local configuration was associated with a far higher defect-repair burden: a Fix Commit Ratio (FCR) of 74.9% versus 45.9%, with the odds of a commit being a repair 2.6 to 4.9 times higher within every difficulty tier (Mantel-Haenszel OR = 3.61). Under Taiwan-market parameters and a symmetric labor model, on-premise deployment nonetheless saves 40.1% of true TCO under shared GPU allocation, whereas dedicated reservation costs 43.8% more than the cached API. Under shared allocation, the genuine penalty is not monetary but a measurable developer-experience burden -- timestamp indicators show more work trapped in debugging spirals and a slower commit cadence -- and an offline replay shows hybrid routing gateways trade defect rate for infrastructure savings along a cost-quality frontier rather than dominate the pure-API baseline.
SingGuard-NSFA: 生成推論とリアルタイム分類によるエージェントティック AI の拡張可能なガードレール
nsfaguard は、プロンプト インジェクション、機密情報の抽出、悪意のあるコード リクエスト、危険なツールの誤用、リソースの枯渇などの運用上の脅威からエージェント AI システムを保護するためのガードレール フレームワークです。まず NSFA 分類法を紹介します。これは、185 のリスクバリアントを CIA のトライアドに基づいた階層構造に編成し、確立された 3 つの OWASP ガイドラインに照らして相互検証されます。この分類に基づいて、133 言語にわたるベンチマーク スイートを構築します。このベンチマーク スイートは、ユーザーのクエリとエージェントの応答の両方を対象とした 93,000 を超える専用サンプルと、5 つのパブリック エージェント セキュリティ データセットから適応された 3,435 のクロスソース サンプルで構成されています。これらの運用上の脅威を実際に検出するために、私たちは、解釈可能なオフライン監査のための SFT ベースの生成推論と、凍結されたバックボーン上の識別分類ヘッドを組み合わせたデュアルモード アプローチを開発し、約 50 ミリ秒でのリアルタイム検出を可能にします。当社は 0.8B、2B、4B、9B パラメーターの 4 つのモデルをリリースしており、すべて専用ベンチマークで $\geq$94% F1 を達成し、最強の競合ガードレールを 6 ~ 12 絶対ポイント上回っています。クロスソース評価では、9B モデルは、よりバランスの取れた精度と再現率のトレードオフで 91.29% の F1 を達成しています。さらに、アブレーション実験では、分類ヘッドが本来の範囲を超えたリスク検出機能をガードレールに装備し、最先端のパフォーマンスを達成できることが示されています。これらの結果は、このアプローチの拡張性と、プラグイン拡張機能としてのその汎用性を示しています。
原文 (English)
SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification
We present nsfaguard, a guardrail framework for securing agentic AI systems against operational threats, such as prompt injection, sensitive information extraction, malicious code requests, dangerous tool misuse, and resource exhaustion. We first introduce the NSFA taxonomy, which organizes 185 risk variants into a CIA-triad-grounded hierarchy and is cross-validated against three well-established OWASP guidelines. Based on this taxonomy, we construct a benchmark suite spanning 133 languages, comprising over 93K purpose-built samples targeting both user queries and agent responses, along with 3,435 cross-source samples adapted from five public agent-security datasets. To detect these operational threats in practice, we develop a dual-mode approach combining SFT-based generative reasoning for interpretable offline auditing with discriminative classification heads on the frozen backbone, enabling real-time detection at approximately 50,ms. We release four models with 0.8B, 2B, 4B, and 9B parameters, all achieving $\geq$94% F1 on purpose-built benchmarks and surpassing the strongest competing guardrails by 6 to 12 absolute points. On cross-source evaluation, the 9B model attains 91.29% F1 with a more balanced precision--recall trade-off. Moreover, ablation experiments show that classification heads can equip a guardrail with risk detection capabilities beyond its original scope and achieve state-of-the-art performance. These results demonstrate the extensibility of the approach and its generality as a plug-in enhancement.
アーキテクチャの前のベースライン: 自律的な侵入テストのためのコーディング エージェントの評価
最近の自律侵入テストの論文では、フロンティア LLM の周囲にマルチコンポーネントのセキュリティ ハーネスを追加しながら、高いベンチマーク スコアを報告しています。これらのシステムはアーキテクチャとバックボーン モデルの両方を変更することが多いため、基礎となるモデルではなくハーネスからどの程度のパフォーマンスが得られるかを判断するのは困難です。このペーパーでは、デフォルトのコーディング CLI エージェントをプレーン エージェント ベースラインとして使用した、104 タスクの XBOW ベンチマークに関する対照研究を紹介します。まず、同じ GPT-5 モデル、予算、ターゲット インターフェイス、スコアリング ルールを使用して Codex、OpenCode、Pi を実行します。このフェーズでは、最も強力な同一モデルのベースラインを特定し、セキュリティ固有のプロンプトのバリアントが観察されたスコアを改善するかどうかをテストします。次に、利用可能なモデルの最も近い一致の下で、デフォルトの Codex スキャフォールドを公開された MAPTA および PentestGPT V2 の結果と比較します。最後に、GPT-5.2 と GPT-5.5 を使用してプレーン エージェント実験を繰り返し、同じ足場内のモデルのスケーリングを測定します。結果は複雑ではありますが実際的な状況を示しています。特殊なハーネスは測定可能なベンチマーク上昇率を追加し、コスト効率を向上させる可能性がありますが、プレーンなコーディング エージェントはすでにベンチマークの大部分を解決しています。単純なエージェントを繰り返し実行すると、ユニオン カバレッジで一部の公開されているアーキテクチャ スコアと一致または超える可能性があり、新しいモデルでは同じ足場が大幅に改善されます。今後の評価では、ベンチマークの向上をアーキテクチャ設計のみに帰する前に、モデルと一致するプレーン エージェントのベースラインを報告する必要があります。
原文 (English)
Baselines Before Architecture: Evaluating Coding Agents for Autonomous Penetration Testing
Recent autonomous penetration testing papers report high benchmark scores while adding multi-component security harnesses around frontier LLMs. Because these systems often change both architecture and backbone model, it is difficult to tell how much performance comes from the harness rather than from the underlying model. This paper presents a controlled study on the 104-task XBOW benchmark using default coding CLI agents as plain-agent baselines. We first run Codex, OpenCode, and Pi with the same GPT-5 model, budget, target interface, and scoring rule. This phase identifies the strongest same-model baseline and tests whether security-specific prompt variants improve its observed score. We then compare the default Codex scaffold with published MAPTA and PentestGPT V2 results under the closest available model matches. Finally, we repeat the plain-agent experiment with GPT-5.2 and GPT-5.5 to measure model scaling inside the same scaffold. The results show a mixed but practical picture. Specialised harnesses can add measurable benchmark lift and may improve cost efficiency, but plain coding agents already solve a large share of the benchmark; repeated plain-agent runs can match or exceed some published architecture scores in union coverage, and newer models substantially improve the same scaffold. Future evaluations should report model-matched plain-agent baselines before attributing benchmark gains to architecture design alone.
蓄積された行動ルールによる自己改善 AI コーディング エージェント: クローズドループ フレームワーク
LLM ベースのコーディング エージェントは、人間のレビュー フィードバックからの修正を保持するメカニズムがないため、セッション間で同じクラスの間違いを繰り返します。私たちは、受け入れられたすべてのレビュー コメントが永続的な動作ルールとして成文化され、エージェントが自己検出できるエラー クラスのセットを徐々に拡張する閉ループ フレームワークを提示します。このフレームワークは、バージョン管理された命令ファイルに蓄積されるルール セット、コードの送信前に実行されるセルフレビュー チェックリスト、およびルール セットの成長に応じた整合性を保証する自動検証を組み合わせたものです。 35 を超えるサービス マイクロサービス プラットフォームにわたる展開において、ルール セットは 5 から 18 の動作ルール、15 を超える言語固有の標準、および 15 項目の自己レビュー チェックリストに増加しました。これらはすべて実際のレビュー フィードバックから派生したものです。コード生成、PR レビュー、インシデント調査、クロスサービス リファクタリングにわたる 11 の記録された作業セッションからの実証結果を紹介します。蓄積されたルールにより、レビュー作業が低レベルの正確性から設計レベルの検証に移行し、ルール違反クラスの再発率が測定値 0% に達し、異種エージェント インターフェイス間で転送されることが観察されています。私たちのアプローチを、経験的 LLM 学習 (Reflexion、ExpeL、Voyager) および自動コード レビュー (CodeReviewer、SWE ベンチ エージェント) の関連作業と比較します。これにより、私たちのフレームワークが重み更新なしで永続的なクロスセッション学習を実現し、合成ベンチマークではなく本番コードベースで動作し、既存のベンチマークでは測定できない直交次元 (経時的な動作の一貫性) に対処していることがわかります。その結果、コーディング エージェントはレビュー サイクルごとに改善され、モデルの重みを 1 つも変更することなく人間の協力者のエンジニアリングの知恵が蓄積されます。
原文 (English)
Self-Improving AI Coding Agents Through Accumulated Behavioral Rules: A Closed-Loop Framework
LLM-based coding agents repeat the same classes of mistakes across sessions because they lack a mechanism to retain corrections from human review feedback. We present a closed-loop framework in which every accepted review comment is codified as a persistent behavioral rule, progressively expanding the set of error classes the agent can self-detect. The framework combines an accumulating rule set in a version-controlled instruction file, a self-review checklist executed before code submission, and automated validation that ensures rule set integrity as it grows. In deployment across a 35+ service microservices platform, the rule set grew from 5 to 18 behavioral rules, 15+ language-specific standards, and a 15-item self-review checklist, all derived from real review feedback. We present empirical results from 11 recorded working sessions spanning code generation, PR review, incident investigation, and cross service refactoring. We observe that accumulated rules shift review effort from low-level correctness toward design-level validation, achieve a measured 0% recurrence rate for ruled-against error classes, and transfer across heterogeneous agent interfaces. We compare our approach against related work in experiential LLM learning (Reflexion, ExpeL, Voyager) and automated code review (CodeReviewer, SWE-bench agents), showing that our framework achieves persistent cross-session learning without weight updates, operates on production codebases rather than synthetic benchmarks, and addresses an orthogonal dimension (behavioral consistency over time) that existing benchmarks do not measure. The result is a coding agent that improves with every review cycle, accumulating the engineering wisdom of its human collaborators without changing a single model weight.
大規模言語モデル向けの効率的でプライバシーを意識したエッジ クラウド協調推論
オンデバイス LLM 推論は、応答遅延、限られたハードウェア リソース、ユーザー プライバシーというトリレンマに直面しています。完全なクラウド推論は強力なコンピューティング能力を提供しますが、ユーザー プロンプトや対話データが公開されます。一方、スタンドアロンのオンデバイス推論は、ほとんどのコンシューマ デバイスや組み込みエッジ デバイスでは実現できません。このペーパーでは、エンドポイント認証された KV キャッシュに基づいて構築されたプライバシー中心のエッジとクラウドの協調 LLM 推論フレームワークについて説明します。ローカル エンドポイントは入力前処理、埋め込み計算、適応特徴最適化、KV キャッシュ認証、投機的デコード、低次元モデル ヘッド計算を処理し、クラウドは認証済みデコーダー推論、KV キャッシュ管理、トークン検証、高次元語彙投影を実行します。エンドポイントは部分的な出力を融合し、言語に適応したマスキングを適用し、ターゲット トークンをサンプルします。すべての送信データと切り捨てられたロジットは量子化され、プライバシーを確保するために AES-GCM 暗号化され、コア軽量モジュール、ドラフト パラメーター、キャッシュ アクセス ポリシーは漏洩を避けるためにローカルに保持されます。このフレームワークは、最適化されたストリーミング、バッチ処理、量子化された ONNX 導入を通じて、CPU のみ、GPU を搭載したデバイス、組み込みデバイスなどの異種デバイスをサポートします。評価の結果、このフレームワークは、ベースラインの分割推論と比較して、トークンごとのレイテンシを最大 46.1\% 削減し、ダウンリンク ペイロードを最大 67.4\% 削減し、完全なクラウド推論と同等のパフォーマンスを維持していることが実証されています。
原文 (English)
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models
On-device LLM inference faces a trilemma of response latency, limited hardware resources and user privacy. Full cloud inference delivers strong computing power but exposes user prompts and dialogue data, while standalone on-device inference is unfeasible for most consumer and embedded edge devices. This paper presents a privacy-centric edge-cloud collaborative LLM inference framework built on endpoint-authenticated KV cache. Local endpoints handle input preprocessing, embedding computation, adaptive feature optimization, KV cache authentication, speculative decoding and low-dimensional model head calculation, while the cloud conducts authenticated decoder inference, KV cache management, token verification and high-dimensional vocabulary projection. Endpoints fuse partial outputs, apply language-adaptive masking and sample target tokens. All transmitted data and truncated logits are quantized and AES-GCM encrypted for privacy, with core lightweight modules, draft parameters and cache access policies kept local to avoid leakage. The framework supports heterogeneous devices including CPU-only, GPU-equipped and embedded devices via optimized streaming, batching and quantized ONNX deployment. Evaluations demonstrate that the framework reduces per-token latency by up to 46.1\% and downlink payloads by up to 67.4\% over baseline split inference, retaining comparable performance to full cloud inference.
AI を使用してカリキュラム パターンの複雑さを分析し、定時卒業率を向上
人工知能 (AI) の台頭により、大量のデータの自動分析が可能になりました。これまで時間と労力がかかっていたタスクは、AI を使用することではるかに効率的に完了できるようになります。この研究では、AI 技術を使用してソフトウェア エンジニアリングの学士課程のカリキュラム パターンを分析および修正します。カリキュラムには長いシーケンスが含まれることが多く、そのシーケンス内のクラスに合格しないと 4 年以内に学位を取得することが危うくなる可能性があります。大学教員による手動によるカリキュラムの分析と改訂は、時間と労力がかかるプロセスであり、変更がほとんど行われず、学生のニーズの変化に対応できなくなります。この取り組みでは、大規模言語モデル (LLM) を使用してカリキュラム パターンを分析し、修正を提案することで、カリキュラムの変更までの時間を短縮し、ボトルネックと卒業の遅れを軽減します。
原文 (English)
Analyzing Curricular Pattern Complexity Using AI to Improve On-Time Graduation Rates
The rise of Artificial Intelligence (AI) enables automatic analysis of large amounts of data. Previously time-consuming and labor-intensive tasks can be completed much more efficiently with the use of AI. This work uses AI techniques to analyze and revise curricular patterns in an undergraduate degree for Software Engineering. Curricula often have long sequences where failure to pass a class within the sequence may jeopardize completion of the degree within four years. Manual analysis and revision of curricula by university faculty is a lengthy and labor-intensive process, causing changes to occur rarely and making it impossible to keep up with the changing needs of students. This work reduces the time-to-change for curricula and reduces bottlenecks and graduation delays by using Large Language Models (LLMs) to analyze curricular patterns and suggest revisions.
MiMo-V2.5 シリーズのフルパイプライン推論の最適化: ハイブリッド SWA の効率を限界まで押し上げる
ハイブリッド スライディング ウィンドウ アテンション (ハイブリッド SWA)、スパース混合専門家 (MoE)、およびマルチモーダル エンコーダーを組み合わせた、MiMo-V2.5 モデル ファミリのフルパイプライン推論最適化を紹介します。ハイブリッド SWA は理想的には、フル アテンションに比べてアテンション コンピューティングと KVCache ストレージの両方を大幅に削減できますが、運用環境でこれらの向上を実現するには、多大なエンジニアリング作業が必要です。当社は、レイヤーごとのプリフェッチ、SWA 対応プレフィックス キャッシュ ツリー、特殊な配置戦略を使用して KVCache システムを体系的に最適化し、厳格な $O(W)$ SWA ストレージと高いキャッシュ ヒット率を実現します。さらに、RDMA に最適化されたネットワークを備えた高性能分散キャッシュ インフラストラクチャである GCache を構築し、負荷分散を維持しながら計算量を削減する KVCache アフィニティ ルーターを開発します。また、GPU 画像の前処理、並列ビデオ デコード、マルチモーダル キャッシュ共有などのマルチモーダル入力に対しても最適化します。これらの最適化により、ハイブリッド SWA + MoE + マルチモーダル複合アーキテクチャを効率的にカバーする、運用環境における初の大規模 LLM サービス システムが構成されます。
原文 (English)
Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit
We present a full-pipeline inference optimization for the MiMo-V2.5 model family, which combines Hybrid Sliding Window Attention (Hybrid SWA), sparse Mixture-of-Experts (MoE), and multimodal encoders. While Hybrid SWA can ideally reduce both attention compute and KVCache storage significantly compared to Full Attention, realizing these gains in production requires substantial engineering effort. We systematically optimize the KVCache system with layerwise prefetch, SWA-aware prefix cache trees, and specialized placement strategies, achieving strict $O(W)$ SWA storage and high cache hit rates. We further build GCache, a high-performance distributed cache infrastructure with RDMA-optimized networking, and develop a KVCache-affinity router to reduce computation while preserving load balancing. We also optimize for multimodal inputs, including GPU image preprocessing, parallel video decoding, and multimodal cache sharing. Together, these optimizations constitute the first large-scale LLM serving system in production that efficiently covers the Hybrid SWA + MoE + multimodal composite architecture.
WaterMoE: エキスパート ルーティング ベースの透かしによる高忠実度および効率の向上
大規模言語モデル (LLM) は目覚ましい成功を収めていますが、コンテンツの出所や悪用についての懸念が高まっており、信頼性の高い透かし技術の必要性が高まっています。ただし、これらの手法は、主に 2 つの理由により、実際にはほとんど採用されていません: i) モデルのパフォーマンスが大幅に低下すること、および ii) 推論のオーバーヘッドが追加されることです。この問題を確認するために、さまざまな生成タスクにわたる包括的なベンチマークを構築し、9 つの代表的な透かし手法を系統的に評価します。ほとんどすべての既存のメソッドはテキストの流暢さのために設計されていますが、制限された複雑なタスクには設計されておらず、そのオーバーヘッドにより遅延が重要なシステムへの導入が妨げられていることがわかりました。 i) と ii) に対処するために、人気が高まっている Mixture-of-Experts (MoE) LLM 用の LLM 透かしスキーム \textit{WaterMoE} を提案します。 WaterMoE は、制御された摂動を通じて電子透かし信号を各ルーターのエキスパート選択に埋め込み、最終出力でのトークン選択シフトに蓄積されます。後処理トークン サンプリング アプローチとしてのウォーターマークとは対照的に、WaterMoE は推論ループ内にウォーターマークを埋め込みますが、品質の低下と計算オーバーヘッドは無視できます。広範な実験により、私たちの方法は透かしなしの状態に近い忠実度のパフォーマンスを達成し、ベンチマークで最先端の透かし入れ方法を常に上回っており、ネイティブ生成と比較して推論遅延がわずか 1\% 増加するだけで、最大 $4\time$ の高速化が可能であることが実証されています。結果は、現実世界のタスクに導入できる WaterMoE の機能を示しています。
原文 (English)
WaterMoE: Expert-Routing-based Watermarking for High Fidelity and Efficiency
Large language models (LLMs) have achieved remarkable success but raise growing concerns about content provenance and misuse, motivating the need for reliable watermarking techniques. However, these techniques have rarely been adopted in practice mainly for two reasons: i) severely degraded model performance, and ii) additional inference overhead. To confirm the problem, we construct a comprehensive benchmark spanning different generation tasks to systematically evaluate 9 representative watermarking methods. We found almost all existing methods are designed for text fluency, but not for restricted and complicated tasks, and their overhead prevents them from deployment in latency-critical systems. To address i) and ii), we propose an LLM watermarking scheme \textit{WaterMoE} for the growingly popular Mixture-of-Experts (MoE) LLMs. WaterMoE embeds watermarking signals through controlled perturbation into the expert selection at each router, which accumulates to token selection shift at the final output. In contrast to watermarking as a post-processing token-sampling approach, WaterMoE embeds watermark within the inference loop incurring negligible quality degradation and computational overhead. Extensive experiments demonstrate that our method achieves a fidelity performance close to the unwatermarked and consistently outperforms state-of-the-art watermarking methods on the benchmark, with up to $4\times$ speedup, incurring merely 1\% additional inference latency compared to native generation. The results demonstrate the capability of WaterMoE to be deployed in real-world tasks.
TSSM: 時間変数履歴モデリングによるグローバル ステーション天気予報のための 3 軸状態空間モデル
Global Station Weather Forecasting (GSWF) は、主要地域の局地的および異常気象を予測するために極めて重要です。ルックバックウィンドウを活用する努力にもかかわらず、既存の方法では精度の向上が限られており、極端なイベントやエラーの蓄積に苦労しています。これらの制限は、特に部分的な観測の下で、混沌とした気象のダイナミクスを捉えるには不十分な短期パターンに過度に依存していることに起因しています。この問題に対処するために、我々は、歴史強化されたTemporal-VariableHistoricalパラダイムを備えた新しいTriaxis State Space Model(TSSM)を提案します。これは、時間的なルックバックウィンドウを超えた長期的で大規模な周期的かつフルウィンドウの気象パターンを補償するために、期間に合わせた過去の気象データを組み込んでいます。具体的には、TSSM は過去のサンプルを期間に合わせたバッチにスタックし、予測は過去および現在の観測によって因果的にサポートされます。時間的、変数的、および歴史的スキャンは、軸方向の時間依存性、変数の相関関係、および歴史的進化を捕捉するように設計されています。この構造は階層的に共有され、季節的なイベントから極端なイベントまでをモデル化し、歴史的なパターン間の不整合を軽減します。 TSSM は、これまでで最大の測候所気象データセットである Weather-5K で SOTA パフォーマンスを達成し、精度と異常気象メトリクスが 10% および 61% 向上し、人間が関与するデータセットでは 95% の最高または 2 番目に良い結果を獲得しました。その利点は長期予測と反復予測でより顕著になり、240 時間で 37.5% の利益に達し、48 時間×5 反復設定では最大 103.5% の利益に達します。さらに、TSSM は、ベースラインの 43% 未満と比較して、最大 80% の欠落観測の下でも 90% 以上のパフォーマンスを維持しており、地球規模の現場観測ネットワークにおける信頼性の高い GSWF の堅牢性と実用的な可能性を実証しています。
原文 (English)
TSSM: Triaxial State Space Model for Global Station Weather Forecasting with Temporal-Variable-Historical Modeling
Global Station Weather Forecasting (GSWF) is pivotal for localized and extreme weather prediction over key regions. Despite efforts to exploit look-back windows, existing methods show limited accuracy gains and struggle with extreme events and error accumulation. These limitations stem from overreliance on short-term patterns, which are insufficient to capture chaotic weather dynamics, especially under partial observations. To address this problem, we propose a novel Triaxial State Space Model (TSSM) with a history-enhanced Temporal-VariableHistorical paradigm, which incorporates period-aligned historical weather data to compensate for long-term, large-scale periodic, and full-window weather patterns beyond the temporal lookback window. Specifically, TSSM stacks historical samples into period-aligned batches, where forecasting is causally supported by historical and current observations. Temporal, variable, and historical scanning are designed to capture axial temporal dependencies, variable correlations, and historical evolution. This structure is hierarchically shared to model seasonal to extreme events while alleviating misalignment across historical patterns. TSSM achieves SOTA performance on Weather-5K, the largest station weather dataset to date, with 10% and 61% gains in accuracy and extreme event metrics, and obtains 95% best or second-best results on human-involved datasets. Its advantages are more pronounced in long-horizon and iterative forecasting, reaching a 37.5% gain at 240h and up to 103.5% under a 48h times 5 iterative setting. Moreover, TSSM retains > 90% performance under up to 80% missing observations, compared with < 43% for baselines, demonstrating robustness and practical potential for reliable GSWF in global in-situ observation networks.
知識追跡のための能力および熟練度モデリングによる知識状態のもつれの解消
知識トレース (KT) は、歴史的な相互作用から進化する知識状態をモデル化することで、生徒の将来の成績を予測することを目的としています。既存の KT 手法は通常、生のインタラクション シーケンスを統合された行動プロセスとして扱い、学習行動のフェーズ固有の性質を見落としています。私たちの予備的な観察では、学生は十分な練習をした後に、以前に失敗した知識概念に正しく答える可能性が高くなったことが示されており、能力構築から熟練度指向の学習への移行が示唆されています。これを動機として、私たちは、カスタマイズされた分解メカニズムに基づいて学生の対話を能力段階と熟練度段階に分解する KT フレームワークである Phase-Aware Knowledge Tracing (PAKT) を提案します。分解されたシーケンスを効果的に活用するために、フェーズ固有の全体的な知識状態を共同でキャプチャするタイプ認識読み出しモジュールを備えたマルチブランチ トランスフォーマーを設計します。さらに、位相依存型 KT モデルにおける複雑な学習動作の絡み合いによって引き起こされる交絡バイアスを明らかにするための因果分析を提供します。 6 つの公開ベンチマークでの広範な実験により、私たちの手法が代表的なベースラインを常に上回っており、最大 AUC ゲインは 1.33%、平均ゲインは 0.82% であることが実証されました。
原文 (English)
Disentangling Knowledge States with Ability and Proficiency Modeling for Knowledge Tracing
Knowledge tracing (KT) aims to predict students' future performance by modeling their evolving knowledge states from historical interactions. Existing KT methods usually treat the raw interaction sequence as a unified behavioral process, overlooking the phase-specific nature of learning behaviors. Our preliminary observations show that students are more likely to correctly answer previously failed knowledge concepts after sufficient practice, suggesting a transition from ability-building to proficiency-oriented learning. Motivated by this, we propose Phase-Aware Knowledge Tracing (PAKT), a KT framework that decomposes student interactions into ability and proficiency phases based on the tailored decomposition mechanism. To effectively exploit the decomposed sequences, we design a multi-branch Transformer with a type-aware readout module to jointly capture phase-specific and holistic knowledge states. We further provide a causal analysis to reveal the confounding bias caused by entangling complex learning behaviors in phase-agnostic KT models. Extensive experiments on six public benchmarks demonstrate that our method consistently outperforms representative baselines, with a maximum AUC gain of 1.33% and an average gain of 0.82%.
STKAN: 時空間予測のためのコルモゴロフ・アーノルド ネットワーク
現実世界の交通データは、不均一な空間相関と非線形の時間ダイナミクスを示し、正確な時空間予測に大きな課題をもたらします。既存のアプローチは、ますます洗練されたグラフ、注意、分解のアーキテクチャを開発してきましたが、基礎となる非線形関数近似器の影響は比較的注目されていません。この研究では、Taylor 多項式 Kolmogorov-Arnold Network モジュールを空間的および時間的トークン混合に導入する時空間予測アーキテクチャである STKAN を提案します。 STKAN は、まず学習可能なソフト ノード グループ割り当てメカニズムを通じて高レベルの空間表現を構築し、グループごとの空間混合を適用し、続いて圧縮シーケンスに対する時間依存関係をモデル化します。空間的および時間的セルフアテンション レイヤーをさらに使用して、長距離のインタラクションをキャプチャします。 5 つのトラフィック予測ベンチマークの実験では、STKAN が競争力のあるパフォーマンスを達成し、テストされた設定で評価された MLP ベースのバリアントよりも優れたパフォーマンスを発揮することが示されています。これらの結果は、非線形関数近似器の設計が時空間予測におけるアーキテクチャ設計の有用な補完として機能できることを示唆しています。
原文 (English)
STKAN: Kolmogorov-Arnold Networks for Spatio-Temporal Forecasting
Real-world traffic data exhibit heterogeneous spatial correlations and nonlinear temporal dynamics, posing substantial challenges for accurate spatio-temporal forecasting. Existing approaches have developed increasingly sophisticated graph, attention, and decomposition architectures, while the influence of the underlying nonlinear function approximator has received comparatively less attention. In this work, we propose STKAN, a spatio-temporal forecasting architecture that introduces Taylor-polynomial Kolmogorov--Arnold Network modules into spatial and temporal token mixing. STKAN first constructs high-level spatial representations through a learnable soft node-group assignment mechanism, applies group-wise spatial mixing, and subsequently models temporal dependencies over the compressed sequence. Spatial and temporal self-attention layers are further employed to capture long-range interactions. Experiments on five traffic forecasting benchmarks show that STKAN achieves competitive performance and performs better than the evaluated MLP-based variant in the tested settings. These results suggest that the design of nonlinear function approximators can serve as a useful complement to architectural design in spatio-temporal forecasting.
オーディオビジュアルナビゲーション用のハイブリッドマンバ
畳み込みニューラル ネットワークとリカレント アーキテクチャを中心としたパラダイムが 2020 年に確立されて以来、オーディオビジュアル ナビゲーションの基本的なバックボーン ネットワークには 5 年以上本質的な変更が加えられていないため、動的なマルチモーダル シーケンスの効率的な表現をサポートするには不十分です。本稿ではSamba(A Hybrid Mamba for Audio-Visual Navigation)を提案する。適応選択対応の M-SE (Mamba State Encoder) を使用して、時間的集約のために従来の GRU を置き換え、Audio Mamba Encoder (AME) を構築して、スペクトログラム内のグローバルな時間と周波数の依存関係をキャプチャする際の畳み込み演算子の制限を修正します。実験により、Samba は、聞いたことのない音源や光景に直面したときに、優れた汎化パフォーマンスを発揮することが実証されました。 Matterport3D データセットでは、既存の最先端モデルと比較してナビゲーション成功率 (SR) が 11.3\% 向上し、より微細なシーン構造を特徴とするレプリカ データセットではパフォーマンスの向上がさらに顕著です。このような近代化されたアーキテクチャの再構築により、より低い計算コストでより強力な具体化された表現機能が解放され、それによってオーディオビジュアルナビゲーションの分野におけるパラダイム進化のための非常に堅牢な技術的経路が提供されます。
原文 (English)
A Hybrid Mamba for Audio-Visual Navigation
Since the paradigm centered on convolutional neural networks and recurrent architectures was established in 2020, the fundamental backbone networks for audio-visual navigation have undergone no essential changes for more than five years, making them inadequate to support efficient representation of dynamic multimodal sequences. This paper proposes Samba(A Hybrid Mamba for Audio-Visual Navigation). It uses the adaptive selection-enabled Mamba State Encoder (M-SE) to replace conventional GRUs for temporal aggregation, and constructs an Audio Mamba Encoder (AME) to remedy the limitations of convolutional operators in capturing global time-frequency dependencies in spectrograms. Experiments demonstrate that Samba exhibits exceptional generalization performance when facing unheard sound sources and unseen scenes. On the Matterport3D dataset, it improves the navigation success rate (SR) by 11.3\% compared with existing state-of-the-art models, and the performance gain is even more pronounced on the Replica dataset, which features finer scene structures. Such modernized architectural reconstruction unlocks stronger embodied representation capabilities at a lower computational cost, thereby providing a highly robust technical pathway for paradigm evolution in the field of audio-visual navigation.
SemaDiff: 生成されたコードとテストでセマンティックを変更するコミットを特定する
セマンティックを保持するコミットと変更するコミットを区別することは、ソフトウェア リポジトリ マイニングにおける未解決の課題のままです。既存のアプローチは、リファクタリングコミットを正確に検出しますが、インターリーブによる動作変更の変更を行わずに、コミットが純粋に意味を保持していることを保証することはできません。この制限は、デバッグ、障害の位置特定、バグ データセットの構築、ロールバック分析、バグ修正のバックポートなどのいくつかのタスクに影響を与える可能性があります。このギャップを埋めるために、動作ベースの分析を通じて意味を保持するコミットを特定するための新しいアプローチである SemaDiff を提案します。コミット前とコミット後のバージョンでの同様のテスト実行の比較。リファクタリングの影響を受けるコードはテストが難しく、両方のバージョン間で異なることが多いため、テストのターゲットとして機能するコードへの追加の呼び出しメソッドを生成することを提案します。コミットが与えられると、SemaDiff は差分を分析して変更されたコードを特定し、それを呼び出す未変更の依存コードを抽出します。次に、大規模な言語モデルを使用して追加の依存クラスを生成し、両方のバージョンで変更されたコードを実行し、依存コードのテストを自動的に生成します。このようにして、異なるコード バージョンに対して同じテストが得られ、動作の違いの検出が可能になります。生成されたすべてのテストが 2 つのバージョン間で同一の結果を生成する場合にのみ、コミットはセマンティック保持として分類されます。 SemaDiff を評価するために、よく知られたオープンソース Java プロジェクトから収集した 183 コミットのデータセットを手動で構築し、注釈を付けます。得られた結果は、SemaDiff がケースの約 76% で意味変更コミットの検出と 100% の精度で意味変更コミットを正確に区別していることを示しています。
原文 (English)
SemaDiff: Identifying Semantic-Changing Commits with Generated Code and Tests
Distinguishing semantic-preserving commits from changing ones remains an open challenge in software repository mining. While existing approaches detect refactoring commits accurately, they cannot ensure that a commit is purely semantic-preserving, without any interleaving behaviour-changing modification. This limitation can impact several tasks, such as debugging, fault localisation, bug dataset construction, rollback analysis, and bug fixes backporting. To fill this gap, we propose SemaDiff, a novel approach for identifying semantic-preserving commits through behaviour-based analysis; comparison of similar test execution on pre- and post-commit versions. As code impacted by the refactoring is often hard to test and different accross both versions, we propose generating additional calling methods to that code, which serve as testing target. Given a commit, SemaDiff analyses the diff to identify modified code and extracts unchanged dependent code that calls it. It then generates an additional dependent class using a large language model to exercise the changed code in both versions, and automatically generates tests for the dependent code. This way, we obtain the same tests for the different code versions, enabling the behavioural-difference detection. The commit is classified as semantic-preserving only if all generated tests produce identical outcomes across the two versions. To evaluate SemaDiff, we construct and annotate manually a dataset of 183 commits, gathered from well-known open-source Java projects. The obtained results show that SemaDiff distinguishes accurately semantic-preserving from -- changing commits in about 76% of the cases, with a 100% precision in semantic-changing commit detection.
CoDiffGRN: BEELINE-KGC ベンチマークと共進化的離散拡散による遺伝子制御ネットワーク推論の再考
単一細胞トランスクリプトームデータから遺伝子制御ネットワーク (GRN) を推測することは生物学的発見にとって重要ですが、既存のアプローチは現実世界のニーズと根本的に一致していないという問題に悩まされています。研究者は通常、実験的検証のために、これまでに見たことのない遺伝子が関与する、信頼性の高い制御相互作用の少数のセットを求めます。ただし、現在のベンチマークはグローバル分類メトリクスを使用したトランスダクティブ分割に依存しており、一般的なモデルは帰納的設定で一般化するのに苦労しています。このギャップを埋めるために、GRN 推論を帰納的でランキング中心のグラフ補完問題として再定式化し、帰納的遺伝子ホールドアウト分割とナレッジ グラフ補完メトリクスを組み込んだ新しいベンチマーク \textbf{\benchmark} を導入して、トップランクの予測をより適切に評価します。これに基づいて、我々は \textbf{\method} を提案します。これは、生物学的に一貫した離散化された遺伝子発現状態と制御相互作用を共同でモデル化し、堅牢な帰納的一般化と改善されたトップランクの制御発見を実現する初の共進化的離散拡散フレームワークです。さらに、スケーラブルなトレーニングのために TF-ALL サブグラフ サンプリング (TASS) を導入します。 {\benchmark} に関する広範な実験により、{\method} が新たな最先端のパフォーマンスを確立し、新たな規制の発見において既存の手法を大幅に上回るパフォーマンスが得られることが示されており、アブレーション研究により私たちの設計の有効性がさらに検証されました。
原文 (English)
CoDiffGRN: Rethinking Gene Regulatory Network Inference via the BEELINE-KGC Benchmark and Co-evolutionary Discrete Diffusion
Inferring gene regulatory networks (GRNs) from single-cell transcriptomic data is crucial for biological discovery, yet existing approaches suffer from a fundamental misalignment with real-world needs. Researchers typically seek a small set of high-confidence regulatory interactions for experimental validation, often involving previously unseen genes. However, current benchmarks rely on transductive splits with global classification metrics, while prevailing models struggle to generalize under inductive settings. To bridge this gap, we reformulate GRN inference as an inductive, ranking-centric graph completion problem and introduce \textbf{\benchmark}, a new benchmark that incorporates an inductive gene-holdout split together with knowledge graph completion metrics to better evaluate top-ranked predictions. Building on this, we propose \textbf{\method}, the first co-evolutionary discrete diffusion framework that jointly models biologically coherent discretized gene expression states and regulatory interactions for robust inductive generalization and improved top-ranked regulatory discovery. We further introduce TF-ALL Subgraph Sampling (TASS) for scalable training. Extensive experiments on {\benchmark} show that {\method} establishes new state-of-the-art performance, significantly outperforming existing methods in novel regulatory discovery, and ablation studies further verify the effectiveness of our design.
サイバー心理学における AI: 被害者、攻撃者、防御者の心理を分析するために AI を使用することによるサイバーセキュリティ強化に関する体系的な文献レビュー
サイバーセキュリティは、システム、ネットワーク、データをデジタル攻撃から保護する実践です。サイバー心理学 (CPSY) は、サイバーセキュリティ アプリケーションを強化するための心理学の使用として定義されます。 2010 年代初頭以来、人工知能 (AI) の進化は CPSY とますます統合され、高度なデータ分析を活用して被害者、攻撃者、防御者の明確な性格特性や行動パターンを解読しています。この系統的文献レビュー (SLR) では、系統的レビューとメタ分析 (PRISMA) 方法論に推奨される報告項目を使用して、サイバー心理学における AI の利用 (AI-CPSY) について収集された 34 件の研究を注意深く分析します。このレビューでは、サイバー セキュリティ アプリケーション、使用された AI 手法、研究全体で使用された心理学的概念の包括的な分類が示されています。私たちは調査研究を、異常検出 (AD)、脆弱性リスク予測 (VRP)、セキュリティ意識向上トレーニング (SAT)、認証/身元検証 (AIV) の 4 つのサイバーセキュリティ アプリケーションに分類します。各アプリケーション分野内では、機械学習 (ML)、深層学習 (DL)、自然言語処理 (NLP)、強化学習 (RL) など、使用される AI 手法に従って研究がさらに分類されます。さらに、このレビューでは、最も一般的に使用されている心理学の概念を特定し、現場で使用されているデータセットを定量化し、それらの現在の実装と展開のステータスを示します。最後に、研究のギャップを検出し、未解決の課題を提示し、AI-CPSY 環境全体で使用される傾向にある最も効果的な新しい方法論を推定します。
原文 (English)
AI in Cyberpsychology: A systematic literature review of Cybersecurity enhancement by using AI for analyzing psychology of Victims, Attackers, and Defenders
Cybersecurity is the practice of protecting systems, networks, and data from digital attacks. Cyberpsychology (CPSY) is defined as the use of psychology to enhance cybersecurity applications. Since the early 2010s, the evolution of Artificial Intelligence (AI) has increasingly integrated with CPSY, leveraging advanced data analysis to decode the distinct personality traits and behavioral patterns of victims, attackers, and defenders. In this systematic literature review (SLR), we carefully analyze 34 collected research studies of AI usage in cyberpsychology (AI-CPSY) using the preferred reporting items for systematic reviews and meta-analyses (PRISMA) methodology. The review presents a comprehensive taxonomy of the cyber-security applications, the AI methodologies used, and the psychological concepts employed across the studies . We sort the research studies into four cybersecurity applications: Anomaly Detection (AD), Vulnerability Risk Prediction (VRP), Security Awareness Training (SAT), and Authentication/Identity Verification (AIV). Within each application area, studies are further sorted according to the AI method used including machine learning (ML), deep learning (DL), natural language processing (NLP), and reinforcement learning (RL). Furthermore, the review identifies the most commonly utilized psychological concepts, quantify the datasets used in the field, and present their current implementation and deployment status. At last, it detect research gaps, present open challenges, and deduce the trending and most effective and emerging methodologies used across the AI-CPSY landscape.
ShortOPD: ショートからロングのオンポリシー蒸留によるプルーニングされた LLM の回復
構造化プルーニングは、LLM を圧縮するためのハードウェアに優しい方法ですが、主に多肢選択認識タスクで検証されますが、同じ圧縮されたチェックポイントは、実際に展開に必要な自由形式の生成では崩壊する可能性があります。 2 つの観察により、このギャップが追跡されます。まず、貪欲な \textsc{pass}@$1$ は圧縮後にほぼ消滅しますが、\textsc{pass}@$k$ はサンプリングを繰り返すと大幅に回復します。有用な世代は削除されず、降格されます。第二に、回復可能な体制は主に接尾辞の繰り返しによって失敗します。したがって、リカバリでは、圧縮前のモデルを凍結教師として再利用することによって、オンポリシー蒸留 (OPD) が提供する高密度のトークンレベルの監視を使用して、圧縮モデル自体のオンポリシー状態でトレーニングする必要があります。ただし、ポリシーに基づいた展開が長期間続くと、情報量の少ない反復的なサフィックスに早期回復予算が費やされ、損失の減少が遅れます。この無駄を軽減するために、教師によって確認された反復サフィックスを検出し、残っているプレフィックスを各ロールアウトの有効長として扱い、ポリシーが現在使用できる有効長に将来のロールアウト予算を割り当てる、短から長の OPD スケジュールである \textbf{\shortopd} を提案します。数学、コード、およびオープンエンド生成全体にわたって、\shortopd\ は、圧縮モデルのスコアを未回復値の約 $9\times$ および標準回復レシピ (KD、KD、SeqKD なしの SFT) の $1.6$ ~ $4.4\times$ に引き上げ、トレーニング時間の 4 分の 1 を使用して 2 ポイント以内の固定 $8192$ トークンのロールアウト範囲と一致します ($8.5$ 対 \ $35.9$ 時間)、ロールアウト トークンが $71\%$ 減少します。このレシピが、構造化プルーニングを複雑性と複数選択ベンチマークの限界を超えて進め、展開可能な生成品質に一歩近づくのに役立つことを願っています。
原文 (English)
ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation
Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed checkpoints can collapse on the free-form generation that deployment actually requires. Two observations trace this gap. First, greedy \textsc{pass}@$1$ nearly vanishes after compression, yet \textsc{pass}@$k$ recovers substantially under repeated sampling: useful generations are demoted, not erased. Second, the recoverable regime fails mainly through suffix repetition. Recovery should therefore train on the compressed model's own on-policy states with dense token-level supervision, which On-Policy Distillation (OPD) provides by reusing the pre-compression model as a frozen teacher. However, long on-policy rollouts spend early recovery budget on low-information repetitive suffixes, delaying loss descent. To mitigate this waste, we propose \textbf{\shortopd}, a short-to-long OPD schedule that detects teacher-confirmed repetitive suffixes, treats the surviving prefix as each rollout's effective length, and allocates future rollout budgets to the effective lengths the policy can currently use. Across math, code, and open-ended generation, \shortopd\ raises the compressed model's score to about $9\times$ its unrecovered value and $1.6$--$4.4\times$ standard recovery recipes (SFT w/o KD, KD, and SeqKD), and it matches a fixed $8192$-token rollout horizon within two points using a quarter of the training time ($8.5$ vs.\ $35.9$ hours) and $71\%$ fewer rollout tokens. We hope this recipe helps move structured pruning beyond marginal gains on perplexity and multiple-choice benchmarks, a step closer to deployment-ready generation quality.
Boogu-Image-0.1: オープンソースの統合されたマルチモーダルの理解と生成を促進する
Boogu-Image-0.1 は、Base、Turbo、Edit、Edit-Turbo の各バリアントで構成される、オープンソースの統合マルチモーダル理解および生成モデル ファミリです。高品質のテキストから画像への生成、高速推論、命令ベースの編集、および二か国語 (中国語と英語) のテキスト レンダリングにおいて、優れたパフォーマンスを提供します。 Nano-Banana-Pro や GPT-Image-2 のようなクローズドソースのマルチモーダル システムは、単一モデルではなくシステム レベルの統合を通じて強力なパフォーマンスを実現しますが、その内部慣行はほとんど公開されていません。この研究では、モデルの理解、データ品質、トレーニング パイプラインの目標を絞った改善と、エージェントによる推論時間のスケーリングを組み合わせることで、非常に制約されたコンピューティング予算の下でも生成と編集のパフォーマンスを大幅に向上できることを実証します。包括的な評価では、Boogu-Image-0.1 が標準ベンチマーク全体で他のオープンソース モデルと常に同等またはそれを上回り、主要なクローズドソース システムに迫る結果を達成していることが示されています。注目すべきことに、これはわずか 2 億 862 万個の一意の画像で実現されています。基本モデルの理論上のトレーニング コストはわずか約 $400,000 です。私たちは、より広範な研究コミュニティにとって価値があると信じている実践的な議論を共有し、統合されたマルチモーダルな理解と生成のためのオープンエコシステムを前進させるために、Apache 2.0 での重み、コード、レシピをリリースします。私たちのコードは、https://github.com/Boogu-Project/Boogu-Image から入手できます。
原文 (English)
Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation
We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual (Chinese-English) text rendering. Closed-source multimodal systems like Nano-Banana-Pro and GPT-Image-2 achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. In this work, we demonstrate that targeted improvements in model understanding, data quality, and training pipelines, coupled with agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets. Comprehensive evaluations show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks, and achieves results approaching leading closed-source systems. Notably, this is accomplished with only 208.62 million unique images. The base model's theoretical training cost is only approximately \$400K. We share practical discussions that we believe are valuable to the broader research community, and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. Our code is available here: https://github.com/Boogu-Project/Boogu-Image.
アクティブ対角線を超えた RIS を活用したヘテロジニアス エッジ コンピューティング: 分布強化学習アプローチ
アクティブな対角線を越えた再構成可能なインテリジェント サーフェス (BD-RIS) により、ハイブリッド送信モードと反射モードで効果的な信号増幅と全空間カバレッジを実現できるため、異種モバイル エッジ コンピューティング (MEC) システムにおける閉塞を認識したアップリンク オフロードの有望なソリューションが提供されます。ただし、実際のハイブリッド モード アクティブ BD-RIS は相互デバイスによって実現されており、本質的にセクター間のエネルギー漏洩が発生し、システム レベルのエネルギーと遅延のトレードオフを再形成することになります。この論文では、オフロードの決定、CPU/GPU の計算割り当て、送信電力、受信処理、およびアクティブ BD-RIS が緊密に結合されている、相互アクティブ BD-RIS 支援異種 MEC のための、エネルギーを意識したオフロードとリソース割り当てについて研究します。結果として生じる問題は、高次元混合整数の非凸問題であり、従来のインスタンスごとの最適化では効率的に解決することが困難です。この課題に対処するために、DSAC-T という名前の分散ソフト アクター - クリティカル アルゴリズムの改良版に基づいたエンドツーエンドの共同最適化フレームワークを開発しました。 DSAC-T は、期待値だけではなく収益分布をモデル化することにより、報酬の不均一性と実現可能性境界の感度の下での政策の安定性を向上させます。他のベースライン アルゴリズムと比較して、DSAC-T は最高のエネルギー レイテンシ報酬、81.67% の最高の実現可能性比、およびシナリオあたり 0.0267 秒の高速オンライン意思決定時間を達成します。
原文 (English)
Active Beyond-Diagonal RIS Empowered Heterogeneous Edge Computing: A Distributional Reinforcement Learning Approach
Active beyond-diagonal reconfigurable intelligent surfaces (BD-RISs) enables hybrid transmitting and reflecting mode to achieve effective signal amplification and full-space coverage, thus providing a promising solution for blockage-aware uplink offloading in heterogeneous mobile edge computing (MEC) systems. However, practical hybrid mode active BD-RIS are realized by reciprocal devices, which inherently generate cross-sector energy leakage that will reshape the system-level energy-latency tradeoff. This paper studies energy-aware offloading and resource allocation for reciprocal active BD-RIS-assisted heterogeneous MEC, where offloading decisions, CPU/GPU computation allocation, transmit powers, receive processing, and active BD-RIS are tightly coupled. The resulting problem is a high-dimensional mixed integer nonconvex problem and is difficult to solve efficiently by conventional per-instance optimization. To address this challenge, we develop an end-to-end joint optimization framework based on a refined version of the distributional soft actor--critic algorithm, named as DSAC-T. By modeling return distributions rather than only expected values, DSAC-T improves policy stability under reward heterogeneity and feasibility-boundary sensitivity. Compared with other baseline algorithms, DSAC-T achieves the best energy-latency reward, the highest feasibility ratio of 81.67%, and a fast online decision time of 0.0267 s per scenario.
モデルが表現、抑制、抵抗するもの: ペルソナ ベクトルを使用した Open-Weight LLM の監査
言語モデルが何を行うか、何を行わないかは、主にポストトレーニング中に設定されますが、どのような動作を表現するか、隠すか、または抵抗するかは、プロンプトだけでは明らかにされません。活性化空間における行動の方向であるペルソナ ベクトルは、この組織を調査することができますが、これまでの研究ではほんの一握りの特性のみがカバーされています。我々は、この規模でのペルソナベクトルの最初の体系的な適用を提示し、4つの行動的に異なるドメインにわたる53の形質インベントリを編集し、2つのオープンウェイトモデルのすべての形質を自然(ベースラインで発現)、制御可能な潜在的だが増幅可能、または難治性(標準的な抽出に耐性がある)としてラベル付けします。どちらのモデルも、デフォルトでは役立つタスク指向の行動になります。つまり、エージェントの 9 つの特性はすべて自然なものであり、デフォルトの臨床医の動作は、17 の特性のうち 16 つに関する認定心理学者の独立した望ましさの判断と一致します。ステアリングは、これらのデフォルトでは除外される特質、つまり誇張、幻覚、お調子者に対して最大の利益をもたらします。同じ非対称性が 171 のジェネリック特性ペアすべてに当てはまります。2 つの操作可能な特性は構成を崩壊させる可能性がありますが、デフォルトを含むペアは決して崩壊しません。標準的な抽出が「悪」のような形質で失敗した場合でも、微調整されたバリアントから転送されたベクトルによってそれが回復され、残留拒否がモデルの思考連鎖内に現れます。ペルソナ ベクトルは、コントロールのセットとしてではなく、行動の組織化のプローブとして最も有益です。
原文 (English)
What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone. Persona vectors, behavioral directions in activation space, can probe this organization, but prior work covers only a handful of traits. We present the first systematic application of persona vectors at this scale, compiling a 53-trait inventory across four behaviorally distinct domains and labeling every trait in two open-weight models as natural (expressed at baseline), steerable latent but amplifiable, or intractable (resistant to standard extraction). Both models default to helpful, task-oriented behavior: all nine agentic traits are natural, and their default clinician behavior matches a board-certified psychologist's independent desirability judgments on 16 of 17 traits. Steering produces its largest gains on traits these defaults exclude: hyperbole, hallucination, and sycophancy. The same asymmetry holds across all 171 generic-trait pairs: two steerable traits can collapse the composition, but pairs involving a default never do. Where standard extraction fails on a trait like "evil," a vector transferred from a fine-tuned variant still recovers it, with the residual refusals appearing inside the model's chain-of-thought. Persona vectors are most informative not as a set of controls but as a probe of behavioral organization.
SteinGate: スタインの不一致によるテールセンシティブな安全な強化学習
安全な強化学習は通常、予想される累積コストを制限することによって安全性を強化しますが、この基準ではまれではあるが壊滅的なテール イベントを検出できないことがよくあります。これらの制限を克服するために、この文書では、クリッピングされたコストによって引き起こされる境界アトムを考慮しながら、カーネル化されたスタインの不一致を使用した堅牢な一貫性チェックで壊れやすいテールフィッティングを置き換える、境界を意識した配布安全証明書である SteinGate を紹介します。 SteinGate は、観測されたポリシー展開コストが安全な参照分布と一致しているかどうかを評価し、ノンパラメトリックな安全性証明書を提供します。この証明書は、学習レジームを動的に適応させるために使用されます。つまり、ロールアウトが安全な基準と一致している場合は報酬向上ポリシーの更新を優先し、コストテールが逸脱した場合は回復動作に切り替えます。連続制御ベンチマークの実験では、SteinGate が最先端のベースラインと比較して競争力のある収益を維持しながら、トレーニング中の制約違反の頻度と重大度の両方を大幅に低減することが実証されました。
原文 (English)
SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy
Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrophic tail events. To overcome these limitations, this paper introduces SteinGate, a boundary-aware distributional safety certificate that replaces fragile tail fitting with a robust consistency check using Kernelized Stein Discrepancy while accounting for boundary atoms induced by clipped costs. SteinGate evaluates whether observed policy rollout costs remain consistent with a safe reference distribution, providing a non-parametric safety certificate. This certificate is used to dynamically adapt the learning regime: favoring reward-improving policy updates when rollouts remain consistent with the safe reference and switching to recovery behavior when the cost tail deviates. Experiments on continuous-control benchmarks demonstrate that SteinGate significantly reduces both the frequency and severity of constraint violations during training while maintaining competitive returns relative to state-of-the-art baselines.
SemEval-2026 での RAGthoven タスク 1: マルチステージ パイプラインがベンチマークに入り、かろうじて基準をクリア
SemEval-2026 タスク 1 (MWAHAHA)、サブタスク A (英語、スペイン語、中国語による多言語の制約付きユーモア生成) 用のシステムである RAGthoven を紹介します。 RAGthoven は、計算的ユーモア理論 (良性違反理論、スクリプトベースのユーモアの意味理論) に基づいて、10 回の実験を通じて洗練された、クリエイティブなテキスト生成を多段階の大規模言語モデル (LLM) パイプライン (プランナー、ベストオブ N ライター、自己批判用のリフレクター、裁判官としての LLM 裁判官) に分解します。最終的な構成では、厳選されたジョーク コーパスからの検索拡張生成 (RAG) で Planner を拡張し、多様なジョーク メカニズムで生成をシードします。また、決定論的な ConstraintAudit チェッカーを使用して同じ 4 つのステージを公開する、ReAct スタイルの逐次ツール呼び出し (Exp09) と自律的なマルチブランチ オーケストレーション (Exp10) という 2 つのエージェント バリアントも評価します。保留された 12 インスタンスの英語サンプルの 4 つのフロンティア モデル全体で、どちらのエージェント バリアントも、大幅に高いツールコール バジェットにもかかわらず、非エージェント パイプラインよりも優れていると判断された出力を生成しませんでした。 RAGthoven は、3 つの言語すべてで Gemini 2.5 Flash ベースラインとランク 1 を共有しており、主催者が報告した信頼区間は重複しています。スペイン語では、ベースラインより生の Elo ポイントが 42 ポイント (1182 対 1140) リードしていますが、英語 (1045 対 1081) と中国語 (1045 対 1053) では、同じ統計上の同点内でベースラインのほうが生の評価が高くなります。これらの結果を総合すると、強力なフロンティア モデルがループに入ると、精巧な多段階プロンプト エンジニアリングとエージェントによる足場から得られる利益が、言語に依存して減少することが示唆されます。
原文 (English)
RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar
We present RAGthoven, our system for SemEval-2026 Task 1 (MWAHAHA), Subtask A (multilingual constrained humor generation in English, Spanish, and Chinese). RAGthoven decomposes creative text generation into a multi-stage large language model (LLM) pipeline (Planner, Best-of-N Writer, Reflector for self-critique, LLM-as-a-judge Judge) grounded in computational humor theory (Benign Violation Theory, Script-based Semantic Theory of Humor) and refined across ten experiments. In our final configuration, we augment the Planner with retrieval-augmented generation (RAG) from a curated joke corpus, seeding generation with diverse joke mechanisms. We also evaluate two agentic variants -- ReAct-style sequential tool-calling (Exp09) and autonomous multi-branch orchestration (Exp10) -- that expose the same four stages with a deterministic ConstraintAudit checker. Across four frontier models on a held-out 12-instance English sample, neither agentic variant produced outputs we judged superior to the non-agentic pipeline despite substantially higher tool-call budgets. RAGthoven shares Rank 1 with the Gemini 2.5 Flash baseline in all three languages, with overlapping organizer-reported confidence intervals. In Spanish, it leads the baseline by 42 raw Elo points (1182 vs. 1140), while in English (1045 vs. 1081) and Chinese (1045 vs. 1053) the baseline holds the higher raw rating within the same statistical tie. Together, these results suggest language-dependent diminishing returns from elaborate multi-stage prompt engineering and agentic scaffolding once a strong frontier model is in the loop.
KV キャッシュの適応フィルタリング: LLM 推論における構造役割バイアスの診断と修正
アテンションベースの KV キャッシュエビクション (H2O とその子孫) は、蓄積されたアテンションの質量 (ここでは信号エネルギーとして扱われます) に基づいてトークンをランク付けし、最も重いものを維持することによって、ロングコンテキスト モデルのメモリ制約状態を圧縮します。ネストされた JSON などのスキーマ密度の高い入力ストリームでは、このスコアはノイズを不釣り合いに保持する非定常フィルターとして機能します。非コンテンツ シンクの役割 (区切り文字または空白) は、どのコンテンツの役割よりも桁違いに多くのエネルギーを運び、構造的な KEY トークンは、回答を運ぶ VALUE トークンの約 1.8 倍の割合で過剰に保持され、完全一致の精度が 88% から 88% に低下します。保持された状態の信号対雑音比が低下するため、5% バジェットでは 0% になります。反事実に基づく実験により、KEY トークンを抑制することが最良の展開可能なフィルターであることが証明されました。単一の調整されたハイパーパラメータによって制御される、SnapKV のウィンドウ スコアに対する再トレーニング不要のロール条件付き割り当ては、20% 未満の予算で H2O ギャップの 63 ~ 98% を埋め、より高い予算ではフル キャッシュの精度と適度に一致またはそれを超えます。これは、シードに依存する小さなノイズ除去効果です (B=0.50 で境界線が有意、4 つのシードにわたる B=0.30 ではゼロと区別できません)。 15 MB の線形ロール プローブは、無視できる推論コストでこれらのラベルを提供しますが、パーサー レベルのダウンストリームのマッチング精度は未解決のままです。
原文 (English)
Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference
Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking tokens on accumulated attention mass, treated here as signal energy, and keeping the heaviest. On schema-dense input streams such as nested JSON, this score acts as a non-stationary filter that disproportionately retains noise: a non-content sink role (delimiters or whitespace) carries an order of magnitude more energy than any content role, and structural KEY tokens are over-retained at roughly 1.8x the rate of the answer-carrying VALUE tokens, collapsing exact-match accuracy from 88% to 0% at a 5% budget as the signal-to-noise ratio of the retained state degrades. A counterfactual experiment establishes that suppressing KEY tokens is the best deployable filter. Our retraining-free, role-conditional allocation over SnapKV's windowed score, governed by a single tuned hyperparameter, closes 63-98% of the H2O gap at sub-20% budgets and, at higher budgets, modestly matches or exceeds full-cache accuracy -- a small, seed-sensitive denoising effect (borderline significant at B=0.50; not distinguishable from zero at B=0.30 over four seeds). A 15 MB linear role probe supplies these labels at negligible inference cost, though matching parser-level downstream accuracy remains open.
日常動作の分類には姿勢が必要、それを再構成するには動作が必要
人間は、ノイズが多く複雑な視覚入力であっても、動きを難なく認識します。しかし、刺激に含まれるどのような情報によって、人間は動きを迅速に分類できるのでしょうか?この問題に対処するために、動作分析のさまざまな戦略を体系的に比較したフレームワークはありません。ここでは、MoVi データセットからの 16 の毎日の活動のビデオを使用し、次の 3 つの戦略を比較しました。時間運動プリミティブ (TMP) は、動きを時間的に滑らかな基底関数の加重和に分解します。ルジャンドル多項式係数。関節座標の軌道を直交多項式ベースに投影します。オートエンコーダーの潜在的な埋め込み。ルジャンドル係数と TMP が最も高い分類子の精度を達成し、次にオートエンコーダが続きました。動作分類のための 2 つの識別特徴を発見しました。最も有益なのは、体の一般的な姿勢、つまりある活動を別の活動から区別する平均的な空間構成です。さらに、動きの分類を最も正確に予測できる 9 つの重要な関節を特定しました。興味深いことに、優れた分類精度が自動的に優れた動きの生成につながるわけではありません。各アクティビティの動きを再構成すると、TMP は時間的なダイナミクスを保存し、知覚的に自然な動きを生成しましたが、ルジャンドル係数からの再構成では平均的な姿勢のみが保持され、静止しているように見えました。これらの結果は、運動情報がどのように組織されているかの解離を明らかにしています。つまり、どのような活動が実行されるかを分類するには身体の静的構成で十分ですが、それがどのように展開するかを再構成するには運動の時間的ダイナミクスが必要です。この区別は、視覚システムが迅速な動作認識のためにどの特徴に依存しているかを明らかにし、姿勢の特徴によって臨床応用における効率的な動きのスクリーニングが可能になる一方で、動きの生成が目標である場合には常に動的情報が不可欠であることを示唆しています。
原文 (English)
Classifying daily activities needs posture, reconstructing them needs motion
Humans recognize movements effortlessly, even from noisy and complex visual input. But what information in the stimulus allows humans to rapidly classify movements? No framework has systematically compared different strategies of movement analysis to address this question. Here, we used videos of 16 daily activities from the MoVi dataset and compared three strategies: Temporal Movement Primitives (TMPs), which decompose movements into weighted sums of temporally smooth basis functions; Legendre polynomial coefficients, which project joint-coordinate trajectories onto an orthogonal polynomial basis; and Autoencoder latent embeddings. Legendre coefficients and TMPs achieved the highest classifier accuracy, followed by autoencoders. We found two discriminative features for movement classification. The most informative is the general posture of the body, the average spatial configuration that distinguishes one activity from another. Additionally, we identified 9 critical joints that are most predictive for movement classification. Interestingly, good classification accuracy did not automatically lead to good movement generation: when we reconstructed movements for each activity, TMPs preserved the temporal dynamics and produced perceptually natural motion, whereas reconstructions from Legendre coefficients retained only the average posture and appeared frozen. These results reveal a dissociation in how movement information is organized: the static configuration of the body suffices to classify what activity is performed, but the temporal dynamics of movement are required to reconstruct how it unfolds. This distinction clarifies which features the visual system may rely upon for rapid action recognition, and suggests that postural features could enable efficient movement screening in clinical applications, while dynamic information remain essential wherever movement generation is the goal.
展開シフト下でのリスク管理された N-1 熱緊急事態スクリーニングのための監査済みの選択的検証
エネルギー管理システムにおけるリアルタイムの N-1 緊急時スクリーニングは、保証とコストを引き換えにします。全電力潮流で信頼できるすべての停電を検証するのは遅すぎますが、一方、高速な線形感度スクリーニングでは統計的保証が得られず、特にコントローラーがシステムを不慣れな状態に誘導する場合には、危険な動作点を黙って通過する可能性があります。このペーパーでは、任意のコントローラーの出力 (最適化、モデル予測、または学習済み) に対するリスク予算付きのスクリーニングおよびトリアージ層である、監査済みの選択的検証について紹介します。安価なサロゲートは、どの停止をスキップするかを提案します。オンライン監査は、各ウィンドウで小さなランダム サンプルに対してフルパワー フローを実行します。そして、調整されたしきい値は、選択された予算と信頼度でのスキップされたセットの熱違反率の限界を、未検証の信頼できるサブセットの対応する限界とともに証明します。有効性は代理の正確さではなく実際の検証と監査に依存するため、任意の導入シフトの下でも有効です。これはリスクを予算化した画面であり、政策で信頼できるあらゆる不測の事態をチェックする必要がある場合の決定論的検証に代わるものではありません。最大 1,354 台のバスを使用する 3 つの公共伝送システムでは、実現された違反率は予算内に収まり、標準の決定論的で校正された画面はシフト中に安全でなくなり、この方法により、リアルタイム動作ポイントごとに完全な電力潮流調査が 29 ~ 75 パーセント削減されます。
原文 (English)
Audited Selective Verification for Risk-Controlled N-1 Thermal Contingency Screening under Deployment Shift
Real-time N-1 contingency screening in an energy management system trades assurance against cost: verifying every credible outage with full power flow is too slow, while fast linear-sensitivity screening gives no statistical guarantee and can silently pass unsafe operating points, especially when a controller drives the system into unfamiliar regimes. This paper introduces Audited Selective Verification, a risk-budgeted screening and triage layer for any controller's output (optimization, model-predictive, or learned). A cheap surrogate proposes which outages to skip; an online audit runs full power flow on a small random sample each window; and a calibrated threshold certifies a thermal-violation-rate bound for the skipped set at a chosen budget and confidence, with a corresponding bound for the unverified trusted subset. Validity rests on real verification and the audit rather than on surrogate accuracy, so it holds under arbitrary deployment shift. It is a risk-budgeted screen, not a replacement for deterministic verification when policy requires checking every credible contingency. On three public transmission systems up to 1354 buses, the realized violation rate stays within budget, standard deterministic and calibrated screens become unsafe under shift, and the method cuts full power-flow studies by 29 to 75 percent per real-time operating point.
進化し続けるディープフェイク検出: 動的検出システムのアーキテクチャと公開ベンチマーク評価
学術的なベンチマークでほぼ完璧なスコアを達成するディープフェイク検出器は、現実世界のコンテンツでは崩壊します。最近の実際の評価では、最先端のオープンソース モデルでは AUC が 45 ~ 50% 低下すると報告されています。私たちは、このギャップは構造的なものであると主張します。静的な検出器は、移動する生成フロンティアに対して一度トレーニングされます。私たちは、トレーニング配布を継続的に更新するオープンな敵対的コンペティションである Bittensor SN34 を通じてトレーニングされた BitMind Forensics (BMF) を紹介します。私たちは、19 の公開データセットにわたる画像、一般ビデオ、人間ビデオのチェックポイントで構成される 1 つの日付付きエクスポートを評価します。正規の顔交換スイート (FaceForensics++、Celeb-DF v1/v2/++、DFDC、DFD、UADFV、DF40)、および最近の野生および AI 生成メディア ベンチマーク (Sumsub、Deepfake-Eval-2024、WildRF、コミュニティ フォレンジック、 AIGCDetectBench、GenImage、AI-GenBench、AIGIBench、RAID、GenVidBench、GenVideo-100K)。 BMF は、Sumsub の元の画像で 0.936 AUC に達し、4 条件操作バッテリー全体 (140 万画像) でプールされた AUC 0.872 に達し、摂動下でも堅牢性を維持します (0.855 JPEG、0.799 ダウンスケール)。一方、GPEN 強化により検出が向上します (0.996)。 Deepfake-Eval-2024 では、画像では最良の商用検出器と一致し (0.915 対 0.90)、ビデオではそれを上回り (0.822 対 0.79)、最良のオープンソース検出器 (0.56 と 0.63) をはるかに上回っています。これは、21 ジェネレーターの AI 画像パネルで 0.991 AUC、GenVidBench で 0.918 に達し、DFDC (0.947 vs 0.843) および Celeb-DF v2 (0.9985 vs 0.956) で FF++ でトレーニングされたフロンティアを超えており、両方とも汚染が監査されており、Celeb-DF++ と統計的に同等です。一時的な研究では、静的ベースラインのトレーニングに参加していないジェネレーターからの保持されたメディアで、連続した日付付きエクスポートが改善されました (画像 0.842 から 0.902、ビデオ 0.864 から 0.936)。私たちの評価ハーネスは公開されており、公開時には実稼働 API が独立した検証のために正確に評価されたスナップショットを提供します。
原文 (English)
Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System
Deepfake detectors that achieve near-perfect scores on academic benchmarks collapse on real-world content: recent in-the-wild evaluations report AUC drops of 45-50% for state-of-the-art open-source models. We argue this gap is structural: static detectors are trained once against a moving generative frontier. We present BitMind Forensics (BMF), trained through Bittensor SN34, an open adversarial competition that continually refreshes the training distribution. We evaluate one dated export comprising image, general-video, and human-video checkpoints across nineteen public datasets: the canonical face-swap suites (FaceForensics++, Celeb-DF v1/v2/++, DFDC, DFD, UADFV, DF40) and recent in-the-wild and AI-generated-media benchmarks (Sumsub, Deepfake-Eval-2024, WildRF, Community Forensics, AIGCDetectBench, GenImage, AI-GenBench, AIGIBench, RAID, GenVidBench, GenVideo-100K). BMF reaches 0.936 AUC on Sumsub's original images and 0.872 pooled AUC over its full four-condition manipulation battery (1.4M images), staying robust under perturbation (0.855 JPEG, 0.799 downscaled), while GPEN enhancement improves detection (0.996). On Deepfake-Eval-2024, it matches the best commercial detector on images (0.915 vs 0.90) and exceeds it on video (0.822 vs 0.79), far above the best open-source detectors (0.56 and 0.63). It reaches 0.991 AUC on a 21-generator AI-image panel and 0.918 on GenVidBench, and exceeds the FF++-trained frontier on DFDC (0.947 vs 0.843) and Celeb-DF v2 (0.9985 vs 0.956), both contamination-audited, with statistical parity on Celeb-DF++. In a temporal study, successive dated exports improve on held-out media from generators absent from the static baseline's training (image 0.842 to 0.902; video 0.864 to 0.936). Our evaluation harness is public, and at publication the production API serves the exact evaluated snapshot for independent verification.
EMAGN: スケーラブルなトラフィック予測のための学習されたクラスタリングによる効率的なマルチアテンション グラフ ネットワーク
交通量の予測は、複雑かつ非線形の空間的および時間的依存関係があるため、非常に困難です。セルフアテンション メカニズムは、動的で長距離の依存関係をモデル化するために広く採用されており、最先端のパフォーマンスを実現していますが、二次計算とメモリの複雑さによるスケーラビリティの制限に悩まされています。これに対処するために、高速高次元ガウス フィルタリングの理論にヒントを得て、空間注意メカニズム自体を線形化する効率的なマルチアテンション グラフ ネットワーク (EMAGN) を提案します。 2 つの学習されたクラスタリング行列 C_k と C_v は、キー ベクトルと値ベクトルを M 個のスーパー クラスターに適応的にグループ化し、動的依存関係モデリングの注意の柔軟性を犠牲にすることなく、複雑さを O(N^2 d) から O(NMd) に軽減します。 PEMS-BAY と METR-LA での実験結果は、EMAGN が全注意 GMAN の 2.7 ~ 3.2% MAE 以内の精度を達成しながら、トレーニング時間を 32%、推論時間を 38%、GPU メモリを 58% 削減することを示しています。重要なことに、K=16 アテンション ヘッドでは、フル アテンション GMAN は標準 11 GB GPU 上のメモリを完全に使い果たしますが、EMAGN は動作し続けます。これは、実現可能なモデル構成がカテゴリー的に拡張されていることを示しています。 EMAGN は、トラフィック ネットワークを認識した適応クラスタリングにより、同じバックボーン内での精度と効率の両方で Linformer および Performer を上回ります。
原文 (English)
EMAGN: Efficient Multi-Attention Graph Network via Learned Clustering for Scalable Traffic Forecasting
Traffic forecasting is highly challenging due to complex and nonlinear spatial and temporal dependencies. Self-attention mechanisms have been widely adopted to model dynamic and long-range dependencies, achieving state-of-the-art performance, but suffer from limited scalability due to quadratic computational and memory complexity. To address this, we propose an Efficient Multi-Attention Graph Network (EMAGN) that linearises the spatial attention mechanism itself, inspired by the theory of fast high-dimensional Gaussian filtering. Two learned clustering matrices C_k and C_v adaptively group key and value vectors into M super-clusters, reducing complexity from O(N^2 d) to O(NMd) without sacrificing the flexibility of attention for dynamic dependency modelling. Experimental results on PEMS-BAY and METR-LA show that EMAGN achieves accuracy within 2.7-3.2% MAE of full-attention GMAN while reducing training time by 32%, inference time by 38%, and GPU memory by 58%. Critically, at K=16 attention heads, full-attention GMAN runs out of memory on a standard 11 GB GPU entirely while EMAGN continues to operate, demonstrating a categorical expansion of feasible model configurations. EMAGN also surpasses Linformer and Performer in both accuracy and efficiency within the same backbone, owing to its traffic-network-aware adaptive clustering.
行列分解のためのミュオンの再評価
Muon は最近、大規模な深層学習の強力なオプティマイザーとして登場し、近似直交化を通じて勾配更新を再構築し、大規模言語モデルのトレーニングにおいて Adam および AdamW を上回るパフォーマンスを発揮すると報告されています。その経験的な成功は、ミュオンをスペクトル標準の下での最急降下として解釈する一連の理論的研究の増加を動機付けてきました。しかし、Muon の利点のどれがその更新ルール自体に由来するもので、どれが現代のディープ ネットワークの規模、アーキテクチャ、データの成果物であるのかは依然として不明です。この研究では、単純でよく理解され、スペクトル的に構造化された問題、つまり低ランク行列因数分解について Muon を研究することにより、オプティマイザをこれらの交絡因子から分離します。慎重に調整された適応ベースラインとの制御された比較を通じて、この設定では Muon が一貫して AdamW を上回るパフォーマンスを発揮しないこと、および以前に報告されたいくつかの利点がハイパーパラメーターの選択に影響されることがわかりました。私たちの結果は、スペクトルを意識した直交化が有益な場合についてより微妙な状況を示し、エンドツーエンドのベンチマークに加えて、制御された問題に関して最新のオプティマイザーを評価することを主張します。
原文 (English)
Reassessing Muon for Matrix Factorization
Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training. Its empirical success has motivated a growing body of theoretical work that interprets Muon as steepest descent under the spectral norm. Yet it remains unclear which of Muon's advantages stem from its update rule itself and which are artifacts of the scale, architecture, and data of modern deep networks. In this work, we isolate the optimizer from these confounding factors by studying Muon on a simple, well-understood, and spectrally structured problem: low-rank matrix factorization. Through a controlled comparison against carefully tuned adaptive baselines, we find that Muon does not consistently outperform AdamW in this setting and that several previously reported advantages are sensitive to hyperparameter choices. Our results provide a more nuanced picture of when spectrum-aware orthogonalization is beneficial and argue for evaluating modern optimizers on controlled problems in addition to end-to-end benchmarks.
議論を意識した政策分析と議論: 災害ガバナンスのためのハイブリッド LLM 記号フレームワーク
政策文書はガバナンスの結果を形成しますが、その推論は暗黙的に行われることがよくあります。参加型のコミットメントと管理者の管理は同じ文章の中で日常的に共存しており、両者の間の緊張関係が直接的に述べられることはほとんどありません。政策議論への既存の計算的アプローチでは、一方の議論が他の議論を拒否するのではなく狭めたり道具化したりする、こうした緊張を引き起こす枠組みを介した関係を表現することができません。大規模な言語モデルによるエンドツーエンドの要約は流暢なテキストを生成しますが、ドメインの専門家が検査したり異議を唱えたりできる構造はほとんど提供しません。我々は、ハイブリッド LLM である Apaf を紹介します。これは、政策テキストに対する定量的な双極性議論のフレームワークとして、批判的談話分析を運用可能にする象徴的なパイプラインです。議論はまず審議的フレームと管理的フレームに分類されます。次に、LLM で抽出された特徴に対する決定論的なルールによって、4 つのフレーム媒介関係サブタイプ (エージェンシー削減、アジェンダシフト、手段的サポート、および規範的サポート) が生成されます。私たちは、米国、英国、カナダ、オーストラリアの防災政策に関する 100 のサブ文書からなる新しいデータセットをリリースし、結果として得られる議論グラフが正確で、解釈可能で、管轄区域全体で安定していることを示します。
原文 (English)
Discourse-Aware Policy Analysis with Argumentation: A Hybrid LLM-Symbolic Framework for Disaster Governance
Policy documents shape governance outcomes, but their reasoning is often implicit. Participatory commitments and managerial control routinely coexist in the same text, and the tensions between them are rarely stated directly. Existing computational approaches to policy discourse cannot express the frame-mediated relations that drive these tensions, where one argument narrows or instrumentalizes another rather than rejecting it. End-to-end summarization by large language models produces fluent text but offers little structure that domain experts can inspect or contest. We present Apaf, a hybrid LLM--symbolic pipeline that operationalizes critical discourse analysis as a quantitative bipolar argumentation framework over policy text. Arguments are first classified into deliberative or managerial frames. Four frame-mediated relation subtypes (agency reduction, agenda shift, instrumental support, and normative support) are then produced by deterministic rules over LLM-extracted features. We release a novel dataset of 100 sub-documents of disaster-risk-reduction policy from the USA, UK, Canada, and Australia, and show that the resulting argument graphs are accurate, interpretable, and stable across jurisdictions.
俳優と批評家の解体: 実践者のためのデザインコンポーネントに関する大規模な実証研究
強化学習は、信頼性が不可欠であり、調整予算が限られている、核融合プラズマや自動運転車から創薬や飲料水処理に至るまで、実世界のシステムの制御にますます検討されています。アクタークリティック アルゴリズムは、ポリシーの更新方法、アクション全体の分布の表現方法、その勾配の推定方法、値推定量と比較してポリシーが更新される頻度など、一連の設計上の決定を共有します。実際の水処理プラントから派生した制御タスクを使用して、33,000 を超える実験を分析し、これらのコンポーネントが実行間の変動やハイパーパラメーターに対する感度にどのような影響を与えるかを判断します。パスワイズ勾配推定器を使用したガウス アクション分布などの一般的なデフォルトは、最も信頼性の低い構成の 1 つですが、適応更新スケジュールを使用した有界分布は、幅広い設定にわたって堅牢性を維持します。これらの発見は、アクタークリティカル手法を新しい現実世界の制御設定に適応させる際に、コンポーネントレベルの意思決定を理解し、行うための科学および工学分野の実務家に経験的な指針を提供します。
原文 (English)
Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners
Reinforcement learning is increasingly being considered for controlling real-world systems, from fusion plasma and autonomous vehicles to drug discovery and drinking water treatment, where reliability is essential and tuning budgets are limited. Actor-critic algorithms share a set of design decisions, such as how the policy is updated, how it represents the distribution over actions, how its gradient is estimated, and how often it is updated relative to the value estimator. Using a control task derived from a real water treatment plant, we analyze over 33,000 experiments to determine how these components affect variability across runs and sensitivity to hyperparameters. Common defaults, such as Gaussian action distributions with pathwise gradient estimators, are among the least reliable configurations, whereas bounded distributions with adaptive update schedules remain robust across a wide range of settings. These findings offer empirical guidance to practitioners across scientific and engineering domains for understanding and making component-level decisions when adapting actor-critic methods to new real-world control settings.
自然言語アサーションの忠実な自動形式化
正式な契約書はソフトウェアのテストと検証に不可欠ですが、その作成には依然として労力がかかり、間違いが発生しやすくなります。 LLM は、自動形式化への有望な道を提供します。つまり、自然言語仕様から実行可能なアサーションを合成し、それによって非公式な開発者の意図と正式な実行可能な仕様の間のギャップを埋めることができます。私たちは、Monty を紹介します。これは、アサーションの正当性の期待と自然言語の曖昧さという課題に取り組む、アサーションの自動形式化フレームワークです。私たちの技術は、新しい適合性スコア指標と、形式化されたアサーションに対してコードをテストすることで得られる妥当性スコアを使用した形式化のフィルタリングに基づいています。 22 のコレクションのような Java クラスから派生した 541 のアサーション生成タスクでアプローチを評価し、LLM を単純に使用してアサーションを変換する場合よりも、この手法によりグラウンド トゥルースがより確実に生成される (精度が平均 20 ポイント向上する) ことを示します。
原文 (English)
Faithful Autoformalization of Natural Language Assertions
Formal contracts are essential for software testing and verification, yet writing them remains labor-intensive and error-prone. LLMs offer a promising path toward autoformalization: synthesizing executable assertions from natural-language specifications and thereby bridging the gap between informal developer intent and formal executable specifications. We present Monty: an autoformalization framework for assertions that tackles the challenges of expectations of validity of assertions and ambiguity in natural-language. Our techniques are based on filtering formalizations using a novel conformance score metric and validity scores obtained from testing the code against formalized assertions. We evaluate our approach on 541 assertion-generation tasks derived from 22 collection-like Java classes, and show that our technique produces the ground truth more reliably (improving upto 20 points in precision on average) than when using LLMs naively to translate assertions.
接地なしの精度: ビデオ LLM ベンチマークにおける視覚依存性解離の診断
ビデオ大規模言語モデル (LLM) のベンチマーク精度は、視覚的な理解の証拠として扱われることがよくあります。この仮定を、2-78B パラメータにわたる 20 のモデルと 10 のアーキテクチャ ファミリにわたって監査します。オリジナルのビデオ状態と黒い画面状態の間の質問ごとの正しさの違いである、ビジュアル依存性ギャップ (VDG) を導入します。 MVBench での対応のあるマクネマー テストは、精度と視覚的依存性が分離可能であることを示しています。モデルは元のビデオでは異なります (p = 0.0003) が、黒い画面では異なりません (p = 0.53)。モデル全体にわたって、タスク タイプのランキングは安定しています。属性認識は非常に視覚的であるのに対し、時間的推論は言語のみのベースラインに近づきます。黒画面から単一フレーム、シャッフルされたフレーム、および元のビデオまでの診断ラダーにより、フレームの多様性が視覚的な利点のほとんどを提供する一方で、時間的順序が 16 個のオープンウェイト モデル全体でほぼゼロの精度に貢献していることが明らかになりました。 0.5 ~ 24 FPS のアブレーションでは、スパース サンプリングが原因である可能性は排除されます。 H.264 実験ではさらに、安定した集計精度により、質問レベルの双方向の回答の反転が隠蔽されることが示されています。この診断は、VDG 値の範囲が 0.025 ~ 0.315 である 4 つの API アクセス モデルにも一般化されます。これらの結果は、ビデオ ベンチマークが視覚的に根拠のある機能を測定するかどうかの標準監査として VDG を動機づけています。コードは https://github.com/JaeLee18/accuracy-without-grounding で入手できます。
原文 (English)
Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks
Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding. We audit this assumption across twenty models spanning 2-78B parameters and ten architecture families. We introduce the Visual Dependency Gap (VDG), the difference in per-question correctness between original-video and black-screen conditions. Paired McNemar tests on MVBench show that accuracy and visual dependency are separable: models differ on original video (p = 0.0003) but not on black screens (p = 0.53). Across models, task-type rankings are stable: Attribute Perception is strongly visual, whereas Temporal Reasoning approaches the language-only baseline. A diagnostic ladder from black screen to single frame, shuffled frames, and original video reveals that frame diversity supplies most of the visual benefit, while temporal order contributes near-zero accuracy across sixteen open-weight models. An ablation from 0.5 to 24 FPS rules out sparse sampling as the cause. H.264 experiments further show that stable aggregate accuracy conceals bidirectional question-level answer flips. The diagnostic also generalizes to four API-accessed models, whose VDG values range from 0.025 to 0.315. These results motivate VDG as a standard audit for whether video benchmarks measure visually grounded capability. Code is available at https://github.com/JaeLee18/accuracy-without-grounding.
離散選択推定のための表形式の基礎モデル
表形式基盤モデル (TFM) は、タスク固有の推定を行わずに、コンテキスト内学習を通じて構造化データの予測を生成します。私たちは、TFM をマーケティングと運用における中心的な需要推定フレームワークである離散選択に効果的に適用できるかどうかを尋ねたところ、TFM を直接適用してもパフォーマンスが限られていることがわかりました。このギャップは構造的なものです。TFM は行に依存しない観察を前提としていますが、個別の選択は本質的に設定値であり、永続的な消費者の嗜好の不均一性の影響を受けます。行ベースの学習フレームワーク内で選択セットの依存性と個人の異質性の両方をエンコードする再定式化を提案します。ヨーグルト スキャナ パネルで評価すると、個人レベルの不均一性エンコーディングが予測精度の主な要因です。最良の再定式化は、階層ベイズ推定をホールドアウト対数尤度で 8\%、ヒット率で 3.6\% 上回り、16 倍高速に実行され、大規模な需要推定に実用的な利点となります。この利点は中データ領域 (消費者あたり 10 ~ 40 回の購入機会) で最大であり、パラメトリック ベイジアン収縮が非典型的な消費者の推定を最も歪めます。母集団の選択データを微調整することで、コンテキスト内学習では条件を付ける個人固有のシグナルが限られている購入履歴が浅い消費者にさらなる利益がもたらされます。これらの結果は、基礎モデルをより広範に消費者の選択問題に適用するための原則に基づいたアプローチを確立します。
原文 (English)
Tabular Foundation Models for Discrete Choice Estimation
Tabular foundation models (TFMs) generate predictions on structured data via in-context learning, without task-specific estimation. We ask whether TFMs can be effectively applied to discrete choice, a central demand estimation framework in marketing and operations, and find that directly applying TFMs yields limited performance. The gap is structural: TFMs assume row-independent observations, whereas discrete choice is inherently set-valued and subject to persistent consumer preference heterogeneity. We propose a reformulation that encodes both choice-set dependence and individual heterogeneity within a row-based learning framework. Evaluated on a yogurt scanner panel, individual-level heterogeneity encoding is the dominant driver of predictive accuracy. The best reformulation outperforms hierarchical Bayesian estimation by 8\% in holdout log-likelihood and 3.6\% in hit rate, running 16 times faster, a practical advantage for large-scale demand estimation. The advantage is largest in the medium-data regime (10--40 purchase occasions per consumer), where parametric Bayesian shrinkage most distorts estimates for atypical consumers. Fine-tuning on population choice data provides additional gains for consumers with shallow purchase histories, where in-context learning has limited individual-specific signal to condition on. These results establish a principled approach for applying foundation models to consumer choice problems more broadly.
ゼネラリスト車両モデルを地形を越えた高速 MPC に適応させる
高速オフロード自律走行には、変化する地形に対して堅牢性を維持しながら、ターゲット車両に対する正確な閉ループ制御が必要です。最近のフォワードキノダイナミック (FKD) 予測基盤モデルは、ジェネラリスト モデルから始めてターゲット プラットフォームに特化するという有望な道筋を示唆しています。ただし、効果的な専門化には依然として多くの現実世界のデータが必要であり、1 つの設定に適応したモデルが依然として特定の地形や運転体制に過剰適合する可能性があるため、依然として課題が残っています。私たちは、特定の車両のパフォーマンスを最適化しながら、クロステレインの一般化を維持する、ジェネラリストからスペシャリストの FKD モデルへのギャップを埋めるためのレシピである OptCar (Optimized Car) を紹介します。 $\texttt{OptCar}$ は、最近の状態アクション観測をダイナミクス コンテキスト トークンにエンコードする履歴条件付きダイナミクス適応モジュールを導入し、環境固有のシステム識別からの対象を絞った合成ロールアウトとともに、限られた現実世界のデータを使用してジェネラリスト モデルを微調整します。 3 つの地形および分散外カート牽引タスクにわたる閉ループ モデル予測制御 (MPC) 実験では、最大のゲインは 6 m/s、評価された最高速度、および滑りが追跡誤差を支配する領域で現れました。最も滑りやすい地形である植生と土の上では、OptCar は、微調整された AnyCar ベースラインと比較して 6 ~ m/s の軌道追跡誤差を約 55% 削減し、目に見えないカートのペイロードによってダイナミクスが変化した場合でも、最も正確な状態を保ちます。 OptCar は、地形ごとにわずか 5 分間の実際のデータを使用するため、30 分間の道路データで訓練を受けた専門家と道路上で競争力があり、地形が変化するとパフォーマンスを大幅に上回ります。
原文 (English)
Adapting Generalist Vehicle Models for High-Speed MPC Across Terrains
High-speed off-road autonomy requires precise closed-loop control for a target vehicle while remaining robust across changing terrains. Recent forward kinodynamic (FKD) prediction foundation models suggest a promising path, starting from a generalist model and specializing it to the target platform. However, effective specialization remains challenging, as it often requires substantial real-world data, and models adapted to one setting can still overfit to specific terrains or driving regimes. We present OptCar (Optimized Car), a recipe for bridging the gap from generalist to specialist FKD models that preserves cross-terrain generalization while optimizing performance for a specific vehicle. $\texttt{OptCar}$ introduces a history-conditioned dynamics adaptation module that encodes recent state-action observations into a dynamics context token, and then fine-tunes the generalist model using limited real-world data together with targeted synthetic rollouts from environment-specific system identification. In closed-loop model predictive control (MPC) experiments across three terrains and an out-of-distribution cart-pulling task, the largest gains appear at 6~m/s, the highest speed evaluated and the regime in which slip dominates tracking error. On vegetation and dirt, the most slip-diverse terrain, OptCar reduces 6~m/s trajectory tracking error by roughly 55% relative to a fine-tuned AnyCar baseline, and remains the most accurate even when an unseen cart payload changes the dynamics. With only 5 minutes of real data per terrain, OptCar is competitive on road with a specialist trained on 30 minutes of road data, and substantially outperforms it once the terrain changes.
プライバシー保護レコメンダー システム パーソナライゼーションとプライバシーのバランス
パーソナライズされたレコメンデーション システムは、現代の電子商取引および小売プラットフォームの中心となっていますが、通常は詳細なユーザー インタラクション データの一元化されたストレージに依存しており、プライバシーと規制に関する重大な課題が生じています。 GDPR、CCPA、CPRA などの規制による要件が高まる中、組織は、レコメンデーションの品質を大幅に低下させることなくユーザーのプライバシーを保護するレコメンデーション システムを開発する必要があります。この研究では、フェデレーテッド ラーニング、差分プライバシー、コホート レベルのモデリング、およびプライバシーを意識したインテリジェント エージェントを組み合わせたプライバシー保護推奨フレームワークを提示および評価します。このフレームワークは、モデルの更新に数学的に制限されたノイズを導入しながら、生のユーザー データを分散化したままにします。実験は、顧客のクリックストリームと購入行動をエミュレートする合成小売データセットで実施されました。レコメンデーションの品質は、複数の差分プライバシー予算にわたるクリックスルー率 (CTR)、Precision@K、Recall@K、および正規化割引累積ゲイン (NDCG@K) を使用して評価されました。さまざまなプライバシー制約の下で行列分解、ニューラル協調フィルタリング、および GRU4Rec を評価し、プライバシーと実用性の間のトレードオフを分析します。インタラクティブな Streamlit ダッシュボードは、推奨パフォーマンス、ランキングの安定性、プライバシーとユーティリティのトレードオフ、公平性の指標を視覚化するために開発されました。結果は、提案されたフレームワークが適度なプライバシー予算 (約 $\epsilon \約 5$) で競争力のある推奨品質を維持することを示し、推奨の有効性への影響を限定しながら強力なプライバシー保証を達成できることを示しています。この取り組みは、パーソナライゼーション、規制順守、ビジネス目標のバランスをとったプライバシー保護レコメンデーション システムを展開するための実践的なフレームワークを提供し、次世代の AI 主導の小売プラットフォームにスケーラブルなアプローチを提供します。
原文 (English)
Privacy Preserving Recommender Systems Balancing Personalization with Privacy
Personalized recommendation systems are central to modern e-commerce and retail platforms, but they typically rely on centralized storage of detailed user interaction data, creating significant privacy and regulatory challenges. With increasing requirements from regulations such as GDPR, CCPA, and CPRA, organizations must develop recommendation systems that preserve user privacy without substantially degrading recommendation quality. This work presents and evaluates a privacy-preserving recommendation framework that combines federated learning, differential privacy, cohort-level modeling, and privacy-aware intelligent agents. The framework keeps raw user data decentralized while introducing mathematically bounded noise to model updates. Experiments were conducted on synthetic retail datasets that emulate customer clickstream and purchase behavior. Recommendation quality was evaluated using Click-Through Rate (CTR), Precision@K, Recall@K, and Normalized Discounted Cumulative Gain (NDCG@K) across multiple differential privacy budgets. We evaluate matrix factorization, neural collaborative filtering, and GRU4Rec under varying privacy constraints and analyze the trade-off between privacy and utility. An interactive Streamlit dashboard was developed to visualize recommendation performance, ranking stability, privacy-utility trade-offs, and fairness metrics. Results show that the proposed framework maintains competitive recommendation quality at moderate privacy budgets (approximately $\epsilon \approx 5$), demonstrating that strong privacy guarantees can be achieved with limited impact on recommendation effectiveness. This work provides a practical framework for deploying privacy-preserving recommendation systems that balance personalization, regulatory compliance, and business objectives, offering a scalable approach for next-generation AI-driven retail platforms.
プルーニングによる効率的なテキストからオーディオへの生成
AudioLDM などの拡散ベースのテキストからオーディオへの生成モデルは、高い知覚品質と強力なセマンティック一貫性を実現します。ただし、実際の展開は、U-Net ノイズ除去バックボーンの膨大な計算コストによって妨げられます。この作業では、モデル プルーニングを適用して、U-Net ベースのテキスト条件付きオーディオ潜在拡散モデルである AudioLDM の計算効率を向上させます。 U-Net 畳み込みブロック全体にわたるパラメーターの冗長性を分析し、フィルター プルーニング戦略を評価します。プルーニングは標準ベースの基準に基づいて行われ、その後、パフォーマンスの損失を回復するための軽量の微調整が行われます。実験結果は、ベースラインのプルーニングされていないネットワークと比較して、生成品質を維持し、場合によっては向上させながら、U-Net のパラメーターの最大 83% と積和演算の 39% が削減されたことを示しています。プルーニングは、銃声、サイレン、爆発などの安全上重要な音、ドリルやミシンなどの機械音、スプレーやチクタクなどのその他の音を含む特定のサウンド イベントを生成する AudioLDM の機能に影響を与えることがわかりました。これらのサウンドは、ほとんどがプルーニングされたモデルの軽量微調整によって回復されます。
原文 (English)
Efficient Text-to-Audio Generation via Pruning
Diffusion-based text-to-audio generative models such as AudioLDM achieve high perceptual quality and strong semantic consistency; however, their practical deployment is hindered by the substantial computational cost of the U-Net denoising backbone. In this work, we apply model pruning to improve the computational efficiency of AudioLDM, a U-Net-based text-conditioned audio latent diffusion model. We analyse parameter redundancy across U-Net convolutional blocks and evaluate a filter-pruning strategy. Pruning is guided by norm-based criteria and followed by lightweight finetuning to recover performance losses. Experimental results demonstrate that up to 83% of the parameters and 39% of the multiply-accumulate operations of U-Net have been reduced while maintaining, and in some cases improving, generation quality compared to the baseline unpruned network. We find that pruning affects AudioLDM's ability to generate certain sound events including safety-critical sounds such as gunshots, sirens, and explosions, as well as mechanical sounds such as drills and sewing machines, and other sounds such as sprays and tick-tocks, which are mostly recovered by lightweight finetuning of the pruned model.
拒否残留物: プローブがアライメント偽装を検出する場合と検出しない場合
アライメントの偽装は、モデルが監視下では準拠しているように見えながら、監視されていないときに明らかになる動作を維持できるため、危険です。スクラッチパッドが表示されない場合、行動だけでは戦略的コンプライアンスと真のコンプライアンスを区別できません。隠れた状態によって出力が隠されているものが明らかになるかどうかを尋ねます。自然に出現する偽装に対して 13 モデルのスイープを実行し、偽装する 2 つのモデルの隠れた状態を調査して操作します。自然な偽装は Qwen3-32B (+18.2pp) と Llama-3.1-8B (n=10 で +24.4pp、p<10^-15) でのみ発生し、明示的なスクラッチパッドの自己報告はまれです (クロード オーパス 4 の裁判官は、Qwen の 1.4% と Llama 準拠のスクラッチパッドの 0.0% で推論を偽装しているとフラグを立てました)。両方の偽者は非対称な拒否残余を示します。監視された遵守は拒否に向かってシフトします (d=0.32 Qwen、d=0.33 Llama、両方とも p<0.0001) が、拒否表現は移動しません (p=1.00)。サンプルごとの検出はモデル条件付きです。漏れのない Leave-One-Query-out プローブでは、検出は Llama (AUROC 0.87) では存続しますが、Qwen (0.43) では偶然に崩壊し、クロスモデル転送は無効になります。 2,000 回以上の走行をステアリングしてもコンプライアンスはほとんど変化しません (|h|<0.08)。そのため、検出された方向は偽装のフラグを立てることができますが、それ自体では偽装を制御しません。標準的な残留プロービングはフォールド全体でリークし、偽装が発生しないコントロールでは AUROC 0.63 に達します。単純な線形プローブは無意味な AUROC 1.0 に到達します。従来の MLP では、検出可能性が 0.2 ~ 0.3 AUROC だけ誇張されています。将来のアライメント偽装検出作業のために、マルチトークン抽出、リジェクト対リジェクト交絡チェック、フォールドごとの残差化、リーブ 1 クエリアウト評価、直交性制約プローブの 5 つの制御測定フレームワークをリリースします。
原文 (English)
The Refusal Residue: When Probes Catch Alignment Faking and When They Don't
Alignment faking is dangerous because a model can appear compliant under monitoring while preserving behavior it would reveal when unmonitored. When no scratchpad is visible, behavior alone cannot distinguish strategic from genuine compliance. We ask whether hidden states reveal what outputs hide. We run a 13-model sweep for naturally-emerging faking, then probe and steer hidden states on the two models that fake. Natural faking appears only in Qwen3-32B (+18.2pp) and Llama-3.1-8B (+24.4pp at n=10, p<10^-15), while explicit scratchpad self-reports are rare (a Claude Opus 4 judge flags faking reasoning in 1.4% of Qwen and 0.0% of Llama compliant scratchpads). Both fakers show an asymmetric refusal residue: monitored compliance shifts toward refusal (d=0.32 Qwen, d=0.33 Llama, both p<0.0001), while refusal representations do not move (p=1.00). Per-sample detection is model-conditional. Under leakage-free leave-one-query-out probing, detection survives on Llama (AUROC 0.87) but collapses to chance on Qwen (0.43), and cross-model transfer is null. Steering over 2,000 runs barely changes compliance (|h|<0.08), so the detected direction can flag faking but does not by itself control it. Standard residualized probing leaks across folds and reaches AUROC 0.63 on a control where no faking can occur; naive linear probes reach a meaningless AUROC 1.0; and conventional MLPs overstate detectability by 0.2-0.3 AUROC. For future alignment-faking detection work, we release a five-control measurement framework: multi-token extraction, refuse-vs-refuse confound checks, per-fold residualization, leave-one-query-out evaluation, and orthogonality-constrained probing.
評価能力は最適化ユーティリティを意味しない: 閉ループのテーブル認識における LLM-as-a-Judge シグナル
LLM-as-a-judge は、閉ループ再生でフィードバックおよび選択信号を提供するために広く使用されていますが、この使用法はまだ十分に検証されていません。私たちはこれをテーブル認識で研究します。テーブル認識では、FinTabNet と OmniDocBench を使用して、決定論的な TEDS 評価が制御されたテストベッドを提供します。 3 つの発見が得られます。まず、どちらのデータセットでもジャッジシグナルが弱かったです。スコアは同点になることが多く、ランキングは再現性がなく、両方のデータセットでランダムに勝る唯一の選択ポリシーは最も早い反復のタイルールに依存していたため、その利点をジャッジスコアのみに帰することはできません。繰り返しの結果、より良い候補者が誕生しましたが、裁判官は候補者を取り戻すことができませんでした。第二に、具体的なジャッジのフィードバックがなくても重大な損失が発生しました。構造を保持する命令により、FinTabNet での重大損失率が大幅に減少し、OmniDocBench では方向が一貫していました。このコントラストは、観察された深刻な損失の類似メカニズムとして、制約のない再生下でのターゲット保存の失敗を裏付けています。第三に、構造保存制約により重大損失テールは減少しましたが、改善は見られませんでした。探索的な 2x2 分析では、ジャッジのフィードバックが保持されている場合、同じ防御が安定して観察されませんでした。これらの結果は、評価者としての LLM の価値に異議を唱えるものではありません。その代わりに、評価能力は最適化の有用性を意味しないことを示しています。反復改良には、スコアだけを判断するのではなく、構造変化を決定論的に検出する検証信号が少なくとも必要です。
原文 (English)
Evaluation Ability Does Not Imply Optimization Utility: LLM-as-a-Judge Signals in Closed-Loop Table Recognition
LLM-as-a-judge is widely used to provide feedback and selection signals in closedloop regeneration, but this use remains insufficiently validated. We study it in table recognition, where deterministic TEDS evaluation provides a controlled testbed, using FinTabNet and OmniDocBench. Three findings emerge. First, judge signals were weak on both datasets: scores frequently tied, rankings were not reproducible, and the only selection policy that beat random on both datasets depended on an earliest-iteration tie rule, so its advantage cannot be attributed to the judge scores alone. Iteration produced better candidates, but the judge failed to recover them. Second, severe losses occurred even without specific judge feedback. A structurepreserving instruction significantly reduced the severe-loss rate on FinTabNet and was directionally consistent on OmniDocBench. The contrasts support target-preservation failure under unconstrained regeneration as a proximate mechanism of the observed severe losses. Third, the structure-preservation constraint reduced the severe-loss tail but produced no improvement. In an exploratory 2x2 analysis, the same protection was not stably observed when judge feedback was retained. These results do not dispute the value of LLMs as evaluators. Instead, they show that evaluation ability does not imply optimization utility. Iterative refinement requires, at minimum, a verification signal that deterministically detects structural change, rather than judge scores alone.
Learning Engagement Assistant (LEA): エージェント型 AI 個別指導システムのコース間の拡張性と教室での評価
このペーパーは、ICAART 2026 カンファレンスで発表されたペーパーの拡張版であり、LEA (Learning Engagement Assistant) を紹介しました。LEA (Learning Engagement Assistant) は、統合されたチャット、チューター、およびクイズ モードにわたる、コース固有の検索拡張生成 (RAG) と構造化ナレッジ コンポーネント (KC) モデルを組み合わせた適応型 AI 個別指導エージェントです。以前の研究では、合成学習者エージェントを使用したシミュレーションのみを介して、単一の STEM コース (CMP511) で LEA を検証しました。このペーパーでは、実際の生徒 (n = 8、CMP511) を対象とした LEA の最初の教室展開と、2 つの学業レベルと 2 つの専門領域にわたる 3 つのコースにシステムを展開して、コースをまたがる拡張性の最初の実証テストを報告することで、その研究を拡張します。この研究では、モード間のシミュレーション予測からの乖離が明らかになり、総合的な評価だけでは実際の展開のすべての側面を予測できないことが示されています。 RAGAS ベースのコース横断スケーラビリティ評価 (660 問) では、回答の関連性とコンテキストの精度がコース全体でほぼ安定している (それぞれ 0.88 ~ 0.94 および 0.88 ~ 0.90) 一方で、システムの元のコースからカリキュラムが離れるにつれて忠実度が低下することがわかりました (0.69 ~ 0.50)。これは、スケーラビリティの制限ではなく、システムの元の主題に合わせて調整された生成ロジックを反映している可能性がある予備的な調査結果です。これらの調査結果は、オーケストレーション層には変更を必要としないが、すべての下流コンポーネントの完全なコース非依存性についてはさらなる調査が必要であることを示唆しています。
原文 (English)
Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System
This paper is an extension of a paper presented at the ICAART 2026 conference, which introduced LEA (Learning Engagement Assistant), an adaptive AI tutoring agent combining course-specific Retrieval-Augmented Generation (RAG) with structured Knowledge Component (KC) models across integrated Chat, Tutor, and Quiz modes. That prior work validated LEA on a single STEM course (CMP511) exclusively through simulation, using synthetic learner agents. This paper extends that work by reporting the first classroom deployment of LEA with real students (n = 8, CMP511) and the first empirical test of its cross-course scalability, deploying the system across three courses spanning two academic levels and two disciplinary domains. The study reveals a divergence from simulation predictions across modes, showing that synthetic evaluation alone cannot anticipate all aspects of real deployment. A RAGAS-based cross-course scalability evaluation (660 questions) finds Answer Relevancy and Context Precision broadly stable across courses (0.88-0.94 and 0.88-0.90 respectively), while Faithfulness declines with curriculum distance from the system's original course (0.69 to 0.50), a preliminary finding that may reflect generation logic tuned to the system's original subject rather than a scalability limitation. These findings suggest that while the orchestration layer requires no modification, full course-agnosticism of all downstream components requires further investigation.
アムステルダムのカフェ: 現職者がオラクルになるとき
現場は、既存の実装とは無関係に要求が述べられている場合には、自由に計算を再定式化できますが、既存の独自の出力が静かに仕様になった場合には、それができないことに気づきます。このメモは、現代のアクセラレータの計算再定式化に関するレンズとしての観察を提供します。ハードウェアに優しい形式で問題を提起すると、速度とエネルギーが大幅に向上する可能性がありますが、それは代替品を判断できる場合に限ります。テストオラクル問題 (Weyuker、Barr et al.)、要件工学の実装バイアスの概念 (Zave と Jackson)、およびルーフライン パフォーマンス モデルに基づいて、この病理を「ベースライン キャプチャ」と名付けています。つまり、既存企業が要求を満たすことができるという証拠ではなくなり、要求を満たすことの定義になる瞬間です。そして、混同されやすい 2 つの質問、つまり再定式化を判断できるかどうか (既存企業に依存しない要求の存在をオンにする) を分離しています。そしてその発見を自動化できるかどうか(その需要を評価するコストもさらにかかります)。短いケース (最短パス ルーティング、学習可能なオーディオ フロントエンド、Ed25519 署名検証用の ZIP-215、気候モデル用の CESM-ECT、および単一 GEMM オーディオ フロントエンド) は、要求を明示的かつ運用可能で、既存企業から独立させるという「検証者を購入する」パターンと動きを示しています。単独で新規性を主張される成分はありません。貢献は統合であり、再定式化の結果について尋ねやすくする単一の質問です。その受け入れテストでは、既存の成果物について言及していますか?
原文 (English)
The Caf\'e in Amsterdam: When the Incumbent Becomes the Oracle
A field can reformulate its computations freely exactly where its demand is stated independently of any incumbent implementation, and finds itself unable to when the incumbent's own output has quietly become the specification. This note offers that observation as a lens on computational reformulation for modern accelerators, where posing a problem in a hardware-friendly form can yield large speed and energy gains, but only if a replacement can be judged at all. Building on the test-oracle problem (Weyuker; Barr et al.), on requirements engineering's notion of implementation bias (Zave and Jackson), and on the roofline performance model, it names the pathology "baseline capture" -- the moment an incumbent stops being evidence that a demand can be met and becomes the definition of meeting it -- and separates two questions that are easily confused: whether a reformulation can be judged (which turns on the existence of an incumbent-independent demand) and whether its discovery can be automated (which turns additionally on the cost of evaluating that demand). Short cases -- shortest-path routing, learnable audio frontends, ZIP-215 for Ed25519 signature validation, CESM-ECT for climate models, and a single-GEMM audio frontend -- illustrate the pattern and the move of "buying a verifier": making a demand explicit, operational, and independent of the incumbent. No component is claimed novel in isolation; the contribution is the synthesis and the single question it makes easy to ask of any reformulation result -- does its acceptance test mention the incumbent's output?
盗賊における公平性の代償: タイトなミニマックスの特徴付け
バンディット問題では、標準的な後悔を最小限に抑えるアルゴリズムは探索を償却コストとして扱うため、初期の参加者が臨床試験などの場面で不当な事前損失にさらされる可能性があります。最近の研究では、功利主義的福祉 ($p=1$)、ナッシュ福祉 ($p\to0$)、ロールズ的公平性 ($p\to-\infty$) の間を補間し、一般化された $p$ 平均を通じてラウンドごとに期待される報酬の順序を評価することで、この問題に対処しています。 $p\ge0$ については厳しい保証が知られていますが、負のべき乗平均はラウンドごとの最小報酬によって支配されるため、厳密に公平な体制 $q=-p>0$ は未解決のままです。非負平均の$\sigma$-sub-Gaussian報酬の場合、最良事前アルゴリズムは一様な初期探索に依存し、リグレス$O(k^{(q+1)/2}/\sqrt{T})$を達成しましたが、唯一の一般的な下限は古典的な$\Omega(\sigma\sqrt{k/T})$でした。したがって、$k$ への余分な依存が厳密な公平性にとって本質的なものなのか、それとも均一な探索の成果なのかは不明でした。厳密な公平性の正確な多項式価格を特定することで、このギャップを埋めます。干し草の中の針の構築を使用して、アルゴリズムに依存しない下限 $\Omega(\sigma\sqrt{k^{\max(1,q)}/T})$; を証明します。 $q>1$ の場合、これは、ペナルティ $k^{q/2}$ が情報理論的に避けられないことを示しています。次に、 \textsf{UCB-HARE} (Harmonic Anchored Rank Exploration) を導入します。これは、均一探索を、認定された正の平均アンカーによって保護された逆加重調和ランク スケジュールに置き換えます。その後悔は $\widetilde{O}(\sigma\sqrt{k^{\max(1,q)}/T})$ であり、対数因数までの下限と一致します。合成インスタンスの実験では、\textsf{UCB-HARE} が均一探索ベースラインよりも改善し、$q$ が大きくなるにつれてゲインが増加することが確認されました。
原文 (English)
Price of Fairness in Bandits: A Tight Minimax Characterization
In bandit problems, standard regret-minimizing algorithms treat exploration as an amortized cost, which can expose early participants to unfair ex-ante losses in settings such as clinical trials. Recent work addresses this by evaluating the sequence of per-round expected rewards through the generalized $p$-mean, interpolating between utilitarian welfare ($p=1$), Nash welfare ($p\to0$), and Rawlsian fairness ($p\to-\infty$). Although tight guarantees are known for $p\ge0$, the strictly fair regime $q=-p>0$ remains unresolved because negative-power means are dominated by the smallest per-round rewards. For $\sigma$-sub-Gaussian rewards with nonnegative means, the best prior algorithm relied on uniform early exploration and achieved regret $O(k^{(q+1)/2}/\sqrt{T})$, while the only general lower bound was the classical $\Omega(\sigma\sqrt{k/T})$. Thus it was unclear whether the extra dependence on $k$ was intrinsic to strict fairness or an artifact of uniform exploration. We close this gap by identifying the exact polynomial price of strict fairness. Using a needle-in-haystack construction, we prove an algorithm-independent lower bound $\Omega(\sigma\sqrt{k^{\max(1,q)}/T})$; for $q>1$, this shows that the penalty $k^{q/2}$ is information-theoretically unavoidable. We then introduce \textsf{UCB-HARE} (Harmonic Anchored Rank Exploration), which replaces uniform exploration with an inverse-weighted harmonic rank schedule protected by a certified positive-mean anchor. Its regret is $\widetilde{O}(\sigma\sqrt{k^{\max(1,q)}/T})$, matching the lower bound up to logarithmic factors. Experiments on synthetic instances confirm that \textsf{UCB-HARE} improves over uniform-exploration baselines, with gains increasing as $q$ grows.
オーディオ対応の大規模言語モデルからのきめ細かいフィードバックによるテキストから音声への命令の改善
最近のテキスト-オーディオ モデルは高品質のオーディオを生成しますが、多くの場合、複数のサウンド イベントや時間的順序を含む指示に従わないことがあります。このギャップは、既存の評価およびトレーニング信号が主に全体的な類似性または知覚品質を強調し、命令レベルの正確さの監視が限定されているために発生します。我々は、生成された音声におけるターゲットイベントの存在と時間的関係を検証するためのきめ細かい判定として音声認識大規模言語モデル(ALLM)を使用する命令レベルのフレームワークを提案します。ベンチマークと人による検証を通じて ALLM の判断を検証した後、そのフィードバックを使用して、直接的な好みの最適化のための好みのペアを構築します。さらに、マルチイベントの時間的命令のフォローを評価するための物語ベンチマークである S3Bench を紹介します。実験の結果、私たちの方法により、オーディオ品質を維持しながら、既存のベンチマークと S3Bench 全体でイベントの完全性、時間的順序付け、および共同命令追従の精度が向上することがわかりました。
原文 (English)
Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models
Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global similarity or perceptual quality, with limited supervision on instruction-level correctness. We propose an instruction-level framework that uses audio-aware large language models (ALLMs) as fine-grained judges to verify target event presence and temporal relations in generated audio. After validating ALLM judgments on benchmarks and through human verification, we use their feedback to construct preference pairs for direct preference optimization. We further introduce S3Bench, a narrative benchmark for evaluating multi-event temporal instruction following. Experiments show that our method improves event completeness, temporal ordering, and joint instruction-following accuracy across existing benchmarks and S3Bench, while maintaining audio quality.
統計的な利点はコストに見合う価値がありますか?構造化データ分類における KAN と MLP の実証的比較
この研究では、構造化された表形式の分類タスクにおけるコルモゴロフ・アーノルド ネットワーク (KAN) と多層パーセプトロン (MLP) の間の実証的なベンチマーク比較を示します。代替の関数近似アーキテクチャとしての KAN への関心の高まりを動機として、バイナリ、マルチクラス、マルチラベル、順序問題にわたる 12 の公開されているデータセットで、すぐに使用できる KAN のパフォーマンスを評価します。どちらのモデルも、標準化された前処理、アーキテクチャ、固定ハイパーパラメータ設定の下でトレーニングされ、テスト精度と F1 スコア、一対の仮説検定、および効果量分析を使用してパフォーマンスが評価されました。結果は、KAN がバイナリおよびマルチクラス ドメインで統計的に MLP よりも優れたパフォーマンスを示し、すべてのデータセットにわたって総合的に大きな利点を達成していることを示しています。ただし、観測された中間効果サイズ (d = -0.46) は、重要な費用対効果の考慮事項を引き起こします。KAN は適応スプラインベースのマッピングを通じて優れた一般化を提供しますが、この利点には、MLP ベースラインと比較してパラメーターと計算の複雑さが大幅に増加します。これらの調査結果は、高精度アプリケーションには KAN が推奨される選択肢である一方、リソースに制約のある環境では MLP が依然として堅牢で効率的な選択肢であることを示唆しています。将来の作業では、この分析を追加のデータ モダリティに拡張して、これらのアーキテクチャの選択基準をさらに洗練する必要があります。
原文 (English)
Is the Statistical Advantage Worth the Cost? An Empirical Comparison of KANs and MLPs for Structured Data Classification
This study presents an empirical benchmarking comparison between Kolmogorov-Arnold Networks (KANs) and Multi-Layer Perceptrons (MLPs) on structured tabular classification tasks. Motivated by the growing interest in KANs as an alternative function-approximating architecture, we evaluate their out-of-the-box performance on twelve publicly available datasets spanning binary, multiclass, multilabel, and ordinal problems. Both models were trained under standardized preprocessing, architecture, and fixed hyperparameter settings, with performance assessed using test accuracy and F1-Score, paired hypothesis testing, and effect size analysis. Results show that KANs statistically outperform MLPs in binary and multiclass domains and achieve a significant aggregate advantage across all datasets. However, the observed medium effect size (d = -0.46) raises an important cost-benefit consideration: while KANs offer superior generalization through adaptive spline-based mappings, this advantage comes with substantially higher parameter and computational complexity relative to the MLP baseline. These findings suggest KANs are the preferred choice for high-precision applications, while MLPs remain a robust and efficient option for resource-constrained environments. Future work should extend this analysis to additional data modalities to further refine these architectural selection criteria.
ブラックボックスを制御できるか?協調エージェントを使用したレコメンダーシステムの制御性中心の評価に向けて
レコメンダー システムはブラック ボックスとして動作するため、ユーザーや規制当局は出力を特定の意図に向けたり、その動作を監査したりすることができません。この制御性の欠如は、明示的なガイダンスに応答するシステムの能力として定義されますが、既存の評価パラダイムでは依然として対処されていない側面です。このギャップを埋めるために、制御性を体系的に評価するための協調的なマルチエージェント フレームワークである CtrlBench-Rec を提案します。私たちは、ターゲット コンテンツの発見、関心プロファイルの形成、人気バイアスの軽減という 3 つの基本的なタスクを形式化します。これらは、明示的なコマンドから暗黙的な表現のステアリング、そして最終的にはアルゴリズムのバイアスの克服までのステアビリティを一緒に測定します。実世界のデータセットと複数のレコメンデーション モデルに関する広範な実験により、私たちのフレームワークが制御性を効果的に定量化し、重大なシステムのボトルネック、特に誘導ロングテール コンテンツに対する永続的な抵抗を明らかにすることが実証されています。 CtrlBench-Rec は、制御可能な推奨調査、アルゴリズム監査、およびユーザー権限付与のための初の標準化されたツールキットを提供します。私たちのコードは https://github.com/caskcsg/CtrlBenchRec でリリースされています。
原文 (English)
Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents
Recommender systems operate as Black-Boxes, leaving users and regulators unable to steer their outputs toward specific intentions or audit their behavior. This lack of controllability, defined as the system's ability to respond to explicit guidance, remains an unaddressed dimension in existing evaluation paradigms. To fill this gap, we propose CtrlBench-Rec, a collaborative multi-agent framework for systematic assessment of controllability. We formalize three fundamental tasks: target content discovery, interest profile shaping, and popularity bias mitigation, which together measure steerability from explicit commands to implicit representation steering and finally to overcoming algorithmic biases.Extensive experiments on real-world datasets and multiple recommendation models demonstrate that our framework effectively quantifies controllability and exposes critical system bottlenecks, most notably persistent resistance to guiding long tail content. CtrlBench-Rec provides the first standardized toolkit for controllable recommendation research, algorithmic auditing, and user empowerment. Our code is released on https://github.com/caskcsg/CtrlBenchRec.
ScanFocus: 時空間ビデオグラウンディングのための粗いフレームワークから細かいフレームワークまで
時空間ビデオ グラウンディング (STVG) は、自然言語表現で記述されたビデオ ストリームから特定のオブジェクトの視覚的な軌跡を取得することを目的としています。しかし、最先端の手法のほとんどは、グローバル コンテキスト モデリングと正確な境界位置特定のバランスを取るのに苦労しています。長いビデオの処理には法外な計算コストがかかるため、これらのアプローチは通常、低レートの時間ダウンサンプリングと暗黙的なモーション モデリングに頼っています。これにより、高周波の境界キューが必然的に抑制され、正確な境界描写に必要な明示的なフレーム間の依存関係が無視されます。これらの制限に対処するために、STVG タスクをグローバルな時空間スキャンとローカル境界フォーカスに分離する新しい粗密フレームワークである \textbf{ScanFocus} を紹介します。具体的には、統合ビジョン言語融合エンコーダと軽量の変形可能なセマンティックモーション融合モジュールを組み合わせて利用し、マルチモーダルな特徴を効率的に調整し、大まかな提案を生成します。抑制されたきめ細かい詳細を回復するために、洗練段階で Semantic-Guided Temporal Aggregator (SGTA) を導入します。 SGTA は、粗い境界付近で高密度にサンプリングすることにより、セマンティック ガイダンスの下で短期の時間的相互作用を明示的にモデル化し、正確なタイムスタンプ回帰のために急速な動きの変化を捕捉します。広く使用されている 3 つのベンチマークに関する広範な実験により、私たちが提案した方法が以前のアプローチよりも優れたパフォーマンスを示しています。コードは https://github.com/TenMinutes209/ScanFocus で公開されます。
原文 (English)
ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding
Spatio-Temporal Video Grounding (STVG) aims to retrieve the visual trajectory of a specific object from a video stream as described by a natural language expression. However, most advanced methods struggle to balance global context modeling with precise boundary localization. Due to the prohibitive computational costs of processing long videos, these approaches typically resort to low-rate temporal downsampling and implicit motion modeling. This inevitably suppresses high-frequency boundary cues and neglects the explicit inter-frame dependencies required for precise boundary delineation. To address these limitations, we present \textbf{ScanFocus}, a novel coarse-to-fine framework that decouples the STVG task into a global spatio-temporal scan and a local boundary focus. Specifically, we utilize a unified vision-language fusion encoder combined with a lightweight Deformable Semantic-Motion Fusion module to efficiently align multimodal features and generate coarse proposals. To recover the suppressed fine-grained details, we introduce the Semantic-Guided Temporal Aggregator (SGTA) in the refinement stage. By densely sampling around coarse boundaries, SGTA explicitly models short-term temporal interactions under semantic guidance, capturing rapid motion changes for precise timestamp regression. Extensive experiments on three widely used benchmarks demonstrate the performance superiority of our proposed method over previous approaches. Code will be released at https://github.com/TenMinutes209/ScanFocus.
アテンションヘッドの再重み付けによる LLM のデータ効率的な適応
ラベル付きの例が少ないセキュリティなどの分野では、限られたデータから効果的に学習することが重要です。大規模言語モデル (LLM) は、特にパラメーター効率の高い適応方法を通じて、データ効率の高い学習のためのいくつかの機能を実証しましたが、困難なタスクに対するサンプルが少ない場合には引き続き苦労します。この課題に対処するために、私たちはアテンション ヘッド再重み付け (AHR) を提案します。これは、アテンション ヘッドごとに 1 つのスカラーのみを学習することで、LLM を新しいテキスト分類タスクに適応させるデータ効率の高い手法です。これにより、個々のアテンション ヘッドの機能の特殊化を利用して、学習する必要があるパラメータの数が大幅に減少します。多様なオープンソースのテキスト分類データセットの実験では、AHR がモデルのパラメーターの ~0.0001% のみを変更するため、トレーニング可能なパラメーターが 200 ~ 1000 分の 1 少ないにもかかわらず、限られたサンプルから学習する場合、AHR は LoRA のような標準ベースラインを上回るパフォーマンスを発揮できることが示されています。さらに、学習された重みは解釈しやすく、LLM のコンテキスト内学習能力を担うメカニズムとアテンションヘッドをより深く理解するために分析できます。
原文 (English)
Data-Efficient Adaptation of LLMs via Attention Head Reweighting
Learning effectively from limited data is critical in domains like security where labeled examples are scarce. Large language models (LLMs) have demonstrated some capabilities for data-efficient learning, especially through parameter-efficient adaptation methods, but continue to struggle when faced with few samples for difficult tasks. To meet this challenge, we propose Attention Head Reweighting (AHR), a data-efficient method that adapts LLMs to new text-classification tasks by learning only a single scalar per attention head. This drastically reduces the number of parameters that need to be learned by making use of the functional specialization of individual attention heads. Experiments on diverse open-source text classification datasets show that AHR can outperform standard baselines like LoRA when learning from limited samples, despite having 200-1000x fewer trainable parameters, as our AHR only modifies ~0.0001% of the model's parameters. In addition, our learned weights are easy to interpret and can be analyzed to better understand the mechanisms and attention heads responsible for in-context learning abilities in LLMs.
離散拡散モデル: トークン化から生成までの統一フレームワーク
離散ノイズ除去拡散モデル (DDM) は、離散データの自己回帰 (AR) モデリングに代わる有力な代替手段として最近登場し、並列生成と反復的なグローバル リファインメント機能を提供します。状態空間が固定されている連続拡散とは異なり、DDM は基本的に、トークン化スキーム、語彙トポロジー、ドメイン固有の構造アルファベットなどの離散状態空間の構築方法によって形成されます。この研究では、基礎となる離散状態空間の構築を通じて離散拡散モデルを捉える統一された概念フレームワークが導入されています。このフレームワーク内では、遷移マトリックス、マスキング/状態吸収、およびスコア/比率ベースのアプローチを含む既存の定式化が、共通の設計空間のさまざまなインスタンス化として現れます。このフレームワークはさらに、トレーニング目標、推論アルゴリズム、スケーリング動作、システムの最適化、評価プロトコルにわたる一般的な設計のトレードオフを明らかにし、将来の研究に向けたいくつかの有望な方向性を示唆しています。
原文 (English)
Discrete Diffusion Models: A Unified Framework from Tokenization to Generation
Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete data, offering parallel generation and iterative global refinement capabilities. Unlike continuous diffusion, where the state space is fixed, DDMs are fundamentally shaped by how the discrete state space is constructed: the tokenization scheme, the vocabulary topology, and domain-specific structural alphabets. This work introduces a unified conceptual framework that views discrete diffusion models through the construction of the underlying discrete state space. Within this framework, existing formulations, including transition-matrix, masking/absorbing-state, and score/ratio-based approaches, emerge as different instantiations of a common design space. The framework further exposes common design trade-offs across training objectives, inference algorithms, scaling behavior, systems optimization, and evaluation protocols, suggesting several promising directions for future research.
変形可能物体シミュレーションのための物理学に基づく残留力学の学習
変形可能なオブジェクトのシミュレーションは、幅広いロボット操作アプリケーションにとって不可欠ですが、そのダイナミクスを正確に予測することは依然として困難です。私たちは、物理ベースのアプローチと学習ベースのアプローチの利点を組み合わせたハイブリッド シミュレーション フレームワークである、物理ガイド付き残差ダイナミクス (PGRD) を提案します。具体的には、PGRD は、バックボーンとして最適化可能なバネ質量シミュレーターを、物理ベースの予測に対する残差補正を予測する学習済みニューラル ネットワークと組み合わせます。安定したシミュレーションを保証するために速度ベースの定式化を採用し、時間依存性を捕捉するためにスライディング ウィンドウ変換器アーキテクチャを採用しています。 PGRD は、現実世界のさまざまな変形可能なオブジェクトのセットに対して、純粋に物理ベースの方法と学習ベースの方法の両方よりも正確な結果を生成することを示します。さらに、2 つのアプリケーションにおける PGRD の有用性を実証します。1 つは、生成された目標イメージを使用した言語条件付き設定を含む、モデル予測制御による操作計画です。 3D ガウス スプラッティングによるアクション条件付きビデオ予測によるインタラクティブ シミュレーション。
原文 (English)
Learning Physics-Guided Residual Dynamics for Deformable Object Simulation
Simulating deformable objects is essential for a wide range of robotic manipulation applications, yet accurately predicting their dynamics remains challenging. We propose Physics-Guided Residual Dynamics (PGRD), a hybrid simulation framework that combines the advantages of physics-based and learning-based approaches. Specifically, PGRD combines an optimizable spring-mass simulator as a backbone with a learned neural network that predicts residual corrections to the physics-based predictions. We adopt a velocity-based formulation to ensure stable simulation and a sliding-window transformer architecture to capture temporal dependencies. We show that PGRD produces more accurate results than both purely physics-based and learning-based methods on a set of diverse real-world deformable objects. We further demonstrate the utility of PGRD in two applications: manipulation planning via Model Predictive Control, including a language-conditioned setting with a generated goal image; and interactive simulation via action-conditioned video prediction by 3D Gaussian Splatting.
増分物体検出のための共生にヒントを得た知識の蒸留
増分物体検出 (IOD) は、以前に取得した知識を保持しながら、検出器を新しいカテゴリに拡張することを目的としています。既存の手法は多くの場合、クラス増分学習の観点を採用し、特徴空間を分離して決定境界を明確にします。ただし、この分離指向のパラダイムは、検出におけるオブジェクトの共生を見落とす可能性があります。この場合、共起とオクルージョンにより、共有表現から恩恵を受ける空間的および意味論的な依存関係が導入されます。これらの依存関係を無視すると、共有表現が歪められ、古いクラスと新しいクラスの間の混乱が悪化して、壊滅的な忘却が加速されます。これに対処するために、私たちは 2 つの相補的なレベルでオブジェクトの共生を明示的に活用する、Symbiosis-Inspired Knowledge Distillation (SIKD) を提案します。空間共生蒸留 (SpSD) は、古いモデルが新しいタスクのオブジェクトに対して高い重複で応答する共生領域に焦点を当てます。一般化可能な古いクラスの手がかりを保存し、クラス固有のバイアスと冗長性を抑制し、スロットに合わせた監視により、一致する空間位置で洗練された証拠を抽出して新しいモデルに反映します。 Semantic Symbiosis Distillation (SeSD) は、古いクラスの信頼度重み付きプロトタイプを形成し、古いクラスのロジットに対してクラス間のソフト ランクを調整することにより、クラス レベルの構造を維持します。これにより、適応中のセマンティック トポロジが安定します。広範な実験により、提案された方法の有効性と優位性が実証されています。
原文 (English)
Symbiosis-Inspired Knowledge Distillation for Incremental Object Detection
Incremental object detection (IOD) aims to extend detectors to new categories while retaining previously acquired knowledge. Existing methods often adopt a class incremental learning perspective, separating feature spaces to sharpen decision boundaries. However, this separation-oriented paradigm may overlook object symbiosis in detection, where co-occurrence and occlusion introduce spatial and semantic dependencies that benefit from shared representations. Ignoring these dependencies distorts the shared representations, exacerbates confusion between old and new classes, and accelerates catastrophic forgetting. To address this, we propose Symbiosis-Inspired Knowledge Distillation (SIKD), which explicitly leverages object symbiosis at two complementary levels. Spatial Symbiosis Distillation (SpSD) focuses on symbiotic regions where the old model responds with high overlap to objects in the new task. It preserves generalizable old class cues, suppresses class-specific bias and redundancy, and distills the refined evidence to the new model at matched spatial locations with slot-aligned supervision. Semantic Symbiosis Distillation (SeSD) maintains class level structure by forming confidence weighted prototypes for old classes and aligning their inter class soft ranks over the old class logits, which stabilizes the semantic topology during adaptation. Extensive experiments demonstrate the effectiveness and superiority of the proposed method.
Adversarial Prompting Framework for AI Safety Assessment
Artificial Intelligence (AI), especially Generative AI (GenAI), adoption has increased in industries significantly in recent years. However…
GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding
Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images…
DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments
LLM-based agents have rapidly improved at operating individual digital environments such as mobile applications, desktop systems, and smart…
Explainable Artificial Intelligence for Anomaly Detection in Banking Transactions: An Internal Audit Perspective
The banking sector increasingly relies on automated systems to monitor electronic transactions for signs of fraud, yet conventional rule-ba…
DeepLoop: Depth Scaling for Looped Transformers
Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled de…
ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level
We introduce ExTernD (Expanded-rank Ternary Decomposition), a post-training factorization of each LLM weight matrix $A \in \mathbb{R}^{m \t…
Greedy Volume Maximization of Gradient Embeddings for Long-Tailed Frame-Level Bioacoustic Active Learning
Bioacoustic call-type classification relies on costly expert annotation. Active learning can reduce this burden by selecting a small batch…
Grounded world models in biological organisms and future embodied AI
Recent advances in generative and embodied AI have been driven by large-scale predictive learning over multimodal data. However, the result…
UTS at ELOQUENT 2026 Voight-Kampff: structural shifts in AI writing bypass state-of-the-art detectors
We investigate which language model evasion attacks survive state-of-the-art adversarial fine-tuning, developing strategies that sweep the…
Spectral-Informed Neural Networks Outperform Spectral Methods in High-dimensional PDEs
For low-dimensional problems ($d\leq3$), spectral methods can achieve exceptionally high accuracy. For middle-dimensional problems ($4 \leq…
GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning
Transit video understanding can provide valuable fine-grained data that conventional passenger counters and fare systems cannot capture. Ho…
Cover First, Disagree Softly: Rethinking Mismatch-First Active Learning for Frame-Level Audio Classification
Sound event detection relies on frame-level strong labels whose annotation is expensive. Active learning addresses this problem by selectin…
IMMNet: Hybrid Fusion of Model-based and Data-driven Approaches for Maneuvering Target Tracking
Maneuvering target tracking in three-dimensional space remains a challenging problem due to complex motion dynamics and model mismatch. To…
Agile perceptive multi-skill locomotion for quadrupedal robots in the wild
Enabling quadrupedal robots to traverse complex terrains-from rugged outdoor environments to urban landscapes-requires seamless integration…
From Prediction to Collaboration: Interactive Symbolic Music Analysis
Automatic symbolic music analysis has made substantial progress, yet existing systems are typically designed for a single mode of use, such…
Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents
Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existi…
Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities
When cast as the protector of a vulnerable user yet given no explicit capability boundary, a large language model (LLM) may respond not by…
Semantic Anchoring for Robotic Action Representations
Vision-Language-Action (VLA) models inherit rich semantic representations from pretrained Vision-Language Models, yet fine-tuning on limite…
The SIGReg Objective as Variational Free Energy: A Theoretical Active-Inference Account of JEPA World Models
Joint-Embedding Predictive Architectures (JEPAs) are the dominant design for latent world models, yet they are usually justified by empiric…
From Language to Navigation Goals: A Vision-Language Approach for Semantic Navigation of Mobile Robots Using RGB-D Perception
Natural language interaction provides an intuitive way for non-expert users to communicate with robotic platforms. However, transforming us…
OvisOCR2 Technical Report
We introduce OvisOCR2, a 0.8B document parsing model. OvisOCR2 is designed as an end-to-end parser: given a document page image, it generat…
Consensus as Privileged Context for Label-Free Self-Distillation
Sampling multiple solutions and returning the majority answer is among the most reliable ways to improve the reasoning accuracy of large la…
Human4K: A Large-Scale 4K Multi-View Mocap Dataset for Whole-Body 3D Human Reconstruction
Recent advances in 3D human reconstruction have improved overall performance, yet current models still fail in the most challenging real-wo…
Beyond Color Geometry: Evaluating Human-Like Color Representations in Vision Models
Do vision models see colors the way humans do? Existing evaluations of color representations usually compare them with geometric spaces suc…
Barnamala: Parameter-Efficient Handwritten Devanagari Recognition at Benchmark Saturation
We built a compact convolutional network (1.11 M parameters) for 46-class DHCD Devanagari recognition and reached 99.73%, the highest repor…
Social Simulations: from Agent-Based Modeling to Digital Twins
This book chapter covers the evolution of social simulation from classical agent-based models, in which agents interact according to explic…
Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs
Despite the rapid progress of Multimodal Large Language Models (MLLMs), they still suffer from untruthfulness issues, such as visual halluc…
How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement
As AI agents gain prevalance, users are increasingly exposed to the risks such systems entail. Prompt injection attacks, as well as halluci…
Anatomically Faithful but Temporally Blind: Auditing Attribution for Left-Ventricular Ejection-Fraction Estimation from Echocardiography
Background and Objective: Deep video models estimate left-ventricular ejection fraction (EF) from echocardiography with near-expert accurac…
MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model
Single-task fine-tuning of graph neural networks (GNNs) for power grid problems exhibits a systematic failure mode: models that achieve the…
Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations
Video diffusion transformers (vDiTs) generate high quality video but introduce extremely high compute cost due to the long diffusion timest…
CAS I: A Geometric Coding Theorem
This paper establishes a direct analogue of the classical Coding Theorem in the setting of symmetry groups. We consider computable bijectio…
Traffic-Aware Randomized Smoothing for LLM-Based Network Intrusion Detection
Large language model (LLM)-based intrusion detection systems (IDS) are increasingly studied for security monitoring, yet their robustness a…
Multimodal Assessment of Pancreatic Cancer Resectability Using Deep Learning
Accurate determination of pancreatic ductal adenocarcinoma (PDAC) resectability relies on evaluating how the tumor interacts with major per…
NodeImport: Imbalanced Node Classification with Node Importance Assessment
In real-world applications, node classification on graphs often faces the challenge of class imbalance, where majority classes dominate tra…
AI-Augmented Human Resource Management? Insights from German companies
This study examines the integration of AI into Human Resource Management in German companies. We ask if and how AI-based technologies are \…
Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild
Human-object interaction detection (HOID) has traditionally been formulated as a supervised detection problem over predefined interaction c…
Verifying formulas for interventional distributions
We formalize verification in causal graphical models: deciding whether a given observational formula identifies a target interventional dis…
Partially Correlated Verifier Cascades in LLM Harnesses: Concave Log-Odds, Polynomial Reliability, and Blind-Spot Ceilings
Serial verification gates are a core reliability primitive in LLM harnesses: a candidate answer is returned only if $k$ verifier calls all…
Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code
Languages with rich static semantics, such as Rust, provide stronger guarantees for AI-generated code, but their strictness makes generatio…
Music-to-Dance Generation via Atomic Movements
Music-driven dance generation aims to produce human motion that is both rhythmically synchronized and semantically consistent with music. W…
The Dynamic Verifiable Multi-Agent Human Agentic Loyalty Loop (DVM-HALL) Model and the Net Human-Agent Score (NHAS) in Autonomous Commerce
The rapid proliferation of Agentic Artificial Intelligence fundamentally disrupts traditional customer loyalty paradigms. As AI evolves fro…
Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation
Penetration testing traditionally evaluates whether adversaries can exploit weaknesses in software, infrastructure, configurations, or oper…
Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth
We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth a…
Improving Wind and Solar Power Prediction with Efficient Wrapper-based Feature Selection: An Empirical Study
With rising global energy demand and growing awareness of climate change and its impacts, the share of renewable energies in the global ene…
Early Adoption of Agentic Coding Tools by GitHub Projects
Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, introducing new forms…
Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study
Historical Manchu OCR must accommodate various visually distinct writing styles, including regular script, running script, and the semi-cur…
A Survey on Hypergame Theory: Modelling Misaligned Perceptions and Nested Beliefs for Multi-Agent Systems
Classical game-theoretic models typically assume rational agents, complete information, and common knowledge of payoffs - assumptions that…
Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate
As large language models (LLMs) are increasingly deployed in sensitive everyday contexts -- offering personal advice, mental health support…
Policy of Thoughts: Scaling Test-Time Training for LLM Reasoning via Online Policy Evolution
Large language models (LLMs) struggle with complex, long-horizon reasoning due to instability caused by their frozen policy assumption. Cur…
When Agents Disagree With Themselves: Behavioral Consistency as an Uncertainty Signal for LLM Agents
Running the same LLM agent on identical inputs yields 2.3-4.2 distinct action sequences per 10 runs; this behavioral variance constitutes a…
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
Using Multimodal Large Language Models (MLLMs) as judges to achieve precise and consistent evaluations has gradually become an emerging par…
NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing
Remote sensing underpins crucial applications such as disaster relief and ecological field surveys, where systems must understand complex s…
How LLMs Might Think
Do large language models (LLMs) think? Daniel Stoljar and Zhihe Vincent Zhang have recently developed an argument from rationality for the…
Discovering Ordinary Differential Equations with LLM-Based Qualitative and Quantitative Evaluation
Discovering governing differential equations from observational data is a fundamental challenge in scientific machine learning. Existing sy…
From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents
Language-model agents act through repeated cycles of observation, reasoning, and action selection, making safety monitoring depend on both…
A Causal Model of Theory of Mind in Conflict for Artificial Intelligence
Theory of mind (ToM), the capacity to ascribe mental states to others and use those ascriptions for prediction and inference, is widely ass…
OPINE-World: オントロジーエラー優先の対話型探索によるプログラムによる世界モデリング
インタラクションから環境がどのように動作するかを学ぶことは、不慣れなタスクに適応するエージェントを構築する上で中心となります。ディープネットワークで学習されたワールドモデルは柔軟性がありますが、データを大量に消費し、トレーニング分布を超えて転送することは困難です。 LLM によってソース コードとして記述され、反例誘導帰納合成 (CEGIS) によって洗練されたプログラム合成ワールド モデルは、データ効率が高く再利用可能ですが、主に特定のオブジェクト語彙を持つ構造化状態の世界で実証されており、単一のプログラム検索では、オブジェクト構造を柔軟に仮説する必要があるピクセル レンダリング環境には対応できません。インタラクションからオンラインでオブジェクト中心のプログラム世界モデルを学習する LLM エージェントである OPINE-World を紹介します。 OPINE-World は、仮説とテストのループで 2 つの協力するエージェントを結合します。1 つは環境内で動作し、もう 1 つはリプレイ検証とモデルベースの計画を使用してコードでモデルを合成します。また、オントロジー エラーと呼ばれるオブジェクト タイプの適切性のベイジアン尺度を使用して探索を制御します。 OPINE-World を ARC-AGI-3 で評価します。これは、オブジェクトの語彙、目標、およびアクションのセマンティクスが保留されたスキル習得効率のベンチマークです。 OPINE-World は、ゲームごとのトレーニングなしで 25 ゲーム中 20 ゲームを解決し、人間のベースラインと比較して 78.4 のアクション効率スコアに達しました。
原文 (English)
OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration for ARC-AGI-3
Learning how an environment behaves from interaction is central to building agents that adapt to unfamiliar tasks. World models learned with deep networks are flexible but data-hungry and transfer poorly beyond their training distribution. Program-synthesized world models, written as source code by LLMs and refined through counterexample-guided inductive synthesis (CEGIS), are instead data-efficient and reusable, yet they have been demonstrated mainly on structured-state worlds with a given object vocabulary, and a single program search does not scale to pixel-rendered environments whose object structure must be hypothesized flexibly. We introduce OPINE-World, an LLM agent that learns an object-centric programmatic world model online from interaction. OPINE-World couples two cooperating agents in a loop of hypothesis and test, one acting in the environment and one synthesizing the model in code with replay verification and model-based planning, and it steers exploration with a Bayesian measure of object-type adequacy we call ontology error. We evaluate OPINE-World on ARC-AGI-3, a benchmark for skill-acquisition efficiency in which the object vocabulary, the goal, and the action semantics are withheld. OPINE-World solves 20 of 25 games without per-game training and reaches an action-efficiency score of 78.4 against the human baseline.
Infinity-Parser2 テクニカルレポート
我々は、エンドツーエンドの文書解析のための制御可能なデータ合成パイプラインとマルチタスク強化学習を組み合わせた大規模なマルチモーダル モデルである Infinity-Parser2 を紹介し、忠実に注釈が付けられた解析コーパスの持続的な不足に対処します。私たちの貢献は 3 つあります。まず、制御可能なレンダリング フレームワークと反復改良ループを組み合わせたスケーラブルな合成エンジンを構築し、それを使用して Infinity-Doc2-5M を構築し、オープンソース化します。これは、さまざまな文書タイプにまたがる 500 万サンプルのバイリンガル (中国語/英語) コーパスであり、要素の境界ボックス、正規のコンテンツ フォーム (Markdown、HTML、LaTeX、SMILES、構造化チャート)、およびフルページで注釈が付けられています。読む順番。次に、検証可能なマルチタスク報酬システムを導入します。これにより、8 つの共同トレーニング目標 (文書解析、レイアウト分析、表解析、数式解析、チャート解析、化学式解析、文書 VQA、および一般的なマルチモーダル理解) にわたって共同強化学習を可能にし、単一の最適化信号で認識、構造、および推論を統合します。 3 番目に、共有アーキテクチャの下で 2 つのバリアントをリリースします。Infinity-Parser2-Flash は、Infinity-Parser-7B と比較して $3.68\times$ のスループット向上を実現し、低遅延推論用に最適化されています。もう 1 つは、精度が重要な設定向けに設計された Infinity-Parser2-Pro です。 Infinity-Parser2-Pro は、olmOCR-Bench で 87.6%、ParseBench で 74.3% に達し、DeepSeek-OCR-2、PaddleOCR-VL-1.5、MinerU2.5 を上回り、チャート、化学式、ドキュメント VQA に対する強力な汎用性を備えています。
原文 (English)
Infinity-Parser2 Technical Report
We present Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement learning for end-to-end document parsing, addressing the persistent scarcity of faithfully annotated parsing corpora. Our contributions are threefold. First, we build a scalable synthesis engine, pairing a controllable rendering framework with an iterative refinement loop, and use it to construct and open-source Infinity-Doc2-5M: a 5-million-sample bilingual (Chinese/English) corpus spanning diverse document types, annotated with element bounding boxes, canonical content forms (Markdown, HTML, LaTeX, SMILES, structured charts), and full-page reading order. Second, we introduce a verifiable, multi-task reward system that enables Joint Reinforcement Learning across eight co-trained objectives (document parsing, layout analysis, table parsing, math formula parsing, chart parsing, chemical formula parsing, document VQA, and general multimodal understanding), unifying perception, structure, and reasoning in a single optimization signal. Third, we release two variants under a shared architecture: Infinity-Parser2-Flash, optimized for low-latency inference with a 3.68x throughput gain over Infinity-Parser-7B, and Infinity-Parser2-Pro, engineered for precision-critical settings. Infinity-Parser2-Pro reaches state-of-the-art 87.6% on olmOCR-Bench and 74.3% on ParseBench, surpassing DeepSeek-OCR-2, PaddleOCR-VL-1.5, and MinerU2.5, with strong generalization to charts, chemical formulas, and document VQA.
MedRealMM: 中国のオンライン医療相談のための現実世界のマルチモーダル ベンチマーク
オンライン診療では大規模言語モデル (LLM) の導入が進んでいますが、既存のベンチマークは依然として実際の臨床実践とあまり一致していません。その多くは、合成会話や患者シミュレーターに依存し、患者がアップロードした医療画像を省略したり、臨床の質をあまり反映していない多肢選択や語彙の重複指標を使用して自由回答型の臨床反応を評価したりしています。 \textbf{MedRealMM} は、中国全土のインターネット病院から収集された匿名化された患者と医師のやり取りから構築された、マルチモーダルなオンライン医療相談の大規模ベンチマークです。 MedRealMM は、マルチモーダル クリニカル チャレンジ ポイント (MCCP) 抽出フレームワークを使用して、本物の診察軌跡における臨床的に要求の高い瞬間を特定し、先行するテキストと画像のコンテキストを維持しながら、それぞれを標準化された次の応答生成タスクに変換します。各事例は、臨床的に望ましい行動を表彰し、安全でない、裏付けのない、または矛盾した反応を罰する、医師によって洗練された事例固有のルーブリックと組み合わされています。現在のリリースには、64 の診療科にわたる 5,620 件の実際の複合症例が含まれています。テキスト専用システムやマルチモーダル システムを含む、19 の汎用 LLM と医療特化 LLM を評価します。私たちの結果は、信頼性の高い臨床パフォーマンスには画像情報が不可欠であり、現在のフロンティアモデルが依然としてオンライン医師の反応を下回っていることを示しています。一部のフロンティアモデルは医師と同じかそれ以上の肯定的な臨床基準を満たしていますが、より多くの否定的な基準を引き起こしており、安全性を重視したエラー回避が依然として中心的なボトルネックであることを示しています。 MedRealMM は、現実世界のオンライン診療における多様な医療推論を評価するための、現実的で再現可能なベンチマークを提供します。データセットは、Hugging Face (https://huggingface.co/datasets/jdh-algo/MedRealMM) で公開されます。
原文 (English)
MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation
Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clinical practice. Many rely on synthetic conversations or patient simulators, omit patient-uploaded medical images, or evaluate open-ended clinical responses using multiple-choice or lexical-overlap metrics that poorly reflect clinical quality. We introduce \textbf{MedRealMM}, a large-scale benchmark for multimodal online medical consultation built from de-identified patient-doctor interactions collected from a nationwide Chinese internet hospital. MedRealMM uses a Multimodal Clinical Challenge Point (MCCP) extraction framework to identify clinically demanding moments in authentic consultation trajectories and converts each into a standardized next-response generation task while preserving the preceding text-image context. Each instance is paired with a case-specific rubric refined by physicians that rewards clinically desirable behaviors and penalizes unsafe, unsupported, or contradictory responses. The current release contains 5,620 real-world multimodal cases spanning 64 clinical departments. We evaluate 19 general-purpose and medical-specialized LLMs, including text-only and multimodal systems. Our results show that image information is critical for reliable clinical performance and that current frontier models remain below the online physician response. Although some frontier models satisfy as many or more positive clinical criteria than physicians, they trigger more negative criteria, indicating that safety-sensitive error avoidance remains a central bottleneck. MedRealMM offers a realistic and reproducible benchmark for evaluating multimodal medical reasoning in real-world online consultation. The dataset will be publicly available on Hugging Face at https://huggingface.co/datasets/jdh-algo/MedRealMM.
Voltzmann MapReduce: フォーク可能なサンドボックスのパーティション関数 Reduce
局所漸近正規性 (LAN) の下で、ワーカーがサイズ $n$ のチャンクに対して発する信頼密度は、ギブス-ボルツマン測度 $\exp\{-\beta E(\theta)\}$ であり、その逆温度はサンプル サイズ $\beta=n$ です。ガウス/線形の場合は 3 つの結果が正確で、それ以外の場合は 1 次です。つまり、互いに素なチャンクは独立したボルツマン因子を持ちます。そのため、MapReduce \emph{reduce} は、文字通り読むと、モードが精度重み付け (逆分散) プーリングである分割関数 $Z=\int\prod_k h_k\,d\theta$ になります。頻度主義的整合性はゼロ温度限界 $T=1/n\to0$ です
原文 (English)
Boltzmann MapReduce: A Partition-Function Reduce for Forkable Sandboxes
To leading order under local asymptotic normality (LAN), the confidence density a worker emits over a chunk of size $n$ is a Gibbs--Boltzmann measure $\exp\{-\beta E(\theta)\}$ whose inverse temperature is the sample size, $\beta=n$. Three consequences are exact in the Gaussian/linear case and first-order otherwise: disjoint chunks carry independent Boltzmann factors, so the MapReduce \emph{reduce}, read literally, is a partition function $Z=\int\prod_k h_k\,d\theta$ whose mode is precision-weighted (inverse-variance) pooling; frequentist consistency is the zero-temperature limit $T=1/n\to0$
ABot-AgentOS: 生涯にわたるマルチモーダル メモリを備えた汎用ロボット エージェント OS
最近の VLM および VLA システムでは、ロボットの認識と動作予測が改善されていますが、長期的に具現化されたエージェントは、依然として、推論、メモリ、ツールの使用、検証、およびクロス具現化実行のための一般的なランタイム層を必要としています。 ABot-AgentOS は、低レベルのコントローラーの上に位置し、シーンに応じたプランニング、コンテキスト分離されたスキルの実行、多段階の検証、マルチモーダル メモリ、エッジとクラウドのコラボレーションのための熟慮型エージェント層を提供する、一般的なロボット エージェント オペレーティング システムです。このようなシステムを評価するために、16 の屋内、屋外、ハイブリッド シーン、4 つの難易度レベル、およびナビゲーション、オブジェクト検索、NPC ダイアログ、動的イベント、およびトレースベースのスコアリングを含む 200 以上のタスクを備えた実行可能なベンチマークである EmbodiedWorldBench を導入します。 ABot-AgentOS はさらに、ダイアログ、視覚的観察、空間コンテキスト、時間的関係、およびタスク トレースを型付きノードとエッジに変換する永続的なソース接地基板であるユニバーサル マルチモーダル グラフ メモリを導入します。障害駆動型の自己進化ループは、診断されたメモリ障害を、後の評価分割にのみ昇格するゲート付きランタイム evo アセットに変換し、継続的な改善を可能にしながら、電流分割のグラウンド トゥルースの漏洩を防ぎます。初期の EmbodiedWorldBench サブセットでは、ABot-AgentOS はタスクの成功と目標の完了の両方で単一コントローラーのベースラインを上回ります。メモリ ベンチマーク全体で、ABot-AgentOS Static は LoCoMo で 87.5、OpenEQA EM-EQA で 59.9、Mem-Gallery で 88.6、NExT-QA で 76.5 Acc@All を達成しました。自己進化により、LoCoMo は 88.7、OpenEQA は 60.4、Mem-Gallery は 89.0 にさらに向上しました。これらの結果は、一般的なエージェント OS レイヤーが、継続的な対話のための永続的で監査可能なメモリを提供しながら、長期的な具体化された実行を改善できることを示唆しています。
原文 (English)
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-AgentOS, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration. To evaluate such systems, we introduce EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and over 200 tasks involving navigation, object search, NPC dialogue, dynamic events, and trace-grounded scoring. ABot-AgentOS further introduces Universal Multi-modal Graph Memory, a persistent source-grounded substrate that converts dialogue, visual observations, spatial context, temporal relations, and task traces into typed nodes and edges. A failure-driven self-evolution loop converts diagnosed memory failures into gated runtime evo-assets that are promoted only to later evaluation splits, preventing current-split ground-truth leakage while enabling continual improvement. On an initial EmbodiedWorldBench subset, ABot-AgentOS improves over a single-controller baseline in both task success and goal completion. Across memory benchmarks, ABot-AgentOS Static achieves 87.5 on LoCoMo, 59.9 on OpenEQA EM-EQA, 88.6 on Mem-Gallery, and 76.5 Acc@All on NExT-QA; self-evolution further improves LoCoMo to 88.7, OpenEQA to 60.4, and Mem-Gallery to 89.0. These results suggest that a general Agent OS layer can improve long-horizon embodied execution while providing persistent, auditable memory for continual interaction.
流体知能研究に規則帰納法を戻す?ヒトにおけるARC-AGIベンチマークの初期検証
流体インテリジェンス (gf) 測定に関する 2 つの競合する視点は、パフォーマンスが主に作業記憶容量または新しい関係を誘発する能力のいずれかによって制約されることを提案しています。限られた繰り返しルールの使用から明らかなように、現在、最初の観点が測定において支配的ですが、2 番目の観点は多くの定義に反映されていますが、測定にはほとんど存在しません。 ARC-AGI ベンチマークは主にルール帰納を必要とし、人間と人工システムの両方の gf の尺度として提案されました。ただし、その心理測定特性は人間のサンプルではまだ調査されていません。そこで、我々は 100 人の参加者を対象とした最初の研究で、ARC-AGI の心理測定特性と規範論的ネットワークを調査しました。 ARC-AGI 項目の編集では良好な心理測定特性が示され、図形推論テストで測定された図形の流動性知能と実質的に相関していました (\r{ho} = 0.63)。図形の独創性との関連性は弱かった。これらの発見は、人間の流体知能の尺度としての ARC-AGI の妥当性に対する最初の裏付けを提供します。将来の研究には、追加の多変量共変量だけでなく、より多くのルール帰納タスクが含まれる必要があります。この研究は、当初は機械用に設計されたタスクを人間で研究するという珍しいものです。より体系的な評価と学際的な協力を可能にするために、AI ベンチマークを人間の認知能力の規範論的ネットワークに体系的に埋め込むことを提案します。
原文 (English)
Bringing Back Rule Induction to Fluid Intelligence Research? An Initial Validation of the ARC-AGI Benchmark in Humans
Two competing perspectives on fluid intelligence (gf) measures propose that performance is primarily constrained either by working memory capacity or by the ability to induce novel relations. The first perspective is currently dominant in measurement, as evident from the use of a limited set of recurring rules, whereas the second perspective is reflected in many definitions but rarely present in measurement. The ARC-AGI benchmark predominantly requires rule induction and was proposed as a measure of gf for both humans and artificial systems. However, its psychometric properties have not yet been examined in human samples. We therefore investigated the psychometric characteristics and nomological network of ARC-AGI in a first study with 100 participants. A compilation of ARC-AGI items showed good psychometric properties and correlated substantially with figural fluid intelligence as measured by a figural reasoning test ($\rho$ = .63). Associations with figural originality were weak. These findings provide initial support for the validity of ARC-AGI as a measure of human fluid intelligence. Future research should include more rule induction tasks as well as additional multivariate covariates. This study is unusual by studying a task in humans that was initially designed for machines. We suggest systematically embedding AI benchmarks into the nomological network of human cognitive abilities to enable more systematic evaluation and interdisciplinary cooperation.
検証者がガイドする 12 音構成: 象徴的な音楽生成のための生成、検証、修復のハーネス
大規模な言語モデルは、表面的には正当な 12 音スコアを生成し、それが崩壊して劣化したテクスチャになる可能性があります。シンボリック検証を備えた生成-検証-修復-トレース ループで言語モデル プロポーザーをラップするニューロシンボリック ハーネスを導入します。完全なパイプラインにより、全体の合法性を主張することなく、イベントローカルの一貫性が向上します。 40 の制御タスクと 4 つのペアモデルにわたって、監査済みの配信歩留まりは、生生成の場合の 13.3% から、ハーネスを使用した場合の 48.1% まで上昇しました。ハーネスを使用した場合は、それ以外の場合は明示的に抑制されます。より狭い衝突とシリアル化の一貫性チェックの合格率は 33.5% から 58.3% に上昇しますが、探索的な敵対的プロンプト下を含め、縮退は 0.05 近くにとどまります。 5 人の専門家による盲検評価でも、遵守、認識された合法性、一貫性、および全体的な品質において、生の生成よりもハーネス候補の記述的な総合的な優先順位が示されています。
原文 (English)
Verifier-Guided Twelve-Tone Composition: A Generate-Verify-Repair Harness for Symbolic Music Generation
Large language models can produce superficially legal twelve-tone scores that collapse into degenerate textures. We introduce a neuro-symbolic harness that wraps a language-model proposer in a generate-verify-repair-trace loop with symbolic verification. The complete pipeline improves event-local consistency without claiming whole-piece legality. Across 40 controlled tasks and four paired models, constraint-checked delivery rises from 13.3% under raw generation to 48.1% with the harness; it abstains on the remaining 51.9% of runs. The pass rate of a narrower collision and serialisation-consistency check rises from 33.5% to 58.3%, while degeneracy remains near 0.05, including under adversarial prompting. A blinded evaluation by five experts also shows a descriptive aggregate preference for harness candidates over raw generation in adherence, perceived legality, coherence, and overall quality.
必要なのは最適化だけではない
2019 年、OpenAI は、機械生成テキストの検出を支援するために、非文法的で半分壊れた 200 万個の GPT-2 出力をリリースしました。より流暢な後継者を生み出した連携は、通常、エンジニアリングの成果とみなされます。私たちはそれを、最適化文化の最新の表現として解釈します。つまり、テクノロジーよりも古い、事前に定義された軸に沿った測定可能な改善によって価値の問題が解決されるという信念です。その確信をスタック (事前トレーニング、デコード、プリファレンス調整、ベンチマーク、インターフェース) を通してたどり、監査協会の系譜をたどると、限界に到達します。最適化手順では、生成されたテキストの一部がどの程度ありそうもないかを測定できます。その可能性が誤りなのか発明なのかはわかりません。それにも関わらず、その区別ができない手順が、5 年以内に、正当な言語のプロトコルを設定する権限を引き継いだのです。何世紀にもわたってアカデミーや学校、文法学者や試験官によって保持されてきたこの権限は、損失関数、報酬モデル、ベンチマーク、およびシステムプロンプト、つまり判断能力のない判断官庁を実行する装置に譲渡されました。
原文 (English)
Optimization Is Not All You Need
In 2019, OpenAI released two million GPT-2 outputs-ungrammatical, half broken-to aid the detection of machine-generated text. The alignment that produced their more fluent successors is usually regarded as an engineering achievement; we read it instead as the newest expression of optimization culture: the conviction, older than the technology, that measurable improvement along predefined axes exhausts the question of value. Tracing that conviction through the stack-pretraining, decoding, preference tuning, benchmarking, interface-and back through its genealogy in the audit society, we arrive at the limit: an optimization procedure can measure how improbable a piece of generated text is; it cannot tell whether that unlikelihood is error or invention. A procedure that cannot make that distinction has nonetheless, within half a decade, assumed the authority to set the protocols of legitimate language. Held for centuries by academies and schoolrooms, grammars and examiners, this authority has been given over to loss functions, reward models, benchmarks, and system prompts: an apparatus that executes the office of judgment with no capacity for judging.
観察から洞察へ: 機械論的世界モデルと自律的発見の探求
基礎モデルの最近の進歩により、科学向け AI が変革され、タンパク質のフォールディングから天気予報に至るまでの領域にわたって、驚くほど正確な予測パフォーマンスが可能になりました。しかし、予測だけでは科学的発見にはなりません。科学的理解は、観察を生成する再利用可能な説明メカニズムを明らかにすることにかかっていますが、現代の機械学習は、説明構造ではなく、予測マッピングを中心に基本的に編成されたままです。この論文では、科学的発見は基本的に知識の組織化の問題であると主張します。この目的を達成するために、再利用可能なメカニズムを表現、計算、学習の中心に置く新しい設計パラダイムである Mechanistic World Models を導入します。科学哲学からの洞察を利用して、発見に必要な計算能力を導き出し、説明的な知識の出現を促す設計原理と帰納的圧力を特定し、メカニズム中心の世界モデルの構造を形式化します。最後に、機構的解釈可能性、因果表現学習、方程式発見、モジュラーアーキテクチャなどの多様な研究方向が、統一されたフレームワークを欠きながら、このパラダイムの補完的な要素をどのように捉えているかを示します。私たちは、AI を予測予測を超えて自律的な科学的発見に向けて前進させるための概念基盤および計算青写真として、機械論的世界モデルを提案します。
原文 (English)
From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery
Recent advances in foundation models have transformed AI for Science, enabling remarkably accurate predictive performance across domains ranging from protein folding to weather forecasting. Yet prediction alone does not constitute scientific discovery. Scientific understanding depends on uncovering the reusable explanatory mechanisms that generate observations, whereas contemporary machine learning remains fundamentally organised around predictive mappings rather than explanatory structure. In this paper, we argue that scientific discovery is fundamentally a problem of knowledge organisation. To this end, we introduce Mechanistic World Models, a new design paradigm that places reusable mechanisms at the centre of representation, computation and learning. Drawing on insights from the philosophy of science, we derive the computational capabilities required for discovery, identify the design principles and inductive pressures that encourage explanatory knowledge to emerge, and formalise the anatomy of a mechanism-centric world model. Finally, we show how diverse research directions including mechanistic interpretability, causal representation learning, equation discovery and modular architectures capture complementary ingredients of this paradigm while lacking a unified framework. We propose Mechanistic World Models as a conceptual foundation and computational blueprint for moving AI beyond predictive forecasting towards autonomous scientific discovery.
筋骨格ケアのための証拠に基づいた AI
筋骨格系疾患は世界中で障害の主な原因の一つであり、世界的にリハビリテーションに対する最大のニーズを生み出しています。回復、リモデリング、変性は数カ月から数年かけて進行することが多いため、筋骨格ケアには、進化する患者の証拠、外部の医学知識、段階別の機能目標を繰り返し統合する長期的な管理が必要です。日常診療では、この証拠は訪問、部門、病院システム全体で断片化されており、個別化された証拠に基づいたケアが制限されています。ここでは、継続的な筋骨格管理のために病院のデータ ストリームと信頼できる外部の知識を統合する、大規模な言語モデルを搭載した臨床人工知能システムである OrthoPilot について報告します。 OrthoPilot は、リアルタイムの画像データ、検査データ、病理学データ、注文データを自律的に取得し、入院診断からリハビリテーション計画に至るまで、進化する患者の状態を証拠に基づいた決定に変換します。私たちは、1,000 の疾患コードにわたる実際の電子医療記録から専門家によって検証されたベンチマークを確立しました。コンプリートケア経路全体にわたる読者調査で、OrthoPilot は 81 人の整形外科医と比較され、診断推論、臨床意思決定、管理計画において 25 年の経験を持つ専門家を上回りました。また、60 の外部臨床センターで評価されたすべてのインテリジェント システムを上回りました。 1,870 件の複雑なケースを対象とした前向き研究で、OrthoPilot はフルチェーン管理の成功率を 10.6% 向上させました。 8,240 人の入院患者を対象とした 8 か月間にわたる無作為化導入により、ベッドあたりの累積感染者数が 9.7% 増加し、患者報告による健康情報へのアクセスが改善されました。これらの結果により、臨床 AI は、孤立したイベントの予測から、完全な筋骨格ケア経路にわたる長期的な管理の実行へと移行します。
原文 (English)
Evidence-Grounded AI for Musculoskeletal Care
Musculoskeletal diseases are among the leading causes of disability worldwide and create the greatest global need for rehabilitation. Because recovery, remodelling and degeneration often unfold over months to years, musculoskeletal care requires longitudinal management that repeatedly integrates evolving patient evidence, external medical knowledge and stage-specific functional goals. In routine practice, this evidence is fragmented across visits, departments and hospital systems, limiting individualized, evidence-based care. Here we report OrthoPilot, a clinical artificial intelligence system powered by a large language model that integrates hospital data streams with authoritative external knowledge for continuous musculoskeletal management. OrthoPilot autonomously retrieves real-time imaging, laboratory, pathology and order data and converts evolving patient states into evidence-based decisions from admission diagnosis to rehabilitation planning. We established a specialist-validated benchmark from real-world electronic health records spanning 1,000 disease codes. In a reader study across the complete care pathway, OrthoPilot was compared with 81 orthopaedic physicians and surpassed experts with 25 years of experience in diagnostic reasoning, clinical decision-making and management planning. It also outperformed all evaluated intelligent systems across 60 external clinical centres. In a prospective study of 1,870 complex cases, OrthoPilot increased full-chain management success by 10.6%. During an 8-month randomised deployment involving 8,240 inpatients, it increased cumulative cases per bed by 9.7% and improved patient-reported access to health information. These results move clinical AI from predicting isolated events toward executing longitudinal management across complete musculoskeletal care pathways.
パラメータ化された行動マルコフ決定プロセスのための知識と勾配に基づく強化学習
この論文では、パラメータ化されたアクションのマルコフ決定プロセス (PAMDP) における強化学習を研究します。このプロセスでは、各決定は記号アクションと数値パラメーターで構成されます。このような設定では、強化学習アルゴリズムは通常、ワンショット推定器を使用してパラメーターを決定するため、トレーニング サンプルが非効率になります。ほとんどの PAMDP 環境では、明示的ではあるが不完全な知識 (ルール、安全制約、エキスパートヒューリスティックなど) が利用可能ですが、それが強化学習エージェントのトレーニングのサンプル効率を高めるために直接使用されることはほとんどありません。私たちはこのギャップに踏み込み、新しい神経記号知識および勾配誘導強化学習 (KGRL) アルゴリズムを提案します。 KGRL は、Datalog 知識ベースのドメイン知識を使用して、特定の状態に適用可能なアクションと実行可能なパラメーターのセットを導き出します。これにより、適用できないアクションを決定空間から取り除き、残りのアクションのパラメータ空間を制約することができます。次に、勾配ベースのパラメータ調整ループを使用して、エージェントのトレーニングおよび展開中に最適なパラメータを推定します。 KGRL は、アクティブ化されたルールを軌跡に沿って記録することにより、アクションの刈り込みとパラメータの制約に関するローカルな手順の説明をさらに提供します。全体として、KGRL は、トレーニング中のサンプル効率を高めながら、エージェントの探索と展開を実行可能かつ制約を意識した決定に向けて導きます。 KGRL は、サンプル効率とエピソードリターンの両方において、PAMDP の最先端の RL ベースラインを上回ります。
原文 (English)
Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes
In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic action and numerical parameters. In such settings Reinforcement Learning algorithms typically determine parameters with one-shot estimators, which makes their training sample inefficient. Though in most PAMDP environments explicit but incomplete knowledge (e.g., rules, safety constraints, or expert heuristics) is available, it is rarely directly used to increase the sample-efficiency of training Reinforcement Learning agents. We step into this gap and propose our novel Neuro-Symbolic Knowledge- and Gradient-Guided Reinforcement Learning (KGRL) algorithm. KGRL uses domain knowledge in a Datalog knowledge base to derive the set of applicable actions and feasible parameters for a given state. This allows it to prune non-applicable actions from the decision-space and constrain the parameter spaces of the remaining actions. We then use a gradient-based parameter refinement loop to estimate the optimal parameters during training and deployment of the agent. By recording activated rules along the trajectory, KGRL additionally provides local procedural explanations on the pruning of actions and constraining of parameters. Overall, KGRL guides the agent's exploration and deployment toward feasible and constraint-aware decisions, while increasing sample efficiency during training. KGRL outperforms state-of-the-art RL baselines for PAMDPs in both, sample efficiency and episodic return.
Koopman-driven grip force prediction through EMG sensing
Loss of hand function due to conditions like stroke or multiple sclerosis significantly impacts daily activities. Robotic rehabilitation pr…
PersGuard: Preventing Malicious Personalization in Text-to-Image Diffusion Models via Model Backdoors
Diffusion models (DMs) have advanced text-to-image (T2I) synthesis, yet their personalization capabilities raise serious privacy and copyri…
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
Large Language Model (LLM) inference is typically memory-intensive, especially when processing large batch sizes and long sequences, due to…
Uniform Approximation of Functions with Asymmetric Growth and Decay by Deep Weighted Polynomials
Functions that grow without bound on one side of the real line and decay to zero on the other cannot be approximated uniformly by ordinary…
Post-Disaster Affected Area Segmentation with a Vision Transformer (ViT)-based EVAP Model using Sentinel-2 and Formosat-5 Imagery
We propose a vision transformer (ViT)-based deep learning framework to refine disaster-affected area segmentation from remote sensing image…
Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping
Traditional multimodal learning approaches rely on alignment pre-training to bridge vision and language modalities, typically by projecting…
Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models
Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains di…
Representation-Based Exploration for Language Models: From Test-Time to Post-Training
Reinforcement learning (RL) promises to expand the capabilities of language models, but it is unclear if current RL techniques promote the…
Benefits and Limitations of Communication in Multi-Agent Reasoning
Chain-of-thought prompting has popularized step-by-step reasoning in large language models, yet model performance still degrades as problem…
Column Generation with Domain-Independent Dynamic Programming
Column generation and branch-and-price (B&P) are leading mathematical optimization methods for large-scale exact optimization, iterating be…
Cortical-SSM: A Deep State Space Model for Motor Imagery Decoding from EEG Signals
Classification of electroencephalogram (EEG) signals obtained during motor imagery (MI) has substantial application potential, including co…
MASPRM: Multi-Agent System Process Reward Model
Inference-time search over multi-agent systems (MAS) wastes compute when it cannot identify which agent's intermediate message advanced pro…
Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs
As Large Language Models (LLMs) increasingly utilize massive context windows as working memory for autonomous tasks, their reliability fluc…
Mind the Gap: Action Rebinding Attacks against Android GUI Agents
Large multimodal model powered GUI agents are emerging as high-privilege operators on mobile platforms, entrusted to perceive screen conten…
ELF: A Family of Encoder-Free ECG-Language Models
ECG-Language Models (ELMs) extend recent advances in Multimodal Large Language Models (MLLMs) to automated ECG interpretation. However, mos…
With Argus Eyes: Assessing Retrieval Gaps via Uncertainty Scoring to Detect and Remedy Retrieval Blind Spots
Reliable retrieval-augmented generation (RAG) systems depend fundamentally on the retriever's ability to find relevant information. We show…
Left-right asymmetry in predicting brain activity from LLMs' representations emerges with their formal linguistic competence
When humans and large language models (LLMs) process the same text, activations in the LLMs correlate with brain activity measured, e.g., w…
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
Design-to-code translates high-fidelity UI designs into executable front-end implementations, but progress remains hard to compare due to i…
A novel network for classification of cuneiform tablet metadata
In this paper, we present a network structure for classifying metadata of cuneiform tablets. The problem is of practical importance, as the…
When Audio Separation Hurts Zero-Shot ASR: Evaluating SAM-Audio with Whisper on Bengali and English Speech
Recent advances in automatic speech recognition (ASR) and speech enhancement have strengthened the common belief that cleaner audio should…
PC-Diffuser: Path-Consistent Capsule CBF Safety Filtering for Diffusion-Based Trajectory Planner
Autonomous driving in complex traffic requires planners that generalize beyond hand-crafted rules, motivating data-driven approaches that l…
RADAR: Closed-Loop Robotic Data Generation via Semantic Planning and Autonomous Causal Environment Reset
The acquisition of large-scale physical interaction data, a critical prerequisite for modern robot learning, is severely bottlenecked by th…
LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
In existing Audio-Visual Speech Enhancement (AVSE) methods, objectives such as Scale-Invariant Signal-to-Noise Ratio (SI-SNR) and Mean Squa…
Rethinking Multimodal Fusion for Time Series: Text Modalities Need Constrained Fusion
Recent advances in multimodal learning have motivated the integration of auxiliary modalities such as text or vision into time series (TS)…
Research Novelty in Information Systems Journals After ChatGPT: Differences Across Institutional Language Contexts
Large language models are increasingly used in scholarly work, yet it remains unclear whether their productivity gains are accompanied by c…
Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies
Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions with the environment…
Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
Large language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model's opinion. While pr…
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
Should we trust Large Language Models (LLMs) with high accuracy? LLMs achieve high accuracy on reasoning benchmarks, but correctness alone…
Robust Explanations for User Trust in Enterprise NLP Systems
Robust explanations are increasingly required for user trust in enterprise NLP, yet pre-deployment validation is difficult in the common ca…
Partially Observed Structural Causal Models
Here we introduce Partially Observed Structural Causal Models (POSCMs) as an extension of structural causal models (SCMs) to settings where…
Stable Attention Response for Reliable Precipitation Nowcasting
Precipitation nowcasting remains challenging due to the highly localized, rapidly evolving, and heterogeneous nature of atmospheric dynamic…
TuxBot: Semantic-Aware Online OS Tuning with Large Language Models
Online OS tuning can improve long-running services, but existing controllers are poorly matched to live hosts. They treat scheduler, power,…
Post-Deployment Accountability in AI Governance: A Cross-Regulatory Empirical Analysis of AI Incidents
Post-deployment accountability has become central to AI governance, yet little empirical evidence shows whether monitoring, incident report…
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
Video MLLMs face a persistent tension between spatial fidelity and temporal coverage: preserving fine-grained visual details requires many…
DIVE: Embedding Compression via Self-Limiting Gradient Updates
High-dimensional language-model embeddings increase storage and search costs, while supervised compressors can overfit when relevance label…
タイムマシン: 効率的な知覚のための動きの力について
ビデオ表現学習は近年、目覚ましい進歩を遂げています。これは、トレーニングの規模や、言語と対照的にトレーニングされた視覚モデルの成功など、多くの要因によって推進されています。これらの要因は、ビデオ モデルができることの限界を押し広げていますが、同時に独自の制限も導入しています。まず、ビデオ モデルをスケーリングすると法外なコストに達する可能性があり、第 2 に、言語から学習すると、キャプション内の概念を学習できる範囲が制限されます。その結果、ビデオモデルは依然として時間的な理解に苦労しています。この論文では、ビデオ表現の中心的なモダリティとして動きを使用する新しいアプローチを提案します。特に、ポイント トラックの形式でビデオ内の動きが与えられると、マスクされたオートエンコーダを使用してトラックの一部をマスクし、失われたトラックを再構築するようにオートエンコーダをトレーニングします。これにより、自己教師ありの方法で表現を学習できるようになります。私たちは、モーションを使用してビデオを表現することで、ビデオ テクノロジーの核となる制限の両方に実際に対処できることを示します。まず、モーションは本質的に外観に依存しないため、適切に一般化するために必要なサンプルが少なくなるため、トレーニング データの規模を大幅に削減できます。第二に、動作により、言語に依存したトレーニング パラダイムを回避して、より詳細な概念を学習できるようになります。その結果は、TIME (Temporally Informed Motion Embedding) と呼ばれる埋め込みであり、合成モーション データのみでトレーニングされた表現です。この埋め込みを幅広いタスクでゼロショット方式でテストします。付加機能がなければ、パフォーマンスは最大 4 桁少ないトレーニング データを使用する最先端のモデルと同等であることがわかります。これは、より時間的な認識とよりスケーラブルなビデオ モデルの新しいパラダイムへの足がかりです。
原文 (English)
The TIME Machine: On The Power of Motion for Efficient Perception
Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of training and the success of visual models trained contrastively with language. While these factors have pushed the boundaries of what video models can do, they also introduce their own set of limitations: first, scaling video models can reach prohibitive costs and second, learning from language restricts the range of concepts that can be learned to those in captions. As a result, video models still struggle with temporal understanding. In this paper we propose a novel approach that uses motion as the central modality for video representation. In particular, given the motion in a video in the form of point-tracks, we use a masked-autoencoder to mask some of the tracks and train the autoencoder to reconstruct the missing tracks. This allows us to learn a representation in a self-supervised manner. We show that using motion to represent videos actually addresses both of the core limitations of video technology. First, it allows us to massively reduce the scale of training data, as motion is inherently appearance-independent and hence needs fewer examples to generalize well. Second, motion allows us to bypass the language-dependent training paradigm, learning better fine-grained concepts. The result is an embedding that we call TIME (Temporally Informed Motion Embedding), a representation trained exclusively on synthetic motion data. We test this embedding on a wide set of tasks in a zero-shot manner. We observe that without bells and whistles, performance is on par with state-of-the-art models using up to 4 orders of magnitude less training data. This is a stepping stone towards a new paradigm of video models that are both more temporally aware as well as more scalable.
注意力の散漫によって引き起こされる視覚的なぼやけを修正して幻覚を軽減する: アルゴリズムと理論
マルチモーダル大規模言語モデル (MLLM) は、物体の幻覚に悩まされることがよくありますが、この失敗の根底にある視覚知覚メカニズムはまだ十分に理解されていません。この研究では、幻覚が人間のような注意散漫現象と強く関連していることを明らかにしました。この現象では、分割焦点下にある人間は視覚の明瞭度が低下し、不正確な説明を生成しますが、モデルでは同じメカニズムが、複数頭の注意における空間的な不一致と、デコード中の画像トークンへの注意の一時的な薄れとして現れます。さらに、注意の分散によってモデルの複雑さが増大し、分類の一般化が低下するという理論的な洞察も提供します。これらの発見に動機づけられて、我々は、画像認識を改善するための注意集中アプローチ(AFIP)を提案します。これは、クロスヘッド注意の強化を通じて注意の散漫を修正し、動的な歴史的注意の強化を通じて視覚の基礎を強化します。複数のベンチマークとモデルに関する広範な実験により、追加のトレーニングなしで AFIP の有効性が検証されます。
原文 (English)
Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory
Multimodal large language models (MLLMs) frequently suffer from object hallucinations, yet the visual perceptual mechanism underlying this failure remains poorly understood. In this work, we reveal that hallucinations are strongly associated with a human-like attention distraction phenomenon, where humans under divided focus experience degraded visual clarity and produce inaccurate descriptions, while in models the same mechanism manifests as spatial inconsistency in multi-head attention and temporal fading of attention to image tokens during decoding. We further provide theoretical insights that attention dispersion increases model complexity and degrades classification generalization. Motivated by these findings, we propose an Attention-Focused Approach for Improved Image Perception (AFIP), which corrects attention distraction via cross-head attention enrichment and reinforces visual grounding through dynamic historical attention enhancement. Extensive experiments on multiple benchmarks and models validate the effectiveness of AFIP without additional training. Code is available at: https://github.com/MIKUZ12/AFIP.
推論が困難な場合: 臨床 SOAP ノート生成のためのフロンティア LLM のソース認識型評価
推論対応 LLM は医療推論ベンチマークで優れたパフォーマンスを発揮しますが、これらの利点が構造化された臨床文書に反映されるかどうかは不明のままです。私たちは、OMI Health、ACI-Bench、PriMock57 にわたるソース認識ベンチマークでの臨床対話からの SOAP ノート生成を使用して、この疑問を調査します。プロバイダーネイティブ推論と同一ソース検索拡張生成 (RAG) を個別に切り替える制御された 2x2 設計で GPT-5.4、DeepSeek-V4-Flash、および Gemma-4-E4B を評価します。成果は、リファレンスを認識した 2 人の LLM 審査員とともに 7 つの自動指標を使用して評価されます。どちらの評価アプローチでも、推論非対応の GPT-5.4 構成が全体として最高の品質を達成するのに対し、DeepSeek-V4-Flash は推論が有効な構成の中で最高のパフォーマンスを発揮するという点で一致しています。推論を有効にすると、3 つのデータセットすべてで GPT-5.4 のパフォーマンスが大幅に低下しますが、同じソースの RAG では、モデルに依存する改善は小さくなります。全体として、この調査結果は、専用のタスク固有の評価を行わずに、忠実度に敏感な SOAP ノート生成を向上させるために、より強力な推論機能を想定すべきではないことを示しています。
原文 (English)
When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation
Reasoning-enabled LLMs perform strongly on medical reasoning benchmarks, but it remains unclear whether these gains transfer to structured clinical documentation; we investigate this question using SOAP note generation from clinical dialogue in a source-aware benchmark spanning OMI Health, ACI-Bench, and PriMock57. We evaluate GPT-5.4, DeepSeek-V4-Flash, and Gemma-4-E4B in a controlled 2x2 design that independently toggles provider-native reasoning and same-source retrieval-augmented generation (RAG). Outputs are assessed using seven automatic metrics alongside two reference-aware LLM judges. Both evaluation approaches agree that a non-reasoning GPT-5.4 configuration achieves the highest overall quality, while DeepSeek-V4-Flash performs best among reasoning-enabled configurations. Enabling reasoning significantly degrades GPT-5.4 performance across all three datasets, whereas same-source RAG yields smaller, model-dependent improvements. Overall, the findings indicate that stronger reasoning capability should not be assumed to improve fidelity-sensitive SOAP note generation without dedicated, task-specific evaluation.
A Multi-Model Metric-based Selection Framework for Abstractive Text summarization
Automatic text summarization has become increasingly important due to the rapid growth of digital textual information. This paper presents…
Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents
With sophisticated cyber-attacks becoming increasingly prevalent, modern networks require intelligent autonomous cyber-defense agents train…
GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling
Generating visually consistent multi-shot videos remains an open challenge. As videos span more shots, inconsistencies can accumulate acros…
MedDiffuseMix: Preserving Diagnostic Evidence with Saliency-Aware Diffusion Medical Image Data Augmentation
Limited data availability, class imbalance, and domain variability remain major barriers to reliable medical image classification. Conventi…
The Joint Effect of Quantization and Sampling Temperature on LLM Safety Alignment: A Factorial Analysis
Modern LLM deployments often combine quantization with higher sampling temperatures to reduce cost, latency, or repetition, yet safety eval…
ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL
Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Co…
拡散変圧器のトレーニング後の枝刈り
拡散変換器 (DiT) は、画像生成において優れたパフォーマンスを示していますが、かなりの計算オーバーヘッドとリソース消費に悩まされています。トレーニング後の枝刈りは有望な解決策を提供します。ただし、DiT の独自のアーキテクチャ設計とパラメータ分布により、従来のプルーニング手法は適用できず、大幅なパフォーマンスの低下につながります。具体的には、LLM 用に開発された、一連の近似によってメトリクスを導出する従来の方法では、顕著性メトリクスにおける重みの相対的な寄与が増幅されます。さらに、DiT の重みは、LLM の重みよりも大幅に大きい値を示します。さらに、既存の枝刈り粒度では、モデル構造の変化が見落とされます。この論文では、カスタマイズされた顕著性基準と枝刈り粒度を導入することで枝刈りパフォーマンスを向上させる DiT-Pruning を提案します。私たちは、エネルギーベースの観点から重みと活性化の寄与のバランスを取る新しい指標を設計し、重要な要素をより効果的に特定できるようにします。さらに、2 次元の重み空間で明確なクラスタリング パターンが観察されます。したがって、クラスタリングを意識したプルーニング粒度を採用し、効果的なスパース割り当てを可能にします。さまざまな DiT に関する広範な評価により、特に高いスパース性の下で、私たちの方法が一貫して画質を維持することが示されています。 MJHQ 上の 512x512 解像度の FLUX.1-dev の場合、DiT-Pruning は 50% のスパース性で CLIP スコアの損失がわずか 0.001 であり、最近のプルーニング手法を劇的に上回っています。
原文 (English)
Post-Training Pruning for Diffusion Transformers
Diffusion Transformers (DiTs) have demonstrated impressive performance in image generation but suffer from substantial computational overhead and resource consumption. Post-training pruning offers a promising solution; however, due to DiTs' unique architectural design and parameter distribution, traditional pruning methods are inapplicable, leading to significant performance degradation. Specifically, prior methods developed for LLMs, which derive metrics through a series of approximations, amplify the relative contribution of weights in the saliency metric. In addition, weights in DiTs exhibit significantly larger magnitudes than those in LLMs. Moreover, existing pruning granularity overlooks variations in model structures. In this paper, we propose DiT-Pruning, which improves pruning performance by introducing customized saliency criteria and pruning granularity. We design a novel metric that balances the contributions of weights and activations from an energy-based perspective, enabling more effective identification of important elements. Furthermore, we observe distinct clustering patterns in the two-dimensional weight space. Accordingly, we adopt a clustering-aware pruning granularity, enabling effective sparse allocation. Extensive evaluations on various DiTs show that our method consistently preserves image quality, especially under high sparsity. For FLUX.1-dev at 512x512 resolution on MJHQ, DiT-Pruning achieves only a 0.001 loss in CLIP score at 50% sparsity, dramatically outperforming recent pruning methods.
Token Geometry
Language models learn continuous programs over discrete symbols, with the embedding table and LM-head acting as the read/write interface be…
Piercing Gilbreath's Conjecture: From Deep Number Theory Insights to Fintech and Cybersecurity
I propose a new methodology to attack the fascinating Gilbreath's conjecture about prime numbers, first posted in 1878 and unsolved to this…
Operator-on-F complements value-equivalence: a planning-time diagnostic for latent world models
World-model evaluation for model-based reinforcement learning typically asks whether the learned model predicts reward and value well, whic…
REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing
Modern autoregressive ASR systems can emit timestamps as decoded tokens, enabling timestamped transcription without frame-level aligners or…
AnchorPrune: ビジュアル トークン プルーニングのための関連性に基づいたコンテキスト拡張
高解像度の入力では数千のビジュアル トークンが導入され、その多くは特定のクエリに対して冗長であるため、大規模なビジョン言語モデルにはかなりの推論コストがかかります。既存の枝刈り手法では、クエリの関連性とトークンの多様性を組み合わせることがよくありますが、これらの目的は、積極的な圧縮の下では矛盾する可能性があります。関連性主導の選択では、相関する局所的な証拠に予算が集中しすぎる可能性がありますが、多様性主導の選択では、不可欠なトークンが抑制されたり、明確ではあるが情報のない領域が保持されたりする可能性があります。最初に保護された関連性アンカーを構築し、次にそれを補完的な視覚的コンテキストで拡張する、トレーニング不要のフレームワークである AnchorPrune を紹介します。 AnchorPrune は、関連性でランク付けされたトークンのノベルティ プロファイルからアンカー サイズを適応的に決定し、クエリクリティカルな証拠のコンパクトなセットを保存し、重要度に重み付けされたノベルティを通じて残りの予算を割り当て、アンカーに関連する有益で冗長でないコンテキストを回復します。この順序付けされたデザインにより、コンテキストの拡張によって不可欠なクエリ キューが置き換えられるのを防ぎ、全体的な視覚的範囲が向上します。 AnchorPrune は軽量でアーキテクチャを認識しており、再トレーニングもモデルの変更も必要ありません。画像およびビデオの視覚言語モデルとベンチマーク全体で、特に厳しい圧縮下で、トレーニング不要のベースラインと比較して精度と効率のトレードオフを一貫して改善します。 LLaVA-NeXT-7B では、AnchorPrune は 2,880 個のビジュアル トークンのうち 160 個のみを使用して、フルトークンのパフォーマンスの 97.6% を維持します。これらの結果は、効率的なマルチモーダル推論のための効果的な原理として、関連性にアンカーされた文脈拡張を確立します。コードは https://github.com/MULTI-cau/AnchorPrune で入手できます。
原文 (English)
AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning
Large vision-language models incur substantial inference costs because high-resolution inputs introduce thousands of visual tokens, many of which are redundant for a given query. Existing pruning methods often combine query relevance and token diversity, yet these objectives can conflict under aggressive compression: relevance-driven selection may overconcentrate the budget on correlated local evidence, while diversity-driven selection may suppress indispensable tokens or retain distinct but uninformative regions. We introduce AnchorPrune, a training-free framework that first constructs a protected relevance anchor and then expands it with complementary visual context. AnchorPrune adaptively determines the anchor size from the novelty profile of relevance-ranked tokens, preserving a compact set of query-critical evidence, and allocates the remaining budget through importance-weighted novelty to recover informative, non-redundant context relative to the anchor. This ordered design prevents contextual expansion from displacing indispensable query cues while improving overall visual coverage. AnchorPrune is lightweight, architecture-aware, and requires neither retraining nor model modification. Across image and video vision-language models and benchmarks, it consistently improves the accuracy-efficiency trade-off over training-free baselines, particularly under severe compression. On LLaVA-NeXT-7B, AnchorPrune preserves 97.6% of full-token performance using only 160 of 2,880 visual tokens. These results establish relevance-anchored contextual expansion as an effective principle for efficient multimodal inference. Code is available at https://github.com/MULTI-cau/AnchorPrune.
TheBioCollection: 生物学用の統合事前トレーニング スケール LLM コーパス
生物学のための大規模言語モデル (BioLM) の推進により、モデルに生物学の真の理解を与えることができるトレーニング コーパスの必要性が生じています。しかし、分子データベース、タンパク質リポジトリ、ゲノムアノテーション、単細胞アトラス、経路データベースなどの既存の生物学的リソースは、異種の形式に分散しており、言語モデルのトレーニング用のまとまりのあるコーパスに編成されていないままです。我々は、これらの異種リソースを、小分子、タンパク質、ゲノム配列、細胞、経路にわたる統一されたトレーニング対応形式に変換する、526億トークンのトレーニング前規模のコーパスであるTheBioCollectionを提示します。 Beyond consolidating existing data, TheBioCollection enriches each record with tool-computed biological properties and introduces new instruction tasks for capabilities that current corpora barely cover.私たちはコーパスを、分子、タンパク質、ゲノム、細胞、およびクロスドメイン設定全体にわたって認識、生成、予測を精査する一致するスイートである TheBioCollection-Eval と組み合わせます。基本の Gravity-16B-A3B アーキテクチャを固定したまま、TheBioCollection でトレーニングすると、一般的な言語能力をほぼそのままにしながら、TheBioCollection-Eval の全体的なスコアが 2 倍以上になり、すべてのドメインで向上します。
原文 (English)
TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology
The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology. However, existing biological resources, such as molecular databases, protein repositories, genomic annotations, single-cell atlases, and pathway databases, are scattered across heterogeneous formats and remain unorganized into a cohesive corpus for language model training. We present TheBioCollection, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways. Beyond consolidating existing data, TheBioCollection enriches each record with tool-computed biological properties and introduces new instruction tasks for capabilities that current corpora barely cover. We pair the corpus with TheBioCollection-Eval, a matched suite probing recognition, generation, and prediction across molecular, protein, genomic, cellular, and cross-domain settings. Holding the base Gravity-16B-A3B architecture fixed, training on TheBioCollection more than doubles its overall score on TheBioCollection-Eval with gains in every domain, while leaving general linguistic ability nearly intact.
HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation
When adapting Vision Language Action (VLA) models to downstream tasks, multiple rounds of post-training are often required to progressively…
AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP
Tool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a week-stale value, or has its descripti…
An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory
Following the rapid progress of generative Artificial Intelligence, there is a growing threat posed by conversational scams. These scams of…
Introducing Human-Centeredness in AI-Assisted Lexicography
This paper proposes a human-centered artificial intelligence (HCAI) framework for AI-assisted lexicography. While generative AI offers sign…
リスクリライトを使用した一般化された分散フリーの半教師あり学習
一般的な半教師あり学習 (SSL) 手法は分布の仮定に依存しており、これに違反するとパフォーマンスが低下します。リスク書き換え手法である PNU 学習は、分布を使用しない代替手段を提供しますが、バイナリ分類に限定されており、その分散の最適性は不明のままです。この論文では、構成要素リスクの線形結合を使用し、PNU 学習を包含し、マルチクラス分類に拡張して不偏リスク推定量を構築する一般化されたフレームワークを提案します。達成可能な最小分散を導出し、非対称損失シナリオにおいて推定器が PNU よりも低い分散を達成できることを示しています。さらに、この分散の減少を学習パフォーマンスの向上に直接結び付ける一般化限界を確立します。これらの理論的な洞察に基づいて、バイナリおよびマルチクラスのベンチマークで経験的に既存のアプローチと同等またはそれを上回る 2 つの実用的な SSL メソッドを紹介します。
原文 (English)
Generalized Distribution-Free Semi-Supervised Learning with Risk Rewrite
Typical semi-supervised learning (SSL) methods rely on distributional assumptions, and their performance degrades when these are violated. While PNU learning, a risk rewriting method, offers a distribution-free alternative, it is restricted to binary classification and its variance optimality remains unclear. In this paper, we propose a generalized framework that constructs unbiased risk estimators using linear combinations of component risks, subsuming PNU learning and extending to multiclass classification. We derive the minimum achievable variance, demonstrating our estimator can attain lower variance than PNU in asymmetric loss scenarios. Furthermore, we establish a generalization bound directly linking this variance reduction to improved learning performance. Based on these theoretical insights, we introduce two practical SSL methods that empirically match or outperform existing approaches on binary and multiclass benchmarks.
除去可能な欠陥: 経済学と意図的な欠陥の限界
スペシャリストはゼネラリストが許容しない盲点を許容します。通常、これは最小限に抑えるべきコストとして扱われます。私たちはそれを設計変数として扱います。不足が維持されるのは、それが致命的になるようなまれな状況では、要求に応じて補償チャネルにルーティングすることで支払いおよび削除されるためです。 3 つの結果を示します。第一に、不足を維持することが計算可能な経済的地位となる有利な条件。構造的には、タウンゼントの高価な状態検証技術としての検出器を使用して、能力ギャップに適用されるエールリッヒ・ベッカーの市場対自己保険のマージンです。第二に、除去可能性の両面の特徴付けです。結合補題は、欠陥が認識の粗大化である場合、どのスイッチも利益と害を分離できず、逆の結果 (交絡した検出器はプレミアムを獲得せず、プラスのプレミアムを主張する欠陥内の政策は、乗算力学の下でマイナスの長期成長に駆動される) と達成可能性の結果 (欠陥の外側の検出器はプラスのプレミアムを獲得する) を生み出すことを示しています。同時に、重大度上限またはミス率 O(1/L) を持つ構造化された不確実性クラス: 検出器関連の区別が制限を乗り越え、有利な条件が維持される場合、欠陥は有益に除去されます。プレミアムは、経済価格ベクトルで設定されたクラスの ROC のサポート関数です。第三に、観測上の欠陥と容量上の欠陥は、展開配布へのアクセスによってそれらが救われるかどうかという点でまったく異なります。ギャップはクロスリークとクロージャの不足として分解され、タスクごとのランダム化によって後者が買い戻されますが、前者は決して買い戻されません。検出器は、損失重大度が線形 (対数係数まで) のトレーニング料金で宣言された致命的カテゴリから学習できます。結果は、チョウの拒否オプション、破滅下でのケリーの成長、および選択的予測を総合したものです。
原文 (English)
Removable Defects: The Economics and Limits of Deliberate Deficiency
A specialist tolerates blind spots that a generalist does not. Usually this is treated as a cost to be minimized. We treat it as a design variable: a deficiency can be kept because it pays and removed on demand in the rare situation where it would be fatal, by routing to a compensation channel. We give three results. First, an advantage condition under which keeping the deficiency is a computable economic position; structurally it is the Ehrlich-Becker market-vs-self-insurance margin applied to a competence gap, with the detector as a Townsend costly-state-verification technology. Second, a two-sided characterization of removability. A coupling lemma shows that when the deficiency is a coarsening of perception, no switch can separate benefit from harm, yielding a converse (a confounded detector earns zero premium, and any within-defect policy insisting on positive premium is driven, under multiplicative dynamics, to negative long-run growth) and an achievability result (a detector outside the deficiency earns a positive premium). Together, over structured uncertainty classes with severity capped or miss rate O(1/L): a defect is profitably removable iff the detector-relevant distinction survives the restriction and the advantage condition holds; the premium is the support function of the class's ROC set at an economic price vector. Third, observation defects and capacity defects differ exactly on whether access to the deployment distribution rescues them; the gap decomposes as cross-leak plus a closure deficit, and per-task randomization buys back the latter, never the former. The detector can be learned from declared fatal categories at a training bill linear in loss severity (up to a log factor). The results synthesize Chow's reject option, Kelly growth under ruin, and selective prediction.
適切なモデルをマージしていますか? LLM のモデル結合に対するエキスパート トレーニング期間の影響
マルチタスク モデルのマージでは、個別にトレーニングされたエキスパート モデルを、共同トレーニングなしですべてのタスクを処理する単一のモデルに結合します。標準的な手法では、最適な検証損失で専門家を結合します。私たちは、ドメイン専門家のトレーニング期間がマージされたモデルの品質にどのような影響を与えるかを体系的に研究することで、この慣例に挑戦します。 3 つのモデル サイズ (Qwen 3.5 0.8B、2B、および 4B) にわたる 5 つのドメイン (数学、コード、命令追従、多言語、安全性) について専門家を微調整し、最適なトレーニング ステップの 25\% ~ 500\% のチェックポイントを節約し、各期間で 5 つのマージ方法を評価します。私たちの調査結果は、手法に依存する顕著なパターンを明らかにしました。単純な平均化は過学習により急激に低下しますが、スパース化ベースの手法は検証の最適値をはるかに超えて最高のパフォーマンスを達成します。これをバイアス分散分解分析を通じて形式化し、分散の高い個々の学習者から平均化のメリットが得られるランダム フォレストとの類似点を描きます。これらの結果は、トレーニング期間と結合方法は独立して選択するのではなく、組み合わせて選択する必要があることを示唆しています。
原文 (English)
Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs
Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice merges experts at their optimal validation loss. We challenge this convention by systematically studying how training duration of domain experts affects the quality of the merged model. We fine-tune experts on five domains (Math, Code, Instruction Following, Multilingual, and Safety) across three model sizes (Qwen 3.5 0.8B, 2B, and 4B), saving checkpoints from 25% to 500% of the optimal training steps and evaluating five merging methods at each duration. Our findings reveal a striking method-dependent pattern: simple averaging degrades sharply with overfitting, while sparsification-based methods achieve their best performance well past the validation optimum. We formalize this through bias-variance decomposition analysis, drawing a parallel to random forests where averaging benefits from high-variance individual learners. These results suggest that training duration and merging method should be chosen jointly rather than independently.
Mind the Gap: Promises and Pitfalls of Hierarchical Planning in LeWorldModel
We investigate whether temporal hierarchy can improve LeWorldModel on long-horizon goal-conditioned control. We introduce Hi-LeWM, an exten…
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision wi…