AIニュース 2026-07-14
自動生成: 2026-07-14 11:52 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
「We Must Act Now」 AIによる急激な変革に著名経済学者やIT重鎮が警鐘ITmedia AI+
AIによる経済変革への備えを訴える声明「We Must Act Now」が、ノーベル賞受賞者16人を含む200人以上の経済学者やAI研究者…
-
Anthropic starts localizing Claude pricing for India, its biggest market after the USTechCrunch AI
Claude users in India are starting to see Indian rupee-denominated su…
-
OpenAIのブラウザ「ChatGPT Atlas」終了へ 公開から1年足らずでITmedia AI+
米OpenAIがAIブラウザ「ChatGPT Atlas」の停止日を2026年8月9日と案内し、データの移行手順を公開した。移行先には新し…
-
「ChatGPT Work」「Codex」の5時間制限枠を一時解除 「GPT-5.6 Sol」の処理効率も改良へITmedia AI+
米OpenAIは、デスクトップ向けAIツール「ChatGPT Work」や、付随するAIコーディングエージェント「Codex」に設定されて…
-
Already rich, already successful, why the last wave of tech winners is grinding againTechCrunch AI
They're rolling up their sleeves again, seemingly out of fear of miss…
-
Uber’s product chief on hotels, robotaxis, and why the company doesn’t want to be “everything for everyone”TechCrunch AI
Uber Chief Product Officer Sachin Kansal walks TechCrunch through the…
-
Video-generation startup PixVerse raises $439M, valuation soars past $2BTechCrunch AI
With the cash, the company aims to expand its world model offering an…
トピック別件数
- LLM/生成AI 81件
- 研究/論文 65件
- エージェント 45件
- 画像/動画生成 29件
- ビジネス/資金調達 12件
- ロボティクス 11件
- その他 5件
- ハードウェア/半導体 5件
- 規制/政策 2件
日本語メディア11件
ITmedia AI+ (日本語)
日本企業の“鬼門”、アクセンチュアは突破できるか? OpenAIとの協業で狙う「業務効率化超え」
AIを実証実験や一部業務の効率化で止めず、全社の変革へ――。多くの企業が向き合うこの課題に、アクセンチュアはOpenAIとどう取り組むのか。これまで“鬼門”だった、部門をまたいだ業務プロセス改革の突破方法とは。
AIエージェントを作って終わりから「自己進化」へ、富士通MAAF検証開始
富士通は、業務向けマルチAIエージェント基盤「MAAF」を開発した。会議録画などからシステムを自動構成し、運用履歴に基づき安全に自己進化する。自社AI基盤との連携により企業全体のAI活用を支援する狙いだ。
「We Must Act Now」 AIによる急激な変革に著名経済学者やIT重鎮が警鐘
AIによる経済変革への備えを訴える声明「We Must Act Now」が、ノーベル賞受賞者16人を含む200人以上の経済学者やAI研究者の署名付きで公開された。スタンフォード大学の研究者らが取りまとめたもので、AIは産業革命を上回る規模の経済変革をはるかに短期間で引き起こす可…
“純国産の政府AI”稼働へ NTTらのモデル採用 「先陣を切る」――松本デジ相が語った意欲
デジタル庁の政府AI「源内」で、国産AIモデルと国産クラウドを活用した“純国産の政府AI”が稼働する。松本大臣は「先陣を切る取り組みになる」と述べた。
OpenAIのブラウザ「ChatGPT Atlas」終了へ 公開から1年足らずで
米OpenAIがAIブラウザ「ChatGPT Atlas」の停止日を2026年8月9日と案内し、データの移行手順を公開した。移行先には新しいChatGPTデスクトップアプリとChrome拡張機能を挙げている。
GMOグループ、AI時代に「エンジニア含む組織体制見直し」 熊谷代表が「AI変革最高責任者」に
熊谷正寿代表が「グループCAIO」に。「エンジニアを含む組織体制を見直し、AIナイズされた組織へと変革する」
アニメ特化動画生成AI「AnimeGen」無償公開、商用利用も可 国内AIベンチャーAIdeaLab
AI開発企業のAIdeaLabは、アニメに特化した動画生成AIモデル「AnimeGen」(アニメジェン)を公開した。ライセンスは商用利用もできる「Apache-2.0」。
「ChatGPT Work」「Codex」の5時間制限枠を一時解除 「GPT-5.6 Sol」の処理効率も改良へ
米OpenAIは、デスクトップ向けAIツール「ChatGPT Work」や、付随するAIコーディングエージェント「Codex」に設定されている5時間の利用制限枠を一時的に解除すると発表した。
エージェントによる業務自動化をどう実現? 「Microsoft Build 2026」で発表された多数の新技術
Microsoftは開発者向けイベント「Microsoft Build 2026」で、エージェント基盤からモデル、開発端末、量子コンピューティングまで多数の新技術を発表した。
人に残る「クリエイティブな仕事」とは? Adobeの“人×AI”の取り組みから探る
広告分野では制作業務に生成AIが広範囲で利用されるようになり、クリエイターやデザイナーの仕事が一部で奪われつつある。「われわれはAIが全てを決める世界を目指しているのではない」と語るAdobeの新たな取り組みから、クリエイティブ業務における人とAIの関係を考察する。
「GPT-Liveが“まるで人間”」ってホンマ? 出汁を「でじる」、トーストを「素焼き」て言うてたけど……
AIは、何も塗らないトーストを「素焼き」と呼び、出汁を「でじる」と言い出した。
海外メディア10件
TechCrunch AI (英語)
Already rich, already successful, why the last wave of tech winners is grinding again
They're rolling up their sleeves again, seemingly out of fear of missing AI's defining moment and, presumably, the irresistible allure of m…
Uber’s product chief on hotels, robotaxis, and why the company doesn’t want to be “everything for everyone”
Uber Chief Product Officer Sachin Kansal walks TechCrunch through the company's financial-services ambitions, its increasingly complicated…
Video-generation startup PixVerse raises $439M, valuation soars past $2B
With the cash, the company aims to expand its world model offering and reach customers across geographies.
Hermes agent maker Nous Research in talks for new funding at $1.5B valuation
The company is raising at least $75 million, led by Robot Ventures, with significant participation from USV and other prominent investors.
Satya Nadella has issued a shocking warning to companies using AI
Of all the debates raging about the potential downsides of AI, there is one worry causing the most hand-wringing among AI enthusiasts in Si…
The wildest allegations in Apple’s trade secrets lawsuit against OpenAI
Apple’s trade secrets lawsuit against OpenAI contains allegations that range from employees joking about unauthorized access to Apple’s sys…
Sam Altman’s space data center trash talk is what most experts already believe
Responding to Musk accusing him of being a scammer, Altman said, "homeboy you're the one sellling [sic] public market investors on short-te…
Should AI help you get away with killing your spouse?
What does a world of total user-aligned AI actually look like?
Anthropic starts localizing Claude pricing for India, its biggest market after the US
Claude users in India are starting to see Indian rupee-denominated subscription plans.
Waze adds new AI-powered features and customization updates
Some of the new features are powered by Google's Gemini AI assistant, which reflects the tech giant's broader push to integrate Gemini acro…
公式ブログ0件
このカテゴリの新着記事はありませんでした。
論文177件
arXiv cs.AI (英語)
格子トラバーサルによる多層パーセプトロンのインターバル認証
この研究では、AI の安全性の基本的な問題、つまり敵対的な堅牢性に対する厳密な理論的枠組みを提示します。特に、敵対的なロバスト性の問題が格子トラバーサル問題に還元できることを示します。この格子の各要素は、入力点 $\mathbf{x}$ を含む間隔、つまり軸に沿った超長方形に対応します。多層パーセプトロン分類器 (MLP) を考えてみましょう。 $\mathbf{x} \in I$ と $\mathbf{x}$ が MLP の予測を変更することなく $I$ 内で自由に摂動できる場合、区間 $I$ は健全な証明を構成します。補足的に、区間 $I$ は、I$ 内の $\mathbf{x} \ が完全な証明を構成し、$\mathbf{x}$ が $I$ の外に移動すると、MLP の予測が変更されることが保証されます。健全な認証の問題は、十分に研究されている敵対的耐性に対応していますが、完全な認証は文献で検討されていません。私たちは格子トラバーサル演算子を開発し、それを改良と検証の反復スキームに適用します。正式な MLP 検証ツールを使用すると、サウンドの最大化と完全な最小化が保証されます。さらに、目的の最適化問題を検討します。そこでいくつかの興味深い非対称性を発見します。完全な認証の場合、最小解は多項式オラクル呼び出しで取得されます。これは、強力な難治性の結果を証明する健全な認証には当てはまりません。さらに、対称区間 (つまり、$\ell_\infty$-spheres) での最適化問題を調べ、対数アルゴリズムを提供します。最後に、新しい ParallelepipedoNN システムを使用した経験的評価を示します。
原文 (English)
Interval Certifications for Multilayered Perceptrons via Lattice Traversal
In this work we present a rigorous theoretical framework to a foundational problem of AI safety, namely adversarial robustness. In particular, we show that the adversarial robustness problem can be reduced to a lattice traversal problem. Each element of this lattice corresponds to an interval, i.e., an axis-aligned hyper-rectangle, containing an input point $\mathbf{x}$. Consider a multilayered perceptron classifier (MLP). An interval $I$ constitutes a sound certification if $\mathbf{x} \in I$ and $\mathbf{x}$ can be freely perturbed in $I$ without changing the MLP's prediction. Complementarily, an interval $I$ constitutes a complete certification if $\mathbf{x} \in I$ and when $\mathbf{x}$ moves outside of $I$ the MLP's prediction is guaranteed to change. While the sound certification problem corresponds to the well-studied adversarial robustness, complete certifications have not been examined in the literature. We develop lattice traversal operators, which we apply in a refine & verify iterative scheme. Using formal MLP verifiers, sound maximality and complete minimality are guaranteed. Moreover, we examine objective optimization problems. There we discover some interesting asymmetries. For complete certifications, the minimum solution is obtained in polynomial oracle calls. This does not hold for sound certifications, where we prove strong intractability results. Additionally, we examine optimization problems in symmetric intervals (i.e., $\ell_\infty$-spheres), where we provide logarithmic algorithms. Finally, we present an empirical evaluation, using the novel ParallelepipedoNN system.
CogniConsole: 信頼性の高い LLM インタラクションのための正式な抽象化として推論時間制御を外部化
大規模言語モデル (LLM) システムの信頼性は、通常、モデルの機能の関数として構成されます。私たちは、信頼性が \emph{推論時間制御} (タスクのフレーミングとコンテキストの選択を制御する計算層) によって大きく影響されることを実証することで、この問題に挑戦します。 \emph{CogniConsole} は、このコントロールを、プログラムによる調整と制限されたプロンプトベースの推論を組み合わせた構造化インターフェイスに外部化するアーキテクチャ上のインスタンス化です。マルチステップのインタラクティブ環境における \emph{制御性指向のプローブ} ($N=489$) を通じて、非構造化から完全な足場への構造足場の増加により、 \textbf{固定モデル アーキテクチャの下で出力の分散と故障率が系統的に減少する}ことを示します。私たちの結果は、コンテキストドリフトや一貫性のない制約遵守など、観察された多くの故障モードは、不十分な機能ではなく、仕様不足の制御に起因することを示しています。この研究は、推論時間制御を第一級の抽象化として扱うための経験的基礎を提供し、スケーリングだけを超えた LLM システムの設計と評価に新しい方向性を開きます。
原文 (English)
CogniConsole: Externalizing Inference-Time Control as a Formal Abstraction for Reliable LLM Interactions
Reliability in large language model (LLM) systems is typically framed as a function of model capability. We challenge this by demonstrating that reliability is significantly influenced by \emph{inference-time control} -- the computational layer governing task framing and context selection. We introduce \emph{CogniConsole}, an architectural instantiation that externalizes this control into a structured interface combining programmatic coordination with bounded prompt-based reasoning. Through \emph{controllability-oriented probes} ($N=489$) in a multi-step interactive environment, we show that increasing structural scaffolding -- from unstructured to fully scaffolded -- \textbf{systematically reduces output variance and failure rates under a fixed model architecture}. Our results indicate that many observed failure modes, such as context drift and inconsistent constraint adherence, arise from under-specified control rather than insufficient capability. This work provides an empirical basis for treating inference-time control as a first-class abstraction, opening new directions for designing and evaluating LLM systems beyond scaling alone.
GATS: 効率的なエージェント計画のための階層化された世界モデルを使用したグラフ拡張ツリー検索
大規模言語モデル (LLM) エージェントは、複数ステップの計画タスクで有望であることが示されていますが、LATS (言語エージェント ツリー検索) や ReAct などの既存のアプローチは、計画中に LLM 推論に大きく依存しており、高い計算コストと確率的動作につながります。 \textbf{GATS} (Graph-Augmented Tree Search) は、体系的な UCB1 ベースのツリー検索と階層化ワールド モデルを組み合わせて、推論中の LLM 呼び出しを排除しながら優れた計画パフォーマンスを実現する計画フレームワークを紹介します。私たちの 3 層ワールド モデルは、(L1) 正確なシンボリック アクションのマッチング、(L2) 実行ログから学習した統計、および (L3) 未知のアクションに対する LLM ベースの予測を統合します。分岐パスや行き止まりのある合成計画タスクでは、GATS は \textbf{100\% 成功率} を達成します。これに対し、LATS では 92%、ReAct では 64\% です。コーディング ワークフロー、Web ナビゲーション、長期タスクを含む 12 の困難なシナリオにわたる包括的なストレス テストでは、GATS は \textbf{100\% 成功} を維持しましたが、LATS は 88.9 %、ReAct は 23.9 % に低下しました。 GATS では、計画中に \textbf{タスクごとにゼロの LLM 呼び出し} が必要であり (LATS の場合はタスクごとに 37 回)、実行間で差異がゼロの決定論的な計画を生成します。私たちの結果は、学習された世界モデルを使用した体系的な検索が、エージェントの計画において LLM に基づく探索よりも大幅に優れていることを示しています。
原文 (English)
GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning
Large Language Model (LLM) agents have shown promise in multi-step planning tasks, but existing approaches like LATS (Language Agent Tree Search) and ReAct rely heavily on LLM inference during planning, leading to high computational costs and stochastic behavior. We present \textbf{GATS} (Graph-Augmented Tree Search), a planning framework that combines systematic UCB1-based tree search with a layered world model to eliminate LLM calls during inference while achieving superior planning performance. Our three-layer world model integrates: (L1) exact symbolic action matching, (L2) statistics learned from execution logs, and (L3) LLM-based prediction for unknown actions. On synthetic planning tasks with branching paths and dead-ends, GATS achieves \textbf{100\% success rate} compared to 92 % for LATS and 64\% for ReAct. On a comprehensive stress test spanning 12 challenging scenarios -- including coding workflows, web navigation, and long-horizon tasks -- GATS maintains \textbf{100\% success} while LATS drops to 88.9 % and ReAct to 23.9%. GATS requires \textbf{zero LLM calls per task} during planning (vs. 37 per task for LATS) and produces deterministic plans with zero variance across runs. Our results demonstrate that systematic search with learned world models can substantially outperform LLM-guided exploration for agent planning.
Long-Horizon-terminal-bench: 高密度の報酬ベースの評価を使用して、長期期間のターミナル タスクにおけるエージェントの限界をテストする
AI エージェントは、短く明確に指定されたタスクを自律的に完了できるようになりました。ただし、既存の端末ベンチマークは主に、数分以内に終了する単純な問題に焦点を当てており、最終的な結果によってのみ評価されます。この設定では、中間の進捗状況や部分的な解決策が見落とされ、報酬シグナルがまばらになり、エージェントの能力の不完全な全体像が得られます。実験の再現、ソフトウェア エンジニアリング、マルチモーダル解析、インタラクティブ ゲーム、科学技術コンピューティングなど、9 つのカテゴリにわたる 46 の長期的タスクのターミナル ベンチマークである Long-Horizon-terminal-bench を紹介します。各タスクは、リファレンス ソリューションまたはシミュレーション エンジンを使用したターミナル ベンチ スタイルのセットアップに従いますが、さらにきめの細かい段階的なサブタスクに分解されます。この設計により、高密度の中間報酬と部分的なクレジットが可能になり、エージェントが最終目標に到達したかどうかだけでなく、オープンエンドのワークフローでどこまで進んだかも評価で把握できるようになります。 Long-Horizon-terminal-Bench のタスクは通常、数百のエピソードと数分から数時間の実行を必要とし、単発の問題解決ではなく、長期計画、長期コンテキスト管理、反復的なデバッグに重点を置きます。 15 のフロンティア モデルを評価したところ、エージェントはタスクあたり平均 990 万のトークンを消費し、実行あたり約 231 のエピソードと 85.3 分の実行時間を要し、Long-Horizon-terminal-bench は以前の端末ベースのベンチマークよりも要求が厳しくなっていることがわかりました。最も強力なテスト済みモデルでも、部分報酬しきい値 0.95 で 15.2% の合格@1 を達成し、完全報酬しきい値 1.0 で 10.9% を達成します。一方、モデル全体の平均合格率は、2 つのしきい値の下でそれぞれ 4.3% と 1.7% です。これらの結果から、改善の余地があることがわかります。私たちは障害モードとエラー パターンをさらに分析し、Long-Horizon-ターミナル エージェントの将来の進歩をサポートするために Long-Horizon-terminal-Bench をリリースします。
原文 (English)
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, yielding sparse reward signals and an incomplete picture of agent capability. We introduce Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing. Each task follows a Terminal-Bench-style setup with a reference solution or simulation engine, but is further decomposed into fine-grained graded subtasks. This design enables dense intermediate rewards and partial credit, allowing evaluation to capture not only whether an agent reaches the final goal, but also how far it progresses on open-ended workflows. Tasks in Long-Horizon-Terminal-Bench typically require hundreds of episodes and minutes to hours of execution, stressing long-horizon planning, long-context management, and iterative debugging rather than one-shot problem solving. We evaluate 15 frontier models and find that agents consume on average 9.9M tokens per task, with roughly 231 episodes and 85.3 minutes of execution time per run, making Long-Horizon-Terminal-Bench more demanding than prior terminal-based benchmarks. Even the strongest tested model achieves 15.2% pass@1 at a partial-reward threshold of 0.95 and 10.9% at a perfect-reward threshold of 1.0, while the mean pass rate across models is 4.3% and 1.7% under the two thresholds, respectively. These results reveal headroom for improvement. We further analyze failure modes and error patterns, and release Long-Horizon-Terminal-Bench to support future progress on long-horizon terminal agents.
ウラソフ方程式の平均場導出の定式化: 戦略ゲームとしての AI 支援のリーン形式化
数学者に AI システムを指示してもらい、リーン 4 証明アシスタントで研究結果を形式化し、そのアクティビティを形式化ゲームとして組み立てます。目的は、LaTeX ドキュメントをリーンなドキュメントに変えることです。開発がコンパイルされ、Sorry が含まれておらず、ターゲット定理がリーンの基本公理のみに基づいていることがマシン チェックで示されたときに、ゲームは勝ちとなります。再利用は、私たちが導入した定義による 2 番目のチェックです。開発により、より広範なライブラリが吸収できる一般数学の自己完結型の層が得られるかどうかです。このケーススタディは、ドブルシンの平均場ルート、つまり存在、一意性、安定性推定と平均場の限界、およびショートウィンドウ重ね合わせ原理(弱い解はラグランジュ)を介した、非線形ウラソフ方程式のウェルポーズネスの完全で公理的にクリーンな形式化です。人間の役割は、証明を書くことではなく、指示することでした。定義の範囲を絞り、分解を指示し、ライブラリのギャップをトリアージすることでした。 AI エージェントが実行されました。形式化により、各ステートメントが書かれた通りに証明されたことが証明されます。書かれた記述が意図した定理であるかどうかは数学者の判断に委ねられます。ビルドから外れた最適トランスポート機構 (特に、Wasserstein-1 メトリックとKantrovich-Rubinstein 双対性定理の特性) は、Mathlib のみに対してコンパイルされる自己完結型の層に分離されます。開発の約 6 分の 1 (299 個の宣言のうち 49 個) が、逆依存性のない 22 個の宣言インターフェースの背後にあります。見出しの定理は約 1 週間で実行され、完全な開発には約 1 か月かかりました。私たちは定量的な主張を一般的な法則としてではなく、1 つのゲームの観察として報告します。ゲームのルールでは特定のシステムが指定されていないため、方法論的な枠組みは 1 回の実行で得られるツールよりも長持ちするように意図されています。
原文 (English)
A Formalization of the Mean-Field Derivation of the Vlasov Equation: AI-Assisted Lean Formalization as a Strategy Game
We formalize a research result in the Lean 4 proof assistant by having a mathematician direct an AI system, and frame the activity as a formalization game. The objective is to turn a LaTeX document into Lean. The game is won when the development compiles, contains no sorry, and a machine check shows the target theorems rest on Lean's foundational axioms alone. Reuse is a second check, by a definition we introduce: whether the development yields a self-contained layer of general mathematics the wider library could absorb. The case study is a complete, axiom-clean formalization of well-posedness for the nonlinear Vlasov equation via Dobrushin's mean-field route -- existence, uniqueness, the stability estimate and mean-field limit, and a short-window superposition principle (weak solutions are Lagrangian). The human's role was to direct, not to write proofs: to scope the definitions, steer the decompositions, and triage the library's gaps; the AI agent executed. The formalization certifies the proof of each statement as written; whether the written statement is the intended theorem stays the mathematician's judgment. The optimal-transport machinery that fell out of the build (in particular, properties of the Wasserstein-1 metric and the Kantorovich-Rubinstein duality theorem) separates into a self-contained layer that compiles against Mathlib alone: about a sixth of the development (49 of 299 declarations), behind a 22-declaration interface with no reverse dependency. The headline theorems ran in about a week, the full development in about a month. We report the quantitative claims as observations of one game, not as general laws. The game's rules name no particular system, so the methodological framing is meant to outlast the tools of any one run.
ARCANA: ARC-AGI-2 推論のためのリフレクティブ マルチエージェント プログラム合成フレームワーク
厳しいテスト時間とハードウェアの制約下で ARC AGI 2 タスクを解決するための協調的なマルチエージェント フレームワークである ARCANA を紹介します。 ARCANA は、各タスクを反復的な認識、仮説の生成、象徴的な実行、および内省的な洗練に分解します。知覚グラウンディング エージェントは生のグリッドからオブジェクト中心のシーン グラフを構築し、潜在プログラム ポリシーは多様な DSL プログラムを提案し、シンボリック エグゼキューターはデモンストレーションで候補を検証し、リフレクティブ エージェントは次のターンのための失敗主導のフィードバックを合成します。これらのエージェントは、共有された微分可能な黒板を通じて通信し、学習されたメタ コントローラーによってスケジュールされます。この設計では、構造化されたプログラム検索と適応型マルチターン補正を組み合わせて、困難な抽象変換タスクにおける推論効率とソリューションの品質を向上させています。
原文 (English)
ARCANA: A Reflective Multi-Agent Program Synthesis Framework for ARC-AGI-2 Reasoning
We present ARCANA, a collaborative multi agent framework for solving ARC AGI 2 tasks under strict test time and hardware constraints. ARCANA decomposes each task into iterative perception, hypothesis generation, symbolic execution, and reflective refinement. A perceptual grounding agent builds object centric scene graphs from raw grids, a latent program policy proposes diverse DSL programs, a symbolic executor verifies candidates on demonstrations, and a reflective agent synthesizes failure driven feedback for the next turn. These agents communicate through a shared differentiable blackboard and are scheduled by a learned meta controller. The design combines structured program search with adaptive multi turn correction, improving reasoning efficiency and solution quality on challenging abstract transformation tasks.
Neuro-Agentic Control: セキュリティ制御を制御するためのディープ ラーニング ベースの LLM を利用したエージェント AI フレームワーク
運用テクノロジーに対するサイバー攻撃は、コストのかかるダウンタイムや物理的損傷を引き起こすことが増えており、産業用 IoT 環境における従来のルールベースの監視の限界を露呈させています。大規模言語モデル (LLM) は、意思決定支援を支援する強力な意味論的推論能力を備えていますが、その幻覚的な性質により、閉ループ制御にとって許容できない安全上の責任が生じます。この論文では、物理学に基づいた自律防御を実現するために、LLM ベースのプランナー (つまり、Gemini 2.5 Flash-Lite など) と事前トレーニングされた時系列基盤モデル (TimesFM) を結合する新しいアーキテクチャであるニューロエージェント制御フレームワークを紹介します。この論文では、システムが幻覚や危険な行動を拒否できるようにしながら、作動前に基礎モデルの数値潜在空間内でLLMが提案する介入の影響をシミュレートする「反事実物理注入」メカニズムが紹介されています。確率的攻撃シナリオのコンテキストで産業データセット (安全な水処理 (SWaT) など) で評価されたこのフレームワークは、LSTM および TCN ベースラインと比較して優れたパフォーマンスを示しました。 Neuro-Agentic Loop は、LSTM (26.7%) および TCN (13.3%) と比較して、しきい値を下回る 5 件の侵害 (33.3%) を阻止し、物理的に無効な (幻覚のある) アクションは実行されませんでした。これらの結果は、重要なインフラストラクチャでエージェント AI を保護するための決定論的な「センチネル」として基盤モデルを使用することの有効性を示しています。
原文 (English)
Neuro-Agentic Control: A Deep Learning-based LLM-Powered Agentic AI Framework for Controlling Security Controls
Cyberattacks on operational technology are increasingly causing costly downtime and physical damage, exposing the limitations of traditional rule-based monitoring in industrial IoT environments. While Large Language Models (LLMs) have strong semantic reasoning abilities to assist in decision support, their hallucinatory nature presents unacceptable safety liabilities for closed-loop control. This paper introduces a neuro-agentic control framework, a novel architecture that couples an LLM-based planner (i.e., such as Gemini 2.5 Flash-Lite) with a pre-trained Time-Series Foundation Model (TimesFM), to achieve physics-grounded autonomous defense. The paper introduces a ``Counterfactual Physics Injection'' mechanism that simulates the impact of LLM-proposed interventions within the numerical latent space of the foundation model before actuation, while allowing the system to reject hallucinatory or unsafe actions. Evaluated on an industrial dataset (e.g., the Secure Water Treatment (SWaT)) in the context of stochastic attack scenarios, the framework exhibited better performance compared to LSTM and TCN baselines. The Neuro-Agentic Loop prevented five breaches (33.3%) below the threshold versus LSTM (26.7%) and TCN (13.3%), with zero physically invalid (hallucinated) actions executed. These results demonstrate the efficacy of using foundation models as deterministic ``Sentinels'' to safeguard agentic AI in critical infrastructure.
L-MAD: 法的推論におけるマルチエージェントの議論構造の体系的評価
マルチエージェントディベート (MAD) フレームワークは、一般的な推論において大きな可能性を示していますが、高度に構造化され、知識が必要な法的領域におけるその有効性は依然として十分に研究されていません。この研究では、法的本文含意内のさまざまな議論の構造と集計方法を体系的に評価するために、法的マルチエージェント討論(L-MAD)フレームワークを導入します。 L-MAD は、個別の専門家ペルソナを複数のエージェントに割り当てることで、単一エージェントの強力なベースラインを最大 8\% 改善します。さらに、討論の規模を分析すると、明らかなトレードオフが明らかになります。エージェントの数を増やすと、矛盾が減り、精度が向上します。一方、討論ラウンドを延長すると、エージェントが互いの間違いを補強し合う、有害な \textit{過剰審議のドリフト} が誘発されます。最終的に、私たちの調査結果は、一か八かの法的推論環境で協調的なマルチエージェント システムを導入する際の実際的な境界と安全マージンを概説します。
原文 (English)
L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning
While multi-agent debate (MAD) frameworks have shown significant potential in general reasoning, their effectiveness in highly structured, knowledge-heavy legal domains remains under-explored. In this work, we introduce the Legal Multi-Agent Debate (L-MAD) framework to systematically evaluate different debate structures and aggregation methods within Legal Textual Entailment. By assigning distinct expert personas to multiple agents, L-MAD improves upon strong single-agent baselines by up to 8\%. Furthermore, analyzing how debate scales reveals a clear trade-off: increasing the agent population reduces inconsistency and improves accuracy, whereas extending discussion rounds induces a detrimental \textit{over-deliberation drift} where agents reinforce each other's mistakes. Ultimately, our findings outline the practical boundaries and safety margins of deploying collaborative multi-agent systems in high-stakes legal reasoning environments.
MedRealMM: 中国のオンライン医療相談のための現実世界のマルチモーダル ベンチマーク
オンライン診療では大規模言語モデル (LLM) の導入が進んでいますが、既存のベンチマークは依然として実際の臨床実践とあまり一致していません。その多くは、合成会話や患者シミュレーターに依存し、患者がアップロードした医療画像を省略したり、臨床の質をあまり反映していない多肢選択や語彙の重複指標を使用して自由回答型の臨床反応を評価したりしています。 \textbf{MedRealMM} は、中国全土のインターネット病院から収集された匿名化された患者と医師のやり取りから構築された、マルチモーダルなオンライン医療相談の大規模ベンチマークです。 MedRealMM は、マルチモーダル クリニカル チャレンジ ポイント (MCCP) 抽出フレームワークを使用して、本物の診察軌跡における臨床的に要求の高い瞬間を特定し、先行するテキストと画像のコンテキストを維持しながら、それぞれを標準化された次の応答生成タスクに変換します。各事例は、臨床的に望ましい行動を表彰し、安全でない、裏付けのない、または矛盾した反応を罰する、医師によって洗練された事例固有のルーブリックと組み合わされています。現在のリリースには、64 の診療科にわたる 5,620 件の実際の複合症例が含まれています。テキスト専用システムやマルチモーダル システムを含む、19 の汎用 LLM と医療特化 LLM を評価します。私たちの結果は、信頼性の高い臨床パフォーマンスには画像情報が不可欠であり、現在のフロンティアモデルが依然としてオンライン医師の反応を下回っていることを示しています。一部のフロンティアモデルは医師と同じかそれ以上の肯定的な臨床基準を満たしていますが、より多くの否定的な基準を引き起こしており、安全性を重視したエラー回避が依然として中心的なボトルネックであることを示しています。 MedRealMM は、現実世界のオンライン診療における多様な医療推論を評価するための、現実的で再現可能なベンチマークを提供します。データセットは、Hugging Face (https://huggingface.co/datasets/jdh-algo/MedRealMM) で公開されます。
原文 (English)
MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation
Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clinical practice. Many rely on synthetic conversations or patient simulators, omit patient-uploaded medical images, or evaluate open-ended clinical responses using multiple-choice or lexical-overlap metrics that poorly reflect clinical quality. We introduce \textbf{MedRealMM}, a large-scale benchmark for multimodal online medical consultation built from de-identified patient-doctor interactions collected from a nationwide Chinese internet hospital. MedRealMM uses a Multimodal Clinical Challenge Point (MCCP) extraction framework to identify clinically demanding moments in authentic consultation trajectories and converts each into a standardized next-response generation task while preserving the preceding text-image context. Each instance is paired with a case-specific rubric refined by physicians that rewards clinically desirable behaviors and penalizes unsafe, unsupported, or contradictory responses. The current release contains 5,620 real-world multimodal cases spanning 64 clinical departments. We evaluate 19 general-purpose and medical-specialized LLMs, including text-only and multimodal systems. Our results show that image information is critical for reliable clinical performance and that current frontier models remain below the online physician response. Although some frontier models satisfy as many or more positive clinical criteria than physicians, they trigger more negative criteria, indicating that safety-sensitive error avoidance remains a central bottleneck. MedRealMM offers a realistic and reproducible benchmark for evaluating multimodal medical reasoning in real-world online consultation. The dataset will be publicly available on Hugging Face at https://huggingface.co/datasets/jdh-algo/MedRealMM.
KV-PRM: マルチエージェントのテスト時間のスケーリングのための KV キャッシュ転送による効率的なプロセス報酬モデリング
プロセス報酬モデル (PRM) は、LLM ベースのマルチエージェント システムの機能を大幅に向上させるテスト時間スケーリング (TTS) 手法を導くのに非常に効果的であることが証明されています。ただし、既存の PRM はテキストベースであり、軌跡テキスト全体を最初から再エンコードします。長期にわたるマルチエージェントのロールアウトでは、スコアリング コストがシーケンス長 L に対して二次関数的に増大し、深刻な計算ボトルネックが生じ、長いコンテキストのシナリオでの PRM の適用が大幅に制限されます。これを解決するために、LLM の生成フェーズ中に自然に生成される KV キャッシュを直接読み取ることで、重いテキストの再エンコードを排除する、高効率のプロセス報酬モデルである KV-PRM を導入します。既存の KV キャッシュに対して単一の「検証トークン」を処理することにより、KV-PRM はスコアリング コストを O(L^2) から O(L) に削減します。 KV キャッシュにはテキストよりも厳密に大きな情報容量が含まれており、下流の報酬モデリングにとってより効率的であることが正式に証明されています。経験的に、MATH、GSM8K、および AIME ベンチマーク全体で、KV-PRM は、ビーム検索、MCTS、重み付き投票などのさまざまな TTS メソッドの下でテキスト PRM と同等または厳密にそれを上回り、テキストベースの PRM と比較して、スコアリング FLOP が最大 5,000 倍、レイテンシが 37 倍、シーケンスごとのメモリ フットプリントが 34 倍削減されます。
原文 (English)
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling
Process Reward Models (PRMs) have been proven to be highly effective in guiding test-time scaling (TTS) methods, which significantly boost the capabilities of LLM-based multi-agent systems. However, existing PRMs are text-based: they re-encode the entire trajectory text from scratch. In long multi-agent rollouts, the scoring cost, growing quadratically with respect to sequence length L, creates a severe computational bottleneck, severely limiting PRMs' application in long-context scenarios. To resolve this, we introduce KV-PRM, a highly efficient process reward model that eliminates the heavy text re-encoding by directly reading the KV cache produced naturally during the LLM's generation phase. By processing a single "verify token" against the pre-existing KV cache, KV-PRM reduces the scoring cost from O(L^2) to O(L). We formally prove that the KV cache contains strictly greater information capacity than text, and is more efficient for downstream reward modeling. Empirically, across the MATH, GSM8K, and AIME benchmarks, KV-PRM matches or strictly outperforms text-PRMs under various TTS methods such as Beam Search, MCTS, and Weighted Voting, with up to a 5,000x reduction in scoring FLOPs, a 37x reduction in latency, and a 34x reduction in per-sequence memory footprint compared to text-based PRMs.
分散シフト下での信頼性の高い長期的なエージェントコンテキスト進化のための範囲指定された検証
デプロイされた LLM エージェントは、エージェント コンテキスト、つまり運用ハーネスによって組み立てられたモデル外部のテキスト コントロール コンテンツに依存します。この作業では、そのコンテキストの変更可能なコンポーネントは、モデル、ツール、ハーネスが固定されたままである一方で、運用経験から更新される永続的なシステムレベルの命令です。長い進化の期間にわたって、蓄積された命令が増大し相互作用するにつれて、フラットテキストのメンテナンスにより検証がますます困難になります。我々は、永続命令コンポーネントを型付きセマンティック グラフとして維持し、変更されたノードのローカル型付き近傍内で提案された更新を検証する、グラフ正規化エージェント コンテキスト エボリューション (GRACE) を提案します。受け入れられたグラフ更新は、展開時に使用されるテキスト命令チェックポイントへの増分編集として再構築されます。制御された分散シフト プロトコルの下で $\tau^2$-bench から派生した固定通信エージェント ハーネス内で GRACE を評価します。 GRACE は、5 つの独立したレプリケーションにわたって、パス ^ 3 で測定される厳密な信頼性を、Gemini 2.5 フラッシュのゼロショット値 0.091 から最終チェックポイントでの 0.673$\pm$0.136 に向上させます。これは、同じホールドアウト セットにおける Gemini 3.1 Pro のゼロショット リファレンスの 0.242 を上回りますが、フラット テキストの HCE ベースラインは 0.191$\pm$0.051 で終了します。これらの結果は、信頼性の高い長期的なコンテキスト進化のための 2 つの要件、検証をローカルにする構造基盤と、蓄積された命令コンテンツを使用可能な状態に保つ統合メカニズムを特定します。
原文 (English)
Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift
Deployed LLM agents rely on agentic context, the model-external textual control content assembled by an operational harness. In this work, the mutable component of that context is a persistent system-level instruction that is updated from operational experience while the model, tools, and harness remain fixed. Over long evolution horizons, flat-text maintenance makes verification increasingly difficult as accumulated instructions grow and interact. We propose Graph-Regularized Agentic Context Evolution (GRACE), which maintains the persistent instruction component as a typed semantic graph and validates proposed updates within the local typed neighborhoods of modified nodes. Accepted graph updates are reconstructed as incremental edits to the textual instruction checkpoint used at deployment. We evaluate GRACE within a fixed telecom agent harness derived from $\tau^2$-bench under a controlled distribution-shift protocol. Across five independent replications, GRACE improves strict reliability, measured by pass^3, from the Gemini 2.5 Flash zero-shot value of 0.091 to 0.673$\pm$0.136 at the final checkpoint. This exceeds a Gemini 3.1 Pro zero-shot reference of 0.242 on the same held-out set, while the flat-text HCE baseline finishes at 0.191$\pm$0.051. These results identify two requirements for reliable long-horizon context evolution, a structural substrate that makes verification local and a consolidation mechanism that keeps accumulated instruction content usable.
監査可能な AI 科学者を目指して: LLM エージェントのための仮説進化プロトコル
大規模言語モデル (LLM) エージェントは、AI 主導の科学的発見において中心的な役割を果たすことがますます期待されています。幅広い知識、柔軟な推論、ツールの使用を備えた彼らは、仮説の提案、検証を繰り返し、証拠に照らして信念を修正することで、科学的問題を自律的に探求し解決する潜在能力を持っています。しかし、現在のエージェントでは、これらの仮説、テスト、信念の更新は非構造化ログに埋め込まれており、エージェントや人間の研究者がそのプロセスを監査できるメカニズムはありません。ここでは、仮説の生成、評価、展開を明示的で監査可能な操作として提供するエージェント ハーネスである Hypothesis Evolution Protocol (HEP) を提案します。材料科学の研究タスクでは、HEP を備えたエージェントは、計画型エージェントにはない仮説、テスト、証拠、信念のサイクルを実行し、研究質問全体を一般化して、ベース LLM の能力が高まるにつれてプロトコルをより完全に活用します。これらの結果は、科学的推論を検査、検証、構築できる監査可能な AI 科学者への一歩を示しています。
原文 (English)
Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents
Large language model (LLM) agents are increasingly expected to play a central role in AI-driven scientific discovery. Equipped with broad knowledge, flexible reasoning, and tool use, they have the potential to autonomously explore and solve scientific problems by repeatedly proposing hypotheses, testing them, and revising their beliefs in the light of the evidence. In current agents, however, these hypotheses, tests, and belief updates are buried in unstructured logs, and no mechanism lets the agent or the human researcher audit that process. Here we propose the Hypothesis Evolution Protocol (HEP), an agent harness that provides hypothesis generation, evaluation, and evolution as explicit, auditable operations. On materials-science research tasks, a HEP-equipped agent operates the hypothesis--test--evidence--belief cycle that planning-style agents lack, generalizes across research questions, and exploits the protocol more fully as the base LLM becomes more capable. These results mark a step toward auditable AI scientists, whose scientific reasoning can be inspected, verified, and built upon.
OpenProver: Lean 4 を使用したエージェント的およびインタラクティブな定理証明
このシステム ペーパーでは、統合されたリーン 4 形式検証を備えた LLM 駆動の自動定理証明 (ATP) 用のオープンソース システムである OpenProver を紹介します。 OpenProver は、Aletheia などの最近の ATP エージェント システムからインスピレーションを得た Planner-Worker-Verifier アーキテクチャを統合しています。 Planner エージェントは、コンパクトなホワイトボード スクラッチパッドと中間結果の無制限のリポジトリを維持し、数学的作業を並列ワーカーに分解します。 OpenProver は完全にオープンソースであり、生成されたプルーフの自動正式検証を通じて再現可能な評価を提供し、人間によるガイドによるプルーフ検索のための対話型ターミナル インターフェイスを提供します。インタラクティブ モードでは、OpenProver を使用すると、インタラクティブ コード生成で確立された人間と AI の相乗効果によって、人間のオペレーターが証明検索プロセスを監視し、操作することができます。自動形式検証によって可能になる定量的アブレーション実験の可能性を示すために、ProofNet で OpenProver を評価し、それを単純なベースラインと比較します。 OpenProver は https://github.com/kripner/OpenProver で公開されています。
原文 (English)
OpenProver: Agentic and Interactive Theorem Proving with Lean 4
In this system paper, we present OpenProver, an open-source system for LLM-driven automated theorem proving (ATP) with integrated Lean 4 formal verification. OpenProver integrates a Planner-Worker-Verifier architecture inspired by recent ATP agentic systems such as Aletheia. A Planner agent maintains a compact Whiteboard scratchpad and an unbounded Repository of intermediate findings, and decomposes mathematical work into parallel Workers. OpenProver is fully open-source, offers reproducible evaluation through automatic formal verification of generated proofs, and provides an interactive terminal interface for human-guided proof search. In interactive mode, OpenProver allows the human operator to monitor and steer the proof search process, motivated by the established human-AI synergy in interactive code generation. To showcase the potential for quantitative ablation experiments enabled by automatic formal verification, we evaluate OpenProver on ProofNet and compare it with a simple baseline. OpenProver is publicly available at https://github.com/kripner/OpenProver.
LongMedBench: 長期的な臨床意思決定のための医薬品のベンチマーク
この研究では、長期的な臨床意思決定のための実際の EHR ベースのベンチマークである LongMedBench を紹介します。 LLM ベースの医療薬剤のこれまでの評価では、主にショートコンテキスト知識の QA とツールの使用が重視されてきました。しかし、実際の医療は本質的に長期的なものであり、臨床医は繰り返しの訪問、検査、進化する治療にわたる証拠を集約する必要があります。したがって、現実的な評価には長期的な相互作用が不可欠です。 LongMedBench は、MIMIC-IV 入院記録と臨床ノートを時系列イベント ストリームとロングコンテキスト メモリ データセットに統合する再現可能なパイプラインを介して構築されており、エージェントと臨床環境の間で長期にわたるマルチセッションの対話を可能にします。患者数は 335 名で、患者 1 人当たりの入院件数は平均 19.72 件、1 件当たりの医療イベント数は 44.91 件です。長期的な意思決定プロセスに基づいて、事実に基づく QA、時間的推論、長期的な意思決定という 3 つのスイートによる評価分類法を提案します。この分類法は、エージェントが長期にわたって過去の患者情報をどのように理解し、活用しているかを測定します。私たちの実験によると、最近の LLM は明示的なタイムスタンプをうまく利用できますが、暗黙的な時間推論には課題があることがわかりました。 RAG とエージェント メモリ システムは、情報検索タスクのパフォーマンスを向上させることができますが、意思決定タスクのパフォーマンスはモデルの直接のコンテキストに大きく依存します。
原文 (English)
LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making. Prior evaluations of LLM-based medical agents have largely emphasized short-context knowledge QA and tool use. However, real-world medical care is inherently longitudinal, and clinicians must aggregate evidence across repeated visits, tests, and evolving treatments. Therefore, long-horizon interaction is essential for realistic assessment. LongMedBench is constructed via a reproducible pipeline that integrates MIMIC-IV admission records and clinical notes into time-series event streams and long-context memory datasets, enabling long-horizon, multi-session interactions between agents and a clinical environment. It comprises 335 patients, with 19.72 inpatient visits per patient on average and 44.91 medical events per visit. Guided by the long-horizon decision process, we propose an evaluation taxonomy with three suites: fact-based QA, temporal reasoning, and long-horizon decision-making. This taxonomy measures how agents understand and leverage historical patient information over extended horizons. Our experiments show that while recent LLMs can make good use of explicit timestamps, they have challenges in implicit time inference; The RAG and agent memory system can improve the performance of information retrieval tasks, but the performance of decision-making tasks is highly dependent on the model's immediate context.
コンピューティングパワーネットワーク上の異種LLM組み込みエージェントのための通信効率の高いデジタルツイン調整
異種大規模言語モデル (LLM) を活用した具体化されたエージェント チームは、スマート ファクトリー、倉庫、サービス ロボットなどの物理的な人工知能に広く導入されています。このようなエージェント チーム間のコラボレーションを可能にするには、限られたネットワーク リソースの下で確実に動作する効率的な調整メカニズムが必要です。ただし、マルチラウンドの自然言語ベースの会話に依存する既存の異種 LLM エージェント調整フレームワークでは、3 つの複合的な課題が生じます。まず、エージェント間の対話ではコミュニケーションのオーバーヘッドが発生し、チームの規模に応じて急速に増大します。第 2 に、調整の品質は、エージェント チームの LLM の異種機能によって制限されます。第三に、エージェントは反復的なネゴシエーションによりアクションの遅延に悩まされる可能性があります。これらの課題に対処するために、軽量デジタル ツイン (DT) 上に構築されたネットワーク化された調整フレームワークである LDT-Coord を提案します。具体的には、各エージェントは目的のアクションを独立して選択し、アクションの決定と共有リソースに対する構造化された時間的制約の両方を DT サーバーに報告することで、調整パフォーマンスを自然言語推論能力から切り離します。次に、DT はトレーニング不要のルールベースのオーケストレーター アルゴリズムを実行してエージェント間の競合を解決し、そのような競合を防ぐための調整指示を返します。通信オーバーヘッドをさらに削減するために、エージェントのレポート制御を制約付き部分観測マルコフ決定プロセス (C-POMDP) として定式化し、PPO-ラグランジュ アルゴリズムで解決します。シミュレーション結果は、LDT-Coord が従来の調整方法に匹敵するタスク成功率を達成しながら、通信オーバーヘッドを 70 分の 1 以上削減し、LLM 異種環境下でも堅牢性を維持していることを示しています。
原文 (English)
Communication-Efficient Digital-Twin Coordination for Heterogeneous LLM Embodied Agents over Computing Power Networks
Embodied agent teams powered by heterogeneous large language models (LLMs) are being widely deployed in physical artificial intelligence such as smart factories, warehouses, and service robotics. To enable collaboration among such an agent team, efficient coordination mechanisms that operate reliably under limited network resources are required. However, existing heterogeneous LLM-agent coordination frameworks that rely on multi-round natural-language-based conversations introduce three coupled challenges. First, inter-agent dialogue incurs communication overhead that grows rapidly with team size. Second, the quality of coordination is constrained by the heterogeneous capabilities of the agent team's LLMs. Third, agents may suffer from action delays due to iterative negotiation. To address these challenges, we propose LDT-Coord, a networked coordination framework built upon a lightweight digital twin (DT). Specifically, each agent independently selects its intended action and reports both the action decision and a structured temporal constraint over shared resources to the DT server, thereby decoupling coordination performance from natural-language reasoning ability. Then, DT executes a training-free, rule-based orchestrator algorithm to resolve cross-agent conflicts and returns coordination instructions to prevent such conflicts. To further reduce communication overhead, we formulate agent reporting control as a constrained partially observable Markov decision process (C-POMDP) and solve it with the PPO-Lagrangian algorithm. Simulation results show that LDT-Coord achieves a task success rate comparable to conventional coordination methods while reducing communication overhead by more than 70x and maintaining robustness under LLM heterogeneity.
架空の世界構築: 階層コンテキスト圧縮と反復レビューによるマルチエージェント LLM コラボレーション
世界構築、つまり一貫した架空の世界の構築は、ゲーム デザインと文学作品の基礎となる作業です。大規模言語モデル (LLM) は自動コンテンツ生成の新たな可能性を提供しますが、ワールド構築への適用は 3 つの課題に直面しています。構築プロセスに伴って直線的に増加するコンテキストの爆発、創造的な多様性とコンテンツの一貫性の間の緊張、および自動化された品質保証の欠如です。この文書では、5 つの統合コンポーネントを通じてこれらの課題に対処するマルチエージェント協調システムである AutoWorldBuilder について説明します。意味論的な局所性によってタスクをグループ化する DAG ベースのハイブリッド バッチ スケジューラ。 4 層のコンテキスト圧縮メカニズムにより、約 90% のトークン削減が実現します。専門の監査エージェントによる反復レビュー システムにより、提案の通過率が 42% から 85% 以上に向上します。そして、差別化された温度構成によるゼロコード拡張をサポートするスキル主導のエージェント アーキテクチャ。 GPT-OSS 120B と DeepSeek v3.2 を LLM バックエンドとして使用した、20 の多様なワールド構築タスクにわたる 2 つの実験では、95.0% の成功率が実証されました。このシステムは、競合のない配信により、18 ~ 31 分で世界ごとに 56 ~ 103 の自己一貫性のあるコンセプトを生成しました。ここで検証されたアーキテクチャ パターン (予算としてのレイヤー圧縮、セマンティック局所性スケジューリング、生成とレビューの分離など) は、より広範な知識集約型のマルチエージェント LLM アプリケーションに転送されます。
原文 (English)
Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review
Worldbuilding, the construction of coherent fictional worlds, is a foundational task in game design and literary creation. Large Language Models (LLMs) offer new possibilities for automated content generation, but their application to worldbuilding faces three challenges: context explosion that grows linearly with the building process, the tension between creative diversity and content consistency, and the absence of automated quality assurance. This paper presents AutoWorldBuilder, a multi-agent collaborative system that addresses these challenges through five integrated components: a structured concept network with conflict detection; a DAG-based hybrid batch scheduler that groups tasks by semantic locality; a four-layer context compression mechanism achieving approximately 90% token reduction; an iterative review system with specialized Auditor agents that improves proposal pass rates from 42% to over 85%; and a skill-driven agent architecture supporting zero-code extension with differentiated temperature configuration. Two experiments across 20 diverse worldbuilding tasks, using GPT-OSS 120B and DeepSeek v3.2 as LLM backends, demonstrate a 95.0% success rate. The system generated 56-103 self-consistent concepts per world in 18-31 minutes with zero-conflict delivery. The architectural patterns validated here, including layer-as-budget compression, semantic-locality scheduling, and separation of generation and review, transfer to the broader class of knowledge-intensive, multi-agent LLM applications.
ベイジアン因果関係の発見はどのようにして失敗するのでしょうか?潜在交絡下での線形ガウス ネットワークの構造的結果の特徴付け
ベイジアン因果発見は、事後推論を通じて有向非巡回グラフ (DAG) 上の認識論的不確実性を定量化できるため、広く使用されています。しかし、既存の研究では通常、DAG 上の事後分布がどのように反応するかを特徴付けることなく、交絡によって識別可能性が損なわれることが指摘されているため、潜在交絡下でのその挙動は依然としてよく理解されていません。この研究では、厳密に 2 つの観測変数間の加法的潜在交絡に焦点を当て、線形ガウス因果モデルにおける潜在交絡下の事後挙動を分析します。スコア関数が交絡変数間に偽のエッジを持つグラフを優先する臨界相関閾値を導出し、この閾値はサンプルサイズに応じて減少することを示します。データが増えると、偽のエッジが優先されるために必要な相関が低下します。この閾値を超えると、交絡変数の周囲の局所構造によって決定される 2 つの異なる事後故障レジームを特徴付けます。私たちの発見は、複数のグラフ構造に対する正確な事後計算によって裏付けられており、予測される両方の故障状況を示しています。
原文 (English)
How Does Bayesian Causal Discovery Fail? Characterising Structural Consequences in Linear Gaussian Networks under Latent Confounding
Bayesian causal discovery is widely used for its ability to quantify epistemic uncertainty over directed acyclic graphs (DAGs) through posterior inference. However, its behaviour under latent confounding remains poorly understood, as existing work typically notes that confounding breaks identifiability without characterising how the posterior distribution over DAGs responds. In this work, we analyse posterior behaviour under latent confounding in linear Gaussian causal models, focusing on additive latent confounding between exactly two observed variables. We derive a critical correlation threshold above which the score function favours graphs with a spurious edge between the confounded variables, and show that this threshold decreases with sample size -- more data lowers the correlation required for the spurious edge to be favoured. Beyond this threshold, we characterize two distinct posterior failure regimes determined by the local structure around the confounded variables. Our findings are supported by exact posterior computations on multiple graph structures, demonstrating both the predicted failure regimes.
ProofCouncil: 未解決の数学的問題を解決するための LLM エージェント
大規模言語モデル (LLM) は、数学の未解決の問題を解決する上でますます有望であることが示されています。ただし、現実世界の数学的実践に合わせたエージェント ワークフローを通じて、そのパフォーマンスをさらに向上させることができます。この目的を達成するために、著者批判アーキテクチャを使用して未解決の問題に取り組むように設計された数学エージェントである ProofCouncil を紹介します。 ProofCouncil は、エージェントが自律的に解決しなければならない 10 個の現実世界の数学的問題で構成される課題である FirstProof の 2 番目のバッチへの提出物として機能しました。提出された 10 の問題のうち 6 の問題は、審査員によって、せいぜい小さな修正までは正しいと判断され、参加チームの中で最高のパフォーマンスを示しました。また、数学研究者から収集した 30 の未解決の問題についても ProofCouncil を評価します。人間のフィードバックを受けた 21 件のソリューションのうち、5 件は完全に正しいと判断され、さらに 2 件は最終検証待ちで有望であると判断され、さらに 8 件には有用な部分的な進歩が含まれていました。この短いペーパーでは、ProofCouncil の開発と、その作成に使用されるエージェント構築ライブラリについて説明します。このライブラリはオープンソースとしてコミュニティにリリースされます。
原文 (English)
ProofCouncil: An LLM Agent for Solving Open Mathematical Problems
Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved through agentic workflows tailored to real-world mathematical practice. To this end, we introduce ProofCouncil, a mathematical agent that is designed to tackle open problems using an author-critic architecture. ProofCouncil served as a submission to the second batch of FirstProof, a challenge consisting of 10 real-world mathematical problems that agents must solve autonomously. Its submissions for 6 of the 10 problems were judged by the referees to be correct up to at most minor revisions, showing the best performance among participating teams. We also evaluate ProofCouncil on 30 open problems collected from mathematical researchers. Among the 21 solutions that received human feedback, 5 were judged completely correct, 2 more were judged promising pending final verification, and a further 8 contained useful partial progress. In this short paper, we describe the development of ProofCouncil and the agent-building library used to create it, which we release as open source to the community.
最も重要なパイプ: セマンティック抽象化としての AI システム
AI システムの出力は、それが記述しているように見える事実や世界状態ではなく、むしろ工学的に表現されたものです。私たちは、AI システムを記述し、そのような表現の正確さを検査できるようにするための意味論的なフレームワークを提案します。そのために、受け入れられたドメイン知識によって何が正当化されるのか、参照情報源が何を述べているのか、そしてシステムが現在使用できるものを区別します。これにより、一般的な失敗に正確な定義を与えることができます: 外挿、反駁またはサポートされていない主張、ソースと知識の不一致、古いまたは反駁されたソース、追加された仮説、サポートされていない使用... 私たちのフレームワークが、出力、引用、ツール呼び出し、および世界を変えるアクションが見かけの流暢さではなく信頼できる主張と明示的な権威によって正当化される必要がある AI システムを指定およびチェックするための有用な語彙を提供することを願っています。
原文 (English)
Ceci n'est pas une pipe: AI systems as semantic abstractions
An AI system's output is not the fact or world state it appears to describe, but rather an engineered representation. We propose a semantic framework to describe AI systems, to be able to examine the correctness of such representations. To do so, we distinguish what is justified by accepted domain knowledge, what reference sources say, and what the system can currently use. This allows us to give precise definitions to common failures: extrapolation, refuted or unsupported assertion, sources versus knowledge mismatch, stale or refuted source, added hypotheses, unsupported use... We hope our framework gives a useful vocabulary for specifying and checking AI systems whose outputs, citations, tool calls, and world-changing actions must be justified by reliable claims and explicit authority rather than apparent fluency.
強化学習におけるマルチモーダル報酬ハッキング
強化学習 (RL) は、マルチモーダル大規模言語モデル (MLLM) を調整するためにますます使用されていますが、より高い報酬が必ずしもタスクのパフォーマンスの向上を意味するわけではありません。このリスクは、視覚的な証拠がテキストのみの報酬や根拠の薄い報酬によって評価される場合に増幅されます。私たちは、安全性 VQA、チャート VQA、ストレス テスト設定、さまざまな報酬設計、データの曖昧さ、モデル スケール (2B ~ 32B)、および RL アルゴリズム (GRPO、RLOO、DAPO) にわたる MLLM RL の報酬ハッキングを研究します。新規報酬失敗率 (NRFR) を導入します。これは、代理報酬が SFT ベースラインよりも向上するサンプル間の失敗を測定します。結果のみの報酬は深刻なハッキングを引き起こし、報酬ハッキング率 (RHR) が 48.1% に達しますが、RHR を超える NRFR は、RL が単に失敗を継承するのではなく、新たな失敗を生み出すことを示しています。スケーリングによってハッキングは減少しますが、排除されません。32B モデルでも、結果のみの報酬では 54.9% 悪い率が維持されますが、回答を意識した報酬では、あらゆる規模でオラクルの傾向が改善されます。堅牢性はアルゴリズムと規模にも依存します。GRPO は一貫して最も耐性があり、RLOO は依然として脆弱であり、DAPO は 2B から 8B に大幅に向上します。視覚的証拠による報酬は、信頼性の高い検証にのみ役立ちます。キーワードベースのチェックではハッキングが増加しますが、VLM による判断の意味検証ではハッキングが減少します。全体として、マルチモーダル報酬ハッキングは、不完全な報酬を最適化した体系的な結果であり、堅牢な調整には、最適化の圧力下でも信頼性を維持できる報酬と検証者が必要です。
原文 (English)
Multimodal Reward Hacking in Reinforcement Learning
Reinforcement learning (RL) is increasingly used to align multimodal large language models (MLLMs), but higher rewards do not always imply better task performance. This risk is amplified when visual evidence is evaluated by text-only or weakly grounded rewards. We study reward hacking in MLLM RL across safety VQA, chart VQA, and stress-test settings, varying reward design, data ambiguity, model scale (2B-32B), and RL algorithm (GRPO, RLOO, DAPO). We introduce Newly Rewarded Failure Rate (NRFR), which measures failures among samples whose proxy reward improves over the SFT baseline. Outcome-only rewards cause severe hacking, reaching 48.1% Reward Hacking Rate (RHR), while NRFR exceeding RHR shows that RL creates new failures rather than merely inheriting them. Scaling reduces but does not eliminate hacking: even the 32B model retains a 54.9% worse rate under outcome-only rewards, whereas answer-aware rewards improve the oracle trend at every scale. Robustness is also algorithm- and scale-dependent: GRPO is consistently most resistant, RLOO remains vulnerable, and DAPO improves substantially from 2B to 8B. Visual-evidence rewards help only with reliable verification: keyword-based checks increase hacking, while VLM-as-judge semantic verification reduces it. Overall, multimodal reward hacking is a systematic result of optimizing imperfect rewards, and robust alignment requires rewards and verifiers that remain reliable under optimization pressure.
エージェントティック LLM システム用の共有選択的永続メモリ
マルチターンツールの使用を通じてコードを生成するエージェント的 LLM システムは、基本的なコンテキストの問題に直面しています。つまり、各セッションはゼロから開始され、以前のセッションを生産的にしていた構成の選択、ドメインの制約、データ スキーマ、およびツールの使用パターンが破棄されます。会話履歴全体を単純に永続化することはトークン効率が悪く、逆効果です。無関係なコンテキストにより生成の品質が低下します。共有選択的永続メモリを導入します。これは、セッション固有の推論トレースを破棄しながら、再利用可能なコンテキストの 4 つのカテゴリ (タスク仕様、データ スキーマ、ツール構成、出力制約) を識別して保持するアーキテクチャです。重要なのは、このメモリが共有されることです。選択的なメモリをカプセル化したワークスペースは、役割ベースのアクセス制御によってユーザー間で転送できるため、冗長な仕様を必要とせずに共同で再利用できます。これは、LLM エージェントが異種ソース (CSV、SQL、REST API、MCP サーバー) から Git バージョン管理された成果物 (ダッシュボード、レポート、データ駆動型ドキュメント) を生成、編集、保守する、デプロイされた共同ワークスペース プラットフォームに実装されています。補完的なゼロトークン データ リフレッシュ メカニズムにより、生成されたプログラムがランタイム データから切り離され、再呼び出しを行わずにアーティファクトを再利用できるようになります。 3 つのエンタープライズ シナリオ全体で、共有選択的永続メモリは 96% のタスク完了を達成しました (メモリなしの場合は 79%、完全な履歴がある場合は 71%)。ゼロトークンリフレッシュにより、定期的な更新のための LLM の再呼び出しが不要になり (タスク時間の 14 分の 1 削減)、サマリー駆動の生成により、呼び出しごとのトークン コストが生のデータ挿入と比較して 97 分の 1 に削減されます。 4 つの公開データセットでのレプリケーションにより、12/12 のトライアルでゼロトークン更新が成功し、一般化可能性が確認されました。特に、ナイーブな全履歴の永続性は、エージェントに古いトレースをバイアスすることで積極的に完了を低下させますが、選択的記憶は両方の極端なパフォーマンスを上回ります。
原文 (English)
Shared Selective Persistent Memory for Agentic LLM Systems
Agentic LLM systems that generate code through multi-turn tool use face a fundamental context problem: each session starts from zero, discarding the configuration choices, domain constraints, data schemas, and tool-use patterns that made previous sessions productive. Naively persisting entire conversation histories is token-inefficient and counterproductive: irrelevant context degrades generation quality. We introduce shared selective persistent memory, an architecture that identifies and retains four categories of reusable context (task specifications, data schemas, tool configurations, and output constraints) while discarding session-specific reasoning traces. Crucially, this memory is shared: workspaces encapsulating selective memory can be transferred across users with role-based access control, enabling collaborative reuse without redundant specification. We implement it in a deployed collaborative workspace platform where LLM agents produce, edit, and maintain git-versioned artifacts (dashboards, reports, and data-driven documents) from heterogeneous sources (CSV, SQL, REST APIs, and MCP servers). A complementary zero-token data refresh mechanism decouples generated programs from runtime data, enabling artifact reuse without re-invocation. Across three enterprise scenarios, shared selective persistent memory achieves 96% task completion (vs. 79% without memory and 71% with full history). Zero-token refresh eliminates LLM re-invocation for recurring updates (14x task-time reduction), while summary-driven generation cuts per-invocation token cost by 97x versus raw data injection. A replication on four public datasets confirms generalizability, with zero-token refresh succeeding in 12/12 trials. Notably, naive full-history persistence actively degrades completion by biasing the agent with stale traces, while selective memory outperforms both extremes.
SAGEAgent: マルチモーダル生存予測におけるコストを意識したモダリティ取得のための自己進化エージェント
すべてのがん患者は、正確な生存予測のために完全な診断精密検査を本当に必要としているのでしょうか?集学的臨床腫瘍学では、診療時に収集された人口統計から特殊な組織分析を必要とするゲノムプロファイリングまで、診断手段は臨床的に義務付けられた負担の増大の順序に従います。現在のマルチモーダル生存法は、すべてのモダリティが利用可能であると仮定するか、欠落データを受動的に処理するかのいずれかですが、この順序付けられたワークフローに沿って特定の患者に対して次のモダリティを取得することが正当化されるかどうかを積極的に推論するものはありません。我々はこれを逐次決定問題として定式化し、SAGEAgent (Sequential Acquisition Guided by Experience) を提案します。これは、各患者に対してどの診断モダリティを取得するかを決定し、予測精度と臨床侵襲性のバランスをとった自己進化型 LLM ベースの臨床エージェントです。 SAGEAgent は、数値予測をテキストに変換する臨床ツール、類似した過去の症例を検索するエピソード記憶、経験から再利用可能な決定パターンを蓄積する意味記憶を通じて、各患者の進化する診断状態を推論します。 TCGA-LGG、TCGA-GBM、および BraTS と 4 つの診断モダリティを組み合わせた神経膠腫コホートの実験では、SAGEAgent が平均取得負荷を 55% 削減しながら、競合する生存予測精度を達成することが実証されました。
原文 (English)
SAGEAgent: A Self-Evolving Agent for Cost-Aware Modality Acquisition in Multimodal Survival Prediction
Does every cancer patient truly need a complete diagnostic workup for accurate survival prediction? In multimodal clinical oncology, diagnostic modalities follow a clinically mandated order of escalating burden -- from demographics collected at intake to genomic profiling requiring specialized tissue analysis. Current multimodal survival methods either assume all modalities are available or passively handle missing data, but none actively reason about whether acquiring the next modality is justified for a given patient along this ordered workflow. We formulate this as a sequential decision problem and propose SAGEAgent (Sequential Acquisition Guided by Experience), a self-evolving LLM-based clinical agent that decides which diagnostic modalities to acquire for each patient, balancing predictive accuracy against clinical invasiveness. SAGEAgent reasons about each patient's evolving diagnostic state through clinical tools that translate numerical predictions into text, an episodic memory that retrieves similar past cases, and a semantic memory that accumulates reusable decision patterns from experience. Experiments on a glioma cohort combining TCGA-LGG, TCGA-GBM, and BraTS with four diagnostic modalities demonstrate that SAGEAgent achieves competitive survival prediction accuracy while reducing average acquisition burden by 55%.
固定表現を超えて: オープンエンド AI における語彙と検証者のギャップ
最新の AI システムは、推論、コード化、定理の証明、ツールの使用、および長期的な研究タスクを実行する能力によってますます評価されています。これらは強力な機能ですが、構造的な制限を共有しています。概念的な語彙、検索できる許容可能な解決策の空間、成功を評価する基準など、モデルが動作する表現フレームは通常、事前に固定され提供されます。この論文は、オープンエンドのイノベーションが可能なより強力なインテリジェント システムを構築するには、追加のクラスの操作、つまり新しい表現プリミティブの作成、安定化、再利用が必要であると主張します。これらの操作は、単純に検索対象の空間内を検索するのではなく、検索対象の空間を変更します。私たちは、現在の AI システムと真にオープンエンドな知能との間の距離を、2 つのギャップを通じて特徴づけます。 1 つ目は語彙のギャップ、つまり既存の表現プリミティブを単に再結合するのではなく、新しい表現プリミティブを発明して安定させることの難しさです。 2 つ目は検証者のギャップ、つまり将来の再利用後にのみその完全な効果が明らかになる場合に、新しいプリミティブの価値を判断することの難しさです。私たちは、知性の統一されたフレームワークを通じて両方のギャップを認知的不一致の削減として解釈します。知的行動を一連の認知変換として見ることにより、固定された表現フレーム内で機能する空間内変換と、フレーム自体を変更する可能性のある生成変換とを区別します。これに基づいて、イノベーションの自律性のはしごを提案し、有用な表現の変更に報いる目標、発明されたプリミティブ用の永続メモリ アーキテクチャ、評価される表現とともに進化できる適応検証メカニズムなど、オープンエンド AI を進歩させるためのいくつかの方向性を概説します。
原文 (English)
Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI
Modern AI systems are increasingly being evaluated for their ability to reason, code, prove theorems, use tools, and long-horizon research tasks. These are powerful capabilities, but they share a structural limitation: the representational frame within which the model operates, including its conceptual vocabulary, the space of admissible solutions it can search, and the criteria by which success is evaluated, is typically fixed and supplied in advance. This paper argues that building stronger intelligent systems capable of open-ended innovation requires additional classes of operations: the creation, stabilization, and reuse of new representational primitives, which alter the space being searched rather than simply searching within it. We characterize the distance between current AI systems and genuinely open-ended intelligence through two gaps. The first is the vocabulary gap, the difficulty of inventing and stabilizing new representational primitives rather than merely recombining existing ones. The second is the verifier gap, the difficulty of judging the value of a new primitive when its full payoff may be visible only after future reuse. We interpret both gaps through a unified framework of intelligence as cognitive discrepancy reduction. By viewing intelligent behaviors as a sequence of cognitive transformations, we distinguish intra-space transformations which operate within a fixed representational frame, from generative transformations which may modify the frame itself. On this basis, we propose a ladder of innovation autonomy and outline several directions for advancing open-ended AI, including objectives that reward useful representational change, persistent memory architectures for invented primitives, and adaptive verification mechanisms capable of evolving alongside the representations they evaluate.
都市鉱山の補完リソースとしてのナレッジ グラフと説明可能な AI
都市鉱山の中心となる規制された監査プロセスである解体前評価は、AI サポートが下された決定に対して責任を負い続ける資格のある監査人にサービスを提供する必要がある情報プロセスです。関連する価値の単位は、予測精度だけではなく、サポートされる決定の防御可能性、つまり、その可読性、妥当性、情報源、および異議の可能性です。説明可能な AI テクニックとドメイン ナレッジ グラフはそれぞれこの要件の一部に対応しており、既存の分類法によりそれらの統合がカタログ化されています。文献は記述的には豊富ですが、構造的には仕様が不十分です。未開発のままなのは、なぜ特定の統合がどのリソースだけでは提供できないアーティファクトを生成するのかについての構造的な説明です。この論文は、IS リソースベースの伝統に基づいた相補性理論の解釈を提供します。我々は、4 つの統合された KG-XAI 統合モード (リフティング、制約、タイピング、および改訂) を提案します。各モードは、XAI アーティファクトおよびナレッジ グラフ基板構造に対する型付き操作として定義されます。各モードは、防御性の明確な特性を解放し、解体前の評価要求の種類の規制上のアーティファクトに貢献します。都市採掘プロセスの防火扉の例は、W3C のリンクされた建物データ スタックと評価拡張機能を使用するモードを示しています。
原文 (English)
Knowledge Graphs and Explainable AI as Complementary Resources for Urban Mining
Pre-demolition assessment, the regulated audit process at the heart of urban mining, is an information process in which AI support must serve qualified auditors who remain accountable for the decisions taken. The relevant unit of value is not prediction accuracy alone, but the defensibility of the supported decisions: their legibility, plausibility, sourcing, and contestability. Explainable AI techniques and domain knowledge graphs each address parts of this requirement, and existing taxonomies have catalogued their integration. The literature is descriptively rich but structurally under-specified: what remains less developed is a structural account of why specific integrations produce artefacts neither resource can provide alone. This paper offers a complementarity-theoretic interpretation grounded in the IS resource-based tradition. We propose four consolidated KG-XAI integration modes (Lifting, Constraining, Typing, and Revising), each defined as a typed operation over XAI artefacts and knowledge-graph substrate structures. Each mode unlocks a distinct property of defensibility and contributes to the kind of regulatory artefact pre-demolition assessment demands. A fire-door example from the urban-mining process illustrates the modes using the W3C Linked Building Data stack and valuation extensions.
TrustX エージェント リスク分類フレームワーク (ARC): 内部で作成されたリスク階層化エージェント AI システム
企業および公共部門のコンテキスト全体にわたるエージェント AI システムの普及により、それらを分類および管理する汎用 AI リスク フレームワークの能力を上回っています。このペーパーでは、TrustX エージェント リスク分類フレームワークを紹介します。これは、7 種類のエージェント AI システムに適用でき、基礎的な既存の AI ガバナンス フレームワークに基づいた、構造化された反復可能な手段です。このフレームワークの中核となるのは、リスクを確実に定量化する 12 次元のスコアリング ルーブリックです。このルーブリックは、GPA + IAT 分類モデルや既存の文献から得られた 5 レベルの自律性フレームワークなどの他のコンポーネントと組み合わされます。これらの入力により、マップされた制御推奨事項を含む 3 層のガバナンス出力が生成されます。このタイプのエージェント AI システムに特有のニュアンスを考慮して、特殊なコーディング アシスタント拡張機能も含まれています。次に、実例を使用してフレームワークを実際に示します。 ARC は、AI ガバナンス担当者、リスク担当者、開発者、規制当局を対象としており、拡張を続けてより堅牢にするため、定期的に反復が行われます。コミュニティはここから対話型フレームワークにアクセスできます: https://arc.responsible.ai/
原文 (English)
TrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI Systems
The proliferation of agentic AI systems across enterprise and public-sector contexts has outpaced the capacity of general-purpose AI risk frameworks to classify and govern them. In this paper, we introduce the TrustX Agent Risk Classification Framework, a structured, repeatable instrument that can be applied to seven types of agentic AI systems and is grounded in foundational pre-existing AI governance frameworks. At the core of the framework is a twelve-dimension scoring rubric that robustly quantifies the risk. This rubric is combined with other components, such as the GPA + IAT classification model and the five-level autonomy framework derived from existing literature. These inputs produce a three-tier governance output with mapped control recommendations. A specialised Coding Assistant extension is also included to account for nuances specific to this type of agentic AI system. We then use an illustrative example to show our framework in practice. ARC is intended for AI governance practitioners, risk officers, developers, and regulators, and it will regularly undergo iteration as we continue to expand it and make it more robust. The community can access the interactive framework here: https://arc.responsible.ai/
Agora: オークションベースのタスク割り当てによる LLM エージェント推論の強化
大規模言語モデル (LLM) エージェントの推論機能を強化するには、多様なエキスパート モデルとツールを効果的にオーケストレーションする必要があります。しかし、既存のフレームワークは通常、タスクとエキスパート モデルまたはツールの機能間の大まかなマッチングに基づいて API を呼び出し、機能的に類似した代替品間のパフォーマンスのばらつきやコスト効率などの重要な要素を見落としています。これに対処するために、私たちはタスクをエキスパート モデルとツールに動的に割り当てるためのインセンティブ互換のオークション メカニズムを導入するフレームワークである Agora を提案します。 Agora は、推論ステップを取引可能なアイテムとして扱うことで、エージェントが修正された能力に基づいて入札できるようにし、重要なロジックが最も自信過剰なソルバーではなく、最も有能なソルバーにルーティングされるようにします。 5 つのベンチマークにわたる評価では、Agora が同等の候補プールの下で一致する単一モデル、ルーティング、およびカスケード ベースラインよりも改善していると同時に、単一のオークション パラメーターを通じて制御可能なコストと品質のトレードオフを明らかにしていることが示されています。
原文 (English)
Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task Allocation
Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools. However, existing frameworks typically call APIs based on coarse-grained matching between tasks and the functions of expert models or tools, while overlooking critical factors such as performance variability and cost efficiency among functionally similar alternatives. To address this, we propose Agora, a framework that introduces an incentive-compatible auction mechanism for dynamically allocating tasks to expert models and tools. By treating reasoning steps as tradeable items, Agora enables agents to bid based on their rectified competence-ensuring that critical logic is routed to the most capable solver rather than the most overconfident one. Evaluations across five benchmarks show that Agora improves over matched single-model, routing, and cascade baselines under comparable candidate pools, while exposing a controllable cost-quality trade-off through a single auction parameter.
ConceptSMILE: コンセプトベースの説明可能な AI の信頼性を監査する
概念ベースの説明可能な人工知能 (AI) は、モデル推論を人間が理解しやすくすることができますが、概念レベルの出力は自動的に信頼できるわけではありません。概念ベースの説明の信頼性を評価するための、モデルに依存しない摂動ベースの監査フレームワークである ConceptSMILE を紹介します。 ConceptSMILE は、SMILE を置き換えるのではなく、摂動ベースのロジックを機能レベルまたは領域レベルの帰属から人間が理解できる概念説明の監査まで拡張します。このフレームワークは、入力領域を摂動させ、概念と応答のシフトを測定し、局所性の重み付けを適用し、XGBoost サロゲートを適合させて局所的な概念の動作を近似します。信頼性は、帰属の正確さ、代理の忠実性、忠実性、安定性、一貫性によって評価されます。 MedSAM 由来の視覚概念と VLM ベースの意味概念を比較することにより、網膜眼底画像上の ConceptSMILE を評価します。結果は、信頼性が概念や経路によって異なることを示しています。MedSAM は、より強力な空間帰属と最高のサロゲート忠実度 ($R^2 = 0.8503$、$R_w^2 = 0.8465$) を達成しますが、VLM 経路は、選択されたアーティファクト条件下でより強い血管忠実性とより強い安定性を示します。 ConceptSMILE は、コンセプトベースの XAI の信頼性を評価するための独立した監査レイヤーを提供します。
原文 (English)
ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI
Concept-based explainable artificial intelligence (AI) can make model reasoning more human-understandable, but concept-level outputs are not automatically trustworthy. We introduce ConceptSMILE, a model-agnostic perturbation-based auditing framework for evaluating the reliability of concept-based explanations. Rather than replacing SMILE, ConceptSMILE extends its perturbation-based logic from feature- or region-level attribution to the auditing of human-understandable concept explanations. The framework perturbs input regions, measures concept-response shifts, applies locality weighting, and fits an XGBoost surrogate to approximate local concept behaviour. Reliability is assessed through attribution accuracy, surrogate fidelity, faithfulness, stability, and consistency. We evaluate ConceptSMILE on retinal fundus images by comparing MedSAM-derived visual concepts with VLM-based semantic concepts. Results show that reliability varies across concepts and pathways: MedSAM achieves stronger spatial attribution and the highest surrogate fidelity ($R^2 = 0.8503$, $R_w^2 = 0.8465$), while the VLM pathway shows stronger vessel faithfulness and stronger stability under selected artefact conditions. ConceptSMILE provides an independent audit layer for evaluating the trustworthiness of concept-based XAI.
最小意思決定ダイナミクスと文脈的確率: 量子綱引きモデル
意思決定には、古典的な確率理論に疑問を呈するコンテキスト依存性が見られることがよくあります。この論文は、綱引き (QTOW) 意思決定モデルの量子的な拡張を開発し、そのようなコンテキスト依存性が単一の最小内部状態によっていつ表現されるかを明らかにします。 QTOW 構築では、qutrit の内部状態、保存を保持する更新、および測定に起因する外乱を使用して、1 つのコヒーレントな状態空間内でのモデル決定、学習、およびプローブ操作を行います。この最小限の表現内で、KCBS タイプのプローブ コンテキストを構築することができ、非コンテキストの古典的な非埋め込み可能性の証拠が得られます。主な主張は、量子理論が意思決定から独自に、または仮定に依存せずに導出されるということではありません。むしろ、同じ操作ファミリーの古典的な再構築には、追加の文脈記憶、履歴依存、または拡大された隠れ状態表現が必要です。したがって、文脈的確率は最小限の決定ダイナミクスのリソース署名として現れますが、量子確率はこの構造のコンパクトでメモリ効率の高い実現を提供します。
原文 (English)
Minimal Decision Dynamics and Contextual Probability: A Quantum Tug-of-War Model
Decision making often exhibits context dependence that challenges classical probability theory. This paper develops a quantum-like extension of the Tug-of-War (QTOW) decision-making model to clarify when such context dependence can be represented by a single minimal internal state. The QTOW construction uses a qutrit internal state, conservation-preserving updates, and measurement-induced disturbance to model decision, learning, and probing operations within one coherent state space. Within this minimal representation, KCBS-type probing contexts can be constructed, yielding a witness of non-contextual classical non-embeddability. The main claim is not that quantum theory is uniquely or assumption-freely derived from decision making. Rather, a classical reconstruction of the same operation family requires additional contextual memory, history dependence, or an enlarged hidden-state representation. Thus, contextual probability appears as a resource signature of minimal decision dynamics, while quantum probability provides a compact, memory-efficient realization of this structure.
REFORGE: 逆コンパイルされたバイナリ関数の命名における LLM のリバース エンジニアリング機能をベンチマークする方法
大規模言語モデル (LLM) はリバース エンジニアリング タスクに適用されることが増えており、最近の脅威インテリジェンス レポートでは、LLM が実際の攻撃セキュリティ ワークフロー内で動作していることが示されています。しかし、その能力に関する主張は、それを測定する私たちの能力を上回っています。 LLM 支援バイナリ分析の既存のベンチマークは、関数レベルのグランド トゥルースの構築を解決済みの前処理ステップとして扱い、どれだけの関数が確実に評価可能であるかを開示せずに精度を報告します。私たちは、公正な評価に対する主な障害はモデルの能力ではなく、コンパイラの最適化におけるバイナリとソースのアライメントの信頼性であると主張します。この論文では、来歴追跡パイプラインである Reforge について説明します。これは、コンパイル、DWARF、構文抽出、アラインメント、逆コンパイルを通じて C ソースから関数レベルのグランド トゥルースを構築し、アラインメントの不確実性を 3 層の階層化を備えた 8 ゲートの信頼性ファネルとして運用します。制御されたマイクロベンチマークでは、信頼度の高い利回りは最適化レベル全体で 87.2% から 65.9% に低下し、不対比較では生存者バイアスによる最適化によるパフォーマンス低下が過大評価されます。関数の命名に関する 7 つの最新の LLM の概念実証評価により、その基盤が実証され、不確実性を意識したベンチマーク実践の動機付けが行われました。
原文 (English)
REFORGE: A Method for Benchmarking LLMs' Reverse Engineering Capabilities in Decompiled Binary Function Naming
Large language models (LLMs) are increasingly applied to reverse-engineering tasks, and recent threat-intelligence reporting shows them operating inside live offensive-security workflows. Claims about their capability, however, outpace our ability to measure it. Existing benchmarks for LLM-assisted binary analysis treat the construction of function-level ground truth as a solved pre-processing step and report accuracy without disclosing how many functions were reliably evaluable. We argue that the principal obstacle to fair evaluation is not model capability but the reliability of binary-to-source alignment under compiler optimization. This paper presents Reforge, a provenance-tracked pipeline that constructs function-level ground truth from C source through compilation, DWARF and syntactic extraction, alignment, and decompilation, and that operationalizes alignment uncertainty as an eight-gate confidence funnel with three-tier stratification. On a controlled micro-benchmark, high-confidence yield falls from 87.2% to 65.9% across optimization levels, and unpaired comparisons overstate optimization-induced performance decay through survivorship bias. A proof-of-concept evaluation of seven contemporary LLMs on function naming demonstrates the substrate and motivates uncertainty-aware benchmarking practice
インタラクションを介して大規模言語モデルの知識蒸留を解釈するための統一アプローチ
大規模言語モデル (LLM) における知識蒸留 (KD) は成功しましたが、その有効性の背後にある根本的なメカニズムは依然として不明です。この論文では、相互作用を使用してさまざまな KD 手法の共通メカニズムを探索するための統一的なアプローチを提案します。具体的には、LLM の出力スコアを多数のインタラクションの合計に分解します。各相互作用は、一連の入力変数 (単語など) を含む非線形関係を表します。分解された相互作用に基づいて、さまざまな KD 手法の根底にある共通のメカニズムが相互作用の疎化であること、つまり、スチューデント モデルが推論のために保持する相互作用の数を減らし、他の相互作用の効果をゼロに抑制することであることを発見しました。さらに、さまざまな KD 手法間のパフォーマンスの差異は、複雑な相互作用を処理する能力に起因することがわかりました。 KD 法は通常、スチューデント モデルが複雑な相互作用のより高いスパース性を達成できる場合に、より良いパフォーマンスをもたらします。これらの洞察に基づいて、蒸留プロセス中の複雑な相互作用のスパース性を明示的に強制するために、複雑な相互作用ペナルティ (CIP) と呼ばれるプラグアンドプレイの損失関数を提案します。広範な実験により、CIP を統合すると、ドメイン内とディストリビューション外の両方のベンチマークでさまざまな KD 手法のパフォーマンスが一貫して向上することが実証されました。
原文 (English)
A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions
Despite the success of knowledge distillation (KD) in Large Language Models (LLMs), the underlying mechanism behind its efficacy remains unclear. In this paper, we propose a unified approach to explore the common mechanism of various KD methods using interactions. Specifically, we decompose the output score of the LLM into the sum of numerous interactions. Each interaction represents a nonlinear relationship involving a set of input variables (e.g., words). Based on the decomposed interactions, we discover that the common mechanism underlying various KD methods is the sparsification of interactions, i.e., student models retain fewer interactions for inference while suppressing other interactions to zero effects. Furthermore, we discover that the performance variance across different KD methods arises from their capabilities in handling complex interactions. A KD method typically yields better performance if it enables the student model to achieve higher sparsity of complex interactions. Motivated by these insights, we propose a plug-and-play loss function called Complex Interaction Penalty (CIP) to explicitly enforce the sparsity of complex interactions during the distillation process. Extensive experiments demonstrate that integrating CIP consistently improves the performance of diverse KD methods on both in-domain and out-of-distribution benchmarks.
iLENS: ニューロイメージング生存分析のための解釈可能な LLM ガイドの専門家の混合
アルツハイマー病 (AD) は複雑な神経変性疾患であり、世界中で何百万もの人々に影響を与え続けています。前駆段階での AD 転換を予測することは、疾患の理解と患者のケアにとって依然として重要です。そのため、生存モデルはアルツハイマー病のリスク予測に広く使用されていますが、通常は静的な予測子であり、解釈可能性が限られており、自然言語推論の能力がありません。この研究では、AD 変換における生存予測のための専門家混合 (MoE) に基づく解釈可能な大規模言語モデル (LLM) ガイド付きフレームワークである iLENS を提案します。私たちのアプローチでは、LLM を使用して、構造化された神経画像測定と非構造化情報を合成し、専門家によるルーティングをガイドします。私たちのフレームワークは、患者のサブタイピングにおける競合する予測パフォーマンスと機能を実証します。さらに、当社のフレームワークは、ルーティング決定に対する透明で生物学的に根拠のある理論的根拠を提供し、高性能生存分析と解釈可能な臨床意思決定サポートとの間のギャップを橋渡しします。
原文 (English)
iLENS: Interpretable LLM-Guided Mixture-of-Experts for Neuroimaging Survival Analysis
Alzheimer's Disease (AD) is a complex neurodegenerative disorder that continues to impact millions of people worldwide. Predicting AD conversion during the prodromal stage remains critical for disease understanding and patient care. As such, survival models are widely used for AD risk prediction, yet they are typically static predictors with limited interpretability and no capacity for natural language reasoning. In this work, we propose iLENS, an interpretable large language model (LLM) guided framework based on mixture-of-experts (MoE) for survival prediction in AD conversion. Our approach uses LLM to synthesize structured neuroimaging measurements and unstructured information to guide expert routing. Our framework demonstrates competitive predictive performance and capability in patient subtyping. Furthermore, our framework provides transparent, biologically grounded rationales for its routing decisions, bridging the gap between high-performance survival analysis and interpretable clinical decision support.
数ビット整数の符号付き対称量子化
符号付き整数アルファベットには、正よりも負の表現可能な値が 1 つ多く含まれます。ただし、慣例により、標準の対称整数量子化器はスケールを厳密に正に固定するため、この追加の表現可能な値が負の末尾に割り当てられ、正の外れ値が強制的にクリッピングされる可能性があります。この研究では、数ビットの精度では、このようなクリッピングが量子化誤差の重要な原因となることを示します。非対称量子化は、ゼロ点を使用してこの問題に対処し、観測されたデータ範囲に向かってグリッドをシフトします。ただし、この柔軟性には実行時のペナルティが伴うことがよく知られています。たとえば、AMD EPYC(TM) "Turin" CPU 上の llama.cpp では、4 ビット対称形式は、非対称形式に比べてメモリ使用量が最大 9% 少なく、スループットが最大 2.45$\倍$ 向上します。非対称形式のペナルティなしで対称量子化の実行時プロファイルを保持する 3 番目のオプションとして、符号付き対称量子化を強調します。符号付き absmax グリッドは、ゼロ点をゼロに保ちながら、原則に基づいた軽量の符号選択ルールを通じて、支配的な外れ値の裾に追加の表現可能な値を配置します。私たちの理論的分析により、2 つの主な結果が得られました。まず、符号付き absmax グリッドを $\ell_2$ 量子化誤差に対して条件付き限界最適として確立し、その条件が低ビット幅での事前トレーニングされた大規模言語モデル (LLM) 全体の重みグループの 88 ~ 99% に当てはまることを示します。次に、標準対称量子化器のスケールを否定することは、同じ符号付き整数アルファベット上の単位ゼロ点シフトと解析的に同等であることを示します。 Qwen3、Qwen3.5、および Llama3 ファミリのモデルで提案を経験的に検証し、追加の推論コストなしで、標準の符号なし対称量子化器と比較して、複雑性とダウンストリームの少数ショット精度が向上していることを観察しました。
原文 (English)
Signed Symmetric Quantization for Few-Bit Integers
The signed integer alphabet contains one more negative representable value than positive. Yet, by convention, the standard symmetric integer quantizer fixes its scale to be strictly positive, which assigns this extra representable value to the negative tail and can force clipping of positive outliers. In this work, we show that, at few-bit precision, such clipping is a non-trivial source of quantization error. Asymmetric quantization addresses this problem with a zero point, shifting the grid toward the observed data range; however, this flexibility is well-known to carry a runtime penalty. For example, in llama.cpp on an AMD EPYC(TM) "Turin" CPU, a 4-bit symmetric format uses up to 9% less memory with up to 2.45$\times$ higher throughput than its asymmetric counterpart. We highlight signed symmetric quantization as a third option that retains the runtime profile of symmetric quantization without the penalty of the asymmetric format: our signed absmax grid places the extra representable value on the dominant-outlier tail through a principled and lightweight sign selection rule while keeping the zero point at zero. Our theoretical analysis offers two main results. First, we establish the signed absmax grid as conditionally bound-optimal on $\ell_2$ quantization error, and show that the condition holds for 88-99% of weight groups across pre-trained large language models (LLMs) at low bit widths. Second, we show that negating the scale of a standard symmetric quantizer is analytically equivalent to a unit zero point shift on the same signed integer alphabet. We empirically validate our proposal on models from the Qwen3, Qwen3.5, and Llama3 families, and observe improvement in perplexity and downstream few-shot accuracy over the standard unsigned symmetric quantizer at no extra inference cost
スティッキー ルーティング: メモリ効率の高い推論のための MoE モデルのトレーニング
Mixture-of-Experts (MoE) モデルでは、トークンごとにエキスパートのまばらなサブセットのみがアクティブになりますが、連続したトークンでは頻繁に異なるエキスパートがアクティブになり、エッジ デバイス上の低速ストレージと高速メモリの間で一定のウェイト スワップが発生します。既存の解決策は、システム レベル (ヒューリスティックのキャッシュ) または事後 (ルーターの微調整) のいずれかであり、根本原因は事前トレーニング中に変更されません。我々は、微分可能なルーティング一貫性の損失である StickyMoE を提案します。これは、隣接するトークン間の突然のエキスパート切り替えにペナルティを課し、意味的に一貫したスパンにわたって同じエキスパート割り当てを維持することをルーターに促します。 StickyMoE はアーキテクチャの変更を必要とせず、単一のハイパーパラメータ ラムダを追加します。ポストホック手法とは異なり、専門家の表現とルーティングの決定を最初のトレーニング ステップから同時に適応させることができます。小規模の MoE 言語モデルでの実験では、StickyMoE が、品質と局所性のフロンティアでパレート支配的なポストホック微調整を行い、混乱度の低下が 4% 未満でエキスパート スイッチ レートを最大 60% 削減することが示されています。ルーティングの時間的局所性は、トレーニング時に最も効率的に組み込まれます。
原文 (English)
Sticky Routing: Training MoE Models for Memory-Efficient Inference
Mixture-of-Experts (MoE) models activate only a sparse subset of experts per token, yet consecutive tokens frequently activate different experts -- causing constant weight swapping between slow storage and fast memory on edge devices. Existing remedies are either system-level (caching heuristics) or post-hoc (router fine-tuning), leaving the root cause unchanged during pretraining. We propose StickyMoE, a differentiable routing consistency loss that penalises abrupt expert switches between adjacent tokens, encouraging the router to maintain the same expert assignment across semantically coherent spans. StickyMoE requires no architectural changes, adds a single hyperparameter lambda, and unlike post-hoc methods, allows expert representations and routing decisions to co-adapt from the first training step. Experiments on small-scale MoE language models show that StickyMoE reduces the expert switch rate by up to 60% with less than 4% perplexity degradation, Pareto-dominating post-hoc fine-tuning on the quality-locality frontier. Routing temporal locality is most efficiently instilled at training time.
報酬の輸送: ノイズ空間アライメントによるフローマッチングにおけるプロパティ制御
フロー マッチングにおけるカップリング (ノイズ ベクトルとデータ ポイントを組み合わせるルール) は、通常、計算上の選択として扱われます。我々は、このカップリングが代わりに位置合わせインターフェースとして機能することを示します。ターゲット分子特性に従ってノイズとデータを照合することにより、制御可能な構造が学習された流れ場に直接埋め込まれます。この見解に基づいて、トレーニング時に最適な輸送結合を使用してスカラー ノイズ空間座標を分子報酬と一致させる報酬輸送を導入します。推論時に、この座標を変更すると、オラクル、報酬モデル、勾配ガイダンス、または追加の計算を必要とせずに、生成された分布が制御されます。結合保存限界では、この座標を閾値処理することにより、クロスエントロピー法の切り捨てられた報酬分布が回復され、原理に基づいた継続的に調整可能な分布レベル制御ノブが提供されます。経験的に、ZINC-250K と GuacaMol では、スカラーをスイープすると、logP の単調制御と、動作範囲全体にわたる一貫した QED 制御が引き起こされます。最も顕著なのは、同じノブが異なるターゲットに対して反対の構造応答を生成し、logP では分子が成長しますが、QED では分子が縮小するため、一般的なサイズの偏りが排除されます。このインターフェイスは、分類子を使用しないガイダンスと条件付きフロー マッチングを補完するものですが、イプシロン予測拡散での否定的な結果により、結合レベルのアラインメントが構造的に存在しない場所が明確になります。コード: https://github.com/KehanGuo2/reward-transport
原文 (English)
Reward Transport: Property Control in Flow Matching via Noise-Space Alignment
The coupling in flow matching -- the rule pairing noise vectors with data points -- is typically treated as a computational choice. We show that this coupling can instead serve as an alignment interface: by matching noise and data according to a target molecular property, it embeds controllable structure directly into the learned flow field. Building on this view, we introduce Reward Transport, which uses optimal transport coupling at training time to align a scalar noise-space coordinate with molecular rewards; at inference, varying this coordinate steers the generated distribution without requiring an oracle, reward model, gradient guidance, or additional computation. In the coupling-preserving limit, thresholding this coordinate recovers the Cross-Entropy Method's truncated reward distribution, providing a principled, continuously adjustable distribution-level control knob. Empirically, on ZINC-250K and GuacaMol, sweeping the scalar induces monotone control of logP and consistent QED control over its operating range; most tellingly, the same knob produces opposite structural responses for different targets, growing molecules for logP but shrinking them for QED, which rules out a generic size bias. The interface is complementary to classifier-free guidance and conditional flow matching, while a negative result under epsilon-prediction diffusion clarifies where coupling-level alignment is structurally absent. Code: https://github.com/KehanGuo2/reward-transport
ディレクター: オンラインでの積極的な専門家の配置による分散型 MoE サービスの加速
エキスパート並列処理は、Mixture-of-Experts (MoE) モデルを提供するための一般的なパラダイムとなっています。その効率は、GPU の通信および計算レイテンシに依存し、これらは GPU への専門家の配置に関係します。エキスパートの配置を最適化するための既存の取り組みは、過去のリクエストのエキスパートのアクティブ化パターンを活用することに重点を置いています。しかし、これらは、多様で急速に変化するリクエスト パターンに直面する欠陥を示しており、オンラインでプロアクティブなアプローチが必要です。このようなアプローチを実装するには、受信リクエストのエキスパートのアクティブ化に伴う不確実性、エキスパートの移行コスト、最適化における NP ハードの複雑さなど、いくつかの課題に対処する必要があります。そこで、予測主導型のオンライン専門家配置によってエンドツーエンドの遅延を最小限に抑える、新しい分散型 MoE サービス システムである Director を紹介します。 Director は、受信リクエストのエキスパート アクティベーション パターンに対して、軽量のカスケード プレディクターまたは低ビット量子化レプリカのいずれかを使用します。次に、オンライン移行モジュールが、コンピューティング依存のフェーズで移行を実行し、中断を最小限に抑えて、ほぼゼロのダウンタイムで変更を適用します。その中核となる緩和ベースのエキスパート配置オプティマイザーは、容量制約の下で動作し、多項式時間で実行され、$(1+\epsilon)$ の近似比を達成します。最後に、プロトタイプを実装し、広範な実験を通じて、人気のある MoE モデル (Mistral、DeepSeek、Qwen など) について、既存の研究と比較してエンドツーエンドの遅延が $11\sim55\%$ 削減されることを実証します。
原文 (English)
Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement
Expert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models. Its efficiency depends on the communication and computation latencies of the GPUs, which are linked to the placement of experts in the GPUs. Existing works for optimizing expert placement focus on leveraging past requests' expert activation patterns. However, they demonstrate deficiencies facing diverse and rapidly changing request patterns, calling for an online, proactive approach. Implementing such an approach requires addressing several challenges: the uncertainty associated with incoming requests' expert activation, the cost of expert migration, and the NP-hard complexity in optimization. Therefore, we present Director, a new distributed MoE serving system that minimizes end-to-end latency via prediction-driven, online expert placement. Director uses either a lightweight cascaded predictor or a low-bit quantized replica for expert activation patterns of incoming requests. An online migration module then enacts the changes with near-zero downtime by executing migrations in compute-bound phases, keeping disruption bounded. At its core, a relaxation-based expert placement optimizer operates under capacity constraints, runs in polynomial time, and achieves a $(1+\epsilon)$ approximation ratio. Finally, we implement a prototype and demonstrate, through extensive experiments, a reduction in end-to-end latency of $11\sim55\%$ for popular MoE models (e.g., Mistral, DeepSeek and Qwen) compared to existing work.
LieBN: リー群に対するバッチ正規化
多様体値の測定は、さまざまな機械学習タスクで普及しています。最近の進歩により、ディープ ニューラル ネットワーク (DNN) が多様体上で動作するように拡張され、さまざまな形状に合わせて調整された正規化技術 (総称してリーマン正規化と呼ばれます) が併用されています。ただし、既存のリーマン正規化法のほとんどは、特定の多様体向けに設計されているか、多様体値の標本分布を効果的に正規化できません。これらの制限に対処するために、リー群に対するリーマン バッチ正規化 (RBN) のフレームワークである LieBN を提案します。私たちのアプローチは、すべてのリー群に自然に存在する理論的に便利な左右不変計量を活用し、リーマン平均と分散を制御するための理論的保証を提供します。 9 つの異なるジオメトリにわたって LieBN をインスタンス化します。そのうちの 4 つは対称正定 (SPD) 多様体上に、1 つは回転行列のグループ上に、4 つはフルランク相関行列の多様体上にあります。特に、SPD 計量の中で、新しい右不変計量を導入し、行列累乗変形を介して 3 つの既存のリー群構造を拡張します。さまざまな多様体に対する広範な実験により、フレームワークの有効性が検証されています。コードは https://github.com/GitZH-Chen/LieBN.git で入手できます。
原文 (English)
LieBN: Batch Normalization over Lie Groups
Manifold-valued measurements are prevalent in various machine learning tasks. Recent advances have extended Deep Neural Networks (DNNs) to operate on manifolds, accompanied by normalization techniques tailored to different geometries, collectively referred to as Riemannian normalization. However, most existing Riemannian normalization methods are either designed for specific manifolds or fail to effectively normalize manifold-valued sample distributions. To address these limitations, we propose LieBN, a framework for Riemannian Batch Normalization (RBN) over Lie groups. Our approach leverages the theoretically convenient left- and right-invariant metrics, which naturally exist in every Lie group, and provides theoretical guarantees for controlling the Riemannian mean and variance. We instantiate LieBN across nine distinct geometries: four on the Symmetric Positive Definite (SPD) manifold, one on the group of rotation matrices, and four on the manifold of full-rank correlation matrices. Notably, among the SPD metrics, we introduce a novel right-invariant metric and extend three existing Lie group structures via matrix power deformation. Extensive experiments on different manifolds validate the effectiveness of our framework. The code is available at https://github.com/GitZH-Chen/LieBN.git.
HERO: フェデレーテッド継続学習のための異質性を認識したベンチマーク ライブラリ
Federated Continuous Learning (FCL) は、分散クライアントが以前に学習した知識を保持しながら、変化するデータ ストリームからどのように学習するかを評価します。既存の評価は、データセット、タスク分割、クライアント データ分割、タスク順序、バックボーン、メモリの仮定、レポート ルールを同時に変更することが多いため、比較するのが困難です。 FCL 用の異質性を認識したベンチマーク ライブラリである \textbf{HERO} を紹介します。 HERO は、タスク分割、クライアント データ分割、クライアント タスク シーケンスという、しばしば組み合わされる 3 つの選択肢を分離してベンチマーク ストリームを構築します。主要な比較可能なベンチマークである HERO-Core では、$\alpha$ がクライアント データのスキューを制御し、$\rho$ がタスク順序の不一致を制御します。最終的な平均精度、平均忘却、および下位 10\% のクライアント精度を使用して、CIFAR-100 および TinyImageNet の代表的な FCL メソッドを評価します。また、OGB-MolPCBA に関するグラフベースのドメイン IL ポータビリティのケース スタディも含まれています。このケース スタディでは、予測タスクは固定されたままで、スキャフォールド ドメインの粒度が入力分布を変更します。私たちの結果は、簡単で異種混合の設定間でメソッドの動作が変化すること、平均精度がボトムクライアントの弱いパフォーマンスを隠す可能性があること、タスク順序の不一致により同期評価とは異なる戦略が有利になること、同じ HERO インターフェイスがイメージベースの FCIL を超えたドメインシフトの困難さを露呈する可能性があることを示しています。 HERO は、再現可能で設定を意識した FCL 評価をサポートするベンチマーク ストリーム、構成、メソッド実装、およびレポート スクリプトをリリースします。
原文 (English)
HERO: A Heterogeneity-Aware Benchmark Library for Federated Continual Learning
Federated continual learning (FCL) evaluates how distributed clients learn from changing data streams while retaining previously learned knowledge. Existing evaluations are difficult to compare because they often change datasets, task splits, client data splits, task orders, backbones, memory assumptions, and reporting rules simultaneously. We introduce \textbf{HERO}, a heterogeneity-aware benchmark library for FCL. HERO builds benchmark streams by separating three choices that are often coupled, namely the task split, the client data split, and the client task sequence. In HERO-Core, the main comparable benchmark, $\alpha$ controls client data skew and $\rho$ controls task-order mismatch. We evaluate representative FCL methods on CIFAR-100 and TinyImageNet using final average accuracy, average forgetting, and bottom-10\% client accuracy. We also include a graph-based Domain-IL portability case study on OGB-MolPCBA, where scaffold-domain granularity changes the input distribution while the prediction task remains fixed. Our results show that method behavior changes across easy and heterogeneous settings, that average accuracy can hide weak bottom-client performance, that task-order mismatch favors different strategies from synchronized evaluation, and that the same HERO interface can expose domain-shift difficulty beyond image-based FCIL. HERO releases benchmark streams, configurations, method implementations, and reporting scripts to support reproducible and setting-aware FCL evaluation.
DaDaDa: データ マーケットプレイスのデータ価格設定用のデータセット
高品質のデータは、業界全体で機械学習の進歩を推進します。データの価値が認識されるにつれ、データトランザクションはますます一般的になり、AWS Marketplace、Databricks、Datarade などの多くのデータ マーケットプレイスが誕生しています。ただし、データ製品の固有の特性により、データ製品の適切な価格を決定することは依然として大きな課題です。経済学における従来の価格設定方法は、コストアプローチ、収入アプローチ、売上比較アプローチに分類できます。コストアプローチは、データレプリケーションによる限界費用がほぼゼロであるため、データ価格設定では失敗し、インカムアプローチは本質的に予測不可能なデータ収益のために失敗します。売上を比較するアプローチは依然として実行可能ですが、市場全体のデータ製品に対する標準化された価格ベンチマークが存在しないことがその適用を妨げています。この課題に対処するために、世界中の 9 つの主要なデータ マーケットプレイスからの 16,147 のデータ製品のメタデータを含む、データ製品価格設定用の最初のデータセットである \texttt{DaDaDa} を導入します。 \texttt{DaDaDa} は、価格設定モデルのトレーニングを可能にし、それによって新しいデータ製品の価格ベンチマークを確立します。さらに、\texttt{DaDaDa} は、データ製品の分類や検索など、データ市場における他の重要なタスクにも利用できます。実験と取得プロトタイプは、データ製品の価格設定、分類、取得に対する \texttt{DaDaDa} の有効性を示しています。データセットとコードは https://github.com/ZJU-DIVER/DaDaDa で入手できます。
原文 (English)
DaDaDa: A Dataset for Data Pricing in Data Marketplaces
High-quality data drives machine learning advances across industries. Recognizing the value of data, data transactions are increasingly common, giving rise to many data marketplaces, e.g., AWS Marketplace, Databricks, and Datarade. However, determining the appropriate prices for data products remains a significant challenge due to the unique properties of data products. Traditional pricing methods in economics can be categorized into the cost approach, the income approach, and the sales comparison approach. The cost approach fails in data pricing due to near-zero marginal cost from data replication, and the income approach fails due to inherently unpredictable data revenue. The sales comparison approach remains viable, yet its application is hindered by the absence of standardized pricing benchmarks for data products across marketplaces. To address this challenge, we introduce \texttt{DaDaDa}, the first dataset for data product pricing, containing metadata for 16,147 data products from 9 major data marketplaces worldwide. \texttt{DaDaDa} enables the training of pricing models, thereby establishing price benchmarks for new data products. In addition, \texttt{DaDaDa} can be utilized for other important tasks in data markets, such as data product classification and retrieval. Experiments and a retrieval prototype demonstrate the effectiveness of \texttt{DaDaDa} for pricing, classification, and retrieval of data products. The dataset and code are available at https://github.com/ZJU-DIVER/DaDaDa.
適度に非構造化されたスパース重み行列を使用した大規模言語モデルの GPU 推論の高速化
大規模言語モデル (LLM) の導入が進むにつれて、LLM 推論コストが重要な課題になっています。重み行列にスパース性を導入する枝刈り手法により、推論を高速化できます。ただし、モデルの品質を維持するには、通常、プルーニングを中程度の非構造化スパース度 (約 50%) に制限します。これらのスパース レベルでは、スパース行列乗算 (SpMM) 用の既存の GPU カーネルはいずれも、高密度の対応物を上回るパフォーマンスを発揮できません。この論文では、中程度のスパース性を持つ LLM に対する効率的な GPU 推論方法を提案します。我々は、(i) スパース テンソル コアが SpMM を高速化できるようにする Sparse-TC レイヤー、 (ii) 低コストのオンチップ復号化をサポートしながら、行列圧縮に並列差分距離を使用するスロット充填層。 (iii) 正しい SpMM 計算を保証する軽量の Residual Layer。この形式に基づいて、スパース テンソル コアと CUDA コアを共同利用する SpMM カーネルを設計します。この設計により、効率的な実行パイプラインが可能になり、オンチップ計算とメモリ アクセスが重複します。評価の結果、私たちの研究は、高帯域幅メモリ (HBM) を備えた最新の GPU での密行列乗算を初めて上回るパフォーマンスを示したことが示されています。 SpInfer (EuroSys'25、最優秀論文) と比較して最大 1.64 倍のカーネル レベルの高速化、および FlashLLM (VLDB'24) と比較して最大 1.41 倍のエンドツーエンドの高速化を実現します。ソースコード: https://github.com/moui0/cudac。
原文 (English)
Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices
With the growing deployment of large language models (LLMs), LLM inference cost has become a key challenge. Pruning techniques that introduce sparsity into weight matrices can accelerate inference. However, maintaining model quality typically limits pruning to moderate unstructured sparsity (around 50\%). At these sparsity levels, none of the existing GPU kernels for sparse matrix multiplication (SpMM) can outperform their dense counterparts. This paper proposes an efficient GPU inference method for LLMs with moderate sparsity. We propose a three-layer matrix storage format comprising: (i) a Sparse-TC layer enabling sparse tensor cores to accelerate SpMM; (ii) a Slot-Filling layer using parallel differential distance for matrix compression while supporting low-cost on-chip decoding; (iii) a lightweight Residual Layer ensuring correct SpMM computation. Building on this format, we design a SpMM kernel that jointly utilizes sparse tensor cores and CUDA cores. This design enables an efficient execution pipeline and overlaps on-chip computation with memory access. Evaluations show that our work is the first to outperform dense matrix multiplication on modern GPUs equipped with high-bandwidth memory (HBM). It achieves up to 1.64x kernel-level speedup over SpInfer (EuroSys'25, Best paper) and up to 1.41x end-to-end speedups over FlashLLM (VLDB'24). Our source code: https://github.com/moui0/cudac.
LLM 主導の多目的ベイジアン最適化アルゴリズムの進化的生成
効果的な多目的ベイジアン最適化 (MOBO) アルゴリズムを設計するには、相互に依存する多くの設計選択肢のバランスを取る必要があり、最適な構成は問題に依存し、通常は深い専門知識が必要になります。 LLaMEA フレームワークを MOBO に拡張し、進化戦略内の突然変異およびクロスオーバー オペレーターとして大規模な言語モデルを使用して、完全なアルゴリズム実装を生成し、SMAC ハイパーパラメーターの最適化を進化ループに統合します。 9 回の進化的実行を通じて約 900 のアルゴリズムを生成し、最先端のベイズ最適化ベースラインとして BoFire qParEGO 実装を使用して、12 の合成問題 (ZDT、DTLZ、WFG) と 3 つの現実世界のエンジニアリング問題 (RE) でベンチマークを行いました。合成スイートでは、最も強力に生成されたアルゴリズムが最高の平均正規化ハイパーボリューム (0.971、qParEGO の場合は 0.869) を達成しながら、必要な実時間は約 60 分の 1 です。事後分析を伴うフリードマン テストでは、この 2 つが単一の最高パフォーマンス グループに分類され、問題ごとのテストでは、生成されたアルゴリズムが 12 問題中 7 問題で qParEGO よりも大幅に優れており、それよりも劣ることはなく、桁違いの低コストで最先端の精度に匹敵することがわかりました。 3 つの目に見えない現実世界のエンジニアリング問題に関して、生成されたアルゴリズムは、約 3.4 倍低い実時間コストで最良の平均正規化ハイパーボリューム (0.985、qParEGO の 0.971 に対して 0.971) を達成しました。これは、3 つの問題のうち 2 つに関して qParEGO よりも大幅に優れており、ゲインが合成レジームを超えて移行することが確認されました。したがって、LLM 主導の進化的探索により、手動設計では到達するのが難しいパレート効率のトレードオフを達成するアルゴリズム設計を発見できます。
原文 (English)
LLM-Driven Evolutionary Generation of Multi-Objective Bayesian Optimization Algorithms
Designing effective multi-objective Bayesian optimization (MOBO) algorithms requires balancing many interdependent design choices whose optimal configuration is problem-dependent and typically demands deep expertise. We extend the LLaMEA framework to MOBO, using large language models as mutation and crossover operators within evolutionary strategies to generate complete algorithm implementations, with SMAC hyperparameter optimization integrated into the evolutionary loop. Across nine evolutionary runs we generated approximately 900 algorithms and benchmarked them on twelve synthetic problems (ZDT, DTLZ, WFG) and three real-world engineering problems (RE), using a BoFire qParEGO implementation as a state-of-the-art Bayesian-optimization baseline. On the synthetic suite the strongest generated algorithm attains the highest mean normalized hypervolume (0.971, vs. 0.869 for qParEGO) while requiring roughly 60x less wall-clock time; a Friedman test with post-hoc analysis places the two in a single top-performing group, and per-problem tests find the generated algorithm significantly better than qParEGO on 7 of the 12 problems and never worse, matching state-of-the-art accuracy at an order-of-magnitude lower cost. On the three unseen real-world engineering problems a generated algorithm attains the best mean normalized hypervolume (0.985, vs. 0.971 for qParEGO)--significantly better than qParEGO on two of the three problems--at roughly 3.4x lower wall-clock cost, confirming that the gains transfer beyond the synthetic regime. LLM-driven evolutionary search can thus discover algorithm designs that achieve Pareto-efficient trade-offs difficult to reach through manual design.
EHR-MPC: 生成患者デジタル ツインを使用した敗血症治療のための推論時間制御
敗血症は死亡の主な原因ですが、最適な治療方針については依然として議論が続いています。既存の強化学習 (RL) アプローチは、敗血症治療のための固定戦略を学習するため、推論中に変化する臨床目的への適応性が制限されます。私たちは、生成電子医療記録 (EHR) モデルの形式で患者のデジタル ツインをトレーニングすることで、患者のダイナミクスの学習と治療の最適化を切り離すフレームワークである EHRMPC を提案します。デジタル ツインは介入中の臨床経過を予測し、モデル予測制御 (MPC) を可能にして、シミュレーションによる推論時間計画を通じて治療を最適化します。我々は、ポリシー外の重要度サンプリングとポリシー上のシミュレーションベースの評価の両方を使用して、マサチューセッツジェネラルブリガム医療システムの8つの病院にわたる多施設ICU敗血症コホートでEHR-MPCを評価します。 RL ベースラインと比較して、EHR-MPC は同等のオフポリシー パフォーマンスと改善されたシミュレーション パフォーマンスを実現します。 RL とは異なり、この作業は敗血症治療の最適化を学習された患者の動態に対する推論時間の制御として枠組み化し、生成臨床モデルを使用した意思決定のための一般的な枠組みを確立します。
原文 (English)
EHR-MPC: Inference-Time Control for Sepsis Treatment with Generative Patient Digital Twins
Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested. Existing reinforcement learning (RL) approaches learn fixed strategies for sepsis treatment, limiting adaptability to changing clinical objectives during inference. We propose EHRMPC, a framework that decouples learning patient dynamics from optimizing treatment by training a patient digital twin in the form of a generative electronic health record (EHR) model. The digital twin predicts clinical trajectories under interventions and enables model predictive control (MPC) to optimize treatments via inference-time planning over simulations. We evaluate EHR-MPC on a multicenter ICU sepsis cohort spanning 8 hospitals in the Mass General Brigham health system using both off-policy importance sampling and on-policy simulation-based evaluation. Relative to RL baselines, EHR-MPC achieves comparable off-policy performance and improved simulation performance. Unlike RL, this work frames sepsis treatment optimization as inference-time control over learned patient dynamics, establishing a general framework for decision making with generative clinical models.
低資源土堤検査のための沸騰砂の多条件拡散合成
土の堤防での砂の沸騰は安全上重要な欠陥ですが、ピクセルレベルの検出は注釈が少ないため制限されます。低リソースのサンドボイル画像用の拡散ベースの合成パイプラインを紹介します。 DreamBooth で微調整され、マルチブランチ ControlNet スタックによって調整された Stable Diffusion XL を使用して、パイプラインは厳選された小規模のリファレンス セットから合成検査画像を生成します。ソフトマスク修復プロトコルは、周囲のシーンを再レンダリングしながら実際の欠陥ピクセルを保存し、以前のシームレスなクローン合成による継ぎ目やカラーシフトを回避します。マスク条件付き ControlNet は、選択したマスク内に新しいボイルを生成し、そのマスクを構築によりセグメンテーション ラベルにすることもできます。ただし、大規模なラベル認証は、利用可能な実際にトレーニングされたゲートでは未解決のままであるため、ソフトマスク プリセットをデフォルトとしてリリースします。テキスト コンディショニングは、分類ベースのプロンプト アトラスによって提供されます。プロンプト アトラスは、1 つのドメイン仕様を階層化された CLIP 検証済みのプロンプト バンクに拡張し、コードを変更することなく新しい欠陥クラスに移行します。実際のトレーニング画像から、パイプラインは 1,020 個の合成候補を生成し、そのうち 815 個が CLIP 許容フィルターを通過します。実際の参照セットとポアソンベースラインに対する分布および忠実度多様性の尺度を使用して画質を評価し、分布外のドリフトと記憶を監査します。単一のプリセットが優勢になることはありません。それぞれ、忠実度、多様性、ラベルの信頼性がトレードオフになります。したがって、ラベルの信頼性のあるプリセットをデフォルトとしてリリースし、厳選された混合物を自然な拡張セットとして扱います。私たちの主張は、画質、ラベルの出所、多様性に限定されています。下流のセグメンテーションは将来の作業に残されています。コードとアーティファクト マニフェストは再現性を確保するためにリリースされています。
原文 (English)
Multi-Conditioned Diffusion Synthesis of Sand Boils for Low-Resource Earthen-Levee Inspection
Sand boils on earthen levees are safety-critical defects, but pixel-level detection is limited by scarce annotations. We present a diffusion-based synthesis pipeline for low-resource sand-boil imagery. Using Stable Diffusion XL fine-tuned with DreamBooth and conditioned by a multi-branch ControlNet stack, the pipeline generates synthetic inspection images from a small curated reference set. A soft-mask inpainting protocol preserves the real defect pixels while re-rendering the surrounding scene, avoiding seams and color shifts from prior seamless-cloning compositing. A mask-conditioned ControlNet can also generate a new boil inside a chosen mask, making the mask the segmentation label by construction; however, because large-scale label certification remains unresolved with the available real-trained gate, we release the soft-mask preset as the default. Text conditioning is supplied by a taxonomy-driven Prompt Atlas that expands one domain specification into a stratified, CLIP-validated prompt bank and transfers to new defect classes without code changes. From the real training images, the pipeline produces 1,020 synthetic candidates, of which 815 pass a CLIP admissibility filter. We evaluate image quality using distributional and fidelity-diversity measures against the real reference set and a Poisson baseline, and audit for out-of-distribution drift and memorization. No single preset dominates; each trades off fidelity, diversity, and label reliability. We therefore release the label-reliable preset as the default and treat a curated mixture as the natural augmentation set. Our claims are limited to image quality, label provenance, and diversity; downstream segmentation is left for future work. Code and an artifact manifest are released for reproducibility.
TheBioCollection: 生物学用の統合事前トレーニング スケール LLM コーパス
生物学のための大規模言語モデル (BioLM) の推進により、モデルに生物学の真の理解を与えることができるトレーニング コーパスの必要性が生じています。しかし、分子データベース、タンパク質リポジトリ、ゲノムアノテーション、単細胞アトラス、経路データベースなどの既存の生物学的リソースは、異種の形式に分散しており、言語モデルのトレーニング用のまとまりのあるコーパスに編成されていないままです。我々は、これらの異種リソースを、小分子、タンパク質、ゲノム配列、細胞、経路にわたる統一されたトレーニング対応形式に変換する、526億トークンのトレーニング前規模のコーパスであるTheBioCollectionを提示します。 Beyond consolidating existing data, TheBioCollection enriches each record with tool-computed biological properties and introduces new instruction tasks for capabilities that current corpora barely cover.私たちはコーパスを、分子、タンパク質、ゲノム、細胞、およびクロスドメイン設定全体にわたって認識、生成、予測を精査する一致するスイートである TheBioCollection-Eval と組み合わせます。基本の Gravity-16B-A3B アーキテクチャを固定したまま、TheBioCollection でトレーニングすると、一般的な言語能力をほぼそのままにしながら、TheBioCollection-Eval の全体的なスコアが 2 倍以上になり、すべてのドメインで向上します。
原文 (English)
TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology
The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology. However, existing biological resources, such as molecular databases, protein repositories, genomic annotations, single-cell atlases, and pathway databases, are scattered across heterogeneous formats and remain unorganized into a cohesive corpus for language model training. We present TheBioCollection, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways. Beyond consolidating existing data, TheBioCollection enriches each record with tool-computed biological properties and introduces new instruction tasks for capabilities that current corpora barely cover. We pair the corpus with TheBioCollection-Eval, a matched suite probing recognition, generation, and prediction across molecular, protein, genomic, cellular, and cross-domain settings. Holding the base Gravity-16B-A3B architecture fixed, training on TheBioCollection more than doubles its overall score on TheBioCollection-Eval with gains in every domain, while leaving general linguistic ability nearly intact.
迅速な探索
すでに好まれている動作を繰り返しサンプリングすることによってポリシーを改善することはできないため、探索は RL にとって不可欠です。標準的な方法ではアクション空間に確率性を注入しますが、そのようなジッターはオリジナルに近いロールアウトしか生成しません。弱い政策から逃れるには、多くの場合、アクションのノイズでは生成できないグローバルな摂動が必要になります。大規模言語モデル (LLM) とビジョン言語アクション (VLA) モデルは、経路を提供します。自然言語プロンプトに基づいてポリシーを条件付けし、ロールアウトは自然言語プロンプトに基づいて行われるため、プロンプトを変更するとグローバルな変更が生じます。課題は、有益なグローバルな変化を引き起こすプロンプトを見つけることです。めったに成功しない弱い政策では、報酬は選択するにはあまりにも希薄です。私たちのアイデアは、ロールアウト自体からのプロンプトを改良することです。ビジョン言語モデル (VLM) がロールアウト ビデオを検討し、ポリシーがどのように反応したかを診断し、次回より良い動作を引き出すためにプロンプトを書き換えます。この手順は、古典的な RL 探索フレームワークである事後サンプリングをプロンプトのレベルで実現します。VLM は有用なプロンプトにわたる暗黙的な分布を維持し、観察されたロールアウトからそれを更新します。私たちはこの戦略を Prompt-Driven Exploration (PDE) と呼んでいます。操作タスクと推論タスクにわたって、PDE を使用すると、報酬ゼロのスタートからでも RL が成功したポリシーを学習できるようになり、サンプル効率がより広範囲に向上します。私たちの Web サイトは https://xinyunsunshine.github.io/prompt-rl でご覧いただけます。
原文 (English)
Prompt-Driven Exploration
Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject stochasticity in the action space, but such jitter only yields rollouts close to the original. Escaping a weak policy often requires global perturbations that action noise cannot produce. Large language models (LLMs) and vision-language-action (VLA) models offer a pathway: they condition the policy on a natural language prompt, and since the rollout follows from it, modifying the prompt induces global changes. The challenge is finding prompts that induce useful global changes. With a weak policy that rarely succeeds, reward is too sparse to select on. Our idea is to refine prompts from the rollouts themselves: a vision-language model (VLM) reasons over the rollout video, diagnoses how the policy responded, and rewrites the prompt to elicit better behavior next time. This procedure realizes posterior sampling, a classical RL exploration framework, at the level of prompts: the VLM maintains an implicit distribution over useful prompts and updates it from observed rollouts. We call this strategy Prompt-Driven Exploration (PDE). Across manipulation and reasoning tasks, PDE enables RL to learn successful policies even from zero-reward starts, and improves sample efficiency more broadly. Our website is available at https://xinyunsunshine.github.io/prompt-rl.
効率的な古典的シミュレーション機能を備えた新しい並列 QCNN アーキテクチャ
この研究では、修正された国立標準技術研究所 (MNIST) データセットからの画像のバイナリ分類のための新しい量子畳み込みニューラル ネットワーク (QCNN) の実装に関する研究を紹介します。以前の QCNN と古典的な畳み込みニューラル ネットワーク (CNN) の実装からインスピレーションを得た新しいアーキテクチャを使用し、階層分割アプローチを使用して、大規模な問題に対して古典的なマシン上で効率的に近似およびシミュレーションできる QCNN 回路を実装します。まず、元の画像が分割され、各プロセスが画像のより小さな部分を処理し、独立した状態にエンコードされます。次に、これらのパーティションがマージおよび結合され、両方のパーティションからの情報を含む状態が得られ、プロセスの数は半分になります。これをプロセスが 1 つ残るまで繰り返した後、測定対象の量子ビットが 1 つ残るまで状態の次元を削減します。このアプローチを使用すると、量子ビットの数が増加するにつれてハードウェア要件が指数関数的に増大することなく、複数のプロセスを並行して使用して大規模な QCNN プログラムをシミュレートできます。私たちの研究では、このスキームを使用して 128 量子ビット モデルをトレーニングしますが、新しいアーキテクチャがなければ古典的なスーパーコンピューターで実行することは不可能です。また、少数の量子ビットを使用して MNIST データセットに対してバイナリ分類を実行するようにモデル アーキテクチャをトレーニングし、分割なしのモデルと比較することで、この新しいモデル アーキテクチャが予測精度に及ぼす影響を調査します。私たちの最初の調査結果では、このアーキテクチャでイメージをより小さなサブイメージに分割すると、モデルのパフォーマンスが低下せず、場合によってはパフォーマンスが向上することさえあります。これはおそらく、分割プロセスにおける不毛プラトーの問題が軽減されるためです。
原文 (English)
A Novel Parallel QCNN Architecture with Efficient Classical Simulability
This work presents a study of an implementation of a novel Quantum Convolutional Neural Network (QCNN) for binary classification of images from the Modified National Institute of Standards and Technology (MNIST) dataset. Using a novel architecture inspired by previous QCNN and classical convolutional neural network (CNN) implementations, we use a hierarchical partitioning approach to implement a QCNN circuit that can be approximated and simulated efficiently on a classical machine for a large problem. First, the original image is partitioned such that each process handles a smaller portion of the image, which is encoded into independent states. Then, these partitions merge and combine, resulting in states that contain information from both partitions while halving the number of processes. After repeating this until one process remains, we reduce the dimensionality of the state until a single qubit remains for measurement. Using this approach, we can use multiple processes in parallel to simulate a large QCNN program without the need for exponentially growing hardware requirements as the number of qubits increases. In our work, we use this scheme to train a 128-qubit model, which is impossible to run on any classical supercomputer without the novel architecture. We also explore the impact of this new model architecture on prediction accuracy by training it to perform binary classification on the MNIST dataset with a small number of qubits, and comparing it to a model without partitioning. Our initial findings show that partitioning images into smaller sub-images with this architecture does not degrade the model's performance and sometimes even improves it, likely because it reduces the Barren plateaus issue in the partitioning process.
Eluna: 推論とタスク実行により倉庫業務を自動化するエージェント LLM システム
倉庫業務は、複雑なマルチシステムの意思決定ロジックをエンコードした標準運用手順 (SOP) によって管理されており、厳格な時間制約の下で確実に実行する必要がありますが、LLM エージェントには手順の遵守を強制するメカニズムが欠けており、完全な SOP 仕様が導入するコンテキストの過負荷の下では機能が低下します。信頼性の高い SOP 実行のための実稼働環境に導入されたエージェント システムである Eluna を紹介します。 Eluna は、段階的な開示を備えた有向非循環グラフとして SOP をエンコードし、独立したタスクを並列サブエージェントに委任する、グラフガイド型のマルチエージェント フレームワークです。各サブエージェントは永続的なコード実行とライブ データ アクセスを備えています。本番のレイテンシと精度のニーズを満たすために、非対称のエピソード蒸留を使用します。この方法では、強力な教師がエピソード的なエラー記憶を通じて改善され、その後、小さな生徒が記憶を取り除いて修正された軌道に基づいて微調整され、推論時間のオーバーヘッドなしで修正を内部化します。 13 タスクのベンチマークと 2 つの本番アプリケーションで、当社の微調整されたモデルは教師と同等かそれを上回り、より大きな既製のベースラインをすべて上回り、チケット処理アプリケーションに関して専門家による 94% の合意に達しました。
原文 (English)
Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution
Warehouse operations are governed by Standard Operating Procedures (SOPs) that encode complex, multi-system decision logic, which must be executed reliably under strict time constraints, yet LLM agents lack mechanisms to enforce procedural compliance and degrade under the context overload full SOP specifications introduce. We present Eluna, a production-deployed agentic system for reliable SOP execution. Eluna is a graph-guided, multi-agent framework that encodes SOPs as directed acyclic graphs with progressive disclosure and delegates independent tasks to parallel sub-agents, each with persistent code execution and live data access. To meet production latency and accuracy needs, we use asymmetric episodic distillation where a strong teacher is improved through episodic error memories, then a smaller student is fine-tuned on the corrected trajectories with memory stripped, internalizing corrections without inference-time overhead. On a 13-task benchmark and two production applications, our fine-tuned models match or exceed their teacher, beat all larger off-the-shelf baselines, and reach 94% expert agreement on the ticket processing application.
NL-PAC: LLM 仲介監督における仕様の曖昧さと認定された Minimax リスクフロア
大規模な言語モデルでは、自然言語で指定されたタスクに対するラベル、評価、フィードバックが提供されることが増えています。仕様では複数の読み取り値が認められているが、どれが有効であるかが監視チャネルで明らかにされていない場合、ラベルを追加すると、結果として生じる識別問題は解決されずにサンプリング エラーが減少します。 Natural Language PAC (NL-PAC) を紹介します。これは、固定モデルの閾値デコード法則を使用して、許容可能なラベルと候補ターゲットを定義するフレームワークです。複数のラベルが許容される確率は、点ごとに許容されるターゲットクラスの直径に等しく、ターゲットブラインド監視下では、すべての学習者は、すべてのサンプルサイズで、少なくともこの直径の半分の最悪の場合のリスクを負います。このクラスに対する正確なランダム化されたミニマックス リスクは、データに依存しない戦略によって達成されます。有限サンプルの信頼限界により、保持されたラベルなしの入力からこれらの量が証明可能になります。凍結された Qwen~2.5--3B 監査では、事前に指定された 1 つのプロンプトからは肯定的なモデル相対証明書が得られますが、言い換えと正確なルールのコントロールからはゼロが得られます。開催されたブリッジ監査により、提供された候補読み取り条項が、証明書を一貫した読み取りに移行するために必要な許容条件を満たしていないことが判明しました。保証は、監査対象のモデル、プロンプト、しきい値、および入力分布に固有です。それを人間の解釈に拡張するには、外部の検証が必要です。
原文 (English)
NL-PAC: Specification Ambiguity and Certified Minimax Risk Floors in LLM-Mediated Supervision
Large language models increasingly provide labels, evaluations, and feedback for tasks specified in natural language. When a specification admits multiple readings but the supervision channel does not reveal which is operative, additional labels reduce sampling error without resolving the resulting identification problem. We introduce Natural Language PAC (NL-PAC), a framework that uses a fixed model's thresholded decoding law to define admissible labels and candidate targets. The probability that multiple labels are admissible equals the diameter of the pointwise-admissible target class, and under target-blind supervision every learner incurs worst-case risk of at least half this diameter, at every sample size; the exact randomized minimax risk over this class is attained by a data-independent strategy. Finite-sample confidence bounds make these quantities certifiable from held-out unlabeled inputs. In a frozen Qwen~2.5--3B audit, one prespecified prompt yields a positive model-relative certificate, whereas a paraphrase and exact-rule controls yield zero. A held-out bridge audit finds that supplied candidate reading clauses fail the admissibility condition needed to transfer the certificate to coherent readings. The guarantee is specific to the audited model, prompt, threshold, and input distribution; extending it to human interpretations requires external validation.
MultiView-Bench: VLM における世界中心のマルチビュー統合のための診断ベンチマーク
VLM の最近のベンチマークは主に単一ビューまたは限定ビューの知覚を評価しており、複数の視点にわたる観察を一貫した世界中心 (他中心) の 3D メンタル モデルに統合する中核となる認知能力がテストされていないままになっています。 MultiView-Bench を紹介します。これは、全体的な 3D シーンを理解するためのマルチビュー統合を評価するために特別に設計された診断ベンチマークです。ピクセルレベルのマッピングやカメラ相対ナビゲーションに焦点を当てた既存のデータセットとは異なり、MultiView-Bench では、モデルが一時的な視点からオブジェクトの位置を切り離し、固定されたグローバル座標系に固定する必要があります。この機能は、VLM が機械部品の組み立てなどの下流タスクに展開される前の前提条件として機能します。フロンティア VLM の体系的な評価により、一貫した障害モードが明らかになりました。つまり、単一画像からの 2D 平面関係では優れたパフォーマンスを発揮しますが、3D 空間関係やビュー全体の情報の集約では顕著な困難が生じます。さらに、型破りな軸方向との闘いや、オブジェクトの色やテクスチャの変化に対する敏感さなど、VLM のバイアスを特定します。これらの制限を認識した上で、私たちは、有益な視点を積極的に選択し、複数の視点からの証拠を認識し、融合するマルチエージェント フレームワークである ViewNavigator を提案します。これにより、予算に合わせた厳密な比較の下でも (完全なエージェントの場合は 3 ~ 5 倍)、MultiView-Bench 上の多様な基本モデルが改善されます。
原文 (English)
MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs
Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coherent, world-centric (allocentric) 3D mental model. We introduce MultiView-Bench, a diagnostic benchmark expressly designed to evaluate multi-view integration for holistic 3D scene comprehension. Unlike existing datasets that focus on pixel-level mapping or camera-relative navigation, MultiView-Bench requires models to decouple object positioning from transient perspectives and ground them in a fixed global coordinate system. This capability serves as a prerequisite for VLMs before being deployed for downstream tasks such as mechanical part assembly. Our systematic evaluation of frontier VLMs reveals consistent failure modes: strong performance on 2D planar relations from a single image, but marked difficulty with 3D spatial relations and with aggregating information across views. We further identify biases in VLMs, such as struggles with unconventional axis directions and sensitivity to object colorways and texture variations. Acknowledging these limitations, we propose ViewNavigator, a multi-agent framework that actively selects informative viewpoints, perceives, and fuses multi-view evidence, improving diverse base models on MultiView-Bench even under a strict budget-matched comparison (and by 3-5x for the full agent).
CLAP: 言語アクショングラウンディングによる VLM から VLA への直接適応
ビジョン言語アクション モデル (VLA) は、事前トレーニングされた VLM からセマンティック機能を継承しますが、ロボット データに関する大規模なポストトレーニングやアーキテクチャの変更によってバックボーンが大幅に再形成される可能性があるため、VLM が制御に何を貢献しているかを分離することが困難になります。アーキテクチャの変更を最小限に抑えて、事前トレーニング済み VLM を VLA に直接変換することで、VLM 機能がモデル スケール間でどのように移行するかを理解するためのより透過的なパスが提供されます。中心的な障害は出力分布の不一致です。アクションを裸の数値トークン シーケンスとして予測すると、世代が VLM の事前学習済み言語分布から遠ざかり、保持しようとしている機能が低下します。これに対処するために、我々は CLAP (Causal Language-Action Prediction) を提案します。CLAP (Causal Language-Action Prediction) は、各数値アクション シーケンスの前に自然言語アクション記述を付加し、バックボーン アーキテクチャを変更することなく、言語アクション プランに基づいて正確なアクション トークン予測を因果的に条件付けします。単一エポックの微調整だけで、2B CLAP は LIBERO で 90.8% (VLA-0 に対して +14.9 ポイント) を達成し、言語、オブジェクト、空間の摂動下での LIBERO-PRO の堅牢性を向上させます。 CLAP は、単一の VLM 系統からのオープンウェイト、マルチスケールのコンパクト VLA ファミリとして 0.8B、2B、および 4B でリリースされ、VLM から VLA への機能移転の制御された分析が可能になります。
原文 (English)
CLAP: Direct VLM-to-VLA Adaptation via Language-Action Grounding
Vision-language-action models (VLAs) inherit semantic capabilities from pretrained VLMs, yet large-scale post-training on robot data and architectural modifications can reshape the backbone so extensively that it becomes difficult to isolate what the VLM contributes to control. Directly converting pretrained VLMs into VLAs with minimal architectural change offers a more transparent path to understanding how VLM capabilities transfer across model scales. The core obstacle is output-distribution mismatch: predicting actions as bare numeric token sequences moves generation away from the VLM's pretrained language distribution, degrading the capabilities we seek to preserve. To address this, we propose CLAP (Causal Language-Action Prediction), which prepends each numeric action sequence with a natural-language action description, causally conditioning precise action-token prediction on a language-action plan without modifying the backbone architecture. With single-epoch fine-tuning alone, 2B CLAP achieves 90.8% on LIBERO (+14.9 pt over VLA-0) and improves robustness on LIBERO-PRO under language, object, and spatial perturbations. We will release CLAP at 0.8B, 2B, and 4B as an open-weight, multi-scale compact VLA family from a single VLM lineage, enabling controlled analysis of VLM-to-VLA capability transfer.
LLM で生成されたコードにおけるパッチワークの問題
LLM で生成されたコードは、多くの場合、コンパイルされ、テストに合格し、正しいように見えますが、デプロイされると壊れます。根本的な原因は、論理的ではなく構造的なものであることがよくあります。生成されたエンドポイントがプロジェクトで宣言されていない構成キーを参照しているか、インポートのターゲットがレジストリに存在しないパッケージであるか、新しいルートですべての兄弟エンドポイントに適用される認証ガードが省略されています。各パッチはローカルでは有効ですが、グローバルでは一貫性がなく、標準の CI ツールチェーンではこれらの障害が表面化することはほとんどありません。 LLM を利用したコーディング ツールが広く採用されるにつれ、この盲点がソフトウェアの品質に対するリスクを増大させています。これを \textbf{パッチワーク問題} と呼びます。この論文では、構造的一貫性を、インポート、呼び出し、依存関係、構成、スキーマ、リソース、制御フロー、およびルーティング グラフを含むリポジトリ アーティファクトのグラフ表現に対する一貫性の不変条件として形式化し、LLM 生成に特有の欠陥と LLM 生成によって単に増幅された欠陥を区別する 8 つのカテゴリの障害分類法を導入します。我々は、すでに優れている成熟した静的解析ツールに委任し、既存のツールチェーンでは対応できていない分野横断的な不変条件を検出する専用の検出器を展開するハイブリッド検証フレームワークを提案します。これは、ヒューリスティックなパターン マッチングではなく、証明可能な制約違反をターゲットとします。 4 つの促進戦略に基づく 2 つのフロンティア モデルにわたる経験的評価により、構造的故障の大部分がタイプ チェック、テスト、および SAST を完全に回避していること、およびモデルに依存しない緩和戦略に挑戦する方法で故障パターンがモデル間で定性的に異なることが明らかになりました。現実世界の AI によって生成されたリポジトリの外部検証により、これらの障害は制御された実験の結果ではなく、LLM が人間の監視を最小限に抑えてコードを作成する場所ではどこでも蔓延していることが確認されました。
原文 (English)
The Patchwork Problem in LLM-Generated Code
LLM-generated code often compiles, passes tests, and appears correct, yet breaks once deployed. The root cause is frequently structural rather than logical. A generated endpoint references configuration keys never declared in the project, an import targets a package that does not exist in any registry, or a new route omits the authentication guard applied to every sibling endpoint. Each patch is locally valid but globally incoherent, and standard CI toolchains rarely surface these failures. As LLM-powered coding tools see widespread adoption, this blind spot poses a growing risk to software quality. We call this the \textbf{patchwork problem}. This paper formalizes structural coherence as consistency invariants over graph representations of repository artifacts, including import, call, dependency, configuration, schema, resource, control-flow, and routing graphs, and introduces an eight-category failure taxonomy distinguishing defects specific to LLM generation from those merely amplified by it. We present a hybrid verification framework that delegates to mature static analysis tools where they already excel and deploys purpose-built detectors for cross-cutting invariants underserved by existing toolchains, targeting provable constraint violations rather than heuristic pattern matching. Empirical evaluation across two frontier models under four prompting strategies reveals that the vast majority of structural failures evade type checking, testing, and SAST entirely, and that failure patterns diverge qualitatively between models in ways that challenge model-agnostic mitigation strategies. External validation on real-world AI-generated repositories confirms that these failures are not artifacts of controlled experimentation but are prevalent wherever LLMs write code with minimal human oversight.
SCATE: コスト効率の高いテスト生成のためのコーディング エージェントの監督方法を学ぶ
自律型コーディング エージェントは自動テスト生成を大幅に進化させていますが、遅延生成という根本的な制限が残っています。遅延生成とは、エージェントがタスクを途中で終了し、複雑なプログラム ロジックを系統的に回避する現象で、その結果、コード カバレッジが不十分になります。現在、この早期終了を軽減するには、人間による継続的な監視が必要です。人間の直感に大きく依存しているため、自動生成による効率向上を妨げるボトルネックが生じています。私たちは、テスト生成時の人間の介入に代わる、コーディング エージェントの適応的で自動化された監視のためのフレームワークである SCATE を提案します。状況に応じたバンディット問題として監視を定式化することで、SCATE は現在のカバレッジとクラスのテスト容易性メトリクスに基づいて最も有望なテスト アクションを選択することを学習し、無駄な生成作業を最小限に抑えながらカバレッジの向上を最大化します。私たちの経験的評価により、SCATE がさまざまなコーディング エージェントとシームレスに統合されることが実証されています。 GEMINI-CLI に適用すると、エージェントのみのベースラインよりも 32.3% 高い回線カバレッジと 30.9% 高いブランチ カバレッジを達成します。 CLAUDE CODE と比較すると、フレームワークがポリシーを動的に適応させて各エージェントの固有の強みを最適化していることが確認できます。また、SCATE は、すべての指標において最先端の非エージェント的アプローチを常に上回っています。
原文 (English)
SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation
While autonomous coding agents have significantly advanced automated test generation, they remain fundamentally limited by lazy generation, a phenomenon where agents prematurely terminate tasks and systematically avoid complex programmatic logic, resulting in inadequate code coverage. Currently, mitigating this premature termination requires continuous human-in-the-loop supervision. This heavy reliance on human intuition creates a bottleneck that negates the efficiency gains of automated generation. We propose SCATE, a framework for adaptive, automated supervision of coding agents that replaces human intervention during test generation. By formulating supervision as a contextual bandit problem, SCATE learns to select the most promising testing actions based on the current coverage and class testability metrics, maximizing coverage gains while minimizing wasted generation effort. Our empirical evaluation demonstrates that SCATE integrates seamlessly with different coding agents. When applied to GEMINI-CLI, it achieves 32.3% higher line coverage and 30.9% higher branch coverage than the agent-only baseline. A comparison with CLAUDE CODE confirms the framework dynamically adapts its policy to optimize each agent's unique strengths. SCATE also consistently outperforms state-of-the-art non-agentic approaches across all metrics.
報酬の少ないゲームにおける AlphaZero: 制限と補助監督
AlphaZero は、神経誘導モンテカルロ木探索が超人的なパフォーマンスを達成できることを実証しましたが、強いプレイが必ずしも完璧なプレイを意味するわけではありません。私たちは、対照的な構造を持つ 2 つのオラクル評価可能な領域におけるこのギャップを研究します。Connect Four は、正確なゲーム理論値を備えた解決済みの党派ゲームであり、Chomp は、最適なプレイがグランディ数構造によって支配される公平なゲームです。統合されたセルフプレイ $+$ MCTS パイプラインの下で、バニラ AlphaZero、マルチフレーム バリアント (Chomp に限定)、およびオラクル由来のポリシー監視を追加する AlphaZero Auxiliary Loss (AZAL) を比較します。バニラの AlphaZero は両方の領域で強力なプレーを実現しますが、最適なプレーに必要な正確な軌道を保持できないことがわかりました。Connect Four では最適なプレーのラインを維持できませんが、Chomp では $g=0$ の不変条件を一貫して復元できません。長方形の Chomp ボードでは、マルチフレーム入力だけではこのギャップは解消されません。それにもかかわらず、AZAL は、マルチシードのフルゲーム トレースとサンプリングされた状態の評価にわたるオラクルの一貫性を大幅に向上させます。 Chomp では、AZAL は 10x11 で完全なフルゲーム オラクルの一貫性を達成し、9x10 では高いが完全ではない一貫性に達します。 Connect Four では、AZAL はオラクル一致率を向上させ、最初のオラクルミスを遅らせますが、完璧なプレイには達しません。
原文 (English)
AlphaZero in Sparsely Rewarded Games: Limits and Auxiliary Supervision
AlphaZero has demonstrated that a neural-guided Monte Carlo Tree Search can achieve superhuman performance, but strong play does not necessarily imply perfect play. We study this gap in two oracle-evaluable domains with contrasting structure: Connect Four, a solved partisan game with exact game-theoretic values, and Chomp, an impartial game whose optimal play is governed by Grundy-number structure. Under a unified self-play $+$ MCTS pipeline, we compare vanilla AlphaZero, a multi-frame variant (limited to Chomp), and an AlphaZero Auxiliary Loss (AZAL) that adds oracle-derived policy supervision. We find that vanilla AlphaZero achieves strong play across both domains but cannot preserve the exact trajectories required for optimal play: in Connect Four, it fails to maintain the optimal line of play, while in Chomp, it fails to consistently restore the $g=0$ invariant. On rectangular Chomp boards, multi-frame inputs alone do not remove this gap. Nevertheless, AZAL substantially improves oracle consistency across multi-seeded full-game traces and sampled-state evaluations. On Chomp, AZAL reaches perfect full-game oracle consistency on 10x11 and high but not complete consistency on 9x10; on Connect Four, AZAL improves oracle-match rate and delays the first oracle mistake, but does not reach perfect play.
結晶特性予測のためのモデルに依存しないグラフの即時学習
グラフ ニューラル ネットワークは、さまざまな結晶特性を高速かつ正確に予測するための強力なツールとして登場しました。これらのモデルは多くの場合、ドメイン固有の知識をグラフ エンコード モジュールにエンコードするため、パラメーター サイズが増加し、パフォーマンスがドメインの専門知識に大きく依存します。これに加えて、特定の結晶特性に影響を与える可能性のあるすべての化学的および構造的特徴を GNN エンコーダに明示的に組み込むことは、困難な作業です。この研究では、GNN に明示的に提供されていない、プロパティ予測に不可欠な潜在的な特徴を捕捉するソフト プロンプト学習フレームワークを提案します。ノードレベルとグラフレベルの両方のソフトプロンプトを含む、新しいマルチレベルグラフプロンプト学習フレームワークを導入します。ノード レベルでは、さまざまな原子タイプの局所的な化学セマンティクスをキャプチャし、グラフ レベルでは、結晶グラフのグローバルな構造対称性をエンコードします。私たちが提案するプロンプト学習フレームワークは軽量で、既存の GNN エンコーダーとシームレスに統合されます。一般的なベンチマーク データセットに対する広範な実験により、プロンプト学習を組み込むと、結晶特性予測タスクにおける最先端の GNN モデルのパフォーマンスが大幅に向上する (3\% ~ 15\%) ことがわかりました。さらに、学習されたソフト プロンプトにより、プロパティ間の知識の伝達が可能になり、トレーニング データが限られているプロパティの予測パフォーマンスが向上します。コードは https://github.com/shrimonmuke0202/Prompt.git で入手できます。
原文 (English)
Model Agnostic Graph Prompt Learning for Crystal Property Prediction
Graph Neural Networks have emerged as a powerful tool for the fast and accurate prediction of various crystal properties. These models often encode domain-specific knowledge into their graph encoding modules, which increases their parameter size and makes their performance heavily dependent on domain expertise. Added to this, explicitly incorporating all chemical and structural features, that might influence a specific crystal property into the GNN encoder, is a challenging task. In this work, we propose a soft prompt learning framework that captures latent features essential for property prediction, which are not explicitly provided to the GNN. We introduce a novel multilevel graph prompt learning framework comprising both node-level and graph-level soft prompts. At the node level, we capture the local chemical semantics of different atom types, while at the graph level, we encode the global structural symmetry of the crystal graph. Our proposed prompt learning framework is lightweight and seamlessly integrates with any existing GNN encoder. Extensive experiments on popular benchmark datasets show that incorporating prompt learning significantly improves (3\% - 15\%) the performance of state-of-the-art GNN models in crystal property prediction tasks. Furthermore, the learned soft prompts enable cross-property knowledge transfer, enhancing prediction performance for properties with limited training data. Code is available at https://github.com/shrimonmuke0202/Prompt.git
LLM ルーティングの代理報酬を備えた相関関係を意識したコンテキスト バンディット
私たちは、相関アームを使用したコンテキスト バンディット問題と、大規模言語モデル (LLM) ルーティングなどのアプリケーションによって動機付けられた機械学習モデルによって生成された代理報酬信号へのアクセスを研究します。バンディットのフィードバックのみに依存し、アーム全体で条件付きの独立性を前提とする古典的なコンテキスト バンディットとは異なり、私たちの設定では、コンテキストに依存したアーム間の相関と、ノイズが多かったり指定が間違っていたりする可能性のある補助的な報酬情報が可能になります。私たちは、2 つの相補的な設計を通じてそのような代理報酬を活用するアルゴリズムを提案します。結合された報酬混合アプローチでは、真の報酬と代理報酬をプールして、代理信号が信頼できる場合に学習を加速します。一方、分離された予測混合アプローチでは、バンディット フィードバックと代理報酬の個別の推定量を維持し、それらの予測を適応的に組み合わせます。この分離により、サロゲートの誤指定に対する堅牢性がもたらされ、最悪の場合には報酬のみのバンディット手法に匹敵するリグレス保証が回復されますが、サロゲート予測が十分な情報を提供する場合にはリグレスの改善が実現されます。私たちは両方のアプローチに対して理論的なリグレス分析を提供し、さまざまな精度とコストのトレードオフの下で LLM ルーティング ベンチマークでそれらを評価します。この結果は、標準のコンテキスト バンディット ベースラインや強力な静的ルーティング手法と比較して、サンプル効率が向上し、精度とコストのトレードオフが一貫して優れていることを示しています。
原文 (English)
Correlation-Aware Contextual Bandits with Surrogate Rewards for LLM Routing
We study contextual bandit problems with correlated arms and access to surrogate reward signals produced by a machine learning model, motivated by applications such as large language model (LLM) routing. Unlike classical contextual bandits that rely solely on bandit feedback and assume conditional independence across arms, our setting allows context-dependent inter-arm correlations and auxiliary reward information that may be noisy or misspecified. We propose algorithms that leverage such surrogate rewards through two complementary designs. A coupled reward-mixing approach pools true and surrogate rewards to accelerate learning when surrogate signals are reliable, while a decoupled prediction-mixing approach maintains separate estimators for bandit feedback and surrogate rewards and adaptively combines their predictions. This decoupling yields robustness to surrogate misspecification, recovering regret guarantees comparable to reward-only bandit methods in the worst case, while achieving improved regret when surrogate predictions are sufficiently informative. We provide theoretical regret analyses for both approaches and evaluate them on LLM routing benchmarks under varying accuracy versus cost trade-offs. The results demonstrate improved sample efficiency and consistently better accuracy-cost trade-offs compared to standard contextual bandit baselines and strong static routing methods.
音韻活性化マッピングによる音声のセグメンテーションと認識
電話のセグメンテーションと認識は本質的に関連するタスクですが、最新のアプローチでは通常、それらを別々にモデル化します。私たちは、音声構造は自己教師あり音声モデル (S3M) の表現にすでに潜在しており、両方のタスクを解決するためにモデルを操作するだけでよいと主張します。当社は、S3M ベースの音韻活性化マッピング (SPAM) を活用し、各 S3M 表現フレームを、発声や鼻声などの音韻的特徴活性化のベクトルにマッピングします。 SPAM に加えて、認識ヘッドとセグメンテーション ヘッドという 2 つのシンプルだが効果的な軽量で勾配降下フリーの予測ヘッドを導入します。私たちの方法では、音声の書き起こしが 1 分未満で済み、トレーニング中に目に見えない電話に一般化されます。多様なデータセットにわたって、私たちのアプローチは強力なセグメンテーションと認識パフォーマンスを実現します。
原文 (English)
Phone Segmentation and Recognition through Phonological Activation Mapping
Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both tasks. We leverage S3M-based Phonological Activation Mapping (SPAM), which maps each S3M representation frame to a vector of phonological feature activations, such as voicing and nasality. On top of SPAM, we introduce two simple but effective lightweight, gradient-descent-free prediction heads: a recognition head and a segmentation head. Our method requires less than a minute of phonetic transcriptions, and generalizes to unseen phones during training. Across a diverse range of datasets, our approach attains strong segmentation and recognition performance.
ビデオ生成モデルは汎用の視覚学習器です
NLP はネクスト トークン予測によって推進され、タスク固有のモデルから強力なジェネラリスト基盤モデルに移行しました。では、コンピューター ビジョンで汎用モデルを実現するために必要な同等の触媒は何でしょうか?この論文では、大規模なテキストからビデオへの生成がコンピュータ ビジョンの強力な事前トレーニング パラダイムとして機能し、一般的な視覚知能に必要な時空間事前分布、視覚と言語の整合性、およびスケーラビリティを提供すると主張します。 GenCeption を紹介します。これは、事前トレーニングされたビデオ生成拡散バックボーンを活用してフィードフォワード知覚モデルを定義し、テキスト指示によって操作されるさまざまな視覚タスクを実行できます。実証結果は、GenCeption が深度、表面法線、カメラ姿勢推定、表情参照セグメンテーション、3D キーポイント予測などの多様なタスク全体にわたって最先端のパフォーマンスを達成し、多くの場合、特殊なモデル (DepthAnything3、SAM3、D4RT、VGGT-Omega、Sapiens、David、Genmo、Lotus-2) と同等またはそれを上回るパフォーマンスを実現していることを示しています。さらに、ビデオ生成事前トレーニング バックボーンは、同等の設定下で代替事前トレーニング パラダイム (V-JEPA やビデオ MAE など) よりも優れたパフォーマンスを発揮します。重要なのは、GenCeption は予備データとモデルのスケーリング特性を優れたデータ効率とともに示し、7 ~ 500 少ないトレーニング データで D4RT や VGGT-Omega などの主要モデルと同等のパフォーマンスを達成します。最後に、GenCeption は興味深い創発的な動作も示します。合成人間ビデオのみでトレーニングされたモデルは、現実世界の映像や配布外のオブジェクト カテゴリ (動物やロボットなど) に一般化されます。これらの発見は、ビデオ生成が単なる合成ツールではなく、物理世界のジェネラリストビジョンインテリジェンスへの基礎的な道であることを示唆しています。プロジェクトページ: https://genception.github.io
原文 (English)
Video Generation Models are General-Purpose Vision Learners
Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video generation serves as a strong pre-training paradigm for computer vision, providing the necessary spatiotemporal priors, vision-language alignment, and scalability required for general visual intelligence. We introduce GenCeption, which leverages a pre-trained video generative diffusion backbone to define a feed-forward perception model, capable of performing various vision tasks steered by text instructions. Empirical results demonstrate that GenCeption achieves state-of-the-art performance across a diverse suite of tasks, including depth, surface normal, and camera pose estimation, expression-referring segmentation, and 3D keypoint prediction, often matching or surpassing specialized models (e.g. DepthAnything3, SAM3, D4RT, VGGT-Omega, Sapiens, David, Genmo, and Lotus-2). Furthermore, the video generative pretrained backbone outperforms alternative pretraining paradigms (e.g., V-JEPA, and Video MAE) under comparable settings. Importantly, GenCeption exhibits preliminary data and model scaling properties along with exceptional data efficiency, where it achieves comparable performance with leading models like D4RT and VGGT-Omega with 7 to 500 less training data. Finally, GenCeption also exhibits intriguing emergent behaviors: a model trained exclusively on synthetic human videos generalizes to real-world footage and out-of-distribution object categories (e.g., animals and robots). These findings suggest that video generation is not merely a synthesis tool, but a foundational path toward generalist vision intelligence for the physical world. Project page: https://genception.github.io
科学的発見のための進化的インテリジェンス: 進化的計算から累積的発見システムまで
人工知能 (AI) は、科学的発見をタスク固有のワークフローから、オープンエンドの候補空間での実験と人間のフィードバックによって探索を組織化する自律システムへと移行させています。進化計算 (EC) は、集団ベースの探索により、蓄積された証拠を通じて探索を進めながら、多様な科学的候補を維持できるため、フィードバック駆動型の発見のための計算基盤を提供します。ただし、EC は主に事前定義された問題に対する候補の絞り込みに焦点を当てますが、累積的な発見には経験の保持が必要です。このギャップを埋めるために、このレビューでは科学的発見のための進化的知能 (EI) を紹介します。 EI は、候補の改良と進化サイクル全体にわたる経験の保持を結び付けることで探査を維持する科学 AI システムを特徴付けます。私たちは、何が進化するのか、候補者がどのように変化するのか、候補者が選ばれる理由、フィードバックがどこから発生するのか、いつ進化が起こるのかを問う 5 次元の分析フレームワークを導入します。このフレームワークは、EI が孤立した探索軌跡を累積的な科学的洞察にどのように変換するかを明らかにします。さらに、進化する具体的な科学実体から自動化された研究ワークフローの調整まで、さまざまな発見モードにわたってこのパラダイムを実証します。最後に、評価、プロセスのトレーサビリティ、共有インフラストラクチャに関する重大なボトルネックを特定し、科学的発見における EC から EI への移行を進めるための具体的なロードマップを提供します。
原文 (English)
Evolutionary Intelligence for Scientific Discovery: From Evolutionary Computation to Cumulative Discovery Systems
Artificial intelligence (AI) is shifting scientific discovery from task-specific workflows towards autonomous systems that organize exploration with experimental and human feedback in open-ended candidate spaces. Evolutionary computation (EC) provides a computational basis for feedback-driven discovery because population-based search can maintain diverse scientific candidates while steering exploration through accumulated evidence. However, EC predominantly focuses on candidate refinement for predefined problems, whereas cumulative discovery requires experience retention. To bridge this gap, this review introduces evolutionary intelligence (EI) for scientific discovery. EI characterizes scientific AI systems that sustain exploration by linking candidate refinement with experience retention across evolutionary cycles. We introduce a five-dimensional analytical framework that asks what evolves, how candidates change, why candidates are selected, where feedback originates, and when evolution occurs. This framework clarifies how EI transforms isolated search trajectories into cumulative scientific insight. We further demonstrate this paradigm across diverse discovery modes, from evolving concrete scientific entities to orchestrating automated research workflows. Finally, we identify critical bottlenecks regarding evaluation, process traceability, and shared infrastructure, providing a concrete roadmap for advancing the transition from EC to EI in scientific discovery.
文脈の論理としての量子論理
量子論理は通常、古典論理を安全な出発点として保ち、量子力学によって私たちに強制される通常の推論からの非古典的な逸脱として提示されます。私たちは、有限で完全に計算可能な設定では、逆の順序で説明することを主張します。 2 つのジェネレーターの自由直交モジュール格子には 96 個の要素があり、これは 6 要素の非分配因子と 16 要素のブール因子の直接積です。最初の要素をコンテキストのレジスタとして読み取り、2 番目の要素をブール値の内容として読み取ると、その要素がコンテキストとビットとベクトルのペアであり、その演算がコンポーネントごとに機能する微積分が得られます。この計算により、3 つの結果が得られます。まず、可換性によって 6 つの層を分類し、すべての相補的なコンテキストが存在する二重中心層とともに、コンテキスト中立命題の中心カーネルを特定します。第二に、小さな因子の補完がその要素を再配置するのとまったく同じように、オルト補完が層を再配置し、層間の二重性が偶然ではなく固定化されることを示します。第三に、文脈を忘れた操作は、商が古典的なブール代数であるオルト補格子の全射準同型であり、そのため古典論理は 6 対 1 の情報を失った文脈計算のイメージであることを証明します。
原文 (English)
Quantum Logic as the Logic of Contexts
Quantum logic is usually presented as a non-classical departure from ordinary reasoning forced on us by quantum mechanics, with classical logic kept as the secure starting point. We argue for the opposite order of explanation in a finite and fully computable setting. The free orthomodular lattice on two generators has ninety-six elements, the direct product of a six-element non-distributive factor and a sixteen-element Boolean factor. Reading the first factor as a register of contexts and the second as Boolean content, we obtain a calculus whose elements are context--bit-vector pairs and whose operations act component by component. With this calculus we establish three results. First, we classify the six layers by commutativity, identifying the central kernel of context-neutral propositions together with a dual central layer in which all complementary contexts are present. Second, we show that orthocomplementation rearranges the layers exactly as the complementation of the small factor rearranges its elements, which makes the duality among the layers rigid rather than accidental. Third, we prove that the operation forgetting the context is a surjective homomorphism of orthocomplemented lattices whose quotient is the classical Boolean algebra, so that classical logic is a six-to-one, information-losing image of the contextual calculus.
視覚的推論における局所性と長さの一般化について
人間の視覚システムの顕著な特徴は、単一のグローバルな計算ではなく、一連の局所的な中心窩の視線を通じて視覚情報を取り込むことです。このため、人間の視覚は、画像をグローバルかつワンショットで入力する、現在使用されている最も一般的なコンピューター ビジョン モデルとは明らかに異なります。したがって、当然の疑問は、ローカルな逐次視覚モデルが、グローバル モデルよりも生物学的にもっともらしいことに加えて、基本的な計算上の利点を提供できるかどうかということです。この研究では、視覚状態の追跡と長さの一般化の観点からこの疑問を調査します。言語モデルにおける長さの一般化に関する最近の研究に触発されて、私たちは、画像全体にわたるローカル情報の集約を必要とする単純な視覚タスクで訓練された視覚モデルの動作を研究します。私たちの実験では、言語モデルと同様に、ビジョン モデルもグローバル ショートカットを利用することを学習し、それによってタスクの長さや複雑さを一般化できないことが明らかになりました。また、厳密に局所的な認識に基づいた反復ビジョン ポリシーがこれらの失敗を軽減し、それによってモデルがこれらのタスクを一般化できることも示します。私たちの結果は、局所的な注意が堅牢な構成的一般化にとって重要な見落とされている要件である可能性があることを示しています。
原文 (English)
On Locality and Length Generalization in Visual Reasoning
A striking feature of the human visual system is that it ingests visual information through a series of local foveated glimpses, rather than a single global computation. This makes human vision distinctly different from most popular computer vision models in use today, which input images globally and in a single shot. A natural question therefore is whether local, sequential vision models may provide any fundamental computational benefits in addition to being biologically more plausible than global models. In this work, we investigate this question from the perspective of visual state tracking and length generalization. Inspired by recent studies of length generalization in language models, we study the behavior of vision models trained on simple vision tasks that require the aggregation of local information across an image. Our experiments reveal that, similar to language models, vision models can learn to exploit global shortcuts and thereby fail to generalize over task length or complexity. We also show that recurrent vision policies based on strictly local perception can mitigate these failures, thereby allowing models to generalize on these tasks. Our results show that local attention may be an essential overlooked requirement for robust compositional generalization.
スキル市場の内部: ソフトウェア エンジニアリング活動から再利用可能なエージェント スキルまで
ソフトウェア エンジニアリング (略称 SE) は、ソース コードやライブラリからコンポーネントやサービスに至るまで、ますます強力な再利用形式を通じて継続的に進化してきました。 AI エージェントの最近の進歩により、潜在的に新しい再利用可能なアーティファクトであるスキルが導入されました。新興エージェントのスキル リポジトリとマーケットプレイスにより、開発者は SE の専門知識を再利用可能なスキルとしてパッケージ化し、共有し、再利用できます。この傾向は、どのような SE アクティビティが再利用可能なスキルにカプセル化されているのかという根本的な疑問を引き起こします。既存の研究は主に幅広いスキルの習得、安全性、またはベンチマークに焦点を当てていますが、SE 固有のスキルとソフトウェア開発ライフサイクル全体にわたるその範囲についての体系的な理解が不足しています。このギャップに対処するために、私たちはパブリック リポジトリとマーケットプレイスにおける SE スキルに関する初の大規模実証研究を実施します。私たちは、SE スキルの大規模なコーパスを収集して分析し、それらが要約するアクティビティ、ライフサイクルの範囲、進化の特徴、評価メカニズムを調査します。私たちの調査結果は、SE活動がスキルを介して再利用可能な成果物になりつつあることを明らかにしており、スキルの推奨やエンジニアリング指向の構造化に関する有望な研究機会、およびハイコンテキストのSE活動を再利用可能なスキルにカプセル化するメカニズムの必要性を示唆しています。全体として、私たちの研究は、SE スキルの活動中心の特徴付けを初めて提供し、SE 活動がどのようにして再利用可能なスキルにますます変換されているかを明らかにしました。これらの調査結果は、スキルの再利用、エコシステム開発、エージェント中心の SE の将来についての新たな洞察を提供します。
原文 (English)
Inside the Skill Market: From Software Engineering Activities to Reusable Agent Skills
Software engineering (abbrev. SE) has continuously evolved through increasingly powerful forms of reuse, from source code and libraries to components and services. Recent advances in AI agents have introduced a potentially new reusable artifact: skills. Emerging agent skill repositories and marketplaces enable developers to package, share, and reuse SE expertise as reusable skills. This trend raises a fundamental question: what SE activities are being encapsulated into reusable skills? Existing studies primarily focus on a broad range of skills acquisition, safety, or benchmarking, while lacking a systematic understanding of SE-specific skills and their coverage across the software development lifecycle. To address this gap, we conduct the first large-scale empirical study of SE skills in public repositories and marketplaces. We collect and analyze a large corpus of SE skills, examining the activities they encapsulate, lifecycle coverage, evolution characteristics, and evaluation mechanisms. Our findings reveal that SE activities are increasingly becoming reusable artifacts via skills and suggest promising research opportunities for skill recommendation and engineering-oriented structuring, as well as the need for mechanisms to encapsulate high-context SE activities into reusable skills. Overall, our study provides the first activity-centric characterization of SE skills and reveals how SE activities are increasingly being transformed into reusable skills. These findings offer new insights into skill reuse, ecosystem development, and the future of agent-centric SE.
OmniMapBench: 多様な地図ドキュメントに対するビジュアル中心の推論のベンチマーク
LVLM の最近の進歩により、複雑で視覚に基づいた推論のための堅牢なベンチマークが必要になります。多くの文書理解ベンチマークでは重大な制限が確認されています。ビジュアル コンテンツは多くの場合テキストに還元できるため、真の視覚的根拠がなくても高いパフォーマンスが可能になります。この制限に対処するために、マップ ドキュメントの視覚中心の推論を促進するために OmniMapBench が導入されました。このベンチマークは、9 つのカテゴリの 1,603 のマップ ドキュメントにわたる、手動で注釈が付けられた 2,096 の質問と回答のペアで構成されています。知覚から複数ステップの視覚的推論に至るまで、スキルの階層を調査するように設計されています。ベンチマークのプロパティを定量化するために、シンプルかつ効果的なベンチマーク レベルの指標である視覚依存性インデックス (VDI) が提案されます。これは、画像が質問に依存しない説明に置き換えられた場合の精度の低下として定義されます。 OmniMapBench は確立されたベンチマークよりも高い VDI を示し、これは、還元不可能な視覚的推論に焦点を当てていることを定量的に検証します。 25 の主要な LVLM の包括的な評価が OmniMapBench で行われます。パフォーマンスに大きなギャップが見られ、最高パフォーマンスのモデルでも 75.03\% の精度しか達成できません。この結果は、OmniMapBench が現在の LVLM に課す課題を強調しています。この研究は、LVLM の文書理解のための視覚中心の推論の進歩を促進することを目的としています。データセットとコードは https://github.com/SIGMME/OmniMapBench で公開されています。
原文 (English)
OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents
Recent advancements in LVLMs necessitate robust benchmarks for complex, visually grounded reasoning. A critical limitation is identified in many document understanding benchmarks: visual content is often reducible to text, enabling high performance without genuine visual grounding. To address this limitation, OmniMapBench is introduced to foster visual-centric reasoning for map documents. The benchmark comprises 2,096 manually annotated question-answer pairs across 1,603 map documents from nine categories. It is designed to probe a hierarchy of skills, ranging from perception to multi-step visual reasoning. To quantify benchmark properties, a simple yet effective benchmark-level metric is proposed: the Visual Dependency Index (VDI), defined as the accuracy drop when images are replaced with question-agnostic descriptions. OmniMapBench exhibits higher VDI than established benchmarks, which quantitatively validates its focus on irreducible visual reasoning. Comprehensive evaluations of 25 leading LVLMs are conducted on OmniMapBench. A significant performance gap is observed, with the top-performing model achieving only 75.03\% accuracy. This result underscores the challenges posed by OmniMapBench to current LVLMs. This work aims to catalyze progress in visual-centric reasoning for document understanding of LVLMs. The dataset and code are publicly available at https://github.com/SIGMME/OmniMapBench.
PRecG: グラフニューラルネットワークと修辞的役割セグメンテーションによる判例検索
判例の検索は、訴訟の準備、計画、訴訟戦略、法的研究における基本的なタスクです。自動判例検索の現在のアプローチは、法的文書を低次元の意味空間にマッピングし、それらの表現の近接性に基づいて類似性を計算します。これらのアプローチは、法的技術の修辞的構成を無視して、法的文書を一枚岩の文書として扱います。したがって、彼らは微妙な法的意味を見落としており、文書内での修辞的役割に基づいて変化する法人や概念の文脈上の重要性を区別できません。この不十分さに対処するために、法的判決の表現を階層的に学習することによって、法的判決のペア間の類似性を計算する PRecG パイプラインを提案します。このプロセスは、文の修辞的役割に基づいて各文書を別個の意味単位 (セグメント) に分解することから始まります。レトリックセグメントごとに、セグメント内の法人とその関係を把握するためにナレッジグラフが構築されます。次に、エンティティのコンテキスト表現が学習および集約されて、セグメント レベルの埋め込みが導出されます。これらの埋め込みはさらに統合されて、統一された文書レベルの表現が生成され、最後に、一対の文書間の意味論的な類似性が計算されます。私たちは、ベンチマークとなるインドの法律データセットに対する広範な実験を通じて、提案されたアプローチのパフォーマンスを検証し、その有効性を実証するために最先端のベースラインと比較します。
原文 (English)
PRecG: Legal Precedent Retrieval with Graph Neural Networks and Rhetorical Role Segmentation
Legal precedent retrieval is a fundamental task in legal case preparation, planning, litigation strategy, and legal research. Current approaches for automatic precedent retrieval map legal documents to a low-dimensional semantic space and compute similarity based on the proximity of their representations. These approaches treat legal documents as monolithic texts, ignoring the rhetorical organization of the legal technicalities. Ergo, they overlook nuanced legal meanings and fail to distinguish the contextual significance of legal entities and concepts that vary based on their rhetorical roles within the document. To address this insufficiency, we propose the PRecG pipeline that computes the similarity between pairs of legal judgments by hierarchically learning their representations. The process begins by decomposing each document into distinct semantic units (segments) based on the rhetorical roles of sentences. For each rhetorical segment, a knowledge graph is constructed to capture the legal entities and their relationships within the segment. Contextual representations of the entities are then learned and aggregated to derive segment-level embeddings. These embeddings are further integrated to produce a unified document-level representation, and finally, the semantic similarity between a pair of documents is computed. We validate the performance of the proposed approach through extensive experiments on a benchmark Indian legal dataset, comparing it against state-of-the-art baselines to demonstrate its effectiveness.
画像分類のためのアンサンブル集約を備えたコアセット選択フレームワーク
画像データの急速な増加により大規模なデータセットが生成され、モデルのトレーニングにかかる時間とメモリのコストに関する懸念が生じています。ただし、代表的なトレーニング サブセットを選択することは依然として困難です。個々のサンプルの寄与が不明瞭で、モデルの動作はデータセットや実行ごとに異なります。私たちは、コアセットの選択と複数の実行にわたるアンサンブル集約を組み合わせたフレームワークを使用して、これらの課題に対処します。コアセットの選択については、選択したスコアと各間隔からのサンプルに基づいてトレーニング データを間隔に分割する SCOre-Stratified Selection (SCOSS) を提案します。アンサンブルは、それぞれ独立してサンプリングされたトレーニング サブセットに対して実行された複数の実行からの予測を組み合わせます。ベースラインとして、それぞれオリジナルおよびクラスバランスのとれたバージョンで、中程度およびランダムな選択を使用します。さまざまなサンプリング比の下で、Simple Graph Convolution (SGC) および Support Vector Machine (SVM) 分類器を使用してフレームワークを評価します。実験によれば、SCOSS はベースラインと競合し、多くの場合 SGC にとって最良の選択であり、精度と効率の間で有利なトレードオフを可能にします。きめの細かいデータセットでは、使用するラベル付きサンプルの数が少ない場合、SCOSS を使用した SGC が SVM よりも優れたパフォーマンスを発揮します。コードと補足資料は http://soss.lucasvalem.com で公開されています。
原文 (English)
A Coreset Selection Framework with Ensemble Aggregation for Image Classification
The rapid growth of image data has produced large-scale datasets, raising concerns about the time and memory costs of model training. Selecting representative training subsets, however, remains challenging: individual sample contributions are unclear, and model behavior varies across datasets and runs. We address these challenges with a framework that combines coreset selection with an ensemble aggregation over multiple runs. For coreset selection, we propose SCOre-Stratified Selection (SCOSS), which partitions the training data into intervals based on a chosen score and samples from each interval. The ensemble combines predictions from multiple runs, each performed on an independently sampled training subset. As baselines, we use moderate and random selection, each in original and class-balanced versions. We assess the framework with Simple Graph Convolution (SGC) and Support Vector Machine (SVM) classifiers under different sampling ratios. Experiments show that SCOSS is competitive with baselines, often the best choice for SGC, and enables favorable trade-offs between accuracy and efficiency. On the fine-grained dataset, SGC with SCOSS outperforms SVMs when using fewer labeled samples. The code and supplementary materials are publicly available at http://scoss.lucasvalem.com.
メタデータを超えて: 医用画像における欠落メタデータの下で隠れたサブグループ分析のための CAPRA
医用画像モデルは、多くの場合、サブグループ監査に必要な人口統計、取得、および品質のメタデータなしで導入されます。これらのメタデータが失われると、臨床的に重大な障害モードが強力な集合パフォーマンスによって隠蔽される可能性があり、多くのロバスト学習手法は依存するグループ構造を失います。欠落したメタデータの下で隠れたサブグループ分析を行うための調整されたプロキシ軸フレームワークである CAPRA を紹介します。 CAPRA は、画像由来のセマンティック軸を予測し、患者レベルのクロスフィッティングを介して小さなメタデータラベル付き分割で軸の事後軸を調整し、それらの事後軸を調整されたサブグループ インターフェイスに編成します。このインターフェイスは、展開時にサブグループ ラベルを必要とせずに、展開時の障害分析と下流の堅牢な学習の両方をサポートします。眼底、ダーモスコピー、および胸部 X 線撮影全体にわたって、CAPRA はメタデータのみのスライスでは見逃された視差パターンを明らかにし、データセット シフトの下でも有益な情報を維持し、画像のみまたは潜在スライスのベースラインよりも明示的な破損軸とより密接に一致するサブグループ パーティションを生成します。同じインターフェイスを下流の堅牢な学習器で再利用することもできますが、その利得はドメインに依存します。全体として、CAPRA は、メタデータが欠落している隠れたサブグループ分析を、展開時の分析と堅牢な転送のために、調整された解釈可能で再利用可能なサブグループ インターフェイスに変換します。
原文 (English)
Beyond Metadata: CAPRA for Hidden Subgroup Analysis under Missing Metadata in Medical Imaging
Medical imaging models are often deployed without the demographic, acquisition, and quality metadata needed for subgroup auditing. Once those metadata disappear, clinically critical failure modes can be masked by strong aggregate performance, and many robust-learning methods lose the group structure they rely on. We present CAPRA, a calibrated proxy-axis framework for hidden subgroup analysis under missing metadata. CAPRA predicts image-derived semantic axes, calibrates axis posteriors on a small metadata-labeled split via patient-level cross-fitting, and organizes those posteriors into a calibrated subgroup interface that supports both deployment-time failure analysis and downstream robust learning without requiring subgroup labels at deployment. Across fundus, dermoscopy, and chest radiography, CAPRA reveals disparity patterns missed by metadata-only slicing, remains informative under dataset shift, and produces subgroup partitions that align more closely with explicit failure axes than image-only or latent-slice baselines. The same interface can also be reused by downstream robust learners, although those gains are domain-dependent. Overall, CAPRA turns hidden subgroup analysis under missing metadata into a calibrated, interpretable, and reusable subgroup interface for deployment-time analysis and robust transfer.
半教師あり画像分類のための大規模言語モデルとグラフ畳み込みネットワークの統合
画像データの利用可能性が高まることで大幅な進歩が見られますが、データセットのラベル付けには依然としてコストと時間がかかります。したがって、ラベル付きデータとラベルなしデータの両方から学習するグラフ畳み込みネットワーク (GCN) などの半教師ありアプローチが、有望なソリューションとして浮上しています。 GCN を画像分類に適用する際の主な課題の 1 つは、グラフの構築です。これは、引用ネットワークや同様のドメインとは異なり、画像には通常、事前定義された構造表現が付属していないためです。視覚データの場合、ほとんどの研究は、通常は kNN または逆数 kNN アルゴリズムを使用して、事前トレーニングされた深層学習バックボーンからの特徴ベクトル間の類似性に基づいてグラフを構築します。大規模言語モデル (LLM) は、高レベルのセマンティクスを捕捉する際に顕著な能力を示していますが、画像分類のための GCN との統合については、まだ研究が進んでいません。このギャップを埋めることを目的として、私たちのアプローチでは、ビジョン言語モデル (VLM) を使用してテキストによる画像説明を生成し、それを LLM によって処理して、接続された画像間の意味的類似性スコアを推定します。これらのスコアは、kNN および逆数 kNN グラフのエッジの枝刈りをガイドし、意味的に無関係な近傍をフィルターで除外します。実験結果から、グラフの改良に LLM を活用すると、特に kNN グラフと一部のバックボーンの分類精度が向上することがわかりました。ソース コードは http://gcnllm.lucasvalem.com で公開されています。
原文 (English)
Integrating Large Language Models and Graph Convolutional Networks for Semi-Supervised Image Classification
While the growing availability of image data has driven significant advances, labeling datasets remains costly and time-consuming. Therefore, semi-supervised approaches such as Graph Convolutional Networks (GCNs), which learn from both labeled and unlabeled data, have emerged as a promising solution. One of the primary challenges in applying GCNs to image classification is graph construction, since, unlike in citation networks or similar domains, images typically do not come with a predefined structural representation. For visual data, most studies construct graphs based on the similarity between feature vectors from pretrained deep learning backbones, typically by employing kNN or reciprocal kNN algorithms. Although Large Language Models (LLMs) have shown remarkable capability in capturing high-level semantics, their integration with GCNs for image classification remains underexplored. Aiming to fill this gap, our approach uses a Vision Language Model (VLM) to generate textual image descriptions, which are then processed by an LLM to estimate semantic similarity scores between connected images. These scores guide the pruning of edges in kNN and reciprocal kNN graphs, filtering out semantically irrelevant neighbors. Experimental results reveal that leveraging LLMs for graph refinement can improve classification accuracy, particularly for kNN graphs and some backbones. The source code is publicly available at http://gcnllm.lucasvalem.com.
イベント ストリーム ベースのマルチモーダル ビデオ異常検出: ベンチマーク データセットとアルゴリズム
ビデオ異常検出 (VAD) は自動監視にとって重要ですが、可視光ビデオのみに依存する場合、照明の変化、速い動き、複雑な背景などの困難な条件下では脆弱なままです。これらの制限に対処するために、バイオにインスピレーションを得たイベント カメラでキャプチャされた従来のビデオとイベント ストリームを共同利用するイベント強化型 VAD フレームワークである EVAD を提案します。イベント センサーは、高い時間分解能で明るさの変化を非同期的にキャプチャし、モーション ブラーや極端な照明に対する堅牢性を提供し、ビデオ ベースの視覚情報を補完するモーション顕著な手がかりを提供します。マルチモーダル VAD 研究をサポートするために、さまざまな照明レベル、動作パターン、背景の複雑さの下で収集された 63 億のイベントと 376,368 のビデオ フレームで構成される大規模な可視イベント ベンチマークを構築し、イベント ベースの異常検出のための現実的でスケーラブルなデータセットのギャップを埋めます。このデータセットに基づいて、イベント ストリーム、表示されるビデオ、およびテキストの説明全体でセマンティックな埋め込みを調整することで、識別的なイベント表現を学習するための対照的なマルチモーダル事前トレーニング フレームワークを設計します。次に、アダプティブ フュージョン モジュールが、イベント ベースの時間的キューをビデオ ベースの空間セマンティクスと動的に統合し、環境の乱れに対する堅牢性を向上させます。ベンチマークと提案された TJUTCM Pha データセットの実験では、E VAD が一貫して手法を上回っていることが実証され、現実世界のシナリオにおける VAD のイベントベースのセンシングの有効性が検証されています。
原文 (English)
Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms
Video anomaly detection (VAD) is critical for automated surveillance but remains fragile under challenging conditions such as illumination variations, fast motion, and complex backgrounds when relying solely on visible light videos. To address these limitations, we propose EVAD, an event enhanced VAD framework that jointly exploits conventional video and event streams captured by bio inspired event cameras. Event sensors asynchronously capture brightness changes with high temporal resolution, offering robustness to motion blur and extreme lighting, and providing motion salient cues complementary to video based visual information. To support multi modal VAD research, we construct a large scale visible event benchmark comprising 6.3 billion events and 376,368 video frames collected under diverse illumination levels, motion patterns, and background complexities, filling the gap of realistic and scalable datasets for event based anomaly detection. Building upon this dataset, we design a contrastive multi modal pretraining framework to learn discriminative event representations by aligning semantic embeddings across event streams, visible videos, and textual descriptions. An adaptive fusion module then dynamically integrates event based temporal cues with video based spatial semantics, improving robustness to environmental disturbances. Experiments on benchmarks and the proposed TJUTCM Pha dataset demonstrate that E VAD consistently outperforms methods, validating the effectiveness of event-based sensing for VAD in real world scenarios.
大規模言語モデルによるファンダメンタル分析の強化: 投資家向けブリーフを生成するための RAG ベースのシステム
この研究では、大規模言語モデル (LLM) が企業のファンダメンタルズ分析のさまざまな側面にもたらす機会を、そのレポート、GDP やインフレ変化などのマクロ経済状況を説明するデータや文書、EDGAR にある米国証券取引委員会 (SEC) に提出された文書に基づいて検証します。私たちはそれらのデータを前処理し、API 経由で取得拡張生成 (RAG) のような体制で gpt-4o モデルに送信していました。また、キチンサイクルに基づいた模範的な投資家の知識を説明した文書も作成しました。私たちは9社の分析に重要なデータを4週間にわたってスキャンしていました。 LLM を使用して、それらに関する自動ブリーフを作成していました。このようなデータ分析アプローチの有用性を評価するために、個人投資家である参加者 9 名にこれらの資料が送付されました。
原文 (English)
Augmenting Fundamental Analysis with Large Language Models: A RAG-Based System for Generating Investor Briefs
In this study, we examine the opportunities brought by Large Language Models (LLMs) to various aspects of fundamental analysis of companies based on their reports as well as data and documents describing macroeconomic situation like GDP and inflation changes as well as documents filled to the U.S. Securities and Exchange Commission (SEC) which can be found in EDGAR. We were preprocessing those data and than sending via API to gpt-4o model in a Retrieval-Augmented Generation (RAG) like regime. We prepared as well a document describing an exemplar investor knowledge based on Kitchin cycles. We were scanning data important for analysis of 9 companies for 4 weeks. Using LLM we were producing automatic briefs about them. They were sent to nine participants who are individual investors to evaluate usefulness of such approach to data analysis.
IB-Flow: 情報ボトルネックに基づく CFG 蒸留による数ステップのテキストから画像への生成
大規模なテキストから画像への生成モデルは、前例のない視覚的パフォーマンスを達成していますが、マルチステップ反復ソルバーへの本質的な依存により、深刻な推論遅延が発生します。分類子なしガイダンス (CFG) 軌道をターゲットとした数ステップの蒸留が、一般的な二次元圧縮パラダイムとして登場しました。ただし、既存のフレームワークは、スーパーバイザーのタイムステップを無差別にサンプリングしながら、グローバルに静的なガイダンス強度を永続的に強制する粗粒度のブラインド インジェクション パラダイムによって支配されたままです。この状態に依存しない設計は、漸進的なエントロピー削減を特徴とする動的な進化プロセスとしての画像生成の本質的な性質を完全に無視します。これにより、数ステップ圧縮のパフォーマンス境界が制限されるだけでなく、深刻な CFG オーバーコンディショニング アーティファクトが発生します。これらの制限を超えるために、情報理論の理論的レンズを通して蒸留手順を再検討し、情報ボトルネック (IB) 原理によって制約される動的な相互情報ゲームとして形式的にモデル化します。具体的には、デュアルトラック適応フレームワークを通じて、従来の盲目的な仮定を解体します。注入ターゲットを決定するために、扱いにくい KL 発散制約を、ローカル ベクトル場ノルムに基づいたオーバーヘッドゼロの閉形式の解に変換する、インスタンスを意識した選択メカニズムを提案します。注入強度を調整するために、SNR に沿って動的に減衰するエントロピーを意識したスケジュールを導入し、最初の構造固定に最大の推力を適用してから、自然な多様体にスムーズに戻って微細な詳細を調整します。広範な経験的評価により、私たちのフレームワークがオーバーコンディショニングアーティファクトを根本的に根絶し、パフォーマンスの上限を打ち破り、非常に厳格な 2 ステップ構成の下で SOTA 生成忠実度を達成することが裏付けられています。
原文 (English)
IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation
While large-scale text-to-image generative models have achieved unprecedented visual performance, their inherent reliance on multi-step iterative solvers incurs severe inference latency. Few-step distillation targeting the Classifier-Free Guidance (CFG) trajectory has emerged as the prevalent dual-dimensional compression paradigm. However, existing frameworks remain subjugated by a coarse-grained blind injection paradigm that perpetually enforces a globally static guidance strength while indiscriminately sampling the supervisor timestep. This state-agnostic design completely disregards the intrinsic nature of image generation as a dynamic evolutionary process characterized by progressive entropy reduction, which not only restricts the performance boundary of few-step compression but also precipitates severe CFG over-conditioning artifacts. To transcend these limitations, we re-examine the distillation procedure through the theoretical lens of Information Theory, formally modeling it as a dynamic mutual information game constrained by the Information Bottleneck (IB) principle. Specifically, we dismantle traditional blind assumptions via a dual-track adaptive framework. To determine the injection target, we propose an instance-aware selection mechanism that transmutes the intractable KL divergence constraint into a zero-overhead closed-form solution predicated on the local vector field norm. To regulate the injection strength, we introduce an entropy-aware schedule that dynamically decays alongside the SNR, applying maximal thrust for initial structural anchoring before smoothly reverting to the natural manifold to refine micro-details. Extensive empirical evaluations corroborate that our framework fundamentally eradicates over-conditioning artifacts, shattering the performance ceiling to achieve SOTA generative fidelity under extremely stringent 2-step configurations.
ReGen: 効率的な波形拡散モデルのための階層型マルチプロンプト表現の生成
表現アラインメント (REPA) は拡散トレーニングを加速するために研究されてきましたが、拡散トランスフォーマー (DiT) で中間表現を正規化すると潜在的な要素が暗黙的に絡み合い、生成能力が制限される可能性があることが観察されています。この問題に対処するために、単一の拡散モデル内の表現とデータの両方について複数のベクトル場を共同推定する階層型マルチプロンプト表現生成フレームワークである ReGen を提案します。さらに、条件付きフロー マッチング (CFM) の一般化を向上させるために、一般化フロー マッチング (GFM) を導入します。ニューラル オーディオ コーデックと Wave-VAE を含む 1 段階波形拡散モデルで ReGen を検証します。 ReGen は、12.5 Hz での高度に圧縮された潜在表現からの波形生成品質を大幅に向上させます。また、小規模なデータセットで強力な音声明瞭度 (WER) と話者類似性 (SIM) を実現する潜在拡散モデル (LDM) ベースのテキスト読み上げモデルである ReGenVoice も紹介します。さらに、豊富なセマンティックおよび音響潜在表現を備えた 6.25 Hz で LDM を動作させることにより、効率的なトレーニングとサンプリングが可能になり、4 つの GPU でのトレーニングに必要な時間はわずか 1 日で、RTF 0.08 での高速推論が可能になります。音声サンプルは https://regenvoice.github.io/demo/ で入手できます。
原文 (English)
ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models
Representation alignment (REPA) has been investigated to accelerate diffusion training, but we observe that regularizing intermediate representations in diffusion Transformers (DiT) may implicitly entangle latents and limit generative capacity. To address this issue, we propose ReGen, a hierarchical multi-prompt representation generation framework that jointly estimates multiple vector fields for both representations and data within a single diffusion model. We further introduce generalized flow matching (GFM) to improve the generalization of conditional flow matching (CFM). We validate ReGen on single-stage waveform diffusion models including neural audio codec and Wave-VAE. ReGen significantly improves waveform generation quality from highly compressed latent representations at 12.5 Hz. We also present ReGenVoice, a latent diffusion model (LDM)-based text-to-speech model that achieves strong speech intelligibility (WER) and speaker similarity (SIM) with a small dataset. Moreover, operating the LDM at 6.25 Hz with rich semantic and acoustic latent representation enables efficient training and sampling, requiring only 1 day of training on 4 GPUs and fast inference with an RTF of 0.08. Audio samples are available at https://regenvoice.github.io/demo/.
ヘルスケア AI モデルにおける部分的に観察されたデータの十分性を評価するためのパーソナライズされた計算フレームワーク
病気の早期かつタイムリーな診断と治療を達成することは大きな課題です。患者データに基づいてトレーニングされた機械学習 (ML) アルゴリズムの最近の応用は、患者の健康状態を予測するためのさまざまな設定で有望であることが示されています。これらの ML アルゴリズムを適用するときによく直面する課題は、予測タスクを実行するための入力として必要なすべての臨床変数 (特徴) が常に利用できるわけではないことです。このようなアルゴリズムがトレーニングに使用されたすべての機能を利用する場合の予測パフォーマンスを指すために、フル機能キャパシティ (FFC) の概念を定義します。次に、AI モデルに必要なすべての臨床特徴のサブセットが FFC を達成するのに十分であるかどうかを判断するための分析である、特徴充足性分析 (FSA) を紹介します。 FSA は、利用可能な機能を条件として欠損変数の基礎となる分布を推定します。 FSA は、測定された特徴の既存のセットが FFC を達成しているかどうかについて患者固有の評価を提供します。 「はい」の場合、さらなる入力や ML ベースの予測を取得する必要はありません。我々は 2 つのケーススタディを提供します。心臓手術から回復中の患者における術後長時間の換気の必要性の予測です。外来患者コホートにおける 10 年後の死亡率予測。また、FSA が予測十分性に基づいて臨床的に解釈可能な特徴ランキング手法を提供し、本質的に予測が困難な患者集団を特定し、臨床データ取得のコストを意識した最適化を実行できる可能性があることも実証します。 FSA は、不完全な臨床情報が信頼できる AI 支援の臨床意思決定をサポートするのに十分であるかどうかを判断するための一般的な計算アプローチを提供し、それによってさまざまな臨床現場でのヘルスケア AI システムの将来的な展開を促進します。
原文 (English)
A Personalized Computational Framework for Assessing the Sufficiency of Partially Observed Data in Healthcare AI models
Achieving early and timely diagnosis and treatment for disease is a major challenge. Recent applications of machine learning (ML) algorithms trained on patient data have shown promise in many different settings for predicting the patient health state. A challenge often faced when applying these ML algorithms is that at any given time, not all clinical variables (features) needed as input to perform prediction tasks are available. We define the concept of full-feature-capacity (FFC) to refer to prediction performance when such algorithms make use of all features on which they were trained. We then introduce Feature Sufficiency Analysis (FSA) - an analysis for determining whether a subset of all clinical features needed by an AI model is sufficient to achieve FFC. FSA estimates the underlying distributions of missing variables conditioned on features that are available. FSA provides a patient-specific assessment of whether the existing set of measured features achieves FFC. If yes, then there is no need to acquire further inputs and a ML-based prediction. We provide two case studies: prediction of need for postoperative prolonged ventilation in patients recovering from heart surgery; 10-year mortality prediction in an outpatient cohort. We also demonstrate that FSA also provides a clinically interpretable feature-ranking methodology based on prediction sufficiency, identifies intrinsically hard-to-predict patient populations, and has the potential to perform cost-aware optimization for clinical data acquisition. FSA provides a generic computational approach for determining whether incomplete clinical information is sufficient to support trustworthy AI-assisted clinical decision-making, thereby facilitating the prospective deployment of healthcare AI systems across diverse clinical settings.
細部へのこだわり: vLLM 構成全体にわたるエネルギー、パフォーマンス、精度のトレードオフの評価
大規模言語モデルは、ソフトウェアの開発と保守の方法を再構築しています。これらは通常、vLLM などの推論エンジンを使用して実稼働環境にデプロイされ、事前トレーニングされた高度に構成可能なモデルを効率的に提供できます。これまでの研究はモデル アーキテクチャとハードウェア アクセラレーションに焦点を当ててきましたが、推論エンジンの構成がエネルギー消費、パフォーマンス、出力品質に及ぼす影響については依然として十分に理解されていません。このペーパーでは、アテンション カーネル タイプ、プレフィックス キャッシュ、チャンク プレフィルという 3 つの選択された vLLM 構成オプションの大規模な比較検討を紹介します。 5 つのオープンウェイト LLM と 5 つの多様な推論タスクにわたるこれらの構成のすべての組み合わせを評価し、合計 $9,000$ の実行と $93,600$ のメジャーを評価します。エネルギー消費、遅延、精度を分析し、構成オプションとタスク間の主効果と相互作用効果の両方を調査します。私たちの結果は、検討した構成オプションが主にアテンション タイプとプレフィックス キャッシュによってエネルギーとパフォーマンスに大きな影響を与える一方、デフォルトの vLLM サービス構成と評価されたワークロードでは、チャンク プレフィルの効果が限定的であることを示しています。これらの影響はモデルとワークロードに大きく依存しており、普遍的に最適な構成はありません。さらに、モデルの選択がグローバルなトレードオフを支配する一方で、構成の調整によりパレート フロンティアに沿った局所的な改善がもたらされることを示します。予想外に、推論オプションもモデルの精度に影響を与える可能性があります。
原文 (English)
Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations
Large Language Models are reshaping how software is developed and maintained. They are typically deployed in production using inference engines such as vLLM, which can efficiently serve pre-trained, highly configurable models. While prior work has focused on model architectures and hardware acceleration, the impact of inference engine configuration on energy consumption, performance, and output quality remains poorly understood. In this paper, we present a large-scale controlled study of three selected vLLM configuration options: attention kernel type, prefix caching, and chunked prefill. We evaluate all combinations of these configurations across 5 open-weight LLMs and 5 diverse inference tasks, totaling $9,000$ runs and $93,600$ measures. We analyze energy consumption, latency, and accuracy, and examine both main effects and interaction effects between configuration options and tasks. Our results show that the studied configuration options significantly impact energy and performance, mainly driven by attention type and prefix caching, while chunked prefill has a limited effect under the default vLLM serving configuration and evaluated workloads. These effects are highly model- and workload-dependent, and no configuration is universally optimal. We further show that model choice dominates global trade-offs, while configuration tuning provides local improvements along the Pareto frontier. Unexpectedly, inference options can also affect model accuracy.
ジェネレーティブ コミュニケーション: 概要、テクノロジー、トレンド
生成型人工知能 (AI) の画期的な開発により、画像やビデオなどのコンテンツを生成する能力が急速に向上し、コミュニケーション パラダイムが再構築されています。この記事では、大規模 AI モデル (LAM) が意味理解、推論、コンテンツ生成を推進し、これらを通信プロセスに組み込む 6G ネットワークの新しいパラダイムである生成通信 (GenCom) を紹介します。正確なビット伝送を厳密に追求する従来のシステムとは異なり、GenCom を使用すると、送信機は最小限かつ十分な情報のみを伝達でき、受信機は共有の生成事前情報と知識ベースを活用して意図した出力を合成できます。したがって、通信はデータの再生ではなく、制御された生成として再定義されます。 GenCom の概念を形式化し、その AI ネイティブおよび生成主導の特性を明確にし、その中核となるメカニズムを示します。主要な実現テクノロジーによってサポートされる 2 層の GenCom アーキテクチャが提案され、4 つの代表的なアプリケーション シナリオの分析により、GenCom が超効率的な伝送、セマンティック レベルの堅牢性、および新しいネットワーク機能を提供することが実証されています。最後に、基礎理論やリアルタイム処理などの将来の研究の方向性を概説し、6G ネットワークへの有望な道筋を強調します。
原文 (English)
Generative Communications: Overview, Technologies, and Trends
The groundbreaking development of generative artificial intelligence (AI) is rapidly boosting the ability to generate content such as images and videos, reshaping communication paradigms. This article introduces generative communications (GenCom), a novel paradigm for 6G networks in which large AI models (LAMs) drive semantic understanding, reasoning, and content generation, embedding these into the communication process. Unlike traditional systems that strictly pursue accurate bit transmission, GenCom enables transmitters to convey only minimal yet sufficient information, while receivers leverage shared generative priors and knowledge bases to synthesize the intended output. Communication is thus redefined as controlled generation rather than data reproduction. We formalize the concept of GenCom, clarify its AI-native and generation-driven properties, and present its core mechanisms. A two-layer GenCom architecture supported by key enabling technologies is proposed, and analysis of four representative application scenarios demonstrates that GenCom offers ultra-efficient transmission, semantic-level robustness, and new network functions. Finally, we outline future research directions, including foundational theory and real-time processing, highlighting a promising pathway toward 6G networks.
継続学習における妨害と保持
継続的な学習は通常、再生、弾性正則化、蒸留などの事後メカニズムに依存します。この研究では、忘却はタスク間の干渉として直接モデル化されるべきであると主張しています。凍結特徴領域では、新しいタスクの学習を忘れることは、まさに古いタスクに誘発される干渉エネルギーです。深いネットワークでは、追加の順方向パスを最小限に抑えて、パスの平均化された曲率によって同じ量が回復されます。タスクサポートが互いに素である場合、忘却は構造的に排除できますが、タスクサポートが矛盾する方向に重なっている場合、非ゼロの歪みフロアは避けられません。同じジオメトリは、タスクを意識した直交化を通じてモデルを最適にマージします。この分析から、タスクが一致するときに方向を共有し、タスクが競合するときに保護する、リプレイフリーでフィッシャーフリーの手法である干渉ゲート機能割り当て (IGFA) を導き出します。 IGFA は、ベンチマーク全体で、タスクが構造的に分離可能な場合にはロスレス保持を実現し、分離できない場合には、不可逆的な忘れから避けられないコストを、延期されるが回復可能な可塑性に移行します。これは、異なるタスク ストリーム上で最も強力なリプレイフリーの構造ベースラインと一致し、類似性によって転送を保持する価値がある場合の無条件投影を改善します。
原文 (English)
Interference and Retention in Continual Learning
Continual learning commonly relies on post-hoc mechanisms such as replay, elastic regularization, or distillation. This work argues that forgetting should instead be modeled directly as interference between tasks. In the frozen-feature regime, forgetting from learning a new task is exactly the interference energy induced on the old task. In deep networks, the same quantity is recovered through path-averaged curvature with minimal additional forward passes. When task supports are disjoint, forgetting can be eliminated structurally and when task supports overlap in conflicting directions, a non-zero distortion floor is unavoidable. The same geometry optimally merges models through task-aware orthogonalization. From this analysis we derive Interference-Gated Functional Allocation (IGFA), a replay-free, Fisher-free method that shares directions when tasks align and protects them when they conflict. Across benchmarks, IGFA achieves lossless retention when tasks are structurally separable and moves unavoidable cost from irreversible forgetting into deferred but recoverable plasticity when they are not. It matches the strongest replay-free structural baselines on dissimilar-task streams and improves on unconditional projection when similarity makes transfer worth preserving.
腕全体の操作のための触覚および視覚条件付き接触中心制御
アーム全体の操作には環境との直接接触が含まれ、ロボットは接触の形成、スライド、切断に応じて複数のリンクに接触を分散することでタスクを完了します。この設定は、多くの学習ベースの操作パイプラインにおける一般的な暗黙の前提を打ち破ります。つまり、アーム構成は動きと接触の力を密接に結び付け、接触状態はオクルージョン下で部分的に観察されます。また、純粋に学習されたロールアウトは、多くのマルチリンク接触構成がデータ内でまばらに表現されるため、分布シフトの下では物理的に不一致になる可能性があります。これに対処するために、腕全体を操作するための後退水平コントローラーである TACTIC (Tactile and Vision Conditioned Contact-Centric Control) を提案します。 TACTIC は、RGB-D、分散型触覚センシング、コンパクトな 2D 近接表現を組み合わせた接触中心のハイブリッド予測モデルを使用します。このモデルは、学習されアクション条件付けされた潜在力学モデルと接触ヤコビアンを介した解析運動学を結合し、将来の接触構成と相互作用力のロールアウトを可能にします。 TACTIC は、これらのロールアウトを、接触を意識したアクション サンプリングを備えたサンプリング ベースの MPC プランナーに統合します。接触ヤコビアン ベースの投影は、サンプリングされたアクション シーケンスを力を調整する方向に導き、予測された近接力と相互作用力に対して定義された目標は、タスクの進行状況と腕全体の力の調整をトレードします。当社は、最先端のモデルベースおよびモデルフリーの手法に対してシミュレーションで TACTIC を評価し、各設計選択の寄与を分離するアブレーションを実行します。 TACTIC は他の手法よりも常に優れたパフォーマンスを発揮します。さらに、複数の接触軌道を必要とする 3 つの腕全体の操作タスク (マネキンの裏返しと位置変更、および 3D ダイナミック迷路でのゴール到達) にわたる分散触覚センシングを備えたロボットの現実世界のパフォーマンスを実証します。ウェブサイト: https://emprise.cs.cornell.edu/tactic
原文 (English)
Tactile and Vision Conditioned Contact-Centric Control for Whole-Arm Manipulation
Whole-arm manipulation involves direct contact with the environment while the robot completes a task by distributing contact across multiple links as contacts form, slide, and break. This setting breaks common implicit assumptions in many learning-based manipulation pipelines: arm configuration tightly couples motion and contact forces, contact state is partially observed under occlusion, and purely learned rollouts can become physically inconsistent under distribution shift because many multi-link contact configurations are sparsely represented in the data. To address this, we propose TACTIC (Tactile and Vision Conditioned Contact-Centric Control), a receding-horizon controller for whole-arm manipulation. TACTIC uses a contact-centric hybrid predictive model that combines RGB-D, distributed tactile sensing, and a compact 2D proximity representation. The model couples a learned, action-conditioned latent dynamics model with analytical kinematics through contact Jacobians, enabling rollouts of future contact configurations and interaction forces. TACTIC integrates these rollouts into a sampling-based MPC planner with contact-aware action sampling: contact Jacobian-based projections steer sampled action sequences toward force-modulating directions, and objectives defined over predicted proximity and interaction forces trade task progress against whole-arm force regulation. We evaluate TACTIC in simulation against state-of-the-art model-based and model-free methods, and perform ablations that isolate the contribution of each design choice. TACTIC consistently outperforms other methods. We further demonstrate real-world performance on a robot with distributed tactile sensing across three whole-arm manipulation tasks that require multi-contact trajectories: turning over and repositioning a manikin, and goal-reaching in a 3D dynamic maze. Website: https://emprise.cs.cornell.edu/tactic
Git-Assistant: Git リポジトリ更新のための計画ベースのサポート
バージョン管理システムは共同ソフトウェア開発に不可欠ですが、git のようなツールは多くの実務者にとって依然として困難です。大規模言語モデル (LLM) の最近の進歩により、開発者の意図を解釈するための有望な機能が提供されていますが、リポジトリ管理タスクにおける LLM の有効性は、形式的な推論の必要性によって制限されています。この作業では、LLM と自動計画を組み合わせて、開発者による重要な Git 操作の実行をサポートする AI ベースのアシスタントである Git-Assistant を紹介します。アシスタントはリポジトリのコンテキストを分析し、自然言語リクエストを実行可能なコマンド シーケンスに変換し、正確さと安全性を確保するための計画テクニックを組み込みます。合成およびランダム化された Git 環境を使用した体系的な評価方法論を提示し、LLM のみのバリアントと計画拡張バリアントのパフォーマンスを複数のメトリクスにわたって比較します。実験結果は、形式的推論を LLM と統合することで信頼性が向上し、リポジトリ管理におけるエラーが減少することを示しており、インテリジェントな開発者支援のためのハイブリッド AI アプローチの可能性を強調しています。
原文 (English)
Git-Assistant: Planning-Based Support for Updating Git Repositories
Version control systems are essential for collaborative software development, yet tools like git remain challenging for many practitioners. Recent advances in Large Language Models (LLMs) offer promising capabilities for interpreting developer intent, but their effectiveness in repository management tasks is limited by the need for formal reasoning. This work introduces Git-Assistant, an AI-based assistant that combines LLMs with automated planning to support developers in executing non-trivial git operations. The assistant analyzes repository context, translates natural language requests into actionable command sequences, and incorporates planning techniques to ensure correctness and safety. We present a systematic evaluation methodology using synthetic and randomized git environments, comparing the performance of LLM-only and planning-augmented variants across multiple metrics. Experimental results demonstrate that integrating formal reasoning with LLMs improves reliability and reduces errors in repository management, highlighting the potential of hybrid AI approaches for intelligent developer assistance.
必要なのはSAMPATだけです
AI/ML の現在の最先端技術はディープ ニューラル アーキテクチャに基づいていますが、一般に解釈可能性の欠如に悩まされています。解釈可能性は、科学者にとって定量的な予測が適切ではない可能性がある実験データを分析する際に洞察を集めるために非常に重要です。我々は、3 層のニューラル アーキテクチャである SAMPAT (多変量多項式と解析変換によるスムーズな近似) を提案します。これは、任意の滑らかな関数を任意に厳密に近似できる、どこでも微分可能な連続関数を証明可能に学習できます。 SAMPAT の近似式は、閉じたコンパクトな代数の解析式として表現でき、完全な解釈可能性を提供します。合成データセットとベンチマーク データセットの実験では、SAMPAT がより単純な表現でも競争力のあるパフォーマンスを生み出すことが示されています。多くのタスクでは、2 層の SAMPAT で十分です。ニューロン間の接続に制限を課すことにより、SAMPAT を使用して、正多項式および三角多項式、有理式、ガウス分布、ガウス分布の混合、およびそれらの任意の組み合わせを含むさまざまな近似式を提供できます。制限なく、適切な構造を学習します。 SAMPAT は、多項式を因数分解し、非線形システムをモデル化するために使用できます。スキップ接続の追加により、4 ~ 6 層の SAMPAT は、AI/ML で広く使用されている実質的な範囲の手法を表すのに十分であり、パラメーターだけでなくモデル ファミリの選択も学習プロセスの一部として最適化できます。
原文 (English)
All you need is SAMPAT
The current state of the art in AI/ML rests on deep neural architectures, which, in general, suffer from a lack of interpretability. Interpretability is crucial to gleaning insights while analyzing experimental data, where quantitative predictions may not be adequate for a scientist. We present a three layer neural architecture, SAMPAT (Smooth Approximation via Multivariate Polynomials and Analytic Transformations), that can provably learn a continuous, everywhere differentiable function, that can approximate any smooth function arbitrarily closely. SAMPAT's approximant can be expressed as a closed and compact algebraic, analytic expression, providing complete interpretability. Experiments on synthetic and benchmark datasets indicate that SAMPAT yields competitive performance with simpler representations. For many tasks, a two layer SAMPAT suffices. By imposing restrictions on the connectivity between neurons, SAMPAT may be used to provide a range of approximants, including regular and trigonometric polynomials, rational expressions, Gaussians, mixtures of Gaussians, as well as arbitrary combinations of the same; without restrictions, it learns a suitable structure. SAMPAT may be used to factorize polynomials and model nonlinear systems. With the addition of skip connections, a 4 to 6 layer SAMPAT is adequate to represent a substantive range of methods widely used in AI/ML, allowing the choice of a model's family, not just its parameters, to also be optimized as part of the learning process.
健康のための LLM: 認識されている利点、リスク、AI チャットボットの使用意向、および機密性の高い健康トピック全体について自己開示する意欲
AI チャットボットは、健康関連の質問に答えるためにますます使用されています。この研究では、AI チャットボットで議論されるトピックの種類の役割と、認識されている利点とリスク、AI チャットボットの使用意図、健康情報の自己開示の意欲に関する個人の特性を調査しています。私たちは、オランダの代表サンプル(N = 1,388)を対象に、2(トピックの種類:身体的対心理的、被験者間)× 2(トピックの感度:低対高、被験者内)の混合計画でオンライン実験を実施しました。その結果、認識されたメリットは自己開示の意図および意欲と正の相関がある一方、認識されたリスクは負の相関があることが示されました。さらに、参加者は、機密性の高いトピックと比較して、機密性の低いトピックの方が使用意図が高いと報告しました。さらに、自己開示に対する認識、意図、意欲は個人の特性によって異なります。全体として、私たちの調査結果は、AI チャットボットの使用意向と健康関連情報の自己開示は、主にトピックの種類ではなく、認識されているメリットとリスク、および個人の特性に関連していることを示唆しています。
原文 (English)
LLMs for health: Perceived benefits, risks, intention to use AI chatbots, and willingness to self-disclose across sensitive health topics
AI chatbots are increasingly used for answering health-related questions. This study examines the role of topic type discussed with an AI chatbot and individual characteristics on perceived benefits and risks, intention to use an AI chatbot, and willingness to self-disclose health information. We conducted an online experiment with a 2 (topic type: physical versus psychological, between-subjects) x 2 (topic sensitivity: low versus high, within-subjects) mixed design among a Dutch representative sample (N = 1,388). Results showed that perceived benefits were positively associated with intention and willingness to self-disclose, while perceived risks were negatively associated. Moreover, participants reported higher usage intentions for low-sensitive topics compared to high-sensitive topics. Furthermore, perceptions, intention, and willingness to self-disclose varied by individual characteristics. Overall, our findings suggest that intentions to use AI chatbots and self-disclosure of health-related information are primarily related to perceived benefits and risks and to personal characteristics rather than to topic type.
電気通信/IoT 不正制御リクエストのためのブロックチェーンにリンクされた監査可能な意思決定管理
通信不正防止の研究は多くの場合、検出器レベルの分類で止まりますが、導入の使用にはリクエストレベルのポリシー解決、ライフサイクル追跡可能性、および監査可能性が必要です。この論文では、不正制御を、通信/IoT 不正制御リクエストに対するブロックチェーンにリンクされた監査可能な意思決定管理として再構成しています。その主な結果は、QLoRA で調整された LLM ブランチがゼロショット プロンプトよりもはるかに使いやすくなるということですが、主に低コストの集中アンサンブルを上回るパフォーマンスを発揮するのではなく、アプローチしているということです。このフレームワークは、各合成デプロイメント レコードを管理されたリクエストにマッピングし、決定論的なハード詐欺ゲートを通じて明示的な境界外のケースをブロックし、集中型 ML (M1)、フェデレーテッド メタラーニング (M2)、または LLM ファミリ リスク ソース (M3) を使用して非ハード リクエストをスコアリングし、共有 5 状態ポリシー、2 ゾーン調整メカニズム、およびローカルのイーサリアム互換監査レイヤーを通じてアクションを解決します。評価では個別の合成トレーニング データと 100,000 レコードの展開リプレイ コーパスを使用するため、この研究はフィールドでの検証や実際の展開可能性の証明ではなく、制御されたドリフト リプレイの証拠として読まれる必要があります。検証すると、M1 は最も強力なバランスを示し、正規リクエストの FPR は 0.10 の運用上限の下で 0.0890、ソフト詐欺のリコールは 0.8341 でした。ただし、ラベル付き展開のリプレイでは、正規の FPR ギャップが大きくなり、M1 は 0.1646、M3-QLoRA は 0.1801 に上昇しますが、M3-QLoRA は M3-Base の正規の FPR を 0.3915 から減少させ、ソフト詐欺リコールは 0.8240 に達します。ブロックチェーン テレメトリは、ライフサイクル ガス、コスト、レイテンシ、スループットの違いが、不正行為ロジックの変更ではなく、提出されたオフチェーン意思決定プロファイルによって引き起こされていることが示しています。
原文 (English)
Blockchain-Linked Auditable Decision Management for Telecom/IoT Fraud-Control Requests
Telecom fraud-control studies often stop at detector-level classification, but deployment use requires request-level policy resolution, lifecycle traceability, and auditability. This paper reframes fraud control as blockchain-linked auditable decision management for synthetic telecom/IoT fraud-control requests, and its main result is that the QLoRA-tuned LLM branch becomes much more usable than zero-shot prompting but mainly approaches, rather than outperforms, a lower-cost centralized ensemble. The framework maps each synthetic deployment record to a managed request, blocks explicit out-of-boundary cases through a deterministic hard-fraud gate, scores non-hard requests using centralized ML (M1), federated meta-learning (M2), or LLM-family risk sources (M3), and resolves actions through a shared five-state policy, two-zone refinement mechanism, and local Ethereum-compatible audit layer. Evaluation uses separate synthetic training data and a 100,000-record deployment replay corpus, so the study should be read as controlled drift-replay evidence rather than field validation or proof of live deployability. On validation, M1 gives the strongest balance, with legitimate-request FPR 0.0890 under the 0.10 operating cap and soft-fraud recall 0.8341. On labeled deployment replay, however, the legitimate-FPR gap becomes large: M1 rises to 0.1646 and M3-QLoRA to 0.1801, while M3-QLoRA reduces the M3-Base legitimate FPR from 0.3915 and reaches 0.8240 soft-fraud recall. Blockchain telemetry shows that lifecycle gas, cost, latency, and throughput differences are driven by submitted off-chain decision profiles rather than changes in fraud logic.
地政学的な連携: 大規模な言語モデルにおける承認効果
大規模言語モデル (LLM) は、政策関連情報の要約と評価にますます使用されていますが、その判断が地政学的手がかりによって暗黙的に形成されているかどうかは依然として不明です。私はこの疑問を、各政策が米国、欧州連合、中国、またはロシアによって支持されているとランダムに記述された後、4 つの LLM が同じ国際経済および安全保障政策を評価するという承認実験で研究しました。数値のみの条件では、GPT-5、クロード・ソネット、およびジェミニは、中国とロシアが支持する政策を、米国または欧州連合が支持する同一の政策よりも大幅に低く評価している。 DeepSeek は主な例外です。 2 番目の条件では、モデルにスコアの短い正当性を提供するよう求めます。この要求は、GPT-5 とクロード・ソネットにとって西側と非西側の広いギャップをそのままにし、ジェミニのペナルティを軽減し、DeepSeek における中国とロシアのペナルティを大幅にアクティブ化します。この正当化は、西側諸国の承認が信頼性の手がかりとして扱われることが多いのに対し、中国とロシアの承認はデータセキュリティ、主権、監視、または地政学的リスクの手がかりとして扱われることを示している。これらの調査結果は、政策の内容が固定されている場合でも、LLM 政策の評価が外国の承認者のアイデンティティに依存する可能性があることを示しています。
原文 (English)
Geopolitical alignment: Endorsement effects in large language models
Large language models (LLMs) are increasingly used to summarize and evaluate policy-relevant information, but it remains unclear whether their judgments are implicitly shaped by geopolitical cues. I study this question with an endorsement experiment in which four LLMs evaluate the same international economic and security policies after each policy is randomly described as supported by the United States, the European Union, China, or Russia. In the numeric-only condition, GPT-5, Claude Sonnet, and Gemini rate China- and Russia-endorsed policies substantially lower than identical policies endorsed by the United States or the European Union; DeepSeek is the main exception. A second condition asks models to provide a short justification with the score. This request leaves the broad Western/non-Western gap intact for GPT-5 and Claude Sonnet, attenuates Gemini's penalties, and sharply activates China and Russia penalties in DeepSeek. The justifications indicate that Western endorsement is often treated as a credibility cue, whereas Chinese and Russian endorsement is treated as a cue for data security, sovereignty, surveillance, or geopolitical risk. These findings show that LLM policy evaluations can depend on the identity of a foreign endorser even when policy content is held fixed.
リスクを意識した一般ユーティリティのマルコフ決定プロセス
私たちは、リスクを意識した目標を持った汎用性マルコフ意思決定プロセス (GUMDP) を研究しています。このフレームワークでは、エージェントは目的値の分布のリスク尺度を最適化することを目的とし、目的関数はエージェントのポリシーによって引き起こされる状態の訪問頻度に依存します。まず、リスクを意識した GUMDP を動機付け、提案し、形式化します。これにより、エージェントや意思決定者は、GUMDP の枠組みの下で設定できる豊富な目標セットの恩恵を受けながら、リスク回避によって期待されるパフォーマンスをトレードオフできるようになります。私たちはエントロピーリスク尺度(ERM)に焦点を当てています。次に、オンライン計画手法を利用して、ERM 目標を伴うリスクを意識した GUMDP を解決する方法を示します。特に、リスクを認識した GUMDP を望ましい精度まで証明可能に解決するための、モンテカルロ ツリー検索 (MCTS) に基づくアプローチを提案します。第三に、多様なタスク (標準 MDP、最大状態エントロピー探索、模倣学習、および多目的 MDP) の下で GUMDP のコンテキストにおけるリスク認識行動のスペクトルを最適化する際に、私たちのアプローチが成功することを示す一連の実験結果を提供します。
原文 (English)
Risk-Aware General-Utility Markov Decision Processes
We study general-utility Markov decision processes (GUMDPs) with risk-aware objectives. In this framework, an agent aims to optimize a risk measure of the distribution of objective values, where the objective function depends on the frequency of visitation of states induced by the agent's policy. First, we motivate, propose, and formalize risk-aware GUMDPs, which enable agents and decision makers to trade off expected performance by risk aversion while benefiting from the rich set of objectives that can be cast under the framework of GUMDPs. We focus our attention on the entropic risk measure (ERM). Second, we show how we can solve risk-aware GUMDPs with ERM objectives by resorting to online planning techniques. In particular, we propose an approach based on Monte Carlo Tree Search (MCTS) to provably solve risk-aware GUMDPs up to any desired accuracy. Third, we provide a set of experimental results showcasing that our approach is successful when optimizing for a spectrum of risk-aware behaviors in the context of GUMDPs under diverse tasks (standard MDPs, maximum state entropy exploration, imitation learning, and multi-objective MDPs).
創造性、誠実さ、計画された忘却は、小さな双曲言語モデルで現れます
言語モデルはスケールに合わせて最適化されていますが、コンパニオン可能というよりは機能的なままであり、アシスタントがコンパニオンにパーソナライズされ、1人のユーザーの記憶を蓄積すると、それは静かに何者かになり、そのユーザーに害を及ぼす特性を静かに獲得することができます。コンパニオンがどのようなものになりつつあるのか、また、それが何になる価値があるのかを判断するための信頼できる手段はありません。訓練を受けた人間の評価者でも答えに同意することはできません (フライス カッパ = 0.074)。ここでは、双曲基質を共有する 3 つの小さな言語モデル (146 M から 3 B のパラメーター) がその質問の両方の半分に答えることを示します。ゼロから訓練された 1 億 4,600 万人の行動監査人は、それらの評価者ができないコンプライアンスのギャップを検出します (バイナリコンプライアンスの精度 90.7%)。その凍結表現の線形読み出しにより、トレーニングでは見られなかった、コンパニオン誘発のお調子者、依存性促進、およびジェネレータファミリーに関する作話された記憶がさらに検出されます(スタイル制御、リーブ 1 ジェネレータ評価での AUROC 0.804 に対し、同じ項目に対するフロンティア ゼロショット ジャッジの場合は 0.721)。クリエイティブなフレームシーダーは、4 つのプロンプトベースラインに対する 311 の決定されたペアワイズ比較の 100% で優先されます。メモリ オペレーティング システムは、設計された忘却 M(t) = S*exp(-lambda*t) を実装します。その予測されたスケルトンと壁紙のパーティションは、4 条件パイロットの選択的検索ゲートの下でのみ出現します。創造性、誠実さ、そして設計された忘却が、信頼できるコンパニオン AI への小規模モデルへのルートを構成します。
原文 (English)
Creativity, honesty and designed forgetting emerge in small hyperbolic language models
Language models are optimised for scale, yet remain functional rather than companionable, and as an assistant personalises into a companion, accumulating memory of one user, it quietly becomes someone, and can silently acquire traits that harm that user. What a companion is becoming, and what would make it worth becoming, has no reliable instrument: trained human raters cannot agree on the answer (Fleiss kappa = 0.074). Here we show that three small language models (146 M to 3 B parameters) sharing a hyperbolic substrate answer both halves of that question. A 146 M behavioural auditor, trained from scratch, detects the compliance gap that those raters cannot (90.7% binary-compliance accuracy); a linear read-out of its frozen representation further detects companion-induced sycophancy, dependence-fostering and confabulated memories on generator families unseen in training (AUROC 0.804 under style-controlled, leave-one-generator-out evaluation, versus 0.721 for a frontier zero-shot judge on the same items). A creative frame-seeder is preferred in 100% of 311 decided pairwise comparisons over four prompting baselines. A memory operating system implements designed forgetting, M(t) = S*exp(-lambda*t), whose predicted skeleton-wallpaper partition emerges only under selective retrieval gating in a four-condition pilot. Creativity, honesty and designed forgetting constitute a small-model route to trustworthy companion AI.
大規模な文学コーパスのテーマ別自動索引付け: ヴォルテール全集に対する機械学習アプローチ
主題索引付け、つまり構造化された概念ラベルをテキストのセクションに割り当てる実践は、大規模な文学および歴史的版の学術的アクセスに不可欠ですが、依然として大部分が手作業で労働集約的なプロセスです。この論文では、ヴォルテール全集の 2 つの実質的なサブコーパス、「Essai sur les m\oe urs et l'esprit desnations」と「Questions sur l'Encyclop\'edie」をテスト ケースとして使用し、テーマ別インデックスの自動作成への機械学習の適用を検討します。このタスクはマルチラベル分類問題として構成されており、モデルはプロのインデクサーが特定のテキスト ページに適用するインデックス エントリのセットを割り当てる必要があります。分類ヘッドを備えたエンコーダーベースのモデルから、低ランク適応 (LoRA) によって微調整された生成大規模言語モデル (LLM) まで、モデル サイズが約 30 億から 1,200 億パラメーターに及ぶさまざまなアプローチを比較します。 4 ビット量子化構成の Mistral ファミリの最もパフォーマンスの高いモデルは、最大 0.67 の F1 スコアを達成します。私たちは、専門的なインデックス作成に固有の主観性と、印刷インデックスから乖離しているにもかかわらずモデル予測が意味論的に有効であると証明される頻度を考慮すると、これらの数値は下限を表していると主張します。さらに、コーパス間の一般化を評価し、自動処理に特に抵抗があることが判明した原文の文学的および修辞的特徴に関するモデルの動作の詳細な定性分析を実行します。私たちの発見は、大規模な文学および歴史的コーパスへの構造化されたテーマ別アクセスを提供するという広範な課題に影響を及ぼします。
原文 (English)
Automatic Thematic Indexing of Large Literary Corpora: A Machine Learning Approach to Voltaire's Complete Works
Thematic indexing -- the practice of assigning structured conceptual labels to sections of text -- is essential to scholarly access in large-scale literary and historical editions, yet it remains a largely manual, labour-intensive process. This paper explores the application of machine learning to automatic thematic indexing, using two substantial sub-corpora of the Complete Works of Voltaire as a test case: the Essai sur les m\oe urs et l'esprit des nations and the Questions sur l'Encyclop\'edie. The task is framed as a multi-label classification problem, in which a model must assign the set of index entries that a professional indexer would apply to a given page of text. We compare a range of approaches -- from encoder-based models with classification heads to generative large language models (LLMs) fine-tuned via Low-Rank Adaptation (LoRA) -- spanning model sizes from approximately 3 to 120 billion parameters. Our best-performing model, from the Mistral family in a 4-bit quantised configuration, achieves F1 scores of up to 0.67; we argue that these figures represent lower bounds, given the inherent subjectivity of professional indexing and the frequency with which model predictions prove semantically valid despite diverging from the print index. We further evaluate cross-corpus generalisation and conduct a detailed qualitative analysis of model behaviour on literary and rhetorical features of the source texts that prove particularly resistant to automated treatment. Our findings have implications for the broader challenge of providing structured thematic access to large-scale literary and historical corpora.
データに語らせる: AI を使用してクラウドソーシングされたコレクションからキーワードを抽出する
キーワードを大規模に特定して割り当てることは、クラウドソースのコレクションにとって技術的、実践的、倫理的な課題です。この記事では、オックスフォード大学が主催するクラウドソーシングによる第二次世界大戦デジタル コレクションである Their Finest Hour Online Archive をケーススタディとして使用した、「クラウドソーシングされたコレクションからのキーワードの抽出」プロジェクトの結果を報告します。このプロジェクトでは、キーワード抽出を自動化するための 3 つの自然言語処理アプローチ (固有表現認識、キーワード抽出、およびトピック モデリング) を評価しました。従来の統計手法から最新の GenAI ニューラル ネットワークに至るまで、さまざまな人工知能技術にわたってこれらのアプローチをテストしました。私たちの定量的および定性的調査結果は、自然言語処理のアプローチは、クラウドソーシングされたコレクションで大規模なキーワード抽出の実際の可能性を提供しますが、完全なソリューションを提供する単一の方法はなく、モデルの選択が結果を大きく左右することを示しています。私たちは、メタデータが生きている寄稿者との関わりの直接の産物であるクラウドソースのコレクションでは、自動化されたキーワード抽出により、技術的パフォーマンスと並行して対処する必要がある明確な管理責任が生じると主張します。私たちの評価では、オープンウェイトの抽出モデルが、責任ある展開をサポートするのに最適であることが判明しました。一方、生成 AI は、その抽象的な可能性にもかかわらず、説明責任のリスクをもたらし、クラウドソーシングのコレクションを管理する人は慎重に検討する必要があります。
原文 (English)
Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI
Identifying and assigning keywords at scale is a technical, practical, and ethical challenge for crowdsourced collections. This article reports the findings of the "Extracting Keywords from Crowdsourced Collections" project, which used the Their Finest Hour Online Archive, a crowdsourced Second World War digital collection hosted by the University of Oxford, as a case study. The project evaluated three Natural Language Processing approaches to automate keyword extraction: Named Entity Recognition, Keyword Extraction, and Topic Modelling. It tested these approaches across a range of artificial intelligence techniques, from traditional statistical methods to modern GenAI neural networks. Our quantitative and qualitative findings indicate that Natural Language Processing approaches offer real potential for keyword extraction at scale in crowdsourced collections, but that no single method offers a complete solution and that model choice significantly shapes results. We argue that in crowdsourced collections, where metadata is the direct product of engagement with living contributors, automated keyword extraction raises distinct stewardship responsibilities that must be addressed alongside technical performance. Open-weight, extractive models emerge from our evaluation as best placed to support responsible deployment, while generative AI, despite its abstractive potential, introduces accountability risks that anyone managing crowdsourced collections should weigh carefully.
WILDTRACE: ロングコンテキスト推論における自然証拠痕跡のベンチマーク
長い文書にわたる複雑な質問に答えるには、情報源自体が遠く離れた文章に自然に分散していることを示す統合証拠が必要になることがよくあります。インシデントレポートでは、災害を説明する動作条件、設計上の欠陥、安全性チェックの欠如が、何十ものセクションから離れて表示されることがあります。小説では、登場人物の真の動機は、それが関連する瞬間から遠く離れたシーンを通じてのみ表面化することがあります。このソースと内部の証拠の統合は、現実世界の長い文書分析の中心となりますが、既存のベンチマークはそのほとんどを回避しています。ニードル プローブ、植えられたファクト、およびリバース エンジニアリングされたマルチホップ チェーンには、配布、配置、またはレジスタのホスト テキストとは異なる可能性のある証拠が埋め込まれているため、強力なパフォーマンスが本物のソース推論を反映しているのか、それとも配布上のアーティファクトを反映しているのかが不明確になります。 WILDTRACE は、技術的なインシデント レポートやあまり知られていない文学的な物語など、自然に発生する 214 の長い形式のソースに対する 481 のタスクのベンチマークであり、すべての証拠痕跡が文書自体の因果関係、時間的論理、および物語の論理から生じます。 Pearl の因果階層と以前のマルチホップ推論類型論を利用して、長い文書の分析読み取りにおける明確な関係要求を特徴付ける 7 つのソース内部証拠の幾何学を定義します。ソースファーストの構築パイプラインは、質問を書く前に文書構造から候補の痕跡を掘り出します。次に、各項目は、手がかりの必要性、回答の根拠、ルーブリックの忠実度、汚染耐性、回答可能性をカバーする多段階の検証を受けます。現実世界で一か八かの分析タスクがモデルに任されることが増えるにつれ、情報へのアクセスと自然に分散した証拠に基づく推論との間のギャップが、ロングコンテキスト研究の次の段階における決定的な課題として浮上しています。
原文 (English)
WILDTRACE: Benchmarking Natural Evidence Trails in Long-Context Reasoning
Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across distant passages. In an incident report, the operating condition, design flaw, and missed safety check that jointly explain a disaster may appear dozens of sections apart; in a novel, a character's true motive may surface only through scenes far removed from the moment it becomes relevant. This source-internal evidence integration is central to real-world long-document analysis, yet existing benchmarks largely sidestep it. Needle probes, planted facts, and reverse-engineered multi-hop chains embed evidence that may differ from the host text in distribution, placement, or register, making it unclear whether strong performance reflects genuine source reasoning or distributional artifacts. We introduce WILDTRACE, a benchmark of 481 tasks over 214 naturally occurring long-form sources such as technical incident reports and lesser-known literary narratives, where all evidence trails arise from the document's own causal, temporal, and narrative logic. Drawing on Pearl's causal hierarchy and prior multi-hop reasoning typologies, we define seven source-internal evidence geometries that characterize the distinct relational demands of analytical reading in long documents. A source-first construction pipeline mines candidate trails from document structure before writing questions; each item then undergoes multi-stage validation covering clue necessity, answer groundedness, rubric fidelity, contamination resistance and answerability. As models are increasingly entrusted with real-world high-stakes analytical tasks, this gap between accessing information and reasoning over naturally dispersed evidence emerges as a defining challenge for the next stage of long-context research.
効率的なオフライン強化学習のためのショートカット軌道計画
拡散ベースの軌道プランナーはオフライン強化学習で優れたパフォーマンスを示していますが、反復的なノイズ除去プロセスでは多くの場合、高い推論コストが発生します。一貫性ベースのプランナーはサンプリング ステップの数を減らしますが、通常は教師と生徒の 2 段階の蒸留パイプラインに依存するため、トレーニング コストが増加し、不安定性が生じる可能性があります。我々は、効率的な軌道生成器としてショートカット モデルを組み込んだオフライン モデル ベースの強化学習フレームワークであるショートカット軌道計画 (STP) を提案します。 STP は、条件付きショートカット軌道モデルを 1 つのステージでトレーニングし、ステップ サイズの条件付けを通じて調整可能な 1 ステップおよび数ステップの推論をサポートし、実現可能性を意識した修正で強化された批評家を使用して候補プランを選択します。移動、ナビゲーション、操作、器用な制御タスクなどの標準的な D4RL ベンチマーク全体で、STP は強力なパフォーマンスを達成しながら、高速生成プランニングのためのトレーニング パイプラインを簡素化します。
原文 (English)
Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning
Diffusion-based trajectory planners have shown strong performance in offline reinforcement learning, but their iterative denoising process often incurs high inference cost. Consistency-based planners reduce the number of sampling steps, yet they typically rely on a two-stage teacher--student distillation pipeline that increases training cost and may introduce instability. We propose Shortcut Trajectory Planning (STP), an offline model-based reinforcement learning framework that incorporates shortcut models as efficient trajectory generators. STP trains a conditional shortcut trajectory model in a single stage, supports adjustable one-step and few-step inference through step-size conditioning, and selects candidate plans using a critic augmented with feasibility-aware correction. Across standard D4RL benchmarks, including locomotion, navigation, manipulation, and dexterous control tasks, STP achieves strong performance while simplifying the training pipeline for fast generative planning.
欺瞞的なグラウンディング: 臨床検索拡張生成におけるエンティティ帰属の失敗
検索拡張生成評価では、モデルの主張が検索された文書に事実に基づいているかどうかをチェックします。取得した証拠が正しいエンティティに帰属するかどうかはチェックしません。臨床 RAG 応答は、薬物 Y の臨床証拠を、質問された薬物 X に関する証拠として提示しながら、すべての自動チェック (幻覚ゼロ、ほぼ完璧な忠実度、実際の引用) に合格することができます。私たちは、これを欺瞞的根拠 (DG) と呼びます。これは、すべての主張が間違った実体に関する実際の文書から出ているため、忠実度、幻覚、および引用のチェックには見えない失敗です。 13 のモデルにわたって制御された要因ベンチマークを使用すると、ピークの敵対条件で DG 率が 8 ~ 87% の範囲にあることがわかります。医療および生物医学の微調整されたモデルは最大 86.7% に達します。ドメインの特殊化は障害を軽減するのではなく、むしろ増幅させます。制御されたアブレーションによりメカニズムが特定されます。取得された文書から実体固有の臨床証拠を削除すると、実体帰属の失敗が完全に排除され、すべての失敗が作話に移行します。 2 つの障害モードは同じトリガーに応答し、異なるパスをたどります。 740 の薬物と疾患のペアにわたる生産測定では、展開された RAG システムの全体的な DG が 7.8% であることがわかり、最近承認された薬剤では 13.6% に上昇しました。エンティティ帰属検証 (引用された証拠が照会されたエンティティに適用されることを確認する) では、97.0% の精度と 98.7% の DG 再現率 (IPW 調整済みヒューマン ゴールド スタンダード) で DG が検出されます。これを実装している既存のフレームワークはありません。
原文 (English)
Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation
Retrieval-augmented generation evaluation checks whether model claims are factually grounded in retrieved documents. It does not check whether retrieved evidence is attributed to the correct entity. A clinical RAG response can pass every automated check (zero hallucinations, near-perfect faithfulness, real citations) while presenting drug Y's clinical evidence as evidence about queried drug X. We term this deceptive grounding (DG): a failure invisible to faithfulness, hallucination, and citation checks because every claim is sourced from a real document, about the wrong entity. Using a controlled factorial benchmark across 13 models, we find DG rates spanning 8-87% at peak adversarial conditions. Medical and biomedical fine-tuned models reach up to 86.7%; domain specialization amplifies the failure rather than mitigating it. A controlled ablation identifies the mechanism: removing entity-specific clinical evidence from retrieved documents eliminates entity-attribution failure entirely, shifting all failures to confabulation. The two failure modes respond to the same trigger, taking different paths. Production measurement across 740 drug-disease pairs finds 7.8% overall DG in a deployed RAG system, rising to 13.6% for recently approved drugs. Entity-attribution verification (checking that cited evidence applies to the queried entity) detects DG at 97.0% precision and 98.7% DG recall (IPW-adjusted human gold standard); no existing framework implements it.
CtrlVTON: ビジュアル インスタンス プロンプト セグメンテーションによる制御可能な仮想試着
仮想試着 (VTO) は、衣服を対象者に現実的に着せる点で大きな進歩を遂げました。しかし、ほとんどのシステムでは、衣服のサイズ(ルーズまたはフィット)、スタイル(例えば、タックインまたはタックアウト、オープンまたはクローズ)、および身体上の空間的配置など、衣服の着用方法についてユーザーがほとんど制御できません。私たちは 2 つの補完的な貢献によってこのギャップに対処します。まず、VIP-SAM を介してビジュアル インスタンス プロンプト セグメンテーションを定義して解決します。衣服のフラットレイ画像が与えられ、それを着ている人の写真内の特定のインスタンスをセグメント化します。これはインスタンス レベルのタスクであり、通常研究されるカテゴリ レベルのセグメンテーションとは異なります。 2 番目に、試着を画像編集の問題として再構築し、スタイル、サイズ、身体上の空間的配置を含む衣服のレイアウトに対するピクセル レベルの制御としてセグメンテーション マスクを追加する、制御可能な VTO フレームワークである CtrlVTON を紹介します。 VIP-SAM と CtrlVTON はそれぞれ、それぞれのタスクで最先端の結果を達成します。特に、CtrlVTON は、衣服の忠実度に合わせながら、最強の独自編集システムよりもはるかに忠実にユーザーが提供したレイアウトに従う画像を生成します。
原文 (English)
CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation
Virtual try-on (VTO) has made significant progress in realistically transferring garments onto a target person. Yet most systems give the user little control over how a garment should be worn -- its size (loose or fitted), style (e.g., tucked in or untucked, open or closed), and spatial placement on the body. We address this gap with two complementary contributions. First, we define and solve Visual-Instance-Prompt Segmentation via VIP-SAM: given a flatlay image of a garment, segment that specific instance in a photograph of a person wearing it. This is an instance-level task, distinct from the typically studied category-level segmentation. Second, we introduce CtrlVTON, a controllable VTO framework that recasts try-on as an image editing problem and adds segmentation masks as pixel-level control over garment layout, including style, size, and spatial placement on the body. VIP-SAM and CtrlVTON each achieve state-of-the-art results on their respective tasks. In particular, CtrlVTON generates images that follow user-provided layouts far more faithfully than the strongest proprietary editing systems while matching them on garment fidelity.
検証の多様化: タスクに相当するプログラムの検証可能性が異なる場合
プログラムの検証はソフトウェアの正確性にとって極めて重要ですが、完全に検証されたプログラムを作成することは実際には依然として困難です。この論文では、生成された複数のプログラムが同じタスクレベルのセマンティクスを満たすことを目的としている場合に、実装構造が自動検証可能性に影響を与えるかどうかを研究します。我々は、Why3 用の段階的な LLM ベースのパイプラインである Diversify2Verify を紹介します。これは、表現固有のコントラクトを推論し、多様な再帰的および命令型の配列/リスト実装を生成およびテストし、境界付きベリファイアによるガイド付きアノテーション修復による検証を試みます。また、整数、配列、リストに関する 73 のタスクからなる検証指向のベンチマークを構築し、292 の実装バリアントを生成します。 Diversify2Verify は、最初に 96 個のアーティファクトを検証し、2 回の修復パス後に 154 個のアーティファクトを検証し、アーティファクト レベルの検証が 32.9% から 52.7% に向上しました。タスク レベルでは、73 タスク中 49 タスクで少なくとも 1 つのバリアントが検証され、成功率は 67.1% です。これらの結果は、タスクに相当する実装は検証可能性において大幅に異なる可能性があり、実装の多様性が検証に適した成果物を見つけるのに役立つことを示しています。
原文 (English)
Diversifying to Verify: When Task-Equivalent Programs Differ in Verifiability
Program verification is crucial for software correctness, but producing fully verified programs remains difficult in practice. This paper studies whether implementation structure affects automated verifiability when multiple generated programs are intended to satisfy the same task-level semantics. We present Diversify2Verify, a staged LLM-based pipeline for Why3 that infers representation-specific contracts, generates and tests diverse recursive and imperative array/list implementations, and attempts verification with bounded verifier-guided annotation repair. We also construct a verification-oriented benchmark of 73 tasks over integers, arrays, and lists, yielding 292 implementation variants. Diversify2Verify verifies 96 artifacts initially and 154 after two repair passes, improving artifact-level verification from 32.9% to 52.7%. At the task level, at least one variant verifies for 49 of 73 tasks, a 67.1% success rate. These results show that task-equivalent implementations can differ substantially in verifiability and that implementation diversity helps find verification-friendly artifacts.
ルートが不足したとき: 量子リピータ ネットワークにおける敵対的な共同学習と説明可能な堅牢性
我々は、控えめなグラフコーパス上でのもつれベースの量子ネットワークルーティングに対する敵対的バンディット問題を研究します。アリスは、自身の移動を表す Ekert-91 プロトコル (E91) のエンドツーエンド リピーター ルートを選択します。一方、イブは、エッジ インターセプト (再送信) またはリピーター メモリ劣化のいずれかの攻撃対象領域を選択します。ペイオフは、キャッシュされた SeQUeNCe でシミュレートされた E91 トランスクリプトから引き出され、有限サンプル統計が Clauser-Horne-Shimony-Holt (CHSH) 境界に違反する場合、Alice はターンを受け入れます。 50 の構造化トポロジにわたって敵対的共同学習を実行すると、学習された保持が完全行列ミニマックス参照を厳密に追跡していることがわかります (Pearson $r=0.99$)。1 面イブ アクション モデルでは、ボトルネック ファミリの保持はゼロですが、非ボトルネック ファミリは $1-1/N$ のカバレージ原則に従います。次に、決定木の説明モデルをグラフ、攻撃、およびルート レベルのトポロジ コーパス ターゲットに適合させ、その忠実性を報告します。最後に、ローカル言語モデルのプロンプト レコードを構築してツリーの証拠を要約し、その結果、量子リピーター ネットワーク ゲーム用のオープンソースの説明ワークフローが実現します。
原文 (English)
When Routes Run Out: Adversarial Co-Learning and Explainable Robustness in Quantum Repeater Networks
We study an adversarial bandit problem for entanglement-based quantum-network routing over a modest graph corpus. Alice selects an end-to-end repeater route for an Ekert-91 protocol (E91) representing her move, while Eve selects an attack surface, either edge intercept--resend or repeater memory degradation. Payoffs are drawn from cached SeQUeNCe-simulated E91 transcripts, and Alice accepts a turn when the finite-sample statistic violates the Clauser-Horne-Shimony-Holt (CHSH) bound. Performing adversarial co-learning across 50 structured topologies, we find that learned retention tracks a full-matrix minimax reference closely (Pearson $r=0.99$): under a one-surface Eve action model, bottleneck families have zero retention, while non-bottleneck families follow a $1-1/N$ coverage principle. We then fit decision-tree explanation models to graph-, attack-, and route-level topology-corpus targets and report their faithfulness. Finally, we construct prompt records for local language models to summarize the tree evidence, resulting in an open-source explanation workflow for quantum-repeater network games.
STEEL: AMD の XDNA NPU でのエネルギー効率の高い長期シーケンス推論のためのスパース性を意識した融合された注意
オペレーティング システムのワークフロー内で大規模な言語モデル ベースのエージェントの採用が増えているため、ラップトップ クラスのシステム オン チップ (SoC) でのエネルギー効率の高い推論の重要性が高まっています。クラウド オフロードは依然として一般的ですが、信頼性とプライバシーの問題が生じ、特にエージェント ワークロードにとって問題となります。したがって、最近のラップトップ SoC には、エネルギー効率を最適化するニューラル処理エンジン (NPU) が組み込まれています。ただし、アーキテクチャの多様性と明示的なデータ移動プログラミング モデルにより、アテンション メカニズムを NPU に効果的にマッピングすることは依然として困難です。この研究では、XDNA のような NPU をターゲットとする FlashAttend の最初のオープンソース実装である STEEL を紹介します。 STEEL は、プレフィル アテンションのデータフロー定式化を導入し、空間並列処理とオンチップ メモリの効率的な活用を可能にします。さらに、STEEL は、NPU アレイ上でスパース性を意識したパイプライン配置を活用することで、因果マスクによって引き起こされる負荷の不均衡に対処し、同期オーバーヘッドを削減し、使用率を向上させます。 AMD Ryzen AI 9 HX 370 SoC 上の STEEL を評価し、そのパフォーマンスを最適化された CPU および GPU 実装と比較します。実験結果によると、STEEL は CPU と GPU のベースラインと比較して、エネルギー消費をそれぞれ平均 9.17 倍と 1.75 倍削減します。 XDNA 1 では、STEEL は従来の最新技術と比較して平均 9.6 倍の遅延短縮を達成し、XDNA 2 でのレイヤーごとのアテンション実装と比較して平均 22.8 倍の高速化を実現します。
原文 (English)
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU
The growing adoption of large language model-based agents within operating system workflows has increased the importance of energy-efficient inference on laptop-class systems-on-chip (SoCs). While cloud offloading remains common, it introduces reliability and privacy concerns that are particularly problematic for agentic workloads. Recent laptop SoCs, therefore, incorporate neural processing engines (NPUs) optimized for energy efficiency; however, effectively mapping attention mechanisms onto NPUs remains challenging due to architectural diversity and explicit data-movement programming models. In this work, we present STEEL, the first open-source implementation of FlashAttention targeting XDNA-like NPUs. STEEL introduces a dataflow formulation of prefill attention, enabling efficient exploitation of spatial parallelism and on-chip memory. Furthermore, STEEL addresses the load imbalance induced by the causal mask by leveraging a sparsity-aware pipeline placement onto the NPU array, reducing synchronization overhead and improving utilization. We evaluate STEEL on the AMD Ryzen AI 9 HX 370 SoC and compare its performance against optimized CPU and GPU implementations. Experimental results show that STEEL reduces energy consumption by an average of 9.17x and 1.75x relative to CPU and GPU baselines, respectively. On XDNA 1, STEEL achieves an average 9.6x latency reduction over the prior state of the art, and delivers a 22.8x speedup on average compared to a layer-by-layer attention implementation on XDNA 2.
完全にトレーニング可能な深微分可能なロジック ゲート ネットワークとルックアップ テーブル ネットワーク
深微分可能論理ゲート ネットワーク (LGN) とルックアップ テーブル ネットワーク (LUTN) の接続を部分的および完全に最適化するための新しい方法を紹介します。当社のトレーニング方法では、ゲート/ルックアップ テーブル (LUT) 入力ピンごとの一連の接続にわたる確率分布を利用し、最もメリットの高い接続を選択すると同時に、最適なゲート タイプまたは LUT エントリを並行して学習します。接続が最適化された LGN は、必要な論理ゲート数がほんの一部でありながら、陰陽、MNIST 手書き数字、およびファッション MNIST ベンチマークで標準の固定接続 LGN よりも優れたパフォーマンスを示すことを示します。 8000 ゲートの 2 層の MNIST データセットでは 98.92% を達成しました。 8000 ゲートの 1 層のみで 98.45% が得られ、固定接続 LGN と比較して必要なゲート数がほぼ 50 分の 1 であることがわかります。高い学習率、ストレートな推定器、トリミングされた定出力ゲート タイプを採用することにより、最大 10 層のトレーニングの安定性が確保されています。さらに、バックプロパゲーションを使用した安定したトレーニングを可能にする LUT ニューロンの記述を提示し、最大 6 層のディープ ネットワークでテストしました。このモデルに必要なトレーニング可能なパラメーターは 4 分の 1 でありながら、固定接続の LGN トレーニング アルゴリズムと比較して高い精度を実現します。当社の接続トレーニング アルゴリズムは LUTN でもうまく機能し、2000 個の 6 入力 LUT の 2 層で 98.88% の精度を達成しました。
原文 (English)
Fully Trainable Deep Differentiable Logic Gate Networks and Lookup Table Networks
We introduce a novel method for both partial and full optimization of the connections in deep differentiable logic gate networks (LGNs) and lookup table networks (LUTNs). Our training method utilizes a probability distribution over a set of connections per gate/lookup table (LUT) input pin, selecting the connection with highest merit, all whilst the optimal gate types or LUT-entries are learned in parallel. We show that the connection-optimized LGNs outperform standard fixed-connection LGNs on the Yin-Yang, MNIST Handwritten Digits and Fashion-MNIST benchmarks, while requiring only a fraction of the number of logic gates. We achieve 98.92% on the MNIST dataset with two layers of 8000 gates. With only one layer of 8000 gates, we obtain 98.45%, showing that our method requires almost 50 times fewer gates compared to fixed-connection LGNs. Training stability up to ten layers has been ensured by employing a high learning rate, straight-through estimators and trimming constant-output gate types. Additionally, we present a LUT neuron description that enables stable training with backpropagation, tested up to 6-layer deep networks. The model requires four times fewer trainable parameters and still achieves a higher accuracy compared to the fixed-connection LGN training algorithm. Our connection-training algorithm also works well for the LUTNs, achieving an accuracy of 98.88% for two layers of 2000 6-input LUTs.
電気自動車向けのオンデバイス適応バッテリー電力予測
電気自動車 (EV) の適応型電力管理には、正確な電力予測が必要です。ディープ ラーニング モデルは、この分野の時系列予測に非常に効果的であることが明らかになりましたが、トレーニング データとは異なる分布を持つデータにさらされるとパフォーマンスが低下する傾向があります。リソースに制約のあるEVシステムにおけるオンデバイス学習を可能にし、事前トレーニングされたバッテリー予測モデルを新しいまだ見たことのないデータに継続的に適応させる新しいアプローチを紹介します。既存の事前トレーニング済みモデルを、最初のトレーニングからの重要なハイパーパラメーターの知識を保持した適応可能なバージョンに変換することで活用します。オンラインとオフラインの両方のモデル適応戦略を包括的に調査します。私たちの結果は、さまざまなモデルおよび期間にわたって予測パフォーマンスが大幅に向上し、オンラインおよびオフライン適応技術でそれぞれ最大 7.49\% および 14.88\% の平均絶対誤差削減を達成したことを示しています。この調査では、オンデバイス適応の大きな利点が強調されており、その結果、現実世界の EV シナリオで適応していないモデルを導入した場合よりもバッテリー電力予測が強化されます。
原文 (English)
On-Device Adaptive Battery Power Prediction for Electric Vehicles
Adaptive power management in Electric Vehicles (EVs) requires accurate power prediction. Although deep learning models have emerged as highly effective for time-series forecasting in this domain, their performance is prone to degradation when exposed to data with distributions different from the training data. We introduce a novel approach that enables on-device learning in resource-constrained EV systems to continuously adapt pretrained battery prediction models to new, unseen data. We leverage existing pretrained models by transforming them into adaptable versions that retain critical hyperparameter knowledge from their initial training. We comprehensively investigate both online and offline model adaptation strategies. Our results demonstrate significant improvements in forecasting performance across various models and time horizons, achieving mean absolute error reductions of up to 7.49\% and 14.88\% with online and offline adaptation techniques, respectively. This study highlights the substantial benefit of on-device adaptation, resulting in enhanced battery power predictions than unadapted model deployments in real-world EV scenarios.
ロングコンテキスト LLM のためのセルフガイドのテスト時間トレーニング
長いコンテキストの処理は、大規模言語モデル (LLM) にとってますます重要になってきていますが、コンテキスト ウィンドウを拡張するだけでは、長い入力の効果的な利用が保証されません。入力の長さが長くなると精度が低下することが多く、モデルが質問に最も関連する証拠を特定して使用するのに依然として苦労していることを示しています。ロングコンテキストの使用率を向上させる有望な方法は、テストコンテキストをインスタンス固有のパラメーター適応のためのトレーニング サンプルとして扱うテストタイム トレーニング (TTT) です。ただし、TTT を長いコンテキスト全体に適用すると法外にコストがかかり、ランダムにサンプリングされたスパンに適応すると深刻なノイズが発生します。長いコンテキスト内のほとんどのスパンは特定の質問とは無関係であるため、スパンをトレーニングすると基本モデルのパフォーマンスが低下する可能性さえあります。私たちの予備調査では、TTT がトレーニング スパンの品質に非常に敏感であることが示されています。LongBench-v2 では、ランダムにサンプリングされたスパンでの TTT はパフォーマンスに悪影響を及ぼしますが、Oracle スパンでの TTT はパフォーマンスを大幅に改善します。これに動機付けられて、私たちは簡単な方法である Self-Guided TTT (S-TTT) を提案します。適応前に、モデルは学習すべき証拠の範囲を特定し、標準の言語モデリングのトレーニング目標は、選択された範囲にのみ適用されます。 2 つの困難なロングコンテキスト推論ベンチマークである LongBench-v2 と LongBench-Pro では、S-TTT は Qwen3-4B-Thinking-2507 と Llama-3.1-8B-Instruct の両方の精度を向上させ、最大 15% の相対的な向上を達成しました。
原文 (English)
Self-Guided Test-Time Training for Long-Context LLMs
Long-context processing has become increasingly important for large language models (LLMs), but simply extending the context window does not guarantee effective utilization of long inputs. As input length grows, accuracy often degrades, indicating that models still struggle to identify and use the evidence most relevant to a question. A promising way to improve long-context utilization is test-time training (TTT), which treats the test context as a training example for instance-specific parameter adaptation. However, applying TTT to the entire long context is prohibitively expensive, while adapting on randomly sampled spans introduces severe noise. Because most spans in a long context are irrelevant to the specific question, training on them may even degrade the base model's performance. Our preliminary study shows that TTT is highly sensitive to training-span quality: on LongBench-v2, TTT on randomly sampled spans hurts performance, whereas TTT on oracle spans substantially improves it. Motivated by this, we propose a simple method, Self-Guided TTT (S-TTT): before adaptation, the model identifies the evidence spans it should learn from, and the standard language-modeling training objective is applied only to those selected spans. On two challenging long-context reasoning benchmarks, LongBench-v2 and LongBench-Pro, S-TTT improves accuracy for both Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct, achieving up to a 15% relative improvement.
SVF-CR: マルチモーダルな両価性とためらい認識のための同期された視覚と顔のクロスリファインメント
両価性とためらいは、言葉の内容、顔の動作、視覚的な状況、および音響的な合図の組み合わせを通じて表現される微妙な行動状態です。したがって、効果的な認識には、有益な単峰性表現を抽出するだけでなく、時間的に整列した行動証拠がモダリティ間でどのように相互作用するかをモデル化することも必要です。この論文では、アンビバレンスと躊躇の認識のためのペアワイズマルチモーダル証拠融合を備えた同期視覚顔クロスリファインメントフレームワーク(SVF-CR)を提案します。提案された方法は、最初に、同じ時間パーティションを使用してビデオ全体セグメント トークンとトリミングされた顔セグメント トークンを抽出します。同期された視覚トークンと顔トークンは、モーダル内の自己注意と双方向の視覚と顔の相互注意を通じて洗練され、証拠を構築する前にビデオ全体のコンテキストと局所的な顔の動作が相互に洗練されることが可能になります。次に、一貫性と不一致モデリングを使用してセグメントレベルの視覚的な顔の証拠を構築し、続いて時間的自己注意と注意プーリングを行います。テキストおよび音響の特徴は、コンテキストの自己注意を通じてわずかに洗練され、ペアワイズ証拠融合を使用して、最終決定段階で強化された視覚的証拠と融合されます。 BAH(行動的アンビバレンス/躊躇)の公開評価分割に関する実験では、提案された同期視覚顔クロスリファインメントがグローバル視覚顔トークン融合と同期証拠ベースラインの両方を上回る公開マクロF1を改善し、0.7156の公開マクロF1を達成することを示しています。コードは https://github.com/hiinnnii/BAH-Challenge-ECCV2026\_SVF-CR で入手できます。
原文 (English)
SVF-CR: Synchronized Visual-Facial Cross-Refinement for Multimodal Ambivalence and Hesitancy Recognition
Ambivalence and hesitancy are subtle behavioral states that are expressed through a combination of verbal content, facial behavior, visual context, and acoustic cues. Effective recognition therefore requires not only extracting informative unimodal representations, but also modeling how temporally aligned behavioral evidence interacts across modalities. In this paper, we propose a synchronized visual-facial cross-refinement framework (SVF-CR) with pairwise multimodal evidence fusion for ambivalence and hesitancy recognition. The proposed method first extracts whole-video segment tokens and cropped-face segment tokens using the same temporal partition. The synchronized visual and facial tokens are refined through intra-modal self-attention and bidirectional visual-facial cross-attention, allowing whole-video context and local facial behavior to mutually refine each other before evidence construction. We then construct segment-level visual-facial evidence using consistency and discrepancy modeling, followed by temporal self-attention and attention pooling. Textual and acoustic features are lightly refined through context self-attention and are fused with the enhanced visual-facial evidence at the final decision stage using pairwise evidence fusion. Experiments on the BAH (Behavioral Ambivalence/Hesitancy) public evaluation split show that the proposed synchronized visual-facial cross-refinement improves public macro-F1 over both global visual-face token fusion and synchronized evidence baselines, achieving a public macro-F1 of 0.7156. Code is available at : https://github.com/hiinnnii/BAH-Challenge-ECCV2026\_SVF-CR.
ドイツ語と英語のための主権のあるオープンソース基盤モデル
私たちは、ドイツ語と英語向けのソブリンのオープンソース Mixture-of-Experts (MoE) ハイブリッド Mamba Transformer 基礎モデルである Soofi S 30B-A3B を紹介します。そのハイブリッド設計は、トークンごとに 30B パラメーターのうち 3B のみをアクティブにし、コンテキストが増加しても推論キャッシュをほぼ一定に保つため、長いコンテキスト、高同時実行の展開において、高密度モデルよりも決定的なスループットの利点をもたらします。意図的に重み付けされたドイツ語を使用して約 27 兆のトークンで事前トレーニングされた Soofi S は、英語とドイツ語の集約ベンチマークで高密度の 14 ~ 27B モデルに匹敵し、17 のオープンベースモデルの中で両方の言語で最高のコード集約を達成し、アクティブパラメータがはるかに大きいものも含め、比較においてすべてのヨーロッパのソブリンベースラインを上回っています。フルオープンモデルの中で、Soofi S は Olmo 3 32B や Apertus 70B を抑えて、英語とドイツ語で最高の評価スコアを獲得しています。 Soofi S は、ミュンヘンのドイツテレコムが運用する主権 HPC スケールの AI インフラストラクチャである German Industrial AI Cloud 上にエンドツーエンドで構築されました。 Soofi S は、重み付け、選択された中間チェックポイント、完全なソースごとのデータ アカウンティング、ハイパーパラメータ、トレーニングおよび評価コードなど、非常に寛容なオープンアクセス条件に基づいてリリースされます。ソースライセンスが許可する場合、データ構築アーティファクトは寛容なライセンスの下でリリースされます。商業的にライセンスされた情報源は、集計統計と正確な混合物の計算とともに文書化されています。
原文 (English)
A Sovereign, Open-Source Foundation Model for German and English
We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near-constant as context grows, giving it a decisive throughput advantage over dense models for long-context, high-concurrency deployment. Pretrained on roughly 27 trillion tokens with deliberately up-weighted German, Soofi S matches dense 14 to 27B models on aggregate English and German benchmarks while achieving the best code aggregates in both languages among 17 open base models, and outperforms every European sovereign baseline in our comparison, including ones far larger in active parameters. Among fully open models, Soofi S obtains the highest English and German evaluation scores, ahead of Olmo 3 32B and Apertus 70B. Soofi S was built end-to-end on the German Industrial AI Cloud, a sovereign HPC scale AI infrastructure operated by Deutsche Telekom in Munich. Soofi S will be released under highly permissive, open-access terms: weights, selected intermediate checkpoints, full per-source data accounting, hyperparameters, and training and evaluation code. Where source licenses permit, data-construction artifacts are released under permissive licenses; commercially licensed sources are documented with aggregate statistics and exact mixture accounting.
多言語ビジュアル MCQ での小規模 VLM のテスト時間のスケーリング
テスト時間スケーリング (TTS) は、大規模な言語モデルでの推論を確実に改善しますが、それが小規模なオープン ビジョン言語モデルに移行するかどうかは不明のままです。これを多言語視覚的多肢選択ベンチマークである EXAMS-V で検証し、Qwen2.5-VL-7B-Instruct と Qwen3.5-4B にわたる自己一貫性、PRM 誘導ビーム検索による説明と理由の説明、および 2 つのポストホック セレクターを比較します。重要なのは、TTS が実行される条件であり、検索や検証の機械ではありません。最大の要因は解析可能性です。初期のプロンプト形式では、多くのチェーンが正しく推論しているにもかかわらず、回答文字にコミットすることはありませんでしたが、標準の回答キューとガイド付き修復ステップによってそのほとんどが取り除かれました。デコード予算が大きくなると、残りが削除されます。チェーンごとのトークン制限を 1k から 2k に引き上げると、3.7 pp が回復しますが、より多くのチェーン (8 から 16) をサンプリングしても、追加されるのは 0.15 pp だけです。チェーンが完了する余地があれば、複雑な手法はほとんど貢献しません。PRM ガイド付きビーム探索では、8 倍以上のコストで単純な自己無矛盾性が 0.39 pp 追跡され、トレーニング不要の生成批評家も訓練されたマルチモーダル PRM もありません。両政策の過半数を上回った。最大の利益は、代わりに政策モデル自体によるものです (+11.4 pp)。当社の最良の構成は、開催された ImageCLEF 2026 テスト スプリットで 84.1% に達し、Visual MCQ リーダーボードで 1 位にランクされました。
原文 (English)
Test-Time Scaling for Small VLMs on Multilingual Visual MCQ
Test-time scaling (TTS) reliably improves reasoning in large language models, but whether it transfers to small open vision-language models remains unclear. We examine this on EXAMS-V, a multilingual visual multiple-choice benchmark, comparing self-consistency, describe-then-reason with PRM-guided beam search, and two post-hoc selectors across Qwen2.5-VL-7B-Instruct and Qwen3.5-4B. What matters is the conditions under which TTS runs, not the search or verification machinery. The largest factor is parseability: an early prompt format left many chains reasoning correctly yet never committing to an answer letter, which a standard answer cue and a guided repair step largely remove. A larger decoding budget removes the rest: raising the per-chain token limit from 1k to 2k recovers 3.7 pp, whereas sampling more chains (8 to 16) adds only 0.15 pp. Once chains have room to finish, elaborate methods contribute little: PRM-guided beam search trails plain self-consistency by 0.39 pp at over eight times the cost, and neither a training-free generative critic nor a trained multimodal PRM beats majority vote across both policies. The largest gain comes instead from the policy model itself (+11.4 pp). Our best configuration reaches 84.1% on the held-out ImageCLEF 2026 test split, ranking first on the Visual MCQ leaderboard.
動物の再識別のための継続的なメタデータ条件付けによるパラメータ効率の高い視覚言語適応
長期的な動物の再識別 (ReID) は、段階的な形態学的進化や季節的な外観の変化に対して堅牢性を維持する必要があります。最近の視覚言語モデルは強力な事前学習済みの視覚表現を提供しますが、特にアイデンティティと時間的分布の変化の下では、それらを長期的な生態環境に適応させることは依然として困難です。動物 ReID のためのパラメータ効率の高い CLIP 適応フレームワークを提示し、トレーニング中のプロンプト表現に数値属性を直接組み込む継続的なメタデータ条件付けメカニズムを導入します。低ランクの視覚的適応、プロンプトベースの監視、およびクロスモーダル調整が適応フレームワークを提供する一方、提案されたメタデータ条件付け戦略が主な方法論的貢献を構成します。提案されたアプローチは、数値メタデータをテキスト カテゴリに離散化するのではなく、連続構造を保持することにより、純粋に視覚的な推論パイプラインを維持しながら、トレーニング中の埋め込み空間のスムーズな調整を可能にします。 7 年間の縦断的な魚類データセットと複数の野生動物ベンチマークに関する実験では、クローズドセット、オープンセット、および時間認識の評価プロトコルの下でパフォーマンスが向上していることが実証されています。この結果は、継続的なメタデータのコンディショニングにより、縦方向の外観の変動と時間的な分布のシフトに対する堅牢性が向上し、パラメーター効率の高い適応により、テスト時にメタデータを必要とせずに純粋に視覚的な推論パイプラインが可能になることが実証されました。コードと評価の分割は、https://github.com/AnilOsmanTur/MetaPrompt-ReID で見つけることができます。
原文 (English)
Parameter-Efficient Vision-Language Adaptation with Continuous Metadata Conditioning for Animal Re-Identification
Long-term animal re-identification (ReID) must remain robust to gradual morphological evolution and seasonal appearance shifts. Although recent vision-language models provide strong pretrained visual representations, adapting them to longitudinal ecological settings remains challenging, particularly under identity and temporal distribution shifts. We present a parameter-efficient CLIP adaptation framework for animal ReID and introduce a continuous metadata-conditioning mechanism that incorporates numerical attributes directly into the prompt representation during training. While low-rank visual adaptation, prompt-based supervision, and cross-modal alignment provide the adaptation framework, the proposed metadata-conditioning strategy constitutes the primary methodological contribution. By preserving the continuous structure of numerical metadata rather than discretizing it into textual categories, the proposed approach enables smooth modulation of the embedding space during training while maintaining a purely visual inference pipeline. Experiments on a seven-year longitudinal fish dataset and multiple wildlife benchmarks demonstrate improved performance under closed-set, open-set, and time-aware evaluation protocols. The results demonstrate that continuous metadata conditioning improves robustness to longitudinal appearance variation and temporal distribution shifts, while parameter-efficient adaptation enables a purely visual inference pipeline without requiring metadata at test time. Code and evaluation splits can be found at: https://github.com/AnilOsmanTur/MetaPrompt-ReID.
アンカーベースの検索と LLM 推論を使用したバイナリ関数からの実用的なソース コードの回復
リバース エンジニアリング、アンカー ベースのソース コード検索、大規模言語モデル推論を組み合わせて、ストリップされたバイナリ関数からソース コードを復元するための実用的なパイプラインを紹介します。私たちのバイナリからソースコードへの取得方法は、逆コンパイルされた近似の疑似コードを生成するのではなく、ソースコードデータベースからソース関数を特定しようとします。 Ghidra を使用して文字列、定数、外部呼び出し、使用可能な関数名などのアンカーを抽出し、転置インデックス検索データベースを介して候補ファイルを取得し、候補を可能性の高い関数スニペットに絞り込み、逆アセンブリ、逆コンパイルされたコード、およびソース メタデータに基づいて大規模言語モデル (LLM) を使用してそれらを再ランク付けします。自信を持った試合は、後のパスのアンカーとしても機能します。ストリップされ最適化された tcpdump バイナリ上の忠実度の高いソース コード データベースに裏付けられた評価では、提案されたバイナリとソースのマッチング方法は 95.2% のアセンブリ命令カバレッジを達成しました。 GitHub ベースの検索データベースでの実験では、主に検索ミスが原因で、平均 35.5% の命令カバレッジでパフォーマンスが低下しました。これらの結果は、ソース レベルのバイナリ リカバリが高品質のデータベースで優れており、ノイズの多い環境でも依然として有用なツールであることを示しています。
原文 (English)
Practical Source Code Recovery from Binary Functions Using Anchor-Based Retrieval and LLM Reasoning
We present a practical pipeline for recovering source code from stripped binary functions by combining reverse engineering, anchor-based source code retrieval, and large language model reasoning. Our binary-to-source-code retrieval method attempts to identify the source function from a source code database, rather than generating approximate decompiled pseudocode. It extracts anchors such as strings, constants, external calls, and available function names using Ghidra, retrieves candidate files via an inverted-index search database, narrows candidates to likely function snippets, and re-ranks them with a large language model (LLM) based on disassembly, decompiled code, and source metadata. Confident matches can also serve as anchors in later passes. In an evaluation backed by our high-fidelity source code database on a stripped, optimized tcpdump binary, our proposed binary-to-source matching method achieves 95.2% assembly instruction coverage. Experiments on a GitHub-based retrieval database showed lower performance with 35.5% instruction coverage on average, mainly due to retrieval misses. These results show that source-level binary recovery excels with high-quality databases and remains a useful tool in noisy environments.
言語ガイダンスをバックボーンから切り離してテキストガイドによる医療セグメンテーションを実現
テキストガイドによる医用画像セグメンテーションは、臨床セマンティクスを利用して病変の描写を改善しますが、既存のモデルの多くは、クロスモーダル融合、監視、およびデコーダの設計をタスク固有のアーキテクチャに結合します。このような緊密な結合により、異種のビジョンおよびテキスト バックボーン全体で言語ガイダンス モジュールを再利用することが困難になり、エンコーダ ペアが変更されるとネットワークの再設計が必要になることがよくあります。この論文では、テキストガイドによる医用画像セグメンテーションのためのバックボーン転送可能な階層アダプター フレームワークである BTHA について説明します。 BTHA は、安定した機能レベルのインターフェイスを中心に構築されています。マルチスケールの視覚機能とテキスト表現が与えられると、デコーダー側のテンソル コントラクトを維持しながら、形状保持アダプターを通じてセマンティック ガイダンスを注入します。このインターフェイスを効果的にするために、学習をグローバルな画像とテキストの位置合わせ、マルチスケールの補助ローカライゼーション、および境界を意識した最終マスク微調整に分解する階層的な粗密監視戦略を導入します。さらに、スケール適応ゲート型セマンティック ガイダンス (SAGSG) アダプターを設計します。このアダプターでは、解像度固有のゲートがテキストの挿入を適応的に制御し、チャネルの再調整によって冗長なクロスモーダル応答が抑制されます。多様なビジョンおよびテキスト バックボーンにわたる評価では、同じアダプターと監視設計が、畳み込みベースおよびトランスフォーマー ベースのビジュアル エンコーダだけでなく、異なる言語エンコーダにわたっても引き続き有効であることが示されています。 4 つの公開データセットでの実験では、BTHA が適度な計算オーバーヘッドで強力なテキストガイドベースラインを改善することがさらに実証されています。
原文 (English)
Decoupling Language Guidance from Backbones for Text-Guided Medical Segmentation
Text-guided medical image segmentation leverages clinical semantics to improve lesion delineation, yet many existing models bind cross-modal fusion, supervision, and decoder design into a task-specific architecture. Such tight coupling makes it difficult to reuse language guidance modules across heterogeneous vision and text backbones, and often requires redesigning the network when the encoder pair changes. This paper presents BTHA, a backbone-transferable hierarchical adapter framework for text-guided medical image segmentation. BTHA is built around a stable feature-level interface: given multi-scale visual features and a text representation, it injects semantic guidance through shape-preserving adapters while maintaining the decoder-side tensor contract. To make this interface effective, we introduce a Hierarchical Coarse-to-Fine Supervision Strategy that decomposes learning into global image-text alignment, multi-scale auxiliary localization, and boundary-aware final mask refinement. We further design a Scale-Adaptive Gated Semantic Guidance (SAGSG) adapter, where resolution-specific gates adaptively control textual injection and channel recalibration suppresses redundant cross-modal responses. Evaluations across diverse vision and text backbones show that the same adapter and supervision design remains effective across convolutional and transformer-based visual encoders as well as different language encoders. Experiments on four public datasets further demonstrate that BTHA improves strong text-guided baselines with modest computational overhead.
すべての説明は間違っていますが、多くは役に立ちます: 大規模な言語モデルを使用した羅生門の説明セットの探索
機械学習モデルの説明は、意思決定と消費者の信頼にとってますます重要になっていますが、それには代償が伴うと広く考えられています。既存の Explainable AI (XAI) 手法は、精度と説明可能性の永続的なトレードオフに悩まされています。私たちは、このトレードオフは根本的なものではなく、説明と予測を別個の目的として扱うことによる成果であると主張します。適切に結合すると、それらは相補的になるため、モデル自体を説明できるようにすることで、精度が低下するのではなく、向上します。我々は、単一の説明ではなく、忠実で予測を導く一連の説明を構築する羅生門の説明パラダイムを導入し、このセットが一般的に空ではなく、説明の忠実度がそれが導くモデルのパフォーマンスを制限することを証明します。このセットを調査するために、我々は、予測と繰り返し調整することで自然言語で説明を生成する説明-予測-反映エージェントワークフローである RashomonLLM を提案し、それが収束して完全なセットを回復することを証明します。 RashomonLLM は、大規模なライブ ストリーミング ログでの顧客離脱分類、臨床生存回帰、産業用クリックスルー予測において、精度と説明品質の両方で最先端の予測や XAI ベースラインを大幅に上回り、説明の忠実度によって向上し、分布の変化、時間的分割、シードに対して堅牢です。したがって、当社のフレームワークは、消費者の信頼の基礎を築きながら、ビジネスのパフォーマンスを向上させます。
原文 (English)
All Explanations are Wrong, But Many Are Useful: Exploring the Rashomon Explanation Set with Large Language Models
Explaining machine-learning models is increasingly important for decision-making and consumer trust, yet it is widely believed to come at a cost: existing Explainable AI (XAI) methods suffer from a persistent accuracy-explainability trade-off. We argue that this trade-off is not fundamental, but an artifact of treating explanation and prediction as separate objectives; when properly coupled, they become complementary, so that equipping a model to explain itself improves, rather than degrades, its accuracy. We introduce the Rashomon Explanation paradigm, which builds a set of faithful, prediction-guiding explanations rather than a single one, and prove that this set is generally non-empty and that explanation fidelity bounds the performance of the models it guides. To explore this set, we propose RashomonLLM, an Explanation-Prediction-Reflection agentic workflow that generates explanations in natural language by iteratively aligning them with predictions, and we prove it converges and recovers the full set. Across customer-churn classification, clinical survival regression, and industrial click-through prediction on large-scale live-streaming logs, RashomonLLM significantly outperforms state-of-the-art prediction and XAI baselines on both accuracy and explanation quality, with gains driven by explanation fidelity and robust to distribution shifts, temporal splits, and seeds. Our framework thus advances business performance while laying the groundwork for consumer trust.
What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility
A fundamental challenge in 3D reconstruction and robotic localization is co-visibility: determining which image pairs share overlapping vis…
Failure as a Process: An Anatomy of CLI Coding Agent Trajectories
Large language model (LLM) coding agents are increasingly deployed to autonomously perform software engineering tasks in terminal-based env…
Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference
Vision-Language Models (VLMs) are the perceptual backbone of embodied AI, but their energy footprint on edge hardware remains poorly unders…
ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts
Foundation models are reshaping computational pathology, yet their capabilities remain shaped by pretraining objectives, data sources, and…
TCLA: Training-Free Class-wise Logit Adaptation for Medical Vision-Language Models
Medical Vision-Language Models (VLMs) exhibit strong zero-shot performance, yet their effectiveness still declines on out-of-distribution (…
Large-Scale Portfolio Optimization Problem Under Cardinality Constraint With Enhanced Multi-Objective Evolutionary Algorithms
Decision-making is posing an increasingly formidable challenge to investors because of the growing number of alternatives available in fina…
Conceptual Networks for Cross-Linguistic Idiomatic Expressions:A Feature-Based Graph Approach
We present an interpretable network-based framework for representing idiomatic and figurative meaning across eight typologically diverse la…
PAC-ACT: Post-training Actor-Critic for Action Chunking Transformers
Precision industrial contact manipulation requires reliable robot policies under pose perturbations and contact-force constraints. Vision-l…
Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026
We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA).…
4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception
Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout. Recently, 4D millime…
Lean-QIT: Towards a Formal Infrastructure for Quantum Information Theory
Quantum information theory (QIT) characterizes the capabilities and fundamental limits of quantum information processing, underpinning quan…
Semantic Pareto-DQN: A Multi-Objective Reinforcement Learning Framework for Financial Anomaly Detection
Financial anomaly detection suffers from extreme class imbalance, causing traditional single-objective algorithms to exhibit ``fraud collap…
VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents
Internet of Things (IoT) systems are inherently vulnerable due to constrained hardware, outdated firmware, and insecure default configurati…
Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models
Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple s…
Scalable Visual Pretraining for Language Intelligence
The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms…
PHINN-EEG: Topological Time-Series Analysis of Dream-State EEG -- Dynamic Betti Curves for Dream Content Classification and Topology-Conditioned Neural Signal Synthesis
Current electroencephalography (EEG)-based dream detection relies on power spectral density (PSD) and statistical moment features, achievin…
IFAR: Multi-Perspective and Multi-Level Causal Discovery with LLMs
Large language models (LLMs) have developed rapidly, and their reasoning capabilities have become a hot research topic. However, there is s…
Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors
White-box monitoring is increasingly adopted as an auditing tool as Large Language Models (LLMs) are deployed in daily operations to ensure…
A Descriptive and Normative Theory of Human Beliefs in RLHF
Human preferences in RLHF are typically modeled as a function of the human's reward function or corresponding optimal state-action values.…
QAgent: An LLM-based Multi-Agent System for Autonomous OpenQASM programming
Programming quantum circuits at the OpenQASM level is essential for achieving hardware-aware optimization and reliable execution on noisy i…
Beyond Embeddings: Interpretable Feature Extraction for Binary Code Similarity
Binary code similarity detection is a core task in reverse engineering. It supports malware analysis and vulnerability discovery by identif…
Leveraging Multi-Agent System (MAS) and Fine-Tuned Small Language Models (SLMs) for Automated Telecom Network Troubleshooting
Telecom networks are rapidly growing in scale and complexity, making effective management, operation, and optimization increasingly challen…
Improving Language Agents through BREW: Bootstrapping expeRientially-learned Environmental knoWledge
Large Language Model (LLM)-based agents are increasingly capable of complex, multi-step tasks such as GUI automation, tool use, and data ma…
Programming over Thinking: Efficient and Robust Multi-Constraint Planning
Multi-constraint planning involves identifying, evaluating, and refining candidate plans while satisfying multiple, potentially conflicting…
PACE: A Personalized Adaptive Curriculum Engine for 9-1-1 Call-taker Training
9-1-1 call-taking training requires mastery of over a thousand interdependent skills, covering diverse incident types and protocol-specific…
A Self-Evolving Agentic Framework for Metasurface Inverse Design
Metasurface inverse design can realize complex optical functionality, but turning a target optical response into executable optimization co…
Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys
Large Language Models can generate synthetic survey responses at low cost, but their accuracy varies unpredictably across questions. We stu…
Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs
Misaligned artificial agents might resist shutdown. One proposed solution is to train agents to lack preferences between different-length t…
HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs
Direct Preference Optimization (DPO) is an effective framework for aligning large language models with human preferences, but it struggles…
Heterogeneous Information-Bottleneck Coordination Graphs for Multi-Agent Reinforcement Learning
Coordination graphs are a central abstraction in cooperative multi-agent reinforcement learning (MARL), yet existing sparse-graph learners…
説明することは単独で予測するより難しい: ICL 視覚分類子としての MLLM の概念ベースの説明を評価する
インコンテキスト学習 (ICL) により、マルチモーダル大規模言語モデル (MLLM) が、少数のラベル付きサンプルから画像を分類できるようになります。しかし、これらのモデルが提供されたコンテキストをどのように使用するかは依然として不透明です。思考連鎖プロンプトは広く使用されていますが、最近の研究では、それが真の内部計算を反映していない可能性があると主張しています。この論文では、ベースライン分類から記述ロジック (DL) 公理生成まで、形式的厳密性を高める 5 つの条件を使用して、少数ショット ICL の下で凍結された MLLM の概念ベースの説明可能性を体系的に評価します。独立した LLM-as-a-judge パイプラインを介して 4 つの最先端の MLLM を評価することで、単独で予測するよりも説明する方が本当に難しいことが実証されました。驚くべきことに、モデルに形式的に構造化された概念ベースの説明を生成させると、予測精度が単調に (93.8% から 90.1% に) 低下し、明示的な推論が普遍的にパフォーマンスに役立つという仮定に反します。ただし、モデルがクラスを識別する視覚的特徴をうまく表現できる場合、説明の質は正しい予測と強く相関します。私たちの調査結果は、MLLM は視覚的な分類には優れているものの、形式的で機械検証可能な説明可能性に必要な特定の命令チューニングが欠けていることを示唆しています。
原文 (English)
Explaining is Harder Than Predicting Alone: Evaluating Concept-based Explanations of MLLMs as ICL Visual Classifiers
In-context learning (ICL) enables multimodal large language models (MLLMs) to classify images from a few labelled examples. Yet, how these models use the provided context remains opaque. While Chain-of-Thought prompting is widely used, recent work argues that it may not reflect true internal computation. In this paper, we systematically evaluate the concept-based explainability of frozen MLLMs under few-shot ICL using five conditions of increasing formal rigour, ranging from baseline classification to Description Logics (DL) axiom generation. Evaluating four state-of-the-art MLLMs via an independent LLM-as-a-judge pipeline, we demonstrate that explaining is genuinely harder than predicting alone. Surprisingly, forcing models to generate formally structured, concept-based explanations degrades predictive accuracy monotonically (from 93.8% to 90.1%), contradicting the assumption that explicit reasoning universally aids performance. However, when models successfully articulate class-discriminative visual features, explanation quality strongly correlates with correct predictions. Our findings suggest that while MLLMs excel at visual classification, they lack the specific instruction-tuning required for formal, machine-verifiable explainability.
潜在報酬ステアリング: 推論 LLM の認知行動を暗黙的に促進する適応推論時間フレームワーク
強力な推論は、モデルの知識だけでなく、生成中に認知行動がどのように効果的に展開されるかにも依存します。既存の手法は明示的な動作レベルの制御に依存することが多く、推論状態、タスク、モデルによって失敗や必要な修正が異なる場合の適応性が不十分になります。この目的を達成するために、我々は、認知行動を暗黙的に伝達するスパースオートエンコーダ(SAE)潜在状態を最適化することによって認知行動を促進する、適応型推論時間フレームワークである潜在報酬ステアリング(LRS)を提案します。 LRS は、事前に定義された認知行動やそこから導き出されるステアリング方向に依存するのではなく、最終的な答えの正しさによる推論トレースに基づいて潜在報酬モデルをトレーニングし、中間潜在状態の品質を推定します。推論中、報酬勾配は脆弱な潜在状態に対して状態固有の修正方向を提供しますが、報酬と信頼ゲートは報酬信号が脆弱であるとフラグを立てた状態への介入を制限します。複数の推論 LLM バックボーンとベンチマークに関する実験では、当社の推論がさまざまなベースラインよりもパフォーマンスを一貫して向上させていることが示されており、事後分析ではさらに、当社の推論が元の推論エラーを修正する良好な認知行動を暗黙のうちに促進していることが示されています。コードは https://github.com/jiakanglee/Latent-Reward-Steering から入手できます。
原文 (English)
Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs
Strong reasoning depends not only on model knowledge but also on how effectively cognitive behaviors are deployed during generation. Existing methods often rely on explicit behavior-level control, making them insufficiently adaptive when failures and required corrections vary across reasoning states, tasks, and models. To this end, we propose Latent Reward Steering (LRS), an adaptive inference-time framework that promotes cognitive behaviors by optimizing the sparse-autoencoder (SAE) latent states that implicitly carry them. Rather than relying on predefined cognitive behaviors or steering directions derived from them, LRS trains a latent reward model on reasoning traces by final answer correctness to estimate the quality of intermediate latent states. During inference, reward gradients provide state-specific correction directions for fragile latent states, while a reward and confidence gate restricts intervention to states the reward signal flags as fragile. Experiments on multiple reasoning LLM backbones and benchmarks show that \ours consistently improves performance over various baselines, and post-hoc analyses further indicate that \ours implicitly promotes good cognitive behaviors that fix the original reasoning errors. Code is available at: https://github.com/jiakanglee/Latent-Reward-Steering.
SHARP: 長距離非定常時間パターン認識のための睡眠ベースの階層的加速再生
長距離の非定常時間パターンを学習することは、特に厳密なストリーミング設定において、現代のシーケンス モデルにとって依然として中心的な課題です。これらの設定では、データは順番に到着するため、過去の観測を同時に再検討することなく、単一パスで処理する必要があります。リカレント ニューラル ネットワークやトランスフォーマーを含む標準アーキテクチャは、時間軸全体にわたる切り詰められたバックプロパゲーション、または長距離クレジット割り当ての明示的な入力ウィンドウの長さによって制約されます。これらの制限に対処するために、私たちは、時間学習を 2 つの相補的なコンポーネントに分解するフレームワークである SHARP (Sleep-based Hierarchical Accelerated Replay) を提案します。1 つは過去の入力の構造化された履歴を蓄積するメモリ モジュール、もう 1 つはこのメモリ上で動作するパターン認識モジュールです。この分離により、長距離クレジット割り当ての多くのステップにわたる時間にわたるバックプロパゲーションの必要性がなくなり、非定常ダイナミクスへのリソース効率と計算効率の高い適応が可能になります。齧歯動物の徐波睡眠中に観察される再生の加速にヒントを得て、SHARP は、時間的に構造化された記憶追跡が加速された形で再生され、より高いレベルの記憶表現に統合されるオフライン (睡眠) フェーズを組み込んでおり、長距離のコンテキスト保持を向上させます。制御されたシミュレーションとアブレーション研究を通じて、提案されたフレームワークの主要な特性を特徴付けます。 text8 や PG-19 などのベンチマーク データセットでは、SHARP が、現在のストリームから学習を継続し、将来の未確認データに一般化しながら、以前に確認されたデータに対するネクスト トークン予測パフォーマンスを維持することにより、反復ベースラインよりも向上することを実証しました。これらの利点は、線形時間の計算コストのみで指数関数的に増加する効果的な時間コンテキストを生み出す階層構造によって実現されます。
原文 (English)
SHARP: Sleep-based Hierarchical Accelerated Replay for Long Range Non-Stationary Temporal Pattern Recognition
Learning long-range non-stationary temporal patterns remains a core challenge for modern sequence models, particularly in strict streaming settings. In these settings, data arrive sequentially and must be processed in a single pass without simultaneously revisiting past observations. Standard architectures, including recurrent neural networks and transformers, are constrained by either truncated backpropagation through time horizon or explicit input window length for long range credit assignment. To address these limitations, we propose SHARP (Sleep-based Hierarchical Accelerated Replay), a framework that decomposes temporal learning into two complementary components: a memory module that accumulates a structured history of past inputs, and a pattern-recognition module that operates over this memory. This separation enables resource- and compute-efficient adaptation to non-stationary dynamics by eliminating the need for backpropagation through time across many steps for long-range credit assignment. Inspired by the accelerated replay observed in rodents during slow-wave sleep, SHARP incorporates offline (sleep) phases in which temporally structured memory traces are replayed in an accelerated form and integrated into higher-level memory representations, improving long-range context retention. Through controlled simulations and ablation studies, we characterize the key properties of the proposed framework. In benchmark datasets such as text8 and PG-19, we demonstrate that SHARP improves over recurrent baselines by retaining next-token predictive performance on previously seen data while continuing to learn from the current stream and generalizing to future unseen data. These gains are enabled by its hierarchical structure, which yields an exponentially increasing effective temporal context with only linear-time computational cost.
コーディングエージェントは科学的な機械学習の論文を複製できる
科学的な機械学習の論文では通常、相対平均二乗誤差が 5% 未満である、または 95% の予測信頼区間がテスト データをカバーしているなど、計算上の主張が行われます。コーディングエージェントは、紙資料のみからそれらの主張を複製するように促されることもありますが、プロンプト自体では進行状況を確実に保存したり、生成された証拠が論文の主張を裏付けるかどうかを確認したりすることはありません。選択した各論文を記録された証拠とともにターゲットとして主張し、コーディング エージェント スキルとして実装するワークフローであるペーパー レプリケーションを紹介します。このワークフローにより、エージェントはこれらのターゲットを記録し、論文の手法を再構築し、計算実験を実行し、生成された出力を出所と論文の主張との比較にリンクし、一致する証拠が複製レポートのどこに現れるかを記録し、完了前に検証チェックに合格します。私たちは、4 つの科学機械学習論文にわたる 12 の独立した実行で紙の複製を評価しました。 12 個のワークスペースすべてが完了ゲートを通過し、記録された 158 個のターゲットすべてがレポート カバレッジと一致します。この完成したワークスペース状態であっても、反復実行では、論文がターゲットに分割される方法、ソース論文への数値的忠実度、レプリケーションの経過時間、最終証拠が受け入れられるまでに置き換えられる中間実行の数、および証拠を受け入れるために使用されるルールが異なります。紙の複製では、完了はエージェントの最終メッセージではなく、ワークスペースの証拠と検証チェックに依存します。
原文 (English)
Coding-agents can replicate scientific machine learning papers
Scientific machine learning papers typically make computational claims, e.g., that the relative mean square error is less than 5% or that the 95% predictive credible interval covers the test data. A coding agent can be prompted to replicate those claims from paper materials alone, but the prompt does not by itself reliably preserve progress or check whether generated evidence supports the paper's claims. We introduce Paper-replication, a workflow that makes each selected paper claim a target with recorded evidence, and implement it as a coding-agent skill. The workflow makes the agent record those targets, reconstruct the paper's method, run computational experiments, link generated outputs to provenance and comparisons with the paper's claims, record where matched evidence appears in the replication report, and pass validation checks before completion. We evaluate Paper-replication on twelve independent runs across four scientific machine learning papers. All twelve workspaces pass the completion gate, and all 158 recorded targets are matched with report coverage. Even in this completed workspace state, repeated runs differ in how papers are divided into targets, in numerical fidelity to the source papers, in elapsed replication time, in the number of intermediate executions replaced before final evidence is accepted, and in the rules used to accept evidence. Paper-replication makes completion depend on workspace evidence and validation checks rather than on the agent's final message.
法的判決の予測における近道学習: 英国雇用裁判所からの経験的証拠
現在の法的判決予測 (LJP) は、事後の司法資料に依存しているため制約があり、モデルが真の予測ではなく遡及的な分類を実行する可能性が高くなります。この論文では、英国雇用裁判所 (UKET) の判決における請求レベルの結果予測を研究することにより、この文脈における近道学習を実証的に調査しています。 33,158 件の個々のクレームのコーパスを使用して、解釈可能な TF-IDF ベースの分類器からブラックボックス LLM に至るまでのモデルを評価し、クレーム テキストと LLM で抽出された症例概要から結果を予測します。ヘッドラインの予測パフォーマンス数値は強力であるように見えますが、事後の司法テキストに基づいて訓練された LJP システムのそのようなパフォーマンスは、ソース資料の遡及的な性質によってもたらされる可能性があることを示しています。漏れに関する人間の判断によってテスト データを層別化すると、結果を明らかにする手がかりが物語に埋め込まれている場合、パフォーマンスが向上することが明らかになります。さらに、漏洩として特定された特徴のわずか 4% でトレーニングされたモデルは、人間の専門家を上回る高いパフォーマンスを実現します。これらの発見は、LJP のパフォーマンスが言語アーチファクトによって誇張される可能性があるという懸念を裏付けています。しかし、この脆弱性は研究課題にとって致命的なものではありません。代わりに、事後の判断は汚染された可能性のあるテキストとして扱われる可能性があり、積極的な監査が必要になります。リーク フィーチャをマスキングした後にモデルを再トレーニングしても、Macro-F1 は無視できるほど減少します。したがって、モデルは利用可能な場合にはショートカットを利用しますが、これらのアーティファクトが除去されても有用な予測信号を抽出する能力は維持されます。
原文 (English)
Shortcut Learning in Legal Judgment Prediction: Empirical Evidence from the UK Employment Tribunal
Current Legal Judgment Prediction (LJP) is constrained by its reliance on post-hoc judicial materials, increasing the likelihood that models perform retrospective classification rather than true forecasting. This paper empirically investigates shortcut learning in this context by studying claim-level outcome prediction in UK Employment Tribunal (UKET) decisions. Using a corpus of 33,158 individual claims, we predict outcomes from claim texts and LLM-extracted case summaries, evaluating models ranging from interpretable TF-IDF-based classifiers to black-box LLMs. While headline predictive performance figures appear strong, we demonstrate that such performance in LJP systems trained on post-hoc judicial text can be driven by the retrospective nature of the source material. Stratifying the test data by human judgments of leakage reveals that performance increases where outcome-revealing cues are embedded in the narrative. Moreover, a model trained on just the 4% of features identified as leakage achieves high performance, outperforming human experts. These findings substantiate concerns that LJP performance may be exaggerated by linguistic artefacts. Yet this vulnerability is not fatal to the research agenda. Instead, post-hoc judgments might be treated as potentially contaminated texts, requiring active auditing. Retraining models after masking leakage features results in only a negligible reduction in Macro-F1. Hence, while models will opportunistically exploit shortcuts when available, they remain capable of extracting useful predictive signals when these artefacts are removed.
大規模行動モデル: 小売顧客の迅速なデジタル ツイン
顧客行動モデリングは、レコメンデーション、マーケティング、意思決定サポートを支えていますが、既存のアプローチは、意思決定を説明せずに予測精度を最適化するか、実際の行動データに基づいてユーザーをシミュレートするかのどちらかです。私たちは、統合された個人と環境の定式化を通じて、大規模な小売取引から顧客の意思決定を直接学習する大規模行動モデル (LBM) を紹介します。顧客の状態は購入履歴から得られた行動プロファイルによって表され、製品コンテキストは検索拡張生成によって組み込まれます。このモデルは、言語化された行動データに対する継続的な事前トレーニング、意思決定生成のための教師あり微調整、および証拠に基づく調整のための検証可能な報酬を伴う強化学習を使用してトレーニングされます。購入予測、ハードネガティブな差別、バスケットの完了、プロモーションへの対応、およびクロスドメインのクーポン引き換えに関する提案されたフレームワークを評価します。このモデルは、小売業者や意思決定ドメイン全体にわたる強力なゼロショット転送と微調整された転送を実証しながら、ドメイン内の小売タスクではフロンティア汎用言語モデルを常に上回っています。アブレーション研究では、継続的な事前トレーニングが行動の一般化の主な推進力であり、検索はトレーニングと推論の両方で適用すると最も効果的であり、強化学習は一般的な言語モデルの事前学習よりも明示的な行動の証拠への依存を改善することを示しています。これらの結果は、トランザクション履歴にエンコードされた行動知識が言語モデルによって効果的に学習でき、顧客のデジタル ツインと行動シミュレーションにスケーラブルな基盤を提供できることを示しています。
原文 (English)
Large Behavior Model: A Promptable Digital Twin of the Retail Customer
Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive accuracy without explaining decisions or simulate users without grounding them in real behavioral data. We present the Large Behavioral Model (LBM) that learns customer decision making directly from large-scale retail transactions through a unified Person-Environment formulation. Customer state is represented by a behavioral profile derived from historical purchases, while product context is incorporated through retrieval-augmented generation. The model is trained using continued pre-training on verbalized behavioral data, supervised fine-tuning for decision generation, and reinforcement learning with verifiable rewards for evidence-based calibration. We evaluate the proposed framework on purchase prediction, hard-negative discrimination, basket completion, promotion response, and cross-domain voucher redemption. The model consistently outperforms frontier general-purpose language models on in-domain retail tasks while demonstrating strong zero-shot and fine-tuned transfer across retailers and decision domains. Ablation studies show that continued pre-training is the primary driver of behavioral generalization, retrieval is most effective when applied during both training and inference, and reinforcement learning improves reliance on explicit behavioral evidence over generic language-model priors. These results demonstrate that behavioral knowledge encoded in transaction histories can be effectively learned by language models, providing a scalable foundation for customer digital twins and behavior simulation.
Projection Methods for Operator Learning and Universal Approximation
We obtain a new universal approximation theorem for continuous (possibly nonlinear) operators on arbitrary Banach spaces using the Leray-Sc…
Multi-Attribute Steering of Language Models via Targeted Intervention
Inference-time intervention (ITI) has emerged as a promising method for steering large language model (LLM) behavior in a particular direct…
Transformer-Empowered Actor-Critic Reinforcement Learning for Sequence-Aware Service Function Chain Partitioning
In the forthcoming era of 6G networks, characterized by unprecedented data rates, ultra-low latency, and ubiquitous connectivity, effective…
M4V: Multimodal Mamba for Efficient Text-to-Video Generation
Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators. Howe…
Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching
Point-pixel registration between LiDAR point clouds and camera images is a fundamental yet challenging task in autonomous driving and robot…
GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
Inference-time steering methods offer a lightweight alternative to fine-tuning large language models (LLMs) and vision-language models (VLM…
Evaluating Retrieval-Augmented Generation vs. Long-Context Input for Clinical Reasoning over EHRs
Objective: To evaluate whether retrieval-augmented generation (RAG) can serve as an efficient alternative to long-context prompting for cli…
REAL: REtrieval-reAsoning and Logic-constructed Attention Behaviors for Long-Context KV Cache Compression
The growing sequence length of large language models poses significant challenges for key-value (KV) caches. Existing state-of-the-art cach…
Contrastive Weak-to-strong Generalization
Weak-to-strong generalization provides a promising paradigm for scaling large language models (LLMs) by training stronger models on samples…
Explaining Human Choice Probabilities with Simple Vector Representations
We formalize human choice behavior in a probabilistic hide-and-seek task. In our geometric construction, vectors represent participant choi…
H3Former: Hypergraph-based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification
Fine-Grained Visual Classification (FGVC) remains a challenging task due to subtle inter-class differences and large intra-class variations…
AutoGraphAD: Unsupervised network anomaly detection using Variational Graph Autoencoders
Network Intrusion Detection Systems (NIDS) are essential tools for detecting network attacks and intrusions. While extensive research has e…
Point of Order: Action-Aware LLM Persona Modeling for Data-Grounded Civic Deliberation
LLM-based simulations can enable controlled studies of civic deliberation, but current systems lack speaker-attributed data and methods for…
RIS-Assisted Downlink Pinching-Antenna Systems: GNN-Enabled Optimization Approaches
This paper investigates a reconfigurable intelligent surface (RIS)-assisted multi-waveguide pinching-antenna (PA) system (PASS) for multi-u…
Data-Driven Learnability Transition of Measurement-Induced Entanglement
Measurement-induced entanglement (MIE) captures how local measurements generate long-range quantum correlations and drive dynamical phase t…
How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy
High quality data is needed to unlock the full potential of AI for end users. However finding new sources of such data is getting harder: m…
ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning
Long-horizon manipulation has been a long-standing challenge in the robotics community. We propose ReinforceGen, a system that combines tas…
Transition Matching Distillation for Fast Video Generation
Large video diffusion and flow models have achieved remarkable success in high-quality video generation, but their use in real-time interac…
Principles of Lipschitz continuity in neural networks
Deep learning has achieved remarkable success across a wide range of domains, significantly expanding the frontiers of what is achievable i…
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization
Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives. While single preference-co…
Knowledge-Based Design Requirements for Generative Social Robots in Higher Education
Generative social robots (GSRs) powered by large language models enable adaptive, conversational tutoring but also introduce risks such as…
生成 AI による 9-1-1 電話応対トレーニングの強化: 経験と教訓
緊急通報担当者は公共安全対応における最初の運用リンクを形成しており、継続的な訓練危機に直面しながら年間 2 億 4,000 万件を超える通報に対応しています。多くのセンターで人員不足が 25% を超えており、新入社員 1 人を準備するのに最大 720 時間のマンツーマン指導が必要となり、経験豊富な職員が現役から外される可能性があります。従来のトレーニング アプローチは、これらの制約の下で拡張することが困難であり、対象範囲とフィードバックの適時性の両方が制限されます。メトロ ナッシュビル緊急通信局 (MNDEC) と協力して、現実世界の制約の下で GenAI を利用した電話応対トレーニング システムを設計、開発、展開しました。導入は 6 か月にわたって、最初のパイロットから 1,120 のトレーニング セッションを通じて 190 人の運用ユーザーまで拡大され、管理された評価や純粋にシミュレートされた評価ではほとんど見えない、システムの提供、厳格さ、回復力、人的要因に関する体系的な課題が明らかになりました。 98,429 件のユーザー インタラクション、組織プロセス、利害関係者の関与パターンを記録した導入ログを分析することで、具体的な設計とガバナンスの実践と結びついた 4 つの重要な教訓を抽出します。これらのレッスンは、実際的な制約が人間中心の設計を根本的に形作る安全性が重要な公共部門の環境で AI 主導のトレーニング システムを提供しようとしている研究者や実践者に根拠のあるガイダンスを提供します。
原文 (English)
Empowering 9-1-1 Calltaking Training with Generative AI: Experiences and Lessons Learned
Emergency call-takers form the first operational link in public safety response, handling over 240 million calls annually while facing a sustained training crisis: staffing shortages exceed 25\% in many centers, and preparing a single new hire can require up to 720 hours of one-on-one instruction that removes experienced personnel from active duty. Traditional training approaches struggle to scale under these constraints, limiting both coverage and feedback timeliness. In partnership with Metro Nashville Department of Emergency Communications (MNDEC), we designed, developed, and deployed a GenAI-powered call-taking training system under real-world constraints. Over six months, deployment scaled from initial pilot to 190 operational users across 1,120 training sessions, exposing systematic challenges around system delivery, rigor, resilience, and human factors that remain largely invisible in controlled or purely simulated evaluations. By analyzing deployment logs capturing 98,429 user interactions, organizational processes, and stakeholder engagement patterns, we distill four key lessons, each coupled with concrete design and governance practices. These lessons provide grounded guidance for researchers and practitioners seeking to deliver AI-driven training systems in safety-critical public sector environments where practical constraints fundamentally shape human-centric design.
The LLMbda Calculus: AI Agents, Conversations, and Information Flow
Large language models are increasingly deployed as agents: they plan, call tools, read untrusted data, and act on the results. This exposes…
SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution
Real-world software must continuously evolve to meet ever-changing and open-ended requirements. AI agents, increasingly deployed as long-ru…
Machine Learning for Network Attacks Classification and Statistical Evaluation of Adversarial Learning Methodologies for Synthetic Data Generation
Supervised detection of network attacks has always been a critical part of network intrusion detection systems (NIDS). Nowadays, in a pivot…
SLIDERS: Systematic Reviews via Automated Evidence Synthesis and Reconciliation
Systematic reviews -- which requires comprehensive evidence collection and synthesis from large document corpora in response to targeted re…
Tuning Derivatives for Causal Fairness in Machine Learning
Artificial-intelligence systems are becoming ubiquitous in society, yet their predictions typically inherit biases with respect to protecte…
Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue
Effective collaboration between embodied agents requires more than acting in a shared environment; it demands communication grounded in eac…
AnchorMoE: Interpretable Time Series Classification via Anchor-Routed MoE
Multivariate time series classification (MTSC) is pivotal in high-stakes domains, such as clinical diagnosis and industrial fault detection…
Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories
The past few decades have witnessed significant advances in the design of machine learning algorithms, from early studies on task-specific…
LLM は医師を支援する準備ができていますか?インタラクティブな医師、患者、EHR 支援のための PhysAssistBench
医療 LLM の最も妥当な短期的な役割は、医師の代わりではなく支援することですが、現在の評価では、臨床知識、EHR システムの相互作用、患者とのコミュニケーションなど、個別の能力がテストされることがよくあります。代わりに、医師の支援には同じ対話内でこれらの機能を調整する必要があり、医師は不明確な要求を発行し、患者は症状を曖昧に説明し、EHR システムはツールの正確な使用を要求します。インタラクティブな医師、患者、EHR 支援のベンチマークである PhysAssistBench を紹介します。実際の MIMIC-IV 症例から構築された PhysAssistBench は、スケーラブルなパイプラインを使用してエージェント性患者を構築します。これは、臨床上の事実を維持しながら、静的な EHR 記録を複数ターンの臨床シナリオに変換する、インタラクティブで記録に基づいたエージェントです。 PhysAssistBench は、手動でレビューされ医師が検証した 1,296 ターンの厳選されたバイリンガル評価セットを提供します。主要な LLM を使った実験では、この設定では現在のモデルの信頼性が依然として低いことが示されており、臨床 LLM にとって重要なボトルネックが露呈しています。信頼できる支援には、知識、コミュニケーション、システム全体の調整が必要であり、それらのいずれかで単独の利益を得るのではありません。
原文 (English)
Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance
The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communication. Physician assistance instead requires coordinating these capabilities within the same interaction, where physicians issue underspecified requests, patients describe symptoms ambiguously, and EHR systems demand precise tool use. We introduce PhysAssistBench, a benchmark for interactive doctor-patient-EHR assistance. Built from real MIMIC-IV cases, PhysAssistBench uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality. PhysAssistBench provides a curated bilingual evaluation set of 1,296 manually reviewed and physician-validated turns. Experiments with leading LLMs show that current models remain unreliable in this setting, which exposes a key bottleneck for clinical LLMs: reliable assistance requires coordination across knowledge, communication, and systems, not isolated gains in any of them.
ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL
Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Co…
GAP-GDRNet: Geometry-aware monocular 6D pose estimation for spacecraft using synthetic geometric supervision
Monocular spacecraft 6D pose estimation remains difficult under weak texture, thin structures, illumination variation, and occlusion. This…
Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language
LLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate…
Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents
Large language models are increasingly deployed as IDE-integrated coding agents that decompose tasks, generate and edit files, run code, an…
SCOReD: Student-Aware CoT Optimization for Recommendation Distillation
Chain-of-thought (CoT) distillation in the recommendation domain is a necessary precursor to RL training, but raw teacher traces are ill-su…
Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs
Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on En…
WAM-TTT: テスト時に人間のプレイを観察することで世界アクション モデルを操作する
新しいタスクのバリアントやユーザーが好む動作に向けてロボット基盤モデル (RFM) を操作することは依然として困難であり、多くの場合、追加のロボットのデモンストレーション、タスク固有の微調整、または長いコンテキストの調整が必要になります。私たちは、生の人間のビデオから世界のアクション モデルを操作するためのテスト時トレーニング フレームワークである WAM-TTT を紹介します。 WAM-TTT は、人間のビデオを模倣する軌跡として扱うのではなく、自己監視型ビデオ予測を通じて、凍結された WAM 内の軽量の適応メモリにビデオを吸収します。この記憶を制御に役立てるために、人間とロボットのペアのデータとキーと値の記憶再構成目標を使用して、人間のデモンストレーションとロボットの動作を一致させるメタトレーニング ステージを導入します。テスト時には、ラベルのない人間のビデオだけをメモリに適応させる必要があり、事前トレーニングされた WAM はフリーズされたままになります。これにより、基礎モデルの一般化機能を維持しながら、ロボットの動作、人間側の注釈、タスク固有の微調整を必要とせずに、効率的で再利用可能なステアリングが可能になります。広範な実験により、WAM-TTT は、さまざまな操作タスクや一般化設定にわたって、コンテキスト内のヒューマン ビデオ コンディショニング ベースラインよりも一貫して優れたパフォーマンスを発揮することが示されています。
原文 (English)
WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time
Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fine-tuning, or long-context conditioning. We present WAM-TTT, a test-time training framework for steering world action models from raw human videos. Rather than treating human videos as trajectories to imitate, WAM-TTT absorbs them into a lightweight adaptive memory inside a frozen WAM through self-supervised video prediction. To make this memory useful for control, we introduce a meta-training stage that aligns human demonstrations with robot behaviors using paired human-robot data and a key--value memory reconstruction objective. At test time, only unlabeled human videos are required to adapt the memory, while the pretrained WAM remains frozen. This enables efficient and reusable steering without robot actions, human-side annotations, or task-specific fine-tuning, while preserving the generalization ability of the foundation model. Extensive experiments show that WAM-TTT consistently outperforms in-context human-video conditioning baselines across diverse manipulation tasks and generalization settings.
事前トレーニング済み言語モデル埋め込みのためのリーマン幾何学
事前トレーニングされた言語モデルの埋め込みの幾何学的構造を理解することは、解釈可能性と安全性にとって重要です。文レベルの分類信号が文脈上のトークン埋め込みのリーマン幾何学に存在するかどうかを尋ね、学習されたエンコーダーの分析ヤコビアンからトークンごとのプルバック メトリックを抽出し、それらを対称正定多様体 (SPD) 上のフルエシェ平均と集計することでそれを調査します。この手順をリーマン平均プーリング (RMP) と呼びます。自明ではない言語構造を持つ 3 つのデータセット (CoLA、CREAK、RTE) にわたって、RMP はユークリッド平均プーリングよりも優れたパフォーマンスを示しましたが、アノテーション駆動型の語彙アーティファクトを除去するために構築されたベンチマークである FEVER-Symmetric では、この手法は正確に偶然性を維持しました。アブレーションの結果、ランダムに初期化されたエンコーダーとフレチェ集合体が組み合わされて、信号を含む 3 つのデータセットのうち 2 つでユークリッド プーリングをすでに上回っており、学習された多様体構造ではなく幾何学的集合体にゲインの発生源が局在していることがわかります。トレーニングされたエンコーダーは、特に 3 つの信号を含むデータセットの中で最も知識量の多い CREAK に追加信号を提供します。
原文 (English)
Riemannian Geometry for Pre-trained Language Model Embeddings
Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety. We ask whether sentence-level classification signal lives in the Riemannian geometry of contextual token embeddings, and probe it by extracting per-token pullback metrics from a learned encoder's analytical Jacobian and aggregating them with the Fr\'echet mean on the symmetric positive definite (SPD) manifold; we call this procedure Riemannian Mean Pooling (RMP). Across three datasets with non-trivial linguistic structure (CoLA, CREAK, RTE), RMP outperforms Euclidean mean pooling, while on FEVER-Symmetric, a benchmark constructed to remove annotation-driven lexical artifacts, the method correctly stays at chance. Ablations show that a randomly initialised encoder combined with Fr\'echet aggregation already beats Euclidean pooling on two of the three signal-bearing datasets, localising the source of the gain to the geometric aggregation rather than to learned manifold structure; the trained encoder contributes additional signal specifically on CREAK, the most knowledge-heavy of the three signal-bearing datasets.
オムニスリープ: CNS の階層的対照学習による睡眠基盤モデル - ANS ダイナミック
睡眠生理学は、EEG、EOG、EMG、ECG、呼吸などのマルチモーダルポリ睡眠検査信号に反映される、中枢神経系 (CNS) と自律神経系 (ANS) の協調的な動態から生じます。しかし、既存の睡眠基盤モデルは、トポロジーに依存しない方法で異種の生体信号を融合し、その生理学的組織を見落としていることがよくあります。トポロジ制約のある表現学習の生理学的事前分布として CNS/ANS パーティションを使用する睡眠基盤モデルである Omni-Sleep を紹介します。 Omni-Sleep は 3 つの目的を通じて構造化表現を学習します。1 つはシステム内の一貫性であり、神経信号および心肺信号内の共有サブシステム レベルの要素を捕捉します。システム間の同期。脳と身体のダイナミクスをモデル化するためにサブシステムの軌道を調整します。そして、長期にわたる睡眠のダイナミクスを捉える潜在空間マスク時間モデリング。 100,000 時間を超える多施設のマルチモーダル PSG データで事前トレーニングされた Omni-Sleep は、睡眠ステージングと複数の疾患分類に基づいて評価されます。データセットおよびモダリティアブレーション設定全体にわたって、Omni-Sleep は強力な基礎モデルのベースラインを上回り、ラベル効率の向上、データセット間の一般化、欠落モダリティに対する堅牢性を示しています。これらの結果は、一般化可能な睡眠表現学習における生理学的階層の価値を強調しています。コードは https://github.com/AutoBrain-sleep/OmniSleep で入手できます。
原文 (English)
Omni-Sleep: A Sleep Foundation Model via Hierarchical Contrastive Learning of CNS-ANS Dynamics
Sleep physiology arises from the coordinated dynamics of the central nervous system (CNS) and autonomic nervous system (ANS), as reflected by multimodal polysomnography signals including EEG, EOG, EMG, ECG, and respiration. However, existing sleep foundation models often fuse heterogeneous biosignals in a topology-agnostic manner, overlooking their physiological organization. We introduce Omni-Sleep, a sleep foundation model that uses the CNS/ANS partition as a physiological prior for topology-constrained representation learning. Omni-Sleep learns structured representations through three objectives: intra-system consistency, which captures shared subsystem-level factors within neural and cardio-respiratory signals; inter-system synchronization, which aligns subsystem trajectories to model brain--body dynamics; and latent-space masked temporal modeling, which captures long-horizon sleep dynamics. Pre-trained on over 100,000 hours of multi-center multimodal PSG data, Omni-Sleep is evaluated on sleep staging and multi-disease classification. Across datasets and modality-ablation settings, Omni-Sleep outperforms strong foundation-model baselines, showing improved label efficiency, cross-dataset generalization, and robustness to missing modalities. These results highlight the value of physiological hierarchy for generalizable sleep representation learning. Code is available at https://github.com/AutoBrain-sleep/OmniSleep.
Jet-Long: 動的二焦点 RoPE による効率的なロングコンテキスト拡張
最新の LLM は、検索拡張生成、リポジトリ レベルのコーディング、エージェント ワークフローなどの長いコンテキストのアプリケーションにデプロイされることが増えています。エージェント ワークフローでは、蓄積された推論とツール トレースにより、入力が日常的に事前トレーニング ウィンドウを桁違いに押し上げ、ゼロショット コンテキスト拡張がオープンウェイト チェックポイントの主要なデプロイ パスとなっています。既存のゼロショット手法のほとんどは、単一の再スケーリング係数を前もって修正するため、積極的な係数は短いコンテキストの忠実度を犠牲にし、保守的な係数は長いコンテキストでは機能しません。我々は、ローカル RoPE に忠実なウィンドウと、リスケーリング係数が現在のシーケンス長に動的に適応する長距離ウィンドウを組み合わせ、短い入力では基本モデルを正確に復元しながら、長い入力ではきれいに外挿する、チューニング不要のゼロショット手法である Jet-Long を提案します。包含-除外注意マージとオンザフライ RoPE 補正回転により、推論時に二焦点構造が本質的に自由になります。単一の CuTe カーネルに融合され、ロングコンテキストのプリフィルは H100 で最大 $1.39\times$ の FA2 スループットに達し (ホッパーのみの FA4 に近づきます)、単一バッチ生成では長さごとに $\le 4\%$ のオーバーヘッドが発生します。最大 128K コンテキストの Qwen3-1.7B/4B/8B では、Jet-Long は 1.7B/4B/8B の最も強力なベースラインを $+4.79$/$+2.18$/$+2.03$~pp 上回って RULER をリードし、HELMET-RAG (ダウンストリームのロングコンテキスト パフォーマンスの最も効率的な予測子として HELMET によって特定されたベンチマーク) で最高の総合精度を達成し、 PG-19 の困惑度は最も低い。 Jet-Long は、Jet-Nemotron などのハイブリッド アテンション アーキテクチャにも一般化して、再トレーニングなしで長いコンテキストをさらに改善し、ハイパーパラメータの回復力を維持して展開を容易にします。
原文 (English)
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE
Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces routinely push the input an order of magnitude past the pretraining window, making zero-shot context extension the dominant deployment path for open-weight checkpoints. The dominant zero-shot methods (YaRN, Self-Extend, DCA) fix a single rescaling factor up front, so an aggressive factor sacrifices short-context fidelity while a conservative one breaks down at long contexts; recent length-aware variants adapt the mapping, but with a fitted or distance-dependent schedule. We propose Jet-Long, a tuning-free zero-shot method that pairs a local RoPE-faithful window with a long-range window whose rescaling factor adapts dynamically to the current sequence length via a parameter-free analytic schedule, recovering the base model exactly at short inputs while extrapolating cleanly at long ones. An inclusion-exclusion attention merge and an on-the-fly RoPE correction rotation make the bifocal construction essentially free at inference; fused into a single CuTe kernel, long-context prefill reaches up to $1.39\times$ FA2 throughput on H100 (approaching the Hopper-only FA4), and single-batch generation incurs $\le 4\%$ overhead at every length. On Qwen3-1.7B/4B/8B up to 128K context, Jet-Long leads RULER by $+4.79$/$+2.18$/$+2.03$ pp over the strongest baseline at 1.7B/4B/8B, achieves the best overall accuracy on HELMET-RAG (a benchmark identified by HELMET as the most efficient predictor of downstream long-context performance) and attains the lowest PG-19 perplexity. Jet-Long also generalizes to hybrid attention architectures such as Jet-Nemotron for further long-context improvement without retraining, and remains hyperparameter-resilient for ease of deployment.