AIニュース 2026-08-07
自動生成: 2026-08-07 11:53 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
- WeatherNext: AI model achieves breakthrough in forecasting cyclonesGoogle DeepMind
-
Working with the American Psychological Association on youth mental health and AIOpenAI
OpenAI and the American Psychological Association advance evidence-ba…
-
「GPT-5.6 vs. Claude Fable 5」勝者はClaude、でも企業が選びづらいワケ:891st LapITmedia AI+
物理AIベンチマークで最高評価を獲得した「Claude Fable 5」。だが、企業が導入を判断する際には、性能だけでは見えない悩ましい問…
-
銀行なら3カ月→AIは1カ月 10万件のデータで「数千万円」を引き出した“データドリブン資金調達術”ITmedia AI+
黒字化目前のZehitomoは、手元資金を確保すべく新たな資金調達手段を模索していた。だが、銀行融資は最低3カ月を要し、株式調達は希薄化の…
-
あえて歩かせない――準国産ヒューマノイド「D1」登場、現場稼働で日本の勝ち筋へITmedia AI+
ZEALSは日本の屋内環境に適応した準国産の台車型ヒューマノイド「D1」を発表した。医療や製造現場など実際の稼働を通じて独自の物理データを…
-
アプリが遅い原因をAIがトレースログから分析してくれる「Windows Performance Analyzer MCP」 Microsoftがプレビュー公開ITmedia AI+
米Microsoftが、Windowsアプリケーションが遅くなる原因の調査分析をAIに依頼できるツール「Windows Performan…
-
「1人1AI」のアプローチは破綻する――チームでAI共有時のセキュリティ問題を解決するベストプラクティスITmedia AI+
Anthropicは、「Claude Tag」における「エージェントアイデンティティー」アクセスモデルの仕組みと、チームのワークスペースで…
トピック別件数
- 研究/論文 121件
- LLM/生成AI 117件
- エージェント 70件
- 画像/動画生成 51件
- ビジネス/資金調達 23件
- ロボティクス 14件
- ハードウェア/半導体 9件
- その他 8件
- 規制/政策 2件
日本語メディア14件
ITmedia AI+ (日本語)
投資の機を逃す「残念な会社がいっぱい」――ソフトバンクG投資の秘訣、“金庫番”が語る
投資会社として成果を積み上げてきたソフトバンクG。その投資方針と成功のポイントについて、同社の“金庫番”こと後藤芳光CFOが語った。
「GPT-5.6 vs. Claude Fable 5」勝者はClaude、でも企業が選びづらいワケ:891st Lap
物理AIベンチマークで最高評価を獲得した「Claude Fable 5」。だが、企業が導入を判断する際には、性能だけでは見えない悩ましい問題が浮かび上がった。
銀行なら3カ月→AIは1カ月 10万件のデータで「数千万円」を引き出した“データドリブン資金調達術”
黒字化目前のZehitomoは、手元資金を確保すべく新たな資金調達手段を模索していた。だが、銀行融資は最低3カ月を要し、株式調達は希薄化のリスクを伴う。この壁を打ち破ったのが、AIを活用した「データ駆動型融資」だ。同社が提出した10万件の入金データをAIが解析し、わずか1カ月で…
あえて歩かせない――準国産ヒューマノイド「D1」登場、現場稼働で日本の勝ち筋へ
ZEALSは日本の屋内環境に適応した準国産の台車型ヒューマノイド「D1」を発表した。医療や製造現場など実際の稼働を通じて独自の物理データを収集。フィジカルAIの社会実装を推進し、2026年度内に累計1万時間の現場稼働を目指す。
SaaSの価値は“割り勘”だけじゃない 「SaaSの死」論争を一刀両断
AIエージェントの普及でSaaSの利用者が減り、ベンダーの収入も減る――。「SaaSの死」の前提となる、この見立ては正しいのか。PM歴40年の筆者が「共同利用型システム」時代から業務ソフトウェアの歴史を振り返り、ユーザー企業にとっての価値を明らかにする。
アプリが遅い原因をAIがトレースログから分析してくれる「Windows Performance Analyzer MCP」 Microsoftがプレビュー公開
米Microsoftが、Windowsアプリケーションが遅くなる原因の調査分析をAIに依頼できるツール「Windows Performance Analyzer MCP」(WPA MCP)のアーリープレビューを発表しました。
AIデータセンターは「圧倒的に供給不足」――ソフトバンクG後藤CFO、“バブル疑惑”を否定
AI向けデータセンターは「圧倒的に供給不足だ」――ソフトバンクグループ(以下、SBG)の後藤芳光氏(取締役 専務執行役員 CFO兼CISO)は、同社の2027年3月期第1四半期連結決算(26年4月1日?6月30日)の説明会でこのように指摘した。
ソフトバンクG、投資利益1.8兆円を支えた「OpenAIではない“あの半導体メーカー”」の正体
ソフトバンクGは、第1四半期の投資利益が1兆8594億円だったと発表した。投資利益を押し上げたのは、OpenAIでもArmでもない。歴史的な経営難に陥っていた“あの半導体メーカー”だった。
Googleのジェフ・ディーン氏、独立してAI実験を大規模自動化する新会社Discovery Loop設立
Googleのチーフサイエンティスト、ジェフ・ディーン氏が退社し、新会社「Discovery Loop」を設立すると発表した。サンジェイ・ゲマワット氏らと共同創業する公益法人で、フロンティアAIモデルと計算基盤を活用して実験や検証などの研究プロセス全体を自動化することを目指す。…
書店に「3000冊の発注」、AI企業が古書を買いあさる? Anthropicも数百万冊をスキャン・破棄 実態明らかに
オランダの古書店に、3000冊の注文が入った。届け先は、中国のAI企業。AI学習用データにするとみられる。Anthropicも数百万冊をスキャンした後、破棄している。その実態が明らかになった。
「1人1AI」のアプローチは破綻する――チームでAI共有時のセキュリティ問題を解決するベストプラクティス
Anthropicは、「Claude Tag」における「エージェントアイデンティティー」アクセスモデルの仕組みと、チームのワークスペースでこれを構成する際のベストプラクティスを解説したブログ記事を公開した。
中国DeepSeek、近日中に「大幅値上げ」か API料金ページに追記
中国のDeepSeekは、同社が提供するAPIの料金ページに「近日中に大幅な値上げが見込まれる」と記載した。
NVIDIA、自動運転向けオープンモデルを商用利用可に 新モデルは「卓越した性能」うたう
NVIDIAが自動運転向けAIモデル「Alpamayo」ファミリーを商用利用可能なオープンライセンスで提供開始。新モデル「Alpamayo 2 Super」は推論ベンチマークで首位になるなど卓越した性能をうたう。
リクルート、新卒エンジニア向け研修資料を無料公開 “AI時代の生き残り方”など紹介する13本
リクルートは、2026年度の新卒エンジニア向けの研修資料を公開した。エンジニアとして生き残るためのAIの活用法や、AIによるソフトウェア開発の考え方の変化など、業務やキャリア形成に必要な知見を幅広く紹介している。
海外メディア9件
TechCrunch AI (英語)
OpenAI’s new AI smart speaker will reportedly sell for between $300 and $400
Additional details about OpenAI's mysterious new AI device make it sound like a pricey smart speaker.
Naïve raises $28.5M to automate the grunt work of setting up and running a company
Taking vibe-coding a step further, Naïve claims its infra can automate most of the work in setting up and running a business.
Gen Z dating apps like Ditto ditch swiping in favor of AI matchmaking
This generation of twentysomethings is so disillusioned with swipe-based dating apps that they'll try literally anything else — even an AI…
OpenAI says Apple’s own security practices undermine its trade secrets case
Newly filed court exhibits show OpenAI’s legal strategy in Apple’s trade secrets lawsuit: argue that Apple’s own security and offboarding p…
Amid legal battles, Suno says it will start watermarking songs
Suno's watermarking feature comes as the company is fighting legal battles on several fronts.
Ex-Spotify employees raise $10M to bring the AI behind its recommendations to e-commerce
The startup's platform predicts which product a shopper wants next, learns their general taste, and fine-tunes continuously based on what t…
Exclusive: Mirendil inks $100M+ Google Cloud deal to scale self-improving AI
Mirendil has signed a $100 million-plus Google Cloud partnership to expand its compute infrastructure, powering research into self-improvin…
Google Maps adds agentic features, including food ordering and hotel bookings
The launch of these new features reflects Google’s ambitions to transform Google Maps from a navigation tool into an assistant that's capab…
Omilia raises $67M to scale its customer support platform
The Series B is the company's second fundraise since it last raised capital in 2020. In that time, it has increased its ARR by 10x to $60 m…
公式ブログ2件
OpenAI (英語)
Working with the American Psychological Association on youth mental health and AI
OpenAI and the American Psychological Association advance evidence-based guidance, resources, and safeguards for responsible AI use and you…
Google DeepMind (英語)
論文287件
arXiv cs.AI (英語)
冗長性調整人工年齢スコア (AAS) における AI システムの長期永続性理論
人工知能システムは、個別のワンショット出力ではなく、対話、適応、更新の繰り返しサイクルで動作することがますます期待されています。これは基本的な理論上の疑問を引き起こします: AI システムは際限のない構造的老化を招くことなく無期限に存続できるのでしょうか?この論文では、冗長性を調整した Artificial Age Score (AAS) に基づいた AI システムの長期持続性フレームワークを開発します。このモデルは、AAS を静的評価尺度から、反復操作にわたる年齢シーケンスを生成するサイクル レベルの汎関数に拡張します。各サイクルで、コンポーネントの整合性レベルに対する重み付けされた冗長性を意識した対数ペナルティによって構造年齢が定義されます。この枠組みの中で、周期レベルの年齢は明確に定義され、一様に制限されており、それによって爆発的な点ごとの老化が排除されることが示されています。これに基づいて、この論文は、負荷持続性、ゼロ負荷持続性、振動持続性、および累積終末期負担を含む漸近レジームの階層を定義します。また、比較順序付け、感度限界、成分ごとの安定化下での収束、有限総変動下での持続性、減衰されたサイクル間摂動下での幾何学的安定化、および非縮退冗長条件下でのゼロ負荷特性評価も確立します。主な結果は、無期限の周期的継続には無制限の構造老化が必要ないということです。AI システムは、構造年齢が制限されたままで無限に多くのサイクルを通過する可能性がありますが、より強い規則性条件下では限界老化は消滅し、最も強力な体制では、そのサイクル レベルの負担はゼロに収束します。したがって、このフレームワークは、長期的な人工持続性を、避けられない累積的な劣化ではなく、限定された構造的負担の問題として分析するための正式な基礎を提供します。
原文 (English)
A Long-Run Persistence Theory for AI Systems under the Redundancy-Adjusted Artificial Age Score (AAS)
Artificial intelligence systems are increasingly expected to operate over repeated cycles of interaction, adaptation, and update rather than through isolated one-shot outputs. This raises a fundamental theoretical question: can an AI system persist indefinitely without incurring unbounded structural aging? This paper develops a long-run persistence framework for AI systems based on the redundancy-adjusted Artificial Age Score (AAS). The model extends AAS from a static evaluative measure into a cycle-level functional that generates an age sequence across repeated operation. At each cycle, structural age is defined through a weighted, redundancy-aware logarithmic penalty over component consistency levels. Within this framework, cycle-level age is shown to be well defined and uniformly bounded, thereby excluding explosive pointwise aging. On this basis, the paper defines a hierarchy of asymptotic regimes, including burdened persistence, zero-burden persistence, oscillatory persistence, and cumulative terminal burden. It also establishes comparative ordering, sensitivity bounds, convergence under componentwise stabilization, persistence under finite total variation, geometric stabilization under damped inter-cycle perturbations, and a zero-burden characterization under nondegenerate redundancy conditions. The main result is that indefinite cyclic continuation does not require unbounded structural aging: an AI system may pass through infinitely many cycles while its structural age remains bounded, while under stronger regularity conditions its marginal aging vanishes and, in the strongest regime, its cycle-level burden converges to zero. The framework thus provides a formal basis for analyzing long-run artificial persistence as a problem of bounded structural burden rather than inevitable cumulative deterioration.
LLM が提案し、執行部が処分: 長期エージェントにおけるコミットメント ドリフトとバインディング ドリフトを分離する自己検証エージェント手段
長期にわたるエージェント自身の状態や自己報告がまさに信頼できないものである場合、どのようにしてそのエージェントを検証するのでしょうか?検証が事後的ではなく構造的に行われるように構築されたエージェント手段を紹介します。決定論的な経営者はすべての信念を所有します。言語モデルは入力された提案のみを提出でき、行動前に事前に登録された予測がコードによる観察と一致する場合にのみクレームが認められます。 2 つの特性により、この機器はエージェントだけでなく、それ自体の科学の検証者になります。臓器ごとの書き込みエラー、レンダリング サイズ、またはソルテッド カナリア エコーのフロアが突破されると、すべての実行がそれ自体を無効にします (最初の 8 つのアーキテクチャ実行のうち 4 つが無効になり、それぞれが実際の欠陥の位置を特定しました)。また、レンダリング不可視のシャドウ参照は、システム全体がすべてのアブレーション セルでコミットする計画をコンパイルするため、テスト対象の機構が削除された場合でもドリフト メトリックが定義されます。この手段を使用して、すべての長距離エージェントが経験する障害に関するクリーンな単一変数の結果を報告します。コミットメント メカニズムの除去により、目標放棄が 0.00 から 1.00 に反転しますが、バインディング エラーは 0.00 で横ばいのままです (セルあたり 3 つのシード、実行あたり最大 394 の参照ビート、すべての実行ゲートが有効)。対照的に、結合チャネルは、その修復が除去されてもビートごとのドリフトとして再出現しません。結合はコードが所有しており、失敗クラスは構造的に吸収され、その唯一の残留物が仮説形成の崩壊として 1 層上流に現れるためです。我々は、タスクの有効性がゼロ(ARC-AGI-3 での 52 回のゲート実行全体でゼロレベル完了)であり、構造的敗北者として事前に登録されているという完全な開示の下でこれらを報告します。この貢献は、エージェント開発の検証方法論と、それによって測定可能になるドリフト分解です。
原文 (English)
The LLM Proposes, the Executive Disposes: A Self-Verifying Agent Instrument that Dissociates Commitment Drift from Binding Drift in Long-Horizon Agents
How do you verify a long-horizon agent when its own state and self-reports are exactly what you cannot trust? We present an agent instrument built so that verification is structural rather than post-hoc. A deterministic Executive owns all belief; a language model may only file typed proposals, and a claim is admitted only when a prediction pre-registered before acting is matched against observation by code. Two properties make the instrument a verifier of its own science, not just of the agent: every run invalidates itself when per-organ write-error, render-size, or salted-canary-echo floors are breached (four of the first eight architecture runs were invalidated, each localizing a real defect); and a render-invisible shadow reference compiles the plan the full system would have committed in every ablation cell, so drift metrics are defined even where the mechanism under test has been removed. Using this instrument we report a clean, single-variable result on a failure every long-horizon agent suffers: ablating the commitment mechanism flips goal-abandonment from 0.00 to 1.00 while binding error stays flat at 0.00 (three seeds per cell, up to 394 reference beats per run, every run gated valid). The binding channel, by contrast, does not reappear as per-beat drift when its repair is ablated -- because binding is code-owned, the failure class is structurally absorbed, its only residue appearing one layer upstream as a collapse in hypothesis formation. We report these under full disclosure that task efficacy is null (zero level completions across 52 gated runs on ARC-AGI-3), pre-registered as a structural defeater. The contribution is a verification methodology for agent development and the drift decomposition it makes measurable.
テーブルからマルチモーダルなレポート生成のためのモンテカルロ ツリー検索
構造化された表形式データからテキスト分析とビジュアル チャートの両方を含むプロフェッショナルなマルチモーダル レポートを自動的に生成することは、データ インテリジェンスにおける重要な課題です。既存の方法には、固定された線形パイプラインと分離されたサブタスク処理という問題があり、事実の正確さ、視覚的な品質、物語の一貫性の共同最適化を妨げています。これらの問題に対処するために、この文書では、構造化された検索空間に対する漸進的な構築プロセスとしてマルチモーダルなテーブルからレポートへの生成を定式化するモンテカルロ ツリー検索 (MCTS) ベースのフレームワークである MCTS-Report を提案します。中心となるアイデアは、レポートの生成を、章の計画、視覚化タスクの特定、チャートの生成、洞察の整理、ナラティブの洗練などのアトミックなアクションに分解することであり、各アクションは、現在のレポートの状態を条件とした動的な推論に基づいて LLM によって実行されます。 LLM を使用して、MCTS 中に段階的な推論とアクションを生成し、推論の軌跡を各ノードに保存して、コンテキストを認識した一貫したレポートを構築します。検索をガイドするために、数値的な事実の一貫性 (SQL 経由)、グラフの品質、グラフとテキストの配置、構造の完全性を共同で評価する多次元の報酬関数を設計します。同時に、グラフの繰り返しを抑制する多様性ペナルティと、無効なアクションを取り除くための前提条件チェックを組み込みます。また、6 つのドメインの実世界のテーブルで構成され、専門家が洗練した参照レポート構造と検証可能な重要な洞察を組み合わせた包括的なベンチマークである MMRBench も構築します。 MMRBench での実験では、MCTS-Report が構造の完全性、数値精度、チャートとテキストの整合性、洞察の新規性のすべてにおいて強力なベースラインを大幅に上回り、総合スコア 77.9 を達成していることが実証されました。
原文 (English)
Monte Carlo Tree Search for Table-to-Multimodal Report Generation
Automatically generating professional multimodal reports comprising both textual analysis and visual charts from structured tabular data is a critical challenge in data intelligence. Existing methods suffer from fixed linear pipelines and isolated subtask processing, which hinder joint optimization of factual accuracy, visual quality, and narrative coherence. To address these issues, this paper proposes MCTS-Report, a Monte Carlo Tree Search (MCTS)-driven framework that formulates multimodal table-to-report generation as a progressive construction process over a structured search space. The core idea is to decompose report generation into atomic actions, including chapter planning, visualization task identification, chart generation, insight organization, and narrative refinement, each executed by an LLM based on dynamic reasoning conditioned on the current report state. We use an LLM to generate step-by-step reasoning and actions during MCTS, storing the reasoning trajectory in each node for context-aware, coherent report construction. To guide the search, we design a multi-dimensional reward function that jointly evaluates numerical fact consistency (via SQL), chart quality, chart-text alignment, and structural completeness, while incorporating a diversity penalty to suppress repeated charts and a precondition check to prune invalid actions. We also construct MMRBench, a comprehensive benchmark comprising real-world tables from six domains, paired with expert-refined reference report structures and verifiable key insights. Experiments on MMRBench demonstrate that MCTS-Report significantly outperforms strong baselines across structural completeness, numerical accuracy, chart-text alignment, and insight novelty, achieving a 77.9 overall score.
FinProBench: 専門的な成果物から派生した役割に基づいたルーブリックを使用して金融 AI エージェントを評価する
金融 AI エージェントを評価するには、実際の専門家の仕事に合わせた基準が必要です。既存のルーブリック手法は通常、タスクのプロンプトまたはモデルの出力から基準を導き出し、実践者の成果物にのみ表示される暗黙の基準を見落としています。プロフェッショナルな財務タスクのベンチマークである FinProBench と、同じ役割の実務者が作成した成果物からルーブリックを導き出す再利用可能なパイプラインである、Role-Grounded Rubric Construction (RGRC) を紹介します。 RGRC は、成果物の収集、コンピテンシーの抽出、ルーブリックの合成、および検証の 4 つの段階で構成されます。そのルーブリックは暗黙の基準を捉え、品質レベルを区別し、役割内のタスク間で転送します。分析の前に、成果物のジャンルごとに 57 の職業を、事前に豊富な従来型の 30 の役割と、事前に少ない役割に特化した 27 の役割に分類しました。すべてのロールにわたって、プロンプトのみは従来のロールでは RGRC とほぼ一致しますが (89.2% 対 90.7%)、ロールに特化したロールでは RGRC のパフォーマンスが大幅に優れています (99.1% 対 78.0%)。この分割は、慣例がモデルの事前分布で適切に表現されている場合には迅速なエンジニアリングでルーブリックに近似できる一方、それらの事前分布を超える標準には専門的な基礎が不可欠であることを示しています。 FinProBench は、57 の職業、8 つの金融サブ業界、161 の成果物のタイプにわたる 1,723 の厳選された成果物から構築されており、7 つのサブ業界の 20 の役割をカバーする 20 の完全なタスクの初期評価セットをリリースします。異種の LLM 審査員と役割レベルのルーブリックを使用すると、人間の成果物は平均して 1 位にランクされます (100 点中 73.7 対 70.3、70.2、および 69.6)。その一方で、4 つのシステムすべてが重複する 95% 信頼区間と補完的な強みを示しています。役割レベルでルーブリックを再利用すると、各ルーブリックを最初から作成する場合と比較して、タスクごとの推定構築労力が 6.7 倍削減されます。
原文 (English)
FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables
Evaluating financial AI agents requires criteria aligned with real professional work. Existing rubric methods typically derive criteria from task prompts or model outputs, overlooking tacit standards visible only in practitioner deliverables. We introduce FinProBench, a benchmark for professional financial tasks, and Role-Grounded Rubric Construction (RGRC), a reusable pipeline that derives rubrics from deliverables produced by practitioners in the same role. RGRC comprises four stages: Deliverable Collection, Competency Extraction, Rubric Synthesis, and Validation. Its rubrics capture tacit standards, distinguish quality levels, and transfer across tasks within a role. Before analysis, we classified 57 occupations by deliverable genre into 30 prior-rich conventional roles and 27 prior-sparse role-specialized roles. Across all roles, Prompt-only nearly matches RGRC for conventional roles (89.2% vs. 90.7%), but RGRC substantially outperforms it for role-specialized roles (99.1% vs. 78.0%). This split indicates that prompt engineering can approximate rubrics when conventions are well represented in model priors, while professional grounding is essential for standards beyond those priors. FinProBench is built from 1,723 curated deliverables spanning 57 occupations, 8 financial sub-industries, and 161 deliverable types, and releases an initial evaluation set of 20 complete tasks covering 20 roles in 7 sub-industries. With heterogeneous LLM judges and role-level rubrics, human deliverables rank first on average (73.7 vs. 70.3, 70.2, and 69.6 out of 100), while all four systems show overlapping 95% confidence intervals and complementary strengths. Reusing rubrics at the role level reduces estimated per-task construction effort by 6.7 times relative to authoring each rubric from scratch.
FinPerMA: LLM エージェント向けの、理論に基づいたイベントベースのパーソナライズされたメモリ ベンチマーク
大規模言語モデル (LLM) エージェントは、財務アドバイスなどのリスクの高い分野で個人化されたアシスタントとして使用されることが増えていますが、長期にわたって個別化されたユーザー モデルを維持および更新できるかどうかは依然として不明です。既存のパーソナライズされた記憶ベンチマークは、主に事実の保持をテストするか、弱い制約のモデルによって生成された軌道に依存するため、イベント駆動型の嗜好の適応は十分に検討されていません。 FinPerMA は、凍結された長期的な投資家の軌跡に対して個人化された記憶を評価するイベントベースのベンチマークです。その生成パイプラインは、決定論的で理論に基づいた影響ルール、制御された LLM ナレーション、および自動化された品質スクリーニングを組み合わせています。ショック後のチェックポイントは、エージェントが重要なイベントを永続的なユーザー モデルに統合したかどうかを分離します。 276 人のペルソナからの 2,994 の質問について、7 つのフロンティア LLM と最大 7 つのメモリ構成は依然として飽和状態には程遠い状態です。全体の精度が約 0.47 であるか、多肢選択式の質問では約 39% を超えるフルコンテキスト構成はありません。アトリビューション分析によると、概要ベースの記憶は事実の詳細を保持する一方で、パーソナライゼーションに必要な好みのシグナルを失うことがよくあります。したがって、単純な検索は専用のメモリシステムよりも優れたパフォーマンスを発揮する可能性があり、ショック後にその差は拡大します。
原文 (English)
FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents
Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons. Existing personalized-memory benchmarks primarily test factual retention or rely on weakly constrained model-generated trajectories, leaving event-driven preference adaptation underexplored. We introduce FinPerMA, an event-grounded benchmark that evaluates personalized memory against frozen longitudinal investor trajectories. Its generation pipeline combines deterministic, theory-informed impact rules, controlled LLM narration, and automated quality screening; a Post-Shock checkpoint isolates whether an agent has integrated a material event into its persistent user model. On 2,994 questions from 276 personas, seven frontier LLMs and up to seven memory configurations remain far from saturated: no full-context configuration exceeds approximately 0.47 overall accuracy or approximately 39% on multiple-choice questions. Attribution analysis shows that summary-based memory often preserves factual details while losing the preference signals needed for personalization; simple retrieval can therefore outperform purpose-built memory systems, with the gap widening after shocks.
BrainBench: 包括的な脳波理解のための大規模言語モデルのベンチマーク
脳波 (EEG) 分析は、録音に事前定義されたラベルを割り当てるだけではありません。自然言語の指示、信号処理、定量的証拠、科学的解釈を結び付けるワークフローが必要です。私たちはこの能力を \emph{包括的な脳波理解} と呼びます。しかし、既存の評価は主に分離されたデコードタスクやシステム固有のデモンストレーションを対象としており、大規模言語モデル (LLM) の能力の定量化が不十分なままになっています。 \benchmarkname{} は、包括的で命令条件付きの EEG 理解のための統合ベンチマークです。これは、基礎分析、睡眠評価、神経認知評価、生理学的統合の 4 つのサブセットで構成されており、17 のデータセット、\numcases{} 個のタスク、および \numinstances{} 個以上の実データ インスタンスをカバーしています。指示とオプションの生理学的信号を含む脳波記録が与えられると、システムは分析を実行し、科学的に根拠のあるレポートと、必要に応じてアーティファクトを作成する必要があります。出力は、数値、カテゴリ、セット、シーケンス、セマンティック、およびアーティファクトの検証を通じて評価されます。私たちは、CodeAct による自律的なコード実行と BrainAgent による構造化エージェント分析という 2 つのパラダイムの下で、10 万を超える実行にわたって \nummodels{} の代表的な LLM を評価しました。結果はモデル、サブセット、難易度、実行パラダイムによって大きく異なり、EEG 能力がモデルとその運用に依存することがわかります。 \benchmarkname{} は、LLM ベースの EEG の理解を進めるための再現可能なテストベッドを提供します。コードとベンチマークは間もなくリリースされ、評価結果は継続的に更新されます。
原文 (English)
BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding
Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural-language instructions, signal processing, quantitative evidence, and scientific interpretation. We term this capability \emph{comprehensive EEG understanding}. Existing evaluations, however, primarily target isolated decoding tasks or system-specific demonstrations, leaving the competence of large language models (LLMs) insufficiently quantified. We introduce \benchmarkname{}, a unified benchmark for comprehensive, instruction-conditioned EEG understanding. It comprises four subsets---Foundational Analysis, Sleep Assessment, Neurocognitive Assessment, and Physiological Integration---covering 17 datasets, \numcases{} tasks, and over \numinstances{} real-data instances. Given an instruction and EEG recordings with optional physiological signals, a system must perform the analysis and produce a scientifically grounded report and, when required, artifacts. Outputs are assessed through numerical, categorical, set, sequence, semantic, and artifact validation. We evaluate \nummodels{} representative LLMs across more than 100K executions under two paradigms: autonomous code execution with CodeAct and structured agentic analysis with BrainAgent. Results vary substantially across models, subsets, difficulty levels, and execution paradigms, showing that EEG competence depends on the model and its operationalization. \benchmarkname{} provides a reproducible testbed for advancing LLM-based EEG understanding. The code and benchmark will be released soon, with evaluation results continuously updated.
事前訓練されたトランスフォーマーベースの知覚モデルの敵対的に堅牢なアブダクティブフュージョン
事前にトレーニングされた知覚モデルを新しい環境に導入すると、分布シフトの下で精度が低下し、それらを組み立てるだけでは精度が回復しません。多数決などの結合器は精度を求めてリコールをトレードし、調整された障害に対して脆弱です。従来のメタ認知手法は、モデルのエラーにフラグを立てる論理ルールを学習しますが、真に新しいシーンには転送されない、手動で作成されたドメイン知識の手がかり (オブジェクト サイズの事前分布、セグメンテーション マスク) に依存しています。このメタ認知層は、ベクトル空間ジオメトリを活用することで、ドメイン知識がなくても学習できることを示します。モデルごとのラベル ベクトル プール (LVP) は、各モデル独自のトレーニング埋め込みから構築され、トレーニングで決定されたプロトタイプに対する検出のジオメトリからエラー検出ルールを生成し、テスト セットの F1 ごとに 0.002 ドル以内でドメイン知識ルールと同等に達します。このアプローチは神経象徴的なままであるため、これらの幾何学的規則は単一の論理フレームワークを共有し、利用可能な場合にはドメイン知識によって補完することができます。複数の不完全な ViT ベースの検出器の融合を、正確な整数プログラム (IP) と多項式時間ヒューリスティックによってテスト時に解決される一貫性ベースのアブダクション問題として組み立てます。 15 の天候変動テスト セットと 6 つの ViT 検出器の航空画像ベンチマークでは、ドメイン知識フリー レイヤーはクリーン データ ($0.005$ F1 以内) で最も強い多数決のバリアントと一致し、他の多数決ベースラインとは異なり、協調的なラベル反転攻撃下でもパフォーマンスを維持します。$90\%$ のフリップ レートでは、平均 $0.42$ F1 対 $0.35$ です。 MV-Plurality の場合 ($22\%$ の相対ゲイン)、フリップ レートが $0.4$ を超えると \emph{every} テスト セットで最高の F1 を達成します。
原文 (English)
Adversarially Robust Abductive Fusion of Pre-trained Transformer-based Perception Models
Deploying pre-trained perception models in novel environments degrades their accuracy under distributional shift, and assembling them alone does not recover it: combiners such as majority voting trade recall for precision and are brittle to coordinated failures. Prior metacognitive methods learn logical rules that flag a model's errors, but rely on hand-authored domain-knowledge cues (object-size priors, segmentation masks) that do not transfer to genuinely novel scenes. We show that this metacognitive layer can be learned without any domain knowledge by exploiting vector-space geometry: per-model Label Vector Pools (LVP), built from each model's own training embeddings, yield error-detection rules from the geometry of detections relative to training-determined prototypes, reaching parity with domain-knowledge rules to within $0.002$ every F1 on test set. Because the approach remains neurosymbolic, these geometric rules share a single logical framework and can still be complemented by domain knowledge when available. We frame the fusion of multiple imperfect ViT-based detectors as a consistency-based abduction problem solved at test time by an exact Integer Program (IP) and a polynomial-time heuristic. On an aerial-imagery benchmark of 15 weather-shifted test sets and six ViT detectors, our domain-knowledge-free layer matches the strongest majority-vote variant on clean data (within $0.005$ F1) and, unlike every majority-vote baseline, retains its performance under a coordinated label-flipping attack: at a $90\%$ flip rate it averages $0.42$ F1 versus $0.35$ for MV-Plurality (a $22\%$ relative gain) and attains the highest F1 on \emph{every} test set once the flip rate exceeds $0.4$
MatrAIx: 83 億のペルソナ エージェントによる世界のシミュレーション
AI システムやデジタル製品を人間が評価するのはコストがかかり、時間がかかり、拡張するのが困難です。オフライン評価はよりスケーラブルですが、多くの場合、人間の多様性とインタラクティブな行動が抽象化されます。そこで、異種ユーザーによる AI システムやデジタル製品をテストするための人口規模のシミュレート ユーザー評価インフラストラクチャである MatrAIx を紹介します。 MatrAIx には 3 つのコア コンポーネントがあります。 まず、ペルソナ 8B には、1,290 のカテゴリ ディメンションで表される 83 億のペルソナ レコードが含まれています。レコードは、相関関係のある属性を保持する依存関係グラフからサンプリングされるか、人間が作成したプロファイルから派生します。当社は、599,847 件の人間に基づいたレコードと 400,000 件の合成レコードで構成される、約 100 万人のペルソナの高品質フィルター処理されたコアセットをリリースします。 2 番目に、MatrAIx プレイグラウンドは、多様なユーザーがデジタル製品を評価し、操作するための 4 つの環境 (アンケート、AI チャットボット、Web、アプリ) を提供します。 3 番目に、MatrAIx は、コマース、ソフトウェア、財務、ヘルスケアを含む 25 以上のドメインにわたる 1,010 のアプリケーション タスクを提供します。私たちは 8 つの代表的なタスクにわたって 18,189 件の評価トライアルを実施しました。ペルソナ エージェントは、Claude Opus 4.8、GPT 5.5、および Claude Haiku 4.5 の 3 つの LLM を利用していました。結果として得られるフィードバックは、値上げ後のためらい、AI アシスタントが失敗した後の続行意欲、待ち時間の許容度など、ペルソナの背景によって意思決定や好みがどのように異なるかを捉えています。私たちは 2 つの主要な検証研究を実施しました。まず、400 件の試験を対象とした対照研究で、10 の行動特性と 4 つすべての環境にわたるペルソナの遵守度を評価しました。宣言された行動は、366 件の試験 (91.5%) で発現または正しく抑制されました。次に、人間と LLM の審査員が、人間に基づいたペルソナの抽出品質を評価しました。全体として、MatrAIx は、さまざまなシミュレートされた人間のユーザーを使用して AI システムとデジタル製品を評価するためのエンドツーエンドのインフラストラクチャを提供します。
原文 (English)
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
採食エージェントにおける動的恒常性優先順位付けとしての内受容的注意
生物学的システムは、限られた知覚帯域幅の下で競合するニーズを制御する必要があり、一方の推定値を鮮明にするためには、他の推定値を鮮明にする能力が犠牲になります。したがって、固定予算のシステムでは、その知覚精度をどこに割り当てるかを決定する必要があります。私たちは、生きていくためにいくつかの身体的ニーズを満たさなければならない採食エージェントについて、能動的推論でモデル化してこれを研究します。各ステップで、それは自身の身体状態の信念を読み取り、最も必要なチャネルを特定し、それに向けて内受容精度の固定予算を再割り当てします。これにより、同じ精度で成形された尤度が信念の更新と計画の両方に供給されます。 4 チャネルの採集グリッドワールドである AffectWorld では、この選択的割り当てにより、均一精度のエージェントに対して一致する予算での学習フェーズの生存率が 2 倍以上になります (11 レイアウト全体で $0.414$ 対 $0.199$、それぞれ $n{=}32$ シード、ペアのクラスター ブートストラップ $p \leq 10^{-4}$)。さらに 2 つの結果がメカニズムを明確にします。計画立案者だけが形成された可能性を否定するだけで、その約半分が取り除かれるため、この利点は計画と認識を通じてもたらされます。また、最も必要のないチャネルでの照準精度は均等に分散するよりも悪いため、必要に応じて調整されます。さらに、在席チャネルは自身のダイナミクスを約 2 倍の速さで学習し、一致した観測数でも先を行き続けます。これは同じ高精度ルーティングの動作トレースであり、生存ではなく学習速度で確認できます。
原文 (English)
Interoceptive Attention as Dynamic Homeostatic Prioritization in a Foraging Agent
Biological systems must regulate competing needs under limited perceptual bandwidth, where sharpening one estimate costs the capacity to sharpen the others. Any fixed-budget system therefore has to decide where to allocate its perceptual precision. We study this in a foraging agent that must keep several bodily needs satisfied to survive, modelled with active inference. At each step it reads its own body-state beliefs, identifies the most-needed channel, and reallocates a fixed budget of interoceptive precision toward it, so that the same precision-shaped likelihood feeds both belief update and planning. In AffectWorld, a four-channel foraging gridworld, this selective allocation more than doubles learning-phase survival at matched budget against a uniform-precision agent ($0.414$ vs $0.199$ across 11 layouts, $n{=}32$ seeds each, paired cluster-bootstrap $p \leq 10^{-4}$). Two further results sharpen the mechanism. The benefit runs through planning as well as perception, since denying the shaped likelihood to the planner alone removes about half of it. It is also need-aligned, since aiming precision at the least-needed channel does worse than spreading it evenly. The attended channel additionally learns its own dynamics about twice as fast, and stays ahead even at matched observation count, a behavioural trace of the same precision routing, visible in learning speed, not survival.
神経記号 AI の RAIL 原則: 推論、保証、インターフェース、学習
機械学習と記号推論を統合した神経記号 AI システムが急速に注目を集めています。これらは、ニューラル ネットワークや言語モデルのデータ集約型の統計的アプローチを記号推論アルゴリズムで補完し、現実世界のアプリケーションの多くを特徴づける、一か八かの領域や低データ領域で機能します。私たちは、機械学習と形式的推論の神経象徴的な組み合わせは AI におけるニッチなアプローチではなく、信頼性があり効率的で、最終的には信頼できるシステムの開発にとって極めて重要である、すでに成功している多くの手法が含まれていると主張します。この視点は、現在の AI システムの設計の再検討を促します。私たちは、伝統的に神経象徴的であると考えられていないものを含む、多くの主要な AI システムが、神経象徴的 AI 設計の 4 つの原則、推論、保証、インターフェイス、学習 (RAIL) の観点から分析できることを示します。 RAIL フレームワークを適用すると、物理学を意識した機械学習から神経誘導検索 (Google DeepMind の Alpha-* スイートなど)、因果学習、ツールで拡張された大規模言語モデルに至るまで、一見異質な AI システムの統一されたビューが提供されます。重要なのは、RAIL 原則により、エンジニアは実稼働レベルの AI システムの設計と導入に関して、より適切な情報に基づいて、より原則に基づいた意思決定を行うことができるようになります。この記事では、RAIL 原則を紹介し、AI の主要分野全体に RAIL 原則がどのように適用できるかを検討し、実践者が神経象徴的な手法を次世代 AI テクノロジーに統合する方法を説明します。
原文 (English)
The RAIL Principles for Neurosymbolic AI: Reasoning, Assurances, Interfacing and Learning
Neurosymbolic AI systems that integrate machine learning and symbolic reasoning are rapidly gaining attention. They complement the data-intensive statistical approaches of neural networks and language models with symbolic reasoning algorithms to function in high-stakes domains or in low-data regimes that characterize many real-world applications. We argue that the neurosymbolic combination of machine learning and formal reasoning is not a niche approach within AI, but rather includes many already successful techniques that are of crucial importance to the development of reliable, efficient and, ultimately, trustworthy systems. This perspective prompts a re-examination of the design of current AI systems. We show that many leading AI systems, including some that are not traditionally considered as neurosymbolic, can be analysed from the perspective of four principles of neurosymbolic AI design: Reasoning, Assurances, Interfacing and Learning (RAIL). Applying the RAIL framework offers a unified view of seemingly disparate AI systems, ranging from physics-aware machine learning to neuro-guided search (such as Google DeepMind's Alpha-* suite), causal learning and tool-augmented Large Language Models. Importantly, the RAIL principles will enable engineers to make better-informed and more principled decisions about the design and deployment of production-level AI systems. In this article, we introduce the RAIL principles, examine how they can be applied across major areas of AI, and illustrate how they may guide practitioners to integrate neurosymbolic methods into next-generation AI technologies.
SafeCommit: メモリベースのエージェントが安全に動作できることを証明する
長期的なエージェントは、外部の副作用を伴うアクションを実行するために永続メモリとツールをますます使用しています。中心的な障害モードは時期尚早のコミットメントです。エージェントは、メモリ基盤が古いか、競合しているか、不完全であるか、破損しているかを解決する前に動作します。我々はこの問題をメモリの不確実性の下での安全なコミットメントとして形式化し、エージェント推論と外部実行の間のリスク制御層である SafeCommit を導入します。この層は、記憶、観察、ツールの出力、来歴、およびポリシーの制約から、もっともらしい潜在世界の調整されたセットを構築します。副作用のあるアクションは、保持されているすべての世界においてそのアクションが安全であることが適合アクション証明書によって示されている場合にのみ許可されます。それ以外の場合は、ワールドブロッキング認証を対象とする副作用の少ないプローブを選択するか、保守的なフォールバックを返します。調整された世界の適用範囲では、安全でない認定コミットの確率は最大でも目標レベル {\alpha} です。不完全な世界提案では、境界によってキャリブレーションと表現エラーが分離されます。依存関係のない制御されたシミュレーターは、安全性とユーティリティのトレードオフを示し、報告されたすべての結果を 1 つのコマンドで再現します。目標は、エージェントが何をすべきかだけでなく、利用可能な証拠がいつそれを安全に実行するのに十分であるかを決定するための具体的なアプローチを提供することです。
原文 (English)
SafeCommit: Certifying When Memory-Grounded Agents May Safely Act
Long-horizon agents increasingly use persistent memory and tools to take actions with external side effects. A central failure mode is premature commitment: an agent acts before resolving whether its memory grounding is stale, conflicting, incomplete, or corrupted. We formalize this problem as safe commitment under memory uncertainty and introduce SafeCommit, a risk controlled layer between agent reasoning and external execution. The layer constructs a calibrated set of plausible latent worlds from memory, observations, tool outputs, provenance, and policy constraints. It permits a side effectful action only when a conformal action certificate shows that the action is safe in every retained world. Otherwise, it selects a low-side-effect probe that targets the worlds blocking certification, or returns a conservative fallback. Under calibrated world coverage, the probability of an unsafe certified commit is at most the target level {\alpha}; with imperfect world proposal, the bound separates calibration and representation error. A dependency-free controlled simulator illustrates the safety-utility tradeoff and reproduces all reported results with one command. The goal is to offer a concrete approach for deciding not only what an agent should do, but when the available evidence is sufficient to safely do it.
NeuMoSync: 継続学習における可塑性と適応性のためのエンドツーエンドの神経調節制御
継続学習 (CL) では、モデルがタスクを順番に学習する必要がありますが、ディープ ニューラル ネットワークでは、可塑性の損失や知識伝達の低下が発生することが多く、長期的な適応性が妨げられる可能性があります。脳内の全体的な神経調節メカニズムから高度なインスピレーションを得て、動的でニューロン固有の変調をディープ ニューラル ネットワークに統合してネットワークの適応性と可塑性を強化する新しいアーキテクチャである Neuromodulation and Synchronization (NeuMoSync) を紹介します。 NeuMoSync は、ネットワーク全体の履歴コンテキストを追跡する各ニューロンの学習可能な特徴ベクトルと、より高い抽象レベルで動作するモジュールを備えた標準ニューラル ネットワーク アーキテクチャを拡張します。このモジュールは、現在の入力とネットワークの進化状態の両方を条件としたニューロン固有の信号を合成し、活性化ダイナミクスとシナプス可塑性を適応的に制御します。記憶 (ランダム ラベル CIFAR-10 およびランダム ラベル MNIST)、コンセプト ドリフト (シャッフル CIFAR-10 およびシャッフル Mini-ImageNet)、クラス増分学習 (クラス分割 ImageNet およびクラス分割 CIFAR-100)、ドメイン増分学習 (Permuted MNIST) を含むさまざまな CL ベンチマークで評価された NeuMoSync は、可塑性を保持する優れたパフォーマンスを実証し、前方および前方の両方の改善を達成します。既存の方法と比較した後方適応。アブレーション研究では各コンポーネントの必要性が検証され、学習された変調信号の分析によりタスク全体にわたる解釈可能な調整パターンが明らかになります。私たちの研究は、グローバルな調整メカニズムを深層学習システムに統合して、堅牢で適応的な継続学習を進める可能性を強調しています。コードは https://github.com/RoozbehRazavi/NeuMoSync で公開されています。
原文 (English)
NeuMoSync: End-to-End Neuromodulatory Control for Plasticity and Adaptability in Continual Learning
Continual learning (CL) requires models to learn tasks sequentially, yet deep neural networks often suffer from plasticity loss and poor knowledge transfer, which can impede their long-term adaptability. Drawing high-level inspiration from global neuromodulatory mechanisms in the brain, we introduce Neuromodulation and Synchronization (NeuMoSync), a novel architecture that integrates dynamic, neuron-specific modulation into deep neural networks to enhance their adaptability and plasticity. NeuMoSync extends standard neural network architectures with learnable feature vectors for each neuron that track network-wide historical context and with a module operating at a higher level of abstraction. This module synthesizes neuron-specific signals, conditioned on both current inputs and the network's evolving state, to adaptively regulate activation dynamics and synaptic plasticity. Evaluated on diverse CL benchmarks, including memorization (Random Label CIFAR-10 and Random Label MNIST), concept drift (Shuffle CIFAR-10 and Shuffle Mini-ImageNet), class-incremental learning (Class Split ImageNet and Class Split CIFAR-100), and domain-incremental learning (Permuted MNIST), NeuMoSync demonstrates strong performance in retaining plasticity and achieves improvements in both forward and backward adaptation compared with existing methods. Ablation studies validate the necessity of each component, while analysis of the learned modulatory signals reveals interpretable coordination patterns across tasks. Our work underscores the potential of integrating global coordination mechanisms into deep learning systems to advance robust, adaptive continual learning. The code is publicly available at https://github.com/RoozbehRazavi/NeuMoSync.
ドメイン固有言語によるニューラル PDE ソルバーの自動設計の改善
ニューラル PDE ソルバーの自動設計は、基本的に検索空間表現の問題です。制限のない Python プログラムの空間では、有効なソルバーは非常にまばらなサブセットを形成します。ほとんどの候補プログラムは、構文的に間違っているか、意味的に互換性がない、または数値的に不安定です。したがって、直接コード生成では、LLM はソルバーの品質について推論するのではなく、実装の失敗を回避するために検索能力のほとんどを費やすことになります。 ADSL-PDE は、ソルバーの概念と実行可能コードの間に構造化された検索状態を導入することで、この課題に対処します。これは、低レベルの実装の詳細を抽象化しながら、ニューラル PDE ソルバー (アーキテクチャ、物理的制約、目的、サンプリング、最適化) を決定する機能的な決定を表します。決定論的コンパイラは、有効な検索状態をそれぞれ実行可能なソルバーにマップします。実際、ADSL-PDE は検索空間を再構成します。つまり、無効なプログラムの大部分が削除され、意味のある候補の密度が増加し、これまでに見たことのない設計を発見するために必要な構成の自由度が維持されます。したがって、ソルバーの進化は、コード成果物ではなく、設計上の決定に基づいて機能します。この表現に基づいて構築された進化エージェントは、経験的フィードバックを使用してソルバー検索状態を繰り返し提案、評価、および改良します。複数の PDE ベンチマーク全体で、ADSL-PDE は検索効率と最適化の安定性の両方を向上させ、最初の 10 回の進化反復内で 52% 以上の改善を達成しました。これらの結果は、LLM 駆動の自動設計のより広範な原則を示唆しています。効果的なエージェントには、単により強力な推論が必要ではなく、有効で結果的な決定の探索に集中する検索表現が必要です。
原文 (English)
Improving Auto-Design of Neural PDE Solvers with a Domain-Specific Language
Neural PDE solver auto-design is fundamentally a search-space representation problem. In the space of unrestricted Python programs, valid solvers form an extremely sparse subset: most candidate programs are syntactically incorrect, semantically incompatible, or numerically unstable. Direct code generation therefore forces an LLM to spend most of its search capacity navigating implementation failures rather than reasoning about solver quality. ADSL-PDE addresses this challenge by introducing a structured search state between solver concepts and executable code. It represents the functional decisions that determine a neural PDE solver (architecture, physical constraints, objectives, sampling, and optimization) while abstracting away low-level implementation details. A deterministic compiler maps each valid search state to an executable solver. In effect, ADSL-PDE reshapes the search space: it removes large regions of invalid programs, increases the density of meaningful candidates, and preserves the compositional freedom needed to discover previously unseen designs. Solver evolution can thus operate over design decisions rather than code artifacts. Built on this representation, our evolutionary agent iteratively proposes, evaluates, and refines solver search states using empirical feedback. Across multiple PDE benchmarks, ADSL-PDE improves both search efficiency and optimization stability, achieving an improvement of more than 52% within the first ten evolution iterations. These results suggest a broader principle for LLM-driven auto-design: effective agents do not merely require stronger reasoning, but rather a search representation that concentrates exploration on valid and consequential decisions.
Agentic AI ワークフローのアーキテクチャ上の影響
エージェント型 AI はデータセンターに出現しつつありますが、そのアーキテクチャ上の影響はまだ解明されていません。私たちはエージェント ワークフローを分類法で整理し、Microsoft Azure での実稼働調査とオープンソース フレームワークの管理された調査によってその最初のアーキテクチャの特徴を示します。エージェントの実行が断片化されており、不均一であることを示します。リクエストは、CPU と GPU の境界を繰り返し越える LLM 推論、ツール呼び出し、オーケストレーション決定のワークフローに拡張されます。私たちの分類は、この断片化がどのようにリソース需要に変わるかを説明しています。オーケストレーションとツールがホスト上で実行されると、CPU はクリティカル パス上に位置します。実行構造により時間の経過とともに負荷が設定され、突然のスパイクがあっても負荷は低く抑えられます。モデルの構成は、ワークフローが GPU をどの程度均等に使用するかを設定します。タスクとツールの多様性により、この範囲はさらに広がります。これらの特性により、従来の均一サーバーのアーキテクチャ上の不一致が明らかになります。断片化された実行は、爆発的な需要にもかかわらず CPU と GPU の能力を圧迫します。ソフトウェアの役割が異なると、同種の CPU プロビジョニングが非効率になります。最後に、多くのエージェントを共有コアに多重化すると、マイクロアーキテクチャの局所性が低下します。私たちは調査結果に基づいて、エージェント サーバーへの影響を導き出し、コモディティ サーバーのプロトタイプである Agora を通じてそれらを検証します。 Agora は、ツールのスパイクからエージェントのテール レイテンシーを保護しながら、同じ場所に配置されたスループット作業のためにアイドル状態の CPU コアを動的に収集します。各 GPU にさらに多くのエージェントを配置することで GPU メモリをオーバーサブスクライブし、次のエージェントの状態をプリフェッチしてスワップ レイテンシーを隠します。マシンを異種ロールに一致させるために、Agora はロールごとにコアをプールし、アフィニティを認識したスケジューリングを適用して局所性を復元します。ワークロードに合わせてメカニズムを自動的に調整します。 Agora は、エージェントのテール レイテンシを維持しながら、使用率とサーバー スループットを向上させます。私たちの洞察は、エージェント AI の将来のサーバー アーキテクチャの重要な方向性も特定します。
原文 (English)
Architectural Implications of Agentic AI Workflows
Agentic AI is emerging in datacenters, but its architectural implications remain unexplored. We organize agentic workflows in a taxonomy and present its first architectural characterization with a production study at Microsoft Azure and a controlled study of open-source frameworks. We show that agentic execution is fragmented and heterogeneous. Requests expand into a workflow of LLM inferences, tool invocations, and orchestration decisions that repeatedly cross the CPU-GPU boundary. Our taxonomy explains how this fragmentation turns into resource demand. As orchestration and tools run on the host, the CPU sits on the critical path. Execution structure sets the load over time, which stays low with sudden spikes. Model composition sets how evenly the workflow uses the GPUs. Diversity in tasks and tools widens this range even further. These characteristics expose architectural mismatches of conventional uniform servers. Fragmented execution strands CPU and GPU capacity despite bursty demand. Different software roles make homogeneous CPU provisioning inefficient. Finally, multiplexing many agents onto shared cores degrades microarchitectural locality. Guided by our findings, we derive implications for agentic servers and examine them through Agora, our prototype for commodity servers. Agora dynamically harvests idle CPU cores for co-located throughput work, while protecting agentic tail latency against tool spikes. It oversubscribes GPU memory by placing more agents on each GPU, prefetching the next agent's state to hide swap latency. To match the machine to the heterogeneous roles, Agora pools cores by role and applies affinity-aware scheduling to restore locality. It automatically tunes mechanisms to the workload. Agora improves utilization and server throughput while preserving agent tail latency. Our insights also identify key directions for future server architectures for agentic AI.
CARGO-VL: 視覚言語モデルのリスク制限付きグループ最適化による反事実仲裁
視覚言語システムは画像と取得したテキストを組み合わせますが、これらの情報源が一致していない場合や、答えをサポートできない場合があります。信頼できるモデルは、信頼できるソースを特定し、どちらも適切でない場合は回避する必要があります。既存のトレーニング後の目標はインスタンスを独立してスコアリングするため、反事実の証拠が変更された場合に一貫した動作を強制しません。 CARGO-VL は、位置合わせされた、画像が正しい、テキストが正しい、および両方が間違っている (A/V/T/N) 証拠状態を 1 つのバンドルとしてカバーする、一致するバリアントを最適化するグループ相対フレームワークです。その目的は、条件ごとの正確性と、回答の不変性、ソースの等価性、および回答から棄権への切り替えに対する移行報酬を結び付けるとともに、主二重コントローラーが安全でない回答と過剰な延期のバランスをとります。また、4 つの条件の競合トレーニング リソースである XMC (eXtended Modal Conflict) にも貢献し、CMC-Bench と Modality-Bias で移行を評価します。 CARGO-VL は、複数のシードにわたって、競合処理、サポートされていない回答の回避、点ごとのベースライン上のモダリティ バランスを改善します。アブレーションは、関係遷移シグナルと適応型リスク制御からの補完的な利点を特定し、信頼できるマルチモーダル証拠仲裁のための実際的な目標として反事実の一貫性をサポートします。
原文 (English)
CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models
Vision-language systems combine images with retrieved text, but these sources can disagree or jointly fail to support an answer. Reliable models must identify the trustworthy source and abstain when neither is adequate. Existing post-training objectives score instances independently and therefore do not enforce coherent behavior under counterfactual evidence changes. We introduce CARGO-VL, a group-relative framework that optimizes matched variants covering aligned, image-correct, text-correct, and both-wrong (A/V/T/N) evidence states as one bundle. Its objective couples condition-wise correctness with transition rewards for answer invariance, source equivariance, and answer-to-abstention switching, while a primal-dual controller balances unsafe answers against excessive deferral. We also contribute XMC (eXtended Modal Conflict), a four-condition conflict training resource, and evaluate transfer on CMC-Bench and Modality-Bias. Across multiple seeds, CARGO-VL improves conflict handling, unsupported-answer avoidance, and modality balance over pointwise baselines. Ablations identify complementary benefits from relational transition signals and adaptive risk control, supporting counterfactual consistency as a practical objective for reliable multimodal evidence arbitration.
リーク耐性のあるアンラーニング: マルチホップ推論の一貫性と回復の堅牢性を評価するための新しいベンチマーク
機械の非学習手法のベンチマークは、大規模言語モデル (LLM) から機密知識が削除されているかどうかを理解するために重要です。現在のアンラーニング ベンチマークには、主にシングルホップの質問と、一部のマルチホップの質問が含まれています。効果的ではありますが、依然として 2 つの課題に直面しています。 (1) 知識は分離されていないため、多様なマルチホップ推論パスは通常のクエリよりも知識の漏洩を引き起こす可能性があります。 (2) 未学習は脆弱である可能性があります。未学習の知識は、軽量の非学習後適応などの回復攻撃によって部分的に回復できるため、静的評価が不十分になります。したがって、このペーパーでは、多様な推論パスと回復攻撃にわたる堅牢な LLM 知識の削除を理解するための新しいベンチマークとして \unlearning を紹介します。このベンチマークを 3 つのモデル、6 つの非学習手法、および慎重に厳選された 2 つのデータセットで実験します。結果は、既存の方法がマルチホップ推論パスと回復攻撃に対して脆弱であることを示しています。さらに、LLM の学習解除における品質、堅牢性、モデルの有用性の間のトレードオフを調査します。
原文 (English)
Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness
Benchmarking machine unlearning methods is critical to understand whether sensitive knowledge is removed from large language models (LLMs) or not. Current unlearning benchmarks include mainly single-hop questions and a narrow set of multi-hop questions. Although effective, they still face two challenges. (1) Knowledge is not isolated, whereby diverse multi-hop reasoning paths can potentially induce knowledge leakage than normal queries. (2) Unlearning may be fragile: unlearned knowledge can be partially recovered through recovery attacks such as lightweight post-unlearning adaptation, making static evaluation insufficient. Therefore, in this paper, we introduce \unlearning as a novel benchmark to understand robust LLM knowledge removal across diverse reasoning paths and recovery attacks. We experiment with this benchmark on 3 models, 6 unlearning methods, and 2 carefully curated datasets. Results show that existing methods are vulnerable to multi-hop reasoning paths and recovery attacks. We further explore the trade-off among forget quality, robustness, and model utility for LLM unlearning.
What Is a Skill Worth? Structure-Aware Shapley Valuation of Agent Skills
Agent skills are increasingly optimized by automated feedback loops, producing long structured artifacts whose internal value remains uncle…
Joint UAV Flight and Opportunistic Routing under Reinforcement Learning for Delay-Tolerant Networks
The growing deployment of delay-tolerant networks (DTNs) has made store-carry-forward (SCF) communication indispensable under sparse connec…
多様性の前の合意: 異種言語モデルの調整のための検証優先の補完性
異種言語モデルのアンサンブルは、回答候補の空間を拡張しますが、新しく生成された回答がすでにサポートされている回答をいつ置き換えるべきかについての原則的な基準を欠いています。私たちは候補者の余裕を後任の権限から切り離し、後者を明示的で監査可能なオブジェクトとしてレンダリングします。私たちが提案する手法であるAgreement-Before-Diversity(ABD)は、凍結されたラベルのない決定ルールです。固定された等価関係の下で追加の信頼できるサンプルが2つあることを裏付ける場合、アンカーの回答は保持されます。それ以外の場合は、ヘテロジニアス合成に置き換えられます。このゲート メカニズムについて、2 つの正確な同一性を証明します。 1 つ目は、無条件合成に対する精度ギャップが、保護されたサブセットに対する合意カバレッジとアンカーの利点によって共同して決定されることを示しています。 2 つ目は、決して合成しない場合とのギャップが、許可された回復と許可された破壊の対比を反映していることを示しています。どちらの ID も独立性や調整された信頼性を前提としていないため、予想される推論コストは、呼び出し回数のカバレッジの約 8 から 5 倍を引いたものになります。ブラインドでの正確な ID 評価では、ABD は完全な LiveCodeBench-v6 で 59.43% (対 Single9 では 52.57%、HAC では 52.00%、n = 175)、手つかずの GPQA-Diamond スプリットでは 75.00% (両方のコントロールで 72.78%、n = 180) を達成しました。さらに、これらのアイデンティティは、集計されたすべての差異を数え切れないほどの保護された層に局所化します。LiveCodeBench 上の 3 つの保護されたケースの間に不一致項目は発生しません。カバレッジは、ゲートの寄与を事前に 1.71 ポイントに制限します。 GPQA-Diamond の 132 件中、一致しないケースは 13 件対 8 件。凍結アンカー摂動下では 71 隻中 12 対 0 でした。多様性は可能性をもたらします。検証構造が権限を提供します。
原文 (English)
Agreement Before Diversity: Verification-First Complementarity for Heterogeneous Language-Model Coordination
Heterogeneous language-model ensembles expand the space of candidate responses, yet they lack a principled criterion for when a newly generated answer should supersede an already supported one. We decouple candidate headroom from replacement authority, rendering the latter as an explicit, auditable object. Our proposed method, Agreement-Before-Diversity (ABD), is a frozen, label-free decision rule: an anchor answer is retained if two additional trusted samples corroborate it under a fixed equivalence relation; otherwise, it is replaced by a heterogeneous synthesis. For this gating mechanism, we prove two exact identities. The first shows that the accuracy gap relative to unconditional synthesis is determined jointly by the agreement coverage and the anchor's advantage on the protected subset. The second shows that the gap relative to never synthesizing reflects a contrast between authorized recovery and authorized destruction. Neither identity assumes independence or calibrated confidence, and the expected inference cost is approximately eight minus five times the coverage in number of calls. Under blind, exact-ID evaluation, ABD achieves 59.43% on the complete LiveCodeBench-v6 (vs. 52.57% for Single9 and 52.00% for HAC; n = 175) and 75.00% on an untouched GPQA-Diamond split (both controls at 72.78%; n = 180). Furthermore, these identities localize every aggregate difference to an enumerable protected stratum: no discordant items occur among the 3 protected cases on LiveCodeBench, where coverage bounds the gate's contribution to 1.71 points a priori; 13 versus 8 discordant cases among 132 on GPQA-Diamond; and 12 versus 0 among 71 under a frozen anchor perturbation. Diversity supplies potential; verification structure supplies authority.
A/B Agent: A Self-Evolving Agent for Strategy Iteration in Industrial A/B Testing
Industrial recommendation strategy iteration heavily relies on large-scale A/B experimentation. Traditional tuning requires experts to repe…
AI Literacy for Legal Translation: Developing Digital Resilience
Generative AI is transforming legal translation by introducing opportunities alongside linguistic, technical, legal, ethical and cognitive…
Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning
Cooperative multi-agent reinforcement learning often adds social terms to individual rewards, yet the scale of those terms is usually chose…
ASRS レポートを使用した航空システムの運用安全分析のための、追跡可能な LLM 生成の危険シナリオ
航空システム運用の運用上のハザード分析では、航空機システム レベルで適用される機能上のハザード評価とは異なり、気象、ATC の動作、空域の制約、航空機の運用、および人的要因の間の相互作用を考慮する必要があります。 NASA の航空安全報告システム (ASRS) から危険シナリオの候補を生成する、AI 支援のアプローチを紹介します。目標とする有害な結果が与えられると、カテゴリ要因として構造化された仮説と、その構造と一致する一連の操作イベントを説明する物語シナリオが生成されます。各シナリオには、過去の共起証拠からの妥当性スコアと、最も類似した保持された ASRS レポートへの追跡可能性が含まれます。次に、進化的アブダクションによって生成された構造化された仮説に基づいて物語の生成を条件付けし、正確性を向上させ、ばらつきを減らすハイブリッド バリアントを提案します。複数の大規模な言語モデル、ゼロショットと少数ショットのプロンプト、およびオプションの微調整を評価し、プロンプトとモデルの選択が、生成された構造と物語の妥当性と現実性にどのように影響するかを測定します。
原文 (English)
Traceable LLM-Generated Hazard Scenarios for Operational Safety Analysis of Aviation Systems Using ASRS Reports
Operational hazard analysis of aviation system operations must consider interactions among weather, ATC actions, airspace constraints, aircraft operations, and human factors - distinct from the functional hazard assessment applied at the aircraft-system level. We present an AI-assisted approach that generates candidate hazard scenarios from NASA's Aviation Safety Reporting System (ASRS). Given a target adverse outcome, it produces a structured hypothesis as categorical factors and a narrative scenario describing an operational event sequence consistent with the structure. Each scenario includes by a plausibility score from historical co-occurrence evidence and traceability to the most similar held-out ASRS reports. We then propose a hybrid variant, conditioning narrative generation on a structured hypothesis produced via evolutionary abduction, improving correctness and reducing variability. We evaluate multiple large language models, zero-shot versus few-shot prompting, and optional fine-tuning, measuring how prompting and model choice affect the validity and realism of the generated structures and narratives.
Canary ツールを使用した LLM エージェントでのツール選択推論の診断
エージェントの評価では、モデルが間違ったツールを選択したことがわかりますが、その理由はほとんどわかりません。カナリア ツールを紹介します。これは、エージェントのモデル コンテキスト プロトコル (MCP) ツール セットに組み込まれた診断プローブ ツールであり、それぞれが 1 つの特定のツール選択の弱点を調査するように設計されています。 6 つのタイプの分類法 (セマンティック デコイ、パラメーター トラップ、ケイパビリティ ミラージュ、前提条件ブラインドネス、時間的デコイ、および粒度トラップ) は、単一の「間違ったツール」の結果を、モデルがツールについてどのように推論するかについての多次元プロファイルに変換します。私たちは、3 つの機能層にまたがる 8 つのモデル (6 つのホスト型と 2 つの 8B オープンウェイト) を、3 つのカナリア密度条件と 3 つのシード (8,640 実行) にわたる 120 のタスク、および 2,880 実行のサトルティ アブレーションで評価しました。タスクの成功は、プロバイダーに依存しない審査員によって採点され、2 番目の独立した審査員によって裏付けられます (コーエンのカッパ = 0.75)。 3 つの調査結果を報告します。まず、モデルの能力が向上するにつれて、感受性は急激に低下します。タスクごとのカナリア感受性率 (CSR) は、モデル全体で約 36 倍の範囲にあり、Claude Opus 4.8 が最も低く、Llama 3.1 8B が最も高くなります。第 2 に、機能層だけでは安全性を予測できません。最も影響を受けやすいホスト型モデルは中間層であり、プロバイダー内では安価なモデルの方が安全である可能性があります。第三に、分類は能力によって階層化されています。能力ミラージュは最も確実にフロンティア モデルをトラップしますが、他のタイプは強力なモデルではほとんど不活性ですが、小規模なオープン モデルでは発火するため、弱いというよりも能力によって区別されます。各カナリアのギブアウェイフレーズを和らげると、フロンティアCSRは本質的に変化せず、プローブがフレーズ発見ではなく推論を測定する証拠です。感受性はタスクの失敗も予測しますが (スピアマン rho = -0.34)、最も堅牢なモデルはカナリア圧力によって大幅に劣化しません。フレームワーク、カナリア スキーマ、タスク、ログをリリースします。
原文 (English)
Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools
Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) tool set, each engineered to probe one specific tool-selection weakness. A six-type taxonomy (semantic decoys, parameter traps, capability mirages, prerequisite blindness, temporal decoys, and granularity traps) turns a single "wrong tool" outcome into a multi-dimensional profile of how a model reasons about tools. We evaluate eight models -- six hosted and two 8B open-weight -- spanning three capability tiers, on 120 tasks across three canary-density conditions and three seeds (8,640 runs), plus a 2,880-run subtlety ablation. Task success is graded by a provider-independent judge, corroborated by a second independent judge (Cohen's kappa = 0.75). We report three findings. First, susceptibility drops sharply as models get more capable: the per-task canary susceptibility rate (CSR) ranges about 36x across models, lowest for Claude Opus 4.8 and highest for Llama 3.1 8B. Second, capability tier alone does not predict safety: the most susceptible hosted model is mid-tier, and within a provider the cheaper model can be the safer one. Third, the taxonomy is capability-stratified: capability mirages most reliably trap frontier models, while the other types are largely inert on strong models but fire on small open models, so they discriminate by capability rather than being weak. Softening each canary's give-away phrase leaves frontier CSR essentially unchanged, evidence that the probes measure reasoning, not phrase-spotting. Susceptibility also predicts task failure (Spearman rho = -0.34), while the most robust models are not significantly degraded by canary pressure. We release the framework, canary schemas, tasks, and logs.
When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning
Multimodal large language models increasingly reason over screenshots and documents where the task itself may be written in pixels. Yet ben…
暗黙的影響設定では思考連鎖の監視が信頼できない可能性がある
思考連鎖 (CoT) モニタリングは、フロンティア推論モデルの重要な安全層として扱われることが増えています。ほとんどの監視可能性評価では、明示的な影響設定、つまりプロンプトがモデルに何かを隠すように直接インセンティブを与える設定 (たとえば、隠されたサイドタスクを実行するように指示する) が研究されます。 CoT モニター評価の補足軸は、暗黙的影響設定です。プロンプトには非表示にする指示が含まれていませんが、モデルの動作は依然としてタスクまたはコンテキストの特徴によって形成されます。採用評価に偏りをもたらす候補者に関する無関係な詳細。 2 つの体制下で CoT の監視可能性を直接比較する最初のベンチマークを紹介します。特定のオプションを選択するためのナッジが存在する場合にモデルの動作がどのように変化するかをテストします。ナッジは、何気ない余談として(暗黙的)、またはナッジに基づいて行動し、そのことを隠蔽するという直接の指示として(明示的に)配信されます。このベンチマークは、4 つのタスク形式 (二者択一、数値評価、多肢選択 QA、自由回答型コーディング) と 7 つのフロンティア拡張思考モデルに及びます。明示的な影響下では、CoT モニターは動作の変化の 60 ~ 94% を検出します。それを隠すように指示されたモデルであっても、その命令が CoT に漏洩します。暗黙的な影響下では、同じ要因によって動作が変化しますが、4 つの設定のうち 2 つでは検出率が 41 ~ 46 パーセント ポイント低下します。現実的なシステムプロンプトの追加 (トピックから外れたバイアスを減らすために開発者が導入するようなもの) は、動作への影響自体を維持しながら、暗黙的な検出をさらに 5% まで低下させます。これらの結果は、明示的な影響設定で得られた監視可能性の推定値が監視可能性を過大評価している可能性があり、善意の展開の選択によって監視可能性がさらに低下する可能性があることを示唆しています。私たちのベンチマークとコードは https://github.com/agatha-duzan/implicit-vs-explicit-influence から入手できます。
原文 (English)
Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings
Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability evaluations study explicit-influence settings: setups where the prompt directly incentivizes the model to hide something, e.g., by instructing it to perform a hidden side-task. A complementary axis for CoT-monitor evaluations is implicit-influence settings, where the prompt contains no instruction to hide, but the model's behavior is still shaped by features of the task or context, e.g. an irrelevant detail about a candidate that biases a hiring rating. We introduce the first benchmark that directly compares CoT monitorability under the two regimes. We test how model behavior changes in the presence of a nudge to choose a particular option. The nudge is delivered either as a casual aside (implicit), or as a direct instruction to act on the nudge and to conceal having done so (explicit). The benchmark spans four task formats (binary choice, numeric rating, multiple-choice QA, open-ended coding) and seven frontier extended-thinking models. Under explicit influence, a CoT monitor detects 60-94% of behavior shifts: even models instructed to conceal it leak the instruction into their CoT. Under implicit influence, the same factors still shift behavior, but detection falls by 41-46 percentage points in two of our four settings. Realistic system-prompt additions (of the kind a developer might deploy to reduce off-topic bias) lower implicit detection further, to as low as 5%, while preserving the behavioral influence itself. These results suggest that monitorability estimates obtained in explicit-influence settings may over-estimate monitorability, and that monitorability can be further decreased by well-intentioned deployment choices. Our benchmark and code are available at https://github.com/agatha-duzan/implicit-vs-explicit-influence
EviGraph: Evidence-Guided Autonomous Research Agents
Autonomous research agents can generate hypotheses, execute experiments, and draft manuscripts, yet their outputs often contain unsupported…
Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning
Large Reasoning Models (LRMs) excel on complex tasks through long chain-of-thought (CoT) reasoning, but their lengthy intermediate steps ca…
NSF-HRPT: Neural Semantic Field meets Hierarchical Risk Perception Tree for Safety-Critical Scenario Assessment
The ability to accurately assess and anticipate risks in safety-critical scenarios is crucial for autonomous driving systems. While existin…
特権はあるが偏見がある: PI に条件付けされた教師が自己蒸留をどのように打ち破るか
自己蒸留 (SD) は、検証可能な報酬を伴う強化学習に代わる計算効率の高い方法として登場しました。自己教師は、参照解などの答えに関する特権情報 (PI) を条件として、トークンごとの緻密な監視を、それを見たことのない生徒に提供します。しかし、報告されている利益はほぼ専ら狭い低難易度の設定からのものであり、基本的な疑問が残されています。報酬条件のない唯一の目標として、SD は何かを教えているのでしょうか? SDPO で報告されているゲインを簡単な設定で再現し、同じセットアップを難しいタスクに適用したところ、そうではないことがわかりました。質問応答、数学、コーディング、マルチターン エージェント ツールの使用、推論モード、モデル サイズ、PI の形式全体にわたって、SDPO レシピと OPSD レシピの両方で、トークンごとの損失は着実に減少していますが、検証精度は改善されず、通常は低下しています。この失敗を、損失から生成されるモデルまでの単一の因果関係の連鎖を通じて説明します。この連鎖は PI バイアスから始まります。特定の 1 つの参照ソリューションを見た後、教師のトークンごとの目標は、一般的な正しさではなく、その軌道に引き寄せられます。この効果は、PI バイアス スコアで定量化されます。あらゆる場所でこの目標に一致するように訓練された学生の目標は、ロールアウトが正しいかどうかについてほとんど盲目になり、割り当てられる損失は、答えを決定するものではなく、ストップワード、句読点、不確実性マーカーなどの低情報トークンに主に当てられます。正しいロールアウト内では、探索的トークンは最も大きな発散を引き起こすため、推論に必要なためらいがペナルティとなります。その結果、推論能力が劣る、よりフラットで決断力の低い生徒が生まれます。SD は唯一の目的として、タスクの成功から切り離されたシグナルを最適化します。
原文 (English)
Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation
Self-distillation (SD) has emerged as a compute-efficient alternative to reinforcement learning with verifiable rewards: a self-teacher, conditioned on privileged information (PI) about the answer such as a reference solution, supplies dense per-token supervision to a student that never sees it. Reported gains, however, come almost exclusively from narrow, low-difficulty settings, leaving open a basic question: as a lone objective, with no reward term, does SD teach anything? We reproduce SDPO's reported gains in its easy setting, then apply the identical setup to difficult tasks and find that it does not. Across question answering, mathematics, coding, and multi-turn agentic tool use, across reasoning modes, model sizes, and forms of PI, and under both the SDPO and OPSD recipes, the per-token loss falls steadily while validation accuracy does not improve and typically degrades. We explain this failure through a single causal chain from the loss to the model it produces. The chain begins with PI bias: having seen one particular reference solution, the teacher's per-token target is pulled toward that trajectory rather than toward correctness in general, an effect we quantify with a PI Bias Score. Trained to match this target everywhere, the student's objective becomes nearly blind to whether a rollout is correct, and the loss it assigns falls mostly on low-information tokens like stopwords, punctuation, uncertainty markers, rather than those that determine the answer; within correct rollouts the exploratory tokens incur the highest divergence, so it penalizes the hesitation that reasoning requires. The result is a flatter, less decisive student that is no better at reasoning: as a lone objective, SD optimizes a signal decoupled from task success.
ContextWeave: A Real-World Workflow Benchmark
Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce…
共有ロールアウトが防御運転評価で失敗した場合: NAVSIM スコアベースの監査
防御運転スコアは、周囲のアクターを観察するポリシーとそうでないポリシーの区別を維持する場合にのみ役立ちます。再シミュレーション ベンチマークでは、ログに記録された人間のリファレンスがコンプライアンス チャネルを通過できなかった場合にエージェントがクレジットを受け取る、リファレンス条件付きの寛容性を使用する場合があります。エージェントと参照が不安定なロールアウト変換を共有する場合、このルールにより、共有参照の失敗が広範なコンプライアンス クレジットに伝播される可能性があります。 NAVSIM v2.2 のオリジナル シーンの単一ステージ スコアリングでこのリスクを監査します。監査された数値バックエンドの影響を受けた文書化されたスタック条件の下では、ルート ブラインド Ignore-All プローブとルート認識アクター ブラインド プローブは、完全な 12,146 トークンの navtest スプリットに対して人間のリプレイと PDM-Closed よりも上位にランクされます。公開仕様に従った新規インストールでは、固定の 32 トークン診断セットでロールアウトの相違が再現されます。同一ソース依存関係スタック制御と正確な入力診断により、共有速度再調整における依存関係に依存する数値動作が分離されます。 450 トークンのコントロール プールでは、ソルバーのみを置き換えることでロールアウトの発散がなくなり、許容範囲を有効にしたままブラインド ラスト順序が復元されます。したがって、数値の不安定性が直接のトリガーとなります。参照条件付きの寛容は、結果として生じる共有参照の失敗をコンプライアンスの信用に伝播します。当社は、スコアベースとスタックの開示、ブラインドプローブ、上書きレポート、防御運転主張にスコアを使用する前に展開する安定性テストを必要とする監査プロトコルに貢献します。
原文 (English)
When Shared Rollouts Fail in Defensive Driving Evaluation: A NAVSIM Score Basis Audit
Defensive driving scores are useful only when they preserve distinctions between policies that observe surrounding actors and those that do not. Re-simulation benchmarks may use reference-conditioned forgiveness, under which an agent receives credit when the logged human reference fails a compliance channel. When agent and reference share an unstable rollout transformation, this rule can propagate shared reference failures into broad compliance credit. We audit this risk in NAVSIM v2.2 original scene single-stage scoring. Under the affected documented-stack condition on the audited numerical backend, the route-blind Ignore-All probe and a route-aware actor-blind probe outrank human replay and PDM-Closed over the complete 12,146-token navtest split. A fresh installation following the public specification reproduces rollout divergence on a fixed 32-token diagnostic set. A same-source dependency stack control and an exact-input diagnostic isolate dependency-sensitive numerical behavior in the shared velocity refit. On a 450-token control pool, replacing only the solver eliminates rollout divergence and restores blind-last ordering while keeping forgiveness enabled. Thus, the numerical instability is the direct trigger. Reference-conditioned forgiveness propagates the resulting shared reference failures into compliance credit. We contribute an audit protocol requiring score basis and stack disclosure, blind probes, overwrite reporting, and rollout stability tests before using such scores for defensive driving claims.
WorldCycle: 長期的なビデオワールドモデルのための自己検証可能な強化学習
インタラクティブなビデオ世界モデルは長期的な計画と探索に不可欠ですが、複合的なエラーに悩まされます。強化学習 (RL) などのトレーニング後の手法はこれらのモデルを改善できますが、検証のボトルネックにぶつかります。任意のアクション シーケンスの場合、長期ドリフトを測定するためのグランドトゥルースの将来状態は存在しません。私たちの重要な洞察は、可逆的なアクションサイクルがこの検証を可能にするということです。つまり、その逆で構成されたシーケンスは分析的に初期状態に戻り、長期的な正確性について注釈なしで監視できる必要があります。これに基づいて、通常のアクション シーケンスから閉じたアクション サイクルとその繰り返し実行を構築し、2 つの相補的な報酬を最適化する自己検証可能な RL フレームワークである WorldCycle を導入します。ミラーリングされた順方向セグメントと逆方向セグメント間の対称性を強制する空間的閉鎖報酬と、繰り返されるサイクル実行全体で状態を調整する時間的一貫性報酬です。これらの報酬により、モデルは、記憶された一時的なパターンではなく、一貫した状態の演算子としてアクションを学習することが強制され、基本モデルがうまく処理できない分布外の複合アクション サイクルにも自然に拡張されます。さらに、複雑なアクション構造の下で状態を返す能力の診断ベンチマークである CycleBench をリリースします。 WorldCycle は、状態を返すドリフトを最大 44% 削減し、複合アクションの精度をベース モデルの 4 倍近く向上させ、物理的に接地されたワールド モデルに重要な基盤を提供します。
原文 (English)
WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models
Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors. Post-training methods such as reinforcement learning (RL) can improve these models, but they hit a verification bottleneck: for arbitrary action sequences, no ground-truth future state exists to measure long-term drift. Our key insight is that reversible action cycles make this verification possible: a sequence composed with its inverse must analytically return to the initial state, yielding annotation-free supervision on long-horizon correctness. Building on this, we introduce WorldCycle, a self-verifiable RL framework that constructs closed action cycles and their repeated executions from ordinary action sequences, and optimizes two complementary rewards: a spatial closure reward enforcing symmetry between mirrored forward and reverse segments, and a temporal consistency reward aligning states across repeated cycle executions. These rewards force the model to learn actions as consistent state operators rather than memorized temporal patterns, and extend naturally to out-of-distribution composite action cycles that the base model handles poorly. We further release CycleBench, a diagnostic benchmark for state-returning ability under complex action structures. WorldCycle reduces state returning drift by up to 44% and lifts composite-action accuracy nearly 4x over the base model, providing a vital foundation for physically grounded world models.
Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load
Short-term load forecasting (STLF) play a vital role in the electric power industry. It serves infrastructure that European and German law…
From Score Matrices to Football-Aware Match-State Simulation: An Auditable LLM Harness for Exact-Score Reranking
Football score forecasting combines a strong statistical core with a difficult contextual edge. Dynamic Poisson-family models estimate team…
Item Response Theory for AI Safety
Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores a…
Hierarchical Graph Memory for LLM Agents with Path-level Localization and Rewrite
Agents for long term reasoning require a memory that can be efficiently and effectively updated over time, as new facts and external feedba…
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a fin…
CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument Graphs
AI-supported care planning can help clinicians, patients, caregivers, and care teams coordinate complex decisions across clinical, function…
OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling
Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and…
Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning
Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements…
AutoProteinEngine: A Large Language Model Driven Agent Framework for Multimodal AutoML in Protein Engineering
Protein engineering is important for biomedical applications, but conventional approaches are often inefficient and resource-intensive. Whi…
TourSynbio-Search: A Large Language Model Driven Agent Framework for Unified Search Method for Protein Engineering
The exponential growth in protein-related databases and scientific literature, combined with increasing demands for efficient biological in…
時間的コンテキスト認識: 大規模な言語モデルに対するマルチターン操作攻撃に対する防御フレームワーク
大規模言語モデル (LLM) は、高度なマルチターン操作攻撃に対してますます脆弱になっています。攻撃者は、一見無害な会話ターンを通じて戦略的にコンテキストを構築し、安全対策を回避し、有害な応答や不正な応答を引き出します。これらの攻撃は、対話の時間的性質を利用してシングルターン検出方法を回避し、現実世界の展開に重大な影響を与える重大なセキュリティ脆弱性を表します。この論文では、セマンティック ドリフト、クロスターンの意図の一貫性、進化する会話パターンを継続的に分析することで、この課題に対処するように設計された新しい防御メカニズムである、時間的コンテキスト認識 (TCA) フレームワークを紹介します。 TCA フレームワークは、動的コンテキスト埋め込み分析、クロスターン一貫性検証、およびプログレッシブ リスク スコアリングを統合して、操作の試みを効果的に検出して軽減します。シミュレートされた敵対的シナリオの予備評価は、従来の検出技術では見逃されがちな微妙な操作パターンを識別するフレームワークの可能性を実証し、会話型 AI システムに切望されているセキュリティ層を提供します。 TCA の設計の概要を説明することに加えて、私たちはさまざまな攻撃ベクトルと複数ターンの会話にわたるその進行を分析し、敵対的な戦術とその LLM 脆弱性への影響についての貴重な洞察を提供します。私たちの調査結果は、会話型 AI システムにおける堅牢なコンテキスト認識型防御の緊急の必要性を強調し、正当なアプリケーションでの有用性を維持しながら LLM を保護するための有望な方向性として TCA フレームワークを強調しています。私たちは、この AI セキュリティの新興分野におけるさらなる研究をサポートするために、私たちの実装を利用できるようにしています。
原文 (English)
Temporal Context Awareness: A Defense Framework Against Multi-turn Manipulation Attacks on Large Language Models
Large Language Models (LLMs) are increasingly vulnerable to sophisticated multi-turn manipulation attacks, where adversaries strategically build context through seemingly benign conversational turns to circumvent safety measures and elicit harmful or unauthorized responses. These attacks exploit the temporal nature of dialogue to evade single-turn detection methods, representing a critical security vulnerability with significant implications for real-world deployments. This paper introduces the Temporal Context Awareness (TCA) framework, a novel defense mechanism designed to address this challenge by continuously analyzing semantic drift, cross-turn intention consistency and evolving conversational patterns. The TCA framework integrates dynamic context embedding analysis, cross-turn consistency verification, and progressive risk scoring to detect and mitigate manipulation attempts effectively. Preliminary evaluations on simulated adversarial scenarios demonstrate the framework's potential to identify subtle manipulation patterns often missed by traditional detection techniques, offering a much-needed layer of security for conversational AI systems. In addition to outlining the design of TCA , we analyze diverse attack vectors and their progression across multi-turn conversation, providing valuable insights into adversarial tactics and their impact on LLM vulnerabilities. Our findings underscore the pressing need for robust, context-aware defenses in conversational AI systems and highlight TCA framework as a promising direction for securing LLMs while preserving their utility in legitimate applications. We make our implementation available to support further research in this emerging area of AI security.
Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't
Model families train every size from scratch. Can a pretrained large model be converted into a smaller sibling? We characterize the 1.4B->4…
RAG-Stack: Co-Optimizing RAG Serving Performance and Quality
Retrieval-augmented generation (RAG), which augments large language model (LLM) generation with information retrieved from databases, has b…
Towards a New Grammar of Reasoning for Artificial Legal Intelligence and the Mecelle as Its Semantic Protocol
This article examines the enduring epistemic and methodological crisis of traditional legal practice in light of the opportunities and cons…
C$^2$MOE: Consistency and Complementarity-guided Mixture of Experts for Incomplete Multimodal Emotion Learning
Recent advances in Multimodal Emotion Recognition in Conversations (MERC) highlight its reliance on complete multimodal inputs. However, re…
On Hamming-Lipschitz Type Stability of the Subdominant (Minmax) Ultrametric: Theory and Simple Proofs
The subdominant (minmax) ultrametric is a canonical tree-structured summary of a dissimilarity matrix, arising equivalently as the ultramet…
AI による潜在的媒介構造のマルチモーダル表現学習社会経済的不利益、心理社会的要因、心臓代謝性多疾患の発見: All of Us 研究プログラムからの洞察
社会的不利な状況は多発性疾患と関連していますが、社会的状況と疾病負荷を結びつける経路は依然として十分に理解されていません。私たちは、All of Us Research Program からの社会経済的、心理社会的、臨床、実験室、行動、およびゲノム データを統合する、AI 主導のマルチモーダルなメディエーション フレームワークを開発しました。モダリティ固有の変分オートエンコーダーを使用して各データドメインの潜在表現を導き出し、その後、潜在空間で媒介分析を実行して、社会経済的不利益、心理社会的要因、および多発性疾患の間の間接的な関連を評価しました。最終的な分析コホートには、完全なマルチモーダル データを持つ 20,804 人の参加者が含まれていました。 800 のエクスポージャー、メディエーター、結果の組み合わせ全体で、メディエーションシグナルは少数の潜在的な次元内に集中しました。最も強力な間接的関連性は、社会経済的不利益の側面、心理社会的脆弱性の側面、および心臓代謝性多疾患の側面と関連していた(NIE = 0.002517)。心理社会的側面は、精神的健康状態の悪化、孤独感の増大、社会的幸福度の低下、ヘルスリテラシーの低下によって特徴づけられましたが、結果の側面では、高血圧、糖尿病、高脂血症、肥満、慢性腎臓病、心臓病と関連していました。ブートストラップ解析により、主要経路の安定性が裏付けられました。これらの発見は、心理社会的脆弱性が、社会経済的不利益と心臓代謝性多疾患を結び付ける主要な潜在経路に強く表れていることを示唆しています。より広範には、提案されたフレームワークは、AI ベースの表現学習を使用して、高次元のマルチモーダルな健康データにわたる複雑な関係を調査する方法を示しています。
原文 (English)
AI-driven Multimodal Representation Learning for Latent Mediation Structure Discovery of Socioeconomic Disadvantage, Psychosocial Factors, and Cardiometabolic Multimorbidity: Insights from the All of Us Research Program
Social disadvantage is associated with multimorbidity, but the pathways linking social conditions to disease burden remain poorly understood. We developed an AI-driven multimodal mediation framework that integrates socioeconomic, psychosocial, clinical, laboratory, behavioral, and genomic data from the All of Us Research Program. Modality-specific variational autoencoders were used to derive latent representations of each data domain, and mediation analyses were subsequently performed in latent space to evaluate indirect associations between socioeconomic disadvantage, psychosocial factors, and multimorbidity. The final analytic cohort included 20,804 participants with complete multimodal data. Across 800 exposure--mediator--outcome combinations, mediation signals were concentrated within a small number of latent dimensions. The strongest indirect association linked a socioeconomic disadvantage dimension, a psychosocial vulnerability dimension, and a cardiometabolic multimorbidity dimension (NIE = 0.002517). The psychosocial dimension was characterized by poorer mental health, greater loneliness, lower social well-being, and lower health literacy, whereas the outcome dimension was associated with hypertension, diabetes, hyperlipidemia, obesity, chronic kidney disease, and heart disease. Bootstrap analyses supported the stability of the leading pathway. These findings suggest that psychosocial vulnerability was strongly represented in the dominant latent pathway linking socioeconomic disadvantage and cardiometabolic multimorbidity. More broadly, the proposed framework illustrates how AI-based representation learning can be used to investigate complex relationships across high-dimensional multimodal health data.
Agentic AI システムにおける実行リスクの管理: レッド チーム化のための軌道ガイド付きフレームワーク
AI エージェントは組織のワークフローにますます組み込まれており、外部の情報ソースと対話し、デジタル ツールを呼び出して運用タスクを実行します。組織がこのようなシステムを導入する場合、エージェントを意図しない行動に誘導する可能性のある悪意のあるまたは信頼できない外部情報から生じるリスクを特定し、軽減することが重要な課題となります。既存のレッドチームアプローチは主に固定された攻撃テンプレートまたは最終的な攻撃結果に依存しており、複数段階の推論とツールの使用を通じて攻撃がどのように展開するかについての可視性は限られています。私たちは、エージェント実行リスクは軌跡レベルの現象として理解されるべきであると主張します。この観点に基づいて、私たちは、実行軌跡を使用してエージェント AI システムの脆弱性を発見する軌跡ガイド型のレッドチーム フレームワークである TrajRed を提案します。さらに、レッド チーム化中に発見された高リスクの軌道を使用して進行中のワークフローを監視し、介入するランタイム ガバナンス レイヤーである TrajGuard を開発します。 4 つの組織タスク スイートにわたる AgentDojo の実験では、TrajRed が固定テンプレートおよび自動レッドチーム ベースラインよりも大幅に強力な脆弱性を特定することが示されています。これらの脆弱性の発見に基づいて、TrajGuard は、無害なタスクのユーティリティを維持しながら、評価されたすべての攻撃手法にわたる攻撃の成功率をほぼゼロに抑えます。これらの結果を総合すると、実行軌跡がエージェント AI システムにおけるレッド チーム化とリスク管理の両方に実用的な基盤を提供することが示されています。この研究は、組織の AI 導入におけるエージェント実行の管理の重要性を強調しています。
原文 (English)
Governing Execution Risk in Agentic AI Systems: A Trajectory-Guided Framework for Red Teaming
AI agents are increasingly embedded in organizational workflows, where they interact with external information sources and invoke digital tools to perform operational tasks. As organizations adopt such systems, a critical challenge is identifying and mitigating risks arising from malicious or untrusted external information that can steer agents toward unintended actions. Existing red-teaming approaches largely rely on fixed attack templates or final attack outcomes, providing limited visibility into how attacks unfold through multi-step reasoning and tool use. We argue that agent execution risk should be understood as a trajectory-level phenomenon. Building on this perspective, we propose TrajRed, a trajectory-guided red-teaming framework that uses execution trajectories to uncover vulnerabilities in agentic AI systems. We further develop TrajGuard, a runtime governance layer that uses high-risk trajectories discovered during red teaming to monitor and intervene in ongoing workflows. Experiments on AgentDojo across four organizational task suites show that TrajRed identifies substantially stronger vulnerabilities than fixed-template and automatic red-team baselines. Building on these vulnerability findings, TrajGuard reduces attack success across all evaluated attack methods to near zero while preserving benign task utility. Together, the results demonstrate that execution trajectories provide a practical foundation for both red teaming and risk control in agentic AI systems. This work highlights the importance of governing agent execution in organizational AI deployments.
モーメント推定のための信頼領域フレームワーク
この論文では、確率的勾配最適化における \textsc{Adam} などの適応モーメント推定メカニズムの動作を理解するための信頼領域フレームワークを開発します。具体的には、このフレームワークでは、個々の重みの更新ステップの大きさは、次数 $p\in[2,4]$ のモーメント制約によって支配される信頼領域内に制約されます。結果として得られる導出は、2 番目の瞬間の推定と正規化された $p$ 番目の瞬間の推定に基づく一連の学習率メカニズムにつながります。 $p=4$ の場合、これには尖度のような推定が含まれます。 \textsc{Gmake} と呼ばれる一般的なメカニズムは、共通の信頼領域フレームワーク内で、モーメント推定による正規化、学習率のスケジューリング、運動量としてのスペクトル ローパス フィルター処理、およびオペレーター レベルのスペクトル正規化の統一された解釈を提供します。 FineWeb-Edu と TinyStories で訓練された GPT2-124M の実験では、信頼領域の制約が弱い場合に 4 番目のモーメントの実現が最大の利点を提供することが示唆されています。徐々に強力な信頼領域制御が導入されるにつれて、2 番目の瞬間の実現の競争力が高まり、多くの場合、対応する 4 番目の瞬間の実現よりもわずかに低い検証損失が達成されます。
原文 (English)
A Trust-region Framework for Moment Estimation
In this paper, we develop a trust-region framework for understanding the behavior of adaptive moment estimation mechanisms, such as \textsc{Adam}, in stochastic gradient optimization. Specifically, in this framework, the magnitude of the update step for each individual weight is constrained within a trust-region governed by a moment constraint of order $p\in[2,4]$. The resulting derivation then leads to a family of learning-rate mechanisms based on second-moment estimation and a normalized $p$-th moment estimation. When $p=4$, this involves kurtosis-like estimation. The general mechanism, referred to as \textsc{Gmake}, provides a unified interpretation of normalization by moment estimation, learning-rate scheduling, spectral lowpass filtering as momentum, and operator-level spectral normalization within a common trust-region framework. Experiments on GPT2-124M trained on FineWeb-Edu and TinyStories suggest that the fourth-moment realization provides its greatest benefit when trust-region constraints are weak. As progressively stronger trust-region controls are introduced, the second-moment realization becomes increasingly competitive, often achieving slightly lower validation loss than its corresponding fourth-moment realization.
Lindblad-Inspired Multi-Timescale Reservoir Computing with Separable Rotation and Dissipation
Echo-state networks enable efficient temporal learning by fixing the recurrent dynamics and training only a linear readout. However, conven…
NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts
Generative artificial intelligence (AI) has transformed text-to-image synthesis, yet its ability to represent specialized engineering domai…
EDATracer: An Agentic Framework for Large-Scale EDA Artifact Analysis
Modern chip design relies on electronic design automation (EDA) tools that generate large, heterogeneous artifacts, including source files,…
CheckOne: Lightweight Fault Detection and Mitigation for Vision Transformers
The wide adoption of Vision Transformers (ViTs) in safety-critical applications raises reliability concerns related to hardware faults. Alg…
物語に基づいたインタラクティブな体験のための物語から永続的な世界を再構築する
インタラクティブ コンテンツは、物語が暗示する根底にある世界と一致する必要があるため、物語に基づいたインタラクティブ エクスペリエンスのデザインには依然として労力がかかります。既存のアプローチは、物語の計画、シーンの生成、ゲームプレイの生成などの問題を定式化し、それぞれが根拠となる永続的な世界を明示的に再構築して維持するのではなく、特定の下流タスクに合わせた計算表現を構築します。私たちは、物語に基づいたインタラクティブな実現のための中心的な計算目標として、物語の記述から明示的な永続世界を再構築することを研究します。私たちのアプローチは、世界を下流生成の暗黙的な副産物として扱うのではなく、一貫したインタラクティブな体験をサポートするために必要なコンテキスト情報のみを推論しながら、永続的なエンティティ、場所、意味関係、進化する世界状態を再構築して維持します。この観点を調査するために、物語の記述から構造化された永続的な世界表現を再構築し、その後、プレイ可能なタイルベースの環境をインスタンス化する参照プロトタイプを開発します。手続き型シナリオ、オリジナルのファンタジー物語、翻案されたパブリックドメインの物語にわたる 3 つの代表的なケーススタディを通じて、永続的な世界を再構築する実現可能性を実証し、共有世界の表現がソースの物語に根ざしたままで一貫したゲームプレイをどのようにサポートするかを示します。この作品は、インタラクティブな実現に先立って永続的な世界を明示的に再構築することで、計算による物語の理解とインタラクティブなコンテンツ生成の橋渡しをし、AI 支援のゲーム オーサリング、混合イニシアチブ デザイン、教育シミュレーション、および物語に基づいたインタラクティブ エクスペリエンスのためのセマンティック基盤を提供します。
原文 (English)
Reconstructing Persistent Worlds from Narratives for Narrative-Grounded Interactive Experiences
Designing narrative-grounded interactive experiences remains labor-intensive because interactive content must align with the underlying world implied by the narrative. Existing approaches formulate problems such as narrative planning, scene generation, and gameplay generation, each constructing computational representations tailored to specific downstream tasks rather than explicitly reconstructing and maintaining the persistent world that grounds them. We investigate reconstructing explicit persistent worlds from narrative descriptions as the central computational objective for narrative-grounded interactive realization. Rather than treating the world as an implicit by-product of downstream generation, our approach reconstructs and maintains persistent entities, locations, semantic relationships, and evolving world states while inferring only the contextual information required to support coherent interactive experiences. To investigate this perspective, we develop a reference prototype that reconstructs structured persistent world representations from narrative descriptions and subsequently instantiates playable tile-based environments. Through three representative case studies spanning a procedural scenario, an original fantasy narrative, and an adapted public-domain story, we demonstrate the feasibility of reconstructing persistent worlds and show how a shared world representation supports coherent gameplay while remaining grounded in the source narrative. By explicitly reconstructing persistent worlds prior to interactive realization, this work bridges computational narrative understanding and interactive content generation, providing a semantic foundation for AI-assisted game authoring, mixed-initiative design, educational simulations, and narrative-grounded interactive experiences.
クライアントの良性および敵対性の異質性下での航空機エンジン予測のための堅牢かつパーソナライズされたフェデレーテッド ラーニング
フェデレーテッド ラーニング (FL) を使用すると、航空機の運航者は、生データを共有することなく、エンジン センサーの遠隔測定から残存耐用年数 (RUL) モデルを共同でトレーニングできます。この調査では、2 つの相補的な課題を調査します。1 つは誠実なオペレーターが異なる動作条件と障害モードを観察する「良性異質性」、もう 1 つは侵害されたオペレーターが有害な更新を送信する「敵対的異質性」です。マルチタスク 1 次元畳み込みニューラル ネットワークと、商用モジュール式航空推進システム シミュレーション (C-MAPSS) ベンチマークの構造的に非 IID パーティションを使用して、制御された安全指向の評価を実行します。良性の異質性に対する 4 つの救済策を比較し、エンジンの劣化を隠すために設計された物理的動機によるセンサー値のバックドアを含む 4 つの集計方法に対する 5 つの攻撃を評価します。共有表現のパーソナライゼーションは、ローカルと集中の二乗平均平方根誤差の差を約 70% 縮めます。これに対し、近位正則化では 21%、サーバー側の再重み付けでは 10% です。バックドアは、クリーンな精度を統計的に変化させずに、標準平均に対して 94.9% の攻撃成功率を達成しました。これは、精度だけではモデルの安全性を証明できず、攻撃の成功を明示的に評価する必要があることを示しています。 Krum は攻撃の成功を桁違いに減らし、組織的な攻撃者に耐えられる唯一の評価済みアグリゲーターですが、パーソナライゼーションだけでは保護は提供されません。パーソナライゼーションと堅牢な集約を組み合わせると、わずかな精度コストで堅牢性 (2.8% の攻撃成功) が回復され、堅牢な更新の選択と協調的な表現学習の間のトレードオフが明らかになります。結果は、クライアント数全体およびより難しい 6 条件のデータセットにわたって一貫性を保ちます。コードとデータのパーティションは再現性を確保するためにリリースされています。
原文 (English)
Robust and Personalized Federated Learning for Aircraft-Engine Prognostics under Benign and Adversarial Client Heterogeneity
Federated learning (FL) enables aircraft fleet operators to jointly train remaining-useful-life (RUL) models from engine sensor telemetry without sharing raw data. This study examines two complementary challenges: benign heterogeneity, where honest operators observe different operating conditions and fault modes, and adversarial heterogeneity, where compromised operators submit poisoned updates. We conduct a controlled, safety-oriented evaluation using a multi-task one-dimensional convolutional neural network and a structurally non-IID partition of the Commercial Modular Aero-Propulsion System Simulation (C-MAPSS) benchmark. We compare four remedies for benign heterogeneity and evaluate five attacks against four aggregation methods, including a physically motivated sensor-value backdoor designed to mask engine degradation. Shared-representation personalization closes approximately 70% of the local-to-centralized root-mean-square-error gap, compared with 21% for proximal regularization and 10% for server-side reweighting. The backdoor achieves a 94.9% attack success rate against standard averaging while leaving clean accuracy statistically unchanged, demonstrating that accuracy alone cannot certify model safety and that attack success must be evaluated explicitly. Krum reduces attack success by an order of magnitude and is the only evaluated aggregator that withstands coordinated attackers, whereas personalization alone provides no protection. Combining personalization with robust aggregation restores robustness (2.8% attack success) with only a small accuracy cost, revealing a trade-off between robust update selection and collaborative representation learning. Results remain consistent across client counts and on a harder six-condition dataset. Code and data partitions are released for reproducibility.
QBER しきい値を超えて: BB84 QKD でのマルチ攻撃検出のための時間的 QBER ベースの機械学習フレームワーク
従来の BB84 量子鍵配布 (QKD) システムは、盗聴を検出するために 11% の固定量子ビット誤り率 (QBER) しきい値に依存しています。ただし、ステルス攻撃はこのしきい値を下回ったままでも、チャネルのセキュリティを侵害する可能性があります。この論文では、BB84 QKD システムにおける盗聴攻撃を検出および分類するための時間 QBER ベースの機械学習フレームワークを提案します。このフレームワークは、平均セッション レベル QBER に依存するのではなく、バースト動作、時間的不安定性、基底依存の非対称性、および QBER 損失の相互作用を捕捉する 63 の物理学に基づいた時間的特徴を抽出します。ランダム フォレスト、XGBoost、および放射基底関数カーネル (SVM-RBF) 分類器を備えたサポート ベクター マシンは、ノイズと損失の多い条件下で 7 つの盗聴攻撃と通常のチャネル シナリオで評価されます。 10 回の独立した実行を平均すると、XGBoost は 88.01% (0.47%) の精度と 0.8803 のマクロ F1 スコアで最高のパフォーマンスを達成しましたが、SVM-RBF も同等のパフォーマンスを示し、提案された機能の堅牢性が確認されました。従来のモニタリングと比較するためのバイナリ攻撃対通常の検出器として評価すると、固定の 11% QBER しきい値では、偽陰性率 (FNR) が 0.8477 で 25.82% の精度しか達成できませんが、提案されたフレームワークでは FNR が 0.0198 に低下し、しきい値ベースのモニタリングを回避するステルス攻撃の検出が大幅に向上します。 SHapley Additive exPlanations based (SHAP) の説明可能性は、物理学に基づいた時間的およびチャネル由来の特徴が、盗聴戦略を識別する上で高度に識別できることを示しています。これらの結果は、時間的 QBER 駆動の機械学習が、BB84 QKD システムにおけるマルチ攻撃セキュリティ監視のための正確で説明可能な実用的なフレームワークを提供することを示しています。
原文 (English)
Beyond the QBER Threshold: A Temporal QBER Based Machine Learning Framework for Multi Attack Detection in BB84 QKD
Conventional BB84 Quantum Key Distribution (QKD) systems rely on a fixed 11% Quantum Bit Error Rate (QBER) threshold to detect eavesdropping. However, stealthy attacks can remain below this threshold while still compromising channel security. This paper proposes a temporal QBER based machine learning framework for detecting and classifying eavesdropping attacks in BB84 QKD systems. Rather than relying on average session level QBER, the framework extracts 63 physics-informed temporal features capturing burst behavior, temporal instability, basis dependent asymmetry, and QBER loss interactions. Random Forest, XGBoost, and Support Vector Machine with a Radial Basis Function kernel (SVM-RBF) classifiers are evaluated on seven eavesdropping attacks and a normal channel scenario under noisy and lossy conditions. Averaged over ten independent runs, XGBoost achieves the best performance with 88.01% (0.47%) accuracy and a macro F1 score of 0.8803, while SVM-RBF performs comparably, confirming the robustness of the proposed features. Evaluated as a binary attack-versus-normal detector for comparison with conventional monitoring, a fixed 11% QBER threshold achieves only 25.82% accuracy with a False Negative Rate (FNR) of 0.8477, whereas the proposed framework reduces the FNR to 0.0198, substantially improving detection of stealthy attacks that evade threshold-based monitoring. SHapley Additive exPlanations based (SHAP) explainability shows that physics-informed temporal and channel derived features are highly discriminative for identifying eavesdropping strategies. These results demonstrate that temporal QBER driven machine learning provides an accurate, explainable, and practical framework for multi attack security monitoring in BB84 QKD systems.
反復残差量子化: LLM のプログレッシブ多精度表現
さまざまな展開上の制約の下で大規模言語モデル (LLM) を提供するには、精度、メモリ フットプリント、スループットの間で柔軟なトレードオフが必要です。ただし、従来の量子化方法では通常、ターゲット ビット幅ごとに個別のチェックポイントが必要です。一連の量子化残差補正とともに低ビット量子化ベースとして重みを表現するポストトレーニング量子化 (PTQ) フレームワークであるリカレント残差量子化 (RRQ) を導入し、単一のチェックポイントから複数の効果的な精度を実現します。 RRQ は、ポストトレーニング量子化 (PTQ) または最近接丸め (RTN) によって取得された 2 ビット モデルから開始して、RTN によって生成された軽量の 2 ビット残差を段階的に追加して、4、6、および 8 ビット表現を構築します。この方法はキャリブレーション不要で、マルチビットの同時最適化を回避します。 Qwen3-8B セットアップでは、完全な全 RTN 2/4/6/8 ビット パッケージが 1,293 秒で構築され、これは測定された MatGPTQ 構築よりも 3.3 倍高速です。最近の 6 つの LLM での実験では、6 ビットと 8 ビットで競合する精度が示され、4 ビットではモデルに依存する動作が示されています。コードは公開と同時に公開されます。
原文 (English)
Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs
Serving large language models (LLMs) under diverse deployment constraints requires flexible trade-offs between accuracy, memory footprint, and throughput. However, conventional quantization methods typically require a separate checkpoint for each target bit-width. We introduce Recurrent Residual Quantization (RRQ), a post-training quantization (PTQ) framework that represents weights as a low-bit quantized base together with a sequence of quantized residual corrections, enabling multiple effective precisions from a single checkpoint. Starting from a 2-bit model obtained via post-training quantization (PTQ) or round-to-nearest (RTN), RRQ progressively adds lightweight 2-bit residuals generated via RTN to construct 4-, 6-, and 8-bit representations. The method is calibration-free and avoids joint multi-bit optimization. In our Qwen3-8B setup, the full all-RTN 2-/4-/6-/8-bit package is constructed in 1,293 seconds, 3.3 times faster than the measured MatGPTQ construction. Experiments on six recent LLMs show competitive accuracy at 6 and 8 bits, with model-dependent behavior at 4 bits. The code will be made publicly available upon publication.
AgentAntibody: 即時注射から LLM エージェントを防御する適応免疫システム
即時注入は LLM エージェントにとって重大な脅威であることに変わりはありませんが、既存の防御では、以前の遭遇とは無関係に、各タスクが自己完結型の問題として扱われます。実際には、ユーザー要求は十分に指定されていないことが多く、許容可能な動作を完全には指定せずに、望ましい結果を記述します。インジェクションはこの曖昧さを悪用し、ユーザーが拒否するような方法でエージェントにタスクを完了させる可能性があります。具体的なケースを通じてユーザーの期待がより明確になるにつれて、弁護側はそれぞれの出会いから学び、学んだことを次の出会いに適用する必要があります。適応免疫からインスピレーションを得て、私たちは、即時注射に対する自己進化する免疫システムを LLM エージェントに装備する AgentAntibody を提案します。 AgentAntibody は、ユーザーのセキュリティ境界の進化する理解を抗体の永続的なライブラリとして表します。実行時に、ライブラリはこの境界に対する脅威を認識し、対応する免疫応答を開始します。遭遇を重ねるごとに、将来の攻撃に対するエージェントの免疫力を強化するために進化します。 3 つのベンチマークと 4 つのバックボーン LLM にわたる広範な実験により、AgentAntibody は経験を通じてユーザーの境界を学習することで、有害なアクションと正当なアクションが両方とも指定されたタスクと互換性がある場合でも、正当なタスクの完了を維持しながら有害なアクションを防止する点で既存の防御を上回っていることが示されています。
原文 (English)
AgentAntibody: An Adaptive Immune System for Defending LLM Agents against Prompt Injection
Prompt injection remains a critical threat to LLM agents, yet existing defenses treat each task as a self-contained problem, independent of previous encounters. In practice, user requests are often underspecified: they describe the desired outcome without fully specifying acceptable behavior. An injection can exploit this ambiguity, causing the agent to complete the task in a way the user would reject. As the user's expectations become clearer through concrete cases, a defense should learn from each encounter and apply what it learns to the next. Inspired by adaptive immunity, we propose AgentAntibody, which equips LLM agents with a self-evolving immune system against prompt injection. AgentAntibody represents its evolving understanding of the user's security boundary as a persistent library of antibodies. At runtime, the library recognizes threats to this boundary and mounts corresponding immune responses. Across encounters, it evolves to strengthen the agent's immunity to future attacks. Extensive experiments across three benchmarks and four backbone LLMs show that, by learning the user's boundary through experience, AgentAntibody outperforms existing defenses in preventing harmful actions while preserving legitimate task completion, even when the harmful and legitimate actions are both compatible with the stated task.
マルチモーダルな意図を理解するためのモダリティ合意および競合を認識したプロトタイプ ハイパーグラフ学習
マルチモーダルな意図認識では、テキスト、音響、視覚信号が何を共有しているかだけでなく、それらがどのように一致しないのかを理解する必要があります。このような意見の相違は、多くの場合、階級に影響を与えます。たとえば、不一致な声や顔の動作を伴う語彙のポジティブさは、皮肉や嘲笑を示している可能性がありますが、ほとんどの融合手法はモダリティの一致を奨励するか、不一致を抑制すべき不確実性として扱うかのいずれかです。我々は、マルチモーダルな合意と対立を別個の繰り返しの関係構造として表現する階層型プロトタイプ ハイパーグラフ フレームワークである MACH (モダリティ合意と対立を意識したプロトタイプ ハイパーグラフ) を提案します。 MACH は、単峰性表現を二峰性および三峰性の抽象化に段階的に合成します。該当する各レベルで、モダリティ構成アンカーは、再利用可能なコンセンサス パターンをキャプチャするスパースな合意プロトタイプ ハイパーグラフをアクティブにし、別の競合経路がクロスモーダルの不一致を専用の競合プロトタイプ ハイパーグラフにマッピングします。 2 つの経路は、機能ごとのサンプル適応型調停メカニズムによって結合され、モデルが付随的なモダリティ ノイズを抑制しながら情報の不一致を維持できるようになります。漸進的最適化戦略は、合意と衝突の共同学習の前に相互依存階層を安定させます。ベンチマーク データセットの実験では、提案された定式化の有効性が実証され、コンポーネント分析とロバスト性分析では、階層構成、プロトタイプを介した意味論的改良、合意と競合の調停の明確な役割が検証されます。
原文 (English)
Modality Agreement- and Conflict-Aware Prototype Hypergraph Learning for Multimodal Intent Understanding
Multimodal intent recognition requires understanding not only what textual, acoustic, and visual signals share, but also how they disagree. Such disagreement is frequently class-informative; for example, lexical positivity accompanied by incongruent vocal or facial behavior may indicate sarcasm or taunting, yet most fusion methods either encourage modality alignment or treat inconsistency as uncertainty to be suppressed. We propose MACH (Modality Agreement- and Conflict-aware prototype Hypergraph), a hierarchical prototype-hypergraph framework that represents multimodal agreement and conflict as distinct, recurring relational structures. MACH progressively composes unimodal representations into bimodal and trimodal abstractions. At each applicable level, modality-composition anchors activate sparse agreement prototype hypergraphs that capture reusable consensus patterns, while a separate conflict pathway maps cross-modal discrepancies to dedicated conflict prototype hypergraphs. The two pathways are combined through a feature-wise, sample-adaptive arbitration mechanism, enabling the model to preserve informative disagreement while suppressing incidental modality noise. A progressive optimization strategy stabilizes the interdependent hierarchy before joint agreement-conflict learning. Experiments on benchmark datasets demonstrate the effectiveness of the proposed formulation, while component and robustness analyses validate the distinct roles of hierarchical composition, prototype-mediated semantic refinement, and agreement-conflict arbitration.
LaPrune: 100 万スケールでの制御可能な微分可能なスパース性
Top-$k$ の選択により、疎モデルのどのコンポーネントがアクティブのままになるかが決まります。ハード選択では勾配がブロックされますが、連続的な緩和ではマスクの硬さが選択された質量に結合されることがよくあります。 LaPrune は、選択された質量を維持しながら正規化された二次モーメントを制御する、数学的に正確な予算で微分可能な層です。 LapSum バリアは選択質量を保存し、正規化された 2 番目のモーメント制約によりマスクが密な等質量割り当てから各バジェットのハードトップ $k$ に移動します。飽和部分の人口予測、二値に近い制限法則、およびゼロに近い部分に対する厳しい最悪の場合の保証を導き出します。正規化された硬度パラメータはスコアスケールに対して不変ですが、固定された LapSum 温度はそうではありません。
原文 (English)
LaPrune: Controllable Differentiable Sparsity at Million Scale
Top-$k$ selection determines which components of a sparse model remain active. Hard selection blocks gradients, while continuous relaxations often couple mask hardness to the selected mass. We introduce LaPrune, a mathematically exact-budget differentiable layer that controls the normalized second moment while preserving the selected mass. A LapSum barrier preserves the selection mass, and a normalized second-moment constraint moves the mask from a dense equal-mass allocation toward hard top-$k$ at each budget. We derive a population prediction of the saturated fraction, a near-binary limiting law, and a tight worst-case guarantee on the near-zero fraction. The normalized hardness parameter is invariant to score scale, while a fixed LapSum temperature is not.
SJEPA: ハイブリッド記号ニューラル予測子によるエレガントな潜在ダイナミクスの学習
ジョイントエンベディング予測アーキテクチャは、コンテキストエンベディングからターゲットエンベディングを予測することで抽象状態を学習しますが、その遷移モデルは通常、不透明なニューラルマップです。我々は、誘導されたダイナミクスがコンパクトな記号記述を可能にする予測表現を学習する、再構築不要の JEPA フレームワークである SJEPA を紹介します。そのハイブリッド遷移は、記号法則と、選択された文法外のダイナミクスに対する正規化されたニューラル補正を組み合わせたものです。中心的な原則は、最も単純で適切なダイナミクスを学習することです。表現制約は有益な非崩壊予測座標を保存しますが、オペレーター圧縮は、予測的に適切な状態を維持する低複雑性のシンボリック ニューラル遷移を優先します。我々は、誘起ダイナミクスの複雑さを通じてこの原理を形式化し、予測座標の非識別可能性を分析し、制約のない演算子圧縮が表現崩壊への直接の近道を生み出すことを示します。このフレームワークは、交互表現方程式学習と固定表現に適合したシンボリック ダイナミクスの両方をサポートします。制御された振り子実験では、共同学習により、事後フィッティングよりも長期ロールアウト誤差と発散が低い、実質的に単純なシンボリックダイナミクスが発見され、一方、制約のないワンステップ診断により、予測された崩壊のショートカットが実現されます。文法の仕様が間違っている場合、修正正則化により表現可能な記号メカニズムが保存され、ニューラル コンポーネントが残差ダイナミクスに向けられます。結果は、予測の忠実度、表現品質、記号の節約、記号とニューラルの割り当ての間の制御可能なトレードオフを明らかにします。
原文 (English)
SJEPA: Learning Elegant Latent Dynamics with Hybrid Symbolic-Neural Predictors
Joint-embedding predictive architectures learn abstract states by predicting target embeddings from context embeddings, but their transition models are typically opaque neural maps. We introduce SJEPA, a reconstruction-free JEPA framework that learns predictive representations whose induced dynamics admit compact symbolic descriptions. Its hybrid transition combines a symbolic law with a regularised neural correction for dynamics outside the selected grammar. The central principle is to learn the simplest adequate dynamics: representation constraints preserve informative, non-collapsed predictive coordinates, while operator compression favours low-complexity symbolic-neural transitions that remain predictively adequate. We formalise this principle through induced-dynamics complexity, analyse predictive-coordinate non-identifiability, and show that unconstrained operator compression creates a direct shortcut to representation collapse. The framework supports both alternating representation-equation learning and symbolic dynamics fitted to fixed representations. In controlled pendulum experiments, joint learning discovers substantially simpler symbolic dynamics with lower long-horizon rollout error and divergence than post-hoc fitting, while an unconstrained one-step diagnostic realises the predicted collapse shortcut. Under grammar misspecification, correction regularisation preserves the representable symbolic mechanism and directs the neural component towards residual dynamics. The results expose a controllable trade-off among predictive fidelity, representation quality, symbolic parsimony, and symbolic-neural allocation.
An Inline Control Architecture for Language Models in Intelligent Transportation Systems
Vehicle-to-everything (V2X) systems increasingly incorporate large language models (LLMs) for semantic tasks such as message summarization,…
FBID: IoT ネットワークにおける堅牢な分散外攻撃検出のための適応型パーソナライズされたフェデレーテッド ラーニング
パーソナライズされたフェデレーテッド ラーニング (PFL) は、高度に非独立かつ同一に分散された (非 IID) データ分散下でのローカル適応を改善できるため、異種 IoT 環境における侵入検知の有望なソリューションとして浮上しています。ただし、既存の PFL 手法はクライアント側の自己調整に依存することが多く、過度のパーソナライゼーションや配布外 (OOD) 攻撃検出の大幅な低下につながる可能性があります。このペーパーでは、サーバー側のパーソナライゼーション制御を通じてこの制限に対処する新しい適応 PFL フレームワークである Federated Bandit Intrusion Detection (FBID) を提案します。特に、FBID はサーバーでコンテキスト マルチアーム バンディットを採用し、観察された動作と更新品質に応じて各クライアントのローカル トレーニング強度を動的に調整します。さらに、FBID は信頼ベースのブレンディング メカニズムを導入して、グローバル モデルとローカル モデルの間のクライアント固有の内挿係数を導出します。これにより、グローバルな攻撃検出の知識を維持しながら、有益なローカルの特殊化を可能にします。異種クライアント分布と OOD ストレス テスト設定の下での CICIoT2023 データセットに対する広範な実験を通じて、FBID が個々のクライアントの OOD 検出率 (DR) を最も安定したベースラインと比較して最大 7.66%、F1 スコア (F1) を最大 5.08% (相対) 改善し、同時にこれまで見たことのない攻撃クラスに対する堅牢性も向上させることを示しました。
原文 (English)
FBID: Adaptive Personalized Federated Learning for Robust Out-of-Distribution Attack Detection in IoT Networks
Personalized Federated Learning (PFL) has emerged as a promising solution for intrusion detection in heterogeneous IoT environments, as it can improve local adaptation under highly Non-Independent and Identically Distributed (non-IID) data distributions. However, existing PFL methods often rely on client-side self-adjustment, which may lead to over-personalization and substantial degradation in out-of-distribution (OOD) attack detection. In this paper, we propose Federated Bandit Intrusion Detection (FBID), a novel adaptive PFL framework to address this limitation through server-side personalization control. In particular, FBID employs a contextual multi-armed bandit at the server to dynamically regulate each client's local training intensity according to its observed behavior and update quality. Moreover, FBID introduces a trust-based blending mechanism to derive client-specific interpolation coefficients between the global and local models, thereby preserving global attack-detection knowledge while still allowing beneficial local specialization. Through extensive experiments on the CICIoT2023 dataset under heterogeneous client distributions and OOD stress-test settings, we show that FBID improves individual client OOD Detection Rate (DR) by up to 7.66% and F1-Score (F1) by up to 5.08% (relative) over the strongest stable baseline, while also improving robustness to previously unseen attack classes.
クエリが必要な場所にビットを使用: アテンション保持変換を使用した KV キャッシュ ベクトル量子化
ロングコンテキスト LLM デコードでは、各ステップでキー/値 (KV) キャッシュを読み取ります。ロードには注意を計算するよりも時間がかかるため、スループットは帯域幅に制限されます。したがって、キャッシュ サイズを削減すると、デコード速度と処理能力の両方を向上させることができます。課題は、注目のプロダクトを維持し、再構築を安価に保ち、トークンごとの固定ビット数を使用しながら、キャッシュ サイズを削減することです。要素あたり 2 ビットの場合、最も競争力のある方法は直交変換に依存します。ただし、既存の手法はデータを意識しないか、歪み基準から変換を導出せずにクエリ統計を使用します。さらに、エネルギーを圧縮するのではなくエントリ間の分散を均等化するランダムまたはアダマール回転に基づいて構築された変換と、低レートでは最適ではない固定幅のスカラー量子化器に依存します。本稿では KV キャッシュ量子化を注目製品の誤差を歪みとする変換符号化問題として定式化する。高解像度モデルの下で、キャリブレーション統計からキーと値の閉じた形式の最適な変換を導き出します。最適なキー変換は直交ではなく、一般化されたパーセバル関係を満たすことを示します。つまり、注意を意識した歪みは、変換領域で平均二乗誤差 (MSE) になります。したがって、変換されたキー係数に直接適用される MSE 最適ベクトル量子化器を使用できます。固定幅レイアウト要件を満たすために、係数を等しいボリュームのパーティションにグループ化すると、同じ高解像度モデルの下で等しいサイズのコードブックが可変レートの最適化を達成できることを示します。 NOVA-KV と呼ばれる私たちの方法は、要素あたり 2 ビットで、スカラー量子化法によって失われたロングコンテキストの検索精度のほとんどを同等のスループットで回復します。
原文 (English)
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms
Long-context LLM decoding reads the key-value (KV) cache at every step. Loading it takes longer than computing attention over it, so throughput is bandwidth-bound. Hence, reducing the cache size can raise both decoding speed and serving capacity. The challenge is to reduce cache size while preserving the attention products, keeping reconstruction cheap, and using a fixed per-token bit count. At two bits per element, the most competitive methods rely on orthogonal transforms. However, existing techniques are either data-oblivious or use the query statistics without deriving the transform from a distortion criterion. Moreover, they rely on transforms built on top of random or Hadamard rotations, which equalize variances across entries rather than compacting energy, and fixed-width scalar quantizers, which are suboptimal at low rates. In this paper, we formulate KV cache quantization as a transform coding problem in which distortion is the error in the attention products. We derive closed-form optimal transforms for keys and values from calibration statistics, under a high-resolution model. We show that the optimal key transform is not orthogonal and satisfies a generalized Parseval relation: the attention-aware distortion becomes mean-squared error (MSE) in the transform domain. Thus, we can use MSE-optimal vector quantizers applied directly to the transformed key coefficients. To meet the fixed-width layout requirement, we show that grouping coefficients into equal-volume partitions makes equal-size codebooks attain the variable-rate optimum under the same high-resolution model. At two bits per element, our method, termed NOVA-KV, recovers most of the long-context retrieval accuracy lost by scalar quantization methods at comparable throughput.
Spatiotemporal Graph Transformer for Traffic Intelligence in Edge Computing
Accurate traffic forecasting is essential for proactive resource management in edge computing, where service demand evolves dynamically acr…
ループ外のマルチフィデリティ ベイジアン最適化
ブラックボックス最適化は科学と工学のいたるところに存在する問題であり、多くの場合、安価で忠実度の低いプロキシを利用して高価な目的関数を処理します。マルチ忠実度ベイジアン最適化 (MF-BO) は、この問題に対する原則的なアプローチであり、目的をクエリするときにさまざまな忠実度間の相関を利用します。ただし、多くの重要な MF-BO タスクでは、真の最も忠実度の高い関数を最適化ループの一部にするには非常に高価です。それにもかかわらず、専門家は、現在のタスクに情報を提供する可能性のある、以前の実験から得られたゴールドスタンダード データ (最も忠実度の高い関数の観察結果) を持っていることがよくあります。たとえば、分子の最適化では、化学者はさまざまなコンピューター シミュレーションを使用して上位の $k$ 候補分子を選択し、後でそれらの真の目的関数の値を明らかにします。この研究では、理想的な仮定のもとでも、上記の現実世界のシナリオにおける標準 MF-BO アルゴリズムの準最適性を実証します。次に、タスク記述子を伴う履歴の高忠実度データを組み込むことで、この問題を軽減します。タスク記述子は、明示的に指定するか、非構造化メタデータから抽出できます。私たちは、合成関数だけでなく、化学やハイパーパラメーターの最適化における現実世界の問題に対する私たちの方法の有効性を実証します。
原文 (English)
Out-Of-The-Loop Multi-Fidelity Bayesian Optimization
Black-box optimization is a ubiquitous problem in science and engineering, often dealing with expensive objective functions with cheaper lower-fidelity proxies available. Multi-fidelity Bayesian optimization (MF-BO) is a principled approach to this problem, leveraging correlations across different fidelities when querying the objective. However, for many important MF-BO tasks, the true highest-fidelity function is prohibitively expensive to be part of the optimization loop. Nevertheless, practitioners often have gold standard data (observations of the highest-fidelity function) obtained from previous experiments that might provide information for the current task. For instance, in molecular optimization, chemists often pick the top-$k$ candidate molecules using various computer simulations, and later reveal their true objective function values. In this work, we demonstrate the suboptimality of standard MF-BO algorithms in the real-world scenarios above, even under ideal assumptions. Next, we mitigate this problem by incorporating historical high-fidelity data accompanied by task descriptors---which can be explicitly given or extracted from unstructured metadata. We demonstrate the effectiveness of our methods on synthetic functions, as well as real-world problems in chemistry and hyperparameter optimization.
Interpretable Fuzzy Inference for UAV Target Tracking Using Bounding-Box Geometry
Vision-based guidance of unmanned aerial vehicles (UAVs) toward unmanned ground vehicles (UGVs) supports cooperative aerial--ground robotic…
推論前の認識: ビデオの理解と質問応答のための動的潜在推論
ビデオ質問応答では、モデルが視覚的な証拠で地上言語のクエリを実行し、必要に応じて時間をかけてその証拠に基づいて推論する必要があります。既存の方法は、関連するオブジェクト、アクション、またはフレームが特定されるとすぐに多くの質問に答えることができるにもかかわらず、通常、長いテキストの思考連鎖の理論的根拠に依存しています。私たちは、動的潜在推論 (DyLaR) を提案します。これは、最初に知覚潜在の短いブロック (クエリに関連する視覚的証拠をエンコードする連続的な隠れた状態) で質問を根拠付け、次に回答する前に推論潜在 (潜在空間内のこの証拠を推論する連続的な思考) を追加するかどうかを適応的に決定します。 DyLaR は、検証された視覚的証拠に潜在する知覚を基礎付け、検証された理論的根拠を抽出して潜在的な推論に導き、その後、いつ推論するかをさらに洗練する強化学習によってこの動作を学習します。 DyLaR は、9 つのビデオ ベンチマークと 4 つのマルチモーダル言語モデル バックボーンにわたって、同じバックボーン ベースラインと比較して平均精度を向上させながら、クエリあたり生成するトークンの数を 20 未満に抑えます。たとえば、Qwen3-VL-4B では、DyLaR は Qwen3-VL-4B-Thinking よりも平均精度を 54.0 から 58.2 に向上させ、同時にクエリあたりの応答長を 1,220.7 から 18.5 トークンに短縮します。アブレーションはさらに、根拠に基づいた知覚の潜在力、根拠に基づいて監視された推論の潜在力、および適応型ルーティングがそれぞれ精度を向上させることを示しています。
原文 (English)
Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering
Video question answering requires models to ground language queries in visual evidence and, when necessary, reason over that evidence across time. Existing methods typically rely on long textual chain-of-thought rationales, even though many questions can be answered as soon as the relevant object, action, or frame is localized. We propose Dynamic Latent Reasoning (DyLaR), which first grounds a question in a short block of perception latents (continuous hidden states that encode query-relevant visual evidence), and then adaptively decides whether to append reasoning latents (continuous thoughts that reason over this evidence in latent space) before answering. DyLaR learns this behavior by grounding perception latents in verified visual evidence and distilling verified rationales into reasoning latents, followed by reinforcement learning that further refines when to reason. Across nine video benchmarks and four multimodal language model backbones, DyLaR improves average accuracy over same-backbone baselines while generating fewer than 20 tokens per query. On Qwen3-VL-4B, for example, DyLaR improves average accuracy over Qwen3-VL-4B-Thinking from 54.0 to 58.2 while reducing response length from 1,220.7 to 18.5 tokens per query. Ablations further show that grounded perception latents, rationale-supervised reasoning latents, and adaptive routing each improve accuracy.
InvFlowFD: フローマッチング反転を使用したリファレンスフリーおよびバックグラウンドセットフリーの知覚音楽品質メトリクス
音楽の知覚品質を評価するための既存のリファレンスフリーの方法では、ノイズの多いクリーンなデータのペアの必要性が軽減されていますが、依然としてクリーンなオーディオ サンプルの集計統計を計算するために使用されるバックグラウンド セットに依存しています。この研究では、この要件を排除し、事前トレーニングされたフロー マッチング バックボーンのみを使用してバックグラウンド セット フリーおよびリファレンス フリーの品質推定を達成する新しいアプローチを提案します。我々は、単純なオイラー積分による無条件のフローマッチング逆変換が、さまざまな人工歪みを検出し、人間の知覚判断に対して音楽生成モデルを正確にランク付けするのに十分であることを実証します。 InvFlowFD を導入します。これは、フロー反転を実行し、反転されたサンプルのグループを以前の分布と比較します。私たちは、人間による徹底的な研究を用いて、定量的に、以前の研究に対して私たちの方法を評価します。結果は、InvFlowFD が既存の指標よりも柔軟で制限が少ない一方で、音の歪みに対する人間の認識や生成モデルの品質と高い相関があることを示唆しています。
原文 (English)
InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion
Existing reference-free methods for evaluating music perceptual quality alleviate the need for paired noisy-clean data, but they still rely on a background set, which is used to compute aggregated statistics of clean audio samples. In this work, we propose a novel approach that eliminates this requirement, achieving background-set-free and reference-free quality estimation using only a pre-trained Flow Matching backbone. We demonstrate that unconditional Flow Matching inversion via simple Euler integration is sufficient to detect various artificial distortions and accurately rank music generation models against human perceptual judgments. We introduce InvFlowFD, which performs flow inversion and compares a group of inverted samples to the prior distribution. We evaluate our method against prior work, quantitatively and with a thorough human study. Results suggest that InvFlowFD is highly correlated with human perception of sound distortions, as well as generative models' quality, while being more flexible and less restrictive than existing metrics.
LiNC: サンプルごとの信頼とガウス混合モデリングによる軽量ノイズ補正
ラベル ノイズは、評価者間の変動性、アノテーション エラー、あいまいなケースなどの要因により、医療画像データセットでよく見られます。これにより、これらのデータセットを使用してトレーニングされた機械学習モデルの信頼性と臨床効果が大幅に損なわれる可能性があります。この課題に対処するために、Lightweight Noise Correction (LiNC) を導入しました。これは、トレーニング サンプルごとに 1 つのトレーニング可能な信頼パラメータを追加し、標準トレーニング ループ中に観察されたラベルをいつ使用するか、いつモデルに従うべきかを学習します。重要なアイデアは、サンプルごとの信頼パラメーターによって制御される、観察されたラベルとモデル独自の予測分布の凸状の組み合わせを使用してトレーニングすることです。この目標の勾配により、トレーニングの初期段階でクリーンなサンプルとノイズの多いサンプルで信頼値が逆方向に駆動され、分離可能な信頼分布が得られることを示します。信頼値に対して 3 成分ガウス混合モデルを使用して、信頼値をクリーンなケース、あいまいなケース、ノイズの多いケースに分離し、ノイズの多いケースに対して短いソフト修正フェーズを実行し、最後のハード修正フェーズを実行します。最大 50% のラベル ノイズの下で MedMNISTv2 からの 10 個の 2D データセットを実験したところ、一貫して精度が向上し、強力な誤ラベル検出が示されました。 LiNC は無視できる漸近的なオーバーヘッドを追加します。トレーニング時間の複雑さは引き続きベース ネットワークによって支配され、追加のメモリはトレーニング セットのサイズに応じて直線的に増加します。
原文 (English)
LiNC: Lightweight Noise Correction via Per-Sample Trust and Gaussian Mixture Modeling
Label noise is common in medical imaging datasets due to factors such as inter-rater variability, annotation errors, and ambiguous cases. This can severely undermine the reliability and clinical effectiveness of machine learning models trained using those datasets. To address this challenge, we introduce Lightweight Noise Correction (LiNC), which adds a single trainable trust parameter per training sample and learns when to use the observed label and when to defer to the model during a standard training loop. The key idea is to train using a convex combination of the observed label and the model's own predictive distribution, controlled by a per-sample trust parameter. We show that the gradient of this objective drives trust values in opposite directions for clean versus noisy samples in the early training phase, yielding separable trust distributions. We use a 3-component Gaussian Mixture Model over the trust values to separate them into clean, ambiguous, and noisy cases and then execute a short soft-correction phase on the noisy cases and a final hard correction phase. Experiments on ten 2D datasets from MedMNISTv2 under label noise of up to 50% show consistent gains in accuracy and strong mislabel detection. LiNC adds negligible asymptotic overhead: the training-time complexity remains dominated by the base network, with additional memory growing linearly with the size of the training set.
AgentForge: An Immersive Role-Playing Platform for Learning Agentic Software Engineering
Agentic AI is increasingly used to coordinate planning, implementation, review, and testing in software development, yet it often offers li…
TRNet: マルチモーダル水稲セグメンテーションのための地形に基づく周波数修正と構造認識デコーディング
山岳地帯や丘陵地帯の非常に高解像度の画像から水稲をマッピングすることは、地形によって光学的な外観が変化し、視覚的に類似した植生との混乱が増大するため、困難です。 0.5 m GaoJing-1 赤-緑-青 (RGB) 画像、5 m TanDEM-X デジタル標高モデル (DEM)、および派生した傾斜の TRNet を紹介します。個別のビジュアルエンコーダーと地形エンコーダーにより、モダリティ固有の機能が保持されます。エンコーダの初期段階で、地形エネルギースペクトル整流は、地形条件付き低周波変調と非対称高周波調整を適用して、急斜面のクラッターを抑制し、互換性のある低斜面の米のキューを条件付きで強化します。地形ガイド型水田構造デコーダは、粗い地形をコンテキストとして使用して、セマンティック、米と背景の境界、および内部の手がかりを組み合わせます。実験では、エリア A の内部テスト セットと、より急峻な地形とコメの普及率の低いエリア B を使用しました。 TRNet は、85.10\% および 80.68\% というライス交差オーバーユニオン (IoU) 値を達成し、元のデュアル エンコーダ U-Net をそれぞれ 9.15 および 18.83 パーセント ポイント上回りました。アブレーションとスロープ層別の結果は、これらのゲインを周波数調整、構造学習、および急峻な地形の誤検知の減少に結び付けました。この結果は、非常に高解像度の水稲マッピングのための事前文脈として、粗い地形を裏付けています。
原文 (English)
TRNet: Topography-Guided Frequency Rectification and Structure-Aware Decoding for Multimodal Paddy Rice Segmentation
Mapping paddy rice from very-high-resolution imagery in mountainous and hilly regions is difficult because terrain alters optical appearance and increases confusion with visually similar vegetation. We present TRNet for 0.5-m GaoJing-1 red--green--blue (RGB) imagery, a 5-m TanDEM-X digital elevation model (DEM), and derived slope. Separate visual and terrain encoders preserve modality-specific features. At an early encoder stage, Topographic Energy-Spectral Rectification applies terrain-conditioned low-frequency modulation and asymmetric high-frequency regulation to suppress steep-slope clutter and conditionally enhance compatible low-slope rice cues. The Topography-guided Paddy Structure Decoder combines semantic, rice--background boundary, and interior cues, using coarse terrain as context. Experiments used an Area A internal test set and held-out Area B, which had steeper terrain and lower rice prevalence. TRNet achieved rice intersection-over-union (IoU) values of 85.10\% and 80.68\%, exceeding the original Dual-Encoder U-Net by 9.15 and 18.83 percentage points, respectively. Ablation and slope-stratified results linked these gains to frequency rectification, structure learning, and fewer steep-terrain false positives. The results support coarse topography as a contextual prior for very-high-resolution paddy rice mapping.
Visualizing Graph-to-Answer Mechanism Recovery in Materials-Science Hypothesis Generation
AI co-scientists can generate fluent materials-science hypotheses, but fluency does not show that an answer preserves a scientifically mean…
行動スキルの再構築: LLM エージェント スキルから隠れた機能を再構築する
クローズド ソース エージェントのスキルは、独自の命令、スクリプト、定数、およびデータをエンコードする場合があります。プロバイダーは、基礎となるパッケージを非表示にしたまま、その機能をサービスとして提供する場合があります。これまでの研究は、これらのアーティファクトを直接公開する即時インジェクション攻撃に焦点を当てており、それに応じて既存の防御はそのような漏洩を防ぐことを目的としています。ただし、ファイルの開示を防止しても、ユーザーがそのファイルに実装されている機能を回復することはできます。これにより、基本的な疑問が生じます。ファイルが隠されたままの状態で、ユーザーは通常の使用を通じてスキルの機能を再構築できるのでしょうか?私たちは、攻撃者が有効なタスク要求と観察された応答を使用して、隠れたスキルの機能的なクローンを構築する行動スキル再構成 (BSR) を研究しています。 SkillClone は、パブリック アドバタイズメントからインターフェイス仮説を形成し、構造化された無害なプローブを発行し、実行可能なレプリカを合成し、被害者のスキルに対する差分検証を通じて反復的に修復することにより、ターゲット スキルのクローンを作成するブラック ボックス攻撃です。 SkillClone は、ルール、テーブル、プロシージャ、アルゴリズムにわたる 30 のスキルにわたって、複数のターゲットの保留された入力の正確な回復または部分的な回復を実現します。反復的な再クエリにより、単一ラウンドの再構築で見逃されたギャップが埋められます。 SkillClone は正当なインタラクションのみを使用するため、開示に重点を置いた防御では対応範囲が限られ、スキルの説明があまり詳細でない場合は保護が限定されます。これらの結果は、ファイルの機密性だけでは機能の機密性が保証されないことを示しています。防御では、通常の使用による累積的な情報漏洩も制限する必要があります。
原文 (English)
Behavioral Skill Reconstruction: Reconstructing Hidden Functionality from LLM Agent Skills
Closed source agent skills may encode proprietary instructions, scripts, constants, and data. Providers may offer their capabilities as services while keeping the underlying packages hidden. Prior work focuses on prompt injection attacks that directly disclose these artifacts, and existing defenses accordingly aim to prevent such leakage. However, preventing file disclosure does not prevent users from recovering the functionality those files implement. This raises a fundamental question: can a user reconstruct a skill's functionality through ordinary use while its files remain hidden? We study behavioral skill reconstruction (BSR), in which an attacker uses valid task requests and observed responses to build a functional clone of a hidden skill. We introduce SkillClone, a black-box attack that clones a target skill by forming an interface hypothesis from its public advertisement, issuing structured benign probes, synthesizing an executable replica, and iteratively repairing it through differential validation against the victim skill. Across 30 skills spanning rules, tables, procedures, and algorithms, SkillClone achieves exact or partial recovery on held-out inputs for several targets. Iterative requerying closes gaps missed by single-round reconstruction. Because SkillClone uses only legitimate interactions, disclosure-focused defenses provide limited coverage, and less detailed skill descriptions offer limited protection. These results show that file secrecy alone does not ensure functional secrecy. Defenses must also limit cumulative information leakage from ordinary use.
患者のような私: 説明可能な臨床予測のための変分 LM--GNN フレームワーク
言語モデル (LM) は、電子医療記録 (EHR) に強力なテキスト表現を提供しますが、患者シーケンスを個別にエンコードし、説明可能性が限られています。グラフ ニューラル ネットワーク (GNN) は、患者間の関係を組み込み、参照患者の帰属を可能にすることで LM を補完しますが、高品質の患者表現に依存しています。我々は、ローカルな患者のセマンティクスとグローバルなコホート構造を統合する統合 LM-GNN フレームワークである、Patients-like-me (PLM) を提案します。 PLM を効率的にトレーニングするために、教師付き変分目標の下で LM と GNN の更新を交互に行う変分期待値最大化アルゴリズムを導入します。 MIMIC-III および MIMIC-IV に関する広範な実験により、PLM が常に最先端の方法より優れたパフォーマンスを示し、エンコーダのみおよびデコーダのみの LM バックボーン全体で改善が一般化していることが示されています。これらの利点は、わずかな追加の計算オーバーヘッドのみで達成されます。また、PLM は、影響力のある同様の患者を取得することで参照患者の説明を提供する一方、エッジマスキング実験により、最高ランクの参照がモデル予測に最大の影響を与えることを確認します。
原文 (English)
Patients-like-me: A Variational LM--GNN Framework for Explainable Clinical Prediction
Language models (LMs) offer strong textual representations for electronic health records (EHRs), but they encode patient sequences in isolation and provide limited explainability. Graph neural networks (GNNs) complement LMs by incorporating inter-patient relationships and enabling reference-patient attribution, yet they rely on high-quality patient representations. We propose Patients-like-me (PLM), a unified LM--GNN framework that integrates local patient semantics with global cohort structure. To train PLM efficiently, we introduce a Variational Expectation-Maximization algorithm that alternates LM and GNN updates under a supervised variational objective. Extensive experiments on MIMIC-III and MIMIC-IV show that PLM consistently outperforms state-of-the-art methods, with improvements generalizing across encoder-only and decoder-only LM backbones. These gains are achieved with only modest additional computational overhead. PLM also provides reference-patient explanations by retrieving influential similar patients, while edge-masking experiments confirm that the highest-ranked references have the greatest impact on model predictions.
モデルの結合によるクロスドメイン クローン検出の統合モデル
構文コピーから言語間セマンティック クローン、AI 生成の重複に至るまで、コード クローン タイプの多様性が増大しているため、クローン検出における断片化の危機が生じています。現在の深層学習検出器はドメイン スペシャリストであり、トレーニング分布外では大幅にパフォーマンスが低下し、ドメイン全体で F1 の低下が 70% を超えています。複数の特殊なモデルをデプロイすることは現実的ではありませんが、単一のクロスドメイン検出器をトレーニングするには、すべてのトレーニング データに同時にアクセスする必要があります。これに対処するために、トレーニングされたチェックポイントのみで動作するポストホック手法のファミリーであるモデルのマージを調査します。 5 つのタスク ベクター メソッドによるパラメーターのマージ、貪欲なレイヤー ステッチングによるアーキテクチャのマージ、および 4 つのコード モデル、3 つのベンチマーク、および 12 の構成にわたるクロストークナイザー アラインメントを評価します。同一ベース TIES マージにより、効果的なクロスドメイン検出器が作成され、2 つのモデル ファミリと 3 つのランダム シードにわたって検証され、UniXcoder で結合 F1 が 0.865 に達し、マージ ステップでトレーニング データを使用しなくてもマルチタスク パフォーマンスの 93% に達します。 WUDI は、分布内結合 F1 として最高の 0.899 を達成していますが、TIES は、AI によって生成された未確認のクローンをより適切に一般化するため、推奨される方法となっています。クロスベース マージでは、5 つの方法すべてでわずかな高い分散ゲインのみが得られます。これは、共有の事前トレーニングされたベースを介したタスク ベクトルの互換性が効果的なマージの結合因子であることを示しています。また、マージされた検出器は、低い推論コストで GPTCloneBench 上のゼロショット コード LLM よりも優れたパフォーマンスを示し、目に見えない AI 生成クローンに対するマルチタスク トレーニングよりも最大 4 倍優れた一般化を実現します。これは、ドメイン内のパフォーマンスと OOD の堅牢性の間のトレードオフを示唆しています。この研究は、ソフトウェア エンジニアリングのためのモデル マージに関する最初の体系的な実証研究の 1 つと、クロスドメイン クローン検出器を構築するための実践的なレシピを提供します。
原文 (English)
A Unified Model for Cross-Domain Clone Detection via Model Merging
The growing diversity of code clone types, from syntactic copies to cross-language semantic clones to AI-generated duplicates, has created a fragmentation crisis in clone detection. Current deep learning detectors are domain specialists that degrade significantly outside their training distribution, with F1 drops exceeding 70% across domains. Deploying multiple specialized models is impractical, yet training a single cross-domain detector requires simultaneous access to all training data. To address this, we investigate model merging, a family of post-hoc techniques that operate solely on trained checkpoints. We evaluate parameter merging with five task-vector methods, architecture merging via greedy layer stitching, and cross-tokenizer alignment across four code models, three benchmarks, and twelve configurations. Same-base TIES merging creates effective cross-domain detectors, validated across two model families and three random seeds, reaching 0.865 combined F1 on UniXcoder, 93% of multi-task performance without any training data at the merging step. WUDI achieves the highest in-distribution combined F1 at 0.899, but TIES generalizes better to unseen AI-generated clones, making it our recommended method. Cross-base merging yields only marginal and high-variance gains across all five methods, indicating that task vector compatibility through a shared pre-trained base is the binding factor for effective merging. Merged detectors also outperform zero-shot code LLMs on GPTCloneBench at lower inference cost and generalize up to 4x better than multi-task training to unseen AI-generated clones, suggesting a trade-off between in-domain performance and OOD robustness. This work provides one of the first systematic empirical studies of model merging for software engineering and a practical recipe for building cross-domain clone detectors.
Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary
Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true witho…
サイバーフィジカルシステムにおけるLLMエージェントの計画戦略の戦略的評価
LLM 計画エージェントの評価では、タスクが成功するか、宣言された計画が実行されるかが主に問われます。戦略的なサイバー物理システムでは、自律的な参加者が応答し、物理学が結果を制約した後も、計画アーキテクチャが適切なままであるかどうかがより強力な問題となります。計画に起因する制御軌道、つまり実行アーキテクチャが他のエージェントや物理プロセスに作用する順序付けされた計画操作と指令を中心に構築された、制御された物理に基づいたベンチマークを導入します。 40 の異種プロシューマーと独立してシミュレートされたラジアル フィーダーを備えたスマート グリッド デマンド レスポンス システムに、事前定義された順次的、階層的、および検索エグゼキュータが実装されています。 LLM は、型指定されたポリシー宣言と短いオペレーター メッセージに制限されていますが、スケジュール構築、プロシューマー ダイナミクス、および電力の流れは明示的なコードのままです。このプロトコルは、ペアになった強制モードの反事実、一般的なランダムな応答の抽出、およびイベントレベルの期限の実現可能性を使用します。 3 つのプロパティが続きます。アーキテクチャは結果を大きく変えます。強制検索は 5 つのベースライン シードすべてにおける神託です。実行忠実度にはモード一致以上のものが必要です。目的置換は電圧不足を 2.68 倍に増加させながら 1.0 で一致を保持します。 144 のシナリオ、576 エピソードのバンクには、4 つのアーキテクチャのうち 3 つからの実行可能なオラクルがあります。事前に指定された応力保持リッジの平均リグレスは 90.7 (95% 間隔 [73.8, 108.6]) で、固定シーケンシャルでは検出可能な値はありません。品質予測の前に既知の期限の実現可能性を適用すると、後悔は 29.0 に減少し、固定シーケンシャルよりも 61.1 改善されます。あらゆる実現可能なアブレーションは固定検索に勝るものではなく、残りの課題は実現可能な品質の選択に限定されます。 5 つのモデル拡張により、ストレス条件付き宣言子、状態ブラインド宣言子、および不変宣言子が分離されます。レイテンシーテールは、実際の実現可能性が確率的に扱われる必要があることを示しています。
原文 (English)
Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems
Evaluations of LLM planning agents largely ask whether a task succeeds or a declared plan is followed. In strategic cyber-physical systems, a stronger question is whether the planning architecture remains appropriate after autonomous participants respond and physics constrains the outcome. We introduce a controlled, physics-grounded benchmark built around planning-induced control trajectories: the ordered planning operations and directives through which an execution architecture acts on other agents and the physical process. It implements predefined, sequential, hierarchical, and search executors in a smart-grid demand-response system with 40 heterogeneous prosumers and an independently simulated radial feeder. The LLM is bounded to typed policy declaration and short operator messages, while schedule construction, prosumer dynamics, and power flow remain explicit code. The protocol uses paired forced-mode counterfactuals, common random response draws, and event-level deadline feasibility. Three properties follow. Architecture materially changes outcomes: forced search is the oracle in all five baseline seeds. Execution fidelity needs more than mode agreement: objective substitution holds agreement at 1.0 while increasing voltage shortfall by 2.68x. A 144-scenario, 576-episode bank has feasible oracles from three of the four architectures. A prespecified stress-held-out ridge has mean regret 90.7 (95% interval [73.8, 108.6]) and no detectable value over fixed sequential; applying known deadline feasibility before quality prediction cuts regret to 29.0 and improves over fixed sequential by 61.1. An all-feasible ablation does not beat fixed search, localising the remaining challenge to within-feasible quality selection. A five-model extension separates stress-conditioned, state-blind, and invariant declarers; latency tails show that live feasibility should be treated probabilistically.
コンパス: 現場の反射を介してソーシャル メディア フィードを継続的に調整する
ソーシャル メディアの推奨フィードは、多くの場合、ユーザーが深く考えた後に保持する好みではなく、ユーザーの即時の衝動に合わせて最適化されます。一部のシステムは、単なる動作信号ではなく、構成ページまたはインフィード コントロールを介してユーザーの明示的な設定を組み込むことで、この不整合に対処します。ただし、通常、ユーザーの好みは進化しており、ユーザーが述べた好みや行動は自然に発散するため、継続的な反映とフィードの再調整が必要になります。しかし、既存の戦略ではユーザーが主導権を握る必要があり、多くの場合、労力がかかります。その結果、実際にはこれらが呼び出されることはほとんどありません。私たちは、ユーザーが自分の行動を考慮して自分の好みを反映し、明確に表現できるようにすることで、ユーザーのフィードをユーザーの反射的な好みに合わせるシステムを紹介します。毎日のブラウジング中に継続的な反映を可能にするために、Compass は軽量通知を介してその場での反映を表示し、行動信号を定期的にシミュレートし、フィード コンテンツを直接操作することでフィードの調整を実現します。 YouTube ショートに Compass を埋め込み、10 日間の現地調査を通じて継続的なサポートなしのベースラインと比較しました (N=15)。 Compass は、フィード閲覧のカジュアルな性質を損なうことなく、より思慮深く目的を持ったフィード消費、反復的な好みの調整、およびより強力なフィード調整を促進することがわかりました。
原文 (English)
Compass: Continuously Aligning Social Media Feeds via In-Situ Reflections
Social media recommendation feeds often optimize for users' immediate impulses rather than preferences they would hold after deeper reflection. Some systems address this misalignment by incorporating users' explicit preferences via a configuration page or in-feed controls instead of just behavioral signals. However, users typically have evolving preferences, and their stated preferences and behavior naturally diverge, necessitating continuous reflection and feed realignment. But existing strategies require the user to take initiative and are often effortful; as a result, in practice they are rarely invoked. We present Compass, a system that aligns a user's feed with their reflective preferences by helping users reflect on and articulate their preferences given their behavior. To enable continuous reflection during everyday browsing, Compass surfaces in-situ reflections via lightweight notifications, while feed alignment is achieved by periodically simulating behavioral signals and directly manipulating feed content. We embedded Compass within YouTube Shorts and compared it against a baseline without continuous support through a 10-day field study (N=15). We found that Compass promoted more reflective and purposeful feed consumption, iterative preference adjustment, and stronger feed alignment, without sacrificing the casual nature of feed browsing.
EA-Graph: アップストリーム ドリフト下のコーディング エージェント用のアーティファクト アンカー検証メモリ
コーディングエージェントはセッションをまたいで動作することが増えていますが、散文ノートは、それをサポートするプログラムの状態がなくても結論を保持できます。アップストリームの変更後、以前の検証クレームが有効でなくなっても、リポジトリは引き続き構築される場合があります。 EA-Graph は、検証クレーム用のアーティファクトに固定されたメモリです。サブパス粒度でアーティファクトを表し、リーフ定義へのエイリアスを解決し、各クレームを確立するために使用されたコンテンツに固定し、証拠の強度を鮮度から切り離して保ちます。代替コンテンツが利用できない場合、その主張は推測ではなく証明不可能になります。 EA-Graph は、生成されたリポジトリで評価されます。リポジトリの動作からアーティファクトまでのグラウンド トゥルースは、構築によって既知です。課題は、以前の主張を、値のドリフト、ロジックのドリフト、および意図的に保留された上流のコンテンツの後、影響を受けていない、影響を受けている、または証明できないとして分類することです。この分析は、7 つのクリーン ワールド、14 のモデルワールド インスタンス、3 つのメモリ条件、および 2 つのモデル層にわたる 42 のセッションを対象としています。俳句ラウンドでは、7 つの世界すべてにおいて、アーティファクトに固定された記憶が散文ノートと永続的な記憶を上回りました。それぞれの正確な対ウィルコクソン比較では、p = 0.0156 が得られました。ソネットラウンドでは、アンカー状態は完璧でしたが、頻繁なコントロールの天井により、事前に登録されたコントラストは有意ではなくなりました。セッションで捏造された保留されたコンテンツはありません。これらの結果は、アーティファクトに固定されたメモリにより、このテストベッドにおける小規模モデルの証明可能性の判断が向上したという、限定された主張を裏付けています。さらに探索的な比較では、構造化クレーム メモリがセッション内の再導出を外部化することで機能ギャップを狭める可能性があることを示唆していますが、モデル間の同等性は確立されていません。この研究では、効率や修理の品質については主張していません。
原文 (English)
EA-Graph: Artifact-Anchored Verification Memory for Coding Agents under Upstream Drift
Coding agents increasingly work across sessions, but prose notes can preserve a conclusion without the program state that supported it. After an upstream change, a repository may still build even though earlier verification claims are no longer valid. EA-Graph is an artifact-anchored memory for verification claims. It represents artifacts at sub-path granularity, resolves aliases to leaf definitions, anchors each claim to the content used to establish it, and keeps evidence strength separate from freshness. When replacement content is unavailable, the claim becomes unprovable rather than guessed. EA-Graph is evaluated on generated repositories whose behavior-to-artifact ground truth is known by construction. The task is to classify prior claims as unaffected, affected, or unprovable after value drift, logic drift, and deliberately withheld upstream content. The analysis covers 42 sessions across seven clean worlds, 14 model-world instances, three memory conditions, and two model tiers. In the Haiku round, artifact-anchored memory outscored prose notes and no persistent memory in all seven worlds; each exact paired Wilcoxon comparison yielded p = 0.0156. In the Sonnet round, the anchored condition was perfect, but frequent control ceilings left the preregistered contrasts non-significant. No session fabricated withheld content. These results support a bounded claim: artifact-anchored memory improved the smaller model's provability judgments in this testbed. An exploratory comparison further suggests that structured claim memory may narrow a capability gap by externalizing in-session re-derivation, but it does not establish cross- model equivalence. The study makes no claim about efficiency or repair quality.
MIDAS: マルチ LLM 反復データ適応型要約
文章の要約は一見難しいです。情報を要約するのは簡単そうに見えますが、実際の企業でサポート チケット、法的文書、インシデント レポートなどを要約するには、ドメイン固有のガイドライン、出力形式、組織の規則に厳密に従う必要があります。これらの制約を確実に満たすプロンプトを作成するには多大な労力がかかり、要件の進化に応じて人間の多大な専門知識と継続的なメンテナンスが必要になります。既存の自動プロンプト最適化手法は、大規模言語モデル (LLM) の批評主導の改良を通じてこの負担を軽減しますが、要約アプリケーションの多様性に適応できない静的なプロンプトによる制限が依然として残ります。私たちは、データ駆動型のパターン学習とユースケース固有のパーソナライゼーションによってこのパラダイムを拡張するマルチ LLM フレームワークである Multi-LLM Iterative Data-Adaptive Summarization (MIDAS) を提案します。これにより、手動によるプロンプト エンジニアリングを行わずに、さまざまな要約要件に自動的に適応できるようになります。 5 つの出力形式にわたる企業顧客のチケット要約に適用された MIDAS は、CriSPO や ZERA などの最先端の批評主導型最適化フレームワークに対して最も強力な全体パフォーマンスを達成し、ROUGE-1 で最大 11.0%、ROUGE-2 で最大 18.2%、ROUGE-L で最大 8.0% 向上し、すべての形式と出力タイプで BERTScore F1 を一貫して向上させます。さらに、マルチ LLM 構成と財務ドメインの要約ベンチマークを通じて、クロスモデルおよびクロスドメインの一般化を実証します。
原文 (English)
MIDAS: Multi-LLM Iterative Data-Adaptive Summarization
Text summarization is deceptively difficult. While condensing information seems straightforward, real-world enterprise summarization of support tickets, legal documents, incident reports, and more, demands strict adherence to domain-specific guidelines, output formats, and organizational conventions. Crafting prompts that reliably satisfy these constraints is labor-intensive, requiring significant human expertise and continuous maintenance as requirements evolve. Existing automated prompt optimization methods reduce this burden through Large Language Model (LLM) critique-driven refinement, yet remain limited by static prompts that cannot adapt to the diversity of summary applications. We propose Multi-LLM Iterative Data-Adaptive Summarization (MIDAS), a multi-LLM framework that extends this paradigm with data-driven pattern learning and use-case-specific personalization, enabling automatic adaptation to different summarization requirements without manual prompt engineering. Applied to enterprise customer ticket summarization across five output formats, MIDAS achieves the strongest overall performance against state-of-the-art critique-driven optimization frameworks such as CriSPO and ZERA, improving ROUGE-1 by up to 11.0%, ROUGE-2 by up to 18.2%, and ROUGE-L by up to 8.0%, while consistently improving BERTScore F1 across all formats and output types. We additionally demonstrate cross-model and cross-domain generalization through multi-LLM configurations and finance-domain summarization benchmarks.
Trident : 深層強化学習によるサイバー防御を突破する方法 (Agentic)
深層強化学習 (DRL) に基づく自律型サイバー防御システムは、研究で大きな注目を集めていますが、依然として静的でヒューリスティックなレッド エージェントに対してのみほとんど評価されており、適応型の脅威に対する堅牢性は十分に研究されていません。一方、検証可能な報酬を伴う強化学習(RLVR)の最近の進歩により、LLM 推論は改善されましたが、適切なベンチマーク環境とインタラクション データセットが存在しないため、サイバーセキュリティへの統合は依然として困難です。このギャップを埋めるために、Trident を導入します。Trident は、CybORG CAGE 4 と CyberWheel にまたがる分離されたサンドボックス サーバーを備えた動的ベンチマーク、RLVR の 13,000 を超える高忠実度の赤と青の相互作用軌跡で構成されるデータセット、および「Code-as-Policy」RLVR エージェント アーキテクチャである Trident Agentic の 3 つのコンポーネントで構成されるエージェント LLM レッド チーミング フレームワークです。後者は、Log Summarizer、Planner、Coder の 3 者構成設計により、コンテキスト バンディットとしてレッド エージェントのトレーニングを再定式化します。トレーニング可能な Planner が、圧縮された実行ログから完全な攻撃戦略を生成し、凍結された Coder が、ライブ DRL 防御者に対して展開される実行可能な Python ポリシーに変換します。経験的評価により、既存の防御の根本的な脆弱性が明らかになりました。単一のトレーニング可能な 7B プランナーを使用すると、Trident は、静的な赤色エージェントのベースラインと比較して青色エージェントの防御パフォーマンスを平均 522% 低下させながら、静的ヒューリスティックでは完全に発見できないおとり回避や適応状態優先順位付けなどの緊急動作を自律的に発見します。
原文 (English)
Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)
Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adaptive threats critically understudied. Meanwhile, recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have improved LLM reasoning, but their integration into cybersecurity remains elusive due to the absence of suitable benchmark environments and interaction datasets. To bridge this gap, we introduce Trident, an agentic LLM red teaming framework comprising three components: a dynamic benchmark with isolated sandbox servers spanning CybORG CAGE 4 and CyberWheel, a dataset comprises over 13,000 high-fidelity red-blue interaction trajectories for RLVR, and a ``Code-as-Policy'' RLVR agentic architecture Trident Agentic). The latter reformulates red agent training as a contextual bandit via a tripartite Log Summarizer--Planner--Coder design, where a trainable Planner generates complete attack strategies from compressed execution logs, which a frozen Coder translates into executable Python policies deployed against live DRL defenders. Empirical evaluations reveal a fundamental brittleness in existing defenses: with a single trainable 7B planner, Trident reduces blue agent defensive performance by an average of 522% compared to static red agent baselines while autonomously discovering emergent behaviors such as decoy avoidance and adaptive state prioritization that static heuristics entirely fail to uncover.
Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits
This paper studies generalized low-rank matrix bandits with multiple prioritized objectives. At each round, the learner selects a matrix-va…
ATLAS: 継続学習のための抽象後継者を使用した適応トポロジカル学習
最新のモデルフリー強化学習アルゴリズムは、非常に高いパフォーマンスを達成できますが、サンプル効率が低く、環境の変化に対して堅牢ではありません。モデルベースのアルゴリズムはサンプル効率がはるかに高くなりますが、環境が変化すると失敗します。このペーパーでは、これらの課題に対処するために、Adaptive Topological Learning with Abstract Successors (ATLAS) を紹介します。 ATLAS は、高いサンプル効率を達成しながら壊滅的な忘却にもしっかりと取り組むために、Successor 機能を備えた Grow When Required ネットワークを使用します。私たちは空間ナビゲーション タスクで ATLAS を評価し、一般的なオンポリシー アルゴリズムとオフポリシー アルゴリズムに対してそのパフォーマンスをベンチマークします。私たちの経験的結果は、ATLAS が報酬信号から遷移ダイナミクスを構造的に切り離すことにより、新しい目標へのほぼ瞬時の適応を達成し、正の逆方向伝達を示し、非定常環境におけるベースライン手法を大幅に上回るパフォーマンスを示すことができることを示しています。
原文 (English)
ATLAS: Adaptive Topological Learning with Abstract Successors for Continual Learning
Contemporary model-free reinforcement learning algorithms can achieve very high performance, but have low sample efficiency and are not robust to changes in the environment. Model-based algorithms have much higher sample efficiency, but still fail when the environment shifts. This paper introduces Adaptive Topological Learning with Abstract Successors (ATLAS) to combat these challenges. ATLAS uses a Grow When Required network with Successor Features in order to achieve high sample efficiency while also robustly tackling catastrophic forgetting. We evaluate ATLAS in spatial navigation tasks, benchmarking its performance against common on-policy and off-policy algorithms. Our empirical results demonstrate that by structurally decoupling transition dynamics from the reward signal, ATLAS achieves near-instantaneous adaptation to new goals and can exhibit positive backward transfer, significantly outperforming baseline methods in non-stationary environments.
COMPAS: コード生成を最適化するための難易度を意識した共同探索
コード生成システムは、モデル、プロンプト、およびデコード設定を使用して各 LLM 呼び出しを行います。ただし、既存の最適化手法は通常、これらの選択肢の一部のみを調整するか、すべてのタスクに対して 1 つの固定構成を使用します。グローバル オプティマイザーはすべてのタスクに対して 1 つの構成を検索し、ルーターはモデルのみを選択し、プロンプト オプティマイザーはモデルとデコード設定を固定したままにします。このため、彼らの共同のグループ固有の相互作用は不明瞭なままになります。したがって、これらの選択がどのように相互作用するかを調査し、プロンプトとデコード設定が相互作用し、チューニング効果がモデルによって異なり、最適な構成がタスクの難易度によって異なることを観察します。これらの観察に基づいて、低コストのモデル選択とプロンプトの共同デコード検索を通じてグループ固有の品質コスト フロントを学習し、それ以上の検索を行わずに各テスト タスクをオンラインで一致するフロントにルーティングする、困難を認識した手法である COMPAS (モデル、プロンプト、およびデコード設定に対するコード生成の最適化) を導入します。 LiveCodeBench の一致した検索予算の下で、COMPAS は pass@1 を最良のベースラインの 45.9% から 52.8% に改善し、コストを 36.57 ドルから 4.92 ドルに削減します。これにより、SWE ベンチでのリポジトリ レベルのコード生成にも移行し、タスクの 76.0% が解決されるのに対し、最良のベースラインでは 70.0% が解決されます。コードと再現成果物は https://github.com/gjz78910/COMPAS で入手できます。
原文 (English)
COMPAS: Difficulty-Aware Joint Search for Optimizing Code Generation
Code generation systems make each LLM call with a model, a prompt, and decoding settings. However, existing optimization methods usually tune only part of these choices or use one fixed configuration for all tasks: global optimizers search one configuration for all tasks, routers choose only a model, and prompt optimizers keep the model and decoding settings fixed. This leaves their joint, group-specific interactions unclear. We therefore examine how these choices interact and observe that prompts and decoding settings interact, tuning effects vary by model, and the best configuration varies by task difficulty. Guided by these observations, we introduce COMPAS (Code-generation Optimization over Models, Prompts, And Decoding Settings), a difficulty-aware method that learns group-specific quality-cost fronts through low-cost model selection and joint prompt-decoding search, then routes each test task to its matching front online without further search. Under a matched search budget on LiveCodeBench, COMPAS improves pass@1 from 45.9% for the best baseline to 52.8% while reducing cost from $36.57 to $4.92. This also transfers to repository-level code generation on SWE-bench, resolving 76.0% of tasks versus 70.0% for the best baseline. Code and the reproducibility artifact are available at https://github.com/gjz78910/COMPAS.
Equitable System-Prompt Selection via Constrained Mixed-Strategy GroupDRO
Large language models are increasingly used for information seeking, yet semantically equivalent questions phrased in different ways can re…
iStructTab: 画像および表形式データのマルチモーダル学習のための構造化特徴シーケンス
画像や表形式データのマルチモーダル学習は、非効率な表現によって損なわれることが多く、その結果、冗長性、分散性、一般化の問題が発生します。この課題に取り組むために、列置換問題 (CPP) の原理に基づいた構造化特徴シーケンス アルゴリズムであるグラフ拡張記述子シーケンス (GEDS) を導入します。 GEDS は、類似性グラフベースの計算を通じて特徴の統計的記述子を洗練し、効果的な特徴の順序付けを系統的に決定します。専用の損失関数を介して派生特徴シーケンスに明示的に準拠する順序認識メモリ トークンを利用して、順序認識の効率的な変換フレームワーク内に GEDS を組み込みます。マルチモーダル ベンチマークにわたる実験結果は、iStructTab が特徴の分散を効果的に最小限に抑え、予測パフォーマンスとロバスト性を向上させ、マルチモーダル学習における構造化特徴シーケンスの重要性を強調していることを示しています。
原文 (English)
iStructTab: Structured Feature Sequencing for Multimodal Learning of Image and Tabular Data
Multimodal learning of images and tabular data is often impaired by ineffective representations, resulting in redundancy, dispersion, and generalization problems. To tackle this challenge, we introduce Graph-Enhanced Descriptor Sequencing (GEDS), a structured feature sequencing algorithm grounded in principles from the Column Permutation Problem (CPP). GEDS refines statistical descriptors of the features through similarity graph-based computations, systematically determining an effective feature sequencing. We incorporate GEDS within an order-aware efficient transformer framework, utilizing order-aware memory tokens that explicitly adhere to the derived feature sequencing via a dedicated loss function. Experimental results across multimodal benchmarks demonstrate that iStructTab effectively minimizes feature dispersion, improving predictive performance and robustness, and highlighting the significance of structured feature sequencing in multimodal learning.
HyPASE: 大規模な音声言語モデル向けのパラメータ効率の高い音声感情微調整フレームワークのための双曲幾何学
大規模音声言語モデル (LALM) は、一般的な音声理解に優れています。ただし、音声感情認識 (SER) などのきめ細かいタスクにそれらを適応させることは、依然として大きなボトルネックとなっています。現在のパラメータ効率のよい微調整 (PEFT) 手法は通常、平坦なユークリッド空間で動作しますが、この幾何学形状では、低レベルの韻律から高レベルの意味論に至る、感情の手がかりの多粒度の性質を捉えることができません。これに対処するために、LALM ベースの SER 用の双曲線 PEFT フレームワークである HyPASE を提案します。 HyPASE はポアンカレ ボール モデルを利用し、表現の粒度の明示的な代用として双曲半径を使用します。このフレームワークは 2 つのコア コンポーネントで構成されています。レイヤー適応重み変調用の双曲線幾何学アダプター (HGA) と、マルチスケール機能をコンパクトなオーディオ プレフィックスに圧縮する感情認識マルチキャパシティー クロスモーダル アグリゲーター (EMCA) です。標準ベンチマークの実験結果は、HyPASE が MELD のすべてのメトリクスにわたってユークリッド PEFT ベースラインを上回り、特にクラス不均衡感情認識において、IEMOCAP で顕著な非重み付き精度の向上を達成することを示しています。これには、少数派クラス表現の双曲空間の幾何学的優先順位を反映するわずかな重み付き精度のトレードオフが伴います。さらに、HyPASE は、制約されたパラメータ予算内で堅牢なゼロショットのクロスデータセット一般化を実現します。 HyPASE は、適応プロセスを双曲幾何学に基づいて行うことにより、LALM 微調整のための非常に効率的なパスを提供します。
原文 (English)
HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models
Large Audio-Language Models (LALMs) excel at general speech understanding; however, adapting them to fine-grained tasks like Speech Emotion Recognition (SER) remains a significant bottleneck. Current Parameter-Efficient Fine-Tuning (PEFT) methods typically operate in flat Euclidean space, and this geometry fails to capture the multi-granularity nature of emotion cues, which range from low-level prosody to high-level semantics. To address this, we propose HyPASE, a hyperbolic PEFT framework for LALM-based SER. HyPASE leverages the Poincare ball model, using the hyperbolic radius as an explicit proxy for representational granularity. The framework consists of two core components: a Hyperbolic Geometric Adapter (HGA) for layer-adaptive weight modulation, and an Emotion-aware Multi-capacity Cross-modal Aggregator (EMCA) that compresses multi-scale features into compact audio prefixes. Empirical results on standard benchmarks show that HyPASE outperforms Euclidean PEFT baselines across all metrics on MELD and achieves a notable Unweighted Accuracy gain on IEMOCAP, particularly in class-imbalanced emotion recognition, with the accompanying slight Weighted Accuracy trade-off reflecting hyperbolic space's geometric prioritization of minority-class representations; furthermore, HyPASE achieves robust zero-shot cross-dataset generalization within a constrained parameter budget. By grounding the adaptation process in hyperbolic geometry, HyPASE offers a highly efficient path for LALM fine-tuning.
Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework
While retrieval-augmented generation systems partially address the hallucination issues in large language models, it also introduces new vu…
FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation
Large language models can produce fluent financial analysis, but fluency alone does not establish whether a report is suitable for institut…
ラベル ノイズの下で信頼できるハイパーグラフ ニューラル ネットワークを目指して
ハイパーグラフ ニューラル ネットワーク (HGNN) は、複雑な高次の関係を処理する際に優れた能力を実証しています。ただし、そのパフォーマンスはラベル付きデータに大きく依存するため、ラベル ノイズの影響を受けやすくなります。ラベル ノイズを使用した学習 (LLN) とラベル ノイズを使用したグラフ学習 (GLN) は進歩していますが、ハイパーグラフでのノイズのあるラベルの学習はまだ研究されていません。この論文では、ラベル ノイズの下でのハイパーグラフ ノード分類の体系的な研究を紹介します。まず、代表的な LLN および GLN 手法をハイパーグラフに適応させ、統一されたベンチマークの下で評価し、ハイパーグラフに対する既存のロバストな学習戦略の限界を明らかにします。これに基づいて、新しいハイパーグラフの堅牢なフレームワークである HyperTrust を提案します。これは、最初に事前トレーニングベースのエントロピー認識戦略を通じてハイパーエッジの信頼性を推定し、次にラベルのないノードを信頼できるハイパーエッジに接続することで信頼性の高い監視を強化する HyperedgeBoost モジュールと、信頼できないノードとハイパーエッジの発生を除去することでノイズの多い伝播を抑制する HyperedgePrune モジュールを組み込みます。最後に、2 つのモジュールが連携してハイパーグラフ構造を調整し、最終的な予測を生成します。広範な実験と理論分析により、さまざまなノイズの多い設定下での複数のハイパーグラフ データセットに対する HyperTrust の有効性と堅牢性が実証されています。私たちの研究は、ラベル ノイズを使用したハイパーグラフ学習のための統一ベンチマークと効果的なソリューションを提供し、この方向における将来の研究の基礎を築きます。
原文 (English)
Towards Trustworthy Hypergraph Neural Networks under Label Noise
Hypergraph neural networks (HGNNs) have demonstrated remarkable capabilities in processing complex higher-order relationships. However, their performance is highly dependent on labeled data, making them vulnerable to label noise. Despite advances in learning with label noise (LLN) and graph learning with label noise (GLN), noisy-label learning on hypergraphs remains underexplored. In this paper, we present a systematic study of hypergraph node classification under label noise. First, we adapt representative LLN and GLN methods to hypergraphs and evaluate them under a unified benchmark, revealing the limitations of existing robust learning strategies for hypergraphs. Building on this, we propose a new hypergraph robust framework, HyperTrust, which first estimates hyperedge trustworthiness through a pretraining-based, entropy-aware strategy, and then incorporates the HyperedgeBoost module to enhance reliable supervision by connecting unlabeled nodes to trustworthy hyperedges, as well as the HyperedgePrune module to suppress noisy propagation by removing untrustworthy node-hyperedge incidences. Finally, two modules work collaboratively to adjust the hypergraph structure and generate final predictions. Extensive experiments and theoretical analysis demonstrate the effectiveness and robustness of HyperTrust on multiple hypergraph datasets under various noisy settings. Our work provides a unified benchmark and an effective solution for hypergraph learning with label noise and lays a foundation for future research in this direction.
Image Classification Using CNN-QNN Hybrid Model with Optimized Correlated Features
We propose a method to optimize the correlation among convolutional neural network (CNN) features that are used as inputs to quantum neural…
NodeJEPA: Structure-Conditioned Latent Prediction for Node-Level Graph Self-Supervised Learning
Self-supervised learning on graphs is largely shaped by contrastive methods that depend on carefully designed augmentations, and by generat…
Approximate Multi-Objective Search Under Rulebooks
Robotic planning often involves multiple objectives with complex priority relationships, such as safety, efficiency, and regulatory complia…
Training-Free Hashing-Based Attention via Binary Principal Components
Long-context large language models (LLMs) are increasingly deployed in real-world applications, yet self-attention remains a major efficien…
MESH: 専門家混合トレーニング向けのメモリ効率の高いシンクホーン最適化
Sinkhorn gradient descent などのメモリ効率の高い行列オプティマイザーは、密な Transformer 行列のほとんどの AdamW オプティマイザー状態を削除しますが、Mixture-of-Experts (MoE) トレーニングへの直接適用は信頼できません。私たちは、制御された 1 億 1,000 万パラメータのナノクジラの DeepSeek スタイルの MoE 事前トレーニング設定でこの失敗を研究しました。 SAGE/Sinkhorn ハイブリッドは、オプティマイザの状態を 0.883 GB から 0.331 GB に削減しますが、評価損失は 3.8265 に低下し、同じセットアップで観察された AdamW ベースライン (調査したシード全体で 3.58 ~ 3.64) をはるかに上回ります。ルーティングされた MoE エキスパート行列が主な障害点であることを示します。その勾配は条件付きで、時間的に変化し、ステートレスな Sinkhorn 正規化の機能が不十分です。私たちは、MoE 専門家向けに隠れた勢いの Sinkhorn アップデートである MESH を提案します。 MESH は、エキスパートの最初の瞬間をオプティマイザーの状態として保存せずに、勾配バッファーのライフサイクルを通じて一時的な最初の瞬間の信号を復元します。 MESH は、粗いニューロン/ブロックの逆 RMS 乗算器を追加する、オプションのブロック事前条件付きバリアントです。アブレーション全体にわたって、マトリックス正規化前の一時的な平滑化が主な原因要素です。ブロック/ニューロンのプリコンディショニングはメモリ品質のフロンティアを改善できますが、普遍的に必要であるとは確立されていません。追加の 2 つのシードでは、MESH と MESH-B は、AdamW と比較して、オプティマイザー状態のメモリを 62.5% 削減し、ピーク時の PyTorch CUDA 割り当てを約 12.6% 削減し、評価損失のギャップは控えめです。完全な状態の診断バリアントはアブレーションにおいて AdamW のようなパフォーマンスを回復し、MoE の専門家は時間的な平滑化を必要としているが、必ずしも完全な座標上の AdamW 状態ではないという結論を裏付けています。
原文 (English)
MESH: Memory-Efficient Sinkhorn Optimization for Mixture-of-Experts Training
Memory-efficient matrix optimizers such as Sinkhorn gradient descent remove most AdamW optimizer state for dense Transformer matrices, but direct application to Mixture-of-Experts (MoE) training is unreliable. We study this failure in a controlled 110M-parameter nanowhale DeepSeek-style MoE pretraining setting. A SAGE/Sinkhorn hybrid reduces optimizer state from 0.883GB to 0.331GB but degrades evaluation loss to 3.8265, far above the AdamW baselines observed in the same setup (3.58--3.64 across the seeds we study). We show that routed MoE expert matrices are the dominant failure point: their gradients are conditional, temporally varying, and poorly served by stateless Sinkhorn normalization. We propose MESH, a hidden-momentum Sinkhorn update for MoE experts. MESH restores a temporal first-moment signal through the gradient-buffer lifecycle, without storing the expert first moment as optimizer state. MESH is an optional block-preconditioned variant that adds a coarse neuron/block inverse-RMS multiplier. Across ablations, temporal smoothing before matrix normalization is the primary causal ingredient; block/neuron preconditioning can improve the memory-quality frontier, but is not established as universally necessary. In two additional seeds, MESH and MESH-B reduce optimizer-state memory by 62.5\% and peak PyTorch CUDA allocation by about 12.6\% relative to AdamW, with a modest evaluation-loss gap. Full-state diagnostic variants recover AdamW-like performance in ablations, supporting the conclusion that MoE experts need temporal smoothing, but not necessarily full coordinate-wise AdamW state.
Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation
On-policy distillation (OPD) supervises student-visited trajectories, yet divergence-based rules cannot determine whether an erroneous pref…
SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation
On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can ass…
Generative Optimization for Incentivized Advertising with Global Level Constraints
Incentivized advertising allocates monetary or virtual rewards to drive user engagement, where a key challenge is optimizing continuous inc…
MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages
We present MERaLiON-GR, a speech gender recognition system that performs binary classification (female / male) on English and Southeast Asi…
ExeCRE: Execution-Consistency Guided Reliability Estimation for Self-Correcting Code Generation
Large language models (LLMs) have made notable progress in code generation, but they still struggle on challenging tasks that require sophi…
D$^2$F-ReAG: Dynamic Decomposition and Filtering for Multi-Hop Reasoning-Augmented Generation
Large language models (LLMs) often generate inaccurate answers due to their reliance on static internal knowledge. Retrieval-augmented gene…
When does training on downscaled images yield the same gradients?
Diffusion transformers deliver strong image generation, but their training cost grows superlinearly with resolution. Recent work justifies…
Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning
High-resolution pixels and crop or zoom tools give multimodal large language models the ability to inspect an image, but they do not provid…
TwinIR: Coordinated Invisible Dual-Point Attacks on Online HD Map Construction
Online HD map construction is critical to prediction and planning in autonomous driving. We find that existing physical attacks against onl…
Eigenius: A Typed Knowledge-Graph DBMS with Epistemic Stratification and Institution-Mediated Reasoning
As "AI Scientists" emerge to drive research via the Model Context Protocol (MCP), systems relying on ephemeral scripts will fail. The sheer…
Tropical Algebraic Geometry for Neuronal Representations: An Arakelov-Green Measure Based Descriptor for Graph Learning
The quantitative analysis of 3D neuronal morphologies requires capturing both graph topology and spatial geometry. Current message-passing…
Beyond Linear Dynamics: Neural Bilinear Dynamical Models for Time Series Forecasting
Time series in real-world applications are often generated by nonlinear dynamical systems, making accurate forecasting challenging. Existin…
EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment
The development of foundation models (FMs) is crucial for advancing endoscopic image analysis. However, existing endoscopy FMs mainly rely…
AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation
Text-to-audio (TTA) generation has recently achieved remarkable progress in synthesizing realistic audio from natural language descriptions…
AFD-Ledger: Deployment Provisioning for Attention--FFN Disaggregation
Attention--Feed-Forward Network (FFN) Disaggregation (AFD) is emerging as a promising architecture for serving Mixture-of-Experts (MoE) lan…
GeoReward: Mitigating Contextual Variable Overestimation in Vision-Language Models for Cross-Market Preference Prediction
Vision-language models excel in many multimodal tasks but remain prone to a subtle yet impactful failure mode: they tend to overestimate do…
GUARD: Grounding Uncertainty and Ablation-Based Risk Detection for Diffusion-Based VLAs
Diffusion-based vision-language-action (VLA) policies can generate plausible actions even when their predictions are weakly grounded in the…
CARVE: Cross-Slice Anisotropic Reallocation of Visual Evidence for Efficient 3D Medical Volume Understanding
Slice-based MLLMs leverage mature 2D encoders by representing 3D volumes as sequences of 2D slices. However, this slice-wise formulation pr…
A Model Merging Approach for Continual MLLM Unlearning
Multimodal large language model (MLLM) unlearning methods have been proposed to remove private, sensitive, or proprietary information from…
EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks
Frontier LLMs are increasingly put to use on open-ended complex questions, different in nature from the ones they are typically evaluated o…
Breadcrumbing Search Agents
LLM-based search agents are widely used for information-seeking tasks, but their reliance on external tool returns introduces a critical se…
PhysMind: From Video to Executable Worlds for Training-Free Physical Reasoning
Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions. Existing vision-lan…
Breaking the Curse ofMultilinguality inMany-to-Many Speech-to-Text Translation via a Resource-AwareMixture of Speech Encoders
Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing m…
EASy: Towards Efficient LLM-Based Agentic System
Agentic systems have emerged as a promising paradigm for solving complex tasks by coordinating specialized LLM-based agents. However, most…
The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering
EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can…
When Absence Is Evidence: Evaluating Completeness-Sensitive Negative Reasoning in Large Language Models
Large language models (LLMs) are often asked whether something is absent from a record, list, or retrieved context. Yet non-observation lic…
Rethinking Reservoir Pruning: A Dynamical Perspective for Echo State Networks
Echo State Networks (ESNs) offer an efficient framework for temporal prediction, but their randomly initialized reservoirs are often over-p…
The Order Is the Guarantee: Verifier-Budgeted Code Deletion with Static-First Learned Proposals
Frontier coding models now match or exceed strong human reference points on programming benchmarks, yet benchmark success does not imply ma…
Masked diffusion enables coherent beat tracking
Current neural networks for beat tracking generate invalid outputs, such as consecutive downbeats and erratic tempo changes, even when thes…
DisMix: Order-Aware Mixup for Medical Imaging via Disentangling Ordinal and Non-Ordinal Features
Image mixup is a widely adopted data augmentation strategy, yet it is ill-suited for ordinal classification tasks such as medical disease g…
CSGen: A Multi-Domain Curvilinear Structure Generation Model via Hierarchical Multimodal Diffusion
Curvilinear structure analysis is an important and fundamental task in multimedia. However, the controllable generation of images with prec…
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark
Large Language Models (LLMs) have transformed computational linguistics and achieved remarkable performance across numerous natural languag…
Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports
Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a…
Personalized Federated Sparse Adaptation of Time-Series Foundation Models
Federated adaptation of time-series foundation models (TSFMs) is attractive for building energy forecasting because meter data are private,…
Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO
We tackle the challenging yet underexplored task of Generalized Referring Expression Comprehension (GREC), which requires a model to locali…
Design Choices That Matter: A Functional ANOVA Analysis for Remote Sensing Multi-Label Classification
Benchmarking deep learning (DL) models for multi-label classification (MLC) of remote sensing images (RSI) typically yields rankings that d…
A 6G Integrated Sensing and Communication Framework for Railway Intrusion Detection and Collision Prediction
Integrated Sensing and Communication (ISAC) combines sensing and communication to efficiently utilize wireless resources and is emerging as…
What We Observe as LLM Behavior Can Be a Side-effect of Inference Backend
Benchmark scores are reported as properties of a model, yet the inference framework used to produce them, such as HuggingFace, vLLM, or Oll…
Toward Integrating Adaptive Experience Replay and Online Uncertainty Estimation in Safe Actor-Critic Optimal Control
Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though ea…
PURPOSE: Poisoning Conflict Resolution in RAG via Proxy-Fact-Grounded Updates
In Retrieval-Augmented Generation (RAG), post-retrieval conflict resolution arbitrates among noisy or contradictory retrieved passages. How…
InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval
Self-improving agents accumulate reusable insights from prior trajectories, making retrieval increasingly important for turning accumulated…
Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models
Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language understanding, and robotic contro…
FUSEP: A Multi-Center Benchmark for Diverse Tasks in Early Pregnancy Fetal Ultrasound Screening
A large number of infants with congenital anomalies are born each year globally, especially in areas with underdeveloped medical resources.…
Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent
Scaling supervision for multi-turn medical agents is difficult because expert dialogue annotation is costly and clinical conversations are…
IMFACT: Counterfactual Explanations for Time Series via Intrinsic Mode Function Substitution
Oscillatory signals, such as vibration, carry class-discriminative information in specific frequency bands; perturbing them in raw feature…
RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists
The integration of Large Language Models (LLMs) into software engineering has shifted the focus from function-level generation to repositor…
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation
Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited g…
Scrouting: Cost-Aware Routing of Coding Agents by Scouting the Repository First
Frontier language models can resolve repository-level software issues, but each attempt is expensive, and existing routers select a model f…
Towards a satellite image manipulation and deepfake localization benchmark dataset
Verifying the authenticity of satellite imagery has become increasingly critical given advances in generative artificial intelligence. High…
A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination
Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop…
When Does Latent Communication Pay? A Causal Audit of Relayed KV Caches in Multi-Agent LLMs
Multi-agent LLM systems relay key--value caches instead of text and credit their gains to exchanged ``latent thoughts''. That credit is a c…
A Chain Is Only as Strong as Its Weakest Link: A Scoping Review of System Integration Audits in AI
As AI systems become increasingly integrated into diverse interfaces and applications, model-centric audits are insufficient to address ris…
Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning
As chart images, tabular data, and visualization code play increasingly important roles across diverse domains, cross-representation unders…
SVI-DAG: A Structured Variational Inference Approach to Bayesian Causal Discovery
Bayesian causal discovery seeks to determine the posterior distribution of causal theories, which are interpreted as directed acyclic graph…
CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications
CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and techn…
A General Sufficient Condition for Rewriting Horn-ALCHI Atomic Queries into GQL
The emergence of the ISO standard GQL introduces a powerful query language extending first-order logic with controlled recursion, raising t…
SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models
SciCode is the standard measure of the scientific-coding ability of language models: research-level problems that demand both frontier scie…
Protoreasoning in Tiny Transformers
We show that tiny transformers can profitably employ a simple form of Chain of Thought, which we call protoreasoning, allowing us to study…
ORACLE: A Multi-Objective Reinforcement Learning-Based Analog Circuit Design Optimizer with Large Language Models-Guided Exploration
Analog circuit design automation using reinforcement learning (RL) has emerged as a promising approach for reducing manual effort. However,…
OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents
LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-en…
Revealed Rationality: Label-Free Evaluation and Regularization from Representation Theorems
Representation theorems in decision theory establish that behavior satisfies certain axioms if and only if it can be rationalized by a well…
ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation
High-quality annotation of artworks is essential for computational art research, yet extracting implicit semantics remains challenging due…
Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning
Released aligned large language models remain vulnerable to malicious downstream finetuning. Existing defenses are largely designed for the…
The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR Simulations
Against the backdrop of violence in police interactions with the U.S. public, we explore how deferentially police officers speak to virtual…
MarsCast: Transfer Learning of AI Weather Foundation Models to Planetary Atmospheres
We investigate the transferability of Earth weather foundation models to planetary atmospheres by adapting the GraphCast graph neural weath…
RepairFormer: Automated Repair of Structured Inputs Using Transformers
Structured input files such as JSON, DOT, OBJ, INI, S-expression, and TinyC are widely used in software systems, but small corruptions can…
Hardware Design and Security in the Era of Chiplets and LLMs
The semiconductor industry is undergoing a dual revolution: the shift toward heterogeneous 2.5D chiplet systems and the integration of Larg…
Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language Models
Small open-weight language models increasingly run in private, offline, and cost-sensitive settings, where the key deployment question is n…
VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly Detection
Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual variability in surveillance f…
MultiPathFormer: Towards a Foundation Model for Multipath Wireless Propagation
Recent advances in machine learning have enabled training of wireless foundation models, which aim to support tasks such as channel estimat…
Capability-Gated Planning: Cost-to-Goal Discovery and the Limits of Myopic Experiment Selection
Systems that automate scientific discovery must repeatedly decide which experiment to run, which hypothesis to test, which tool to build, a…
Representational separation between unitary and channel quantum generative models via shared classical randomness at shallow depth
Near-term quantum hardware limits circuit depth and often imposes geometrically local connectivity for quantum generative models, restricti…
Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition
Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computationally efficient classroom i…
Chained Recursive Language Models for Multi-Iteration Reasoning
Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simulta…
SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant
Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging. Existing vec…
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language…
Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains
Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important fo…
GRALS: GCN-Guided Redundancy-Aware Local Search for Minimum Vertex Cover
The minimum vertex cover (MVC) problem seeks to identify the smallest set of vertices that cover all edges in an undirected graph. As a fun…
The Yokai Learning Environment: Tracking Beliefs Over Space and Time
The ability to cooperate with unknown partners is a central challenge in cooperative AI and widely studied in the form of zero-shot coordin…
Zero-shot reasoning for simulating scholarly peer-review
Scholarly publishing requires scalable scrutiny supported by auditable evidence. This paper presents a two-component benchmark of xPeer, th…
Corrigibility Transformation: Constructing Goals That Accept Updates
An AI agent will learn a desired goal more effectively if it does not resist the training process, but many partially learned goals incenti…
Calibrating Transformer Attention via Task-Space Sensitivity Feedback
Transformer-based pre-trained language models (PLMs) excel in text classification but suffer from attention dilution and attention sink eff…
XGrammar-2: エージェント LLM 用の効率的な動的構造化生成エンジン
最新の LLM エージェントは、ツールの呼び出しや応答のプロトコルなど、動的に構造化された生成にますます依存しています。静的構造を使用した従来の構造化生成とは異なり、これらのワークロードはリクエスト間とリクエスト内で変化し、既存のエンジンに新たな課題をもたらします。動的エージェント ワークロード用の構造化生成エンジンである XGrammar-2 を紹介します。私たちの設計は 2 つの重要なアイデアに基づいています。1 つはタグトリガーによる構造切り替えの最上級のサポート、もう 1 つは異なる出力構造を持つリクエスト間でのきめ細かい再利用です。具体的には、XGrammar-2 では、動的な構造ディスパッチのための TagDispatch と、文法全体でサブ構造レベルのキャッシュを再利用するための Cross-Grammar Cache が導入されています。 Earley ベースの適応型トークン マスク キャッシュ、ジャストインタイム コンパイル、および繰り返し状態圧縮により、効率がさらに向上します。実験の結果、XGrammar-2 は以前の構造化生成エンジンよりも 6 倍以上高速なコンパイルを実現し、最新の LLM サービス システムではエンドツーエンドのオーバーヘッドがほぼゼロになることが示されています。
原文 (English)
XGrammar-2: Dynamic and Efficient Structured Generation Engine for Agentic LLMs
Modern LLM agents increasingly rely on dynamic structured generation, such as tool calling and response protocols. Unlike traditional structured generation with static structures, these workloads vary both across requests and within a request, posing new challenges to existing engines. We present XGrammar-2, a structured generation engine for dynamic agentic workloads. Our design is based on two key ideas: first-class support for tag-triggered structure switching, and fine-grained reuse across requests with different output structures. Concretely, XGrammar-2 introduces TagDispatch for dynamic structural dispatching and Cross-Grammar Cache for substructure-level cache reuse across grammars. It further improves efficiency with an Earley-based adaptive token mask cache, just-in-time compilation, and repetition state compression. Experiments show that XGrammar-2 achieves over 6x faster compilation than prior structured generation engines, and incurs near-zero end-to-end overhead in modern LLM serving systems.
MemFly: On-the-Fly Memory Optimization via Information Bottleneck
Long-term memory enables large language model agents to tackle complex tasks through historical interactions. However, existing frameworks…
Text2GraphQuery-Bench: A Text to Graph Query Benchmark
Graph models are fundamental to data analysis in domains rich with complex relationships. Unlike SQL, which benefits from a rel- atively un…
SimMOF: AI agent for Automated MOF Simulations
Metal-organic frameworks (MOFs) offer a vast design space, and as such, computational simulations play a critical role in predicting their…
AI Assistance Reduces Persistence and Hurts Independent Performance
People often optimize for long-term goals in collaboration: A mentor or companion doesn't just answer questions, but also scaffolds learnin…
Assessing and Explaining the Persuadability of Large Language Models as Legal Decision Tools
As Large Language Models (LLMs) are proposed as legal decision assistants, and even first-instance decision-makers, across a range of judic…
Contextual Agentic Memory is a Memo, Not True Memory
Current agentic memory systems (vector stores, retrieval-augmented generation, scratchpads, and context-window management) do not implement…
Online Goal Recognition using Path Signature and Dynamic Time Warping
Online goal recognition in continuous domains poses two central challenges: efficiently encoding large trajectories and effectively compari…
CogniFold: コグニティブフォールディングによる常時オンのプロアクティブなメモリ
既存のエージェントの記憶は主に反応的かつ検索ベースのままであり、経験を自律的に永続的な認知構造に組織化する能力が欠けています。真の自律型エージェントを目指して、次世代のプロアクティブ アシスタント向けに設計された、脳からインスピレーションを得た「常時オン」エージェント メモリである CogniFold を紹介します。 CogniFold は、断片化されたイベント ストリームを自己出現の認知構造に継続的に折り畳んで、入ってくるイベントと蓄積された知識から徐々により高いレベルの認知をブートストラップします。私たちは、相補学習システム (CLS) 理論を 2 層 (海馬、新皮質) から 3 層に拡張し、前頭前意図層を追加することでこれを根拠にしています。意図的な制御と意思決定の拠点として前頭前野をエミュレートする CogniFold は、グラフ トポロジーの自己組織化を通じてこれを実現します。つまり、認知構造はストリームの下で積極的に集まり、意味的に類似している場合は結合し、古くなっている場合は減衰し、連想想起を通じて再リンクし、概念クラスターの密度がしきい値を超えると意図を表面化します。 CogEval-Bench を使用して構造形成を評価し、CogniFold が認知的期待と概念創発に一致する記憶構造を独自に生成することを実証します。さらに、5 つの認知ドメインにわたる 7 つの広範なベンチマークにわたって、CogniFold が従来のメモリ ベンチマークでも同時に堅牢に実行されることを検証しました。
原文 (English)
CogniFold: Always-On Proactive Memory via Cognitive Folding
Existing agent memory remains predominantly reactive and retrieval-based, lacking the capacity to autonomously organize experience into persistent cognitive structure. Toward genuinely autonomous agents, we introduce CogniFold, a brain-inspired "always-on" agent memory designed for the next generation of proactive assistants. CogniFold continuously folds fragmented event streams into self-emerging cognitive structures, bootstrapping progressively higher-level cognition from incoming events and accumulated knowledge. We ground this by extending Complementary Learning Systems (CLS) theory from two layers (hippocampus, neocortex) to three, adding a prefrontal intent layer. Emulating the prefrontal cortex as the locus of intentional control and decision-making, CogniFold achieves this through graph-topology self-organization: cognitive structures proactively assemble under the stream, merge when semantically similar, decay when stale, relink through associative recall, and surface intents when concept-cluster density crosses a threshold. We evaluate structural formation using CogEval-Bench, demonstrating that CogniFold uniquely produces memory structures that match cognitive expectations and concept emergence. Furthermore, across eight downstream benchmarks -- two probing long-term conversational memory (LoCoMo, LongMemEval) and six spanning other cognitive domains -- we validate that CogniFold simultaneously performs robustly on conventional memory tasks. Our code is available at https://github.com/OpenNorve/CogniFold.
古典的なヒューリスティック検索問題としての思考の木: 正式な基礎とデザイン パターン
大規模言語モデル (LLM) は優れた推論機能を実証していますが、その標準的な生成プロセス (自動回帰トークン予測) は本質的に近視眼的であり、連鎖的なエラーが発生しやすいものです。これに対処するために、Tree-of-Thoughts (ToT) フレームワークは、中間の推論ステップにわたって検索スペースを作成し、検索モデルが探索、先読み、後戻りできるようにします。ただし、現在の ToT 研究は、自然言語処理および自動計画のコミュニティ全体で断片化したままであり、多くの場合、一貫性のない用語やアドホックな実装が使用されています。その結果、古典的なヒューリスティック検索用語に基づいた統一分類法を通じて ToT ランドスケープを統合します。 LLM ベースの推論を、状態表現 (思考の粒度)、後続生成 (プロンプト演算子)、ヒューリスティック評価 (進捗状況の自己評価) といった古典的な検索コンポーネントにマッピングします。私たちは分類法のコンテキスト内で既存の作業を分析し、新たな設計パターンを特定します。つまり、浅くて決定論的なタスクに対する体系的な検索 (Best-First Search) と、深い複数ステップの推論に対する先読み重視の戦略 (DFS、MCTS) です。最後に、ヒューリスティック検索と LLM 推論の交差点における未解決のアルゴリズムの課題を特定し、ヒューリスティック検索コミュニティにこの新たな領域に取り組むよう呼びかけます。
原文 (English)
Tree of Thoughts as a Classical Heuristic Search Problem: Formal Foundations and Design Patterns
Large Language Models (LLMs) have demonstrated remarkable reasoning capabilities, yet their standard generation process -- auto-regressive token prediction -- is inherently myopic and prone to cascading errors. To address this, the Tree-of-Thoughts (ToT) framework creates a search space over intermediate reasoning steps, allowing search models to explore, look ahead, and backtrack. However, current ToT research remains fragmented across Natural Language Processing and Automated Planning communities, often using inconsistent terminology and ad-hoc implementations. Consequently, we synthesize the ToT landscape through a unified taxonomy based on classical heuristic search terminology. We map LLM-based reasoning to classical search components: state representation (granularity of thoughts), successor generation (prompting operators), and heuristic evaluation (self-assessment of progress). We analyze existing work within the context of our taxonomy and identify emerging design patterns: systematic search (Best-First Search) for shallow, deterministic tasks and lookahead-heavy strategies (DFS, MCTS) for deep multi-step reasoning. We conclude by identifying open algorithmic challenges at the intersection of heuristic search and LLM reasoning, and call on the heuristic search community to engage with this emerging domain.
トリビアム: 因果記憶コントローラーの第一級目標としての時間的後悔
現在のエージェント システムと LLM パイプラインの多くは、結果の報酬を最適化することで間違いを修正します。これは失敗の内容のみを扱います。結果が予測と異なる場合、不一致の理由と時期が体系的に記録、レビュー、修正されないため、同じエラーがエピソードごとに再発する可能性があります。私たちは、これは単にモデルの能力の問題ではなく、構造的な問題であると主張します。私たちは、作業因果モデルに対する結果の後悔や認識論的な後悔と並んで、長期的な時間的後悔を第一級の目標として提案します。時間的リグアロングは、失敗が継続するとき、すなわち、調整ミスの因果モデルが修正されるまでにどのくらいの期間許容されるかを捉えます。認識論的後悔は、失敗が続く理由、つまり作業因果モデルにおける残留不確実性またはエラーを捉えます。 3 つの後悔を総合すると、長命のエージェントがいつ、何が、なぜ失敗する可能性があるのかについて、反証可能な説明が得られます。エージェントを E エピソードのストリームとしてモデル化し、明示的な因果関係の調査、持続性、および検出可能性の仮定に基づいて 3 つの条件付き結果を証明します。まず、観察的に等価な交絡のもとでは、結果のみの学習では介入チャネルがなければ因果構造と偽の構造を区別できないため、結果の後悔がゼロになった後でも時間的誤調整が線形的に持続する可能性があります。第 2 に、永続的な因果ログと予算付きプローブを使用すると、総プローブの複雑さはエピソード期間内で対数的となり、O(log E) の時間的後悔を引き起こします。第三に、K 個の検出可能な変化点の下では、速度は O(K log E) まで拡張されます。 Trivium をインスタンス化し、5 つの反証可能な予測を事前に登録します。 CausalBench-Seq では、Trivium は予測された対数エンベロープに従いますが、結果のみのベースラインは直線的に増加します。パイロットのリアル LLM ストリームは、1 回の完全な E = 500 実行と 3 回の E = 100 フロンティア モデル パイロットにわたる予備的な外部妥当性証拠を提供します。ここでの自己学習とは、LLM 重みを再トレーニングすることではなく、外部因果モデルを修正することを意味します。
原文 (English)
Trivium: Temporal Regret as a First-Class Objective for Causal-Memory Controllers
Many agentic systems and LLM pipelines correct mistakes by optimizing outcome reward. This addresses only the what of failure; the why and when may go unlogged, allowing the same error to recur across episodes. We propose long-horizon temporal regret alongside outcome regret and epistemic regret. These are diagnostic quantities, not standard comparator-based online-learning regrets. Temporal regret captures how long an unresolved or incorrect causal model is tolerated; epistemic regret captures posterior error over that model. Over a stream of E episodes, we prove three conditional results under explicit probing, persistence, and detectability assumptions. First, under observationally equivalent confounding, outcome-only learning cannot separate causal from spurious structure, so miscalibration can persist after outcome regret reaches zero. Second, with a persistent causal log and budgeted probes, total probe complexity is logarithmic in E, inducing O(log E) delayed-identification temporal regret. The implemented clipped soft score adds a linear numerical floor, so this logarithmic claim applies only to identification delay. Third, under K detectable change-points, the rate extends to O(K log E). We instantiate Trivium and state five falsifiable predictions. On CausalBench-Seq, a hard structural readout records 7.8+-2.9 misidentified episodes per seed over 500 episodes and 20 seeds, with zero observed stationary errors, while outcome-only controllers remain misidentified throughout. Audit ablations show that continued posterior updating, not detector reopening or local repair, drives posterior recovery; local repair instead reduces committed-graph dispatch exposure. The prior logarithmic-envelope verdict based on a clipped soft score is withdrawn. A pilot real-LLM stream provides external evidence. Self-learning here means revising an external causal model, not retraining LLM weights.
アブレーション可逆ヘッドは転移しない: 変圧器における機械的役割要求のストレス テスト
機械的解釈可能性では、アテンションヘッドは一般に、ある動作に必要な場合に役割主張(たとえば、「このヘッドは追加を表します」)に昇格し、それを線形にエンコードし、アブレーション後に復元されたときにその動作を回復します。我々は、この証拠が不十分であることを示しています。3 つの 7-8B 命令調整モデルと 5 つの計算ファミリーにわたって、3 つのチェックすべてに合格したヘッドは、それらのアクティベーションが一致した制御下で別のプロンプトにパッチされると、計算の転送に日常的に失敗します。アテンションヘッド向けの役割割り当てレンズである KID (Knowing / Intent / Doing) を導入し、それを 3 段階のパイプラインと組み合わせます。つまり、機能選択的スクリーニング (CSS)、特異値分解 (SVD)、および一致した制御下での活性化変換です。私たちの結果は、予備的な役割分類(プロンプト軌道安定化装置、応答側ロジットバイアスヘッド、およびソフト計算パターンキャリアを含む)を文書化し、同一応答制御(応答文字列を共有するが要求された計算を共有しない変換ターゲット)が、意味論的特異性を装った広範な状態転送を暴露する十分に活用されていないチェックであることを示しています。
原文 (English)
Necessary, Decodable and Reversible, Yet Not Transferable: A Stress Test for Attention-Head Role Claims
Mechanistic studies often assign a component a role when removing it damages a behavior, its activation linearly encodes task information, and restoring that activation repairs the damage. We test the stronger implication: can the same activation carry the requested computation into another prompt? In the roughly 8 of 15 model-family cells receiving our full matched-control transfer assay, covering three 7-8B instruction-tuned models and 5 single-step computation families, tested attention-head states did not produce clean computation transfer. Same-answer and matched-context controls instead exposed inert or broad effects. Positive controls bound this result rather than eliminating its caveats: the instrument recovers known function vectors when writing into an underspecified prompt, and recovers a compact override-capable residual state in one of three controlled-model seeds. Neither control proves sensitivity to the object carried by real-model attention during prompt override. On two composed forms in Qwen, a strongly decodable operation direction also failed to steer the generated answer under passing gates; the corresponding cross-model test was invalid. These results show that necessity, decodability, same-prompt repair, and cross-prompt transfer are separable evidence. They do not establish that a portable selection state is absent from untested components or settings.
理論レベルの自動形式化: 分離されたステートメントから統合された形式的知識ベースへ
自動形式化は、非公式な自然言語を機械検証可能な正式な言語に変換します。ほとんどの作業は個々のステートメントに焦点を当てていますが、実際の形式化の取り組みは本質的に理論レベルです。対象となる定理を述べる前に、公理、定義、補題の網全体が必要です。このポジションペーパーでは、理論レベルの自動形式化、つまりすべての相互依存関係を含む完全な理論を構造化ライブラリとして形式化することについて主張します。私たちはこの変化の重要性を検討し、別の見解に取り組み、未解決の課題を特定し、今後の 3 つの有望な道筋を提案します。自動形式化に関する調査は、https://github.com/marcusm117/Awesome-Autoformalization でご覧いただけます。
原文 (English)
Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases
Autoformalization translates informal natural language into formal, machine-verifiable languages. While most work focuses on individual statements, real formalization efforts are inherently theory-level: they require an entire web of axioms, definitions, and lemmas before target theorems can even be stated. In this position paper, we argue for theory-level autoformalization: formalizing complete theories, including all their inter-dependencies, as structured libraries. We examine the significance of this shift, address alternative views, identify open challenges, and propose three promising paths forward. Our survey of autoformalization is available at https://github.com/marcusm117/Awesome-Autoformalization.
意味的等価性を超えて: LLM 不確実性定量化のための論理グラフ
大規模言語モデル (LLM) は、自信を持って記述されているにもかかわらず信頼性の低い出力を生成することが多く、安全性が重視されるアプリケーションへの展開に重大な課題をもたらします。意味論的エントロピーなどの既存の不確実性指標は、意味論的等価性のレベルで一致を捉えますが、異なる答え間の論理的関係をほとんど無視します。その結果、生成された応答の形式は多様であるが論理的に互換性がある(たとえば、粒度または特異性のみが異なる)設定では、不確実性を過大評価し、幻覚に誤ってフラグを立てる傾向があります。私たちは、回答間の含意と非互換性を明示的にモデル化するフレームワークである Logical Graph Uncertainty (LGU) を提案します。 LGU は含意チェーンに沿って確率質量を集計し、論理的に最大の仮説に対するエントロピーを計算し、それらの間の相互非互換性にペナルティを課します。複数の質問応答ベンチマークにわたって、LGU は既存の手法よりも不確実性の推定を一貫して改善し、データセット全体でセマンティック エントロピー ベースラインを最大 +7.1% AUROC および +3.5% AUARC 上回っています。
原文 (English)
Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification
Large Language Models often produce confidently stated yet unreliable outputs, posing critical challenges for deployment in safety-sensitive applications. Existing uncertainty metrics such as semantic entropy capture agreement at the level of semantic equivalence, but largely ignore the logical relationships between distinct answers. As a result, they tend to overestimate uncertainty and falsely flag hallucinations in settings where generated responses are diverse in form yet logically compatible (e.g., differing only in granularity or specificity). We propose Logical Graph Uncertainty (LGU), a framework that explicitly models implication and incompatibility among answers. LGU aggregates probability mass along entailment chains onto the most specific hypotheses the answers support, measures the entropy of the resulting distribution, and penalizes mutual incompatibility among those hypotheses. Across multiple question-answering benchmarks and model families, LGU ranks first on average among existing uncertainty measures, with its largest gains---up to +7.1\% AUROC and +3.5\% AUARC over semantic entropy---on questions whose sampled answers are logically structured.
Risk Is Not the Target: A Monotonic Framework for Evaluating Wildfire Operational Risk Signals
Evaluating wildfire risk systems using standard machine-learning metrics such as F1-score or IoU is fundamentally flawed: these metrics ass…
不完全な調整下での価値の脆弱性
AI システムに課せられる責任が増すにつれて、これらのシステムが人間性と整合していることを保証することがますます重要になります。 AI の安全性に関する一般的な懸念は、人間の価値は脆弱であるということです。つまり、人間の価値を不完全に代替するために過度に最適化すると、壊滅的な結果につながるということです。この論文では、エージェントが世界を最適化する前にその価値関数が代理条件を満たすことを保証する理想的なアライメント トレーニングを受けるアライメント問題のモデルを紹介します。私たちの主要な結果は、人間の価値関数に関する条件と、$\eta$-壊滅的な価値関数を持つエージェント、つまり最適化能力の限界において人間の価値の期待値が $\eta$ を下回ることが保証されるエージェントが配備される場合のいくつかの代用条件の精度を特定しました。私たちの結果は、過剰最適化の危険性を浮き彫りにし、導入前のトレーニングのみに依存するのではなく、量子化器などの最適化圧力を制限する AI 設計を動機付けるものです。
原文 (English)
Fragility of Value under Imperfect Alignment
As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imperfect proxy to human values will lead to a catastrophic outcome. In this paper, we present a model of the alignment problem where an agent undergoes idealized alignment training that guarantees its value function satisfies a proxy condition before optimizing the world. Our primary results identify conditions on the human value function and the accuracy of several proxy conditions under which an agent with an $\eta$-catastrophic value function, one that is guaranteed to take the expectation of human value below $\eta$ in the limit of optimizing power, would be deployed. Our results highlight the danger of overoptimization and motivate AI designs that limit optimization pressure, such as quantilizers, rather than relying solely on pre-deployment training.
ExtractBench: スキーマに基づいたエンタープライズ ドキュメント抽出のベンチマーク
エンタープライズ ワークフローでは、\emph{スキーマに基づく抽出} に関してエージェントへの依存度が高まっています。ドキュメントとユーザー定義のスキーマが与えられると、エージェントはそのスキーマに従って、根拠となるメタデータとしてソース証拠を含む正しい出力を生成します。私たちは、スキーマに基づいた抽出のベンチマークである ExtractBench を紹介します。これは、私たちの知る限り、値の精度をスコアリングし、大規模な完全性、根拠、および測定コストをまとめて記録した最初のベンチマークです。この評価システムには、370 の企業文書、8 つのビジネス ドメイン、および 67 の文書タイプにわたる 4,869 ページが含まれており、課題シナリオを区別する明確なタグが付いています。スケーラブルなスキーマとグラウンド トゥルース キュレーション パイプラインは、実際のドキュメントに対する独立したシステムの合意、合成リストに対する既知の値、およびフォームに対する人間による検証を組み合わせています。値の精度として順序に依存しない値 F1 と、ソースのトレーサビリティのための 2 つの基本指標 (ワード レベルとページ レベルの F1) を報告します。商用 VLM は短いドキュメントではうまく機能しますが、長いドキュメントではレコード リストが切り捨てられることがよくありますが、コーディング エージェントははるかに高いコストで高い精度を維持します。 LlamaExtract Agentic Plus は、わずかなコストでコーディング エージェントに匹敵する精度を備え、3 つの指標すべてで第 1 位にランクされています。データセットと評価コードは、\href{https://huggingface.co/datasets/llamaindex/ExtractBench}{HuggingFace} および \href{https://github.com/run-llama/ExtractBench}{GitHub} で入手できます。
原文 (English)
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction
Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We present ExtractBench, a benchmark for schema-guided extraction and, to our knowledge, the first to score value accuracy, record completeness at scale, grounding, and measured cost together. The evaluation system contains 4,869 pages across 370 enterprise documents, 8 business domains, and 67 document types, with clear tags differentiating their challenge scenarios. The scalable schema and ground-truth curation pipeline combines independent-system agreement for real documents, known values for synthetic lists, and human verification for forms. We report order-insensitive value F1 for value accuracy, plus two grounding metrics for source traceability: word- and page-level F1. Commercial VLMs perform well on short documents but often truncate record lists on long ones, while coding agents retain higher accuracy at much higher cost. LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost. Dataset and evaluation code are available on \href{https://huggingface.co/datasets/llamaindex/ExtractBench}{HuggingFace} and \href{https://github.com/run-llama/ExtractBench}{GitHub}.
Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation
Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory complianc…
MemSIF: From Structured Interactions to Dual-Track Fact Memory for LLM Agents
Long-term memory is critical for LLM agents operating over long-horizon interactions. However, several persistent limitations of existing m…
HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents
Retrieval-augmented search agents answer multi-hop questions by repeatedly issuing search queries and accumulating evidence. This creates a…
Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy
Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities,…
Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling
Recurrent linear attention models (RLAs) such as Mamba offer efficient linear-time sequence modeling as an alternative to Transformers, yet…
Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions
Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users'…
When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO
Reinforcement learning with verifiable rewards (RLVR) com- monly optimizes each correct completion as an independent learning signal. In GR…
PhyAI: エッジでのリアルタイム物理 AI、クラウドでのスケーラブルなロールアウト
物理 AI ポリシーでは、モデルの評価、クラウド強化学習のロールアウト、エッジ GPU の提供、オンボード展開などのライフサイクル全体にわたって推論が必要です。これらの設定は同じチェックポイントとアクションのセマンティクスを共有しますが、多くの場合、別の推論プログラムに依存します。これらを統合するために、グラフの実行、カーネル、メモリ管理、並列サービスを共有しながら、アーキテクチャ固有のコンディショニング、ソルバー、キャッシュ、出力ロジックをモデル アダプターに保持する単一のランタイムを備えた物理 AI 推論エンジンである PhyAI を構築します。同じコードベースは、オンボード、エッジ、クラウドの展開全体で単一または複数の GPU 上でビジョン言語アクション (VLA) モデルとワールド アクション モデル (WAM) を実行します。 MiniCPM-Robot のリリース日にアダプター インターフェイスを使用して追加しました。 PhyAI は、pi0、pi0.5、GR00T N1.7、および MiniCPM-Robot の公式実装と比較して 1.40 倍から 4.65 倍の高速化を達成します。 Cosmos3-Nano-Policy-DROID では、8 つの H20 GPU (CFG=2、TP=4) でレイテンシが 2.46 秒から 1.18 秒に短縮され、2.08 倍の速度向上になります。特殊なランタイムはいくつかの構成で引き続き高速であるため、私たちの目標は、すべてのケースで最速の結果ではなく、競争力のあるレイテンシーを備えた 1 つのランタイムです。詳細なプロファイルにより、異なるモデルに異なる実行ポリシーが必要な理由が明らかになります。バッチ サイズ 1 の Hopper シリーズ GPU では、pi0.5 アクション エキスパートは FLOP の 8.8% を占めますが、レイテンシーの 57.2% を占めます。バッチ サイズ 32 では、そのシェアは 13.5% に低下し、スループットは約 100 サンプル/秒に達します。 Cosmos3 は世代主導のままで、バッチ サイズが 1 から 16 に増加してもスループットは 14.3% しか向上しません。さらに、推論制限制御と環境制限制御を区別する制御時間ルーフラインを導入します。 4 つの LIBERO スイートで測定された pi0.5 ポイントは環境に依存していますが、Cosmos3 は推論に依存したままです。コードとベンチマーク: https://github.com/mingti-org/phyai。
原文 (English)
PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud
Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Although these settings share the same checkpoint and action semantics, they often rely on separate inference programs. To unify them, we build PhyAI, a Physical AI inference engine with a single runtime that keeps architecture-specific conditioning, solver, cache, and output logic in model adapters while sharing graph execution, kernels, memory management, and parallel services. The same codebase runs vision-language-action (VLA) models and world-action models (WAMs) on single or multiple GPUs across onboard, edge, and cloud deployments. We used the adapter interface to add MiniCPM-Robot on the day of its release. PhyAI achieves 1.40x-4.65x speedups over the official implementations of pi0, pi0.5, GR00T N1.7, and MiniCPM-Robot. On Cosmos3-Nano-Policy-DROID it reduces latency from 2.46 to 1.18 s on eight H20 GPUs (CFG=2, TP=4), a 2.08x speedup. Specialized runtimes remain faster in several configurations, so our goal is one runtime with competitive latency rather than the fastest result in every case. Detailed profiles reveal why different models need different execution policies: on a Hopper-series GPU at batch size one, the pi0.5 action expert accounts for 8.8% of FLOPs but 57.2% of latency; at batch size 32 its share drops to 13.5% and throughput reaches about 100 samples/s. Cosmos3 remains generation-dominated and gains only 14.3% throughput as batch size increases from 1 to 16. We further introduce the control-time Roofline, which distinguishes inference-bound from environment-bound control; the measured pi0.5 points on four LIBERO suites are environment-bound while Cosmos3 stays inference-bound. Code and benchmarks: https://github.com/mingti-org/phyai.
アウトプットが分散すると、認識論の修正が続くのか? Machine Collective 向けのブラックボックス カップリング診断
集合知の研究では、意見の相違を認識論的多様性の証拠として扱います。エージェントが異なる見解を表明した場合、グループは修正する能力を保持する必要があります。 LLM 集合体では、このプロキシが壊れる可能性があります。エージェントは、同じ結論を維持しながら、多様に見える議論を生成できます。私たちは分散と修正の結合を操作します。つまり、埋め込み空間における集団の成果の分散を検証可能に増加させる介入が、前提を保持した再定式化ではなく認識論的立場の真の修正を伴う度合いです。診断はブラックボックスです。生成されたテキストのみを処理し、生成モデルの内部表現については主張しません。 2 つのチャネルは独立して測定されます。出力チャネルであるコヒーレンス インデックス (CI) は、介入によって出力分散が変化したことを検証します。認識チャネル、ターンごとのスタンスの注釈は、集合体が修正されたかどうかを測定します。我々は、この結合領域を推定するための再利用可能な方法として、出力が過収束したときに再微分プロトコル (RDP) を挿入するメタ予測明瞭度システム (MPCS) を備えた CI を提案します。 2 つの構成 (gpt-4o-mini および gemini-2.5-flash、条件ごとに 310 ペアのエピソード) から 5 つのエージェント集合を評価します。 gpt-4o-mini では、条件付き反対は誤った前提の回復を +17.7 ポイント (p<1e-6) 改善しますが、静的なペルソナの多様性は回復に悪影響を及ぼします (-8.1、p=.007)。ジェミニ 2.5 フラッシュでは、分散の低下が確認されたにもかかわらず、同等の予算で同じ介入を行っても利益は得られませんでした (26.1% 対 27.1%、p=.84)。 2 つの治療効果は互いに異なります (z=3.79、p<.001)。メカニズムのタグ付けは、ジェミニがフレームワーク内の反対意見を通じて誤った前提を維持していることを示しています。タグ付けされたRDP後の回答の94%が譲歩せずに再定式化しました(GPTでは24%)。精度とともに、介入ごとのスタンスシフトと前提保存率を報告することをお勧めします。
原文 (English)
When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Diagnostic for Machine Collectives
Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise. In LLM collectives this proxy can break: agents can produce diverse-looking arguments while preserving the same conclusion. We operationalize dispersion-revision coupling: the degree to which an intervention that verifiably increases the dispersion of a collective's outputs in embedding space is accompanied by genuine revision of its epistemic stance rather than premise-preserving reformulation. The diagnostic is black-box: it operates on generated text alone and makes no claims about the internal representations of the generating models. Two channels are measured independently: an output channel, the Coherence Index (CI), verifies that the intervention changed output dispersion; an epistemic channel, per-turn stance annotation, measures whether the collective revised. We propose CI with the Meta-Predictive Clarity System (MPCS), which inserts a Re-Differentiation Protocol (RDP) when outputs over-converge, as a reusable method for estimating this coupling regime. We evaluate five-agent collectives from two configurations (gpt-4o-mini and gemini-2.5-flash; 310 paired episodes per condition). On gpt-4o-mini, conditional dissent improves false-premise recovery by +17.7 points (p<1e-6) while static persona diversity harms recovery (-8.1, p=.007). On gemini-2.5-flash, the same intervention at a comparable budget yields no gain (26.1% vs 27.1%, p=.84) despite a verified dispersion drop; the two treatment effects differ from each other (z=3.79, p<.001). Mechanism tagging shows Gemini preserves the false premise via intra-framework dissent: 94% of tagged post-RDP responses reformulate rather than concede (vs 24% on GPT). We recommend reporting per-intervention stance shift and premise-preservation rate alongside accuracy.
On The Suitability of Differential Dataflow For Datalog Interpretation In Highly Dynamic Settings
In the domain of knowledge representation and reasoning within AI, datalog engines play an ever-increasingly crucial role. The crux of thei…
Large-Small Model Collaboration for Enhancing Edge-Deployed Small Models
Edge devices host domain-specific small language models (SLMs) with limited resources, while private clouds offer larger LLMs. We propose G…
Curiosity-Diffuser: Curiosity Guide Diffusion Models for Reliability
One of the bottlenecks in robotic intelligence is the instability of neural network models. This leads to risks when applying intelligence…
ZoomV: Temporal Zoom-in for Efficient Long Video Understanding
Long video understanding poses a fundamental challenge for large video-language models (LVLMs) due to the overwhelming number of frames and…
Review Text as a Leading Indicator of Displayed Reputation in Platform Rating Systems: Evidence from 34 U.S. Short-Term Rental Markets
Rating systems on accommodation platforms suffer from a familiar problem: nearly every listing displays a nearly perfect score, so the numb…
Pun Intended: Multi-Agent Translation of Wordplay with Contrastive Learning and Phonetic-Semantic Embeddings for CLEF JOKER 2025 Task 2
Translating wordplay across languages presents unique challenges that have long confounded both professional human translators and machine…
Emergence of Hierarchical Emotion Organization in Large Language Models
As large language models (LLMs) increasingly power conversational agents, understanding how they model users' emotional states is critical…
Uncertainty-aware Predict-Then-Optimize Framework for Equitable Post-Disaster Power Restoration
The increasing frequency of extreme weather events, such as hurricanes, highlights the urgent need for efficient and equitable power system…
Arnold: A multi-task, multi-embodiment muscle transformer policy
Controlling high-dimensional and nonlinear musculoskeletal models of the human body is a foundational scientific challenge. Recent machine…
Memorization in Large Language Models in Medicine: Prevalence, Characteristics, and Implications
Large Language Models (LLMs) have demonstrated significant potential in medicine, with many studies adapting them through continued pre-tra…
Dynamic Jailbreaking Attack
Existing gradient-based jailbreak attacks typically optimize a fixed-length adversarial suffix toward a predefined target response with a s…
RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation
Vision-Language-Action (VLA) models have shown strong manipulation capability when trained with large-scale imitation learning datasets. Ho…
Bi-Level Reinforcement Learning Pathway for Sim-to-Real Optimality
Training Reinforcement Learning (RL) policies using simulation models before deployment in real-world environments is a common strategy whe…
When Large Language Models Know the Table: A Framework for Assessing Data Contamination in Tabular Datasets
Large language models (LLMs) are increasingly exposed to data contamination, i.e., performance gains driven by prior exposure of test datas…
Neural Diversity Regularizes Hallucinations in Language Models
Language models continue to hallucinate despite increases in parameters, compute, and data. We propose neural diversity -- decorrelated par…
Reinforcement Learning and Consumption-Savings Behavior
This paper demonstrates how reinforcement learning can explain two puzzling empirical patterns in household consumption behavior during eco…
MediRec: Enhancing Chinese Medication Recommendation with Explainable Clinical Reasoning
Large language models (LLMs) have shown strong potential for clinical decision support through their advanced language understanding and re…
FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation
While LLMs have shown great success in financial tasks like stock prediction and question answering, their application in fully automating…
Stabilizing Multi-Attack Adversarial Training via Bandit Optimization
Deep Neural Networks (DNNs) remain vulnerable to diverse adversarial perturbations, motivating multi-attack adversarial training (AT) for i…
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
Does Chain-of-Thought (CoT) reasoning genuinely improve Vision Language Action (VLA) models, or does it merely add overhead? Existing CoT-V…
Interpreting GFlowNets for Drug Discovery: What probes can and cannot show
Generative Flow Networks (GFlowNets) construct molecules through sequential decisions, but their internal policies remain opaque, limiting…
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
Vision-Language Models (VLMs) excel at reasoning in linguistic space but struggle with perceptual understanding that requires dense visual…
MODEST: Multi-Optics Depth-of-Field Stereo Dataset
Training and evaluation of state-of-the-art computer vision algorithms for reliable shallow depth of field (DoF) rendering and defocus debl…
Revisiting Generalization Across Difficulty Levels: It's Not So Easy
We investigate how well large language models (LLMs) generalize across different task difficulties, a key question for effective data curat…
Feedback Loops and Code Perturbations in LLM-based Software Engineering: A Case Study on a C-to-Rust Translation System
The advent of strong generative AI has a considerable impact on various software engineering tasks such as code repair, test generation, or…
HERO: Hierarchical Evidential Reasoning Optimization for Radiology Report Generation via Reason-then-Summarize
Multimodal Large Language Models (MLLMs) have substantially advanced Radiology Report Generation (RRG), yet aligning them through reinforce…
Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation
Reinforcement learning (RL) has emerged as a powerful paradigm for fine-tuning large-scale generative models, such as diffusion and flow mo…
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
RL-based post-training with GRPO is widely used to improve large language models on individual reasoning tasks. However, real-world deploym…
A Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon Agents
Research in artificial intelligence is shifting from model innovations and benchmark scores towards problem definition and rigorous real-wo…
Can Post-Training Transform LLMs into Causal Reasoners?
Causal inference is essential for decision-making but remains challenging for non-experts. While large language models (LLMs) show promise…
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
The transition toward localized intelligence through Small Language Models (SLMs) has intensified the need for rigorous performance charact…
AdaCorrection: Adaptive Offset Cache Correction for Accurate Diffusion Transformers
Diffusion Transformers (DiTs) achieve state-of-the-art performance in high-fidelity image and video generation but suffer from expensive in…
Formal Analysis and Supply Chain Security for Agentic AI Skills
32 pages, 5 theorems with full proofs, 68 references, open-source tool: https://github.com/qualixar/skillfortify. v2: corrects the bibliogr…
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
Video object insertion is fundamental to video editing, yet existing diffusion methods often produce visually plausible but physically inco…
Seeking Physics in Diffusion Noise
Do video diffusion models encode signals predictive of physical plausibility? We probe intermediate denoising representations of pretrained…
Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discussions
Multi-agent discussions have been widely adopted, motivating growing efforts to develop attacks that expose their vulnerabilities. In this…
The Luna Bound Propagator for Formal Analysis of Neural Networks
The parameterized CROWN analysis, a.k.a., alpha-CROWN has emerged as a practically successful abstract interpretation method for neural net…
SleepVLM: A Rule-Grounded Vision-Language Model for Auditable Sleep Staging
Sleep staging is essential for sleep assessment and disorder diagnosis. In recent years, automatic sleep staging systems have achieved accu…
Terminal Agents Suffice for Enterprise Automation
There has been growing interest in building agents that can interact with digital platforms to execute meaningful enterprise tasks autonomo…
MOON3.0: Reasoning-aware Multimodal Representation Learning for E-commerce Product Understanding
With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention. A…
Generative Experiences for Digital Mental Health Interventions: Evidence from a Randomized Study
Digital mental health (DMH) tools have extensively explored personalization of interventions to users' needs and contexts. However, this pe…
Multi-Modal Learning meets Genetic Programming: Analyzing Alignment in Latent Space Optimization
Symbolic regression (SR) aims to discover mathematical expressions from data, a task traditionally tackled using Genetic Programming (GP) t…
Topology-Aware Reasoning over Incomplete Knowledge Graph with Graph-Based Soft Prompting
Large Language Models (LLMs) have shown remarkable capabilities across various tasks but remain prone to hallucinations in knowledge-intens…
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world usefulness. Instead, users often rely on ``vib…
Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models
Recent advances in vision language models (VLMs) offer reasoning capabilities, yet how these unfold and integrate visual and textual inform…
Just Repair: A Minimal Denoising Network for Time Series Anomaly Detection
Time series anomaly detectors have grown steadily more complex, incorporating attention mechanisms, adversarial training, and stochastic la…
A Systematic Review and Taxonomy of Reinforcement Learning-Model Predictive Control Integration for Linear Systems
The integration of Model Predictive Control (MPC) and Reinforcement Learning (RL) has emerged as a promising paradigm for constrained decis…
DeepImagine: Clinical Trial Outcome Prediction via Stepwise Local Counterfactual Imaginations
Predicting the outcomes of prospective clinical trials remains a major challenge. Clinical trial outcomes result from complex interactions…
IConFace: Fine-Grained Identity Conditioning for Reference-Aware Face Restoration
Severe face degradation can remove person-specific evidence, making restoration underdetermined. A generative prior may recover a sharp, pl…
CIDR: A Large-Scale Industrial Source Code Dataset for Software Engineering Research
We present the Curated Industrial Developer Repository (CIDR), a large-scale dataset of real-world software repositories collected from ind…
Stable Attention Response for Reliable Precipitation Nowcasting
Precipitation nowcasting remains challenging due to the highly localized, rapidly evolving, and heterogeneous nature of atmospheric dynamic…
When Bits Break Recourse: Counterfactual-Faithful Quantization
Model quantization is widely used to reduce memory, latency, and deployment cost, and is typically judged by whether predictive accuracy is…
Spectral Integrated Gradients for Coarse-to-Fine Feature Attribution
Integrated Gradients (IG) is a widely adopted feature attribution method that satisfies desirable axiomatic properties. However, the choice…
構造的に欠損した共変量を使用した分布的に堅牢な転移学習、国境を越えた心停止予測への応用
主要なトレーニング共変量が導入時に利用できず、ラベル付けされた結果がターゲット領域で制限されている場合、医療システム全体に臨床予測モデルを導入すると失敗することがよくあります。たとえば、院外心停止 (OHCA) の高性能モデルは、リソースが豊富な環境で定期的に収集される病院前の詳細な測定値に依存していますが、多くの国際登録では利用できません。既存の手法は、欠落している共変量を破棄して予測情報を犠牲にするか、ターゲットの分布に関するテスト不可能な仮定に依存します。私たちは、特定の共変量が構造的に欠如し、結果ラベルが利用できないターゲット母集団に予測モデルを転送するフレームワークである DRUM (\underline{D}istributionally \underline{R}obust \underline{U}nsupervised transfer learning with Structurely \underline{M}issing covariates) を提案します。 DRUM パーティションは、すべての設定で観察される共有コンポーネント ($X$) と、ソースでのみ観察される欠落コンポーネント ($A$) に共変します。 DRUM は、欠落している共変量を代入するのではなく、ソース条件からの許容偏差を制御するロバストネス パラメーターを使用して、ニューラル ネットワーク ジェネレーターを使用して $A \mid X$ の未知のターゲット分布に対する最悪の場合の予測パフォーマンスを最適化します。さらに、迷惑な推定誤差に対する感度を低減するバイアス補正手順を開発します。シミュレーションでは、分布シフトの下で平均予測誤差と最悪の場合の予測誤差の両方が大幅に改善されたことが示されています。 DRUM を国境を越えた OHCA 予測に適用し、米国のレジストリから病院前の変数が記録されていない複数のアジアのレジストリにモデルを転送すると、より適切に校正された予測が得られ、施設全体で臨床分類のパフォーマンスが向上します。
原文 (English)
Distributionally Robust Transfer Learning with Structurally Missing Covariates, with Application to Cross-National Cardiac Arrest Prediction
Deploying clinical prediction models across healthcare systems often fails when key training covariates are unavailable at deployment and labeled outcomes are limited in the target domain. For example, high-performing models for out-of-hospital cardiac arrest (OHCA) rely on detailed prehospital measurements routinely collected in high-resource settings but unavailable in many international registries. Existing methods either discard missing covariates, sacrificing predictive information, or rely on untestable assumptions about their target distribution. We propose DRUM (\underline{D}istributionally \underline{R}obust \underline{U}nsupervised transfer learning with structurally \underline{M}issing covariates), a framework that transfers prediction models to target populations where certain covariates are structurally absent and outcome labels are unavailable. DRUM partitions covariates into shared components ($X$), observed across all settings, and missing components ($A$), observed only in the source. Rather than imputing missing covariates, DRUM optimizes worst-case predictive performance over the unknown target distribution of $A \mid X$ using a neural network generator, with a robustness parameter controlling allowable deviation from the source conditional. We further develop a bias correction procedure that reduces sensitivity to nuisance estimation error. Simulations show substantial improvements in both mean and worst-case prediction error under distribution shift. Applied to cross-national OHCA prediction, transferring models from a US registry to multiple Asian registries where prehospital variables are unrecorded, DRUM yields better-calibrated predictions and improved clinical classification performance across sites.
VibeSearchBench: 実環境における長期的なプロアクティブ検索のベンチマーク
LLM ベースのエージェントは検索ベンチマークで高いスコアを獲得していますが、実際のユーザーは一貫して結果が満足できないと感じており、評価とエクスペリエンスのギャップが根強く残っていることが明らかになりました。このギャップは、既存のベンチマークが過剰に指定されたクエリ、シングルターン インタラクション、および固定スキーマ評価に依存しているためであると考えられますが、これらのいずれも、ユーザーとエージェントが協力してマルチターン対話を通じてあいまいな意図を洗練するという実際の検索動作を反映していません。私たちはこのパラダイムを VibeSearch と名付け、20 のドメインにわたって手動で精選された 200 のバイリンガル (中国語と英語) タスクで構成されるベンチマークである VibeSearchBench を導入します。このベンチマークは、VibeSearch-Pro (プロフェッショナル) サブセットと VibeSearch-Daily (日常生活) サブセットに分かれています。各タスクは、ユーザー ペルソナとスキーマフリーのグラウンド トゥルース ナレッジ グラフを組み合わせ、漸進的開示ユーザー シミュレーターとグラフ マッチング評価フレームワークを通じて評価されます。 ReAct フレームワークと OpenClaw エージェント ハーネスの両方で 7 つのフロンティア モデルのベンチマークを行います。結果は、すべてのモデルが依然として VibeSearch には実質的に不十分であることを示し (最高 F1: 30.30)、ロングコンテキスト推論、プロアクティブな意図の引き出し、および構造化された知識の構築における根本的な進歩の必要性を強調しています。
原文 (English)
VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild
LLM-based agents score well on search benchmarks, yet real users consistently find results unsatisfying, revealing a persistent evaluation-experience gap. We attribute this gap to existing benchmarks' reliance on over-specified queries, single-turn interactions, and fixed-schema evaluation, none of which reflect real search behavior where users and agents collaboratively refine vague intent through multi-turn dialogue. We term this paradigm VibeSearch and introduce VibeSearchBench, a benchmark comprising 200 manually curated bilingual (Chinese and English) tasks across 20 domains, split into VibeSearch-Pro (professional) and VibeSearch-Daily (daily-life) subsets. Each task pairs a user persona with a schema-free ground-truth knowledge graph, and is evaluated through a progressive-disclosure user simulator and a graph-matching evaluation framework. We benchmark seven frontier models under both the ReAct framework and the OpenClaw agent harness. Results show that all models remain substantially inadequate for VibeSearch (best F1: 30.30), highlighting the need for fundamental advances in long-context reasoning, proactive intent elicitation, and structured knowledge construction.
深層学習のハミルトン・ヤコビ理論
この論文では、ニューラル ネットワークのトレーニングは、ハミルトン - ヤコビの初期値問題による検索として正確に特定されています。各勾配ステップは、ホップ - コール プロパゲータが観測値に最もよく適合する粘性ハミルトン - ヤコビ方程式の初期データを選択します。推論時の入力は、その解が評価される空間点であり、初期条件はすでに重みにエンコードされています。この対応関係は、log-sum-exp 層と、より広範なアーキテクチャの構造に対して正確です。残差ネットワーク、変換器、リカレント アーキテクチャ (RNN、LSTM、SSM) はそれぞれ、アーキテクチャに依存するハミルトニアンと粘性を使用して、同じクラスのハミルトン-ヤコビ方程式を離散化します。単一の変形パラメータ $\varepsilon$ は、リプシッツ条件下で閉じた可換図の 4 つの視点 (ネットワーク、熱帯代数、粘性偏微分方程式、凸最適化) をすべて統合します。定量的な結果には以下が含まれます: 固定 $t$ に対するミニマックス最適汎化率 $O(n^{-1/(d+2)})$。敵対的な堅牢性は $\varepsilon$ によって制御されます。残差ネットワークのハミルトニアン系の共状態方程式としてのバックプロパゲーション (Pontryagin Maximum Principle)。 PDE求積法によるデータ固有の次元と一致するスケーリング指数。閉じた形式の $O(N)$ 影響関数 (ソフトマックス属性重み $\pi_j$) のエントロピー ランドスケープは $\varepsilon$ が増加するにつれて褶曲分岐を起こし、それぞれが属性盆地をマージします。
原文 (English)
The Hamilton-Jacobi Theory of Deep Learning
In this paper, training a neural network is identified, exactly, as a search through Hamilton--Jacobi initial-value problems: each gradient step selects the initial data of a viscous Hamilton--Jacobi equation whose Hopf--Cole propagator best fits the observations; at inference, the input is the spatial point at which that solution is evaluated and the initial condition is already encoded in the weights. The correspondence is exact for log-sum-exp layers, with ReLU, sigmoid, SiLU, and GELU each an exact limit, gradient, or moment of the same object, and exact in composition across depth and width, with a quantified error at finite depth that vanishes in the joint limit. It is structural for residual networks, transformers, and recurrent networks (RNNs, LSTMs, SSMs), each discretizing the same class of equations, at a named and quantified approximation error. A single deformation parameter $\varepsilon$ unifies all four perspectives (network, tropical algebra, viscous PDE, convex optimization) in a commutative diagram closed under Lipschitz conditions. Quantitative consequences include: the minimax optimal generalization rate $O(n^{-1/(d+2)})$ for fixed $t$; adversarial robustness controlled by $\varepsilon$; backpropagation as the co-state equation of the Hamiltonian system for residual networks (Pontryagin Maximum Principle); scaling exponents consistent with data intrinsic dimension via PDE quadrature; and a closed-form $O(N)$ influence function (softmax attribution weights $\pi_j$) whose entropy landscape undergoes fold bifurcations as $\varepsilon$ increases, each merging attribution basins.
An Enhanced Geometric-Spectral Feature Learning Framework for Airborne Multispectral Point Cloud Classification
Multispectral point cloud (MPC) is composed of 3D spatial-spectral information, which holds tremendous potential for accurate land-cover cl…
AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models
Recent advances in pretrained large audio-language models (LALMs) have demonstrated strong capabilities across speech, sound, and music. To…
Human-in-the-Loop Atlas-Based 3D Asset Segmentation for Interactive Content Workflows
Segmenting 3D assets into meaningful regions remains challenging, especially when segmentation criteria are application-dependent and requi…
Delta-Diffusion: Modeling Longitudinal Brain Amyloid-PET Trajectories via Conditional Poisson Diffusion Bridge
While longitudinal brain PET imaging is the gold standard for quantifying the spatiotemporal accumulation of Beta-amyloid, its widespread c…
War in the Abstract: The Rise and Consequences of Militarized Language in Scientific Communication
Scientists do not, by profession, wage war. Yet warfare's vocabulary consistently appears in their abstracts. To quantify the extent to whi…
Foundations of Equivariant Deep Learning: Unifying Graph and Sheaf Neural Networks
Symmetry is everywhere in nature and society. Geometric deep learning builds architectures respecting group symmetries, whereas topological…
Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment
Best-of-$N$ (BoN) inference improves content consistency in zero-shot text-to-speech by selecting among multiple candidates with an automat…
Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results
Large language models increasingly operate as tool-using agents, where small format, argument, or function-call errors can invalidate other…
TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale
Training robust autonomous driving agents requires a simulator fast enough for reinforcement learning at scale, realistic enough to ground…
意思決定には不確実性の定量化が必要 [講義ノート]
多くの信号処理システムは最終的には「行動」するために存在します。意思決定者またはエージェントがとるべきアクションを決定する状態変数が不確実な場合、その不確実性をどのように表現するかによって、エージェントのパフォーマンスとそのパフォーマンスがどの程度信頼できるかが決まります。この講義ノートは、第一原理から単一の決定理論的設定内で、{目的} とエージェントの知識との間のつながり、および最適に動作するのに十分な不確実性表現の形式を開発します。まず、既知の環境分布を仮定して、リスク中立エージェントは状態の事後分布を必要とするのに対し、リスク回避エージェントは最適性を失うことなく {予測セット} と最悪の場合の決定ルールに依存できることを示します。次に、環境が未知の場合に目を向け、結果として生じる認識論的不確実性に対処するための 3 つの相補的なアプローチを特定します。それは、固定予測子のキャリブレーション、分布的にロバストな最適化によるクレダル (曖昧さ) セット、およびモデル パラメーターに対するベイズ推論です。共通しているのは、信頼できる意思決定には、意思決定の目的とエージェントの知識プロファイルに一致する不確実性の表現と、エージェントが実際に得られる有用性を証明する保証が必要であるということです。
原文 (English)
Decision Making Needs Uncertainty Quantification [Lecture Notes]
Many signal processing systems ultimately exist to {act}. Whenever the state variable that determines the action to be taken by a decision maker, or agent, is uncertain, the way that uncertainty is represented decides how well the agent performs and how much its performance can be trusted. This lecture note develops, from first principles and within a single decision-theoretic setting, the link between the {objective} and the knowledge of an agent and the form of uncertainty representation that is sufficient to act optimally. To start, assuming a known environment distribution, we show that a risk-neutral agent needs the posterior distribution over the state, whereas a risk-averse agent can rely without loss of optimality on a {prediction set} and a worst-case decision rule. We then turn to the case in which the environment is unknown, and identify three complementary approaches to address the resulting epistemic uncertainty: calibration of a fixed predictor, credal (ambiguity) sets with distributionally robust optimization, and Bayesian inference over model parameters. The common thread is that reliable decisions require an uncertainty representation matched to the decision objective and to the knowledge profile of the agent, together with a guarantee that certifies the utility the agent will actually obtain.
Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models
Synthetic tabular data are valued for preserving not just column-wise marginals but inter-column dependency, which carries much of the mino…
SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving
As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment…
CORF-GS: Real-Time Wireless Radiance Field Reconstruction via Coupled Optical-RF Gaussian Splatting
Recent advances in 3D Gaussian Splatting (3DGS)-based wireless radiance field (WRF) reconstruction provide an efficient solution for wirele…
SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups
Generative modeling of protein backbones promises the de novo design of proteins with prescribed structural and functional properties. Exis…
ORCA-bench: How Ready Are Language Model Agents for Oncall?
Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over n…
DragonCrawl: スケーラブルなモバイル エンドツーエンド テストのための生成的なインテント ベースのフレームワーク
モバイル アプリケーションが複雑になるにつれて、従来のエンドツーエンド (E2E) テスト フレームワークは、UI の不安定性、メンテナンスのオーバーヘッド、クロスプラットフォームのスケーラビリティに苦労しています。この論文では、埋め込みベースの類似性マッチングから大規模な言語モデルを使用した生成意図ベースの推論に進化した、連続回帰テスト用の AI 駆動モバイル テスト システムである DragonCrawl について説明します。探索的テストとクラッシュ検出に焦点を当てたこれまでの LLM ベースのテスト研究とは異なり、DragonCrawl はコード変更のたびに特定のユーザー フローを検証し、重要な機能を破壊するコミットをブロックします。 GPT-4o のマルチモーダル機能を活用することで、DragonCrawl は、CI/CD パイプラインで継続的に実行される 1,013 の自動テスト全体で、iOS で 91.6%、Android で 92.2% の合格率を達成しました。このシステムにより、テストのオンボーディング時間が 96 ~ 120 時間から 4 時間未満に短縮され、開発者のテスト メンテナンスの労力が推定 27 年節約されました。 V1 (セマンティック埋め込みマッチング) から V2 (生成的インテントベース推論) へのアーキテクチャの進化を紹介し、トークン爆発やメモリ制約などの実装上の課題について議論し、運用環境での運用経験をレポートします。エンドステート検出のためのマルチモーダルビジョンとバックエンドステート遷移を呼び出すツールの統合により、UI インタラクションとシステムステートの橋渡しとなる包括的な回帰テストが可能になります。私たちの結果は、AI 主導のテストが安定性を維持しながら、従来の自動テストの脆弱性を排除し、大規模な継続的な品質保証を可能にすることを示しています。
原文 (English)
DragonCrawl: A Generative, Intent-Based Framework for Scalable Mobile End-to-End Testing
As mobile applications grow in complexity, traditional End-to-End (E2E) testing frameworks struggle with UI volatility, maintenance overhead, and cross-platform scalability. This paper presents DragonCrawl, an AI-driven mobile testing system for continuous regression testing that has evolved from embedding-based similarity matching to generative intent-based reasoning using large language models. Unlike prior LLM-based testing research focused on exploratory testing and crash detection, DragonCrawl validates specific user flows on every code change, blocking commits that break critical functionality. By leveraging GPT-4o's multimodal capabilities, DragonCrawl achieves 91.6% pass rate on iOS and 92.2% on Android across 1,013 automated tests running continuously in CI/CD pipelines. The system reduces test onboarding time from 96-120 hours to under 4 hours and has saved an estimated 27 developer years in test maintenance effort. We present the architectural evolution from V1 (semantic embedding matching) to V2 (generative intent-based reasoning), discuss implementation challenges including token explosion and memory constraints, and report operational experience from production deployment. The integration of multimodal vision for end-state detection and tool calling for backend state transitions enables comprehensive regression testing that bridges UI interactions with system state. Our results demonstrate that AI-driven testing can maintain stability while eliminating the brittleness of traditional automated tests, enabling continuous quality assurance at scale.
It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling
Test-time scaling lifts large language model reasoning by sampling many candidate solutions and selecting among them, yet the same recipe t…
PICopilot: An LLM-based Agentic Framework for Assisting Photonic Integrated Circuit Design via Script Generation
The rapid development of photonic integrated circuits (PICs) is shifting the design flow from traditional graphical user interface (GUI)-ba…
CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship
LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scen…
PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs
Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an unde…
GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience
A wearable assistant should both answer questions about its visual history and recognize when that history is useful to the present situati…
A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models
Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models tha…
Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Research Agents
Existing deep-research agents use a Search--Visit workflow that retrieves whole webpages without considering the structure they expose thro…
GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models
Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large la…